LTX-2 ships the weights, and they make sound
LTX-2 finally ships the weights, so picture and sound can leave the same pass on a GPU you already own. Hunyuan stays the small local cutter. The stills start taking orders. A few famous names are still rent.
31 January 2026Team COEY

LTX-2 put the weights on the table. Not a teaser, not an API waitlist, the actual checkpoint: picture and sound from one model, native 4K, up to 50 frames a second, on hardware a studio already owns. December asked whether you could keep a still. This month asked whether the clip could talk back without booking a vendor.
That is the test that matters. A silent hero shot is a screensaver. A campaign needs the door close, the line, the bed. If those arrive in the same pass, you stop stitching a voiceover onto a mute file at midnight and calling it a workflow.
If the sound is a second product, you do not have a video model. You have homework.
The clip that finally speaks
LTX-2 is a dual-stream model: a fat video path and a thinner audio path, talking to each other so the foley follows the picture. Speech, room, and a bit of score, not just a beep when someone opens their mouth. You can call it through Fal, Replicate, ComfyUI, or the LTX API if you want someone else to hold the GPU. You can also download it.
Read the licence before you point it at a client that bills more than ten million a year. Under that, commercial use is open. Over it, you buy a seat. That is not a trick. It is the sentence legal will ask for, and it is better than "open" meaning a screenshot.
The other local cutter is HunyuanVideo 1.5. Smaller. Quieter. We already wrote HunyuanVideo 1.5 Makes Local AI Video Practical when it became obvious you did not need a cluster to get a first pass. LTX is the one that wants to be the film. Hunyuan is the one that fits next to the edit. If you only have one card, start there.
Grok Imagine grew a ten second clip with audio. PixVerse R1 sold "real time" like it was a director in the room. Fine. Those are products. They are not a stack you own. Watch three shots of the same face. If it drifts, you are renting a slot machine.
| Path | You keep it | You call it |
|---|---|---|
| LTX-2 | Weights, local GPU | API, Fal, Comfy |
| HunyuanVideo 1.5 | Weights, one card | Your own box |
| Hosted clips | Never | Per second |
Can you automate it
Yes, if the output is a file and a receipt, not a vibe. Text in, clip out, check the face, check the stem, refuse the take that wandered. An agent can queue that. A person still has to say whether the line landed.
The receipt is the part most demos skip. Duration. Codec. Whether audio actually arrived. If you cannot measure the file, you cannot put it in a folder a later session will trust. A pretty preview in a browser tab is not a deliverable.
Do not automate taste. Automate the grind around taste: the queue, the naming, the reject that saves a producer from watching twenty near-misses.
Stills that take an order
Video stole the month. The stills did the useful work.
GLM-Image is the one that remembers letters. Ads die on type before they die on vibes. If the headline comes back as alphabet soup, you do not have a generator. You have a mood board. A marketer who cannot read the offer in the frame will not ship the frame, and they are right.
FIBO Edit is the revision model. Change the hat. Keep the room. That is the same test we used last month, and it is still the only one a designer will accept. A full regen is not an edit. It is a new job, with a new round of legal, a new round of "why does she have a different nose," and a new reason the campaign slips a week.
HunyuanImage 3.0 Instruct is Tencent trying to make the prompt behave like a brief: do this, not that, and stop guessing the brand colour. ByteDance NextFlow promised five second stills. Speed is not the feature. Repeatability is. If you cannot call it twice and get the same kit, you cannot put it in a workflow.
What a marketer can actually plug in
Pick one engine you can run or call. Lock a style kit. Refuse any output you cannot edit without starting over. That is the whole image stack. Everything else is a demo reel.
If the team already has a Qwen edit path from last month, do not rip it out because a new stills model posted a prettier raccoon. Add a second engine only when the first one fails a job you can name.
Agents you can leave on a box
MiniMax M2.1 shipped open weights aimed at agents that stay on a tool chain. Not a smarter chatbot. A model that can take a brief, call a function, and not wander off to write you a poem about synergy.
GLM-4.7 Flash is the speed cut of the planner we already liked. Smaller. Faster. Good enough to sit next to a queue and score a draft before a human sees it. TranslateGemma 27B is the boring one, which is a compliment: offline translation you can deploy, not a tourist widget that phones home with the campaign copy.
None of these replace a person who knows the brand. They replace the intern who copied the wrong SKU into twelve sizes.
The brief still starts with a person
An agent that can call a tool is only as good as the list of tools you give it. Keep that list short. Name the file. Name the reject. Name the person who gets the ugly ones.
Hosted clips are still a vendor
Seedance, Runway, Veo, Sora. They were louder. They still will not give you the weights. Price them as a service, the way you price a colourist you do not employ.
Wan 2.6 is still the clip everyone forwards. It is still an API. Last month already said that. This month did not change it. If the film has to live on your machines, you are on Wan 2.2 or on LTX. If you are willing to rent the continuity, pay for it like a vendor and stop calling it your stack.
Hosted video is not a moral failure. It is a line item. The failure is pretending a login is an archive.
One local path is enough for February
One local video path that makes sound. LTX-2 if you can hold it. Hunyuan if you cannot. One stills engine that survives a revision. One small agent model on your side of the wall.
Do not add a new video vendor because a trailer was pretty. Do not fine-tune anything until a plain checkpoint has failed a real brief twice. Do not let the chat window become the production of record.
GroundSlate is where that work sits on a Mac. The models stay local. The footage does not have to leave the desk. If you do not want to operate the channel yourselves, that is a services conversation. It is not a new vendor for every modality.