Open models finally earned a first real job
December is the first month we would start a real job on open models. Stills you can revise, sound you can select, and a brief that stays on the tools. A few famous names still make you rent.
31 December 2025Team COEY

We spent most of 2025 trying open models. December is the first month we would start a job on them.
Not because the pictures got prettier. Because you can change the hat without remaking the room. You can pull a keyboard out of a take with a sentence. You can hand a brief to a model and it stays on the tools instead of wandering off to chat. A few of the names everyone forwarded still will not give you the weights. That is rent. Price it that way.
Stills
The still has to survive a revision. That is the whole test.
LongCat-Image is six billion parameters, Apache 2.0, with an edit path and a Dev checkpoint you can fine-tune. It fits on a box a studio already has. The type holds in English and in Chinese. Pretty is not the point. You can run it.
FLUX.2 has been around since late November. [dev] is the one people already have wired. Read the licence before it touches a client catalogue. It is open weights, not Apache. [max] showed up mid-month as the paid top end. We wrote about the family in FLUX.2 Makes Image Generation Finally Automatable when it started behaving like something you can call twice and get the same kit.
Then Qwen spent the last two weeks finishing the job. Layered splits a still into RGBA, so the hat and the room are different files. Edit-2511 keeps a face when you move the light. 2512, today, is the generator that does not fall over on letters and skin. All Apache 2.0. Generate, edit, layers. One house.
Pick one engine. Lock a kit. If the designer cannot change a thing without a full regen, you do not have an image model yet.
Cuts
Wan 2.6 is the clip everyone sent around. Multi-shot, 1080p, a voice that stays with the face. It looks like a campaign. The weights did not come with it. Wan 2.1 and 2.2 are still what you host. From 2.5 on, Alibaba sold the better story as an API.
If the film has to live on your machines, you are on 2.2. If you are willing to rent the continuity, 2.6 will take the money. We said as much in Wan 2.6 Makes AI Video Multi Shot Ready, and we should have said the second sentence louder.
The open work sat next to that. LongCat-Video-Avatar downloads. You give it audio and a person, and it keeps going. That is a talking character, not a fifteen second ad. TalkVerse opened the talking-head corpus the field had been sitting on: hours of synced clips and a small model on Wan 2.2. Useful if you train. Not useful if the spot is due Tuesday. JavisGPT is a seven billion parameter experiment that understands sounding video and tries to make it. Early. The signal is the interesting part: picture and sound from one checkpoint is leaving the closed labs.
Seedance and Runway were louder. They are vendors. Watch three shots of the same person. If the face drifts, stop. If you cannot download the model, you are not building a stack. You are booking a service.
Sound
SAM Audio is Segment Anything for a mix. You say the thing and you get a stem. The voice. The crowd. The keyboard in the take you already paid to record.
Generation makes more noise. This is how you clean the noise you have. We have been treating speech as a timeline for a while: words, music maps, stems you can duck. December is when "take the keyboard out" became a tool call instead of an afternoon in a DAW. The write-up from the week it shipped is Meta SAM Audio Makes Editing Promptable.
If someone is still soloing tracks by hand to start the day, that is the first thing to replace.
Briefs
The language models were the part of the month that did not fit in a screenshot.
DeepSeek-V3.2 thinks inside a tool call. It used to think, then call, then forget why it called. That is the difference between an agent and a chatbot with extra steps. They also showed off Speciale on the olympiad papers. That one was API-only, no tools, and they turned it off on the 15th. Keep V3.2.
Mistral dropped a whole local range in a week. Large 3 if you have the cluster. Ministral 3 at 3B, 8B and 14B if you do not, all with vision, all Apache 2.0. Devstral Small 2 is the 24B coding agent you can actually leave on a Mac. Read the licence on the big Devstral if a company that size is going to serve it.
GLM-4.6V can look at a deck and call a tool in the same loop. That is the one we wanted more than the text follow-up that came two weeks later. GLM-4.7 stays on a chain once it starts. Both belong on the desk.
Nemotron-3 Nano is NVIDIA's small sparse model for long briefs, with the recipes. FunctionGemma is 270 million parameters that only exist to turn a sentence into a tool name. Use it for that. Do not spend a frontier model deciding which function to call.
Olmo 3.1 is not the strongest. It is the one that still publishes the data and the checkpoints. Use it when someone asks how the model was made.
The model is not the product. The product is the thing that calls it, checks the output, and stops.
January
One image engine you can edit. One video path that holds a face, and an honest answer about whether you own it. SAM Audio on the library you already have. One planner on your side of the wall.
GroundSlate is where that work sits on a Mac. The models stay local. The footage does not have to leave the desk. If you do not want to operate the channel yourselves, that is a services conversation. It is not a new vendor for every modality.