Gemma 4 puts Apache agents on the desk
Gemma 4 puts an Apache agent on a laptop and a cluster. Nemotron 3 Nano Omni adds eyes and ears to the small NVIDIA stack. Kimi K2.6 is the open-weight cousin. Hosted video got louder. The useful work is still the model you can leave on a desk.
30 April 2026Team COEY

Gemma 4 is the open drop that made this month feel like a tooling month instead of a trailer month. Apache 2.0. A ladder: tiny edge models, a mid-size mixture of experts, a dense 31B for the desk that can hold it. Vision on the grown-up sizes. Audio on the edge pair. Google is no longer pretending "open" means a research PDF and a waitlist.
We already wrote the operator's case in Gemma 4 Is Open for Business. The month-end version is the one you tell a marketing lead. You can put a planner next to the brand kit. You can let it see a layout. You do not have to send the campaign to a hosted chatbot because the open one "cannot see."
An agent that cannot see the layout is a very confident intern with the lights off.
The ladder you actually wanted
E2B and E4B are the pocket pair. Small enough to sit on a device. Audio in, text out, which is how a field team talks to a system without opening a laptop. They are not writing your brand film. They are routing a request. That is a real job. Most "agents" fail because they try to be the film and the receptionist at once.
26B MoE is the efficient middle. A lot of parameters on paper, a working slice per token. That is the one you serve if you want throughput without buying a new rack.
31B dense is the desk model. Vision. Longer jobs. The one you fine-tune on a style of brief so it stops sounding like a press release. This is the Gemma that replaces a junior who copies the wrong SKU into twelve sizes.
The agentic pitch is real enough if you keep it boring. Function calls. A schema. A stop. Do not ask it to "be creative" and then act shocked when it invents a product name. Creativity is the human's job. Routing is the model's.
Can you automate it
Yes. Brief in, tool call out, check the arguments, refuse the wander. Gemma 4 will do that on your metal. The edge pair will do the first hop. The 31B will do the thinking. That is a stack, not a mascot.
Give it a real layout, not a stock photo of a laptop. Ask it to list the claims, the legal risks, and the missing product shot. If it hallucinates a badge you do not have, it is not ready to sit in the queue. If it points at the empty space and waits, you might have a coworker.
| Cut | Lives where | Sees what |
|---|---|---|
| E2B / E4B | Device, edge | Audio, text |
| 26B MoE | Small server | Text, vision |
| 31B dense | Desk, cluster | Text, vision |
Eyes and ears on the other open desks
Nemotron 3 Nano Omni is the other open-weight story. Same family as the Nano that landed last winter. This cut can take image and audio, not just a prompt. NVIDIA wants it as the agent that sits next to tools on a workstation. If you already run Nemotron, this is the upgrade. If you do not, Gemma is the friendlier Apache path. Do not run both "to be safe." That is how you get two wrong answers and a bigger bill.
Kimi K2.6 is the other download, not a walled garden. Open weights. A hosted API if you do not want to serve it. Long context. A swarm story that sounds like a keynote until you remember you still have to stop the swarm. We already covered the operator's version in Kimi K2.6 Pushes Open Models Closer to Real Agent Work. Use it if the desk already speaks Moonshot. Do not move the brand kit for a new religion.
What a marketer can actually plug in
One open planner. Gemma 4 if you want Apache and a ladder. Nemotron Omni if the desk is already NVIDIA. Kimi if you already live there. One vision pass on layouts and product stills. Everything else waits.
The vision pass is the part a creative director will feel. A model that can look at a layout and say "the logo is in the gutter" is worth more than a model that can write a paragraph about disruption. Seeing the work is how collaboration starts. Generating more words is how meetings get longer.
The hosted video that will not sit down
PixVerse V6 and HappyHorse-1.0 spent the month sounding like cinema. They are products. They are fast. They will not send you a checkpoint. Use them for a burst you do not need to keep. Do not build a pipeline on a model that can vanish in a pricing slide.
Llama 4 coverage was a mess of launches, reversals, and people quoting screenshots. We are not going to pretend a rumor is a stack. If it is not a file you can serve, it is not this month's story.
Qwen and LTX from the last two months are still the local picture path. This month did not replace them. It gave the text-and-vision layer something grown-up to sit on.
Speed is not custody
A hosted clip that arrives in five seconds is a gift for a mood board. It is a trap for an archive. If the brand film has to be recut in a year, you need the engine or you need the project files. A login is neither.
Pretty demos are not a pipeline
PixVerse. HappyHorse. Anything with a trailer and a waitlist. Gemma, Nemotron, and Kimi are the exception, and even those need a person who can serve them. "Open" is not "it runs itself."
An Apache licence is a start. A machine, a queue, and a person who can say no are the rest.
One planner is enough for May
One Apache agent on a box you control. Gemma 4 31B if you can hold it, a smaller cut if you cannot. One Omni or vision pass for layouts. Keep LTX for the film. Do not add a new video vendor because the demo was pretty.
If you already have Qwen3.5 from February, you do not need a second brain this month. You need a job for the one you have.
GroundSlate is where the local half sits. The planner, the still, the clip. This month made the planner honest. The rest of the year will try to sell you a replacement. You do not need one yet.