---
title: Gemma 12B sees and hears on your laptop
summary: Gemma 4 12B puts a seeing, hearing Apache model on a laptop. JoyAI-Echo tries to hold a story across shots. Cosmos 3 opens a world model for physical work. LTX Trainer lets you teach the local film. Hosted video got longer. Custody still lives on your metal.
lede: A laptop model that can see and hear, without a cloud hop.
date: 2026-06-30
authors: Team COEY
image: /blog/what-june-2026-opened-for-production.webp
image_alt: An open laptop throwing a still one way and a sound ring the other
image_credit: COEY
keywords: AI Open Source
software: groundslate
---

[Gemma 4 12B](https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12b/) is the open drop that made this month feel like someone finally built the middle of the ladder. Apache 2.0. Dense. Encoder-free, which is a fancy way of saying pictures and sound go straight into the same brain instead of through a pair of bolted-on translators. Native audio. Small enough to sit on a laptop with about 16GB of memory. Google is filling the gap between the pocket Gemma 4 cuts from April and the 26B mixture that wants a server.

That is a production sentence, not a research one. A planner that can look at a layout and hear a scratch read, on the same machine as the footage, is how you stop shipping briefs to a chatbot "just this once."

> A model that needs a cloud hop to see the work is a guest. A model that can see it on the desk is staff.

## The middle of the ladder

April gave you E2B, E4B, a 26B mixture, and a 31B dense. Useful. Awkward, if the job was "see this, hear that, stay on a laptop." 12B is the cut that job was waiting for. Multi-token prediction if you care about latency. A 256K window if your serving stack can hold it. Hugging Face, Ollama, Google's on-device stack. The distribution is the product as much as the weights.

Encoder-free is the part the internet will over-explain. The useful version: fewer moving pieces, less memory spent on translators, one set of weights to fine-tune when you teach it a kit. If you have ever babysat a vision encoder that refused to learn the logo, you already know why that matters.

It is not a film model. It will not replace LTX. It will sit next to LTX and tell you whether the take matches the brief. That is the unsexy win. Most "multimodal" launches want to be the movie. This one wants to be the assistant director.

### Can you automate it

Yes. Image in, audio in, a schema out. Claims, risks, a yes or a no. An agent can refuse a take before a producer watches it. A person still has to hear the ones that almost work.

Do not ask 12B to generate the campaign. Ask it to look at the campaign you already have. Seeing is the job. Generating is how you get another folder of almosts.

| Model | You keep it | It is for |
| --- | --- | --- |
| Gemma 4 12B | Weights, Apache 2.0 | See, hear, route |
| JoyAI-Echo | Weights, LTX licence | Long A/V stories |
| Cosmos 3 | Weights, OpenMDW | Worlds, not ads |

## A story that tries to remember

[JoyAI-Echo](https://github.com/jd-opensource/JoyAI-Echo) is the other open-weight swing. JD's team wants a multi-shot audio-video story that holds a face and a voice across minutes, not seconds. Memory bank. Distilled generator. Built on the LTX family, which means you inherit that licence, including the commercial ceiling. Read it. This is not Apache just because the repo is public.

The pitch is the one every producer has been shouting at short clips: stop giving me a beautiful five seconds that cannot remember the jacket. Echo is an attempt. It is also a research stack that wants a serious GPU and a structured prompt, not a magic button. Codes and weights are out. "Ready for a client Tuesday" is a different sentence.

If you try it, try it as a first assembly, not a finish. Hold the face. Hold the voice. Then hand the cut to a person. A five minute generation that drifts at minute three is still a slot machine, just a longer one.

## A world model, not a brand film

[NVIDIA Cosmos 3](https://blogs.nvidia.com/blog/cosmos-3-physical-ai-open-world-foundation-model/) is the physical-world cousin. Open weights. Nano and Super. Text, image, video, ambient sound, action. Built for robots, vehicles, and synthetic data, not for your summer campaign. The licence is OpenMDW, which is a real licence, not a vibe. Useful if you simulate a warehouse. Overkill if you need a hero shot of a shoe.

We are mentioning it because the industry will try to sell you world models as the new video models. Sometimes the overlap is real. Usually it is a keynote. If your job is a spot, stay on LTX. If your job is teaching a machine how a forklift moves, Cosmos is the one that belongs on the list.

## Teaching the local film

[LTX Trainer](https://ltx.io/blog/introducing-the-new-ltx-trainer-one-framework-every-training-mode) is the quiet piece that will matter longer than a trailer. One framework. Video, audio, cross-modal, reference LoRAs. You describe what you want to teach. You keep the adapter. That is how a studio stops prompting "in our brand style" like it was a personality, and starts owning a kit.

Fine-tuning is not a personality transplant. It is a habit. Feed it good examples. Hold out a few. Refuse the run that overfits to one hero face. The trainer makes that possible. It does not make it free, and it does not make it wise on the first try.

### What a marketer can actually plug in

Gemma 12B on the laptop that already holds the deck. Echo only if you have the GPU and a producer who can kill a long take. Cosmos only if someone on the team actually builds physical systems. LTX Trainer if you already run LTX and you are tired of describing the kit in adjectives.

Hosted video got longer this month. ByteDance showed a Seedance cut that wants to be a thirty second story. We wrote the operator's version in [Seedance 2.5 Pushes AI Video Toward Longer, Workflow-Ready Clips](/resources/blog/2026/06/25/bytedances-seedance-2-5-pushes-ai-video-toward-longer-workflow-ready-clips). It is still rent. Longer rent is still rent.

## Longer hosted clips are still rent

The long hosted clip. The world-model demos that live on a vendor stage. Anything that will not send you a file you can serve. Gemma, Echo, Cosmos, and the trainer are the exceptions, each with a licence you should read before a client asks.

## Teach a kit only after a named miss

One laptop multimodal. Gemma 4 12B if you can hold it. One local film path you already have, plus a LoRA only if a plain checkpoint has failed a named job twice. Do not start a world-model workstream because the reel was pretty.

GroundSlate is where the local half sits. The planner that can see. The clip that can speak. The kit you are finally allowed to teach. The rest is a vendor, and this month made that easier to say without sounding bitter.
