One Model or Many: What Multimodal Architecture Actually Changes
Published October 8, 2026
Unified models run text, image and audio through one network. Pipelines chain separate models. The difference is not only technical.

One model or many models: what actually changes
Modern multimodal systems follow one of two patterns. The first is a pipeline: each input type has its own model, and the outputs are handed to the next stage. The second is unified, or omni: a single network processes text, images, audio and video through the same path.
The difference is not only technical. A unified model builds one representation across modalities, so the same model can write about an image or describe a video without a separate translation step.
Gemini 3.5: multimodal and agentic in one system
On 1 October 2026, Google introduced Gemini 3.5 Flash, described as an enterprise-grade model combining agentic and multimodal capability. The interesting claim is not the word "multimodal", since that is now common, but the assertion that several abilities ship as one package.
What is more notable is how Google documented the result. An official post that month stated the model reached 77.9 percent on DeepSWE, and a post about a competing model recorded a gap of roughly 12 points on the same benchmark. Numbers this large usually come from a specific configuration rather than the defaults users get.
The practical consequence is clear. If a vendor measures agentic and multimodal ability together, the resulting number cannot be compared directly against a model tested in only one area. That is another reason cross-model comparison always requires reading the method.
Reka Rho-1: discrete tokens meet continuous state
In early October 2026, Reka introduced the Rho-1 model on a design that differs from most multimodal models. Instead of turning images into discrete tokens the way vision-language models usually do, Rho-1 keeps visual information as continuous tokens inside the model's internal state.
Status note: unverified. The Rho-1 architecture details here come from Reka's pre-release material. The continuous token design, the two transformer parts sharing a KV cache, and the performance claims around them have not been retested by a third party, and the release is still an early preview.
The architecture has two transformer parts that take turns. One handles language input and writes instructions. The other stays focused on visual information. Both share a single KV cache, so visual context does not have to be copied into the language space.
The consequence is not only speed. Continuous tokens are harder to compress, so a model with fewer parameters can still retain visual detail that disappears when images are compressed into discrete tokens. This is why several teams found that a smaller model could match a larger one on multimodal performance.
Analysis: where pipelines still win
Pipelines are not obsolete technology. For many production cases, this pattern remains tidier and easier to control.
| Aspect | Unified | Pipeline |
|---|---|---|
| Models used | One large model | Several smaller models |
| Compute cost | High, one full context | Low, callable separately |
| Output control | Less granular | Very granular per stage |
| Cache | One large KV cache | Separate caches can be dropped |
| Errors | Hard to trace | Easy to isolate |
| Typical use | Chat that can see video | Document summary plus image |
The difference reduces to one question: does the problem need a model that genuinely understands both input types at once, or would two models passing data to each other be enough?
If "passing data" is enough, a pipeline is almost always cheaper. If the system has to understand both at once, as an agent deciding from what it sees and what it hears, unified has an advantage two separate models cannot reproduce.
Why unified does not mean unified has won
The statement "multimodal has converged" appears often even though two things are still separate.
First, the inputs have merged while the outputs have not. Many systems accept an image as input but still produce only text. That is multimodal in one direction, not omni in the full sense.
Second, unified models pay a capacity tax. One model has to spend its capacity across text, vision, audio and video. For a smaller model, this often means weaker performance per modality compared to a model that concentrates on one modality with the same parameter count.
What to watch
First, whether independent teams start testing unified models on mixed tasks rather than on a single modality. Without that, omni claims stay hard to verify.
Second, whether open-weight teams start releasing unified architectures. Open models such as Qwen-VL and MiniCPM already combine vision and language, but audio and video are still rarely included in a genuinely open package.
Third, whether inference cost for unified models drops sharply. Attention cost over a multimodal context is currently the main reason unified models are expensive. Architectural progress here matters more for cost than a few benchmark points.
Additional context: how to read a vendor comparison
Vendor comparisons usually report the best run. Recording the median over several runs, and reporting which tasks were excluded, tells a much more useful story.
| What to ask | Why it matters |
|---|---|
| Median or best run | Best run hides instability |
| Number of runs | One run cannot show variance |
| Which tasks were dropped | Cherry-picking is common |
| Who paid for the test | Independent results differ |
A short history: from pipelines to unified
Pipelines are not a leftover from an older era. They grew precisely because data centres had to work. When a team needs to read a document, search inside it, then summarise, the cheapest approach is running a small model for each task in sequence. Using one large model for everything means a cost that makes no sense for work repeated thousands of times a day.
The shift toward unified happened because some tasks a pipeline solves poorly. If a system has to choose an action based on what it sees and what it hears at once, the information has to sit in one place. Passing it from one model to another always loses something at the boundary: a text description of an image is never as informative as the image.
The boundary is worth naming. Converting a screenshot into words loses colour, layout and small text. Converting speech into text loses tone and background noise that carried the actual meaning. Unified models avoid that loss because the signal never has to be written down as text in the first place.
The turn is that unified models are arriving at sizes smaller than people expect. Models from 3B to 7B parameters now accept images and text in one context, and on several simple tasks they approach models twice their size. That matters more than any frontier model, because real usage sits at the smaller end.
Additional context: what each pattern costs
A cost comparison cannot be read off a spec sheet. Three things determine the real number: how often the context is reprocessed, how long the cache survives, and whether stages can run in parallel.
| Factor | Pipeline | Unified |
|---|---|---|
| Context reprocessed | Once per stage | Once for all modalities |
| Execution order | Can be serialised | Waits for the full context |
| Cache | Can be dropped per stage | One large cache |
| Memory pressure | Dominated by the largest stage | The combined context |
| Cost scaling | Linear per stage | Multiplied by context length |
In a pipeline the logic is simple: when the first stage finishes, the second can start. In a unified model there is no such split. Visual context has to enter before language can be generated, so one large input determines the whole cost.
The practical consequence is straightforward. If your work breaks into clearly defined steps, a pipeline almost always wins on cost. If the work requires a judgement that looks at several input types at once, unified is much harder to replace.
That table is worth reading twice, because the last row is the one that decides most projects. A strict budget usually settles the question before capability does. A team that has counted its own cost per thousand requests will pick a pipeline even when a unified model scores better, because the gap keeps recurring every single day while the score difference is only visible on one test.
There is also a middle option that gets skipped in these comparisons. Run a pipeline for the bulk of the work, then hand only the hard cases to a unified model. Most documents never need cross-modal reasoning, and the ones that do are a small minority. This hybrid keeps the cost of the pipeline while paying the unified premium only where it changes the outcome.
| Hybrid step | What it handles | Why it is worth it |
|---|---|---|
| Small models | Extraction, classification, cleanup | Cheap and predictable |
| One multimodal model | Images inside documents | Rare in normal traffic |
| Unified model | Cases needing sound and vision | Changes the answer |
| Human review | Ambiguous cases | Cheaper than a wrong answer |
Two common mistakes
The first mistake is assuming one large model is always better than several small ones. For work with clearly defined steps, a pipeline of three small models can match one large model at a fraction of the cost. What decides this is whether the steps separate cleanly, not how impressive the headline looks.
The second mistake is assuming unified makes every modality equal. In practice a unified model is usually strongest on text, decent on images, and noticeably weaker on audio or video. Providers most often support video because it is expensive to serve, not because it is the most mature capability.
Numbers that need their status attached
The "77.9 percent on DeepSWE" figure for Gemini 3.5 needs its status read alongside it. That number comes from Google's own testing, and a gap of roughly 12 points against another model on the same benchmark does not automatically mean one model is smarter. What to look for is whether the configuration matched, whether the test matched, and whether a third party has reproduced the result.
Reka's figures need the same treatment. The Rho-1 distillation claim that turns 99 steps into 8 comes from the vendor's own testing. Technically plausible, interesting, and not independently verified.
Related tools
Free browser tools that apply to this topic.
- AI Writing TellsFind generic patterns in your own writing. Not an AI detector.
- AI Text HumanizerRewrite text into more natural, readable language while preserving its meaning.
- AI Text SummarizerCreate an abstractive summary locally in your browser with AI.
- Text and Line ToolsCount, dedupe, sort, filter, reformat, and convert case across twelve line operations in one box.
Share this article
Share to
Related articles

October 9, 2026
Why Non-English Costs More: What a Tokenizer Actually Decides
One tokenizer can turn the same word into 5 tokens or 15. That difference decides your API bill and how much context you actually have.

October 8, 2026
Why AI Benchmark Scores Are So Often Misleading
The numbers carrying a vendor's biggest claims are often measured on tests users never run. Four reasons, and how to read them.

October 8, 2026
Why Images Use Diffusion and Text Does Not
Two ways to make an image out of noise share the same endpoint and opposite processes. Why the choice has flipped since 2015.



