Why Images Use Diffusion and Text Does Not
Published October 8, 2026
Two ways to make an image out of noise share the same endpoint and opposite processes. Why the choice has flipped since 2015.

Language models make text in an obvious way: one token at a time, left to right, each token built on everything before it. Image models cannot work that way. An image has no natural order. There is no left to right inside a photograph of a cat.
So the opposite route was taken. Image models start from pure noise and repeatedly remove that noise, step by step, until what is left is an image. The name is diffusion, borrowed from physics, and it dates to 2015.
Two opposite directions
The difference fits in one table, and that table explains almost every technical consequence that follows.
| Language model | Image model | |
|---|---|---|
| Starts from | empty text | full noise |
| Direction | adds tokens | removes noise |
| Nature | irreversible per token | reversible at every step |
| Technical term | autoregressive | diffusion |
The third row is the key. Writing a token is irreversible: once "sek" is written, "s" can no longer replace it. Adding a layer of noise to an image, on the other hand, is reversible, and that reversibility is what makes the whole algorithm possible.
Why reversibility matters
The denoising diffusion probabilistic model, settled in 2020, works like this. Take a real training image. Add noise gradually until it cannot be told apart from pure noise. The model now has two processes: one that turns an image into noise, and one that learns to reverse it.
To make a new image: start from pure noise, run the reverse process a number of times, and at each step predict the noise present and subtract it. After thirty to a hundred steps, what remains is an image.
That asymmetry is what makes it work. Training needs a forward process that destroys the image, because there is no reverse. Inference runs the reverse, because by then it has been learned.
Numbers that explain why this moved so fast
Three figures from 2026 show the direction, and all three come from each company's own testing, so their status matters.
Status note: unverified. All three figures in this section are company claims about their own testing. No third party has repeated the measurement, and the numbers may differ on a different configuration, resolution or device.
- A 14.5x speedup. Video DeltaNet completes denoising for a 14.3 second 768p video in 6.70 seconds on eight NVIDIA B200 GPUs, against a 50-step baseline taking 91 seconds on the same GPU count.
- Ninety-nine steps down to eight. The distilled variant of Rho-1 cuts denoising steps by more than ten to one.
- A 75 percent cut in communication. Sparse Sequence Parallelism reduces inter-rank communication volume by 75 percent and drops per-block communication steps from four to one.
Those three numbers come from three different problems: attention that costs too much, too many steps, and communication that piles up across devices. Not one of them was solved by making the model bigger.
That is the interesting pattern. Since 2022 the gains have come from architecture and from the serving pipeline, not from parameter count. It is the opposite of the first two years of the language model boom.
A short history: why the direction was flipped
In 2015 diffusion work focused on physics simulation: modelling how particles move, then learning to reverse it. The researchers came from physics rather than from computer vision.
It was 2020 before the approach moved to images, and the result surprised people. GANs at the time struggled to produce genuinely sharp images, while diffusion produced images that were immediately usable, by improving rather than by generating in one pass.
Two years later Stable Diffusion made the approach accessible. The same road leads to video, and the thing that had to change was not the idea but the number of frames that must stay consistent.
Why video is harder
A video is not one image but hundreds of them that must stay consistent with each other. Noise removed in the first frame has to produce motion that remains coherent through the last one.
That is why generative video is so much slower than generative images. In JoyAI-Video-Edit, the pipeline reaches 30.19 frames per second at 720x1280, and that came from four specific changes: an autoregressive design, long horizon optimisation, bounded KV-state inference, and scheduling that favours decode throughput.
Three mechanisms work together there. The first makes frames arrive sequentially instead of all at once. The second stops errors accumulating, because a generated video without it will shift colour after a few seconds. The third bounds how much state the model has to remember.
From denoising to flow matching
The 2020 algorithm had a practical problem: too many steps. Each step is a full pass over the entire network, and thirty steps already feels slow.
Flow matching replaces it. Instead of removing noise in small increments, the model predicts the velocity of the change between noise and image. One step means one large update rather than many small ones.
The effect is visible in the numbers. Reka reports that a distilled variant of Rho-1 cuts denoising from 99 steps to 8, and claims a 5.3 second clip in about one second. That figure comes from the company's own testing, so it reads as a claim rather than a third-party verified result.
Eight steps against ninety nine is a difference of an order of magnitude. That is what moved generative video from lab demo to product.
The newest video architecture: two paths, one context
One framework is more interesting still. Vega separates two jobs that are usually tangled together.
An autoregressive transformer, based on Qwen2.5-3B, treats every input as discrete tokens, including images through a vision tokenizer that maps frames into a 65,536 token vocabulary. That transformer does not produce images. It produces semantic tokens describing content and motion.
A diffusion decoder, based on Wan2.1-1.3B, then renders the actual pixels. The split exists for a reason: one path reads and writes in a shared representation and keeps things coherent, while the other renders in the space where detail lives.
It scores 82.58 on VBench with a semantic score of 82.35, reported by the authors themselves, so treat it as an unverified claim.
What this means for readers in Indonesia
For anyone who has only ever consumed AI images from a server, there is one thing worth noticing: cost and control are not the same problem.
The drop from 99 steps to 8 lowers compute cost directly, because video generation cost is dominated almost entirely by the number of denoising passes. A model that needed several GPUs can now sit on one workstation. That is what makes small local video tools possible at all.
The more relevant angle for readers here is the prompt language. Diffusion reads its prompt through a text encoder, and that encoder is trained on data dominated by a handful of languages. An Indonesian prompt often yields a composition that is literally correct but quietly wrong: the number of people, the direction of motion, or the relationship between objects does not follow the sentence.
Three things worth trying now:
- Write the prompt in two passes. Describe the layout briefly in English, then add the details in your own language.
- State the composition explicitly. Number of objects, position, direction of motion. Diffusion does not infer the way a language model does.
- Use a model that exposes a seed. The same seed gives the same result, so iteration is repeatable rather than a matter of luck.
The most common complaint from users here is the number of people and the shape of hands. Both are known limitations that no video model has solved.
Analysis: which approach wins
Neither one wins outright, and the reason is that they are good at different things.
- Diffusion excels at fine refinement. Correcting the pixel distribution one small step at a time is what makes plausible detail.
- Autoregressive excels at consistency. Because every token sees all previous tokens, there is no "frame that forgot" -- which is the central problem in video.
- Video needs both, and that is why recent architectures move to two paths: one for understanding and holding something coherent, one for rendering it smoothly.
What has changed since 2015 is not that these models are new. What changed is that denoising which needed 99 steps now needs 8, and that number is what made it practical.
What to watch
Three things. First, whether distilled steps can drop further without losing fine detail. If it reaches four, the cost halves again.
Second, whether native resolution is still the limit. Rho-1 is still capped at 672x384 while other video models reach 720p at 24 frames per second. The limit is video memory, not the algorithm.
Third, whether this reaches consumer devices. For now video models need tens of GPUs. What makes it reachable for ordinary people is not the algorithm but hardware cheap enough to run it.
Related tools
Free browser tools that apply to this topic.
- Image Resizer & ConverterResize, compress, and convert JPG, PNG, and WEBP images in your browser.
- Resize ImageResize an image to custom dimensions in your browser.
- Text and Line ToolsCount, dedupe, sort, filter, reformat, and convert case across twelve line operations in one box.
- Media ConverterConvert video, audio, and image formats in the browser with FFmpeg.wasm.
Share this article
Share to
Related articles

October 9, 2026
Why Non-English Costs More: What a Tokenizer Actually Decides
One tokenizer can turn the same word into 5 tokens or 15. That difference decides your API bill and how much context you actually have.

October 8, 2026
One Model or Many: What Multimodal Architecture Actually Changes
Unified models run text, image and audio through one network. Pipelines chain separate models. The difference is not only technical.

October 8, 2026
Why AI Benchmark Scores Are So Often Misleading
The numbers carrying a vendor's biggest claims are often measured on tests users never run. Four reasons, and how to read them.



