Kimi K2.6: A Trillion Parameters, Only 32B Active Per Token
Published October 10, 2026
Kimi K2.6 carries 1 trillion parameters but only activates 32 billion per token. That is the MoE secret making it cheap like a small model while still topping benchmarks.

The Secret Behind 384 Experts
K2.6 has 384 expert sub-networks. Every incoming token passes a router, like a football coach reading the opponent and pressing the substitution button. Only 8 experts are picked — plus 1 shared expert that always shows up — per token. The rest never join the compute.
The result: compute per token can be squeezed dramatically, roughly one-fifteenth of a dense model with the same total parameter count. Think of it as having the national library in your garage, but when asked, you only pull the books from the 384 aisles you actually need. Nobody forces you to carry the whole library to answer one question.
Numbers That Break Expectations
Open-weight models used to be the punchline. "Only good enough for toy tasks." K2.6 slaps that idea. In the HLE with tools benchmark it scores 54.0, passing its sibling K2.5 at 50.2, and edging past projected GPT-5.4 and Claude Opus 4.6. SWE-Bench Pro: 58.6. SWE-Bench Multilingual: 76.7 — that one speaks to multi-language coding. BrowseComp returns to 83.2, a leap from 60.2. For context, BrowseComp tests how well a model navigates the web — K2.5 could wander a little, K2.6 can walk for hours.
The GPT-5.4 and Claude Opus 4.6 numbers from the AI Daily table are projections based on public September 2026 datasets. K2.6 is the primary-source number. Treat the table as a map, not a verdict.
$0.60 per Million Tokens: Warteg Prices
The most tempting figure is not 1 trillion — it is $0.60 per million input tokens. That is roughly one-sixth of comparable closed-source alternatives. For a startup owner who reads token cost as food cost, K2.6 feels like a favorite Warteg: filling, cheap, does not rattle your wallet.
Why so cheap? Because only 32 billion are active. When one token passes, only a fraction of the 1-trillion warehouse is touched. Picture ordering a single slice from a big kitchen: you do not bake the whole warehouse of dough.
Able to Run 300 Agents at Once
What separates K2.6 from earlier open-weight generations is its scale at the agent level. It can run 300 parallel sub-agents on a project, each with its own context, totaling 4,000 tool calls across runs longer than 12 hours. That is not a single prompt collision. It is orchestration architecture.
The clearest analogy: agents here are not "standing alone." Each sub-agent is like a cook in a busy restaurant kitchen. Each has its own stove, its own tools, its own ingredients. The head chef (router/orchestrator) coordinates them. One fails? It reassigns instructions, instead of panicking and throwing up its hands.
Expert Collapse: MoE's Public Enemy
Classic MoE trouble is expert collapse: all tokens pass through one or two experts, like a football player forced to cover three positions. The team limps. Kimi adds two regularizers — routing loss and load balancing loss — so the 384 experts share the load evenly. Analogy: a healthy router is like a fair referee, distributing the ball across all lines. If every ball goes to the striker, the match becomes boring.
MIT License: The Door for Commercial Use
K2.6 is licensed under Modified MIT. You can use it commercially, modify it repeatedly, and do not pay royalties per token to Moonshot. DeepSeek and Qwen will feel the pressure: competitive pressure pushes more aggressive open-weight MoE releases. That is the flywheel: one player matching closed-source at one-sixth the price forces the closed players to cut prices or release more aggressively. Eventually, the user wins. Like a marathon where every participant keeps improving the route.
What It Means for You
A decade ago, building your own AI was making a rocket in the kitchen. Now K2.6 hands you the rocket parts — a 1-trillion-parameter model that can still sit on a handful of GPU chips — at one-fifteenth of the old price. For the developer who always wanted to self-host but was priced out, K2.6 loosens that tight belt.
But remember: "trillion parameters" should now always be followed by "...but which ones are active?". If you cannot name the active number, you might be reading a brochure, not a spec sheet.
A Note on Authority
I need to be honest: the GPT-5.4 and Claude Opus 4.6 rows here are estimates, not official claims. The primary numbers come from Kimi K2.6 — HLE 54.0, SWE-Bench Pro 58.6, BrowseComp 83.2, Toolathlon 50.0, and $0.60 per million tokens. If you want exact figures for the competitors, I can pull their official papers. For now, treat them as a floor for conversation, not a verdict.
Reasonable Expectations
One thing worth noting: a release like Kimi K2.6 does not change the fact that you still need infrastructure. Open-weight models can be downloaded, but running 1 trillion parameters -- even with just 32 billion active per token -- needs distributed memory, fast inter-node networks, and an ops team that understands queues. If your infrastructure feels like a packed highway at rush hour, latency can spike even when the math is light. K2.6's real impact shows up when you already have a decent GPU farm and want to run agents overnight -- not when you try to turn a garage into a factory overnight.
Last week, a friend looked at K2.6 at $0.60 per million tokens and immediately recalculated what he already paid for inference last month. That math felt like stumbling on a decent coffee shop charging pocket change for a real cup. The ratio is not new, but the numbers change habits.
A practical example: if a news site like this wants to store summaries of 10,000 articles in Sanity and surface the most relevant to each reader, hitting the full model on every click is not practical. But routing 85% of queries to a cheap open-weight model like K2.6 and caching the top results can cut costs from tens of dollars a month to roughly a dollar a day. That small margin is often all that separates a product worth scaling from one that stalls at the demo stage.
Even if you have no plan to spin up a trillion parameters overnight, one lesson still applies: never judge a model by total parameter count alone. Ask three things first -- how much is active? what does it cost? and can it run on the hardware you have today? For most teams, K2.6 is not a punchline or a comfort blanket, it is a reasonable bet for next year.
If you pick one benchmark from the table above, which number going cheap and runnable in your own office would actually change your work?
Additional Context on Pricing
| Fact | Figure | Source |
|---|---|---|
| HLE with tools | 54.0 | Moonshot |
| BrowseComp | 83.2 | Moonshot |
| Input cost | $0.60/1M tokens | Moonshot |
Add this pricing context: per-input cost has dropped to classic level, no longer a closed-source premium tier. That changes the calculus -- running inference cheaply is no longer a dream, it is the new baseline.
| Metric | Before | After |
|---|---|---|
| Active endpoint | 1 doc | Hundreds of sub-agents |
| Cost per token | High | $0.60 / 1M |
| Focus read time | One session | 4,000 calls |
We deliberately line the numbers up in a single table so the debate does not drift. If you say the model is cheaper without showing the cost per million tokens, your audience hears a slogan. To expand the analysis, just compare: with the reference figure above, a small everyday task can finish dramatically sooner. On top of that, active efficiency is far more convincing than paying for silence. That is why the newest release moves from weight, to how many sailors you can manage at once -- each sub-agent gets its very own context, then all outputs get rechecked before delivery. You get results that do not swing day to day, and that is the most alive parameter in production. The field relevance is plain: when per-token cost drops sharply, projects once postponed by a quarter can run on a shorter review; when task breakdown can be orchestrated, you no longer wait on one giant context window -- you use many smaller windows in parallel. The combination of those two points is the real additional context that makes this release worth recording, not the release name alone.
Related tools
Free browser tools that apply to this topic.
Share this article
Share to
Related articles

October 10, 2026
Opus 4.8: The Frontier Is Orchestration, Not Weights
Claude Opus 4.8 arrived with no new benchmark headlines. Its power shifted to hundreds of agents running in parallel. The signal: AI's frontier is now in how it runs, not how big it is.

October 9, 2026
Why Non-English Costs More: What a Tokenizer Actually Decides
One tokenizer can turn the same word into 5 tokens or 15. That difference decides your API bill and how much context you actually have.

October 8, 2026
One Model or Many: What Multimodal Architecture Actually Changes
Unified models run text, image and audio through one network. Pipelines chain separate models. The difference is not only technical.






