Kimi K3 Is the Biggest Open-Weight Model Ever Released. The Headline Number Isn't the Story.
Moonshot AI's 2.8-trillion-parameter model is really a bet on memory engineering over raw compute — and a preview of what it takes to actually run something this size.
2.8T: Parameters — the largest open-weight model to date
1.4TB: Model size at 4-bit precision
64+: Accelerators recommended to serve
Kimi K3 Pricing Analysis: Why the Biggest Open-Weight Model Exited the Budget Tier.
Since Moonshot AI released Kimi K3 on July 16, nearly every headline has led with the same number: 2.8 trillion parameters, the largest open-weight model ever shipped. That number is also the least useful thing about it — because K3 wasn't built to win a size contest. It was built to solve a memory problem.
K3 leaps from Moonshot's previous 1T-class flagship straight past the 3T threshold, clearing DeepSeek's 1.6T V4 Pro by a wide margin. But Moonshot's own technical blog tells a more specific story than "bigger model": K3 doesn't dodge US compute restrictions so much as trade compute cost for memory cost at nearly every layer of its design — and those two constraints are not equally available to a Chinese lab.
1: Why This Is a Memory Story, Not a Compute Story:
Running a large model costs two different things: compute, the calculation needed to produce each word, and memory, how much of the model has to sit loaded and instantly reachable the entire time. Chip export controls have squeezed China's compute hardest, and K3's architecture reads as a sustained effort to spend less of it.
Its main lever is mixture-of-experts: K3 splits into 896 specialised sections and activates just 16 at a time — about 1.8% of the total — cutting the calculation per word sharply. But that trick doesn't touch the memory bill, since all 2.8 trillion parameters still have to stay loaded in case they're called next. So Moonshot went after memory directly, training K3 at four bits of precision per parameter instead of the usual sixteen. Independent analysis puts the resulting model at roughly 1.4TB, versus 5.6TB at full precision.
2: The Second Lever — And the Commercial Bet Behind It:
A second change, called Kimi Delta Attention, targets a different memory cost: as a model works through a long document, it builds up a running store of everything it's already read. At K3's advertised million-token limit — a few thousand pages — that store becomes the single biggest thing sitting in memory, bigger than the model itself.
Moonshot has been unusually direct that this is a commercial play. It contributed caching code to the open-source vLLM serving project and says the combination of quantisation and attention design is what lets it price K3 competitively despite its size — recommending it be served across 64 or more tightly-linked accelerators, the same pooling approach behind Huawei's CloudMatrix systems.
"Large-scale pre-training combined with architectural work can still deliver step-change gains for flagship Chinese models despite compute constraints." — Alex Liu, Bank of America analyst

The Hidden AI War
Nobody Is Telling You About
Our latest documentary deep-dive into the geopolitical struggle for machine intelligence dominance. Explore the two paths of AI development: open source vs. closed architecture.
3: What It Actually Costs to Deploy:
K3's weights land July 27, and any organisation with the hardware can download, modify, and self-host it. The real question is how many can. At roughly 1.4TB and a recommended 64-plus accelerators wired as one pool, this is a data-centre commitment, not a server-room one — which means most enterprises will end up renting dedicated capacity rather than owning it.
That still satisfies in-country data requirements, but it doesn't deliver the infrastructure independence that draws many buyers to open weights in the first place.
Pricing has shifted too: $3 per million input tokens (dropping to $0.30 on recently-seen input) and $15 per million output tokens — well under Fable 5's $50 output price, but far above z.ai's GLM-5.2 at $4.40 and DeepSeek V4 at $0.87. K3 has left the budget tier its predecessors occupied, and with only a maximum-reasoning-effort setting available at launch, long reasoning chains can add up fast.
The tooling isn't fully ready either — K3's two architectural changes are new enough that standard open-source serving tools don't yet support them, so teams should treat the launch date and the usable date as separate milestones.
4: What Enterprises Should Take From This:
Moonshot itself is candid that K3 still trails Claude Fable 5 and GPT 5.6 Sol on overall performance, and flags real limitations around unstable generation and unexpected autonomous decisions when intent is ambiguous. Open-weight models overall are gaining ground fast — up to 29% of tokens routed through Vercel's production gateway in June, from about a ninth in April — but K3's real test starts on July 27, when its claims become independently checkable for the first time.
The bigger lesson for any enterprise evaluating open-weight AI isn't about parameter counts. It's that the deployability of a model — the infrastructure, tooling maturity, and true cost per completed task — matters far more than what tops a benchmark chart.
Skip the Data-Centre Arms Race:
K3 is a reminder that raw model size doesn't equal usable AI — deployability, cost per task, and infrastructure fit are what actually determine ROI.
Otherworlds AI's Agent+ platform gives you enterprise-grade AI capability without a 64-accelerator commitment, and our team also builds custom enterprise AI solutions sized to what your business can actually run.
Learn more at otherworldsai.com







