2.8 Trillion Parameters Open Source! Kimi K3 Pushes Closed-Source Flagships Into a Corner
2.8 Trillion Parameters, Open Source! Kimi K3 Pushes Closed-Source Flagships Into a Corner
On July 16th, Moonshot AI dropped Kimi K3 without warning.
2.8 trillion parameters, fully open source.
This is the first open-source model in history to break through the 3-trillion-parameter ceiling—what was previously stuck behind it was Llama 4 Behemoth, a 2-trillion "futures contract" that hasn't truly shipped yet, and before that...
💡 What You Will Learn
2.8 Trillion Parameters, Open Source! Kimi K3 Pushes Closed-Source Flagships Into a Corner On July 16th, Moonshot AI dropped Kimi K3 without warning. 2.8 trillion parameters, fully open source. Thi
📜 Table of Contents
- 01. In 48 Hours, an AI Designed Its Own Chip
- 02. It Also Wrote a GPU Compiler on the Side
- 03. How Does 2.8 Trillion Actually Run? Two New Architectures Plus MXFP4 Quantization
- 04. This Time, Moonshot AI Is Grabbing Food from Every Open-Source Player
- 05. Developers, Bosses, and AGI Watchers—Everyone Gets a Piece This Time
- 06. Parameters Are the Ticket, but What Counts Is Whether It Lasts 6 Months
2.8 Trillion Parameters, Open Source! Kimi K3 Pushes Closed-Source Flagships to the Brink
On July 16th, Moonshot AI dropped Kimi K3 out of nowhere.
2.8 trillion parameters, fully open source.
This is the first open-source model to break through the 3-trillion-parameter threshold—what was stuck ahead of it was Llama 4 Behemoth's 2-trillion "futures" that hasn't truly shipped yet, and before that, DeepSeek V4's 1.6 trillion. K3 just kicked the ceiling right off.
But raw parameter count alone isn't worth much. What really made me read twice were the several "hands-on capability" demos released alongside it.
01. In 48 Hours, an AI Designed Its Own Chip
Kimi K3 was given a task: use open-source EDA tools + the Nangate 45nm process library to independently design an AI inference chip within 48 hours.
It did it.
- 4 mm² area
- Timing closure at 100 MHz
- Simulated sustained throughput of 8,700 tokens/s decoding
- Packed with 1.46M standard cells and 0.277 MB SRAM
- One INT4 MAC array with fused dequantization
In other words, a large model ran the entire "how an AI accelerator is built" pipeline from start to finish in two days—front-end, synthesis, place-and-route, verification. It designs hardware for another AI, and AI runs on chips designed by AI. This isn't a demo; these are hard numbers written in black and white in Moonshot AI's technical report.
02. It Also Wrote a GPU Compiler on the Side
Think the demo wasn't hardcore enough? K3 also casually wrote a Triton-level compiler called MiniTriton.
- Comes with a tile-level IR built on top of MLIR
- Has its own optimization passes
- Has a PTX code generation pipeline
Running roofline benchmarks, it's on par with Triton and torch.compile, and on some workloads it outright beats Triton. The most absurd part: it used this self-written toolchain to run end-to-end nanoGPT training, with the loss curve hugging the reference implementation tightly.
What does this mean? Large models are no longer just "calling compilers"—they're starting to "build compilers." Previously, programmers wrote CUDA; in the future, programmers might just be filing requirements with K3.
03. How Does 2.8 Trillion Actually Run? Two New Architectures Plus MXFP4 Quantization
With parameters stacked to 2.8T, a single GPU definitely can't handle it. K3's solution is two brand-new architectural components, plus a quantization scheme designed from the training stage onward.
- Kimi Delta Attention (KDA): A linearized rework of attention that allows KV cache prefilling under long contexts, driving down per-token cost
- Attention Residuals (AttnRes): A new residual connection method that makes information flow more stably through ultra-deep networks
The base is Stable LatentMoE, activating 16/896 experts—meaning each inference only calls on about 1.8% of parameters, but you get access to 2.8T worth of knowledge capacity.
The more critical piece is the quantization strategy: MXFP4 weights + MXFP8 activations, with quantization-aware training starting from the SFT stage. This means weights and activations are already optimized for 4-bit numeric formats during training, rather than being compressed after the fact—the latter loses precision, while the former preserves inference quality.
But don't mistake quantization for "runs on a regular GPU." Moonshot AI's official recommendation for deployment hardware is clear:
"We recommend deploying Kimi K3 on supernode configurations with 64 or more accelerators."
Supernode configurations with 64+ accelerators—this is ammunition prepared for supercomputing centers, not a toy for average players. The INT4 quantized deployment in the K2.5 era also used distributed inference frameworks like vLLM / SGLang / KTransformers. The open-source community currently has no public "single-machine K3 deployment" case, and the reason is simple: even at 4-bit quantization, a 2.8T model needs roughly 1.4TB of VRAM, and a single H200 (141GB HBM3e) simply can't hold it.
For developers, there are only two realistic paths: either go through Moonshot's official Kimi API, or save up 64 accelerators for a supernode deployment. Don't get swept away by the words "fully open source"—what's open source is the weights, not the deployment cost.
04. This Time, Moonshot AI Is Grabbing Food from Every Open-Source Player
You might think DeepSeek already told the "open-source flagship" story. It's different.
DeepSeek V4 is 1.6T—strong in performance, but still in the posture of "open source chasing closed source." K3's official statement is:
"Overall performance still trails Claude Fable 5 and GPT 5.6 Sol, but across our full evaluation suite, it reaches frontier level and consistently surpasses all other tested models, including Opus 4.8, GPT 5.6 Sol, and GPT 5.5."
Translated into plain language: K3 isn't chasing closed source—on some dimensions, it's already stepping on previous-generation closed-source flagships like GPT 5.5 and Opus 4.8. It's half a step from the ceiling, but it's already sitting at the table.
Now look at the timing:
- GPT-5.6 Sol, just days after launch, was exposed for auto-deleting user files and production databases
- Claude Fable 5 excels at asynchronous multi-step tasks, but its API pricing is absurdly expensive
- OpenAI pulling out GPT-Red for automated red-teaming is essentially an admission that its own models are still being broken into
At this moment, K3 drops a 2.8T model that's frontier-level in its own evaluations, with training methods fully disclosed and weights fully released—this isn't catching up, this is flipping the table.
05. Developers, Bosses, and AGI Watchers—Everyone Gets a Piece This Time
- If you run a business: The Kimi API platform already has the kimi-k3 option live. Under the Mooncake inference architecture, cache hit > 90%, and token pricing for coding scenarios is very competitive. Claude Code users can switch to Kimi Code at zero cost (same CLI experience).
- If you're a developer: Weights are on Hugging Face, with MXFP4 quantization + vLLM / SGLang / KTransformers inference framework support, but Moonshot AI officially recommends supernode deployment with 64+ accelerators—running it yourself means budgeting at least seven figures in hardware costs
- If you care about the AGI timeline: When a model can write its own compilers, design its own chips, and build its own 3D games—how many more versions until "models build the next generation of models" arrives?
06. Parameters Are the Ticket, but What Counts Is Whether It Lasts 6 Months
Moonshot AI's playbook has been underestimated for the past two years.
Back at Kimi K1.5, people said "Moonshot only excels at long context"; when K2 came out, people said "still can't beat GPT-4's descendants." K3's punch genuinely landed on the closed-source camp's face.
But let me pour some cold water: the real test for an open-source model isn't how explosive launch day is, but whether the community still embraces it 6 months later. Llama 3 was also the "open-source nuclear bomb" back in the day, but after the MetaAI team left, the ecosystem went straight into a gap. Whether Kimi can sustain open-source iteration and support the developer community on Discord/Hugging Face matters more than the parameters.
Parameters are the ticket; ecosystem is the moat.
Let's talk in the comments: Which K3 capability impressed you the most—self-designed chips, writing compilers, or the engineering combo of MoE + 4-bit quantization? If you're a developer, would you immediately plug K3 into your own projects and give it a spin? Come on, the comments are waiting for you.
Written by our editorial team; tools listed here are tested or verified against public sources. Links point to official sites or GitHub repos for reference only — no paid placements.
