Kimi K3 Is Live: 2.8T Open Frontier Intelligence
> Kimi K3 is live with 2.8T parameters, a 1M-token context window, native vision, KDA and AttnRes. Architecture, reported benchmarks, pricing and deployment caveats.
🎧 Listen — ~14 min
Ready · Kimi K3 Is Live: 2.8T Open Front
Kimi K3 is now live on Kimi.com, Kimi Work, Kimi Code and the Kimi API. Moonshot AI says the model has 2.8 trillion total parameters, a 1,048,576-token context window, native visual understanding and a sparse Mixture-of-Experts design that activates 16 of 896 experts. Full model weights are planned for release by July 27, 2026.
That makes Kimi K3 one of the biggest open models ever announced—and a serious test of the belief that open Chinese models must remain six to eight months behind US closed models. The early results are striking, especially in software engineering, browser work, visual agents and long-running knowledge tasks. But the accurate conclusion is not that K3 wins every test. It is that the gap between open and proprietary frontier models is now much harder to describe with a simple leaderboard slogan.
This deep review explains what Kimi K3 changes, what Moonshot's reported release results indicate, how it compares with Claude Fable 5, GPT-5.6 Sol, GPT-5.5, GLM-5.2 and Opus 4.8, and what developers should verify before moving production workloads.
Kimi K3 at a glance
| Feature | Kimi K3 |
|---|---|
| Developer | Moonshot AI |
| Model size | 2.8 trillion total parameters |
| Architecture | Sparse MoE with Stable LatentMoE, Kimi Delta Attention and Attention Residuals |
| Active experts | 16 of 896, according to Moonshot AI |
| Context window | 1,048,576 tokens |
| Input modes | Text, images and video; native vision |
| Reasoning | Always-on thinking; max effort at launch |
| Access | Kimi.com, Kimi Work, Kimi Code and Kimi API |
| API price | $0.30/M cached input, $3/M uncached input, $15/M output |
| Weights | Planned by July 27, 2026 |
Quick answer: Kimi K3 is a long-context, native-multimodal, agent-focused model built for tasks that continue for hours rather than a few conversational turns. Its strongest early story is coding and tool use. Its open-weight significance will depend on the eventual release, license, hardware requirements and independent reproduction of the reported scores.
On this page
- Frontend Code Arena claim
- Benchmark screenshots and results
- Kimi K3 architecture
- K3 versus top-tier models
- API access and pricing
- Developer checklist
- FAQ
Frontend Code Arena: a 17-place jump
A public Frontend Code Arena announcement says Kimi K3 reached #1 with 1,679 points, ahead of Claude Fable 5, after Kimi K2.6 was ranked #18. It also says K3 led six of seven frontend domains—Brand & Marketing, reference-based design, data & analytics, consumer product, simulations, and content-creation tools—and ranked second in gaming.
That is an unusually strong frontend claim. It is relevant because UI coding increasingly involves a complete visual loop: interpret a reference, write the app, run it, inspect the result and correct the next iteration. A benchmark win in that setting is evidence of useful capability, but it is not a guarantee that a production design system, accessibility requirements, browser matrix or deployment pipeline will work without supervision.
Evidence note: This ranking is attributed to the public Frontend Code Arena announcement. The linked post is the source for the 1,679-point figure and domain breakdown. We have not independently audited the arena methodology or re-run the leaderboard, so it should be treated as a third-party reported snapshot rather than a permanent universal ranking.
The reported benchmarks: what they actually say
Moonshot's release materials give Kimi K3 a particularly strong showing in coding and agent evaluations. They should be read as a map of strengths, not as a universal ranking of intelligence.
Coding results
The coding chart reports these scores:
| Benchmark | Kimi K3 | Best score shown | What it tests |
|---|---|---|---|
| DeepSWE | 67.5 | GPT-5.6 Sol, 73.0 | Software-engineering problem solving |
| FrontierSWE | 81.2 | Fable 5, 86.6 | Longer software tasks |
| Kimi Code Bench 2.0 | 72.9 | Fable 5, 76.9 | Kimi Code’s internal coding evaluation |
| Terminal Bench 2.1 | 88.3 | Kimi K3 and GPT-5.6 Sol, 88.3 | Terminal-based agent work |
| Program Bench | 77.8 | Kimi K3, 77.8 | Program reconstruction and coding ability |
| SWE Marathon | 42.0 | Kimi K3, 42.0 | Long-horizon software engineering |
The release figures are presented below as vendor-reported results. Evaluation setups, model settings and test dates should be checked before making a direct purchasing decision.

Courtesy: Moonshot AI. Screenshot captured from the official Kimi K3 technical blog, July 17, 2026 UTC. Results are vendor-reported and harness/settings differ by benchmark.
The pattern is more interesting than the headline. K3 is not first on every software test: the chart places GPT-5.6 Sol ahead on DeepSWE and Fable 5 ahead on FrontierSWE and Kimi Code Bench 2.0. K3 does, however, lead the displayed Program Bench and SWE Marathon results, while tying GPT-5.6 Sol on Terminal Bench 2.1. That points to a model that is especially competitive when the job requires tool use, persistence and a sequence of connected actions.
Agent and visual results
The second chart reports K3 near the top across general agent tasks:
| Evaluation | Kimi K3 score | Leader shown |
|---|---|---|
| GDPval-AA v2 Elo | 1668.0 | Fable 5, 1760.0 |
| AA-Briefcase Elo | 1548.0 | Fable 5, 1583.0 |
| Automation Bench | 30.8 | Kimi K3, 30.8 |
| JobBench | 52.9 | Fable 5, 57.4 |
| SpreadsheetBench 2 | 34.8 | Kimi K3, 34.8 |
| BrowseComp | 91.2 | Kimi K3, 91.2 |
| CharXiv (RQ) with tool | 91.3 | Fable 5, 93.5 |
| Zerobench with tool (Pass@5) | 41.0 | Fable 5, 46.0 |
K3’s native vision matters here. A coding agent that can inspect a screenshot, modify the interface, run the app and inspect the next screenshot is operating in a feedback loop. That is a different workflow from producing code from a text-only specification. It is also why K3 may be more useful in frontend, game-development and CAD tasks than a model’s text-only coding score would suggest.

Courtesy: Moonshot AI. Screenshot captured from the official Kimi K3 technical blog, July 17, 2026 UTC. Scores are a release-time snapshot, not independent testing.
Internal knowledge-work results
Moonshot’s internal chart shows Kimi K3 ahead of the listed GPT-5.5 and Claude Opus 4.8 results on three production-oriented evaluations:
| Internal evaluation | Kimi K3 | GPT-5.5 | Claude Opus 4.8 |
|---|---|---|---|
| Online Exp Bench | 75.5 | 70.6 | 65.9 |
| DECK-Bench | 73.5 | 68.2 | 66.9 |
| Finance-Bench | 62.6 | 58.4 | 60.7 |
These numbers are useful signals, but they are not independent evidence. Moonshot designed or supplied the comparison, and the public documentation does not yet provide enough detail for readers to reproduce every internal task. Treat them as vendor-reported results until the technical report and external evaluations arrive.

Courtesy: Moonshot AI / Kimi K3 release materials. Screenshot supplied by the article source and preserved locally. The comparison is Moonshot-reported internal evaluation data, not an independently audited leaderboard.

Courtesy: Moonshot AI. Screenshot captured from the official Kimi K3 technical blog, July 17, 2026 UTC.

Courtesy: Moonshot AI. Source: Kimi K3 API quickstart. Accessed: July 17, 2026 UTC.
This official developer page documents the K3 API surface. It is useful evidence of availability and supported workflow; it is not independent validation of benchmark results.
Why Kimi K3 is technically different
Kimi K3 is not simply Kimi K2 scaled up. Moonshot combines three ideas that attack different limits: attention cost across long sequences, information flow through deep networks and expert routing at extreme scale.
Kimi Delta Attention: less pressure from million-token context
Kimi Delta Attention, or KDA, is a hybrid linear-attention mechanism. The basic idea is to avoid treating every position in a very long sequence as if it must interact with every other position in the same expensive way. Moonshot reports up to 6.3× faster decoding in million-token contexts.
That claim matters because a million-token context is only useful if the system can search, retrieve and generate within it at a tolerable cost. A large window can otherwise become a marketing number: technically available, but too slow or expensive for real projects. KDA is intended to make long code repositories, research archives, videos and multi-step agent state more practical.

Courtesy: Moonshot AI. Source: Introducing Kimi K3 — Open Frontier Intelligence. Accessed: July 17, 2026 UTC.
Attention Residuals: controlling information across depth
Traditional residual connections pass information from one layer to the next by addition. Attention Residuals, or AttnRes, let the network selectively retrieve earlier representations instead of giving every earlier layer the same implicit treatment. In plain language, the model gets more control over which previous computations remain useful as depth increases.
Moonshot reports about 25% higher training efficiency for less than 2% additional cost. The production-scale kernel example supplied with the announcement is even more concrete: after more than 15 hours of autonomous iteration, K3 reduced an AttnRes forward-and-backward runtime from 283.6 ms to 114.4 ms without changing the numerics. That is a company case study, not a neutral benchmark, but it is a compelling example of the model improving the infrastructure behind its own architecture.
Stable LatentMoE: huge total capacity, sparse per-token work
K3 uses a Mixture-of-Experts design with 896 experts and activates 16 for a token, according to Moonshot. This is why “2.8 trillion parameters” does not mean every token runs through 2.8 trillion parameters at once. Sparse routing allows a model to contain a very large pool of specialist capacity while keeping the per-token computation lower than a dense model of the same total size.
The trade-off is engineering complexity. Expert routing, memory placement, inter-device traffic, load balance and quantization become central deployment problems. Moonshot recommends supernode configurations with 64 or more accelerators for K3 inference, so an eventual open-weight release is unlikely to mean ordinary laptop deployment.
Kimi K3 vs GPT-5.6 Sol, Fable 5 and Opus 4.8
The attached charts support a careful conclusion: K3 has entered the same competitive conversation as the leading proprietary systems in several agent and coding tasks. They do not prove that K3 is simply “better” than every model.
K3 vs Fable 5
Fable 5 remains ahead on the displayed FrontierSWE, Kimi Code Bench 2.0, GDPval-AA v2, AA-Briefcase and visual-agent scores. K3 leads the shown Automation Bench, BrowseComp, Program Bench and SWE Marathon results. For a team building a long-running coding or research agent, the right question is which task distribution resembles its workload—not which model has the most winning rows.
K3 vs GPT-5.6 Sol
GPT-5.6 Sol leads the displayed DeepSWE score and ties K3 on Terminal Bench 2.1. K3 leads Program Bench and SWE Marathon in the supplied chart. The practical difference may come down to access, cost, privacy, tool reliability and whether a million-token context is valuable to the project.
K3 vs Claude Opus 4.8
K3 is ahead of Opus 4.8 on the displayed Program Bench, SWE Marathon, Automation Bench, BrowseComp and internal knowledge-work scores. Opus 4.8 remains ahead on several other rows, including FrontierSWE and visual-agent tests. The phrase “K3 beats Opus” is therefore too broad; “K3 beats Opus on several reported agent and coding evaluations” is defensible.
K3 vs GLM-5.2 and GPT-5.5
K3 has a clear lead over the displayed GLM-5.2 results in the coding and agent charts. It also outperforms GPT-5.5 on many of the shown tasks, while GPT-5.5 remains ahead on some rows. This is a useful reminder that model generations do not form a single straight line: one system may be better at browsing, another at code repair, another at visual reasoning and another at low-latency chat.

Courtesy: Moonshot AI. Source: Official Kimi K3 API pricing. Accessed: July 17, 2026 UTC.
This official pricing page confirms the active K3 product listing and identifies the input, output and caching cost model. Pricing and rate limits should be rechecked before a production rollout.
Is Kimi K3 a “DeepSeek 2.0 moment”?
The comparison makes sense as a signal, not as a final verdict. DeepSeek’s biggest impact was not just a benchmark score; it changed expectations about the cost, openness and global distribution of frontier capability. Kimi K3 creates a similar conversation because it pairs frontier-level reported results with a planned open-weight release and an architecture designed around extreme scale.
The open-weight date is the important checkpoint. Until July 27, the public cannot fully inspect the weights, run the model independently or measure hardware requirements. Even after release, “open” will need a precise definition: weights, code, data, training recipe and commercial rights may not all be equally open.
Sanctions and export controls also do not translate into a simple story of permanent technical delay. Chinese labs can still improve algorithms, data, training systems and inference software. K3 does not establish that China has reached parity across every AI capability, but it does weaken the claim that open Chinese models are automatically half a year behind the best closed American models.
How to access Kimi K3 today
Kimi K3 is available through:
- Kimi.com for the consumer-facing experience.
- Kimi Work for research, documents, dashboards and agentic knowledge work.
- Kimi Code for software engineering workflows.
- Kimi API quickstart for developers.
The API is OpenAI-compatible at https://api.moonshot.ai/v1. The launch documentation says K3 uses always-on thinking and currently exposes reasoning_effort="max"; lower effort levels are planned for later updates. It also supports tool calls, structured output, JSON mode, context caching and vision input.
At the listed rate, one million cached input tokens costs $0.30, one million uncached input tokens costs $3.00 and one million output tokens costs $15.00. Long-context projects should estimate output carefully: a cheap cache hit does not make an uncontrolled agent loop cheap.
What developers should test before production
- Long-context retrieval: put a real repository or document archive into the context and measure whether K3 can find the right evidence, not just quote nearby text.
- Tool-call reliability: test malformed arguments, retries, parallel calls, timeouts and state restoration.
- Visual feedback: give it screenshots from the actual product, browser or game and count how many iterations are needed to reach an acceptable result.
- Cost and latency: compare cache-hit rates, output length, time to first token and total task time against the model you use now.
- Safety and privacy: check data retention, access controls, logging and whether external tools can move data outside your boundary.
- Weight-release readiness: if you plan to self-host, wait for the license, reference implementation, quantization formats, hardware guidance and independent inference results.
FAQ
What is Kimi K3?
Kimi K3 is Moonshot AI’s flagship model with 2.8 trillion total parameters, a 1,048,576-token context window, native visual understanding and a sparse Mixture-of-Experts architecture. It is designed for long-horizon coding, reasoning, knowledge work and tool-using agents.
Is Kimi K3 open source?
Moonshot AI says full Kimi K3 weights will be released by July 27, 2026. Until the weights, license and supporting code are available, it is more precise to call K3 an announced open-weight model rather than a fully reproducible open-source stack.
Is Kimi K3 better than GPT-5.6 Sol or Claude Fable 5?
Not across every benchmark. K3 leads several reported coding and agent tests, while GPT-5.6 Sol and Fable 5 lead other evaluations. Model choice should follow the task, tool environment, price, latency and privacy requirements.
What is Kimi Delta Attention?
Kimi Delta Attention is a hybrid linear-attention mechanism intended to make long-sequence processing more efficient. Moonshot reports up to 6.3× faster decoding in million-token contexts, but independent tests are still needed to establish how that result transfers across hardware and workloads.
What does 2.8 trillion parameters mean for local deployment?
It refers to total model capacity, not the number of parameters used for every token. K3 routes each token through 16 of 896 experts, but the full system still has demanding memory and communication needs. Moonshot recommends large accelerator clusters for high-performance inference.
Final verdict
Kimi K3 is a major release because it combines scale, long context, native multimodality and agent-focused engineering in one system. Moonshot's reported results show a model that is competitive with top proprietary models in coding, browsing and long-horizon tasks, even though it does not lead every evaluation.
The July 27 weight release will decide how historic this launch becomes. If developers can run K3 with a practical stack, a clear license and reasonable hardware economics, it could be a genuine open-frontier turning point. If deployment proves too expensive or the results are difficult to reproduce, K3 will still matter as a demonstration of how far model architecture and training efficiency have moved.
For now, the fairest headline is also the most interesting one: Kimi K3 has not ended the frontier-model race—but it has made the open side of that race impossible to ignore.
Related guides
- July 2026 Frontier Model Comparison: 16 New Models from GPT-5.6 to Bonsai 27B — Complete Guide with Pricing & Benchmarks
- GLM-5.2: Z.ai's 753B Open-Weight Model That Tied GPT-5.5 at 1/7 Cost — June 2026
- GPT-5.6 Sol vs Terra vs Luna: OpenAI's 3-Tier Frontier Family Explained — July 2026 Deep Dive
Sources and further reading
- Moonshot AI: Introducing Kimi K3 — Open Frontier Intelligence
- Kimi K3 API quickstart and model documentation
- Official Kimi K3 API pricing
- Artificial Analysis: Kimi K3 intelligence, speed and cost tracking
- Arena AI code leaderboard
- Frontend Code Arena Kimi K3 announcement
- Moonshot AI Kimi Code on GitHub
Editorial note: Benchmark values in this guide are Moonshot-reported unless explicitly identified as third-party. Recheck scores, model names and rankings before publication because frontier leaderboards change quickly.
Related reading
Keep reading
Related reading
⚡ Daily AI Model Drop — Get Kimi K3 benchmarks before Twitter
Join 2,400+ AI engineers. 1 email/day, no spam, unsubscribe anytime