Moonshot AI dropped Kimi K3 on July 17, 2026, and the numbers hit hard. The 2.8 trillion-parameter model is the largest open-weight system ever published, and it posts scores that sit within striking distance of Anthropic's Claude Fable 5 and OpenAI's GPT-5.6 Sol on the benchmarks that matter for coding and reasoning work (Moonshot AI, 2026). Within hours, shares of rival Chinese AI firms Zhipu and MiniMax tumbled 27.7% and 16.5% respectively on the Hong Kong exchange (Reuters, 2026).

The model uses Kimi Delta Attention and Attention Residuals, activates 16 of 896 experts through a Mixture of Experts architecture, and delivers a 1-million-token context window with native vision capabilities. It is the first open model in the 3-trillion-parameter class, and the full weights dropped on July 27. The gap between Chinese open-weight labs and US frontier systems has closed faster than most analysts expected.

How does Kimi K3 perform on the benchmarks that matter?

On Terminal-Bench 2.1, an agentic coding test, Kimi K3 scores 88.3 percent, just behind GPT-5.6 Sol's 88.8 and ahead of Claude Fable 5's 88.0 (Moonshot AI, 2026). On SWE-Marathon, a harder repo-level test, Kimi K3 posts 42.0 percent, beating Claude Fable 5's 35.0 and GPT-5.5's 14.0. On DeepSWE, a software engineering benchmark, Kimi K3 scores 67.5, trailing GPT-5.6 Sol's 73.0 but ahead of Claude Opus 4.8's 59.0.

Arena.ai, a blind human-preference benchmarking platform, ranked Kimi K3 first in six of its seven frontend domains, including brand and marketing, reference-based design, data and analytics, consumer product, simulations, and content creation tools (The Deep View, 2026). The model only ranked second in the gaming domain, slotting just behind Claude Fable 5. Artificial Analysis placed Kimi K3's overall intelligence on par with OpenAI's GPT-5.5 and Anthropic's Claude Opus 4.8 (Artificial Analysis, 2026).

Kimi K3 benchmark scores vs frontier models
BenchmarkKimi K3Claude Fable 5GPT-5.6 Sol
Terminal-Bench 2.188.3%88.0%88.8%
SWE-Marathon42.0%35.0%39.0%
DeepSWE67.5%70.0%73.0%
ProgramBench77.8%76.8%77.6%

Why does Kimi K3 matter for the open-weights race?

Because it is open. Claude Fable 5 and GPT-5.6 Sol are closed systems you can only access through APIs. Kimi K3 is downloadable, modifiable, and self-hostable. That means enterprises can run it on their own infrastructure, fine-tune it for their data, and avoid vendor lock-in. The open-weights tier just got a flagship that competes with the best closed models.

  • Kimi K3 is the first open model in the 3-trillion-parameter class with 2.8 trillion total parameters
  • It matches frontier US systems on coding and reasoning benchmarks across 6+ independent evaluations
  • The weights released July 27, 2026 under an open license with MXFP4 quantization support
  • It uses a 1-million-token context window and native vision for multimodal tasks
  • Moonshot AI claims approximately 2.5x improvement in scaling efficiency over Kimi K2 (Moonshot AI, 2026)

What about the distillation accusations?

Anthropic accused Moonshot AI, along with DeepSeek and MiniMax, of running industrial-scale distillation campaigns against Claude models in February 2026. The allegation is that Moonshot used approximately 3.4 million exchanges with Claude through hundreds of fraudulent accounts to extract agentic reasoning, coding, and tool-use capabilities (Anthropic, 2026). DeepSeek generated over 150,000 exchanges, while MiniMax produced over 13 million exchanges across similar campaigns.

Moonshot has not publicly responded to the specific allegations, but the timing of Kimi K3's release, one month after Anthropic's Fable and Mythos models were temporarily withdrawn by the US government over security concerns, adds fuel to the fire. Brian Jackson, principal research director at Info-Tech Research Group, told The Deep View that Moonshot is marketing K3 with new architecture work unique to its lab, but that could be used in conjunction with distillation techniques as part of training a model (The Deep View, 2026).

Anthropic's report details how these distillation campaigns used commercial proxy services with hydra cluster architectures, sprawling networks of fraudulent accounts that distribute traffic across Anthropic's API and third-party cloud platforms. In one case, a single proxy network managed more than 20,000 fraudulent accounts simultaneously, mixing distillation traffic with unrelated customer requests to make detection harder (Anthropic, 2026). The labs used carefully crafted prompts to extract specific capabilities, targeting agentic reasoning, coding, and tool use, the most commercially valuable skills in frontier models.

Kimi K3 may not be evidence enough that China is ahead. But it is evidence that the gap has closed.

The Deep View

Can US labs ignore open-weight competition now?

No. Kimi K3 is not the first Chinese open-weight model to post strong scores. DeepSeek R1 shook the market in early 2025, and DeepSeek V4-Pro posted SWE-Bench Verified at 80.6 percent in April 2026 (DeepSeek, 2026). The pattern is clear: Chinese labs are shipping frontier-class open models at a pace that forces US labs to either open their own weights or risk losing the developer ecosystem. Meta's Llama 5.1 405B shipped in the same week as DeepSeek R2, the first time in over a quarter that the open-weights frontier matched closed-model release cadence.

Gavin Baker, managing partner and CIO of Atreides Management, called the release an inflection point for AI that could spell trouble for Anthropic and OpenAI by breaking up their dominance (The Deep View, 2026). Cisco CEO Jeetu Patel posted that improving models like these could lead to more competitive industry dynamics by creating a gross margin drop for frontier models. David Sacks, co-chair of President Trump's Council of Advisors on Science and Technology, called the release concerning in light of calls to regulate AI in the US.

The broader pattern is unmistakable. Lian Jye Su, chief analyst at Omdia, said Chinese models are gaining traction because they can be deployed far more cheaply than leading US systems, but cautioned that scale does not necessarily mean the best performance by default (Reuters, 2026). Moonshot itself acknowledges that Kimi K3 trails Claude Fable 5 and GPT-5.6 Sol in overall performance. The difference is that developers can download Kimi K3, run it on their own hardware, and modify it however they want, something they cannot do with either of those closed models.

How much does it cost to run Kimi K3?

Through Moonshot's official API, Kimi K3 costs $0.30 per million tokens for cache-hit input, $3.00 per million tokens for cache-miss input, and $15.00 per million tokens for output (Moonshot AI, 2026). Mooncake's disaggregated inference architecture achieves a cache hit rate above 90% in coding workloads, which keeps real-world costs well below the sticker price for most development teams.

Running it locally is a different story. Ryan Fedasiuk, a fellow at the American Enterprise Institute, noted that running a 2.8 trillion-parameter model locally would require hundreds of thousands of dollars of computing equipment (Reuters, 2026). Kimi K3 uses MXFP4 weights with MXFP8 activations for hardware compatibility, and Moonshot recommends deploying on supernode configurations with 64 or more accelerators. For most organizations, the API route makes more economic sense than self-hosting at this scale.

For context, Claude Fable 5 charges $15 per million input tokens and $75 per million output tokens through Anthropic's API. GPT-5.6 Sol costs $10 per million input tokens and $30 per million output tokens through OpenAI. Kimi K3's $3.00 input and $15.00 output pricing puts it roughly 3 to 5 times cheaper than the closed alternatives for raw token throughput, before you factor in cache hits that drop the input cost to just $0.30 per million tokens (Moonshot AI, 2026). That pricing advantage alone will pull developer attention.

$0.30per million tokens for cache-hit input through the Kimi API · /MTok

What happens next in the open-weights race?

Three things. First, enterprises that paused open-weights deployment in Q1 2026 resume. Kimi K3 ships into production pipelines within 60 days, and the open license lets teams fine-tune without asking permission. Second, DeepSeek R2 becomes the default cheap-reasoning API for agent systems on a budget. Third, the next closed-lab move is more aggressive pricing on agent-runtime. Premium per-call billing, not per-token, becomes the new closed-lab moat. The open-weights category just reset the trajectory.

The real signal is not that Kimi K3 beats GPT-5.6 Sol on every benchmark. It does not. The signal is that a Chinese lab published a model with 2.8 trillion parameters, open weights, and competitive scores across the board. That changes the economics for every company building on AI. When the best open model is this close to the best closed model, the conversation shifts from capability to cost, flexibility, and control. And those are arguments the open-weights side tends to win.

Key takeaways

  • Kimi K3 is the world's first open-weight model in the 3-trillion-parameter class, released by Moonshot AI on July 17, 2026
  • It scores within 0.5 percentage points of GPT-5.6 Sol on Terminal-Bench 2.1 and beats Claude Fable 5 on SWE-Marathon
  • Arena.ai ranked it first in 6 of 7 frontend domains, a 17-place jump from Kimi K2.6
  • Anthropic alleges Moonshot ran 3.4 million distillation exchanges with Claude through fraudulent accounts
  • API pricing starts at $0.30 per million tokens for cache-hit input, making it one of the cheapest frontier-class models available
  • Self-hosting requires hundreds of thousands of dollars in hardware; most teams should use the API

Written by

AI Correspondent

Covers frontier models and the humans behind them. Former ML engineer, reformed speedrunner.

Bottom line

The real signal is not that Kimi K3 beats GPT-5.6 Sol on every benchmark. It does not. The signal is that a Chinese lab published a model with 2.8 trillion parameters, open weights, and competitive scores across the board. That changes the economics for every company building on AI. When the best open model is this close to the best closed model, the conversation shifts from capability to cost, flexibility, and control. And those are arguments the open-weights side tends to win.

What we still don't know

This is a fast-moving story. We update the post as new facts land — and we'll flag it when we do.

Enjoyed this? Pay it forward

A sharp story is worth passing on. Share it with the people who read tech like it matters.

Read moreShare on X