Back to Blog
Dense glowing neural network cluster rising above a stylized Shanghai skyline at night
AI & ML
Jul 19, 2026
9 Min Read

What Kimi K3 Is, and Why Moonshot AI's New Model Is a Big Deal

Moonshot AI took the stage in Shanghai on Thursday to announce its 2.78-trillion-parameter model. Three days later, the AI community is still talking about it. The launch video has already passed ten million views, the servers crashed within two days, and the company admitted that K3 received more attention than expected, putting a strain on their GPUs. As a result, new subscriptions have been paused.

Kimi.ai post announcing that new subscriptions are temporarily paused due to high demand, while existing subscribers are unaffected and additional capacity is being added.

It’s being called an open-weight release, but the weights haven’t been released yet. Right now, all that’s available is the announcement, a paid API, and a technical report. Everyone online is discussing a model they can only access by renting.

So what’s actually in it?

Let’s start with scale, since that’s what everyone noticed first. K3 uses a mixture-of-experts approach, with 2.78 trillion parameters in total, but only about 104 billion are active for any given token. For comparison, Kimi K2 had 1.04 trillion total and 32 billion active. So, K3 is about two and a half times larger overall and has more than three times the active compute per token. The context window also jumped from 128,000 tokens to a full million.

The architecture is where things get really interesting, and it’s where the main story is. Moonshot designed the model around what they call Kimi Delta Attention, a type of linear attention that runs 69 of the model’s 93 layers. These are mixed three-to-one with 24 standard latent-attention layers, which help keep precise global recall. Linear attention has been considered a future goal in this field for years because it’s efficient but has usually been less accurate. No one had built a model this large that relies on it so much. Moonshot claims this approach allows for decoding up to 6.3 times faster when working with million-token inputs.

On top of that, there’s a feature called Attention Residuals, which allows any layer to access information from all previous layers. This reportedly adds less than 2% extra compute but improves training efficiency by about 25%. The model also uses expert routing, with 896 experts in total, 16 of which are active at a time plus two shared ones. This setup is very sparse and uses a balancing method based on score quantiles instead of manual tuning. For vision, the model includes a 401-million-parameter encoder that was trained from scratch alongside the language model, rather than being added later.

Diagram of Kimi K3's architecture showing attention layers, attention residuals, and expert routing

Altogether, Moonshot says K3 delivers about 2.5 times more capability per unit of training compute compared to K2. That result is more important than just the parameter count. It shows they believe smart techniques can still outperform sheer size.

What’s missing from the paper is just as noticeable. There’s no mention of the pretraining token count, total FLOPs, GPU-hours, or the hardware used for the cluster. For a report with so much detail on architecture, the lack of cost information stands out. That’s the number everyone wants, since the debate about whether China did this more cheaply depends on it.

The post-training process is also unusual. Instead of running one long reinforcement-learning session, they trained nine separate specialists: general tasks, general agents, and coding agents, each at three different levels of reasoning effort. These were then combined into a single model using multi-teacher distillation. All the agent training took place in a sandbox environment built on microVMs. In total, they created over 51 million of these sandboxes during training and evaluation. That’s 51 million temporary computers, just to teach the model how to use a terminal.

Now, the scores.

Artificial Analysis ranked K3 fourth on its intelligence index, behind Anthropic’s Fable 5 and OpenAI’s GPT-5.6 Sol, but ahead of Claude Opus 4.8. That’s the highest ranking ever for a model that might release its weights publicly. In LMArena’s frontend coding arena, where developers choose between two outputs without knowing which is which, K3 took first place with about 1,679 Elo. It beat Fable 5 in most direct matchups. On a private long-term knowledge-work test, it scored 732 Elo points higher than its own predecessor from just a few months ago.

Bar chart ranking Kimi K3 fourth on the Artificial Analysis intelligence index, behind Claude, Fable 5, and GPT-5.6 Sol

It’s important to keep two things in mind, though. Many of the coding results came from Moonshot’s own agent harness, while competitors used different setups. Just changing the harness can shift scores by about twenty points. There’s also a cost to K3’s approach: it always runs with maximum reasoning, with no way to adjust it. Artificial Analysis measured it using about 130 million output tokens during evaluation, compared to a median of 63 million. Simon Willison did his usual fun test, asking it to draw a pelican riding a bicycle, and saw it use about 13,000 reasoning tokens and spend about 25 cents. The hallucination rate was around 51%, which actually increased even as accuracy improved.

A post by Simon Willison about Kimi K3, noting lessons from the pelican benchmark and long-conversation agentic tool use

This leads us to the price, which is why regular developers were so surprised, not just researchers.

It costs three dollars per million tokens in and fifteen dollars per million tokens out. For a Chinese lab, that’s a big jump. K2 was 55 cents and $2.20, so this is about five to seven times higher. K3 is now the most expensive model any Chinese company has sold. Still, it’s only about a third of the price of the model above it on the leaderboard, even though it beats that model in blind coding tests. The numbers speak for themselves. This isn’t ‘cheap Chinese AI’ anymore. It’s more like: we’re at the cutting edge now, and we’ll charge for it, just not as much as you do.

Bar chart comparing Kimi K2 and Kimi K3 API pricing per million input and output tokens

On r/LocalLLaMA, reactions fell into three groups. Some people celebrated that downloadable models are now only days behind closed labs instead of months. Others joked that almost no one can actually run a 2.8-trillion-parameter model at home. The third group, who are actually using it, said the real draw isn’t beating Fable 5; it’s the price and the fact that it doesn’t refuse requests. The old debate about whether OpenAI and Anthropic have a real advantage came back, with the best arguments saying the real advantage isn’t the architecture, but the data pipeline and the computing power to run it at scale.

On Hacker News, people quickly focused on the benchmarks, questioning whether test data had leaked into training, which is a common suspicion when results are this strong. One discussion suggested that the real lesson is to invest in the companies making the tools, not just those using them. The gap between open and closed models has shrunk from six or nine months to just three to five. Moonshot also impressed with tough demos, like having the model write a Triton-style GPU compiler from scratch and design a chip on its own.

The markets reacted as you might expect. Chip stocks dropped sharply, the semiconductor index lost nearly 10% and briefly entered bear-market territory, and the total dollar amounts involved reach into the trillions, depending on how you count. If this sounds familiar, it’s because DeepSeek’s R1 wiped $589 billion off Nvidia in one day eighteen months ago, the largest single-day loss for any company. This time, investors seemed to learn a bit from the past, and the recovery was quicker.

Some argue that the selloff was actually the “wrong reaction”. Running K3 requires more than a terabyte and a half of high-bandwidth memory and many accelerators connected together. This model doesn’t reduce the need for expensive hardware. In fact, making inference cheaper usually leads to more inference, not less.

The financial side is where things get uncomfortable for the American tech industry. Moonshot was valued at about 4.3billionlastDecember.ByMay,thatjumpedto4.3 billion last December. By May, that jumped to20 billion. Now, there’s talk of a pre-IPO round aiming for 50billionbeforeaHongKonglisting,andreportssaydailysalesincreasedsixfoldafterThursday.Theirannualrecurringrevenuerosefromabout50 billion before a Hong Kong listing, and reports say daily sales increased sixfold after Thursday. Their annual recurring revenue rose from about100 million in March to over 300millionbymidJune,withAPIlicensingmakingupmorethan70300 million by mid-June, with API licensing making up more than 70%. At a50 billion valuation on 300millioninrevenue,thatsamultipleofover160x,whichisextreme.ButcomparedtoAnthropics300 million in revenue, that's a multiple of over 160x, which is extreme. But compared to Anthropic's65 billion round at a 965billionvaluationinMay,orOpenAIsvalueabove965 billion valuation in May, or OpenAI's value above850 billion, the comparison raises tough questions. A smaller lab, working under export controls, reached fourth place on the leaderboard.

Moonshot AI valuation and ARR shown as side-by-side charts, highlighting how fast Moonshot's growth has been while still far below US frontier-lab valuations.

It’s still unclear whether this means American valuations are unrealistic or if catching up is simply less expensive than leading. It’s probably a bit of both. Catching up usually costs less.

The policy debate also intensified that week. Xi used the Shanghai conference to position China as the leader in open AI, which stands out since American labs only offer API keys. Behind the scenes, there’s a real divide: AMD, Vercel, and Ollama are reportedly advocating for downloadable models to remain freely available, no matter where they come from. Meanwhile, Anthropic and OpenAI are pushing in the opposite direction. If small teams and startups around the world end up relying on Chinese open models by default, that’s a gradual shift that no one officially chose.

For now, everything depends on whether the weights are actually released. Until that happens, we don’t know the license terms, how the model works outside of Moonshot’s own setup, or if the benchmark numbers hold up when tested independently. Three days in, the truth is that something important has happened, but we’re still waiting for all the details.

It will be interesting to see how the price changes once other providers can host the model themselves. That figure will reveal more than any leaderboard.

Join the Conversation

This dispatch is part of an ongoing series on the future of intelligence. Share your perspective or subscribe for more.

Weekly dispatches. No spam. Ever.