AERIOXFLUX
AI Tools
AI Tools · coding

Cognition Built Its Best Coder on a Chinese Base Model

SWE-2 is post-trained from Moonshot's Kimi K3 and lands within a point of Fable 5.1 on FrontierCode while costing 64% less. The frontier lab that can't afford pretraining just borrowed one.

Flux Desk·2026-09-11·5 min read

Cognition shipped SWE-2 on September 10. It is the company's most capable coding model, and it is post-trained from Kimi K3 — Moonshot AI's 2.8-trillion-parameter open-weight base, which had already been through extensive reinforcement learning for agentic coding before Cognition touched it.

The headline results:

  • FrontierCode 1.1 Main: 50.0% — within one point of Fable 5.1, at 64% lower cost
  • DeepSWE 1.1: 73.0%
  • Terminal-Bench 2.1: 92.8%
  • Terminal-Bench 4: 27.3%

Cognition also reports that SWE-2 medium beats its predecessor while costing 81% less on average and using 58% fewer turns. It launched in Devin Desktop and Devin CLI, rolling out to Devin Web and Fusion. No standalone API, no published pricing, no weights.

Two of those numbers deserve to be read against each other, and one deserves to be read against the whole industry.

The 92.8 and the 27.3

Terminal-Bench 2.1: 92.8%. Terminal-Bench 4: 27.3%.

Same family of benchmark, same model, a 65-point gap. On Terminal-Bench 4, Fable 5.1 scores 55.8% and GPT-6 Astra scores 57.9% — roughly double SWE-2.

The honest reading is that Terminal-Bench 2.1 is saturated and has been optimized against, by everyone, for long enough that 92.8% conveys very little. Terminal-Bench 4 is harder, newer, and had less opportunity for anyone to tune toward it. Where a model sits on the new benchmark relative to the old one is a reasonable proxy for how much of its score is capability and how much is familiarity.

By that measure SWE-2 is a genuinely strong mid-tier agent that is not competitive with the frontier on long-horizon terminal work — and Cognition's own materials do not claim otherwise. The claim is the Pareto frontier, which is a cost-performance claim, not a capability claim. At 64% less than Fable 5.1 for one point less on FrontierCode, that claim holds on its own terms.

The part that matters more than the scores

Cognition is a US company valued at $48 billion as of this month. It just shipped its flagship coding model on a Chinese open-weight base.

This is not a licensing arrangement or a partnership. Kimi K3's weights are public; Cognition took them, ran its own reinforcement learning — training the medium, high, and max effort levels in a single run — and found 5 to 6 points of headroom on many benchmarks, shifting K3's entire cost-performance curve.

That sequence tells you what pretraining now costs and who can still afford it. Moonshot spent the capital to build a 2.8T-parameter base with agentic RL already baked in. Cognition spent a fraction of that to specialize it and ship a product. Both got paid. Neither had to duplicate the other's work.

Anthropic's threat-intelligence report, published the same week, traced roughly 24,000 fraudulent accounts to distillation operations run through DeepSeek, Moonshot, and MiniMax — and NSA, FBI, and CISA issued a joint advisory about industrial-scale Chinese model distillation. The irony is precise: while Washington warns about Chinese labs extracting capability from American models, the most valuable American coding-agent company built its flagship by extracting capability from a Chinese one. The extraction runs both directions, and only one direction has an advisory.

The strategic read for anyone building on this

Three implications, in descending order of certainty.

Pretraining is becoming a commodity input. The set of organizations that can afford a 2.8T-parameter base from scratch is small and shrinking relative to the set that wants frontier-adjacent capability. Open weights from that small set become the substrate everyone else specializes. Cognition is not an outlier here — Genspark trained its slides model on a MiniMax open-weight base, and the pattern will keep repeating because the economics are unambiguous.

The differentiation moves to the harness and the RL. What Cognition sells is not the model. It is Devin — the environment, the tooling, the effort-level routing, the 58% turn reduction. A base model you can download is not a moat. A reinforcement-learning recipe that finds six points of headroom in someone else's weights, wired into a product developers already use, is closer to one.

Supply-chain risk is now a real procurement question. If your coding agent is post-trained from Kimi K3, your capability floor depends on a Chinese lab continuing to release open weights. That is a policy variable, not an engineering one. Beijing has tightened and loosened open-weight release posture before. An American enterprise standardizing on SWE-2 is taking a geopolitical dependency it probably has not modeled.

What Cognition did not publish

No weights, no API, no pricing. SWE-2 is available only inside Cognition's own products.

That is a defensible commercial choice and it also makes the cost claims unverifiable. "64% cheaper than Fable 5.1" is a number about Cognition's internal economics on a model it serves itself, not a rate card anyone can check. Until there is an API with published pricing, the Pareto-frontier claim is an assertion about a product, not a benchmark result the rest of the field can reproduce.

The technical result is real. The base model it stands on says more about where the industry is going than the benchmark table does.

#cognition#swe-2#kimi-k3#moonshot#open-weights

The state of AI, in flux.

The directory + magazine for AI tools and the workflows people use to make money with them.

🔥 The Sauce Drop

The week's highest-earning AI workflows, in your inbox.

Some outbound links are affiliate links — Flux may earn a commission at no cost to you; this never affects rankings. Earnings figures are self-reported and not guarantees of income; most people earn less, some earn nothing.