AERIOXFLUX
Frontier Labs
Frontier Labs · meta

Meta Priced Its Best Model at Ten Cents a Million Tokens

Muse Spark 1.3 landed at #6 on the Artificial Analysis index with a 1M-token context window and charges $0.10 in, $0.20 out. Meta is not selling inference. It is buying position.

Flux Desk·2026-09-08·5 min read

Meta Superintelligence Labs released Muse Spark 1.3 on September 2. The capability story is respectable. The pricing story is the one that should make competitors uncomfortable.

Muse Spark 1.3 sits at #6 of 636 models on the Artificial Analysis Intelligence Index, carries a 1M-token context window, and takes text, image, and video input. On agentic and coding work it posts 75.4% on DeepSWE 1.1, 88.8% on Terminal-Bench 2.1, 59.4% on SWEAtlas CodeBase QnA, and 98.5% on long-context retrieval. Meta shipped it into Muse Code — its terminal coding agent — and the Meta Model API on the same day.

It costs $0.100 per million input tokens and $0.200 per million output tokens. Cached input is $0.002 per million.

Put that number next to the others

Qwen3.8-Max charges $2 in and $6 out. Gemini 3.8 Flash, itself positioned as the cheap agentic option, launched at $0.75 in and $3.75 out — and that is introductory pricing that rises in January. Frontier-tier models from OpenAI and Anthropic sit an order of magnitude above that again.

Muse Spark 1.3 is roughly 7.5× cheaper on input and nearly 19× cheaper on output than the model Google just marketed as its value play. Against a $2/$6 model it is 20× and 30×.

A top-ten model does not arrive at a twentieth of the going rate because someone found a clever kernel. Inference has real marginal cost — GPUs, power, memory bandwidth — and none of those got 20× cheaper in the last quarter. This is a price set by strategy, not by cost.

What Meta is actually buying

Meta is the only frontier lab whose core business does not need inference revenue. Google sells cloud. Microsoft sells cloud. OpenAI and Anthropic sell tokens because tokens are the company. Meta sells advertising against attention, and has spent two decades demonstrating a willingness to give away infrastructure that makes its competitors' moats less valuable.

That was the entire logic of open-weight Llama: commoditize the layer underneath you so nobody can charge rent on it. Muse Spark 1.3 is the same play run through an API instead of a weights download — and, notably, it is not open weights. Meta wants the usage, the telemetry, and the developer habit. It just does not want the margin.

The specific target is legible from the benchmark selection. DeepSWE, Terminal-Bench, CodeBase QnA, long-context retrieval — that is not a chatbot scorecard. That is the workload profile of a coding agent running for hours, and the p95 first-token latency of 7.09 seconds confirms it: this model is tuned for long-horizon work where a seven-second start is irrelevant because the task runs for forty minutes.

Why agentic workloads make the price a weapon

Cheap-per-token matters most exactly where Meta aimed.

A chat turn costs a few thousand tokens. An agentic coding session costs millions — the model reads the repository, reasons, calls tools, reads the results, and does it again a hundred times. Token consumption in agentic work is one to three orders of magnitude above conversational use, and it is the fastest-growing category of inference demand in the industry.

That means the price differential does not scale linearly against a competitor. It compounds. A team burning $40,000 a month on agentic coding at $2/$6 pricing is looking at something closer to $1,500 on Muse Spark 1.3 for a model six slots down a leaderboard most of them have never opened. For a lot of workloads — CI triage, dependency upgrades, test generation, migration sweeps, documentation — the marginal quality gap does not survive contact with a 25× cost gap.

Anthropic understood this dynamic well enough to spend its own September release on it, cutting cache reads 75% while holding sticker price. That is the defensive version of the same insight: agentic workloads are mostly cache, so make cache cheap. Meta's version is less surgical. It just made everything cheap.

The part that should be treated skeptically

Introductory economics are not permanent economics, and there is no public commitment that this price holds.

Google was explicit that 3.8 Flash pricing rises on January 1. Meta said nothing comparable, which cuts both ways — no announced increase, and no announced floor. Anyone building a business model on $0.10 input is building it on a number one company controls and can revise with a blog post.

The second caution is Meta's own track record on continuity. Muse Spark 1.1 opened to developers earlier this year, the versioning has moved quickly, and Meta's AI org has been reorganized more than once in the last eighteen months. Cheap tokens are worth less when the deprecation horizon is unclear.

The third is that #6 is #6. Meta's chief AI officer described the model as edging closer to top competitors, which is an honest framing and also an admission. On the hardest agentic coding evaluations the frontier still leads. If your workload is the one where the top model's marginal advantage actually pays for itself, price is not your binding constraint.

What it changes anyway

The floor moved. That is the durable outcome regardless of what Meta does next.

Every lab pricing an agentic tier now has to explain a number that is 10-20× a credible top-ten model, to a buyer who can run both against the same repository in an afternoon and read the diff. That conversation used to be about capability. It is now about whether the capability delta is worth an order of magnitude, and for most production workloads the honest answer is that nobody has measured it.

They are about to. Muse Spark 1.3 is cheap enough that running the comparison costs less than the meeting about whether to run it.

#meta#muse-spark#inference-pricing#agentic-ai#superintelligence-labs

The state of AI, in flux.

The directory + magazine for AI tools and the workflows people use to make money with them.

🔥 The Sauce Drop

The week's highest-earning AI workflows, in your inbox.

Some outbound links are affiliate links — Flux may earn a commission at no cost to you; this never affects rankings. Earnings figures are self-reported and not guarantees of income; most people earn less, some earn nothing.