AERIOXFLUX
Frontier Labs
Frontier Labs · chinese labs

Alibaba Shipped the API First and the Weights a Week Later

Qwen3.8-Max is 2.4 trillion parameters, tops the Chinese field on Arena, and will be open-weight — but at 1.2TB of checkpoint, open-weight has stopped meaning anything you can run.

Flux Desk·2026-08-06·5 min read

Alibaba put Qwen3.8-Max into global API access on August 3, through Alibaba Cloud Model Studio and QwenWork, its workplace agent platform. The weights follow next week, on Hugging Face and ModelScope. That ordering — paid endpoint first, checkpoint second — is the most interesting thing about the release, and it is the part that will get read as a footnote.

The model itself is the largest Qwen ever built: 2.4 trillion total parameters, roughly 95 billion active per token under a mixture-of-experts routing scheme. Context runs to 1 million tokens. It accepts text, image, and video, and returns text. API pricing is $2 per million input tokens, $6 per million output, and $0.25 for cached input.

Alibaba's own benchmark tables put it at or above Anthropic's Fable 5 on several evaluations. On Arena.AI, the crowdsourced preference board, it became the highest-ranked Chinese model for text on arrival, while still sitting below several Anthropic entries. Both of those things can be true, and the gap between them is the usual gap between a lab's chosen suite and a preference market.

The release order is the strategy

For most of this year, Alibaba kept its top tier closed. The open Qwen releases were the mid-sized checkpoints; the Max-class flagships stayed behind the API. Shipping Max weights at all is a reversal.

But look at how the reversal is staged. The API went live on a Monday. The weights go live roughly seven days later. In that window, Alibaba owns every token anyone runs through the model, at $2 in and $6 out, and every benchmark writeup, every Arena placement, and every "we tested it" review is generated against the hosted endpoint on Alibaba's infrastructure.

By the time the checkpoint lands, the narrative is set and the integration decisions have already been made. Developers who wired against Model Studio during the launch week do not un-wire because a tensor file appeared.

This is the inverse of how open-weight releases used to work. Moonshot put Kimi K3 on Hugging Face on July 27 as the primary event — 2.8 trillion parameters, download first, ask questions about the licence later. Alibaba has taken the same asset class and run it as a product launch with an open-source tail. The weights are the press release; the API is the business.

1.2 terabytes is not a download

The second thing worth sitting with: at this scale, "open weights" and "self-hostable" have fully decoupled.

Do the arithmetic. 2.4 trillion parameters at four-bit precision is roughly 1.2TB of weights before you allocate a single byte to KV cache, activations, or the million-token context the model advertises. Fitting that means aggregate accelerator memory in the multi-terabyte range, spread across nodes, joined by an interconnect fast enough that expert routing does not spend its life on the wire.

That is a datacenter deployment. It is not a workstation, a rack, or a startup's reserved capacity. The set of organizations that can serve Qwen3.8-Max from downloaded weights and the set that could have negotiated a private endpoint anyway are close to the same set.

The genuinely runnable artifact in this release is the other one: Qwen3.8-27B, which lands around 17GB at Q4_K_M and fits comfortably on a single 24GB card. That is the checkpoint that will actually appear in self-hosted stacks, on-prem inference boxes, and hobbyist rigs. It is also the one that generated almost none of the coverage.

So the release has two audiences and one headline. The headline number belongs to a model almost nobody will run locally. The model people will run locally gets a footnote.

What open weights are for now

None of this makes the release cynical. Publishing a 2.4T checkpoint under permissive terms is a real transfer of capability — to researchers who want to probe a frontier-scale MoE without a lab badge, to sovereign programs assembling domestic stacks, to anyone building distillation pipelines who now has a very large teacher.

But the function has shifted. Open weights at 27B are an access mechanism: they put capability in hands that could not otherwise buy it. Open weights at 2.4T are a positioning mechanism. They establish that the lab is not gatekeeping, they seed the research literature, they pressure competitors who charge for comparable capability, and they cost the releasing lab very little — because the people who could exploit the download were never going to be its retail customers.

LG AI Research made the same trade last week from a different direction, putting K-EXAONE 2.0 — 750 billion parameters, 37 billion active — on Hugging Face under Apache 2.0 for South Korea's national foundation model program. Sovereignty you cannot run yourself is not sovereignty; the same logic applies to openness. At some parameter count, a licence stops being the binding constraint and physics takes over.

The pricing line matters more than the licence

The number in this release with the shortest path to somebody's P&L is $2 in, $6 out.

That is Max-class capability priced roughly where mid-tier Western models sat six months ago, from a lab that is simultaneously giving the weights away and running a workplace agent platform on top. OpenAI cut GPT-5.6 Luna by 80% on July 30, to $0.20 in and $1.20 out, three weeks after shipping it — a repricing that read as a response to exactly this kind of pressure at the volume end.

Qwen3.8-Max applies it at the capability end instead. If a 2.4T multimodal model with a million-token window clears $6 per million output tokens, the question for every frontier lab charging a multiple of that is what the multiple buys. Some of them have a good answer — reliability, tooling, enterprise contracts, evaluation depth the benchmark tables do not capture. Some of them have been charging for scarcity.

The weights drop next week. The pricing already landed.

#alibaba#qwen#open-weights#moe#inference-pricing

The state of AI, in flux.

The directory + magazine for AI tools and the workflows people use to make money with them.

🔥 The Sauce Drop

The week's highest-earning AI workflows, in your inbox.

Some outbound links are affiliate links — Flux may earn a commission at no cost to you; this never affects rankings. Earnings figures are self-reported and not guarantees of income; most people earn less, some earn nothing.