AERIOXFLUX
Tech & Culture
Tech & Culture · chips compute

Positron Raised $875M to Bet Against HBM

A $5 billion valuation for an inference chip whose entire thesis is that high-bandwidth memory is the wrong place to spend money — and that commodity LPDDR5X is good enough to break Nvidia's grip on serving.

Flux Desk·2026-09-11·5 min read

On September 10, Positron AI announced $875 million at a $5 billion post-money valuation. In February the same company raised $230 million at $1 billion. Seven months, five times the price.

The round arrived in two pieces — a $375 million Series C followed by a Series C-1 of up to $500 million — co-led by NEA, Atreides Management, Valor Equity Partners, Andra Capital, Dylan Patel's SemiAnalysis Capital, and Jim Clark, who founded Silicon Graphics and Netscape. The money funds three things: the tapeout of a chip called Asimov, a 2 MW-plus engineering data center and emulation platform, and the production ramp of a system called Titan.

What makes this interesting is not the multiple. It is the architectural bet the multiple is paying for.

The bet: skip HBM entirely

Asimov pairs Positron's compute architecture with 288 GB to 2,304 GB of memory per chip using commodity LPDDR5X — the same class of memory in phones and laptops — instead of high-bandwidth memory.

Every serious AI accelerator on the market does the opposite. HBM is the default because it delivers bandwidth per watt that nothing else touches, and bandwidth is what keeps thousands of matrix units fed. The cost of that default is brutal: HBM requires advanced 2.5D packaging, it is supply-constrained to three vendors, and it has been the single hardest line item to secure for three straight years.

Positron's argument is that for inference specifically, the binding constraint is not bandwidth per watt — it is capacity. A model you cannot fit in memory has to be sharded across more chips, and every shard boundary costs interconnect, latency, and utilization. If one chip can hold 2.3 terabytes, a workload that needed a rack of HBM parts fits in a fraction of the hardware, and the lower per-byte bandwidth stops mattering because you stopped paying the sharding tax.

Titan, the system built around this, is pitched at 16 trillion-plus parameter models or 10 million-plus token contexts. Those are the two workloads where capacity dominates most decisively.

Why the timing is not a coincidence

On the same day Positron announced, Reuters reported that Chinese AI chipmakers had raised prices 20% to 50% — explicitly because of a worldwide HBM shortage. Huawei's Ascend 950DT is now quoted above 250,000 yuan. A week earlier, CXMT was sampling HBM3E to Chinese accelerator vendors roughly a year ahead of forecast, because the demand is that desperate.

HBM scarcity is not a rumour investors are pricing. It is a spot-market fact visible in accelerator quotes across two continents. A company whose bill of materials routes around the scarce part is, right now, worth a premium for exactly that reason.

The other half of the timing is Oracle. Positron already has 50-plus Atlas racks deployed at Oracle Cloud Infrastructure. That is not a pilot with a logo attached — it is a hyperscaler running production silicon from a company that, in February, was worth $1 billion. Revenue-adjacent proof is what separates this round from the last three years of inference-chip announcements.

The part that is still a promise

Asimov tapes out on TSMC N3P at the end of 2026, with production targeted for the second half of 2027.

That is eighteen months of execution risk on a leading-edge node, and the graveyard of AI silicon is full of companies that had a good architectural argument and a bad tapeout. Positron's current deployment is on the previous-generation Atlas part. The thing that justifies $5 billion does not exist in silicon yet.

There is also a real technical objection to the thesis, and it should be stated plainly: LPDDR5X trades bandwidth for capacity, and some inference workloads are genuinely bandwidth-bound — short-context, high-throughput serving of a small model, where the weights already fit comfortably and every millisecond is a token. On that shape of work, HBM wins and will keep winning. Positron is not claiming the whole market. It is claiming the part of the market that is getting bigger fastest: long contexts and very large sparse models.

Whether that claim holds depends on a fight nobody has resolved. If the industry converges on trillion-parameter mixture-of-experts models with million-token contexts, capacity-first is the correct architecture and Positron is early. If inference consolidates around small distilled models served at extreme throughput — which is the direction price competition has been pushing all year — the bandwidth-first incumbents keep the volume.

What to watch

Three markers, in order.

The tapeout date. End of 2026 on N3P is the first falsifiable claim. Slips here cascade into the 2027 ramp and into whatever Positron has told Oracle.

Whether Oracle expands. Fifty racks is a serious pilot. Five hundred is a platform decision. The gap between those two numbers is where inference-chip startups usually die — not from technical failure, but from a customer who evaluates, publishes a nice case study, and then re-ups with Nvidia because the software stack is where the real switching cost lives.

Whether the HBM shortage eases. This is the uncomfortable one. Positron's differentiation is sharpest while HBM is scarce and expensive. Samsung, SK hynix, and Micron are all adding HBM capacity, and CXMT is coming up underneath them. A world with abundant, cheaper HBM in 2028 is a world where "we route around the scarce part" stops being a moat and becomes a footnote about memory economics.

The $875 million is a bet that the second half of 2027 still looks like right now. That is the whole trade.

#positron#inference#hbm#lpddr5x#nvidia

The state of AI, in flux.

The directory + magazine for AI tools and the workflows people use to make money with them.

🔥 The Sauce Drop

The week's highest-earning AI workflows, in your inbox.

Some outbound links are affiliate links — Flux may earn a commission at no cost to you; this never affects rankings. Earnings figures are self-reported and not guarantees of income; most people earn less, some earn nothing.