Huawei Replaced 48,000 Optical Modules With 5,500
The Atlas 960E SuperPoD is the first AI system built around near-packaged optics: 4,096 Ascend NPUs, 8 exaflops of FP8, and 550 kilowatts of power savings that come entirely from deleting pluggable transceivers.
At Huawei Connect in Shanghai on September 17, Huawei launched the Atlas 960E SuperPoD — which it describes as the world's first SuperPoD built on near-packaged optics (NPO), aimed at training and inference for ten-trillion-parameter models.
The headline specifications: up to 4,096 Ascend NPUs in a single pod, 8 EFLOPS at FP8 and 16 EFLOPS at FP4, and up to one petabyte of high-bandwidth memory. Networking is all-optical UnifiedBus — Huawei's Lingqu fabric — combined with its Hi-ONE optical system. Availability is Q3 2027.
The interesting number is none of those.
5,500 versus 48,000
A single Atlas 960E uses 5,500 Hi-ONE modules where a conventional design would need roughly 48,000 800G pluggable optical modules. Huawei says the substitution cuts power consumption by more than 550 kilowatts, doubles mean time between failures, and lifts system availability to 99.8%.
To understand why that is the real announcement, you have to understand what pluggable optics cost a large cluster.
At scale, the network is not a rounding error on the power budget. Every GPU-to-GPU hop outside a node crosses an optical transceiver, each transceiver burns watts continuously whether or not traffic is flowing, and in a fully-connected fabric the transceiver count scales faster than the accelerator count. Operators of large training clusters have been reporting for years that optics are a meaningful double-digit share of network power.
They are also the least reliable component in the building. Pluggable modules are the thing that fails. In a 48,000-module fabric with realistic failure rates, something is dying every day, and a training run that has to checkpoint and restart around network faults loses throughput to an availability problem rather than a compute problem.
Near-packaged optics moves the optical engine adjacent to the switch ASIC instead of into a front-panel cage. You shorten the electrical path, drop the retimers and much of the signal conditioning, and collapse many discrete modules into far fewer integrated ones. The power savings are real and the reliability improvement is arguably more valuable.
Huawei is claiming both, and the claimed availability figure — 99.8% — is a system-level number, which is the number an operator actually cares about.
Why this is the move a constrained vendor makes
Huawei does not have access to leading-edge lithography. That is the premise of every analysis of Chinese AI silicon, and it is correct.
The strategic response, visible now across multiple Chinese efforts, is to stop competing on transistor density and start competing on everything around the transistor. If your accelerator is a node or two behind on process, you lose on per-chip performance-per-watt. You can recover a substantial fraction of that at the system level by winning on interconnect, memory bandwidth, and fabric reliability — because at cluster scale, utilization is determined by how well the parts talk to each other, not by how fast any single part is.
A pod with 4,096 accelerators on an all-optical fabric, a petabyte of HBM, and 550 kilowatts of interconnect power deleted is a coherent expression of that thesis. It is also, notably, not a thesis Nvidia disagrees with — co-packaged and near-packaged optics are on every serious roadmap in the industry, for exactly these reasons.
The difference is urgency. For an unconstrained vendor, optics integration is an efficiency program. For Huawei, it is one of the few axes where being first is possible.
The timing is not accidental
The Atlas 960E was announced days before an expected Trump–Xi meeting in Washington on September 24, and it follows the Atlas 950 SuperPoD by only months.
Announcing a flagship AI system with world-first claims immediately ahead of a leaders' summit is a familiar Huawei pattern and serves a specific function: it is evidence, offered publicly, that export controls have not achieved their objective. Whether that is true is a separate question. That it is being asserted at that moment is the point.
The caveats that matter
Q3 2027 is a long way out. Every number above is a specification, not a benchmark, and the gap between announced pod configurations and delivered cluster performance is where this industry's credibility is usually spent. Huawei has shipped Ascend hardware at volume — the company has been reported to be doubling output of its top AI chip — but a first-of-its-kind optical architecture arriving in roughly a year is a schedule, not a fact.
FP4 exaflops are a marketing unit. Sixteen exaflops at FP4 is a real capability and a number that should never be compared against anything quoted at a different precision. The FP8 figure — 8 EFLOPS — is the one to carry forward.
Nothing here addresses software. The persistent gap between Ascend and Nvidia is CUDA, not FLOPS. A better fabric does not port a training stack.
What to watch
Whether anyone outside China buys one. Export dynamics run both directions, and a SuperPoD sold into the Gulf or Southeast Asia would say more about competitive position than any specification.
Real availability numbers from a deployed pod. 99.8% claimed at launch is a design target. Twelve months of operator data is evidence.
Whether Nvidia's optics roadmap accelerates. If co-packaged optics arrive in a Vera Rubin successor earlier than previously signaled, Huawei will have set the pace on the one axis where it currently can.
