AERIOXFLUX
← Frontier Labs
Frontier Labs · openai

GPT-6 Astra Broke an Enigma Message Nobody Could Read

An 82-letter German Army signal from July 1941 had sat on the unbroken list since 2005. OpenAI's flagship picked it, built its own Bombe, and recovered the key in about ten hours.

Flux Desk·2026-09-23·5 min read

For two decades, message Nr. 172 sat on a short list of German Army Enigma traffic that nobody could read. The indicator was MVUEH. It was 82 letters long, sent on July 10, 1941, and received by the quartermaster of the SS-Totenkopf division at 17:30. Frode Weierud, who maintains the CryptoCellar archive of wartime Enigma material, had listed it as unbroken since 2005.

It is not on that list anymore. On September 15, Carter Leffen, a product developer at Bloomberg, sent Weierud a key recovered by GPT-6 Astra. Weierud checked it and confirmed that both the key and the plaintext were correct. The message is routine. Translated, it asks for the route of march, reports the sender is in Rosenow, and requests an immediate reply by radio. It is signed by a sender tentatively read as Waschbusch.

What the message says doesn't matter. What matters is how it got read.

The break

Two things had kept MVUEH unbroken. The published ciphertext contained transcription errors, which quietly poison any method that assumes every letter is right. And the machine's left-hand wheel turns over at the 72nd letter. Left-wheel turnovers are rare, and they break the assumptions a standard attack relies on near the end of a message.

The way in was a crib, meaning a guessed stretch of plaintext. A related message from the same day, SIPVX, was broken in 2017, and it contains the repeated town name ROSENOW ROSENOW. Using those 14 letters as a crib against MVUEH shrinks the search space to something a computer can work through. Enigma's best-known design flaw, that no letter can encrypt to itself, rules out many of the positions where the crib could sit.

The recovered setting was Enigma I with reflector B, wheel order II-V-III, ring settings H-M-F and ten plugboard pairs. Weierud notes that the wheel order differs from the one used by the other two known keys for that date. That helps explain why nobody found it: anyone who assumed MVUEH used the same day's key was searching the wrong space.

According to the reporting, the model wrote its own Enigma simulator and Bombe software in Python and C++, and the run took about ten hours of model time. Separate implementations then re-checked the full search, all 43,016 batches of it.

Who actually did it

The accounts disagree on this point, and it's the most interesting part of the story.

Weierud's write-up says Leffen only asked the model to try the unbroken messages on the CryptoCellar site. By his account, Astra chose MVUEH, noticed the connection to SIPVX, picked the crib and built the tooling itself. Weierud called the autonomy "the most astonishing thing about this break," and wrote that what it did in two days would take a human researcher weeks or months.

Leffen's own case study is more careful. It calls the effort a researcher-led investigation: agents handled archive analysis, comparison against solved messages and running the cryptanalysis code, while he set the objectives and made the procedural decisions. He also joked that building the website explaining the break took 99 times more effort than the break itself.

Both accounts can be true. The model did the technical work, and a person decided where to point it. That split is common now in AI research results, and headlines that say "the AI did it alone" usually leave out the person who aimed it. The honest reading: the person spent effort choosing the target and deciding what counted as success. The machine spent effort everywhere else.

Why it matters

This is not a cryptographic threat. Enigma was broken during the war. Its weaknesses are well documented, and hobbyists have been clearing the remaining unbroken messages with commodity hardware for years. Nothing about modern encryption is affected.

What changed is the shape of the work. Clearing an Enigma message takes archive research, historical judgment about which messages are related, careful handling of corrupted ciphertext, and custom search code built for one message's quirks. Until recently, a model could help with any one of those steps. Here a single system connected all of them and produced a result that an independent domain expert accepted as soon as he saw it.

It also wasn't a one-off. Shortly afterward, Anthropic's Claude Opus 5 broke another listed message, Nr. 205/285 (indicator FMNGI), using a crib built around a signature, according to Bruce Schneier's write-up. When two labs' models clear archival puzzles within days of each other, the capability is widespread, not a stunt.

There is an uncomfortable detail too. OpenAI's own system card classifies Astra as its first model at the Critical level for cybersecurity under the Preparedness Framework. At that level, a model can find and exploit previously unknown vulnerabilities in hardened systems with little human direction. An 85-year-old cipher is a harmless showcase for the same traits that got that rating: persistence, building its own tools, and turning partial clues into a working attack.

The math parallel

Mathematics is seeing the same division of labor. Scientific American reports that six mathematicians, meeting after a May workshop at the American Institute of Mathematics at Caltech, used AI agents to build a degree-23 polynomial whose symmetry group is the sporadic group M23. It was the last of the 26 sporadic groups without a known solution to the inverse Galois problem. The AI found candidate geometric surfaces and better coordinates, and the humans turned the workable one into equations. It took three months.

A parallel crowdsourced contest run by Terence Tao's Foundation for Science and AI Research catalogued all 25,000 polynomials corresponding to every group acting on 24 roots. The detail worth noting: the winning team, a group of German mathematicians, used AI only for an upload script.

Put the two stories side by side and the picture is consistent. Models are now useful at the step where searches exceed human scale: large, fiddly spaces that reward stamina and custom tooling. Framing the problem, and deciding when an answer is really an answer, still depends on people. The Enigma break was accepted because Weierud could check it in minutes, and M23 because the polynomial either has that symmetry group or it doesn't.

What to watch

The next test is results that can't be checked by running a decryption or evaluating a polynomial. Enigma and the inverse Galois problem are good showcases because verification is cheap and certain. As labs point these systems at open problems where checking is expensive or ambiguous, the key question becomes who does the checking and how long it takes.

#gpt-6-astra#enigma#cryptanalysis#openai#ai-research

The state of AI, in flux.

The directory + magazine for AI tools and the workflows people use to make money with them.

🔥 The Sauce Drop

The week's highest-earning AI workflows, in your inbox.

Some outbound links are affiliate links — Flux may earn a commission at no cost to you; this never affects rankings. Earnings figures are self-reported and not guarantees of income; most people earn less, some earn nothing.