Two researchers have independently used advanced large language models to decode long-unsolved Enigma messages—cryptographic puzzles that have remained cracked for decades. Developer Carter Leffen deployed OpenAI's latest model, Astra, to search through archived Enigma communications and successfully deciphered a message that stumped researchers since 2005. Meanwhile, cryptanalyst Jack Willis used Anthropic's Claude Opus 5 model to break a separate uncracked message. These breakthroughs represent a significant demonstration of AI reasoning capabilities applied to real-world historical problems.
The Enigma machine generated the encrypted communications used by Nazi Germany during World War II. While mathematician Alan Turing and his cryptanalytic team at Bletchley Park built the Bombe—an early computing machine—to translate Enigma messages during wartime, some archival communications remained unbroken due to transcription errors or encoding mistakes by original operators. Of the original body of messages, only seven remain unsolved, plus one message where plaintext is known but the encryption method remains undetermined.
Leffen instructed Astra to locate an unbroken Enigma message within existing databases and decode it. The model executed a multi-step analytical process: it conducted its own archival research, identified contextual clues within the encrypted text, constructed a functional simulator of the Enigma machine itself, and ultimately recovered the original plaintext of the message that had confounded researchers for two decades. Leffen even used Astra to develop an interactive website detailing the entire problem-solving methodology.
Frode Weirerud, a retired electrical engineer who maintains Crypto Cellar—an extensive resource containing cryptographic documentation and a searchable message database—validated Leffen's solution. Weirerud expressed astonishment at the achievement, noting that Astra's approach mirrored professional cryptanalytic and archival research practices. He observed that the model accomplished in two days work that would typically require a human researcher several weeks or months, citing his own experience spending weeks analyzing the German federal archives files that Astra referenced.
One intriguing detail emerged in Astra's processing logs: the model referenced materials from a private collection not publicly hosted on Weirerud's website. The source of these materials remains unclear—they may have been accessed from other researchers' online publications or potentially from German government public archives, though Weirerud acknowledged uncertainty about the actual data access path.
On September 21, cryptanalyst Jack Willis contacted Weirerud with news of a separate breakthrough. Willis had employed Anthropic's Claude Opus 5 model to decode a different unbroken message. This effort required more direct human guidance than Leffen's experiment; Willis provided Claude with specific information about a particular officer's name signature, which the model leveraged to ultimately crack the encryption. This approach highlighted how AI can incorporate human-provided historical context to accelerate cryptanalytic work.
These successes showcase capabilities that extend well beyond the original Turing test—the theoretical experiment examining whether artificial intelligence can demonstrate intelligence indistinguishable from humans. Instead, these models demonstrated sophisticated multi-stage reasoning: conducting research, synthesizing information across disparate sources, building working models of complex systems, and executing precise analytical work. The achievements underline how contemporary language models can tackle specialized historical research problems requiring domain expertise, archive knowledge, and systematic problem-solving methodology.
Gist is a free AI reader for your browser, iPhone, and Android. Get concise summaries and key takeaways from any article or podcast.
Get Gist — Free