Carter Leffen has announced that he used OpenAI's GPT-6 Astra to restore "MVUEH," a German radio message from 1941. The message, encrypted on the electromechanical Enigma cipher machine, is 82 characters long, and CryptoCellar, which publishes historical records of such messages, now lists it as solved by Leffen. A human led the investigation, while AI agents handled source comparison and the writing of search programs. Our own separate implementation also reproduced the published key's results. To judge what that means, however, we need to separate the information used to guide the search from the information used to check the answer.
Restoring a short request preserved in the records
MVUEH is a message dated July 10, 1941, and both the sender's copy and the receiver's record survive. According to Leffen's public research page, the restored content is a request: "Please give us the marching route. We are at Rosenow. Reply immediately by radio." The place name is repeated, and the message ends with what appears to be a personal name. Which Rosenow is meant, though, has not been determined.
The decryption result is kept distinct from the German as it would be corrected for readability. For example, "Bitte" ("please") appears as "BTTE" in the raw computed output, and the apparent signature remains "WASCHBBSCH." Reading the latter as "Waschbusch" is only a provisional interpretation. In the message, Q stands for CH, and X is used to separate phrases. Tidying the text into proper sentences would make it impossible to check against the original ciphertext.
In CryptoCellar's current list, both the sender's entry for MVUEH ("88 NF/61") and the receiver's entry ("172") are marked as solved by Leffen on September 14, 2026. In a LinkedIn post, Leffen also thanked Frode Weierud, who runs the site, for reviewing the results and offering advice. So this is not a claim that stands only on the presenter's own website; the result has also been recorded by the source's publisher.
As for the time taken, Leffen says he used GPT-6 Astra on its "Extra High" setting and solved the message in about 10 hours. The public research, however, gives an investigation period of two days, September 14–15, with work from source checking through verification overlapping. It also states that no complete log exists that would allow human and model reasoning time to be totaled. His roughly 10-hour figure therefore cannot be equated with the total working time or computing resources used across the whole investigation.
Eighty-five years separate the message's date from this investigation. The research page does not claim a world-first decryption, however, and does not prove that no one could read the message in the meantime.
A repeated place name narrowed the search
The search began with a difficulty that predates the cryptography itself. Some characters differed between the faint handwritten original on the sender's side and the published transcription of the receiver's side. If an incorrect character is entered as fixed, even the correct key will not produce the expected text.
For 12 positions where the reading was uncertain, the research team recorded a finite set of candidates before searching. The combinations came to 13,824. Rather than freely adding characters afterward to make the text fit, they searched within that candidate set. The final restoration differs from the initial independent transcription by 3 characters, and from the separately published receiver-side transcription by 8 characters. These two differences are measured against different comparison sources and should not be confused.
The breakthrough came from a solved message of the same date, "SIPVX." The team assumed that the repeated place name "ROSENOWROSENOW" in that message might also appear in MVUEH. In cryptanalysis, a guessed fragment of the original text like this is called a crib. Here, they used a 14-character crib to search for where it might fit and for the machine settings.
On the Enigma, the rotors advance with each keystroke, changing how letters are substituted. The result is also affected by the plugboard, which swaps pairs of letters, and by the ring settings, which change the relative position of the rotors' internal wiring and their letter display. Matching a single letter correspondence is not enough; the same settings must explain the whole message.
The Enigma also has the property that it never encrypts a letter as itself. If the R in the crib would line up with an R in the ciphertext at that position, that placement can be ruled out. Using constraints like this, the search found the repeated place name at characters 40–53, with text requesting a marching route and an immediate reply appearing around it. The 68 characters outside the assumed 14 were not text supplied as a search condition.
The supporting evidence is best read in four parts
In the search of the message body, the team says it found the candidate without using the separately recorded header. Because the result could then be compared with the header afterward, it could be tested without relying only on the place name fitting neatly.
Verifying MVUEH requires evaluating three things separately: the assumed 14 characters, the meaning of the remaining 68, and the separately recorded header.
| Verification stage | Information used / result obtained | What it confirms and its limits |
|---|---|---|
| Input to the search | The 14-character place name from another message, plus character candidates fixed before the search | Limits the search range. The place name appearing in the output is not independent corroboration in itself |
| Text outside the input | The remaining 68 characters read as an instruction for a marching route and a request for an immediate radio reply | The text holds together beyond the assumed phrase. Linguistic naturalness alone, however, is not conclusive |
| Comparison with a separate record | Header GTA/KCI matches the message-body start position RWD | Narrows the settings that behave equivalently for the body. Does not mean it is the only key in the whole search range |
| Computational reproduction | Recomputing decryption and re-encryption of the 82 characters, and the header, with the published key | Confirms computational consistency. Does not establish the handwriting or that it is the historically unique answer |
The table classifies the search and verification explanation on the public research page, checked on September 18, 2026, by role: input, text outside the input, separate record, and computational reproduction. The figure of 68 characters is the 82-character body minus the assumed 14. What shapes the evaluation of the result is less whether multiple AIs agreed than whether the same hypothesis can be tested against independent information.
The header's role also calls for care. According to the research team, there were 26 combinations of ring settings and start positions that behave equivalently to the restored body, and among them HMF/RWD was the one consistent with the header. Across the full search range the study declared, however, 923 physical keys remained that matched the header. "One of 26 could be chosen" cannot be read as "the header alone proves a unique solution among all candidates."
The published key is reflector B, rotors II, V and III from left to right, ring settings HMF, and a body start position of RWD. The plugboard has 10 pairs, including AC, BE and DG, and the full settings are on the research page. To read the header, set the start position to GTA, enter KCI, then reset to the resulting RWD and process the body.
We also recomputed the published restored ciphertext and key in a separate implementation based on the Crypto Museum wiring tables. The 82-character decryption matched the published verbatim plaintext, and the reverse re-encryption and the recovery of RWD from the header also matched. This confirms the computation of the published answer; it is not a replication of the full key search or of the handwriting in the source.
After all, any key can convert a ciphertext into some string of letters and convert it back with the same settings. A matching re-encryption becomes strong evidence only because it comes with meaningful text that was not assumed and agreement with a separate record.
The AI's role was moving between sources and computation
The difficulty of breaking short messages has been studied since before AI. A paper that Olaf Ostwald and Weierud published in Cryptologia in 2017 explains that for short messages, or those with many errors, scoring methods such as letter frequency become unstable. Wrong candidates can even score higher statistically than the correct plaintext.
With longer text, many word biases and combinations appear, making it easier to judge whether a candidate is close to correct. Short text offers few such clues, and military-specific abbreviations and mistakes get mixed in. So the fact that an AI produced plausible German is no proof of decryption.
In this case, the human set the goal and the course of the investigation, while the agents compared sources, built search programs, and ran experiments in parallel. Unsuccessful searches were also recorded, and results were checked in a separate implementation. The method of using a known phrase is itself old, but the practical role of AI agents lies in being able to move back and forth between reading sources and testing those hypotheses computationally.
That said, the published material is an individual case study, and we could not confirm that the work has undergone peer review in an academic journal. Listing on CryptoCellar and thanks to Weierud are factors in assessing the result, but they must be distinguished from a peer-review process. Uncertainties also remain over the reading of the signature, the identification of the place name, and the relationship to resends of other messages.
With the cryptographic computation now reproducible, there is a foundation for investigating the remaining questions from the side of the sources and history. If other researchers can follow the published inputs and verification methods and examine even the ambiguities in the readings, it will be possible to judge further how far a string restored by AI can be trusted as a historical document.
