GPT-6 Astra took 18 hours 12 minutes to become champion in Pokémon, according to Clad3815, who runs the game-streaming project. Meanwhile, the AI evaluation company Vals AI reports that in Minecraft, Astra was gathering supplies for its run when a creeper explosion destroyed its stored items, after which it spent several hours doing little but farming potatoes. The model has clearly improved at chaining long sequences of actions toward a goal, but its ability to return to the original plan after a setback remains unstable. Reading the two experiments alongside ARC Prize's official evaluation suggests that, beyond completion time, the conditions of operation and the way memory is carried over are key to interpreting the results.
Pokémon in 18 hours 12 minutes: the benchmark is reaching champion
In a report from Clad3815, who runs "GPT Plays Pokémon," Astra reached champion in far less time than GPT-5.6 Sol. Lining up the reported figures shows both the size of the gap and the conditions that complicate the comparison.
| Model | Reasoning setting | Reported time | Result at that point |
|---|---|---|---|
| GPT-6 Astra | high | 18h 12m | Reached champion |
| GPT-5.6 Sol | max | 96h 35m | Reached champion |
| GPT-5.5 | Not confirmed in the report referenced here | 218 hours | Not completed |
Converted to minutes, Astra's and Sol's times are 1,092 and 5,795, respectively. Using Sol as the baseline, the reduction is (5,795 − 1,092) ÷ 5,795, or about 81%. However, this is a difference in reported times to reach champion; it does not mean Astra finishes every kind of task about 81% faster. Nor should GPT-5.5's 218 hours be compared as a completion time.
Clad3815 explains that the model's only input was screenshots, with no access to the game's internal memory, no hints, and no walkthrough provided, and that it progressed autonomously. In other words, the report describes a model reading the situation from the screen, acting, and checking the result on screen again over a long stretch. That no walkthrough was supplied on the spot does not prove the model never saw walkthrough information during training.
The reasoning settings also differ: high for Astra, max for Sol. Because this is not a controlled experiment, with unknowns such as the number of attempts, variance in results, and what the timing includes, additional conditions would be needed to measure a general performance gap between models. Even so, as an operator's record against the same goal, it shows progress in the ability to carry a long procedure through to the end.
Note that the explanations from Clad3815 and Vals AI on X are operator reports cross-checked against public reposts, not results from independent verification of the full play footage.
How to read the potato farming after the explosion
According to a Vals AI post dated September 15, Astra in Minecraft built a semi-automatic blaze-collection setup and gathered six blaze rods. It then defeated six or more Endermen in a warped forest and obtained three ender pearls. It was linking steps together, from gathering to building a facility to seeking needed materials in another location.
Then, after it stored its valuable items in a chest, a nearby creeper exploded, destroying the chest and the bed. What was lost was the stored supplies and similar items; the game world itself did not vanish, and all previous actions were not rolled back. Still, losing the items it was meant to use for the next stage of the run was a heavy blow.
Vals AI reports that Astra spent the following several hours almost exclusively farming potatoes. There was also a note in which it saw something tall and green and confirmed it was sugar cane, not a creeper. A human might describe such behavior as being scared, but this observation does not show that the AI was traumatized. What can be confirmed is the operator's explanation that its behavior became more cautious after the failure and that the way its run progressed changed.
The notes Astra left after the failure also contain something that matters for reading the experimental conditions. Astra had reportedly reminded itself, on the assumption that items are kept on death, to carry important items with it and not leave them in unprotected chests. Vals AI also said in an explanation on September 8 that Astra had enabled keeping inventory on death, which sped up its progress.
With that setting, the risk of losing carried items on death is reduced. But it does not protect the chest in which items were stored. So the reminder to carry important items was itself a reasonable response to the conditions. That the note was left is separate from whether that caution caused the stall into potato farming, and the public descriptions cannot establish the internal cause.
In an early demo on September 8, Astra reportedly built a Nether portal in under three hours. It was also reported dealing with zombies from high ground, trapping a skeleton in a pit, and returning to base by boat. Vals AI explains that this was ordinary computer operation without a game-specific harness. However, this early demo's description alone cannot confirm the full settings of the later long-running experiment or whether there was any intervention.
Even with the same Astra, how memory is handed over changes the score
In the evaluation ARC Prize published on September 3, the same Astra at the high setting scored 54.8% versus 99.9% on ARC-AGI-3 Semi-Private. What differed was the harness, the execution layer around the model that handles observation, operation, and memory management.
| Harness | Reasoning setting | ARC-AGI-3 Semi-Private score | How memory is carried over |
|---|---|---|---|
| Standard | high | 54.8% | The model leaves what to carry over in visible notes |
| Provider Adapter | high | 99.9% | Reasoning state that is invisible from outside is kept between calls, and long conversations are compressed and managed |
With the same high setting, the gap is 45.1 points. The frequently cited Standard score of 62.7% is for the max setting, so comparing it with 99.9% changes both the harness and the reasoning setting. Aligning the reasoning settings, as in the table, makes it clear that the model name alone cannot explain the score.
Standard does have notes. This is not an experiment comparing a model with no memory against one with memory. The conditions for carrying over earlier work differ between a method that leaves it to the model to decide which information to keep as visible notes and one that uses a provider-supplied function for preserving reasoning state. In long tasks, what is tested is not only the ability to stack up judgments but also how past judgments are passed to the next action.
ARC Prize also observed Astra converting the mechanics of an unfamiliar game into concise symbols and recording object positions, rules, and remaining plans. Rather than merely re-viewing the screen each time, it organizes what it has understood into a reusable form. The explanation is that such behavior helps it choose its next action.
The idea of saving and reusing experience has precedents. The 2023 study "Voyager" used GPT-4 to automatically select tasks in Minecraft, write programs for actions, and store successful code in a skill library. It also incorporated a mechanism for fixing code in response to execution errors and environmental feedback. The method differs from the screen-based operation in this case, so the two are not suitable for a head-to-head comparison of completion times, but it shows that designs for handing past successes to the next task have long been studied.
However, ARC scores are not an experiment explaining the cause of the Minecraft stall. Also, the action efficiency measured by ARC-AGI-3 is a measure of how many times the model acted on the environment to solve it, which differs from how many hours it actually took. ARC Prize itself states that the scores come from environments with limited rules and goals where results are fixed under the same conditions, that they do not represent real-world complexity, and that they are not proof of AGI.
Recovery from failure is what long-running AI needs
Reaching champion in Pokémon is a record of stacking up steps and advancing to the finish. Farming after the loss of supplies in Minecraft shows that even while the system keeps running, progress toward the original goal can slow. Collapsing both into a single impression of whether it is good at games erases the differences needed to evaluate long-duration operation.
For developers considering practical applications, it would be useful to record separately the outcome achieved, the time it took, and whether the system recovered after a failure. For example, whether it repeats the same action after a failed operation, re-examines the situation, or picks a different route changes how far it can be trusted to see a job through. This is an evaluation implication drawn from game experiments, and these results alone cannot be used to estimate success rates across business tasks.
What to check in Astra's next run records is whether they can trace the time from losing supplies to regathering the needed items, and even the number of times a human helped. If the operating conditions and memory-management method are made clear, and recovery after failure can be repeatedly tested under the same conditions, it will be possible to judge how much work can be entrusted to a long-running AI more concretely than by how impressive a completed run looks.
