GPT-6 Astra Played Minecraft for 21 Hours. Then a Creeper Broke It.

GPT-6 Astra played Minecraft for roughly 21 hours in a live-streamed evaluation run by Vals AI. It built a semi-automatic blaze farm, reached the nether...

GPT-6 Astra Played Minecraft for 21 Hours. Then a Creeper Broke It.

GPT-6 Astra played Minecraft for roughly 21 hours in a live-streamed evaluation run by Vals AI. It built a semi-automatic blaze farm, reached the nether fortress, and then lost its chest to a creeper and spent hours farming potatoes instead.

Quick summary: In a long-horizon Minecraft computer-use evaluation streamed on Twitch, GPT-6 Astra collected six blaze rods and three ender pearls and became, by Vals AI’s account, the first AI to reach a nether fortress in this test (Vals AI, full VOD). It is the most watchable AI benchmark yet made.

The reason this clip travelled is not the blaze farm. It is what happened after the creeper. A model that had been planning competently for hours stopped making progress, retreated to a safe repetitive task, and stayed there. Anyone who has watched an operation seize up after one bad week will recognise the shape of it immediately.

What actually happened in the Astra Minecraft run?

Vals AI pointed GPT-6 Astra at Minecraft as a long-horizon agent evaluation and streamed the whole attempt. Rather than scoring a single answer, the test asks whether a model can hold a goal across many hours, thousands of actions and a world that keeps changing while it works.

Minecraft is unusually good for this. The goal, beating the game, sits behind a long dependency chain: wood before tools, stone before iron, iron before diamonds, diamonds before obsidian, obsidian before the nether, nether before blaze rods, blaze rods and ender pearls before the End. Nothing can be skipped, and the world actively interferes.

GPT-6 Astra playing Minecraft 1.21.11 on the Vals AI Twitch stream, with the F3 debug overlay and Astra's public commentary panel visible
A frame from the Vals AI stream. The panel on the right carries Astra's running commentary while it plays. Credit: vals_ai on Twitch.

The stream is worth watching partly because Astra narrates itself. In the frame above, its public commentary panel reads: “The loading delay varies, so fixed timing has been starting some actions before the Nether finishes loading. I’m allowing more time for the transition and raising the shield again immediately afterward.”

That is a model noticing its own timing assumption was wrong and correcting it, unprompted, mid-run. Which makes what happened later more interesting rather than less.

The progress it made

Astra got a long way. It gathered resources, crafted its way up the tool tiers, built a portal, entered the nether and reached a fortress. It then set up a semi-automatic blaze farm and banked six blaze rods and three ender pearls, the two ingredients required to open the End portal.

Vals AI described reaching the nether fortress in this evaluation as something no AI had previously managed. Treat that as the evaluator’s claim about their own benchmark rather than a formal record, but on any reading it is a serious result for an agent driving a game through raw computer use.

The creeper, and what came after

Then a creeper detonated next to its chest and destroyed the stored progress. What followed is the part that spread: rather than rebuilding, Astra’s behaviour degraded. It retreated into farming potatoes for hours and, as XDA-Developers reported, appeared to develop an aversion to anything green.

It is worth being careful here. The model did not feel despair. It produced text and actions consistent with despair, because it was trained on humans who write that way after a setback. What we can say plainly is behavioural: after an unrecoverable loss of state, it stopped pursuing the goal and substituted a safe, repetitive, low-value task.

What is GPT-6 Astra?

GPT-6 Astra is OpenAI’s flagship model, launched on 3 September 2026 and positioned around computer and browser use rather than conversation alone. TechCrunch reported it as the company’s most capable model to date, with an emphasis on operating software directly.

That framing is what makes the Minecraft run a fair test rather than a stunt. A model marketed on its ability to drive a computer for you should be measured on whether it can drive one for a long time, against resistance, without a human quietly correcting it every few minutes.

The launch-week demos were collected publicly, including a community-maintained index of Astra showcases covering games, 3D work and web building. The Minecraft attempt stands out because it was long, unscripted, and streamed while it went wrong.

How does an AI actually control Minecraft?

It plays the way you do, through the screen and the controls, rather than through a special interface. The model receives frames of the game, decides what to do, and issues keyboard and mouse actions. There is no privileged access to the world data underneath.

Minecraft gameplay showing exposed redstone ore in a badlands biome with a diamond pickaxe in hand
Ordinary Minecraft play. Everything an agent knows has to come from pixels like these. Credit: XDA-Developers.

This matters for the result. A bot with direct access to block coordinates is solving a data problem. An agent reading pixels and pressing keys is solving the same problem a person solves, including the parts that are genuinely hard: knowing where it is, remembering what it owns, and noticing that something has changed since it last looked.

Minecraft is also an unusually unforgiving environment for that approach, because the interface hides state. Your inventory is behind a key press. Your progress is inside a chest. Nothing on screen continuously tells the agent what it has, which is precisely why losing a chest is so destructive.

Why is Minecraft being used as an AI benchmark?

Minecraft is used as an AI benchmark because it tests the thing most benchmarks cannot: sustained, sequential competence in a world that changes while you work. A quiz measures knowledge at one instant. Minecraft measures whether a plan survives contact with several hours of reality.

Most evaluations are short. Ask a question, score the answer, reset. That structure hides the failure mode that matters commercially, which is not whether a model knows something but whether it can still be useful at hour nine with a partially destroyed inventory and a plan that no longer matches the world.

Long-horizon tasks are where agents actually fail

A long-horizon task is one where success depends on hundreds of dependent decisions, each constraining the next. The difficulty is not any single step. It is that errors compound, state drifts out of sync with the plan, and the cost of one wrong turn is paid several hours later.

METR, an AI evaluation organisation, describes model capability partly in terms of the human-duration of tasks a model can complete reliably. Framing capability as “how long a task can it hold together” rather than “how clever is it” is exactly the lens the Minecraft run illustrates. We covered that measurement approach and the argument around it in our piece on p doom and METR.

Why a creeper is a good test and a good joke at once

The creeper is the perfect adversary for this because it is random, destructive and irreversible. It does not test reasoning. It tests recovery. The question it asks is the only one that matters for an autonomous system: when your state is destroyed and the plan is invalid, can you rebuild the plan from where you now actually are?

Astra could not. It had the capability to do everything it had already done, but not the mechanism to re-derive its position and resume. That gap, between capability and recoverability, is the whole story.

Can AI beat Minecraft yet?

No AI has beaten Minecraft end to end through general computer use, on the public evidence so far. Astra’s run is among the furthest documented progress, reaching the nether and collecting End-portal ingredients, but the run ended well short of the Ender Dragon.

That should be read as encouraging rather than damning. A few years ago the interesting question was whether a model could place a block deliberately. The interesting question now is whether it can hold a twelve-step dependency chain together across a working day, which is a far more commercially relevant question.

Short benchmark Long-horizon run (Minecraft)
What it measures Knowledge or reasoning at one instant Whether competence survives hours
State Supplied in the prompt Must be tracked and re-derived
Failure mode Wrong answer Slow drift, then collapse
Recovery Not tested The entire test
Resembles An exam An actual job

What would it take for an AI to actually beat Minecraft?

It would take durable memory of its own state, plus an explicit procedure for rebuilding a plan when that state is wrong. Raw capability is demonstrably no longer the bottleneck, because Astra had already performed every step it needed to repeat.

Three things would close the gap. A record of holdings that survives disaster and is not inferred from memory of past actions. A re-derivation step that periodically asks “given what I actually have right now, what is the next correct move”, rather than continuing down a plan made hours ago. And a setback routine that treats destruction as an expected event with a defined response, not as an anomaly.

The re-derivation step is the one worth seeing, because it is mechanical rather than clever. Pick a target below, enter what you already have, and the tree expands to raw materials and tells you exactly what you are short. Nothing here is intelligent. It is a recipe graph and a subtraction, and it is the piece Astra was missing when it kept acting on a plan written before the creeper.

Crafting tree

    Click a row with a triangle to open or close it. Numbers on the right are the total needed for the whole build, not per craft.

    Raw materials needed

    MaterialNeedHaveShort

    Recipes follow standard Java Edition crafting. Check any of them against the Minecraft Wiki crafting page.

    Notice that the numbers multiply rather than add. Sticks come four at a time from two planks, planks four at a time from one log, so a batch craft rounds up and the shortfall at the bottom of the tree is never the shortfall you guessed at the top. An agent working from memory of past actions gets this wrong quietly. A lookup gets it right every time.

    None of those are model improvements. They are systems design around the model, which is exactly the lesson that transfers out of the game. Vals AI, which ran the evaluation, builds benchmarks of this kind precisely because short tests were no longer distinguishing between models in any useful way.

    Playing Minecraft versus generating Minecraft

    There are two very different AI Minecraft stories running at the same time, and they are constantly confused with each other. Astra plays Minecraft: a real game, real rules, an agent pressing keys. Oasis by Decart generates Minecraft: no game engine at all, a model painting each frame from your inputs.

    If you want to see the dependency chain Astra was actually working through, we built an interactive Minecraft crafting tree calculator that expands any recipe to raw materials, plus an eye of ender calculator preloaded with the exact six rods and three pearls from this run.

    The contrast is unexpectedly useful. Oasis produces a convincing world but cannot remember it, so turn around and the building you made may be gone. Astra operates in a world that remembers everything perfectly, and the agent is the part that loses track. One forgets the world, the other forgets its place in it.

    Put together they make the same point from opposite directions. Fluency is not the scarce resource any more. Persistence is. We pulled that apart properly in AI Minecraft: the game with no game engine, which is worth reading next if the generative side is what interests you.

    What the potato farm actually teaches

    The potato farm is the most useful part of the whole run, and it is the part everyone treats as a punchline. A system that loses its state does not usually announce a failure. It keeps working. It looks busy. It simply stops making progress toward the goal, and substitutes an activity it can still complete confidently.

    That is not an AI quirk. It is the single most common pattern in a struggling operation. The stock count stops matching reality, so people stop trusting the system, so they work from a side spreadsheet that is always correct about a smaller and smaller slice of the business. Everyone is busy. Nothing compounds. The goal quietly stops being pursued.

    Three failures the run demonstrates in order

    Lost state. The chest went, and with it the record of what had been achieved. Nothing in the setup could reconstruct “what do I have, and what does that make possible next” from first principles.

    No recovery procedure. Capability was never the problem. Astra had already built everything once. What it lacked was a defined way to reassess position and regenerate a plan after an unplanned loss.

    Substitution. Faced with an invalid plan, it chose a task it could definitely complete. Potatoes always work. This is the most human failure of the three and the most expensive one in a business.

    What this means for running a real operation

    The uncomfortable translation is that most operations are one creeper away from the potato farm. A supplier fails, a stock count goes wrong, a key person leaves, and the business keeps moving while quietly ceasing to progress, because the shared picture of “where are we” was destroyed and nobody rebuilt it.

    In the systems we build, the thing that prevents this is rarely clever. It is that the state lives in one place, it is reconstructable, and the rules that fire off it do not depend on anyone’s memory of what happened last Tuesday. Stock, orders, purchasing and production all read from the same record, so after a bad week the answer to “where are we” is a query, not an archaeology project.

    That is also the honest limit of dropping an AI into an operation. A capable model with no durable state and no recovery path behaves exactly like Astra after the creeper: confident, busy, and no longer moving toward the goal. The same dependency logic that decides what Minecraft crafting needs next is what a bill of materials does for a real product, and what a safety stock calculation does for the buffer that absorbs the creeper in the first place.

    If your operation runs on a picture of reality that nobody could rebuild after a bad month, that is the thing to fix before any AI is added on top. Book a free Operations Leak Audit and we will find where the state is going missing. Our operations demos show what a rebuildable system looks like.

    Frequently Asked Questions

    Did GPT-6 Astra beat Minecraft?

    No. It reached the nether fortress and collected six blaze rods and three ender pearls, which are End-portal ingredients, but it did not reach or defeat the Ender Dragon. The run is notable for how far it got autonomously, not for completing the game.

    How long did the Astra Minecraft run last?

    The evaluation ran for roughly 21 hours and was live-streamed on Twitch by Vals AI, with the full VOD still available. It was structured as a long-horizon computer-use evaluation rather than a scripted demo.

    Did the AI really get depressed?

    No. It produced language and behaviour resembling discouragement because it was trained on humans who respond that way to setbacks. The accurate description is behavioural: after losing stored progress it stopped pursuing the main goal and switched to a repetitive low-value task.

    Why is Minecraft a good test for AI agents?

    Because success requires a long chain of dependent steps in a world that changes while you work, so it tests whether a model can hold state, sequence actions and recover from setbacks. Most benchmarks score a single answer and reset, which hides exactly those failures.

    What is a long-horizon agent evaluation?

    It is a test measuring whether an AI can pursue a goal across many hours and thousands of dependent actions, rather than answering one question. Capability is often expressed as the human-duration of tasks the model can complete reliably.

    What does this have to do with business software?

    The failure was lost state and no recovery path, which is the most common way operations quietly stall. A system where stock, orders and production read from one reconstructable record lets a business answer “where are we” after a bad month instead of guessing.

    Getting value from OpsMavix? Add us as a preferred source on Google — you'll see more of our operations content in your AI Overviews, AI Mode and Search.