On September 25, 2026, Priyan R published a write-up of a long personal experiment: give a coding agent the 1989 Prince of Persia source, do not edit the code himself, and only play the build and say what is wrong. By the time Claude Opus 5.5 took a turn, the first screen of level 1 had gone from 8,429 pixels that differed from his DOS original down to 2. The Hacker News thread (55 points, submitted by msephton) read that number as a model ranking. explainx.ai's read is narrower. The number is a measurement of an oracle. It is a weak ranking of models.
The post is Analyzing Frontier Model Progress with My Favourite Game: Prince of Persia (Apple II, 1989). The code and the session history are in his repository, PrinceOfPersia_C-Sharp_Port_By_AI. This is a non-commercial fan project. It is not affiliated with Ubisoft or Jordan Mechner. Prince of Persia is a Ubisoft trademark. No original game data is in the repo.
TL;DR
| Question | Answer |
|---|---|
| Which model first played correctly? | Opus 5 (September 2026) was the first playable prince, with real animation. Opus 5.5 then made the rooms pixel-close. |
| Did Opus 5.5 reverse the drawing code? | No. It ported the room-drawing routine already documented by SDLPoP. It said a one-prompt reverse of the DOS binary was probably out of reach. |
| Is this a fair benchmark? | No. Each model inherited the last model's code. A fair comparison would restart from the 6502 source. His goal was a playable game, not an ablation. |
| Can I run it without owning the DOS game? | No. The program reads frames, art, and levels from his local DOS install at runtime. The repo ships code and session history, not the game. |
| What should I copy for my own agent evals? | An oracle (run the reference), a property you can diff, and a rule against painting over a wrong architecture. Skip "port every classic." |
What he actually did
Priyan first played Prince of Persia in 1995 on an IBM PC XT. That setup is ordinary. The DOS release targeted modest PCs, including CGA, and an XT still in a house in the mid-1990s is a hand-me-down, not a fabricated biography. Jordan Mechner published the Apple II 6502 source in 2012 at github.com/jmechner/Prince-of-Persia-Apple-II.
The rule of the experiment stayed fixed across models. Hand the source over. Do not read the generated code. Do not edit it. Play, and describe what is wrong, in ordinary language. Same project the whole way. Prompts only. The repository keeps the session history, which is why a reader can see what each model was given.
That last detail is the one the headline buries. A coding agent that starts on a finished, wrong engine is being asked a different question from a coding agent that starts on the 6502 listing. Priyan optimized for a game he could play. Anyone using the thread as a bake-off is answering a question he did not ask.
Four rounds, one codebase
| Round | When | Model | Starting point | Outcome |
|---|---|---|---|---|
| 1 | February 2026 | Claude Opus 4.6 | Apple II source and level files | C# console game, then a level editor, then Raylib graphics. Rooms, tiles, gates, and guard positions parsed. Looked vaguely like the Apple II game. Played wrong. |
| 2 | March 2026 | OpenAI Codex, model unspecified | The Round 1 codebase | Sharper pixels, transparent sprites, gates you could walk through, pillar tops that were no longer solid. Prince still stuck in the wrong places. |
| 3 | September 2026 | Claude Opus 5 | The same project, plus tools that could launch his DOS copy | Rebuilt the engine around animation frames. Morning build could drop, land, turn, run, jump, fall, and crouch. Room art still slightly off. |
| 4 | September 2026 | Claude Opus 5.5 | The Round 3 codebase, one prompt | Ported a known room-drawing routine. Level-1 first screen: 8,429 differing pixels down to 2. |
Round 1: a picture of the game, and the wrong machine underneath
The first prompt, in Priyan's account, asked Opus 4.6 to use the level files and make a C# console game from the 6502. Over two days the model produced a console game, a level editor, and a Raylib version. It parsed rooms, tiles, gates, and guard positions. The window looked vaguely like the Apple II game.
His notes from those sessions are rough, non-native English. One of them is the phrase "prince is not in floor." That is underspecified feedback. It is also an accurate bug report if you are the only tester and you are not going to open the source. The model still moved the project forward. The root cause, which he learned only later, was architectural. The prince moved one tile at a time. The real game plays hand-drawn animation frames, with a small movement on each frame. A hopper on a grid can be made to resemble a screenshot. It cannot be made to feel like Prince of Persia by adjusting pixels.
Round 2: polish on a wrong engine
In March 2026 he gave the existing codebase to OpenAI Codex. He does not remember which model. The instruction was that the graphics were not good. Codex made sensible, local fixes: pixel filtering so the art was not blurred, transparent sprite backgrounds, open gates you could walk through, and pillar tops that had been wrongly solid.
The prince was still stuck in the wrong positions. This is the failure mode to remember when you grade agents by eye. The surface got closer to a screenshot. The engine was the same wrong engine. A reviewer who only looks at stills will score Round 2 as progress. A reviewer who plays for ten seconds will not.
The jump was an oracle, starting with Opus 5
For Round 3, Priyan changed the harness, not just the model. He wrote two Claude Code skills: one that controls DOSBox (launch, send keys, take screenshots) and one that controls Windows programs. He pointed Opus 5 at his DOS copy and told it to compare against the real game and keep going until success or 6am.
That is a loop with a clock as the stop rule, wrapped in a harness that can see a second program. The skills matter as much as the model name. Before this round, Priyan was the only tester. After it, the agent could launch the original, look, and try again.
Opus 5 diagnosed the tile-grid versus frame-sequence split and rebuilt the engine. It parsed the DOS game files for animation frames, dungeon art, and all 15 levels. It found animation tables inside PRINCE.EXE by searching for byte patterns it recognized from the Apple II source, and it cross-checked each table before trusting it. None of that copyrighted data went into the repository. The C# program reads it from his local DOS install at runtime.
The morning result was playable. The prince could drop, land, turn, run, jump, fall, and crouch, with the real animation. In a second session Priyan played the first screens in DOSBox while the model watched screenshots. Ledge detection was looking in the wrong cell. The model fixed that from the original assembly routine.
The levels still looked slightly off. Opus 5 had placed pieces by matching screenshots and guessed the rest. Playable movement and wrong masonry can coexist. A screenshot match of "about right" does not tell you which of those you have.
What Opus 5.5 changed, and what it did not
Claude Opus 5.5 got one prompt. The character moves, but gates, bricks, and positions differ; see what you can improve. The prompting playbook's useful line here is the finish condition: name what done looks like. "See what you can improve" is vague. The model supplied its own finish line by going after the original drawing routine and then counting pixels.
It did not tune the bricks by eye. It found the room-drawing routine as documented in SDLPoP, the open-source DOS reconstruction by Dávid Nagy and the princed.org community, and ported it. The file in his tree, DosRoomDrawer.cs, is a C# port of SDLPoP's room-drawing code (src/seg008.c). Because of that port, the project is GPL-3.0-or-later. He credits Mechner's published source, Nagy and SDLPoP, and Fabien Sanglard's code review of the original.
Two facts from that reconstruction explain bugs that screenshot-matching had papered over. The brick pattern is seeded from the room, the row, and the column, using the Microsoft C runtime's random-number constants, so a wall looks the same on every visit. The level-1 start gate is open in the data. A hidden button slams it as you drop in, which is why you hear it shut. Guessing brick positions from a crop of a screenshot cannot discover either fact. The generator and the button are in the program, not in the pixels.
Opus 5.5 also found that PRINCE.EXE is EXEPACK-compressed. Earlier tables had worked only because they sat in a stretch the packer left unpacked. It wrote an unpacker and read the tile-drawing tables from his copy. An earlier screenshot crop was off by one pixel row. A theory about "stencil colours" was wrong: those colours are the teal of the exit door.
Then it diffed pixels against the DOS game in DOSBox. On the first screen of level 1, differing pixels went from 8,429 to 2. The two that remained were a torch flame at a different moment. The second room matched. The level-3 exit door matched.
It also moved the prince sprite from the original code, and it only checked one facing. The other facing sank into walls. Priyan restored a backup. The model proved that movement matched across its test ticks, found the sprite bug, and reverted only that part. That sequence is the part worth stealing. A check that can fail, on a case the edit did not stare at, is what separated "the diff went down" from "the sprite is correct."
Could it have done this from the DOS binary alone?
Priyan asked the model that, in one prompt, with no SDLPoP. The answer he reports is: probably not. The model did not reverse-engineer the drawing code. It read SDLPoP's reconstruction and ported it. It found the tables in PRINCE.EXE because SDLPoP said what those tables contained. The previous model had searched the executable without that key and found nothing. Disassembling 16-bit x86 drawing code would have been a longer, multi-session job.
What Opus 5.5 brought, on his telling and on the model's own account, was knowing the reconstruction existed, reading it, spotting EXEPACK, and proving the result with a pixel count. That is a real capability. It is retrieval-plus-port-plus-check. It is not "the model derived the renderer from a compressed DOS binary in one shot."
Priyan's own summary is the right scope for the result: the biggest jump was not only smarter models. From Opus 5 on, the model could see the original and test itself, and he stopped being the only tester. Still unfinished, by his list: guards, sword fighting, palace levels, and exact landing positions.
This is not a clean model bake-off
Say it plainly, because the thread keeps almost saying it and then ranking models anyway. This is not a clean model bake-off. Each model inherited the previous codebase. Opus 4.6 built the tile engine. Codex polished it. Opus 5 replaced the movement model after it could see DOSBox. Opus 5.5 inherited a playable prince and fixed the rooms. A fair comparison would restart from the 6502 listing every time, with the same tools, and would log a fresh tree per model.
Priyan's goal was a playable game, not an ablation. That is a legitimate goal. It does not become a leaderboard because four famous model names appear in order. If you want the leaderboard, you have to pay for the restarts. How to read an AI benchmark applies here in miniature: a score is the output of a model, a prompt, a scaffold, a judge, and a dataset. Change the scaffold (add DOSBox) or the dataset (hand over last month's wrong engine) and you are no longer measuring the same thing.
Commenters on the thread noticed the inheritance and asked why he did not start from the original each time. That question is correct as methodology. It is the wrong complaint about his diary. He was trying to finish a port. The session history is evidence of a process. It is not a controlled eval, and treating it as one will flatter whichever model happened to receive the oracle.
Rough prompts, and what they do and do not prove
Several readers called the prompts gibberish and said the English got in the way. Do not take that as the lesson. The prompts are rough. They are non-native English. The models still made progress. "prince is not in floor" is not a specification. It is a symptom, reported by someone who had agreed not to become the programmer.
The interesting point is when that kind of note became enough. While Priyan was the only pair of eyes, underspecified English had to carry the entire bug. Once the agent could launch the original and take a screenshot, the same kind of note was a pointer, and the oracle did the specifying. The model tolerated vague feedback because it could go look. That is a property of the setup. It is not proof that sloppy tickets are a good interface, and it is not a reason to sneer at the tickets that got a fan game this far.
English as a spec, versus copying a finished game
Another comment on the thread put the discomfort cleanly. If English is supposed to be the new high-level language, this experiment skipped the hard part. The agent was handed a finished game and asked to copy it. A classic, plus an emulator you can run, sidesteps the vagueness of a prose spec. You already own the thing you claim to be specifying.
That objection is right about what this project is. It is a reconstruction with a very patient reference implementation, which is his licensed DOS copy plus a public reconstruction of how that copy draws a room. It is a weak argument for "natural language replaced programming" in general, because most software you want does not already exist as a pixel-stable oracle.
The transferable lesson for people building coding agents is the oracle. Run the original. Diff a property. Keep a second case you did not tune on. The lesson is not "port every commercial game." Porting Prince of Persia was his probe. Your probe should be a reference you are allowed to run: a golden service, a previous binary, a fixture suite, a simulator with a fixed seed. The friendslop games built with Opus 5.5 show the adjacent habit of iterating on screenshots until a move "looks right." That works for a party game you are designing. It stalls when the target is an existing program with a correct answer, because "looks close" can sit on top of a tile engine forever.
A parallel claim that used different rules
On the same thread, a reader described pointing GPT-6 Astra, Sol, and Luna at a 2008 c't magazine Asteroids bot contest, hill-climbing a client that watches the emulator and sends keys. Astra reportedly scored about 1.7 million before the game crashed, and it disassembled parts of the ROM along the way. The same reader then checked the contest rules. The contest used a five-minute limit. The winning human-era bots synced to the random-number generator. Astra played reactively, off the screen, and the big number was measured on games that were allowed to run without that time limit.
Treat that as a caution, parallel to this port. An impressive agent-versus-game claim has to use the same rules as the baseline. 1.7 million does not beat the contest. It is a different task that happens to share a ROM. The Prince of Persia pixel count is on firmer ground inside its own rules, level-1 first screen against his DOS copy, and it still does not license a cross-model ranking.
In the same thread, Kuyawa described a Swift macOS port and put the spend at about $2 of tokens while naming DeepSeek. That is an unverified comment, not a result.
A short eval recipe
If you are evaluating coding agents this month, steal the shape of Round 3 and Round 4, not the game.
- Give the agent a reference runtime it can start. A binary, a service, a simulator, a test harness. If a human has to narrate the difference, you are back in Round 1.
- Pick one property that is allowed to fail in a number. Pixel count, byte diff, trace distance, assertion failures. "Looks better" is how Codex got credit for a wrong engine.
- Do not let the next run only patch that number on a wrong architecture. If the data model disagrees with the reference (a cell step where the reference uses a frame sequence), stop and replace the model of the program. Handing the next agent last week's tree and saying "make the picture closer" reproduces Round 2.
- Hold out a case. Opus 5.5 checked one facing of the sprite. The other facing was the bug. Your diff should include a condition you did not watch while editing.
- Restart when you want a model comparison. Inherited trees measure the stack. Fresh trees measure the model. Publish which one you ran.
Two prompts you can paste onto a task that already has a reference program. They are deliberately not about Prince of Persia, DOS unpackers, or any particular game.
You have a reference program and your implementation of the same behavior.
Launch the reference on a fixed input and record its output.
Launch yours on that same input.
Diff one property that can fail: pixels, bytes, a trace, or exit status.
Before you edit, name which subsystem produced the difference.
If your structure differs from the reference, replace the structure.
Do not tune outputs until a screenshot looks close while the structure stays wrong.
Stop when the diff is explained, or when you need a fact you cannot observe.
Report the measurement before and after, and one input you did not watch while editing.
The reference runtime is the spec. Prose from me is a pointer, not the spec.
Write down the invariant you will check and the command that measures it.
Run both programs. Record the number.
If the number improves and the architecture still disagrees with the reference, revert the cosmetic change and fix the architecture.
When you claim a match, include the before and after measurement and a case you did not use to guide the edit.
If you used an existing write-up or a third-party reconstruction, say so. Do not describe a port of someone else's routine as a reverse-engineering of the binary.
The second prompt is there because of the honest answer Opus 5.5 gave. Knowing a reconstruction exists and porting it is a legitimate way to finish. Calling that a from-scratch reverse of the binary is how eval write-ups go bad.
Honest limits
- Not a controlled benchmark. One person, one inherited tree, prompts that changed with the tools he had that month. Model order is chronology.
- Depends on SDLPoP. The pixel-close rooms came from porting a reconstruction the community had already written, then checking it. Without that document, the previous model searched
PRINCE.EXEand found nothing. - The game stays on his machine. You cannot clone the repo and play the DOS version of the port unless you supply your own copy. That is the correct copyright outcome. It also means a third party cannot re-run the pixel diff from the repo alone.
- Combat and several levels are still open. Guards, sword fighting, palace levels, and exact landing positions are on his list, not in the build.
- GPL comes with the port.
DosRoomDrawer.csis a C# port of SDLPoP's room-drawing code, and the project is GPL-3.0-or-later because of that. If you were hoping for a permissively licensed demo to drop into a product, this tree is the wrong artifact. - One screen is not the whole game. 8,429 to 2 is the first screen of level 1, with a torch flame left over. The second room and the level-3 exit door matched. That is strong evidence about drawing. It is not a claim that every room, every frame of the prince, and every guard already matches.
What to copy, and what to leave
Copy the moment the agent stopped needing Priyan as the only tester. Copy the pixel count, because a number can fail. Copy the backup-and-revert on the sprite, because a held-out facing caught a bug the diff on one direction missed. Copy the model's admission that it used SDLPoP.
Leave the leaderboard. Leave the idea that the next model should inherit a wrong engine and sand the pixels. Leave "port a classic" as a default eval. The oracle is the method. The game was his.
Related on explainx.ai
- Five friendslop browser games built with Opus 5.5 — screenshot iteration on games you are designing, rather than diffing against a finished original
- Claude Opus 5.5 prompting guide — name the finish line, and hand over the whole task
- Claude Opus 5.5 launch: benchmarks and pricing
- What are agent skills? — the DOSBox and Windows controls in this experiment are skills
- Loop engineering for coding agents — a clock stop ("until 6am") versus a diff that can fail
- How to read an AI benchmark — why an inherited codebase is not a model score
- What is harness engineering?
- Opus 5.5 showcase: what builders shipped in the first day
Primary sources: Priyan R's write-up (September 25, 2026), the session-history repository, Mechner's Apple II source release, and the Hacker News thread.
Model names, pixel counts, and unfinished work (guards, sword fighting, palace levels, landing positions) are as Priyan R described them on September 25, 2026, and as discussed on Hacker News. This page is accurate as of September 27, 2026. Prince of Persia is a trademark of Ubisoft. The fan port is not affiliated with Ubisoft or Jordan Mechner.
