Behind the scenes
The autopilot
A controller that plays Xenon 2 by itself — written almost entirely by an AI coding assistant, under the direction of a human who kept score of what went wrong.
What it is
Every emulated frame, the emulator hands over a canonical snapshot of the game's memory: every live object with its position, type, sprite and update routine, the tile map, the camera, the ship's exact sub-pixel state. From that the autopilot maintains a world model — a persistent map of walls it has seen, tracked enemies with predicted motion, destructible gates, pickups worth collecting — plans a route through the level, picks a tactic (advance, retreat, intercept, suppress, escape) and presses the joystick.
It started as a Python program talking to the desktop emulator over a socket, one frame at a time. As the observations and tactics settled, the hot paths were moved into dependency-free C that now runs inside the emulator itself, with the Python kept as a reference and an inspector.
Where it stands
It finishes the game. On 5 October 2026 four uninterrupted runs flew from Level 1 to the destroyed Level 5 ship without losing a life, with level-specific route, boss and shop policies (the runs). It still takes damage, which pickups and shop repairs make good, and it still misses some pickups and a boss reward.
Getting there took far longer than it should have. It took countless experiments before each level could be completed, and the AI kept repeating the same mistakes. Most of the trouble was terrain and navigation: for weeks the ship got stuck in front of open passages for no visible reason, or kept bumping into the walls, while the AI patched one waypoint or one wall rule after another. Where the ship stopped, a valid route was often already there; it was lost on the way from the plan to the joystick, and no amount of waypoint patching could fix that (the navigation review).
Why the mistakes are the interesting part
The retrospective is the most honest document on this site. It records, in order, every wrong turn the AI took while building the controller: guessing values that could have been measured, mixing screen and world coordinates, letting a pickup silently take over the joystick, patching symptoms of a broken planner contract, declaring success before replaying the failing frame. Again and again the entries come back to walls, terrain and routes. Each entry names the root cause, and the patterns section distils them into rules that changed how the rest of the work was done.
228 ways an AI got it wrong
The full retrospective, with a filter box. Some highlights:
- #11 — Used screen-space motion in a world-space game model
- #20 — Pickup collection had unconditional priority until danger was already immediate
- #44 — Declared dead-end fixes successful before replaying the reported failure
- #90 — Invented wall occlusion that the game's projectile code does not implement
- #108 — Let the current scene and the persistent map disagree about an open gate
Plausible inference substituted for direct evidence. Coordinate systems were not enforced architecturally. Diagnostics were added after failures instead of before policy. Completion was communicated too early or at the wrong layer.
— from "Patterns behind these"
A much shorter log covers the renderer: 27 mistakes from one shop-screen session.
Replay inspector
Offline route maps
Routes the autopilot plans over the complete level maps before play. Click through for the interactive versions.
Interactive explorers
Autopilot decision cost
Where each frame of autopilot thinking time goes, and how it scales.
Space-time distance field
Two ways to ask "is this move safe": eager flood fill versus lazy search, measured.
Level 2 navigation flow field
The offline flow field the autopilot follows through Level 2, phase by phase.
Level 2 route comparison
Candidate routes through Level 2 side by side.
Level 2 offline route
The planned route over the full Level 2 map.
Level 3 offline route
The planned route over the full Level 3 map, including destructible gates.
Level 4 offline route
The planned route over the full Level 4 map.
Level 5 offline route
The planned route over the full Level 5 map, up to the tank boss.
Start here
These notes are snapshots: each describes the controller as it was when written. The earlier ones cover the Python implementation; the native C port that followed changed a good deal of it, so prefer the more recent documents where they disagree.
The autopilot
The implementation guide: how it reads the game's memory, builds a world model, plans routes, picks tactics and drives the joystick. Start here.
228 ways an AI got it wrong
The full retrospective of every mistake the AI made while building the autopilot, most of them about walls, terrain and navigation, with root causes, the patterns behind them and the rules that came out of it.
Finishing the game
Four uninterrupted campaigns from the first level to the destroyed Level 5 ship with no lives lost (2026-10-05), and the damage, missed pickups and rewards that still went wrong.
Why the ship kept getting stuck
The navigation review behind that result: where the ship stopped, a valid route was often already there, lost between planning and the joystick. Widening waypoints or forcing inputs did not fix it.
Level 2 postmortem
Why many individually plausible fixes did not make Level 2 reliable: the work was done in the wrong order.
The Level 5 tank
The tank boss's model, the many experiments that failed against it, and the C pilot's run that finally beat it without cheats (2026-09-29).
The Level 3 carrier
The articulated carrier boss: its tails, cores and health words, and the gun lane and refuges the autopilot uses against it.
The Level 5 final ship
The final ship's damage sequence read from the original code, and the firing lanes that destroy its 18 mounts and collect all 20 cash drops.
The replay inspector
How recordings are replayed in the browser through the same C autopilot compiled to WebAssembly, and what each overlay shows.
Full-game regression recording
Recording complete playthroughs so that every change can be checked against a real run.
The maneuver planner
The second-generation controller, its architecture, limits and how it compares to the original.
Navigation optimization
The exact grid navigation backend and its live regression evidence.
Architecture review
Reducing decisions and queries per frame before porting the controller to C.
The native C port
Moving the controller from Python into dependency-free C inside the emulator, stage by stage.
Native session inside Hatari
The resident C session that drives the ship straight from the emulator's frame hook.
Running the autopilot in the browser
Options for running the ~50k-line Python autopilot alongside the WASM game, including on phones.
Level 2 strategy
Level 2 compiled into a static mission from the resident tile map.
Level 3 strategy
A deliberately small mission policy for Level 3.
Shop autopilot
Buying the right upgrades from Crispin using the game's own RAM state instead of screen reading.
All autopilot notes (every tweak, kernel and validation run)
Working notes written during development. They are detailed, sometimes superseded by later ones, and kept for the record.
- Architecture experiments (2026-09-06) AUTOPILOT_ARCHITECTURE_EXPERIMENTS.MD
- Batch fixed protocol records before planning AUTOPILOT_BATCH_FRAME_DECODING.MD
- Combined native candidate trials AUTOPILOT_COMBINED_CANDIDATE.MD
- Direct shared script bodies and dispatch bypass AUTOPILOT_DIRECT_SCRIPT_BODIES.MD
- Remove unnecessary velocity fitting AUTOPILOT_EXACT_MOTION_CONSUMERS.MD
- Firing-lane routing: remove repeated terrain work AUTOPILOT_FIRING_LANE_REUSE.MD
- Regenerating embedded autopilot data AUTOPILOT_GENERATED_ASSETS.MD
- Incremental renderer terrain observations AUTOPILOT_INCREMENTAL_TERRAIN.MD
- Mission-controller migration AUTOPILOT_MISSION_MIGRATION.MD
- Batched native body queries (2026-09-06) AUTOPILOT_NATIVE_BODY_KERNEL.MD
- Whole-candidate evaluation (2026-09-06) AUTOPILOT_NATIVE_CANDIDATE_PLAN.MD
- Native candidate-scene preparation AUTOPILOT_NATIVE_CANDIDATE_PREPARATION.MD
- Native combat rollout (2026-09-06) AUTOPILOT_NATIVE_COMBAT_KERNEL.MD
- Native combat-bound preparation AUTOPILOT_NATIVE_COMBAT_PREPARATION.MD
- Native driving coordinator (ABI 81) AUTOPILOT_NATIVE_DRIVER.MD
- Native navigation environment (ABI 82) AUTOPILOT_NATIVE_ENVIRONMENT.MD
- Native linear prediction and collision-shape preparation (2026-09-06) AUTOPILOT_NATIVE_LINEAR_KERNEL.MD
- Native/Python live validation — 2026-09-06 AUTOPILOT_NATIVE_LIVE_VALIDATION.MD
- Native maneuver forecast (2026-09-06) AUTOPILOT_NATIVE_MANEUVER.MD
- Native map observation sessions (ABI 83) AUTOPILOT_NATIVE_MAP_SESSION.MD
- Native evaluation of scripted swarm threats AUTOPILOT_NATIVE_POLICY_PATHS.MD
- Autopilot: removing repeated preparation work (2026-09-06) AUTOPILOT_NATIVE_PREPARATION.MD
- Native scene models (2026-09-06) AUTOPILOT_NATIVE_SCENE_MODELS.MD
- Native canonical object tracking AUTOPILOT_NATIVE_WORLD.MD
- Native world driving and optional Python views (ABI 84) AUTOPILOT_NATIVE_WORLD_VIEWS.MD
- Policy preparation optimization AUTOPILOT_POLICY_PREPARATION.MD
- Prediction reuse (2026-09-06) AUTOPILOT_PREDICTION_REUSE.MD
- Shared maneuver setup and deferred Python results AUTOPILOT_PREPARATION_REUSE.MD
- Shared native prediction and scene preparation AUTOPILOT_SHARED_PREDICTION_SCENE.MD
- Shell firing-lane alignment AUTOPILOT_SHELL_FIRING_LANES.MD
- Level-1 shell stations and spans validation AUTOPILOT_SHELL_VALIDATION.MD
- Direct prediction use of C-owned tracks AUTOPILOT_TRACKED_VELOCITY.MD
- Travel pacing and the Level-1 corridor AUTOPILOT_TRAVEL_PACING.MD
- Remove unused render geometry from world-state preparation AUTOPILOT_WORLD_DRAW_PREPARATION.MD
- World preparation: scalar decoding and numeric kinds AUTOPILOT_WORLD_PREPARATION.MD
- Campaign incident fixes, October 3, 2026 CAMPAIGN_FIXES_20261003.MD
- Sequential campaign incident repairs — 2026-10-04 CAMPAIGN_INCIDENT_FIXES_20261004.MD
- Native campaign from the game start (2026-10-01) FULL_NATIVE_CAMPAIGN_20261001.MD
- Three native campaign attempts — 2026-10-04 FULL_NATIVE_CAMPAIGN_THREE_20261004.MD
- level2-map.html level2-map.html
- level2-worm-terrain-mask.html level2-worm-terrain-mask.html
- Level 2 wall-lip repair and full-level validation — 2026-10-02 LEVEL2_CORRIDOR_STALL_20261002.MD
- Level 2 life-loss analysis — October 2, 2026 LEVEL2_LIFE_LOSS_ANALYSIS_20261002.MD
- Level 2 validation after pulling xenon2 — October 2, 2026 LEVEL2_MERGED_VALIDATION_20261002.MD
- Level 2 r150 worm launch and post-shop swarm fixes LEVEL2_R150_FIX_VALIDATION_20261002.MD
- Level 2 connector ownership and physical retreat recovery LEVEL2_ROUTE_ARBITRATION_20261002.MD
- Level 3 gate regression: historical comparison, 3 October 2026 LEVEL3_GATE_REGRESSION_ANALYSIS_20261003.MD
- Level 4 full validation, September 27, 2026 LEVEL4_FULL_VALIDATION_20260927.MD
- Level 4 full native validation, 28 September 2026 LEVEL4_FULL_VALIDATION_20260928.MD
- Level 4 native pilot validation, September 26–27, 2026 LEVEL4_NATIVE_VALIDATION.MD
- Level 4 temporary shield and first-boss handover LEVEL4_SHIELD_AND_HANDOVER_FIX.MD
- Level 4 resident-C validation, 3 October 2026 LEVEL4_VALIDATION_20261003.MD
- Level 5 diagonal chamber: early stall and local correction LEVEL5_DIAGONAL_GATE_20261004.MD
- Level 5 full-entry C validation, September 30 LEVEL5_FULL_VALIDATION_20260930.md
- Level 5 opening pickups and full-level validation — 2026-10-04 LEVEL5_PICKUPS_AND_COMPLETION_20261004.MD
- Level 5 resident-C validation — 2026-10-04 LEVEL5_VALIDATION_20261004.MD
- Native fast-forward investigation (2026-10-01) NATIVE_FAST_FORWARD_PERFORMANCE.MD
- Fixed wall model and route execution: validation NAVIGATION_FIX_VALIDATION_20261005.MD
- Three native fast-forward campaigns, October 2–3, 2026 THREE_NATIVE_CAMPAIGNS_20261003.MD
- THREE_NATIVE_CAMPAIGNS_20261003_INCIDENTS.csv THREE_NATIVE_CAMPAIGNS_20261003_INCIDENTS.csv


