Xenon 2

Autopilot · write-up

Level 2 autopilot postmortem

xenondoc/LEVEL2_AUTOPILOT_POSTMORTEM.MD · 20 KB · updated 2026-09-17

Date: 2026-08-24

This is a deliberate pause after the long sequence of Level 2 autoplay experiments. The goal is not to defend the latest implementation or propose another threshold adjustment. It is to identify why many individually plausible fixes did not produce the required outcome: reliably completing Level 2 without repeatedly losing lives.

The short conclusion is that the work was performed in the wrong order. I kept improving tactics and search while the complete transition from an issued command to the game's consumed input, next ship state, enemy state, and damage result was not yet authoritative. A sophisticated planner over an unverified transition model produces convincing diagnostics and still chooses unsafe moves.

Sections before Resolution implemented on 2026-08-24 describe the state at the time of the postmortem. They are retained as the failure record; the resolution section records the replacement model and current validation results.

Scope and evidence

xenon_tools/run_logs currently contains 164 level2-*.x2events recordings, totaling about 2.45 GB. Some are renderer or protocol probes rather than independent policy trials, but the volume itself is evidence of excessive live iteration. The attempts fall into these broad phases:

Phase Representative recordings What was learned Why it did not finish Level 2
Initial Level 2 and rendering level2-autoplay-0821-*, level2-artifacts-* Level-specific assets, palettes, draw hooks, backgrounds, flat-color effects and AVI comparison were made observable. Rendering correctness was necessary, but it did not validate control or collision prediction.
First-half navigation and pressure level2-progress-*, level2-pressure-relief-*, level2-completion-*, level2-exploration-* The controller reached new areas, learned to create reaction space, and exposed several wall, projectile and route failures. Each failure prompted a local tactic or priority change; ownership between route, combat, pickups and survival remained unstable.
Three-eye boss level2-boss-controller-*, level2-boss-fix-* The actual objective was encoded: hit the two corridor eyes and then the central eye. The boss and shop became reachable. This solved one scripted encounter, not the general movement/safety problem.
Post-shop exploration level2-postshop-*, level2-continue-*, level2-pre-homing-* The post-shop area introduced emitters, homing enemies, overlapping swarms and severe spawn pressure. Generic swarm and projectile heuristics were applied before the Level 2 overlay routines and damage geometry were fully recovered.
Route and scroll-control tuning level2-prewave-* through level2-prevent-branch-0824-55 Backward scroll can pace spawns; routes can retain a strategic endpoint while avoidance temporarily deviates; known dead branches can be rejected. Many route-owner, corner, re-entry, spawn-gate and pressure-relief patches changed which subsystem won the joystick without proving that the resulting command sequence was executed as modeled.
Homing search level2-homing-static-aware-0824-56 through level2-player-hull-model-0824-62 The live Level 2 homing updater at $050798, its eight-direction motion, retarget period, lifetime, and two collision paths were recovered. Exact enemy displacement alone was insufficient. Ship hull phase, collision call order, walls, scrolling and command timing still disagreed with planner assumptions.
Input-boundary investigation level2-input-entry-0824-63/64, level2-input-fifo-0824-65 $0A00 sampled at the canonical boundary was not necessarily the value later consumed at ship movement entry $00661C. Exact entry telemetry and a provisional FIFO command handoff were added. The provisional FIFO made the race visible but introduced delayed stale actions. It was replaced by the validated level-triggered handoff described in the resolution. One hit was attributed to the homing updater and one remained unresolved at this stage.

The latest stopped run is level2-input-fifo-0824-65.x2events. It began with shield 31 and was stopped at shield 19 with no life lost yet. It recorded two damage events: one exact Level 2 homing association and one unresolved association. This is not a successful validation. The recorder closed the event log and released input, but its wrapper again stalled while requesting the final snapshot, so no final .sav should be assumed to exist.

What repeatedly went wrong in gameplay

The ship entered or remained in tactically bad space

The controller frequently advanced too high on the screen, entered a branch with little reaction space, or remained close to walls while worms or homing waves approached. Later policies added center-band and backward-scroll preferences, but these were still preferences competing with route, target and spawn objectives. They were not safety invariants.

Retreat and pacing became objectives instead of constraints

Backward scrolling can delay a later spawn and create time to clear the current wave. It is bounded, however, and pressing DOWN is useful only while the game's scroll gate permits it. The controller sometimes made spawn pacing the joystick owner even when a current hazard required another action. In other runs, hazard avoidance abandoned the strategic escape suffix and selected the bad branch again. Both failures came from the same design error: mission intent and immediate control were allowed to replace one another.

A visually correct prediction did not imply a safe ship trajectory

The UI could show the homing path or swarm lane correctly while the ship still collided. That is because the overlay mainly demonstrated enemy prediction. Safety requires the joint evolution of:

  • the precise input the game consumes;
  • steering accumulator and ship position;
  • wall collision/replay and camera scrolling;
  • every enemy update in game order;
  • the collision rectangle or point used at each call site; and
  • damage/invulnerability state.

Until the latest protocol work, the first item was not even captured at the authoritative routine entry. Some damage paths are still not attributed. The joint model therefore has not passed an end-to-end proof.

Strategy arbitration kept invalidating locally good plans

The controller has accumulated modes for collection, combat, formations, swarms, retreat, pressure relief, spawn pacing, route recovery, homing preparation and homing hold. Each is reasonable in isolation. Their transitions and overrides are not one coherent state machine. A target, route, safe screen band, spawn barrier or hazard plan can each become the effective movement owner. Adding one more priority rule has repeatedly moved the failure rather than removing it.

The search optimized an approximate simulator

The long-horizon beam search increased confidence without increasing truth. It evaluated many action sequences with approximate ship and collision transitions, then committed the best-looking one. A one-control-step phase error is enough to invalidate the whole trajectory near a homing ring. A large beam cannot compensate for a wrong state transition; it searches the wrong state space more thoroughly.

Recorded regressions were not used as a hard acceptance corpus

Many unit tests assert a helper's intended output or reconstruct a selected historical frame. They are useful, but they often encode the current model rather than compare it with what the Atari did one frame later. Passing 274 tests did not prevent the latest run from taking damage. The tests proved internal consistency, not external fidelity.

Root causes in the engineering approach

1. Tactical patching preceded model validation

I repeatedly reacted to the visible symptom: move lower, hold a lane, prefer a route, stage before a spawn, widen a hull, add a collision path, or increase the horizon. Several changes were individually valid, but there was no gate requiring the complete one-frame emulator transition to match the model before another gameplay strategy was attempted.

2. Observability was added only after a theory failed

Authoritative ship-movement input, the Level 2 overlay bytes, direction-specific homing geometry and exact damage associations should have existed before tuning avoidance. Instead they were added one at a time after failed runs. The field formerly called consumed_input_mask is the clearest example: its name suggested stronger semantics than its capture point guaranteed.

3. Full runs substituted for minimal experiments

Long runs contain scrolling, walls, spawns, pickups, projectiles and multiple tactic transitions. They are poor experiments for determining whether one input is consumed on frame N or N+1. A short branched checkpoint test would have exposed the input race much earlier and much more cheaply.

4. The architecture is monolithic at the decision boundary

World modeling, mission selection, target selection, route following, spawn pacing and survival all feed a final per-frame arbitration path. The system lacks a strict contract saying that the mission layer may propose goals but cannot issue a command that the safety layer has rejected.

5. Protocol and experiment artifacts were not kept stable enough

Recent protocol layouts changed during the investigation. level2-input-entry-0824-64.x2events currently produces a trailing-byte decoder error in the present analyzer, so an important experiment is not a durable regression artifact. The latest recorder wrapper also hung after stopping the log while saving a snapshot. Tooling instability reduced the reliability and comparability of trials.

6. A stale ReVa snapshot was nearly treated as Level 2 truth

Level loading replaces the overlay arena around $04F000 and above. The original ReVa snapshot contains Level 1 bytes there; it cannot be used to infer Level 2 routines such as $050798 without first importing live Level 2 RAM. Resident routines remain useful, but overlay analysis needs a level-specific memory image and explicit database annotations.

7. Success criteria drifted with the latest obstacle

Reaching the boss, leaving a dead end, entering the shop, surviving the first homing ring, and fixing one input transition were each treated as the current milestone. None is the requested acceptance condition. The condition is completing Level 2 reproducibly, with full recording and useful damage metrics.

What is trustworthy now

  • Canonical analysis runs on original Atari game frames, not extrapolated rendering frames.
  • Stable object identities and world-space coordinates are available.
  • The persistent tile map and collision-expanded navigation representation exist.
  • Level-specific rendering assets, the three-eye boss objective, shop detection and shop control are substantially understood.
  • The Level 2 homing update at $050798 has an exact observed eight-direction movement model and retarget schedule for the tested state.
  • Protocol version 11 captured the input and steering accumulator at ship movement entry $00661C.
  • The provisional autopilot step FIFO exposed command-boundary timing, but was not the final input model and was subsequently removed.

These are foundations. They do not establish that collision prediction or the current safety planner is correct.

What is not yet trustworthy

  • Complete attribution of every shield-loss path. The latest run still has an unresolved hit.
  • Exact direction- and animation-dependent player collision geometry at every relevant call site.
  • Joint ship, wall, scroll and hazard rollout over several future frames.
  • Planner revalidation after a scroll change, object spawn, tactic transition or queued input.
  • Hazard spawn forecasting beyond currently active objects.
  • The current arbitration rule between immediate survival, route retention and spawn pacing.
  • Compatibility of all recent recordings with the current decoder.
  • The recorder's stop/save transaction.

Required reset before another full gameplay run

No new Level 2 tactic should be added until the following model-validation gates pass from a small, reproducible checkpoint around the post-shop homing encounter.

  1. Input gate: for a scripted sequence containing all eight directions, reversals and neutral frames, prove which mask is consumed at every $00661C entry. Assert the requested level-triggered mask, captured input and observed steering change.
  2. Ship-transition gate: predict the next ship world position, screen position, steering state, wall-block flags, and scroll delta from the captured pre-state and consumed input. Compare against the next canonical frame. Any unexplained mismatch fails the gate.
  3. Hazard-transition gate: predict every active homing object's next age, direction, world position, lifecycle and collision geometry. Compare against the next frame. Do the same later for each new Level 2 hazard family.
  4. Collision gate: instrument every call path that can reduce shield. For every frame and every ship/hazard pair, assert predicted hit/miss against the actual call-site result. An unresolved damage event fails the gate.
  5. Rollout gate: branch from one saved state, execute several short candidate sequences in Hatari, and compare each actual trajectory with the planner's predicted trajectory. The planner may be enabled only after the sequences match, not merely after the first action matches.
  6. Policy gate: express Level 2 mission intent as constraints and goals, then verify that it cannot override a rejected safety action. Test route retention, bounded backscroll and spawn pacing as secondary preferences.

The intended control separation should be:

Level objective / route / target
             |
             v
      proposed goal region
             |
             v
exact emulator transition model ---> reachable safe states over time
             ^                                  |
             |                                  v
authoritative input, ship,              safe action / committed tube
wall, scroll, hazard and                         |
collision telemetry                              v
                                      level-triggered input handoff

The safety planner should compute a set of reachable safe states (a control tube), not merely attach penalties to unsafe states. Collision is a hard constraint whenever any collision-free sequence exists. Route progress, staying near corridor center, backward-scroll pacing, firing position and pickup value are ordered secondary objectives. If several evasions are safe, choose the one that retains the strategic route; if none is safe, maximize time to impact and minimize expected damage.

Backward scrolling to delay a second swarm is worth a controlled experiment after these gates pass. It can serialize pressure and buy reaction time, but it is not a substitute for a correct avoidance model because the game limits how far a region can scroll backward.

Process changes

  • Freeze strategy tuning while any transition or collision mismatch remains unexplained.
  • Create one minimal regression per newly discovered game mechanism before launching a full run.
  • Keep a level-specific live RAM snapshot for ReVa overlay analysis.
  • Version the frame schema explicitly and keep recorded fixtures decodable by the analyzer used for their acceptance tests.
  • Record a compact machine-readable validation report beside every run: starting checkpoint, code revision, protocol version, first/last frame, shield damage, lives lost, unresolved hits and terminal condition.
  • Do not call a fix successful until the same binary and policy pass both the minimal gate and a fresh normal-speed end-to-end recording.

The appropriate next action is therefore not another autoplay run. It is to build the deterministic model-validation harness and close the unresolved collision path. Only after those are green should we decide whether the existing beam planner can be simplified into a safe-state search or should be replaced.

Resolution implemented on 2026-08-24

The reset above was completed as a model replacement, not as another tactic adjustment.

  • Protocol 12 captures $CCC, $CE0, exact $661C and $7B5A entry/exit state, and the current player's sprite-header collision hull.
  • Step input is now a level-triggered state. Frames without a ship update replace the pending mask instead of accumulating a FIFO of movements that execute later.
  • game_model.py is the single ship/camera/homing transition implementation used by autoplay, unit tests, recording validation, and snapshot branch validation.
  • Level-2 homing world motion no longer contains a camera-reversal impulse. $CD8 screen compensation cancels the camera-origin change exactly.
  • The search carries exact homing animation state and collision-header hulls. It models the generic pre-movement AABB collision and post-movement homing-anchor point test as separate phases.
  • Common damage routine $65A4 is correlated with the authoritative $CC6 write. This captures the formerly unresolved direct 4-point homing path and correctly reports lethal underflow as the remaining shield rather than an unsigned 65531-point hit.
  • Committed avoidance plans are revalidated by replaying their inputs through the same exact model.

Acceptance evidence:

  • 323 Python tests pass.
  • level2-model-validation-0824-04.x2events: 119/119 ship, 120/120 camera, and 311/311 homing transitions match; the formerly damaging 120-frame encounter keeps shield 35→35.
  • level2-damage-hook-validation-0824-08.x2events: all seven shield decreases are attributed and their damage sums match, including direct Level-2 point hits and lethal depletion.
  • Branching level2-homing-exact-beam-0824-58.sav through neutral and all eight directions produces nine exact $661C matches.

These gates validate the corrected mechanisms. They do not by themselves prove completion of the remaining Level-2 route or boss, so future gameplay work must keep the gates green and treat any new unexplained transition as a model defect before changing policy.

Protocol-13 correction to the first resolution

The first protocol-12 acceptance result was incomplete. Its branch validator predicted a $661C transition from the input that $661C had already consumed, but did not prove that this was the mask Python had just requested. A fresh run exposed 76 directional mismatches in 311 steps. At the former damage frame 33142, the controller had already selected a different escape while the ship consumed an older mask. The apparent exact ship and camera counts were therefore internally consistent evidence for the wrong control boundary.

Protocol 13 replaces that boundary contract:

  • every SET/STEP state receives a generation and $661C records the generation it consumed;
  • after a STEP frame is emitted, XenonControl_OnRealFrameComplete synchronously pumps the control socket without returning to the 68000 until the next command installs its input;
  • snapshot branch tests explicitly release that synchronous barrier before restoring RAM, avoiding a reentrant restore from inside the $406 memory-write callback;
  • update-pass world coordinates use screenY + $CD0 + $CD8, while draw-pass tiles use screenY + $CD0; and
  • beam equivalence includes the complete homing-object state, so histories with different future retargets cannot be merged merely because the ship coordinates match.

Current acceptance evidence is level2-model-v13-0824-03.x2events, replayed from level2-homing-exact-beam-0824-58.sav: shield remains 35 from frame 32981 through 33240, including the formerly damaging frame 33142. validate_game_model.py reports 259/259 input-generation, 259/259 ship, 260/260 camera and 626/626 homing transitions with zero failures. Independent snapshot branches also pass neutral plus all eight joystick directions with requested and consumed generation 2/2 in every branch. The Python suite contains 326 tests after adding protocol-13 and coordinate-phase fixtures.