Xenon 2

Autopilot · write-up

Level-1 shell stations and spans validation

xenondoc/AUTOPILOT_SHELL_VALIDATION.MD · 9 KB · updated 2026-09-17

The subsequent top-edge/corridor corrections are documented in AUTOPILOT_TRAVEL_PACING.MD.

This follows AUTOPILOT_NAVIGATION.MD. The changes are isolated in autopilot_v2; the legacy planner remains available for comparison. Use --planner maneuver --navigation-backend spans with autoplay.py, record_run.py or validate_campaign.py run. An explicit backend overrides the backend in a restored controller checkpoint. The compatibility default remains grid; selecting spans is explicit in the recordings below.

Shell encounter

The right shell at frame 7842 is a reachable cannon target. Its narrow animated collision header previously invalidated the attack proposal, and the span route could send the ship back through a corridor centre instead of directly underneath the shell. The distant eye links did not justify abandoning that firing station.

autopilot_v2/shells.py retains the target identity and firing station across animation changes. A local connector tests the actual player/wall mask and takes the furthest reachable waypoint. A dodge reconnects to the same station; a dead target or a station that has become physically blocked releases it. Only the station proposal uses the observed horizontal animation envelope. Actual contact and predicted destruction still use the live target geometry. Narrow endpoints use a half-pixel goal tolerance, so a generic goal band cannot stop the ship just outside the cannon lane. Terrain safety is never relaxed because a shell can die.

The old unconditional shell Y-release rejection also prevented progression past unreachable shells. The new validator considers the three possible emitted projectiles for every eligible release time. The release can happen regardless of horizontal player distance; distance only determines when the shots can reach the ship. A predicted kill cannot remove projectiles already emitted.

The setup at $4F982 uses shell Y+7, directions 1/2/3 on the left and 5/6/7 on the right, with Level-1 speed scale eight. The forecast includes a timing and sprite envelope and assumes no favourable RNG outcome. Burst geometry and shell station proposals are cached; local maneuver validation still runs every frame.

Final eye encounter

Passing the shells exposed a separate prediction error: the eye's articulated links copy the predecessor's same-update position and then integrate again. The old generic follower forecast was incorrect. autopilot_v2/eyes.py shares one chain forecast per controller/decision using the existing boss_eye_motion.py integer model. A recorded 1/4/8/16-frame sample had exact X throughout; Y was exact or one pixel low. The validator retains an additional three-pixel envelope.

The damage callback $503DA tests the fixed body core at X=152..167 and controller Y+68..82. The object historically classified boss_eye_weak_point is the moving tentacle tip, not this damage region. The new mission aims the ship's cannon at the fixed core and stays near the bottom of the screen. Short approach/retreat sequences can exploit a firing window when holding the lane for the entire prediction horizon would collide with the arm. A core attack can persist beyond the ordinary eight-frame refresh, but its complete remaining trajectory is revalidated each frame. Extending that trajectory waits at its retreat station; it does not append an unplanned second attack. Aim opportunities influence only the choice among safe candidates and never grant predicted boss destruction.

Replay navigation checkbox

The Navigation strategy checkbox is off initially and works independently of the ordinary diagnostic overlay. It draws on the scene and world map:

  • Cyan dashed line: mission route.
  • Yellow rectangle: finite movement goal.
  • Orange circle and target identity: retained firing station.
  • Green line: verified maneuver; red line: an unverified fallback.

New decision traces contain the world-coordinate strategy and backend. The .planner.json sidecar identifies the planner, navigation backend and fire schedule. Binary replay loads its companion decision trace lazily when the box is checked and joins on both game frame and VBL. The label says Recorded navigation when this data exists, or Analysis navigation when showing a recomputed plan. Older recordings without strategy data remain usable. Partial pending-spawn labels in old JSONL traces are not interpreted as full motion records.

Recorded evidence

Recordings are under xenon_tools/run_logs, with original and sprite AVIs, decision traces, controller checkpoints and source snapshots in each validation directory. All runs used spans, real-time emulation and no shield cheat.

Run Region Shield damage Lives lost Outcome
shell-station-spans-0905-02 7683 onward 0 0 Right shell #17577 destroyed at 7855; old release guard still stalled progression
shell-burst-spans-0905-03 6247–9246 78 2 Passed shells; exposed eye prediction failure at 8019–8023
shell-eye-spans-0905-04 6247 onward 0 0 Corrected chain forecast avoided damage, but wrong attack objective stalled
shell-core-spans-0905-06 7633–9032 0 0 Core health fell from 30; boss disappeared at 8976; shop entered at 9031 with cash 3500

These checkpoint runs explain the failure modes and fixes. The from-opening Level-1 replay below was recorded in two consecutive parts, split at the mid-level shop and resumed from its saved Atari/controller state.

The from-opening replay level1-spans-shell-fix-0905-02 reached the mid-level shop at 2612 with 39 shield damage and one life lost, 13,130 score gained, 10 wall shooters destroyed and 16 pickups. Swarm hits at 1409, 1800–1803 and 2242 (life lost at 2259) remain an acceptance failure. The continuation through the shop and second half is level1-spans-shell-fix-0905-03.

That continuation completed Level 1, defeated the eye at 6827 and stopped on the first playable Level-2 frame at 8225. It lost no further lives and took 8 shield damage, from directional projectiles at 5784 (swarm/shell passage) and 6500 (eye encounter). It gained another 19,800 score and recorded 33 pickups and six wall-shooter destructions. Shop transactions confirmed SPEEDUP at 3137 and DOUBLE SHOT at 7487; the final purchase left 400 cash. Combined result: 47 shield damage, one life lost, final shield 31, two lives, score 32,930. This is level-completion evidence, not a regression-free or low-damage acceptance pass. No claim is made that the new projectile contacts would or would not occur on an independently driven old-controller trajectory.

Replay files:

From the repository root, open either file with python xenon_tools/replay_ui.py followed by its path, then enable Navigation strategy. The companion .jsonl, .planner.json and AVI files are in the same directory. The emulator was left paused at Level-2 entry with automation input released.

To isolate the opening failure, the committed spans controller at 097b71dee5221a807c4a1c97e6fe2d1e7d163684 and the changed controller were run sequentially against the same canonical observations from frame 200 through 1810. All 1,611 actions and tactics matched, including the damaging swarm. This establishes that the new shell/eye changes did not alter the decisions in that sample; it is not a separate live counterfactual or a claim that the swarm is safe. See shell-change-opening-comparison-0905.json.

The broad Python regression suite passes 688 tests, including station persistence, wall blockage, target identity, checkpoint restoration, shell burst timing, recorded eye motion, core lane/attack retention, backend metadata and overlay selection. Actual Tk smoke checks exercised checkbox on/off with a binary recording and its companion trace. No C source changed, so no emulator rebuild was needed. Browser/phone performance and Levels 2–5 gameplay acceptance remain outside this Level-1 regression result.

After the live run, two display-only guards were added to suppress a stale maneuver when no current rollout exists and to omit gameplay navigation during shop/transition frames. The focused planner/replay suite then passed 61 tests; these guards do not change control inputs or predictions.