Xenon 2

Autopilot · write-up

Three native fast-forward campaigns, October 2–3, 2026

xenondoc/THREE_NATIVE_CAMPAIGNS_20261003.MD · 21 KB · updated 2026-10-05

Outcome and reproducibility

Committed the terrain-route fix as cc14b5f0b36304d07980de1b2c82a48535ca79a0 (Preserve Level 2 terrain routes through retreats and scrolling crossings). Rebuilt the Release Hatari executable and native replay DLL using xenon_tools/hatari_dev.ps1 before starting any campaign. No pilot source was changed between the runs or during their analysis.

All three attempts started at Level 1, completed Levels 1 and 2, and exhausted all three lives in early Level 3. None completed the campaign. Levels 4 and 5 were not reached and are not validated by these runs. Gameplay damage, life-loss frames, equipment changes, missed equipment and shop decisions were identical across the three attempts. They repeat the same starting conditions; they do not measure robustness across different loadouts or starting states.

Recording Last recorded game frame Termination
R1 19477 no_lives_remaining
R2 19463 no_lives_remaining
R3 19486 no_lives_remaining

The recordings are:

  • R1: E:\xenon_runs\campaign-three-fast-20261002-r192-1.validation\campaign-three-fast-20261002-r192-1.x2events
  • R2: E:\xenon_runs\campaign-three-fast-20261002-r192-2.validation\campaign-three-fast-20261002-r192-2.x2events
  • R3: E:\xenon_runs\campaign-three-fast-20261002-r192-3.validation\campaign-three-fast-20261002-r192-3.x2events

Each folder contains its original and sprite AVIs, their VBL indices, manifest.json, incidents.jsonl, driver-result.json, start.sav, start.ram, source snapshot and periodic checkpoints. The filename stem is shared by the recording and its -original.avi / -sprite.avi files. level2-boss.trace contains the enabled Level 2 diagnostics.

All manifests specify resident-c-autopilot, driver_backend=hatari-c, per_frame_controller_calls=false, full_speed=true, paired_avi=true, checkpoint_interval=150 and cheat=false. Hatari was launched visibly through the official launcher. Python supervised recording and audited captured facts; it did not drive the game. There were no monitor-defined stalls, recording finalization errors or cleanup errors. Owned Hatari processes were removed after AVI finalization.

The snapshot SHA-256 was 651d0f8931aa3ce5b5a9cd89aacda5fa7dc6f6aef1ef6792d58dcf1d32a90439; the executable SHA-256 was d4114175c007ca3b44be8f1feb940b80ab3468dd25afb34c9d463bebca4575be in all three manifests. Official build logs are D:\src\hatari\work\l2-source-comparison\campaign-three-kernel-build.log and D:\src\hatari\work\l2-source-comparison\campaign-three-hatari-build.log.

Reproduce one attempt from D:\src\hatari using a fresh run name:

$env:XAP_LEVEL2_TRACE_FILE = 'E:\xenon_runs\NEW-NAME.validation\level2-boss.trace'
python xenon_tools\validate_campaign.py native-run NEW-NAME --output-root E:/xenon_runs --resume D:/src/hatari/assets/xenonplay.sav --build-type Release --fast-forward --checkpoint-interval 150 --continue-after-life-loss --continue-after-stall --port 6903

Damage and lost lives

The following table applies to each of R1, R2 and R3, at the same game frames. The accompanying THREE_NATIVE_CAMPAIGNS_20261003_INCIDENTS.csv expands every incident into a row containing its complete recording filename. Shield losses are actual captured shield decreases, not merely hook events.

Level Frame Shield loss / lives Collision source and short description
2 9381 4 Directional projectile #19183, $4180, hits during the early corridor passage.
2 9449 4 Directional projectile #19422, $4180, hits later in the same passage.
2 14544 8 Lower-labyrinth scripted swarm #36468, $50346, collides with the ship.
2 14546 8 Another member, #36467, $50346, collides two frames later.
2 14829 8 A later lower-labyrinth swarm #37616, $50346, hits the ship.
2 15356 4 Directional projectile #40441, $4180, hits during the spider encounter.
2 15658 4 Directional projectile #41997, $4180, hits before spider defeat.
3 18673 16 Scripted swarm #50091, $4F244, hits near the top of the central passage.
3 18742 11 Scripted swarm #50359, $4F244, removes the first life's remaining shield.
3 18759 Lives 3 → 2 First life lost after the preceding swarm collision.
3 18997 16 Scripted swarm #51538, $4F244, hits after respawn.
3 19060 16 Scripted swarm #51808, $4F244, hits again in the same early section.
3 19149 7 Scripted swarm #52484, $4F244, removes the second life's remaining shield.
3 19166 Lives 2 → 1 Second life lost.
3 19316 8 Expanding-formation leader #52726, $5073C, collides after the next respawn.
3 19376 16 Scripted swarm #53309, $4F244, hits the last ship.
3 19439 15 Scripted swarm #53576, $4F244, removes the last remaining shield.
3 19456 Lives 1 → 0 Third life lost; campaign attempt ends.

Per attempt: Level 1 has no shield loss; Level 2 has seven shield-decrease incidents totaling 40 damage, without losing a life; Level 3 has eight totaling 105 damage and loses three lives. Level 3 begins with shield 27. Cumulative damage includes the replenished shields after respawns, not just the change between a level's starting and final shield values. Both Level 2 bosses were defeated; the three-eye fight caused no shield decrease in these attempts.

What the checks establish

Level 3: danger is predicted, but the selected escape fails

At frame 18673 the player is at world (152,3279), screen Y=26. At the actual collision, swarm #50091 is at screen (169,35). Its collision bounds touch the player's bottom edge. The player is above the approaching body and cannot rely on forward weapons to destroy it from that position. Repeated ceiling holds and short Up/Down reversals precede the next fatal hit at 18742. Captured wall-replay flags are zero throughout 18200–18743. This is before Level 3's first destructible gate, not a gate-egress failure.

The shared C path audit checked surviving $4F244 bodies against later canonical C anchors. All 175 eight-frame and 131 sixteen-frame comparisons were exact. Of 245 one-frame comparisons, one differed by two pixels. No paths were missing. These checks establish useful motion accuracy for this sample; they do not validate every future spawn, collision box or predicted destruction.

Replaying the complete captured history through the current C DLL reproduces all 543 steering masks from 18200–18742 in the next recorded frame. That one-frame offset is the resident input-latch delay. The comparison ignores Fire and does not claim identical combat state solely from a steering match.

The reconstructed decision at 18620 already predicts a body contact in 34 frames. At 18640 it predicts contact in 26 frames and still selects Up. At 18660 the ship is near the ceiling and predicted contact is 12 frames away. These choices have verified=0, contact reason XAP_CANDIDATE_BODY and search purpose XAP_SEARCH_NO_ESCAPE. The formation-hold mission asks for screen Y=120–136; the final search choice abandons that posture to postpone predicted contact. It repeatedly returns an unsafe fallback rather than finding a safe crossing or preserving space earlier.

Therefore, a missing swarm trajectory is not the main explanation for these Level 3 hits. The immediate failure is the escape search and resulting posture. Whether correcting optimistic future kills improves the earlier decisions still needs a controlled live comparison; that causal link is not established by the motion audit.

Evidence: D:\src\hatari\work\l2-source-comparison\r192-1-level3-history.json, r192-1-level3-paths.json and the original AVI screenshot D:\src\hatari\work\l5-boss-campaign-three-fast-20261002-r192-1-18673.png.

Correction recorded after disassembly (October 3)

The $50346 handler below was initially misidentified as a sine-script handler. Saved Level 2 RAM proves it is a ship-aligned diver with climb/chase/dive phases. See CAMPAIGN_FIXES_20261003.MD for the corrected rules and controlled validation. Adding it to the motion-script capture list would be incorrect. The baseline still correctly establishes that fitted linear motion missed its turns.

Level 2: a known curved handler falls back to fitted straight motion

XENON_UPDATE_PROC_LEVEL2_SCRIPTED_SINE_SWARM ($50346) is recognized by xenon_autopilot_world_classify.c, but is absent from both:

  • src/xenonControl.c:1397 — the motion-script capture list;
  • src/autopilot/xenon_autopilot_world_prediction.c:122 — scripted_proc().

Its captured script payload is consequently empty and its C prediction model is XAP_MODEL_POLICY, which uses fitted velocity in src/autopilot/xenon_autopilot_predictions.c:268. Classification as a swarm does not imply that its actual movement script is interpreted.

For the actual collision source #36468 at frame 14535, C predicts position (206,864) one frame later; the observed position is (208,847). Eight frames later C predicts (206,941), whereas the observed position is (230,805). For #37616 at 14820, the eight-frame prediction is approximately (152,435); the observed position is (176,516). The curved/reversing motion makes a straight extrapolation unreliable precisely where these hits occur.

Across 14000–14900, 1680/1685 eight-frame and 1397/1405 sixteen-frame surviving swarm comparisons differ by more than one pixel. Maximum axis errors are 136 and 272 pixels respectively. The audit uses a fixed -1px camera forecast; that limitation cannot explain the missing script dispatch or these large errors. No claim is made that simply registering the handler will suffice: its wrapper's coordinate convention and update timing must also be checked.

Evidence: D:\src\hatari\work\l2-source-comparison\r192-1-level2-maze-paths.json and r192-1-level2-contact-paths.json.

Directional projectile motion and Cannon creation are different issues

The C directional-projectile paths in the early Level 2 corridor and lower maze match captured motion within one pixel at horizons 1, 8 and 16. The same is true in the sampled Level 3 section. This does not validate every collision shape or future firing prediction, but provides no evidence for a blanket direction/velocity repair of existing $4180 rounds. Corridor room and dodge timing should be investigated locally before changing this shared trajectory.

The previously documented Cannon creation error is independently reproduced in Level 1. At frame 5300, the next 56 frames contain 12 actual Cannon births, while the present combat forecast creates 28 rounds from Fire intent times. The observed gaps are four or five frames. Captured attachment animation state predicts every next-tick allocation in the 551-frame audit with zero mismatches. This confirms a shared firing-state error; it is not a Level 2 geometry rule. It can let evaluation assume enemies disappear too early, but the contribution to each individual hit remains to be established.

Keep this correction separate from movement tactics: derive Cannon allocations from each mount's actual internal phase instead of equating a Fire pulse with a new projectile. The prior investigation in FULL_NATIVE_CAMPAIGN_20261001.MD also describes update order, spawn offsets and remaining damage-model limits. The Cannon model was not changed for these runs.

Evidence: D:\src\hatari\work\l2-source-comparison\r192-1-cannon-audit.json, r192-1-level2-corridor-paths.json, r192-1-level2-maze-paths.json and r192-1-level3-paths.json.

Equipment, health and cash

These outcomes are identical in all three recordings. Equipment disappearance is inferred from captured pickup lifetimes and proximity; collection lacks an explicit event. Weapon loadout changes independently confirm Cannon and Side Shot acquisition. Cash counts below are likely collections, not exact cash accounting based on collection hooks.

Level Pickup Appearance–expiry Finding
2 POWERUP #21924 10048–10108 Missed; closest collision-rectangle gap 42px. Double Shot remains at power 0.
2 ZAPPER #31749 13420–13491 Missed; closest gap 18.97px.
3 HEALTH POWER 1 #46824 18010–18079 Missed while shield is 27; closest gap 24.74px.
1 HEALTH POWER 1 #7977 4280–4356 Missed while shield is already full; no useful shield recovery was lost.
2 HEALTH POWER 1 #16727 8770–8838 Missed while shield is full; do not force a dangerous pursuit.

Level 1's Cannon was collected at 5017. Level 2's Side Shot was collected at 13917 and replaced Rear Shot as intended. Two Level 3 Rear Shot pickups were skipped, correctly retaining Side Shot. Side Shot disappears temporarily at 18742, 19149 and 19439 during fatal damage/death sequences; this is not a rear-weapon pickup replacing it. The restored loadout is observed after the first two respawns.

The Level 2 POWERUP acquisition phase remains active through its lifetime. At 10080 and 10100 it still targets #21924, but hazard escape replaces the movement toward it with Fire only. This miss is not carrier visibility flickering or premature eye targeting. The ship is already far to the right when the pickup's short remaining crossing opportunity occurs. The proposed fix is earlier safe staging/interception before the worm blocks the crossing, not suppressing collision avoidance to chase an expiring pickup.

Level Likely cash collected Cash missed Unresolved at recording end
1 44 23 0
2 52 32 0
3, partial 22 14 3

Reward bursts: Level 1's final burst at 6833 is 17/18 likely collected; #14419 expires at 6907. Level 2's three-eye burst at 11461 is 10/10. The spider burst at 15727 is 18/20: #42508 expires at 15821 and #42490 at 15824. Both missed spider coins previously came within a 2px rectangle gap, so the final short interception needs inspection. Do not replace working boss combat merely to repair its remaining two loot pickups.

Shops

Shop visits and loadouts are identical across all three runs. Captured cash and equipment show:

Level / visit Frames Cash before → after Purchase / resulting loadout
1 / first 2041–2919 750 → 750 No purchase; Forward Shot power 1 and Rear Shot retained.
1 / second 6908–7899 3500 → 500 Double Shot at 7496 replaces Forward Shot; Cannon and Rear Shot retained.
2 / first 11530–12496 1550 → 1050 Health, +20 subject to shield cap; equipment unchanged.
2 / second 15835–16944 4550 → 50 Laser at 16429 and Health +20; Double Shot, Cannon and Side Shot retained.

The health purchases are supported by shop shield recovery and expenditure, not an attachment change. Detailed shop audit files are D:\src\hatari\work\l2-source-comparison\r192-1-shops.md, r192-2-shops.md and r192-3-shops.md.

Recording and monitor issues

At 13843, campaign_monitor.py reports a damage incident even though shield remains 39. HEALTH POWER 1 #33212 retires beside the ship at that frame, and the hit hook's other_damage_caller entry has no valid object identity and contains unrelated object fields. This is consistent with the capped healing callback being treated as damage. It must not inflate the actual shield-loss count. A general telemetry fix should distinguish healing from hostile damage; unchanged shield alone is insufficient if healing and a real hit coincide.

All six AVIs have complete indices, continuous VBL tags and decodable first and last frames. They were finalized gracefully. There is nevertheless a general compatibility issue: R2's sprite AVI starts its next RIFF at odd byte 1074392413, and R3's at 1074388859, without a padding byte. The existing hatari_avi.py reader accepts these files. A strict RIFF reader can reject them despite successful finalization.

src/avi_record.c's PNG frame writer does not pad odd PNG payload sizes before subsequent chunks. Fix padding and the affected size/index accounting in that writer, then validate a recording crossing the OpenDML split. This is separate from gameplay. The audit now records unpadded RIFF boundaries as warnings instead of silently treating a compatible, indexed AVI as truncated. It still rejects invalid or out-of-file segment lengths.

Prioritized proposed fixes

  1. Correct shared prediction inputs before further tactical tuning. Register/capture $50346 and verify its exact wrapper convention in C; correct Cannon allocations from internal attachment state. These are model correctness fixes in shared preparation/evaluation. The $50346 handler is level-specific, so its recognition must remain Level 2 scoped; do not apply its semantics to unrelated overlay code. Replay surviving body anchors and shot births, then run the known corridor and boss checkpoint sets. Do not globally remove future shots or disable destruction modeling. Add a coverage check linking known script-handler registration to capture and decoding so classification cannot silently conceal a missing movement model again.

  2. Resolve the fatal early Level 3 swarm encounter. Preserve reaction space before entering its narrow passage, and compare a safe lower firing/escape lane with the current ceiling-seeking fallback. Selection must use world position, live script phase and map geometry, not recorded frame numbers. Begin with a local Level 3 encounter policy. If candidate inspection proves the same search defect elsewhere, repair that specific shared defect with those other encounters included in validation. Do not globally forbid upper refuges required by the Level 2 bosses. Check from multiple checkpoints and from Level 1 with this actual loadout; require no life loss and no repeated body hits through this passage.

  3. Recover the valuable Level 2 POWERUP and Zapper. Stage the ship in the drop lane before the three-eye pickup releases and intercept the post-shop Zapper when it is safe and useful. Preserve temporary hazard overrides. These are local encounter timing changes unless the same pickup scheduling defect is demonstrated elsewhere. Confirm both acquisition and continued clean three-eye/spider fights, not just target selection.

  4. Remove remaining Level 2 damage with local passages/boss policies. Revalidate the lower maze after $50346 motion is repaired before adding another pacing rule. For 9381/9449 and spider 15356/15658, preserve maneuvering room and begin lateral dodges early enough for queued input. Existing bullet direction checks pass, so do not change every level's projectile trajectory without a demonstrated model defect. Verify the terrain-route retreat and both boss checkpoint suites remain intact.

  5. Improve safe pickup and reward completion. Collect Level 3 health while shield is below its cap; diagnose the two near-missed spider coins and one Level 1 reward coin before changing common pickup scoring. Generalize a confirmed shared deadline/collection-envelope defect, but keep boss loot phases local where their geometry requires it. Require reward acquisition to finish before ordinary progression resumes. Do not force health detours at full shield or under predicted collision.

  6. Repair general observability/AVI compatibility. Classify capped healing correctly and pad PNG chunks/RIFF boundaries with matching index offsets. These affect trustworthy reports and playback; they do not solve the campaign's collisions.

After the fatal Level 3 blocker is fixed, repeat the three complete attempts. Only then can the campaign assess Levels 4–5, their boss rewards, shops and end-of-game completion. No proposed gameplay fix above was implemented as part of this analysis request.

Reusable analysis

work/audit_native_campaign.py audits native recordings and paired AVIs. work/audit_native_hazard_paths.py checks the current C predictor against later captured C world anchors, without invoking a Python pilot. It compares only surviving identities with unchanged update handlers and explicitly reports the camera-forecast limitation. For example:

python work\audit_native_hazard_paths.py E:/xenon_runs/campaign-three-fast-20261002-r192-1.validation/campaign-three-fast-20261002-r192-1.x2events --first 14000 --last 14900 --output work/l2-source-comparison/r192-1-level2-maze-paths.json

Per-run machine audits are D:\src\hatari\work\l2-source-comparison\r192-1-audit.json, r192-2-audit.json and r192-3-audit.json. They include source identities, procedure addresses, pickup lifetimes, loadout changes and AVI indexing data.