Xenon 2

Autopilot · write-up

Full-game regression recording

xenondoc/CAMPAIGN_VALIDATION.MD · 13 KB · updated 2026-10-06

Updated 2026-09-02. The driver is xenon_tools/validate_campaign.py. It observes the existing autopilot rather than introducing another tactical controller. No C/protocol changes or extra Python dependencies are required.

Resident C campaign endpoint (2026-10-05)

Use validate_campaign.py native-run for current validation. Its default endpoint includes the complete final shop: timed dialogue, transition/loading screens and the fire acknowledgement, ending only when Level 1 has an active player again. The byte $1178C selects the ending message and remains set after the game loops; it is not evidence that the ending has finished. driver-result.json records final_shop_selected_frame, final_shop_opened_frame and level1_restarted_frame under completion. The compatible stop reason remains final_completion. Explicit --stop-at-shop, --stop-at-level and --frames remain shorter probes.

See FINAL_SHOP_VALIDATION_20261005.MD for the focused continuation and both finalized MP4 recordings. The older Python run mode described below retains its historical endpoint and is not the current gameplay controller.

Start at the beginning

From the repository root:

python xenon_tools\validate_campaign.py run full-game-01 --build

This builds through hatari_dev.ps1 in build2, launches visible Hatari on port6903, loads assets/xenonplay.sav, and records at normal speed with no shield cheat. Omit --build to launch the already-built binary. If port6903 is occupied, either choose another --port, or explicitly use the visible instance:

python xenon_tools\validate_campaign.py run full-game-01 --connect --port 6903

The launcher uses the available PowerShell7 executable when present, otherwise Windows PowerShell. This avoids inheriting PowerShell7 module paths in an incompatible older shell. On Linux, start visible Hatari using the normal local launch instructions and use --connect; the recording/monitoring code itself is platform-independent. --connect assumes the chosen instance is visible.

Run names must be unique. Existing recordings are never overwritten. Ctrl+C requests graceful shutdown: drain both AVI writers, stop the event stream, save the final checkpoint and policy state, and release input. Campaign shutdown does not forcibly kill the controller after the old15-second timeout. Hatari stays open, paused, with keyboard control available. Do not force-close the recorder while it is flushing files.

By default the run continues across lost lives, shops and level transitions. It stops on the final-completion flag, a life-loss event leaving zero lives, low disk space, an error, or operator interruption. A transient zero-lives HUD during loading is not sufficient to stop. --frames 100 makes a bounded test.

Live reports and checkpoint semantics

The console reports damage with captured hit sources, lost lives, boss detection, shop entry/exit and transactions, level changes, cleared local threats, and wall stalls. Status prints at most once per five seconds. The same events are appended immediately to NAME.validation/events.jsonl and report.md; there is no need to wait until the run ends.

Slow autopilot warnings

Only the existing decision_ms is measured for this feature; rendering, AVI and total frame-processing time are excluded. The first decision taking 100ms or more emits slowdown_started immediately. Change the threshold with --slow-frame-ms 80. Nearby slow frames belong to one period; continuing warnings are throttled to once every five seconds. Ten consecutive measured fast frames emit slowdown_ended. Stopping before recovery emits slowdown_unfinished, not a false recovery.

Each warning includes the recording path, first slow game-frame number, last slow frame, peak time/frame, starting tactic, and nearest earlier checkpoint. That checkpoint reference is fixed at onset, including for shop slowdowns. If no checkpoint predates the slowdown, the archived start.sav is used when available; otherwise the reference is explicitly absent. No extra checkpoint or profiling work is triggered by slow decisions.

Warnings appear in the console, events.jsonl, report.md, and the compact slowdowns.jsonl stream. List periods during recording or after stopping:

python xenon_tools\validate_campaign.py slowdowns full-game-01

The list shows the latest state of each period, so refer to it as, for example, full-game-01.x2events, slowdown #2, beginning at frame 12345. Use its start_frame..last_slow_frame range with the offline profiler after stopping. An end warning's own frame is the recovery-detection frame; it does not move the slowdown's start or extend the slow-frame range through the fast tail.

Checkpoints are made every150 canonical frames by default, and at boss detection, shop entry/exit, confirmed/failed shop transaction acknowledgement, life loss, local-clear milestones, and wall-stall detection. Change with --checkpoint-interval 75. These are Atari logic frames, not extrapolated render frames or VBL numbers. The index also includes VBL, level, shield and lives.

Before-boss and before-shop references name the latest earlier playable save. Detection-time saves are honestly named boss-detected or shop-entry, not mislabelled as pre-entry. No driver can retroactively save the state before an unrecognized event. A pre-detection boss save is not proof that no earlier boss attack had begun. All periodic saves are retained, so an earlier one can be selected if necessary. Their frequency bounds the gap; increase frequency for an especially short approach.

Each save has Hatari's identity sidecar .sav.x2meta and .sav.autoplay.json. The latter includes the existing durable route/tactic state plus next input, fire pulse phase, transition state, and shop controller plan/pending acknowledgement. This preserves mid-shop purchase counts instead of starting its recipe again. It is not a claim that every transient Python cache is serialized or that a resumed branch is bit-for-bit identical to uninterrupted play.

A wall_stall requires a sustained, small world-coordinate motion range plus repeated wall-collision indications while trying to move. A quiet combat hold alone is not labelled a wall stall. Default window100 frames; change with --stall-frames. The detector only reports/checkpoints; it does not alter steering. It will not diagnose every oscillation spanning more than12 world pixels. Likewise local_threats_cleared means no tracked enemies for30 frames, not proof every offscreen enemy died. boss_section_completed requires reaching the shop; one missing draw frame never declares a boss defeated.

Resume a chosen checkpoint

python xenon_tools\validate_campaign.py checkpoints full-game-01
python xenon_tools\validate_campaign.py run full-game-02 --connect --port 6903 --resume "D:\src\hatari\xenon_tools\run_logs\full-game-01.validation\checkpoints\L3-f12345-boss-detected.sav"

Use the actual filename from checkpoints, not the illustrative frame above. The new run has its own outputs and archives its starting state, leaving the old recording untouched. Resume may also use an existing checkpoint made by record_run.py; absent policy sidecars are reported and reconstructed from RAM.

Outputs

Everything goes under ignored xenon_tools/run_logs:

Both AVI streams are always enabled by the campaign driver, including with --connect. New runs put all artifacts inside NAME.validation: events, JSONL, AVIs, final save, reports and checkpoint directory. Existing older runs are not moved, and the campaign reader still resolves their legacy sibling recordings. The driver prints their absolute paths at startup and shutdown and stores them in NAME.validation/manifest.json as original_avi and sprite_avi.

All paths in this table are relative to xenon_tools/run_logs/NAME.validation/.

File Purpose
NAME.x2events Full canonical capture for Python replay
NAME-original.avi, NAME-sprite.avi, both .vbl indexes Actual original and C Sprite Stream renderer video
NAME.jsonl Rich decisions, predictions, objects, inputs, hit telemetry, shop diagnostics and decision_ms
NAME.sav + sidecars Final state
report.md, events.jsonl, summary.json Incremental human report, event stream, end summary
slowdowns.jsonl Live slow-period warnings with recording/frame/checkpoint references
checkpoints.json, checkpoints/ Indexed intermediate saves with both sidecars
start.sav, start.ram, optional starting policy sidecar Archived input state
terrain.json, terrain-Ln-fN.ram The same per-level RAM terrain seeds used by live navigation
source/, working-tree.patch, manifest.json Python sources including new files, tracked working-tree changes, revision/runtime/resume metadata
catalog.json, final-map.json Existing behavior catalog/player metrics and discovered wall map
live-hotspots.json Post-stop ranking of live decision times

Paired lossless AVIs can be very large. The default10GiB free-space reserve is checked before starting and during play; low space triggers graceful saving, not deletion. --minimum-free-gb changes it. PNG compression remains level6 by default (--png-level overrides it). Checkpoints add brief pauses while the emulator is already stopped at a canonical boundary; short tests measured approximately 37--47ms per save, not continuous gameplay overhead.

Shopping audit and Double Shot

The current semantic recipes, compatibility/mount checks, MORE-page browsing and spend-remaining policy are reused unchanged. Advice, Autofire and ordinary-shop temporary Nashwan remain excluded. Double Shot is already a primary-weapon goal in the Level1 second-shop recipe, ahead of Side Shot upgrades. It is not the auxiliary Cannon item; logs retain the numeric item type and catalog name to avoid confusing them.

shop_transaction is emitted only when the controller receives the completed menu acknowledgement, not when it first presses FIRE. It records sale/purchase, item name/type, cash before/after, delta, and confirmed. A zero/wrong-direction cash change is explicitly unconfirmed. Animated100-credit installments do not produce duplicate transactions. Entry/exit record all attachment slots; ordinary JSONL still contains the plan, page, cursor, quotes and equipment every frame. This instruments the newer policy for validation; it does not claim it is already proven across every shop or blindly force a different loadout.

Replay and profile offline

python xenon_tools\replay_ui.py xenon_tools\run_logs\full-game-01.validation\full-game-01.x2events
python xenon_tools\validate_campaign.py profile full-game-01
python xenon_tools\validate_campaign.py profile full-game-01 --from-frame 12000 --to-frame 12100 --cprofile --top 20

The replay UI discovers the adjacent original/sprite AVIs. Profiling makes no emulator connection, uses the existing ReplayModel, warms earlier frames, and reapplies the recorded per-level RAM terrain seeds before decisions. It writes .timings.jsonl, .summary.json, and optionally .pstats inside the validation directory. Progress reaches100% at the requested destination.

Choose hotspots from live-hotspots.json after stopping, or run profile_autoplay.py summary NAME.jsonl against a completed trace. Instrumented cProfile milliseconds are not normal-speed performance; make a non-profiled pass too. Long late-frame intervals still require sequential warm-up from the recording start. Profiling observes recorded gameplay, not a counterfactual run. Replay's shop dispatch and speculative fire schedule differ from the live loop; RAM seeding fixes terrain equivalence, not those other documented limitations. See PROFILING.MD for attribution and visual profiler instructions.

Architecture and validation

validate_campaign -> record_run -> autoplay -> one XenonClient -> Hatari

The recorder owns startup/restore, the event log and final cleanup. Autoplay owns the sole live connection, both AVI start/stop calls and canonical STEP. After a decision it passes the already-produced trace row to CampaignMonitor, which reports events and requests a checkpoint callback on that same connection. There is no polling second client, worker AI, scoring adjustment or second world model. Offline profiling runs only the recording decoder and existing replay.

Tests: test_campaign.py, test_shop_strategy.py, test_profile_autoplay.py, test_replay_ui.py, and the existing autoplay suite. Live smoke03 recorded45 frames from the beginning with periodic saves and both AVIs; resume04 started from its frame15 save and recorded12 more frames. All four AVI files decoded. Offline cProfile ran frames15--20 without a connection. These test the driver, not a completed full-game regression run.