Autopilot · write-up
Full-game regression recording
Updated 2026-09-02. The driver is xenon_tools/validate_campaign.py.
It observes the existing autopilot rather than introducing another tactical
controller. No C/protocol changes or extra Python dependencies are required.
Resident C campaign endpoint (2026-10-05)
Use validate_campaign.py native-run for current validation. Its default endpoint
includes the complete final shop: timed dialogue, transition/loading screens and
the fire acknowledgement, ending only when Level 1 has an active player again.
The byte $1178C selects the ending message and remains set after the game loops;
it is not evidence that the ending has finished. driver-result.json records
final_shop_selected_frame, final_shop_opened_frame and level1_restarted_frame
under completion. The compatible stop reason remains final_completion.
Explicit --stop-at-shop, --stop-at-level and --frames remain shorter probes.
See FINAL_SHOP_VALIDATION_20261005.MD
for the focused continuation and both finalized MP4 recordings. The older Python
run mode described below retains its historical endpoint and is not the current
gameplay controller.
Start at the beginning
From the repository root:
python xenon_tools\validate_campaign.py run full-game-01 --build
This builds through hatari_dev.ps1 in build2, launches visible Hatari
on port6903, loads assets/xenonplay.sav, and records at normal speed with no
shield cheat. Omit --build to launch the already-built binary. If port6903 is
occupied, either choose another --port, or explicitly use the visible instance:
python xenon_tools\validate_campaign.py run full-game-01 --connect --port 6903
The launcher uses the available PowerShell7 executable when present, otherwise
Windows PowerShell. This avoids inheriting PowerShell7 module paths in an
incompatible older shell. On Linux, start visible Hatari using the normal local
launch instructions and use --connect; the recording/monitoring code itself is
platform-independent. --connect assumes the chosen instance is visible.
Run names must be unique. Existing recordings are never overwritten. Ctrl+C requests graceful shutdown: drain both AVI writers, stop the event stream, save the final checkpoint and policy state, and release input. Campaign shutdown does not forcibly kill the controller after the old15-second timeout. Hatari stays open, paused, with keyboard control available. Do not force-close the recorder while it is flushing files.
By default the run continues across lost lives, shops and level transitions.
It stops on the final-completion flag, a life-loss event leaving zero lives,
low disk space, an error, or operator interruption. A transient zero-lives HUD
during loading is not sufficient to stop. --frames 100 makes a bounded test.
Live reports and checkpoint semantics
The console reports damage with captured hit sources, lost lives, boss detection,
shop entry/exit and transactions, level changes, cleared local threats, and wall
stalls. Status prints at most once per five seconds. The same events are appended
immediately to NAME.validation/events.jsonl and report.md; there is no need to
wait until the run ends.
Slow autopilot warnings
Only the existing decision_ms is measured for this feature; rendering,
AVI and total frame-processing time are excluded. The first decision taking
100ms or more emits slowdown_started immediately. Change the threshold with
--slow-frame-ms 80. Nearby slow frames belong to one period; continuing
warnings are throttled to once every five seconds. Ten consecutive measured
fast frames emit slowdown_ended. Stopping before recovery emits
slowdown_unfinished, not a false recovery.
Each warning includes the recording path, first slow game-frame number,
last slow frame, peak time/frame, starting tactic, and nearest earlier
checkpoint. That checkpoint reference is fixed at onset, including for shop
slowdowns. If no checkpoint predates the slowdown, the archived start.sav is
used when available; otherwise the reference is explicitly absent. No extra
checkpoint or profiling work is triggered by slow decisions.
Warnings appear in the console, events.jsonl, report.md, and the compact
slowdowns.jsonl stream. List periods during recording or after stopping:
python xenon_tools\validate_campaign.py slowdowns full-game-01
The list shows the latest state of each period, so refer to it as, for example,
full-game-01.x2events, slowdown #2, beginning at frame 12345. Use its
start_frame..last_slow_frame range with the offline profiler after stopping.
An end warning's own frame is the recovery-detection frame; it does not move
the slowdown's start or extend the slow-frame range through the fast tail.
Checkpoints are made every150 canonical frames by default, and at boss detection,
shop entry/exit, confirmed/failed shop transaction acknowledgement, life loss,
local-clear milestones, and wall-stall detection. Change with
--checkpoint-interval 75. These are Atari logic frames, not extrapolated render
frames or VBL numbers. The index also includes VBL, level, shield and lives.
Before-boss and before-shop references name the latest earlier playable save.
Detection-time saves are honestly named boss-detected or shop-entry, not
mislabelled as pre-entry. No driver can retroactively save the state before an
unrecognized event. A pre-detection boss save is not proof that no earlier boss
attack had begun. All periodic saves are retained, so an earlier one can be
selected if necessary. Their frequency bounds the gap; increase frequency for
an especially short approach.
Each save has Hatari's identity sidecar .sav.x2meta and .sav.autoplay.json.
The latter includes the existing durable route/tactic state plus next input,
fire pulse phase, transition state, and shop controller plan/pending acknowledgement.
This preserves mid-shop purchase counts instead of starting its recipe again.
It is not a claim that every transient Python cache is serialized or that a
resumed branch is bit-for-bit identical to uninterrupted play.
A wall_stall requires a sustained, small world-coordinate motion range plus
repeated wall-collision indications while trying to move. A quiet combat hold
alone is not labelled a wall stall. Default window100 frames; change with
--stall-frames. The detector only reports/checkpoints; it does not alter steering.
It will not diagnose every oscillation spanning more than12 world pixels.
Likewise local_threats_cleared means no tracked enemies for30 frames, not proof
every offscreen enemy died. boss_section_completed requires reaching the shop;
one missing draw frame never declares a boss defeated.
Resume a chosen checkpoint
python xenon_tools\validate_campaign.py checkpoints full-game-01
python xenon_tools\validate_campaign.py run full-game-02 --connect --port 6903 --resume "D:\src\hatari\xenon_tools\run_logs\full-game-01.validation\checkpoints\L3-f12345-boss-detected.sav"
Use the actual filename from checkpoints, not the illustrative frame above.
The new run has its own outputs and archives its starting state, leaving the old
recording untouched. Resume may also use an existing checkpoint made by
record_run.py; absent policy sidecars are reported and reconstructed from RAM.
Outputs
Everything goes under ignored xenon_tools/run_logs:
Both AVI streams are always enabled by the campaign driver, including with
--connect. New runs put all artifacts inside NAME.validation: events, JSONL,
AVIs, final save, reports and checkpoint directory. Existing older runs are not
moved, and the campaign reader still resolves their legacy sibling recordings. The driver
prints their absolute paths at startup and shutdown and stores them in
NAME.validation/manifest.json as original_avi and sprite_avi.
All paths in this table are relative to xenon_tools/run_logs/NAME.validation/.
| File | Purpose |
|---|---|
NAME.x2events |
Full canonical capture for Python replay |
NAME-original.avi, NAME-sprite.avi, both .vbl indexes |
Actual original and C Sprite Stream renderer video |
NAME.jsonl |
Rich decisions, predictions, objects, inputs, hit telemetry, shop diagnostics and decision_ms |
NAME.sav + sidecars |
Final state |
report.md, events.jsonl, summary.json |
Incremental human report, event stream, end summary |
slowdowns.jsonl |
Live slow-period warnings with recording/frame/checkpoint references |
checkpoints.json, checkpoints/ |
Indexed intermediate saves with both sidecars |
start.sav, start.ram, optional starting policy sidecar |
Archived input state |
terrain.json, terrain-Ln-fN.ram |
The same per-level RAM terrain seeds used by live navigation |
source/, working-tree.patch, manifest.json |
Python sources including new files, tracked working-tree changes, revision/runtime/resume metadata |
catalog.json, final-map.json |
Existing behavior catalog/player metrics and discovered wall map |
live-hotspots.json |
Post-stop ranking of live decision times |
Paired lossless AVIs can be very large. The default10GiB free-space reserve is
checked before starting and during play; low space triggers graceful saving, not
deletion. --minimum-free-gb changes it. PNG compression remains level6 by default
(--png-level overrides it). Checkpoints add brief pauses while the emulator is
already stopped at a canonical boundary; short tests measured approximately
37--47ms per save, not continuous gameplay overhead.
Shopping audit and Double Shot
The current semantic recipes, compatibility/mount checks, MORE-page browsing and spend-remaining policy are reused unchanged. Advice, Autofire and ordinary-shop temporary Nashwan remain excluded. Double Shot is already a primary-weapon goal in the Level1 second-shop recipe, ahead of Side Shot upgrades. It is not the auxiliary Cannon item; logs retain the numeric item type and catalog name to avoid confusing them.
shop_transaction is emitted only when the controller receives the completed
menu acknowledgement, not when it first presses FIRE. It records sale/purchase,
item name/type, cash before/after, delta, and confirmed. A zero/wrong-direction
cash change is explicitly unconfirmed. Animated100-credit installments do not
produce duplicate transactions. Entry/exit record all attachment slots; ordinary
JSONL still contains the plan, page, cursor, quotes and equipment every frame.
This instruments the newer policy for validation; it does not claim it is already
proven across every shop or blindly force a different loadout.
Replay and profile offline
python xenon_tools\replay_ui.py xenon_tools\run_logs\full-game-01.validation\full-game-01.x2events
python xenon_tools\validate_campaign.py profile full-game-01
python xenon_tools\validate_campaign.py profile full-game-01 --from-frame 12000 --to-frame 12100 --cprofile --top 20
The replay UI discovers the adjacent original/sprite AVIs. Profiling makes no
emulator connection, uses the existing ReplayModel, warms earlier frames,
and reapplies the recorded per-level RAM terrain seeds before decisions. It
writes .timings.jsonl, .summary.json, and optionally .pstats inside the
validation directory. Progress reaches100% at the requested destination.
Choose hotspots from live-hotspots.json after stopping, or run
profile_autoplay.py summary NAME.jsonl against a completed trace. Instrumented
cProfile milliseconds are not normal-speed performance; make a non-profiled
pass too. Long late-frame intervals still require sequential warm-up from the
recording start. Profiling observes recorded gameplay, not a counterfactual run.
Replay's shop dispatch and speculative fire schedule differ from the live loop;
RAM seeding fixes terrain equivalence, not those other documented limitations.
See PROFILING.MD for attribution and visual profiler instructions.
Architecture and validation
validate_campaign -> record_run -> autoplay -> one XenonClient -> Hatari
The recorder owns startup/restore, the event log and final cleanup. Autoplay owns
the sole live connection, both AVI start/stop calls and canonical STEP. After a
decision it passes the already-produced trace row to CampaignMonitor, which
reports events and requests a checkpoint callback on that same connection.
There is no polling second client, worker AI, scoring adjustment or second world
model. Offline profiling runs only the recording decoder and existing replay.
Tests: test_campaign.py, test_shop_strategy.py, test_profile_autoplay.py,
test_replay_ui.py, and the existing autoplay suite. Live smoke03 recorded45
frames from the beginning with periodic saves and both AVIs; resume04 started
from its frame15 save and recorded12 more frames. All four AVI files decoded.
Offline cProfile ran frames15--20 without a connection. These test the driver,
not a completed full-game regression run.