Xenon 2

Autopilot · write-up

Plan: web replay inspector (port of replay_ui.py to the site)

xenondoc/plan-web-replay-ui.md · 93 KB · updated 2026-10-03

Status: implemented on branch web-replay, phases 0-7 plus a completion pass; summary and outstanding items in §0, detailed results and deviations in §12. Written 2026-09-24 against branch xenon2 at 3f64e554 with a clean working tree. Revision 2 (same day): decided the recording/session start rule, added the spans overlay, a full description of what the C autopilot predicts and how each part is shown, mission route/goal/geometry and role visualization, the three.js overlay design, the 12 KB tile-map record instead of RAM snapshots, desktop + web recording (AVIs desktop-only), and re-synced the Level-2 fields with the committed code.

0. Status summary (2026-09-28)

What was implemented

  • C replay core (src/autopilot/xenon_autopilot_replay.c/.h, xrp_*): decodes .x2events, steps the same native autopilot session the live run used, keeps checkpoints for seeking and answers JSON inspector sections (status, decision, mission, tracks, predictions, scene, plan, spans, map, geometry, draws). Built as the native DLL (out/build/autopilot-replay) and as the page's WASM module (hatari_dev.ps1 -Action build-replay-wasm). The autopilot itself only gained additions: read-only accessors and diagnostic exports (Level 2 geometry, span lists, and the arbitration trail the driver writes but never reads).
  • Recording in both builds: the desktop recorder (src/xenonRecorder.c) and the web recorder (Record button, ?record=1, recordings kept in the browser's IndexedDB). Recordings carry the frames, a recording-info record, the 12 000-byte tile map and the autopilot's start/stop/sync events and live status. xenon_tools/pack_recording.py converts older runs.
  • /replay/ page (site/content/replay.html, web/src/replay/): recording library with import and export; a worker holding compressed chunks and checkpoints; transport and timeline with markers; the game drawn with the web game's own three.js sprite renderer (original and HD sprites); the overlay (layers, object categories, presets, look-ahead); grouped state cards; selected-object card with its sprite; sortable object table; level map; deep links (?rec=, ?src=, frame=, sel=); a ?debug=1 hook for scripted checks.
  • Explanations: an "i" button on every layer, card row and level-specific field, with texts checked against the C code. The texts for the level fields, the configuration and the driver stages are marked comments in the C headers. The module returns them (xrp_describe()), and the JSON writers for those fields are generated from the same comments (§12, "Explanations from C").
  • Overlay rules settled on the way: future positions are drawn only up to the end of the autopilot's plan, with the plan's own camera; the goal band is screen rows.
  • Python replay path fixes (replay_ui.py now matches the live capture): the wall-collision replay flag and the protection byte.
  • Verification tools: decoder_parity.py, replay_parity.py, render_parity_frames.py, check_replay_wasm.mjs.

Verified (details in §12): the C decoder and the Python path hand the autopilot identical input on every frame of recordings from all five levels (over 460 000 frames, 458 000 of them from older runs); replayed decisions equal the live ones on every stepped frame of the recordings that carry live status (Levels 2 and 5); the 12 000-byte tile map gives the same map as full RAM (Levels 3 and 5); rendering is pixel-identical to the desktop Sprite Stream AVI (12 of 12 frames on Levels 2 and 5); exported recordings re-import byte for byte; recording costs no measurable emulation time (2.2 ms/VBL either way, about 2 KB per game frame).

Outstanding

  1. Autopilot merges: xenon2 up to 4a03e159 (Level 3 carrier, Level 4 and 5 tactics) and xenone2-l2 (Level 2 fixes) were merged and verified (§12, "Merge with xenon2" and "Merge with xenone2-l2"); xenon2 holds all of it. Shop and level-start screens do not render pixel-identically in the replay (also before the merge; §12), and Level 5 now costs 8-21 ms per replayed frame.
  2. Sample recordings under /replay/samples/, as decided, once the autopilot finishes the game.
  3. Publishing: the page has only been built and served locally (out/site-replay); the public site has not been rebuilt with --replay-build.
  4. Phone performance was not measured (desktop only).
  5. Explanations from C, not done: * the desktop replay_ui.py could show the same texts through xrp_describe(); * the value names (eye plan steps, carrier refuges...) are explained inside their field's text, not one by one; * Level 1's central passage and several other authored areas still have their numbers inline, so they are not drawn yet (§12, "Constants in texts and per-level geometry").
  6. Desktop overlay: autoplay_ui.py (used by replay_ui.py) still draws predicted paths past the plan with the camera extrapolated from the current scroll velocity; decide whether to apply the web page's rule there.
  7. Autopilot leftovers noticed while writing the explanations (autopilot work, not replay): Level 2 refuge_reached, crossing and refuge_hold_frames are reset every frame by the three-eye sequence and only matter to its fallback route; the eye transfer stages climb_left, climb_top and cross_top are never set.

1. Goal

Replace the Tk tool xenon_tools/replay_ui.py (which reuses autoplay_ui.py's AutoplayWindow) with a browser page on the site. It is not a 1:1 port. The requested differences:

# Requirement Where the plan answers it
R1 Drop everything tied to the obsolete Python autopilot; the C autopilot in src/autopilot is the only authority, compiled to WASM §3, §4
R2 Render the game view with the web game's own sprite-stream renderer (three.js) and overlay diagnostics on top §4.4, §6.2
R3 Filters for the crowded top-left overlay (e.g. hide player bullets) §6.2
R4 List what is obsolete with the C autopilot; visualize what the C autopilot uses instead §3, §5
R5 Group the upper-right autopilot state list into categories §7.3
R6 Sortable/filterable object list §7.5
R7 Selected sprite shown inside the "selected object" section §7.4
R8 Site look and feel, integrated into site/ §7.8, §4.5
R9 Record in both the web and the desktop version, access recordings later (library / upload) §6

2. What exists today (facts this plan builds on)

Python replay pipeline. replay_ui.py indexes an .x2events file (Recording), then ReplayModel.seek() walks frames. For native recordings it steps the complete C session through ctypes (_step_native_session → autopilot_v2/native_session_host.Host → native_session.Session → xap_session_step). Everything the Tk window shows is read back through Python views (WorldView, MapView, DriverView) and native_driver.inspect(), which fills old Python controller attributes (recovery_path, movement_goal, navigation_strategy, …) so the legacy AutoplayWindow can draw them. Frame decoding is xenon_client.decode_frame (≈440 lines, protocol versions 6–16) followed by Python packing (native_memory.pack, native_observation.pack).

C autopilot. XapSession owns world, map, mission, driver, shop and input phase (xenon_autopilot_session.h). Read-only inspection handles exist: xap_session_status, xap_session_world/xap_session_observation (tracks, decoded XapObjectFields, walls), xap_session_driver → xap_driver_mission_state (route, destination, per-level state XapLevel1..5MissionState), xap_driver_mission (tactic, target, weapon, goal rectangle, band, lane, proposed mask), xap_driver_decision (winning plan: ship forecast frames, kills, pickups), xap_driver_selection (targets/blockers/pickups/sources), xap_driver_scene (predicted bodies), xap_session_map, xap_session_shop. xap_session_clone gives checkpoints. ABI is at 91 and the mission-state structs change often (git history of xenon_autopilot_mission.h).

A standalone WASM target already exists (xenon_autopilot_wasm in src/autopilot/CMakeLists.txt, createXenonAutopilot) with a hand-maintained 150-name export list. It is only configured when src/autopilot itself is the CMake source directory; there is no preset for it.

Game (desktop and web). The C autopilot is linked into both builds (src/xenonAutopilot.c) and is started from the ESC options menu's AUTOPILOT row; src/xenonOptionsMenu.c has no platform #ifdefs, so the menu is the same on desktop and web. xenonControl.c can write an .x2events stream (s_eventLog, XENON_MSG_START_LOG) but only on request over the TCP control socket (desktop automation: record_run.py, campaign tools), which does not exist in a browser. BuildFrameBatch is built before XenonAutopilot_Frame runs in XenonControl_OnRealFrameComplete. AVI recording (START_AVI: authentic + sprite AVI with .vbl sidecars) exists on desktop only.

Renderer. src/spriteStreamFrame.c turns the frame's draw entries into SpriteRenderQuads (wall-tile occluders re-appended after terrain-masked sprites, separate status-bar pass). web/src/views/spriteStreamView.ts renders them with three.js: it owns a THREE.WebGLRenderer, a scene, a raw-shader material and the shared + level atlas textures, and reads the atlases from Emscripten's preloaded FS.

Site. site/build_site.py generates static pages from content/*.html, templates/base.html and assets/site.css (dark palette, Orbitron/Press Start 2P, amber/orange accents) and copies the game to dist/play/.

Measured recording sizes (two Level-2 native runs, run_logs/*.validation):

run frames VBL per game frame raw bytes/frame gzip bytes/frame gzip ratio
l2-newtactic-r30 803 4.0 27 708 3 154 8.8×
l2-full-r23 1 172 4.0 14 723 1 683 8.7×

At 50 Hz and 4 VBL per game frame that is 12.5 game frames/s: 11–21 MB/min raw, 1.3–2.4 MB/min gzipped, so a 30-minute run is roughly 40–70 MB compressed.

3. Not ported: obsolete with the C autopilot

Everything below exists in replay_ui.py/autoplay_ui.py but is either Python-autopilot-only or never populated when the C session drives. The "Replacement" column is what the web version shows instead; §5 describes the C-side data behind each replacement.

3.1 Python planner selection and analysis modes

Item (Python) Why it goes Replacement
--planner legacy/maneuver, create_controller, ReactiveController/ManeuverController C session is the only autopilot none
--prediction-reuse and the "Replay prediction: Recorded / Optimized / Reference" combobox Python prediction modes none
--architecture feature overrides, --kernel-backend python/c Python kernel ablations none
planner_description() caption (recorded planner, navigation/kernel/session backends) describes Python backends "Configuration" card (§7.3)
recorded_configuration() sniffing of .planner.json, manifest.json, JSONL prefix Python metadata in-stream RECORDING_INFO/AUTOPILOT_EVENT records (§6.1)
JSONL decision sidecar: TraceDecisionSidecar, _trace_recording_info, navigation_trace, "RECORDED MANEUVER", "RECORDED tactic/action/next" Python telemetry; the native runs in run_logs have no such sidecar recorded live C status per frame (AUTOPILOT_STATUS, §6.1)
Recorded boss fast path: _RECORDED_BOSS_TACTICS, _recorded_boss_decision, _apply_recorded_boss_decision Level-4 Python boss decisions none
Autoplay-trace JSONL input: trace_record_to_frame, normalize_record old Python trace format only .x2events (optionally gzipped)
Python world reconstruction: WorldState.update, PersistentWorldMap.update, override application, level_transition_active/shop_visit_level mirroring, assign_fire_schedule/FirePulse the C session owns phase, shop, transition and fire none
Capture-only fallback (_start_capture_only, _step_capture_only) built on the Python world Python world on session failure: keep rendering recorded draws, show an error banner, no tracks
_copy_state deep-copy checkpoint machinery Python objects xap_session_clone checkpoints

3.2 Overlays whose Python data source is gone (C equivalents are visualized)

Python overlay Why the Python version goes C equivalent shown instead
4 px navigation grid (world_map.navigation_visualization(), NAV_RESOLUTION/NAV_COLUMNS, green/red cells, "nav KiB", nav=(x,y) in the map readout) only Python route searches (find_navigation_connector, navigation_flow_field, find_navigation_path_*) fill it. C builds a 4 px free grid only transiently in prepare_graph() (xenon_autopilot_environment.c), turns it into span graphs and frees it spans overlay (§5.5)
Swarm/formation dashed prediction via controller.predicted_hazard_step, lane-crossing markers Python predictor C per-object predicted paths with their motion model (§5.2), optional lane-crossing markers against the mission lane
8-frame motion arrows and "PREDICTOR fittedWorldV" from _predicted_displacement Python predictor C track velocity and fitted world velocity arrows (§5.1)
Formation-path sources from Python caches (formation_trajectories, _scripted_displacement_cache, shared_prediction_scene, _formation_successors) Python caches C predictions + candidate scene bodies (§5.2–5.3)
SWARM FIRE LANE (swarm_anchor_x), projected ship stop (swarm_projected_stop_x) legacy controller fields mission lane and firing station (§5.4)
Level-5 route plan (level5_route_plan: corridor, active route, "L5 ACTIVE" waypoint, contract, hull, invalid reason) Python-only type; C has no separate Level-5 route the shared C mission route plus the Level-5 mission state and geometry (§5.4)
Target-priority labels P{n} and TARGET PRIORITY list (target_priorities) Python ranking role badges from the C selection and mission state (§5.6)
Python ship navigation hull (_player_navigation_hull_at) + navigation_reserve box Python geometry captured player collision bounds + configured decision.reserve; spans already include the hull (§5.5)

3.3 Summary lines that are empty or meaningless in C mode

Item Replacement
maneuver_summary_lines: ~40 Python kernel counters (kernel calls/bytes, native_body_*, linear_*, formation/script frames, nav grid builds, mission cache, shared prediction, terrain observations, prediction mode, architecture) the values the C path actually fills (native_session_host.publish + native_driver.inspect): stage/purpose, verified, horizon, rollouts, transitions, route searches, contact reason/frame/kind/key, predicted kills, pickups, candidate calls/frames, combat steps, step time
L2 STRATEGY (level2_strategy_phase, homing gate, cleared gates) C XapLevel2MissionState card
SWARM MODEL and SWARM collision/commit lines none
HAZARDS "scored/chosenRisk" C clearance
L5 TANK CAM, L5 ROUTE/CONTRACT/PROGRESS, L5 ALTERNATE C XapLevel5MissionState card + geometry
RETREAT, PRESSURE relief none
UNREACHABLE TARGETS (target_rejections) none
MISSION PROPOSALS / CANDIDATE SCORES (scores, risk, hit, exposure, blocked) and the ACTION score proposed vs applied mask and the arbitration list (§5.7)
FIRE LANES (firing_directions) weapon mask + selected mount
AIM y and flight (shot_flight_frames) goal rectangle, band, lane

3.4 Not in the web replay

Item Reason
Recorded AVI panes (original/sprite AVI, VBL sync, AUTHENTIC_AVI_VBL_DELAY) the web replay renders the real sprite stream. Desktop recordings still produce AVIs for the Python tools (§6.3)
--level2-boss-trace text panel (XAP_LEVEL2_TRACE_FILE) every field it prints is shown structured (Decision card with proposed/applied mask and arbitration, Level mission card, driver wall-escape state)
terrain.json RAM snapshot sidecar replaced by the in-stream TILE_MAP record; old runs are converted (§6.5)
"Pause stepping" (live AutoplayWindow only) live mode is not part of the replay

4. Architecture

 desktop Hatari / web hatari.wasm (C autopilot resident)          /replay/  (site page)
 ┌─────────────────────────────────┐                ┌──────────── main thread ───────────────┐
 │ xenonRecorder (shared C):        │  web: EM_JS    │ Library (IndexedDB) · transport · cards│
 │  frame batches + TILE_MAP +      ├─ chunks ──────►│ ReplaySceneView (three.js):            │
 │  INFO + AUTOPILOT_EVENT/STATUS   │  gzip→IndexedDB│   SpriteStreamLayer + overlay scene +  │
 │ desktop: file (+ optional AVIs)  │                │   CSS2D labels · level map (three.js)  │
 └─────────────────────────────────┘                └───────────────▲────────────────────────┘
                                                                     │ quads (transfer), palette,
                                                                     │ inspection JSON, geometry arrays
                                                     ┌───────────────┴──────── Web Worker ─────┐
                                                     │ xenon_replay.wasm                       │
                                                     │  xrp decoder → XapSessionInput          │
                                                     │  xap_session_step / clone (checkpoints) │
                                                     │  xrp inspector → JSON + typed arrays    │
                                                     │  SpriteStreamFrame builder → quads      │
                                                     └─────────────────────────────────────────┘

4.1 Replay core in C (xrp, new, e.g. src/autopilot/replay/)

One small C layer, compiled into a dedicated WASM module and also into the native DLL so Python can use it for parity tests.

typedef struct XrpReplay XrpReplay;
XrpReplay *xrp_create(const XapSessionConfig *config);  /* NULL = XenonAutopilot_Create defaults */
void       xrp_free(XrpReplay *r);
/* One EVENT_BATCH payload (no 16-byte header) and its protocol version. */
int32_t    xrp_step(XrpReplay *r, const uint8_t *payload, int32_t size, int32_t version);
int32_t    xrp_tile_map(XrpReplay *r, int32_t level, int64_t frame, uint32_t base, const uint8_t *words, int32_t size);
int32_t    xrp_peek(const uint8_t *payload, int32_t size, int32_t version, XrpPeek *out); /* frame, vbl, level, shield, lives, hits, native flag */
int32_t    xrp_checkpoint(XrpReplay *r);  int32_t xrp_restore(XrpReplay *r, int32_t id);  void xrp_drop(XrpReplay *r, int32_t id);
const uint8_t *xrp_quads(const XrpReplay *r, int32_t *count, int32_t *stride, const uint32_t **palette);
const char    *xrp_inspect_json(XrpReplay *r, uint32_t sections, int32_t *length);
int32_t        xrp_inspect_points(XrpReplay *r, uint32_t layer, const float **xy, const int32_t **offsets, int32_t *count);
  • Decoder. Parses the big-endian batch straight into XapSessionInput (the same structs xenonAutopilot.c fills live) plus the display-only facts (draw kind/flags/palette index, palette, player hits, game-memory trailers). Supports the protocol versions that carry the trailers the session needs; older files are rejected with a clear message instead of guessed defaults. Doing this in C once avoids a second TypeScript port of decode_frame + pack().
  • Session lifecycle mirrors the live host exactly (xenonAutopilot.c, xenonControl.c): create when the stream says the autopilot started, apply the warm-up skip, do not step outside levels 1–5, pass the recorded tile map on the first stepped frame and on each level change, recreate on run_id change (snapshot restore → XenonAutopilot_Reset), destroy on stop. The recorded AUTOPILOT_EVENT/AUTOPILOT_STATUS records (§6.1) make this deterministic instead of inferred.
  • Checkpoints. xap_session_clone every N frames plus a small LRU of recent frames (Python keeps 120). N is chosen after measuring clone size/time (§8).
  • Quads. spriteStreamFrame.c is refactored so the builder takes an entry iterator rather than only the DrawCommandStream ring; the emulator and the replay then share one implementation of the ordering/occluder rules (tuned against the real game) instead of a TypeScript re-implementation.
  • Inspector returns JSON plus typed arrays, not struct mirrors. Mission-state structs change often and several ints are overloaded (e.g. final_boss_center_mode holds 1 = centre mode and 10–12 = front-attack stages; boss_refuge_reached holds refuge codes 2–5). A TypeScript mirror would drift (the earlier SpriteRenderQuad stride bug in spriteQuad.ts is the precedent). A C function compiled with the structs emits named fields and decodes enums/overloaded codes to names. Per-level mission states use X-macro field lists so the struct and its JSON writer come from one list. Hazard-kind, model, tactic and purpose names also come from C. JSON is produced only for displayed frames, never during bulk seeking. Point-heavy layers (predicted paths, scene bodies, plan frames, spans, map cells) go out as Float32Array/Int32Array via xrp_inspect_points.
  • Sections (bit flags; the UI requests only what is visible): status, decision (+ arbitration), mission, level state, geometry, shop, config, tracks (+ decoded fields, model, roles), walls, predictions, scene bodies, plan, spans, map, raw bytes of the selected object.

4.2 Worker

The worker owns the WASM module, the recording chunk cache and the checkpoints, so long seeks never freeze the page. Messages: open(recordingId) → index (frame count, game-frame → record map, markers); seek(index) → progress… → frame (quads as a transferred buffer, palette, inspection JSON, typed arrays); setSections(mask); select(identity) (adds raw bytes/paths for that object); lookahead(t) (§5.3). Playback asks for the next frame and the worker prefetches.

4.3 Main thread

Library, transport/timeline, ReplaySceneView, the level map, and the right-hand cards (§7). No emulator is loaded on this page.

4.4 Rendering: the same three.js code as the game

The replay uses the game's three.js sprite renderer, not a separate implementation, and draws the overlay with three.js as well (one renderer, one canvas, one coordinate system):

  • Split SpriteStreamView into SpriteStreamLayer (geometry, raw-shader material, atlas textures, update(quadBytes, count, stride, palette)) and the existing thin view that owns a THREE.WebGLRenderer and canvas for the play page. Play-page behaviour stays unchanged (same stride guard, same heap adapter).
  • ReplaySceneView owns one WebGLRenderer: pass 1 renders the SpriteStreamLayer into the 320×200 viewport (setViewport/setScissor); pass 2 renders an overlay THREE.Scene with an OrthographicCamera in logical Atari coordinates including the margins (Python: 48 px left/right, 32 px top, 48 px bottom), so off-screen objects and look-ahead stay visible.
  • Overlay primitives: LineSegments/LineDashedMaterial for boxes, paths and arrows, Line2 (three's fat lines) where width matters, transparent meshes for zones/goal rectangles, instanced quads for spans and tick marks. Text labels via CSS2DRenderer (ships with three), so labels stay crisp and cheap to toggle. Hit-testing is done in logical coordinates in TypeScript.
  • The level map pane is a second ReplaySceneView-style view (instanced tile quads + the same overlay primitives), so route/goal/spans code is shared between scene and map.
  • Atlas loading gets an AtlasSource interface: Emscripten FS for /play/, fetch() for /replay/. The replay page must not download the 290 MB hatari.data; build_site.py copies the low-resolution atlas banks (spriteatlas_shared, spriteatlas_level_01..05, about 2 MB each) to dist/replay/atlases/. HD banks (30–43 MB each) are an opt-in toggle, loaded on demand. Frame interpolation stays off (real frames only).

4.5 Build and site integration

  • New CMake preset wasm-replay producing xenon_replay.js/.wasm (MODULARIZE, EXPORT_NAME=createXenonReplay, ENVIRONMENT=worker,node, ALLOW_MEMORY_GROWTH). Its export list is only xrp_* plus malloc/free, removing the drift risk of the 150-name list in the existing xenon_autopilot_wasm target.
  • web/vite.config.ts gets a second entry (replay app + worker) sharing spriteStreamLayer.ts, atlas.ts, spriteQuad.ts, palette.ts.
  • site/build_site.py: new --replay-build input; content/replay.html (an "app" page variant of base.html: full width, starfield disabled so it does not compete with WebGL); a card on the Autopilot hub and a link from the play page.
  • The build helper remains xenon_tools/hatari_dev.ps1; a new -Action for the replay module follows the existing pattern.

5. What the C autopilot uses, and how the replay shows it

Traced through xap_driver_step (xenon_autopilot_driver.c), xenon_autopilot_predictions.c, xenon_autopilot_driver_selection.c, xenon_autopilot_driver_spawns.c, the scene/planner headers and xenon_autopilot_environment.c. The per-frame pipeline is:

tracks + decoded fields ─► per-object motion model ─► predicted anchor paths (XapPredictions)
        │                                                    │
        ▼                                                    ▼
 selection (contact kinds, relevance culling,        anticipation (pending scroll spawns,
 targets / blockers / pickups / source flags)        off-screen entries, scheduled shots)
        └──────────────────────► candidate scene ◄───────────┘
          body paths · linear bodies · per-frame slices · pickups · shells · shots ·
          homing models · barriers · combat (player bullets, target health)
                                   │
 mission (route over span graph, goal, band, lane, weapon, level state) ─► proposed mask
                                   ▼
                     planner rollouts ─► winning plan (ship forecast, kills, pickups)
                                   ▼
                     driver arbitration (mission/level overrides, escapes) ─► applied mask

5.1 Velocity

XapTrack carries screen vx/vy, ax/ay, world world_vx/world_vy and fitted world velocity (fitted_vx/fitted_vy, xap_fit_world_velocity over up to 12 history samples). The fitted value is what the linear (POLICY) model uses, so the default arrow is fitted world velocity × N frames (N adjustable, default 8, same as Python). Raw screen velocity is an optional second arrow.

5.2 Per-object predicted paths (XapPredictions)

Each object gets a motion model from xap_world_models(); xap_predictions_path produces a world-space anchor path for horizon + 1 frames:

Model What it predicts
POLICY straight line from fitted world velocity; directional projectiles use the game's exact direction table × speed scale
SCRIPT_ROOT the captured motion-script interpreter (sine/script opcodes)
ORDINARY_FOLLOWER formation follower replaying its leader's path with a delay
FOLLOW_TARGET, COMPOSITE_ATTACHMENT linked to a root object's path
COMPOSITE_CHAIN Level-3 composite chains (all members from one root)
EYE_LINK, BOSS_LINK articulated boss eyes and chains (xap_world_prepare_boss_paths)
SHELL_OSCILLATOR boss shells
UNSELECTED/unresolved no path (shown as "no model", which is itself useful)

Shown as a path per object, coloured by model, tick marks every 8 frames labelled +8, +16, … Paths the planner actually used are solid; paths requested only for inspection (objects the selection culled) are dashed and labelled "not used by planner". The model name appears in the object table and the selected-object card. Optional: markers where a path crosses the mission lane (the C counterpart of Python's lane-crossing markers).

5.3 The candidate scene: what the planner checks the ship against (xap_driver_scene)

Scene part Shown as
Body paths (XapHazardPath + points + hull offsets; active_from_frame for anticipated shots) hull outline along the path; anticipated ones dashed from their activation frame
Linear bodies (XapLinearBody, Q32 velocity) hull + velocity ray
Per-frame slices (XapBodySlice, e.g. linked chains) ghost rectangles
Anticipated waves (pending scroll spawns → delayed paths) and off-screen entries (XapCandidateEntry: screen rect + trigger scroll) dashed paths starting off screen; entry rectangle with "spawns at scroll N"
Scheduled enemy shots (XapCandidateShot: spawn point, frame, 8 directions) and shell releases (3 velocities) fan of short rays at the spawn point, labelled with the spawn frame
Homing missiles (simulated per candidate from XapHomingState + tables) ghost boxes for the winning plan's simulation
Barriers (Level 5), pickups (per-frame rectangles) markers / ghost rectangles
Combat (player bullets, target health → predicted kills) kill markers on targets the winning plan destroys

Look-ahead scrubber. A slider t = 0…horizon draws every scene body and the planned ship hull at frame t at the same time. This shows directly why a move was judged safe or unsafe, which static paths cannot.

Winning plan. Ship forecast frames (XapManeuverFrame: position, hull, clearance, contact, input mask, wall-free flag) drawn as the ship's planned path, coloured by clearance; the contact that ended the plan is marked and the responsible body highlighted (candidate reason + contact source/index → object, as native_driver.inspect already resolves it). Verified = green, unverified = red (as in Python's navigation overlay).

5.4 Mission route, goals and level geometry

There is no separate Level-5 route type in C. Every level plans through the shared mission code (xap_mission_route_to, xap_mission_route_goal, span connectors), which fills XapMissionState.route/route_index/destination/route_frame, the goal rectangle, band and lane in XapMissionResult.decision, and the level's own state (Level 5: station_x, refuge_x, train_x, laser group, tank phase …). Shown for every level:

  • route polyline with the current index and next waypoint, destination (firing station), goal rectangle, screen band, preferred lane, mount offset (scene and map);
  • level geometry export (new C, per level): authored zones and dynamic anchors with names. Examples from the current code: Level 2 world-Y zones (left corridor 3360–4208, eye arena 2500–3100, eye approach 3100–3360, lower swarm 376–960, maze swarm 1088–1200), eye refuges and worm-wait lane, final-boss pockets (upper left/right), centre X 156, front station X 192 and cross row Y 385, dodge rows 40/184; Level 5 station/refuge/train positions and laser-gate crossing; Level 3 intercept X; Level 4 pickup position and refuge. Zones show as labelled bands on the map; anchors as labelled markers in the scene.

5.5 Spans overlay (optional layer, replaces the 4 px grid)

C navigates over span graphs: rows of free runs [left, right] in 4 px cells (80 columns), origin = first world grid row. Two graphs exist per environment: physical (reserve 0) and travel (dilated by the configured reserve; Level 3 uses the physical graph for both). Spans are configuration space: where the ship's anchor may be, with the hull already accounted for. New read-only export xap_span_graph_spans(graph, first_row, row_count, out, capacity) returns (row, left, right) for a window around the viewport. Shown as translucent horizontal bars on the map (primary) and optionally in the scene; spans the current route passes through are highlighted; a toggle selects physical/travel.

5.6 Role badges

Badges on scene labels and in the object table/selected card:

  • mission target (always highlighted), damage target / blocker / pickup / scene source with its flags (homing L2/L5, independent, damage supported, scheduled shot) from xap_driver_selection;
  • level-state references: L1 attack/blocked target, L2 powerup/eye/worm target and hazard-escape object, L3 intercept target, L4 pickup/blocked target, L5 station target;
  • "killed by plan", "collected by plan", "plan contact".

5.7 Decision arbitration

The applied mask is not the planner's mask: xap_driver_step passes it through many stages (planner prefix, mission-specific overrides for Level 2/4 pickups and bosses, Level-4 edge/exit recovery and escapes, Level-2 route brake, corridor guard, hazard escape, pocket hold, wall escape). The Level-2 trace file already prints proposed vs applied because this matters. Plan: a small bounded log in the driver (stage id, mask after) recorded only when a stage changes the mask, exported with the decision. The Decision card lists it ("planner 0x88 → L2 hazard escape 0x81 → wall escape 0x84"). This is the only autopilot-behaviour-neutral instrumentation the plan adds to the driver.

6. Recording (desktop and web) and accessing recordings

6.1 What gets recorded

The same .x2events message stream the desktop writes today (16-byte header + payload), so files are interchangeable between desktop, web and the Python tools. Recording._build_index already skips non-EVENT_BATCH messages, so the new record types are backward compatible:

Type Content Why
EVENT_BATCH 0x4001 (existing) unchanged, plus a new draw flag bit 0x08 = terrainMasked without it the replay cannot re-append the wall occluders over the Level-2 worm. Uses a spare bit; layout unchanged
RECORDING_INFO (new) JSON: source (web/desktop), git commit + dirty flag, autopilot ABI, hires/highfps, start time, label, session mode (below) identifies what produced the file; drives the summary and the Configuration card
TILE_MAP (new) level, frame, base address, the level's tile map: 300 rows × 20 × 2 = 12 000 bytes see below
AUTOPILOT_EVENT (new) start / stop / reset, with the XapSessionConfig and whether the warm-up skip applied makes session lifecycle deterministic in replay
AUTOPILOT_STATUS (new) the live XapSessionStatus after each live step (~150 bytes, tagged with ABI) ground truth for "replay == live" per frame (controls, tactic, purpose, verified, contact)

Why TILE_MAP and not RAM. The session uses RAM for exactly one thing: xap_map_session_sync_ram → xap_map_state_sync_ram reads the tile-map words of the level's destructible groups at tile_map + (row*20 + column)*2 (Level 2 has 1 group, Level 3 has 9, Level 5 has 22; Levels 1 and 4 have none). The live host passes the RAM on the first stepped frame and on every level change. Without the same words the replay's map can differ from the live map (gates already opened off screen) and so can its decisions. Recording the whole 1 MB RAM is unnecessary; the 12 000-byte tile map (addresses from xap_level_asset: L1 355666, L2 362302, L3 395416, L4 400600, L5 381466) is written whenever the live autopilot is handed RAM, and once at recording start. The replay passes it through a new xap_map_session_sync_tiles(map, base, words, size) wrapper (same logic, no fake 1 MB buffer). Python can use the same record instead of terrain.json.

6.2 Session start rule (decided)

The live autopilot session is never restarted because recording starts. Normally recording and the autopilot start together (starting the autopilot while RECORD is armed, or starting both from the menu, begins recording on the same frame), which gives an exact replay. If recording starts while the autopilot is already running, the replay's session starts cold while the live session has history (track ages, fitted velocities, latched mission state such as the Level-2 final_boss_python_parity latch or the final_boss_center_mode stage), so decisions can differ for a while or, for latches, permanently.

RECORDING_INFO.session_mode records which case applies: exact or joined_mid_session (frame N). It is shown as a badge in the library list, in the Configuration and Diagnostics cards, and as a timeline marker; the per-frame AUTOPILOT_STATUS shows exactly where replay and live differ.

6.3 Recorder in both builds

A shared C module xenonRecorder.c (extracted from the s_eventLog code in xenonControl.c) with two sinks:

  • Desktop: file sink. Unchanged behaviour for the socket START_LOG (automation keeps working). New interactive start: a RECORD row in the ESC options menu (shared menu code). Output goes to the fixed folder recordings/ relative to Hatari's working directory, one file per recording (<timestamp>.x2events). Optional AVIs, desktop only: an "AVI: off / original / sprite / both" row starts the existing authentic and sprite AVI writers alongside, named <name>-original.avi / <name>-sprite.avi with .vbl sidecars, which replay_ui.py already discovers.
  • Web: chunk sink. An EM_JS callback hands each batch to web/src/recorder.ts (bundled into the existing bridge), which copies it synchronously (the pointer is only valid during the call), groups ~256 frames (≈1 MB compressed), compresses with CompressionStream('gzip') off the emulator slice, and stores independently decompressible chunks in IndexedDB (recordings metadata store + chunks store keyed [recordingId, chunkNo] with first/last game frame). Chunking keeps random access cheap and limits loss on a tab crash to the unflushed tail. No AVI option on web.
  • Web UI on /play/: the same menu RECORD row, plus a record button next to Fullscreen showing a red dot, elapsed time and size; optional ?record=1. It requests navigator.storage.persist() at the first recording and shows the quota estimate. On stop: toast "Saved — open in replay".

localStorage is not usable: it holds ~5 MB of strings per origin, about three minutes of gzipped recording. IndexedDB is shared with /replay/ (same origin), so recordings appear there directly.

6.4 Library and upload (/replay/)

  • List: name, date, source (web/desktop), session mode badge, levels covered, frames, duration, size, autopilot on/off; rename, delete, download (concatenated gzip members = one valid .x2events.gz).
  • Upload = import a local .x2events or .x2events.gz (desktop or web): stream it through DecompressionStream if needed, re-chunk it into IndexedDB, build the index. Nothing leaves the machine. AVI files are not imported.
  • ?src=<url> opens a recording from a URL. Curated samples in site/dist/replay/samples/ come later, once the autopilot finishes the game.
  • Deep links: /replay/?rec=<id>&frame=<gameFrame>.

6.5 Python interoperability

  • Python Recording learns to open gzip (detect 1f 8b, decompress to a temp file), so web recordings open in replay_ui.py, and uses TILE_MAP records when present.
  • A small xenon_tools/pack_recording.py converts an old run: extracts the tile-map words from terrain.json RAM snapshots (or start.ram) into TILE_MAP records and gzips, producing one file to import.

7. UI design

7.1 Layout (desktop, ≥1400 px; stacks on narrow screens)

┌ site topnav ───────────────────────────────────────────────────────────────────────┐
│ [Library ▾ l2-r30 · L2 · 802 frames · exact]  ⏮ ◀ ▶ ⏭  ▷ Play  speed [10]  frame [ ] │
│ timeline ▁▁▁▁│▁▁▁▁▁✕▁▁▁$▁▁▁▁▁•▁▁▁▲▁▁│▁▁▁  (│ level  ✕ death  $ shop  • hit  ▲ diverge) │
│ game frame 14912 · record 381/803 · L2 · PLAYING · input ✓ replay 0x88 = live 0x88 │
├──────────────────────────────────────────────┬────────────────────────────────────┤
│ Filters: [preset ▾] [layers…] [categories…]  │ AUTOPILOT STATE (grouped cards)    │
│ ┌──────────────────────────────────────────┐ │  Decision · Target & weapon        │
│ │ three.js: sprite stream + overlay scene  │ │  Route · Level 2 mission           │
│ │ 320×200 ×3 + margins · look-ahead [t]    │ │  Ship · Input & camera · …         │
│ └──────────────────────────────────────────┘ ├────────────────────────────────────┤
│ Level map (three.js, full level)             │ SELECTED #1234 swarm_enemy [sprite]│
│ ┌──────────────────────────────────────────┐ │  position · motion · model · fields│
│ │ walls · zones · spans · viewport · route │ ├────────────────────────────────────┤
│ │ goal · click = coordinates               │ │ OBJECTS [search][kind chips][zone] │
│ └──────────────────────────────────────────┘ │  sortable table                    │
└──────────────────────────────────────────────┴────────────────────────────────────┘

7.2 Top-left scene: layers and filters

Base: the real sprite stream with a "sprite dimming" slider (Python's brightness slider, applied to the real renderer), so overlays stay readable.

Layers (each a toggle, grouped): * Objects: collision boxes · true collision boxes ($36–$3C) · labels (off / selected + target / all) · role badges. * Motion (§5.1–5.3): velocity arrows · predicted paths (off / selected / used by planner / all) · scene bodies · anticipated spawns and entries · scheduled shots and shell releases · homing ghosts · look-ahead scrubber. * Plan and mission (§5.3–5.4): winning plan path · route · goal rectangle · band · lane · firing station · level geometry · lane-crossing markers. * Terrain: terrain tiles · spans (physical/travel) (§5.5). * Other: recorded hits · screen grid · off-screen margins.

A master "diagnostic overlay" toggle and "isolate selection" remain.

Object categories (names come from C's XapHazardKind; unmapped kinds fall into "Other" so a new kind never disappears). Per category: visible · label · path, as a small matrix.

Category Kinds Default
Player PLAYER, PLAYER_ATTACHMENT box + label
Player shots PLAYER_PROJECTILE hidden
Enemy shots DIRECTIONAL_PROJECTILE, HOSTILE_PROJECTILE, ANTICIPATED_PROJECTILE, LEVEL5_TANK_PROJECTILE box, no label
Homing LEVEL2_HOMING_ENEMY, LEVEL2_HOMING_EMITTER, LEVEL5_CURVED_HOMING_MISSILE box + path
Enemies & formations ENEMY, SWARM_ENEMY, FORMATION_LEADER, FORMATION_FOLLOWER, WALL_SHOOTER box
Gates & traps LASER_GATE, LEVEL4_LASER_JAW, LEVEL4_DIAGONAL_SWEEPER, WORM_HOLE box
Bosses BOSS_, LEVEL2_BOSS_EYE, LEVEL2_FINAL_BOSS(CONTROLLER), LEVEL3_COMPOSITE, LEVEL4_STAGE2_BOSS_, LEVEL4_BOSS_, LEVEL5_TANK_* box + label
Pickups PICKUP, PICKUP_CARRIER box + label
Inert EFFECT, DESTROYING, DESTROYING_HOSTILE, CONTROLLER, UNKNOWN hidden

The mission target, the plan-contact body and the selected object are always drawn, whatever the filters say.

Presets: Clean (sprites only) · Decision (target, route, goal, plan, selected) · Hazards (enemy shots, homing, bosses, scene bodies) · Navigation (route, spans, geometry) · Everything. Filter settings persist per browser (localStorage, wrapped in try/catch).

Clicking the scene selects the smallest box under the cursor (as in Python); hovering shows a tooltip with kind/identity/position/model, so labels can stay off by default.

7.3 Autopilot state panel: grouping

Collapsible cards in this order; open/closed state is remembered. Values that changed since the previous frame are highlighted briefly; cards that do not apply (Shop outside the shop, other levels' mission state) are hidden.

Group Contents Source
Header bar (always visible, not a card) game frame, record i/N, VBL, level, session phase (inactive/playing/respawn/shop/transition/complete), run id, input match badge (replay controls vs live AUTOPILOT_STATUS, else vs recorded input), session-mode badge status, recording
Decision (the selected option/tactic) tactic (level + XapMissionTactic name), purpose (budget_exhausted … no_verified_escape), action (arrows + fire), fire allowed, proposed → applied mask with the arbitration list (§5.7), verified, horizon (and mission horizon limit), rollouts, transitions, route searches, clearance, contact (reason/kind/frame/object), step time status.driving, mission result, decision view, driver log
Target & weapon target (id, kind, position, health), weapon/mount (forward/laser/cannon/side-left/right), mount offset, goal x/y range, band, lane, horizontal-first, predicted kills and pickups of the plan, counts of considered targets/blockers/pickups/sources mission result, plan, selection
Route & navigation route index/count, next waypoint, destination (firing station), route frame, map revision, corridor index/complete, which span graph (physical/travel) mission state, environment
Level mission (current level only, enums decoded to names) L1: attack/blocked target and deadlines · L2 (as committed): phase, waypoint, gate, completed gates, wave/boss active, entry phase (retreat/cross/advance/fight), refuge reached (left upper/lower/wall, worm wait), crossing (generic, left→right, right→left), refuge hold frames, corridor patrol direction, eye plan + transfer stage, powerup phase + missing frames, powerup/eye/worm targets, worm missing frames, final-boss pocket (upper left/right) + frames + threat seen + clear frames, centre mode / front stage (descend, cross, fire), final-boss latch (final_boss_python_parity), hazard escape object/mask/frames · L3: gate, egress waypoint, intercept target/X · L4: bypass, refuge, crossing, retracted frames, progress, blocked target, pickup target/position/escape · L5: group, diagonal crossed, egress, tank phase, station target/X, refuge X, train active/X, laser group/crossed group · driver wall recovery: replay state, blocked mask, contact pending/point, escape mask/frames XapLevelNMissionState, driver
Ship shield, lives, score (+ deltas, life lost), player present, speed level, weapon mask, fire cooldown/reload/rate, damage bonus, attachments (7 slots: code, item, power), auxiliary count, Nashwan power, protection, hits this frame (source kind, object, proc, damage) status, game memory, hits
Input & camera recorded input mask, consumed mask, steering, scroll intent, wall-collision replay countdown, scroll y/velocity, applied delta, forward/backward limits, backtrack window, limits valid, interstitial waiting-for-fire, fire phase, transition phase game memory, status
Shop (shop phase only) captured menu (mode, cursor, cash, level, page, availability) and C shop state (recipe, initial/projected cash, phase, target item, pending item/cash, purchases, last transaction, quote) game memory, xap_session_shop
World (collapsed) object count, wall count, per-category and per-model counts, map revision, solid/destroyed cells status, tracks, map
Configuration (collapsed) recording source, session mode, build commit, ABI (recorded vs replay module; warning on mismatch), session config (reserve, horizon, rollouts, transitions, refresh, tactical/attacks/combat/defer, homing horizon, fire pulse) RECORDING_INFO, AUTOPILOT_EVENT
Diagnostics (banner when non-empty) session rejected a frame, first divergence frame, joined-mid-session notice, decoder warnings replay core

Where each current Python line goes:

Python line New place
AUTOPILOT, REC PLANNER, window caption Configuration (Python planner fields dropped)
MANEUVER rows Decision (surviving counters) / Target & weapon (plan outcome); rest dropped (§3.3)
TACTIC, CONTROL, ACTION Decision (score dropped)
APPLIED Input & camera + header badge
TARGET, AIM Target & weapon (y/flight dropped)
FIRE LANES dropped → weapon mask in Ship
PLAYER Ship
SCROLL, SHIP INPUT, CAMERA Input & camera
FIRE STATE, SHIP SPEED, ATTACHMENTS, HIT SOURCE Ship (hits also as timeline markers)
MAP, OBJECTS World
HAZARDS Decision (clearance)
SHOP, SHOP PLAN, SHOP CHECK Shop
MOVE GOAL Route & navigation (and the goal-rectangle layer)
L2 STRATEGY, L5 * Level mission (from C state) + level geometry layer
SWARM *, RETREAT, PRESSURE, TARGET PRIORITY, UNREACHABLE TARGETS, MISSION PROPOSALS / CANDIDATE SCORES, RECORDED * dropped (§3)
ANALYSIS ERROR Diagnostics
Level-2 boss trace panel Decision (arbitration) + Level mission

7.4 Selected object

A card between the state panel and the object table:

  • Title row: #identity kind with previous kind if it changed, role badges, and the sprite preview on the right (atlas crop drawn nearest-neighbour up to 4×; low/HD bank follows the scene). The Level-5 horizontal-laser special case (composite of the live wall-tile draws) is kept because it is about the game, not the Python autopilot.
  • Position/bounds (screen and world), interaction box, true collision box.
  • Motion: v, a, world v, fitted v, age, history samples; motion model, predicted path length, "used by planner" yes/no; in scene as path/linear/slice.
  • Decoded fields: only the XapObjectFields present for this object (health, fire phase, boss timers, composite links, …) with their names.
  • Procs, sprite, status/lifecycle, drawn, zone; captured script (cursor/phase) when present.
  • Raw object bytes as hex with offsets (replaces Python's fixed $1A/$1C/$5E/$5F line).
  • Buttons: follow (keep selected across frames), show path only, copy as JSON.

7.5 Object table

  • Columns: id, category/kind, model, health, drawn, zone, screen x/y, world x/y, world velocity, age, update proc, sprite, roles. Click a header to sort (shift-click for a secondary key); default order: target, then roles, then category, then id (similar to Python's priority-first sort).
  • Filters: text search (id, kind, proc/sprite hex), category chips, zone (on-screen/above/below/left/right), "drawn only", "used by planner", and "same as scene filters" (on by default, so hiding player shots hides them here too).
  • Row hover highlights the object in the scene; click selects it; the selected row stays pinned.

7.6 Level map

The C map is built from the embedded level asset and updated from observed tile rows and the tile-map sync, so the page shows the whole level with destroyed groups marked. Layers: walls, destructible groups, level zones (§5.4), spans (§5.5), viewport, player, captured collision bounds + reserve, route with the current index, destination, goal rectangle. Click shows L tile=(c,r) world=(x,y) solid/free group=g span=[l,r] with a copy button.

7.7 Transport

Play/pause, frame step, speed (1–200 fps), timeline scrubber with markers (level change, life lost, shop visit, player hit, first divergence, joined-mid-session), go to game frame, keyboard: Space, ←/→, Home/End, Ctrl+G (same as Python), and [/] for the look-ahead scrubber. Seeking shows progress, as the Tk tool does.

7.8 Visual style

Reuse site/assets/site.css tokens (--panel, --line, --amber, --orange, --font-display for headings, --font-mono for values), the site top navigation and footer, .card-style panels, amber section titles. Data text is monospace at a dense size; category colours keep the Python palette (COLORS in autoplay_ui.py) so existing screenshots stay comparable.

8. Verification: measure before building on assumptions

  1. Decoder parity. Run the C decoder and Python pack() over every frame of several recordings (Level 1, Level 2 r30, a shop visit, Level 5) and compare the packed session inputs byte for byte. Harness: a ctypes test in xenon_tools/.
  2. Decision parity, native. Replay r30 through xrp natively and compare controls with what the live run applied. Establish the frame offset by measurement: the batch is built before the live step, so the controls computed at frame N should appear as the input of frame N+1, but this must be confirmed on data. Report the match rate and explain every mismatch (session start timing, tile-map sync) before continuing.
  3. Tile-map subset. On recordings with start.ram, sync the map once from the full RAM and once from the 12 000-byte tile map: identical map state and revision.
  4. WASM vs native. Same recording in node (WASM) and native: identical status sequence.
  5. Render parity. For sample frames, compare quads built from recorded draws with the live SpriteStreamFrame_Build output (a desktop debug dump while recording), and compare screenshots against the sprite AVI of the same run.
  6. Round trip. Record in the browser and on desktop, open each in the other and in replay_ui.py: same frame count and content; desktop AVIs still discovered by replay_ui.py.
  7. Arbitration log neutrality. With the log enabled, the live status sequence of a campaign segment is identical to a build without it.
  8. Performance. Step time per frame in the worker, xap_session_clone size and time, seek time with the chosen checkpoint interval, inspection cost with all motion layers on, memory for a 30-minute recording, IndexedDB write throughput while playing (the emulator must stay at full speed on the phone; see the ?perf=1 overlay).

9. Phases

Phase Deliverable Done when
0. Spikes parity measurements 8.1–8.3 using the existing Python replay + a native xrp prototype; clone cost match rate known and mismatches explained
1. Replay core (C) decoder, lifecycle, checkpoints, inspector JSON (X-macro level states, named enums), typed-array layers (predictions, scene, plan, spans), level geometry export, arbitration log, xap_span_graph_spans, xap_map_session_sync_tiles, refactored sprite-stream builder, wasm-replay preset 8.1–8.4 and 8.7 pass
2. Recording xenonRecorder.c with file and chunk sinks, new record types, terrainMasked bit, menu RECORD row (both builds), desktop recordings folder + optional AVIs, web recorder + IndexedDB + play-page button, gzip + TILE_MAP in Python Recording, pack_recording.py 8.6 passes; no measurable emulator slowdown
3. Replay shell /replay/ page in the site build, library (list/import/download/delete/quota), worker, transport, SpriteStreamLayer + ReplaySceneView with fetched atlases a recording plays back with correct sprites (8.5)
4. Overlay object layers, category filters, presets, selection, tooltips; motion layers (velocity, predicted paths, scene bodies, anticipation, plan, look-ahead scrubber); route/goal/geometry; spans player shots can be hidden; every §5 item is drawable
5. Right column grouped cards, selected-object card with sprite, sortable/filterable table every row of the §7.3 mapping table is shown or intentionally dropped
6. Map level map with zones, spans, route, goal, coordinates —
7. Polish timeline markers, deep links, HD atlas toggle, ?src= loader —
later curated sample recordings in site/dist/replay/samples/ after the autopilot finishes the game

replay_ui.py stays unchanged apart from gzip and TILE_MAP support; retiring it is a later decision.

10. Decisions

All decided (2026-09-24): * The live autopilot session is never restarted when recording starts; the recording summary marks exact vs joined_mid_session (frame N) (§6.2). * AVIs only for desktop recordings; recording works in both builds (§6.3). * Overlay and map use three.js, sharing the game's sprite renderer (§4.4). * Replay page placement: a card on the Autopilot hub and a link from /play/; no top-navigation entry. * HD atlases: off by default, loaded when requested. * No "what-if" analysis and no changing of session settings in the replay; the config always comes from the recording. When the replay module's ABI or commit differs from the recording's, the page says so. * Sample recordings: yes, but not in this work. They are added once the autopilot finishes the whole game (parallel work); ?src= loading is still implemented so adding them is only content. * Desktop recordings go to a fixed folder, recordings/ relative to Hatari's working directory; not configurable.

Implementation happens on branch web-replay, one commit (or more) per phase in §9; no phase is squashed into another.

11. Risks

  • ABI churn. Replay decisions equal the live run only when the replay module is built from the same autopilot source as the recording. RECORDING_INFO + ABI tags make a mismatch visible rather than silent.
  • Open-loop replay. After the first divergence, the recorded game no longer reacts to the replayed decisions; later decisions are "what this build would do given these observations". The divergence marker makes that boundary explicit.
  • Inspection cost of motion layers. Paths for all objects, scene bodies and the look-ahead scrubber are the heaviest exports; they are computed only for the displayed frame and only for enabled layers (8.8 measures it).
  • Storage. 40–70 MB per 30-minute run; the browser may evict non-persistent storage. Persist request, quota display and download/export mitigate this.
  • Seek cost. Unknown until 8.8; checkpoint interval is the lever.
  • Stale docs. AUTOPILOT_BROWSER_PORT_OPTIONS.MD still recommends a TypeScript port; the C route superseded it. This plan does not change that document.

12. Implementation log

Phase 0 + 1 (C replay core)

Built: src/autopilot/xenon_autopilot_replay.{h,c} (decoder, lifecycle, checkpoints, inspector), src/spriteStreamBuild.c (dependency-free quad builder shared by the emulator and the replay), src/replay/xenonReplayWasm.c (WASM glue), the xenon_replay_wasm target, read-only autopilot accessors (xap_driver_arbitration, xap_driver_predictions, xap_driver_wall_recovery, xap_span_graph_spans, xap_environment_spans, xap_level2_geometry), and the tools xenon_tools/replay_parity.py and xenon_tools/check_replay_wasm.mjs. hatari_dev.ps1 gained build-replay-wasm, build-web, -KernelBuildDirectory, -ShowBuildOutput, and prints compiler output when a build fails.

Measured: * The decoder reads every frame of the checked recordings; stepping the session costs 0.06-0.13 ms per frame natively, a checkpoint clone 0.02-0.04 ms. * WASM vs native (r30, 802 stepped frames): identical controls on every frame. WASM cost 0.18-0.21 ms per frame including quad building and status JSON. Full inspection of all sections: ~84 KiB JSON per frame. * Checkpoint -> restore -> re-feed reproduces the same controls exactly. * Parity against live runs cannot be proven with the existing recordings: every run in run_logs except l2-spider-own-r1 was recorded from a dirty tree, and that run did not rebuild Hatari (its binary hash matches neither the current exe nor a known tree). On it the replay matches the live controls exactly up to frame 14801, then diverges. Settled in Phase 2 with a fresh recording from a known build that carries AUTOPILOT_STATUS records. * The live capture sets XAP_PLAYER_HAS_WALL_REPLAY and the wall-collision replay countdown; the Python replay's native_memory.pack() does not, so the existing Python replay can differ from the live run in Level-2 wall recovery. The C decoder follows the live capture.

Deliberate deviations from sections 4-6: * Inspector output is JSON only (no typed arrays yet); measured size is acceptable for displayed frames. Typed arrays remain the fallback if the page shows it matters. * No X-macro conversion of the mission-state structs: they are under parallel development, so the inspector writes their fields by hand and decodes enums and overloaded codes (final_boss_center_mode, boss_refuge_reached, boss_crossing) to names. (Replaced on 2026-10-01 by writers generated from marked comments, which leave the structs as plain C; §12, "Explanations from C".) * Tile-map sync reuses the session's RAM input: the replay builds a sparse RAM image containing only the recorded 12 000-byte tile map, so the session ABI is unchanged. * The WASM module is built by a hatari_dev.ps1 action (standalone src/autopilot configure, as build-kernel does) instead of a CMake preset.

Phase 2 (recording in both builds)

Built: src/xenonRecorder.{h,c} (one .x2events writer for a file on desktop and the page's recorder on the web), the new records (RECORDING_INFO, TILE_MAP, AUTOPILOT_EVENT, AUTOPILOT_STATUS) written by src/xenonControl.c, the options menu rows RECORD (both builds) and RECORD AVI (desktop), XenonControl_StartRecording/StopRecording/IsRecording exported to the web page, web/src/recorder.ts + web/src/recordingStore.ts (IndexedDB library with gzip chunks), a Record button and ?record=1 on the play page, xenon_tools/pack_recording.py, and gzip plus TILE_MAP support in replay_ui.py's Recording. hatari_dev.ps1 gained build-wasm (-WasmBuildDirectory, uses the prebuilt web/dist). Desktop recordings from the menu go to recordings/ in Hatari's working directory (git-ignored).

Measured: * Fresh desktop run from a known build (validate_campaign.py native-run web-replay-parity-l2-r1, 1629 batches, Level 2): replay controls equal the live controls on 1611/1611 stepped frames and every recorded AUTOPILOT_STATUS equals the replayed status field by field; the WASM module is identical to the native DLL on all frames; checkpoint restore is exact. This settles the Phase 0+1 parity question. * Web build (/play/?record=1, autopilot switched on while recording): 968 frames, 10.2 MB raw, 1.15 MB stored (4 gzip chunks); the WASM replay of the stored chunks matched all 644 live statuses. With the autopilot on a recording grows by ~10 KB per frame raw, ~1.2 KB stored. * The recorded TILE_MAP words equal the tile map in the run's start.ram.

Deviations and additions: * A recording that starts in the middle of a frame skips that partial batch. When the autopilot stepped on that skipped frame the recorder writes a SYNC autopilot event, and the replay starts its session there and marks it stepped_unrecorded_frame (it cannot reproduce the missed step). This was the only mismatch in the first parity run. * .mjs is served as JavaScript by site/serve.py (the replay module is an ES module). * Old desktop runs have no TILE_MAP records; pack_recording.py adds them from terrain.json or, for validate_campaign.py runs, from start.ram, and writes a gzip file the web page can import.

Phase 3 (replay page)

Built: /replay/ (site/content/replay.html, site/assets/replay.css, card on the Autopilot hub, "Replays" link on the play page), the app in web/src/replay/ (main.ts page controller, worker.ts, protocol.ts, library.ts, timeline.ts, sceneView.ts), a second vite build (web/vite.replay.config.ts -> web/dist-replay/), SpriteStreamLayer split out of SpriteStreamView (the play page now draws through it), AtlasSource (Emscripten FS or fetch), and build_site.py --replay-build (bundle + module to /replay/app/, low-resolution atlases to /replay/atlases/, 16 MB). hatari_dev.ps1 -Action build-web builds both bundles and copies the replay module next to the page bundle.

Worker: the recording stays as its compressed IndexedDB chunks; opening decompresses each chunk once to index messages and peek every frame (timeline markers), and keeps six decompressed chunks. Frame b is shown after feeding batch b and its trailing AUTOPILOT_STATUS records. Checkpoints are taken the first time playback passes every 256th frame (at most 128 per recording; each costs about 0.34 MB of WASM heap); backward seeks restore the nearest earlier checkpoint. Long seeks report progress and give way to newer requests.

Measured in the browser (the web recording from Phase 2 and the packed Level 2 desktop run loaded with ?src=): real-time playback at 50 fps, forward seeks of 400-640 frames in about 50 ms, replayed controls equal the recorded live status on every sampled frame, including after checkpoint restores. The play page renders unchanged through the shared layer.

Deviations: the right-hand panel shows interim Frame / Autopilot / Replay session cards and the raw inspection JSON until the grouped cards of Phase 5 replace it. library rename uses the browser prompt. The page's .mjs module needs a JavaScript MIME type (site/serve.py sets it).

Phase 4 (overlay)

Built: web/src/replay/overlay.ts (three.js line and triangle batches in logical coordinates, an HTML label layer, hit-testing), overlaySettings.ts (categories from the C hazard-kind names with the Python palette, layers, presets Clean / Decision / Hazards / Navigation / Everything, settings kept in localStorage), overlayControls.ts (toolbar: overlay switch, presets, Layers and Objects popovers, sprite dimming, look-ahead, isolate). Hover shows a tooltip, a click selects the smallest box (kept across frames and sent to the worker, so its predicted path and raw bytes are always inspected), Esc clears, [/] move the look-ahead (Shift = 8 frames). The page asks the worker only for the sections the enabled layers need.

Coordinates, measured from the inspector output: tracks and walls are screen space; predictions, scene bodies, plan frames, route, goal and geometry are world space with screen y = world y - (scroll - 1); spans are (row, left, right) in 4-pixel cells with world y = 4 * row (checked: no free span overlaps any wall tile). Future positions are drawn where they will be on screen at their frame t, with the camera the winning plan forecasts for that frame, so bodies and the ship compare correctly at equal t. The autopilot forecasts the camera only along its plan (56 frames, 80 while homing sources are selected, fewer when the plan is shorter), so the overlay draws nothing past the plan's last frame: predicted paths, scene bodies, homing ghosts, anticipated shots and pickups all stop there. (Until 2026-09-26 paths ran to 80 frames past a 56-frame plan with the camera extrapolated from the current frame's scroll velocity, which is +1 on a frame where the camera steps back: a jump and a wrong direction past +56.) The look-ahead slider spans the configured maximum horizon and keeps its value; on a frame with a shorter plan it reads "+N (plan ends)".

Measured: 1.4 ms per frame in the worker with the Decision preset's sections (tracks, predictions of used bodies, plan, decision, mission).

Deviations: labels use a small HTML layer instead of CSS2DRenderer (same effect, fewer dependencies); lines are 1 device pixel (WebGL line width); the "recorded hits" layer is left to the Ship card and the timeline's hit markers (Phase 5).

Phase 5 (state panel, selected object, object table)

Built: web/src/replay/cards.ts (Diagnostics, Decision with the proposed -> arbitration -> applied masks, Target & weapon, Route & navigation, Level mission with the current level's decoded state and the driver's wall recovery, Ship with attachments and hits, Input & camera, Shop (shop frames only), World, Configuration with recording info, replay module ABI, session config), selected.ts (identity, kind, roles, the sprite cut from the loaded atlas at up to 4x, position and boxes, motion and model, predicted path and scene membership, decoded fields, raw object bytes as hex), objects.ts (sortable table, Shift+click for secondary keys, search, category chips, zone, planner-only, "scene filters" on by default; hover highlights in the scene, click selects, the selected row stays on top). Cards remember open/closed state, changed values flash when stepping, and the panel refreshes at most ten times a second during playback. The first frame where the replayed controls differ from the recorded live status becomes a red timeline marker and a Diagnostics row.

Deviations: the object table sits under the scene (it needs the width) rather than in the right column; the level map (Phase 6) takes the space the plan gave it there. Recordings without AUTOPILOT_STATUS show no input-match badge (comparing with the next frame's recorded input needs the next batch, which the worker does not peek yet). "Follow" is implicit: the selection is an identity and stays across frames; "show path only" is the Isolate button.

Phase 6 (level map)

Built: web/src/replay/mapView.ts, a second three.js view of the C autopilot's map of the whole level: solid, unknown and destructible cells (open groups shaded), level zones, physical or travel spans, the camera viewport, the player's captured collision bounds with the planner's reserve, the mission route with its current index, the destination and the goal rectangle. It follows the viewport by default; the wheel or a drag looks around, "Whole level" fits the level into the pane, and a click reports L tile=(c,r) world=(x,y) screen y=... solid|free|destructible span=[l,r] with a copy button. The map sits next to the object table under the scene.

The worker sends the map (and the spans with it) only when the level or status.map_revision changes; the page asks for it again right away when a frame reveals a new revision. Map row r covers world y (origin + r) * 16, measured the same way as the spans: no solid cell overlaps a free span under this mapping, while the flipped or origin-less mappings collide thousands of times.

Not verified visually in this session: the preview pane stopped painting (requestAnimationFrame paused while the app window was hidden); picking, labels and data flow were checked through the page's DOM.

Built: an "HD sprites" switch (off by default) that downloads the 4x banks for the current level when first switched on and falls back to the original sprites with a message when a site was built with --replay-no-hd; build_site.py publishes every atlas bank's pixels gzip-compressed (<stem>.raw.gz, decompressed with DecompressionStream and sniffed in case the host already decoded them): the original banks shrink from 16 MB to 0.8 MB, the HD banks from 248 MB to 94 MB. Measured before choosing: PNG would be smaller for HD (7.8 vs 11.6 MB for Level 2), but the HD and shared banks use partial alpha, which canvas decoding would not return exactly. Deep links now carry the selection (?rec=&frame=&sel=); a joined-mid-session recording gets a timeline marker where the replay's session begins; a Keys popover lists the shortcuts. ?src= (Phase 3) and the level, life, shop, hit, autopilot and divergence markers (Phases 3 and 5) complete the list.

Still open, as decided: sample recordings under /replay/samples/ once the autopilot finishes the game (parallel work).

Completion pass (plan audit, verification section 8)

An audit of sections 5, 7 and 8 against the build found and closed these gaps:

  • Section 5 drawing: homing missiles as the planner simulates them against the winning plan (the candidate workspace's initial states stepped along the forecast; read-only accessors xap_driver_workspace, xap_workspace_homing_count); scene-source flags; "used by the planner" only for paths the planner collides against (independent sources and boss shells; homing sources become homing requests in xenon_autopilot_world_scene.c); paths coloured by motion model with +8 ticks; fitted world velocity over 8 frames; linear body velocity rays; kill and collect markers; the plan coloured by clearance with its verified state; the weapon mount line; route spans highlighted; a recorded-hits layer.
  • Section 7: per object "drawn" and the captured motion script; a composite preview (the frame's sprites inside the box) for objects without an atlas image such as Level 5 laser gates; the header compares with the next frame's recorded input when no live status exists, and shows the run id.
  • Bugs found on the way: Chrome's DecompressionStream rejects gzip members after the first, so exported recordings could not be re-imported (now split member by member); the Python replay path (native_memory.pack) never set XAP_PLAYER_HAS_WALL_REPLAY, so replay_ui.py skipped the Level 2 wall recovery (fixed; 25 of 3000 frames differed before); overlay and map colours were passed to three.js as linear values and shown lighter than the Python palette; the Python path sent the protection flag as 1 instead of the game's byte (below).

Verification (section 8), with the recordings web-replay-parity-l2-r1 (final boss), web-replay-wall-l2-r1 (three-eye arena, wall contacts), web-replay-homing-l2-r1 (homing rings), l2-shop-r24, l2-full-r23 (Level 1) and web-replay-parity-l3-r1 (Level 3 gates):

  1. Decoder parity: xenon_tools/decoder_parity.py compares the C decoder's session input (xrp_session_input) with native_session.pack() field by field: identical on every frame of all six recordings. It reports the old wall-replay bug on every frame. Levels 4 and 5: see "Level 4 and 5 checks" below.
  2. Decision parity: C replay and live status identical on 1611 + 3000 + 2001 stepped frames; the Python session path (decoder_parity.py --decisions) too, including map_revision.
  3. Tile-map subset: --tile-map-subset syncs one session from the run's full start.ram and one from the 12 000-byte tile map; status and every map cell identical on all 1199 Level 3 frames (Level 2's only group is the boss arena, which the sync skips). A blank RAM image diverges on the first frame.
  4. WASM vs native: identical (Phases 0-2); rechecked after the inspector additions.
  5. Render parity: render_parity_frames.py plus the page's ?debug=1 capture: 12 of 12 frames of the homing run pixel-identical to the desktop Sprite Stream AVI at 640x400.
  6. Round trip: web recording -> library export (4 gzip members) -> import: 968 frames, same bytes; desktop recordings and packed runs open in the page (?src=) and in replay_ui.py.
  7. Arbitration log neutrality: arbitrate() only appends (stage, mask) to a log that nothing but the clone and the diagnostic accessor reads; the masks it receives are already final.
  8. Performance (desktop browser, ?bench=1&perf=1, autopilot on): emulation 2.2 ms/VBL with and without recording (p90 2.3 vs 2.2). The web recorder stores 1.95 KB per game frame (14 KB raw): a 30-minute run is about 44 MB in IndexedDB. Replay of a 7275-frame web recording: open 0.33 s, forward jump of 6000 frames 0.96 s, backward jumps 10-200 ms, one step with the full panel 62 ms. Not measured: a phone.

Level map: checked offscreen (mapCapture()): all 231 solid cells in view at their C map positions and colour; the preview pane was not painting for a screenshot.

Level 4 and 5 checks (the older clone's run_logs, and fresh runs from its saves):

  • Decoder parity on every older recording: 289 Level 5 recordings from the Python autoplay (28 August to 2 September, 163 601 frames) and 134 Level 4 and 5 recordings from the native autopilot (15 September, 294 328 frames). Identical on every frame after one fix: the live capture hands the autopilot the game's protection byte (0xFF while active, src/xenonCapture.c), the C decoder copies it, but decode_frame stored a bool and native_memory.pack sent 1 (8 637 frames after the shop's protection purchase differed). The autopilot only tests it for zero, so no decision changed; pack now sends the recorded byte.
  • Fresh runs with the replay build: web-replay-l5-entry-r1 (Level 5 start, 1078 frames), web-replay-l5-tank-r1 (tank, 169 frames, last life) and web-replay-l5-protection-r1 (protection active on all 970 frames). Decoder parity, the Python session against the live status (--decisions), the tile-map subset and the C replay (replay_parity.py, live status matched on every frame) all identical. Level 5's tile map is not trivial: a blank RAM image diverges on the first frame.
  • Render parity on Level 5 (its own atlas bank, laser gates): 12 of 12 frames of web-replay-l5-entry-r1 pixel-identical to the desktop Sprite Stream AVI.

Explanations (2026-09-27): every layer in the Layers popup, the Objects popup's columns and every row of the state cards have an "i" button that opens a short explanation under the row (web/src/replay/help.ts; texts next to their definitions in overlaySettings.ts and cards.ts, checked against the C code: e.g. "plan: candidate #N won" is the Nth candidate in the search's fixed order, xenon_autopilot_search.c). Found on the way: the goal band is screen rows, not world y (the evaluator compares it with the ship's screen y), so the overlay drew its lines about 1700 px off screen; fixed.

Still open, as decided: sample recordings once the autopilot finishes the game.

Merge with xenon2 (2026-10-01)

origin/xenon2 up to 4a03e159 (three commits: Level 3 carrier and boss cash, Level 4 and 5 tactics and the tank cannon stall, Level 5 progression and the final ship; ABI 91 to 98) was merged into this work as the web-replay-merge branch: a merge commit, no rebase, xenon2 itself untouched.

Conflicts and how they were resolved:

  • xenon_autopilot_driver.c: xenon2's control logic as it is, with the arbitration trail added back on top (the file differs from xenon2 by additions only). The Level 4 recovery stage went with the calls xenon2 removed; six stages follow xenon2's new per-level actions (level 3 carrier, level 4 entry egress, level 4 second bend, level 4 first boss forward, level 4 boss cash, level 5 final cash). The trail's capacity is now the stage count.
  • xap_workspace_homing_count was added on both sides. The header conflicted on its comment only; xenon_autopilot_workspace.c merged silently into two definitions. xenon2's is kept, and the "inspection only" comments are gone: its planner now calls it while driving.
  • native_memory.py imports both new flags; replay_ui.py combines the recording's tile maps with xenon2's validation-run terrain and scene snapshots.

Follow-ups:

  • ABI 99, in C and autopilot_v2/native.py.
  • xenon2's capture trailer 0x5841 (the $CC8 Temporary Shield countdown, 116 bytes) in the C decoder; 0x5840 recordings still decode. The countdown is a row of the capture card.
  • Level 5 final ship: the live pilot hands the session RAM again when the final ship's controller appears (the ship replaces the terrain across its arena). XenonAutopilot_WillSyncLevel, which the recorder asks, and XenonAutopilot_Frame now share one test, so a TILE_MAP record is written on that frame. The C replay, decoder_parity.py --decisions and replay_ui.py (as a scene snapshot) apply it there, and replay checkpoints keep the flag.
  • Level card: the Level 3 carrier refuge (driver state; its names come from xap_level3_carrier_refuge_name next to the enum), Level 5 core_attack_ready and final_flank, with explanations, and texts for the six stages. The Level 3-5 texts were checked against xenon2's code and corrected where it changed them (Level 3 core selection and firing x; Level 4 bypass, second-boss refuge and shield chase; Level 5 diagonal crossing, tank firing x and train hold).
  • Generated data (xenon_tools/generate_native_assets.py, AUTOPILOT_GENERATED_ASSETS.MD): xenon2's two new destructible groups (level4-first-boss-wall, level5-tank-center-wall) were produced properly: level_tile_maps.py rebuilds assets/xenon_level_maps.json byte-identically from xenon_tools/level_capture, and --check finds xenon_autopilot_map_data.c current. xenon_autopilot_collision_data.h was stale, from this branch's side: the intro sprites injected into the shared atlas (ae68e6ed) had not been regenerated into it. Regenerated: the only change is 37 new entries per level for those sprites, every existing sprite's mask is identical, and the replays above still match the live runs on every frame. The page's level map leaves the Level 4 wall unlabelled (it has no anchor cell).
  • Fix in xenon2's code: xap_driver_clone did not copy carrier_refuge, so a replay checkpoint inside the carrier fight would have reset it. Live play never clones.

Verification, with the merged desktop build (out/build/mingw-replay) and fresh recordings under E:\xenon_runs\web-replay-merge:

  • xenon2's 31 changed test modules (392 tests): the same 5 fail on pure xenon2 (four Level 2 eye/pocket mission tests and test_world_decode), nothing else.
  • Driving unchanged: each recording's frames, packed by the same Python, stepped through xenon2's own DLL (a worktree at 4a03e159) give the live merged run's status on every stepped frame ("xenon2 DLL" below). Session input and status layouts are the same in both, so the ABI number check was relaxed for this test only.
Recording From Stepped frames C replay = live, WASM = native, decoder, Python = live, tile-map subset, xenon2 DLL = live New code exercised
merge-l2-r1 l2-newtactic-r30 start 1553 all equal (Level 2 stages)
merge-l3-carrier-r1 level3-carrier-cash-20260924-r1 1503 all equal all 10 carrier refuges; level 3 carrier on 307 frames
merge-l4-r1 native-level3-visible-20260914/end.sav 3002 all equal level 4 entry egress; Temporary Shield countdown on 170 frames
merge-l4-bend-r1 level4-bend-r90a start 1201 all equal level 4 second bend
merge-l4-cash-r1 level4-cash-r101 start 285 all equal level 4 first boss forward
merge-l5-entry-r1 native-level4-to-level5-visible-20260915-final.sav 1003 all equal
merge-l5-tank-r1 level5-tank-aggressive-fire-validation-0901-116.sav 139 all equal
merge-l5-tankfight-r1 level5-whole-staged-r125 f76020 902 all equal tank emplacements and cannons (stopped before the core phase)
merge-l5-final-r1 level5-whole-staged-r125 f81309 1038 all equal final ship map sync at 81369; final_flank left/right; level 5 final cash; game completed

Not reached by these runs: the level 4 boss cash stage and core_attack_ready set. replay_parity.py's other measure (the replayed controls against the next frame's applied input) misses the last frame of runs that stop with the pilot released; the status comparison covers it.

  • Negative control: merge-l5-final-r1 replayed with the pre-merge rule (RAM at level entry only) diverges at frame 81369, the final ship's first frame: 58 of 1038 frames match.
  • Render parity, page against the desktop Sprite Stream AVI: gameplay frames pixel-identical (Level 3 5/5, Level 4 10/10, Level 5 final ship 12/12). Shop and level-start screens are not (Level 3 shop 7 of 7 frames, two Level 4 start frames by 972 px). The same happens on the pre-merge xenon2 recording level2-shop-level3-20260924-r1 (all 10 shop frames and three level-start frames differ, all 23 gameplay frames identical), so it is not from the merge; the earlier checks had no such frames. Open.
  • Cost: with xenon2's tactics a replayed Level 5 frame takes 8 ms (entry) to 21 ms (final ship), natively and in WASM; Levels 2-4 take 0.1-0.4 ms. Playback on the final ship is about real time, and long seeks there are slow.
  • The replay page shows the new stage rows and the carrier refuge with their explanations (merge-l3-carrier-r1, frame 20566).

Explanations from C (2026-10-01)

Before this, a level field lived in three places: the struct member, a hand-written JSON line in xenon_autopilot_replay.c and a text in cards.ts (LEVEL_FIELD_HELP; likewise the driver stages and the configuration). Now the struct member and its marked comment (/*! ... */) are the only source:

  • xenon_tools/generate_replay_descriptions.py reads the comments in xenon_autopilot_mission.h (Level 1-5 states), _session.h / _decision.h (configuration) and _driver.h (XapArbitrationStage). It writes src/autopilot/xenon_autopilot_replay_fields.h, which xenon_autopilot_replay.c includes: the JSON writers for level_state and config, the enum name functions (also hazard kinds, tactics, weapons, replacing hand-kept tables), and the texts with their format and C source, returned by the new xrp_describe().
  • The worker sends them with opened; the cards show them with the C member or enumerator under each text, and choose how to show a value (object, joystick mask, bit list) from the description instead of guessing from the key's name. cards.ts lost about 200 lines of texts.
  • Enforcement: a member or stage without a marked comment stops the generator. The generated header copies each described struct's layout, so a changed struct fails the build until it is regenerated (shown by adding a member: size of array 'xrp_regenerate_replay_fields_XapLevel5MissionState' is negative). test_replay_descriptions.py checks the output is current and the rules hold. CLAUDE.md tells autopilot sessions to keep the texts true.
  • C changes besides the comments, all value-preserving:
  • the level structs declare one member per line;
  • the Level 2 pockets, crossings, final-boss pockets and modes, and the Level 3 carrier positions became public enums, and the private constants in level2.c and level3.c are now aliases of them;
  • three stage enumerators were renamed to their display names (MISSION_OVERRIDES, LEVEL2_CORRIDOR_GUARD, LEVEL2_AUTHORED_ROUTES);
  • xap_level3_carrier_refuge_name was removed.
  • Two display fixes came with it: pickup_escape (0/1/2) had been written as a yes/no, and corridor_patrol_direction now carries its name (off, left, right).

Verified: the full inspection JSON of the old and new library, frame by frame, is identical apart from those two fields and key order. That covers 31 491 frames from 12 recordings: the 9 merge recordings, two Level 2 boss runs and a full-campaign run with 7 899 Level 1 frames. On the page, checked on screen: * Level 1 (full-campaign-native-release-20261001-r136, frame 4000): the attack target shows as an object; * Level 2: corridor patrol shows as left; * Level 3 (merge-l3-carrier-r1, frame 20566): carrier refuge and the level 3 carrier stage, with their C sources; * Level 4 (merge-l4-r1, frame 63921): pickup target and pickup position; * the Configuration card: texts from XapDecisionConfig / XapSessionConfig.

There were no console errors. The native test modules (497 tests) fail the same 9 tests with the library built before the change. One of them, test_native_world_homing, is flaky with either library: its failing subtests change from run to run.

Merge with xenone2-l2 (2026-10-02)

origin/xenone2-l2 (62268dfd) branched from 4a03e159, before any of the replay work. It adds three commits: the Level 2 worm-transfer and post-shop-swarm fixes, uncapped native fast-forward, a -BuildType Release option and campaign validation. It was merged into xenon2 with a merge commit.

Conflicts:

  • xenon_autopilot_driver.c: the branch's gate on the Level 2 hazard escape (skipped while a verified corridor retreat steers) is kept, with the arbitration calls around it.
  • xenon_autopilot_level2.c: both sides appended at the end; the branch's new functions and the geometry export are both kept.
  • hatari_dev.ps1: the branch's -BuildType and -MemoryTrace are kept, beside the kernel, WASM and site folders.

Follow-ups:

  • ABI 100, in C and native.py. Both sides had moved 98 to 99, for different layouts.
  • The branch's new object field level2_worm_origin (the hole a worm-train member came from) is written to the per-object fields, so the selected-object card shows it. It is decoded from the recorded object bytes, so older recordings have it too.
  • Texts updated for the changed behaviour: eye_plan (after the right train's leader dies the ship keeps firing until the next left train is out and the route up is clear), transfer_stage (the launch crosses at screen row 16), hazard_escape_frames, and the "level 2 hazard escape" stage.
  • Overlay: the branch's zones in the Level 2 geometry are the stream-lip retreat, the post-shop stream, its lane at x 160-176 (in phase 2) and the eye-transfer row.

Same driving as the branch, verified three ways:

  • Source: the merged autopilot differs from the branch only by the replay's recording calls (written, never read), read-only exports, comments, constant aliases with the same values, the ABI number, and the intro sprites in the collision table. The host's capture code is equivalent.
  • Closed loop: the branch's eight final validation runs were repeated with the merged Release build from the same snapshots and commands (E:\xenon_runs\l2-merge-check):
  • r172: the post-shop stream, spider, shop and entry to Level 3;
  • r171 and r170: the post-shop stream;
  • r169: five starts through the three-eye boss.

The joystick input is identical on all 12 874 common frames, apart from the final frame of r172 and r171, where the run monitor had already released the autopilot. The whole captured frame (game memory, objects, draws) is identical too, apart from: * the replay recorder's terrain-mask draw flag (8914a1fa), which the branch's capture does not write; * the frames after the stop condition, where the monitor releases control at slightly different moments. * Open loop: the branch's own DLL, fed the merged runs' frames, makes the live merged decision on all 13 270 stepped frames.

The C replay equals the live status on all 13 270 frames, the decoder and Python path agree, and WASM equals native (r169-f10038, r170). The test modules give the same results as on the branch.

Constants in texts and per-level geometry (2026-10-02)

Explanations no longer copy numbers. A marked comment names the value it quotes: {WALL_ATTACK_BUDGET}, {boundaries[0]}, {len(backbone) - 1}. generate_replay_descriptions.py reads the integer constants and numeric tables of src/autopilot and writes in the values; an unknown or ambiguous name stops it.

About 70 measurements were converted. Values the code had written inline got names in their level files first, all with the same values: * Level 1's side range; * Level 2's two 40-frame final-boss rules and the route points' reach distance; * Level 3's intercept, gate station and carrier positions; * Level 4's bypass, refuge, side stations, retraction, wall-gun progress and shield chase; * Level 5's core station and attack, train band and final-ship flanks.

The expanded texts equal the previous ones, except "nine gates" now reads "9 gates", and one addition found on the way: a hazard escape also commits 96 frames for a worm escape on the eye route when the session has no predictions.

Each level now owns its overlay geometry: xap_level1_geometry() to xap_level5_geometry() in the level files, and xap_level_geometry() picks the level. Besides Level 2's list and the items the replay used to choose itself (intercept, pickup chase, station, refuge, train hold), the overlay now draws: * Level 1's corridor waypoints; * Level 3's intercept rows and carrier refuges; * Level 4's shield-chase band, second bend, and second-boss refuge and stations; * Level 5's tank column, core-attack row, train stage band and final-ship flanks.