Xenon 2

Autopilot · write-up

Prediction reuse (2026-09-06)

xenondoc/AUTOPILOT_PREDICTION_REUSE.MD · 13 KB · updated 2026-09-17

Subsequent work adds independently selectable architecture features; see architecture experiments and gameplay gates. prediction_reuse=False controls the original tape/shared-formation optimization. Use --architecture none as well when comparing with the complete pre-architecture maneuver implementation. compare_architecture.py isolates the newer changes.

The maneuver planner now defaults to PlannerConfig(prediction_reuse=True). The legacy controller and its predictor remain unchanged. Set the option to False to retain the maneuver planner's previous prediction/validation path. Navigation, search budgets, missions, travel pacing and destruction rules are held fixed by this optimization.

What work is removed

autopilot_v2/trajectories.py resolves ordinary $502C2 follower chains into shared root predictions and delayed member tracks once per observation. Every link still copies its successor's previous-frame screen anchor, including the camera-origin adjustment. Current follower positions supply the initial prefix; the leader supplies the remaining trajectory. Semantic attachments and Level-3 composite chains retain their separate reference models.

Offscreen body selection first tests a conservative whole-track bound. If it could approach the playfield, the existing four-frame sample test is performed numerically against the current collision offsets. This avoids constructing a rectangle for every irrelevant member/time sample. It preserves the reference selection rule; it does not introduce a new sampling interval or safety margin. The full root horizon is requested once, avoiding repeated cache expansion for each successive horizon. The longer homing horizon is built only when a homing enemy or modeled launcher actually calls for it; ordinary travel uses the ordinary maneuver horizon, including the pending input frame.

autopilot_v2/script_tapes.py exposes immutable intermediate states of the sine script interpreter. At the next canonical frame it compares the captured full state with the expected successor: fixed-point coordinates, phase, acceleration, remaining substeps, script cursor/start/window/bytes and update rate. Object identity, run, updater and consecutive frame number must also match. A match reuses the suffix and computes only the newly requested tail. Otherwise the trajectory is rebuilt. Current world anchors and camera velocity are applied afterward, rather than retained from an old collision rectangle.

Delays, deterministic jumps, metadata and termination retain the reference semantics. Missing or random commands retain its tangent fallback; an unknown successor never certifies cross-frame reuse. Entries are bounded by the scripts requested in the preceding/current decision and their horizons. No mutable trajectory entry is shared between replay snapshots. Checkpoint restore discards these caches.

A bounded 4,096-entry cache also shares the pure sine/phase increment calculation independently of position. Phase, phase delta, acceleration and substep count are its complete inputs. Thus a changed captured camera anchor still invalidates the trajectory, while identical arithmetic can be reused at the new origin with the same signed 32-bit wraparound. This does not relax state validation.

PredictionService additionally retains independent-body clearance checks within one decision. The key includes future frame, exact player collision hull, camera correction and incoming clearance. Shared maneuver prefixes can therefore reuse the same check. Nothing is carried to another observation, and the shortcut is bypassed after combat simulation begins. Walls, homing motion, aimed shots, shell release histories, pickups and destructible terrain continue through their existing validation paths. Emitted shots survive destruction of their source; observed terrain remains authoritative.

Comparing and inspecting

Run a sequential, unprofiled reference/shared comparison with spans fixed:

python xenon_tools/compare_autopilot.py RECORDING.x2events --prediction-only --from-frame FIRST --to-frame LAST --output xenon_tools/run_logs/NEW-DIRECTORY

The tool restores available start-controller and terrain sidecars, warms up earlier observations, fingerprints source/assets and records both raw timing streams. changed_actions compares selected commands; changed_validation compares stage, verification, contact and credited destruction/collection. Fixed observations alone do not establish closed-loop gameplay quality. Each variant runs in a fresh Python process: module-level navigation caches must not inherit the preceding variant's contents or eviction history.

For a function profile use profile_autoplay.py replay with --cprofile and either --prediction-reuse or --no-prediction-reuse. Do not interpret profiler timings as device performance. The new default is included in saved planner configuration and the prediction_reuse decision metric.

Decision metrics now include formation_roots, formation_tracks, formation_bound_rejections, script_rebuilds, script_computed_frames, script_reused_frames and independent_clearance_hits. These count actual work or reuse, rather than speculative enemies destroyed.

Replay UI

replay_ui.py now labels the recorded prediction mode separately from the current replay analysis mode. New .planner.json sidecars record the effective prediction_reuse setting, including a setting restored from the starting controller checkpoint. For recordings made before this metadata was added, replay checks a bounded prefix of the decision trace for the setting. If it was not recorded, the caption says unknown (not recorded) and identifies the analysis mode as the current default; it does not invent a historical mode.

The Replay prediction selector offers Recorded, Optimized and Reference. Changing it clears replay caches and reconstructs the current frame, so retained plans and counters cannot leak between modes. Rewind/reset preserve the selected override. The command-line equivalents are --prediction-reuse and --no-prediction-reuse with the maneuver planner.

The GUI loads the decision sidecar independently of the navigation checkbox. The recorded and analysis maneuver summaries each show formation roots/tracks/ rejections, computed/reused script frames, script rebuilds and reused independent collision checks. Recorded rows are accepted only at matching canonical frame and VBL. Missing historical values remain unavailable rather than becoming invented zeroes. Full sidecar indexing is confined to the GUI.

Swarm / formation paths (analysis) is optional and off by default. It draws cached root anchor trajectories in cyan, delayed follower trajectories in dashed purple, and independent scripted swarm trajectories in gold. Root and script paths have time markers every 16 frames. The paths use the default camera forecast, respect selected-object filtering and are explicitly labeled as replay analysis. This overlay reads existing tracks; it never populates prediction caches or claims to show recorded cache internals. Script paths are also available in reference/legacy analysis when cached; shared leader/follower tracks require optimized maneuver prediction. Empty labels distinguish inactive planning, selection filtering, offscreen paths and missing cached predictions. Not every visible enemy necessarily has a cached path. Script displacement dictionaries are isolated in replay snapshots so backward seeks retain the correct frame. Existing navigation and AVI overlays retain their own meanings.

The initial overlay missed independent scripted swarms: frame 1600 of prediction-reuse-corridor-0906-02 contains 12 swarm enemies and seven cached script paths but no shared leader/follower tracks. It previously showed an empty label despite those predictions being available. The fix passes 52 replay/prediction/recording tests. A real Tk replay check shows seven gold paths at frame 1600, preserves those paths after stepping forward and backward, and shows 16 paths spanning all three types at frame 1682. The updated checkbox fits within the toolbar.

Replay validation: the full Python suite passed 808 tests; 87 focused tests passed after the final legacy/empty-sidecar compatibility checks. A real Tk-window smoke test on prediction-reuse-corridor-0906-02 at frame 1682 verified recorded counters with navigation disabled, 31 formation-overlay canvas items, reference/recorded mode switches at the same frame, and toolbar widget bounds.

Validation

The full Python suite passes 800 tests. Differential tests cover interpreter commands and integer overflow, delayed chains, fast offscreen approaches, live collision extents/link changes, camera conversion, replay-safe immutability, state/identity/frame invalidation and unknown branches. Clearance tests compare cached/reference results, preserve aimed-shot checks and require camera offsets to distinguish otherwise identical queries. One pre-existing protocol test assumed --stop-file was the last CLI argument; it now checks the flag's value independently of the already-existing planner option.

The initial implementation matched all 1,755 corridor decisions but showed essentially no net speedup. Profiling exposed repeated root-cache growth while requesting consecutive horizons. Requesting the longest horizon first improved that intermediate comparison to 1.08x. Final measurements below also include reuse of independent clearance checks.

Final sequential comparisons (prediction-isolated-*-0906-05, Python 3.13.7, same source/assets/fire/terrain/controller state, no profiler or live emulator workload during measurement, fresh Python process per variant):

Recorded observations Frames Reference mean Shared mean Reference/shared p95
cruise-spans-0905-03, corridor 801..2555 1,755 21.989 ms 18.331 ms 55.139 / 47.957 ms
shell-core-spans-0905-06, boss 7633..9030 1,398 15.709 ms 15.567 ms 30.909 / 29.253 ms
level5-barrier-clearance-0830-118, 90160..90459 300 16.773 ms 16.012 ms 18.263 / 17.466 ms

All 3,453 selected actions and validation outcomes match. These measurements include observation reconstruction. Corridor mean time is 16.6% lower (1.200x throughput); its formation-heavy 2400..2555 ending falls from 21.538 to 15.778 ms, 26.7% lower. Mean corridor policy time falls from 5.599 to 3.546 ms. Independent group queries fall from 675 to 447 per decision; 45 clearance queries per decision are reused on average. About 106 script frames per decision reuse an existing suffix. These counters explain which repeated work was actually removed.

Boss average differences are small: treat its roughly 1% mean difference as effectively unchanged. Level-5 mean time is 4.5% lower in this short sample. This does not solve the remaining maneuver-validation/ship-motion cost or establish phone/WASM performance.

Earlier prediction-final-*-0906-04 reports shared an interpreter. Examining a boss spike exposed different navigation LRU histories between variants. Those timings are superseded by the isolated runs above; navigation row-build/hit counts now match between variants on both corridor and boss observations.

Reports with fingerprints and raw timing streams:

Closed-loop retests

Both use maneuver/spans, canonical frames, no shield cheat or fast-forward, paired original/sprite AVI, decision traces and periodic checkpoints.

  • prediction-reuse-corridor-0906-02.x2events: original frame-800 checkpoint through the shop at 2556. One 8-point hit at 1940, zero lives lost, final shield 31, three lives, nine pickups, eight wall shooters destroyed and 9,310 points gained. These outcomes match the accepted cruise-spans-0905-03 baseline exactly.
  • prediction-reuse-boss-0906-01.x2events: original frame-7632 checkpoint through the shop at 8477. Boss objects are present at 7633 and gone by the final recorded frame 8476. Zero damage, zero lives lost, shield 39, three lives, and 17 pickups. The historical boss recording predates the final core-rejoin tactic adjustment; its later shop frame is not evidence of a prediction-optimization gameplay improvement.

The first corridor attempt (prediction-reuse-corridor-0906-01) used the copied start.sav and stopped on a stale shop indication at 801. It is excluded from validation. The accepted retest uses the original checkpoint and its emulator metadata sidecar. Hatari is paused with automation input released.