Autopilot · write-up
Prediction reuse (2026-09-06)
Subsequent work adds independently selectable architecture features; see
architecture experiments and gameplay gates.
prediction_reuse=False controls the original tape/shared-formation optimization.
Use --architecture none as well when comparing with the complete pre-architecture
maneuver implementation. compare_architecture.py isolates the newer changes.
The maneuver planner now defaults to PlannerConfig(prediction_reuse=True).
The legacy controller and its predictor remain unchanged. Set the option to
False to retain the maneuver planner's previous prediction/validation path.
Navigation, search budgets, missions, travel pacing and destruction rules are
held fixed by this optimization.
What work is removed
autopilot_v2/trajectories.py resolves ordinary $502C2 follower chains into
shared root predictions and delayed member tracks once per observation. Every
link still copies its successor's previous-frame screen anchor, including the
camera-origin adjustment. Current follower positions supply the initial prefix;
the leader supplies the remaining trajectory. Semantic attachments and Level-3
composite chains retain their separate reference models.
Offscreen body selection first tests a conservative whole-track bound. If it could approach the playfield, the existing four-frame sample test is performed numerically against the current collision offsets. This avoids constructing a rectangle for every irrelevant member/time sample. It preserves the reference selection rule; it does not introduce a new sampling interval or safety margin. The full root horizon is requested once, avoiding repeated cache expansion for each successive horizon. The longer homing horizon is built only when a homing enemy or modeled launcher actually calls for it; ordinary travel uses the ordinary maneuver horizon, including the pending input frame.
autopilot_v2/script_tapes.py exposes immutable intermediate states of the sine
script interpreter. At the next canonical frame it compares the captured full
state with the expected successor: fixed-point coordinates, phase, acceleration,
remaining substeps, script cursor/start/window/bytes and update rate. Object
identity, run, updater and consecutive frame number must also match. A match
reuses the suffix and computes only the newly requested tail. Otherwise the
trajectory is rebuilt. Current world anchors and camera velocity are applied
afterward, rather than retained from an old collision rectangle.
Delays, deterministic jumps, metadata and termination retain the reference semantics. Missing or random commands retain its tangent fallback; an unknown successor never certifies cross-frame reuse. Entries are bounded by the scripts requested in the preceding/current decision and their horizons. No mutable trajectory entry is shared between replay snapshots. Checkpoint restore discards these caches.
A bounded 4,096-entry cache also shares the pure sine/phase increment calculation independently of position. Phase, phase delta, acceleration and substep count are its complete inputs. Thus a changed captured camera anchor still invalidates the trajectory, while identical arithmetic can be reused at the new origin with the same signed 32-bit wraparound. This does not relax state validation.
PredictionService additionally retains independent-body clearance checks
within one decision. The key includes future frame, exact player collision
hull, camera correction and incoming clearance. Shared maneuver prefixes can
therefore reuse the same check. Nothing is carried to another observation, and
the shortcut is bypassed after combat simulation begins. Walls, homing motion,
aimed shots, shell release histories, pickups and destructible terrain continue
through their existing validation paths. Emitted shots survive destruction of
their source; observed terrain remains authoritative.
Comparing and inspecting
Run a sequential, unprofiled reference/shared comparison with spans fixed:
python xenon_tools/compare_autopilot.py RECORDING.x2events --prediction-only --from-frame FIRST --to-frame LAST --output xenon_tools/run_logs/NEW-DIRECTORY
The tool restores available start-controller and terrain sidecars, warms up
earlier observations, fingerprints source/assets and records both raw timing
streams. changed_actions compares selected commands; changed_validation
compares stage, verification, contact and credited destruction/collection.
Fixed observations alone do not establish closed-loop gameplay quality.
Each variant runs in a fresh Python process: module-level navigation caches must
not inherit the preceding variant's contents or eviction history.
For a function profile use profile_autoplay.py replay with --cprofile and
either --prediction-reuse or --no-prediction-reuse. Do not interpret profiler
timings as device performance. The new default is included in saved planner
configuration and the prediction_reuse decision metric.
Decision metrics now include formation_roots, formation_tracks,
formation_bound_rejections, script_rebuilds, script_computed_frames,
script_reused_frames and independent_clearance_hits. These count actual work
or reuse, rather than speculative enemies destroyed.
Replay UI
replay_ui.py now labels the recorded prediction mode separately from the
current replay analysis mode. New .planner.json sidecars record the
effective prediction_reuse setting, including a setting restored from the
starting controller checkpoint. For recordings made before this metadata was
added, replay checks a bounded prefix of the decision trace for the setting.
If it was not recorded, the caption says unknown (not recorded) and identifies
the analysis mode as the current default; it does not invent a historical mode.
The Replay prediction selector offers Recorded, Optimized and
Reference. Changing it clears replay caches and reconstructs the current
frame, so retained plans and counters cannot leak between modes. Rewind/reset
preserve the selected override. The command-line equivalents are
--prediction-reuse and --no-prediction-reuse with the maneuver planner.
The GUI loads the decision sidecar independently of the navigation checkbox. The recorded and analysis maneuver summaries each show formation roots/tracks/ rejections, computed/reused script frames, script rebuilds and reused independent collision checks. Recorded rows are accepted only at matching canonical frame and VBL. Missing historical values remain unavailable rather than becoming invented zeroes. Full sidecar indexing is confined to the GUI.
Swarm / formation paths (analysis) is optional and off by default. It draws cached root anchor trajectories in cyan, delayed follower trajectories in dashed purple, and independent scripted swarm trajectories in gold. Root and script paths have time markers every 16 frames. The paths use the default camera forecast, respect selected-object filtering and are explicitly labeled as replay analysis. This overlay reads existing tracks; it never populates prediction caches or claims to show recorded cache internals. Script paths are also available in reference/legacy analysis when cached; shared leader/follower tracks require optimized maneuver prediction. Empty labels distinguish inactive planning, selection filtering, offscreen paths and missing cached predictions. Not every visible enemy necessarily has a cached path. Script displacement dictionaries are isolated in replay snapshots so backward seeks retain the correct frame. Existing navigation and AVI overlays retain their own meanings.
The initial overlay missed independent scripted swarms: frame 1600 of
prediction-reuse-corridor-0906-02 contains 12 swarm enemies and seven cached
script paths but no shared leader/follower tracks. It previously showed an empty
label despite those predictions being available.
The fix passes 52 replay/prediction/recording tests. A real Tk replay check shows
seven gold paths at frame 1600, preserves those paths after stepping forward and
backward, and shows 16 paths spanning all three types at frame 1682. The updated
checkbox fits within the toolbar.
Replay validation: the full Python suite passed 808 tests; 87 focused
tests passed after the final legacy/empty-sidecar compatibility checks. A real
Tk-window smoke test on prediction-reuse-corridor-0906-02 at frame 1682 verified
recorded counters with navigation disabled, 31 formation-overlay canvas items,
reference/recorded mode switches at the same frame, and toolbar widget bounds.
Validation
The full Python suite passes 800 tests. Differential tests cover interpreter
commands and integer overflow, delayed chains, fast offscreen approaches, live
collision extents/link changes, camera conversion, replay-safe immutability,
state/identity/frame invalidation and unknown branches. Clearance tests compare
cached/reference results, preserve aimed-shot checks and require camera offsets
to distinguish otherwise identical queries. One pre-existing protocol test
assumed --stop-file was the last CLI argument; it now checks the flag's value
independently of the already-existing planner option.
The initial implementation matched all 1,755 corridor decisions but showed essentially no net speedup. Profiling exposed repeated root-cache growth while requesting consecutive horizons. Requesting the longest horizon first improved that intermediate comparison to 1.08x. Final measurements below also include reuse of independent clearance checks.
Final sequential comparisons (prediction-isolated-*-0906-05, Python 3.13.7,
same source/assets/fire/terrain/controller state, no profiler or live emulator
workload during measurement, fresh Python process per variant):
| Recorded observations | Frames | Reference mean | Shared mean | Reference/shared p95 |
|---|---|---|---|---|
cruise-spans-0905-03, corridor 801..2555 |
1,755 | 21.989 ms | 18.331 ms | 55.139 / 47.957 ms |
shell-core-spans-0905-06, boss 7633..9030 |
1,398 | 15.709 ms | 15.567 ms | 30.909 / 29.253 ms |
level5-barrier-clearance-0830-118, 90160..90459 |
300 | 16.773 ms | 16.012 ms | 18.263 / 17.466 ms |
All 3,453 selected actions and validation outcomes match. These measurements include observation reconstruction. Corridor mean time is 16.6% lower (1.200x throughput); its formation-heavy 2400..2555 ending falls from 21.538 to 15.778 ms, 26.7% lower. Mean corridor policy time falls from 5.599 to 3.546 ms. Independent group queries fall from 675 to 447 per decision; 45 clearance queries per decision are reused on average. About 106 script frames per decision reuse an existing suffix. These counters explain which repeated work was actually removed.
Boss average differences are small: treat its roughly 1% mean difference as effectively unchanged. Level-5 mean time is 4.5% lower in this short sample. This does not solve the remaining maneuver-validation/ship-motion cost or establish phone/WASM performance.
Earlier prediction-final-*-0906-04 reports shared an interpreter. Examining a
boss spike exposed different navigation LRU histories between variants. Those
timings are superseded by the isolated runs above; navigation row-build/hit
counts now match between variants on both corridor and boss observations.
Reports with fingerprints and raw timing streams:
Closed-loop retests
Both use maneuver/spans, canonical frames, no shield cheat or fast-forward, paired original/sprite AVI, decision traces and periodic checkpoints.
- prediction-reuse-corridor-0906-02.x2events:
original frame-800 checkpoint through the shop at 2556. One 8-point hit
at 1940, zero lives lost, final shield 31, three lives, nine pickups,
eight wall shooters destroyed and 9,310 points gained. These outcomes match
the accepted
cruise-spans-0905-03baseline exactly. - prediction-reuse-boss-0906-01.x2events: original frame-7632 checkpoint through the shop at 8477. Boss objects are present at 7633 and gone by the final recorded frame 8476. Zero damage, zero lives lost, shield 39, three lives, and 17 pickups. The historical boss recording predates the final core-rejoin tactic adjustment; its later shop frame is not evidence of a prediction-optimization gameplay improvement.
The first corridor attempt (prediction-reuse-corridor-0906-01) used the copied
start.sav and stopped on a stale shop indication at 801. It is excluded from
validation. The accepted retest uses the original checkpoint and its emulator
metadata sidecar. Hatari is paused with automation input released.