Autopilot · write-up
Architecture experiments (2026-09-06)
These changes optimize gameplay work rather than require identical joystick decisions. Acceptance considers shield/lives lost, cash and useful pickups, necessary kills, progression/stalls, and mean/tail processing time. Fixed-recording replay measures work and decisions; it cannot measure the resulting gameplay.
The previous implementation remains selectable with --architecture none.
The five independent PlannerConfig switches are now enabled by default in the
maneuver planner, with tactical acceptance and mission reuse restricted to the
Level-1 core attack/refuge encounter after the travel experiments failed. all enables
all five; a comma-separated list enables exactly those named features. These
options work in autoplay.py, record_run.py, validate_campaign.py run,
replay_ui.py, and profile_autoplay.py replay. Checkpoints and recording .planner.json sidecars
preserve the effective selection. Replay restores it and displays the selected
features and work counters separately for recorded and current analysis data.
1. Retained validation (retained_validation)
autopilot_v2/certificates.py carries one observation's bounded certificates.
An exact first-step ship/camera/after-hull match permits reusing the deterministic
motion suffix. The first old collision hull always comes from fresh observation;
animation can change it without changing the later predicted motion. Speed,
camera limits/intent, discontinuous frames and namespace changes invalidate reuse.
Wall results require the same terrain revision, level, wall mask and player mask. Only queried edges survive into the next certificate. Independent-body results include exact current nearby group geometry, player hull, camera correction and incoming clearance. Distant groups receive cheap live bound checks before their members could enter the certificate key. New/moved hazards therefore cannot use an unrelated safety verdict. Homing, aimed shots, shell release histories, pickups and destruction-dependent checks continue against the current observation. No enemy is assumed dead because it was expected to die in a previous frame.
This is deliberately narrower than blindly executing several old inputs between replans. It reuses proven components while retaining immediate live safety checks.
2. Tactical candidates (tactical_candidates)
autopilot_v2/tactics.py proposes ordered waiting locations. Safe cash excursions and escape
directions can be accepted immediately rather than scoring every acceptable
alternative. The existing bounded repair search remains the fallback. A candidate
must complete its horizon, meet the safety reserve, and rank at least as well as
the nominal path; budget exhaustion is not success. This prevents a coin from
overriding a better paced continuation. Boss attack-and-retreat and refuge
constraints remain in force. Travel keeps complete candidate ranking: greedy
corridor choices lost a life in two closed-loop attempts. The ordered proposal
generator retains its corridor template for future work, but that path is not
enabled by the production gate.
3. Formation bounds (formation_broadphase)
Before allocating member rectangles, conservative eight-frame group bounds cover every anchor (including the preceding frame), live collision offsets and required eye margins. Only the active planning horizon is requested; ordinary travel does not pay for the longer homing horizon. Camera-dependent bounds transform the query. A far group can be rejected in one check. Nearby groups still use exact individual members, preserving gaps through formations. Unsupported motions and incomplete trajectory blocks use the ordinary predictor. This is a rejecting bound, not a solid formation obstacle.
4. Useful attacks (attack_objectives)
At most two nearby ordinary forward-fire targets propose firing stations using their predicted position at bullet arrival. Only otherwise safe shortlisted paths pay for detailed consumable-shot simulation. Positive modeled damage can select an attack even if no enemy was predicted to hit the ship. Previously combat was mainly invoked to resolve a collision, leaving safe attack opportunities unvalued.
This first implementation covers scripted swarms and ordinary wall shooters. Existing specialized bosses, gates and progression barriers keep their attack missions. Optional attacks yield to nearby pickups and cannot override a refuge or leave the preferred corridor lane. Predicted destruction never opens a wall: the current persistent map remains authoritative until the opening is observed.
5. Persistent missions (persistent_missions)
autopilot_v2/policy_missions.py retains Level-1 core-lane
proposals for at most three intervening frames. A small explicit hook after common
observation/model/wall-recovery preparation bypasses the remaining legacy target
selection and nine-action proposal scoring. Live maneuver validation still runs.
New/removed/changed hazards, pickup identities, equipment, terrain, scroll-region changes, reached waypoints, blocked movement and unsuccessful prior validation force fresh selection. The core lane refreshes its camera-relative goal each observation. Travel and specialized gate/level state machines are not cached. Generic travel reuse kept full shields but missed the corridor rejoin and failed progression; that experiment was rejected. This separates the stationary mission without skipping route updates and encounter transitions elsewhere.
Reproducing the independent measurements
python xenon_tools/compare_architecture.py RECORDING.x2events --from-frame FIRST --to-frame LAST --repeats 2 --output xenon_tools/run_logs/NEW-DIRECTORY
The tool measures baseline, each feature alone, and the combined configuration. Each gets a fresh Python process, the same warm-up, spans navigation, shared predictions, assets, fire metadata and available controller/terrain sidecars. The second pass reverses order. Reports include fingerprints, mean/p95/max, work counters, changed actions and unverified frames. Timing runs are sequential; the emulator remains paused while offline measurements run.
The 0906-01 corridor/boss reports are exploratory and predate removal of
unnecessary distant-member hashing in retained validation. The 0906-02 timing
reports predate the gameplay-driven restrictions; final results use 0906-03.
The comparison tool now warns when repeated baseline mean
time differs by more than 15%.
Final measurements (narrowed implementation)
These 0906-03 measurements average two complete passes in opposite order,
using fresh processes and the same recorded observations. Baseline mean drift
was below 1% in every sample. Values include reconstruction and exclude Tk,
AVI, sockets and the emulator. Each feature is enabled separately against none.
| Configuration | Corridor, 1,755 frames (ms) | Boss, 844 frames (ms) | Level-5 barrier, 300 frames (ms) |
|---|---|---|---|
| Baseline | 18.859 | 15.012 | 15.954 |
| Retained validation | 18.361 | 14.827 | 15.954 |
| Tactical candidates | 18.827 | 15.071 | 15.996 |
| Formation bounds | 16.885 | 14.097 | 16.663 |
| Attack objectives | 18.871 | 14.989 | 15.949 |
| Persistent missions | 18.761 | 14.648 | 16.052 |
| Combined | 16.265 | 12.697 | 16.725 |
The final combination reduces mean processing time by 13.8% in the corridor and 15.4% at the boss, but is 4.8% slower in the Level-5 barrier sample. Average per-pass p95 is 49.994 -> 50.150 ms in the corridor, 27.915 -> 23.978 ms at the boss, and 17.182 -> 20.942 ms at the barrier. Corridor tail latency has not improved. These measurements do not establish a phone/WASM frame budget.
Independent effects:
- Formation bounds provide the clearest standalone gain: 10.5% lower mean in the corridor and 6.1% at the boss. Hazard slices fall from 940.06 to 269.08 and from 512.09 to 149.32 per decision, respectively. The Level-5 sample instead pays 4.4% overhead; this optimization is not beneficial everywhere.
- Retained validation removes about 38% of simulated corridor transitions (103.15 -> 63.48) and 43% at the boss (114.13 -> 64.68). Its standalone time reduction is only 2.6% and 1.2%, respectively, and was not consistent across earlier experiments. Certificate bookkeeping consumes much of the saving.
- Tactical candidates are inactive in the corridor and barrier. At the boss they reduce rollouts from 5.61 to 5.34, but show no standalone mean-time win.
- Attack objectives select three optional attacks in the corridor sample, with four changed actions and essentially unchanged processing time. They are a gameplay feature, not a demonstrated speed or collection improvement.
- Persistent missions are inactive in the corridor and barrier. At the boss they skip about 0.63 legacy selection passes per observation and reduce mean time by 2.4%; treat this small timing difference cautiously.
Effects interact, so the individual percentages cannot be added. Retained validation and formation bounds preserve every selected action in all three samples. Combined selected actions differ on four corridor and 77 boss frames; all barrier actions match. The number of unverified observations is unchanged: 133 corridor, 17 boss and all 300 barrier observations. In particular, the barrier sample is a cost comparison, not a successful safety validation.
Final reports contain source/asset fingerprints, raw per-frame timings, work counters, selected-action comparisons and both passes:
Exploratory measurements (before gameplay-driven restrictions)
Python 3.13.7, same desktop/recordings/assets, no profiler or live gameplay running during timing. Values are mean milliseconds per recorded observation, including reconstruction. Each feature row enables only that feature; the last enables all.
| Configuration | Corridor (1,755 frames) | Boss (844 frames) | Level-5 barrier (300 frames) |
|---|---|---|---|
| Baseline | 18.613 | 15.025 | 16.352 |
| Retained validation | 19.028 | 14.970 | 16.399 |
| Tactical candidates | 18.150 | 14.664 | 16.447 |
| Formation broad phase | 16.996 | 14.144 | 16.523 |
| Attack objectives | 18.833 | 14.923 | 16.124 |
| Persistent missions | 18.195 | 14.510 | 16.070 |
| All five | 15.007 | 12.190 | 16.535 |
Corridor uses the complete second, reversed-order pass. The first pass is excluded: its baseline was 32.773 ms versus 18.613 ms on repetition, despite identical actions, transitions, rollouts, hazard slices and combat steps. Host timing drift makes that pass unsuitable for estimating speedup. Its later combined run (15.019 ms) agrees with the second (15.007 ms). Boss and Level-5 values average both complete passes. Raw runs, including exclusions, remain available in the reports; no individual frames were removed.
The exploratory combined mean time is 19.4% lower in the corridor and 18.9% lower at the boss. That unrestricted combination failed the corridor gameplay gate below; these numbers are not the final production speedup; use the final table above. Corridor p95 is essentially unchanged (49.489 -> 49.159 ms); boss p95 falls from 28.272 to 22.763 ms. The Level-5 sample is essentially unchanged/slightly slower (1.1% higher mean, p95 16.981 -> 18.200 ms). Small single-feature differences of roughly 1–3% should not be treated as established wins. In particular, retained validation has no demonstrated standalone speedup, despite removing work. These are desktop Python results, not phone/WASM performance claims.
The corridor counters explain the actual reduction:
- Retained validation: newly simulated transitions 103.15 -> 63.48 per decision; approximately 49 motion frames and 53 independent-body checks reused. Certificate construction/lookup offsets those standalone savings in Python.
- Tactical candidates: rollouts 5.28 -> 4.63 (12.3% fewer).
- Formation bounds: allocated hazard slices 940.06 -> 269.08 (71.4% fewer).
- Persistent missions: 649 legacy tactical-selection passes skipped; mean policy time 3.54 -> 2.91 ms, though the complete decision improves much less.
- Attack objectives: three attack missions selected on these fixed observations; four action differences. This is a gameplay opportunity, not a speed optimization.
- Combined: 54.36 new transitions, 264.66 hazard slices and 4.70 rollouts per decision. Effects interact; individual timing changes cannot simply be added.
Retained validation and formation bounds each preserve every selected action in these three samples. The tactical and mission changes intentionally differ. One corridor observation (1826) becomes unverified with persistent missions: a formation-leader contact is predicted 30 frames ahead. The next observation must refresh mission selection. This is tracked for closed-loop validation, not silently counted as safe. All Level-5 selected actions match the baseline.
Reports with complete fingerprints, counters, raw timings and per-pass results:
Closed-loop gates
Each run uses the original frame-800 corridor or frame-7632 boss checkpoint, spans navigation, no shield cheat, no fast-forward, paired original/sprite AVI, decision traces and periodic checkpoints. Failed runs and their archived source remain available for inspection.
- Unrestricted combination rejected:
architecture-corridor-0906-01took damage at 1432 and 1803–1807, lost a life at 1824, and was stopped. - Greedy travel rejected:
architecture-tactical-0906-01lost a life at 1861. Adding a nominal-path quality floor prevented the early cash mistake butarchitecture-tactical-0906-02still lost a life at 1842. Greedy acceptance is consequently restricted to the explicit core attack/refuge phase. - Generic mission retention rejected:
architecture-missions-0906-01kept shield 39 and three lives but failed to rejoin the corridor route. It was still attempting that route after frame 3700, well beyond the baseline shop at 2556, and was stopped. Travel mission reuse is disabled by the production gate. - Attack-only corridor passed the principal gates:
architecture-attacks-0906-01reached the shop with 8 damage at 1940, no lost lives, nine pickup detections and eight destroyed wall shooters. Score gain was 9,210 versus the baseline's 9,310; this trial does not demonstrate improved enemy destruction or cash collection. - Narrowed combination, corridor:
architecture-corridor-0906-02reaches the shop with the same 8 damage at 1940, no lost lives, nine pickup detections, eight destroyed wall shooters and 9,210 score gain. - Narrowed combination, boss:
architecture-boss-0906-01completes the encounter with zero damage, zero lost lives and 17 pickup detections, matching the baseline's principal outcome metrics.
Pickup counts use the existing BehaviorCatalog disappearance-near-player detection. They are not newly invented exact collection events. These checkpoint tests do not establish a damage-free five-level campaign or phone performance.
The final corridor run and its baseline both reach the shop at frame 2556 with 750 cash. The final boss run and its baseline both reach the shop at 8477 with 3,400 cash. Shop balances were read from the binary recordings; their last records include two cleanup frames after shop detection. The corridor score is 100 lower despite matching cash, pickups, damage and observed shooter kills. Optional attack selection therefore has not demonstrated a better outcome.
Accepted recordings:
These live runs used explicit all, before the final removal of unused
prediction-tail work and promotion of the defaults. The final fixed-recording
comparison confirms that the structural reuse/bounds changes preserve actions.
No complete five-level campaign was rerun.
Verification
python -m unittest discover -s xenon_tools -p "test_*.py": 821 tests pass.
Coverage includes live-geometry/camera/terrain invalidation, fresh first-step
collision hulls, conservative interval bounds and unknown tails, attack
shortlisting, mission invalidation, travel restrictions, recording metadata and
replay overrides across rewind/reset. A real Tk smoke run opened the accepted
corridor recording, enabled formation paths with scene overlay disabled, and
switched analysis between none and all through reset. The five recorded
feature names and effective analysis selections were verified. No emulator
rebuild was needed for these Python-only changes.