Autopilot · write-up
Maneuver autopilot (2026-09-05)
For subsequent Level-1 shell/eye work and spans replay validation, see AUTOPILOT_SHELL_VALIDATION.MD.
The current maneuver defaults also include the narrowed architecture changes in
AUTOPILOT_ARCHITECTURE_EXPERIMENTS.MD, with
independent measurements and rejected/accepted gameplay runs. Use
--architecture none to retain this document's earlier maneuver implementation.
The experimental new planner is selected with --planner maneuver. The old
controller remains the default (--planner legacy) while full-level gameplay
acceptance is outstanding. Replay uses the recording's selector and defaults to
legacy for older recordings without metadata. ReactiveController itself still
means the old controller, including its original scorer and beam search.
This is a Python architecture implementation, not a browser/WASM port or a claim
that the full campaign is solved. See the validation evidence below before
interpreting verified as a gameplay guarantee.
Navigation update: the maneuver controller now uses exact cached wall-row geometry and flat-index A* over the resident level maps. The default retains reference route costs; a compressed corridor variant remains experimental. See AUTOPILOT_NAVIGATION.MD for the implementation, isolated before/after measurements and comparison switch. Historical measurements below predate this optimization unless explicitly identified otherwise.
Runtime path
- Update the existing canonical world model and persistent level map.
- Obtain an encounter/route/firing-station mission from the policy adapter. Level-specific corridor, gate, boss, retreat and shop policies remain available. Expensive wall-station proposals are retained across small changes of position and animation rather than recomputing full-map A* on every decision.
- Revalidate the previous maneuver's remaining actions against the current world, and extend it to the full horizon. Otherwise validate the mission command plus a concrete continuation through the exact ship/steering/camera model.
- If necessary, try nearby cash excursions, alternate holding stations, a bounded directional maneuver set, then repairs timed shortly before predicted contact. Holding stations let a swarm pass before returning to the firing lane. Horizontal control accounts for steering coast, including the speed upgrades bought at shops.
- Resolve a predicted contact with conservative shot/destruction simulation when the contacted object has a supported damage/movement model. Safe paths do not pay for an unconditional bullet-by-object simulation.
- Execute one selected input and observe again. The live driver's old unvalidated
map_discovery_retreatreplacement is disabled for this controller.
There is no call to the legacy horizon scorer or formation beam. The adapter's
nine cheap _score calls only propose immediate legal mission directions; they
perform no future hazard/comfort rollout. Final choice uses lexicographic safety,
reserve, pickups, credited damage, mission progress and maneuver simplicity.
Predicted shield contact cannot be bought with cash or damage reward.
Every candidate is one actual action sequence. Collision rejection is lazy: a
failed candidate stops at its first contact rather than constructing the rest of
its trajectory. Incomplete work is never marked safe. If no complete safe
continuation is found, the planner explicitly reports no_verified_escape and
chooses the examined prefix with the latest contact. Complete budget exhaustion
falls back to the immediate legal mission proposal, with verified: false.
Shared predictions and event handling
For the subsequent shared follower tracks, rolling script states and within-decision clearance reuse, see AUTOPILOT_PREDICTION_REUSE.MD.
PredictionService builds independent hazard geometry once per future frame for
all maneuvers. Existing exact leader/follower and scripted-path models are reused.
Formation chains and swarms sharing a script are queried as groups; a broad group
rejection avoids visiting members. If a group overlaps the query, individual
rectangles are checked, preserving gaps between followers. Grouping by script is
only a geometric broad phase, not a claim that unrelated waves share a leader.
Camera-dependent groups transform the player query. Homing motion advances against the actual candidate player path. Aimed shots capture the ship position at their predicted allocation frame. The next known launcher/periodic/barrier event is predicted using captured state; later RNG resets are not invented as deterministic events. Dormant worm/emitter and composite-chain uncertainty remains conservative.
Pending bottom-entry scripts reserve a small entry lane around the actual scroll trigger. This anticipates an enemy that could otherwise first appear already overlapping the old ship rectangle. It is an area tactic, not an expansion of every pending swarm member. It needs the captured script/initial-position fields; older reduced text traces may not contain them.
The ordinary object check uses the post-movement player interaction rectangle. Homing retains the separate old-hull/new-anchor phases. Wall legality uses the fixed navigation mask and the persistent map, including the existing monotonic escape rule for an already obstructed start. All intermediate movement updates are checked; no waypoint-only collision test is used.
Destruction and terrain
Known positive health and the shared damage callback are required to credit a kill. Supported targets currently include independent scripted swarms, exact curved missiles, vertical boss shells and supported wall shooters. Linked followers and unknown/boss damage callbacks are not optimistically removed.
Captured ordinary and cannon bullets are finite resources. Forward/Side Shot future allocations use the driver's fire forecast; unknown future cannon/laser shots are not invented. A moved bullet must hit both predicted object phases, with no competing concrete blocker, to credit damage. Animated swarm/shell hits also require an inset to reduce optimistic credit at changing sprite edges. Ambiguous contacts consume the shot without credit. One shot cannot kill multiple targets. The destruction update and an additional final-contact update remain collidable. Already emitted shots remain hazards after their source dies.
Ordinary bullets test object bounds rather than the wall-tile raster, so an object mounted in solid tiles can still be shot. Tiles remain authoritative for ship movement. The Level-1 shell tactic also respects its Y-only proximity burst: moving sideways does not make crossing that trigger safe. A shell must be cleared or the ship must remain below its release band.
Resident maps for all five levels, the Level-2 navigation field and the existing destructible-group metadata are reused. The map is never shortened to the local viewport for strategic routing. Live RAM synchronization and subsequent observed tile updates remain authoritative about open gates. Shooting a destructible controller does not speculatively open its map cells in a candidate rollout. The controller can approach/fire using the level's gate mission, then route through the opening after it is observed.
Work limits and migration boundary
autopilot_v2/types.py:PlannerConfig owns deterministic limits:
| Setting | Default |
|---|---|
| Ordinary horizon | 56 canonical updates |
| Homing horizon | 80 canonical updates |
| Maximum maneuver validations | 32 per decision |
| Maximum unique trajectory-prefix transitions | 2,048 per decision |
| Retained-plan refresh | 8 updates |
| Optional wall-station builds | 2 per decision |
| Station proposal refresh | 64 updates or relevant event |
| Station proposal cache capacity | 64 entries |
| Navigation backend | grid (exact accelerated geometry/A*) |
Terrain revision, identity namespace, target identity/procedure, equipment, coarse player/target/camera area, and blocked approach/staging floor invalidate the relevant proposal. Moving Level-1 shell proposals refresh within eight updates and on oscillation reversals. A stale proposal can cost progress until refresh; it cannot bypass the final current-world movement check. The cache uses copy-on-write at decision boundaries so replay snapshots are not mutated by later queries.
These limits bound maneuver search work, not every inherited policy operation or wall-clock latency. The navigation update reduces grid construction and route acquisition using a cheaper map representation; cold geometry preparation and other mission operations still need measured latency budgets before calling this a phone-ready runtime. No C code has been changed.
| Module | Responsibility |
|---|---|
autopilot_controller.py |
Single selector and recorded configuration reader |
autopilot_v2/controller.py |
Lifecycle, checkpoints, selected input, metrics |
autopilot_v2/adapter.py |
Explicit dependency on old encounter policies/models; overrides old search |
autopilot_v2/missions.py |
Persistent, bounded firing-station proposals |
autopilot_v2/navigation.py |
Exact shared wall rows/flat A*; experimental corridor routing |
autopilot_v2/planner.py |
Continuations, holding stations, cash detours, retained validation and repair |
autopilot_v2/prediction.py |
Shared motion, grouped contact queries and candidate-dependent events |
autopilot_v2/combat.py |
Consumable shots and conservative destruction |
autopilot_v2/types.py, geometry.py |
State/budgets and inclusive geometry |
The new planner is separate, but still deliberately depends on the old module's
proven models and area tactics. Removing ReactiveController._score and its beam
later does not require changing the new search/combat modules. Removing all of
autoplay.py would first require extracting the shared driver, models and tactics.
Do not mistake the adapter for a fully independent portable runtime.
Comparison and replay
From the repository root, with an already running Hatari on port 6903:
python xenon_tools/validate_campaign.py run new-level1 --connect --planner maneuver --stop-at-level 2
python xenon_tools/validate_campaign.py run old-level1 --connect --planner legacy --stop-at-level 2
Use --resume path/to/checkpoint.sav for the same later-level starting state.
Choose only one automatic stopping condition: --frames, --stop-at-shop, or
--stop-at-level. Each campaign archives source, start state, RAM terrain seeds,
trace, checkpoints and both AVIs. No shield cheat is enabled by these commands.
For identical-observation timing and action comparison:
python xenon_tools/compare_autopilot.py path/to/recording.x2events --output xenon_tools/run_logs/new-comparison
python xenon_tools/replay_ui.py path/to/recording.x2events --planner maneuver
python xenon_tools/validate_campaign.py profile campaign-name --planner legacy --cprofile
The comparison output directory must be new. Both planners run sequentially with the same observations, terrain/fire metadata, start-controller state and warm-up. Outputs include raw timings, action changes, stage frequencies, Python source and JSON asset fingerprints, recording hash and environment. Old recordings lacking fire or RAM metadata keep that limitation; the tool does not fabricate metadata. Replay measures the observed world and cannot prove outcomes of changed actions.
*.planner.json records the planner and fire configuration for direct recordings.
Campaign manifest.json also contains them. Replay overrides are explicit and
timing filenames are separated by planner. Checkpoints store planner config and
common semantic policy state; trajectory and mission-query caches are discarded
on restore. Old checkpoints remain readable.
Replay's title and persistent header show Recorded autopilot and Replay
analysis, marking a different analysis planner as an override. Missing recording
metadata is displayed as unknown while analysis defaults to legacy. Standalone
JSONL decision traces can supply their planner from the first row without a sidecar.
Keep the .planner.json or campaign manifest with a binary .x2events recording;
the binary event format itself does not contain this Python planner metadata.
The shared inspector shows maneuver stage, finite-horizon verification, policy and planning times, work/cache counts, predicted kills/pickups and contact cause. JSONL traces retain a separate RECORDED MANEUVER block; binary event recordings without recorded diagnostic rows show that those diagnostics are unavailable. Recomputed metrics are labeled as analysis and must not be read as original-run timing. The old candidate table is labeled MISSION PROPOSALS for this planner and omits unused legacy risk/exposure values.
Trace planner_metrics separates policy and planning times, trajectory work,
group/member queries, optional combat work, station builds/cache hits/deferred
work, stage, contact cause and verified. nearest_hazard is now a bounded
rectangle clearance (24 pixels maximum), not the old global centre distance.
Candidate scores in the inspector are immediate mission proposals, not final
survival rankings. verified means the configured finite model/horizon found no
contact; future RNG, incomplete observations and unmodeled callbacks limit it.
Validation evidence
The focused tests cover exact ship/camera trajectories and pending input, inclusive endpoints, old-hull homing contact, lazy rejection/work limits, formation gaps, camera coupling, pickups, bullet ownership/ambiguity/final death contact, emitted-shot survival, pending bottom entries, persistent maps and gate updates on all levels, cache invalidation, checkpoints and legacy selection. The legacy scorer and beam are patched to raise in a test of the new controller.
The first recorded 1,200-update opening test (maneuver-architecture-0905-smoke1)
had zero shield damage, zero lives lost, 10 pickups collected and 7 wall shooters
destroyed. That test preceded the later combat-on-conflict and mission-query
improvements. It is an opening regression, not level completion.
The broad regression suite passed 648 tests; test_campaign added 15 passing
tests. This includes 36 focused new-planner tests. Commands from xenon_tools:
python -m unittest test_autoplay test_autopilot_v2 test_game_model test_level_tile_maps test_level2_navigation_field test_replay_ui test_profile_autoplay
python -m unittest test_campaign
Live checkpoint regressions (no shield cheat):
Run under xenon_tools/run_logs/*.validation |
Updates | Shield damage | Lives lost | Result |
|---|---|---|---|---|
maneuver-architecture-0905-smoke1 |
1,200 | 0 | 0 | Opening: 10 pickups, 7 wall shooters destroyed |
maneuver-architecture-0905-spawn-fix |
300 | 0 | 0 | Crossed the previously failing bottom-entry trigger; 2 pickups |
maneuver-architecture-0905-timed-repair |
700 | 20 | 0 | Boss checkpoint: shield 39 to 19; 1 pickup, 2 wall shooters destroyed |
maneuver-architecture-0905-boss-exit |
1,564 | 0 | 0 | Continued from the preceding save: 7 pickups; stopped before shop because progression cycled |
Earlier boss iterations lost lives, and an earlier longer run stopped on a shell
station cache-contract KeyError. Publishing cached proposals into the inherited
policy index fixed that exception and has a regression test. The bottom-entry
damage failure led to the pending-entry guard and its successful retest. The
latest boss repair still took two swarm hits and a directional-projectile hit;
full-level low-damage acceptance remains open. These are separate checkpoint
tests, not a single successful full-campaign run.
The final continuation preserved shield 19 through frame 8510, but repeatedly switched between route advance and shell clearance near the remaining edge shells. It was stopped through the normal stop-file mechanism; both AVIs and the checkpoint were finalized. The global Y-only shell-release constraint appears to conflict with route/attack availability here. This is an unresolved progression failure, not a successful boss exit despite the run's name. A future fix needs an explicit choice between acquiring a reachable shell firing station and handling its released projectiles during a bypass; indefinitely waiting is not acceptable.
Final identical-observation timing comparison, with source fingerprints verified unchanged during both sequential runs:
- Recording:
maneuver-architecture-0905-spawn-fix.validation/maneuver-architecture-0905-spawn-fix.x2events. - Report: comparison.json.
- 300 updates, game frames 4914–5213, including reconstruction and cold setup; actual campaign RAM/fire metadata and the same starting controller state.
| Processing time | Legacy | Maneuver |
|---|---|---|
| Mean | 55.69 ms | 28.49 ms |
| Median | 20.07 ms | 10.83 ms |
| 95th percentile | 250.03 ms | 123.32 ms |
| Maximum | 1,164.00 ms | 703.78 ms |
Mean improvement was 1.95× on this sample. Actions differed on 223/300 updates, so this comparison establishes processing cost, not counterfactual survival. The new controller averaged 18.13 ms in inherited policy and 8.40 ms in maneuver planning. It averaged 97 unique trajectory transitions and 65 member queries per decision; maxima were 408 transitions and 32 rollouts. Independent predictions and group rejection avoid the original repeated global scoring, but inherited policy/navigation is now the larger mean cost in this sample.
Outstanding acceptance work
Keep maneuver opt-in until comparable live runs demonstrate level completion,
shield loss and cash collection across all five levels. Old later-level recordings
were exercised as smoke inputs, but missing RAM/fire metadata in some recordings
prevents treating those results as gameplay evidence. In particular, the older
Level-5 barrier trace without a starting RAM seed produced wall rejections.
The navigation work following this initial validation is documented separately in AUTOPILOT_NAVIGATION.MD. The maneuver search budget does not bound all inherited mission operations. Boss swarm holding/escape choices still need further live tuning before increasing aggression. Browser integration and target-phone timing remain unimplemented and unmeasured.