Xenon 2

Autopilot · write-up

Maneuver autopilot (2026-09-05)

xenondoc/AUTOPILOT_MANEUVER.MD · 19 KB · updated 2026-09-17

For subsequent Level-1 shell/eye work and spans replay validation, see AUTOPILOT_SHELL_VALIDATION.MD.

The current maneuver defaults also include the narrowed architecture changes in AUTOPILOT_ARCHITECTURE_EXPERIMENTS.MD, with independent measurements and rejected/accepted gameplay runs. Use --architecture none to retain this document's earlier maneuver implementation.

The experimental new planner is selected with --planner maneuver. The old controller remains the default (--planner legacy) while full-level gameplay acceptance is outstanding. Replay uses the recording's selector and defaults to legacy for older recordings without metadata. ReactiveController itself still means the old controller, including its original scorer and beam search.

This is a Python architecture implementation, not a browser/WASM port or a claim that the full campaign is solved. See the validation evidence below before interpreting verified as a gameplay guarantee.

Navigation update: the maneuver controller now uses exact cached wall-row geometry and flat-index A* over the resident level maps. The default retains reference route costs; a compressed corridor variant remains experimental. See AUTOPILOT_NAVIGATION.MD for the implementation, isolated before/after measurements and comparison switch. Historical measurements below predate this optimization unless explicitly identified otherwise.

Runtime path

  1. Update the existing canonical world model and persistent level map.
  2. Obtain an encounter/route/firing-station mission from the policy adapter. Level-specific corridor, gate, boss, retreat and shop policies remain available. Expensive wall-station proposals are retained across small changes of position and animation rather than recomputing full-map A* on every decision.
  3. Revalidate the previous maneuver's remaining actions against the current world, and extend it to the full horizon. Otherwise validate the mission command plus a concrete continuation through the exact ship/steering/camera model.
  4. If necessary, try nearby cash excursions, alternate holding stations, a bounded directional maneuver set, then repairs timed shortly before predicted contact. Holding stations let a swarm pass before returning to the firing lane. Horizontal control accounts for steering coast, including the speed upgrades bought at shops.
  5. Resolve a predicted contact with conservative shot/destruction simulation when the contacted object has a supported damage/movement model. Safe paths do not pay for an unconditional bullet-by-object simulation.
  6. Execute one selected input and observe again. The live driver's old unvalidated map_discovery_retreat replacement is disabled for this controller.

There is no call to the legacy horizon scorer or formation beam. The adapter's nine cheap _score calls only propose immediate legal mission directions; they perform no future hazard/comfort rollout. Final choice uses lexicographic safety, reserve, pickups, credited damage, mission progress and maneuver simplicity. Predicted shield contact cannot be bought with cash or damage reward.

Every candidate is one actual action sequence. Collision rejection is lazy: a failed candidate stops at its first contact rather than constructing the rest of its trajectory. Incomplete work is never marked safe. If no complete safe continuation is found, the planner explicitly reports no_verified_escape and chooses the examined prefix with the latest contact. Complete budget exhaustion falls back to the immediate legal mission proposal, with verified: false.

Shared predictions and event handling

For the subsequent shared follower tracks, rolling script states and within-decision clearance reuse, see AUTOPILOT_PREDICTION_REUSE.MD.

PredictionService builds independent hazard geometry once per future frame for all maneuvers. Existing exact leader/follower and scripted-path models are reused. Formation chains and swarms sharing a script are queried as groups; a broad group rejection avoids visiting members. If a group overlaps the query, individual rectangles are checked, preserving gaps between followers. Grouping by script is only a geometric broad phase, not a claim that unrelated waves share a leader.

Camera-dependent groups transform the player query. Homing motion advances against the actual candidate player path. Aimed shots capture the ship position at their predicted allocation frame. The next known launcher/periodic/barrier event is predicted using captured state; later RNG resets are not invented as deterministic events. Dormant worm/emitter and composite-chain uncertainty remains conservative.

Pending bottom-entry scripts reserve a small entry lane around the actual scroll trigger. This anticipates an enemy that could otherwise first appear already overlapping the old ship rectangle. It is an area tactic, not an expansion of every pending swarm member. It needs the captured script/initial-position fields; older reduced text traces may not contain them.

The ordinary object check uses the post-movement player interaction rectangle. Homing retains the separate old-hull/new-anchor phases. Wall legality uses the fixed navigation mask and the persistent map, including the existing monotonic escape rule for an already obstructed start. All intermediate movement updates are checked; no waypoint-only collision test is used.

Destruction and terrain

Known positive health and the shared damage callback are required to credit a kill. Supported targets currently include independent scripted swarms, exact curved missiles, vertical boss shells and supported wall shooters. Linked followers and unknown/boss damage callbacks are not optimistically removed.

Captured ordinary and cannon bullets are finite resources. Forward/Side Shot future allocations use the driver's fire forecast; unknown future cannon/laser shots are not invented. A moved bullet must hit both predicted object phases, with no competing concrete blocker, to credit damage. Animated swarm/shell hits also require an inset to reduce optimistic credit at changing sprite edges. Ambiguous contacts consume the shot without credit. One shot cannot kill multiple targets. The destruction update and an additional final-contact update remain collidable. Already emitted shots remain hazards after their source dies.

Ordinary bullets test object bounds rather than the wall-tile raster, so an object mounted in solid tiles can still be shot. Tiles remain authoritative for ship movement. The Level-1 shell tactic also respects its Y-only proximity burst: moving sideways does not make crossing that trigger safe. A shell must be cleared or the ship must remain below its release band.

Resident maps for all five levels, the Level-2 navigation field and the existing destructible-group metadata are reused. The map is never shortened to the local viewport for strategic routing. Live RAM synchronization and subsequent observed tile updates remain authoritative about open gates. Shooting a destructible controller does not speculatively open its map cells in a candidate rollout. The controller can approach/fire using the level's gate mission, then route through the opening after it is observed.

Work limits and migration boundary

autopilot_v2/types.py:PlannerConfig owns deterministic limits:

Setting Default
Ordinary horizon 56 canonical updates
Homing horizon 80 canonical updates
Maximum maneuver validations 32 per decision
Maximum unique trajectory-prefix transitions 2,048 per decision
Retained-plan refresh 8 updates
Optional wall-station builds 2 per decision
Station proposal refresh 64 updates or relevant event
Station proposal cache capacity 64 entries
Navigation backend grid (exact accelerated geometry/A*)

Terrain revision, identity namespace, target identity/procedure, equipment, coarse player/target/camera area, and blocked approach/staging floor invalidate the relevant proposal. Moving Level-1 shell proposals refresh within eight updates and on oscillation reversals. A stale proposal can cost progress until refresh; it cannot bypass the final current-world movement check. The cache uses copy-on-write at decision boundaries so replay snapshots are not mutated by later queries.

These limits bound maneuver search work, not every inherited policy operation or wall-clock latency. The navigation update reduces grid construction and route acquisition using a cheaper map representation; cold geometry preparation and other mission operations still need measured latency budgets before calling this a phone-ready runtime. No C code has been changed.

Module Responsibility
autopilot_controller.py Single selector and recorded configuration reader
autopilot_v2/controller.py Lifecycle, checkpoints, selected input, metrics
autopilot_v2/adapter.py Explicit dependency on old encounter policies/models; overrides old search
autopilot_v2/missions.py Persistent, bounded firing-station proposals
autopilot_v2/navigation.py Exact shared wall rows/flat A*; experimental corridor routing
autopilot_v2/planner.py Continuations, holding stations, cash detours, retained validation and repair
autopilot_v2/prediction.py Shared motion, grouped contact queries and candidate-dependent events
autopilot_v2/combat.py Consumable shots and conservative destruction
autopilot_v2/types.py, geometry.py State/budgets and inclusive geometry

The new planner is separate, but still deliberately depends on the old module's proven models and area tactics. Removing ReactiveController._score and its beam later does not require changing the new search/combat modules. Removing all of autoplay.py would first require extracting the shared driver, models and tactics. Do not mistake the adapter for a fully independent portable runtime.

Comparison and replay

From the repository root, with an already running Hatari on port 6903:

python xenon_tools/validate_campaign.py run new-level1 --connect --planner maneuver --stop-at-level 2
python xenon_tools/validate_campaign.py run old-level1 --connect --planner legacy --stop-at-level 2

Use --resume path/to/checkpoint.sav for the same later-level starting state. Choose only one automatic stopping condition: --frames, --stop-at-shop, or --stop-at-level. Each campaign archives source, start state, RAM terrain seeds, trace, checkpoints and both AVIs. No shield cheat is enabled by these commands.

For identical-observation timing and action comparison:

python xenon_tools/compare_autopilot.py path/to/recording.x2events --output xenon_tools/run_logs/new-comparison
python xenon_tools/replay_ui.py path/to/recording.x2events --planner maneuver
python xenon_tools/validate_campaign.py profile campaign-name --planner legacy --cprofile

The comparison output directory must be new. Both planners run sequentially with the same observations, terrain/fire metadata, start-controller state and warm-up. Outputs include raw timings, action changes, stage frequencies, Python source and JSON asset fingerprints, recording hash and environment. Old recordings lacking fire or RAM metadata keep that limitation; the tool does not fabricate metadata. Replay measures the observed world and cannot prove outcomes of changed actions.

*.planner.json records the planner and fire configuration for direct recordings. Campaign manifest.json also contains them. Replay overrides are explicit and timing filenames are separated by planner. Checkpoints store planner config and common semantic policy state; trajectory and mission-query caches are discarded on restore. Old checkpoints remain readable.

Replay's title and persistent header show Recorded autopilot and Replay analysis, marking a different analysis planner as an override. Missing recording metadata is displayed as unknown while analysis defaults to legacy. Standalone JSONL decision traces can supply their planner from the first row without a sidecar. Keep the .planner.json or campaign manifest with a binary .x2events recording; the binary event format itself does not contain this Python planner metadata.

The shared inspector shows maneuver stage, finite-horizon verification, policy and planning times, work/cache counts, predicted kills/pickups and contact cause. JSONL traces retain a separate RECORDED MANEUVER block; binary event recordings without recorded diagnostic rows show that those diagnostics are unavailable. Recomputed metrics are labeled as analysis and must not be read as original-run timing. The old candidate table is labeled MISSION PROPOSALS for this planner and omits unused legacy risk/exposure values.

Trace planner_metrics separates policy and planning times, trajectory work, group/member queries, optional combat work, station builds/cache hits/deferred work, stage, contact cause and verified. nearest_hazard is now a bounded rectangle clearance (24 pixels maximum), not the old global centre distance. Candidate scores in the inspector are immediate mission proposals, not final survival rankings. verified means the configured finite model/horizon found no contact; future RNG, incomplete observations and unmodeled callbacks limit it.

Validation evidence

The focused tests cover exact ship/camera trajectories and pending input, inclusive endpoints, old-hull homing contact, lazy rejection/work limits, formation gaps, camera coupling, pickups, bullet ownership/ambiguity/final death contact, emitted-shot survival, pending bottom entries, persistent maps and gate updates on all levels, cache invalidation, checkpoints and legacy selection. The legacy scorer and beam are patched to raise in a test of the new controller.

The first recorded 1,200-update opening test (maneuver-architecture-0905-smoke1) had zero shield damage, zero lives lost, 10 pickups collected and 7 wall shooters destroyed. That test preceded the later combat-on-conflict and mission-query improvements. It is an opening regression, not level completion.

The broad regression suite passed 648 tests; test_campaign added 15 passing tests. This includes 36 focused new-planner tests. Commands from xenon_tools:

python -m unittest test_autoplay test_autopilot_v2 test_game_model test_level_tile_maps test_level2_navigation_field test_replay_ui test_profile_autoplay
python -m unittest test_campaign

Live checkpoint regressions (no shield cheat):

Run under xenon_tools/run_logs/*.validation Updates Shield damage Lives lost Result
maneuver-architecture-0905-smoke1 1,200 0 0 Opening: 10 pickups, 7 wall shooters destroyed
maneuver-architecture-0905-spawn-fix 300 0 0 Crossed the previously failing bottom-entry trigger; 2 pickups
maneuver-architecture-0905-timed-repair 700 20 0 Boss checkpoint: shield 39 to 19; 1 pickup, 2 wall shooters destroyed
maneuver-architecture-0905-boss-exit 1,564 0 0 Continued from the preceding save: 7 pickups; stopped before shop because progression cycled

Earlier boss iterations lost lives, and an earlier longer run stopped on a shell station cache-contract KeyError. Publishing cached proposals into the inherited policy index fixed that exception and has a regression test. The bottom-entry damage failure led to the pending-entry guard and its successful retest. The latest boss repair still took two swarm hits and a directional-projectile hit; full-level low-damage acceptance remains open. These are separate checkpoint tests, not a single successful full-campaign run.

The final continuation preserved shield 19 through frame 8510, but repeatedly switched between route advance and shell clearance near the remaining edge shells. It was stopped through the normal stop-file mechanism; both AVIs and the checkpoint were finalized. The global Y-only shell-release constraint appears to conflict with route/attack availability here. This is an unresolved progression failure, not a successful boss exit despite the run's name. A future fix needs an explicit choice between acquiring a reachable shell firing station and handling its released projectiles during a bypass; indefinitely waiting is not acceptable.

Final identical-observation timing comparison, with source fingerprints verified unchanged during both sequential runs:

  • Recording: maneuver-architecture-0905-spawn-fix.validation/maneuver-architecture-0905-spawn-fix.x2events.
  • Report: comparison.json.
  • 300 updates, game frames 4914–5213, including reconstruction and cold setup; actual campaign RAM/fire metadata and the same starting controller state.
Processing time Legacy Maneuver
Mean 55.69 ms 28.49 ms
Median 20.07 ms 10.83 ms
95th percentile 250.03 ms 123.32 ms
Maximum 1,164.00 ms 703.78 ms

Mean improvement was 1.95× on this sample. Actions differed on 223/300 updates, so this comparison establishes processing cost, not counterfactual survival. The new controller averaged 18.13 ms in inherited policy and 8.40 ms in maneuver planning. It averaged 97 unique trajectory transitions and 65 member queries per decision; maxima were 408 transitions and 32 rollouts. Independent predictions and group rejection avoid the original repeated global scoring, but inherited policy/navigation is now the larger mean cost in this sample.

Outstanding acceptance work

Keep maneuver opt-in until comparable live runs demonstrate level completion, shield loss and cash collection across all five levels. Old later-level recordings were exercised as smoke inputs, but missing RAM/fire metadata in some recordings prevents treating those results as gameplay evidence. In particular, the older Level-5 barrier trace without a starting RAM seed produced wall rejections.

The navigation work following this initial validation is documented separately in AUTOPILOT_NAVIGATION.MD. The maneuver search budget does not bound all inherited mission operations. Boss swarm holding/escape choices still need further live tuning before increasing aggression. Browser integration and target-phone timing remain unimplemented and unmeasured.