Autopilot · write-up
Maneuver navigation optimization — 2026-09-05
Follow-up shell/eye fixes, the replay navigation checkbox and new live regression evidence are in AUTOPILOT_SHELL_VALIDATION.MD.
The exact grid backend is enabled for --planner maneuver; --planner legacy
retains the original grid builder and cell A. The default accelerates geometry
and uses flat cell indices while preserving reference A costs and tie-breaking.
The more aggressive spans search remains experimental after a boss regression.
The implementation is separated in
xenon_tools/autopilot_v2/navigation.py, behind small optional hooks in
PersistentWorldMap.
Existing offline data
The source remains assets/xenon_level_maps.json, loaded by the existing
PersistentWorldMap/level_tile_maps code, plus the runtime atlas's exact opaque
wall and ship pixels. No competing map asset or guessed corridor layout was added.
plot_level_route.py already defines the useful topology: complete resident
maps, a four-pixel configuration grid, connected components, and explicit
destructible-gate engagements. The experimental route mode contracts horizontal runs in that
same grid into connected corridor spans. It does not blindly replay the plotting
tool's all-gates-open illustration or replace the existing Level-2 offline field.
Runtime observed gates, current ship mask/reserve, and camera limits still matter.
Runtime changes
- Wall occupancy uses one integer bitset per pixel row. Repeated tile bands share immutable results. Exact player-mask probes no longer build an unused full pixel summed-area table (roughly six MB for a 320×4800 map).
- Collision inflation evaluates all horizontal anchors together using bit shifts. Contiguous runs in the ship mask use logarithmic shift/OR expansion. Identical wall neighborhoods and masks share cached free-space rows across decisions.
- Complete free-space grids remain available to the existing policy, renderer, exact connector validation, and smoothing. Their clearance arrays match the reference capped Manhattan transform.
- Default route search uses flat cell indices and direct grid lookups to eliminate repeated tuple construction and nested coordinate/free-space calls. It retains the original neighbor ordering, costs, heuristic, heap tie-breaking and frontier fallback. Tests compare the actual paths, not just reachability.
- Experimental route search operates on horizontal spans rather than individual four-pixel cells. Overlapping spans on adjacent rows are connected. This preserves reachability without cutting diagonal corners; reconstructed horizontal and vertical edges are legal before the existing smoothing/checks run.
- Existing mission routes and bounded firing-station proposals remain retained. In experimental mode a new acquisition/reconnection searches the cached corridor graph. This changes route cost preferences, so exact old waypoint/action parity is not promised: search favors short, wide passages using span-center costs.
RAM synchronization and observed tile changes still invalidate map geometry. Unchanged row neighborhoods reuse content-addressed results after a destruction; changed neighborhoods are recomputed. Clearance and graph connectivity are rebuilt for a changed grid because their effects can extend beyond the changed tiles. No predicted enemy destruction speculatively opens tiles.
Caches contain immutable derived geometry and are bounded: 512 tile bands, 256 mask profiles, 16,384 horizontal occupancy expansions, 8,192 free-space rows, 1,024 span rows, and eight corridor graphs. Existing per-map grid storage keeps its 32-entry limit. Replay snapshots share immutable values; diagnostic counters belong to the individual decision.
Comparison
To hold the maneuver planner fixed and compare only navigation implementations:
python xenon_tools/compare_autopilot.py path/to/run.x2events --navigation-only --output xenon_tools/run_logs/nav-comparison
For a separate profile, use profile_autoplay.py replay ... --planner maneuver
--navigation-backend legacy, --navigation-backend grid, or the experimental
--navigation-backend spans. Add --navigation-candidate spans to the comparison
command to test that experiment instead of the default exact backend. Supply the same
starting controller state and terrain manifest. PlannerConfig.navigation_backend
is checkpointed and can also select either implementation programmatically.
Trace/replay diagnostics include navigation_backend, searches, expanded spans,
grid builds, and reused/rebuilt rows. A profile comparison includes reconstruction,
not drawing or emulation, and is not a counterfactual gameplay result.
Measured results
Default exact-backend comparison on 300 updates (4914–5213), with zero changed actions, including cold setup and reconstruction:
| Metric | Previous maneuver navigation | Exact accelerated navigation |
|---|---|---|
| Mean total processing | 40.08 ms | 29.01 ms |
| Median | 17.04 ms | 17.42 ms |
| 95th percentile | 189.25 ms | 94.50 ms |
| Maximum | 702.72 ms | 274.90 ms |
| Mean policy/navigation | 24.27 ms | 12.29 ms |
Total processing improved 1.38×, policy/navigation 1.97×; ordinary fast frames were essentially unchanged. Machine timings varied during the session, so compare the paired variants within each report rather than absolute times between experiments. Default-backend report.
A separate quiet comparison of 100 boss-area updates (6247–6346) also selected identical actions on every update:
| Metric | Previous maneuver navigation | Exact accelerated navigation |
|---|---|---|
| Mean total processing | 122.73 ms | 63.20 ms |
| Median | 36.50 ms | 39.78 ms |
| 95th percentile | 596.88 ms | 121.83 ms |
| Maximum | 2,037.86 ms | 457.20 ms |
| Mean policy/navigation | 72.59 ms | 21.58 ms |
Total processing improved 1.94×, policy/navigation 3.36×. This smaller sample deliberately includes the expensive cold target-preposition section. The median did not improve; the gain is concentrated in expensive acquisitions. Default boss-area report. An earlier non-quiet boss measurement overlapped the tail of the test suite and is not used here.
Experimental corridor results (not the default)
On the same 300 updates (frames 4914–5213) from
maneuver-architecture-0905-spawn-fix.validation, with RAM/fire metadata and cold
setup, sequential runs without cProfile produced:
| Metric | Previous maneuver navigation | Corridor navigation |
|---|---|---|
| Mean total processing | 28.30 ms | 14.32 ms |
| Median | 10.58 ms | 10.11 ms |
| 95th percentile | 120.41 ms | 35.40 ms |
| Maximum | 702.84 ms | 117.62 ms |
| Mean policy/navigation | 18.32 ms | 4.09 ms |
| Mean maneuver planning | 8.06 ms | 8.28 ms |
Policy/navigation improved 4.48×, total processing 1.98×. Actions differed on 16/300 updates. These measurements include cold setup; they are not warmed cache-only numbers. The immutable geometry caches were first used by the second variant; the baseline variant does not populate them.
Report with source/asset fingerprints and raw timing streams: nav-comparison-spawn-0905/comparison.json.
A second comparison used 300 boss-area observations (6247–6546), with the same terrain/fire/start-state discipline:
| Metric | Previous maneuver navigation | Corridor navigation |
|---|---|---|
| Mean total processing | 105.11 ms | 50.84 ms |
| Median | 58.36 ms | 36.98 ms |
| 95th percentile | 353.30 ms | 100.45 ms |
| Maximum | 1,367.60 ms | 431.93 ms |
| Mean policy/navigation | 70.72 ms | 9.53 ms |
| Mean maneuver planning | 26.47 ms | 35.27 ms |
Policy/navigation improved 7.42×, total processing 2.07×. Actions differed on 41/300 updates; new proposals increased maneuver repair work in this sample, so the total gain is smaller than the policy gain. Boss comparison report.
The profiled 100-update diagnostic changed navigation route acquisition from 5.33 seconds cumulative to approximately 0.62 seconds. Profiler timings are diagnostic and should not be compared directly to the table above.
Validation and limits
The final broad suite passed 672 tests. New checks compare every free-space and clearance cell against the reference on all five resident levels, at two reserves and both full-map/windowed extents. Additional checks cover bit-mask clipping, randomized reference reachability, camera limits, legal reconstructed edges, and gate opening/reclosing without stale topology. The RAM regression verifies that an actual gate update invalidates the cached wall raster and produces the same new grid/clearance as the reference builder. Randomized exact-path tests also compare all three A* entry points, including frontier fallback, goal regions, varying clearance costs and camera limits.
The final default-backend live test (nav-exact-0905-boss.validation) ran 700
uninterrupted updates from the same boss checkpoint as the reference run.
Every applied input, next input/action, shield value, life count, scroll position
and score matched the reference on all 700 updates. It finished with the same
20 shield damage, no lives lost, 2,110 score gained, two wall shooters destroyed
and one pickup. This verifies preservation of the existing behavior, including
its remaining encounter failures; it is not a claim that those failures are fixed.
Live comparison.
The following live tests used the experimental corridor variant:
A recorded live 300-update replay from the pre-bottom-entry checkpoint
(nav-spans-0905-spawn.validation) finished with zero shield damage, zero lives
lost, and three pickups, increasing score by 3,150. The previous checkpoint
test collected two pickups. Changed trajectories and hazard encounters mean this
is a regression sample, not proof of better collection over the campaign.
A separate 300-update boss checkpoint run (nav-spans-0905-boss.validation)
preserved shield 39 and destroyed two wall shooters. That run ends immediately
before the troublesome swarm and is not a boss-completion result. Older Level
2–5 recordings also completed 60-frame smoke reconstructions each without an
exception; incomplete terrain/fire metadata (especially the old Level-5 trace)
means those are compatibility checks, not live safety evidence.
The 400-update continuation (nav-spans-0905-boss-wave.validation) lost one life
to five swarm contacts. This is worse than the earlier uninterrupted reference
700-update test (20 damage, no life lost), although restarting at the checkpoint
also discards transient plans. The corridor variant was therefore not promoted
to the default. Its speed measurements above must not be presented as the
accepted default's speed or as evidence of successful boss play.
The known Level-1 shell progression issue is an encounter-policy problem and is not declared fixed by this geometry/search optimization. Full campaign acceptance and browser/phone measurements remain outstanding. Search work is much smaller, but this backend still has finite full-map cold construction and is not a hard wall-clock-budgeted runtime.