Xenon 2

Autopilot · write-up

Policy preparation optimization

xenondoc/AUTOPILOT_POLICY_PREPARATION.MD · 11 KB · updated 2026-09-17

Baseline: 5a6ac645 (complete native candidate evaluation, ABI 11).

Order and attribution

  1. Replace single-anchor fallback grid builds with local wall probes. Retain the old fallback behind navigation.LOCAL_FALLBACK for comparison. The local query uses the same four-pixel anchor rounding, hull extents, horizontal clamp, unknown-tile handling and optional wall-pixel rectangle test. It reads the current map, so confirmed destructible-wall changes need no new cache protocol.
  2. Share the neutral swarm forecast between the mission and pickup queries in one observation. Empty swarms need no ship prediction. Longer requests extend the checked prefix; shorter requests reuse its earliest contact. Ship motion is requested as one trajectory, rather than reconstructing every prefix.
  3. Move actual configuration-grid construction into a coarse standalone C call: prepare free cells and their capped Manhattan clearance together. Retain the resident wall raster and Python reference. This is distinct from step 1: point probes should never trigger full grid construction in either language.

Evidence before changes

The Level 1 candidate profile (candidate-final-profile-0907/c-1.pstats) contains 30.95 seconds cumulative in legacy policy selection. Swarm threat queries account for 8.45 seconds, navigation grid preparation 5.60 seconds (including 2.56 seconds in free-row generation and 1.50 seconds in clearance), and target priorities 4.13 seconds. Cumulative times overlap; they are not additive wall-clock totals.

The focused Level 2 profile, frames 32680–32700, found 31.04 of 31.54 profiler seconds in _ensure_navigation_grid. Missing ship-mask data sends point probes through the legacy whole-grid fallback. This occurs in both policy and candidate validation; it is not an expensive homing model.

Step 1 measurement

policy-local-l2-0907/comparison.json: two fresh pinned workers per variant, reversed order on repeat two, 21 measured observations with preceding warm-up. Mean replay processing: 456.839 → 10.069 ms (97.8% lower). No changed verdicts in either repeat. At frame 32690 the first optimized repeat reports policy 0.77 ms and planning 2.76 ms. This small selected window is diagnostic, not a whole-level performance claim.

The local-query differential test covers tile-only and wall-atlas-without-ship-mask maps, unknown cells, negative world coordinates, map boundaries and observed changes.

Controls

  • c-policy-old: the committed candidate kernels and previous policy preparation.
  • c-policy-local: local fallback queries only.
  • c-policy-swarm: local queries plus shared swarm forecasts, Python grid builder.
  • c: all current optimizations.

The earlier policy-stages-l1-0907 artifact predates the C grid builder; its c means local queries plus swarm reuse (now named c-policy-swarm). Its source fingerprints preserve that attribution.

Step 2 measurement

policy-stages-l1-0907/comparison.json: two fresh pinned workers per variant, 7,693 observations (5,425 gameplay) per worker, frames 2367–10059. Every replay verdict matches in both repeats.

Stage Mean processing ms
Previous preparation 6.141
Local fallback probes 6.078
Local probes + shared swarm forecast 5.837

Swarm reuse reduces the preceding stage's mean by 4.0%. The local-probe change has little effect in this recording because authoritative masks are usually available. These timings include decoding and non-gameplay observations.

Native grid boundary

ABI 12 adds XapNavigationGridRequest and xap_prepare_navigation_grid in xenon_autopilot_navigation.c. The request names its relative pixel window and footprint extents; output is an 80-column array of free cells plus capped Manhattan clearance. One invocation builds the complete requested grid. The map-owned cache shares wall and ship-mask buffers with exact candidate wall checks, and uploads changed wall rows only when the observed raster identity changes.

Python stores immutable grid bytes in the existing revision/geometry cache. Destructible walls remain closed until observed clearing is confirmed. No C pointer is retained in controller checkpoints, no callbacks enter Python, and route search and encounter tactics remain in Python. The original grid builder is retained for Python selection and comparison; unsupported ship-mask formats continue through that builder.

Tests cover complete free/clearance buffers, mask shifts across 64-bit words, unknown map rows, tactical windows, undersized output rejection, and confirmed gate destruction. The existing exact-wall cache tests also cover replay raster restoration. Shared forecast tests cover longer/shorter queries, hit precedence and empty scenes.

Step 3 measurement

policy-grid-l1-0907/comparison.json compares the shared-swarm stage with and without native grids over the same 7,693 Level 1 observations. Two fresh pinned workers per variant, reversed order on repeat two, produce zero changed verdicts.

Stage Mean ms Mean repeated p95 ms
Local queries + swarm reuse, Python grid 5.843 13.913
Local queries + swarm reuse, C grid 5.500 13.647

Native grid preparation reduces this stage's mean by 5.9%. Across the staged measurements, previous preparation's 6.141 ms becomes 5.500 ms, approximately 10.4% lower. This combined number spans two benchmark groups; the isolated step-3 comparison above is the direct within-group measurement. The first repeats' gameplay policy timers decrease from 2.515 to 1.879 ms; planning is 3.315 versus 3.236 ms. Most of the gain belongs to the intended preparation layer.

The complete suite passed 888 tests before adding one more grid/gate test; all three native navigation tests then passed. The standalone ABI-12 library builds successfully. New Python functions have typed arguments/returns and parameter documentation; the C entry point separates validation from execution in place.

Full Level 2 recording

policy-final-level2-homing-dual-collision-0824-60-0907/comparison.json compares 961 observations (910 gameplay). Zero changed verdicts. One pinned run per variant gives 168.357 → 7.787 ms mean, with p95 1171.998 → 15.304 ms. This is a 95.4% mean reduction in this older recording, whose unavailable ship-mask data exercises the fallback. It should not be extrapolated to recordings with authoritative masks or to phone/WASM execution. Optimized gameplay averages 291 local fallback probes per observation; none of those probes builds a grid.

Other level replays

Each single-pass comparison has zero changed verdicts, including selected input, authority, clearance, first-contact frame/type/identity, predicted kills and pickups. Small timing differences from these checks should not be treated as precise speedup estimates; the repeated Level 1 measurements establish stage attribution.

Recording Observations / gameplay Previous mean ms Current mean ms
Level 3 stage-2 small shots 225 / 225 14.243 10.913
Level 4 wall model 900 / 752 17.548 8.274
Level 5 tank homing 5,652 / 391 4.613 3.314

Level 5 is mostly non-gameplay; its whole-recording mean is not an active combat frame budget. Its p95 remains approximately 30 ms. Artifacts:

Final separate profile and remaining work

policy-final-profile-0907/c-1.pstats profiles the same 7,693 Level 1 observations. Its decisions/diagnostics also match the uninstrumented reference. Total function calls decrease from 290,393,740 to 264,598,807; profiler time is 92.449 seconds. Instrumented elapsed times are not runtime speedup measurements.

Cumulative profile time Before, seconds After, seconds
Policy adapter selection 31.619 24.367
Swarm threat query 8.445 5.474
Configuration-grid preparation 5.595 0.240
Candidate evaluation 18.768 18.329

The 55 complete native grid builds take 0.145 cumulative seconds, including packing and copying. They reuse the resident terrain buffers; there is no per-cell Python/C traffic. Cumulative policy time is about 23% lower.

Next candidates are shared hazard/neutral-ship preparation and per-observation target reachability: swarm policy still costs 5.47 cumulative seconds, while candidate setup/materialization remains substantial. Investigate sharing existing prepared models before adding another parallel predictor. Existing mission refresh rules remain unchanged; broadening reuse across observations should be validated in closed-loop gameplay, particularly around gates and route rejoins. Frame decoding (17.24 seconds) and world updates (10.95 seconds) remain separate Python-driver costs. Future direct-C integration may remove part of that work; this stage does not port the decoder or integrate with Hatari.

Reproduction commands

Use fresh output directories. Timing workers must run sequentially with unchanged source and DLL fingerprints; run cProfile separately from timing comparisons.

./xenon_tools/hatari_dev.ps1 -Action build-kernel
python xenon_tools/compare_kernels.py xenon_tools/run_logs/native-python-live-c-0906-02.validation/native-python-live-c-0906-02.x2events --output <timings> --variants c-policy-old,c-policy-local,c-policy-swarm,c --from-frame 2367 --to-frame 10059 --repeats 2 --pin-cpu
python xenon_tools/compare_kernels.py xenon_tools/run_logs/level2-homing-dual-collision-0824-60.x2events --output <level2-timings> --variants c-policy-old,c --repeats 1 --pin-cpu
python xenon_tools/compare_kernels.py xenon_tools/run_logs/native-python-live-c-0906-02.validation/native-python-live-c-0906-02.x2events --output <profile> --variants c --from-frame 2367 --to-frame 10059 --repeats 1 --pin-cpu --cprofile
python -m unittest discover -s xenon_tools -p 'test_*.py'

Benchmarks use compare_kernels.py, the spans navigation backend, unchanged planner settings, matching checkpoints and fixed CPU affinity. Function profiles are separate from uninstrumented timings. Validation is offline replay; this work does not establish closed-loop survival, phone or WASM performance.