Xenon 2

Autopilot · write-up

Autopilot: removing repeated preparation work (2026-09-06)

xenondoc/AUTOPILOT_NATIVE_PREPARATION.MD · 12 KB · updated 2026-09-17

This follows the matched native/Python live validation. The committed starting point is bb828893 (implementation 793ad01f). Python remains the driver and the reference kernel remains selectable; this work does not integrate the autopilot into Hatari or profile WASM.

Compact maneuver keys

autopilot_v2/action_keys.py encodes each joystick input as one byte. A maneuver prefix is extended incrementally instead of repeatedly hashing a tuple of Candidate dataclasses. The readable actions and Motion objects are unchanged. Restored motion certificates and newly generated continuations use the same key representation. c-old-keys is a fresh-process reference representation for timing. It keeps the new incremental-prefix scaffolding while using Candidate tuples, so this control isolates key representation rather than replaying the old source.

Two opposite-order full-recording comparisons (7,693 observations, frames 2367–10059) give 12.162 → 11.306 ms mean replay processing time: 7.0% less. There are no decision-verdict differences. Each variant's repeated mean differs by less than 1%. This measures offline replay processing, not phone performance or another live emulator run.

Resident wall lookups and corridor intervals

The wall kernel accepts a resident 320-pixel-wide opaque raster and the current ship's dilated pixel mask. One call prepares a 256-row band of exact one-pixel anchor legality, checking all 320 horizontal anchors together. The Python hot path reads a cached byte; it does not call C per anchor or per motion sample. Motion sampling and the occupied-anchor escape policy remain unchanged.

Raster data is copied only when the observed map raster changes. Bands are lazy, shared across decisions, and bounded to eight distinct mask contents. Equivalent masks from different hull descriptions share an entry. A new observed raster invalidates all bands, including after destructible walls change and after replay seeks. The existing resident map remains authoritative for all levels; prediction never opens a destructible gate. The Python map still serves strategic spans.

An initial identity-keyed mask cache was rejected: equivalent masks churned its eight entries, causing 119,523 preparations and a mean regression from 11.531 to 17.937 ms in its first pass. Sharing by mask content removes that duplicate work. The failed run is retained under run_logs/navigation-kernel-0906 as diagnostic evidence, not included in the final timing estimates.

The corridor implementation projects opaque rendered wall pixels onto one horizontal bitset and derives blocked anchor intervals directly. It avoids the old 291-position scan. It preserves current rendered-wall semantics, which differ from the navigation ship-mask test; fractional custom inputs use the old scan.

The repeated navigation comparison produced the following mean replay times:

Variant Mean ms Change from preceding row
Compact keys, reference navigation 11.686 —
Native wall lookup 11.289 -3.4%
Wall lookup plus corridor projection 11.199 -0.8%

No verdicts changed in either repeat. Wall lookup required only 136 band preparations and 963,840 raster bytes uploaded across gameplay observations, instead of a C call for every anchor. The isolated corridor difference is small compared with its 4.1% repeat spread; treat it as a small/uncertain end-to-end gain, not a large additional speedup. Its policy work decreases in both repeats.

Shared shell paths and direct slice preparation

ABI 8 adds XapShellMotion and xap_prepare_shell_paths. The Python ShellPaths owner captures eligible shell oscillators and prepares them lazily: one C call advances every shell through the forecast horizon. Body prediction reads anchors from this buffer; body scenes bulk-copy the packed points and reuse them for both collision sweeps and combat bounds. C applies the existing four-pixel shell animation margin to body contact only; projectile hit bounds remain narrow.

This removes restarting the shell oscillator for each requested future frame. Captured shell procedure/state and the ordinary unlinked model define eligibility; unsupported state or a query beyond the prepared horizon retains the Python model. Homing, aimed releases, shell burst uncertainty and destruction timing are unchanged.

The first trial also redirected policy anchors to these paths. Full replay caught 32 changed verdicts: the existing policy's shell burst anchor uses a linear estimate, whereas body collision uses the oscillator. This trial is excluded from final comparisons. Burst anchors retain their established model here; reconciling that inconsistency needs a separate gameplay validation, not a performance claim.

prepared_slices.py supplies the remaining specialized envelopes and anticipated shots directly to the C scene's per-frame slice cache. It avoids constructing Group objects and computing group bounds that the C loop never consumes. PredictionService.group_slices remains the Python reference and detailed-contact path. Tests compare the two preparations across hazard kinds, exclusions, safety envelopes, collapsed chains and anticipated shots.

The shared replay/live diagnostics include wall preparation/upload counters, shell paths/points and direct native slice frames. Backend metadata names the new kernels and ABI 8; old recordings continue to show missing counters as unknown. New Python helpers have annotated inputs/outputs and parameter documentation.

Only the standalone DLL is built. The public C APIs and standalone WASM exports are available for future direct callers, but no Hatari driver/build integration is added here.

Measurement protocol

Measurements use the complete .x2events recording from the matched live run, the same terrain seeds and planner settings, fresh sequential Python processes, one pinned logical CPU, and reversed variant order on the second repeat. Sources and the DLL are fingerprinted across every repeat. No compilation, tests or other profiling runs overlap timing. The abandoned JSON-trace run under run_logs/cache-keys-0906 is excluded because it was not the full recording.

Planner counters exclude player-absent/shop rows, which can carry stale metrics. End-to-end replay times include reconstruction and all recorded observations. The known replay fire-schedule/shop limitations still apply; matching replay verdicts are a regression comparison, not a new closed-loop campaign validation.

Reproduction

The input for each command below is xenon_tools/run_logs/native-python-live-c-0906-02.validation/native-python-live-c-0906-02.x2events.

python xenon_tools/compare_kernels.py <recording> --output <new-directory> --variants c-old-keys,c-keys --from-frame 2367 --to-frame 10059 --repeats 2 --pin-cpu
python xenon_tools/compare_kernels.py <recording> --output <new-directory> --variants c-keys,c-walls,c-navigation --from-frame 2367 --to-frame 10059 --repeats 2 --pin-cpu
python xenon_tools/compare_kernels.py <recording> --output <new-directory> --variants c-navigation,c-shells,c --from-frame 2367 --to-frame 10059 --repeats 2 --pin-cpu

Use a different output directory for each command. The first experiment originally named its compact-key variant c; the later c-keys name explicitly disables all subsequently added stages. The latest default c enables all implemented stages.

Preparation results and validation

Variant Mean ms Change from preceding row
Compact keys and navigation changes 10.738 —
Shared native shell body paths 10.525 -2.0%
Paths plus direct native slices 10.480 -0.4%

All 7,693 verdicts match in both repeats after preserving burst-anchor behavior. Shell preparation improves the mean in both orders. Direct slices are slightly slower in one order and faster in the other; the 0.4% aggregate is within measurement noise, so this is not evidence of a material additional speedup. The complete preparation stage reduces mean time by 2.4% in this comparison. Do not add percentages from separate experiments: their controls and timing conditions differ, and the preparation results remain modest.

The gameplay counters show 953 shell calls preparing 2,366 paths / 229,502 points. Native fallback slice uploads drop from 17,766,720 to 14,720,928 bytes (17.1% less). Direct preparation serves 220,346 cached slice frames. These describe removed work and the changed ownership boundary, not an equivalent percentage FPS gain.

Validation: 866 Python/native tests pass, including random shell turn/contact comparisons, all hazard-kind slice comparisons, mask boundaries, cache reuse, five-level atlas corridor comparisons, and an observed destructible gate that opens only after the second confirming observation. Replay diagnostics have tests for the new counters and historical missing values. The standalone DLL builds successfully; no emulator rebuild or WASM profiling was performed.

Final profile and remaining cost

Final profile, using the same full recording and current default C backend, records 501.6 million function calls versus the previous profile's 759.3 million. This is useful evidence of removed Python work, but not a corresponding FPS improvement: cProfile disproportionately inflates tiny Python calls, especially the old Candidate hashes.

Current instrumented cumulative times (nested; do not add them):

Function / layer Calls Cumulative seconds
Maneuver evaluation 42,264 108.065
Python group queries 484,556 22.016
Python group construction/cache 484,556 15.890
Native body packing/query wrapper 104,925 12.682
Maneuver continuation 1,144,766 11.103
Motion-key lookup/simulation 2,234,099 10.448
Wall motion checks 1,135,755 9.973
Direct slice generator, including resumes 527,032 9.041
Ship/camera arithmetic 604,323 4.841
Corridor calculation wrapper 5,162 0.564

Motion service cumulative time was 34.023 seconds in the earlier profile; corridor calculation was 13.614 seconds. Those large instrumented reductions explain the change in call structure, not the much smaller uninstrumented gains.

There are still 637,165 Python body-reference frames with a native scene present. Fresh maneuver branches generally retain Python validation; the native batching primarily covers already-restored motions, plus combat checks. This is the next useful boundary to move: continuation, player/camera stepping and validation together in a whole-maneuver C call, using the resident terrain and prepared scene. Porting another scalar helper would leave these repeated Python loops and boundary handoffs intact. Independent hazards already consume C paths; player-dependent prediction and level tactics still need an explicit handoff.

Replay decoding (17.114 cumulative seconds) and world reconstruction (11.044) also remain. Do not build a native JSON/replay decoder solely for these timings: the planned direct C observation path removes that transport later.