Autopilot · write-up
Autopilot: removing repeated preparation work (2026-09-06)
This follows the matched native/Python live validation.
The committed starting point is bb828893 (implementation 793ad01f).
Python remains the driver and the reference kernel remains selectable; this work
does not integrate the autopilot into Hatari or profile WASM.
Compact maneuver keys
autopilot_v2/action_keys.py encodes each joystick input as one byte. A maneuver
prefix is extended incrementally instead of repeatedly hashing a tuple of
Candidate dataclasses. The readable actions and Motion objects are unchanged.
Restored motion certificates and newly generated continuations use the same key
representation. c-old-keys is a fresh-process reference representation for timing.
It keeps the new incremental-prefix scaffolding while using Candidate tuples,
so this control isolates key representation rather than replaying the old source.
Two opposite-order full-recording comparisons (7,693 observations, frames 2367–10059) give 12.162 → 11.306 ms mean replay processing time: 7.0% less. There are no decision-verdict differences. Each variant's repeated mean differs by less than 1%. This measures offline replay processing, not phone performance or another live emulator run.
Resident wall lookups and corridor intervals
The wall kernel accepts a resident 320-pixel-wide opaque raster and the current ship's dilated pixel mask. One call prepares a 256-row band of exact one-pixel anchor legality, checking all 320 horizontal anchors together. The Python hot path reads a cached byte; it does not call C per anchor or per motion sample. Motion sampling and the occupied-anchor escape policy remain unchanged.
Raster data is copied only when the observed map raster changes. Bands are lazy, shared across decisions, and bounded to eight distinct mask contents. Equivalent masks from different hull descriptions share an entry. A new observed raster invalidates all bands, including after destructible walls change and after replay seeks. The existing resident map remains authoritative for all levels; prediction never opens a destructible gate. The Python map still serves strategic spans.
An initial identity-keyed mask cache was rejected: equivalent masks churned its
eight entries, causing 119,523 preparations and a mean regression from 11.531 to
17.937 ms in its first pass. Sharing by mask content removes that duplicate work.
The failed run is retained under run_logs/navigation-kernel-0906 as diagnostic
evidence, not included in the final timing estimates.
The corridor implementation projects opaque rendered wall pixels onto one horizontal bitset and derives blocked anchor intervals directly. It avoids the old 291-position scan. It preserves current rendered-wall semantics, which differ from the navigation ship-mask test; fractional custom inputs use the old scan.
The repeated navigation comparison produced the following mean replay times:
| Variant | Mean ms | Change from preceding row |
|---|---|---|
| Compact keys, reference navigation | 11.686 | — |
| Native wall lookup | 11.289 | -3.4% |
| Wall lookup plus corridor projection | 11.199 | -0.8% |
No verdicts changed in either repeat. Wall lookup required only 136 band preparations and 963,840 raster bytes uploaded across gameplay observations, instead of a C call for every anchor. The isolated corridor difference is small compared with its 4.1% repeat spread; treat it as a small/uncertain end-to-end gain, not a large additional speedup. Its policy work decreases in both repeats.
Shared shell paths and direct slice preparation
ABI 8 adds XapShellMotion and xap_prepare_shell_paths. The Python
ShellPaths owner captures eligible shell oscillators and prepares them lazily:
one C call advances every shell through the forecast horizon. Body prediction reads
anchors from this buffer; body scenes bulk-copy the packed points and reuse them
for both collision sweeps and combat bounds. C applies the existing four-pixel
shell animation margin to body contact only; projectile hit bounds remain narrow.
This removes restarting the shell oscillator for each requested future frame. Captured shell procedure/state and the ordinary unlinked model define eligibility; unsupported state or a query beyond the prepared horizon retains the Python model. Homing, aimed releases, shell burst uncertainty and destruction timing are unchanged.
The first trial also redirected policy anchors to these paths. Full replay caught 32 changed verdicts: the existing policy's shell burst anchor uses a linear estimate, whereas body collision uses the oscillator. This trial is excluded from final comparisons. Burst anchors retain their established model here; reconciling that inconsistency needs a separate gameplay validation, not a performance claim.
prepared_slices.py supplies the remaining specialized envelopes and anticipated
shots directly to the C scene's per-frame slice cache. It avoids constructing
Group objects and computing group bounds that the C loop never consumes.
PredictionService.group_slices remains the Python reference and detailed-contact
path. Tests compare the two preparations across hazard kinds, exclusions, safety
envelopes, collapsed chains and anticipated shots.
The shared replay/live diagnostics include wall preparation/upload counters, shell paths/points and direct native slice frames. Backend metadata names the new kernels and ABI 8; old recordings continue to show missing counters as unknown. New Python helpers have annotated inputs/outputs and parameter documentation.
Only the standalone DLL is built. The public C APIs and standalone WASM exports are available for future direct callers, but no Hatari driver/build integration is added here.
Measurement protocol
Measurements use the complete .x2events recording from the matched live run,
the same terrain seeds and planner settings, fresh sequential Python processes,
one pinned logical CPU, and reversed variant order on the second repeat. Sources
and the DLL are fingerprinted across every repeat. No compilation, tests or
other profiling runs overlap timing. The abandoned JSON-trace run under
run_logs/cache-keys-0906 is excluded because it was not the full recording.
Planner counters exclude player-absent/shop rows, which can carry stale metrics. End-to-end replay times include reconstruction and all recorded observations. The known replay fire-schedule/shop limitations still apply; matching replay verdicts are a regression comparison, not a new closed-loop campaign validation.
Reproduction
The input for each command below is
xenon_tools/run_logs/native-python-live-c-0906-02.validation/native-python-live-c-0906-02.x2events.
python xenon_tools/compare_kernels.py <recording> --output <new-directory> --variants c-old-keys,c-keys --from-frame 2367 --to-frame 10059 --repeats 2 --pin-cpu
python xenon_tools/compare_kernels.py <recording> --output <new-directory> --variants c-keys,c-walls,c-navigation --from-frame 2367 --to-frame 10059 --repeats 2 --pin-cpu
python xenon_tools/compare_kernels.py <recording> --output <new-directory> --variants c-navigation,c-shells,c --from-frame 2367 --to-frame 10059 --repeats 2 --pin-cpu
Use a different output directory for each command. The first experiment originally
named its compact-key variant c; the later c-keys name explicitly disables all
subsequently added stages. The latest default c enables all implemented stages.
Preparation results and validation
| Variant | Mean ms | Change from preceding row |
|---|---|---|
| Compact keys and navigation changes | 10.738 | — |
| Shared native shell body paths | 10.525 | -2.0% |
| Paths plus direct native slices | 10.480 | -0.4% |
All 7,693 verdicts match in both repeats after preserving burst-anchor behavior. Shell preparation improves the mean in both orders. Direct slices are slightly slower in one order and faster in the other; the 0.4% aggregate is within measurement noise, so this is not evidence of a material additional speedup. The complete preparation stage reduces mean time by 2.4% in this comparison. Do not add percentages from separate experiments: their controls and timing conditions differ, and the preparation results remain modest.
The gameplay counters show 953 shell calls preparing 2,366 paths / 229,502 points. Native fallback slice uploads drop from 17,766,720 to 14,720,928 bytes (17.1% less). Direct preparation serves 220,346 cached slice frames. These describe removed work and the changed ownership boundary, not an equivalent percentage FPS gain.
Validation: 866 Python/native tests pass, including random shell turn/contact comparisons, all hazard-kind slice comparisons, mask boundaries, cache reuse, five-level atlas corridor comparisons, and an observed destructible gate that opens only after the second confirming observation. Replay diagnostics have tests for the new counters and historical missing values. The standalone DLL builds successfully; no emulator rebuild or WASM profiling was performed.
Final profile and remaining cost
Final profile, using the same full recording and current default C backend, records 501.6 million function calls versus the previous profile's 759.3 million. This is useful evidence of removed Python work, but not a corresponding FPS improvement: cProfile disproportionately inflates tiny Python calls, especially the old Candidate hashes.
Current instrumented cumulative times (nested; do not add them):
| Function / layer | Calls | Cumulative seconds |
|---|---|---|
| Maneuver evaluation | 42,264 | 108.065 |
| Python group queries | 484,556 | 22.016 |
| Python group construction/cache | 484,556 | 15.890 |
| Native body packing/query wrapper | 104,925 | 12.682 |
| Maneuver continuation | 1,144,766 | 11.103 |
| Motion-key lookup/simulation | 2,234,099 | 10.448 |
| Wall motion checks | 1,135,755 | 9.973 |
| Direct slice generator, including resumes | 527,032 | 9.041 |
| Ship/camera arithmetic | 604,323 | 4.841 |
| Corridor calculation wrapper | 5,162 | 0.564 |
Motion service cumulative time was 34.023 seconds in the earlier profile; corridor calculation was 13.614 seconds. Those large instrumented reductions explain the change in call structure, not the much smaller uninstrumented gains.
There are still 637,165 Python body-reference frames with a native scene present. Fresh maneuver branches generally retain Python validation; the native batching primarily covers already-restored motions, plus combat checks. This is the next useful boundary to move: continuation, player/camera stepping and validation together in a whole-maneuver C call, using the resident terrain and prepared scene. Porting another scalar helper would leave these repeated Python loops and boundary handoffs intact. Independent hazards already consume C paths; player-dependent prediction and level tactics still need an explicit handoff.
Replay decoding (17.114 cumulative seconds) and world reconstruction (11.044) also remain. Do not build a native JSON/replay decoder solely for these timings: the planned direct C observation path removes that transport later.