Autopilot · write-up
Native maneuver forecast (2026-09-06)
This extends the preparation kernels committed as 53fbb53b with ABI 9.
--kernel-backend python keeps the reference implementation; c enables
the new forecast. Only the standalone library is built: no Hatari integration
or WASM profiling is part of this change.
Boundary and ownership
xap_forecast_maneuver receives named structs for the current ship/camera,
feedback mission, explicit input prefix, steering-dependent interaction hulls,
resident wall raster and prepared independent hazards. One call generates the
continuation, advances ship and camera, samples physical wall-mask motion and
queries independent bodies through the forecast horizon. Pending input precedes
selected input and affects feedback immediately. A confirmed wall collision ends
the C loop immediately; a possible enemy contact does not, because combat may
remove that conflict. The return value is the number of written frames.
The C loop follows game_model.py: steering decay, arithmetic negative halves,
pre-movement vertical clamps, scroll override acceleration and bounded backtracking.
World Y uses the camera's entry scroll, while camera-coupled hazards use the
post-camera correction. Wall masks retain four-pixel motion sampling and
ties-to-even anchor rounding. The native body query and maneuver kernel share
one internal contact/clearance implementation.
Python owns the backing ctypes arrays. C allocates nothing, retains no pointers,
and calls no Python callbacks. Immutable hazard paths, slice ranges and terrain
are shared across candidate maneuvers within an observation. The wall raster
continues to belong to NativeWallCache; observed destructible-wall changes
replace it through the existing map invalidation path. C never predicts a wall
as destroyed merely because a shot is expected to hit it.
Ordered checks retained in Python
The C result contains provisional all-alive contact/clearance and wall legality
for each frame. Python consumes reached frames in order, materializing ordinary
Motion objects only as needed. Existing prefix caches and transition budgets
remain active. This avoids adding a second set of serialized trajectory types
to recordings and retained plans.
Homing, aimed and barrier shots, shell bursts, imminent bottom spawns, refuge
constraints, pickup checks and destruction conflict resolution retain their
existing order in PredictionService.evaluate. Possible native body contacts
request the detailed reference check, which can invoke batched native combat.
Once combat is active, death-aware body validation uses that rollout's target
state. Anticipated projectiles remain independent of their shooter's later death.
An occupied starting wall anchor requests the existing monotonic escape test; missing wall-mask data requests reference wall validation. Neither case grants unchecked safety. Unsupported forecast limits use the Python continuation.
Comparison and diagnostics
compare_kernels.py --variants c-preparation,c isolates this change with all
previous optimizations enabled in both variants. Earlier ablation variants keep
the new maneuver port disabled. The shared live/replay summary displays native
maneuver calls, forecast frames and newly materialized frames; historical absent
fields remain unknown. Kernel metadata records ABI 9 and maneuver_forecast.
Tests compare feedback/physics over speed levels, steering reversal, camera clamps, override ages and pending inputs, plus fresh-branch body verdicts, crossed wall samples, occupied-start handoff and insufficient output capacity.
Reproduction
The source recording is
xenon_tools/run_logs/native-python-live-c-0906-02.validation/native-python-live-c-0906-02.x2events.
The comparison loads the same controller/terrain seeds, spans navigation and
architecture settings in fresh sequential workers, pinned to one logical CPU.
Repeat two reverses the order. Sources and DLL are fingerprinted across runs;
tests, compilation and profiling do not overlap the timing experiment.
python xenon_tools/compare_kernels.py <recording> --output xenon_tools/run_logs/maneuver-final-0906 --variants c-preparation,c --from-frame 2367 --to-frame 10059 --repeats 2 --pin-cpu
Use a new output directory when repeating the command. These are offline replay
processing times, including decoding and world reconstruction. Matching decisions
are a regression comparison, not a new closed-loop campaign or phone benchmark.
The early maneuver-first-0906 experiment predates the final input review and wall
early termination; it is retained as development evidence, not the final estimate.
Results
Final comparison: all 7,693 verdicts match in both repeats, frames 2367–10059.
| Kernel | Mean replay ms | Mean of repeated p95 ms |
|---|---|---|
Previous C preparation (c-preparation) |
10.578 | 38.649 |
Native maneuver (c) |
7.886 | 23.581 |
This is 25.4% lower mean processing time and 39.0% lower p95. The control's means differ by 0.1% across repeats; native means differ by 1.1%. It is a substantial measured reduction on this Level 1 recording, not a claim about every level, a live emulator campaign, or phone/WASM performance.
Gameplay-only work counters (5,425 observations with a player outside the shop):
| Work | Previous C preparation | Native maneuver |
|---|---|---|
| Python independent-body reference frames | 637,165 | 19,223 |
| Python wall-motion queries | 339,130 | 28 |
| Native maneuver calls | 0 | 42,264 |
| Native maneuver forecast frames | 0 | 1,864,820 |
| Newly materialized native motions | 0 | 476,648 |
| Direct slice preparation frames | 220,346 | 303,800 |
| Native slice bytes | 14,720,928 | 18,574,464 |
The new loop removes 97.0% of Python body-reference frames. It does forecast cheap native frames that an earlier dependent conflict or a reused prefix may make unnecessary. Early termination at confirmed walls removes another 501,964 native frames compared with the initial development version. Transition-budget counts and rollout counts are unchanged. Remaining wall queries handle occupied starts or unavailable physical mask data.
Full-horizon specialized slice preparation increases by 83,454 frames and about 3.85 MB across the recording. This is the explicit tradeoff for a coarse call without per-frame callbacks; the measured whole-replay gain includes this cost. Scene ownership leaves room to move those remaining specialized models later.
The standalone DLL builds without warnings. 870 tests pass, including the new physical gate/maneuver integration and replay counter checks. New Python functions have type annotations, one-sentence descriptions and parameter docs.
Final profile and remaining work
Final profile uses
the same full recording in a separate pinned worker, with --variants c
--repeats 1 --cprofile. It records 368.6 million calls, down from 501.6 million
in the preceding preparation profile. Profiler overhead is large: use the
uninstrumented paired results above for the performance claim.
Selected instrumented cumulative times are nested and must not be added:
| Layer | Calls | Cumulative seconds |
|---|---|---|
| Planner evaluation including native preparation | 44,225 | 57.268 |
| Ordered Python evaluation | 42,264 | 37.295 |
| Native maneuver preparation/call wrapper | 42,264 | 19.202 |
| Observation scene preparation (inside that wrapper) | 5,425 | 15.495 |
| Native fallback slice packing | 381,868 | 14.904 |
| Specialized slice generator including resumes | 690,768 | 12.781 |
| Native-motion materialization/cache lookup | 1,135,755 | 3.554 |
| Python group queries | 19,223 | 2.267 |
| Python group construction/cache | 19,223 | 1.645 |
The next boundary is now clearer: specialized hazard and anticipated-shot preparation, which accounts for most of the new wrapper's cost. The ordered Python evaluation loop also remains substantial (8.911 seconds self time), but porting it requires moving its dependent prediction and combat handoffs together. Do not restore per-frame ctypes calls for those models. Keep level tactics and candidate ranking outside that kernel until measurements justify moving them.
Replay decode (17.399 cumulative seconds) and world reconstruction (11.145) remain outside this port. Those are desktop transport/reconstruction costs; their eventual direct C observation interface is a separate task, not a reason to build a native replay-file parser. No additional port is implemented here.