Autopilot · write-up
Batched native body queries (2026-09-06)
This document records the ABI-2 body-query port committed as 16334ce0 and its
isolated measurements. The subsequent
linear prediction extension uses ABI 3,
adds compact linear descriptors and removes some generic frame preparation.
The measurements below remain the historical body-port results.
Scope and selection
--kernel-backend c now selects both the existing articulated-eye forecast and
batched independent-body validation. --kernel-backend python retains the Python
reference. The body implementation is separate in
src/autopilot/xenon_autopilot_bodies.c and
xenon_tools/autopilot_v2/native_bodies.py; Python group validation remains the
reference and destruction fallback in prediction.py.
This port accelerates validation of existing predictions. It does not change level tactics, map ownership, pickups, homing, shell bursts, weapon damage rules, or the navigation backend. Use spans navigation and the existing architecture features for the measured configuration. Body batching currently requires the formation broadphase, at least one supported complete hazard trajectory and an exactly restored motion suffix of at least eight steps. Fresh search branches retain Python validation.
Boundary and ownership
HazardKind(IntEnum)andXapHazardKindhave matching explicit numeric values. Convert observed classification strings once per observation. Prediction groups use enums; observations, recordings, contact diagnostics and replay retain readable names. Unknown labels use numericUNKNOWNand keep their original diagnostic label; they remain colliders.- ABI 2 uses named
ctypes.Structure/C structs. The existing eye-state layout is unchanged. Rebuild the shared library; an old ABI-1 library is rejected. - Python owns an observation's packed path and slice buffers. C borrows them for a call and retains no pointers. No native handle enters a retained plan, checkpoint or replay snapshot. Each query buffer is reused within that observation, and returned results are copied to immutable Python values.
- Supported ordinary formation followers, independent scripted swarms and eye chains use their existing full anchor trajectories plus current hull offsets. C constructs swept rectangles directly, including the eye margin. Synthetic pending allocations and incomplete paths retain their specialized predictor.
- Other independent bodies and anticipated shots supply typed frame slices. Each requested frame is predicted and packed once, then reused by subsequent maneuver checks. These models still run in Python; C does their collision and clearance scan together with the supported paths.
- One call checks up to eight future player hulls against the entire supplied body scene. It returns clearance capped at 24 pixels or a possible-contact marker/type for each frame. There are no per-member or per-rectangle ctypes calls. Scene bytes are packed once, not transferred again for each query.
- Terrain stays with the Python mission/navigation layer. This kernel never needs maps, so copying them into C would add ownership work without a benefit.
The ABI is shared by the Python-driven native build and standalone WASM module. The whole browser autopilot/driver is not implemented by this change, and desktop native/Node measurements do not establish phone performance.
Lazy behavior, ordering and destruction
Feedback-based maneuvers retain their original lazy generation and transition
budget. After cross-frame certificates restore an exactly matching motion suffix,
the native scene indexes those immutable Motion objects. When validation reaches
one of them, C checks up to eight already-known future hulls in one call. The result
is reused only for that exact motion object in the same observation. An early
rejection never generates speculative actions or ship/camera states, and body
prediction advances only through the requested block.
Other search branches retain Python's lazy validator, broadphase and certificates. This deliberately exploits geometry the planner already owns; it does not prepare extra candidate trajectories merely to fill a native batch. Returned contact markers are conservative and do not alter the caller's incoming clearance.
C assumes every supplied body is alive. A clear result can therefore certify the ordinary body pass without invoking combat. Any possible contact re-enters the full ordered Python group pass with the incoming clearance unchanged. That pass may advance combat and credit a destruction only according to the existing rules. After combat advances, subsequent body checks use Python's death-aware pass. Independent projectiles remain dangerous after their source is destroyed.
Homing, wall checks, bottom entries, refuge boundaries, shell bursts, aimed shots, barrier shots and pickup contacts retain their original per-frame ordering. C results for a later frame are consumed only if all earlier checks passed. Packed body results never certify terrain or assume a destructible wall disappeared.
Validation
- Full tools suite: 838 tests passed, with the native library built.
- C/Python enum values, struct sizes and field offsets are checked.
- Native differential tests cover swept bounds, inclusive touching, camera correction, eye margins, mixed specialized slices, empty scenes, malformed ranges/capacities, and unchanged output on invalid input.
- Integration tests compare full maneuver results and transition counts for early rejection, pending input and destruction resolved by an earlier shot.
- Native and WASM match 160 scenes / 1,280 body queries generated by Python.
- The rebuilt WASM module also matches 654,496 existing eye points in 1,316 cases. This verifies the shared C source/ABI, not a live browser campaign.
./xenon_tools/hatari_dev.ps1 -Action build-kernel
python xenon_tools/native_body_vectors.py --output xenon_tools/run_logs/native-body-vectors-0906-01.json
cmake --build out/build/autopilot-wasm
node xenon_tools/check_native_bodies_wasm.mjs out/build/autopilot-wasm/xenon_autopilot_wasm.js xenon_tools/run_logs/native-body-vectors-0906-01.json
Measurement method
compare_kernels.py --variants c-eye,c isolates this port: c-eye is a fresh-
process benchmark control that keeps the C eye forecast and numeric Python group
types, while disabling the new body scene/batching. It is not a production backend.
The ordinary python variant tests the complete Python kernel selection.
Both controls retain the same numeric group types, so body-kernel gains are not
attributed to replacing strings or to the earlier eye port.
Each timing includes observation reconstruction, policy, planning, path/slice packing, ctypes calls and output conversion, excluding Tk/AVI/emulator work. Use two opposite-order passes pinned to the same permitted logical CPU. Source, asset and DLL hashes, sidecars, raw timings and per-frame counters are recorded. Actions, tactics, verification, contact frame/kind/key, kills, pickups and reported clearance are compared. cProfile is a separate run; its cumulative times overlap and its instrumented elapsed times are not whole-controller speed estimates.
Final repeated measurements and profiles are below.
Diagnostics
The replay/autoplay maneuver panel displays body paths, calls, queried frames, clear frames, reference fallback frames and packed/input/output bytes, separately for recorded diagnostics and current analysis. Recordings retain backend, ABI, DLL hash and kernel names. Timing rows also include clearance so ranking regressions cannot be hidden by an unchanged first joystick input.
Rejected implementation experiments
The first implementation speculatively generated eight motions to fill each batch.
The initial mapping overlay repeatedly hashed long Candidate tuples; cProfile
showed motion service time growing from about 8 seconds to 21 seconds. Replacing
that overlay with dictionary copies removed much of the regression, but two
opposite-order plain runs still showed no convincing overall gain: corridor
16.921 vs 16.687 ms, boss eye-only C 9.204 vs 9.372 ms. That version also packed
geometry for rejected tails. It was removed, rather than enabled on the strength
of fewer Python group-query counters. Raw experiments remain under
native-bodies-*-0906-02 and native-bodies-corridor-profile-0906-01.
The final implementation batches only already-restored geometry. Its results must be assessed using the later profiles/timings below, not those rejected runs.
Final repeated timings and remaining cost
All runs use spans, all architecture features, identical recordings/sidecars, and two opposite-order passes pinned to one permitted logical CPU. Means include all preparation and ctypes overhead; p95 below is the average of each pass's p95.
| Sample / measured frames | Existing C eye + Python bodies | New C selection | Mean reduction | p95 before / after |
|---|---|---|---|---|
| Corridor, 801–2555 (1,755) | 17.000 ms | 16.425 ms | 3.38% | 53.766 / 52.226 ms |
| Boss, 7633–8476 (844) | 9.052 ms | 8.737 ms | 3.48% | 22.727 / 22.456 ms |
| Level 5, 90160–90459 (300) | 16.724 ms | 16.729 ms | unchanged (0 native body calls) | 19.185 / 18.720 ms |
The full Python boss kernel measured 12.847 ms, versus 8.737 ms for both C kernels together (32.0% lower). Most of that is the previously implemented eye port; the new body work contributes the incremental 3.5% above. Level 5 makes zero body-kernel calls, and its tiny mean difference is timing noise. The existing Level-5 unverified outcomes remain unchanged; this is not evidence of new Level-5 tactics or successful gameplay progression.
Every comparison matches actions, tactics, verification, contact frame/kind/key, kills, pickups and clearance across 2,899 sampled observations, in both passes. No new live campaign or phone/browser-driver benchmark was run.
These are modest gains, not an order-of-magnitude improvement. Both corridor passes improve in the same direction, but the effect is small enough that device and workload changes can alter it. Keep the Python option and compare on target hardware before choosing a deployment default.
Raw timing reports (including hashes, frame differences and repeated-run spread):
Separate final cProfile run
Corridor profile: native-bodies-corridor-profile-0906-03/{c-eye,c}-1.pstats.
All 1,755 compared decisions and clearances match. cProfile total
function time is 80.390 vs 78.805 seconds; use the plain timings above for speedups.
| Corridor cumulative function cost | Python bodies | C body selection |
|---|---|---|
query_groups |
17.340 s / 202,922 calls | 7.357 s / 131,102 calls |
group_slices (all callers) |
13.896 s | 12.783 s |
motion |
8.465 s / 511,161 calls | 8.534 s / 511,161 calls |
Native batch check (including frame preparation) |
— | 8.297 s / 7,617 calls |
frame_slices (inside check) |
— | 7.938 s |
The limiting cost is still Python geometry preparation. Moving the final
contact/distance scan to C removes many Python queries, but specialized
group_slices prediction and packing consume most of that saving. C does not
yet own those motion models. Do not interpret the reduction in query calls as
an equivalent reduction in total processing time. Motion calls and transition
budgets are unchanged; the discarded speculative version failed that work-cost
objective. Cumulative costs overlap and must not be added.
Average body boundary traffic per observation (zero-call observations included):
| Sample | C body calls | Queried hulls | Packed paths + slices | Query + result bytes |
|---|---|---|---|---|
| Corridor | 4.34 | 34.13 | 16235 | 2458 |
| Boss | 5.61 | 44.07 | 9665 | 3173 |
Packed scene bytes remain in caller-owned native buffers for the observation; only pointers/counts cross the call boundary. The byte counters measure prepared payload, not additional per-call copying or peak Python heap use. The standalone WASM module is 11,179 bytes in this build, excluding its JavaScript loader.
Boss profile: native-bodies-boss-profile-0906-04/{c-eye,c}-1.pstats.
All 844 decisions/clearances match. cProfile total is 20.543 vs 19.763 seconds.
query_groups falls from 4.960 to 1.331 seconds, while native batch preparation
and checking costs 3.040 seconds (2.888 in frame_slices). Motion calls remain
exactly 273,199. This independently confirms the same limitation: Python slice
preparation consumes most of the collision-query saving.