Xenon 2

Autopilot · write-up

Batched native body queries (2026-09-06)

xenondoc/AUTOPILOT_NATIVE_BODY_KERNEL.MD · 13 KB · updated 2026-09-17

This document records the ABI-2 body-query port committed as 16334ce0 and its isolated measurements. The subsequent linear prediction extension uses ABI 3, adds compact linear descriptors and removes some generic frame preparation. The measurements below remain the historical body-port results.

Scope and selection

--kernel-backend c now selects both the existing articulated-eye forecast and batched independent-body validation. --kernel-backend python retains the Python reference. The body implementation is separate in src/autopilot/xenon_autopilot_bodies.c and xenon_tools/autopilot_v2/native_bodies.py; Python group validation remains the reference and destruction fallback in prediction.py.

This port accelerates validation of existing predictions. It does not change level tactics, map ownership, pickups, homing, shell bursts, weapon damage rules, or the navigation backend. Use spans navigation and the existing architecture features for the measured configuration. Body batching currently requires the formation broadphase, at least one supported complete hazard trajectory and an exactly restored motion suffix of at least eight steps. Fresh search branches retain Python validation.

Boundary and ownership

  • HazardKind(IntEnum) and XapHazardKind have matching explicit numeric values. Convert observed classification strings once per observation. Prediction groups use enums; observations, recordings, contact diagnostics and replay retain readable names. Unknown labels use numeric UNKNOWN and keep their original diagnostic label; they remain colliders.
  • ABI 2 uses named ctypes.Structure/C structs. The existing eye-state layout is unchanged. Rebuild the shared library; an old ABI-1 library is rejected.
  • Python owns an observation's packed path and slice buffers. C borrows them for a call and retains no pointers. No native handle enters a retained plan, checkpoint or replay snapshot. Each query buffer is reused within that observation, and returned results are copied to immutable Python values.
  • Supported ordinary formation followers, independent scripted swarms and eye chains use their existing full anchor trajectories plus current hull offsets. C constructs swept rectangles directly, including the eye margin. Synthetic pending allocations and incomplete paths retain their specialized predictor.
  • Other independent bodies and anticipated shots supply typed frame slices. Each requested frame is predicted and packed once, then reused by subsequent maneuver checks. These models still run in Python; C does their collision and clearance scan together with the supported paths.
  • One call checks up to eight future player hulls against the entire supplied body scene. It returns clearance capped at 24 pixels or a possible-contact marker/type for each frame. There are no per-member or per-rectangle ctypes calls. Scene bytes are packed once, not transferred again for each query.
  • Terrain stays with the Python mission/navigation layer. This kernel never needs maps, so copying them into C would add ownership work without a benefit.

The ABI is shared by the Python-driven native build and standalone WASM module. The whole browser autopilot/driver is not implemented by this change, and desktop native/Node measurements do not establish phone performance.

Lazy behavior, ordering and destruction

Feedback-based maneuvers retain their original lazy generation and transition budget. After cross-frame certificates restore an exactly matching motion suffix, the native scene indexes those immutable Motion objects. When validation reaches one of them, C checks up to eight already-known future hulls in one call. The result is reused only for that exact motion object in the same observation. An early rejection never generates speculative actions or ship/camera states, and body prediction advances only through the requested block.

Other search branches retain Python's lazy validator, broadphase and certificates. This deliberately exploits geometry the planner already owns; it does not prepare extra candidate trajectories merely to fill a native batch. Returned contact markers are conservative and do not alter the caller's incoming clearance.

C assumes every supplied body is alive. A clear result can therefore certify the ordinary body pass without invoking combat. Any possible contact re-enters the full ordered Python group pass with the incoming clearance unchanged. That pass may advance combat and credit a destruction only according to the existing rules. After combat advances, subsequent body checks use Python's death-aware pass. Independent projectiles remain dangerous after their source is destroyed.

Homing, wall checks, bottom entries, refuge boundaries, shell bursts, aimed shots, barrier shots and pickup contacts retain their original per-frame ordering. C results for a later frame are consumed only if all earlier checks passed. Packed body results never certify terrain or assume a destructible wall disappeared.

Validation

  • Full tools suite: 838 tests passed, with the native library built.
  • C/Python enum values, struct sizes and field offsets are checked.
  • Native differential tests cover swept bounds, inclusive touching, camera correction, eye margins, mixed specialized slices, empty scenes, malformed ranges/capacities, and unchanged output on invalid input.
  • Integration tests compare full maneuver results and transition counts for early rejection, pending input and destruction resolved by an earlier shot.
  • Native and WASM match 160 scenes / 1,280 body queries generated by Python.
  • The rebuilt WASM module also matches 654,496 existing eye points in 1,316 cases. This verifies the shared C source/ABI, not a live browser campaign.
./xenon_tools/hatari_dev.ps1 -Action build-kernel
python xenon_tools/native_body_vectors.py --output xenon_tools/run_logs/native-body-vectors-0906-01.json
cmake --build out/build/autopilot-wasm
node xenon_tools/check_native_bodies_wasm.mjs out/build/autopilot-wasm/xenon_autopilot_wasm.js xenon_tools/run_logs/native-body-vectors-0906-01.json

Measurement method

compare_kernels.py --variants c-eye,c isolates this port: c-eye is a fresh- process benchmark control that keeps the C eye forecast and numeric Python group types, while disabling the new body scene/batching. It is not a production backend. The ordinary python variant tests the complete Python kernel selection. Both controls retain the same numeric group types, so body-kernel gains are not attributed to replacing strings or to the earlier eye port.

Each timing includes observation reconstruction, policy, planning, path/slice packing, ctypes calls and output conversion, excluding Tk/AVI/emulator work. Use two opposite-order passes pinned to the same permitted logical CPU. Source, asset and DLL hashes, sidecars, raw timings and per-frame counters are recorded. Actions, tactics, verification, contact frame/kind/key, kills, pickups and reported clearance are compared. cProfile is a separate run; its cumulative times overlap and its instrumented elapsed times are not whole-controller speed estimates.

Final repeated measurements and profiles are below.

Diagnostics

The replay/autoplay maneuver panel displays body paths, calls, queried frames, clear frames, reference fallback frames and packed/input/output bytes, separately for recorded diagnostics and current analysis. Recordings retain backend, ABI, DLL hash and kernel names. Timing rows also include clearance so ranking regressions cannot be hidden by an unchanged first joystick input.

Rejected implementation experiments

The first implementation speculatively generated eight motions to fill each batch. The initial mapping overlay repeatedly hashed long Candidate tuples; cProfile showed motion service time growing from about 8 seconds to 21 seconds. Replacing that overlay with dictionary copies removed much of the regression, but two opposite-order plain runs still showed no convincing overall gain: corridor 16.921 vs 16.687 ms, boss eye-only C 9.204 vs 9.372 ms. That version also packed geometry for rejected tails. It was removed, rather than enabled on the strength of fewer Python group-query counters. Raw experiments remain under native-bodies-*-0906-02 and native-bodies-corridor-profile-0906-01.

The final implementation batches only already-restored geometry. Its results must be assessed using the later profiles/timings below, not those rejected runs.

Final repeated timings and remaining cost

All runs use spans, all architecture features, identical recordings/sidecars, and two opposite-order passes pinned to one permitted logical CPU. Means include all preparation and ctypes overhead; p95 below is the average of each pass's p95.

Sample / measured frames Existing C eye + Python bodies New C selection Mean reduction p95 before / after
Corridor, 801–2555 (1,755) 17.000 ms 16.425 ms 3.38% 53.766 / 52.226 ms
Boss, 7633–8476 (844) 9.052 ms 8.737 ms 3.48% 22.727 / 22.456 ms
Level 5, 90160–90459 (300) 16.724 ms 16.729 ms unchanged (0 native body calls) 19.185 / 18.720 ms

The full Python boss kernel measured 12.847 ms, versus 8.737 ms for both C kernels together (32.0% lower). Most of that is the previously implemented eye port; the new body work contributes the incremental 3.5% above. Level 5 makes zero body-kernel calls, and its tiny mean difference is timing noise. The existing Level-5 unverified outcomes remain unchanged; this is not evidence of new Level-5 tactics or successful gameplay progression.

Every comparison matches actions, tactics, verification, contact frame/kind/key, kills, pickups and clearance across 2,899 sampled observations, in both passes. No new live campaign or phone/browser-driver benchmark was run.

These are modest gains, not an order-of-magnitude improvement. Both corridor passes improve in the same direction, but the effect is small enough that device and workload changes can alter it. Keep the Python option and compare on target hardware before choosing a deployment default.

Raw timing reports (including hashes, frame differences and repeated-run spread):

Separate final cProfile run

Corridor profile: native-bodies-corridor-profile-0906-03/{c-eye,c}-1.pstats. All 1,755 compared decisions and clearances match. cProfile total function time is 80.390 vs 78.805 seconds; use the plain timings above for speedups.

Corridor cumulative function cost Python bodies C body selection
query_groups 17.340 s / 202,922 calls 7.357 s / 131,102 calls
group_slices (all callers) 13.896 s 12.783 s
motion 8.465 s / 511,161 calls 8.534 s / 511,161 calls
Native batch check (including frame preparation) — 8.297 s / 7,617 calls
frame_slices (inside check) — 7.938 s

The limiting cost is still Python geometry preparation. Moving the final contact/distance scan to C removes many Python queries, but specialized group_slices prediction and packing consume most of that saving. C does not yet own those motion models. Do not interpret the reduction in query calls as an equivalent reduction in total processing time. Motion calls and transition budgets are unchanged; the discarded speculative version failed that work-cost objective. Cumulative costs overlap and must not be added.

Average body boundary traffic per observation (zero-call observations included):

Sample C body calls Queried hulls Packed paths + slices Query + result bytes
Corridor 4.34 34.13 16235 2458
Boss 5.61 44.07 9665 3173

Packed scene bytes remain in caller-owned native buffers for the observation; only pointers/counts cross the call boundary. The byte counters measure prepared payload, not additional per-call copying or peak Python heap use. The standalone WASM module is 11,179 bytes in this build, excluding its JavaScript loader.

Boss profile: native-bodies-boss-profile-0906-04/{c-eye,c}-1.pstats. All 844 decisions/clearances match. cProfile total is 20.543 vs 19.763 seconds. query_groups falls from 4.960 to 1.331 seconds, while native batch preparation and checking costs 3.040 seconds (2.888 in frame_slices). Motion calls remain exactly 273,199. This independently confirms the same limitation: Python slice preparation consumes most of the collision-query saving.