Autopilot · write-up
Native linear prediction and collision-shape preparation (2026-09-06)
This document records the linear port and its ABI-4 cleanup. The subsequent combat extension uses ABI 5, adds target indices to body descriptors and retains the Python driver. Measurements below remain the historical linear-port results.
The previous typed body-query implementation was committed first as
16334ce0 (Add typed batched native autopilot body validation). This extension
removes repeated constant-velocity fitting and per-frame body-shape packing.
It retains the Python backend and the ordered Python destruction fallback.
Two separately measured changes
- Reuse motion parameters in Python.
autopilot_v2/linear_prediction.pycomputes an eligible hazard's observed collision rectangle and stable world velocity once per observation. Future steps reuse those parameters with the original floating-point arithmetic. This is enabled byprediction_reusefor both kernel backends. The cache is replaced at each observation; replay snapshots cannot accidentally acquire another frame's parameters. - Prepare linear sweeps in C. The C backend packs a 40-byte
XapLinearBodyonce per eligible body per observation. C constructs its future swept bounds while checking the existing batches of up to eight player hulls. Those bodies no longer require Python Rect/Group construction and slice packing for each native frame. Existing scripted formation and eye paths keep their path model.
Eligibility follows the original motion and collision-shape dispatch, through
two small overridable hooks in autoplay.py. A generic velocity alone is not
enough: an animated or specialized collision hull must retain its own model.
Scripted motion, linked predictions, active shell motion and extra safety
envelopes keep their reference paths. Probing eligibility does not eagerly build
scripted trajectories already covered by the formation broadphase.
The body-only exclusion in PredictionService.group_slices deliberately keeps
the source's anticipated shots. Moving a shooter body to C must not remove its
projectiles. Possible contacts still trigger the complete ordered Python pass,
where existing weapon simulation can resolve destruction. Walls, homing, level
tactics, cash and pickup decisions remain with their existing implementation.
Representation and ABI
The shared native/WASM ABI is now 4; rebuild the native DLL. Older DLLs are
rejected by the wrapper. All public inputs use named C structs mirrored by
ctypes.Structure.
XapLinearBody contains an integer XapIntRect, signed 64-bit Q32 X/Y velocity,
numeric hazard kind and camera-coupling flag. Q32 means 32 fractional bits: for
example, 1.5 pixels/frame is stored as 1.5 * 2**32. Static bodies use zero speed.
Bounds must be integral and within the documented coordinate/extent limits;
velocity is limited to 64 pixels/frame and the horizon to 1,024 frames.
Fitted velocities are rounded to the nearest Q32 value when packing the descriptor.
A value such as 1/3 is handled in C; it does not trigger Python preparation.
The maximum rounding drift is less than 0.00000012 pixels per axis over the
1,024-frame limit, so no extra margin or special precision fallback is needed.
Geometry generation uses integer/fixed-point arithmetic, then converts to the
common floating-point clearance representation. Existing
XapPoint, XapRect, generic slices and distance results retain doubles. This
extension does not claim that every predictor or distance calculation is integer.
Python owns the packed observation buffers. C borrows them for a call, allocates nothing and retains no pointers. No map transfer is needed: this kernel does not read maps. The caller still controls when an observation is replaced.
Simplified validation and results
xap_validate_body_scene(paths, path_count, points, point_count, linear, linear_count)
checks the immutable descriptors and all referenced trajectory points once after
packing. BodyScene calls it once per observation. Those buffers must remain
unchanged until the scene is discarded; another caller modifying them must
revalidate before querying. There is no opaque context, allocation, registration
token or retained native pointer.
xap_query_bodies retains pointer/count, path indexing, frame, output-capacity,
query-value and dynamic-slice checks per batch. It does not rescan the immutable
scene's coordinate values, velocity ranges or type flags. Overflow-sensitive
fixed-point ranges are established during scene validation. The caller's
immutability obligation is part of the native API contract.
XapBodyResult now contains only clearance and an int32 contact flag (0/1).
The unused contact index and hazard kind are removed. Actual contact identity
still comes from the ordered Python destruction-aware check when needed.
The result remains 16 bytes because of native alignment; this cleanup is not a
claimed reduction in returned bytes. Both ctypes and WASM use the same layout.
A prepared-buffer comparison against the saved ABI-3 DLL used 12 paths (972 points), 5 linear bodies, eight queries per batch and four opposite-order passes of 2,000 observations pinned to one CPU. It includes ctypes and the new validation call, but excludes packing and the rest of the autopilot:
| Batches per observation | Previous ABI 3 | Simplified ABI 4 |
|---|---|---|
| 1 | 2.587 microseconds | 4.105 microseconds |
| 7 | 21.518 microseconds | 21.136 microseconds |
The one-time call/full-path scan adds roughly 1.5 microseconds for a scene used only once. Seven batches are effectively unchanged at this scale. This is an interface/validation cleanup, not a demonstrated whole-controller speedup. Both versions returned identical contact flags and clearances for these queries. Raw samples and DLL hashes: body-simplification-micro-0906-01.json.
Native batching requires the formation broadphase and an exactly restored motion suffix of at least eight steps. A supported linear body is now sufficient to create a scene even when there is no supported full path. New search branches remain lazy Python checks; the port does not invent motions to fill batches.
Validation
- 849 tools tests pass with the native library available.
- Tests cover fractional fitted motion, per-observation invalidation, specialized
animated hulls, preservation of shots, C/Python struct layouts, rounded Q32
velocities (including
1/3using C), negative speeds, camera correction, horizon limits and unchanged output after invalid native input. - Scene validation covers future points beyond the first query batch and rejects invalid fixed-point ranges before querying. Multiple batches validate their shared scene only once; changing slice/query data still receives per-call checks.
- Native and WASM agree with Python on 160 mixed scenes / 1,280 queries, including paths, linear descriptors and generic slices.
- The rebuilt WASM module also matches 1,316 eye cases / 654,496 points.
Vectors: xenon_tools/run_logs/native-linear-vectors-0906-01.json and the existing
native-eye-vectors-0906-01.json. The ABI-4 body vectors are
native-simple-body-vectors-0906-01.json. WASM checks use
check_native_bodies_wasm.mjs and check_native_eye_wasm.mjs.
./xenon_tools/hatari_dev.ps1 -Action build-kernel
python -m unittest discover -s xenon_tools -p "test_*.py"
python xenon_tools/autoplay.py --planner maneuver --navigation-backend spans --kernel-backend c
Measurement method
compare_kernels.py --variants c-baseline,c-reuse,c separates the two changes:
c-baseline: previous C eye/body behavior, new parameter reuse disabled and no linear descriptors.c-reuse: Python parameter reuse enabled, no linear descriptors.c: parameter reuse plus native linear sweep preparation.
These are fresh-process benchmark controls, not extra production backends. The
two controls use the same current DLL with zero linear descriptors. Thus the timing
comparison isolates the new stages instead of crediting them for previous C
ports. Use --kernel-backend python or --kernel-backend c in production.
Plain runs use two opposite-order passes pinned to one permitted logical CPU, spans navigation, all architecture features, identical recordings/checkpoints and terrain sidecars. Timings include reconstruction, policy, planning, packing and ctypes overhead; they exclude emulator/rendering work. Source, DLL and input hashes are recorded. cProfile is run separately and is used only to attribute work, not to estimate end-to-end gains.
For example, from xenon_tools, using a new output directory for each run:
python compare_kernels.py run_logs/prediction-reuse-corridor-0906-02.validation/prediction-reuse-corridor-0906-02.x2events --from-frame 801 --to-frame 2555 --variants c-baseline,c-reuse,c --repeats 2 --pin-cpu --output run_logs/linear-comparison
Add --cprofile, use --repeats 1 and a separate output directory for attribution.
Repeated plain timings
These ABI-3 measurements precede the simplification to rounded velocity conversion and the ABI-4 result/validation cleanup. They used an exact-conversion gate, since removed. Historical timings and exact replay matches below are retained as measured; they have not been remeasured for the rounding change.
Means below average the two opposite-order passes. Percentages compare each stage with the preceding stage; combined reduction compares full C with baseline.
| Sample / frames | Baseline | Parameter reuse | Reuse + native sweeps | Reuse reduction | Additional C reduction | Combined reduction |
|---|---|---|---|---|---|---|
| Corridor, 801-2555 (1,755) | 15.975 ms | 15.375 ms | 14.838 ms | 3.76% | 3.49% | 7.12% |
| Boss, 7633-8476 (844) | 8.558 ms | 8.072 ms | 7.569 ms | 5.68% | 6.24% | 11.56% |
| Level 5, 90160-90459 (300) | 16.317 ms | 16.210 ms | 15.746 ms | inconclusive | not exercised | inconclusive |
Both corridor and boss passes improve at each stage. The Level-5 sample makes
zero native body calls: c-reuse and c execute the same relevant work there,
yet their timing differs. The reuse pass also varies by 3.7%. Treat those Level-5
differences as noise, not a native speedup or a demonstrated parameter-reuse gain.
The average pass p95 changes from 51.218 to 49.664 ms in the corridor and 22.607 to 22.940 ms at the boss. This is a mean-time improvement, not evidence of consistently improved tail latency. The remaining slower search/repair frames still need Python validation and specialized preparation.
Actions, tactics, authority, verification, contact frame/kind/key, kills, pickups and reported clearance match across 2,899 observations, in both passes. This is replay comparison against the observed trajectories, not a new closed-loop campaign or phone/browser-driver benchmark. Existing Level-5 unverified outcomes remain unchanged; its special tactics are still separate work.
Reports with per-frame timings, counters, differences and fingerprints:
Work and boundary size
Per-observation averages include frames that do not use the native scene.
| Counter | Corridor baseline / new | Boss baseline / new |
|---|---|---|
| Python hazard slices | 235.911 / 173.958 | 125.308 / 57.364 |
| Native generic slice bytes | 4989.698 / 2736.957 | 3368.246 / 137.441 |
| Native linear descriptors | 0 / 1.632 | 0 / 1.521 |
| Native linear descriptor bytes | 0 / 65.276 | 0 / 60.853 |
| Body kernel calls | 4.340 / 5.076 | 5.607 / 5.617 |
| Queried player hulls | 34.133 / 39.922 | 44.073 / 44.158 |
Generic slice bytes fall 45% in the corridor and 96% at the boss. Existing path bytes remain; these reductions concern the generic slice payload, not the entire native scene or total Python memory. More corridor frames can use C because a linear-only scene is now supported. Calls remain batched, never one per body.
With Python reuse alone, there are 1.867 parameter builds and 98.696 reuse hits per corridor observation (boss: 1.827 builds and 98.577 hits). Adding native preparation reduces those hits to 37.151 and 31.916: much of the repeated Python step work itself is now absent. These counters refer to the generic motion fallback, not every use of velocity estimation elsewhere in the policy.
Separate cProfile attribution
Corridor profile: linear-kernel-corridor-profile-0906-01/{c-baseline,c-reuse,c}-1.pstats.
All compared verdicts match. The table reports cumulative times, which overlap
and must not be added. Use the plain timing table for controller speedups.
| Corridor function / work | Baseline | Reuse | Reuse + C |
|---|---|---|---|
_stable_world_velocity |
2.743 s / 361,397 calls | 0.129 s / 11,699 | 0.166 s / 11,699 |
group_slices |
11.869 s | 9.160 s | 7.227 s |
Native frame_slices |
7.183 s | 5.578 s | 3.865 s |
Native check, including preparation |
7.464 s | 5.845 s | 4.162 s |
| Reference generic sweep calls | 374,086 | 374,086 | 266,075 |
Player motion calls |
511,161 | 511,161 | 511,161 |
| Total profiled function time | 70.766 s | 67.908 s | 65.736 s |
Parameter reuse removes repeated fits. Native preparation then removes 108,011 generic Python sweep calls and more packing work, while player-motion work stays identical. The remaining native preparation time includes specialized models and anticipated shots still supplied as Python slices. This port reduces a measured cost; it does not eliminate heterogeneous Python prediction or search costs.
Boss profile: linear-kernel-boss-profile-0906-01/{c-baseline,c-reuse,c}-1.pstats.
All 844 compared verdicts match. It independently confirms the same pattern:
| Boss function / work | Baseline | Reuse | Reuse + C |
|---|---|---|---|
_stable_world_velocity |
0.955 s / 169,780 calls | 0.019 s / 1,840 | 0.020 s / 1,840 |
group_slices |
3.708 s | 2.524 s | 1.463 s |
Native frame_slices |
2.947 s | 1.932 s | 0.533 s |
Native check, including preparation |
3.099 s | 2.080 s | 0.661 s |
| Reference generic sweep calls | 102,669 | 102,669 | 46,407 |
Player motion calls |
273,199 | 273,199 | 273,199 |
| Total profiled function time | 19.680 s | 18.541 s | 17.391 s |
Replay diagnostics
The shared autoplay/replay maneuver summary now includes
linear_parameter_builds, linear_parameter_hits, native_linear_bodies and
native_linear_bytes. Detailed counters also record unsupported native inputs
and the existing body calls, frames and slice bytes. Recording metadata identifies
ABI 4 and the linear_bodies kernel alongside eye_chain and body_queries.
Historical recorded diagnostics remain distinct from current replay analysis.
The shared summary also shows native_body_scene_validations alongside query
calls, so repeated batches can be distinguished from scene preparation.