Autopilot · write-up
Native combat-bound preparation
Baseline: 42334e6f; native ABI 15 adds XapCombatPrediction and
xap_prepare_combat_bounds. The Python driver remains supported. This does not
integrate with Hatari's driver or profile WASM.
Work removed
PredictionService.body_bounds reuses immutable rectangles when motion or camera
correction is zero. packed_bounds avoids validating and constructing the same
before/after rectangle twice. These skips use the existing observation cache.
An exploratory separate Python stationary-bound cache was slower and was removed.
For remaining combat sources, Python supplies one descriptor per supported object: its observed hull, target/camera flags, output slot, and either constant velocity (including stationary) or a borrowed shared trajectory. The existing prediction dispatch establishes the motion model. Unsupported linked, changing-hull and specialized motion models retain their Python preparation.
One C call fills supported slots for the normal planning horizon when combat is first needed. It uses the caller's bounds buffer, never calls back into Python, and retains no pointers. All candidates share these immutable prepared bounds; target health, bullets and death frames remain per-rollout state.
Source order is unchanged, including holes left for Python fallback sources. Ordinary combat bodies keep separate before/after hulls. Combat-only blockers retain the union of their observed, before and after hulls. These extra blockers are never added to the ship-collision scene. Candidate camera correction is still applied by combat execution, not baked in twice during preparation.
Bounds use fixed frame-major slots when a native batch exists. Requests beyond the prepared horizon use reference preparation, and buffer growth preserves previous slots. No caller-owned path buffers need to stay alive after this preparation call; the resulting bounds buffer belongs to the observation.
Controls and diagnostics
c-combat-geometry-old: previous translations and per-frame Python packing.c-combat-geometry-reuse: zero-motion/zero-camera and identical-hull skips only.c: those skips plus native batch preparation.
These fresh-process variants are available in compare_kernels.py. Implementation
controls are combat_preparation.REUSE and combat_preparation.NATIVE.
Replay/live summaries show native_combat_preparation_calls,
native_combat_geometry_sources, and native_combat_prepared_bounds.
native_combat_bound_bytes continues to describe consumed frame slots, so it does
not count native preparation as another copy or omit the remaining fallback slots.
Build through xenon_tools/hatari_dev.ps1 -Action build-kernel. The new function is
also listed in standalone WASM exports; no WASM timing was run.
Validation
The full suite passes 907 tests. New tests compare mixed linear/shared paths, conservative unions, target indices and camera flags with Python; verify unchanged output-slot order and untouched fallback holes; reject invalid ranges before writing; and exercise zero translations, buffer growth, out-of-order requests and requests beyond the native horizon. The older synthetic combat test explicitly selects reference geometry because its fixture overrides the Python bounds method.
Restarted unprofiled measurements
The user reported concurrent PC work during the initial runs. Those timings are excluded from the final estimate. After the user authorized restarting, fresh workers ran sequentially on logical CPU 4, with reverse order on the second repeat. All variants used the same source/library fingerprints and replay inputs.
| Level 1 preparation | Run 1 mean ms | Run 2 mean ms | Average mean ms |
|---|---|---|---|
| Previous preparation | 4.077 | 4.087 | 4.082 |
| Translation/identical-hull reuse only | 4.079 | 4.052 | 4.065 |
| Reuse plus native batch | 3.904 | 3.902 | 3.903 |
The complete replay pipeline is 4.4% faster, saving about 0.179 ms per observation. Mean p95 falls from 10.517 to 9.939 ms (5.5% lower). Reuse alone shows only a small 0.4% average difference, within the scale of run variation; most of the measured improvement comes from native batch preparation. The largest within-variant spread is 0.68%, and the native runs differ by only 0.04%.
Mean planning time over gameplay frames drops from 2.393 to 2.174 ms (9.2%). The native batch averages 0.322 calls and 0.499 supported source descriptors per gameplay frame, preparing 27.964 bounds. Combat remains lazy: observations that do not need combat do not pay for this preparation.
All 7,693 Level 1 observations (5,425 gameplay frames) retain identical verdicts across both repeats: action, tactic, movement authority, clearance, contacts, predicted kills and pickups. Available Level 2-5 recordings also retain identical verdicts in single old/new comparisons:
| Recording | Previous mean ms | Native preparation mean ms |
|---|---|---|
| Level 2 | 6.764 | 6.391 |
| Level 3 | 8.873 | 8.492 |
| Level 4 | 6.992 | 6.780 |
| Level 5 | 2.719 | 2.558 |
These are desktop sequential replay measurements, without emulator execution, drawing, sockets, or WASM. Cross-level single-run differences are validation samples, not repeated speedup estimates or new live level-completion claims.
Separate restarted profile
All instrumented variants also retain identical replay verdicts. Cumulative times below overlap; they must not be added or substituted for unprofiled timings.
| Measurement | Previous | Reuse only | Native batch |
|---|---|---|---|
| Total recorded function calls | 152,517,829 | 153,469,918 | 146,303,947 |
| Python combat-bound packing calls | 159,208 | 159,208 | 7,504 |
| Python body-bound requests | 159,208 | 159,208 | 7,504 |
| Rectangle translation calls | 1,204,027 | 1,195,515 | 898,942 |
| Combat frame-bounds, cumulative s | 3.685 | 3.605 | 0.294 |
| Candidate combat preparation, cumulative s | 3.821 | 3.750 | 0.602 |
Native batch setup adds 0.149 cumulative seconds, including descriptor discovery and the C call. Python packing calls fall by 95.3%, and total recorded calls fall by about 6.21 million. Remaining unsupported bounds still use Python; there are no per-frame Python callbacks from the new C preparation function. The reuse-only changes do not establish a meaningful standalone speedup; the native batch supplies the useful reduction in work.
Artifacts
xenon_tools/run_logs/combat-geometry-restarted-timing-0907/comparison.json: fresh three-way stage isolation, two repeats after the user's restart.xenon_tools/run_logs/combat-geometry-restarted-level*-0907/comparison.json: Level 2-5 verdict comparisons.xenon_tools/run_logs/combat-geometry-restarted-profile-0907/: separate instrumented comparisons.xenon_tools/run_logs/combat-geometry-final-tests-0907.txt: 907 passing tests.
Pre-restart timing folders are retained for provenance but are not used for the reported final speedup. The exploratory Python stationary cache was removed before these restarted measurements.