Xenon 2

Autopilot · write-up

Native combat-bound preparation

xenondoc/AUTOPILOT_NATIVE_COMBAT_PREPARATION.MD · 7 KB · updated 2026-09-17

Baseline: 42334e6f; native ABI 15 adds XapCombatPrediction and xap_prepare_combat_bounds. The Python driver remains supported. This does not integrate with Hatari's driver or profile WASM.

Work removed

PredictionService.body_bounds reuses immutable rectangles when motion or camera correction is zero. packed_bounds avoids validating and constructing the same before/after rectangle twice. These skips use the existing observation cache. An exploratory separate Python stationary-bound cache was slower and was removed.

For remaining combat sources, Python supplies one descriptor per supported object: its observed hull, target/camera flags, output slot, and either constant velocity (including stationary) or a borrowed shared trajectory. The existing prediction dispatch establishes the motion model. Unsupported linked, changing-hull and specialized motion models retain their Python preparation.

One C call fills supported slots for the normal planning horizon when combat is first needed. It uses the caller's bounds buffer, never calls back into Python, and retains no pointers. All candidates share these immutable prepared bounds; target health, bullets and death frames remain per-rollout state.

Source order is unchanged, including holes left for Python fallback sources. Ordinary combat bodies keep separate before/after hulls. Combat-only blockers retain the union of their observed, before and after hulls. These extra blockers are never added to the ship-collision scene. Candidate camera correction is still applied by combat execution, not baked in twice during preparation.

Bounds use fixed frame-major slots when a native batch exists. Requests beyond the prepared horizon use reference preparation, and buffer growth preserves previous slots. No caller-owned path buffers need to stay alive after this preparation call; the resulting bounds buffer belongs to the observation.

Controls and diagnostics

  • c-combat-geometry-old: previous translations and per-frame Python packing.
  • c-combat-geometry-reuse: zero-motion/zero-camera and identical-hull skips only.
  • c: those skips plus native batch preparation.

These fresh-process variants are available in compare_kernels.py. Implementation controls are combat_preparation.REUSE and combat_preparation.NATIVE. Replay/live summaries show native_combat_preparation_calls, native_combat_geometry_sources, and native_combat_prepared_bounds. native_combat_bound_bytes continues to describe consumed frame slots, so it does not count native preparation as another copy or omit the remaining fallback slots.

Build through xenon_tools/hatari_dev.ps1 -Action build-kernel. The new function is also listed in standalone WASM exports; no WASM timing was run.

Validation

The full suite passes 907 tests. New tests compare mixed linear/shared paths, conservative unions, target indices and camera flags with Python; verify unchanged output-slot order and untouched fallback holes; reject invalid ranges before writing; and exercise zero translations, buffer growth, out-of-order requests and requests beyond the native horizon. The older synthetic combat test explicitly selects reference geometry because its fixture overrides the Python bounds method.

Restarted unprofiled measurements

The user reported concurrent PC work during the initial runs. Those timings are excluded from the final estimate. After the user authorized restarting, fresh workers ran sequentially on logical CPU 4, with reverse order on the second repeat. All variants used the same source/library fingerprints and replay inputs.

Level 1 preparation Run 1 mean ms Run 2 mean ms Average mean ms
Previous preparation 4.077 4.087 4.082
Translation/identical-hull reuse only 4.079 4.052 4.065
Reuse plus native batch 3.904 3.902 3.903

The complete replay pipeline is 4.4% faster, saving about 0.179 ms per observation. Mean p95 falls from 10.517 to 9.939 ms (5.5% lower). Reuse alone shows only a small 0.4% average difference, within the scale of run variation; most of the measured improvement comes from native batch preparation. The largest within-variant spread is 0.68%, and the native runs differ by only 0.04%.

Mean planning time over gameplay frames drops from 2.393 to 2.174 ms (9.2%). The native batch averages 0.322 calls and 0.499 supported source descriptors per gameplay frame, preparing 27.964 bounds. Combat remains lazy: observations that do not need combat do not pay for this preparation.

All 7,693 Level 1 observations (5,425 gameplay frames) retain identical verdicts across both repeats: action, tactic, movement authority, clearance, contacts, predicted kills and pickups. Available Level 2-5 recordings also retain identical verdicts in single old/new comparisons:

Recording Previous mean ms Native preparation mean ms
Level 2 6.764 6.391
Level 3 8.873 8.492
Level 4 6.992 6.780
Level 5 2.719 2.558

These are desktop sequential replay measurements, without emulator execution, drawing, sockets, or WASM. Cross-level single-run differences are validation samples, not repeated speedup estimates or new live level-completion claims.

Separate restarted profile

All instrumented variants also retain identical replay verdicts. Cumulative times below overlap; they must not be added or substituted for unprofiled timings.

Measurement Previous Reuse only Native batch
Total recorded function calls 152,517,829 153,469,918 146,303,947
Python combat-bound packing calls 159,208 159,208 7,504
Python body-bound requests 159,208 159,208 7,504
Rectangle translation calls 1,204,027 1,195,515 898,942
Combat frame-bounds, cumulative s 3.685 3.605 0.294
Candidate combat preparation, cumulative s 3.821 3.750 0.602

Native batch setup adds 0.149 cumulative seconds, including descriptor discovery and the C call. Python packing calls fall by 95.3%, and total recorded calls fall by about 6.21 million. Remaining unsupported bounds still use Python; there are no per-frame Python callbacks from the new C preparation function. The reuse-only changes do not establish a meaningful standalone speedup; the native batch supplies the useful reduction in work.

Artifacts

  • xenon_tools/run_logs/combat-geometry-restarted-timing-0907/comparison.json: fresh three-way stage isolation, two repeats after the user's restart.
  • xenon_tools/run_logs/combat-geometry-restarted-level*-0907/comparison.json: Level 2-5 verdict comparisons.
  • xenon_tools/run_logs/combat-geometry-restarted-profile-0907/: separate instrumented comparisons.
  • xenon_tools/run_logs/combat-geometry-final-tests-0907.txt: 907 passing tests.

Pre-restart timing folders are retained for provenance but are not used for the reported final speedup. The exploratory Python stationary cache was removed before these restarted measurements.