Autopilot · write-up
Native combat rollout (2026-09-06)
The c kernel selection now includes ordinary-shot combat and destruction-aware
body checks. The Python kernel remains selectable and the reference combat
implementation stays in autopilot_v2/combat.py. This extends commit f739676d.
Profiling confirms substantially less Python combat work. Overall controller speedup is not yet established reliably: the first corridor comparison improved, but the final repeat had substantial control timing drift (details below).
Subsequent clean live testing after the ABI-6 simplification is documented in AUTOPILOT_NATIVE_LIVE_VALIDATION.MD: across Level 1, all current C kernels together reduce mean gameplay decision time by 15% with identical observed gameplay. This measures the complete native backend, not the isolated contribution of the combat port in the earlier comparisons.
Scope and ownership
src/autopilot/xenon_autopilot_combat.c provides xap_advance_combat. It accepts
ordinary named C structs and pointers to caller-owned buffers. Python calls it
through ctypes; a later C driver can construct the same structs directly. There
is no JavaScript-specific interface, hidden allocation, callback into Python,
global combat state, or C-side ownership transfer.
Only the standalone kernel library/module is built in this work. No new Hatari driver integration, emulator call sites or emulator rebuild are included. WASM is compile-checked for portability, not profiled.
XapCombatState refers to the caller's bullet, target and scratch arrays. C
compacts surviving bullets in place and updates health, damage and death frames.
The caller can read those arrays directly. XapCombatScene borrows the observation's
existing path and linear descriptors plus cached specialized before/after bounds.
XapCombatFrame supplies the requested frame, camera/player coordinates, geometry
ranges and actual forecast fire edges. Maneuver-dependent homing bounds remain
separate from the shared observation data.
The ABI is 6. Body descriptors carry an observation-local target index;
path/linear/slice sizes are 56/48/48 bytes. A negative index means no predicted
damageable target. Native pointer-containing scene/state structs follow the
calling platform's ordinary pointer layout; ctypes mirrors that layout. They
are not a serialized file format. XapCombatTarget is now 12 bytes; the unused
kill-order field has been removed. Rebuild old kernel DLLs.
Simplification after the initial port
Validation and execution sections are marked inside the C API functions. Combat validates initial target/bullet state on the first advance, then trusts state maintained by C. Every call still checks buffer capacities, ranges, frame inputs and target indices before mutation. Cached combat bounds are checked when Python packs each new entry; a future direct C caller must meet the same documented preparation contract. New maneuver-dependent extras still receive C validation.
Python no longer mirrors updated health or rebuilds and sorts a death dictionary after every call. Membership uses a fixed target-key set, absence checks read the ctypes target array directly, and the kill list is constructed only for the final evaluation. Kill ordering is stable by target key, not Python insertion order.
Already-expired targets are filtered during frame preparation rather than checked again for every bullet. Enemies killed during the current frame remain blockers through their final destroy update. Each bullet's coupled/uncoupled world Y is computed once outside the body loop. The simple bullet/body scan is retained.
Avoiding unnecessary work
- Combat is still lazy. Safe candidates create no native combat state. A collision conflict or shortlisted attack requests only the frames already needed by the maneuver; one native call catches up all of those frames.
- Paths and linear bodies generate raw damage bounds in C, using descriptors already packed for body queries. Specialized independent bounds are prepared once per requested observation/frame and reused across combat rollouts.
- C prepares the frame's geometry once for all bullets and skips that preparation if there are no bullets to move. It stops searching blockers after a second overlapping blocker makes damage attribution ambiguous.
- There are no per-bullet or per-target ctypes calls. After combat starts, a later single-frame request can still require one call; no speculative future actions or combat frames are generated to fill a batch.
- Body queries now read the rollout's death frames instead of reverting every subsequent frame to Python. Death-aware results never enter shared all-alive caches. Fresh motions can use this path after combat has started.
Preserved gameplay rules
Bullets are consumed by any possible blocker. Damage is credited only for a unique supported target hit in both collision phases, with the existing inset for animated targets. New shots move on the next frame. Actual forecast fire edges and observed weapon upgrades determine shot creation/damage.
A destroyed enemy remains a collider through its final destroy update. Independent projectiles have no target index and remain dangerous after their source dies. Unknown colliders still block shots without receiving predicted damage. Predicted kills do not remove wall tiles. Navigation, destructible-wall observation, level tactics, homing motion, aimed shots and pickups remain under the existing Python logic.
A native possible-contact flag still permits the ordered Python conflict resolver to identify the collision and request combat when needed. Thus this port does not claim to remove all Python fallback or all specialized geometry preparation.
Validation and measurement
After simplification, the complete tools suite passes 857 tests, including eight combat tests, with the native DLL built. The standalone native and WASM sources compile cleanly. Native and WASM body queries match Python on 160 scenes / 1,280 queries with the ABI-6 layouts. WASM combat is compile-checked; no WASM performance measurements are made. The combat differential tests exercise the native DLL through ctypes.
Native tests compare batched and incremental C combat with Python across shot movement, firing, overlapping blockers, animated hitboxes, health, deaths and surviving bullets. A mixed-geometry case checks path, linear and maneuver-dependent bounds together with camera scrolling. Integration tests verify a shot resolving an earlier contact, continued native body checks after combat, final destroy-update timing, projectiles outliving their source and no combat allocation on a safe candidate. All four combat geometry sources retain the final destroy update and omit expired targets. Invalid initial state, batches and insufficient caller capacity are rejected before state mutation; invalid newly packed cached bounds are rejected in Python.
The native ABI supports at most 1,024 forecast frames. The driver retains Python
combat for configurations exceeding that limit. Ordinary supported forecasts use
native combat with --kernel-backend c; fractional fitted velocities do not cause
a precision-based fallback.
The fresh-process control compare_kernels.py --variants c-no-combat,c retains
the existing C eye/body/linear kernels in both cases and disables only this combat
port in the control. Production choices remain --kernel-backend python and
--kernel-backend c. Plain timings and cProfile attribution are separate.
The shared autoplay/replay summary shows native_combat_rollouts,
native_combat_calls, native_combat_frames and native_combat_body_queries.
Detailed counters include cached combat-bound bytes and frame/extra input bytes.
Recording metadata identifies ABI 6 and the combat kernel.
The post-simplification corridor check (frames 801-2555) matches all 1,755
observation verdicts against c-no-combat. Its single plain pass measures
15.266 ms for Python combat versus 14.588 ms for native combat, with p95
50.401 versus 41.343 ms. This is a correctness revalidation with indicative
timing, not an isolated before/after measurement of the simplifications or a
replacement for the repeated-timing caveat below.
Report
./xenon_tools/hatari_dev.ps1 -Action build-kernel
python xenon_tools/autoplay.py --planner maneuver --navigation-backend spans --kernel-backend c
python xenon_tools/autoplay.py --planner maneuver --navigation-backend spans --kernel-backend python
First repeated plain comparison
The following measurements describe the original ABI-5 port, before the above simplifications; they do not measure the incremental effect of this cleanup.
Two opposite-order passes pinned to one permitted logical CPU use spans and all architecture features. Timings include observation reconstruction, policy, planning, preparation and ctypes. They exclude rendering/emulator work.
| Sample / frames | C kernels with Python combat | With native combat | Interpretation |
|---|---|---|---|
| Corridor, 801-2555 (1,755) | 16.417 ms | 15.000 ms | 8.63% lower mean |
| Boss, 7633-8476 (844) | 7.724 ms | 7.869 ms | zero combat calls; no demonstrated change |
| Level 5, 90160-90459 (300) | 19.289 ms | 16.677 ms | zero combat calls; timing drift, not a gain |
The corridor improves 7.57% and 9.56% in the individual passes. Average pass p95 falls from 56.144 to 43.907 ms. Absolute times vary between passes (14.2% for the control); the final repeat below prevents treating this as a reliable speedup. Level 5 exceeds the comparator's 15% drift threshold and cannot support a timing conclusion. No benefit is attributed to workloads where combat did not execute.
Every comparison matches the selected action, tactic, authority, verification, contact frame/kind/key, clearance, predicted kills and pickups across all 2,899 observations, in both passes. This is recorded-observation replay, not a newly run live campaign.
Corridor averages: 0.979 native combat rollouts and 1.334 calls advancing 28.272 combat frames per observation (about 21 frames per call). The total combat-step count is identical to Python. Shared fallback damage-bound payload averages 683 bytes per observation; frame/extra input averages 1,583 bytes. Those are prepared buffer sizes, not repeated copies across the C boundary. Native body checks after combat average 0.356 calls per observation.
Raw reports include per-frame results, counters, input/source/DLL hashes and both timing passes:
Final repeated plain comparison
The final-source corridor repeat again matches all 1,755 observation verdicts in both passes, but its timings do not reproduce a consistent overall gain:
| Pass | C kernels with Python combat | With native combat |
|---|---|---|
| 1 | 20.662 ms | 19.173 ms |
| 2 (reverse order) | 16.017 ms | 19.187 ms |
| Average | 18.340 ms | 19.180 ms |
The control varies by 29%, exceeding the comparator's drift threshold. The first pass improves, the second regresses, and the average is 4.6% slower. These results cannot establish either a reliable overall gain or its size. The earlier 8.63% improvement should therefore be treated as provisional, not the final performance claim. The demonstrated result is removal of repeated Python combat work; stable end-to-end timing remains to be established.
Separate native/Python cProfile attribution
native-combat-corridor-profile-0906-01/{c-no-combat,c}-1.pstats compares the
same 1,755 observations. All compared verdicts match. Cumulative times overlap;
use the plain timings for controller speedups.
| Function / work | Python combat control | Native combat |
|---|---|---|
| Combat catch-up, including preparation | 6.906 s / 2,374 calls | 0.690 s / 2,374 calls |
Python CombatRollout.advance |
3.853 s / 49,617 calls | 0 calls |
Python body_bounds |
1.944 s / 938,984 calls | 0.131 s / 16,639 calls |
| Native combat adapter, including C calls | - | 0.531 s / 2,342 calls |
group_slices |
7.523 s | 7.380 s |
Native body frame_slices |
4.061 s | 4.056 s |
Player motion calls |
511,161 | 511,161 |
| Total profiled function time | 68.568 s | 63.076 s |
The port removes the Python bullet loop and most repeated damage-bound requests. It simulates exactly the same 49,617 combat frames, without speculative player motions. Specialized body-query slice preparation remains a substantial Python cost; the combat port does not remove it. Minor final input-validation/comment cleanup followed this profile; final plain measurements use the final source.