Xenon 2

Autopilot · write-up

Native evaluation of scripted swarm threats

xenondoc/AUTOPILOT_NATIVE_POLICY_PATHS.MD · 7 KB · updated 2026-09-17

Baseline: ba88bec8; native ABI 13 adds xap_first_path_contact.

Expensive work selected

The preceding Level 1 profile spent 5.308 cumulative seconds in the shared policy swarm query. It still traversed future frames and swarm objects in Python, invoking the hazard predictor and constructing expanded swept rectangles. This is prediction evaluation used by tactical policy, rather than host decoding or display geometry.

The new path packs supported scripted swarm trajectories once per observation and passes all unchecked neutral-player frames to C in one call. Script displacements come from the existing observation cache/tapes, so scripts are not reinterpreted for each query. The C evaluator first rejects complete paths outside the lane envelope, then checks nearby paths only until the earliest known contact.

Boundaries and semantics

  • Python owns named XapHazardPath, XapPoint and XapBodyQuery buffers; C retains no pointers and requires no Python callbacks, allocation or emulator symbols.
  • Ship and enemy safety margins remain 3 and 2 pixels respectively. Policy contact uses strict rectangle overlap; candidate collision's inclusive touching rule is intentionally separate.
  • World-fixed and camera-coupled paths use separate lane envelopes. The query returns the first future contact frame, zero for clear and -1 for invalid input.
  • Only scripted swarm procedures whose displacement contract is known use this path. Pending spawns, linked sources and unsupported models keep the reference predictor. Mixed queries select the earliest native or reference contact.
  • The policy remains an all-alive threat estimate. Ordered candidate evaluation still handles weapon fire and destruction; this step changes no combat rules.
  • Previously checked prefixes and shorter repeated queries preserve existing caching behavior. Buffers belong to the observation and are replaced on the next policy decision.

native_policy_threat.ENABLED = False retains the Python evaluation loop. compare_kernels.py --variants c-policy-paths-old,c isolates this change while keeping the other native optimizations enabled. The Python kernel remains selectable and uses the reference threat loop. Replay/live metrics include native_policy_swarm_paths, native_policy_swarm_queries and native_policy_swarm_query_frames; these count native preparation and submitted queries, not Python transition-budget charges.

The library is built through hatari_dev.ps1 -Action build-kernel. The export is available to ctypes and listed for future WASM builds. This task does not integrate the kernel into Hatari's driver or profile WASM.

Validation

The full suite passes 896 tests. New tests compare random paths with strict Python rectangle intersections, including camera correction, inactive prefixes, query gaps and distant paths. Additional checks cover touching edges, invalid indices, mixed native/reference sources and repeated/extended horizons.

Timing and replay comparison

xenon_tools/run_logs/native-policy-paths-l1-0907/comparison.json contains two fresh pinned runs per variant, with the order reversed for the second repeat. All 7,693 observations (5,425 gameplay frames) retain the same planner verdicts in all four runs.

Policy swarm evaluator Run 1 mean ms Run 2 mean ms Average mean ms
Python loop with native planner 4.723 4.736 4.730
Native path query 4.595 4.590 4.592

Complete replay processing is 2.9% faster (approximately 0.138 ms per observation). Mean policy selection over gameplay frames decreases from about 1.861 to 1.723 ms, 7.4% lower. The optimization is useful but does not remove the rest of tactical policy, script execution, scene preparation or candidate evaluation. Timings exclude drawing, sockets and emulator execution.

Each Level 1 run prepares 16,600 native swarm paths and submits 4,723 queries covering 52,214 future frames. These counts describe submitted work, not the number of contacts or native rectangle comparisons after early rejection.

Single-pass cross-level comparisons also retain all planner verdicts:

Recording level Observations / gameplay Old mean ms New mean ms
2 961 / 910 6.887 6.916
3 225 / 225 10.008 9.662
4 900 / 752 7.103 7.050
5 5,652 / 391 2.753 2.735

Levels without many supported swarm queries show essentially flat results; one pass is insufficient to establish small differences. Level 5 is mostly inactive observations, so its overall mean is not its active-combat cost. Artifacts are in run_logs/native-policy-paths-<recording-stem>-0907/comparison.json, using stems level2-homing-dual-collision-0824-60, level3-stage2-smallshot-0826-213, level4-wall-model-0827-07 and level5-tank-exact-homing-dev-0901-157. These are offline replay checks, not fresh closed-loop completion or phone tests.

Profile evidence

The separate native-policy-paths-profile-0907 comparison also retains all planner verdicts. Instrumented cumulative times describe where work moved, not the real-time gain:

Profile measure Reference policy loop Native policy paths
Swarm query cumulative seconds 5.330 2.853
Policy selection cumulative seconds 23.987 21.505
Adapter hazard-step calls 631,082 393,604
Total recorded function calls 178,051,821 171,429,641

Native swarm preparation costs 1.546 cumulative seconds over 3,477 scene builds; 4,723 batched queries, including Python lane preparation, cost 0.307 seconds. Script forecasts and native buffer preparation therefore remain material work. This stage roughly halves the profiled swarm-query cost, while the measured total processing gain remains 2.9%. The 23 replay UI tests also pass with the final counter display.

Reproduction

Run fresh timing workers sequentially, reversing order in the second repeat. Keep profiling separate from unprofiled timings.

powershell -ExecutionPolicy Bypass -File xenon_tools/hatari_dev.ps1 -Action build-kernel
python xenon_tools/compare_kernels.py xenon_tools/run_logs/native-python-live-c-0906-02.validation/native-python-live-c-0906-02.x2events --output <timings> --variants c-policy-paths-old,c --from-frame 2367 --to-frame 10059 --repeats 2 --pin-cpu
python xenon_tools/compare_kernels.py xenon_tools/run_logs/native-python-live-c-0906-02.validation/native-python-live-c-0906-02.x2events --output <profile> --variants c-policy-paths-old,c --from-frame 2367 --to-frame 10059 --repeats 1 --pin-cpu --cprofile
python -m unittest discover -s xenon_tools -p 'test_*.py'