Xenon 2

Autopilot · write-up

Native/Python live validation — 2026-09-06

xenondoc/AUTOPILOT_NATIVE_LIVE_VALIDATION.MD · 12 KB · updated 2026-09-17

Code under test: 793ad01f, ABI 6. Both optimized maneuver planners use spans, prediction reuse and all architecture features. This compares all currently ported kernels with their Python implementations, not with the legacy planner.

Method

The native and Python runs were sequential, at normal emulator speed with no shield cheat, both original/sprite AVIs, full decision traces and checkpoints every 500 canonical frames plus encounter milestones. They used the same existing Hatari executable; no emulator rebuild or new Hatari integration was involved. Python remained the driver for both runs. No WASM profiling was performed.

The first native attempt (native-python-live-c-0906-01) was stopped cleanly and excluded after the user reported heavy background processes. All results here are from the restarted pair after those processes were stopped.

The native run loaded assets/xenonplay.sav. Python loaded the native recording's archived start.sav, including its identity sidecar, to preserve the exact starting canonical frame counter and object identities. Both began at frame 2367 / VBL 13666, with shield 39, three lives and zero score/cash. The runs stopped on playable Level 2 at frame 10060. Each contains 7,693 decision rows, frames 2367–10059.

The initial snapshot has no saved controller sidecar; both planners reconstruct from the same Atari state. Per-level terrain RAM seeds and the native DLL are archived by the existing validation tooling. The isolated validation emulator was closed only after both recordings and final saves had been finalized.

Gameplay result

Both runs completed Level 1, its boss and both shops, reaching playable Level 2. The comparison found no differences on any of the 7,693 common frames in next action/input, tactic, movement authority, player identity/position, VBL, shield, lives, score, cash, level or scroll position. This is a live action comparison, not just replay proposals against fixed future observations.

Result Python Native C
Lives lost 0 0
Final shield 27 27
Shield damage / hits 12 / 3 12 / 3
Pickups collected 50 50
Score gained 33,350 33,350
Wall shooters destroyed 16 16
Final cash 450 450

Both bought Speedup at frame 5254 (600 → 100 cash) and Double Shot at frame 9323 (3,450 → 450 cash). Final cash includes the game's shop/bonus accounting; it should not be equated with the sum of observed pickup cash changes alone.

Both took identical directional-projectile hits at 7246, 8303 and 8323, each costing four shield points. Useful replay regions are 7220–7260 and 8270–8340. The first occurs during swarm_escape / maneuver:station; the later hits occur during boss_shell_clearance / maneuver:no_verified_escape and wall_escape / maneuver:repair. These remain gameplay weaknesses shared by both backends, not new native-only regressions.

Live timing

decision_ms measures the live decision branch. It excludes most world/map updates, socket/emulator stepping, rendering and recording. Gameplay timing excludes shop/intermission rows so their near-zero decisions cannot dilute it.

Gameplay decisions (5,425 frames) Python Native C
Mean 18.431 ms 15.667 ms
Median 9.043 ms 7.576 ms
p95 64.371 ms 52.120 ms
Maximum 320.412 ms 270.422 ms
Decisions over 40 ms 673 505
Decisions over 100 ms 108 53

The native run has approximately 15.0% lower mean decision time, 19.0% lower p95, and 50.9% fewer decisions over 100 ms. Across all 7,693 rows, total measured decision time falls from 100.057 to 85.068 seconds. These are observed savings for this matched Level 1 pair, not total emulator throughput or a phone performance estimate. The remaining slow tail is still substantial. This validation does not establish gameplay or performance coverage for Levels 2–5, and one sequential pair is not a repeated machine-independent benchmark.

For gameplay rows only, mean policy time changes from 3.758 to 3.574 ms and mean planning time from 14.370 to 11.767 ms. Most of the measured benefit is in the maneuver planner. Do not sum planner metrics from shop rows: those rows can carry the previous gameplay decision's counters even though the shop controller ran.

After both live runs finished, the native recording was profiled offline over all 7,693 decision frames, with its recorded kernel configuration and terrain seeds. The profile contains 759.3 million calls and 237.1 seconds of instrumented function time. These numbers locate work; they are not uninstrumented speed estimates. Cumulative times overlap and must not be added together.

Function/work with C enabled Calls Cumulative time
Maneuver PredictionService.evaluate 42,264 146.775 s
PredictionService.motion 2,234,099 34.023 s
group_slices 705,209 31.111 s
Generated dataclass __hash__ (merged profile entry) 80,998,422 17.562 s
wall_free 1,135,755 15.330 s
Native body frame_slices preparation in Python 287,854 14.585 s
corridor_center 5,162 13.614 s
is_player_mask_anchor_free 1,055,755 13.121 s
Native combat adapter, including C execution 89,975 6.492 s
Actual advance_ship_and_camera arithmetic/allocation 604,323 5.178 s

The 34-second motion total does not imply a 34-second opportunity from translating ship physics. Its dictionary lookups account for 19.630 seconds, including repeated candidate hashing. Retained certificate restoration, maneuver evaluation and motion-cache access hash action tuples. The generated hash profile entry is shared by Candidate, Rect and other dataclasses, so its full count must not be attributed exclusively to Candidate objects. The explicit action-cache call sites still establish substantial repeated candidate hashing.

The remaining group cost has two entry paths: 13.387 seconds from native frame_slices preparation, and 17.724 seconds from ordinary Python group queries. The C backend still prepares specialized shapes in Python and uses Python validation for fresh maneuvers before native combat takes over. Existing native body queries are mainly batched over already-restored motions. Thus extending scene preparation/validation coverage matters more than making rectangle math inside the existing C kernel slightly faster.

The live native counters record 330,666 combat frames in 89,975 native calls: 3.68 frames per call across the full level. The earlier corridor sample's 21 frames/call does not generalize to this boss-heavy run. Continuing combat one frame at a time while the Python validator proceeds is a real remaining boundary cost. However, the complete adapter costs only 6.492 profiled seconds, including 4.486 seconds in cached/fallback frame_bounds preparation; further combat-only micro-optimization is not the highest priority.

Recommended order (proposals only; none implemented by this validation):

  1. Simplify action-cache keys in Python before porting more work. Use compact action IDs or prefix IDs instead of repeatedly hashing tuples of Candidate objects. Preserve shared prefixes and lazy early rejection. Measure this separately; the merged hash profile is evidence of repeated work, not a promise that all 17.562 seconds belong to these keys or disappear in normal execution.
  2. Port batched wall-mask and corridor queries. Keep level wall rows resident in a small native terrain context, initialized from the existing map data and updated only when observed tiles change. Query batches of anchors/segments and corridor intervals in one call. This targets the mask loops and the current corridor_center scan over 291 horizontal anchors without one ctypes call per anchor. Reuse existing span navigation; do not write another full map builder or assume destructible walls are open before their destruction is observed. This is the recommended first new C kernel: bounded scope, substantial measured work and a useful resident map for a later direct Hatari caller.
  3. Extend native hazard scene preparation and validation. Prepare/cache independent shapes once per observation/frame, then apply query exclusions without rebuilding Python Group/Rect collections for each exclusion set. Start with the common scripted/formation and projectile cases still appearing in the Python path. Let C consume its prepared shapes directly; returning every rectangle to Python would preserve much of the packing/allocation overhead. This is a larger port than wall queries and needs a clear supported-model boundary, with the existing Python implementation retained for comparison.
  4. Later, combine maneuver continuation, player/camera transitions and native validation into a whole-maneuver call. This can eliminate repeated Python cache traversal, temporary motions and single-frame combat handoffs. Do it after the scene and terrain boundaries are ready. Porting only advance_ship_and_camera now would move about five profiled seconds while retaining much of the surrounding 34-second motion cost and many callbacks.

Keep level/area tactics and shopping in Python during these stages. Avoid investing in a new native JSON/replay decoder just to accelerate the tooling: that transport is expected to disappear when observations come directly from the C game integration. These recommendations preceded the implementation in AUTOPILOT_NATIVE_PREPARATION.MD; the live measurements in this document remain the committed ABI-6 baseline.

The offline profile uses recorded observations; replay still has the documented fire-schedule/shop handoff limitations in PROFILING.MD. Its matching live counters for major rollout/combat totals help establish representative work, but its instrumented timing must not replace the live measurements above.

Full native profile

Reproduction and artifacts

Each .validation directory also contains paired AVIs, .planner.json, the source archive, manifest, gameplay catalog, live hotspots, events, terrain seeds and resumable checkpoints. The native metadata lists ABI 6 and the eye-chain, body-query, linear-body and combat kernels. Recordings remain under ignored run_logs; this document records the durable findings.

python xenon_tools/validate_campaign.py run native-python-live-c-0906-02 --connect --port 6913 --build-directory out/build/mingw-debug --planner maneuver --navigation-backend spans --architecture all --kernel-backend c --stop-at-level 2 --checkpoint-interval 500
python xenon_tools/validate_campaign.py run native-python-live-python-0906-02 --connect --port 6913 --build-directory out/build/mingw-debug --resume xenon_tools/run_logs/native-python-live-c-0906-02.validation/start.sav --planner maneuver --navigation-backend spans --architecture all --kernel-backend python --stop-at-level 2 --checkpoint-interval 500
python xenon_tools/run_logs/analyze_native_python_live.py native-python-live-python-0906-02 native-python-live-c-0906-02
python xenon_tools/validate_campaign.py profile native-python-live-c-0906-02 --from-frame 2367 --to-frame 10059 --cprofile --top 40

Use new output names when repeating the captures; existing recordings are never overwritten. The commands assume an already-started validation emulator.