Xenon 2

Autopilot · write-up

Whole-candidate evaluation (2026-09-06)

xenondoc/AUTOPILOT_NATIVE_CANDIDATE_PLAN.MD · 12 KB · updated 2026-09-17

Baseline: 68310253, native scene preparation. Keep the Python evaluator intact and selectable. Build the standalone ctypes library only; no Hatari integration or WASM profiling in this step.

Boundary and implementation

  1. Add a plain C candidate API over the already prepared maneuver frames and observation-owned body scene. C walks frames in gameplay order: walls, pending entries/refuge, dependent bodies, shell releases/bursts, independent bodies, aimed shots, pickups. Return the first contact, clearance, collected flags and score aggregates. Keep readable named structs and caller-owned buffers.
  2. Resolve destruction inside that loop using the existing combat implementation. Start bullet simulation only for a damageable conflict, a possible shell release, or an explicitly scored attack. Once started, preserve frame ordering and collision through death-frame + 1. Reuse an internal combat step without crossing ctypes or repeating public-call validation each frame.
  3. Cache shell anchors, pickup geometry and generic combat bounds per observation. Remaining progression-barrier motion models produce packed per-candidate geometry in Python. C owns their collision ordering, clearance and destruction checks; no callbacks during evaluation. Periodic and composite eight-direction shots are generated directly in C using the player position on the allocation frame. Their numerical model was added after profiling the first implementation (below). Level-2 enemies and Level-5 homing missiles also advance in the reached-frame loop, including animation collision headers, retargeting and known allocations.
  4. Materialize only reached Motion objects for compatibility with retained plans, repairs and replay. Aggregate pacing, top exposure and lane penalties in C so ranking does not repeatedly traverse those objects. Preserve transition budgets and the pending-input convention. Occupied-start wall escape remains an explicit Python geometry preparation step until its special map semantics are ported. A provisional C pass bounds that work and remaining barrier samples to reached frames. Confirmed wall/barrier input triggers at most one recheck from fresh mutable state; speculative kills and pickup flags never leak into that result.

Validation and measurement

  • Differential tests against Python cover contacts, death grace, lazy combat, pickups, dependent old/new hull order, shell bursts, pending input and budgets.
  • Build standalone C and run the relevant tests, then the complete Python suite.
  • Compare c-scene (baseline ordered Python loop) with c in fresh sequential workers, one pinned CPU, two repeats in reversed order. Keep sources and DLL unchanged during timing. Use the long Level 1 recording and focused recordings for other supported encounter models; inspect any decision differences.
  • Collect cProfile separately from uninstrumented timings. Report end-to-end processing time, evaluator coverage, remaining preparation cost and limitations; do not use instrumented cumulative reductions as runtime speedup claims.

Results

The final Level 1 comparison uses frames 2367 through 10059 of native-python-live-c-0906-02.x2events: 7,693 observations, including 5,425 gameplay observations. Both repetitions have identical actions and diagnostics.

Stage Mean processing ms Mean of repeated p95 ms
Committed scene kernel (c-scene) 7.068 20.876
Whole candidate (c) 6.078 14.300

Mean time decreases 14.0% and p95 31.5%. Repeat spread is at most 1.9%. These are full replay-processing measurements on the host, not phone timings or closed-loop gameplay results. Comparison.

The native evaluator handles all 42,264 rollouts and 1,135,755 reached frames. Fifteen provisional rechecks add only 214 frame checks and 192 combat frames. Final evaluations account for 330,666 combat frames, exactly matching the baseline's demand, and preserve all 481,501 charged transitions. The observation scenes prepare 1,289 periodic/composite shot descriptors. No Python dependent samples are needed in this Level 1 recording after the periodic-shot port.

The separate final profile also matches all 7,693 Level 1 verdicts.

Final profile

Profile was collected after the timing workers finished. Total profiled calls decrease from 334,752,873 to 290,393,740. Cumulative times below are nested and instrumented; they must not be added together or used as runtime/FPS speedup estimates.

Layer Previous scene kernel seconds Whole candidate seconds
PredictionService.evaluate, including preparation 36.150 18.768
Native maneuver forecast wrapper, including scene preparation 7.924 7.919
Reached Motion materialization/cache access 3.556 3.614
Planner ranking 0.711 0.554
Policy adapter choose 31.588 31.619

The candidate evaluator's cumulative cost falls by about 48%. Its final C-call wrapper accounts for 0.849 seconds cumulative across 42,279 calls; much of the remaining evaluator time is Python scene/combat preparation and compatibility Motion objects. The discarded first version's dependent-sample constructor cost 14.555 seconds; the final Level 1 constructor costs 0.210 seconds with no samples.

The largest remaining autopilot layer is policy/model preparation. Outside that layer, the Python host still spends 17.038 seconds cumulative decoding captured frames and 11.043 seconds updating WorldState. These results explain the modest end-to-end gain despite halving evaluation cost. Future work should separate policy preparation and host decoding costs before selecting another numerical port. The current implementation retains full-horizon native maneuver forecasting and Python result materialization for compatibility; it does not integrate the autopilot into Hatari or claim phone performance.

Reproduction (choose new output directories):

python xenon_tools/compare_kernels.py xenon_tools/run_logs/native-python-live-c-0906-02.validation/native-python-live-c-0906-02.x2events --output <new-timing-directory> --variants c-scene,c --from-frame 2367 --to-frame 10059 --repeats 2 --pin-cpu
python xenon_tools/compare_kernels.py xenon_tools/run_logs/native-python-live-c-0906-02.validation/native-python-live-c-0906-02.x2events --output <new-profile-directory> --variants c --from-frame 2367 --to-frame 10059 --repeats 1 --pin-cpu --cprofile
python -m unittest discover -s xenon_tools -p 'test_*.py'

Cross-level replay checks

Each of these single-pass comparisons has zero changed verdicts, including input, authority, clearance, contact frame/type/identity, kills and pickups. Their timing is a smoke measurement; small differences are not robust speedup claims.

Recording Observations / gameplay Previous mean ms Current mean ms
Level 2 homing dual collision 961 / 910 169.346 167.930
Level 3 stage-2 small shots 225 / 225 14.802 14.253
Level 4 wall model 900 / 752 18.870 17.470
Level 5 tank homing 5,652 / 391 4.971 4.626

Means include the entire recording, including non-gameplay observations. This is particularly significant for Level 5; the table is not its active-combat frame budget. Level 2 is effectively flat end-to-end: at the regression frame 32690, final planning is 2.9 ms while policy preparation is 241 ms. Escape queries at that frame are now 17 instead of 1,360. Native homing preparation covers 1,502 observed/future models in Level 2 and 2,055 in Level 5; neither recording needs Python dependent samples.

Artifacts:

The older level3-composite-single-trajectory-0826-133.x2events fixture fails in the baseline policy at autoplay.py:7012 (min() on an empty sequence), before the candidate port is reached. It is excluded; composite numerical behavior is covered by differential tests and Level 3 replay uses the previously valid small-shot fixture. The first Level 4 comparison matched actions and safety but omitted anticipated-shot owner IDs in diagnostics. Restoring those IDs produces the fully matching comparison above.

The final implementation passes the complete 884-test suite. Tests cover pending input, transition exhaustion, first-contact ordering, lazy damage and death grace, pickup flags, native periodic/composite shots, shell bursts, both homing families and bounded escape/barrier preparation. The standalone library build and whitespace checks pass. Python remains selectable; live/replay summaries show candidate calls, reached/work frames, combat frames, rechecks and model counts.

Profiling-driven revision

The first full-candidate implementation retained Python numerical preparation for periodic and composite shots. It passed 879 tests but improved the long Level 1 replay mean by only 1.6% in one trial. Its separate profile attributed 14.555 seconds cumulative to dependent-sample construction, including 632,836 calls to the periodic-shot predictor. Packing the entire candidate horizon was generating geometry beyond its first rejection; moving only the checks into C was insufficient.

The revised scene stores each next periodic/composite spawn once, with its exact carry frame, quantized spawn point, eight velocities and aimed/random flag. C quantizes the candidate's allocation-frame aim using the integer $39A4 compass rule, then evaluates the $4180 swept header only on reached frames. Random shots reserve all eight directions. Spawned shots persist after their source dies. Progression-barrier numerical models remain visibly in Python.

c-scene retains the committed ordered evaluator. c-candidate enables the new C evaluator with the initial Python periodic-shot preparation; c additionally enables native periodic/composite shots. All retain the earlier scene/kernel ports.

The first Level 2 trial exposed eager preparation around homing encounters: over the first 215 comparable observations, mean processing was about 604 ms versus 262 ms for the baseline. That trial was stopped and is excluded from final performance claims. The next revision adds caller-owned homing state and a shared Level-2 animation/header table. C advances and checks both homing families only on reached frames. Previous/new bounds are stored in the candidate's combat history so lazy bullet catch-up sees the same blockers, without forecasting the remainder of a rejected candidate. A future Level-5 missile is checked on its allocation frame and enters bullet bounds on its next update, matching the reference.

Homing alone did not remove the Level 2 regression. Detailed counters at frame 32690 showed 1,360 Python escape-map queries for only 17 reached candidate frames: the adapter evaluated unknown/occupied-start terrain for the entire forecast before C could reject its first frame. The final implementation marks these checks provisional, lets C bound the candidate, then prepares only reached escape geometry. An earlier wall failure causes one clean recheck. The same two-pass boundary limits remaining Python progression-barrier samples. A focused 21-frame replay after this correction matched decisions and removed the slowdown (457.6 to 449.1 ms mean); those absolute costs are mostly outside candidate evaluation.

The differential tests exercise all eight directions, animation countdown/cursor states, retarget boundaries, lifetime expiry and near/far contacts for both families. NATIVE_HOMING retains the initial Python sample-preparation control for attribution; normal c evaluation enables the native model.

The initial trial also exposed omitted partial clearance on a rejected body frame. The C result now includes distances accumulated before that contact. Body buffer order can still select a different contact identity among overlapping hazards; the first contact frame and safe/unsafe outcome are the validation priorities.