Autopilot · write-up
Combined native candidate trials
Architecture and staged implementation
ABI 17 retains the existing forecast and prepared-candidate APIs, adding
xap_forecast_actions, xap_run_candidate, and xap_record_prefix.
candidate_work.CandidateWorkspaceowns reusable homing/sample buffers, collected-pickup flags, shell-release flags and combat scratch for one observation. Initial targets and observed player bullets are captured once on first combat use, then restored for each sequential trial. Candidate results capture values before the workspace is reused. Unsupported dependent sample preparation keeps its existing path.- Native forecasts return contiguous joystick masks. Python creates keys from those bytes and decodes only selected actions requested by result consumers. Full decoded forecasts remain available lazily for reference callers.
- The combined loop chooses an input, checks the shared-prefix budget, advances the ship, and evaluates that frame before generating another. Body checks run once, after earlier hazard checks and destruction updates. No future frames are generated after contact or budget exhaustion.
XapPrefixBudget is an observation-owned tree of reached input prefixes. Node
zero is the root; a zero child means an unseen transition. Shared prefixes count
once. The C call receives the remaining Python transition allowance, consumes
only new reached nodes, and returns the next input when the budget stops before
its transition. Reached frames are registered with Python's existing lazy motion
reservations, preserving later ranking, trace output and cache reuse. Reference
fallbacks register their reached prefix with one C call. Earlier Python motions
seed the tree on first use; only maximal cached paths are registered, since each
registration includes its parents. Restored certificate nodes have separate
capacity from this observation's remaining allowance. Later Python motions
synchronize their added prefix.
The combined path requires native aimed/homing support, compact keys, lazy motion results, and an available resident wall mask. Progression barriers and unsupported motion preparation retain the prepared-frame path. An occupied initial wall position returns a retry status before admitting any transition; Python runs its existing escape check after restoring mutable state. It cannot double-charge the budget. Geometry belongs to the observation and mutable scratch to the caller; no C callback, allocation or retained pointer is introduced.
For ordinary body contacts, combat preparation is also deferred until a supported
target actually blocks the provisional trial. That case resets the workspace and
retries in C with combat enabled, preserving the allowance remaining after the
first pass. Its reached prefixes are already known and are not charged again.
Shell, damageable homing-missile and explicit attack trials prepare combat up
front because their earlier checks can require destruction prediction. The
c-trial-eager variant disables only this additional optimization.
The large frame buffer for each result remains independently owned because lazy motion reservations can still refer to earlier candidates. Reusing that buffer would corrupt retained results. Only mutable scratch with no surviving result references is reused.
Independent comparison controls
compare_kernels.py runs fresh sequential workers:
c-trial-old: previous buffers, eager action conversion, separate C calls.c-trial-buffers: reusable workspace only.c-trial-actions: workspace plus packed action handling.c: all stages, including combined trials and shared-prefix accounting.
The Python kernel remains selectable. WASM exports are maintained, but this work does not integrate the Hatari driver or measure WASM performance.
Validation and measurement
Tests cover shared branching prefixes, reuse at exhausted budgets, pending input, contact, occupied-start retries, atomic rejection at prefix capacity, and combat scratch restoration. The full existing suite also covers supported hazard models, combat destruction grace, pickups, barriers and wall check ordering.
All 917 tests pass. The final test log is
xenon_tools/run_logs/combined-candidate-final-tests-0908.txt.
The initial replay smoke test exposed insufficient prefix-tree capacity for restored certificate paths, which are known without consuming new transition allowance. That issue is fixed and covered by a regression test; the aborted smoke run is not used for the reported speedup.
Stage isolation
Sequential workers pinned to logical CPU 4, two repeats in opposite orders. All four variants retain identical verdicts across 7,693 observations per repeat.
| Preparation | Run 1 ms | Run 2 ms | Mean ms | Mean p95 ms |
|---|---|---|---|---|
| Previous candidate pipeline | 3.761 | 3.781 | 3.771 | 9.552 |
| Reusable buffers | 3.589 | 3.679 | 3.634 | 8.777 |
| Buffers and packed actions | 3.536 | 3.592 | 3.564 | 8.546 |
| Combined C loop, eager combat | 3.513 | 3.542 | 3.527 | 7.975 |
Incremental mean improvements are approximately 3.6% for reusable buffers, 1.9% for packed actions, and 1.0% for the combined loop: 6.5% overall before on-demand combat preparation. The small incremental mean differences should be read alongside the run variation and p95 improvements, not as guaranteed fixed savings on every level.
A separate two-repeat comparison of eager versus deferred combat averages 3.522 versus 3.457 ms (about 1.8% faster), with identical verdicts. The final old/new measurement below includes this additional optimization directly.
The combined-trial diagnostics count completed calls, occupied-start fallbacks,
and combat retries. native_maneuver_frames includes every generated frame,
including retry work. native_candidate_avoided_frames is the requested horizon
minus final reached frames; it is not an incremental saving versus the old
forecast, which already stopped at confirmed walls.
Final end-to-end comparison
The final implementation (including on-demand combat and actual retry frame counters) was compared directly with the old pipeline in two fresh reversed-order runs over the same Level 1 observations. All verdicts match.
| Pipeline | Run 1 mean ms | Run 2 mean ms | Average mean ms | Average p95 ms |
|---|---|---|---|---|
| Old | 3.798 | 3.729 | 3.763 | 9.525 |
| Final | 3.513 | 3.511 | 3.512 | 8.041 |
Mean replay time falls 6.7%, and p95 falls 15.6%. Stage estimates above come from separate comparisons; do not add them to claim a larger final speedup. Small incremental gains are less certain than the combined result. These are fixed-observation desktop replay timings, without emulator execution, live level completion, drawing, sockets or WASM.
Separate final profile
The old path records 137,216,289 Python calls versus 127,020,269 for the final implementation: 10,196,020 fewer calls (7.4%). This instrumented run is separate from the unprofiled timing comparison.
| Profile measure | Old | Final |
|---|---|---|
| Python prefix append calls | 2,126,866 | 263,065 |
| Python combat initialization calls | 27,060 | 1,799 |
| Native combat wrapper initialization calls | 26,444 | 1,718 |
| Candidate evaluation cumulative seconds | 7.379 | 5.232 |
| Maneuver preparation cumulative seconds | 4.991 | 3.531 |
Cumulative profiler times overlap and must not be added to estimate speedup. Prefix-tree seeding/synchronization makes 4,902 registration calls over the recording, taking 0.063 cumulative seconds. Mutable workspace resets preserve initial health, player bullets and homing state between candidate trials.
The unprofiled counters show 343.56 versus 240.20 generated maneuver frames per gameplay observation, including combat retries: 30.1% fewer. Combined forecast/evaluation calls fall from 15.56 to 9.30 per gameplay observation. Combat wrapper construction falls from 4.87 to 0.32 per gameplay observation; ordinary contacts need 1.51 combat retries on average. Occupied-start fallback occurs in 49 trials across the complete recording.
Cross-level replay validation
The available Level 2–5 recordings retain identical old/new verdicts. These single-run timings are validation samples, not repeated speedup estimates or new live level-completion results.
| Recording | Observations | Old mean ms | Final mean ms | Changed verdicts |
|---|---|---|---|---|
| level2-homing-dual-collision-0824-60 | 961 | 6.490 | 5.926 | 0 |
| level3-stage2-smallshot-0826-213 | 225 | 8.502 | 8.404 | 0 |
| level4-wall-model-0827-07 | 900 | 7.061 | 6.886 | 0 |
| level5-tank-exact-homing-dev-0901-157 | 5652 | 3.122 | 2.314 | 0 |
Measurement artifacts
Under xenon_tools/run_logs/:
combined-candidate-timing-0908/comparison.json: independent three-stage isolation.combined-candidate-lazy-combat-0908/comparison.json: eager/deferred combat isolation.combined-candidate-final-timing-0908/comparison.json: final repeated old/new timing.combined-candidate-final-profile-0908/: separate instrumented old/new comparison.combined-candidate-level*-0908/comparison.json: cross-level verdict comparisons.combined-candidate-final-tests-0908.txt: 917 passing tests.