Autopilot · write-up
Shared maneuver setup and deferred Python results
Baseline: 47b2ce34 (ABI 12, native configuration grids).
Final implementation
- Deferred motion objects: retain checked motion frames in their native forecast buffers. Reserve and charge every newly reached input prefix immediately, preserving work budgets and prefix reuse. Construct Python objects only for requested endpoints, repair anchors and the final selected path. Reading an endpoint uses its preceding native hull directly, without recursively allocating the whole prefix. The chosen path becomes an ordinary tuple before entering diagnostics and replay certificates; native references do not enter those snapshots.
- Shared maneuver origin: the existing observation-owned scene stores the observed ship/camera state, three steering collision headers and initial hull. All candidates borrow this immutable setup. Candidate-specific goals, inputs, forecast outputs, collision evaluation and combat state remain independent.
autopilot_v2/preparation_reuse.py exposes LAZY_MOTIONS and SHARED_ORIGIN
for fresh-process comparisons. The original eager materialization and per-candidate
setup remain selectable. This stage changes no C ABI, integrates nothing with Hatari,
and does not profile WASM. The Python kernel stays selectable.
Replay/live metrics expose native_motion_objects and native_maneuver_origins.
native_maneuver_used_frames continues to count newly charged native prefixes;
it is no longer the number of allocated Python Motion objects.
Experiments removed after measurement
Raw hazard-step caching already exists in the reference controller. We tested additional reuse of admitted linear parameters, stationary combat bounds, observed bullet/health snapshots and firing-ray geometry. These reduced repeated calls but showed no measurable end-to-end improvement, so their implementation was removed. Stateful weapon selection, target priority and mission refresh were never cached.
reuse-stages-l1-0907/comparison.json records two fresh pinned runs per experimental
stage over 7,693 Level 1 observations, reversing order on repeat two. All verdicts
matched. The experimental controls in that artifact predate the simplified final
implementation and are identified by its source fingerprints.
- Baseline mean: approximately 5.582 ms.
- Additional preparation caches: approximately 5.576 ms (effectively unchanged).
- Adding deferred motion results: approximately 5.219 ms.
- Adding targeting geometry caches: approximately 5.222 ms (effectively unchanged).
Deferred results reduced Python Motion construction from 476,648 to 113,193 values, with identical 481,501 transition charges. The discarded snapshot experiment reduced combat reference construction from 27,116 to 1,856, illustrating why fewer calls alone are not sufficient evidence of a useful speedup.
Validation and reproduction
Tests compare deferred paths with eager motion values under budget exhaustion, check repeated-prefix reuse, and verify that shared maneuver origins produce the same independent candidate forecasts. Existing native/reference tests cover pending input, collision ordering, combat, homing, navigation and replay restoration. The final complete suite passes 891 tests.
Final Level 1 timing and work counts
reuse-final-stages-l1-0907/comparison.json compares 7,693 observations (5,425
gameplay), frames 2367–10059. Two fresh pinned workers per variant, reversed order
on repeat two, produce zero changed verdicts in every stage.
| Stage | Mean processing ms |
|---|---|
| Eager motions and per-candidate origin | 5.869 |
| Deferred motions only | 5.531 |
| Deferred motions and shared origin | 5.469 |
The combined mean is 6.8% lower. Paired-run reductions range from approximately 4% to 10%; baseline timings vary by about 9%, so this is an approximate runtime gain, not a precise hardware-independent guarantee. The small additional origin timing difference is not established separately from drift. Origin sharing is a simple removal of invariant setup from the per-candidate loop; it avoids adding another geometry cache or changing prediction ownership.
Work counts are identical in both repeats:
- Python motion construction: 476,648 → 113,193 (76.3% fewer).
- Observed origin/header preparation: 42,264 → 5,425 (87.2% fewer).
- Transition charges: 481,501, unchanged.
- Evaluated rollouts: 42,264, unchanged.
Timing includes host decoding and non-gameplay observations. Work counters show the eliminated preparation independently of timing drift.
Cross-level replay checks
Each recording below ran once per variant in fresh sequential pinned workers. Every compared verdict matched, including selected input, tactic, movement authority, clearance, contact details and predicted kills/pickups.
| Level | Observations / gameplay | Old mean ms | Final mean ms | Old / final p95 ms |
|---|---|---|---|---|
| 2 | 961 / 910 | 7.828 | 7.469 | 15.433 / 14.342 |
| 3 | 225 / 225 | 10.989 | 10.883 | 17.338 / 17.380 |
| 4 | 900 / 752 | 8.445 | 7.794 | 18.650 / 17.609 |
| 5 | 5,652 / 391 | 3.304 | 3.328 | 29.956 / 30.480 |
Levels 3 and 5 show no established timing benefit. These single-pass measurements are regression checks, not precise estimates of speedup. Level 5 is mostly inactive observations, so its overall mean does not represent active combat cost.
Artifacts are under xenon_tools/run_logs/, with comparison.json in each directory:
reuse-final-level2-homing-dual-collision-0824-60-0907reuse-final-level3-stage2-smallshot-0826-213-0907reuse-final-level4-wall-model-0827-07-0907reuse-final-level5-tank-exact-homing-dev-0901-157-0907
Final profile and remaining work
reuse-final-profile-0907/c-1.pstats covers the same 7,693 Level 1 observations.
Its verdicts also match the eager baseline. Profiling records 259.9 million calls;
instrumented elapsed times are diagnostic and must not replace the unprofiled
timing results above.
Deferred reservation costs 0.546 cumulative seconds; requested motion materialization
costs 0.826 seconds including cache hits. Native candidate evaluation falls from
18.043 to 14.922 cumulative seconds relative to policy-final-profile-0907, while
native maneuver preparation falls from 7.849 to 7.173 seconds. Some materialization
now occurs outside candidate evaluation, so that function's reduction alone is not
an end-to-end gain.
The remaining large areas are policy selection (24.928 cumulative seconds), frame decoding (17.375), world-state updates (11.045), maneuver preparation (7.173), and combat bound preparation (4.479). These are nested profile categories and must not be added together. Further work should examine the policy's repeated hazard queries and preparation of native combat geometry, with separate measurements. Host decoding and world updates also deserve attention for the Python driver, though direct C integration will have a different input path. The discarded caches show why simply adding more reuse tables is not sufficient.
Benchmark controls
Current compare_kernels.py controls:
c-work-old: eager Python motions and per-candidate origin preparation.c-work-lazy: deferred motions only.c: deferred motions and shared origin preparation.
Run timings sequentially with fixed CPU affinity and unchanged source fingerprints; use a separate cProfile run. Comparisons validate offline replay decisions and diagnostics, not closed-loop level completion or phone performance.
python xenon_tools/compare_kernels.py xenon_tools/run_logs/native-python-live-c-0906-02.validation/native-python-live-c-0906-02.x2events --output <stage-timings> --variants c-work-old,c-work-lazy,c --from-frame 2367 --to-frame 10059 --repeats 2 --pin-cpu
python xenon_tools/compare_kernels.py xenon_tools/run_logs/native-python-live-c-0906-02.validation/native-python-live-c-0906-02.x2events --output <profile> --variants c --from-frame 2367 --to-frame 10059 --repeats 1 --pin-cpu --cprofile
python -m unittest discover -s xenon_tools -p 'test_*.py'