Autopilot · write-up
Native scene models (2026-09-06)
The maneuver kernel was committed first as 889aa110. This extension uses
ABI 10 and retains the Python kernel and earlier native-stage controls.
It builds only the standalone library callable from Python; no Hatari integration
or WASM profiling is included.
Model attribution and selected work
The previous profile grouped specialized body preparation together. A separate
instrumented replay attributed calls by hazard classification and update routine:
model audit.
The largest remaining body cost was the fallback for boss_eye_segment and
boss_eye_weak_point ($50478): 201,304 generic sweep calls, about 2.25 seconds
inside those instrumented calls. These are not the already-ported articulated
controller forecasts. A captured-link audit confirmed ordinary delayed copies
from an unlinked, fitted-velocity root when the controller forecast is absent.
Shot dispatch also repeatedly inspected incapable owners: for example 279,697
calls on directional projectiles and 165,642 on $4F4CE enemies, all empty in
this Level 1 recording. This is dispatch work to remove, not a numerical loop
worth translating into C.
Source partitioning
prepared_slices.slice_sources separates remaining bodies and shot-capable
owners once per observation. Frames no longer revisit native bodies or run the
wall-shot dispatcher for incapable objects. The owner procedures cover Level 3
oscillating/small/large wall shooters, Level 4 directional wall shooters and the
two Level 5 barrier families. Player-dependent owners remain in their ordered
rollout pass. The old traversal is retained for comparison.
Delay-chain fallback
XapDelayedMember contains an observed anchor, a root velocity and a predecessor
index. xap_prepare_delayed_paths prepares all selected chains in one call.
A root uses its fitted velocity; each follower copies its predecessor at the
preceding frame and adds the default scroll contribution. Every member retains
its own observed frame-zero anchor. No recursive Python prediction or per-frame
Rect construction is needed for these body slices.
The first Level 1 trial exposed an overly narrow input check at frame 7803:
the observed eye reposition produced fitted velocity (-26, -236). Reusing the
ordinary 64-pixel linear-body limit rejected this valid fallback input. Delayed
root arithmetic uses doubles, so it now admits observed jumps up to the coordinate
limit; ordinary scheduled shots retain their existing speed bound. The captured
case is covered by a differential test. The aborted scene-model-stages-0906
run is excluded from final timing claims.
The wrapper rejects cycles, scripted roots, specialized parent relationships and chains already covered by the articulated controller forecast. It does not replace the active controller with a linear approximation. Native body queries retain the existing three-pixel eye animation margin. Combat uses the same anchors with its existing narrow bounds.
Deterministic next enemy shots
XapScheduledShot holds the spawn anchor, fixed-direction velocity, allocation
frame and first movement age. Python selects the owner model and computes the
next allocation once. xap_prepare_shot_paths advances the complete forecast;
the body kernel derives collision sweeps directly from those packed points.
The first two owner models are:
- Level 3 small wall shooter (
$50FF6): captured phase/accumulator, both mount orientations, age zero on allocation, 5-by-8 collision bounds. - Level 4 directional wall shooter (
$512F6): captured phase/accumulator, both directions, age one on allocation, 5-by-5 collision bounds.
Velocity is the binary-exact game table increment in pixels per frame; C floors the resulting anchors as the reference's 16.16 movement does. No extra fixed-point representation or exact-conversion fallback is introduced. Only the next known allocation is predicted; RNG-selected resets and subsequent shots are not invented.
XapHazardPath.active_from_frame suppresses pre-spawn contact. Anticipated paths
have no damageable target index, survive their owner's later predicted death,
and are excluded from player-bullet blockers. These are enemy shots; the existing
native combat kernel still simulates player fire separately.
All buffers are caller-owned, with named C structs mirrored by ctypes.Structure.
The preparation functions allocate nothing and retain no pointers. Prepared paths
are observation-owned and shared by maneuver candidates. Unsupported shot models,
including reflecting fans and aimed releases, keep their Python prediction.
Validation and measurements
Tests cover delayed copy ordering and observed anchors, cyclic-link rejection, articulated-model precedence, all supported firing phases and both orientations, pre-spawn inactivity, player bullets passing through anticipated enemy shots, source partition equivalence and replay counters. The complete suite passes 875 tests; targeted tests also pass after the final eligibility/counter review.
Fresh workers use one pinned logical CPU, identical recording/settings, two repeats with reversed order, and source/DLL fingerprints. Tests, compilation and profiling do not overlap timing. Results are offline replay processing times, not new closed-loop gameplay or phone benchmarks.
c-maneuver disables all changes in this extension, c-sources enables owner
partitioning, c-chains additionally enables fallback eye paths, and c also
enables scheduled shots. Historical controls keep later stages disabled.
The final Level 1 comparison covers frames 2367 through 10059 of
native-python-live-c-0906-02.x2events: 7,693 observations, including 5,425
gameplay observations. All stages have identical verdicts in both repeats.
| Stage | Mean ms | Mean of repeated p95 ms | Incremental mean reduction |
|---|---|---|---|
| Previous maneuver kernel | 7.656 | 22.976 | baseline |
| Source partitioning | 7.428 | 22.435 | 3.0% |
| Native fallback eye chains and scheduled-owner checks | 7.161 | 21.437 | 3.6% |
The combined mean reduction is 6.5%, with p95 reduced by 6.7%. Repeat spread is at most 1.6%. Native preparation handles 3,552 delayed-chain members in 444 batched calls; Python fallback slice bytes fall from 18,574,464 to 9,026,688. Source partitioning alone preserves those bytes while removing repeated owner dispatch.
The Level 3 shot-only comparison uses
level3-stage2-smallshot-0826-213.x2events, all 225 observations:
15.893 → 14.890 ms mean, 6.3% lower, with identical verdicts in both repeats.
It prepares 675 native shot paths in 225 batched calls. Remaining Python slice
bytes decrease from 2,971,872 to 1,589,472.
The Level 4 shot-only comparison uses level4-wall-model-0827-07.x2events,
all 900 observations (752 gameplay observations): 19.501 → 18.868 ms mean,
3.2% lower, again with identical verdicts in both repeats. It prepares 1,639
shot paths in 752 batched calls. Python slice bytes decrease from 5,414,640 to
1,744,512. The mean of the repeated p95 values changes from 89.426 to 88.410 ms;
the gain is principally in average preparation work, not the slowest decisions.
Reproduction (use new output directories):
python xenon_tools/compare_kernels.py xenon_tools/run_logs/level3-stage2-smallshot-0826-213.x2events --output <new-l3-directory> --variants c-chains,c --repeats 2 --pin-cpu
python xenon_tools/compare_kernels.py xenon_tools/run_logs/level4-wall-model-0827-07.x2events --output <new-l4-directory> --variants c-chains,c --repeats 2 --pin-cpu
python xenon_tools/compare_kernels.py xenon_tools/run_logs/native-python-live-c-0906-02.validation/native-python-live-c-0906-02.x2events --output <new-l1-directory> --variants c-maneuver,c-sources,c --from-frame 2367 --to-frame 10059 --repeats 2 --pin-cpu
The Level 1 recording has no supported scheduled-shot owners. Its final stage therefore measures the fallback eye chains plus the small cost of checking for scheduled-shot owners. The Level 3/4 comparisons isolate the actual shot port. Do not add percentages from these different recordings.
Final profile
Profile, collected
separately from timing, completes all 7,693 Level 1 frames after the numeric-range
fix. Its verdicts also match the saved c-maneuver control. Total profiled calls
decrease from 368.6 million to 334.8 million.
Instrumented cumulative times are nested; do not add them or interpret their percentage reduction as an FPS gain:
| Layer | Previous maneuver kernel seconds | Current seconds |
|---|---|---|
| Native maneuver preparation/call wrapper | 19.202 | 8.022 |
| Observation-scene preparation inside that wrapper | 15.495 | 4.392 |
| Native fallback slice packing | 14.904 | 3.794 |
| Specialized slice generator, including resumes | 12.781 | 2.471 |
| Ordered Python evaluation | 37.295 | 36.150 |
The remaining major autopilot cost is now the ordered Python evaluation loop (8.517 seconds self time), followed by policy/model preparation. The newly ported scene construction no longer dominates the maneuver wrapper. Moving the ordered loop further requires treating dependent hazards, pickups and combat handoffs together; translating isolated scalar helpers is unlikely to remove much of that cost. That further port is not included here.