Xenon 2

Autopilot · write-up

Shared native prediction and scene preparation

xenondoc/AUTOPILOT_SHARED_PREDICTION_SCENE.MD · 7 KB · updated 2026-09-17

Baseline: c488910f; native ABI 14.

Script and formation preparation live in src/autopilot/xenon_autopilot_prediction.c. Public declarations remain in xenon_autopilot.h; the existing internal header shares one sine-table definition and the integer-coordinate helper with eye prediction. The file split changes no public API or prediction behavior.

Ownership and removed work

Each native decision now owns a shared point pool. The first script request batches captured, unlinked script roots into xap_prepare_script_paths. C executes the sine recurrence, delay/jump/metadata/stop commands and the existing unknown-command tangent estimate directly into absolute world anchors. Screen-relative roots include the normal camera shift. This is the game's fixed-point motion representation, not a new precision requirement on planner coordinates.

Ordinary formation followers borrow the root path. xap_prepare_formation_paths prepares a predecessor-first chain in one call, copying the preceding frame and adding the camera shift. Policy contacts and body/collision/combat preparation use indices into the same allocation instead of packing those paths again. Other prepared paths can enter the pool once when a native consumer needs them.

The pool is observation-local. Append-only growth preserves offsets and copies old points; existing Python views retain their original allocation. C retains no pointers after a call and needs no Python callback or emulator symbols. The future native driver can allocate the same structures and call these APIs directly.

Python can read an individual anchor or borrow a contiguous view. Only reference consumers explicitly asking for displacement lists materialize those lists. The old script tapes and packing paths remain available. Specialized linked models, pending spawns, and requests beyond the native horizon retain reference handling. This does not move all tactical policy, world decoding, navigation, or every scene model to C. Destruction remains part of ordered candidate combat evaluation.

Comparison controls and diagnostics

shared_prediction_scene.ENABLED = False restores previous prediction and packing. With it enabled, SHARE_BUFFERS = False isolates native prediction execution while retaining consumer packing. compare_kernels.py exposes these as c-shared-old, c-shared-predictions, and c respectively. The Python kernel remains selectable.

Replay/live diagnostics show shared_script_calls, shared_script_paths, shared_formation_paths, shared_native_points, and shared_path_borrows. native_body_scene_bytes counts newly packed body data; borrowed pool bytes are excluded to avoid reporting them as another upload. Pool points include spare paths prepared by the script batch, but exclude unused allocation capacity.

Build with xenon_tools/hatari_dev.ps1 -Action build-kernel. No Hatari driver integration or WASM timing is part of this change.

Validation and measurements

The final full suite passes 900 tests, including 24 replay UI tests. New coverage compares 66 script states across sine segments, stop/unknown/truncated commands, jumps, metadata, delays, signed wrapping and camera shifts. Tests also verify follower delays, point-pool growth, actual policy/body allocation sharing, and native-only overlay rendering without creating a Python displacement cache.

Two fresh pinned Level 1 runs per variant, reversing order in the second repeat:

Variant Run 1 mean ms Run 2 mean ms Average mean ms
Previous prediction and packing 4.579 4.618 4.599
Native prediction, consumer packing retained 4.371 4.452 4.411
Native prediction and shared buffers 4.306 4.290 4.298

The complete sequential replay pipeline is 6.5% faster, saving about 0.301 ms per observation. Native prediction accounts for 0.188 ms and sharing for another 0.113 ms in this comparison. Mean policy time on gameplay frames falls from 1.739 to 1.475 ms (15.2%). These are desktop replay measurements, not phone/WASM performance claims or new autonomous live runs.

All 7,693 Level 1 observations (5,425 gameplay frames) retain identical planner verdicts. Single comparisons on the available Level 2–5 recordings also retain identical verdicts: action, tactic, movement authority, clearance, contact, predicted kills and pickups. Small timing differences on these shorter/different workloads should not be treated as established speedups:

Recording Previous mean ms Shared mean ms
Level 2 6.913 6.764
Level 3 9.657 9.458
Level 4 7.227 7.181
Level 5 2.749 2.761

The Level 1 shared scene averages 4.04 script roots and 9.60 delayed followers per gameplay frame, using 0.79 script calls and 0.93 formation calls. Consumers borrow paths about 8.05 times per frame. Newly packed body-scene data falls from 8,673 to 604 bytes per gameplay frame (93% less); this excludes borrowed shared-pool bytes. The shared pool itself averages 1,428 used points, about 22.8 KB. Old allocations may remain alive through views when the pool grows, so this is not peak memory.

Separate profile and remaining work

The instrumented Level 1 run drops from 170.4 to 162.0 million calls. Cumulative profile times below overlap and must not be added or substituted for unprofiled measurements:

Stage Previous cumulative s Shared cumulative s
FormationTrajectories.get 4.420 1.013
Native policy scene construction 1.620 0.455
Script displacement wrapper 2.911 0.032
Body scene construction 1.477 1.115
Remaining body frame slices 3.952 4.820
Ordered candidate evaluation 14.263 14.302

Native script batch preparation itself takes 0.333 cumulative seconds; native formation preparation takes 0.276 seconds, including Python descriptor work. The first implementation used a ctypes point object for each Python read, offsetting much of the prediction win. Borrowed double memoryviews and direct iteration remove that extra wrapper allocation; the older exploratory timing is retained for evidence.

Remaining frame slices still enter Python hazard dispatch and now sometimes read shared points through that dispatch. Their increased profile cost limits the total gain. A useful next step is to route more eligible constant-hull script roots into the native body scene and bypass native eligibility calls for non-script objects, while retaining specialized collision envelopes and enemy-shot rules. Candidate combat/target preparation and high-level tactical policy also remain substantial.

Artifacts

  • xenon_tools/run_logs/shared-scene-final-timing-0907/comparison.json: final unprofiled Level 1 stage isolation, two repeats.
  • xenon_tools/run_logs/shared-scene-timing-0907/comparison.json: exploratory version before removing ctypes point-wrapper reads.
  • xenon_tools/run_logs/shared-scene-profile-0907/: separate old/new pstats and matching verdict comparison.
  • xenon_tools/run_logs/shared-scene-level*-0907/: Level 2–5 comparisons.
  • xenon_tools/run_logs/shared-scene-final-tests-0907.txt: full test log.

The timing fingerprint precedes an out-of-range formation-horizon fallback guard and overlay/test-only changes. The subsequent profile and cross-level runs include the guard; the measured Level 1 horizons never reach it.