Autopilot · write-up
Remove unused render geometry from world-state preparation
Baseline: 2a10db47 (batched frame decoding; native ABI 12).
Selection
The preceding Level 1 profile attributed 10.885 cumulative seconds to world-state updates and 7.147 seconds to native maneuver preparation. Of the latter, 4.618 seconds were shared scene construction, including independent hazard geometry. Removing that preparation requires preserving collision and combat ordering.
World preparation contained simpler avoidable work: it collected rectangles for every draw, including draws without any surviving captured object. Those rectangles were never read. A sampled Level 1 recording contained 69,002 unmatched draws, 42,758 wall draws and 19,337 matched draws. Every matched object in that sample had only one draw, yet went through list construction and a four-pass union.
Change
WorldState.update filters destroyed identities before collecting geometry and
builds the set of surviving (object address, identity) pairs. Non-wall draws
outside that set are skipped before coordinate/size extraction. Walls still pass
through the existing visibility and tile handling, independently of object records.
One render rectangle is stored directly per matching object. Additional draws
extend it through _merge_render_bounds; no temporary rectangle lists or full
union scans are needed. Objects without a matching draw retain the existing
collision-word and fallback behavior. Object capture, classification, history,
velocity, damage, pickup and weapon models are unchanged.
world_state.COMPACT_DRAW_GEOMETRY = False retains the previous collector.
compare_kernels.py --variants c-world-draws,c changes only this collector while
keeping the current decoder and native kernels enabled. The improvement also
applies with the Python kernel. No C source, ABI, Hatari integration or WASM
profiling changed in this step.
Validation
The complete suite passes 893 tests. A new regression test compares complete world state across multiple draws, allocation identity mismatches, destroyed objects, missing render data, zero-size draws, wall visibility, scroll changes and a new run. The implementation adds a short typed and parameter-documented helper.
xenon_tools/run_logs/world-draw-equivalence-0907.json records complete world-state
comparisons across Level 1–5 recordings. The timing and profile comparisons below
also check planner verdicts. These are offline checks, not fresh closed-loop runs.
All 15,432 complete world states match. Across these recordings, geometry for 1,548,233 unmatched draws is eliminated. The 215,388 matching draws each belong to an object with one draw; the synthetic regression test covers multiple draws. All wall tiles match, including visibility and scroll handling.
Timing
xenon_tools/run_logs/world-draw-l1-0907/comparison.json contains two fresh pinned
runs per variant over 7,693 Level 1 observations (5,425 gameplay frames). The second
repeat reverses worker order. All planner verdicts match in all four runs.
| Collector | Run 1 mean ms | Run 2 mean ms | Average mean ms | Average p95 ms |
|---|---|---|---|---|
| Full draw lists and unions | 4.799 | 4.805 | 4.802 | 11.865 |
| Matching draws and direct bounds | 4.776 | 4.746 | 4.761 | 11.762 |
The end-to-end reduction is 0.84%, approximately 0.040 ms per observation. Paired reductions are 0.47% and 1.21%. This is a small improvement, not a major step toward phone performance. The simpler geometry path removes known unused work, but the allocation count reduction must not be presented as an equivalent runtime reduction. Frame decoding, object-state extraction and planning still run.
The separate world-draw-profile-0907 comparison has zero changed verdicts:
| Profile measure | Full collector | Compact collector |
|---|---|---|
| World update cumulative seconds | 10.647 | 8.728 |
| Render union calls | 210,072 | 0 |
| Total recorded function calls | 182,885,778 | 178,127,034 |
Instrumented world-update time is 18.0% lower, but instrumentation magnifies Python call overhead. Use the unprofiled 0.84% for the total runtime benefit. Remaining world-update cost is predominantly object-state extraction and tracking; further geometry caching is unlikely to provide a large gain. Maneuver scene preparation remains a separate candidate for subsequent work.
Reproduction
Use fresh sequential workers and reverse order for the second timing repeat; keep profiling separate from real-time measurements.
python xenon_tools/compare_kernels.py xenon_tools/run_logs/native-python-live-c-0906-02.validation/native-python-live-c-0906-02.x2events --output <timings> --variants c-world-draws,c --from-frame 2367 --to-frame 10059 --repeats 2 --pin-cpu
python xenon_tools/compare_kernels.py xenon_tools/run_logs/native-python-live-c-0906-02.validation/native-python-live-c-0906-02.x2events --output <profile> --variants c-world-draws,c --from-frame 2367 --to-frame 10059 --repeats 1 --pin-cpu --cprofile
python -m unittest discover -s xenon_tools -p 'test_*.py'