Xenon 2

Autopilot · write-up

Batch fixed protocol records before planning

xenondoc/AUTOPILOT_BATCH_FRAME_DECODING.MD · 5 KB · updated 2026-09-17

Baseline: 225155ee (deferred motions and shared maneuver origin, ABI 12).

Why this target

The preceding Level 1 profile spent 17.375 cumulative seconds in decode_frame, including 20,940,325 calls to its scalar take helper. Each field read constructed a format string, invoked struct.unpack_from, computed a format size and advanced the offset. This is host preparation, not prediction or candidate scoring.

The Python driver needs those records, but does not need a Python call per field. Existing struct.Struct unpacking performs the fixed-layout work in native code; another C ABI or an additional copy into autopilot buffers is unnecessary here. This change benefits live Python driving and binary replay. It does not improve the future direct Hatari C input path or establish browser/phone performance.

Implementation

xenon_client.py defines named wire layouts and corresponding field-name tuples for attachments, shop slots, object headers, lifecycle events and draw records. Each record is unpacked in one call; padding is skipped in the layout. Object raw bytes and protocol-version trailers retain their existing handling. The returned dictionaries and shared shop object retain their existing semantics.

The scalar loop bodies remain available through BATCH_FRAME_RECORDS = False. compare_kernels.py --variants c-field-decode,c changes only decoding; both variants use all current native autopilot optimizations. The Python kernel also uses the faster decoder by default. No C source, ABI or Hatari build changed.

Validation

The complete suite passes 892 tests. The new regression test compares both decoders with negative coordinates, maximum-width values, nonzero padding, variable raw-object lengths, all lifecycle kinds and a player-hit/control trailer. Both paths reject a truncated fixed draw record.

xenon_tools/run_logs/batch-decode-equivalence-0907.json records full decoded dictionary comparisons across five Level 1–5 recordings, independently of planner verdict checks. The timed replay and separate profiles use the Level 1 recording and frames 2367–10059 from the preceding optimization.

All 15,432 decoded frames match: Level 1 7,694; Level 2 961; Level 3 225; Level 4 900; Level 5 5,652. This is offline validation, not a fresh closed-loop run.

Measured result

xenon_tools/run_logs/batch-decode-l1-0907/comparison.json contains two fresh pinned runs per variant over 7,693 observations, including 5,425 gameplay frames. All planner verdicts match in all four runs.

Decoder Run 1 mean ms Run 2 mean ms Average mean ms Average p95 ms
Scalar fields 5.239 5.282 5.261 12.490
Batched records 4.887 4.878 4.882 12.079

Total replay processing is 7.2% lower, approximately 0.38 ms per observation. The paired gains are 6.7% and 7.6%; each variant varies by less than 1% between runs. This measures complete replay processing without drawing, emulator execution or sockets. The planner algorithms and number of evaluated candidates are unchanged.

The separate batch-decode-profile-0907 comparison also has zero changed verdicts:

Profile measure Scalar fields Batched records
Scalar take calls 20,940,325 1,266,640
decode_frame cumulative seconds 17.101 3.069
Total recorded function calls 261,018,255 183,598,712

This removes 94.0% of scalar field reads and reduces instrumented decoder time by 82.1%. cProfile penalizes tiny Python calls, so those percentages are not the real-time speedup; use the unprofiled 7.2% result above. Remaining scalar reads handle the short frame header and semantic trailers. The dominant remaining autopilot work is policy selection and prediction/scene preparation, rather than these now-batched repeated protocol records.

Reproduction

Run timing workers sequentially; profile separately because instrumentation changes relative costs. Reverse variant order on the second timing repeat.

python xenon_tools/compare_kernels.py xenon_tools/run_logs/native-python-live-c-0906-02.validation/native-python-live-c-0906-02.x2events --output <timings> --variants c-field-decode,c --from-frame 2367 --to-frame 10059 --repeats 2 --pin-cpu
python xenon_tools/compare_kernels.py xenon_tools/run_logs/native-python-live-c-0906-02.validation/native-python-live-c-0906-02.x2events --output <profile> --variants c-field-decode,c --from-frame 2367 --to-frame 10059 --repeats 1 --pin-cpu --cprofile
python -m unittest discover -s xenon_tools -p 'test_*.py'