Autopilot · write-up
Batch fixed protocol records before planning
Baseline: 225155ee (deferred motions and shared maneuver origin, ABI 12).
Why this target
The preceding Level 1 profile spent 17.375 cumulative seconds in decode_frame,
including 20,940,325 calls to its scalar take helper. Each field read constructed
a format string, invoked struct.unpack_from, computed a format size and advanced
the offset. This is host preparation, not prediction or candidate scoring.
The Python driver needs those records, but does not need a Python call per field.
Existing struct.Struct unpacking performs the fixed-layout work in native code;
another C ABI or an additional copy into autopilot buffers is unnecessary here.
This change benefits live Python driving and binary replay. It does not improve
the future direct Hatari C input path or establish browser/phone performance.
Implementation
xenon_client.py defines named wire layouts and corresponding field-name tuples
for attachments, shop slots, object headers, lifecycle events and draw records.
Each record is unpacked in one call; padding is skipped in the layout. Object raw
bytes and protocol-version trailers retain their existing handling. The returned
dictionaries and shared shop object retain their existing semantics.
The scalar loop bodies remain available through BATCH_FRAME_RECORDS = False.
compare_kernels.py --variants c-field-decode,c changes only decoding; both
variants use all current native autopilot optimizations. The Python kernel also
uses the faster decoder by default. No C source, ABI or Hatari build changed.
Validation
The complete suite passes 892 tests. The new regression test compares both decoders with negative coordinates, maximum-width values, nonzero padding, variable raw-object lengths, all lifecycle kinds and a player-hit/control trailer. Both paths reject a truncated fixed draw record.
xenon_tools/run_logs/batch-decode-equivalence-0907.json records full decoded
dictionary comparisons across five Level 1–5 recordings, independently of planner
verdict checks. The timed replay and separate profiles use the Level 1 recording
and frames 2367–10059 from the preceding optimization.
All 15,432 decoded frames match: Level 1 7,694; Level 2 961; Level 3 225; Level 4 900; Level 5 5,652. This is offline validation, not a fresh closed-loop run.
Measured result
xenon_tools/run_logs/batch-decode-l1-0907/comparison.json contains two fresh
pinned runs per variant over 7,693 observations, including 5,425 gameplay frames.
All planner verdicts match in all four runs.
| Decoder | Run 1 mean ms | Run 2 mean ms | Average mean ms | Average p95 ms |
|---|---|---|---|---|
| Scalar fields | 5.239 | 5.282 | 5.261 | 12.490 |
| Batched records | 4.887 | 4.878 | 4.882 | 12.079 |
Total replay processing is 7.2% lower, approximately 0.38 ms per observation. The paired gains are 6.7% and 7.6%; each variant varies by less than 1% between runs. This measures complete replay processing without drawing, emulator execution or sockets. The planner algorithms and number of evaluated candidates are unchanged.
The separate batch-decode-profile-0907 comparison also has zero changed verdicts:
| Profile measure | Scalar fields | Batched records |
|---|---|---|
Scalar take calls |
20,940,325 | 1,266,640 |
decode_frame cumulative seconds |
17.101 | 3.069 |
| Total recorded function calls | 261,018,255 | 183,598,712 |
This removes 94.0% of scalar field reads and reduces instrumented decoder time by 82.1%. cProfile penalizes tiny Python calls, so those percentages are not the real-time speedup; use the unprofiled 7.2% result above. Remaining scalar reads handle the short frame header and semantic trailers. The dominant remaining autopilot work is policy selection and prediction/scene preparation, rather than these now-batched repeated protocol records.
Reproduction
Run timing workers sequentially; profile separately because instrumentation changes relative costs. Reverse variant order on the second timing repeat.
python xenon_tools/compare_kernels.py xenon_tools/run_logs/native-python-live-c-0906-02.validation/native-python-live-c-0906-02.x2events --output <timings> --variants c-field-decode,c --from-frame 2367 --to-frame 10059 --repeats 2 --pin-cpu
python xenon_tools/compare_kernels.py xenon_tools/run_logs/native-python-live-c-0906-02.validation/native-python-live-c-0906-02.x2events --output <profile> --variants c-field-decode,c --from-frame 2367 --to-frame 10059 --repeats 1 --pin-cpu --cprofile
python -m unittest discover -s xenon_tools -p 'test_*.py'