‹ XENON 2AutopilotSpace-time distance field
Xenon 2 · autoplay.py prototype

Two ways to ask "is this move safe"

Real hazard geometry from two consecutive decision frames — pulled straight from controller._scoring_hazard_sweeps, not synthesized. Both panels answer the exact same question for the same 9 candidates; they differ only in which cells get touched to answer it. Toggle frames to see how little actually changes in one canonical step — the premise the rollover design depends on.

01

Eager flood vs. lazy search

The left panel floods outward from every occupied cell regardless of whether a candidate will ever ask about it — cells are shaded red→amber by distance only. The right panel does the opposite: it starts from each of the 9 candidates' predicted positions (dots) and searches outward only as far as it needs to, click a dot to see its actual search order. Hazard cells here are coloured by risk_group — the same value _hazard_risk_group already computes, so a hit tells you not just "how far" but "which chain." Cells outlined in white persisted from the other frame unchanged — exactly the cells a rollover would reuse instead of recomputing.

Eager flood-fill
flooded cells this frame: · persisted from other frame:
hazard cell (distance 0) flooded, distance 1 flooded, distance 3
Lazy ring search
cells checked so far: 0 · persisted from other frame:
candidate — safe within radius 3 candidate — danger found
risk_group:
Click a candidate dot above to trace its search.
02

What it cost, measured

Same real recording (full-campaign-flow-field-1), worst-case horizon (56 frames) forced on every decision frame, benchmarked in isolation against the current production _score cost for the same range.

approachbuildquerytotal / framevs. current _score
current _score (production, mixed horizons)——~62msbaseline
occupancy only, no distance signal1.14ms0.11ms1.25ms~50x
eager flood-fill, 4px cells, radius 39.91ms0.12ms10.04ms~6x
eager flood-fill, 4px cells, radius 619.32ms0.14ms19.46ms~3x
separable exact-distance transform (Felzenszwalb-Huttenlocher), radius 346.60ms0.21ms46.81msslower than baseline pieces
lazy ring search + risk_group, 4px cells, radius 31.47ms1.64ms3.10ms~20x
+ shared bbox broad-phase pre-check, radius 31.42ms0.94ms2.36ms~26x
lazy ring search alone at radius 18 (the real 70px alignment-gate threshold)0.82ms32.64ms33.46msworse than half of _score
+ broad-phase pre-check at radius 180.95ms17.24ms18.19msstill too slow to use this way

The theoretically optimal O(n) exact-distance algorithm lost badly in practice: its cost tracks the bounding box around each horizon's hazards, and Python's per-cell overhead in a dense grid dominates over asymptotic complexity at this problem size. Lazy ring search's cost tracks query count instead (504 lookups/frame, fixed), which is why it wins regardless of how spread out the hazards are. Adding risk_group cost ~0.4ms — free at the hit cell, no separate lookup. All 9 candidates query the same per-horizon occupancy set, so a single shared bounding-box pre-check can rule out a whole horizon for all 9 at once before falling back to per-candidate search — a real ~24% win at radius 3, and roughly halves the cost at radius 18. It does not make radius 18 viable on its own, though: _score also uses a 70px threshold check (nearest < 70.0, autoplay.py:25926) to relax a firing-lane-alignment penalty near danger — a tier-5, non-survival gate that doesn't need per-candidate horizon-swept precision. The right fix there is a separate, cheap, frame-global check against hazards' current positions (~13 math.hypot calls), not a wider radius on this field.

horizon forced
56 (worst case)
real hazards this frame
occupied cells (pre-flood)
isolated-benchmark winner
lazy, radius 3
03

Integrated into _score, then reverted

The isolated benchmark above never shipped: it was wired into _score behind a use_hazard_field flag, validated correct against 717 tests and a real live Level 1 playthrough, then reverted after two rounds of fixing still left it slower than the code it was meant to replace. Recorded here because the isolated number was real and the integrated regression was also real — the gap between them is the actual lesson.

Live Level 1, three runs, same start snapshot, same stop condition
mean decision time: 35.5ms (off) → 47.2ms (on, buggy) → 46.0ms (on, fixed twice)
All three runs reached the shop safely, zero lives lost — the survival bar was met every time. Only the speed goal failed.
runmeanp99maxshieldlives lost
off (baseline)35.5ms344ms1012ms39→350
on — hazard rects expanded by 16px player margin47.2ms488ms761ms39→350
on — fix 1: raw rects, margin folded into ring radius46.0ms312ms738ms39→310

Two real bugs were found and fixed along the way, each confirmed by re-measuring rather than assumed:

bug 1 — fixed, insufficient alone

Hazard-side Minkowski expansion

To make a raw center-point query equivalent to a full hull-sized rect query, every hazard rect was expanded by a 16px player-hull margin before rasterizing. That inflated each hazard's footprint by roughly 9x in cell count. Fix: rasterize hazards raw, fold the same margin into the ring-search radius instead (same Minkowski sum, applied at the query side) — free, since the ring search already walks outward.

skip rate 66.3% → 74.3%
bug 2 — fixed, still insufficient

Four query points per candidate, not one

The isolated benchmark checked one point per candidate per horizon. _score actually needs four (predicted_player, swarm_player, formation_player, projectile_player) to stay safe across every classification branch. Fix: one shared bounding-box pre-check across all four before falling back to precise per-point ring search — a real but modest ~4% win, because the 25.7% of checks that DO find something nearby still pay for the full precise search on all four points, and Level 1's swarm/formation combat makes that the common case, not the rare one.

confirmed logically identical: skip rate unchanged at 74.3%

A methodology correction, caught before it became a wrong conclusion: the first attempt at explaining why the regression persisted used validate_campaign.py profile to re-profile the fixed code with cProfile. That tool rebuilds its controller through ReplayModel.__init__ (replay_ui.py:364), which constructs ReactiveController(wall_mask=self.wall_mask) with no way to pass use_hazard_field — a path that was only ever wired through the live record_run.py route. That cProfile pass was silently measuring the flag-off baseline the whole time, not the fixed "on" configuration it was meant to inspect. The resulting "_score self-time unchanged, 803M vs. 772M calls" explanation was retracted once this was caught — it was never evidence about the code being evaluated.

What survives that correction: the live wall-clock numbers above are still valid (measured through the correctly-wired record_run.py --use-hazard-field path), and so are the skip-rate and cell-count diagnostics for both fixes (measured by explicitly setting controller.use_hazard_field = True after construction, bypassing the same gap). The regression itself is real and measured twice over. The specific mechanistic explanation for why it persists after both fixes is not actually diagnosed — that would need a live-path-correct cProfile pass this investigation didn't get to before the decision to revert and return to the architecture question directly.

status
reverted
use_hazard_field flag
removed from autoplay.py
correctness while it existed
hard-block / collision-frame fields exact
next step
architecture question, not another patch