Real hazard geometry from two consecutive decision frames — pulled
straight from controller._scoring_hazard_sweeps, not synthesized.
Both panels answer the exact same question for the same 9 candidates; they differ
only in which cells get touched to answer it. Toggle frames to see how little
actually changes in one canonical step — the premise the rollover design
depends on.
The left panel floods outward from every occupied cell regardless of whether a
candidate will ever ask about it — cells are shaded red→amber by distance only.
The right panel does the opposite: it starts from each of the 9 candidates' predicted
positions (dots) and searches outward only as far as it needs to, click a dot to see
its actual search order. Hazard cells here are coloured by risk_group
— the same value _hazard_risk_group already computes, so a hit tells
you not just "how far" but "which chain." Cells outlined in white persisted from the
other frame unchanged — exactly the cells a rollover would reuse instead of
recomputing.
Same real recording (full-campaign-flow-field-1), worst-case horizon
(56 frames) forced on every decision frame, benchmarked in isolation against
the current production _score cost for the same range.
| approach | build | query | total / frame | vs. current _score |
|---|---|---|---|---|
current _score (production, mixed horizons) | — | — | ~62ms | baseline |
| occupancy only, no distance signal | 1.14ms | 0.11ms | 1.25ms | ~50x |
| eager flood-fill, 4px cells, radius 3 | 9.91ms | 0.12ms | 10.04ms | ~6x |
| eager flood-fill, 4px cells, radius 6 | 19.32ms | 0.14ms | 19.46ms | ~3x |
| separable exact-distance transform (Felzenszwalb-Huttenlocher), radius 3 | 46.60ms | 0.21ms | 46.81ms | slower than baseline pieces |
| lazy ring search + risk_group, 4px cells, radius 3 | 1.47ms | 1.64ms | 3.10ms | ~20x |
| + shared bbox broad-phase pre-check, radius 3 | 1.42ms | 0.94ms | 2.36ms | ~26x |
| lazy ring search alone at radius 18 (the real 70px alignment-gate threshold) | 0.82ms | 32.64ms | 33.46ms | worse than half of _score |
| + broad-phase pre-check at radius 18 | 0.95ms | 17.24ms | 18.19ms | still too slow to use this way |
The theoretically optimal O(n) exact-distance algorithm lost badly in practice: its cost tracks
the bounding box around each horizon's hazards, and Python's per-cell overhead in a dense
grid dominates over asymptotic complexity at this problem size. Lazy ring search's cost tracks
query count instead (504 lookups/frame, fixed), which is why it wins regardless of how
spread out the hazards are. Adding risk_group cost ~0.4ms — free at the hit cell, no
separate lookup. All 9 candidates query the same per-horizon occupancy set, so a single shared
bounding-box pre-check can rule out a whole horizon for all 9 at once before falling back to
per-candidate search — a real ~24% win at radius 3, and roughly halves the cost at radius
18. It does not make radius 18 viable on its own, though: _score also uses
a 70px threshold check (nearest < 70.0, autoplay.py:25926) to relax a
firing-lane-alignment penalty near danger — a tier-5, non-survival gate that doesn't need
per-candidate horizon-swept precision. The right fix there is a separate, cheap, frame-global
check against hazards' current positions (~13 math.hypot calls), not a
wider radius on this field.
_score, then reverted
The isolated benchmark above never shipped: it was wired into _score behind a
use_hazard_field flag, validated correct against 717 tests and a real live
Level 1 playthrough, then reverted after two rounds of fixing still left it slower
than the code it was meant to replace. Recorded here because the isolated number was real and
the integrated regression was also real — the gap between them is the actual lesson.
| run | mean | p99 | max | shield | lives lost |
|---|---|---|---|---|---|
| off (baseline) | 35.5ms | 344ms | 1012ms | 39→35 | 0 |
| on — hazard rects expanded by 16px player margin | 47.2ms | 488ms | 761ms | 39→35 | 0 |
| on — fix 1: raw rects, margin folded into ring radius | 46.0ms | 312ms | 738ms | 39→31 | 0 |
Two real bugs were found and fixed along the way, each confirmed by re-measuring rather than assumed:
To make a raw center-point query equivalent to a full hull-sized rect query, every hazard rect was expanded by a 16px player-hull margin before rasterizing. That inflated each hazard's footprint by roughly 9x in cell count. Fix: rasterize hazards raw, fold the same margin into the ring-search radius instead (same Minkowski sum, applied at the query side) — free, since the ring search already walks outward.
The isolated benchmark checked one point per candidate per horizon. _score actually
needs four (predicted_player, swarm_player, formation_player,
projectile_player) to stay safe across every classification branch. Fix: one shared
bounding-box pre-check across all four before falling back to precise per-point ring search —
a real but modest ~4% win, because the 25.7% of checks that DO find something nearby still pay
for the full precise search on all four points, and Level 1's swarm/formation combat makes that
the common case, not the rare one.
A methodology correction, caught before it became a wrong conclusion: the first attempt at
explaining why the regression persisted used validate_campaign.py profile to
re-profile the fixed code with cProfile. That tool rebuilds its controller through
ReplayModel.__init__ (replay_ui.py:364), which constructs
ReactiveController(wall_mask=self.wall_mask) with no way to pass
use_hazard_field — a path that was only ever wired through the live
record_run.py route. That cProfile pass was silently measuring the flag-off
baseline the whole time, not the fixed "on" configuration it was meant to inspect. The resulting
"_score self-time unchanged, 803M vs. 772M calls" explanation was retracted once this
was caught — it was never evidence about the code being evaluated.
What survives that correction: the live wall-clock numbers above are still valid (measured
through the correctly-wired record_run.py --use-hazard-field path), and so are the
skip-rate and cell-count diagnostics for both fixes (measured by explicitly setting
controller.use_hazard_field = True after construction, bypassing the same gap). The
regression itself is real and measured twice over. The specific mechanistic explanation for why it
persists after both fixes is not actually diagnosed — that would need a live-path-correct
cProfile pass this investigation didn't get to before the decision to revert and return to the
architecture question directly.