Xenon 2

How it works · write-up

AI errors — shop screen investigation session

xenondoc/AIERRORS.MD · 23 KB · updated 2026-10-06

A retrospective on a single long session spent getting the Xenon 2: Megablast shop screen working in Hatari's Sprite Stream reconstruction. Built from the actual session transcript (not just memory), so entries are grounded in what really happened. Roughly chronological. Kept so the same mistakes aren't repeated.

Autoplay/navigation errors are maintained separately in AUTOPILOT_AIERRORS.MD. Its 2026-08-24 Level-2 completion addendum records the excessive full-run iteration, misuse of generic mechanisms for known route facts, late scripted-state analysis, performance bloat, regression causes, and the low-cost operating rules that must be followed before proceeding through Level 3.

Errors

1. Two entire draw systems had no capture hook at all — not bugs, gaps

The shop's small dialog text ("O.K. WHAT DO YOU WANT TO BUY?") and the alien's hand prop both used draw routines nothing had ever hooked, so they simply never appeared — reported live as "the small text... does not render" and (later) "none of [the hands] appear or are animated." Neither was a logic bug; the capture simply didn't exist yet for that code path. Worth checking "is there a hook at all" before assuming an existing one is wrong.

2. Fixing the missing text hook caused a much bigger regression

Adding the text hook alone made things "worse": item-grid slots stopped drawing, the selector highlight disappeared, text appeared then vanished a frame later, and animated-sprite positions looked wrong. Root cause (correctly self-diagnosed by the user, not me): the shop's own draw routines only redraw when something changes, relying on ST framebuffer persistence for everything else — our renderer has no such persistence and redraws from scratch every frame. Fixed by generalizing the existing title-logo persistence-cache pattern into a general shop cache. A large, correct fix, but one that only existed because the first hook was added without this already being accounted for.

3. Fixing the shop caused an unrelated gameplay regression: the background vanished

After enabling shop content, ordinary gameplay's background sprite stopped rendering. Root cause: the debug sprite-catalog capture uses a fixed-size 1280×1000 layout canvas, separate from the atlas packer's own (unlimited) capacity — a region that doesn't fit the fixed canvas is silently dropped before the real packer ever sees it. Enabling the shop's raw96 regions pushed the canvas over capacity and the background (positioned late in the region list) was the casualty. Fixed by bumping canvas height to 2000. A capacity ceiling nobody was watching, tripped by unrelated work.

4. Stack-offset bug in the animated-sprite-loop hook

The hook read its per-iteration discriminator from (SP) directly, but the actual loop counter sat 4 bytes deeper — an intervening move.l A0,-(SP) in the game's own code (preserving a walk pointer across the call) shifted what was "on top of stack" at hook-fire time. Different animated sprites (eyes, mouth, etc.) aliased onto the same cache slot and silently overwrote each other — almost certainly the reported "eye rendered too far right, other facial animations wrong." Found via manual disassembly, not the decompiler (which obscured the stack layout). Fixed by reading SP+4.

5. Hand-decoded a chrome data table from raw bytes — failed twice

Tried to manually decode the shop's static-chrome list format (borders/dividers/button bar) from raw instruction bytes twice, caught a real bit-decoding mistake mid-process both times, and the numbers still didn't resolve to sane addresses. Only worked once the approach changed (at the user's suggestion) to hooking the copy routine's own entry and reading its already-resolved parameters live off a real shop visit. The live approach should have been the first attempt, not the third.

6. Chrome data: wrong byte-interleaving convention, caught by a second AI's review

Even after the chrome data itself was captured correctly, it was decoded using the wrong pixel format — a byte-interleaved "font glyph" convention instead of the standard word-interleaved tile convention every other piece of UI chrome in this codebase uses. Not caught through my own verification; the user independently ran the same question past a different AI (Codex), whose analysis flagged it. I then confirmed Codex was right by reading the actual decoder function used for the closest existing analogue (the logo/status-bar panel) — a check that should have happened before picking a decoder, not after being contradicted.

7. Build failure from an incompletely-enabled code path

After wiring up the corrected chrome decoder, the build failed with "implicit declaration of function" — the function's own definition, not just its call site, was still wrapped in a separate #if 0 block that had been missed. A grep-verified #if 0/#endif balance check before declaring the change complete would have caught this without needing a failed build first.

8. Raw96 hook fired once per row instead of once per call — collapsed the alien head to a sliver

Hooked blitter_DrawRawPanelBitmap96_FUN_00011f6e's shared entry point, following the same pattern used successfully elsewhere. Looked architecturally sound, but this specific function's own per-row copy loop branches back to its own entry instruction every row (dbf D0w,0x11f6e) — so the hook fired 52 times for one real "alien head" draw, once per row, each firing overwriting the same cache slot. Only the last row (a 1-pixel-tall sliver) ever stayed cached — the entire visible cause of "large parts of alien remain invisible." Found unambiguously from a diagnostic log the user pasted back (sourceOffsetRows climbing 0→39, heightRows falling 52→13 across 40 consecutive firings within one real draw). Fixed by hooking the six known caller sites instead of the shared callee entry. A hook placed at a shared entry point needs to be checked for whether that entry point can re-execute itself internally, not just whether the call site "looks like" every other one already hooked.

9. Door/alien position was 7px too far left — found only because the user measured it by hand

After the raw96 fix above, the alien portrait and doors rendered at the wrong X position, badly enough to overwrite part of the left chrome border. Static/disassembly analysis of the destination math didn't reveal a bug (the byte-offset arithmetic checked out). Only resolved once the user manually measured two screenshots pixel-by-pixel and reported the real renderer's door content starting at native x=215 versus this renderer's x=208 — an exact, fixed 7px gap, which turned out to be the atlas crop boundary (the raw96 source columns 7-102 are the actual visible sub-image, not the full 112px stored block). A position bug that direct pixel measurement resolved in one message, after disassembly alone hadn't found it.

10. Door2's progressive reveal was a vertical squash, not a crop

The second shop door's opening animation was implemented so the whole sprite image scaled to fit a shrinking box (first and last atlas rows always visible, everything in between compressed) — reported live as being "squeezed vertically... as if the whole atlas image is being rendered in a continually smaller screen area." The correct behavior is a crop: the bottom rows should stop being drawn as the door closes, with the remaining rows staying at native scale. A scale-vs-crop mixup in the reveal-window math for this specific door, separate from the (correctly cropped) reveal logic used elsewhere.

11. Misattributed an animated-sprite position glitch to extrapolation that wasn't even running

After ruling out several other explanations for "renders correctly, then wrong position," I proposed that spriteStreamInterpolation.c's extrapolation was the cause and "fixed" it by forcing isScaledSize=true. The user's direct correction: extrapolation is gated behind a key press they had not pressed during that capture, so the code path I blamed was never active and couldn't be the explanation. A plausible-sounding theory offered without first checking whether the blamed subsystem was even running.

12. Cash counter position was wrong in two different ways, found one at a time

First bug: the digit-drawing hook assumed the live draw-buffer pointer told it which screen buffer was current, but the real routine writes to both buffers via a hardcoded address delta — using the live pointer as the reference occasionally picked the wrong buffer and pushed the whole counter Y off-screen (reported as "counter not visible"). Second bug, found only after fixing the first: the fix assumed the usual 16px word-granularity every other hook in the file uses, but this routine actually addresses individual bytes (8px granularity) — reported live as "every second digit" / "spaced 2x too wide" (the user's own debugger-supplied register values, alternating +7,+1 byte deltas, revealed the real granularity). Two separate wrong assumptions about the same 20 lines of code, each only found by asking the user to read live register/memory values off their own debugger.

13. Grid-cell chrome collided because x/y alone wasn't actually a unique key

Assumed a tile's screen position (x/y) was enough to tell two different sprites apart in the cache. It wasn't — disassembly showed both sprites for a cell are drawn with identical x/y (confirmed: D0/D1 are saved and restored unchanged between the two draw calls), so they collided under the same synthetic ID and silently overwrote each other, reported as "top and bottom chrome for each cell is not rendered." Fixed by discriminating on a bit derived from which of the two call sites fired, not on position.

14. "Corrected" a threshold based on a stale comment, not fresh disassembly

Changed a door-visibility gate from <= 8u to <= 9u because an old ReVa plate comment claimed > 9. Fresh disassembly showed the real condition was 8 < DAT_4ec (i.e. <= 8u was already correct) — the "fix" was a pure regression. Reverted once re-checked.

15. Misdiagnosed "winged doors don't fully open" on the first attempt

First fix targeted the raw96 door-panel threshold — confirmed a no-op via the user's own live debugger (the door cache entries it was supposed to invalidate were "always null"). The real cause (six item-icon sprites never invalidated once doorStep settled at 0) was a completely different draw call, gated on a different/outer condition, found only after re-reading the full disassembly instead of trusting the first plausible-looking gate.

16. Guessed noise-quad dimensions three times before getting ground truth

32×28 (rough guess) → widened to 38×36 without re-verifying against disassembly (user rejected this, explicitly said "verify this with ReVa") → re-read disassembly, found a missed lea instruction, landed on 34×28 (still a compromise, not verified) → only actually correct once the user clicked the real sprite-table asset with the mouse-click inspector and reported its true size: 32×28. Three iterations to reach a number that live inspection would have given immediately.

17. Vague "frame bucketing" explanation instead of reading the code

Asked why a real-hardware effect wasn't reproduced, answered with a plausible-sounding but unverified theory about frames merging together. The user pushed back ("what bucketing... explain more clearly") — the real answer required actually reading DrawCommandStream_OnMemoryWrite, which took two minutes and gave a precise, code-grounded answer. Should have read the code before answering, not after being challenged.

18. Misread a write-log correlation because I didn't check the cycle delta

Found that a grid slot's state field got written to 0 then reseeded moments later, concluded this was the cause of a visible noise flash. Never checked how many cycles separated the two writes (248 — far too fast for any other code to observe the intermediate value). The user's pushback ("Atari renderer does not replace noise until next frame... how did we draw and hide noise in the same frame?") forced a recheck, which disproved the theory. The real cause was a separate, previously-dismissed code path (see #19).

19. Dismissed a real code path as "nothing to hook here" without checking its callers

An early pass noted a second call site to the noise generator (state 1-2 branch) but reasoned "no sprite draw happens here, nothing to capture" and moved on. That assumption was never checked against the actual caller list. It was the real mechanism behind occupied cells randomly flashing noise, only found much later via a live MemoryAccessLog capture — after two other wrong theories (#18, and an earlier stale-cache-invalidation guess) had already been tried and failed.

20. Coordinate-scaling bug in the mouse-click tool I built myself

Added a debug-window click inspector without checking whether click coordinates arrived in the window's logical render resolution or its actual displayed pixel size. They were raw window pixels; the window was displayed at 2x. Clicks near the origin "worked" by coincidence (small physical and logical coordinates overlap there), which masked the bug for many clicks until one landed far enough away to return "nothing found." The user had to suggest the possibility ("is it possible the click handler uses screen coordinates instead of renderer coordinates") before it was found and fixed.

21. A coordinate-bug-corrupted click sent an entire investigation down the wrong path

Before the bug in #20 was found, a click that should have hit real grid content instead only matched a background decoration (because the click coordinates were silently wrong). This got written up as "no capture hook exists for this element at all" — a full false conclusion, corrected only after #20 was fixed and the same click was retried.

22. Assumed reverse array order in the shop quad list was a reliable z-order

The click-inspector's hit-test stopped at the first match, walking the quad array in reverse on the assumption that later-in-array meant "drawn on top." The shop's own replay order is explicitly not a deliberate z-order (its own code comment says so) — the assumption was never checked against that comment before being built on. Fixed by reporting every overlapping match instead of guessing which one is "correct."

23. isNoise flag leaked across reused cache slots

When a shop cache slot was invalidated and later reused for a different kind of content, its isNoise flag was never reset, so the new content could get replayed as flat-color noise wearing someone else's object ID. Existed for a full round of the shop-cache work before being found via a live click that showed a background quad rendering as type=flatColor — a combination that could only mean this exact leak.

24. Noise never actually animated, for a purely mathematical reason

The noise seed was computed as SDL_GetTicks() % 100000 cast to float — always a whole number. The shader's hash does fract(x + seed) as its first operation, and fract(x + integer) == fract(x) always — the entire time-varying input was being silently discarded by the very first line of the hash, every frame. Survived two rounds of live testing and bug reports before a diagnostic log proved the seed was updating correctly and the bug had to be in the shader math itself.

25. Selector-cursor fix merged two independent UI elements into one shared slot

Fixed a duplicate-cursor bug by routing all 8 of two functions' possible sprites to one shared cache slot, assuming they were all "the same moving cursor at different moments." Two of those sprites were actually a completely different, simultaneously-visible element (the BUY/SELL button label), reached through the same two functions only because the game reuses them for two different UI modes. This introduced a new bug (the label reverting to the wrong text) that the user found and diagnosed largely on their own.

26. Sized the memory-access-log ring buffer for "some capture," not the actual workflow

Built a live-memory-log dump tool without accounting for human reaction time: the default buffer covered under half a second before wrapping, useless for "notice something, then press the dump hotkey." Only caught once the user described their actual capture timing and the covered duration was computed after the fact — should have been sized for the workflow from the start.

27. Declared a Level-4 target passive before following its phase-4 tail call

The first pass over Level 4 routine $51260 correctly identified its world anchor, collision rectangle, tile animation, health callback, and destruction behavior, but concluded that it did not fire. That conclusion stopped at the local phase logic. Phase 4 prepares spawn registers and tail-calls shared allocator $E4E, which installs the ordinary $4180 directional-projectile update procedure. A live identity trace then showed every lethal corridor shot appearing beside the still-live $51260 object, forcing the incomplete interpretation to be reopened.

The routine is now named Level4DirectionalWallShooter_Update; its phase/timer, spawn offsets, facing-dependent trajectories, and speed are modeled symbolically and covered by a pre-allocation regression test. The general lesson is that "this routine does not allocate" cannot be established by searching only for a local object-allocation call: follow tail calls and shared vector thunks through the state that prepares their input registers, then correlate the result with live object births.

Patterns behind these

  • Guessing dimensions/positions/granularity instead of measuring them. #9, #12, #16, and to a lesser extent #20/#21, would have been resolved in one step by reading the real value (a live register, a clicked sprite, a memory dump, a measured screenshot) first, instead of assuming and iterating on user feedback.
  • Trusting an old comment or a first-pass assumption without re-verifying it. #6, #13, #14, #15, #19, #22 all involved treating something already written down or previously assumed (a plate comment, an earlier "nothing to hook here" note, a positional key, an ordering assumption) as settled fact instead of checking it against current disassembly/code.
  • Assuming a hook exists and is correct, rather than checking whether it exists at all. #1 was the clearest case — the actual problem was "there's no capture logic here yet," not "the existing logic is subtly wrong."
  • Hooking a shared entry point without checking what happens inside it. #8 specifically: the hooked function's own internal loop looped back through its own entry instruction, so "hook the entry, same as every other function" silently fired once per row instead of once per call.
  • Explaining before checking. #11, #17, #18 were all cases of answering a "why" question (or proposing a fix) with a plausible narrative instead of reading the actual code, checking whether the blamed subsystem was even active, or checking the actual numbers (cycle deltas) first.
  • Fixing one thing without checking what else it breaks. #3 (chrome work broke gameplay background via a shared fixed-size buffer) and #25 (a cursor fix broke an unrelated button label sharing the same draw functions) are both cases where a locally-correct fix had an un-audited side effect elsewhere.
  • An external second opinion caught what I missed. #6 was only caught because the user independently cross-checked my analysis against a different AI's — worth remembering that disagreement between two analyses is itself useful signal, and worth inviting rather than being defensive about.
  • Building tools without modeling how they'll actually be used. #26 sized a buffer for "a capture" rather than "a human reacting to something they just saw."

Recommendations

  1. When a number (size, position, offset, granularity) is guessable and directly observable via the click-inspector, a live register/memory read, or a captured log — observe it, don't guess. Every guess-then-correct cycle in this session cost a full round trip with the user; live tools were available the whole time.
  2. For any visible game element that isn't rendering, check whether a capture hook exists at all before assuming an existing one is wrong. "Missing entirely" and "present but buggy" look the same from the outside but need different investigation.
  3. When hooking a shared entry point (used by multiple callers, or with its own internal loop), verify what happens on repeat/re-entrant execution before trusting "once per call." A function whose loop body branches back through its own entry instruction will fire a hook placed there once per iteration, not once per invocation.
  4. Before reusing a synthetic ID/cache slot for multiple sprites or call sites, explicitly list every context that can produce each one. If two of them can be true/visible at the same time, they need separate slots, no matter how similar the code drawing them looks.
  5. Treat "nothing to hook here" as a claim to verify, not a conclusion. If a code path calls a function with a real side effect (writes memory, generates noise, sets a flag other code reads), check its actual behavior and every caller before deciding it's a no-op for capture purposes.
  6. Re-verify old comments/notes against fresh disassembly before building on them, especially plate comments and "confirmed" notes written earlier in the same investigation — they can be stale or simply wrong, and citing them isn't the same as re-checking them.
  7. When asked "why does X happen," or before blaming a specific subsystem, read the actual code path (and check whether that subsystem was even active) before answering. A plausible explanation that turns out to be wrong costs more (in user trust and follow-up work) than the two minutes it takes to check.
  8. Check the actual timing (cycle deltas, call frequency) behind a correlation before concluding causation. Two entries in a log being near each other in a dump is not the same as one being observable by the code that would need to see it.
  9. Don't assume a "standard" addressing convention (word-granularity, a fixed buffer base, a consistent stride, a crop-not-scale reveal) applies to a new hook just because it applied to every previous one. Several bugs here came from carrying an assumption from other hooks into a new one without re-checking it against that specific routine's own disassembly.
  10. After a locally-correct fix, check for shared state or shared resources it might affect elsewhere — a shared fixed-size buffer, a shared cache slot, a shared draw function used for more than one purpose. Two regressions in this session were "correct fix, unaudited blast radius."
  11. Size diagnostic/investigative tooling for how a human will actually use it — reaction time, capture windows, buffer depth — not just for "does it technically work once."