How it works · write-up
Plan: Render the background as one scrolling bitmap, not tile-by-tile
Context
The wall-tile capture work (just completed, verified working) only covers tilemap cells that
reference an actual wall tile. Cells with tilemap code 0 ("empty") currently render as black in
RENDER_VIEW_SPRITE_STREAM, because the game fills those cells from a completely separate source: a
pre-rendered background mosaic bitmap, copied in via a scrolling cursor
(PTR_DAT_0436_currentTileMapPointer) rather than looked up per-tile like walls are.
The original 68000 code reconstructs this by copying the mosaic tile-by-tile, 16x16 pixels at a time, once per row-band, because that's how a 68000 blitter has to work. We don't have that constraint. Since the background is already just one big pre-rendered image (not per-cell authored data the way wall tiles are), we can capture it as a single scrolling bitmap and let the GPU handle both the vertical clipping and the wraparound natively — no tile reconstruction loop needed at all, and (per the corrected analysis below) no CPU-side quad-splitting either.
What we found (corrected after a first pass got this wrong — verified against the actual live
update code, not just the offline scanner's assumptions)
Atlas export is already done. RenderWholeBackground (xenonRender.c, already runs today as
part of the existing one-time debug-window scan) exports the mosaic as a single atlas entry — one
RecordRenderedSprite call keyed by g_spriteMemoryRegion[]'s SPRITE_REGION_TYPE_BACKGROUND
entry (stAddr = 0x0006989c), tagged SPRITE_FORMAT_2PLANES. The .NET side already fully decodes
and packs this format (confirmed earlier this session). No atlas/export-side changes needed —
SpriteAtlas_FindRect(0x0006989c, ...) already resolves to the image's atlas rect.
This is genuine parallax — confirmed, not assumed. Traced the actual cursor-update code,
FUN_0000702c (called once per frame from the main loop, before the tile/background draw call).
It reads its step count from a register area (0x00000cda/0xcdc) that is completely separate
from the wall tilemap's own scroll registers (0xcd6/0xccc) — a distinct mechanism, consistent
with it visibly moving slower than the tile layer.
The wrap period is 192 scanlines (0x300 bytes), not 384 — this was the actual bug in the first
draft of this plan. FUN_0000702c's own wrap check is unambiguous in the disassembly:
puVar1 = cursor + 4;
if (puVar1 > background_base + 0x300) puVar1 -= 0x300; // (mirrored for the decrementing direction)
cursor = puVar1;
background_base (PTR_0004f000_points_to_Background) was confirmed via direct memory read to
equal 0x0006989c exactly — the same base RenderWholeBackground uses. So the cursor only ever
ranges over [base, base+0x300) = 768 bytes = 192 scanlines (at 4 bytes/scanline for this
2-bitplane format) — not the full 384-row image RenderWholeBackground exports. The earlier %
384 in this plan's first draft was carried over from that export's dimensions without
cross-checking it against the actual live update logic — the same mistake pattern as the +16
bug fixed earlier in the wall-tile work, caught the same way: by going back to the disassembly
instead of trusting an existing constant.
192 is not a coincidence — it's the game's real playfield height. It exactly matches the wall tile drawing function's own total per-frame row coverage, already established earlier this session: 11 full middle rows × 16px = 176, plus the top and bottom partial row-bands always summing to exactly 16px between them = 192 total. Both layers are built around the same 192-line drawable area; the real 200-line ST screen has ~8 lines reserved for something else (consistent with your earlier observation that the bottom edge is normally covered by a status bar).
Sample height equals wrap period exactly (192 == 192) — this is what makes a single quad with GPU-side wrap possible. Since the visible window is always exactly one full period of the source, this is a true seamless rotation, not a partial/irregular overlap — exactly matching what you described: the sample position slides toward one edge and then resets, rather than needing two simultaneously-visible source regions blended together.
Design
A single quad per frame, with the wrap handled by the shader, not split into two quads on the CPU. This mirrors the same principle just used for the wall-tile row-band fix (compute the correct source position and let the GPU do the rest, rather than reproducing the game's own clipping/wrapping logic in C) — just applied to UV sampling instead of vertex position.
Two ways to achieve GPU-side wrapping were considered:
- Hardware texture repeat-addressing on the sampler — rejected: the shared sprite atlas's
sampler needs clamp behavior for every other packed sprite (to avoid bleeding into neighbors in
the packed texture), and switching it to repeat mode for this one draw would require carving the
background out into its own dedicated texture + sampler + a way to select which texture a given
quad samples from — real new GPU-side plumbing for a problem that has a smaller fix available.
- Manual wrap in the fragment shader, computed from the quad's own known sub-region bounds
(recommended): since we already know the background's atlas sub-rect (atlasY, height 192)
from the existing SpriteAtlas_FindRect lookup, the shader can wrap the sampled V coordinate
within that sub-rect itself (wrappedV = atlasYStart + mod(v - atlasYStart, 192)) without
touching the sampler's address mode at all — every sample stays within the sub-rect, so there's
no bleed risk, and the shared atlas texture/sampler stay exactly as they are today. This is the
same shape of change as the flash-tint feature (one more per-vertex flag, one more small
per-frame uniform) — reusing an already-proven pattern in this exact file rather than introducing
a new one.
New capture hook at the tile-draw function's entry (0x1dd0). Already reserved as an unused
hook point from the wall-tile work. ST_DrawGeneratedBackgroundFromTileMap_FUN_00001dd0 runs once
per frame, before the three row-band tile-fetch hooks, so pushing the background entry here first
guarantees correct z-order (background drawn behind walls/objects) for free, via the same
capture-order-preserves-z-order mechanism already relied on for wall tiles. Just a direct
STMemory_ReadLong(0x436) — the cursor is a stable global, not a transient register tied to a
specific instruction.
Implementation
src/includes/sdlGpuRenderView.h / src/sdlGpuRenderView.c — small extension of the existing flash-tint mechanism
SpriteStreamVertex: add one more field,float bgWrap;(0.0 = normal sample, 1.0 = this quad's V should be wrapped) — same shape as the existingflashfield.UniformBufferSpriteStream: addfloat bgWrapVStart;andfloat bgWrapVPeriod;(both in normalized UV terms:atlasY/atlasHeightand192/atlasHeight) — recomputed each frame inSdlGpuRenderView_Submitthe same wayflashColoralready is, from a value threaded in from the capture side (see below) rather than re-deriving it from a hardcoded region lookup insidesdlGpuRenderView.citself.vertex_sprite_stream.glsl: passbgWrapthrough unchanged (layout(location = 3) in float inBgWrap; ... outBgWrap = inBgWrap;), same pattern asinFlash/outFlash.sprite_atlas_fragment.glsl: addbgWrapVStart/bgWrapVPeriodto the existing uniform block, and:glsl float v = inUV.y; if (inBgWrap > 0.5) { v = bgWrapVStart + mod(v - bgWrapVStart, bgWrapVPeriod); } vec4 texColor = texture(atlasTex, vec2(inUV.x, v));
src/drawCommandStream.c / .h
- New constants:
TILE_DRAW_FUNCTION_ENTRY_PC = 0x00001dd0u,BACKGROUND_SCROLL_CURSOR_PTR = 0x00000436u,BACKGROUND_MOSAIC_BASE = 0x0006989cu,BACKGROUND_PLAYFIELD_HEIGHT = 192. - New function
DrawCommandStream_OnBackgroundDrawInstructionFetch(uint32_t addr): early-out unlessaddr == TILE_DRAW_FUNCTION_ENTRY_PC. Look up the background's atlas rect once viaSpriteAtlas_FindRect(BACKGROUND_MOSAIC_BASE, ...). Read the live cursor, computescrollY = ((cursor - BACKGROUND_MOSAIC_BASE) / 4u) % BACKGROUND_PLAYFIELD_HEIGHT. Push oneDrawMaskedSpriteEntrycovering the full playfield (x=0, y=0, width/height = full atlas width × 192), with its atlas rect's Y set toatlasY + scrollY(this is deliberately allowed to read past the sub-rect's own bottom edge in the V direction the vertex builder computes it — that's fine, the shader's wrap correction is what actually keeps the sample valid, not the rect math). This needs a new bit of information threaded alongside the existing fields: whether this entry should be flaggedbgWrap=1and whatatlasY/192(in pixels) are, so the vertex builder (screentrace.c) can populate the new uniform inputs. Simplest approach: extendDrawMaskedSpriteEntrywithbool isBackgroundWrap(mirrors howisFlashwas added) — the vertex builder already hasatlasY/atlasHon the entry, which is exactlyatlasYStart/192in pixel terms, so no additional fields are needed beyond the one bool. - Call it from the same chokepoint (
ScreenTrace_LogRead,src/screentrace.c) alongside the two existing tile/object hooks. - Update
drawCommandStream.h's doc comments, matching the existing style for the other two hooks.
src/screentrace.c
SpriteStreamVertexBuilder_AppendEntry: set the new bgWrap vertex field from
entry->isBackgroundWrap (same pattern as flash from entry->isFlash). When pushing the
background's quad, note its atlasY/atlasH fields are what the fragment shader will read back
out (via the per-frame uniform, populated in SdlGpuRenderView_Submit from these same values threaded
through) as bgWrapVStart/bgWrapVPeriod — needs a small amount of plumbing to get those two
numbers from "the one background entry in this frame's batch" into the per-frame uniform push,
since today SdlGpuRenderView_Submit only receives the vertex array, not the original entries. Simplest:
have SpriteStreamVertexBuilder also record the background entry's atlas Y/height (there's only
ever at most one per frame) as it walks entries, and thread that one extra (atlasY, height) pair
through to SdlGpuRenderView_Submit alongside the existing vertex array pointer/count.
Known risks / verify empirically
- Scroll direction/rate not fully traced. We deliberately don't need to reproduce
FUN_0000702c's exact step-size math (we just read its result, the live cursor, each frame), but the exact speed ratio vs. the tile layer hasn't been measured — only confirmed to be a separate mechanism. Worth a visual sanity check once rendering, not just a "does it move" check. - Whether rows 192-383 of
RenderWholeBackground's exported 384-row image are meaningful data or unused padding is unconfirmed. This plan only ever samples the first 192 rows (matching the live cursor's real range), so it doesn't matter either way for correctness — but if a future pass wants to shrink the export to match, that's a separate, optional cleanup. - Direction of wrap — confirmed both directions are actually used. The game supports genuine
backward/reverse vertical scroll in addition to forward scroll (not just a theoretical code
path — confirmed), so both the incrementing and decrementing cases in
FUN_0000702care live in normal play. This doesn't complicate the design: since we only ever read the already-wrapped live cursor value (the update function's own code keeps it within[base, base+0x300)at all times, regardless of direction), the C-sidescrollYcomputation needs no direction-awareness. GLSL'smod(x, y)is defined to always return a result in[0, y)for any sign ofx, so the shader-side wrap formula is already correct for both directions with no special-casing. Worth a visual check in both scroll directions once implemented, but not expected to need a design change.
Verification
- Rebuild, recompile shaders (
compilerShaders.bat) — this touches both shader files. - Open "Sprite Stream" during actual gameplay with the level scrolling; confirm previously-black empty-cell areas now show the scrolling background, moving visibly slower than the wall tiles (parallax).
- Confirm no seam/tear/duplicate-row artifact as the scroll position crosses the 192-line wrap
boundary — this is the main thing the shader-side
mod()needs to get exactly right. - Confirm wall tiles and ship/enemies still render on top of the background (z-order intact).
- Confirm the other 3 debug window types are unaffected by the
SpriteStreamVertex/uniform changes (same regression check as after the flash-tint and wall-tile changes).