Xenon 2

How it works · write-up

Plan: Render the background as one scrolling bitmap, not tile-by-tile

xenondoc/plan-background-scrolling-bitmap.md · 13 KB · updated 2026-09-17

Context

The wall-tile capture work (just completed, verified working) only covers tilemap cells that reference an actual wall tile. Cells with tilemap code 0 ("empty") currently render as black in RENDER_VIEW_SPRITE_STREAM, because the game fills those cells from a completely separate source: a pre-rendered background mosaic bitmap, copied in via a scrolling cursor (PTR_DAT_0436_currentTileMapPointer) rather than looked up per-tile like walls are.

The original 68000 code reconstructs this by copying the mosaic tile-by-tile, 16x16 pixels at a time, once per row-band, because that's how a 68000 blitter has to work. We don't have that constraint. Since the background is already just one big pre-rendered image (not per-cell authored data the way wall tiles are), we can capture it as a single scrolling bitmap and let the GPU handle both the vertical clipping and the wraparound natively — no tile reconstruction loop needed at all, and (per the corrected analysis below) no CPU-side quad-splitting either.

What we found (corrected after a first pass got this wrong — verified against the actual live

update code, not just the offline scanner's assumptions)

Atlas export is already done. RenderWholeBackground (xenonRender.c, already runs today as part of the existing one-time debug-window scan) exports the mosaic as a single atlas entry — one RecordRenderedSprite call keyed by g_spriteMemoryRegion[]'s SPRITE_REGION_TYPE_BACKGROUND entry (stAddr = 0x0006989c), tagged SPRITE_FORMAT_2PLANES. The .NET side already fully decodes and packs this format (confirmed earlier this session). No atlas/export-side changes needed — SpriteAtlas_FindRect(0x0006989c, ...) already resolves to the image's atlas rect.

This is genuine parallax — confirmed, not assumed. Traced the actual cursor-update code, FUN_0000702c (called once per frame from the main loop, before the tile/background draw call). It reads its step count from a register area (0x00000cda/0xcdc) that is completely separate from the wall tilemap's own scroll registers (0xcd6/0xccc) — a distinct mechanism, consistent with it visibly moving slower than the tile layer.

The wrap period is 192 scanlines (0x300 bytes), not 384 — this was the actual bug in the first draft of this plan. FUN_0000702c's own wrap check is unambiguous in the disassembly:

puVar1 = cursor + 4;
if (puVar1 > background_base + 0x300) puVar1 -= 0x300;   // (mirrored for the decrementing direction)
cursor = puVar1;

background_base (PTR_0004f000_points_to_Background) was confirmed via direct memory read to equal 0x0006989c exactly — the same base RenderWholeBackground uses. So the cursor only ever ranges over [base, base+0x300) = 768 bytes = 192 scanlines (at 4 bytes/scanline for this 2-bitplane format) — not the full 384-row image RenderWholeBackground exports. The earlier % 384 in this plan's first draft was carried over from that export's dimensions without cross-checking it against the actual live update logic — the same mistake pattern as the +16 bug fixed earlier in the wall-tile work, caught the same way: by going back to the disassembly instead of trusting an existing constant.

192 is not a coincidence — it's the game's real playfield height. It exactly matches the wall tile drawing function's own total per-frame row coverage, already established earlier this session: 11 full middle rows × 16px = 176, plus the top and bottom partial row-bands always summing to exactly 16px between them = 192 total. Both layers are built around the same 192-line drawable area; the real 200-line ST screen has ~8 lines reserved for something else (consistent with your earlier observation that the bottom edge is normally covered by a status bar).

Sample height equals wrap period exactly (192 == 192) — this is what makes a single quad with GPU-side wrap possible. Since the visible window is always exactly one full period of the source, this is a true seamless rotation, not a partial/irregular overlap — exactly matching what you described: the sample position slides toward one edge and then resets, rather than needing two simultaneously-visible source regions blended together.

Design

A single quad per frame, with the wrap handled by the shader, not split into two quads on the CPU. This mirrors the same principle just used for the wall-tile row-band fix (compute the correct source position and let the GPU do the rest, rather than reproducing the game's own clipping/wrapping logic in C) — just applied to UV sampling instead of vertex position.

Two ways to achieve GPU-side wrapping were considered: - Hardware texture repeat-addressing on the sampler — rejected: the shared sprite atlas's sampler needs clamp behavior for every other packed sprite (to avoid bleeding into neighbors in the packed texture), and switching it to repeat mode for this one draw would require carving the background out into its own dedicated texture + sampler + a way to select which texture a given quad samples from — real new GPU-side plumbing for a problem that has a smaller fix available. - Manual wrap in the fragment shader, computed from the quad's own known sub-region bounds (recommended): since we already know the background's atlas sub-rect (atlasY, height 192) from the existing SpriteAtlas_FindRect lookup, the shader can wrap the sampled V coordinate within that sub-rect itself (wrappedV = atlasYStart + mod(v - atlasYStart, 192)) without touching the sampler's address mode at all — every sample stays within the sub-rect, so there's no bleed risk, and the shared atlas texture/sampler stay exactly as they are today. This is the same shape of change as the flash-tint feature (one more per-vertex flag, one more small per-frame uniform) — reusing an already-proven pattern in this exact file rather than introducing a new one.

New capture hook at the tile-draw function's entry (0x1dd0). Already reserved as an unused hook point from the wall-tile work. ST_DrawGeneratedBackgroundFromTileMap_FUN_00001dd0 runs once per frame, before the three row-band tile-fetch hooks, so pushing the background entry here first guarantees correct z-order (background drawn behind walls/objects) for free, via the same capture-order-preserves-z-order mechanism already relied on for wall tiles. Just a direct STMemory_ReadLong(0x436) — the cursor is a stable global, not a transient register tied to a specific instruction.

Implementation

src/includes/sdlGpuRenderView.h / src/sdlGpuRenderView.c — small extension of the existing flash-tint mechanism

  1. SpriteStreamVertex: add one more field, float bgWrap; (0.0 = normal sample, 1.0 = this quad's V should be wrapped) — same shape as the existing flash field.
  2. UniformBufferSpriteStream: add float bgWrapVStart; and float bgWrapVPeriod; (both in normalized UV terms: atlasY/atlasHeight and 192/atlasHeight) — recomputed each frame in SdlGpuRenderView_Submit the same way flashColor already is, from a value threaded in from the capture side (see below) rather than re-deriving it from a hardcoded region lookup inside sdlGpuRenderView.c itself.
  3. vertex_sprite_stream.glsl: pass bgWrap through unchanged (layout(location = 3) in float inBgWrap; ... outBgWrap = inBgWrap;), same pattern as inFlash/outFlash.
  4. sprite_atlas_fragment.glsl: add bgWrapVStart/bgWrapVPeriod to the existing uniform block, and: glsl float v = inUV.y; if (inBgWrap > 0.5) { v = bgWrapVStart + mod(v - bgWrapVStart, bgWrapVPeriod); } vec4 texColor = texture(atlasTex, vec2(inUV.x, v));

src/drawCommandStream.c / .h

  1. New constants: TILE_DRAW_FUNCTION_ENTRY_PC = 0x00001dd0u, BACKGROUND_SCROLL_CURSOR_PTR = 0x00000436u, BACKGROUND_MOSAIC_BASE = 0x0006989cu, BACKGROUND_PLAYFIELD_HEIGHT = 192.
  2. New function DrawCommandStream_OnBackgroundDrawInstructionFetch(uint32_t addr): early-out unless addr == TILE_DRAW_FUNCTION_ENTRY_PC. Look up the background's atlas rect once via SpriteAtlas_FindRect(BACKGROUND_MOSAIC_BASE, ...). Read the live cursor, compute scrollY = ((cursor - BACKGROUND_MOSAIC_BASE) / 4u) % BACKGROUND_PLAYFIELD_HEIGHT. Push one DrawMaskedSpriteEntry covering the full playfield (x=0, y=0, width/height = full atlas width × 192), with its atlas rect's Y set to atlasY + scrollY (this is deliberately allowed to read past the sub-rect's own bottom edge in the V direction the vertex builder computes it — that's fine, the shader's wrap correction is what actually keeps the sample valid, not the rect math). This needs a new bit of information threaded alongside the existing fields: whether this entry should be flagged bgWrap=1 and what atlasY/192 (in pixels) are, so the vertex builder (screentrace.c) can populate the new uniform inputs. Simplest approach: extend DrawMaskedSpriteEntry with bool isBackgroundWrap (mirrors how isFlash was added) — the vertex builder already has atlasY/atlasH on the entry, which is exactly atlasYStart/192 in pixel terms, so no additional fields are needed beyond the one bool.
  3. Call it from the same chokepoint (ScreenTrace_LogRead, src/screentrace.c) alongside the two existing tile/object hooks.
  4. Update drawCommandStream.h's doc comments, matching the existing style for the other two hooks.

src/screentrace.c

SpriteStreamVertexBuilder_AppendEntry: set the new bgWrap vertex field from entry->isBackgroundWrap (same pattern as flash from entry->isFlash). When pushing the background's quad, note its atlasY/atlasH fields are what the fragment shader will read back out (via the per-frame uniform, populated in SdlGpuRenderView_Submit from these same values threaded through) as bgWrapVStart/bgWrapVPeriod — needs a small amount of plumbing to get those two numbers from "the one background entry in this frame's batch" into the per-frame uniform push, since today SdlGpuRenderView_Submit only receives the vertex array, not the original entries. Simplest: have SpriteStreamVertexBuilder also record the background entry's atlas Y/height (there's only ever at most one per frame) as it walks entries, and thread that one extra (atlasY, height) pair through to SdlGpuRenderView_Submit alongside the existing vertex array pointer/count.

Known risks / verify empirically

  1. Scroll direction/rate not fully traced. We deliberately don't need to reproduce FUN_0000702c's exact step-size math (we just read its result, the live cursor, each frame), but the exact speed ratio vs. the tile layer hasn't been measured — only confirmed to be a separate mechanism. Worth a visual sanity check once rendering, not just a "does it move" check.
  2. Whether rows 192-383 of RenderWholeBackground's exported 384-row image are meaningful data or unused padding is unconfirmed. This plan only ever samples the first 192 rows (matching the live cursor's real range), so it doesn't matter either way for correctness — but if a future pass wants to shrink the export to match, that's a separate, optional cleanup.
  3. Direction of wrap — confirmed both directions are actually used. The game supports genuine backward/reverse vertical scroll in addition to forward scroll (not just a theoretical code path — confirmed), so both the incrementing and decrementing cases in FUN_0000702c are live in normal play. This doesn't complicate the design: since we only ever read the already-wrapped live cursor value (the update function's own code keeps it within [base, base+0x300) at all times, regardless of direction), the C-side scrollY computation needs no direction-awareness. GLSL's mod(x, y) is defined to always return a result in [0, y) for any sign of x, so the shader-side wrap formula is already correct for both directions with no special-casing. Worth a visual check in both scroll directions once implemented, but not expected to need a design change.

Verification

  1. Rebuild, recompile shaders (compilerShaders.bat) — this touches both shader files.
  2. Open "Sprite Stream" during actual gameplay with the level scrolling; confirm previously-black empty-cell areas now show the scrolling background, moving visibly slower than the wall tiles (parallax).
  3. Confirm no seam/tear/duplicate-row artifact as the scroll position crosses the 192-line wrap boundary — this is the main thing the shader-side mod() needs to get exactly right.
  4. Confirm wall tiles and ship/enemies still render on top of the background (z-order intact).
  5. Confirm the other 3 debug window types are unaffected by the SpriteStreamVertex/uniform changes (same regression check as after the flash-tint and wall-tile changes).