Xenon 2

How it works · write-up

Xenon 2 renderer: 4x atlas upscale + sprite-stream frame interpolation

xenondoc/plan-atlas-upscale-and-frame-interpolation.md · 71 KB · updated 2026-10-06

Status: implemented and live-tested. Written 2026-08-10; several sections below were revised during implementation after live testing surfaced bugs the original plan didn't anticipate -- this document now describes actual current behavior, not the original (partially incorrect) plan. See each section for what changed and why.

2026-10-06 implementation update

The terrain and background interpolation described later in this document has been replaced where it inferred scrolling from quads. Scroll RAM is now captured with each draw frame; the applied camera delta drives terrain and half-speed background presentation. Background offsets stay fractional through the C/TypeScript quad interface. See HIGH_FPS_SCROLL_JITTER_20261006.MD for the verified causes, assembly references, implementation and before/after recording measurements. Historical statements about retaining integer background offsets and independent background speed registers no longer describe current code.

The gameplay starfield still has integer-anchor jitter, confirmed against r584 video on 2026-10-06. The original records already contain fractional phase; capture currently omits it and interpolation fits rounded positions instead. The verification and proposed phase-based repair are in the same report. The historical claim below that better output resolution alone would resolve this issue is superseded by that finding.

Context

The Xenon 2 custom renderer currently draws from a 1x-resolution sprite atlas (assets/spriteatlas_packed.{idx,raw,png}, 686x685) built offline by hatari_dotnet from a captured sprite catalog. A 4x-upscaled version of that same atlas already exists at assets/sprite_atlas_upscaled.png.png (2744x2740, confirmed via IHDR inspection to be an exact uniform 4x upscale of the same packed layout — not a repack). Nothing in the codebase currently loads it.

Separately, the game's own logic only advances its drawable state once every ~4 hardware VBLs. Live instrumented capture (xenondoc/XENON2.MD §8.1, 82 observed swaps) shows the real cadence is one buffer swap every 4 VBLs (~12.5Hz effective game-frame rate) — not every 2 VBLs (~25Hz) as originally assumed going into this planning session. The GPU render/submit path already runs on every VBL (~50Hz) today, it just resubmits the same static content 3 times out of every 4. The goal is to make those 3 "held" VBLs show smoothly extrapolated in-between positions instead of frozen content, without touching the emulated game logic (and therefore without risking desyncing its RNG-driven state).

The game will support both the existing low-res atlas and the new 4x atlas simultaneously at runtime, switchable with a key, rather than replacing one with the other — this also sets up a later crossfade transition between them without redesign.

This plan covers the desktop renderer only; the WASM/web renderer (web/src/atlas.ts, web/src/views/spriteStreamView.ts) mirrors the same formats/math and will need equivalent changes later, but that's explicitly out of scope for this phase.

Both tasks were scoped by reading the actual C/C#/asset code directly (no reliance on decompiled/reconstructed logic) — the interpolation task additionally drew on xenondoc/XENON2.MD and xenon2_position_update_mechanisms.html, which already document a partial audit of which object-update routines are pure functions of known state vs. which consume the game's shared RNG stream.


Task 1 — Atlas upscale (dual-resolution, runtime-switchable)

Why "just swap the file" isn't quite true

The runtime loader (src/spriteAtlas.c) computes every sprite's UV rect as a fraction: u0 = atlasX / atlasWidth, where atlasX/Y/W/H come from the .idx file and atlasWidth/Height come from the .raw file's own header. Because it's a fraction, uniformly scaling both files by the same factor reproduces identical UVs — but that means the .idx must be regenerated at the new scale, not left as-is (SpriteAtlas_LoadPixels in src/spriteAtlas.c:136-142 also hard-fails if the .raw's declared dimensions don't exactly match the .idx's).

There's one real exception: the scrolling-background quad path in StreamVertex_AppendQuad (src/sdlGpuRenderView.c:~1150) hardcoded const float sourceScaleY = 1.0f; and uses it to convert a logical (original-resolution) scroll offset into atlas V-fraction space. At 4x this would sample completely wrong V coordinates. The struct field comment on sourceOffsetY (src/includes/drawCommandStream.h:89-91) already anticipated this ("may be applied to a higher-resolution replacement asset by the renderer") — it just was never implemented.

Implementation note — the first fix attempt was itself wrong, twice. The obvious-looking fix, sourceScaleY = atlasH / logicalWrapHeight (using the background sprite's own resolved atlas-rect height), shipped first and produced a visibly vertically-squashed background. The background mosaic's atlas rect is taller than one screen — "the background image is larger than a screen, only a scrolled slice of it is ever displayed" — so its own rect height was never a valid stand-in for the atlas's overall scale (the original hardcoded-1.0f code even said so directly: "Do not infer scale from atlasH: the background asset's atlas rectangle can contain more data than its repeating playfield period" — that warning was correct and got missed on the first pass). The actual fix is SpriteAtlas_GetActiveScale(), a new spriteAtlas.c function comparing the active atlas's overall pixel dimensions against the low-res atlas's own (1.0 at 1x, 4.0 at 4x) — independent of any individual sprite's rect size. See the corrected code below.

Both resolutions ship together, loaded simultaneously

Per feedback: don't replace the 1x atlas with the 4x one — carry both, and let the player/dev switch at runtime. This means:

  • hatari_dotnet produces all files in one run, always. No separate "upscale mode" subcommand. The existing default flow (capture → rectangle-pack → export) additionally looks for sprite_atlas_upscaled.png.png alongside its other inputs and, if present, emits a second .idx/.raw pair at the derived scale in the same invocation. If that PNG isn't present (e.g. mid-iteration on a fresh capture before the external upscale step has run), the tool logs a warning and skips only the HD half — the normal 1x pipeline still completes.
  • New runtime asset filenames (both shipped, both loaded at startup): assets/spriteatlas_packed.idx / .raw (existing, 1x, unchanged format) and assets/spriteatlas_packed_hd.idx / .raw (new, 4x). This does require adding the two new filenames to src/CMakeLists.txt's HATARI_RUNTIME_ASSETS list (CMakeLists.txt:108-115) — unlike the original single-atlas plan, this step is no longer avoidable, since both variants need to exist on disk simultaneously.
  • src/spriteAtlas.c/.h becomes instantiable rather than holding one set of file-scope statics: introduce a small SpriteAtlas struct wrapping the existing entries/pixels/dimensions, with SpriteAtlas_LoadIndex/ _LoadPixels/_FindRect/_GetPixels taking an explicit instance pointer. main.c creates two instances at startup (g_spriteAtlasLow, g_spriteAtlasHigh) and loads both eagerly — both are small enough (~1.9MB and ~30MB resident) that keeping both in memory simultaneously is cheap on desktop, and it means switching is instant with no load-time hitch. A module-level "active atlas" pointer (or slot enum) is what sdlGpuRenderView.c's per-frame code reads.
  • GPU side: SdlGpuRenderView_InitializeSpriteStreamResources creates two GPU textures at init (one per atlas), uploads both once. The per-frame render pass binds whichever texture corresponds to the currently active slot — a cheap pointer swap, not a re-upload.
  • The sourceScaleY fix generalizes for free: SpriteAtlas_GetActiveScale() (see fix below) is derived from the active atlas's own pixel dimensions, so switching the active slot automatically produces the right scale on the next frame — no separate code path needed per resolution.
  • Toggle key: R, polled with SDL_GetKeyboardState the same way as this codebase's other keys — accepted regardless of which window has OS focus, matching existing convention (an earlier draft of this plan gated the poll on the debug window's own focus; dropped as unnecessary and inconsistent with how other keys already work here).
  • Future transition effect: noted, not built now. Because both textures are already GPU-resident simultaneously under this design, a later crossfade is just a shader blend between two bound textures over a few frames — this dual-load architecture is what makes that cheap to add later without restructuring.

Steps

  1. New hatari_dotnet/PngReader.cs — a hand-rolled, deliberately narrow PNG decoder (chunk parsing + System.IO.Compression.ZLibStream for inflate + manual scanline unfiltering, all five filter types), the read-side counterpart to the existing hand-rolled PngWriter.cs. No new NuGet dependency: an earlier draft of this plan used SixLabors.ImageSharp, but that package requires accepting/bundling the Six Labors Split License (not MIT, and it demands a license file) — since this decoder only ever needs to read one specific, well-understood format (non-interlaced, 8- or 16-bit truecolor+alpha, i.e. PNG color type 6), hand-rolling it is small and removes the dependency question entirely. Reuses PngWriter.cs's existing Crc32 helper to validate each chunk. Throws InvalidDataException with a specific message for anything outside that narrow format (palette/grayscale color types, other bit depths, Adam7 interlacing, an unrecognized critical chunk, a bad CRC, a truncated file, an unknown scanline filter byte) rather than silently mishandling it. The 16-bit → 8-bit downsample is therefore exactly controlled here (take the high-order byte of each big-endian sample — the PNG spec's own standard bit-depth-reduction technique), removing the ImageSharp-era open question of whether its downsample was byte-exact.

  2. New hatari_dotnet/AtlasIndexReader.cs — counterpart to the existing AtlasIndexWriter.cs; parses the current SATL .idx format back into {AtlasWidth, AtlasHeight, Entries[{StAddress,X,Y,W,H}]}. Refactor AtlasIndexWriter's inner write loop into a small shared helper so both the existing pack-time writer and the new rescale writer share one binary-format implementation.

  3. New hatari_dotnet/UpscaledAtlasExporter.cs — ExportUpscaled(packedIndexPath, upscaledPngPath, outIndexPath, outRawPath): - Read the just-produced 1x .idx (rects for the current run's packed layout). - Decode sprite_atlas_upscaled.png.png via PngReader.ReadFile. - Derive the scale factor from actual dimensions (image.Width / packed.AtlasWidth) rather than hardcoding 4 — assert it's exact and equal on both axes, log what was found. - Multiply every .idx rect's X,Y,W,H by that scale (pure multiplication — this is not a repack; RectpackSharp is not invoked here since the layout is already correct). - Convert the decoded RGBA to BGRA8 (mirror of the existing SpriteAtlasExporter RGBA byte-swap) and write via the existing, unchanged RawAtlasWriter.WriteBgra. - Write to temp files then move into place, so an interrupted run can't leave a half-written pair.

  4. hatari_dotnet/Program.cs — in the existing single flow, right after the current ExportPacked(...) call, additionally check whether sprite_atlas_upscaled.png.png exists next to the other inputs; if so call UpscaledAtlasExporter.ExportUpscaled(...) writing to spriteatlas_packed_hd.idx/.raw; if not, log and continue (1x output is unaffected either way). No new subcommand/mode.

  5. src/CMakeLists.txt — add spriteatlas_packed_hd.idx and spriteatlas_packed_hd.raw to HATARI_RUNTIME_ASSETS (CMakeLists.txt:108-115) alongside the existing three entries.

  6. src/spriteAtlas.c/.h — refactor file-scope statics into an instantiable struct + instance-taking functions (see above). Existing call sites (main.c, sdlGpuRenderView.c) updated to pass the active instance explicitly.

  7. src/main.c — load both atlases at startup (four SpriteAtlas_Load* calls instead of two); keep the existing fatal-if-missing behavior for the low-res pair, decide whether the HD pair missing should be fatal or a soft fallback-to-low-res (recommend soft fallback, so a dev build without the HD assets deployed still runs).

  8. src/sdlGpuRenderView.c — two GPU textures, active-slot bind, and the sourceScaleY fix in StreamVertex_AppendQuad (~line 1150): c const float sourceScaleY = SpriteAtlas_GetActiveScale(); (not atlasH / logicalWrapHeight — see the implementation note above for why that first attempt was wrong.) SpriteAtlas_GetActiveScale() compares the active atlas's overall pixel dimensions against the low-res atlas's own. At 1x this evaluates to 1.0 (unchanged, self-consistent with today); at 4x it evaluates to 4.0 automatically — correct for whichever atlas is currently active, with nothing to hand-maintain per resolution, and correct regardless of any individual sprite's own atlas-rect size.

  9. src/drawCommandStream.c, DrawCommandStream_PushMaskedSprite — not in the original plan; discovered live after Task 2's dual-atlas testing. This function set entry->destW = atlasW; entry->destH = atlasH; directly from whichever atlas was active at capture time. That silently assumed atlas-texel == logical-pixel (true at 1x, where it went unnoticed for a long time), but destW/destH must stay in logical 320x200 units always, exactly like x/y already do — with the HD slot active during capture, this baked in 4x-too-large on-screen sizes, reported live as "sprites rendered in the gameplay screen are 4 times too big" after pressing the resolution-toggle key. (Screens using the zoom- text/logo path were unaffected — those already compute destW/destH independently of atlasW/atlasH, which is exactly why only gameplay sprites showed the bug.) Fixed by re-resolving destW/destH specifically against the low-res atlas slot (which has one texel per logical pixel by definition) regardless of which atlas is active for UV sampling, falling back to the passed-in atlasW/atlasH only if that lookup fails (e.g. no low-res atlas loaded at all). DrawCommandStream_PushMaskedSpriteScaled (zoom-text, logo) was already unaffected -- it takes destW/destH as independent caller-supplied parameters, never derived from atlasW/atlasH at all.

Visual-bleed risk (verification only, not a code fix)

The packer used a 2px border between sprites at 1x (8px at 4x). If the external upscaler processed the whole packed image with a wide receptive field, neighboring sprites' content may have bled into that border band. Verify by cropping a few sprites' own rects from the new atlas and comparing against a naive nearest-neighbor 4x upscale of the same crop from the original spriteatlas_packed.png — a "melted toward its neighbor" edge is the fail signal. If it fails, the fix is regenerating that external asset with more border margin, not application code.

Resolved 2026-09-21 (the opposite failure). The border (now 4px) was empty transparent black, so the upscaler treated every sprite edge as an image edge and produced a dark, partly transparent rim one to two HD texels wide on every sprite -- measured on level-1 wall tiles, source edge rows of brightness ~110-180 came back at 0-30. On the wall grid that rim showed as seams the moment HIRES SPRITES was on. pad_packed_atlas_edges (xenon_tools/build_level_atlases.py, run on the chaiNNer input right before the upscale) now replicates each sprite's outermost pixels into its 4px border (clamp-to-edge, transparent edges stay transparent), so the upscaler sees continuous content past the rect; the HD rects are unchanged and the .NET pass after the upscale rewrites the clean 1x PNG. Verified on the same tiles: edge rows now match their interiors, no transparent corner pixels.

That removed the hard dark rim but not the grid: upscaled in isolation (even edge-padded) every 16x16 wall tile is still shaded as its own object, lighter in the middle and darker towards its edges, so walls and the tile-built bosses read as a grid of pillows. So wall tiles are upscaled a second time in context: build_tile_mosaic lays each level's tiles out as the game does (the 20-column tilemap plus the composite boss grids, using the 1x atlas's own pixels), the mosaic goes through the same chaiNNer run as one image, and apply_tile_mosaic cuts the 4x cells back into the level's upscaled atlas before the HD export. A tile index recurs across the map with different neighbours; choose_tile_placements picks the cell whose 8-neighbour signature is most common for that index (ties: most non-empty neighbours, then the earliest cell). Tiles that appear in no grid (animation frames) keep their per-tile upscale. Regenerated 2026-09-21: 236/254/182/351/157 tiles per level; the walls no longer show the grid -- where the tile sits in the neighbourhood it was upscaled in.

2026-09-22. That is only 24-41 % of the cells: the maps reuse each tile index in 5-9 different neighbourhoods, and everywhere else the mosaic tile carries the wrong neighbour's shapes cut off hard at the edge. Compared side by side with the per-tile upscale (clamp-, wrap- or mirror-padded) the mosaic's continuous texture was still preferred, so it stays the default (--no-tile-mosaic for the per-tile version). Measurements and the open options are in HD_WALL_TILE_SEAMS.md.

Verification plan

  • Numeric spot-check: hand-compute UVs from both the 1x and 4x pairs for a few known stAddress values, confirm they match to float precision.
  • Run with RENDER_VIEW_SPRITE_STREAM, toggle between resolutions live, confirm identical placement at both, higher fidelity only at 4x.
  • Watch the wall-tile mosaic scroll continuously through several full 192-scanline wrap periods at 4x — this is what exercises the sourceScaleY fix specifically; a stale/wrong scale shows up unmistakably as a frozen background or sampling into a neighboring sprite's atlas region.
  • Confirm the existing dimension-mismatch guard in SpriteAtlas_LoadPixels still rejects a deliberately-mismatched pair per instance.
  • Toggle to the HD slot mid-gameplay and confirm ordinary sprites (ship, enemies, bullets) stay the same on-screen size as at 1x — this is what exercises the destW/destH fix in step 9 above; a regression shows up unmistakably as everything suddenly rendering ~4x too large.

Task 2 — Frame interpolation

Three different identity models — this matters a lot for the design

Investigated directly (not assumed) how each quad type's objectId behaves across frames, since the interpolation design lives or dies on whether objectId is a stable match key:

Quad type objectId source Stable across motion?
Gameplay object (enemy/ship/bullet/etc.) Real MyObjectEntry ST address Yes, while the object is alive (may be reused after it dies — see below)
Background mosaic Fixed singleton 0xFFFE0000u (drawCommandStream.c:725) Yes, always — one quad per frame, motion lives in sourceOffsetY, not x/y
Wall tile Synthetic, derived from screen position: 0xFFFF0000u \| (y<<9) \| x (drawCommandStream.c:656-659) No — a tile's "id" changes every frame it moves, by construction

This means a single generic "match by objectId, extrapolate velocity" scheme is right for gameplay objects and (in a modified singleton form) for the background, but would be silently useless for wall tiles: a moving tile's synthetic id never matches its own previous-frame id, so every wall tile quad would register as a "miss" every frame and just hold at its last known position — safe (no wrong guesses, no gaps) but zero smoothing benefit for wall tiles specifically, and wall tiles are a large fraction of the visible playfield.

Wall tiles: tried group-level scroll extrapolation, reverted to always-held. The first implementation extrapolated every wall-tile quad by one shared vertical delta, driven by the live fine-scroll register (FINE_SCROLL_ADDR = 0x00000cd6u, drawCommandStream.c) — reasoning that the real game moves every wall tile in lockstep via that one shared register, so this should be the correct model, not just a workaround. Live testing showed this was wrong: most wall tiles don't actually move continuously at all. Only the top partial row-band has genuine sub-pixel motion; every other row is redrawn each real frame at a fixed 16px-grid Y position, with only its tile content (spriteId) changing as the level scrolls past — a discrete asset swap, not a position change. Applying one shared Y shift to all of them uniformly pushed the grid-fixed rows off their true position during held VBLs, then snapped them back on the next real frame — visible as forward-then-backward jitter, reported live and confirmed by re-reading the capture code (DrawCommandStream_OnTileDrawInstructionFetch's own row-band-Y-derivation comment). Current behavior: wall tiles are never extrapolated at all — always held at their last known real position, recognized by their synthetic objectId range ((objectId & 0xFFFF0000u) == 0xFFFF0000u, which cannot collide with real MyObjectEntry addresses or the background's distinct 0xFFFE0000u). A correctly row-band-aware model (extrapolating only the genuinely-moving top band) is a plausible future refinement but wasn't built — the visual cost of holding tiles for up to ~80ms is low, and getting it wrong was worse than not smoothing them at all.

This directly answers the specific questions originally raised: - Do tiles have unique ids? Yes, but positionally derived, not a stable per-instance identity — see table above. - What happens when a new tile scrolls into view with no velocity yet — is there a gap between it and existing tiles? No gap: it (like every other wall-tile quad) is simply held at its real captured position during synthetic frames, same as any object with no extrapolation. - Should all wall tiles move in sync with a single velocity? That was the original hypothesis and it's half right — the top row-band genuinely does, but the middle/bottom rows don't move continuously at all, they're redrawn at fixed grid rows with changing content. Building and shipping the "all tiles, one velocity" model live is what surfaced this distinction.

Group-shared-delta fix, built in a later session — corrected understanding of why the first attempt jittered. A first pass at revisiting this mis-scoped the fix to "only the top row-band genuinely moves" (reasoning from the row-band table above plus the original revert's symptom) and built a row-band-aware version that only extrapolated top-band quads. That framing was wrong: the whole 16x16 tile grid genuinely scrolls as one, at a constant speed, with every tile sharing the same offset. The game's three-row-band split exists purely because the original 68000 blitter has to manually avoid drawing rows that would land off-screen — a real cost worth avoiding on that CPU — not because different rows move differently. This renderer gets that clipping for free from the GPU, so the row-band split doesn't need to be reproduced at all for positioning purposes.

The fix: spriteStreamInterpolation.c's isWallTile branch derives one shared scroll delta per held tick from the top row-band's own Y alone (a new FindSharedWallTileTopBandY helper, reading newerQuad->y < 0 — the capture side already pushes the top band's unclipped logical Y as -fineScroll, a direct, unambiguous live-register read, unlike trying to infer a delta from the middle/bottom bands' own measured destination addresses) — extrapolated the same wide-baseline, wrap-aware, linear-only way as the background mosaic's sourceOffsetY (reasoning that the wall scroll register plausibly dithers its step count the same way the background's confirmed does, so a quadratic term would amplify that noise rather than track real curvature) — then applies the resulting delta (not an absolute recomputed position) uniformly to every wall-tile quad's own already-correct captured Y, top/middle/bottom band alike. No per-band distinction, no wire-format field, no per-quad identity matching (impossible anyway, since wall-tile objectId is position-derived) — just one shared scalar's own frame-to-frame delta applied everywhere.

This is structurally the same shape as the original reverted attempt (one shared delta for every wall-tile quad) — confirmed live to be correct this time via a temporary diagnostic trace (screentrace.c, logging the topmost and first mid-band wall-tile quad's position every VBL): the real (non-extrapolated, held=0) captured Y for both the top row-band and a representative mid-band row advanced by the identical +1px every real frame, disproving the original revert's own diagnosis ("most wall tiles are redrawn at a fixed 16px-grid position, only content changes" — see FINE_SCROLL_ADDR's corrected comment, drawCommandStream.c). The extrapolated held=1 ticks matched too: e.g. 4.000 → 4.250 → 4.500 → 4.750 across three held VBLs, then the next real frame landed at exactly 5.000 — zero prediction error. So the earlier revert's jitter came from how that delta was derived/applied (exact mechanism not re-examined — the buggy code was already fully removed by the time this was revisited), not from the tiles' actual motion being non-uniform. One accepted trade-off either way: a wall-tile coarse-scroll boundary crossing hands off to a physically different tile, so while the shared Y delta reads smoothly across that boundary, the tile content (never extrapolated, only ever held) still pops to the new tile's sprite at that one instant — bounded to once per full tile-scroll cycle, same "error visible for at most one real-frame interval" bound as every other branch in this module.

Extrapolation model for gameplay objects — linear/quadratic, with a plausibility clamp

Pure linear (2-frame, constant-velocity) extrapolation is exact for the several documented updateProcs that really are constant-velocity between real frames (compass-table enemies, radial-burst projectiles, rigid-follow attachments), but will systematically miss curved motion: it extrapolates along the tangent, undershooting a circle's true path and mistiming sine-wave direction reversals near the peaks/troughs. This is the model actually used for ordinary gameplay objects (ship, enemies, bullets, stars — anything with a real, stable MyObjectEntry objectId, i.e. everything that isn't the background singleton, a wall tile, the status bar, or a scaled/zoom quad — see below): cache 3 real frames and use constant-acceleration (quadratic) extrapolation when all 3 are available and pass the plausibility clamp, falling back to linear (2 known) or hold (1 known / a miss) otherwise:

v = (p[t] - p[t-1]) / Δ
a = (p[t] - 2*p[t-1] + p[t-2]) / Δ²
p(t+τ) = p[t] + v*τ + 0.5*a*τ²

This is a second-order Taylor expansion of the true motion — it tracks gentle curves and sine-like motion meaningfully better than linear, still never touches game logic or RNG, still costs nothing but one more cached frame and a few extra multiplies per object. It's still fundamentally an extrapolation, not a simulation: it will still visibly lag at sudden curvature changes (sharp direction reversal, a collision bounce, a spawn transition) since it has no way to know the next change in curvature before it happens — same self-correcting bound as before (any error is visible for at most one real-frame interval, then overwritten by truth). Final position is rounded to nearest (lroundf), not truncated — see the background section below for why that distinction mattered live.

Zero-velocity guard, per axis (added after live testing of the ship): the quadratic fit above assumes the last-measured acceleration stays valid going forward, which is actively wrong for step-function motion (joystick- driven movement — constant velocity while held, dropping straight to zero on release, no coasting). If an axis has already stopped by the two most recent real frames (p[t] == p[t-1]), any acceleration term derived from an older, still-moving sample is spurious: it reads the stop as deceleration and extrapolates past zero in the opposite direction for the held ticks, snapping back once the next real frame confirms the object never moved — visible live as the ship jittering for ~2 real-frame intervals after releasing the stick. Fix: an axis with p[t] == p[t-1] is held at p[t] unconditionally, skipping both the quadratic and linear formulas for that axis (independently of the other axis). Does not address the milder, unreported symmetric case of an axis just starting to move from rest — a real option if that turns out to need fixing too is reading live input state directly to detect a motion-stopping/starting input change mid-held-interval rather than waiting for it to show up in real-frame history, but that would be a ship-specific special case (this module is otherwise backend/object-agnostic), deferred unless the guard above proves insufficient.

Second baseline added to the same guard (live testing surfaced a different case): a stationary object whose effect animation cycles through sprites with different x_origin/y_origin (e.g. a rear-cannon muzzle flash) produces a genuinely different captured x/y every real frame even though the object never moves — a two-frame A/B/A/B cycle. Consecutive real frames never agree here, so the original guard didn't catch it; the quadratic fit read the middle sample as "acceleration continuing away from it" and visibly exaggerated a true displacement of zero into a swinging motion. Fix: the guard now also fires when the oldest and newest cached frame agree on an axis (net motion over the full 3-frame window is zero regardless of the middle sample), independent of the original newest-vs-older check.

Ring-buffer stale-flag crash (fixed): DrawMaskedSpriteEntry.isPerspectiveStar (added for the star-wrap fix above) was only ever set by DrawCommandStream_PushFlatColorQuad; PushMaskedSprite/PushMaskedSpriteScaled never touched it, so a ring-buffer slot last written by a perspective-star push could leak a stale true into the next entry that reused that slot — including, once, the background singleton, tripping spriteStreamFrame.c's "SPRITE_QUAD_BACKGROUND carries no flags" assert after extended play. Fixed by explicitly resetting it false in both of the other two push functions, matching the existing defensive pattern already used there for isFlatColor/isScaledSize.

Option C remains deferred, but is now the explicit answer for whatever still looks wrong after this: tight circles, high-frequency sine, or any pattern where quadratic extrapolation is visibly insufficient are precisely the cases worth reimplementing that specific updateProc faithfully in C (Option C) as a narrow, targeted follow-up — informed by which patterns actually look bad in practice, not built broadly upfront.

The background mosaic — a different extrapolation model than gameplay objects

The background singleton (objectId = 0xFFFE0000u) doesn't move via x/y at all — its motion lives entirely in source.background.sourceOffsetY, a value that wraps modulo 192 (the mosaic's repeat period). Extrapolating it needed two live-tested corrections beyond what the original plan described:

  1. Wraparound must be unwrapped before doing velocity math. A legitimate wrap (e.g. 190 -> 5) read as a huge spurious jump if fed directly into a linear/quadratic fit, producing a wildly wrong extrapolated position that then snapped back once the next real frame arrived — exactly the "forward then backward" jitter reported live. Fixed with an UnwrapRelative(value, reference, period) helper that re-expresses a sample as the representative of its residue class closest to a reference sample, applied before any velocity/acceleration math.

  2. Quadratic extrapolation was actively wrong for the background, confirmed via ReVa disassembly of BackgroundScrollCursor_UpdatePerFrame_ FUN_0000702c (the cursor's own per-frame update routine). The step count advanced per real frame is (in_D0w & scrollSpeedMask & 1) + (in_D0w >> 1) — a dithered/fractional-speed technique: to make the average scroll rate non-integer (e.g. 3.5 texels/frame), the game alternates between 3 and 4 steps frame-to-frame. So even during genuinely "constant speed" scrolling, the per-frame delta itself isn't constant — it oscillates by ±1 step. Both the 1-interval linear delta and the quadratic acceleration term (a second finite difference) amplify this dither into visible jitter, no matter how correctly wraparound is handled. Fixed by widening the velocity baseline to span 2 real intervals (oldest -> newer, divided by 2) when 3 real frames are cached, which cancels a single-frame ±1 dither step the same way averaging over any longer window cancels noise; falls back to the narrower 1-interval linear estimate only when 2 real frames are cached. Quadratic is never used for the background — the plan originally assumed the same linear/quadratic model as gameplay objects applied uniformly; live testing showed the background specifically needed both a wider baseline and to drop quadratic entirely, for reasons that don't apply to ordinary objects (whose positions are plain integers, not run through an extra scroll-speed dithering step).

  3. Rounding, not truncating, the final wrapped value mattered visibly. A plain (uint32_t)wrapped cast floors toward zero; e.g. a continuous extrapolated value of 94.88 (much closer to a real 95 on both sides than to 94) was being displayed as 94 for the whole held interval, then jumping to 95 on the next real frame — a small but real 1-texel flicker, caught by comparing a live diagnostic log's raw numbers against the displayed result. Fixed with lroundf, matching every other extrapolation path in the module (gameplay objects already rounded correctly; the background path was the one place still truncating).

Zoom-text glyphs and the title-screen logo — also held, not extrapolated

Not in the original plan at all — discovered live after the wall-tile and background fixes, reported as "the zoom-in (and out) text and logo is jittery — after a logo zoom-in finishes, it moves a little after it settles in place." Both use DrawCommandStream_PushMaskedSpriteScaled, where position (x/y) and on-screen size (destW/destH) are both driven by the same live "scale" register, but via different formulas — position is linear in scale (confirmed via the logo's own plate comment: x = ((D0*scale)>>4)+160), size comes from a rotate/popcount lookup table (not necessarily linear). This module only ever extrapolates position, never destW/destH (held at the last real value) — so extrapolating position alone drifts out of sync with the frozen size every held VBL, worst right when the zoom animation stops changing: the predictor doesn't know that yet, so it keeps extrapolating for a couple of held VBLs before the next real frame confirms the truly-settled position and snaps back.

The logo has a stable objectId (spriteId, constant throughout its whole zoom animation) so it was reliably matched and extrapolated every held VBL. Zoom-text glyphs' objectId (destScreenAddr + characterIndex) turned out to also be stable enough to match — the scaled text path computes its real position independently, without touching A0/destScreenAddr, so that value stays roughly constant across the animation rather than moving with it the way a wall tile's synthetic position-derived id does.

Fixed with a new DrawMaskedSpriteEntry.isScaledSize field (threaded through to SpriteRenderQuad as SPRITE_QUAD_FLAG_SCALED_SIZE), set by DrawCommandStream_PushMaskedSpriteScaled and checked (false) by both DrawCommandStream_PushMaskedSprite and DrawCommandStream_PushFlatColorQuad. Flagged quads are always held, never extrapolated — same "hold rather than guess wrong" treatment as wall tiles and the status bar, on the reasoning that these are discrete UI animations stepping once per real frame, not continuous physics-driven motion, so extrapolation doesn't meaningfully help them anyway (modeling the shared scale-to-size/position relationship instead would still have the same end-of-animation overshoot problem, just for a different quantity).

One subtlety this surfaced: the background mosaic also goes through PushMaskedSpriteScaled (it needs an explicit 320x192 size independent of its own, much taller, atlas rect — see the destW/destH fix in Task 1) so isScaledSize is true for it too, but for an unrelated reason (a fixed constant clamp, not a live scale register). A pre-existing assertion in SpriteStreamFrame_AppendEntry requires background quads to carry no flags at all; naively propagating isScaledSize onto the background's own SpriteRenderQuad violated it. Fixed by excluding isBackgroundWrap entries specifically from the flag propagation in spriteStreamFrame.c — behaviorally identical either way, since the interpolation module already special-cases SPRITE_QUAD_BACKGROUND before it would ever check this flag.

objectId reuse — explicit identity tracking (objectIdentity.c/.h)

ST addresses backing MyObjectEntry are reused after an object dies and a new one is allocated at the same address, and the 48-slot starfield record array is reused by two semantically incompatible rendering modes (see "Three different identity models" above). Rather than continue inferring reuse heuristically, capture hooks now detect the two authoritative lifecycle events directly (found via ReVa disassembly, not decompiler guesses — both turned out to be single choke points despite dozens of call sites):

  • Allocate/recycle: AllocateOrRecycleObjectEntry_FUN_00002be4's single shared rts at 0x00002c00 — every one of its ~26 callers converges here, with A0 holding the (re)allocated MyObjectEntry*.
  • Destroy/free: DestroyObject_UnlinkAndFree_FUN_0000107c's entry PC 0x0000107c — installed as a shared "destroy" function-pointer slot at ~30 object-setup sites, but it's one physical function, so hooking its own entry catches every invocation regardless of caller. A0 holds the entry about to be unlinked and freed.
  • Starfield mode switch: already distinguishable at capture time (which of the two starfield hooks fired this frame) — no new hook needed, just a tracked "last mode" flip that resets all 48 record slots.

src/objectIdentity.c/.h is a small host-side registry (never touches ST memory — there's nowhere sanctioned in the guest's own state to put this) of {slotAddr, identityId}, backed by one global monotonic counter. SpriteRenderQuad gained a new identityId field (objectId stays, purely for diagnostics/ReVa correlation). ObjectIdentity_Get(slotAddr) degenerates to slotAddr itself for any slot that's never been through ObjectIdentity_Reset — so untouched quads (background, wall tiles, status bar, an object that simply never dies while observed) need no table entry and no special-casing; identityId is a strict superset of comparing raw objectId, only actually diverging where a reset was recorded. spriteStreamInterpolation.c's FindQuadById/FindMatchingQuad now compare identityId, so a reused-address mismatch is structurally impossible to match rather than merely unlikely — a recycled/destroyed/mode-switched slot simply has no same-identity candidate in the older cached frame, correctly falling through to "no match, hold" instead of needing the plausibility clamp to catch it after the fact.

The plausibility clamp (DisplacementIsPlausible) stays, now as a second layer for the one lifecycle event not yet hooked: the perspective starfield's own internal RNG-driven respawn (a star keeps the same identityId — same slot, same mode — but genuinely jumps to a fresh random position mid-flight; see Starfield_DrawPerspectiveFlythrough_FUN_000085b4's doc comment). Its only hookable point is an inline branch inside a per-star loop, not a clean call/return site like the two above, so it's deliberately out of scope for now. (The clamp's tuning — K=8 size-relative with a 64px absolute cap — was already validated live against the reused-address cases this change now handles structurally; unchanged here.)

Explicitly NOT treated as an identity change: the vertical-scroll starfield's own Y-position wrap (modulo the 192-line playfield, confirmed via disassembly of Starfield_UpdateVerticalPositions) is a continuous, same-instance quantity — a bounded modular correction, not a reset — so it stays a math problem for UnwrapRelative, not an ObjectIdentity_Reset call.

Star Y-wrap, now applied (live testing surfaced a 192px "misprediction" at a 1x1 flat-color quad, confirming this was still needed): the blocker described in the previous revision of this doc — distinguishing vertical-scroll stars from perspective-mode stars on the quad itself, since applying the wrap unconditionally would misread a perspective respawn's genuine discontinuity as a safe wrap — is resolved by a new DrawMaskedSpriteEntry.isVerticalScrollStar bool (drawCommandStream.h), surfaced as SPRITE_QUAD_FLAG_VERTICAL_STAR on the render side. In spriteStreamInterpolation.c's generic gameplay-object path, a quad is treated as a wrap-aware vertical-scroll star purely when that flag is set: its y history is unwrapped via UnwrapRelative(..., BACKGROUND_PLAYFIELD_HEIGHT) — relative to the newest sample, same convention the background branch already uses — before both the plausibility gate and the linear/quadratic formulas, and the final extrapolated y is folded back into [0, 192) the same round-to-nearest-then-adjust way the background's sourceOffsetY already is. Unwrapping before the plausibility gate specifically (not just before the extrapolation math) mattered live: a 1x1 star's own K * width plausibility clamp is tiny (K=8, width=1 → 8px), so a raw near-full-period wrap delta (up to ~191px) would otherwise get rejected as implausible and just held, never reaching the wrap-aware math at all. x is untouched — gameplay-mode stars have no horizontal motion (confirmed via disassembly of Starfield_UpdateAndDraw_48Stars: the X bitmask word is only ever read). DisplacementIsPlausible's signature changed from two quad pointers to explicit (dx, dy, width, height) to make this possible — callers now compute dy themselves (unwrapped where relevant) rather than the function reading raw ->y fields directly.

Positive marker, not an inferred double-negative (corrected after feedback): the flag started life as isPerspectiveStar, with the wrap condition being "FLAT_COLOR and that flag absent" — but FLAT_COLOR is also used for the status-bar weapon meter today (harmless in practice, that case is held earlier via SPRITE_QUAD_FLAG_STATUS_BAR and never reaches this code) and isn't guaranteed to stay starfield-only in general — a hit-highlight flash effect would plausibly also be flat-colored, and inferring "must be a vertical star" from the absence of one unrelated flag would silently wrap-adjust any such quad's position too. Replaced with a dedicated positive isVerticalScrollStar, set true only by DrawCommandStream_OnStarfieldDrawInstructionFetch (the gameplay starfield hook itself, not inferred from what it isn't) and false everywhere else, including explicitly reset in the other two push functions per the ring-buffer-staleness lesson below. isPerspectiveStar was kept alongside it at first, purely descriptive.

Superseded — consolidated into DrawEntryKind/DrawEntryStarKind enums: once there were two mutually-exclusive pairs of bools on DrawMaskedSpriteEntry (isFlatColor/isBackgroundWrap choosing a 3-way draw kind; isPerspectiveStar/isVerticalScrollStar choosing a star sub-kind), both were replaced with dedicated enums — DrawEntryKind kind (DRAW_ENTRY_KIND_TEXTURED/_BACKGROUND/_FLAT_COLOR) and DrawEntryStarKind starKind (DRAW_ENTRY_STAR_KIND_NONE/ _VERTICAL_SCROLL/_PERSPECTIVE, meaningful only when kind == DRAW_ENTRY_KIND_FLAT_COLOR) — making the mutual exclusivity a property of the type instead of a convention every call site has to uphold by hand. isFlash/isStatusBar/isScaledSize stayed bools: they're genuinely independent and combine freely with any kind or with each other (e.g. a textured quad can be both isStatusBar and isScaledSize), so folding them into the same enum would either explode combinatorially or silently forbid a real combination. All push-function call sites, spriteStreamFrame.c's kind dispatch, and every place that gated on the old flag pairs were updated to match; the ring-buffer-reset discipline established by the crash above applies identically to the new enum fields (every push function sets both unconditionally).

Star Y dithered-step overshoot (mitigated, not eliminated — accepted, deferred): live testing after the wrap fix above still showed vertical-scroll stars "progressing down ok, then jumping back up a little". Confirmed via disassembly of Starfield_UpdateVerticalPositions (no new hook needed — the existing plate comment already had the answer): a star's Y is driven by a phase accumulator that advances a fixed sub-pixel amount every frame, with the visible Y only stepping by one scanline when that phase crosses a threshold — so the raw frame-to-frame sequence dithers (0,0,1,0,1,0,0,1,...), only averaging to the star's true speed over several frames. This is structurally identical to the background scroll's own already-fixed dithered-speed bug (see that branch below) — the standard quadratic/narrow-linear formula picks up the noisy single-interval velocity, overshoots on a "step" frame, and snaps back on the next real frame. Mitigated the same way as the background: for isVerticalStar quads specifically, when 3 real frames are cached, use the wide oldest -> newer baseline average velocity with pure linear extrapolation (no acceleration term — the apparent curvature here is the dither noise) instead of quadratic; falls back to the ordinary narrow 1-interval linear estimate when only 2 real frames are cached. Every other quad type is unaffected — this is gated on isVerticalStar exactly like the wrap-unwrap treatment above it.

Residual confirmed still present after the mitigation above — root cause is a shape mismatch, not a rate-estimation error, so it can't be fully fixed by better velocity averaging alone. The star's true position is a step function: it holds at a fixed integer scanline for a variable number of real frames, then instantly jumps exactly one scanline — it never occupies a fractional position. Any smooth extrapolation curve (linear or quadratic, narrow or wide baseline) necessarily predicts fractional in-between positions during the held ticks, including during real intervals that turn out to be "hold" frames where the star doesn't actually move — and that predicted-but-never-real progress is exactly what gets visibly corrected when the next real frame confirms the star stayed put. The wide-baseline fix reduces the size of this mismatch (a better long-run rate estimate means less average error) but cannot eliminate it, since no continuous curve can exactly reproduce a staircase. Also a plausible contributing factor, not yet isolated from the above: the zero-velocity guard's narrow (newest == older) check can itself match a single dither "hold" step and force a full flat prediction for that interval, bypassing the wide-baseline average entirely for that window — worth revisiting together with the shape-mismatch issue if this is picked up again later.

Deliberately left as a known, accepted issue for now (not pursued further this pass): the practical fix — matching the smoothness of extrapolation to the underlying sub-pixel motion — is expected to become naturally easier once gameplay resolution is upscaled beyond the native 320×200 (a separate, already-planned future work item), since a higher-resolution star position has more headroom to represent genuine sub-pixel progress smoothly rather than needing to snap between whole-native-pixel steps. Revisit this specifically once that upscale work begins, rather than continuing to chase it at native resolution.

Origin recovery for textured objects (live testing: a stationary rear cannon's muzzle-flash animation visibly "moved left and right" while firing): confirmed via disassembly of blitting_top_level_function_FUN_010ae that the real game computes on-screen position as anchor - spriteData-> x_origin/y_origin per draw (sub.w (A0)+,D0w right at entry) — exactly what the capture side already replicates into DrawMaskedSpriteEntry.x/y. The problem: different sprites in the same multi-frame effect animation (a muzzle flash, a growing explosion) can carry different origins to stay visually anchored despite changing size, so the origin-adjusted x/y this module was extrapolating can legitimately shift a few px every real frame even while the object's true anchor never moves — and the animation cycle can be longer than the zero-velocity guard's 3-frame window can reliably pattern-match (confirmed live: the rear cannon's cycle is longer than 2 frames).

Fixed at the source rather than pattern-matched: DrawMaskedSpriteEntry gained originX/originY (the origin values already being subtracted, now kept alongside the result instead of discarded — populated only by DrawCommandStream_PushMaskedSprite's real per-object callers; wall tiles/HUD glyphs pass 0/0, no origin concept applies), surfaced as SpriteRenderQuad.source.textured.originX/Y. For SPRITE_QUAD_TEXTURED quads, spriteStreamInterpolation.c now fits velocity/acceleration on the recovered anchor (x + originX) instead of the raw screen position, for every cached sample — then converts back to screen space at the very end by re-applying the current (held, unextrapolated — spriteId/size are never extrapolated either) sprite's own origin once. The zero-velocity guard now operates on this same "fitting-space" value too, so a truly stationary animated object's recovered anchor matches exactly across real frames and guard #1 (consecutive-frame check) already catches it directly — guard #2 (oldest-vs-newest) is kept as a backstop for non-textured cases or cycles that happen to realign within the 3-frame window, but is no longer the primary defense for this class of bug.

This generalizes cleanly to size-changing effects (explosions) the same way: since size/spriteId are held rather than extrapolated regardless, an object that's genuinely stationary while its effect grows now correctly extrapolates zero position change, instead of misreading the growing sprite's shifting origin as motion.

RNG hazard (why we still never re-run game logic)

xenondoc/XENON2.MD §2.8 and xenon2_position_update_mechanisms.html §09 document that several updateProcs consume a shared gameplay RNG stream per real frame. Replaying any such routine to synthesize an in-between frame would advance that shared RNG stream early, silently diverging gameplay from the original. Everything in this plan only ever extrapolates already-observed positions from the draw-command stream — it never re-executes game/CPU logic, so this hazard doesn't apply to any of it, including the quadratic upgrade above (still just arithmetic over cached history, not simulation).

Architecture

New module: src/spriteStreamInterpolation.c + .h, added to src/CMakeLists.txt next to spriteStreamFrame.c. Keeps screentrace.c's RENDER_VIEW_SPRITE_STREAM branch a thin orchestrator; SpriteStreamFrame_Build's existing call, arguments, and output buffer stay completely untouched — the new code is purely additive after it.

State (file-scope static, not exposed via the header): - Three cached SpriteRenderQuad[] buffers + counts, oldest to newest (s_history[0..2]) — used for gameplay-object linear/quadratic extrapolation and the background's own wide-baseline extrapolation (see above). No longer holds a separate wall-tile scroll-register cache — that approach was built, found to cause jitter, and removed (see "Wall tiles" above); wall tiles are simply never extrapolated now. - A monotonic VBL tick counter, incremented once per call (already called exactly once per VBL). - The tick-gap measured at the last real-frame transition (s_measuredIntervalTicks), seeded near the documented ~4-VBL cadence but overwritten by the first observed transition — self-corrects rather than hardcoding "4" anywhere. - An enable/disable flag for the A/B toggle — defaults to false (not the originally-planned "on by default"; changed after live testing showed several of the bugs described above with it on, and it's simplest for a session to start from the known-correct non-interpolated baseline and opt in). - Prediction-error metrics state (not in the original plan): a snapshot of the last extrapolated output, plus running sum/max-error accumulators. See "Prediction-error metrics" below.

Key functions:

void SpriteStreamInterpolation_Update(const SpriteRenderQuad* realQuads,
    uint32_t realQuadCount, uint32_t frameToRender);
  /* Called every VBL right after SpriteStreamFrame_Build. Detects whether
   * frameToRender is new since last call: if so, shifts the 3-frame cache,
   * records the fresh quads + tick gap, and scores the previous held
   * interval's last prediction against this real frame (see metrics
   * below); if not (held VBL), only advances the tick counter. */

bool SpriteStreamInterpolation_IsHeldVbl(void);

uint32_t SpriteStreamInterpolation_Extrapolate(SpriteRenderQuad* outQuads,
    uint32_t outCapacity);
  /* Routes each cached quad through, in order:
   *   - status bar (SPRITE_QUAD_FLAG_STATUS_BAR): always held -- sparse,
   *     content-driven redraws break the "one real interval apart"
   *     assumption every other branch relies on.
   *   - wall tile (0xFFFF0000u objectId range): always held -- see "Wall
   *     tiles" above for why group extrapolation was reverted.
   *   - background singleton (0xFFFE0000u): dedicated wide-baseline linear
   *     extrapolation on sourceOffsetY, with wraparound unwrapping and
   *     round-to-nearest -- see "The background mosaic" above. Never
   *     quadratic.
   *   - scaled-size (SPRITE_QUAD_FLAG_SCALED_SIZE -- zoom-text glyphs, the
   *     logo): always held -- see "Zoom-text glyphs and the title-screen
   *     logo" above.
   *   - everything else (ordinary gameplay objects, stars): matched across
   *     the 3-frame cache by identityId + type + a displacement-plausibility
   *     clamp (deliberately not spriteId -- see "objectId reuse" below);
   *     quadratic extrapolation on a full 3-frame match, linear on a
   *     2-frame-only match, hold on a miss.
   * spriteId/width/height/flags are always held at the newest cached
   * value; only position (or the background's sourceOffsetY) is ever
   * extrapolated. Returns 0 if fewer than two real frames are cached yet
   * (e.g. right after startup) -- caller falls back to plain real quads. */

Target rate is one tunable, not a design fork: since GPU submission already happens every VBL for free, extrapolating on every held VBL (the current behavior) costs nothing extra over today's baseline — every held tick gets its own distinct extrapolated position rather than repeating the same frozen content. (The original plan described this as a step-count constant to make tunable; in practice the code just always steps every held VBL, and the meaningful on/off choice turned out to be the enable flag above, not a rate knob.)

Prediction-error metrics (not in the original plan)

Added after the fact, at request, to gauge extrapolation quality independent of any specific bug: Extrapolate() snapshots its own output every time it's called; when the next real frame lands, Update() matches that snapshot's quads against the real positions that actually arrived (by identityId+type, same matching rule as extrapolation itself) and accumulates |dx|/|dy|/magnitude error. Status bar, wall tile, and scaled-size quads are excluded from scoring (they're never actually extrapolated, so including them would just report a trivial zero error and dilute the mean). Every ~50 real frames (~4s at the ~12.5Hz real cadence), logs a summary and resets:

SpriteStreamInterpolation: prediction error over N samples -- mean |dx|=.. |dy|=.. px, max |d|=.. px

Purely diagnostic — never affects rendering. A handful of bounded, throttled diagnostic logs used to root-cause the bugs described above (raw per-tick background numbers, logo cache push/invalidate events) were removed once those bugs were confirmed fixed; this metric was kept as the one signal worth monitoring on an ongoing basis.

Per-event misprediction log (added alongside the zero-velocity guard, to verify it live): the periodic summary above only ever reports two aggregate numbers, which can't say whether a given max came from one repeat-offending object or many different ones, or let you correlate it back to a specific real frame. Every individual scored misprediction over SPRITE_STREAM_INTERP_MISPREDICT_LOG_THRESHOLD_PX (2px — deliberately low, to catch modest single-object artifacts like a few-pixel overshoot, not just the ~80px scale of a systemic reused-identity bug) gets its own line:

SpriteStreamInterpolation: misprediction |d|=.. px -- frame=N heldTicks=.. objectId=0x.. identityId=0x.. type=.. predictedSpriteId=0x.. realSpriteId=0x.. predicted=(x,y) real=(x,y) size=WxH

frame is the real frame that confirmed/disproved the guess; heldTicks is how many VBLs it was extrapolated across. Bounded only by SPRITE_STREAM_INTERP_MISPREDICT_LOG_LIMIT (5000), a safety valve against runaway spam in a badly-broken state — not a normal-use sampling cutoff the way the earlier 20-event outlier limit this replaced was.

Metric itself needed the same wrap-awareness as the extrapolation (fixed): live testing after the star Y-wrap extrapolation fix still showed ~191px "mispredictions" for vertical-scroll stars, alternating predicted=0/real=191 then predicted=191/real=0 on consecutive real frames — the extrapolation was actually correct (off by ~1px), but RecordPredictionError's own errorY = predicted->y - real->y had no wrap-awareness, so a genuinely tiny error that happened to straddle the wrap boundary got reported as a false ~191px outlier. Fixed by giving RecordPredictionError the same SPRITE_QUAD_FLAT_COLOR + !SPRITE_QUAD_FLAG_PERSPECTIVE_STAR branch the extrapolation path already has, unwrapping predicted->y relative to real->y before diffing — mirrors the background branch just above it, which already did this correctly for sourceOffsetY.

Call site — src/screentrace.c, RENDER_VIEW_SPRITE_STREAM branch (~450-520)

Purely additive, right after the existing SpriteStreamFrame_Build call and before RenderView_Submit, and gated on emulationActive (not in the original plan — added after live testing showed sprites continuing to drift during a pause, F12): while paused, Update()/Extrapolate() are skipped entirely, freezing tick/history state exactly where it was rather than treating the whole paused interval as one long held VBL to extrapolate across.

if (emulationActive)
{
  SpriteStreamInterpolation_Update(frame.spriteStream.quads,
      frame.spriteStream.quadCount, frameToRender);

  if (SpriteStreamInterpolation_IsEnabled() && SpriteStreamInterpolation_IsHeldVbl())
  {
    static SpriteRenderQuad interpolatedQuads[SPRITE_STREAM_MAX_QUADS];
    uint32_t n = SpriteStreamInterpolation_Extrapolate(interpolatedQuads, SPRITE_STREAM_MAX_QUADS);
    if (n > 0)
    {
      frame.spriteStream.quads = interpolatedQuads;
      frame.spriteStream.quadCount = n;
    }
  }
}

Toggle key: I, polled the same unscoped way as the atlas-resolution switch's R key — see that section's note above.

Status bar gameplay-detection was broadened too, for the same reason as the logo (not in the original plan): the HUD's on/off gate (s_gameplayActiveThisFrame) was set only by the wall-tile-draw hook, which shares the exact blind spot that caused the logo bug — a level opening in open space before any wall tiles appear. Unlike the logo, this flag is reset every frame (not a persistent cache), so only the "turning on late" direction was actually at risk. Fixed by setting it from the same three gameplay-exclusive hooks now used for the logo (wall-tile draw, object dispatch, gameplay starfield).

The logo's own persistence bug needed a fourth signal beyond gameplay detection: the title-screen logo persisted visibly through two intermediate screens (player-count select, "GET READY PLAYER n") before disappearing, because neither wall-tile draw, object dispatch, nor the gameplay starfield ever fire on those text/starfield-only interstitial screens. Found via ReVa: ZoomTextInterstitial_MainLoop_FUN_00008530 (the actual driver behind those screens) calls FUN_00007f96 — an unconditional full 32000-byte framebuffer clear — every frame, which is what physically erases the logo on real hardware once that screen is reached. DrawCommandStream_OnZoomTextGlyphInstructionFetch already gates on the caller's return address to distinguish this screen (0x859a) from the title screen's own credits text (0x84a8, where the logo should keep persisting); invalidation was added specifically for the 0x859a case.

Verification plan

  • A/B toggle: flip interpolation on/off live during gameplay (I key) and watch the same objects for smoother-vs-steppier motion, no relaunch needed.
  • No-desync check (most important): play the same save/seed with interpolation ON vs OFF and confirm score, enemy spawn timing, and run outcome are identical. Since this never re-executes game logic, any observed divergence would indicate something unexpectedly leaked game-observable state into the new code.
  • Curved-motion check: specifically watch any enemy/projectile pattern known to curve (circle/sine-style movement) with interpolation on, to judge whether quadratic is sufficient or a case for deferred Option C.
  • Background scroll check: confirm smooth, jitter-free scrolling in both directions (the player can reverse it by flying backward) and while stopped (ship stuck/killed) — this is what exercises the wide-baseline averaging and wraparound-unwrap fixes.
  • Wall-tile check: confirm tiles look identical with interpolation on vs off (held either way) — no seam, no snap-back; wall tiles are a "nothing should visibly change" case now, not a smoothing win.
  • Zoom/logo check: watch the logo zoom in/out and zoom-text screens ("GET READY PLAYER n" etc.) with interpolation on — confirm no settle-then-jump motion at the end of an animation, and that the logo correctly disappears once the player-select/get-ready screens are reached, not lingering into gameplay.
  • Reused-objectId guard check: force heavy spawn/death churn (e.g. a dense enemy wave) and confirm the plausibility clamp suppresses any teleport-streaks between unrelated objects.
  • Pause check: pause (F12) mid-motion and confirm nothing continues to drift; resume and confirm motion continues smoothly rather than jumping.
  • Prediction-error metric: watch the periodic log during normal play; a low, stable mean with occasional higher max values (self-correcting, bounded to one real interval) is expected — a persistently high or climbing mean would indicate a new problem.
  • Dual-atlas switch: confirm switching resolution mid-gameplay is instant (no load hitch, since both are preloaded) and doesn't perturb interpolation state.

Explicit non-goals for this phase

  • No re-execution/simulation of game update logic (Option C) — deferred until quadratic extrapolation ships and specific object classes are shown to still need it.
  • No delayed/buffered true interpolation between two known real frames (rejected earlier: ~80ms added input latency in a fast-paced shooter, for a correctness gain the self-correcting bounded error mostly already delivers).
  • No object spawn/despawn hook for authoritative objectId-reuse detection yet — heuristic clamp only, per above.
  • No modeling of the zoom-text/logo shared scale-to-size/position relationship (would still have the same end-of-animation overshoot problem, just for a different quantity) — held instead, see above.
  • ~~No row-band-aware wall-tile extrapolation~~ — superseded, later session: the underlying premise ("only the top band genuinely moves") turned out to be wrong — the whole tile grid scrolls as one at constant speed; the row-band split is a CPU-clipping artifact of the original 68000 code, not evidence of different per-band motion. All wall tiles (not just the top band) are now extrapolated by one shared, wrap-aware delta. See "Wall tiles" below for the corrected fix.
  • No changes to the WASM/web renderer at first — a later session found the web-side SpriteRenderQuad wire parser (web/src/spriteQuad.ts) had already drifted out of sync with the C struct (missing identityId, silently dropping every sprite-stream frame) independent of anything in this plan, fixed alongside widening x/y to float (see "Coordinate precision" below). The dual-atlas/sourceScaleY UV-math gap described above is still unaddressed on the web side.
  • No attempt to fix atlas border-bleed in code — if present, it's an external-asset regeneration problem.
  • No crossfade/transition effect between atlas resolutions yet — the dual-load architecture is chosen to make this cheap later, but it isn't built now.

Coordinate precision (later session)

Extrapolated positions were computed in float from the start (see above), but SpriteRenderQuad.x/y (renderFrame.h) was int32_t, and the gameplay-object/vertical-star branches rounded to the nearest native pixel (lroundf) before storing — throwing away the fractional part before it ever reached a renderer, capping visible motion smoothness at one of 320 horizontal positions regardless of atlas or backbuffer resolution. Both renderers' NDC math (StreamVertex_AppendQuad in sdlGpuRenderView.c, and its web port in spriteStreamView.ts) already divided by the logical 320×200 output size as plain floats, so this was never the bottleneck. Fixed by widening SpriteRenderQuad.x/y to float and removing the lroundf-to-int32_t casts (the background's sourceOffsetY stays integer — it indexes a discrete texel row in the mosaic asset, a different kind of quantity). DrawMaskedSpriteEntry.x/y (drawCommandStream.h, capture-side) stays int16_t — genuinely correct, since the original 68000 game logic never computes a sub-pixel position; only the extrapolation math synthesizes fractional values, one layer downstream.

This surfaced a pre-existing, unrelated bug on the web side while confirming the fix would carry through there too: web/src/spriteQuad.ts hand-mirrors SpriteRenderQuad's wire layout by byte offset, and had already drifted out of sync with the C struct — missing the identityId field entirely (added in an earlier session for objectIdentity.c, see above), reading 40 bytes where the struct is actually 44. Since webRenderBackend.c passes the real sizeof(SpriteRenderQuad) as the wire stride, this mismatch should have been tripping spriteQuad.ts's own struct-layout-drift guard and silently dropping every sprite-stream frame in the web build already, independent of anything in this coordinate change. Fixed in the same pass: corrected every field offset, added the missing objectId/identityId fields and SpriteQuadFlags bits, and switched x/y to getFloat32.


Critical files

  • hatari_dotnet/Program.cs, hatari_dotnet/AtlasIndexWriter.cs, hatari_dotnet/SpriteAtlasExporter.cs, hatari_dotnet/PngReader.cs — extended single-pass export
  • src/spriteAtlas.c / src/includes/spriteAtlas.h — instantiable refactor for dual-resolution support, SpriteAtlas_GetActiveScale()
  • src/sdlGpuRenderView.c — dual textures, active-slot bind, sourceScaleY fix
  • src/main.c — dual atlas load at startup
  • src/screentrace.c — interpolation call site (pause-gated), toggle keys
  • src/spriteStreamInterpolation.c / .h — the interpolation module itself; also owns the prediction-error metrics
  • src/spriteStreamFrame.c — maps DrawMaskedSpriteEntry.isScaledSize to SpriteRenderQuad's SPRITE_QUAD_FLAG_SCALED_SIZE (excluding the background, see "Zoom-text glyphs" above)
  • src/includes/renderFrame.h — SPRITE_QUAD_FLAG_SCALED_SIZE (new)
  • src/includes/drawCommandStream.h — DrawMaskedSpriteEntry.isScaledSize (new); FINE_SCROLL_ADDR moved back to being private to drawCommandStream.c (only the now-reverted wall-tile group-delta approach needed it public)
  • src/drawCommandStream.c — DrawCommandStream_PushMaskedSprite's destW/destH resolution-independence fix (Task 1); isScaledSize set per push function; logo-cache invalidation from four independent "not the title screen" signals (wall-tile draw, object dispatch, gameplay starfield, the zoom-text interstitial's own 0x859a caller); s_gameplayActiveThisFrame broadened the same way for the status bar
  • src/CMakeLists.txt — new HD asset filenames, new source file