How it works · write-up
Xenon 2 renderer: 4x atlas upscale + sprite-stream frame interpolation
Status: implemented and live-tested. Written 2026-08-10; several sections below were revised during implementation after live testing surfaced bugs the original plan didn't anticipate -- this document now describes actual current behavior, not the original (partially incorrect) plan. See each section for what changed and why.
2026-10-06 implementation update
The terrain and background interpolation described later in this document has been replaced where it inferred scrolling from quads. Scroll RAM is now captured with each draw frame; the applied camera delta drives terrain and half-speed background presentation. Background offsets stay fractional through the C/TypeScript quad interface. See HIGH_FPS_SCROLL_JITTER_20261006.MD for the verified causes, assembly references, implementation and before/after recording measurements. Historical statements about retaining integer background offsets and independent background speed registers no longer describe current code.
The gameplay starfield still has integer-anchor jitter, confirmed against r584 video on 2026-10-06. The original records already contain fractional phase; capture currently omits it and interpolation fits rounded positions instead. The verification and proposed phase-based repair are in the same report. The historical claim below that better output resolution alone would resolve this issue is superseded by that finding.
Context
The Xenon 2 custom renderer currently draws from a 1x-resolution sprite atlas
(assets/spriteatlas_packed.{idx,raw,png}, 686x685) built offline by
hatari_dotnet from a captured sprite catalog. A 4x-upscaled version of that
same atlas already exists at assets/sprite_atlas_upscaled.png.png
(2744x2740, confirmed via IHDR inspection to be an exact uniform 4x upscale
of the same packed layout — not a repack). Nothing in the codebase currently
loads it.
Separately, the game's own logic only advances its drawable state once every
~4 hardware VBLs. Live instrumented capture (xenondoc/XENON2.MD §8.1, 82
observed swaps) shows the real cadence is one buffer swap every 4 VBLs
(~12.5Hz effective game-frame rate) — not every 2 VBLs (~25Hz) as
originally assumed going into this planning session. The GPU render/submit
path already runs on every VBL (~50Hz) today, it just resubmits the same
static content 3 times out of every 4. The goal is to make those 3 "held"
VBLs show smoothly extrapolated in-between positions instead of frozen
content, without touching the emulated game logic (and therefore without
risking desyncing its RNG-driven state).
The game will support both the existing low-res atlas and the new 4x atlas simultaneously at runtime, switchable with a key, rather than replacing one with the other — this also sets up a later crossfade transition between them without redesign.
This plan covers the desktop renderer only; the WASM/web renderer
(web/src/atlas.ts, web/src/views/spriteStreamView.ts) mirrors the same
formats/math and will need equivalent changes later, but that's explicitly
out of scope for this phase.
Both tasks were scoped by reading the actual C/C#/asset code directly
(no reliance on decompiled/reconstructed logic) — the interpolation task
additionally drew on xenondoc/XENON2.MD and
xenon2_position_update_mechanisms.html, which already document a partial
audit of which object-update routines are pure functions of known state vs.
which consume the game's shared RNG stream.
Task 1 — Atlas upscale (dual-resolution, runtime-switchable)
Why "just swap the file" isn't quite true
The runtime loader (src/spriteAtlas.c) computes every sprite's UV rect as
a fraction: u0 = atlasX / atlasWidth, where atlasX/Y/W/H come from
the .idx file and atlasWidth/Height come from the .raw file's own
header. Because it's a fraction, uniformly scaling both files by the
same factor reproduces identical UVs — but that means the .idx must be
regenerated at the new scale, not left as-is (SpriteAtlas_LoadPixels in
src/spriteAtlas.c:136-142 also hard-fails if the .raw's declared
dimensions don't exactly match the .idx's).
There's one real exception: the scrolling-background quad path in
StreamVertex_AppendQuad (src/sdlGpuRenderView.c:~1150) hardcoded
const float sourceScaleY = 1.0f; and uses it to convert a logical
(original-resolution) scroll offset into atlas V-fraction space. At 4x this
would sample completely wrong V coordinates. The struct field comment on
sourceOffsetY (src/includes/drawCommandStream.h:89-91) already
anticipated this ("may be applied to a higher-resolution replacement asset
by the renderer") — it just was never implemented.
Implementation note — the first fix attempt was itself wrong, twice.
The obvious-looking fix, sourceScaleY = atlasH / logicalWrapHeight (using
the background sprite's own resolved atlas-rect height), shipped first and
produced a visibly vertically-squashed background. The background mosaic's
atlas rect is taller than one screen — "the background image is larger
than a screen, only a scrolled slice of it is ever displayed" — so its own
rect height was never a valid stand-in for the atlas's overall scale (the
original hardcoded-1.0f code even said so directly: "Do not infer scale
from atlasH: the background asset's atlas rectangle can contain more data
than its repeating playfield period" — that warning was correct and got
missed on the first pass). The actual fix is SpriteAtlas_GetActiveScale(),
a new spriteAtlas.c function comparing the active atlas's overall pixel
dimensions against the low-res atlas's own (1.0 at 1x, 4.0 at 4x) —
independent of any individual sprite's rect size. See the corrected code
below.
Both resolutions ship together, loaded simultaneously
Per feedback: don't replace the 1x atlas with the 4x one — carry both, and let the player/dev switch at runtime. This means:
hatari_dotnetproduces all files in one run, always. No separate "upscale mode" subcommand. The existing default flow (capture → rectangle-pack → export) additionally looks forsprite_atlas_upscaled.png.pngalongside its other inputs and, if present, emits a second.idx/.rawpair at the derived scale in the same invocation. If that PNG isn't present (e.g. mid-iteration on a fresh capture before the external upscale step has run), the tool logs a warning and skips only the HD half — the normal 1x pipeline still completes.- New runtime asset filenames (both shipped, both loaded at startup):
assets/spriteatlas_packed.idx/.raw(existing, 1x, unchanged format) andassets/spriteatlas_packed_hd.idx/.raw(new, 4x). This does require adding the two new filenames tosrc/CMakeLists.txt'sHATARI_RUNTIME_ASSETSlist (CMakeLists.txt:108-115) — unlike the original single-atlas plan, this step is no longer avoidable, since both variants need to exist on disk simultaneously. src/spriteAtlas.c/.hbecomes instantiable rather than holding one set of file-scope statics: introduce a smallSpriteAtlasstruct wrapping the existing entries/pixels/dimensions, withSpriteAtlas_LoadIndex/_LoadPixels/_FindRect/_GetPixelstaking an explicit instance pointer.main.ccreates two instances at startup (g_spriteAtlasLow,g_spriteAtlasHigh) and loads both eagerly — both are small enough (~1.9MB and ~30MB resident) that keeping both in memory simultaneously is cheap on desktop, and it means switching is instant with no load-time hitch. A module-level "active atlas" pointer (or slot enum) is whatsdlGpuRenderView.c's per-frame code reads.- GPU side:
SdlGpuRenderView_InitializeSpriteStreamResourcescreates two GPU textures at init (one per atlas), uploads both once. The per-frame render pass binds whichever texture corresponds to the currently active slot — a cheap pointer swap, not a re-upload. - The
sourceScaleYfix generalizes for free:SpriteAtlas_GetActiveScale()(see fix below) is derived from the active atlas's own pixel dimensions, so switching the active slot automatically produces the right scale on the next frame — no separate code path needed per resolution. - Toggle key:
R, polled withSDL_GetKeyboardStatethe same way as this codebase's other keys — accepted regardless of which window has OS focus, matching existing convention (an earlier draft of this plan gated the poll on the debug window's own focus; dropped as unnecessary and inconsistent with how other keys already work here). - Future transition effect: noted, not built now. Because both textures are already GPU-resident simultaneously under this design, a later crossfade is just a shader blend between two bound textures over a few frames — this dual-load architecture is what makes that cheap to add later without restructuring.
Steps
-
New
hatari_dotnet/PngReader.cs— a hand-rolled, deliberately narrow PNG decoder (chunk parsing +System.IO.Compression.ZLibStreamfor inflate + manual scanline unfiltering, all five filter types), the read-side counterpart to the existing hand-rolledPngWriter.cs. No new NuGet dependency: an earlier draft of this plan usedSixLabors.ImageSharp, but that package requires accepting/bundling the Six Labors Split License (not MIT, and it demands a license file) — since this decoder only ever needs to read one specific, well-understood format (non-interlaced, 8- or 16-bit truecolor+alpha, i.e. PNG color type 6), hand-rolling it is small and removes the dependency question entirely. ReusesPngWriter.cs's existingCrc32helper to validate each chunk. ThrowsInvalidDataExceptionwith a specific message for anything outside that narrow format (palette/grayscale color types, other bit depths, Adam7 interlacing, an unrecognized critical chunk, a bad CRC, a truncated file, an unknown scanline filter byte) rather than silently mishandling it. The 16-bit → 8-bit downsample is therefore exactly controlled here (take the high-order byte of each big-endian sample — the PNG spec's own standard bit-depth-reduction technique), removing the ImageSharp-era open question of whether its downsample was byte-exact. -
New
hatari_dotnet/AtlasIndexReader.cs— counterpart to the existingAtlasIndexWriter.cs; parses the current SATL.idxformat back into{AtlasWidth, AtlasHeight, Entries[{StAddress,X,Y,W,H}]}. RefactorAtlasIndexWriter's inner write loop into a small shared helper so both the existing pack-time writer and the new rescale writer share one binary-format implementation. -
New
hatari_dotnet/UpscaledAtlasExporter.cs—ExportUpscaled(packedIndexPath, upscaledPngPath, outIndexPath, outRawPath): - Read the just-produced 1x.idx(rects for the current run's packed layout). - Decodesprite_atlas_upscaled.png.pngviaPngReader.ReadFile. - Derive the scale factor from actual dimensions (image.Width / packed.AtlasWidth) rather than hardcoding4— assert it's exact and equal on both axes, log what was found. - Multiply every.idxrect'sX,Y,W,Hby that scale (pure multiplication — this is not a repack;RectpackSharpis not invoked here since the layout is already correct). - Convert the decoded RGBA to BGRA8 (mirror of the existingSpriteAtlasExporterRGBA byte-swap) and write via the existing, unchangedRawAtlasWriter.WriteBgra. - Write to temp files then move into place, so an interrupted run can't leave a half-written pair. -
hatari_dotnet/Program.cs— in the existing single flow, right after the currentExportPacked(...)call, additionally check whethersprite_atlas_upscaled.png.pngexists next to the other inputs; if so callUpscaledAtlasExporter.ExportUpscaled(...)writing tospriteatlas_packed_hd.idx/.raw; if not, log and continue (1x output is unaffected either way). No new subcommand/mode. -
src/CMakeLists.txt— addspriteatlas_packed_hd.idxandspriteatlas_packed_hd.rawtoHATARI_RUNTIME_ASSETS(CMakeLists.txt:108-115) alongside the existing three entries. -
src/spriteAtlas.c/.h— refactor file-scope statics into an instantiable struct + instance-taking functions (see above). Existing call sites (main.c,sdlGpuRenderView.c) updated to pass the active instance explicitly. -
src/main.c— load both atlases at startup (fourSpriteAtlas_Load*calls instead of two); keep the existing fatal-if-missing behavior for the low-res pair, decide whether the HD pair missing should be fatal or a soft fallback-to-low-res (recommend soft fallback, so a dev build without the HD assets deployed still runs). -
src/sdlGpuRenderView.c— two GPU textures, active-slot bind, and thesourceScaleYfix inStreamVertex_AppendQuad(~line 1150):c const float sourceScaleY = SpriteAtlas_GetActiveScale();(notatlasH / logicalWrapHeight— see the implementation note above for why that first attempt was wrong.)SpriteAtlas_GetActiveScale()compares the active atlas's overall pixel dimensions against the low-res atlas's own. At 1x this evaluates to 1.0 (unchanged, self-consistent with today); at 4x it evaluates to 4.0 automatically — correct for whichever atlas is currently active, with nothing to hand-maintain per resolution, and correct regardless of any individual sprite's own atlas-rect size. -
src/drawCommandStream.c,DrawCommandStream_PushMaskedSprite— not in the original plan; discovered live after Task 2's dual-atlas testing. This function setentry->destW = atlasW; entry->destH = atlasH;directly from whichever atlas was active at capture time. That silently assumed atlas-texel == logical-pixel (true at 1x, where it went unnoticed for a long time), butdestW/destHmust stay in logical 320x200 units always, exactly likex/yalready do — with the HD slot active during capture, this baked in 4x-too-large on-screen sizes, reported live as "sprites rendered in the gameplay screen are 4 times too big" after pressing the resolution-toggle key. (Screens using the zoom- text/logo path were unaffected — those already computedestW/destHindependently ofatlasW/atlasH, which is exactly why only gameplay sprites showed the bug.) Fixed by re-resolvingdestW/destHspecifically against the low-res atlas slot (which has one texel per logical pixel by definition) regardless of which atlas is active for UV sampling, falling back to the passed-inatlasW/atlasHonly if that lookup fails (e.g. no low-res atlas loaded at all).DrawCommandStream_PushMaskedSpriteScaled(zoom-text, logo) was already unaffected -- it takesdestW/destHas independent caller-supplied parameters, never derived fromatlasW/atlasHat all.
Visual-bleed risk (verification only, not a code fix)
The packer used a 2px border between sprites at 1x (8px at 4x). If the
external upscaler processed the whole packed image with a wide receptive
field, neighboring sprites' content may have bled into that border band.
Verify by cropping a few sprites' own rects from the new atlas and
comparing against a naive nearest-neighbor 4x upscale of the same crop from
the original spriteatlas_packed.png — a "melted toward its neighbor" edge
is the fail signal. If it fails, the fix is regenerating that external
asset with more border margin, not application code.
Resolved 2026-09-21 (the opposite failure). The border (now 4px) was
empty transparent black, so the upscaler treated every sprite edge as an
image edge and produced a dark, partly transparent rim one to two HD texels
wide on every sprite -- measured on level-1 wall tiles, source edge rows of
brightness ~110-180 came back at 0-30. On the wall grid that rim showed as
seams the moment HIRES SPRITES was on. pad_packed_atlas_edges
(xenon_tools/build_level_atlases.py, run on the chaiNNer input right
before the upscale) now replicates each sprite's outermost pixels into its
4px border (clamp-to-edge, transparent edges stay transparent), so the
upscaler sees continuous content past the rect; the HD rects are unchanged
and the .NET pass after the upscale rewrites the clean 1x PNG. Verified on
the same tiles: edge rows now match their interiors, no transparent
corner pixels.
That removed the hard dark rim but not the grid: upscaled in isolation
(even edge-padded) every 16x16 wall tile is still shaded as its own object,
lighter in the middle and darker towards its edges, so walls and the
tile-built bosses read as a grid of pillows. So wall tiles are upscaled a
second time in context: build_tile_mosaic lays each level's tiles out as
the game does (the 20-column tilemap plus the composite boss grids, using
the 1x atlas's own pixels), the mosaic goes through the same chaiNNer run
as one image, and apply_tile_mosaic cuts the 4x cells back into the
level's upscaled atlas before the HD export. A tile index recurs across the
map with different neighbours; choose_tile_placements picks the cell whose
8-neighbour signature is most common for that index (ties: most non-empty
neighbours, then the earliest cell). Tiles that appear in no grid
(animation frames) keep their per-tile upscale. Regenerated 2026-09-21:
236/254/182/351/157 tiles per level; the walls no longer show the grid --
where the tile sits in the neighbourhood it was upscaled in.
2026-09-22. That is only 24-41 % of the cells: the maps reuse each tile
index in 5-9 different neighbourhoods, and everywhere else the mosaic tile
carries the wrong neighbour's shapes cut off hard at the edge. Compared side
by side with the per-tile upscale (clamp-, wrap- or mirror-padded) the
mosaic's continuous texture was still preferred, so it stays the default
(--no-tile-mosaic for the per-tile version). Measurements and the open
options are in HD_WALL_TILE_SEAMS.md.
Verification plan
- Numeric spot-check: hand-compute UVs from both the 1x and 4x pairs for a
few known
stAddressvalues, confirm they match to float precision. - Run with
RENDER_VIEW_SPRITE_STREAM, toggle between resolutions live, confirm identical placement at both, higher fidelity only at 4x. - Watch the wall-tile mosaic scroll continuously through several full
192-scanline wrap periods at 4x — this is what exercises the
sourceScaleYfix specifically; a stale/wrong scale shows up unmistakably as a frozen background or sampling into a neighboring sprite's atlas region. - Confirm the existing dimension-mismatch guard in
SpriteAtlas_LoadPixelsstill rejects a deliberately-mismatched pair per instance. - Toggle to the HD slot mid-gameplay and confirm ordinary sprites (ship,
enemies, bullets) stay the same on-screen size as at 1x — this is what
exercises the
destW/destHfix in step 9 above; a regression shows up unmistakably as everything suddenly rendering ~4x too large.
Task 2 — Frame interpolation
Three different identity models — this matters a lot for the design
Investigated directly (not assumed) how each quad type's objectId behaves
across frames, since the interpolation design lives or dies on whether
objectId is a stable match key:
| Quad type | objectId source |
Stable across motion? |
|---|---|---|
| Gameplay object (enemy/ship/bullet/etc.) | Real MyObjectEntry ST address |
Yes, while the object is alive (may be reused after it dies — see below) |
| Background mosaic | Fixed singleton 0xFFFE0000u (drawCommandStream.c:725) |
Yes, always — one quad per frame, motion lives in sourceOffsetY, not x/y |
| Wall tile | Synthetic, derived from screen position: 0xFFFF0000u \| (y<<9) \| x (drawCommandStream.c:656-659) |
No — a tile's "id" changes every frame it moves, by construction |
This means a single generic "match by objectId, extrapolate velocity"
scheme is right for gameplay objects and (in a modified singleton form)
for the background, but would be silently useless for wall tiles: a moving
tile's synthetic id never matches its own previous-frame id, so every wall
tile quad would register as a "miss" every frame and just hold at its
last known position — safe (no wrong guesses, no gaps) but zero
smoothing benefit for wall tiles specifically, and wall tiles are a large
fraction of the visible playfield.
Wall tiles: tried group-level scroll extrapolation, reverted to always-held.
The first implementation extrapolated every wall-tile quad by one shared
vertical delta, driven by the live fine-scroll register (FINE_SCROLL_ADDR
= 0x00000cd6u, drawCommandStream.c) — reasoning that the real game moves
every wall tile in lockstep via that one shared register, so this should be
the correct model, not just a workaround. Live testing showed this was
wrong: most wall tiles don't actually move continuously at all. Only the
top partial row-band has genuine sub-pixel motion; every other row is
redrawn each real frame at a fixed 16px-grid Y position, with only its
tile content (spriteId) changing as the level scrolls past — a discrete
asset swap, not a position change. Applying one shared Y shift to all of
them uniformly pushed the grid-fixed rows off their true position during
held VBLs, then snapped them back on the next real frame — visible as
forward-then-backward jitter, reported live and confirmed by re-reading the
capture code (DrawCommandStream_OnTileDrawInstructionFetch's own
row-band-Y-derivation comment). Current behavior: wall tiles are never
extrapolated at all — always held at their last known real position,
recognized by their synthetic objectId range
((objectId & 0xFFFF0000u) == 0xFFFF0000u, which cannot collide with real
MyObjectEntry addresses or the background's distinct 0xFFFE0000u). A
correctly row-band-aware model (extrapolating only the genuinely-moving top
band) is a plausible future refinement but wasn't built — the visual cost of
holding tiles for up to ~80ms is low, and getting it wrong was worse than
not smoothing them at all.
This directly answers the specific questions originally raised: - Do tiles have unique ids? Yes, but positionally derived, not a stable per-instance identity — see table above. - What happens when a new tile scrolls into view with no velocity yet — is there a gap between it and existing tiles? No gap: it (like every other wall-tile quad) is simply held at its real captured position during synthetic frames, same as any object with no extrapolation. - Should all wall tiles move in sync with a single velocity? That was the original hypothesis and it's half right — the top row-band genuinely does, but the middle/bottom rows don't move continuously at all, they're redrawn at fixed grid rows with changing content. Building and shipping the "all tiles, one velocity" model live is what surfaced this distinction.
Group-shared-delta fix, built in a later session — corrected understanding of why the first attempt jittered. A first pass at revisiting this mis-scoped the fix to "only the top row-band genuinely moves" (reasoning from the row-band table above plus the original revert's symptom) and built a row-band-aware version that only extrapolated top-band quads. That framing was wrong: the whole 16x16 tile grid genuinely scrolls as one, at a constant speed, with every tile sharing the same offset. The game's three-row-band split exists purely because the original 68000 blitter has to manually avoid drawing rows that would land off-screen — a real cost worth avoiding on that CPU — not because different rows move differently. This renderer gets that clipping for free from the GPU, so the row-band split doesn't need to be reproduced at all for positioning purposes.
The fix: spriteStreamInterpolation.c's isWallTile branch derives one
shared scroll delta per held tick from the top row-band's own Y alone (a new
FindSharedWallTileTopBandY helper, reading newerQuad->y < 0 — the
capture side already pushes the top band's unclipped logical Y as
-fineScroll, a direct, unambiguous live-register read, unlike trying to
infer a delta from the middle/bottom bands' own measured destination
addresses) — extrapolated the same wide-baseline, wrap-aware,
linear-only way as the background mosaic's sourceOffsetY (reasoning
that the wall scroll register plausibly dithers its step count the same way
the background's confirmed does, so a quadratic term would amplify that
noise rather than track real curvature) — then applies the resulting
delta (not an absolute recomputed position) uniformly to every
wall-tile quad's own already-correct captured Y, top/middle/bottom band
alike. No per-band distinction, no wire-format field, no per-quad identity
matching (impossible anyway, since wall-tile objectId is position-derived)
— just one shared scalar's own frame-to-frame delta applied everywhere.
This is structurally the same shape as the original reverted attempt (one
shared delta for every wall-tile quad) — confirmed live to be correct
this time via a temporary diagnostic trace (screentrace.c, logging the
topmost and first mid-band wall-tile quad's position every VBL): the real
(non-extrapolated, held=0) captured Y for both the top row-band and a
representative mid-band row advanced by the identical +1px every real
frame, disproving the original revert's own diagnosis ("most wall tiles are
redrawn at a fixed 16px-grid position, only content changes" — see
FINE_SCROLL_ADDR's corrected comment, drawCommandStream.c). The
extrapolated held=1 ticks matched too: e.g. 4.000 → 4.250 → 4.500 → 4.750
across three held VBLs, then the next real frame landed at exactly 5.000
— zero prediction error. So the earlier revert's jitter came from how that
delta was derived/applied (exact mechanism not re-examined — the buggy code
was already fully removed by the time this was revisited), not from the
tiles' actual motion being non-uniform. One accepted trade-off either way: a
wall-tile coarse-scroll boundary crossing hands off to a physically
different tile, so while the shared Y delta reads smoothly across that
boundary, the tile content (never extrapolated, only ever held) still pops
to the new tile's sprite at that one instant — bounded to once per full
tile-scroll cycle, same "error visible for at most one real-frame interval"
bound as every other branch in this module.
Extrapolation model for gameplay objects — linear/quadratic, with a plausibility clamp
Pure linear (2-frame, constant-velocity) extrapolation is exact for the
several documented updateProcs that really are constant-velocity between
real frames (compass-table enemies, radial-burst projectiles, rigid-follow
attachments), but will systematically miss curved motion: it extrapolates
along the tangent, undershooting a circle's true path and mistiming
sine-wave direction reversals near the peaks/troughs. This is the model
actually used for ordinary gameplay objects (ship, enemies, bullets,
stars — anything with a real, stable MyObjectEntry objectId, i.e.
everything that isn't the background singleton, a wall tile, the status
bar, or a scaled/zoom quad — see below): cache 3 real frames and use
constant-acceleration (quadratic) extrapolation when all 3 are available
and pass the plausibility clamp, falling back to linear (2 known) or hold
(1 known / a miss) otherwise:
v = (p[t] - p[t-1]) / Δ
a = (p[t] - 2*p[t-1] + p[t-2]) / Δ²
p(t+τ) = p[t] + v*τ + 0.5*a*τ²
This is a second-order Taylor expansion of the true motion — it tracks
gentle curves and sine-like motion meaningfully better than linear, still
never touches game logic or RNG, still costs nothing but one more cached
frame and a few extra multiplies per object. It's still fundamentally an
extrapolation, not a simulation: it will still visibly lag at sudden
curvature changes (sharp direction reversal, a collision bounce, a spawn
transition) since it has no way to know the next change in curvature
before it happens — same self-correcting bound as before (any error is
visible for at most one real-frame interval, then overwritten by truth).
Final position is rounded to nearest (lroundf), not truncated — see the
background section below for why that distinction mattered live.
Zero-velocity guard, per axis (added after live testing of the ship):
the quadratic fit above assumes the last-measured acceleration stays valid
going forward, which is actively wrong for step-function motion (joystick-
driven movement — constant velocity while held, dropping straight to zero
on release, no coasting). If an axis has already stopped by the two most
recent real frames (p[t] == p[t-1]), any acceleration term derived from
an older, still-moving sample is spurious: it reads the stop as
deceleration and extrapolates past zero in the opposite direction for the
held ticks, snapping back once the next real frame confirms the object
never moved — visible live as the ship jittering for ~2 real-frame
intervals after releasing the stick. Fix: an axis with p[t] == p[t-1] is
held at p[t] unconditionally, skipping both the quadratic and linear
formulas for that axis (independently of the other axis). Does not address
the milder, unreported symmetric case of an axis just starting to move from
rest — a real option if that turns out to need fixing too is reading live
input state directly to detect a motion-stopping/starting input change
mid-held-interval rather than waiting for it to show up in real-frame
history, but that would be a ship-specific special case (this module is
otherwise backend/object-agnostic), deferred unless the guard above proves
insufficient.
Second baseline added to the same guard (live testing surfaced a
different case): a stationary object whose effect animation cycles through
sprites with different x_origin/y_origin (e.g. a rear-cannon muzzle
flash) produces a genuinely different captured x/y every real frame
even though the object never moves — a two-frame A/B/A/B cycle. Consecutive
real frames never agree here, so the original guard didn't catch it; the
quadratic fit read the middle sample as "acceleration continuing away from
it" and visibly exaggerated a true displacement of zero into a swinging
motion. Fix: the guard now also fires when the oldest and newest cached
frame agree on an axis (net motion over the full 3-frame window is zero
regardless of the middle sample), independent of the original
newest-vs-older check.
Ring-buffer stale-flag crash (fixed): DrawMaskedSpriteEntry.isPerspectiveStar
(added for the star-wrap fix above) was only ever set by
DrawCommandStream_PushFlatColorQuad; PushMaskedSprite/PushMaskedSpriteScaled
never touched it, so a ring-buffer slot last written by a perspective-star
push could leak a stale true into the next entry that reused that slot —
including, once, the background singleton, tripping
spriteStreamFrame.c's "SPRITE_QUAD_BACKGROUND carries no flags" assert
after extended play. Fixed by explicitly resetting it false in both of
the other two push functions, matching the existing defensive pattern
already used there for isFlatColor/isScaledSize.
Option C remains deferred, but is now the explicit answer for whatever
still looks wrong after this: tight circles, high-frequency sine, or any
pattern where quadratic extrapolation is visibly insufficient are precisely
the cases worth reimplementing that specific updateProc faithfully in C
(Option C) as a narrow, targeted follow-up — informed by which patterns
actually look bad in practice, not built broadly upfront.
The background mosaic — a different extrapolation model than gameplay objects
The background singleton (objectId = 0xFFFE0000u) doesn't move via x/y
at all — its motion lives entirely in source.background.sourceOffsetY, a
value that wraps modulo 192 (the mosaic's repeat period). Extrapolating
it needed two live-tested corrections beyond what the original plan
described:
-
Wraparound must be unwrapped before doing velocity math. A legitimate wrap (e.g.
190 -> 5) read as a huge spurious jump if fed directly into a linear/quadratic fit, producing a wildly wrong extrapolated position that then snapped back once the next real frame arrived — exactly the "forward then backward" jitter reported live. Fixed with anUnwrapRelative(value, reference, period)helper that re-expresses a sample as the representative of its residue class closest to a reference sample, applied before any velocity/acceleration math. -
Quadratic extrapolation was actively wrong for the background, confirmed via ReVa disassembly of
BackgroundScrollCursor_UpdatePerFrame_ FUN_0000702c(the cursor's own per-frame update routine). The step count advanced per real frame is(in_D0w & scrollSpeedMask & 1) + (in_D0w >> 1)— a dithered/fractional-speed technique: to make the average scroll rate non-integer (e.g. 3.5 texels/frame), the game alternates between 3 and 4 steps frame-to-frame. So even during genuinely "constant speed" scrolling, the per-frame delta itself isn't constant — it oscillates by ±1 step. Both the 1-interval linear delta and the quadratic acceleration term (a second finite difference) amplify this dither into visible jitter, no matter how correctly wraparound is handled. Fixed by widening the velocity baseline to span 2 real intervals (oldest -> newer, divided by 2) when 3 real frames are cached, which cancels a single-frame ±1 dither step the same way averaging over any longer window cancels noise; falls back to the narrower 1-interval linear estimate only when 2 real frames are cached. Quadratic is never used for the background — the plan originally assumed the same linear/quadratic model as gameplay objects applied uniformly; live testing showed the background specifically needed both a wider baseline and to drop quadratic entirely, for reasons that don't apply to ordinary objects (whose positions are plain integers, not run through an extra scroll-speed dithering step). -
Rounding, not truncating, the final wrapped value mattered visibly. A plain
(uint32_t)wrappedcast floors toward zero; e.g. a continuous extrapolated value of94.88(much closer to a real95on both sides than to94) was being displayed as94for the whole held interval, then jumping to95on the next real frame — a small but real 1-texel flicker, caught by comparing a live diagnostic log's raw numbers against the displayed result. Fixed withlroundf, matching every other extrapolation path in the module (gameplay objects already rounded correctly; the background path was the one place still truncating).
Zoom-text glyphs and the title-screen logo — also held, not extrapolated
Not in the original plan at all — discovered live after the wall-tile
and background fixes, reported as "the zoom-in (and out) text and logo is
jittery — after a logo zoom-in finishes, it moves a little after it settles
in place." Both use DrawCommandStream_PushMaskedSpriteScaled, where
position (x/y) and on-screen size (destW/destH) are both driven by
the same live "scale" register, but via different formulas — position
is linear in scale (confirmed via the logo's own plate comment:
x = ((D0*scale)>>4)+160), size comes from a rotate/popcount lookup table
(not necessarily linear). This module only ever extrapolates position,
never destW/destH (held at the last real value) — so extrapolating
position alone drifts out of sync with the frozen size every held VBL,
worst right when the zoom animation stops changing: the predictor doesn't
know that yet, so it keeps extrapolating for a couple of held VBLs before
the next real frame confirms the truly-settled position and snaps back.
The logo has a stable objectId (spriteId, constant throughout its whole
zoom animation) so it was reliably matched and extrapolated every held VBL.
Zoom-text glyphs' objectId (destScreenAddr + characterIndex) turned out
to also be stable enough to match — the scaled text path computes its
real position independently, without touching A0/destScreenAddr, so
that value stays roughly constant across the animation rather than moving
with it the way a wall tile's synthetic position-derived id does.
Fixed with a new DrawMaskedSpriteEntry.isScaledSize field (threaded
through to SpriteRenderQuad as SPRITE_QUAD_FLAG_SCALED_SIZE), set by
DrawCommandStream_PushMaskedSpriteScaled and checked (false) by both
DrawCommandStream_PushMaskedSprite and DrawCommandStream_PushFlatColorQuad.
Flagged quads are always held, never extrapolated — same "hold rather than
guess wrong" treatment as wall tiles and the status bar, on the reasoning
that these are discrete UI animations stepping once per real frame, not
continuous physics-driven motion, so extrapolation doesn't meaningfully
help them anyway (modeling the shared scale-to-size/position relationship
instead would still have the same end-of-animation overshoot problem, just
for a different quantity).
One subtlety this surfaced: the background mosaic also goes through
PushMaskedSpriteScaled (it needs an explicit 320x192 size independent of
its own, much taller, atlas rect — see the destW/destH fix in Task 1) so
isScaledSize is true for it too, but for an unrelated reason (a fixed
constant clamp, not a live scale register). A pre-existing assertion in
SpriteStreamFrame_AppendEntry requires background quads to carry no flags
at all; naively propagating isScaledSize onto the background's own
SpriteRenderQuad violated it. Fixed by excluding isBackgroundWrap
entries specifically from the flag propagation in spriteStreamFrame.c —
behaviorally identical either way, since the interpolation module already
special-cases SPRITE_QUAD_BACKGROUND before it would ever check this flag.
objectId reuse — explicit identity tracking (objectIdentity.c/.h)
ST addresses backing MyObjectEntry are reused after an object dies and
a new one is allocated at the same address, and the 48-slot starfield record
array is reused by two semantically incompatible rendering modes (see "Three
different identity models" above). Rather than continue inferring reuse
heuristically, capture hooks now detect the two authoritative lifecycle
events directly (found via ReVa disassembly, not decompiler guesses — both
turned out to be single choke points despite dozens of call sites):
- Allocate/recycle:
AllocateOrRecycleObjectEntry_FUN_00002be4's single sharedrtsat0x00002c00— every one of its ~26 callers converges here, withA0holding the (re)allocatedMyObjectEntry*. - Destroy/free:
DestroyObject_UnlinkAndFree_FUN_0000107c's entry PC0x0000107c— installed as a shared "destroy" function-pointer slot at ~30 object-setup sites, but it's one physical function, so hooking its own entry catches every invocation regardless of caller.A0holds the entry about to be unlinked and freed. - Starfield mode switch: already distinguishable at capture time (which of the two starfield hooks fired this frame) — no new hook needed, just a tracked "last mode" flip that resets all 48 record slots.
src/objectIdentity.c/.h is a small host-side registry (never touches ST
memory — there's nowhere sanctioned in the guest's own state to put this) of
{slotAddr, identityId}, backed by one global monotonic counter. SpriteRenderQuad
gained a new identityId field (objectId stays, purely for
diagnostics/ReVa correlation). ObjectIdentity_Get(slotAddr) degenerates to
slotAddr itself for any slot that's never been through
ObjectIdentity_Reset — so untouched quads (background, wall tiles, status
bar, an object that simply never dies while observed) need no table entry
and no special-casing; identityId is a strict superset of comparing raw
objectId, only actually diverging where a reset was recorded.
spriteStreamInterpolation.c's FindQuadById/FindMatchingQuad now compare
identityId, so a reused-address mismatch is structurally impossible to
match rather than merely unlikely — a recycled/destroyed/mode-switched slot
simply has no same-identity candidate in the older cached frame, correctly
falling through to "no match, hold" instead of needing the plausibility
clamp to catch it after the fact.
The plausibility clamp (DisplacementIsPlausible) stays, now as a
second layer for the one lifecycle event not yet hooked: the perspective
starfield's own internal RNG-driven respawn (a star keeps the same
identityId — same slot, same mode — but genuinely jumps to a fresh random
position mid-flight; see Starfield_DrawPerspectiveFlythrough_FUN_000085b4's
doc comment). Its only hookable point is an inline branch inside a per-star
loop, not a clean call/return site like the two above, so it's deliberately
out of scope for now. (The clamp's tuning — K=8 size-relative with a 64px
absolute cap — was already validated live against the reused-address cases
this change now handles structurally; unchanged here.)
Explicitly NOT treated as an identity change: the vertical-scroll
starfield's own Y-position wrap (modulo the 192-line playfield, confirmed
via disassembly of Starfield_UpdateVerticalPositions) is a continuous,
same-instance quantity — a bounded modular correction, not a reset — so it
stays a math problem for UnwrapRelative, not an ObjectIdentity_Reset
call.
Star Y-wrap, now applied (live testing surfaced a 192px "misprediction"
at a 1x1 flat-color quad, confirming this was still needed): the blocker
described in the previous revision of this doc — distinguishing
vertical-scroll stars from perspective-mode stars on the quad itself, since
applying the wrap unconditionally would misread a perspective respawn's
genuine discontinuity as a safe wrap — is resolved by a new
DrawMaskedSpriteEntry.isVerticalScrollStar bool (drawCommandStream.h),
surfaced as SPRITE_QUAD_FLAG_VERTICAL_STAR on the render side. In
spriteStreamInterpolation.c's generic gameplay-object path, a quad is
treated as a wrap-aware vertical-scroll star purely when that flag is set:
its y history is unwrapped via UnwrapRelative(..., BACKGROUND_PLAYFIELD_HEIGHT)
— relative to the newest sample, same convention the background branch
already uses — before both the plausibility gate and the linear/quadratic
formulas, and the final extrapolated y is folded back into [0, 192) the
same round-to-nearest-then-adjust way the background's sourceOffsetY
already is. Unwrapping before the plausibility gate specifically (not just
before the extrapolation math) mattered live: a 1x1 star's own K * width
plausibility clamp is tiny (K=8, width=1 → 8px), so a raw near-full-period
wrap delta (up to ~191px) would otherwise get rejected as implausible and
just held, never reaching the wrap-aware math at all. x is untouched —
gameplay-mode stars have no horizontal motion (confirmed via disassembly of
Starfield_UpdateAndDraw_48Stars: the X bitmask word is only ever read).
DisplacementIsPlausible's signature changed from two quad pointers to
explicit (dx, dy, width, height) to make this possible — callers now
compute dy themselves (unwrapped where relevant) rather than the function
reading raw ->y fields directly.
Positive marker, not an inferred double-negative (corrected after
feedback): the flag started life as isPerspectiveStar, with the wrap
condition being "FLAT_COLOR and that flag absent" — but FLAT_COLOR is
also used for the status-bar weapon meter today (harmless in practice, that
case is held earlier via SPRITE_QUAD_FLAG_STATUS_BAR and never reaches
this code) and isn't guaranteed to stay starfield-only in general — a
hit-highlight flash effect would plausibly also be flat-colored, and
inferring "must be a vertical star" from the absence of one unrelated flag
would silently wrap-adjust any such quad's position too. Replaced with a
dedicated positive isVerticalScrollStar, set true only by
DrawCommandStream_OnStarfieldDrawInstructionFetch (the gameplay starfield
hook itself, not inferred from what it isn't) and false everywhere else,
including explicitly reset in the other two push functions per the
ring-buffer-staleness lesson below. isPerspectiveStar was kept alongside
it at first, purely descriptive.
Superseded — consolidated into DrawEntryKind/DrawEntryStarKind enums:
once there were two mutually-exclusive pairs of bools on
DrawMaskedSpriteEntry (isFlatColor/isBackgroundWrap choosing a 3-way
draw kind; isPerspectiveStar/isVerticalScrollStar choosing a star
sub-kind), both were replaced with dedicated enums —
DrawEntryKind kind (DRAW_ENTRY_KIND_TEXTURED/_BACKGROUND/_FLAT_COLOR)
and DrawEntryStarKind starKind (DRAW_ENTRY_STAR_KIND_NONE/
_VERTICAL_SCROLL/_PERSPECTIVE, meaningful only when
kind == DRAW_ENTRY_KIND_FLAT_COLOR) — making the mutual exclusivity a
property of the type instead of a convention every call site has to
uphold by hand. isFlash/isStatusBar/isScaledSize stayed bools:
they're genuinely independent and combine freely with any kind or with
each other (e.g. a textured quad can be both isStatusBar and
isScaledSize), so folding them into the same enum would either explode
combinatorially or silently forbid a real combination. All push-function
call sites, spriteStreamFrame.c's kind dispatch, and every place that
gated on the old flag pairs were updated to match; the ring-buffer-reset
discipline established by the crash above applies identically to the new
enum fields (every push function sets both unconditionally).
Star Y dithered-step overshoot (mitigated, not eliminated — accepted,
deferred): live testing after the wrap fix above still showed
vertical-scroll stars "progressing down ok, then jumping back up a
little". Confirmed via disassembly of Starfield_UpdateVerticalPositions
(no new hook needed — the existing plate comment already had the answer):
a star's Y is driven by a phase accumulator that advances a fixed
sub-pixel amount every frame, with the visible Y only stepping by one
scanline when that phase crosses a threshold — so the raw frame-to-frame
sequence dithers (0,0,1,0,1,0,0,1,...), only averaging to the star's
true speed over several frames. This is structurally identical to the
background scroll's own already-fixed dithered-speed bug (see that branch
below) — the standard quadratic/narrow-linear formula picks up the noisy
single-interval velocity, overshoots on a "step" frame, and snaps back on
the next real frame. Mitigated the same way as the background: for
isVerticalStar quads specifically, when 3 real frames are cached, use
the wide oldest -> newer baseline average velocity with pure linear
extrapolation (no acceleration term — the apparent curvature here is the
dither noise) instead of quadratic; falls back to the ordinary narrow
1-interval linear estimate when only 2 real frames are cached. Every other
quad type is unaffected — this is gated on isVerticalStar exactly like
the wrap-unwrap treatment above it.
Residual confirmed still present after the mitigation above — root cause
is a shape mismatch, not a rate-estimation error, so it can't be fully
fixed by better velocity averaging alone. The star's true position is a
step function: it holds at a fixed integer scanline for a variable number
of real frames, then instantly jumps exactly one scanline — it never
occupies a fractional position. Any smooth extrapolation curve (linear
or quadratic, narrow or wide baseline) necessarily predicts fractional
in-between positions during the held ticks, including during real
intervals that turn out to be "hold" frames where the star doesn't
actually move — and that predicted-but-never-real progress is exactly
what gets visibly corrected when the next real frame confirms the star
stayed put. The wide-baseline fix reduces the size of this mismatch
(a better long-run rate estimate means less average error) but cannot
eliminate it, since no continuous curve can exactly reproduce a staircase.
Also a plausible contributing factor, not yet isolated from the above:
the zero-velocity guard's narrow (newest == older) check can itself
match a single dither "hold" step and force a full flat prediction for
that interval, bypassing the wide-baseline average entirely for that
window — worth revisiting together with the shape-mismatch issue if this
is picked up again later.
Deliberately left as a known, accepted issue for now (not pursued further this pass): the practical fix — matching the smoothness of extrapolation to the underlying sub-pixel motion — is expected to become naturally easier once gameplay resolution is upscaled beyond the native 320×200 (a separate, already-planned future work item), since a higher-resolution star position has more headroom to represent genuine sub-pixel progress smoothly rather than needing to snap between whole-native-pixel steps. Revisit this specifically once that upscale work begins, rather than continuing to chase it at native resolution.
Origin recovery for textured objects (live testing: a stationary rear
cannon's muzzle-flash animation visibly "moved left and right" while
firing): confirmed via disassembly of blitting_top_level_function_FUN_010ae
that the real game computes on-screen position as anchor - spriteData->
x_origin/y_origin per draw (sub.w (A0)+,D0w right at entry) — exactly
what the capture side already replicates into DrawMaskedSpriteEntry.x/y.
The problem: different sprites in the same multi-frame effect animation (a
muzzle flash, a growing explosion) can carry different origins to stay
visually anchored despite changing size, so the origin-adjusted x/y this
module was extrapolating can legitimately shift a few px every real frame
even while the object's true anchor never moves — and the animation cycle
can be longer than the zero-velocity guard's 3-frame window can reliably
pattern-match (confirmed live: the rear cannon's cycle is longer than 2
frames).
Fixed at the source rather than pattern-matched: DrawMaskedSpriteEntry
gained originX/originY (the origin values already being subtracted,
now kept alongside the result instead of discarded — populated only by
DrawCommandStream_PushMaskedSprite's real per-object callers; wall
tiles/HUD glyphs pass 0/0, no origin concept applies), surfaced as
SpriteRenderQuad.source.textured.originX/Y. For SPRITE_QUAD_TEXTURED
quads, spriteStreamInterpolation.c now fits velocity/acceleration on the
recovered anchor (x + originX) instead of the raw screen position, for
every cached sample — then converts back to screen space at the very end by
re-applying the current (held, unextrapolated — spriteId/size are never
extrapolated either) sprite's own origin once. The zero-velocity guard now
operates on this same "fitting-space" value too, so a truly stationary
animated object's recovered anchor matches exactly across real frames and
guard #1 (consecutive-frame check) already catches it directly — guard #2
(oldest-vs-newest) is kept as a backstop for non-textured cases or cycles
that happen to realign within the 3-frame window, but is no longer the
primary defense for this class of bug.
This generalizes cleanly to size-changing effects (explosions) the same way: since size/spriteId are held rather than extrapolated regardless, an object that's genuinely stationary while its effect grows now correctly extrapolates zero position change, instead of misreading the growing sprite's shifting origin as motion.
RNG hazard (why we still never re-run game logic)
xenondoc/XENON2.MD §2.8 and xenon2_position_update_mechanisms.html §09
document that several updateProcs consume a shared gameplay RNG stream
per real frame. Replaying any such routine to synthesize an in-between
frame would advance that shared RNG stream early, silently diverging
gameplay from the original. Everything in this plan only ever extrapolates
already-observed positions from the draw-command stream — it never
re-executes game/CPU logic, so this hazard doesn't apply to any of it,
including the quadratic upgrade above (still just arithmetic over cached
history, not simulation).
Architecture
New module: src/spriteStreamInterpolation.c + .h, added to
src/CMakeLists.txt next to spriteStreamFrame.c. Keeps screentrace.c's
RENDER_VIEW_SPRITE_STREAM branch a thin orchestrator; SpriteStreamFrame_Build's
existing call, arguments, and output buffer stay completely untouched — the
new code is purely additive after it.
State (file-scope static, not exposed via the header):
- Three cached SpriteRenderQuad[] buffers + counts, oldest to newest
(s_history[0..2]) — used for gameplay-object linear/quadratic
extrapolation and the background's own wide-baseline extrapolation (see
above). No longer holds a separate wall-tile scroll-register cache — that
approach was built, found to cause jitter, and removed (see "Wall tiles"
above); wall tiles are simply never extrapolated now.
- A monotonic VBL tick counter, incremented once per call (already called
exactly once per VBL).
- The tick-gap measured at the last real-frame transition
(s_measuredIntervalTicks), seeded near the documented ~4-VBL cadence but
overwritten by the first observed transition — self-corrects rather than
hardcoding "4" anywhere.
- An enable/disable flag for the A/B toggle — defaults to false (not
the originally-planned "on by default"; changed after live testing showed
several of the bugs described above with it on, and it's simplest for a
session to start from the known-correct non-interpolated baseline and
opt in).
- Prediction-error metrics state (not in the original plan): a snapshot of
the last extrapolated output, plus running sum/max-error accumulators.
See "Prediction-error metrics" below.
Key functions:
void SpriteStreamInterpolation_Update(const SpriteRenderQuad* realQuads,
uint32_t realQuadCount, uint32_t frameToRender);
/* Called every VBL right after SpriteStreamFrame_Build. Detects whether
* frameToRender is new since last call: if so, shifts the 3-frame cache,
* records the fresh quads + tick gap, and scores the previous held
* interval's last prediction against this real frame (see metrics
* below); if not (held VBL), only advances the tick counter. */
bool SpriteStreamInterpolation_IsHeldVbl(void);
uint32_t SpriteStreamInterpolation_Extrapolate(SpriteRenderQuad* outQuads,
uint32_t outCapacity);
/* Routes each cached quad through, in order:
* - status bar (SPRITE_QUAD_FLAG_STATUS_BAR): always held -- sparse,
* content-driven redraws break the "one real interval apart"
* assumption every other branch relies on.
* - wall tile (0xFFFF0000u objectId range): always held -- see "Wall
* tiles" above for why group extrapolation was reverted.
* - background singleton (0xFFFE0000u): dedicated wide-baseline linear
* extrapolation on sourceOffsetY, with wraparound unwrapping and
* round-to-nearest -- see "The background mosaic" above. Never
* quadratic.
* - scaled-size (SPRITE_QUAD_FLAG_SCALED_SIZE -- zoom-text glyphs, the
* logo): always held -- see "Zoom-text glyphs and the title-screen
* logo" above.
* - everything else (ordinary gameplay objects, stars): matched across
* the 3-frame cache by identityId + type + a displacement-plausibility
* clamp (deliberately not spriteId -- see "objectId reuse" below);
* quadratic extrapolation on a full 3-frame match, linear on a
* 2-frame-only match, hold on a miss.
* spriteId/width/height/flags are always held at the newest cached
* value; only position (or the background's sourceOffsetY) is ever
* extrapolated. Returns 0 if fewer than two real frames are cached yet
* (e.g. right after startup) -- caller falls back to plain real quads. */
Target rate is one tunable, not a design fork: since GPU submission already happens every VBL for free, extrapolating on every held VBL (the current behavior) costs nothing extra over today's baseline — every held tick gets its own distinct extrapolated position rather than repeating the same frozen content. (The original plan described this as a step-count constant to make tunable; in practice the code just always steps every held VBL, and the meaningful on/off choice turned out to be the enable flag above, not a rate knob.)
Prediction-error metrics (not in the original plan)
Added after the fact, at request, to gauge extrapolation quality
independent of any specific bug: Extrapolate() snapshots its own output
every time it's called; when the next real frame lands, Update() matches
that snapshot's quads against the real positions that actually arrived (by
identityId+type, same matching rule as extrapolation itself) and
accumulates |dx|/|dy|/magnitude error. Status bar, wall tile, and
scaled-size quads are excluded from scoring (they're never actually
extrapolated, so including them would just report a trivial zero error and
dilute the mean). Every ~50 real frames (~4s at the ~12.5Hz real cadence),
logs a summary and resets:
SpriteStreamInterpolation: prediction error over N samples -- mean |dx|=.. |dy|=.. px, max |d|=.. px
Purely diagnostic — never affects rendering. A handful of bounded, throttled diagnostic logs used to root-cause the bugs described above (raw per-tick background numbers, logo cache push/invalidate events) were removed once those bugs were confirmed fixed; this metric was kept as the one signal worth monitoring on an ongoing basis.
Per-event misprediction log (added alongside the zero-velocity guard,
to verify it live): the periodic summary above only ever reports two
aggregate numbers, which can't say whether a given max came from one
repeat-offending object or many different ones, or let you correlate it
back to a specific real frame. Every individual scored misprediction over
SPRITE_STREAM_INTERP_MISPREDICT_LOG_THRESHOLD_PX (2px — deliberately low,
to catch modest single-object artifacts like a few-pixel overshoot, not
just the ~80px scale of a systemic reused-identity bug) gets its own line:
SpriteStreamInterpolation: misprediction |d|=.. px -- frame=N heldTicks=.. objectId=0x.. identityId=0x.. type=.. predictedSpriteId=0x.. realSpriteId=0x.. predicted=(x,y) real=(x,y) size=WxH
frame is the real frame that confirmed/disproved the guess; heldTicks
is how many VBLs it was extrapolated across. Bounded only by
SPRITE_STREAM_INTERP_MISPREDICT_LOG_LIMIT (5000), a safety valve against
runaway spam in a badly-broken state — not a normal-use sampling cutoff the
way the earlier 20-event outlier limit this replaced was.
Metric itself needed the same wrap-awareness as the extrapolation
(fixed): live testing after the star Y-wrap extrapolation fix still
showed ~191px "mispredictions" for vertical-scroll stars, alternating
predicted=0/real=191 then predicted=191/real=0 on consecutive real
frames — the extrapolation was actually correct (off by ~1px), but
RecordPredictionError's own errorY = predicted->y - real->y had no
wrap-awareness, so a genuinely tiny error that happened to straddle the
wrap boundary got reported as a false ~191px outlier. Fixed by giving
RecordPredictionError the same SPRITE_QUAD_FLAT_COLOR +
!SPRITE_QUAD_FLAG_PERSPECTIVE_STAR branch the extrapolation path already
has, unwrapping predicted->y relative to real->y before diffing —
mirrors the background branch just above it, which already did this
correctly for sourceOffsetY.
Call site — src/screentrace.c, RENDER_VIEW_SPRITE_STREAM branch (~450-520)
Purely additive, right after the existing SpriteStreamFrame_Build call
and before RenderView_Submit, and gated on emulationActive (not in the
original plan — added after live testing showed sprites continuing to
drift during a pause, F12): while paused, Update()/Extrapolate() are
skipped entirely, freezing tick/history state exactly where it was rather
than treating the whole paused interval as one long held VBL to extrapolate
across.
if (emulationActive)
{
SpriteStreamInterpolation_Update(frame.spriteStream.quads,
frame.spriteStream.quadCount, frameToRender);
if (SpriteStreamInterpolation_IsEnabled() && SpriteStreamInterpolation_IsHeldVbl())
{
static SpriteRenderQuad interpolatedQuads[SPRITE_STREAM_MAX_QUADS];
uint32_t n = SpriteStreamInterpolation_Extrapolate(interpolatedQuads, SPRITE_STREAM_MAX_QUADS);
if (n > 0)
{
frame.spriteStream.quads = interpolatedQuads;
frame.spriteStream.quadCount = n;
}
}
}
Toggle key: I, polled the same unscoped way as the atlas-resolution
switch's R key — see that section's note above.
Status bar gameplay-detection was broadened too, for the same reason as
the logo (not in the original plan): the HUD's on/off gate
(s_gameplayActiveThisFrame) was set only by the wall-tile-draw hook,
which shares the exact blind spot that caused the logo bug — a level
opening in open space before any wall tiles appear. Unlike the logo, this
flag is reset every frame (not a persistent cache), so only the "turning on
late" direction was actually at risk. Fixed by setting it from the same
three gameplay-exclusive hooks now used for the logo (wall-tile draw,
object dispatch, gameplay starfield).
The logo's own persistence bug needed a fourth signal beyond gameplay
detection: the title-screen logo persisted visibly through two
intermediate screens (player-count select, "GET READY PLAYER n") before
disappearing, because neither wall-tile draw, object dispatch, nor the
gameplay starfield ever fire on those text/starfield-only interstitial
screens. Found via ReVa: ZoomTextInterstitial_MainLoop_FUN_00008530 (the
actual driver behind those screens) calls FUN_00007f96 — an unconditional
full 32000-byte framebuffer clear — every frame, which is what physically
erases the logo on real hardware once that screen is reached.
DrawCommandStream_OnZoomTextGlyphInstructionFetch already gates on the
caller's return address to distinguish this screen (0x859a) from the
title screen's own credits text (0x84a8, where the logo should keep
persisting); invalidation was added specifically for the 0x859a case.
Verification plan
- A/B toggle: flip interpolation on/off live during gameplay (
Ikey) and watch the same objects for smoother-vs-steppier motion, no relaunch needed. - No-desync check (most important): play the same save/seed with interpolation ON vs OFF and confirm score, enemy spawn timing, and run outcome are identical. Since this never re-executes game logic, any observed divergence would indicate something unexpectedly leaked game-observable state into the new code.
- Curved-motion check: specifically watch any enemy/projectile pattern known to curve (circle/sine-style movement) with interpolation on, to judge whether quadratic is sufficient or a case for deferred Option C.
- Background scroll check: confirm smooth, jitter-free scrolling in both directions (the player can reverse it by flying backward) and while stopped (ship stuck/killed) — this is what exercises the wide-baseline averaging and wraparound-unwrap fixes.
- Wall-tile check: confirm tiles look identical with interpolation on vs off (held either way) — no seam, no snap-back; wall tiles are a "nothing should visibly change" case now, not a smoothing win.
- Zoom/logo check: watch the logo zoom in/out and zoom-text screens ("GET READY PLAYER n" etc.) with interpolation on — confirm no settle-then-jump motion at the end of an animation, and that the logo correctly disappears once the player-select/get-ready screens are reached, not lingering into gameplay.
- Reused-
objectIdguard check: force heavy spawn/death churn (e.g. a dense enemy wave) and confirm the plausibility clamp suppresses any teleport-streaks between unrelated objects. - Pause check: pause (F12) mid-motion and confirm nothing continues to drift; resume and confirm motion continues smoothly rather than jumping.
- Prediction-error metric: watch the periodic log during normal play; a low, stable mean with occasional higher max values (self-correcting, bounded to one real interval) is expected — a persistently high or climbing mean would indicate a new problem.
- Dual-atlas switch: confirm switching resolution mid-gameplay is instant (no load hitch, since both are preloaded) and doesn't perturb interpolation state.
Explicit non-goals for this phase
- No re-execution/simulation of game update logic (Option C) — deferred until quadratic extrapolation ships and specific object classes are shown to still need it.
- No delayed/buffered true interpolation between two known real frames (rejected earlier: ~80ms added input latency in a fast-paced shooter, for a correctness gain the self-correcting bounded error mostly already delivers).
- No object spawn/despawn hook for authoritative
objectId-reuse detection yet — heuristic clamp only, per above. - No modeling of the zoom-text/logo shared scale-to-size/position relationship (would still have the same end-of-animation overshoot problem, just for a different quantity) — held instead, see above.
- ~~No row-band-aware wall-tile extrapolation~~ — superseded, later session: the underlying premise ("only the top band genuinely moves") turned out to be wrong — the whole tile grid scrolls as one at constant speed; the row-band split is a CPU-clipping artifact of the original 68000 code, not evidence of different per-band motion. All wall tiles (not just the top band) are now extrapolated by one shared, wrap-aware delta. See "Wall tiles" below for the corrected fix.
- No changes to the WASM/web renderer at first — a later session found the
web-side
SpriteRenderQuadwire parser (web/src/spriteQuad.ts) had already drifted out of sync with the C struct (missingidentityId, silently dropping every sprite-stream frame) independent of anything in this plan, fixed alongside wideningx/ytofloat(see "Coordinate precision" below). The dual-atlas/sourceScaleYUV-math gap described above is still unaddressed on the web side. - No attempt to fix atlas border-bleed in code — if present, it's an external-asset regeneration problem.
- No crossfade/transition effect between atlas resolutions yet — the dual-load architecture is chosen to make this cheap later, but it isn't built now.
Coordinate precision (later session)
Extrapolated positions were computed in float from the start (see above),
but SpriteRenderQuad.x/y (renderFrame.h) was int32_t, and the
gameplay-object/vertical-star branches rounded to the nearest native pixel
(lroundf) before storing — throwing away the fractional part before it
ever reached a renderer, capping visible motion smoothness at one of 320
horizontal positions regardless of atlas or backbuffer resolution. Both
renderers' NDC math (StreamVertex_AppendQuad in sdlGpuRenderView.c, and
its web port in spriteStreamView.ts) already divided by the logical
320×200 output size as plain floats, so this was never the bottleneck.
Fixed by widening SpriteRenderQuad.x/y to float and removing the
lroundf-to-int32_t casts (the background's sourceOffsetY stays
integer — it indexes a discrete texel row in the mosaic asset, a different
kind of quantity). DrawMaskedSpriteEntry.x/y (drawCommandStream.h,
capture-side) stays int16_t — genuinely correct, since the original
68000 game logic never computes a sub-pixel position; only the
extrapolation math synthesizes fractional values, one layer downstream.
This surfaced a pre-existing, unrelated bug on the web side while
confirming the fix would carry through there too: web/src/spriteQuad.ts
hand-mirrors SpriteRenderQuad's wire layout by byte offset, and had
already drifted out of sync with the C struct — missing the identityId
field entirely (added in an earlier session for objectIdentity.c, see
above), reading 40 bytes where the struct is actually 44. Since
webRenderBackend.c passes the real sizeof(SpriteRenderQuad) as the wire
stride, this mismatch should have been tripping spriteQuad.ts's own
struct-layout-drift guard and silently dropping every sprite-stream frame
in the web build already, independent of anything in this coordinate
change. Fixed in the same pass: corrected every field offset, added the
missing objectId/identityId fields and SpriteQuadFlags bits, and
switched x/y to getFloat32.
Critical files
hatari_dotnet/Program.cs,hatari_dotnet/AtlasIndexWriter.cs,hatari_dotnet/SpriteAtlasExporter.cs,hatari_dotnet/PngReader.cs— extended single-pass exportsrc/spriteAtlas.c/src/includes/spriteAtlas.h— instantiable refactor for dual-resolution support,SpriteAtlas_GetActiveScale()src/sdlGpuRenderView.c— dual textures, active-slot bind,sourceScaleYfixsrc/main.c— dual atlas load at startupsrc/screentrace.c— interpolation call site (pause-gated), toggle keyssrc/spriteStreamInterpolation.c/.h— the interpolation module itself; also owns the prediction-error metricssrc/spriteStreamFrame.c— mapsDrawMaskedSpriteEntry.isScaledSizetoSpriteRenderQuad'sSPRITE_QUAD_FLAG_SCALED_SIZE(excluding the background, see "Zoom-text glyphs" above)src/includes/renderFrame.h—SPRITE_QUAD_FLAG_SCALED_SIZE(new)src/includes/drawCommandStream.h—DrawMaskedSpriteEntry.isScaledSize(new);FINE_SCROLL_ADDRmoved back to being private todrawCommandStream.c(only the now-reverted wall-tile group-delta approach needed it public)src/drawCommandStream.c—DrawCommandStream_PushMaskedSprite'sdestW/destHresolution-independence fix (Task 1);isScaledSizeset per push function; logo-cache invalidation from four independent "not the title screen" signals (wall-tile draw, object dispatch, gameplay starfield, the zoom-text interstitial's own0x859acaller);s_gameplayActiveThisFramebroadened the same way for the status barsrc/CMakeLists.txt— new HD asset filenames, new source file