Xenon 2

How it works · write-up

AVI recording architecture

xenondoc/AVIRECORDING.MD · 35 KB · updated 2026-10-06

This document describes the recording code added around the Xenon 2 Sprite Stream renderer. It covers the two continuous AVI streams, their control protocol and replay integration, and the related C/F rolling PNG capture. The continuous AVI streams share one bounded-queue/worker implementation, while Sprite AVI and C/F capture also share GPU readback. Their writer state, queues, lifetimes, and output files remain independent.

The desktop SDL3 build can write the same two streams as H.264/AAC MP4 instead, through an external FFmpeg (--record-format mp4). Only the container and encoder change: capture, queues, GPU readback, audio pairing, VBL sidecars and the control protocol are the ones described here. PNG AVI remains the default. See EXTERNAL_FFMPEG_RECORDING.MD for the encoder options, process lifecycle and validation.

The implementation has three capture products:

Product Source Lifetime Output
Authentic display video Hatari's normal SDL display surface Continuous *-original.avi (or .mp4) and its .vbl sidecar
Sprite Stream video Final pixels rendered by the SDL GPU Sprite Stream view Continuous *-sprite.avi (or .mp4) and its .vbl sidecar
Comparison capture Last 16 authentic and last 16 Sprite Stream frames Rolling, flushed on demand Independent PNG files under framecapture/

Both continuous streams carry the game's sound, mono on the ST. The Sprite Stream image is as large as the renderer's drawable: 640x400 at the default --xenon-sprite-scale 2, 1280x800 at 4. The authentic stream is 640x400 in paired recordings.

The event recording (*.x2events) is a fourth, independent stream. It contains game state and decisions rather than pixels. The .avi.vbl files provide the clock mapping that lets the Python replay UI display the corresponding AVI frame beside an event frame.

High-level architecture

flowchart LR
    subgraph Control[Control and lifecycle]
        RR[record_run.py]
        AP[autoplay.py]
        XC[xenon_client.py]
        CT[xenonControl.c]
        HK[Hatari hotkeys / CLI]
        RR -->|human mode| XC
        RR -->|autopilot mode| AP
        AP --> XC
        XC -->|START_AVI / STOP_AVI| CT
    end

    subgraph Authentic[Authentic display path]
        VBL[video.c VBL handler]
        SDL[SDL display surface]
        SND[sound.c audio mix ring]
        AQ[24-job ordered AVI queue]
        AW[Authentic AVI worker]
        AO[AVI writer state: AviParams]
        VBL -->|Avi_RecordVideoStream| AQ
        SDL -->|detached pixel copy| AQ
        SND -->|Avi_RecordAudioStream| AQ
        AQ --> AW --> AO
    end

    subgraph Sprite[Sprite Stream path]
        GPU[SDL GPU renderer]
        TEX[Off-screen color texture]
        RB[4-slot GPU readback ring]
        Q[12-slot detached-frame queue]
        EW[One Sprite AVI worker]
        SO[AVI writer state: SpriteAviParams]
        GPU --> TEX --> RB --> Q --> EW --> SO
    end

    subgraph Diagnostic[C/F comparison capture]
        AR[16-slot authentic ring]
        GR[16-slot GPU ring]
        PQ[32-job PNG queue]
        PW[1 to 4 PNG workers]
        SDL --> AR
        RB --> GR
        AR -->|F: detach frames| PQ
        GR -->|F: detach frames| PQ
        PQ --> PW
    end

    CT --> AO
    CT -->|arm / flush / stop| GPU
    HK --> AO
    HK --> GPU

The continuous streams have separate RECORD_AVI_PARAMS, AVI_FILE_HEADER, and AviAsyncQueue instances. They can therefore record simultaneously without sharing file positions, frame indexes, counters, queues, workers, or FILE* handles. Both instances use the same generic job/worker implementation and share only the selected video codec and global PNG compression level.

Starting a paired recording

record_run.py --avi derives both AVI names from the event-log name. For example, demo.x2events produces:

demo.x2events
demo-original.avi
demo-original.avi.vbl
demo-sprite.avi
demo-sprite.avi.vbl
demo.sav                    # checkpoint written by record_run.py

In human mode, record_run.py sends the AVI request itself. In autopilot mode it hands ownership of Hatari's single automation connection to autoplay.py, which starts the pair using --avi-prefix. This prevents two clients from competing for the same control connection.

The two messages were added in protocol version 10 and are unchanged in the current version 18. A recording started in MP4 mode uses the same messages with .mp4 paths; the PNG level then has no effect.

START_AVI (type 20):
    u8    PNG compression level, 0..9
    u16be original path byte length
    u8[]  original UTF-8 path
    u16be Sprite Stream path byte length
    u8[]  Sprite Stream UTF-8 path

STOP_AVI (type 21):
    empty payload

xenonControl.c handles START_AVI in this order:

  1. Validate the compression level and both paths.
  2. Reject the request if either recorder is already active or armed, or if Hatari was started in MP4 mode and either path does not end in .mp4.
  3. Store the two paths and compression level in Hatari configuration.
  4. Start the authentic display AVI and its asynchronous worker immediately with Avi_StartRecording_WithConfigCropSize(true, 640, 400).
  5. Arm the Sprite Stream recorder with Avi_SpriteRequestStartWithConfig().
  6. If Sprite Stream arming fails, stop the already-started authentic recorder so the pair is not left half-active.
  7. Reply to the client.

For this paired diagnostic command, the authentic stream always excludes Hatari's status bar and is encoded at 640x400, even if the main Hatari window has been resized. This makes its frame geometry identical to the Xenon 2 Sprite Stream recording without changing the user's global screenshot/ordinary-AVI crop preference. The existing GUI, CLI and hotkey authentic-AVI paths still use ConfigureParams.Screen.bCrop and their normal output dimensions.

The Sprite Stream AVI is only armed at step 5. Its size is not known until the GPU produces the first frame, so Avi_SpriteStartForFrame() creates its file and AVI header lazily from that frame's dimensions.

sequenceDiagram
    participant P as Python controller
    participant C as xenonControl.c
    participant A as Authentic AVI
    participant AW as Authentic worker
    participant S as Sprite AVI state
    participant W as Sprite AVI worker
    participant G as GPU renderer

    P->>C: START_AVI(level, originalPath, spritePath)
    C->>A: Avi_StartRecording_WithConfigCropSize(true, 640, 400)
    A->>A: Open AVI and .avi.vbl; write initial headers
    A->>AW: Create shared queue instance and worker
    Note over A,AW: active state 0 -> 2 (recording)
    AW-->>AW: Wait on queue-not-empty condition
    A-->>C: started
    C->>S: Avi_SpriteRequestStartWithConfig()
    S->>W: Create worker and bounded queue
    Note over S,W: active state 0 -> 1 (armed)
    W-->>W: Wait on queue-not-empty condition
    S-->>C: armed
    C-->>P: success

    G->>S: First completed GPU frame
    S->>S: Avi_SpriteStartForFrame(width, height)
    S->>S: Open AVI and .avi.vbl; write initial headers
    Note over S,W: active state 1 -> 2 (recording)
    S->>W: Enqueue detached pixel copy

Other entry points are deliberately narrower:

  • Hatari's existing GUI/CLI authentic-AVI controls operate on the authentic stream.
  • --sprite-avirecord arms the Sprite Stream recorder at startup.
  • The V key toggles only the Sprite Stream AVI.
  • C and F control rolling comparison capture, not either continuous AVI.

Authentic display AVI

The authentic stream preserves Hatari's existing capture points and AVI format, but moves compression and file I/O to an AviAsyncQueue worker:

  1. video.c reaches the VBL handler.
  2. If Avi_AreWeRecording() is true, it calls Avi_RecordVideoStream().
  3. The producer locks the current SDL surface selected by Screen_GetDimension() / Avi_SetSurface(), copies it into a detached video job, unlocks it, and enqueues the job.
  4. Sound_Update_VBL() later reaches sound.c, which passes the VBL's newly generated samples from AudioMixBuffer to Avi_RecordAudioStream(). It does so before resynchronizing the playback ring after a pause, fast-forward or checkpoint; resetting first used to discard one VBL of sound each time (corrected 2026-10-05).
  5. The audio producer linearizes those samples from the mixer ring into a detached audio job and enqueues it after the corresponding video job.
  6. One authentic AVI worker consumes the FIFO, compresses/writes video, writes real PCM, and updates the AVI and VBL indexes in the original order.

The queue has 24 jobs, normally twelve video/audio pairs. The emulation thread pays for bounded pixel and sample copies, but does not wait for PNG compression or file I/O unless all 24 slots are occupied. A full queue deliberately applies backpressure rather than losing frames or consuming unbounded memory.

Authentic and Sprite AVI instantiate the same AviAsyncQueue, AviAsyncJob, worker loop, allocation, backpressure, failure, synchronous-fallback, and drain logic. Their policies differ in queue capacity and in how audio arrives: authentic AVI submits separate PCM jobs in VBL order, while each Sprite video job carries the PCM of its own VBL (see Sprite audio). Sprite AVI is also opened lazily and rejects dimension changes; authentic AVI is opened immediately and preserves Hatari's fixed encoded dimensions if the SDL source changes.

The authentic AVI has real 16-bit PCM audio: one channel on the ST and Mega ST, whose mixer produces identical left and right samples, and two on later machines. The channel count is fixed when the recording starts. For ordinary recording, ConfigureParams.Screen.bCrop determines whether Statusbar_GetHeight() is removed from the bottom. Paired protocol recording overrides only that invocation: it removes the status bar and uses the AVI writer's existing resampling path to normalize the cropped Atari image to 640x400.

sequenceDiagram
    participant E as Emulation / VBL thread
    participant V as video.c
    participant D as SDL display surface
    participant S as sound.c
    participant Q as 24-job shared queue instance
    participant W as Authentic AVI worker
    participant A as AVI and .vbl files

    E->>V: VBL interrupt
    V->>D: Avi_RecordVideoStream: lock and copy pixels
    D-->>Q: Detached VIDEO job with nVBLs tag
    Q-->>W: Signal queue-not-empty
    Note over V,Q: VBL waits only if all 24 jobs are occupied
    V->>S: Sound_Update_VBL()
    S->>Q: Copy new mixer samples into AUDIO job
    Q-->>W: Signal queue-not-empty
    Note over S,Q: Audio update waits only if all 24 jobs are occupied
    loop Ordered worker consumption
        W->>Q: Dequeue next VIDEO or AUDIO job
        alt VIDEO
            W->>W: PNG encode or BGR conversion
            W->>A: Write 00dc/00db chunk, index and VBL tag
        else AUDIO
            W->>A: Write real PCM 01wb chunk and index
        end
        W->>Q: Release slot; signal queue-not-full
    end

Sprite Stream AVI

The Sprite Stream path is asynchronous because its pixels originate in GPU memory and PNG compression is expensive. It consists of two bounded pipelines.

Stage 1: GPU render and readback

While Sprite AVI or C capture needs frames, SdlGpuRenderView_Submit() renders the Sprite Stream view into an off-screen color texture. If a swapchain texture is available, that result is also blitted to the window. A swapchain miss does not lose the recording frame: the renderer can continue off-screen and omit only the visible blit.

The off-screen texture is downloaded through four SpriteCaptureReadback slots. Each slot owns:

SDL_GPUTransferBuffer *transfer
SDL_GPUFence          *fence
width, height
frameTag              # nVBLs at submission
deliverFrameCapture
deliverSpriteAvi
pending

At the start of a later Sprite Stream submit, completed fences are polled in submission order. A completed transfer buffer is mapped once and fanned out to either or both consumers:

  • FrameCapture_PushGpuFrame() for the rolling comparison ring;
  • Avi_SpriteSubmitVideoFrame() for the continuous Sprite AVI.

The renderer normally does not wait for each fence. If all four readback slots are still occupied, it waits for the oldest one before reusing a slot. This bounds GPU memory and latency rather than allocating an unlimited backlog.

Stage 2: detached CPU frames and AVI worker

Avi_SpriteSubmitVideoFrame() uses the generic submit path to copy each mapped GPU image into a reusable AviAsyncJob in the Sprite instance's 12-slot FIFO:

uint8_t  *data         # worker-owned detached pixels or PCM
size_t    capacityBytes
type                    # VIDEO or AUDIO
int       width, height
int       pitch
int       sampleLength
uint32_t  frameTag     # emulator VBL

Copying is required because the GPU transfer buffer is unmapped and recycled as soon as delivery returns. The AVI worker must never retain a pointer into mapped GPU memory.

One SDL worker consumes the FIFO in order. It writes the PNG video chunk, writes the PCM chunk of the same VBL, updates the OpenDML indexes, and appends the VBL tag. Only one encoder worker is used per AVI because all chunks and indexes in that file must remain ordered; multiple workers would require an additional ordered commit stage. The authentic recorder owns a second instance of the same worker and queue implementation and can encode on another core at the same time.

Sprite audio

The Sprite Stream AVI carries the game's sound, the same samples as the authentic AVI. (Before 2026-10-05 it wrote silence.) Its frames arrive from the GPU several VBLs after the emulator generated their sound, so the producer keeps the last eight VBLs of mixed sound (SPRITE_AVI_AUDIO_HISTORY_SIZE), tagged with their VBL, while sprite recording is armed or active:

  • Sound_Update_VBL() stores each VBL's samples in that history (Avi_SpriteRecordAudio()).
  • Avi_SpriteQueueVideo() looks up the sound for the frame's VBL and submits the frame and its PCM as one job, which keeps the one-audio-chunk-per-video-chunk pairing of the OpenDML index.
  • A download that completes within its own VBL, before that VBL's sound exists, is held in s_spritePendingVideo and submitted when the sound arrives.
  • If a frame's sound is no longer in the history, the recording stops with an error rather than writing a frame without its sound.

The history is allocated only while sprite recording is armed or active and is cleared when recording restarts. Comparing the two AVIs' sound must allow for the one-VBL offset between the sprite capture tag and the authentic recorder's tag. Validation: SPRITE_AVI_RESOLUTION_AUDIO_20261005.MD.

Duplicate frameTag values are discarded. This prevents paused redraws or multiple renderer submissions for one VBL from creating repeated AVI frames.

sequenceDiagram
    participant E as Emulation / render thread
    participant G as SDL GPU
    participant R as 4-slot readback ring
    participant Q as 12-slot AVI queue
    participant W as Sprite AVI worker
    participant F as Sprite AVI files

    loop Every rendered VBL while armed/recording
        E->>R: Poll oldest completed fences (non-blocking)
        opt Oldest readback is complete
            R->>G: Map transfer buffer
            G-->>R: RGBA pixels
            R->>Q: Copy frame and enqueue
            R->>G: Unmap and recycle readback slot
            Q-->>W: Signal queue-not-empty
        end
        E->>G: Render to off-screen texture
        E->>G: Queue texture download
        G-->>R: Fence + transfer slot
        opt All 4 GPU slots are occupied
            E->>R: Wait for oldest fence
            Note over E,R: Bounded GPU backpressure
        end
    end

    loop Worker lifetime
        W->>Q: Wait while queue is empty
        Q-->>W: Next FIFO frame
        W->>W: Compress PNG
        W->>F: Write video + its VBL's PCM + indexes + VBL tag
        W->>Q: Release queue slot; signal queue-not-full
    end

    opt All 12 CPU slots are occupied
        E->>Q: Wait for queue-not-full
        Note over E,Q: Encoder/disk backpressure reaches emulation here
    end

Sprite recorder state

The atomic s_spriteAviQueue.active is a small state machine rather than a Boolean:

stateDiagram-v2
    [*] --> Inactive: 0
    Inactive --> Armed: Avi_SpriteRequestStartWithConfig / 1
    Armed --> Recording: first delivered GPU frame / 2
    Armed --> Inactive: stop before first frame
    Recording --> Inactive: stop or writer failure

Avi_SpriteNeedsFrames() is true in both Armed and Recording, so the GPU path starts producing readbacks before the AVI file exists. Avi_SpriteAreWeRecording() is true only in Recording.

If an SDL AVI worker cannot be created, the shared queue code retains a synchronous fallback. In Sprite fallback mode Avi_SpriteSubmitVideoFrame() compresses and writes immediately on the render thread; authentic fallback similarly writes from its VBL/audio calls after first detaching the source data. A resolution change after the Sprite AVI has started is treated as an error and stops that stream; one AVI stream has fixed dimensions.

AVI container and sidecar structures

Both continuous streams use the same generic writer in avi_record.c:

classDiagram
    class RECORD_AVI_PARAMS {
        VideoCodec
        VideoCodecCompressionLevel
        surface_pixels, surface_w, surface_h, surface_pitch
        CropLeft, CropRight, CropTop, CropBottom
        Fps, Fps_scale
        AudioCodec, AudioFreq, AudioChannels
        Width, Height, BitCount
        FileOut, TagFileOut
        External (FFmpeg recorder in MP4 mode)
        TotalVideoFrames, TotalAudioFrames, TotalAudioSamples
        RIFF and MOVI offsets
        OpenDML super-index offsets
        pAviFrameIndex
        IsRecording, LockSurface
    }

    class RECORD_AVI_FRAME_INDEX {
        VideoFrame_Pos
        VideoFrame_Length
        AudioFrame_Pos
        AudioFrame_Length
    }

    class AviAsyncJob {
        data
        capacityBytes
        type
        width, height
        pitch
        sampleLength
        frameTag
    }

    class AviAsyncQueue {
        jobs, capacity
        mutex, notEmpty, notFull
        thread, active
        readIndex, writeIndex, count
        workerStop, workerFailed
        synchronousFallback
        audioAfterVideo
        fixedVideoDimensions
        highWater, producerWaits
        params, header, label
    }

    class SpriteCaptureReadback {
        transfer
        fence
        width, height
        frameTag
        deliverFrameCapture
        deliverSpriteAvi
        pending
    }

    RECORD_AVI_PARAMS "1" o-- "many" RECORD_AVI_FRAME_INDEX
    AviAsyncQueue "1" o-- "many" AviAsyncJob
    AviAsyncQueue --> RECORD_AVI_PARAMS: owns writer policy
    AviAsyncJob --> RECORD_AVI_PARAMS: worker supplies pixels or PCM
    SpriteCaptureReadback --> AviAsyncJob: detached video copy

Video is stored as either uncompressed 24-bit BGR chunks (00db) or lossless PNG chunks (00dc). Audio is 16-bit PCM (01wb), mono on the ST and Mega ST and stereo on later machines; the header's channel count, byte rate, block alignment and sample size follow it, and the audio super-index counts durations in sample frames. The default is PNG with compression level 6. The level changes compression time and file size, not image fidelity.

Both AVI workers encode detached pixels with LockSurface=false. PNG output therefore uses RGB rather than consulting the live emulator palette: this avoids racing ConvertPalette and avoids the additional full-frame palette-eligibility scan. The authentic producer still locks the SDL surface while making its detached copy.

The file is OpenDML AVI. A movi RIFF chunk is closed at approximately 1 GiB, and additional AVIX chunks are created. Per-stream standard indexes and OpenDML super-indexes allow large recordings. The current limits are 256 movi chunks and therefore approximately 256 GiB.

Every successful video write also appends one decimal VBL number to <avi-name>.vbl:

# X2AVI-VBL 1
385321
385322
385323
...

The sidecar is intentionally separate from the AVI container. Ordinary AVI players ignore it, while Xenon tools can synchronize by emulator time without inventing a custom AVI chunk.

Stopping, draining, and file correctness

An AVI is not complete merely because no more frames will be submitted. Pending GPU copies must be delivered, queued frames must be encoded, and the AVI indexes and size fields must be finalized.

STOP_AVI therefore performs a strict drain:

sequenceDiagram
    participant P as Python controller
    participant C as xenonControl.c
    participant G as GPU readback ring
    participant SQ as Sprite AVI queue/worker
    participant S as Sprite AVI writer
    participant AQ as Authentic AVI queue/worker
    participant A as Authentic AVI writer

    P->>C: STOP_AVI
    C->>G: SdlGpuRenderView_FlushSpriteCapture()
    loop Every pending GPU slot
        G->>G: Wait for fence
        G->>SQ: Map, copy and enqueue frame
    end
    C->>SQ: Avi_SpriteStopRecording()
    SQ->>SQ: Set active=0; wake worker
    loop Until queue is empty
        SQ->>S: Encode and write next frame
    end
    C->>SQ: SDL_WaitThread()
    Note over C,SQ: STOP waits for Sprite worker drain and exit
    C->>S: Close MOVI, write indexes/header, close AVI and .vbl
    C->>AQ: Avi_StopRecording()
    AQ->>AQ: Set active=0; wake worker
    loop Until authentic VIDEO/AUDIO queue is empty
        AQ->>A: Encode/write next ordered job
    end
    C->>AQ: SDL_WaitThread()
    Note over C,AQ: STOP waits for authentic worker drain and exit
    C->>A: Close MOVI, write indexes/header, close AVI and .vbl
    C-->>P: success or failure

The order matters: setting the Sprite recorder inactive before flushing GPU readbacks would discard already-rendered frames. For either AVI, finalizing the file before joining its worker would race chunk writes against header/index generation.

Hatari shutdown follows the same principle in Main_UnInit():

  1. stop/drain/finalize the authentic AVI;
  2. flush Sprite GPU readbacks;
  3. stop/drain/finalize the Sprite AVI;
  4. drain and join comparison-capture PNG workers;
  5. tear down the remaining emulator subsystems.

record_run.py additionally pauses emulation before sending STOP_AVI, then stops the event log, saves the emulator checkpoint, restores real-time speed if needed, and releases injected input. Its cleanup path performs these operations even after Ctrl+C or an autopilot failure.

C/F rolling comparison capture

The rolling capture exists for short, pixel-level comparisons and does not write AVI:

  • C toggles capture.
  • The authentic SDL path copies frames into a 16-slot FrameCaptureRing.
  • The Sprite renderer uses the same four-slot GPU readback ring as Sprite AVI and copies completed frames into a second 16-slot FrameCaptureRing.
  • F first drains pending GPU readbacks, then snapshots both rings into heap-owned jobs.
  • One to four SDL worker threads encode the independent PNG files. The worker count is half the detected logical CPU count, clamped to 1..4.

The asynchronous PNG queue has 32 slots, exactly enough for one full 16+16 comparison batch. Because every output is an independent file, jobs may be compressed in parallel without an ordered commit stage.

sequenceDiagram
    participant U as User
    participant H as screentrace.c
    participant G as GPU readback ring
    participant A as Authentic 16-frame ring
    participant S as Sprite 16-frame ring
    participant Q as 32-job PNG queue
    participant W as 1..4 PNG workers

    U->>H: Press C
    H->>H: FrameCapture_SetEnabled(true)
    par On authentic display updates
        H->>A: Copy current SDL pixels
    and On completed Sprite GPU readbacks
        G->>S: Copy mapped GPU pixels
    end
    U->>H: Press F
    H->>G: Flush pending readbacks
    Note over H,G: F waits for outstanding GPU fences
    H->>A: Detach valid slots into jobs
    H->>S: Detach valid slots into jobs
    H->>Q: Enqueue up to 32 jobs
    H-->>U: Return after queuing
    par Independent encodes
        W->>Q: Dequeue job
        W->>W: Encode and write PNG
    end

The rings tag frames with nVBLs, the draw-command game frame, a per-game-frame sequence, and host time. The authentic display is tagged with a one-tick delayed draw-command frame because Hatari's CRT/double-buffer display lags the Sprite Stream command frame by one capture call. The PNG filenames group corresponding authentic and GPU images by game-frame/sequence while retaining the VBL tag.

On desktop, libpng writes the files directly. The WebAssembly path exposes the authentic slots and a one-shot flush request to JavaScript because a browser build cannot use the desktop output directory in the same way.

Where each thread can wait

Caller/thread Wait condition Why it waits Normal or exceptional?
Emulation/VBL thread Authentic queue not full Authentic worker has fallen 24 ordered video/audio jobs behind Backpressure only
Sound/VBL update Authentic queue not full The preceding video and other jobs occupy all 24 slots Backpressure only
Authentic AVI worker Authentic queue not empty No video or PCM job is ready Normal idle state
Sprite render thread Oldest GPU fence All four readback slots are occupied, or an explicit flush is running Backpressure / flush
Sprite render thread Sprite queue not full Encoder has fallen 12 detached frames behind Backpressure
Sprite AVI worker Sprite queue not empty No frame is ready Normal idle state
STOP_AVI caller All GPU fences Preserve every already-submitted frame Normal stop
STOP_AVI caller Sprite and authentic worker termination Drain both FIFOs before finalizing indexes Normal stop
F caller All pending Sprite GPU fences Include frames submitted before the key press Normal flush
F caller PNG queue not full Repeated flushes outpace up to four PNG workers Backpressure only
PNG workers PNG queue not empty No comparison image is waiting Normal idle state
Hatari shutdown PNG worker termination Finish detached comparison images safely Normal shutdown

No capture resources or worker thread are created merely because Hatari is running. The authentic queue/worker exists only while authentic AVI is recording. With Sprite AVI and C capture both disabled, the Sprite renderer uses the swapchain directly; it does not create the off-screen capture texture, GPU transfer ring, Sprite AVI queue, or PNG worker pool. The disabled-path cost is limited to state checks and an empty pending-ring check.

Replay synchronization

xenon_tools/hatari_avi.py is a deliberately small random-access reader for Hatari AVI files. It scans all RIFF AVI and RIFF AVIX containers, indexes 00db/00dc video chunks, and decodes frames on demand with Pillow. An LRU cache avoids repeatedly decoding recently viewed frames.

When a .avi.vbl sidecar is present, frame_index_for_vbl() uses binary search and chooses the AVI frame whose recorded VBL is closest to the selected event frame's VBL. Sprite Stream lookup uses the event VBL directly. The authentic SDL display reaches the same completed game image two VBLs later through Hatari's display/page-flip pipeline, so replay selects authentic VBL N+2 for event VBL N. The sidecars retain physical capture times; the correction belongs in replay rather than rewriting recorded evidence. Original and Sprite panes are synchronized independently, so a missing frame in one stream does not shift every later comparison. If no sidecar exists, replay falls back to proportional position in the recording.

sequenceDiagram
    participant U as Replay UI
    participant E as .x2events model
    participant O as Original HatariAvi
    participant S as Sprite HatariAvi

    U->>E: Seek to event record
    E-->>U: Reconstructed state and VBL N
    par Original pane
        U->>O: frame_index_for_vbl(N)
        O-->>U: Nearest indexed video frame
        U->>O: read(index)
    and Sprite pane
        U->>S: frame_index_for_vbl(N)
        S-->>U: Nearest indexed video frame
        U->>S: read(index)
    end
    U->>U: Display event diagnostics and optional AVI panes

By default replay_ui.py discovers <recording-stem>-original.avi and <recording-stem>-sprite.avi. --original-avi and --sprite-avi override those paths; --no-auto-avi disables discovery.

Dependencies and ownership rules

The C implementation depends on:

  • Hatari timing (nVBLs, refresh-rate calculations), display surface, status bar, audio mixer, configuration, and lifecycle;
  • SDL mutexes, conditions, atomic integers, worker threads, SDL GPU transfer buffers, and GPU fences;
  • libpng for PNG video frames and desktop comparison PNGs;
  • standard C file I/O and the existing Hatari/OpenDML AVI writer.

The Python control and replay tools depend on the Xenon socket protocol. AVI display in the replay UI additionally requires Pillow, but recording itself has no Python image-processing dependency.

The ownership rules that keep the asynchronous path safe are:

  1. The authentic producer copies the live SDL surface while holding Screen_Lock(); the worker never retains the SDL pointer.
  2. Authentic PCM is linearized from AudioMixBuffer before its ring positions can be reused.
  3. A GPU transfer slot owns its mapped pixels only until SdlGpuRenderView_DeliverCaptureReadback() returns.
  4. Avi_SpriteSubmitVideoFrame() copies those pixels before the transfer slot is unmapped.
  5. An AviAsyncJob remains owned by its queue/worker until the worker has written it and advances the read index.
  6. During steady state, only each AVI's own worker writes that AVI's chunks and per-frame indexes.
  7. A stop caller does not finalize either AVI until its worker has joined.
  8. Comparison PNG jobs contain their own pixel copies; the rolling rings may immediately continue receiving new frames after F returns.

Diagnostics and failure behavior

The asynchronous implementation reports these pressure measurements when it stops:

  • SpriteStream capture: ... GPU readbacks, ... forced ring waits identifies GPU/readback pressure.
  • Atari display AVI async queue: high-water .../24, producer waits ... identifies authentic PNG/disk pressure.
  • Sprite Stream AVI async queue: high-water .../12, producer waits ... identifies PNG encoder or disk pressure.
  • FrameCapture: PNG queue high-water .../32, producer waits ... identifies repeated F flush pressure.

A high-water mark alone is harmless. Producer waits mean the bounded consumer could not keep up and emulation was throttled to preserve frames.

Writer errors set the affected queue inactive, stop its worker, and discard frames already queued after the failure because the output can no longer be extended reliably. In MP4 mode an encoder that exits or stops accepting data for 30 seconds fails its stream the same way; STOP_AVI then reports the failure, and the reason is in the output's .ffmpeg.log. Failure to create either AVI worker selects the shared synchronous fallback; failure to create the comparison PNG workers makes that F batch synchronous. A missing .avi.vbl sidecar does not invalidate the AVI, but replay loses exact VBL synchronization.

Performance verification

The shared authentic queue was compared against the immediately preceding build, which already had asynchronous Sprite AVI but still wrote authentic AVI synchronously. Both builds used the same Xenon 2 snapshot, both AVI streams, PNG compression level 6, and the same machine.

Test Synchronous authentic AVI Asynchronous authentic AVI Improvement
125 canonical autoplay frames, graceful protocol stop 14.13 s 9.50 s 32.8%
Same 125-frame autoplay, AVI disabled 9.46 s 9.41 s No measurable regression
500 VBL unthrottled throughput run 16.05 s 8.84 s 44.9%
500 VBL normal-speed run 20.82 s 20.84 s Timing-capped; both kept up

In the autoplay test, dual AVI added about 49% to the preceding synchronous build's no-AVI time, but only about 1% to the new build's no-AVI time. The 0.6% old/new difference with AVI disabled is normal run-to-run noise and confirms that the inactive queues add no measurable overhead.

The gracefully stopped asynchronous run produced 501 authentic frames and 500 Sprite frames, with the same number of VBL sidecar entries for each stream. First, middle, and last PNG frames from both AVIs decoded successfully. The authentic queue reached only 4/24 jobs and recorded zero producer waits; the Sprite queue also reached 4/12 with zero waits. These measurements show that compression and I/O stayed off the emulation path throughout that run.

At the time of this measurement the unthrottled --run-vbls exit called exit(0) without draining or finalizing either AVI, so file integrity was validated on the autoplay/protocol run, which calls STOP_AVI and waits for both workers. Since 2026-10-05 --run-vbls requests a normal shutdown, which drains the GPU readbacks and encoder workers and finalizes both files.

Principal implementation files

File Responsibility
src/avi_record.c Generic AVI/OpenDML writer, shared AviAsyncQueue implementation, authentic/Sprite policies, Sprite audio history and pairing, VBL sidecars, hand-off to FFmpeg in MP4 mode
src/includes/avi_record.h Public authentic and Sprite recorder interface
src/record_ffmpeg.c MP4 mode: the FFmpeg child process and the streaming Matroska feed into it
src/video.c Authentic video call once per VBL
src/sound.c Authentic PCM audio submission
src/sdlGpuRenderView.c Sprite off-screen render, four-slot GPU readback ring, fan-out and flush
src/frameCapture.c Two rolling rings and asynchronous multi-worker PNG flush
src/includes/frameCapture.h Rolling capture contract and WebAssembly accessors
src/sdl/screen.c Authentic rolling-capture source
src/screentrace.c C, F, and V hotkeys
src/xenonControl.c Paired AVI control protocol and ordered stop
src/main.c Shutdown drain/finalization
xenon_tools/xenon_client.py Protocol encoding and start_avi / stop_avi client calls
xenon_tools/record_run.py Human/autopilot recording orchestration and output naming
xenon_tools/hatari_avi.py Random-access Hatari AVI reader and VBL lookup
xenon_tools/audit_sprite_avi.py Checks a Sprite AVI's size, indexes, PCM presence and VBL audio pairing against the original
xenon_tools/audit_mp4_recording.py The same checks for an MP4 pair: frame counts against VBL tags, codecs, audio and duration drift
xenon_tools/replay_ui.py Optional synchronized original/Sprite video panes