Xenon 2

Autopilot · write-up

Native fast-forward investigation (2026-10-01)

xenondoc/NATIVE_FAST_FORWARD_PERFORMANCE.MD · 18 KB · updated 2026-10-03

The initial investigation found effective fast-forward constrained by GPU presentation synchronization and an unoptimized emulator build. The C pilot itself was already optimized. The implementation follow-up below removes those limits and defaults hidden-view memory tracing to automatic. No gameplay tactics or original-game instructions were changed.

Measured effect

All three full Level 5 runs started at frame 72856 and completed at 82438, using the resident C pilot, a visible Hatari, both PNG AVIs, and 150-frame checkpoints. Their damage, inventory, boss progression and cash outcomes match; see the Level 5 validation report.

Run Configuration Recording wall time Recorded VBL delta Mean VBLs / real second
level5-whole-staged-r125 Debug, normal speed, tracing on 1330.983 s (22m11s) 35209 26.45
level5-whole-fast-forward-r126 Debug, fast-forward, GPU VSync, tracing on 832.391 s (13m52s) 35223 42.32
level5-whole-fast-optimized-r134 Release, fast-forward, immediate GPU presentation, tracing auto 201.867 s (3m22s) 35215 174.45

The old fast-forward setup was approximately 1.60 times faster than normal speed. The three fixes together make the full run 4.12 times faster than that old fast-forward setup, reducing elapsed time by 75.7%. It is 6.59 times faster than the original normal-speed validation. Times span original-AVI creation through finalized driver-result.json; they include recording, observation, checkpoints and cleanup, but exclude building. VBL counts come from the sprite AVI's .avi.vbl tags. These sequential full runs are useful validation evidence, not a controlled CPU benchmark. Small differences in final VBL counts reflect stop/cleanup timing.

Recording paths:

  • E:/xenon_runs/level5-whole-staged-r125.validation/level5-whole-staged-r125.x2events
  • E:/xenon_runs/level5-whole-fast-forward-r126.validation/level5-whole-fast-forward-r126.x2events
  • E:/xenon_runs/level5-whole-fast-optimized-r134.validation/level5-whole-fast-optimized-r134.x2events

What fast-forward does

XENON_MSG_SET_FAST_FORWARD sets ConfigureParams.System.bFastForward in src/xenonControl.c. Timing_WaitOnVbl() in src/sdl/timing.c then returns before its normal host sleep/busy wait. It removes the wall-clock pacing limit; it does not multiply the Atari's clock or change Xenon's update rules.

PAL video still produces about 50 VBLs per emulated second. Fast-forward can execute more than 50 VBLs per real second if the host can complete the work. A short early-Level-5 probe below reached 53.94 VBL/s, confirming this. The old full-run average of 42.32 VBL/s was below real-time PAL speed; the optimized run averages 174.45 VBL/s, including recording and checkpoints.

Emulated CPU instructions, interrupts, rendering, sound and recording still execute. The official launcher uses --frameskips 0, so rendering is not skipped. AVI playback uses the emulated video timebase; watching the AVI does not reproduce the acceleration relative to the host clock.

Baseline executable was not a Release build

Actual caches and out/build/mingw-debug/compile_commands.json show:

Component Effective optimization
Hatari main, 68000 interpreter, tracing, GPU view, AVI recorder Debug: -g -O0
C autopilot linked into Hatari -O3, explicitly added by src/autopilot/CMakeLists.txt
Standalone replay DLL Release: -O3 -DNDEBUG

Hatari commands contain an earlier -O, but their later -O0 overrides it. CMakePresets.json explicitly defines mingw-debug as Debug. The official xenon_tools/hatari_dev.ps1 configures that preset even for a custom output directory; naming a directory mingw-release would not change the build type. The follow-up adds a Release preset and official wrapper selector; baseline flags and hashes below identify the original measurements.

Sampled executable: out/build/mingw-debug/src/hatari.exe, SHA256 1b5dcd09c7232931fa59862f755d801c898d92f07dccdc29ae33a34f3e9cd1e4. Replay DLL SHA256: eed43d950e74f3b6a938460bd6927d2f6c5401b52aeb9d9ae43fedc91571fc5b.

Where the sampled time goes

Windows WPR could not enable CPU profiling (0xc5585011, policy/privilege failure). A process-local helper instead briefly suspended only Hatari's main thread, copied its native instruction pointer and occasional small stack windows, then immediately resumed it. Approximately 1900 wall-clock samples were collected over each 30-second active-gameplay window. Main-thread CPU time came from GetThreadTimes, independently of sample attribution.

The first short probe, r127, was insufficiently long to rule out end-of-run pause contamination. It is excluded from the conclusions below. r128 repeated active sampling with checkpoints. r129 and r130 removed periodic checkpoint pauses to isolate the running emulator; both still recorded both AVIs.

Early Level 5 probe Developer memory tracing Mean VBL/s over recording Main-thread CPU during sample Samples in NtDelayExecution
level5-fast-forward-profile-r129 default: on 53.94 58.0% of one core 45.34%
level5-fast-forward-mem-off-r130 off 54.50 37.8% of one core 65.60%

These are approximate samples of the opening section, not an attribution of the entire campaign or all threads. The roughly 1% throughput difference is too small to establish a useful speedup. The CPU reduction is substantial, but the freed time mostly becomes presentation waiting in this section.

GPU submission still synchronizes to host VSync

In r129, 50 of 51 sampled host-delay stack windows contained a validated call return address in SdlGpuRenderView_Submit. That address is 0x1400a56c4 in the sampled binary; addr2line maps it to src/sdlGpuRenderView.c:2033, just after SDL_SubmitGPUCommandBufferAndAcquireFence(commandBuffer). r130 showed the same caller in 78 of 79 delay stack windows. These are stack-window caller clues, not fully unwound call stacks; the repeated exact call site and source configuration together identify GPU submission as the main observed wait.

SdlGpuRenderView_Create calls SDL_ClaimWindowForGPUDevice, but never sets the GPU swapchain's presentation mode. The locally installed SDL header (C:/msys64/ucrt64/include/SDL3/SDL_gpu.h, around line 4100) explicitly says that claiming a window creates a VSync swapchain. Hatari's ordinary screen --vsync setting controls its SDL renderer, not this separate GPU swapchain.

The nonblocking SDL_AcquireGPUSwapchainTexture used during emulation does not prevent the later GPU submission from waiting. Fast-forward bypasses Hatari's timer wait, but does not alter this presentation policy. Capture also retains GPU download/fence work so that every recorded VBL is preserved.

Developer memory tracing is expensive and unnecessary for ordinary play

Desktop originally defaulted to SCREENTRACE_MEMTRACE_ON; WASM defaulted to automatic tracing. With tracing enabled, every emulated memory access updates access logs, address maps and pixel maps for developer views. In r129, leaf samples in ScreenTrace_LogRead_one, AddressMap_LogTouch, PixelMap_TouchAt, and AccessLog_Push alone accounted for about 12% of wall samples, with further work in the tracing entry points.

--xenon-mem-trace off bypasses that bookkeeping while retaining required instruction/memory hooks for sprite reconstruction, canonical observations and resident-C control. Both probe AVIs remained valid. Hidden developer views already skip their rendering; their memory bookkeeping is a separate cost.

Other sampled work included draw-command preparation and sprite lookup. ScreenTrace_LogRead also calls a series of specialized instruction-fetch hooks, including shop hooks, for every guest instruction. Many reject the address immediately. Their cost is amplified by -O0; profile an optimized build before changing the hook architecture. PNG compression and file writes run synchronously for both AVIs, but this opening-section profile does not establish compression as the dominant cost. Busy boss encounters need their own samples after the presentation limit is removed.

The web patch solves a different problem

M68000_PatchXenonRasterWait() in src/m68000.c is compiled only for the WASM patch configuration. It replaces the backwards branch at $2904 with a NOP after verifying the instruction sequence starting at $28F6:

lea     $ff8207,a0
move.l  $040a,d0
add.l   d7,d0
lsr.l   #8,d0
cmp.b   (a0),d0
bne.s   back_to_cmp

This removes a guest CPU busy loop polling the Shifter's raster counter. It does not remove Xenon's four-VBL buffer-swap cadence, or fix host GPU VSync. Existing 82-swap instrumentation documents that the NOP changes swap timing from approximately scanline 192 to scanline 0-2, which can mix buffers/flicker in the authentic ST output; see XENON2.MD, section 8.1.

Simply enabling that NOP on desktop is therefore not the first fix. If guest polling remains expensive after the host/build fixes, accelerate the known idle loop while preserving emulated scanline timing and intervening HBL/timer/audio events. Its individual cost was not measured by this native sampling experiment.

  1. Make GPU presentation honor fast-forward using a supported uncapped presentation mode, restoring normal policy when fast-forward stops. Retain every AVI capture; do not trade missing frames for apparent speed.
  2. Add a Release preset selectable through the official PowerShell wrapper, then compare the same snapshot with the same recording settings.
  3. Use memory tracing off/automatic for ordinary validation, enabling it only when memory-access diagnostics are needed. Required capture hooks stay on.
  4. Profile an expensive boss section again. Optimize recording, remaining guest idle polling or pilot evaluation according to that evidence.

The baseline investigation did not implement these changes; the follow-up below implements and measures recommendations 1-3.

Reproduction and local artifacts

The two principal probe recordings are:

  • E:/xenon_runs/level5-fast-forward-profile-r129.validation/level5-fast-forward-profile-r129.x2events
  • E:/xenon_runs/level5-fast-forward-mem-off-r130.validation/level5-fast-forward-mem-off-r130.x2events

Both completed around frame 74359 with no incidents. All probe AVIs were gracefully finalized, with no cleanup errors; RIFF sizes match file sizes. Their first and last indexed images decode at 640x400, verified using the existing HatariAvi reader. QA is in work/native-fast-forward-probe-avi-qa.json. The finalized Hatari instances were closed.

Baseline probe:

python xenon_tools/validate_campaign.py native-run level5-fast-forward-profile-r129 --output-root E:/xenon_runs --resume E:/xenon_runs/level5-full-from-l4-shop-r1.validation/start.sav --port 6903 --fast-forward --frames 1500 --checkpoint-interval 100000

Tracing-off probe, using the official launcher and immediately attaching the same resident-C validation tool:

& ./xenon_tools/hatari_dev.ps1 run -Snapshot E:/xenon_runs/level5-full-from-l4-shop-r1.validation/start.sav -ControlPort 6903 -HatariArgument @('--xenon-mem-trace','off')
python xenon_tools/validate_campaign.py native-run level5-fast-forward-mem-off-r130 --output-root E:/xenon_runs --connect --port 6903 --fast-forward --frames 1500 --checkpoint-interval 100000

Local diagnostic utilities are work/sample_hatari_process.c and work/summarize_hatari_samples.py. Compile the former with the normalized MSYS2 runtime PATH and gcc -O2 -Wall -Wextra ... -lpsapi, then run:

# Substitute the live PID printed by the launcher; sample while gameplay is active.
work/sample_hatari_process.exe <PID> 30 work/sample.tsv
python work/summarize_hatari_samples.py work/sample.tsv --output work/sample.json

Raw samples/summaries remain under work/l5-debug-fast-stack.* and work/l5-debug-fast-mem-off.*. work/native-fast-forward-performance.json contains timing, AVI checks and both sample summaries. These are local investigation artifacts, not production build inputs.

Implementation and separate measurements (2026-10-01)

The native GPU views now query supported presentation modes once at creation. On a fast-forward toggle they select immediate presentation, or mailbox when immediate is unsupported; stopping fast-forward restores VSync. A platform without either mode keeps VSync and logs that limitation. No per-VBL capability query or swapchain recreation is introduced, and the existing ordered GPU capture ring continues to preserve every recorded VBL.

mingw-release inherits the existing MinGW toolchain and runtime setup but selects Release. The official wrapper accepts -BuildType Release, with a separate default tree out/build/mingw-release. Debug remains the development default. validate_campaign.py native-run --build-type Release --build selects that preset and rebuilds both Hatari and the replay DLL. Its manifest includes the settings and the hash of the executable after building.

Both desktop and browser now default to automatic developer memory tracing. Only access-ring/map bookkeeping is gated; sprite reconstruction, resident C control and observations retain their hooks. Explicit --memory-trace on (-MemoryTrace on in the wrapper) supports hidden-window ring-dump work.

Separate probes use the same snapshot, about 1500 canonical frames, a visible emulator, both PNG-level-6 AVIs, fast-forward, and no periodic checkpoint pauses. The VBL rates account for small stop-poll overshoots. Wall time again spans original-AVI creation through finalized driver-result.json, excluding build.

Probe Emulator / presentation / tracing Wall seconds VBL/s Gain over previous step
r129 baseline Debug / VSync / on 111.497 53.94 -
r131 GPU fix Debug / immediate / on 67.715 89.45 1.66x
r132 Release Release / immediate / on 47.548 127.68 1.43x
r133 tracing auto Release / immediate / auto 40.835 148.60 1.16x

The combined opening-section throughput gain is 2.75x over the previous fast-forward setup. The measurements are sequential probes, not statistical benchmarks under fixed system load. A 20-second r131 sample recorded 98.5% main-thread CPU utilization, versus 58.0% in baseline r129; host-delay stack windows no longer pointed to GPU submission. Required renderer/capture work still runs.

Probe recordings:

  • E:/xenon_runs/level5-fast-forward-gpu-r131.validation/level5-fast-forward-gpu-r131.x2events
  • E:/xenon_runs/level5-fast-forward-release-r132.validation/level5-fast-forward-release-r132.x2events
  • E:/xenon_runs/level5-fast-forward-release-auto-r133.validation/level5-fast-forward-release-auto-r133.x2events

Comparing captured ship positions, consumed inputs and scroll against r129 found no difference over the shared opening-section records for any of these probes. The full-level follow-up retains normal 150-frame checkpoints and the optional Level 5 C trace, so its timing is reported separately.

Full-level follow-up and verification

Run r134 completed Level 5 at frame 82438 in 201.867 seconds. Comparing all shared records against r126 found identical ship positions, consumed inputs and scrolling. The read-only audit also matches exactly after excluding the recording filename: four shield hits, no life loss, the same boss completion, equipment changes and pickup/cash outcomes. The incident monitor reported no stall. The performance changes do not resolve the existing gameplay issues:

  • Frames 76439, 76465, 80774 and 81628: shield damage of 4, 8, 2 and 2.
  • Frame 76425: Homing Missile #603341 missed.
  • Final-boss cash: 17/20 coins collected, worth 1200/1500. Three large coins expire at 82375, 82403 and 82416. Tank cash remains 10/10, worth 750.

The visible resident C run recorded both AVIs and checkpoints with the same requested 150-frame interval. Both AVIs have 35216 indexed frames and VBL tags; their tags are contiguous, their RIFF sizes match their file sizes, and their first/last images decode at 640x400. The three separate implementation probes passed the same AVI checks. Cleanup completed without errors and the finalized Hatari instance was closed.

Debug Hatari, Release Hatari and the standalone replay DLL were rebuilt with the official PowerShell script. Release compilation uses -O3 -DNDEBUG for the emulator as well as the pilot; the DLL remains ABI 98. Eleven targeted campaign launch/monitor tests passed. The r134 manifest records executable SHA-256 ca3bc754bad38f2316892caaa09a8cafedbb72419faa90c4c961a7e41b114758.

Reproduce from the repository root (a new run name is required):

$env:XAP_LEVEL5_TRACE_FILE='E:/xenon_runs/level5-fast-repeat.trace'
python xenon_tools/validate_campaign.py native-run level5-fast-repeat --output-root E:/xenon_runs --resume E:/xenon_runs/level5-full-from-l4-shop-r1.validation/start.sav --port 6903 --build-type Release --build --fast-forward --frames 30000 --checkpoint-interval 150

--build recompiles both Hatari and the replay DLL; omit it only when both are already current. Ordinary tracing is automatic without another option. Use --memory-trace on for a memory-ring investigation. Fast-forward is restored to normal speed during graceful recording cleanup. No desktop game-code speed patch, gameplay decision change or Python autopilot was introduced.

Raw timing and AVI QA are in work/native-fast-forward-optimized-qa.json; the full audit is work/level5-r134-audit.json. A 30-second optimized sample is in work/l5-release-fast-auto.*: main-thread CPU utilization was 67.1%, and no sampled delay-stack windows pointed to GPU submission. This sample includes checkpoint pauses and capture synchronization. Its wall-time function counts are useful leads, not precise attribution of CPU cost. Further optimization should measure capture/recording and required instruction-fetch hooks in an expensive section before changing guest timing or pilot tactics.