Autopilot · write-up
Native fast-forward investigation (2026-10-01)
The initial investigation found effective fast-forward constrained by GPU presentation synchronization and an unoptimized emulator build. The C pilot itself was already optimized. The implementation follow-up below removes those limits and defaults hidden-view memory tracing to automatic. No gameplay tactics or original-game instructions were changed.
Measured effect
All three full Level 5 runs started at frame 72856 and completed at 82438, using the resident C pilot, a visible Hatari, both PNG AVIs, and 150-frame checkpoints. Their damage, inventory, boss progression and cash outcomes match; see the Level 5 validation report.
| Run | Configuration | Recording wall time | Recorded VBL delta | Mean VBLs / real second |
|---|---|---|---|---|
level5-whole-staged-r125 |
Debug, normal speed, tracing on | 1330.983 s (22m11s) | 35209 | 26.45 |
level5-whole-fast-forward-r126 |
Debug, fast-forward, GPU VSync, tracing on | 832.391 s (13m52s) | 35223 | 42.32 |
level5-whole-fast-optimized-r134 |
Release, fast-forward, immediate GPU presentation, tracing auto | 201.867 s (3m22s) | 35215 | 174.45 |
The old fast-forward setup was approximately 1.60 times faster than normal
speed. The three fixes together make the full run 4.12 times faster than
that old fast-forward setup, reducing elapsed time by 75.7%. It is 6.59
times faster than the original normal-speed validation.
Times span original-AVI creation through finalized driver-result.json; they
include recording, observation, checkpoints and cleanup, but exclude building.
VBL counts come from the sprite AVI's .avi.vbl tags. These sequential full
runs are useful validation evidence, not a controlled CPU benchmark. Small
differences in final VBL counts reflect stop/cleanup timing.
Recording paths:
E:/xenon_runs/level5-whole-staged-r125.validation/level5-whole-staged-r125.x2eventsE:/xenon_runs/level5-whole-fast-forward-r126.validation/level5-whole-fast-forward-r126.x2eventsE:/xenon_runs/level5-whole-fast-optimized-r134.validation/level5-whole-fast-optimized-r134.x2events
What fast-forward does
XENON_MSG_SET_FAST_FORWARD sets ConfigureParams.System.bFastForward in
src/xenonControl.c. Timing_WaitOnVbl() in src/sdl/timing.c then returns
before its normal host sleep/busy wait. It removes the wall-clock pacing limit;
it does not multiply the Atari's clock or change Xenon's update rules.
PAL video still produces about 50 VBLs per emulated second. Fast-forward can execute more than 50 VBLs per real second if the host can complete the work. A short early-Level-5 probe below reached 53.94 VBL/s, confirming this. The old full-run average of 42.32 VBL/s was below real-time PAL speed; the optimized run averages 174.45 VBL/s, including recording and checkpoints.
Emulated CPU instructions, interrupts, rendering, sound and recording still
execute. The official launcher uses --frameskips 0, so rendering is not
skipped. AVI playback uses the emulated video timebase; watching the AVI does
not reproduce the acceleration relative to the host clock.
Baseline executable was not a Release build
Actual caches and out/build/mingw-debug/compile_commands.json show:
| Component | Effective optimization |
|---|---|
| Hatari main, 68000 interpreter, tracing, GPU view, AVI recorder | Debug: -g -O0 |
| C autopilot linked into Hatari | -O3, explicitly added by src/autopilot/CMakeLists.txt |
| Standalone replay DLL | Release: -O3 -DNDEBUG |
Hatari commands contain an earlier -O, but their later -O0 overrides it.
CMakePresets.json explicitly defines mingw-debug as Debug. The official
xenon_tools/hatari_dev.ps1 configures that preset even for a custom output
directory; naming a directory mingw-release would not change the build type.
The follow-up adds a Release preset and official wrapper selector; baseline
flags and hashes below identify the original measurements.
Sampled executable:
out/build/mingw-debug/src/hatari.exe, SHA256
1b5dcd09c7232931fa59862f755d801c898d92f07dccdc29ae33a34f3e9cd1e4.
Replay DLL SHA256:
eed43d950e74f3b6a938460bd6927d2f6c5401b52aeb9d9ae43fedc91571fc5b.
Where the sampled time goes
Windows WPR could not enable CPU profiling (0xc5585011, policy/privilege
failure). A process-local helper instead briefly suspended only Hatari's main
thread, copied its native instruction pointer and occasional small stack
windows, then immediately resumed it. Approximately 1900 wall-clock samples
were collected over each 30-second active-gameplay window. Main-thread CPU
time came from GetThreadTimes, independently of sample attribution.
The first short probe, r127, was insufficiently long to rule out end-of-run pause contamination. It is excluded from the conclusions below. r128 repeated active sampling with checkpoints. r129 and r130 removed periodic checkpoint pauses to isolate the running emulator; both still recorded both AVIs.
| Early Level 5 probe | Developer memory tracing | Mean VBL/s over recording | Main-thread CPU during sample | Samples in NtDelayExecution |
|---|---|---|---|---|
level5-fast-forward-profile-r129 |
default: on | 53.94 | 58.0% of one core | 45.34% |
level5-fast-forward-mem-off-r130 |
off | 54.50 | 37.8% of one core | 65.60% |
These are approximate samples of the opening section, not an attribution of the entire campaign or all threads. The roughly 1% throughput difference is too small to establish a useful speedup. The CPU reduction is substantial, but the freed time mostly becomes presentation waiting in this section.
GPU submission still synchronizes to host VSync
In r129, 50 of 51 sampled host-delay stack windows contained a validated call
return address in SdlGpuRenderView_Submit. That address is 0x1400a56c4 in
the sampled binary; addr2line maps it to src/sdlGpuRenderView.c:2033, just
after SDL_SubmitGPUCommandBufferAndAcquireFence(commandBuffer). r130 showed
the same caller in 78 of 79 delay stack windows. These are stack-window caller
clues, not fully unwound call stacks; the repeated exact call site and source
configuration together identify GPU submission as the main observed wait.
SdlGpuRenderView_Create calls SDL_ClaimWindowForGPUDevice, but never sets
the GPU swapchain's presentation mode. The locally installed SDL header
(C:/msys64/ucrt64/include/SDL3/SDL_gpu.h, around line 4100) explicitly says
that claiming a window creates a VSync swapchain. Hatari's ordinary screen
--vsync setting controls its SDL renderer, not this separate GPU swapchain.
The nonblocking SDL_AcquireGPUSwapchainTexture used during emulation does
not prevent the later GPU submission from waiting. Fast-forward bypasses
Hatari's timer wait, but does not alter this presentation policy. Capture also
retains GPU download/fence work so that every recorded VBL is preserved.
Developer memory tracing is expensive and unnecessary for ordinary play
Desktop originally defaulted to SCREENTRACE_MEMTRACE_ON; WASM defaulted to automatic
tracing. With tracing enabled, every emulated memory access updates access
logs, address maps and pixel maps for developer views. In r129, leaf samples
in ScreenTrace_LogRead_one, AddressMap_LogTouch, PixelMap_TouchAt, and
AccessLog_Push alone accounted for about 12% of wall samples, with further
work in the tracing entry points.
--xenon-mem-trace off bypasses that bookkeeping while retaining required
instruction/memory hooks for sprite reconstruction, canonical observations
and resident-C control. Both probe AVIs remained valid. Hidden developer views
already skip their rendering; their memory bookkeeping is a separate cost.
Other sampled work included draw-command preparation and sprite lookup.
ScreenTrace_LogRead also calls a series of specialized instruction-fetch
hooks, including shop hooks, for every guest instruction. Many reject the
address immediately. Their cost is amplified by -O0; profile an optimized
build before changing the hook architecture. PNG compression and file writes
run synchronously for both AVIs, but this opening-section profile does not
establish compression as the dominant cost. Busy boss encounters need their
own samples after the presentation limit is removed.
The web patch solves a different problem
M68000_PatchXenonRasterWait() in src/m68000.c is compiled only for the
WASM patch configuration. It replaces the backwards branch at $2904 with
a NOP after verifying the instruction sequence starting at $28F6:
lea $ff8207,a0
move.l $040a,d0
add.l d7,d0
lsr.l #8,d0
cmp.b (a0),d0
bne.s back_to_cmp
This removes a guest CPU busy loop polling the Shifter's raster counter. It
does not remove Xenon's four-VBL buffer-swap cadence, or fix host GPU VSync.
Existing 82-swap instrumentation documents that the NOP changes swap timing
from approximately scanline 192 to scanline 0-2, which can mix buffers/flicker
in the authentic ST output; see XENON2.MD, section 8.1.
Simply enabling that NOP on desktop is therefore not the first fix. If guest polling remains expensive after the host/build fixes, accelerate the known idle loop while preserving emulated scanline timing and intervening HBL/timer/audio events. Its individual cost was not measured by this native sampling experiment.
Recommended order
- Make GPU presentation honor fast-forward using a supported uncapped presentation mode, restoring normal policy when fast-forward stops. Retain every AVI capture; do not trade missing frames for apparent speed.
- Add a Release preset selectable through the official PowerShell wrapper, then compare the same snapshot with the same recording settings.
- Use memory tracing off/automatic for ordinary validation, enabling it only when memory-access diagnostics are needed. Required capture hooks stay on.
- Profile an expensive boss section again. Optimize recording, remaining guest idle polling or pilot evaluation according to that evidence.
The baseline investigation did not implement these changes; the follow-up below implements and measures recommendations 1-3.
Reproduction and local artifacts
The two principal probe recordings are:
E:/xenon_runs/level5-fast-forward-profile-r129.validation/level5-fast-forward-profile-r129.x2eventsE:/xenon_runs/level5-fast-forward-mem-off-r130.validation/level5-fast-forward-mem-off-r130.x2events
Both completed around frame 74359 with no incidents. All probe AVIs were
gracefully finalized, with no cleanup errors; RIFF sizes match file sizes.
Their first and last indexed images decode at 640x400, verified using the
existing HatariAvi reader. QA is in work/native-fast-forward-probe-avi-qa.json.
The finalized Hatari instances were closed.
Baseline probe:
python xenon_tools/validate_campaign.py native-run level5-fast-forward-profile-r129 --output-root E:/xenon_runs --resume E:/xenon_runs/level5-full-from-l4-shop-r1.validation/start.sav --port 6903 --fast-forward --frames 1500 --checkpoint-interval 100000
Tracing-off probe, using the official launcher and immediately attaching the same resident-C validation tool:
& ./xenon_tools/hatari_dev.ps1 run -Snapshot E:/xenon_runs/level5-full-from-l4-shop-r1.validation/start.sav -ControlPort 6903 -HatariArgument @('--xenon-mem-trace','off')
python xenon_tools/validate_campaign.py native-run level5-fast-forward-mem-off-r130 --output-root E:/xenon_runs --connect --port 6903 --fast-forward --frames 1500 --checkpoint-interval 100000
Local diagnostic utilities are work/sample_hatari_process.c and
work/summarize_hatari_samples.py. Compile the former with the normalized
MSYS2 runtime PATH and gcc -O2 -Wall -Wextra ... -lpsapi, then run:
# Substitute the live PID printed by the launcher; sample while gameplay is active.
work/sample_hatari_process.exe <PID> 30 work/sample.tsv
python work/summarize_hatari_samples.py work/sample.tsv --output work/sample.json
Raw samples/summaries remain under work/l5-debug-fast-stack.* and
work/l5-debug-fast-mem-off.*. work/native-fast-forward-performance.json
contains timing, AVI checks and both sample summaries. These are local
investigation artifacts, not production build inputs.
Implementation and separate measurements (2026-10-01)
The native GPU views now query supported presentation modes once at creation. On a fast-forward toggle they select immediate presentation, or mailbox when immediate is unsupported; stopping fast-forward restores VSync. A platform without either mode keeps VSync and logs that limitation. No per-VBL capability query or swapchain recreation is introduced, and the existing ordered GPU capture ring continues to preserve every recorded VBL.
mingw-release inherits the existing MinGW toolchain and runtime setup but
selects Release. The official wrapper accepts -BuildType Release, with a
separate default tree out/build/mingw-release. Debug remains the development
default. validate_campaign.py native-run --build-type Release --build selects
that preset and rebuilds both Hatari and the replay DLL. Its manifest includes
the settings and the hash of the executable after building.
Both desktop and browser now default to automatic developer memory tracing.
Only access-ring/map bookkeeping is gated; sprite reconstruction, resident C
control and observations retain their hooks. Explicit --memory-trace on
(-MemoryTrace on in the wrapper) supports hidden-window ring-dump work.
Separate probes use the same snapshot, about 1500 canonical frames, a visible
emulator, both PNG-level-6 AVIs, fast-forward, and no periodic checkpoint pauses.
The VBL rates account for small stop-poll overshoots. Wall time again spans
original-AVI creation through finalized driver-result.json, excluding build.
| Probe | Emulator / presentation / tracing | Wall seconds | VBL/s | Gain over previous step |
|---|---|---|---|---|
| r129 baseline | Debug / VSync / on | 111.497 | 53.94 | - |
| r131 GPU fix | Debug / immediate / on | 67.715 | 89.45 | 1.66x |
| r132 Release | Release / immediate / on | 47.548 | 127.68 | 1.43x |
| r133 tracing auto | Release / immediate / auto | 40.835 | 148.60 | 1.16x |
The combined opening-section throughput gain is 2.75x over the previous fast-forward setup. The measurements are sequential probes, not statistical benchmarks under fixed system load. A 20-second r131 sample recorded 98.5% main-thread CPU utilization, versus 58.0% in baseline r129; host-delay stack windows no longer pointed to GPU submission. Required renderer/capture work still runs.
Probe recordings:
E:/xenon_runs/level5-fast-forward-gpu-r131.validation/level5-fast-forward-gpu-r131.x2eventsE:/xenon_runs/level5-fast-forward-release-r132.validation/level5-fast-forward-release-r132.x2eventsE:/xenon_runs/level5-fast-forward-release-auto-r133.validation/level5-fast-forward-release-auto-r133.x2events
Comparing captured ship positions, consumed inputs and scroll against r129 found no difference over the shared opening-section records for any of these probes. The full-level follow-up retains normal 150-frame checkpoints and the optional Level 5 C trace, so its timing is reported separately.
Full-level follow-up and verification
Run r134 completed Level 5 at frame 82438 in 201.867 seconds. Comparing all shared records against r126 found identical ship positions, consumed inputs and scrolling. The read-only audit also matches exactly after excluding the recording filename: four shield hits, no life loss, the same boss completion, equipment changes and pickup/cash outcomes. The incident monitor reported no stall. The performance changes do not resolve the existing gameplay issues:
- Frames 76439, 76465, 80774 and 81628: shield damage of 4, 8, 2 and 2.
- Frame 76425: Homing Missile #603341 missed.
- Final-boss cash: 17/20 coins collected, worth 1200/1500. Three large coins expire at 82375, 82403 and 82416. Tank cash remains 10/10, worth 750.
The visible resident C run recorded both AVIs and checkpoints with the same requested 150-frame interval. Both AVIs have 35216 indexed frames and VBL tags; their tags are contiguous, their RIFF sizes match their file sizes, and their first/last images decode at 640x400. The three separate implementation probes passed the same AVI checks. Cleanup completed without errors and the finalized Hatari instance was closed.
Debug Hatari, Release Hatari and the standalone replay DLL were rebuilt with
the official PowerShell script. Release compilation uses -O3 -DNDEBUG for
the emulator as well as the pilot; the DLL remains ABI 98. Eleven targeted
campaign launch/monitor tests passed. The r134 manifest records executable
SHA-256 ca3bc754bad38f2316892caaa09a8cafedbb72419faa90c4c961a7e41b114758.
Reproduce from the repository root (a new run name is required):
$env:XAP_LEVEL5_TRACE_FILE='E:/xenon_runs/level5-fast-repeat.trace'
python xenon_tools/validate_campaign.py native-run level5-fast-repeat --output-root E:/xenon_runs --resume E:/xenon_runs/level5-full-from-l4-shop-r1.validation/start.sav --port 6903 --build-type Release --build --fast-forward --frames 30000 --checkpoint-interval 150
--build recompiles both Hatari and the replay DLL; omit it only when both are
already current. Ordinary tracing is automatic without another option. Use
--memory-trace on for a memory-ring investigation. Fast-forward is restored
to normal speed during graceful recording cleanup. No desktop game-code speed
patch, gameplay decision change or Python autopilot was introduced.
Raw timing and AVI QA are in work/native-fast-forward-optimized-qa.json;
the full audit is work/level5-r134-audit.json. A 30-second optimized sample is
in work/l5-release-fast-auto.*: main-thread CPU utilization was 67.1%, and
no sampled delay-stack windows pointed to GPU submission. This sample includes
checkpoint pauses and capture synchronization. Its wall-time function counts
are useful leads, not precise attribution of CPU cost. Further optimization
should measure capture/recording and required instruction-fetch hooks in an
expensive section before changing guest timing or pilot tactics.