How it works · write-up
Browser (WASM) build performance
Updated 2026-09-20. Covers the Emscripten build only: how to measure a slow browser session, what the 2026-09-20 investigation found, which fixes landed, and what still costs time. For the Python autopilot's own profiling see PROFILING.MD; for how the browser loop is structured (VBL slices instead of Asyncify) see remove-asyncify.md.
Symptom that started this
On a phone (Samsung Galaxy S23, Android 16, Chrome 152, battery saver on) the game ran at 32-35 emulated VBL/s instead of 50, so music and gameplay were both slow, and the in-game options menu animated at ~18 fps even though it draws almost nothing. Three hypotheses were on the table: a debug rather than release build, wrong host/browser timing, and unnecessary work. Only the third one was real -- the numbers below are why.
Measuring: the PERF line
--xenon-perf-stats on (URL: ?perf=1) prints one line per second into the
browser console and a pinned overlay at the bottom of the page that follows the
game tile into fullscreen
(Module.hatariPerfReport in src/emscripten_shell.html). Implementation:
src/xenonPerf.c, called from Main_WasmRunSlice (src/sdl/main_sdl.c);
the JS-side counters live in Module.hatariPerf (web/src/bridge.ts).
PERF @vbl 14817 (+1152): 60.0 VBL/s (11.3 ms/VBL) | slice avg 11.3 max 15.1 ms x61
(ingest avg 0.06 ms x61) | gap avg 0.3 max 3.2 ms, rearm 0 | raf avg 0.22 ms x61
| busy 69% | audio swallow 0 underruns 3
| Field | Meaning |
|---|---|
@vbl N (+M) |
Hatari's VBL counter, and VBLs since the first report (i.e. since the restored --memstate). Per-VBL cost depends on what the game is doing, so compare runs at similar +M, not at similar wall time. |
VBL/s |
Emulated VBLs completed per wall second. 50 = PAL real time (some screens run at 60). (paused) / (benchmark) mark the mode. |
ms/VBL |
Main-thread time inside the emulation slice per emulated VBL. This is the throughput number -- unlike VBL/s it does not move with browser scheduling. Budget: 20 ms at 50 Hz. |
slice avg/max xN |
Wall time of each Main_WasmRunSlice callback (emulation + ScreenTraceList_UpdateWindows + the synchronous JS ingest it triggers) and how many ran. |
ingest avg xN |
Time inside the onXxxFrame calls in bridge.ts -- heap copies and texture uploads. Already included in slice; shown separately to split C emulation from JS/GPU upload work. |
gap avg/max |
How late each callback fired relative to its setTimeout due time. Large values mean the main thread was busy with something else (rAF work, ingest) or the browser throttled timers. |
rearm |
Callbacks that arrived before Timing_GetHostDelayMs() reached zero and had to be re-armed. |
raf avg xN |
Time per requestAnimationFrame tick in bridge.ts (all views' renderFrame()), and ticks per second. Reads 0 in embedders that do not fire rAF for hidden panes. |
busy |
(slice + raf) / wall -- how close the single browser thread is to saturation. |
audio swallow / underruns |
pulse_swallowing_count (audio-sync pacing correction) and audio callbacks that ran short of generated samples in the last second. |
Benchmark mode
?bench=1 adds --benchmark: Timing_WaitOnVbl never waits, and
Main_WasmRunSlice runs VBL slices back-to-back in ~40 ms bursts before
yielding, so the setTimeout floor (4 ms in Chrome) cannot cap the result. The
ms/VBL figure then is raw emulation cost. The game plays unattended, so it
drifts through phases; sample 30-40 s and compare medians at similar +M.
A/B switches (URL parameters on /play/)
All of these are handled in src/emscripten_shell.html and map to Hatari
options parsed in src/options.c. None are player-facing.
| Parameter | Option | Effect |
|---|---|---|
perf=1 |
--xenon-perf-stats on |
PERF line + overlay |
diagnostics=1 |
--xenon-diagnostics on |
the per-frame sprite-stream and rendering diagnostics in the console (src/includes/wasmDiagnostics.h); off by default |
bench=1 |
--benchmark (+ perf) |
unthrottled throughput, see above |
nodebugviews=1 |
--xenon-debug-views off |
do not create the four developer views (VRAM x2, masked sprites, memory map) at all |
nomemtrace=1 |
--xenon-mem-trace off |
never record the per-access rings/maps (browser default is auto, see below) |
noaudio=1 |
--sound off |
no sound generation and no audio-sync pacing |
noauthentic=1 |
--disable-video on |
never convert/blit the authentic ST display |
What the investigation found
Median ms/VBL in benchmark mode, before the fixes. Desktop = Windows PC in
the Claude desktop app's embedded browser pane; phone = Galaxy S23 with
battery saver on.
| Configuration | Desktop | Phone |
|---|---|---|
| baseline | 13.9 | 20.2 |
nomemtrace |
~10.4 | 17.0 |
nodebugviews |
5.2 | 7.7 |
nodebugviews + nomemtrace + noauthentic |
4.3 | 5.3 |
The phone's real-time baseline was 32-35 VBL/s at 21 ms/VBL with the main thread 97% busy, an 8-10 ms scheduling gap and 25-28 audio underruns/s. With all switches on it ran at 60 VBL/s, 11 ms/VBL, <2 ms gap, 68% busy.
So the 68000 emulation plus the sprite-stream reconstruction cost ~5 ms/VBL on the phone; everything above that was work for hidden developer views:
- Per VBL, JS ingest (~9 ms on the phone):
ScreenTraceList_UpdateWindowsbuilt and submitted frames for all five views regardless of visibility. The memory-map view alone copied 1 MB + 3 x 4 MB out of the wasm heap and uploaded four 1024x1024 textures (~13 MB) every VBL (web/src/views/memoryMapView.ts). - Per rAF (5.7 ms x ~50/s on the phone):
bridge.tscalledrenderFrame()on every view, andneedsRenderwas never cleared, so the hidden 1024x1024 memory-map shader (fourtexelFetchper pixel) and the other debug views re-rendered todisplay:nonecanvases on every frame. - Per memory access (~3 ms/VBL): every CPU read/write/fetch pushed into
the read/write/execute access rings and updated each trace's
address/pixel maps (
ScreenTrace_LogRead/LogWrite,src/screentrace.c). - Per VBL, authentic display (~1-2 ms):
Screen_Drawconverted and blitted the ST frame to the hidden#canvas. - Options menu at ~18 fps: not rendering cost (0.5 ms per redraw) but the
paused poll interval,
WASM_PAUSED_POLL_MS = 50insrc/sdl/main_sdl.c.
What was not the cause: the build is compiled -O3 (link at -O1 affects
only binaryen passes), and the browser timer "gap" was a symptom of the
saturated main thread, not of setTimeout clamping -- it fell to <1 ms once
the waste was gone. One real timing wrinkle remains: with battery saver on,
single 10 ms slices cost about twice what the same work costs in sustained
40 ms benchmark bursts (11 vs 5.4 ms/VBL), because the CPU governor keeps
clocks low for bursty work.
Fixes that landed (2026-09-20)
| Area | Change |
|---|---|
src/screentrace.c ScreenTraceList_UpdateWindows |
Developer views are only built/submitted while RenderView_IsVisible (new, src/renderView.c); the sprite stream always updates. A re-shown view picks up on the next VBL or paused poll. |
web/src/bridge.ts, views/*.ts |
The rAF loop skips hidden tiles; needsRender is cleared after each render, so a view renders once per ingested frame. |
src/screentrace.c, src/options.c |
--xenon-mem-trace on|off|auto (ScreenTraceMemTraceMode). Default auto: the rings/maps are recorded only while a developer view is visible, evaluated once per VBL into a single flag the hot path reads. Native also defaults to auto since 2026-10-01; tooling that dumps the rings with every view hidden must explicitly request on. |
src/conv_st.c ConvST_DrawFrame, src/sdl/screen.c |
Browser-only (Screen_IsAuthenticDrawSkipped): skip the authentic display's conversion/blit while hidden, unless frame capture or AVI recording needs the frames; Screen_SetWindowVisible(true) forces a full update. Desktop unchanged (screenshots/paired AVIs of the hidden window are relied upon). Pitfall found the hard way: STRGBPalette[] -- the palette the sprite stream uses for the hit-flash tint, stars and the weapon meter -- is only filled as a side effect of that conversion, so the skip path must still run Convert_PaletteOnly(); without it the flashed ship rendered as a black silhouette and the starfield vanished. |
src/webRenderBackend.c |
Honours RenderViewDesc.startHidden so the tile state matches RenderView_IsVisible from the first frame. |
src/sdl/main_sdl.c |
WASM_PAUSED_POLL_MS 50 -> 16: the options menu animates at display rate. |
Result on the phone with no URL switches: 60 VBL/s (full speed), 10.5 ms/VBL,
gap <1 ms, rAF 0.25 ms, 65% busy, 3-5 audio underruns/s; ?hires=1&highfps=1
is the same. Showing the memory-map view from the developer menu drops the
phone to ~40 VBL/s while it is open (expected -- it is a desktop diagnostic).
What must never be gated
ScreenTrace_LogRead(..., instruction=true) and ScreenTrace_LogWrite are the
entry points for two different things. The instruction-fetch hooks
(DrawCommandStream_On*InstructionFetch, XenonControl_OnInstructionFetch)
and DrawCommandStream_OnMemoryWrite are what capture the sprite stream --
the update procedures the reconstruction is built from. They run
unconditionally on every access, in every mode. Only what follows them
(AccessLog_Push into g_memoryAccessLog, AddressMap_LogTouch,
PixelMap_TouchAt) is subject to --xenon-mem-trace. The consumers of that
data are: the four developer views, their paused-replay animation
(src/animationState.c), the memory-map click inspector
(src/xenonRender.c), the F-key / XENON_MSG ring dumps, and one debugger
command. Nothing on the play path.
Testing on a phone from the PC
Android only; Chrome exposes the DevTools protocol over adb.
- Developer options -> Wireless debugging -> "Pair device with pairing code".
adb pair <ip>:<pairing-port> <code>; the connection port is discovered by mDNS (adb mdns services, thenadb devicesshows it). adb lives inE:\platform-toolson the dev PC (Google's platform-tools zip). - Serve the site:
python site/serve.pybinds all interfaces on 8732, so the phone openshttp://<pc-ip>:8732/play/?perf=1. Open it from the PC withadb shell am start -a android.intent.action.VIEW -d "<url>" com.android.chrome. adb forward tcp:9222 localabstract:chrome_devtools_remote, thenhttp://localhost:9222/jsonlists tabs with awebSocketDebuggerUrl. A 20-line Node 22 script (WebSocketis built in) sendsRuntime.evaluateand readsdocument.getElementById('perf_overlay').textContentonce a second; the menu can be driven withModule.ccall('XenonOptionsMenu_OnPointerEvent', null, ['number','number'], [x, y])in logical 320x200 coordinates (options icon at 306,12; rows at y = 41 + 22*i, x = 50 -- first tap selects, second tap toggles).- Keep the screen on for the run (
settings put system screen_off_timeout), and note the battery-saver state (settings get global low_power) with the numbers.
Still open
- Audio: 3-5 underruns/s remain at full speed. The SDL2 Emscripten audio
callback runs on the main thread and the mix window is small;
-sAUDIO_WORKLETor a largerSoundBufferSizeare the candidates. Measure with theunderrunsfield before and after. - Battery saver / clocks: the 2x per-slice penalty above. Running more than one VBL per callback when behind (catch-up) would keep the CPU busier in longer bursts, but changes pacing; only worth it if a phone drops below 50 VBL/s with the fixes.
- Console volume: the sprite-stream diagnostics emit two lines per real
game frame (~8000 lines/minute). They are now off unless
?diagnostics=1asks for them; see also the "Diagnostic-output overhead" item in remove-asyncify.md. Video_ScreenCounter_ReadBytewas about a third of a 3 ms callback in an early profile (same doc). With the waste gone it is worth a proper bottom-up profile on the phone (chrome://inspect-> Performance; DWARF symbols are already emitted withHATARI_WASM_DEBUG_SYMBOLS=ON).- Link-time
-O3:src/CMakeLists.txtlinks at-O1to keep link time down. Untested; expected single-digit percent, not a priority. - Developer views on phones: opening the memory map costs ~10 ms/VBL of uploads. Fine as a diagnostic; if it should ever be usable on a phone, upload only the rows that changed or downsample.