Xenon 2

Autopilot · write-up

Mission-controller migration

xenondoc/AUTOPILOT_MISSION_MIGRATION.MD · 17 KB · updated 2026-09-17

The subsequent ABI 81 native driving coordinator now connects these missions directly to native preparation/evaluation, with explicit inspection and checkpoint interfaces. The level-by-level results below remain the historical validation of ABI 76–80.

The native/span path now selects missions for all five levels in separate C modules, borrowing the existing world, prediction and persistent-map owners. The old ReactiveController.choose remains the Python reference. This completes level-mission migration, not the remaining observation/host-driver migration.

Boundaries

  1. Observation preparation owns current map inputs, relationships and prediction owners. It does not select an encounter.
  2. The mission controller selects a target, a route and a goal region. Level tactics belong in separate level modules; common navigation owns route search and waypoint advancement.
  3. The existing C decision session evaluates executable maneuvers against walls, predicted enemies and weapon effects. A mission proposal is not a safety verdict.
  4. Python transports observations and exposes compatibility/replay views. The reference mission selector remains available for comparisons and non-native backends.

Preparation refactor

ReactiveController._prepare_mission_observation now calls named preparation methods for map inputs, per-frame caches, level state, wall indexing, links and risk groups. Player absence has its own lifecycle handler. Link reconstruction has one native/reference branch instead of repeated native_links is None tests.

This is a behavior-preserving migration boundary, not a claim that the remaining reference selector is readable or already native. Its 575 gameplay-policy tests pass after the extraction.

Native Level 1 controller (ABI 76)

xenon_autopilot_mission.h exposes one xap_mission_choose call. It borrows the current C world, shared prediction owner, persistent navigation graphs and captured ship/equipment inputs. It returns XapDecisionMission directly to the existing C decision session. Python does not rebuild the three evaluator requests from mission labels on this path.

xenon_autopilot_mission.c owns reset rules, route storage, span connectors and waypoint advancement. xenon_autopilot_level1.c contains the encounter priority:

  1. Reachable, visible shell firing stations, using individual forward/laser/cannon mounts.
  2. The central boss core, using the existing C attack/refuge evaluator.
  3. Monotonic authored corridor progress, including the right side of the left corridor.
  4. Formation pauses based on shared predictions.
  5. Nearby pickups, protecting installed Side Shot from Homing Missile replacement.
  6. Map-based forward travel with a central screen-height preference.

State and tactic/weapon choices have numeric types. State contains a bounded 256-point route and no borrowed pointers. Oversized routes are rejected rather than truncated. Route searches are refreshed on terrain revision changes, waypoint/destination changes, and a bounded reconnect interval. Python replaces immutable state bytes per observation, so shallow replay snapshots do not share mutable mission state. A transient missing-player frame preserves this state; confirmed death/shop/transition and explicit checkpoint restore discard it.

The old selector remains available through the Python kernel and the c-mission-old comparison variant. Non-span navigation still uses that selector; subsequent sections describe the now-native Levels 2–5. This is not yet a completely Python-free driver: map graph inputs, equipment capture, observation preparation and diagnostic views still cross the Python boundary. There are no per-tactic Python callbacks inside the new C API.

Validation notes

The final Level 1 recorded sample (2367–3367) measured 2.474 ms reference versus 1.677 ms native (32.2% lower), with all 1,000 active frames finding a verified maneuver on both paths. The shell/boss sample (7170–8500), after correcting the scroll input, measured 3.084 versus 2.260 ms (26.7% lower). These are fresh-process, CPU-pinned, opposite-order repeats; mission decisions intentionally differ.

The shell/boss sample had 292 reference and 371 native frames without a verified escape. Recorded observations follow the old driver's actions, so this does not directly measure damage under the new driver, but it prevents claiming gameplay equivalence from the timing results. The live checks below test actual progression.

The new shared-prediction regression test distinguishes absolute camera position from per-frame scripted scroll velocity. Confusing them could lock the shared prediction owner to the wrong convention when the absolute scroll became small.

Final comparisons are in xenon_tools/run_logs/native-mission-final-*-0913/. At the original ABI 76 milestone, Level 2, Level 3, Level 4, Level 5 and barrier samples had identical decisions in both repeat orders. Their small timing differences are not a claimed improvement; those levels still used the reference mission at that milestone. The comparator deliberately returns a decision-difference error for the Level 1 samples after writing reports.

253 native, 67 planner, 575 forced-native policy and 24 replay tests pass (919). Tests prohibit calls to the giant reference selector, unrelated Python level-state updates and duplicate shell-mission preparation on the native Level 1 path. They also cover disjoint weapon lanes, missing mounts, map invalidation, route retention, corridor restores, pickup protection and shallow checkpoint isolation.

Live gameplay

  • native-mission-live-0912-b.validation/native-mission-live-0912-b.x2events: start to the first shop, frames 0–2059; no lives lost, one 4-point directional projectile hit at frame 378, shield 39→35, 7 pickups and 10 wall shooters destroyed.
  • native-mission-shell-live-0912-a.validation/native-mission-shell-live-0912-a.x2events: saved encounter checkpoint to the final shop, frames 7171–9519; zero damage, zero lives lost, shield 39→39 and 23 pickups. Shells and the boss were cleared. Boss completion was slower than in the archived reference run (shop frame 8735).

These are two separate runs, not a continuous full campaign. The second inherits the reference checkpoint's equipment. The projectile hit in the first run remains a gameplay limitation to investigate; the timing improvement is not a claim of better damage avoidance. The existing emulator was rebuilt to refresh stale staged sprite assets; autopilot integration into the Hatari build was not added. No WASM profiling was performed.

The subsequent reference-only preparation removal and generic forward-mount helper rename pass the final test/replay comparisons. Replay navigation data also includes the C-selected shell station and target, and metrics identify the mission backend as c-level1 or python.

Level 2 (ABI 77)

The native dispatcher now covers Levels 1 and 2. Level 2 has separate ordered-eye, homing-ring, authored-backbone and final-arena entry/fight functions. Route and gate coordinates are embedded from assets/xenon_level2_strategy.json; no runtime JSON parser or Python tactic callbacks are added. Common pickups, formation holding and map advance moved from Level 1 into the shared mission module unchanged.

Phase and waypoint progress are monotonic. Eight-pixel waypoint tolerance preserves tight labyrinth turns; the larger tolerance skipped a corner in the regression test. Changing levels resets mission state even with consecutive observation frame numbers. Python reference missions remain selectable through c-mission-old; replay reports c-level2 and native semantic tactic names. Shared observation preparation and graph assembly still have Python consumers.

Validation: 256 native tests passed. Reverse-order, two-repeat recorded comparison on frames 32577-33537 averaged 2.607 ms for the reference mission and 1.984 ms for C (23.9% lower complete decision time). Decisions differ intentionally; replay is not proof of live damage/progression equivalence. Report: xenon_tools/run_logs/native-mission-level2-0913-a/comparison.json.

Bounded live validation native-mission-level2-live-0913-a.validation resumed frame 33361 for 600 observations: shield 39 to 39, zero lives lost, three pickups, score +6850. It exercised preparation, ring clearance and activation of the following ring. This is a homing-gauntlet smoke test, not whole-Level-2 acceptance.

Level 3 (ABI 78)

xenon_autopilot_level3.c owns the reviewed nine-gate order and the latched 24-frame articulated-core intercept. Closed gates remain closed until the native persistent map confirms destruction; object disappearance alone cannot advance the mission. Both gate destruction stages share the authored station. Egress retains its route through the first safe bend. This replaces Python gate dictionaries, per-frame group lookups and duplicate carrier lane assembly on the native path.

The mission receives a borrowed native map owner. Restored map seeds now preserve embedded group ordering; the old sorted-name ordering was valid for Python names but would query the wrong group from a native mission. A regression covers restored Level 3 and Level 5 maps. Level 3 prepares only its physical navigation graph: additional comfort dilation disconnects authored passages. Side Shot is protected from both Homing and Rear pickup replacement during mission selection.

Validation: 259 native tests passed, including the restored-ID regression. Recorded frames 49751-50020: reference mission 4.358 ms, native mission 3.069 ms (29.6% lower complete decision time). Report: native-mission-level3-0913-a/comparison.json under run_logs.

Live native-mission-level3-live-0913-b.validation ran 400 observations from frame 52512: shield 7 to 7, no lives lost, score +1500, route progression through the gate approach. No gate destruction occurred in this bounded window, so this is not proof of whole-level completion. The initial ...-live-0913-a snapshot had no active player and is explicitly excluded from gameplay validation. Earlier Python pacing and pressure-relief branches are not copied verbatim; the C safety evaluator retains control over the route and intercept proposals.

Level 4 (ABI 79)

The level module separates optional wall-shooter stations, first-boss exterior bypass/satellite/main-head phases, and final-boss refuge/retraction/attack phases. Installed side capabilities select exterior and refuge side attacks; forward stations use each installed stream's real offset. Shared station construction moved out of Level 1. The reference Python mission remains intact; set XENON_NATIVE_MISSIONS=0 for a live comparison with C prediction/evaluation still enabled.

A nearby waypoint is no longer skipped if the next leg cuts an occupied corner. This reuses the span connector's line test rather than assembling another geometry. Route goals also supply the initial directional proposal instead of a neutral prefix. Optional wall targets release their approach for 128 frames after 48 frames without progress; mandatory Level 3 gates do not use this rule.

Validation: 264 native tests passed. Final reverse-order recorded comparison, frames 63486-64385: reference 3.728 ms, native 3.860 ms (3.5% slower overall). This is a migration result, not a claimed performance win. Report: native-mission-level4-0913-final/comparison.json under run_logs.

Live tests native-mission-level4-live-0913-a through -d exposed a tight-pocket execution limitation near X47/world Y4050. The corner guard and optional-target release are covered by regressions, but the final 400-frame native run still did not leave the pocket. An identical 400-frame run with the retained Python mission (native-mission-level4-reference-0913-a.validation) also stalled there. Both final runs kept shield 23, lost no lives, and gained no score. Therefore this is explicitly NOT Level 4 progression acceptance. A navigation grid's continuous free path does not prove the discrete ship/camera actuator can execute a narrow approach. Resolving that shared navigation/execution mismatch remains gameplay work, independent of moving mission selection out of Python. Boss phase coverage here is synthetic, not a live claim of complete boss victory.

Level 5 (ABI 80)

All five native/span missions now dispatch through the same C API. Level 5 separates progression barriers, diagonal-chamber ordering, missile allocators and train holds from the tank controller in xenon_autopilot_level5_tank.c. The latter separates entry emplacements, cannon volleys, core stations, camera reveal and pulse refuge. Numeric phases and borrowed native map groups replace Python dictionaries and interleaved level conditions. The existing evaluator retains collision authority.

Tank stations account for all installed forward streams and their sustained damage. A productive cannon/core station survives subsequent observations rather than being rebuilt on every frame. Trains hold one lane for the visible family rather than chasing each delayed follower. Barriers use persistent destruction state, including dormant controllers; laser gates remain hazards rather than destructible objectives.

Shared firing routes search an implicit rectangular region in one native graph query per mount instead of trying many individual positions. Graph scratch is reused, and only a winning route is copied into mission state. Side weapons can approach barriers when no vertical firing region is reachable. This adds no Python route helper API. Synthetic regressions cover disconnected nearest stations with another reachable part of the same firing region, and side-only barrier approaches.

Python still captures observation/map/equipment inputs, drives transport and game lifecycle/shops, and exposes replay/reference views. This change does not integrate with Hatari or claim an independently running C driver. Native missions are enabled for all five levels; XENON_NATIVE_MISSIONS=0 keeps the reference selectable.

Final Level 5 validation

Standalone native build and 935 tests pass: 269 native, 67 planner, 575 forced-native reference-policy and 24 replay tests. Final CPU-pinned comparisons use two fresh-process repeats with reversed execution order and frozen source. They measure complete recorded decision processing, not live TCP or video capture. Decisions deliberately differ; the comparison tool writes its report and then returns its expected decision-difference error.

Recorded window Python mission / C kernel Native mission / C kernel Change
Level 5 barriers, 90160-90459 3.327 ms 2.220 ms 33.3% lower
Level 5 tank, 91740-92130 4.523 ms 3.593 ms 20.6% lower

Reports: native-mission-level5-barriers-0913-final/comparison.json and native-mission-level5-tank-0913-final/comparison.json under run_logs. A short Level 1 shell cross-check after the shared region-search change (7170-7500) also completed: approximately 3.173 ms reference versus 2.204 ms native. This is an execution/timing check, not a new live gameplay or behavioral-equivalence claim.

Bounded live outcomes (recordings are inside matching .validation directories under xenon_tools/run_logs, with the run name as the .x2events filename):

  • native-mission-level5-gates-live-0913-b: 90460-90859, no damage or lives lost, but no barrier cleared. A firing region is now found, while execution still stalls near X100/world Y3190. The previous -a run did not find a station.
  • native-mission-level5-tank-live-0913-a: entry run from 91829, damage at 91993, 92005 and 92006; life loss registered at 92023. Shield damage 17, score +1560.
  • native-mission-level5-tank-reference-0913-a: identical checkpoint with Python mission and C kernel also lost its last life, at 91964 (shield damage 17, score +770). Neither path demonstrates acceptable tank-entry play.
  • native-mission-level5-core-live-0913-a: core snapshot from 92889, curved missile damage at 92905 and 92955, life loss at 92972. Shield damage 15, score +690. There was no paired reference run for this core checkpoint.

Level 5 tactics are enabled as requested, with these failures explicitly retained for follow-up tactical work. Neither Level 5 progression nor tank victory is accepted by these tests. Together with the Level 4 pocket limitation above, these are the outstanding live-gameplay issues; unit coverage and speedups do not resolve them. No whole-campaign completion is claimed.