# Roadmap Groomed 2026-09-08. Replaces the phase-0/phase-2 backlog, all of which shipped. The board (`board.md`) is the ledger of findings; this is the plan. **The product is `sots-engine` — our code.** RE findings are inputs to it, and the measure of progress is how much of the game runs in our engine, verified against the original. Every number here is measured; the source is named. ## 1. The engine today (`sots-engine` main, ~33k LOC, ctest 46/46, 1,031 addresses, clean-room, MIT) | module | LOC | what it is | verification | |---|---:|---|---| | `mars/parse`, `mars/text` | 1,774 | brace-block + flat-KV + CSV readers (the `Mars::Script` pull tokenizer, first-wins keys) | oracle 1,531/1,531 + 64/64 | | `mars/vfs` | 788 | `.gob` ZIP + native-override VFS (the `./Mods` mechanism) | 10,268 files CRC-clean | | `mars/stream` | 5,407 | save reader/writer, 386-class wire schema, `SchemaProbe`, `CoverageArchive` | **100% named coverage, byte-identical round-trip, 11 saves** | | `mars/rng` | 244 | MT19937 + all seven entry points (`NextFloat/NextInt/Chance/NextUInt/FloatRange/IntRangeBell/GaussianRange`) | live: every draw of a turn attributed | | `game/data`, `game/design`, `game/config` | 3,835 | weapons/sections/tech/strings catalogs, design rules, hull class, tuning | 229,042 values 0 diffs; 127/127 designs | | `game/effects`, `game/events` | 1,531 | tech effects, `EventStorage` (turn-bucketed, dedup) | live: research slice 35 calls / 0 div | | `game/sim` | 3,382 | economy (`ComputeBudget`, output term), research + unlock cascade, colony, movement | live: budget 4,437 / output 24,357 / research 35 / movement 45 calls — all 0 div on declared regions | | `game/nav`, `game/combat` | 1,163 | route classifier + failure bits; retreat **planner** | offline vs 58 waypoints / 46 plans, 0 failures; retreat never instrumented | | `app` (`sots_turn`) | 1,754 | **the standalone**: load save → walk all three drivers' phase order → write save; divergence + completion metric | 11/11 saves run; ref pair **209 → 158** leaves, closed 51 / regressed 0 | | `shim` | 769 + hooks | proxy `binkw32.dll` + MinHook trace/compare/replace harness — **scaffolding, to be displaced** | it is the instrument, not the product | What the engine can do end to end today: read every shipped asset and every save byte-for-byte, run the economy/research/movement/output pieces to the original's numbers, and step a turn with **16 of 44** driver phases and **5 of 37** tail phases modelled (10 committed). It cannot yet reproduce a turn. ## 2. The milestone, restated — 2026-09-09 (after the determinism arc) > `sots_turn` loads a save, runs one strategic turn, and writes an autosave that byte-matches > the original's — **for a stated procedure and route**. Two days of lanes went into determinism and the foundation is now laid. What is *measured*: - **The sim is deterministic given the command stream.** Every draw of a turn attributed by return address; the strategic generator's residual is 0 on every measured turn. - **The AI is deterministic given (seeds, visit order).** Its per-process seed is one word per client, minted from the global generator — which does *nothing else* on an encounter-free turn. Pinning the seeds collapses a rich turn's divergence from 94 leaves to **one transposition**, and that transposition is the **order of two heap pointers**: the fleet-assignment pass walks a container keyed on `fleet->Location`. There is **no original order to match**; we choose one member of its outcome set and say so (`orders.h`). The visit order is *written into the block* (list 10), so a captured run carries its own. - **The route is part of the procedure.** The same turn reached by continuation and by load gives different, individually reproducible results, because a `Colonize` task alive in the AI agent by continuation is *absent* after a load. Certified pairs live in `verify/results/saves/certified-pairs.md` as `(save, procedure, route, hashes, processes, masks)`; a hash without its route is not a claim. - **Exposure is a screen, not a decision procedure** (rule 26(c)): it says a control is *likely* to vary; it cannot certify one will agree. **So the rungs stand as written below, with one scope line each:** - **Rung A** — byte-match modulo the command stream, **on a load→turn procedure**. Canonical pair `ref-turn2 → turn3`. 204 leaves open; the RE for them is done (§3-old items 1b/1c/7). - **Rung B** — byte-match given the recorded stream. **Climbed to outcome 3 on the first rich turn (lane CV):** the replay runs and consumes the whole stream; the result is a **ranked leaf list of thirteen named subsystems, zero ungrouped** — 1,092 leaves. Two blockers sit *upstream* of the turn (§3, track 0). The tail's sim draw fires on this workload (`tscr = 253`), so Rung B needs the same generator model as Rung A — that correction is load-bearing here. - **Rung C** — `game/ai` deterministic given (save, seeds, visit order). **C-exact survives literally** on that basis; the comparison side (`--relabel-new-ids`) ships with both acceptance halves. The engine side is not built: `game/ai` mints no fleet ids yet. **What stops now.** Further RNG/determinism lanes, unless a rung is actually blocked by one. Board row count as a success signal — 89% of 444 rows measures campaign activity, not the product. The value-domain census as a work generator — it is a screen; use it when a lane needs a corpus count. **What the headline is.** Displacement (`tools/displacement.py`): **3 replaced / 7 compared / 7 modelled / 3 mapped**, unmoved for two days. And the phase catalogue: 14 of 44 driver phases, 2 of 37 tail. Those two numbers are the product. ## 2-old. The milestone as restated 2026-09-08 (kept for the record) > `sots_turn` loads a save, runs one strategic turn, and writes an autosave that byte-matches > the original's. **Lane L5 showed the original cannot always do this itself.** Three runs of `turn1-state → turn2` gave three different autosaves, differing in exactly four leaves — one AI empire's research pick and the derived checksum — with the instrument exonerated (two runs had identical hooks and still differed; `hooks=off` was a third value). The oracle stands on `ref-turn2 → turn3`, where every AI already holds a target; it fails on the neighbour, where three AI players must *choose* one. This is not a contradiction of the lockstep multiplayer design — it is the reason for it. The AI is a `StrategyClient`, not part of the sim: it decides **once**, on one machine, and its decisions travel as `TurnCommands` through the same host relay as a human's. Peers reproduce the sim's response to a command, never the decision. The sim is deterministic (lane Z's zero-residual ledger is the strategic generator, which the AI never touches); the decision layer is not, and does not need to be. So the rungs are: - **Rung A — byte-match modulo the command stream.** Every leaf except `ModCount`, `Summary.Checksum`, and what the AI's own orders touch. No AI needed. **Canonical reference pair is now `turn2-state → turn3-state`** — deterministic, oracle-hashed. `turn1-state → turn2` stays as a secondary pair with its four non-deterministic leaves masked. - **Rung B — byte-match given a recorded command stream.** `sots_turn --turn-commands ` replays the AI's (and human's) orders captured from a real turn, and the output must match that turn's autosave byte-for-byte. This is what "the engine is the game" means when the game's own decision layer is not reproducible from a save. Lane AI1's fallback becomes the primary path. **Correction 2026-09-09.** I previously told lane RB that Rung B was insulated from the strategic generator "because replay runs no AI". That was wrong, and the resolver caught it. Replay **does** run the post-combat tail (`command_replay.h:83`), and lane AC has now measured a **sim** draw in that tail — on `S+0x16c`, gated on a predicate over save fields, nothing to do with the AI. So **Rung B needs the same generator model as Rung A**: the tail's draws are ours to reproduce in both. The canonical pair is untouched by the predicate (`tscr = 252`, no fleet at a trade sector), so nothing already measured is invalidated — but a replay of any raid turn would have diverged silently under the reasoning I gave. - **Rung C — `game/ai` as a deterministic function of (save, per-client seed).** The original's AI is not non-deterministic: it is MT19937 from **one 32-bit word per AI client**, drawn once at construction from a per-process global generator — and that generator is one we already own bit for bit. Two strengths: - **C-exact.** Given the seeds captured from an original run (`sots_turn --ai-seed =`), our AI's block equals that run's block element for element and the autosave byte-matches, on **every** captured run. This is the user's framing made literal: make the AI deterministic and the old approach still pays. **Precondition:** `game/ai` must consume the client stream in the original's order and count, *including draws whose result never reaches the save* — a per-client draw ledger is the prerequisite (the site table already shows ~8 words/turn on these boards). - **C-set.** Where seeds were never captured — the historical `turn2-state.sav`, which **no process including the original can reproduce** — our canonical pick must be a member of the outcome set of the one seed-sensitive decision. Lane L4 has the XNC arm (k = 6, four observed, all inside) and **one arm unaccounted for (`BIO_GnMod`, 2 of 6)**, so "we can name all k" is **to prove**, not proved. The set is the support of a random draw, not a set of equal-score candidates. Mechanism, **measured** (lane L1, not inferred): the per-client generator at `StrategyClient+0x134` is seeded from `NextUInt` on the global generator at `0x00af6e58`, which is **live and per-process**; the seed is in no save. Only one of three empires *shows* the effect because only one reaches the candidate walk (lane L4) — the other two are stream-insensitive **by code path**. My earlier "pointer-ordered tie under ASLR" inference was wrong and is **withdrawn**. Rung B is unaffected: `--turn-commands` replay stays the no-AI path, and the canonical pair stays `turn2-state → turn3-state` until the pin probe promotes the creation turn. ## 3. The plan — three tracks, in priority order ### Track 0 — make the gate honest, then clear the two upstream blockers The gate reported 59/59 for a day while the corpus tests **skipped** (`SOTS_SAVES_DIR` unset on CT111). With the corpus it is **57/59**: the coverage ratchet broke when the corpus grew 22 → 43 (rule 27, as designed), and a writer defect appeared. `tools/gate.sh --fresh` is now the gate and it fails if a corpus test skips. | # | item | module | evidence | why first | |---|---|---|---|---| | 0a | **`ProjectName` reads `usnc` as one item; the wire is a `usp`/`usc` pair.** 12 of 43 saves round-trip 12 bytes short. | `mars/stream` | lane CV; `ad-turn27` has `usnc == 1` | no byte-match claim is possible on the `ad-*`/`ar-*` family until the writer is exact; re-prove byte-neutrality after | | 0b | **Re-type what the ratchet caught** — the spy saves (`ap-*`) and whatever else dropped `pct` below 99.99. | `mars/stream` | `tools/gate.sh` output | the ratchet did its job; honour it | | 0c | **List-5 applier cannot fire** because the memory capture cannot type 7 of 8 fields. One UI run (two sliders, two values) settles it. | shim capture | lane CV | the only list with an applier, at zero | ### Track 1 — displacement: `compared → replaced`, then `modelled → compared` Only `replaced` means the original's code did not run. The bar, written once: an un-instrumented oracle first (`certified-pairs.md` format), then the same procedure with the hook in `replace` mode and the **same hashes**, with a count from the hook proving the path executed (rule 1). | # | module | today | evidence for the attempt | target | |---|---|---|---|---| | 1a | `TechTree::ProcessResearch` + unlock cascade | compared | 35 calls / 3 workloads / 0 div; advance prediction held on a changed workload | **stays `compared`** (lane CR, 2026-09-09): our code displaced all 13 tech-tree leaves live, oracle missed by **16**, every one written by `OnTechResearched` — a different function. Next: the write-back | | 1b | `ServerSystem::ComputeTotalOutput` + `GroupOutput` | compared | 24,357 calls, 0 undeclared writes — but **13 distinct states** (rule 23) | replaced, after widening the state set | | 1c | `ServerSystem::ProcessTurn` (colony) | compared | 36 calls 0 div; 3 owned systems, gates all zero | replaced, on a save with gate traffic | | 1d | `game/nav`, `game/design`, `game/combat` retreat, `mars/vfs` | modelled, **never hooked** | host-tested only | compared — one lane each | **Ruling (2026-09-09, on lane CR's policy question).** Reaching `replaced` on research needs two event *records* whose text comes from the game's string table. That text is **game data**, not engine code: the engine already reads every catalogue, tech name and design from the user's own install through `$SOTS_DATA_DIR` / the VFS, and a string table is no different. So the answer is **not** "call the game's `PostEvent` and take a QUALIFIED caveat" — it is **load the string table through the same VFS path everything else uses** and construct the record clean. We ship no strings; the user's copy supplies them. That is the same clean-room posture as every other data dependency, and it turns the last research residual into ordinary engine work rather than a policy exception. Track 1a, next lane: `OnTechResearched` write-back (~90 tech-effect fields) + `ObservedTech` element + the two event records via the VFS string table. ### Track 2 — Rung B worklist, ranked against a real rich turn (lane CV, residual 0) | leaves | subsystem | note | |---:|---|---| | 539 | morale event ring `cme2` | **not a module** — a fixed 11-slot ring shifted by 27 appended events of three kinds (colonisation, savings band, population). It is the reporting surface of the next three rows and closes when they do | | 154 | colony growth / repair / bonuses | §3-old 1b/1c — `Pv*` is **mixed**: `PvPop`/`PvInfra` hold the uncapped projection, `PvSuit`/`PvRes` the pre-turn current | | 150 | fleet objects | needs the gather/mint (Rung C engine side) | | **79** | **`T34 RecordObservedDesigns`** | **largest independently closable stub, no upstream dependency — do this one first** | | 30 | player money + counters | §3-old item 1 chain | | 23 | `TShn` | watchpoint specified (§3-old 4) | | 23 | system↔fleet membership | with fleet objects | | 20 / 20 / 15 / 13 / 10 / 8 / 6 | ShipRecs · build queues · other system · crep/CD · id lists+generator · turnstats · trade manager | | | 5 of 6 regressions | one `ntdev` predicate | cheapest fix in the list | `ModCount` residual 2 is localised to lists 1/7/12; the only uniform reading is list 7 at 2 bumps — a prediction fitted to one observation, falsified by any capture with a colonize count ≠ 2. ### §3-old — Rung A leaf ranking (still valid for the items it names) | # | engine deliverable | module | RE input (done unless noted) | closes | |---|---|---|---|---| | 1 | ~~income chain~~ / ~~`ComputeOutput` turn path~~ **BOTH DONE** (lanes E1, C3) — the oracle hit 25/25 and the turn path is modelled (+1 leaf, 0 regressed; unspent construction cascades **two hops** into money). **The blocker moved upstream: `S11` civilian growth.** One input — imperial carrying capacity — gates the human's `Sav`, `PvSav`, `BnkEl` and `BnkPr` on both pairs. | `game/sim` | lanes E1/C3 | the budget leaves + unblocks T36 | | 1b | **`S11` civilian growth** — the new item 1 | `game/sim` | needs imperial carrying capacity | the human budget cluster | | 1c | **`Ship::RepairCost` 0x00815180** — the ninth output input, unread; taken as 0, so a colony with a damaged fleet reads high | `game/sim` | unread | output accuracy | | 2 | ~~wire `ShipCensus` + alliance mask into `BuildTurnRecord`~~ **DONE** (lane E2). T36 still blocked on two named things: the budget (item 1) and **ship construction** — a new blocker found by E2's *falsified* prediction (one census leaf short by exactly one destroyer on both pairs; no phase we run builds a ship). | `app` | lanes D2/A2 | +24 turnstats leaves when item 1 lands | | 3 | **encounter-detection draws** in the generator model | `game/sim` + `app` | lane I bound; **lane H measuring now** | last 2 RNG words | | 4 | ~~`nve` visibility record~~ **DONE** (lane E3): **closed 51, regressed 0** on the reference pair; validated on three pairs it was never fitted to — 128 leaves across five pairs, 0 regressed. The gate is `AFlags`, not `VFlags`/`EFlags`, and all three agree on nearly every corpus system so the wrong one *looks right*. Remainder: `TShn` (10/pair, gate proven NOT `AFlags`, watchpoint specified), `rcex` (6/pair, unexplained). | `game/sim` | lane E3 | **51 leaves closed** | | 5 | **post events into the save's turn bucket** (P11 + tail event phases) | `app` + `game/events` | model exists; wiring only | events subtree | | 6 | ~~`Player.Status` restore~~ **PREMISE REFUTED** (lane T2): there is **no writer between tail 31 and the autosave**. A whole-image scan finds three stores to `+0x164` — `ProcessTurn`'s encounter loop writes 1, `ResumePlaying` writes 0 on load, and `MarkPlayerTurnEnded` writes 4 from the End-Turn **submission** paths, *before* the turn runs. It needs **one predicate**, not a watchpoint — and only two corpus saves carry a non-zero `Status`, so an entry probe should name the set (rule 20). | `app` | lane T2 | S31 regression | | 7 | **tail phases** — bankruptcy decision half, turn results outbox swap, per-player sync | `game/sim` + `app` | lane K map; 34/37 stubs | tail subtree | | 8 | Rung B: **`--turn-commands` replay** from a captured block (then Rung C: `game/ai` behavioural equivalence) | `app`, then `game/ai` | lanes AI1–AI4, L4 | `ModCount`, Checksum given the stream | Rules that hold for every item: formula-held-inputs-missing is **evaluated and reported, not written**; `closed` and `regressed` reported separately, never netted; `verified` in the phase table means compared against the live game and the build fails if it drifts silently. ## 4. Future engine modules (breadth, ranked by what it unblocks) 1. **`game/ai`** — `AIPlayer`, fuzzy `IAIRule`, `StrategyAIAgent`. Essentially unread; **proven to gate Rung B**. Known: `AISystem::Write` emits nothing; the 3 `AIAgent` save blocks are cache state (typed, 10 element layouts unexercised). 2. **`game/combat`** beyond the planner — resolver (7,641 B, orchestration only) and retreat are mapped; the six pipeline callees and the tanker fold are unread; **nothing here has run under an instrument**. Combat *simulation* (`GameCombatSim`) stays parked. 3. **trade + spy managers** — all 11 virtual targets named; four can draw RNG and none has fired (no trade routes / spy program in the corpus). 4. **`net/`** — the lockstep protocol. Direct IP works in the original without GameSpy; our engine inherits determinism to the word, so a lockstep peer is mostly serialization we already own. 5. **UI** and **renderer** — not on any path; DXVK carries rendering for the original. ## 5. Workloads to manufacture (an unexercised path is a hypothesis) - **orders save → End Turn → autosave** — the Rung-A reference (human commands in-file). - **trade routes + spy program** — the four tail draw sites; `spies2`/`SysMem`/`mts` elements. - **a cruiser** (2 census leaves), **a DN platform** (the third). - **a Liir fleet** — waypoint type 2 is the Liir drive, not a node line; **a Hiver game** — types 4/5. - **a battle with `res->+4` clear and a partial retreat** — resolver + retreat instrumented; tests the `NMnx` prediction and the predicted zero-ship-fleet bug in the original. - `ResErrRoll` true entering `ProcessResearch` — save exists; needs one turn funding > ½ a tech. - expense sliders / debtor / aid — likely dead in 1.8; debtor ≈ 35 turns. Low priority. ## 6. Multiplayer revival (lane G2) - **Tier 0 (hours):** `/concurrent /join 127.0.0.1:3369` — two clients on one VM, no server. Predictions written; queued behind lane H for VM140. - **Tier 1:** availability responder, byte-specified; probably a no-op (fails open on DNS failure). - **Tier 3 (config, not code):** self-hosted OpenSpy/UniSpy with the `swordots`/`Z5gR9Z`/1381 row; fix queryport 3369 (not 6500) and add the 16 custom QR2 keys (ids 50–65). - No CD-key code exists; 333networks ruled out (v0 only). ## 7. Method and infrastructure - `guides/method-rules.md` (18 rules) is the contract for every lane brief. Rule 18 — *measure first, the lab exists* — is why the RNG ledger closed in one VM session after three static lanes. - Reachability claims older than lane V2 are **lower bounds**; re-check "no caller" findings with `tools/vtable_map.py`. Never trust Ghidra's function end (five lanes bitten). - Regenerate the determinism-oracle autosaves on VM140 (lane N overwrote them; lane H doing it). - Board rows superseded by later rows are marked `SUPERSEDED` at integration — keep doing it. - Displace the shim: each hook that reaches `replace` with 0 divergences is code the original no longer runs. Track that count; it is the honest "how much of the game is ours" number. ## Parked - Battle-load bug (thread contention) — resume only with a reason. - Renderer — after the sim byte-matches.