sots-re/campaign/backlog.md

156 lines
14 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Roadmap
Groomed 2026-09-08. Replaces the phase-0/phase-2 backlog, all of which shipped. The board
(`board.md`) is the ledger of findings; this is the plan. **The product is `sots-engine` — our
code.** RE findings are inputs to it, and the measure of progress is how much of the game runs in
our engine, verified against the original. Every number here is measured; the source is named.
## 1. The engine today (`sots-engine` main, ~33k LOC, ctest 46/46, 1,031 addresses, clean-room, MIT)
| module | LOC | what it is | verification |
|---|---:|---|---|
| `mars/parse`, `mars/text` | 1,774 | brace-block + flat-KV + CSV readers (the `Mars::Script` pull tokenizer, first-wins keys) | oracle 1,531/1,531 + 64/64 |
| `mars/vfs` | 788 | `.gob` ZIP + native-override VFS (the `./Mods` mechanism) | 10,268 files CRC-clean |
| `mars/stream` | 5,407 | save reader/writer, 386-class wire schema, `SchemaProbe`, `CoverageArchive` | **100% named coverage, byte-identical round-trip, 11 saves** |
| `mars/rng` | 244 | MT19937 + all seven entry points (`NextFloat/NextInt/Chance/NextUInt/FloatRange/IntRangeBell/GaussianRange`) | live: every draw of a turn attributed |
| `game/data`, `game/design`, `game/config` | 3,835 | weapons/sections/tech/strings catalogs, design rules, hull class, tuning | 229,042 values 0 diffs; 127/127 designs |
| `game/effects`, `game/events` | 1,531 | tech effects, `EventStorage` (turn-bucketed, dedup) | live: research slice 35 calls / 0 div |
| `game/sim` | 3,382 | economy (`ComputeBudget`, output term), research + unlock cascade, colony, movement | live: budget 4,437 / output 24,357 / research 35 / movement 45 calls — all 0 div on declared regions |
| `game/nav`, `game/combat` | 1,163 | route classifier + failure bits; retreat **planner** | offline vs 58 waypoints / 46 plans, 0 failures; retreat never instrumented |
| `app` (`sots_turn`) | 1,754 | **the standalone**: load save → walk all three drivers' phase order → write save; divergence + completion metric | 11/11 saves run; ref pair **209 → 158** leaves, closed 51 / regressed 0 |
| `shim` | 769 + hooks | proxy `binkw32.dll` + MinHook trace/compare/replace harness — **scaffolding, to be displaced** | it is the instrument, not the product |
What the engine can do end to end today: read every shipped asset and every save byte-for-byte,
run the economy/research/movement/output pieces to the original's numbers, and step a turn with
**16 of 44** driver phases and **5 of 37** tail phases modelled (10 committed). It cannot yet
reproduce a turn.
## 2. The milestone, restated (again — 2026-09-08 evening)
> `sots_turn` loads a save, runs one strategic turn, and writes an autosave that byte-matches
> the original's.
**Lane L5 showed the original cannot always do this itself.** Three runs of `turn1-state → turn2`
gave three different autosaves, differing in exactly four leaves — one AI empire's research pick and
the derived checksum — with the instrument exonerated (two runs had identical hooks and still
differed; `hooks=off` was a third value). The oracle stands on `ref-turn2 → turn3`, where every AI
already holds a target; it fails on the neighbour, where three AI players must *choose* one.
This is not a contradiction of the lockstep multiplayer design — it is the reason for it. The AI is a
`StrategyClient`, not part of the sim: it decides **once**, on one machine, and its decisions travel
as `TurnCommands` through the same host relay as a human's. Peers reproduce the sim's response to a
command, never the decision. The sim is deterministic (lane Z's zero-residual ledger is the strategic
generator, which the AI never touches); the decision layer is not, and does not need to be.
So the rungs are:
- **Rung A — byte-match modulo the command stream.** Every leaf except `ModCount`,
`Summary.Checksum`, and what the AI's own orders touch. No AI needed. **Canonical reference pair
is now `turn2-state → turn3-state`** — deterministic, oracle-hashed. `turn1-state → turn2` stays
as a secondary pair with its four non-deterministic leaves masked.
- **Rung B — byte-match given a recorded command stream.** `sots_turn --turn-commands <blocks>`
replays the AI's (and human's) orders captured from a real turn, and the output must match that
turn's autosave byte-for-byte. This is what "the engine is the game" means when the game's own
decision layer is not reproducible from a save. Lane AI1's fallback becomes the primary path.
**Correction 2026-09-09.** I previously told lane RB that Rung B was insulated from the strategic
generator "because replay runs no AI". That was wrong, and the resolver caught it. Replay **does**
run the post-combat tail (`command_replay.h:83`), and lane AC has now measured a **sim** draw in
that tail — on `S+0x16c`, gated on a predicate over save fields, nothing to do with the AI. So
**Rung B needs the same generator model as Rung A**: the tail's draws are ours to reproduce in
both. The canonical pair is untouched by the predicate (`tscr = 252`, no fleet at a trade sector),
so nothing already measured is invalidated — but a replay of any raid turn would have diverged
silently under the reasoning I gave.
- **Rung C — `game/ai` as a deterministic function of (save, per-client seed).** The original's AI
is not non-deterministic: it is MT19937 from **one 32-bit word per AI client**, drawn once at
construction from a per-process global generator — and that generator is one we already own bit
for bit. Two strengths:
- **C-exact.** Given the seeds captured from an original run (`sots_turn --ai-seed <netid>=<hex>`),
our AI's block equals that run's block element for element and the autosave byte-matches, on
**every** captured run. This is the user's framing made literal: make the AI deterministic and
the old approach still pays. **Precondition:** `game/ai` must consume the client stream in the
original's order and count, *including draws whose result never reaches the save* — a per-client
draw ledger is the prerequisite (the site table already shows ~8 words/turn on these boards).
- **C-set.** Where seeds were never captured — the historical `turn2-state.sav`, which **no
process including the original can reproduce** — our canonical pick must be a member of the
outcome set of the one seed-sensitive decision. Lane L4 has the XNC arm (k = 6, four observed,
all inside) and **one arm unaccounted for (`BIO_GnMod`, 2 of 6)**, so "we can name all k" is
**to prove**, not proved. The set is the support of a random draw, not a set of equal-score
candidates.
Mechanism, **measured** (lane L1, not inferred): the per-client generator at `StrategyClient+0x134`
is seeded from `NextUInt` on the global generator at `0x00af6e58`, which is **live and per-process**;
the seed is in no save. Only one of three empires *shows* the effect because only one reaches the
candidate walk (lane L4) — the other two are stream-insensitive **by code path**. My earlier
"pointer-ordered tie under ASLR" inference was wrong and is **withdrawn**.
Rung B is unaffected: `--turn-commands` replay stays the no-AI path, and the canonical pair stays
`turn2-state → turn3-state` until the pin probe promotes the creation turn.
## 3. Engine work to Rung A (ranked by leaves closed per lane)
| # | engine deliverable | module | RE input (done unless noted) | closes |
|---|---|---|---|---|
| 1 | ~~income chain~~ / ~~`ComputeOutput` turn path~~ **BOTH DONE** (lanes E1, C3) — the oracle hit 25/25 and the turn path is modelled (+1 leaf, 0 regressed; unspent construction cascades **two hops** into money). **The blocker moved upstream: `S11` civilian growth.** One input — imperial carrying capacity — gates the human's `Sav`, `PvSav`, `BnkEl` and `BnkPr` on both pairs. | `game/sim` | lanes E1/C3 | the budget leaves + unblocks T36 |
| 1b | **`S11` civilian growth** — the new item 1 | `game/sim` | needs imperial carrying capacity | the human budget cluster |
| 1c | **`Ship::RepairCost` 0x00815180** — the ninth output input, unread; taken as 0, so a colony with a damaged fleet reads high | `game/sim` | unread | output accuracy |
| 2 | ~~wire `ShipCensus` + alliance mask into `BuildTurnRecord`~~ **DONE** (lane E2). T36 still blocked on two named things: the budget (item 1) and **ship construction** — a new blocker found by E2's *falsified* prediction (one census leaf short by exactly one destroyer on both pairs; no phase we run builds a ship). | `app` | lanes D2/A2 | +24 turnstats leaves when item 1 lands |
| 3 | **encounter-detection draws** in the generator model | `game/sim` + `app` | lane I bound; **lane H measuring now** | last 2 RNG words |
| 4 | ~~`nve` visibility record~~ **DONE** (lane E3): **closed 51, regressed 0** on the reference pair; validated on three pairs it was never fitted to — 128 leaves across five pairs, 0 regressed. The gate is `AFlags`, not `VFlags`/`EFlags`, and all three agree on nearly every corpus system so the wrong one *looks right*. Remainder: `TShn` (10/pair, gate proven NOT `AFlags`, watchpoint specified), `rcex` (6/pair, unexplained). | `game/sim` | lane E3 | **51 leaves closed** |
| 5 | **post events into the save's turn bucket** (P11 + tail event phases) | `app` + `game/events` | model exists; wiring only | events subtree |
| 6 | ~~`Player.Status` restore~~ **PREMISE REFUTED** (lane T2): there is **no writer between tail 31 and the autosave**. A whole-image scan finds three stores to `+0x164` — `ProcessTurn`'s encounter loop writes 1, `ResumePlaying` writes 0 on load, and `MarkPlayerTurnEnded` writes 4 from the End-Turn **submission** paths, *before* the turn runs. It needs **one predicate**, not a watchpoint — and only two corpus saves carry a non-zero `Status`, so an entry probe should name the set (rule 20). | `app` | lane T2 | S31 regression |
| 7 | **tail phases** — bankruptcy decision half, turn results outbox swap, per-player sync | `game/sim` + `app` | lane K map; 34/37 stubs | tail subtree |
| 8 | Rung B: **`--turn-commands` replay** from a captured block (then Rung C: `game/ai` behavioural equivalence) | `app`, then `game/ai` | lanes AI1–AI4, L4 | `ModCount`, Checksum given the stream |
Rules that hold for every item: formula-held-inputs-missing is **evaluated and reported, not
written**; `closed` and `regressed` reported separately, never netted; `verified` in the phase
table means compared against the live game and the build fails if it drifts silently.
## 4. Future engine modules (breadth, ranked by what it unblocks)
1. **`game/ai`** — `AIPlayer`, fuzzy `IAIRule`, `StrategyAIAgent`. Essentially unread; **proven
to gate Rung B**. Known: `AISystem::Write` emits nothing; the 3 `AIAgent` save blocks are cache
state (typed, 10 element layouts unexercised).
2. **`game/combat`** beyond the planner — resolver (7,641 B, orchestration only) and retreat are
mapped; the six pipeline callees and the tanker fold are unread; **nothing here has run under
an instrument**. Combat *simulation* (`GameCombatSim`) stays parked.
3. **trade + spy managers** — all 11 virtual targets named; four can draw RNG and none has fired
(no trade routes / spy program in the corpus).
4. **`net/`** — the lockstep protocol. Direct IP works in the original without GameSpy; our engine
inherits determinism to the word, so a lockstep peer is mostly serialization we already own.
5. **UI** and **renderer** — not on any path; DXVK carries rendering for the original.
## 5. Workloads to manufacture (an unexercised path is a hypothesis)
- **orders save → End Turn → autosave** — the Rung-A reference (human commands in-file).
- **trade routes + spy program** — the four tail draw sites; `spies2`/`SysMem`/`mts` elements.
- **a cruiser** (2 census leaves), **a DN platform** (the third).
- **a Liir fleet** — waypoint type 2 is the Liir drive, not a node line; **a Hiver game** — types 4/5.
- **a battle with `res->+4` clear and a partial retreat** — resolver + retreat instrumented; tests
the `NMnx` prediction and the predicted zero-ship-fleet bug in the original.
- `ResErrRoll` true entering `ProcessResearch` — save exists; needs one turn funding > ½ a tech.
- expense sliders / debtor / aid — likely dead in 1.8; debtor ≈ 35 turns. Low priority.
## 6. Multiplayer revival (lane G2)
- **Tier 0 (hours):** `/concurrent /join 127.0.0.1:3369` — two clients on one VM, no server.
Predictions written; queued behind lane H for VM140.
- **Tier 1:** availability responder, byte-specified; probably a no-op (fails open on DNS failure).
- **Tier 3 (config, not code):** self-hosted OpenSpy/UniSpy with the `swordots`/`Z5gR9Z`/1381
row; fix queryport 3369 (not 6500) and add the 16 custom QR2 keys (ids 50–65).
- No CD-key code exists; 333networks ruled out (v0 only).
## 7. Method and infrastructure
- `guides/method-rules.md` (18 rules) is the contract for every lane brief. Rule 18 — *measure
first, the lab exists* — is why the RNG ledger closed in one VM session after three static lanes.
- Reachability claims older than lane V2 are **lower bounds**; re-check "no caller" findings with
`tools/vtable_map.py`. Never trust Ghidra's function end (five lanes bitten).
- Regenerate the determinism-oracle autosaves on VM140 (lane N overwrote them; lane H doing it).
- Board rows superseded by later rows are marked `SUPERSEDED` at integration — keep doing it.
- Displace the shim: each hook that reaches `replace` with 0 divergences is code the original no
longer runs. Track that count; it is the honest "how much of the game is ours" number.
## Parked
- Battle-load bug (thread contention) — resume only with a reason.
- Renderer — after the sim byte-matches.