From 4b3cc82efb24fc9a054d841d7991b6e57e861145 Mon Sep 17 00:00:00 2001 From: alex Date: Wed, 9 Sep 2026 09:15:36 -0400 Subject: [PATCH] lane CR: the replace-mode predictions, committed before the build --- findings/subsystems/research-replace.md | 251 ++++++++++++++++++++++++ 1 file changed, 251 insertions(+) create mode 100644 findings/subsystems/research-replace.md diff --git a/findings/subsystems/research-replace.md b/findings/subsystems/research-replace.md new file mode 100644 index 0000000..8c06c5b --- /dev/null +++ b/findings/subsystems/research-replace.md @@ -0,0 +1,251 @@ +# Taking `TechTree::ProcessResearch` + the unlock cascade from `compare` to `replace` (lane CR) + +**Lane CR, 2026-09-09. Guest: VM145 (`sots-re-win10-145`, `re@192.168.10.145`) — held by this lane.** +VM140 is untouched (it is the exclusive reference guest). + +The board's headline metric is `tools/displacement.py`: **3 replaced / 7 compared / 7 modelled / +3 mapped**. `TechTree::ProcessResearch + unlock cascade` is the best-evidenced `compared` row — 35 +calls, three workloads, 0 divergences, `tracecmp` exit 0, and an advance prediction that held on a +*changed* workload. This lane asks the only question that moves the metric: **can our code run +INSTEAD of the original's, live, with a byte-level oracle holding afterwards?** + +Everything above the `## 4.` heading was committed **before** the shim was built or staged +(rule 2). Nothing above that line is edited afterwards; corrections are made below it and named as +corrections (rule 11). + +--- + +## 1. The state, and why this one + +### 1.1 The candidate the brief names, and why it is not sufficient on its own + +`ref-turn2.sav` (`ab4ac2d7…`, = `verify/results/saves/turn2-state.sav`) is the campaign's oldest and +most reproduced oracle: one End Turn by the **load** route, `(Autosave EndTurn)` `bb4fd9ac…` / +`(Autosave)` `978041ac…`, **5 + 1 processes**, certified-pairs row 1. + +**It does not exercise a research completion.** This is not an inference from the exit code — it is +read off lane U's own instruments, both of which are in this repo: + +* `verify/results/compare/unlock-b3-t1.md`: 3 calls, 3 compared, 0 diverged, and + **`0 undeclared write(s) in 0 call(s)`**. Every completion in every measured run produces + undeclared writes on the `player` and `tree_header` guards, because that is where + `OnTechResearched`'s tech effects and the tree's order counter land. Zero undeclared writes is + zero completions. +* `verify/results/shim/unlock-shim.log`: the first three `research: cascade ok=1 …` lines of the + five-turn continuation read `completions=0 unlocked=0 otch_appends=0`. Those three lines *are* + the `ref-turn2` End Turn. + +Reading the trace rather than the verdict (rule 1) says what that turn *does* do: call 0 advances +the generator by one word (`left` 432 → 431), pushes node 144's progress 2879 → 5768, sets its flag +1 → 2 (`OverBudgetNotified`) and posts one event (`next_id` 3 → 4). Calls 1 and 2 change **nothing +at all**. So `ref-turn2` + one End Turn is a real workload for the *pass* and an empty one for the +*completion path*. **A replace-mode oracle taken there would be exactly the green verdict on a hook +that compared nothing that rule 1 exists to catch.** + +### 1.2 The state this lane uses + +**`verify/results/saves/turn3-state.sav`** — 67,219 B, sha256 +`978041acd168b56ed8eb3f5e42e78d5e70eae6e6517d75e659a5eb7ca3d60921`, `Summary.Turn = 3`, +`Sim.ModCount = 24`. It is *the output of certified-pairs row 1*: the `(Autosave).sav` that +`ref-turn2` + one End Turn produces. It is already on VM145 under that hash as `(Autosave).sav`. + +**Procedure: one End Turn. Route: LOAD.** Stated because the route is part of the pair +(rule 26(c) fourth term; lane BS): the same turn reached by continuation and by load is two +different, individually reproducible things, and this lane deliberately does **not** inherit the +continuation's turn-4 numbers as an assumption — it predicts them and lets the load route falsify +them (P1). + +Why this turn: in lane U's five-turn continuation the **fourth-turn** call is the completing one. +The call that enters with `turn = 4` allocates 2,898 points to tech 144, completes it, and runs the +cascade — `completions=1 unlocked=3 otch_appends=1` in the shim log, `next_id` 5 → 7 and +`observed_techs` 440 → 484 bytes in the trace. `turn3-state.sav` is the state that turn starts from. + +### 1.3 The exposure screen (rule 26(c)) — a screen, not a decision procedure + +Read from `turn3-state.sav` with `verify/save-reader/save_reader.py`. Eight players; per lane BP the +engine builds AI clients for net ids **32, 496, 512** only, and the four factions at 528–576 carry +their signature **inertly**. + +| player | species | `ResTNm` | `Status` | AI client | pick-turn signature? | +|---|---|---|---|---|---| +| 16 (human `re`) | 0 | `''` | 4 | **no** (local human) | **false positive** — no client | +| 32 | 2 | `IND_Waldo` | 4 | yes | **no** — mid-research | +| 496 | 0 | `DRV_PlsFiss` | 4 | yes | **no** — mid-research | +| 512 | 2 | `BIO_GnMod` | 4 | yes | **no** — mid-research | +| 528, 544, 560, 576 | 4 | `''` | 0 | **no** | **false positive** — no client, `Status 0` | + +**No player that actually runs enters this turn with an empty research target.** That is the best +the screen can say. It cannot say the control will agree — rule 26(c)'s own retraction is explicit +that `candidatesTried` is not readable from a save, that there is a combat term, and that the load +route is a fourth term that is not a predicate on the save at all. **So the screen is why this turn +is worth trying, and the two-process control in §4.1 is the only thing that decides it.** + +The lineage where exposure is saturated (`ad-…`/`ar-…`, every turn carrying a signature) is +deliberately not used here. + +### 1.4 What is being asked of `hooks=off` + +Three fresh-process runs are planned before any replace result is read: + +| run | config | what it establishes | +|---|---|---| +| **C1**, **C2** | `shim.cfg.croff` (`hooks=off`, the proxy loads and installs nothing) | the control agrees with itself across two fresh processes (rule 26) | +| **N** | `shim.cfg.crcompare` (one detour, `compare`) | the instrument is behaviour-neutral on *this* save and route (rule 19), **and** it is where the completion count comes from the hook itself | + +Only then the replace runs. C1/C2 use the *same proxy DLL* as every measured run, so the only +difference between control and measurement is the config (lane CB's shape). + +--- + +## 2. The four bytes of every float literal on this path (rule 23) + +Two float literals are on the research pass's arithmetic path, and **both are widened `float`s**. +Read here directly out of `dumps/sots.exe` (PE image base `0x00400000`, `.rdata` at `0x009dd000` +file offset `0x5dbe00`), not taken from an earlier note: + +| VA | bytes (LE) | value as `double` | is it `(double)float`? | engine constant | +|---|---|---|---|---| +| `0x009e20c8` | `00 00 00 a0 99 99 e9 3f` | `0.80000001192092896` | **yes** (low 29 bits zero) | `kEarlyCompletionRatio = 0.800000011920929` | +| `0x009e5060` | `00 00 00 a0 99 99 a9 3f` | `0.05000000074505806` | **yes** | `kDecayFraction = 0.05000000074505806` | + +Both source literals were re-parsed and re-packed: `0.800000011920929` → `000000a09999e93f` and +`0.05000000074505806` → `000000a09999a93f`. **Bit-identical to the image.** So the two constants +`sots-engine/src/game/sim/research.cpp` relies on are exactly the image's, not the exact decimals +that bit the money chain twice. + +Two literals that are **not** on this path, and why: + +* the tech-cost multiplier (`1.0 − 0.25·n`, floor `0.25`) — `ours` never evaluates it. It calls the + game's own read-only `TechTree::Cost` for every cost it needs, in both modes (`g_env.cost`). This + is a genuine, declared dependency on the original: the effective cost is **not** displaced. +* `ResearchSpendFloor` / `Ceiling` are integer `×50/100` and `×150/100` with a 32-bit wrapping + multiply — no float involved. + +--- + +## 3. Predictions, committed before the build + +Falsification symptoms are given for each. Predictions are made **per site** and not at a bracket +total (rule 23's 2026-09-09 corollary). + +### P0 — the control agrees with itself +`turn3-state.sav`, one End Turn, load route, `hooks=off`, **two fresh processes** produce identical +`(Autosave EndTurn).sav` and `(Autosave).sav`. +*Falsified if:* the two processes differ. Then this workload is exposed, the screen in §1.3 was +insufficient (which is what rule 26(c)'s retraction predicts is possible), and the honest move is to +report that and pin the source of variation — **not** to pick the run that suits. + +### P1 — the turn is not quiet, and the load route reproduces the continuation's turn-4 call +Exactly **3** `ProcessResearch` calls, one per AI client, in the same owner order. Allocations: +**A → {tech 144, 2898 points}**, **B → {tech 90, 0 points}**, **C → {tech 9, 0 points}**. Exactly +**one completion** (tech 144) and the cascade unlocks exactly **three** nodes: + +| node | `state` | `cost_rp` | `turn_available` | +|---|---|---|---| +| 132 | 0 → 2 | `INT_MAX` → **10000** | −1 → **4** | +| 136 | 0 → 2 | `INT_MAX` → **16000** | −1 → **4** | +| 142 | 0 → 2 | `INT_MAX` → **8000** | −1 → **4** | + +*Falsified if:* a different allocation, a different completion count, or a different unlock set. That +would be a **route** result — the load route not reproducing the continuation — and it is worth as +much as the replace result. It is exactly what rule 26(c)'s fourth term says can happen. + +### P2 — the arithmetic, hand-computed on the boundary +For node 144: `cost = 5000`, `lo = ResearchSpendFloor = 2500`, `hi = ResearchSpendCeiling = 7500`. +Entering progress **5768**, allocation **2898**: + +* `spent = min(2898, 7500 − 5768) = 1732` — the cap binds, so this call lands **exactly on the + ceiling**, which is the boundary rule 23 says to test by hand rather than trust a compare on. +* `progress 5768 → 7500`, `overbudget 0 → 1166`. +* `progress (7500) < hi (7500)` is **false**, so the odds/roll branch is skipped: **zero RNG draws**. + I predict `rng.left` is **unchanged at 413** across all three calls and the `mt` digest does not + move on any of them. +* `ratio = (float)(7500/5000) = 1.5`, and `1.5 < 0.800000011920929` is false → **not** completed + early → `flag` **unchanged**. +* `state 3 → 4`, `turn_researched −1 → 4`, `order −1 → 22`, and the tree's order counter + `22 → 23`. + +*Falsified if:* any single one of these moves. A wrong `hi` shows up as a wrong `overbudget`; a +wrong early-completion literal shows up as `flag` moving. + +### P3 — the instrument is neutral, and the completion is counted by the instrument +Run N (`crcompare`, one detour) produces autosaves **byte-identical to C1/C2**, and `tracecmp` +reports **3 calls / 3 compared / 0 diverged / exit 0**. The hook's own per-call log line reads +`completions=1 … unlocked=3 otch_appends=1` on the completing call and all-zero on the other two. +Undeclared writes: **6 spans in 1 call** — +`player +0x10c/3`, `player +0x110/3`, `player +0x114/3`, `player +0x124/3`, `player +0x294/4` +(`ResT`, the research target being cleared) and `tree_header +0x20/1` (the order counter). + +*Falsified if:* the autosaves move (then rule 19 bites and nothing below is readable), or the +undeclared-write set differs (then the workload or the model moved). + +### P4 — replace with the cascade OFF (`crreplace0`, the shipped behaviour) diverges +The autosave is **not** byte-identical. Predicted diverging leaves, by name: + +1. `Player[32]` tech tree: node 144 `turn_researched` and `order` unstamped (−1/−1); nodes 132, 136 + and 142 unchanged at `state 0`, `cost_rp INT_MAX`, `turn_available −1`. +2. the tree's completion-order counter left at 22. +3. `Player[32]` events: `next_id` short by **2**, two event records absent. +4. `Player[32]` `otch`: one `ObservedTech` element absent (**44 bytes**). +5. `Player[32]` scalar fields at `+0x10c`, `+0x110`, `+0x114`, `+0x124` unchanged, and `ResTNm` still + reading `IND_Waldo` because `ResT` was never cleared. + +What **will** be right even here: node 144's `state = 4`, `progress = 7500`, `flag`, the decay sweep +over every other node, and the `overbudget` accumulator. + +### P5 — replace with the cascade ON (`crreplace1`) diverges, by strictly less +Same binary, one config line different (`research.replace_cascade=on`; see §3.1). Predicted: +**strictly fewer diverging leaves than P4**, with items 1 and 2 of P4 **closed** — node 144 stamped +`turn_researched 4` / `order 22`, nodes 132/136/142 at `state 2` with `cost_rp` 10000/16000/8000 and +`turn_available 4`, counter at 23 — and items 3, 4 and 5 **still open**. + +### P6 — the headline, stated in advance +**`game/sim/research` does NOT move from `compared` to `replaced` this session, and the named cause +is `ServerPlayer::OnTechResearched`, not the research model.** The research pass itself — the spend +cap, the odds branch, the completion, the decay sweep, and (with the flag on) the whole +`SetResearched` cascade — is displaceable and I expect it to reproduce the original's TechTree state +exactly. What blocks the oracle is the *callback*: it posts events whose text comes from the game's +string table, appends an `ObservedTech` element `ours` decides but does not construct, and writes +~90 `ServerPlayer` fields that are a different subsystem's milestone (B2) and a separate `compared` +row on the board. A replace of the research pass is **gated on displacing `OnTechResearched`**, and +that is a statement about a boundary, not about the research model. + +*Falsified if:* `crreplace1`'s autosave is byte-identical to the oracle. That would mean the +callback's writes do not reach the save on this workload, P6 is wrong, and the row moves to +`replaced` — which is the outcome I would rather have and do not expect. + +### P7 — the numbers I expect to be able to quote honestly afterwards +**1** distinct completion, **1** distinct tech (144), **3** unlocked nodes, **3** calls of which +**2** allocate zero points and write nothing at all. That is thin, and it will be reported as +thin (rule 15/23) whatever the verdict. + +### 3.1 The one engine change this lane makes, and why it is not a thumb on the scale + +`sots-engine` worktree `wip/cr` (rule 21), branched from `main` at `e7e2bd6`. + +* **`research.replace_cascade=on|off`, default `off`.** With it off the binary behaves *exactly* as + `main` does today, so P4 measures the shipped behaviour rather than a straw man. With it on, a + **replace**-mode call also writes the four `TechNode` words `SetResearched` stamps and the tree's + completion-order counter. It does **not** post events, does **not** move the `ObservedTech` vector + and does **not** apply a single tech effect — those stay gated on compare mode and stay declared + unmodelled. The flag exists so P4 and P5 differ by a **config line and not by a binary**, which is + what makes the leaf difference between their autosaves attributable to the cascade. +* **A per-call log line in every mode.** The old line was gated on the cascade having run, so a + replace run with the cascade off had no counter at all and "a completion happened" could only be + inferred from the save — the artefact under test. It now prints `steps`, `completions` (counted + from the pass's own step results), `overbudget` and the cascade counters, in compare **and** + replace. This is the instrument answering the brief's "show it fired — a count, from the hook + itself". + +No new binary facts: `ghidra/addresses.d/cr.json` is **not** created, because this lane reads no new +address. Every offset used already exists in the generated header. Host build: 253 targets, **59/59 +tests pass** (the denominator is the healthy 59, per rule 24's signature check); `clean_room_check` +OK; `tools/check_shim_configs.py` OK with 27 registered hooks and all four CR configs +(`croff`, `crcompare`, `crreplace0`, `crreplace1`) — the three `hooks=trace` ones marked +`# exhaustive` and naming all 27. + +--- + +## 4. Results + +*(added after the runs; nothing above this line is edited)*