# B3 — `TechTree::ProcessResearch` old-vs-new, with the RNG state as a declared region **Result (2026-09-08): verified on the live game.** Trace: 3 calls, `tracecmp.py` exit 0. Compare: **15 calls, 13 with zero divergence**; the two that diverged are both completion turns and every diff is in a field `TechTree::SetResearched` owns — the documented scope boundary — plus one extra RNG draw on one of them (below). The RNG post-state matched on 14 of 15 calls, including every roll. Replace: our research pass fed the game for a whole turn from `ref-turn2.sav` and the resulting 609 KB autosave is **item-for-item identical to the oracle except for one missing event** (`EVENT_RESEARCH_OVERBUDGET`), so the End-Turn oracle does **not** pass yet — see "The replace-mode gap". Game left at the main menu in `hooks=trace`. Build `b3-b791392-20260908T0353Z`, staged in `/srv/re-lab/shim/dist-b3`, deployed to the VM from `C:\SOTS\shimdist-b3`. Evidence in `/srv/re-lab/shim/traces/b3-*`. The point of the target: one call exercises the MT19937, the completion-odds formula and the Zuul double roll at once, so a match validates all three together — and its RNG consumption is observable, which makes the draw *count* checkable rather than merely plausible. ## What was hooked | hook name (record `hook`) | RVA | prototype | |---|---|---| | `Game::TechTree::ProcessResearch` | 0x001876c0 | `void (TechTree*, Mars::RNG*, vector<{TechDef*,int}>*, int*)` | `__thiscall`, `[verified]`, so it goes through `Hook<>` with `CallConv::Thiscall` (M2's addition). Source: `src/shim/hooks/research.{h,cpp}`, installed from `src/shim/main.cpp` after the M1/M2 hooks. There is exactly one caller (inside `ServerPlayer::ProcessTurn`), which fires once per player per turn. **The second parameter was a `?` in the address contract; it is the `Mars::RNG` object.** The call site loads it from `StrategyServer+0x16c` and the function re-bases it with `+4` before every draw. That is now in `ghidra/addresses.json` along with the tree/node offsets, the generator layout and `TechTree::Cost`, and regenerated into `sots_addresses.h`. ### Region model Three kinds of region, all snapshotted before the original runs: | region | size | describer | |---|---|---| | `rng` | 0x9cc | `{vptr:ptr, mt:bytes(2496 → sha256+head), left:i32, next_index:i64}` | | `overbudget` | 4 | `{v:i32}` | | `node[i]` | 0x34 each, one per non-null slot | `{def, tech_id, kids_*, unk10, state, cost_rp, progress, turn_available, turn_researched, order, flag, unk30}` | `next` is a heap address, so it is reported as its index into `mt` — which is what it means, and what survives being written by a reimplementation. `left` alone already pins the stream position (`next == &mt[624 - left]` always), so the index is a cross-check, not the evidence. Args carry the evidence a golden log needs to replay offline: `tree`, `owner`, `species`, `node_count`, `rng`, `rng_left_in`, the `alloc` list as `{tech_id, points}`, `overbudget_in`, and **`fpu_cw`** — the x87 control word in force for the call (see "Float mapping" below). Declaring *every* node, not just the ones we expect to change, is deliberate: it is what proves `ours` neither misses a write nor makes an extra one. The cost is size — a compare record is roughly 200 nodes × three snapshots, so the b3 configs turn every other hook off. ### The comparison design `ProcessResearch` consumes RNG, so comparing two implementations that draw from different streams would diverge for a reason that has nothing to do with the formulas. Instead: 1. the generator object is a **declared region**, so its `mt[624]` + `left` are snapshotted before the original runs, alongside the tech-tree state; 2. the original runs and advances the real generator; 3. `ours` runs on the scratch copies, and seeds a `mars::rng::MT19937` with `load_state()` from the *pre-call* snapshot — so both implementations read the identical stream; 4. `ours` writes its final generator state back into the scratch copy, so the diff compares the **post-call RNG state** as well as the outputs. If the post-states match, we consumed the same words in the same order. That is the check with teeth: getting the odds right but drawing twice (or not drawing at all) moves `left` and the hash. `tests/mars_stream/test_rng.cpp` pins the property offline, including the negative case (one draw too few leaves a different state). `next` is rebuilt by `ours` against the **live** generator address so the describer's index arithmetic reads the same on both sides; `ours` never dereferences it. ## What `ours` covers, and what it deliberately does not `ours` is `sots::sim::ProcessResearchTurn` (new, `src/game/sim/research.{h,cpp}`) plus a thin shim adapter. It reproduces exactly the words the hooked function writes itself: * the allocation loop — spend window, spend, `*overbudget`, progress, the roll, the completion decision, the over-budget flag, the "completed early" flag, `state = 4` on completion; * the decay sweep over every available node. It does **not** reproduce `TechTree::SetResearched`, which the original calls on completion: the turn/order stamps, the child-unlock cascade and the owner's tech-effect callback. That is its own milestone, and the callback writes live player state that compare mode must never touch. So: > **A turn on which a tech completes is expected to diverge**, in the completing node's > `turn_researched` / `order` and in the child nodes `SetResearched` unlocks. A turn on which > nothing completes — the overwhelmingly common case — must match everywhere. > *(Borne out: the two divergent calls of the live compare carry exactly those fields and > nothing else. The prediction that the cascade also makes no RNG draw was **half wrong** — one > of the two completions consumed an extra word inside the owner's tech-effect callback; see > "The extra draw" below.)* It also does **not** post the events the pass raises — and that turned out to matter: the over-budget branch sets `node.flag = 2` *and* pushes an `EVENT_RESEARCH_OVERBUDGET` onto the owner's event list, and only the flag is modelled. The event list was never a declared region, so compare mode is blind to it; the End-Turn oracle is what caught it. See "The replace-mode gap". The effective cost of a node comes from the game's own `TechTree::Cost` (read-only: it only reads `costRP`, the def and the owner, and calls the read-only cost-multiplier helper). The cost multiplier is a separate, medium-confidence formula and not what this milestone measures; this is the same delegation M2 makes to `LoadWeapon`. `ours` receives the live tree pointer and treats it as read-only — every node it writes is a scratch copy. `ours` also works in `replace` mode, where no `regions`/`rebind` ran: it then reads the tree's own node vector and the live generator. The per-call statics carry a flag that is cleared at the end of every `ours`, so a replace call can never inherit a stale compare mapping. ## Float mapping — the headline finding **A draw is `y / (2^32 − 1)`, not `y × 2^-32`.** The multiplier in the image is the double `0x3df0000000001000`, which is exactly `1/4294967295`; the constant next to it is the `+2^32` unsigned fix-up applied after a sign-extending integer load. So: * `MT19937::kUnitScale` is now `1.0 / 4294967295.0`; * the range is **closed**: `y == 0xffffffff` maps to exactly `1.0`, not to just below it; * the value is left in `st(0)` and the caller narrows it — every consumer in the strategic sim stores it to a 4-byte float first, which is what `next_float()` models. Honest caveat, because it decides how to read a passing compare: the old and new divisors differ by 2^-32 relative, far below a float32 ulp. Measured over 10^6 draws they give a **different float 0.78 % of the time**, and they flip an actual research completion decision (roll vs an odds of 1/3) **0 times in 10^6** — the expected rate is about one in two billion. So the compare cannot prove the divisor; the disassembly and the constant's bit pattern do, and the compare's job is the rest. The one thing the binary cannot settle is the x87 **precision-control** field at run time. At the MSVC default (53-bit, `cw = 0x027f`) the multiply rounds to double and the caller's store rounds again — that is what `next_float()` does. A Direct3D 9 device created without `FPU_PRESERVE` leaves 24-bit precision, in which the fix-up and the multiply each round to 24 bits; `MT19937::float_from_pc24()` models that, and the two differ for **0.094 %** of words — last-bit only. The hook records `fpu_cw` on every call, so the first trace settles it; if it comes back 24-bit the change is to route `next_float` through `float_from_pc24` and to compute the odds the same way, and nothing else in the milestone moves. `float10` in the decompile is just the i386 float return ABI and was not chased. ## Bugs found and fixed in our implementation All five are read off the instruction sequence, not tuned to make anything match. 1. **The unit divisor** (above): `2^-32` → `1/(2^32 − 1)`. 2. **`NextInt` is inclusive.** The rejection mask is built from `n` itself, not `n − 1`, and the loop re-draws while the masked word is **greater than** `n` — so the result is uniform on `[0, n]`, one value wider than we had. The bound is also passed **by pointer**, which is why the prototype had stayed `[unverified]`. `next_int(n)` is now `next_int_inclusive(n)` and `IRandom::NextInt` is `IRandom::NextIntInclusive`, so every call site had to be re-read. 3. **`spend` has no floor at zero.** The original is a plain signed `min(points, hi − progress)`; ours clamped it at 0, which would have hidden a negative spend (and understated `*overbudget`) whenever progress was already past the 150 % ceiling. 4. **`odds` is a float32.** The original computes `(progress − lo) / hi` on the x87 and stores it to a 4-byte slot before the comparison; ours kept it in double. Likewise the roll, the Zuul minimum, and the `progress / cost` ratio used for the "completed early" flag. 5. **Two constants are widened float literals, not decimals.** The decay fraction is `(double)0.05f = 0.05000000074505806` and the early-completion threshold is `(double)0.8f = 0.800000011920929`. Both sit on a truncation/compare boundary. Two smaller ones in the same pass: the decay guard is `progress != 0`, not `progress > 0`; and a node whose cost is still `INT_MAX` is **not** special-cased by the original — it feeds that straight into the 5 % multiply, which wipes any progress out. The 50 %/150 % bounds are a 32-bit multiply that wraps near `INT_MAX` rather than a widening one. ## Host tests `ctest` 26/26. New coverage: * `mars_rng_unit` — the standard MT19937 vectors (seed 5489, and the 10000th output) were already there; added the unit mapping word by word (`0 → 0`, `0xffffffff → 1.0`, the high-bit fix-up path), the measured rarity of the divisor difference, the PC24 variant's bound, `cover_mask`, the inclusive integer bound (including that `n` itself is reachable and that `n == 0` still consumes a word), and **the compare design end to end**: snapshot → original draws → ours seeded from the snapshot reproduces the values and the post-state, with a negative case that one draw too few does not. * `game_sim_research` — 121 checks. Added the spend window (truncation, the floor/ceiling clamps, the `INT_MAX` edge), that the odds are float32, that `spend` goes negative rather than clamping, the early-completion boundary at exactly 80 % versus one point below, and `ProcessResearchTurn` (order of the passes, the funded node decaying too, hidden slots never touched, an out-of-range entry consuming no draw, `overbudget` accumulating). Cross-build: `b3-81218c7-dirty-20260908T0311Z`, exports 66 names identical to `binkw32.dll`, staged in `/srv/re-lab/shim/dist-b3`. `tools/clean_room_check.sh` OK. ## Gotchas 1. **Two different `this` pointers for one object.** `Twist`, `NextFloat` and `NextInt` take `&mt` — the object **plus 4** — so *their* `this+0x9c0/+0x9c4` are `next`/`left`, while the object's own layout is `{vftable @+0, mt[624] @+4, next @+0x9c4, left @+0x9c8}` = 0x9cc bytes. The address contract used to state both readings as if they were one; it now says which is which. Getting this wrong shifts every generator field by a word. 2. The save blob is 0x9c4 bytes = `mt[624]` + `left`, i.e. it **skips** `next`, which sits between them in the object. `left` is sufficient because `next == &mt[624 - left]` and the original's `Read` recomputes it. 3. `Region::name` is a `const char*` held for the whole call, so the per-node name strings are `resize`d once up front and never grown — a reallocation would dangle every name already pushed. 4. Per-call state is kept in statics between `regions()` → `rebind()` → `ours()` (M1's concession). Safe here because the turn pass is single-threaded and `ProcessResearch` has one caller and never nests; do not copy the pattern to a re-entrant hook. 5. A compare record is large (one region per tech node). The staged configs (`shim.cfg.b3{trace,compare,replace}`, in the dist) switch every other hook off for that reason; `trace.inline_max` stays at 256 so the 2496-byte state block is hashed rather than inlined. 6. `TechDef`'s first word is the tech id and it indexes `TechTree+0x10`; the allocation vector's stride is 8 (`{TechDef*, int}`). Both are in the address contract now rather than inferred. ## Runs (`/srv/re-lab/shim/traces/`, build `b3-b791392-20260908T0353Z`) | file | mode | calls | result | |---|---|---|---| | `b3-trace-golden.jsonl` | trace | 3 | `tracecmp.py` exit 0, 0 invalid | | `b3-compare.jsonl` | compare | 15 | **13 clean, 2 diverged** (both completion turns; see below) | | `b3-replace-autosave.sav` vs `b3-orig-autosave.sav` | replace | 3 | `(Autosave EndTurn).sav` = `bb4fd9ac…` (oracle); `(Autosave).sav` = `fd071d18…` ≠ oracle `978041ac…` | | `b3-replace-save-diff.txt` | | | the whole difference: one missing event | | `b3-shim.log`, `b3-final-menu.png`, `b3-compare-turn7.png` | | | banners, final state | Procedure: `ref-turn2.sav` → End Turn ×1 (trace), ×5 (compare), ×1 (replace and the oracle re-check). Each End Turn produces three calls: the AI Tarkas player (which actually researches), the human, and a second Tarkas tree, the last two with a zero-point allocation. ### What the compare proves Per call, the trace records the draw count as `left` before minus after: | call | species | alloc `{tech, points}` | draws (orig / ours) | outcome | |---|---|---|---|---| | 0 | Tarkas | {144, 2889} | 1 / 1 | over-budget event raised, flag 1 → 2 | | 1, 2 | Human, Tarkas | {90, 0}, {9, 0} | 0 / 0 | zero spend → odds 0, roll 1, no draw | | 3 | Tarkas | {144, 2898} | 0 / 0 | reached the 150 % ceiling → guaranteed, **no draw**, completed | | 6 | Tarkas | {142, 3064} | 1 / 1 | rolled and failed (odds −0.078, below 50 % of cost) | | 9 | Tarkas | {142, 3074} | **2 / 1** | rolled at odds 0.178 and **completed early** (6138/8000 = 0.767) | | 12 | Tarkas | {9, 3086} | 1 / 1 | rolled and failed | Every integer the pass computes matched: progress, the spend cap, `*overbudget`, and the flag on every node, including the three interesting branches — the over-budget event flag (call 0), the guaranteed completion at the ceiling with no draw at all (call 3), and the "completed early" flag 1 → 0 below 80 % of cost (call 9). On call 9 ours drew the same word as the original (the completion decision and the early flag both match), so the extra draw is the original's, not a missing one of ours. **The two divergences.** Both are completion turns, and both carry only `SetResearched`-owned fields: `order` and `turn_researched` on the completing node, and `cost_rp` / `state` / `turn_available` on the children it unlocks (three children on call 3, one on call 9). That is exactly the set this milestone declared out of scope in advance. **The extra draw (call 9).** The original consumed one word more than ours *after* the completion roll. Call 3 also completed a tech and consumed no extra word, so this is not "SetResearched always draws" — it is the owner's tech-effect callback for that particular tech. That callback is `ServerPlayer::OnTechResearched`, which is the B2 lane's target; B2 should expect to find at least one effect that rolls. ### The replace-mode gap — one missing event Replace mode ran our pass for a whole turn. `(Autosave EndTurn).sav` (the pre-turn state) is the oracle's `bb4fd9ac…`, but the post-turn `(Autosave).sav` is `fd071d18…` and not the oracle's `978041ac…`. Dumping both trees and normalising the offsets, the **entire** difference in 40,300 items is: ``` Player > Events > Events > . > Events: one entry present in the original, absent in ours EvEID 3 EvDsc "Research Over Budget" EvMsg "Research for Waldo Units has gone overbudget." EvImg "EVENT_RESEARCH_OVERBUDGET" EvLoc 0 EvPos {inf,inf,inf} EvAct 1 EvCID 0 + the two counters that index it: EvNxID 4 → 3 and the list length 2 → 1 ``` Nothing else — not one number the research pass computes — differs. So the formulas are right and the milestone under-scoped itself: `docs/B3.md` promised the over-budget *flag* and treated the event beside it as part of the same branch without saying it was excluded. It is not modelled, `ours` writes `node.flag = 2` and stops there, and compare mode could not see it because the player's event list was never a declared region. Two consequences worth carrying forward: * **The oracle catches what a region-scoped compare cannot.** Compare mode was clean on exactly the call that produced the missing event, because the event lands somewhere we did not declare. A clean compare bounds the state you declared, nothing more. * To close this, `ours` needs to post the event through the game's own event API (the same call the original makes right after setting the flag), the way M2 delegates to `LoadWeapon` — the message text is composed from the tech name, so it cannot be synthesised from our side alone. Until then, **replace mode is not oracle-clean** and should not be used as a determinism check for this hook. ## Corrections this run made to the notes 1. **The current research target has state 3, not 2 — so it does not decay.** Every funded node in the trace is state 3 (`node[144]`, `node[142]`, `node[9]`), and the decay loop's guard is `state == 2`. `formula-gaps.md` Q8 read the loop correctly but concluded the funded node decays; it does not, because the tree marks the selected target with a different state. Our code was already right (it compares against `Available == 2`); the prose was wrong. The state set observed in one tree: 164 hidden (0), 7 parent-researched (1), 23 available (2), 1 current target (3), 22 researched (4). 2. **`fpu_cw` = `0x127f` at call time** (`0x027f` at DLL init): precision control = 10b = **53-bit double**, rounding = nearest. Bit 12 is the legacy infinity-control flag and has no effect. So the MSVC default path is the right one and `next_float()` — double product, then the caller's narrowing to float32 — models the original exactly. `float_from_pc24()` stays as an unused contingency; nothing needs to change. 3. **Nodes per tree = 293**, not the 196 the tech-name table suggested. ## Coverage this run did not reach * **No Zuul.** `ref-turn2.sav` has two players, `re` (human) and `Fane Lao` (AI), and the trees seen are species 0 (Human) and 2 (Tarkas). The double roll is therefore still verified by disassembly and host tests only, never by behaviour. A compare from a save with a species-5 player would close it; the check is simply that `left` drops by 2 on a roll instead of 1. * **The decay sweep never fired.** No node in any traced tree was state 2 with non-zero progress — idle partially-researched techs do not occur in this save, and the current target is state 3. `DecayAllResearch` is covered by host tests only. * Only one lab-accident-free, single-target allocation shape was seen (one entry per call). ## If the replace-mode gap is closed 1. Add the event post to `ours` (see above), rebuild, restage `/srv/re-lab/shim/dist-b3`. 2. Deploy from `C:\SOTS\shimdist-b3`, copy `shim.cfg.b3replace` over `C:\SOTS\shim.cfg`, relaunch, load `ref-turn2.sav`, End Turn once, and re-check the oracle: `(Autosave).sav` = `978041ac…`, `(Autosave EndTurn).sav` = `bb4fd9ac…`. 3. Restore the default `shim.cfg` (`hooks=trace`) and leave the game at the main menu. Note that replace mode also stops modelling completions the moment one happens on a later turn: `ours` does not run the unlock cascade, so a replace run must not be pushed past the first completion until `SetResearched` is a milestone of its own.