sots-engine/docs/B3.md

20 KiB
Raw Permalink Blame History

B3 — TechTree::ProcessResearch old-vs-new, with the RNG state as a declared region

Result (2026-09-08): verified on the live game. Trace: 3 calls, tracecmp.py exit 0. Compare: 15 calls, 13 with zero divergence; the two that diverged are both completion turns and every diff is in a field TechTree::SetResearched owns — the documented scope boundary — plus one extra RNG draw on one of them (below). The RNG post-state matched on 14 of 15 calls, including every roll. Replace: our research pass fed the game for a whole turn from ref-turn2.sav and the resulting 609 KB autosave is item-for-item identical to the oracle except for one missing event (EVENT_RESEARCH_OVERBUDGET), so the End-Turn oracle does not pass yet — see "The replace-mode gap". Game left at the main menu in hooks=trace.

Build b3-b791392-20260908T0353Z, staged in /srv/re-lab/shim/dist-b3, deployed to the VM from C:\SOTS\shimdist-b3. Evidence in /srv/re-lab/shim/traces/b3-*.

The point of the target: one call exercises the MT19937, the completion-odds formula and the Zuul double roll at once, so a match validates all three together — and its RNG consumption is observable, which makes the draw count checkable rather than merely plausible.

What was hooked

hook name (record hook) RVA prototype
Game::TechTree::ProcessResearch 0x001876c0 void (TechTree*, Mars::RNG*, vector<{TechDef*,int}>*, int*)

__thiscall, [verified], so it goes through Hook<> with CallConv::Thiscall (M2's addition). Source: src/shim/hooks/research.{h,cpp}, installed from src/shim/main.cpp after the M1/M2 hooks. There is exactly one caller (inside ServerPlayer::ProcessTurn), which fires once per player per turn.

The second parameter was a ? in the address contract; it is the Mars::RNG object. The call site loads it from StrategyServer+0x16c and the function re-bases it with +4 before every draw. That is now in ghidra/addresses.json along with the tree/node offsets, the generator layout and TechTree::Cost, and regenerated into sots_addresses.h.

Region model

Three kinds of region, all snapshotted before the original runs:

region size describer
rng 0x9cc {vptr:ptr, mt:bytes(2496 → sha256+head), left:i32, next_index:i64}
overbudget 4 {v:i32}
node[i] 0x34 each, one per non-null slot {def, tech_id, kids_*, unk10, state, cost_rp, progress, turn_available, turn_researched, order, flag, unk30}

next is a heap address, so it is reported as its index into mt — which is what it means, and what survives being written by a reimplementation. left alone already pins the stream position (next == &mt[624 - left] always), so the index is a cross-check, not the evidence.

Args carry the evidence a golden log needs to replay offline: tree, owner, species, node_count, rng, rng_left_in, the alloc list as {tech_id, points}, overbudget_in, and fpu_cw — the x87 control word in force for the call (see "Float mapping" below).

Declaring every node, not just the ones we expect to change, is deliberate: it is what proves ours neither misses a write nor makes an extra one. The cost is size — a compare record is roughly 200 nodes × three snapshots, so the b3 configs turn every other hook off.

The comparison design

ProcessResearch consumes RNG, so comparing two implementations that draw from different streams would diverge for a reason that has nothing to do with the formulas. Instead:

  1. the generator object is a declared region, so its mt[624] + left are snapshotted before the original runs, alongside the tech-tree state;
  2. the original runs and advances the real generator;
  3. ours runs on the scratch copies, and seeds a mars::rng::MT19937 with load_state() from the pre-call snapshot — so both implementations read the identical stream;
  4. ours writes its final generator state back into the scratch copy, so the diff compares the post-call RNG state as well as the outputs.

If the post-states match, we consumed the same words in the same order. That is the check with teeth: getting the odds right but drawing twice (or not drawing at all) moves left and the hash. tests/mars_stream/test_rng.cpp pins the property offline, including the negative case (one draw too few leaves a different state).

next is rebuilt by ours against the live generator address so the describer's index arithmetic reads the same on both sides; ours never dereferences it.

What ours covers, and what it deliberately does not

ours is sots::sim::ProcessResearchTurn (new, src/game/sim/research.{h,cpp}) plus a thin shim adapter. It reproduces exactly the words the hooked function writes itself:

  • the allocation loop — spend window, spend, *overbudget, progress, the roll, the completion decision, the over-budget flag, the "completed early" flag, state = 4 on completion;
  • the decay sweep over every available node.

It does not reproduce TechTree::SetResearched, which the original calls on completion: the turn/order stamps, the child-unlock cascade and the owner's tech-effect callback. That is its own milestone, and the callback writes live player state that compare mode must never touch. So:

A turn on which a tech completes is expected to diverge, in the completing node's turn_researched / order and in the child nodes SetResearched unlocks. A turn on which nothing completes — the overwhelmingly common case — must match everywhere. (Borne out: the two divergent calls of the live compare carry exactly those fields and nothing else. The prediction that the cascade also makes no RNG draw was half wrong — one of the two completions consumed an extra word inside the owner's tech-effect callback; see "The extra draw" below.)

It also does not post the events the pass raises — and that turned out to matter: the over-budget branch sets node.flag = 2 and pushes an EVENT_RESEARCH_OVERBUDGET onto the owner's event list, and only the flag is modelled. The event list was never a declared region, so compare mode is blind to it; the End-Turn oracle is what caught it. See "The replace-mode gap".

The effective cost of a node comes from the game's own TechTree::Cost (read-only: it only reads costRP, the def and the owner, and calls the read-only cost-multiplier helper). The cost multiplier is a separate, medium-confidence formula and not what this milestone measures; this is the same delegation M2 makes to LoadWeapon. ours receives the live tree pointer and treats it as read-only — every node it writes is a scratch copy.

ours also works in replace mode, where no regions/rebind ran: it then reads the tree's own node vector and the live generator. The per-call statics carry a flag that is cleared at the end of every ours, so a replace call can never inherit a stale compare mapping.

Float mapping — the headline finding

A draw is y / (2^32 − 1), not y × 2^-32. The multiplier in the image is the double 0x3df0000000001000, which is exactly 1/4294967295; the constant next to it is the +2^32 unsigned fix-up applied after a sign-extending integer load. So:

  • MT19937::kUnitScale is now 1.0 / 4294967295.0;
  • the range is closed: y == 0xffffffff maps to exactly 1.0, not to just below it;
  • the value is left in st(0) and the caller narrows it — every consumer in the strategic sim stores it to a 4-byte float first, which is what next_float() models.

Honest caveat, because it decides how to read a passing compare: the old and new divisors differ by 2^-32 relative, far below a float32 ulp. Measured over 10^6 draws they give a different float 0.78 % of the time, and they flip an actual research completion decision (roll vs an odds of 1/3) 0 times in 10^6 — the expected rate is about one in two billion. So the compare cannot prove the divisor; the disassembly and the constant's bit pattern do, and the compare's job is the rest.

The one thing the binary cannot settle is the x87 precision-control field at run time. At the MSVC default (53-bit, cw = 0x027f) the multiply rounds to double and the caller's store rounds again — that is what next_float() does. A Direct3D 9 device created without FPU_PRESERVE leaves 24-bit precision, in which the fix-up and the multiply each round to 24 bits; MT19937::float_from_pc24() models that, and the two differ for 0.094 % of words — last-bit only. The hook records fpu_cw on every call, so the first trace settles it; if it comes back 24-bit the change is to route next_float through float_from_pc24 and to compute the odds the same way, and nothing else in the milestone moves.

float10 in the decompile is just the i386 float return ABI and was not chased.

Bugs found and fixed in our implementation

All five are read off the instruction sequence, not tuned to make anything match.

  1. The unit divisor (above): 2^-32 → 1/(2^32 − 1).
  2. NextInt is inclusive. The rejection mask is built from n itself, not n − 1, and the loop re-draws while the masked word is greater than n — so the result is uniform on [0, n], one value wider than we had. The bound is also passed by pointer, which is why the prototype had stayed [unverified]. next_int(n) is now next_int_inclusive(n) and IRandom::NextInt is IRandom::NextIntInclusive, so every call site had to be re-read.
  3. spend has no floor at zero. The original is a plain signed min(points, hi − progress); ours clamped it at 0, which would have hidden a negative spend (and understated *overbudget) whenever progress was already past the 150 % ceiling.
  4. odds is a float32. The original computes (progress − lo) / hi on the x87 and stores it to a 4-byte slot before the comparison; ours kept it in double. Likewise the roll, the Zuul minimum, and the progress / cost ratio used for the "completed early" flag.
  5. Two constants are widened float literals, not decimals. The decay fraction is (double)0.05f = 0.05000000074505806 and the early-completion threshold is (double)0.8f = 0.800000011920929. Both sit on a truncation/compare boundary.

Two smaller ones in the same pass: the decay guard is progress != 0, not progress > 0; and a node whose cost is still INT_MAX is not special-cased by the original — it feeds that straight into the 5 % multiply, which wipes any progress out. The 50 %/150 % bounds are a 32-bit multiply that wraps near INT_MAX rather than a widening one.

Host tests

ctest 26/26. New coverage:

  • mars_rng_unit — the standard MT19937 vectors (seed 5489, and the 10000th output) were already there; added the unit mapping word by word (0 → 0, 0xffffffff → 1.0, the high-bit fix-up path), the measured rarity of the divisor difference, the PC24 variant's bound, cover_mask, the inclusive integer bound (including that n itself is reachable and that n == 0 still consumes a word), and the compare design end to end: snapshot → original draws → ours seeded from the snapshot reproduces the values and the post-state, with a negative case that one draw too few does not.
  • game_sim_research — 121 checks. Added the spend window (truncation, the floor/ceiling clamps, the INT_MAX edge), that the odds are float32, that spend goes negative rather than clamping, the early-completion boundary at exactly 80 % versus one point below, and ProcessResearchTurn (order of the passes, the funded node decaying too, hidden slots never touched, an out-of-range entry consuming no draw, overbudget accumulating).

Cross-build: b3-81218c7-dirty-20260908T0311Z, exports 66 names identical to binkw32.dll, staged in /srv/re-lab/shim/dist-b3. tools/clean_room_check.sh OK.

Gotchas

  1. Two different this pointers for one object. Twist, NextFloat and NextInt take &mt — the object plus 4 — so their this+0x9c0/+0x9c4 are next/left, while the object's own layout is {vftable @+0, mt[624] @+4, next @+0x9c4, left @+0x9c8} = 0x9cc bytes. The address contract used to state both readings as if they were one; it now says which is which. Getting this wrong shifts every generator field by a word.
  2. The save blob is 0x9c4 bytes = mt[624] + left, i.e. it skips next, which sits between them in the object. left is sufficient because next == &mt[624 - left] and the original's Read recomputes it.
  3. Region::name is a const char* held for the whole call, so the per-node name strings are resized once up front and never grown — a reallocation would dangle every name already pushed.
  4. Per-call state is kept in statics between regions() → rebind() → ours() (M1's concession). Safe here because the turn pass is single-threaded and ProcessResearch has one caller and never nests; do not copy the pattern to a re-entrant hook.
  5. A compare record is large (one region per tech node). The staged configs (shim.cfg.b3{trace,compare,replace}, in the dist) switch every other hook off for that reason; trace.inline_max stays at 256 so the 2496-byte state block is hashed rather than inlined.
  6. TechDef's first word is the tech id and it indexes TechTree+0x10; the allocation vector's stride is 8 ({TechDef*, int}). Both are in the address contract now rather than inferred.

Runs (/srv/re-lab/shim/traces/, build b3-b791392-20260908T0353Z)

file mode calls result
b3-trace-golden.jsonl trace 3 tracecmp.py exit 0, 0 invalid
b3-compare.jsonl compare 15 13 clean, 2 diverged (both completion turns; see below)
b3-replace-autosave.sav vs b3-orig-autosave.sav replace 3 (Autosave EndTurn).sav = bb4fd9ac… (oracle); (Autosave).sav = fd071d18… ≠ oracle 978041ac…
b3-replace-save-diff.txt the whole difference: one missing event
b3-shim.log, b3-final-menu.png, b3-compare-turn7.png banners, final state

Procedure: ref-turn2.sav → End Turn ×1 (trace), ×5 (compare), ×1 (replace and the oracle re-check). Each End Turn produces three calls: the AI Tarkas player (which actually researches), the human, and a second Tarkas tree, the last two with a zero-point allocation.

What the compare proves

Per call, the trace records the draw count as left before minus after:

call species alloc {tech, points} draws (orig / ours) outcome
0 Tarkas {144, 2889} 1 / 1 over-budget event raised, flag 1 → 2
1, 2 Human, Tarkas {90, 0}, {9, 0} 0 / 0 zero spend → odds 0, roll 1, no draw
3 Tarkas {144, 2898} 0 / 0 reached the 150 % ceiling → guaranteed, no draw, completed
6 Tarkas {142, 3064} 1 / 1 rolled and failed (odds −0.078, below 50 % of cost)
9 Tarkas {142, 3074} 2 / 1 rolled at odds 0.178 and completed early (6138/8000 = 0.767)
12 Tarkas {9, 3086} 1 / 1 rolled and failed

Every integer the pass computes matched: progress, the spend cap, *overbudget, and the flag on every node, including the three interesting branches — the over-budget event flag (call 0), the guaranteed completion at the ceiling with no draw at all (call 3), and the "completed early" flag 1 → 0 below 80 % of cost (call 9). On call 9 ours drew the same word as the original (the completion decision and the early flag both match), so the extra draw is the original's, not a missing one of ours.

The two divergences. Both are completion turns, and both carry only SetResearched-owned fields: order and turn_researched on the completing node, and cost_rp / state / turn_available on the children it unlocks (three children on call 3, one on call 9). That is exactly the set this milestone declared out of scope in advance.

The extra draw (call 9). The original consumed one word more than ours after the completion roll. Call 3 also completed a tech and consumed no extra word, so this is not "SetResearched always draws" — it is the owner's tech-effect callback for that particular tech. That callback is ServerPlayer::OnTechResearched, which is the B2 lane's target; B2 should expect to find at least one effect that rolls.

The replace-mode gap — one missing event

Replace mode ran our pass for a whole turn. (Autosave EndTurn).sav (the pre-turn state) is the oracle's bb4fd9ac…, but the post-turn (Autosave).sav is fd071d18… and not the oracle's 978041ac…. Dumping both trees and normalising the offsets, the entire difference in 40,300 items is:

Player > Events > Events > . > Events:  one entry present in the original, absent in ours
    EvEID 3   EvDsc "Research Over Budget"
    EvMsg "Research for Waldo Units has gone overbudget."
    EvImg "EVENT_RESEARCH_OVERBUDGET"   EvLoc 0   EvPos {inf,inf,inf}   EvAct 1   EvCID 0
  + the two counters that index it: EvNxID 4 → 3 and the list length 2 → 1

Nothing else — not one number the research pass computes — differs. So the formulas are right and the milestone under-scoped itself: docs/B3.md promised the over-budget flag and treated the event beside it as part of the same branch without saying it was excluded. It is not modelled, ours writes node.flag = 2 and stops there, and compare mode could not see it because the player's event list was never a declared region.

Two consequences worth carrying forward:

  • The oracle catches what a region-scoped compare cannot. Compare mode was clean on exactly the call that produced the missing event, because the event lands somewhere we did not declare. A clean compare bounds the state you declared, nothing more.
  • To close this, ours needs to post the event through the game's own event API (the same call the original makes right after setting the flag), the way M2 delegates to LoadWeapon — the message text is composed from the tech name, so it cannot be synthesised from our side alone. Until then, replace mode is not oracle-clean and should not be used as a determinism check for this hook.

Corrections this run made to the notes

  1. The current research target has state 3, not 2 — so it does not decay. Every funded node in the trace is state 3 (node[144], node[142], node[9]), and the decay loop's guard is state == 2. formula-gaps.md Q8 read the loop correctly but concluded the funded node decays; it does not, because the tree marks the selected target with a different state. Our code was already right (it compares against Available == 2); the prose was wrong. The state set observed in one tree: 164 hidden (0), 7 parent-researched (1), 23 available (2), 1 current target (3), 22 researched (4).
  2. fpu_cw = 0x127f at call time (0x027f at DLL init): precision control = 10b = 53-bit double, rounding = nearest. Bit 12 is the legacy infinity-control flag and has no effect. So the MSVC default path is the right one and next_float() — double product, then the caller's narrowing to float32 — models the original exactly. float_from_pc24() stays as an unused contingency; nothing needs to change.
  3. Nodes per tree = 293, not the 196 the tech-name table suggested.

Coverage this run did not reach

  • No Zuul. ref-turn2.sav has two players, re (human) and Fane Lao (AI), and the trees seen are species 0 (Human) and 2 (Tarkas). The double roll is therefore still verified by disassembly and host tests only, never by behaviour. A compare from a save with a species-5 player would close it; the check is simply that left drops by 2 on a roll instead of 1.
  • The decay sweep never fired. No node in any traced tree was state 2 with non-zero progress — idle partially-researched techs do not occur in this save, and the current target is state 3. DecayAllResearch is covered by host tests only.
  • Only one lab-accident-free, single-target allocation shape was seen (one entry per call).

If the replace-mode gap is closed

  1. Add the event post to ours (see above), rebuild, restage /srv/re-lab/shim/dist-b3.
  2. Deploy from C:\SOTS\shimdist-b3, copy shim.cfg.b3replace over C:\SOTS\shim.cfg, relaunch, load ref-turn2.sav, End Turn once, and re-check the oracle: (Autosave).sav = 978041ac…, (Autosave EndTurn).sav = bb4fd9ac….
  3. Restore the default shim.cfg (hooks=trace) and leave the game at the main menu.

Note that replace mode also stops modelling completions the moment one happens on a later turn: ours does not run the unlock cascade, so a replace run must not be pushed past the first completion until SetResearched is a milestone of its own.