15 KiB
B3 — TechTree::ProcessResearch old-vs-new, with the RNG state as a declared region
Status (2026-09-08): code complete, cross-built and staged; every VM step still owed.
VM140 is held by another lane for the whole of this milestone, so nothing was deployed, the
game was not stopped or relaunched, and C:\SOTS\binkw32.dll / C:\SOTS\shimdist were not
touched. Everything below is offline work plus what the binary says; the run list is at the
end. The build lives in its own tree (/srv/re-lab/build/sots-engine-b3) and its own dist
(/srv/re-lab/shim/dist-b3), not the shared ones.
The point of the target: one call exercises the MT19937, the completion-odds formula and the Zuul double roll at once, so a match validates all three together — and its RNG consumption is observable, which makes the draw count checkable rather than merely plausible.
What was hooked
hook name (record hook) |
RVA | prototype |
|---|---|---|
Game::TechTree::ProcessResearch |
0x001876c0 | void (TechTree*, Mars::RNG*, vector<{TechDef*,int}>*, int*) |
__thiscall, [verified], so it goes through Hook<> with CallConv::Thiscall (M2's
addition). Source: src/shim/hooks/research.{h,cpp}, installed from src/shim/main.cpp after
the M1/M2 hooks. There is exactly one caller (inside ServerPlayer::ProcessTurn), which fires
once per player per turn.
The second parameter was a ? in the address contract; it is the Mars::RNG object. The
call site loads it from StrategyServer+0x16c and the function re-bases it with +4 before
every draw. That is now in ghidra/addresses.json along with the tree/node offsets, the
generator layout and TechTree::Cost, and regenerated into sots_addresses.h.
Region model
Three kinds of region, all snapshotted before the original runs:
| region | size | describer |
|---|---|---|
rng |
0x9cc | {vptr:ptr, mt:bytes(2496 → sha256+head), left:i32, next_index:i64} |
overbudget |
4 | {v:i32} |
node[i] |
0x34 each, one per non-null slot | {def, tech_id, kids_*, unk10, state, cost_rp, progress, turn_available, turn_researched, order, flag, unk30} |
next is a heap address, so it is reported as its index into mt — which is what it means,
and what survives being written by a reimplementation. left alone already pins the stream
position (next == &mt[624 - left] always), so the index is a cross-check, not the evidence.
Args carry the evidence a golden log needs to replay offline: tree, owner, species,
node_count, rng, rng_left_in, the alloc list as {tech_id, points}, overbudget_in,
and fpu_cw — the x87 control word in force for the call (see "Float mapping" below).
Declaring every node, not just the ones we expect to change, is deliberate: it is what proves
ours neither misses a write nor makes an extra one. The cost is size — a compare record is
roughly 200 nodes × three snapshots, so the b3 configs turn every other hook off.
The comparison design
ProcessResearch consumes RNG, so comparing two implementations that draw from different
streams would diverge for a reason that has nothing to do with the formulas. Instead:
- the generator object is a declared region, so its
mt[624]+leftare snapshotted before the original runs, alongside the tech-tree state; - the original runs and advances the real generator;
oursruns on the scratch copies, and seeds amars::rng::MT19937withload_state()from the pre-call snapshot — so both implementations read the identical stream;ourswrites its final generator state back into the scratch copy, so the diff compares the post-call RNG state as well as the outputs.
If the post-states match, we consumed the same words in the same order. That is the check with
teeth: getting the odds right but drawing twice (or not drawing at all) moves left and the
hash. tests/mars_stream/test_rng.cpp pins the property offline, including the negative case
(one draw too few leaves a different state).
next is rebuilt by ours against the live generator address so the describer's index
arithmetic reads the same on both sides; ours never dereferences it.
What ours covers, and what it deliberately does not
ours is sots::sim::ProcessResearchTurn (new, src/game/sim/research.{h,cpp}) plus a thin
shim adapter. It reproduces exactly the words the hooked function writes itself:
- the allocation loop — spend window, spend,
*overbudget, progress, the roll, the completion decision, the over-budget flag, the "completed early" flag,state = 4on completion; - the decay sweep over every available node.
It does not reproduce TechTree::SetResearched, which the original calls on completion:
the turn/order stamps, the child-unlock cascade and the owner's tech-effect callback. That is
its own milestone, and the callback writes live player state that compare mode must never
touch. So:
A turn on which a tech completes is expected to diverge, in the completing node's
turn_researched/orderand in the child nodesSetResearchedunlocks — and in nothing else, including the RNG (the cascade makes no draw; verified statically, though the owner callback itself was not audited). A turn on which nothing completes — the overwhelmingly common case — must match everywhere.
The effective cost of a node comes from the game's own TechTree::Cost (read-only: it only
reads costRP, the def and the owner, and calls the read-only cost-multiplier helper). The
cost multiplier is a separate, medium-confidence formula and not what this milestone measures;
this is the same delegation M2 makes to LoadWeapon. ours receives the live tree pointer and
treats it as read-only — every node it writes is a scratch copy.
ours also works in replace mode, where no regions/rebind ran: it then reads the tree's
own node vector and the live generator. The per-call statics carry a flag that is cleared at
the end of every ours, so a replace call can never inherit a stale compare mapping.
Float mapping — the headline finding
A draw is y / (2^32 − 1), not y × 2^-32. The multiplier in the image is the double
0x3df0000000001000, which is exactly 1/4294967295; the constant next to it is the +2^32
unsigned fix-up applied after a sign-extending integer load. So:
MT19937::kUnitScaleis now1.0 / 4294967295.0;- the range is closed:
y == 0xffffffffmaps to exactly1.0, not to just below it; - the value is left in
st(0)and the caller narrows it — every consumer in the strategic sim stores it to a 4-byte float first, which is whatnext_float()models.
Honest caveat, because it decides how to read a passing compare: the old and new divisors differ by 2^-32 relative, far below a float32 ulp. Measured over 10^6 draws they give a different float 0.78 % of the time, and they flip an actual research completion decision (roll vs an odds of 1/3) 0 times in 10^6 — the expected rate is about one in two billion. So the compare cannot prove the divisor; the disassembly and the constant's bit pattern do, and the compare's job is the rest.
The one thing the binary cannot settle is the x87 precision-control field at run time. At
the MSVC default (53-bit, cw = 0x027f) the multiply rounds to double and the caller's store
rounds again — that is what next_float() does. A Direct3D 9 device created without
FPU_PRESERVE leaves 24-bit precision, in which the fix-up and the multiply each round to 24
bits; MT19937::float_from_pc24() models that, and the two differ for 0.094 % of words —
last-bit only. The hook records fpu_cw on every call, so the first trace settles it; if it
comes back 24-bit the change is to route next_float through float_from_pc24 and to compute
the odds the same way, and nothing else in the milestone moves.
float10 in the decompile is just the i386 float return ABI and was not chased.
Bugs found and fixed in our implementation
All five are read off the instruction sequence, not tuned to make anything match.
- The unit divisor (above):
2^-32→1/(2^32 − 1). NextIntis inclusive. The rejection mask is built fromnitself, notn − 1, and the loop re-draws while the masked word is greater thann— so the result is uniform on[0, n], one value wider than we had. The bound is also passed by pointer, which is why the prototype had stayed[unverified].next_int(n)is nownext_int_inclusive(n)andIRandom::NextIntisIRandom::NextIntInclusive, so every call site had to be re-read.spendhas no floor at zero. The original is a plain signedmin(points, hi − progress); ours clamped it at 0, which would have hidden a negative spend (and understated*overbudget) whenever progress was already past the 150 % ceiling.oddsis a float32. The original computes(progress − lo) / hion the x87 and stores it to a 4-byte slot before the comparison; ours kept it in double. Likewise the roll, the Zuul minimum, and theprogress / costratio used for the "completed early" flag.- Two constants are widened float literals, not decimals. The decay fraction is
(double)0.05f = 0.05000000074505806and the early-completion threshold is(double)0.8f = 0.800000011920929. Both sit on a truncation/compare boundary.
Two smaller ones in the same pass: the decay guard is progress != 0, not progress > 0; and
a node whose cost is still INT_MAX is not special-cased by the original — it feeds that
straight into the 5 % multiply, which wipes any progress out. The 50 %/150 % bounds are a
32-bit multiply that wraps near INT_MAX rather than a widening one.
Host tests
ctest 26/26. New coverage:
mars_rng_unit— the standard MT19937 vectors (seed 5489, and the 10000th output) were already there; added the unit mapping word by word (0 → 0,0xffffffff → 1.0, the high-bit fix-up path), the measured rarity of the divisor difference, the PC24 variant's bound,cover_mask, the inclusive integer bound (including thatnitself is reachable and thatn == 0still consumes a word), and the compare design end to end: snapshot → original draws → ours seeded from the snapshot reproduces the values and the post-state, with a negative case that one draw too few does not.game_sim_research— 121 checks. Added the spend window (truncation, the floor/ceiling clamps, theINT_MAXedge), that the odds are float32, thatspendgoes negative rather than clamping, the early-completion boundary at exactly 80 % versus one point below, andProcessResearchTurn(order of the passes, the funded node decaying too, hidden slots never touched, an out-of-range entry consuming no draw,overbudgetaccumulating).
Cross-build: b3-81218c7-dirty-20260908T0311Z, exports 66 names identical to binkw32.dll,
staged in /srv/re-lab/shim/dist-b3. tools/clean_room_check.sh OK.
Gotchas
- Two different
thispointers for one object.Twist,NextFloatandNextInttake&mt— the object plus 4 — so theirthis+0x9c0/+0x9c4arenext/left, while the object's own layout is{vftable @+0, mt[624] @+4, next @+0x9c4, left @+0x9c8}= 0x9cc bytes. The address contract used to state both readings as if they were one; it now says which is which. Getting this wrong shifts every generator field by a word. - The save blob is 0x9c4 bytes =
mt[624]+left, i.e. it skipsnext, which sits between them in the object.leftis sufficient becausenext == &mt[624 - left]and the original'sReadrecomputes it. Region::nameis aconst char*held for the whole call, so the per-node name strings areresized once up front and never grown — a reallocation would dangle every name already pushed.- Per-call state is kept in statics between
regions()→rebind()→ours()(M1's concession). Safe here because the turn pass is single-threaded andProcessResearchhas one caller and never nests; do not copy the pattern to a re-entrant hook. - A compare record is large (one region per tech node). The staged configs
(
shim.cfg.b3{trace,compare,replace}, in the dist) switch every other hook off for that reason;trace.inline_maxstays at 256 so the 2496-byte state block is hashed rather than inlined. TechDef's first word is the tech id and it indexesTechTree+0x10; the allocation vector's stride is 8 ({TechDef*, int}). Both are in the address contract now rather than inferred.
What remains (needs the VM)
The lane holding VM140 must be finished first; then, in this order:
- Deploy
/srv/re-lab/shim/dist-b3(buildb3-81218c7-dirty-20260908T0311Z):scpit toC:\SOTS\shimdist-b3\and rundeploy.ps1 -Dist C:\SOTS\shimdist-b3— a separate staging directory from the sharedC:\SOTS\shimdist, so the other lane's dist is not overwritten. - Copy
shim.cfg.b3traceoverC:\SOTS\shim.cfg, relaunch, loadref-turn2.sav, press End Turn once, and pullC:\SOTS\shim.trace.jsonl→b3-trace-golden.jsonl.tracecmp.pymust exit 0 with 0 invalid records. Expect one record per player that both has a research target and did not suffer a lab accident that turn — the caller gates the call onif (ResT && !RollResearchAccident())— so a turn with no records at all is a setup problem, not a pass. Read off this trace before going further:fpu_cw(expect0x027f;0x007f/0x003fmeans 24-bit precision and thefloat_from_pc24route), and, per record,rng_left_inminus the after-state'sleft— 1 for a non-Zuul player who rolled, 2 for a Zuul, 0 when the spend was zero or the tech was already at the 150 % ceiling. - Copy
shim.cfg.b3compare, relaunch, loadref-turn2.sav, End Turn →b3-compare.jsonl. Expect 0 divergences on every record where no tech completed. On a record whoseoursshows a completion, the only permitted diffs areturn_researched/orderon that node and the child nodesSetResearchedunlocked; therngregion must still match. Any other diff is a real finding — report it, do not tune the formula. To get more turns, keep pressing End Turn; each turn adds one record per player. - If the compare is clean: copy
shim.cfg.b3replace, relaunch, loadref-turn2.sav, End Turn, and check the determinism oracle —(Autosave).sav=978041ac…,(Autosave EndTurn).sav=bb4fd9ac…. This is the strongest single result available: replace mode means our research pass fed the game, and the save hash means the whole turn still landed byte for byte. - Restore the previous
shim.cfg(hooks=trace) and leave the game at the main menu, as M1/M2 left it.
Not done, and worth saying: the compare above cannot distinguish the two unit divisors (see
"Float mapping"), and it will not exercise the Zuul path at all unless a Zuul player is in the
reference save — check species in the trace args, and if none is 5, run one more compare from
a Zuul save before calling the double roll verified by behaviour rather than by disassembly.