sots-re/findings/control-flow/tail-rng-ledger.md
alex 9c37635ca3 Z: the unnamed counter at S+0x8 is ModCount, and my players flag was my own error
StrategyServer::Write tags both words itself: S+0x8 is ModCount and S+0xc is
Frame. addresses.json has the name on the wrong word and lane T's
PhaseCounter is the one the wire calls ModCount. The saves confirm it
independently -- ModCount 0/12/24 across turn1/2/3-state, 241/412 across Zuul
16/23 -- and those deltas are exactly the 12-44 per turn measured live. So the
'writer nobody has identified' question dissolves: it is a modification
counter, it scales with the empire, and there is no single writer to find.

The players=8 flag is withdrawn. The offset is right, pinned by the ctor's
four-vector enumeration at 0x0085b120 with no frame arithmetic needed, and the
count is right: the vector is empires + one rebel-AI per empire species + four
NPC pseudo-players, so 8 on the Human saves and 7 on the Zuul ones against a
lobby that says 2 in both. My draft claimed the hook read 8 on both saves. It
read 7 on the Zuul one. I generalised from one run without re-reading the
other, and a check aimed at something else caught it.
2026-09-08 10:04:21 -04:00

27 KiB
Raw Blame History

The RNG ledger for one strategic turn — measured, not inferred

Lane Z, 2026-09-08. Program sots / "Sword of the Stars.exe", ImageBase 0x00400000, all addresses VAs. Engine worktree wip/tailrng, sots-engine/docs/Z-tail-rng.md (the prediction, committed before the build).

The question. The milestone is a standalone that loads a save, runs one strategic turn, and writes an autosave that byte-matches the original's. Generator state is part of the saved state. Lane K (combat-done-tail.md §3) found two draw sites in StrategyServer::OnAllCombatDone_Tail that nothing models and that both run before the autosave, and concluded that a reimplementation reproducing both ProcessTurn functions exactly would still diverge. Nobody had measured what a turn actually costs.

The answer, up front. Across eight measured End Turns on two saves, a strategic turn advances the strategic generator by 18–22 words, all of it inside StrategyServer::ProcessTurn, and the residual outside the two turn drivers is exactly zero. The tail's cost on these turns is 0. The defect lane K found is real and latent: it will bite the first turn a node line expires or a real battle resolves, and our saves reach neither.


0. Where to start if you are building the standalone

Three sentences, then the evidence.

  1. A turn costs 18–22 generator words, all of it inside StrategyServer::ProcessTurn; the two autosave files bracket exactly that interval and nothing draws between them outside the two turn drivers.
  2. verify/results/shim/tailrng/z-t6-endturn.sav → z-t6-autosave.sav is a byte-identical oracle pair with a known RNG cost of 18 words, verified from the file bytes independently of the live hook. Test against that pair before any other.
  3. Lane K's warning stands and is now quantified: the tail's cost is 0 today because no save is within ~40 turns of a node-line expiry (§9) and no encounter has ever produced a battle (§8). Both terms are real; both are latent.

1. The instrument, and why it is not an RNG hook

Counting draws by hooking the primitives would have undercounted, and the static work this lane commissioned says by how much. The image has four draw entry points, not three:

VA primitive words per call
0x0047d830 Mars::RNG::NextFloat(&mt) — __thiscall, no stack args, plain ret exactly 1
0x004271c0 Mars::RNG::NextInt(&mt, uint* pMax) — ret 4, rejection loop, bound re-read each iteration 1 or more
0x008e6dd0 Mars::RNG::Chance(RNG*, float p) — ret 4; takes the object base and does the add ecx,4 itself 0 or 1 — see §6.1
0x004f7670 Mars::RNG::NextUInt(RNG*) — plain ret, no stack args, 84 B. In no previous lane's primitive set, and called from 0x007b6700 inside ProcessTurn's closure 1

(Lane J's scan reaches 0x004f7670 from the other direction and calls it one of the fourteen inlined-draw functions rather than a primitive. Both readings are of the same object and neither is wrong: it is a small function whose draw is inlined, so it is a draw site that calls no primitive — which is exactly why a call-graph sweep for NextFloat/NextInt cannot see it.)

plus inlined draws: this lane's scan for the MT tempering immediates found twelve such functions, two of them reachable from the turn roots (0x007aa240 under ProcessTurn, 0x007a7f30 under the combat resolver). Lane J's independent scan, landed while this run was in flight, found fourteen and turned it into method rule 16 — inlined draws are invisible to call-graph sweeps, so any RNG accounting built from the call graph alone is a lower bound. Take lane J's number over this one; the two scans agree on the shape and the point is the same either way. A primitive-counting hook would have missed every one of them silently.

This instrument is not built from the call graph, which is why rule 16 does not apply to it: it reads the generator's state before and after a boundary and reports the difference. An inlined draw, a draw in a function nobody has named, a draw through a vtable — all of them move left, and all of them are counted.

So the instrument reads state, not calls. Mars::RNG is {void* vptr; uint32 mt[624]; uint32* next; int32 left}, sizeof 0x9cc — next at +0x9c4 is a pointer into the block and left at +0x9c8 is the counter, and the three inner primitives are entered with ECX = &mt[0] = RNG+4, which is what lane T's "the generator is entered at rng+4" note was seeing. The draw is if (left == 0) Twist(); y = *next++; --left; — a pre-check against 0, never −1.

The block transform is a pure function, so the blocks a generator visits form a forward-only chain. RngLedger (sots-engine/src/shim/hooks/rng_ledger.{h,cpp}) indexes that chain from the first block it sees and positions any state exactly:

words(block, left) = block * 624 + (624 - left)

Differences between positions are then exact across twists, across rejection loops, and across draws nobody hooked. Nine host tests pin the arithmetic, including the block boundary (left == 0 is a real state and must not be off by 624), a rejection loop counted against a shadow generator, and the out-of-order case below.

One ordering subtlety, and it is load-bearing. Hook<> takes the before snapshot at entry but renders it (calls the region's describe) only after the original returns — so for nested calls the inner snapshots are rendered first, and by the time an outer before is rendered the chain has moved past it. Walking a Mersenne Twister backwards is not possible. Every hook therefore observes at entry from describe_args, which indexes the entry block while it is still current; the render then resolves it from the memo. Without that, every outer-call before would read words: null.

Six nested trace hooks, all watching the object at S+0x16c:

StrategyHost::Autosave (both markers) › StrategyServer::ProcessTurn › OnAllCombatDone_Tail › ApplyEncounterResult (phase 6) › node-line decay (phase 11) › ProcessNodeSpaceTravel (runs twice a turn).


2. The ledger

Save ref-turn2 (2-player Morrigi vs AI), four consecutive End Turns, build z-tailrng-20260908T1314Z, hooks=trace. Words are 32-bit MT outputs consumed by the generator at S+0x16c.

turn ProcessTurn tail ApplyEncounterResult node-line decay ProcessNodeSpaceTravel ×2 bracket total residual
3 19 0 0 0 0, 0 (incomplete — see below) —
4 18 0 0 0 0, 0 18 0
5 20 0 0 0 0, 0 20 0
6 18 0 0 0 0, 0 18 0

"Bracket" is Autosave(endTurn=1) → Autosave(endTurn=0): the pre-turn save file to the post-turn save file, which is exactly the interval a standalone has to reproduce. The turn-3 bracket is incomplete by construction and is reported rather than dropped: the pre-turn autosave of the first End Turn after a load runs before any turn driver, so the hook has no server pointer yet and its record carries words: null.

A second save, zuul-turn16-noderoute (Zuul vs Zuul), build z-tailrng2-20260908T1328Z:

turn ProcessTurn tail ApplyEncounterResult node-line decay ProcessNodeSpaceTravel ×2 bracket total residual
17 20 0 0 0 0, 0 (incomplete) —
18 20 0 0 0 0, 0 20 0
19 22 0 0 0 0, 0 22 0
20 20 0 0 0 0, 0 20 0

So: we consume 18–22 words per turn and model 0 of them.

Not 0 because the modelling is bad — because nothing in the repo models any part of a turn's RNG consumption as a count. B1/B3/B4 compare the generator's state around three specific functions and get it right; no lane has ever stated a turn's total. This table is that statement.

2.1 The generator does not move outside the turn pipeline

Every turn's ProcessTurn entry position equals the previous turn's post-turn autosave position, exactly: 211 → 211, 229 → 229, 249 → 249. The UI, the renderer and the per-frame tick draw nothing from the strategic generator between turns. For the standalone this is worth as much as the total: the interval it must reproduce is closed.


3. An independent check, from the files rather than from memory

The two autosaves of the turn-6 bracket were pulled off the VM and their Sim.RNG blobs parsed by verify/save-reader/save_reader.py (the frame's payload is 2503 bytes: mt[624], then left as int32 at +2496). Twisting the pre-turn block forward until it matches the post-turn block, and applying the same position formula:

z-t6-endturn.sav  (pre-turn)   left = 375
z-t6-autosave.sav (post-turn)  left = 357
twists = 0   ->  words consumed between the two files = 18

The live ledger recorded 18 for that turn, left 375 → 357. Two instruments that share no code path — one reading process memory through a hook, one reading gzip-compressed file bytes through the save reader — agree exactly. Files and the checker in verify/results/shim/tailrng/.

And the two do not share a hidden assumption (the trap of method rule 8, which is live here because both sides know how to twist an MT block). twists = 0: the block is byte-identical in the two files, so the file-side number is left_before − left_after and involves the twist implementation not at all. The agreement is therefore about the game's behaviour, not about two copies of the same algorithm agreeing with each other.

This is the pair a standalone should be tested against first: it is a byte-identical oracle with a known RNG cost attached, which none of the eleven corpus saves has.


4. Lane K's inference, settled

§6, labelled hypothesis: "I did not prove that SNMAllCombatDone is delivered on turns with no combat."

The handler runs on every End Turn. Eight out of eight, across two unrelated saves, OnAllCombatDone_Tail recorded exactly one call per End Turn, at depth 0, between the two autosaves, with the post-turn autosave following it. The determinism-note inference was right.

The stronger claim — that it runs with an empty encounter vector — is not settled by this workload and must not be reported as though it were. On both saves the encounter vector is empty at ProcessTurn entry and holds exactly one encounter by the time the tail runs, on every one of the eight turns: detection (ProcessTurn phase 31) creates it, and the tail's phase 7 clears it. So what is proved is "the tail runs on a turn with no battle", not "on a turn with no encounter at all". See §8 for what closing the remaining gap needs.

Two things fall out of the same records and are worth more than the phrasing:

  • Phase 7 really is a wholesale clear(). encounters reads 1 at the phase-6 ApplyEncounterResult call and 0 at the phase-11 node-line-decay call, on every turn. Lane K read that off the instruction stream against a decompile that reads as a conditional prune; it is now also a behavioural fact.
  • Every encounter on these turns has res->+0x4 != 0 (res_no_battle = 1), the flag that makes ApplyEncounterResult a whole-function no-op. Its measured cost is 0 words, which is what that gate predicts, and which is why the combat resolver has never run under any instrument this campaign has built.

5. S+0x8 has a name, and it is ModCount — correcting combat-done-tail.md §7.1, lane T §0.1, addresses.json, and this lane's own prediction

Lane K wrote that S+0x8 "advances at least twice per turn" and that "the word at S+0x8 has never been named" (lane T §0.1). The first is right and this lane's prediction that it advances exactly twice is wrong. The second is now answered — by StrategyServer::Write's own wire tags:

0079fb2f  lea edx,[edi+0x08] ; push "ModCount"      ; edi = S -- the same edi that indexes
0079fb40  lea eax,[edi+0x0c] ; push "Frame"         ;   the players vector at +0x54

So S+0x8 is ModCount and S+0xc is Frame, the turn number. ghidra/addresses.json has the name on the wrong word: its StrategyServer_off_ModCount = 0x8 is the stored-frame offset of S+0xc, which the wire calls Frame — turn-spine.md was right to call it that and lane T flagged the clash without being able to settle it. The name ModCount belongs to the word lane T recorded as StrategyServer_off_PhaseCounter = 0x4.

Confirmed three ways, and the third is the satisfying one. From the saves:

save Frame ModCount
turn1-state / turn2-state / turn3-state 1 / 2 / 3 0 / 12 / 24
z-t6-endturn / z-t6-autosave (this lane's bracket) 5 / 6 50 / 62
zuul-turn16-noderoute / zuul-turn23-fleet23 16 / 23 241 / 412

Frame is the turn; ModCount moves +12 per turn on the early Human game and averages +24 on the Zuul one. From the live trace, S+0x8 at hook entry:

save turn ProcessTurn entry tail entry tail's callees increments to the next turn
ref-turn2 3 22 23 24 12
ref-turn2 4 34 35 36 14
ref-turn2 5 48 49 50 12
ref-turn2 6 60 61 62 —
zuul-noderoute 17 253 254 255 16
zuul-noderoute 18 269 270 271 21
zuul-noderoute 19 290 291 292 44
zuul-noderoute 20 334 335 336 —

The live deltas (12, 14, 12 on the Human game; 16, 21, 44 on the Zuul one) sit exactly where the saves' ModCount deltas say they should. The two turn drivers account for 2 of 12 to 44 increments; the rest are spread across the turn and mostly fall between the post-turn autosave and the next ProcessTurn.

That is no longer a mystery to be chased — it is what a modification counter is for. S+0x8 is not a turn number, not a driver-invocation counter and not a constant per turn: it counts state mutations, so it scales with the size of the empire, and asking "which writer is responsible" has no single answer. Lane K's operational conclusion stands and is now explained rather than merely observed. S+0xc (Frame) reads 3, 4, 5, 6 and 17, 18, 19, 20 over the same records and is the turn counter.

For the integrator: this is a name collision to reconcile, not a new entry. StrategyServer_off_ModCount (0x8, stored frame) and StrategyServer_off_PhaseCounter (0x4, stored frame) are the two words above with their names swapped; ghidra/addresses.d/lane-z.json records the evidence under StrategyServer_wire_ModCount_vs_Frame rather than adding a third name for either word.

6. Corrections to combat-done-tail.md §3 and §6.1

Both from the instruction stream, both load-bearing for anyone reimplementing these functions.

§3 — the node-line fleet check does not gate the roll. Lane K: "The roll is skipped for a line if any fleet with flag 0x20000 is targeting it." The straight-line order in node-line decay's first loop is

0x007ae088  call NodePath::RemainingLife      ; expiry test
0x007ae08f  jg   0x007ae1e2                   ;   not expired -> next record, NO DRAW
0x007ae095  mov  ecx,[esi+0x16c]              ; THE DRAW
0x007ae0a5  call 0x008e6dd0                   ;   Mars::RNG::Chance(0.5f)
0x007ae0aa  test al,al ; je 0x007ae1e2        ;   roll failed -> next record
0x007ae0b2  ...                               ; THE 0x20000-FLEET SCAN STARTS HERE

The scan begins 0x1d bytes after the Chance call and is reached only when the roll succeeded. It suppresses the collapse (0x007a92e0 / 0x007a4700), never the draw. Lane K's headline — one NextFloat per expired node line per turn — survives intact and is now pinned to a formula.

§6.1 — StrategyHost::Autosave is ret 8 and returns a value. Its epilogue is c2 08 00, and 0x00895b5c mov eax,esi puts the std::string* (the MSVC named-return slot) in EAX. A hook declaring it void drops EAX at both call sites. Neither call site passes this: both hardcode mov ecx,0xb29f98. (The +0x54 candidate on that global was recorded live and is not the StrategyServer — server_agrees is false on all eight autosave records, so the global that the autosave uses is a different object from the StrategyHost whose +0x54 OnMessage reads.)

6.1 The expiry predicate, now concrete

NodePath::RemainingLife 0x006e2130, __thiscall(NodePath*, int turn), ret 4, whole 122-byte body read:

if (npt(+0x04) == 0)            return INT_MAX;   // permanent line, never expires
if (npdtn(+0x1c) == INT_MAX)    return INT_MAX;   // immortal line
aged = (npctm(+0x14) >= 0 && turn >= npctm) ? turn - npctm : 0;
wear = (npdtf(+0x20) != INT_MAX && npdtf > 0) ? nptf(+0x24) / npdtf : 0;   // SIGNED idiv
rem  = npdtn - wear - aged;
return rem > 0 ? rem : 0;                          // callee-side clamp

A line is expired exactly when this returns 0. The lifetime is derived, never ticked — the function writes nothing, and neither does the loop around it — so there is no decrement-ordering question and a snapshot at function entry is a valid prediction basis. nptf is never sign-checked, so the division's signed truncation must be reproduced literally.

Chance 0x008e6dd0 returns false with no draw when p <= 0 and true with no draw when p >= 1; 0.5f takes neither, so it is exactly one word, and the comparison is p > r (equality returns false). A NaN p falls through both early-outs and does draw — irrelevant here, noted because it is the kind of edge a reimplementation gets wrong.


7. One generator, confirmed twice

Static: the image has one persistent strategic Mars::RNG, at S+0x16c, constructed by the StrategyServer ctor 0x007d78d0 (push 0x9cc + Seed(0) at 0x007d7d25/0x007d7d3d) and reseeded only from Read and LoadGame. StrategyClient+0x134 and Mars::CombatSim+0x108 exist but are unreachable from the turn roots; three more are stack temporaries in map generation. The combat resolver draws from the same S+0x16c object — all three RNG entry points in its 750-node direct-call closure load [reg+0x16c].

Behavioural: across 64 ledger observations over eight turns on two saves, every state resolved on a single forward chain — no words: null, no second chain. If a second generator had been in play, the ledger would have said so by construction rather than by anyone noticing.


8. What is not settled, listed as loudly as the results

  • The combat resolver has still never run under an instrument. Every encounter this workload produced had the no-battle flag set, so ApplyEncounterResult was a no-op every time; its measured 0 words says nothing whatever about combat's RNG cost, and the residual-0 result above holds only for turns with no battle.

    Lane J read all 7,641 bytes of it in parallel with this run (combat-resolver.md; and note 7,641, not the 7,499 Ghidra reports — method rule 17). Its conclusion pairs with this one exactly: the resolver has no unconditional draw. All three sites in its subtree are conditional — a node-cannon NextInt, an inlined NextFloat per back-engineering candidate, and a NextInt per successful roll of that. So lane J's cheap first prediction is directly testable with this instrument: a plain fleet battle with no node cannon and no salvage should cost the same 18–22 words as a peaceful turn. That is the next run this hook family should do, and it needs a workload nobody has built yet: a save where two hostile fleets actually meet.

  • Node-line expiry did not fire. See §9 for the quantified distance rather than an absence.

  • A turn with a genuinely empty encounter vector was not observed (§4). Both saves produce exactly one sighting encounter on every turn. The tail-runs-every-turn claim is settled; the no-encounters variant is still an inference, now a much narrower one.

  • players reads 8 on a 2-player save, and may be the S+0x64 bug again. Resolved, and the flag was my own error. The offset is right and so is the count. StrategyServer's base-class ctor 0x0085b120 (entered with ecx = S+4) zero-initialises four consecutive vectors as three-word triples with the fourth word skipped — +0x40/+0x50/+0x60/+0x70 raw, 0x10 apart, allocator-last — which enumerates the players triple as literally {S+0x54, S+0x58, S+0x5c} with no frame arithmetic at all, and puts the fleets vector at S+0x64 exactly where B4 measured it. Five NPC accessors at 0x00788de0ff bounds-check an index against ([S+0x58] − [S+0x54]) >> 2 and then index _Myfirst, which is a third confirmation.

    The vector is not the lobby's player list. It is `#empires + one rebel-AI per distinct empire species

    • 4 NPC pseudo-players(Alien Menace, Peacekeeper Enforcer, Von Neumann, Independent Colony — all species 4).Sim.NumPlrsreads **8** in the Human saves (two species) and **7** in every Zuul save (one species), against aSummary.Players` array with 2 entries in both. Both numbers are right; they count different things.

    And my draft of this document was wrong about my own data. It said the hook read 8 "on both saves". It did not: the trace reads 8 on ref-turn2 and 7 on zuul-turn16-noderoute, matching each save's NumPlrs exactly. I generalised from one run without re-reading the other, and it took a check aimed at something else to catch it.

  • Which of the 12-to-44 S+0x8 increments per turn come from where (§5), and what they scale with.

  • The direct-call sweeps behind "node-line decay's only RNG site is the Chance(0.5f)", "its downstream pair draws nothing" and "ProcessNodeSpaceTravel draws nothing" are worth less than they look, and method rule 16 (landed by lane J while this run was in flight) says why: an inlined draw leaves no call-graph edge at all, so a call sweep is a lower bound. Rule 17 applies too — those sweeps clipped at Ghidra's reported function sizes.

    The behavioural measurement is what carries these claims, not the sweeps. ProcessNodeSpaceTravel moved the generator by 0 words on 16 observations (twice per turn, eight turns) and node-line decay by 0 on 8. That evidence is immune to both rules, because it does not ask which function drew — it asks whether the generator moved.

  • No Guard region is declared by any hook in this family, so nothing here can report an undeclared write. That is deliberate — these hooks make no claim about game state at all, and a guard over the generator would only duplicate the Result region that already covers the whole object — but it means the usual harness-audit safety net is absent by design. Both traces show err = 0, undeclared = 0 on 32 records each; the second number is vacuous and should be read that way.

  • The one record with no ledger position is the first pre-turn autosave of each session, which runs before any turn driver and therefore before the hooks know the server pointer. It declares no region at all rather than declaring one it cannot fill. Every record that did declare the region resolved: 0 words: null across 64 records.

  • The ledger's block-chain machinery has never run live. Every observation in both runs sat inside a single MT block — left walked 432 → 413 → 395 → 375 → 357 on ref-turn2 and 263 → 243 → 223 → 201 → 181 on the Zuul save, never reaching 0. So every live word count reduces to left_before − left_after, and the twist-and-index path that makes the instrument correct across block boundaries is exercised only by the host tests. A turn that crosses a boundary (any turn spending more words than left) is the first real test of it. This is the thinnest part of the instrument and the one to watch.

  • Everything here is one game state per save, two saves, eight turns, with 158 words of generator movement in total (192 → 267 and 361 → 443). It is a thin workload measured precisely, not a broad one. In particular the per-turn total moved only between 18 and 22 across eight turns: the variation is barely sampled, and nothing here says what makes it 18 rather than 22.

9. Node-line expiry — a distance, not an absence

Lane O's zuul-turn16-noderoute.sav was pushed to the VM and played forward. Phase 11 draws one word per expired node line; rather than report "we ran N turns and it never fired", the hook was extended to classify the whole NodePath population at entry, using the same RemainingLife formula the original tests. The classification is what makes the negative result usable:

turn node paths permanent (npt == 0) immortal (npdtn == INT_MAX) mortal min remaining life ≤ 5 expired → words
17 53 51 0 2 40 0 0
18 54 51 0 3 42 0 0
19 56 51 0 5 41 0 0
20 57 51 0 6 43 0 0

Three things follow, none of which was knowable before:

  1. 51 of the 53 node lines on this map can never expire — npt == 0 takes RemainingLife's first early-out. The static map's node network is not a decay candidate at all. Only dug lines are, which is why this is a Zuul save: the mortal count rises by roughly one per turn as the Zuul dig.
  2. Every mortal line is ~40 turns from expiry, and the population's minimum stays in a 40–43 band while new lines are added at full life. So the first phase-11 draw on this save is tens of turns away, not one or two — and it is reachable, which "we saw nothing" would not have told anyone.
  3. It explains why the campaign never noticed: no save in the corpus is within 40 turns of a decay event, and the ones that could get there are the newest saves in it.

Phase 11's draw therefore remains a path no save exercises — a hypothesis, and labelled one. What is now instruction-verified is the predicate that decides it (§6.1), and what is measured is the distance.

9.1 The model is stated but not yet tested

predict_words is computed at hook entry, before the original runs, and recorded on every node-line-decay record. It read 0 on every call and the measured delta was 0 on every call. That agreement is worth exactly nothing as a test of the model — it is the "0 diverged while comparing nothing" shape this campaign has already paid for — and it is reported that way rather than as a green tick.

Compare mode was not run on this hook for the same reason: with a prediction of 0 and a measurement of 0, ours would advance the scratch generator by nothing, diff clean, and prove only that the harness works. The model becomes checkable the first time a line expires, and the descriptor is ready for that day.