# The RNG ledger for one strategic turn — measured, not inferred Lane Z, 2026-09-08. Program `sots` / "Sword of the Stars.exe", ImageBase 0x00400000, all addresses VAs. Engine worktree `wip/tailrng`, `sots-engine/docs/Z-tail-rng.md` (the prediction, committed before the build). **The question.** The milestone is *a standalone that loads a save, runs one strategic turn, and writes an autosave that byte-matches the original's*. Generator state is part of the saved state. Lane K (`combat-done-tail.md` §3) found two draw sites in `StrategyServer::OnAllCombatDone_Tail` that nothing models and that both run **before** the autosave, and concluded that a reimplementation reproducing both `ProcessTurn` functions exactly would still diverge. Nobody had measured what a turn actually costs. **The answer, up front.** Across eight measured End Turns on two saves, a strategic turn advances the strategic generator by **18–22 words**, *all* of it inside `StrategyServer::ProcessTurn`, and the residual outside the two turn drivers is **exactly zero**. The tail's cost on these turns is **0**. The defect lane K found is real and **latent**: it will bite the first turn a node line expires or a real battle resolves, and our saves reach neither. --- ## 0. Where to start if you are building the standalone Three sentences, then the evidence. 1. A turn costs **18–22 generator words**, all of it inside `StrategyServer::ProcessTurn`; the two autosave files bracket exactly that interval and nothing draws between them outside the two turn drivers. 2. `verify/results/shim/tailrng/z-t6-endturn.sav` → `z-t6-autosave.sav` is a **byte-identical oracle pair with a known RNG cost of 18 words**, verified from the file bytes independently of the live hook. Test against that pair before any other. 3. Lane K's warning stands and is now quantified: the tail's cost is 0 **today** because no save is within ~40 turns of a node-line expiry (§9) and no encounter has ever produced a battle (§8). Both terms are real; both are latent. ## 1. The instrument, and why it is not an RNG hook Counting draws by hooking the primitives would have undercounted, and the static work this lane commissioned says by how much. The image has **four** draw entry points, not three: | VA | primitive | words per call | |---|---|---| | 0x0047d830 | `Mars::RNG::NextFloat(&mt)` — `__thiscall`, no stack args, plain `ret` | **exactly 1** | | 0x004271c0 | `Mars::RNG::NextInt(&mt, uint* pMax)` — `ret 4`, rejection loop, bound re-read each iteration | **1 or more** | | 0x008e6dd0 | `Mars::RNG::Chance(RNG*, float p)` — `ret 4`; takes the **object** base and does the `add ecx,4` itself | **0 or 1** — see §6.1 | | **0x004f7670** | **`Mars::RNG::NextUInt(RNG*)`** — plain `ret`, no stack args, 84 B. **In no previous lane's primitive set**, and called from 0x007b6700 inside `ProcessTurn`'s closure | 1 | (Lane J's scan reaches 0x004f7670 from the other direction and calls it one of the fourteen *inlined-draw* functions rather than a primitive. Both readings are of the same object and neither is wrong: it is a small function whose draw is inlined, so it is a draw site that calls no primitive — which is exactly why a call-graph sweep for `NextFloat`/`NextInt` cannot see it.) plus **inlined draws**: this lane's scan for the MT tempering immediates found twelve such functions, two of them reachable from the turn roots (0x007aa240 under `ProcessTurn`, 0x007a7f30 under the combat resolver). Lane J's independent scan, landed while this run was in flight, found **fourteen** and turned it into method rule 16 — *inlined draws are invisible to call-graph sweeps, so any RNG accounting built from the call graph alone is a lower bound.* Take lane J's number over this one; the two scans agree on the shape and the point is the same either way. A primitive-counting hook would have missed every one of them silently. **This instrument is not built from the call graph**, which is why rule 16 does not apply to it: it reads the generator's state before and after a boundary and reports the difference. An inlined draw, a draw in a function nobody has named, a draw through a vtable — all of them move `left`, and all of them are counted. So the instrument reads **state**, not calls. `Mars::RNG` is `{void* vptr; uint32 mt[624]; uint32* next; int32 left}`, `sizeof 0x9cc` — `next` at +0x9c4 is a **pointer** into the block and `left` at +0x9c8 is the counter, and the three inner primitives are entered with `ECX = &mt[0] = RNG+4`, which is what lane T's "the generator is entered at `rng+4`" note was seeing. The draw is `if (left == 0) Twist(); y = *next++; --left;` — a pre-check against **0**, never −1. The block transform is a pure function, so the blocks a generator visits form a forward-only chain. `RngLedger` (`sots-engine/src/shim/hooks/rng_ledger.{h,cpp}`) indexes that chain from the first block it sees and positions any state exactly: ``` words(block, left) = block * 624 + (624 - left) ``` Differences between positions are then exact **across twists, across rejection loops, and across draws nobody hooked**. Nine host tests pin the arithmetic, including the block boundary (`left == 0` is a real state and must not be off by 624), a rejection loop counted against a shadow generator, and the out-of-order case below. **One ordering subtlety, and it is load-bearing.** `Hook<>` takes the `before` snapshot at entry but renders it (calls the region's `describe`) only *after* the original returns — so for nested calls the inner snapshots are rendered first, and by the time an outer `before` is rendered the chain has moved past it. Walking a Mersenne Twister backwards is not possible. Every hook therefore **observes at entry from `describe_args`**, which indexes the entry block while it is still current; the render then resolves it from the memo. Without that, every outer-call `before` would read `words: null`. Six nested trace hooks, all watching the object at `S+0x16c`: `StrategyHost::Autosave` (both markers) › `StrategyServer::ProcessTurn` › `OnAllCombatDone_Tail` › `ApplyEncounterResult` (phase 6) › node-line decay (phase 11) › `ProcessNodeSpaceTravel` (runs twice a turn). --- ## 2. The ledger Save `ref-turn2` (2-player Morrigi vs AI), four consecutive End Turns, build `z-tailrng-20260908T1314Z`, `hooks=trace`. Words are 32-bit MT outputs consumed by the generator at `S+0x16c`. | turn | `ProcessTurn` | tail | `ApplyEncounterResult` | node-line decay | `ProcessNodeSpaceTravel` ×2 | **bracket total** | **residual** | |---|---|---|---|---|---|---|---| | 3 | **19** | 0 | 0 | 0 | 0, 0 | (incomplete — see below) | — | | 4 | **18** | 0 | 0 | 0 | 0, 0 | **18** | **0** | | 5 | **20** | 0 | 0 | 0 | 0, 0 | **20** | **0** | | 6 | **18** | 0 | 0 | 0 | 0, 0 | **18** | **0** | "Bracket" is `Autosave(endTurn=1)` → `Autosave(endTurn=0)`: the pre-turn save file to the post-turn save file, which is exactly the interval a standalone has to reproduce. The turn-3 bracket is incomplete **by construction** and is reported rather than dropped: the pre-turn autosave of the first End Turn after a load runs before any turn driver, so the hook has no server pointer yet and its record carries `words: null`. A second save, `zuul-turn16-noderoute` (Zuul vs Zuul), build `z-tailrng2-20260908T1328Z`: | turn | `ProcessTurn` | tail | `ApplyEncounterResult` | node-line decay | `ProcessNodeSpaceTravel` ×2 | **bracket total** | **residual** | |---|---|---|---|---|---|---|---| | 17 | **20** | 0 | 0 | 0 | 0, 0 | (incomplete) | — | | 18 | **20** | 0 | 0 | 0 | 0, 0 | **20** | **0** | | 19 | **22** | 0 | 0 | 0 | 0, 0 | **22** | **0** | | 20 | **20** | 0 | 0 | 0 | 0, 0 | **20** | **0** | **So: we consume 18–22 words per turn and model 0 of them.** Not 0 because the modelling is bad — because *nothing in the repo models any part of a turn's RNG consumption as a count*. B1/B3/B4 compare the generator's state around three specific functions and get it right; no lane has ever stated a turn's total. This table is that statement. ### 2.1 The generator does not move outside the turn pipeline Every turn's `ProcessTurn` entry position equals the previous turn's post-turn autosave position, exactly: 211 → 211, 229 → 229, 249 → 249. The UI, the renderer and the per-frame tick draw **nothing** from the strategic generator between turns. For the standalone this is worth as much as the total: the interval it must reproduce is closed. --- ## 3. An independent check, from the files rather than from memory The two autosaves of the turn-6 bracket were pulled off the VM and their `Sim.RNG` blobs parsed by `verify/save-reader/save_reader.py` (the frame's payload is 2503 bytes: `mt[624]`, then `left` as int32 at +2496). Twisting the pre-turn block forward until it matches the post-turn block, and applying the same position formula: ``` z-t6-endturn.sav (pre-turn) left = 375 z-t6-autosave.sav (post-turn) left = 357 twists = 0 -> words consumed between the two files = 18 ``` The live ledger recorded 18 for that turn, `left` 375 → 357. **Two instruments that share no code path — one reading process memory through a hook, one reading gzip-compressed file bytes through the save reader — agree exactly.** Files and the checker in `verify/results/shim/tailrng/`. **And the two do not share a hidden assumption** (the trap of method rule 8, which is live here because both sides know how to twist an MT block). `twists = 0`: the block is byte-identical in the two files, so the file-side number is `left_before − left_after` and involves the twist implementation **not at all**. The agreement is therefore about the game's behaviour, not about two copies of the same algorithm agreeing with each other. This is the pair a standalone should be tested against first: it is a byte-identical oracle *with a known RNG cost attached*, which none of the eleven corpus saves has. --- ## 4. Lane K's inference, settled > **§6, labelled hypothesis:** "I did not prove that `SNMAllCombatDone` is delivered on turns with no > combat." **The handler runs on every End Turn.** Four out of four, `OnAllCombatDone_Tail` recorded exactly one call per End Turn, at depth 0, between the two autosaves, with the post-turn autosave following it. The determinism-note inference was right. The stronger claim — that it runs with an *empty encounter vector* — is **not** settled by this workload and must not be reported as though it were. On `ref-turn2` the encounter vector is **empty at `ProcessTurn` entry and holds exactly one encounter by the time the tail runs** on every one of the four turns: detection (`ProcessTurn` phase 31) creates it, and the tail's phase 7 clears it. So what is proved is "the tail runs on a turn with **no battle**", not "on a turn with no encounter at all". See §8 for what closing the remaining gap needs. Two things fall out of the same records and are worth more than the phrasing: * **Phase 7 really is a wholesale `clear()`.** `encounters` reads **1** at the phase-6 `ApplyEncounterResult` call and **0** at the phase-11 node-line-decay call, on every turn. Lane K read that off the instruction stream against a decompile that reads as a conditional prune; it is now also a behavioural fact. * **Every encounter on these turns has `res->+0x4 != 0`** (`res_no_battle = 1`), the flag that makes `ApplyEncounterResult` a whole-function no-op. Its measured cost is 0 words, which is what that gate predicts, and which is why the combat resolver has never run under any instrument this campaign has built. --- ## 5. Correction to `combat-done-tail.md` §7.1 — and to this lane's own prediction Lane K wrote that `S+0x8` "advances **at least twice** per turn". That is correct and this lane's prediction that it advances **exactly** twice is **wrong**. Observed `S+0x8` at hook entry: | turn | `ProcessTurn` entry | tail entry | tail's callees | next turn's `ProcessTurn` entry | |---|---|---|---|---| | 3 | 22 | 23 | 24 | 34 | | 4 | 34 | 35 | 36 | 48 | | 5 | 48 | 49 | 50 | 60 | | 6 | 60 | 61 | 62 | — | The two drivers account for 2 of the **12 to 14** increments per turn. Ten to twelve more happen between the post-turn autosave and the next `ProcessTurn`, from a writer this lane did not identify (both drivers are hooked, so it is neither of them). Lane K's operational conclusion is strengthened, not weakened: **`S+0x8` must never be treated as a turn number.** `S+0xc` (`ModCount`) reads 3, 4, 5, 6 across the same four turns and is the turn counter. ## 6. Corrections to `combat-done-tail.md` §3 and §6.1 Both from the instruction stream, both load-bearing for anyone reimplementing these functions. **§3 — the node-line fleet check does not gate the roll.** Lane K: *"The roll is skipped for a line if any fleet with flag `0x20000` is targeting it."* The straight-line order in node-line decay's first loop is ``` 0x007ae088 call NodePath::RemainingLife ; expiry test 0x007ae08f jg 0x007ae1e2 ; not expired -> next record, NO DRAW 0x007ae095 mov ecx,[esi+0x16c] ; THE DRAW 0x007ae0a5 call 0x008e6dd0 ; Mars::RNG::Chance(0.5f) 0x007ae0aa test al,al ; je 0x007ae1e2 ; roll failed -> next record 0x007ae0b2 ... ; THE 0x20000-FLEET SCAN STARTS HERE ``` The scan begins 0x1d bytes **after** the `Chance` call and is reached only when the roll *succeeded*. It suppresses the collapse (`0x007a92e0` / `0x007a4700`), never the draw. Lane K's headline — one `NextFloat` per expired node line per turn — survives intact and is now pinned to a formula. **§6.1 — `StrategyHost::Autosave` is `ret 8` and returns a value.** Its epilogue is `c2 08 00`, and `0x00895b5c mov eax,esi` puts the `std::string*` (the MSVC named-return slot) in EAX. A hook declaring it `void` drops EAX at both call sites. Neither call site passes `this`: both hardcode `mov ecx,0xb29f98`. (The `+0x54` candidate on that global was recorded live and is **not** the `StrategyServer` — `server_agrees` is false on all eight autosave records, so the global that the autosave uses is a different object from the `StrategyHost` whose `+0x54` `OnMessage` reads.) ### 6.1 The expiry predicate, now concrete `NodePath::RemainingLife` 0x006e2130, `__thiscall(NodePath*, int turn)`, `ret 4`, whole 122-byte body read: ```c if (npt(+0x04) == 0) return INT_MAX; // permanent line, never expires if (npdtn(+0x1c) == INT_MAX) return INT_MAX; // immortal line aged = (npctm(+0x14) >= 0 && turn >= npctm) ? turn - npctm : 0; wear = (npdtf(+0x20) != INT_MAX && npdtf > 0) ? nptf(+0x24) / npdtf : 0; // SIGNED idiv rem = npdtn - wear - aged; return rem > 0 ? rem : 0; // callee-side clamp ``` A line is expired exactly when this returns 0. **The lifetime is derived, never ticked** — the function writes nothing, and neither does the loop around it — so there is no decrement-ordering question and a snapshot at function entry is a valid prediction basis. `nptf` is never sign-checked, so the division's signed truncation must be reproduced literally. `Chance` 0x008e6dd0 returns **false with no draw** when `p <= 0` and **true with no draw** when `p >= 1`; `0.5f` takes neither, so it is exactly one word, and the comparison is `p > r` (equality returns false). A NaN `p` falls through both early-outs and *does* draw — irrelevant here, noted because it is the kind of edge a reimplementation gets wrong. --- ## 7. One generator, confirmed twice Static: the image has one persistent strategic `Mars::RNG`, at `S+0x16c`, constructed by the `StrategyServer` ctor 0x007d78d0 (`push 0x9cc` + `Seed(0)` at 0x007d7d25/0x007d7d3d) and reseeded only from `Read` and `LoadGame`. `StrategyClient+0x134` and `Mars::CombatSim+0x108` exist but are unreachable from the turn roots; three more are stack temporaries in map generation. **The combat resolver draws from the same `S+0x16c` object** — all three RNG entry points in its 750-node direct-call closure load `[reg+0x16c]`. Behavioural: across 32 ledger observations in four turns, **every** state resolved on a single forward chain — no `words: null`, no second chain. If a second generator had been in play, the ledger would have said so by construction. --- ## 8. What is not settled, listed as loudly as the results * **The combat resolver has still never run under an instrument.** Every encounter this workload produced had the no-battle flag set, so `ApplyEncounterResult` was a no-op every time; its measured 0 words says nothing whatever about combat's RNG cost, and the residual-0 result above holds only for turns with no battle. Lane J read all 7,641 bytes of it in parallel with this run (`combat-resolver.md`; and note **7,641**, not the 7,499 Ghidra reports — method rule 17). Its conclusion pairs with this one exactly: **the resolver has no unconditional draw.** All three sites in its subtree are conditional — a node-cannon `NextInt`, an inlined `NextFloat` per back-engineering candidate, and a `NextInt` per successful roll of that. So lane J's cheap first prediction is directly testable with this instrument: **a plain fleet battle with no node cannon and no salvage should cost the same 18–22 words as a peaceful turn.** That is the next run this hook family should do, and it needs a workload nobody has built yet: a save where two hostile fleets actually meet. * **Node-line expiry did not fire.** See §9 for the quantified distance rather than an absence. * **A turn with a genuinely empty encounter vector was not observed** (§3). Every turn of `ref-turn2` in contact produces exactly one sighting encounter. The tail-runs-every-turn claim is settled; the no-encounters variant is still an inference, now a much narrower one. * **`players` reads 8 on a 2-player save.** `(S+0x54)` enumerated as a `ServerPlayer*` vector gives 8 on every record of a game the lobby shows as 2 players. Either the server always allocates a fixed slot count, or this offset enumerates something else (spare capacity is the failure this campaign has already paid for once, at `S+0x64`). **Nothing in the ledger depends on it** — it is decoration on the argument record — but it should not be reused until someone resolves it. * **Which of the twelve-to-fourteen `S+0x8` increments per turn come from where** (§5). * The direct-call sweeps behind "node-line decay's only RNG site is the `Chance(0.5f)`", "its downstream pair draws nothing" and "`ProcessNodeSpaceTravel` draws nothing" are **worth less than they look**, and method rule 16 (landed by lane J while this run was in flight) says why: an inlined draw leaves no call-graph edge at all, so a call sweep is a lower bound. Rule 17 applies too — those sweeps clipped at Ghidra's reported function sizes. **The behavioural measurement is what carries these claims, not the sweeps.** `ProcessNodeSpaceTravel` moved the generator by 0 words on **16** observations (twice per turn, eight turns) and node-line decay by 0 on **8**. That evidence is immune to both rules, because it does not ask which function drew — it asks whether the generator moved. * **The ledger's block-chain machinery has never run live.** Every observation in both runs sat inside a single MT block — `left` walked 432 → 413 → 395 → 375 → 357 on `ref-turn2` and 263 → 243 → 223 → 201 → 181 on the Zuul save, never reaching 0. So every live word count reduces to `left_before − left_after`, and the twist-and-index path that makes the instrument correct across block boundaries is exercised **only by the host tests**. A turn that crosses a boundary (any turn spending more words than `left`) is the first real test of it. This is the thinnest part of the instrument and the one to watch. * Everything here is one game state per save, two saves, eight turns, with **158 words** of generator movement in total (192 → 267 and 361 → 443). It is a thin workload measured precisely, not a broad one. In particular the per-turn total moved only between 18 and 22 across eight turns: the *variation* is barely sampled, and nothing here says what makes it 18 rather than 22. ## 9. Node-line expiry — a distance, not an absence Lane O's `zuul-turn16-noderoute.sav` was pushed to the VM and played forward. Phase 11 draws one word per **expired** node line; rather than report "we ran N turns and it never fired", the hook was extended to classify the whole `NodePath` population at entry, using the same `RemainingLife` formula the original tests. The classification is what makes the negative result usable: | turn | node paths | permanent (`npt == 0`) | immortal (`npdtn == INT_MAX`) | **mortal** | min remaining life | ≤ 5 | expired → words | |---|---|---|---|---|---|---|---| | 17 | 53 | 51 | 0 | **2** | 40 | 0 | 0 | | 18 | 54 | 51 | 0 | **3** | 42 | 0 | 0 | | 19 | 56 | 51 | 0 | **5** | 41 | 0 | 0 | | 20 | 57 | 51 | 0 | **6** | 43 | 0 | 0 | Three things follow, none of which was knowable before: 1. **51 of the 53 node lines on this map can never expire** — `npt == 0` takes `RemainingLife`'s first early-out. The static map's node network is not a decay candidate at all. Only *dug* lines are, which is why this is a Zuul save: the mortal count rises by roughly one per turn as the Zuul dig. 2. **Every mortal line is ~40 turns from expiry**, and the population's minimum stays in a 40–43 band while new lines are added at full life. So the first phase-11 draw on this save is **tens of turns away**, not one or two — and it is reachable, which "we saw nothing" would not have told anyone. 3. It explains why the campaign never noticed: no save in the corpus is within 40 turns of a decay event, and the ones that could get there are the newest saves in it. **Phase 11's draw therefore remains a path no save exercises — a hypothesis, and labelled one.** What *is* now instruction-verified is the predicate that decides it (§6.1), and what is measured is the distance. ### 9.1 The model is stated but not yet tested `predict_words` is computed at hook entry, before the original runs, and recorded on every node-line-decay record. It read **0** on every call and the measured delta was **0** on every call. That agreement is worth exactly nothing as a test of the model — it is the "0 diverged while comparing nothing" shape this campaign has already paid for — and it is reported that way rather than as a green tick. Compare mode was **not** run on this hook for the same reason: with a prediction of 0 and a measurement of 0, `ours` would advance the scratch generator by nothing, diff clean, and prove only that the harness works. The model becomes checkable the first time a line expires, and the descriptor is ready for that day.