# Replaying a recorded turn: what a command stream reproduces, and what it does not Lane RB, 2026-09-08. Host work, no VM. Engine worktree `wip/rb` off `main` `4f25f1e`; predictions committed as `sots-engine` `docs/RB-predictions.md` at `1e6474b`, **before** the module existed (rule 2). Every number below was taken on a **fresh `build-host`** created with `rm -rf` (rule 24). Consumes: `ai-order-emission.md` (AI4, the apply order and the cost table), `ai-order-capture.md` (L4, the live dumps), `turncommands-block.md` (Q, the wire shape), `turn1-to-turn2-nondeterminism.md` (L5). Produces the first end-to-end **turn record** the campaign holds. --- ## 0. Lead: `ModCount` is reachable, and it is the only leaf a stream closes on the reference turn `/Sim/ModCount` has been unreachable from a save for the whole campaign, because it counts a thing a save does not contain. Our standalone wrote **14** where the original writes **24**: the two driver bumps and nothing else. Replaying the turn's recorded command blocks puts it on **24**, exactly, with no fitting and no fudge term. Canonical pair `turn2-state.sav -> turn3-state.sav`, fresh build: | | leaves diverging | closed | regressed | |---|---:|---:|---:| | do nothing | 108 | — | — | | standalone, no stream | 63 | 45 | 0 | | **standalone + recorded stream** | **62** | **46** | **0** | The one leaf the stream closes is `/Sim/ModCount`, and its twenty-four decompose with **zero residual**: ``` 2 turn drivers (S00, T00) -- ours already 4 research-rate gates, one per submitting block (16, 32, 496, 512) 1 list 5 system rates 1 list 3 build order 1 list 10 StrategyServer::OnCommand_TransferShips [named by lane W3; was "(unnamed)" here] 2 list 14 fleet task, modes 0 and 1 against ONE fleet 1 list 8 fleet move 0 list 23 population -- the 17..27 half is free -- 12 delta, on a save that carried 12 ``` **Nothing else closes, and that is the honest result rather than a disappointment.** Of the ten commands, three have a handler this engine holds and all three are no-ops on this workload — the AI re-issues the research rate and the planetary-budget sliders the save already carries. The other seven need subsystems we do not have (§4). The secondary pair is where the stream does work, and it produces the campaign's first **matched triple** — an input save, the stream captured from the run that consumed it, and *that run's own autosave*: | oracle | closed by the stream | still diverging | |---|---:|---| | the recording's own autosave (`l4-turn1to2-instrumented-autosave.sav`) | **7** | — | | the historical `turn2-state.sav`, from a different process | **6** | player 512's `ResTNm` | Both at **regressed 0**. The seven are `ModCount` (0 → 12) plus `ResRate` and `ResTNm` on players 32, 496 and 512. Against the historical oracle only six close, and the one that does not is **exactly** the leaf lanes L4 and L5 showed is decided per-process: the recording picked `XNC_TrnsMorr2`, `turn2-state.sav` holds `BIO_GnMod`. That was written down as a prediction before the run and it is the strongest thing in this lane — a replay that had closed it would have meant something was copying the oracle instead of the capture. --- ## 1. The on-disk form: `.tcb` Line-oriented, whitespace-separated, `#` comments, magic first. A shim can emit it with `fprintf` and a lane can read it without a tool. ``` tcb 1 meta source l4-turn2to3-aiorders.txt seed <32-bit word> -- one per AI client; absent means "not recorded" block -- the batch SLOT, ascending; playerId 0 = a slot no client wrote gate rate|target|boost|group4|group5|civilian list -- required for every non-empty list elem ... -- fields in WIRE order ``` A field is one token: `iN` int, `fN` float, `b0`/`b1` bool, `s:TEXT` string, `vN` a counted vector whose length is known and whose values are not, `vN:a,b,c` one that is fully read, and `?` a scalar the instrument could not reach. Three parts of that are not decoration. **`?` and bare `vN` are the point.** A dump reads a fixed window of each element and cannot follow a pointer, so a route arrives as "one element, contents unknown". That is a different fact from "no element" and from "an element of zeros": the command still costs its bump and its effect still cannot be applied. Recording the ignorance is what lets the counter be right while the state is honestly left alone. The reference turn has **four** such elements out of ten. **A declared count that disagrees with the elements present is an ERROR, not a warning.** A capture that lost an element would otherwise produce a counter quietly one short, and nothing downstream could tell that from a turn that really issued one fewer command. Thirteen malformed captures are in the engine's test as rejection cases. **`seed` is carried even though nothing consumes it.** Lane L1 established that each AI client seeds one MT19937 with a single word at construction, so a decision is a function of *(board, seed)* and a capture that records the seeds can be **re-derived** rather than replayed. The field exists now so a capture taken today is still the right file when `game/ai` lands. A reference save whose seeds were never logged — `turn2-state.sav` — is not reproducible by any process, the original included, and its commands can only ever be replayed. ### 1.1 Two capture tools, one reader — and a defect one of them shares with my first attempt Lane CB was building `tools/turncommands_capture.py` in parallel, emitting **JSON**. That is the better *capture of record* and it should stay: raw element words as ground truth, heap vectors and strings the deep dump followed, the AI seeds, the container self-check, and — the part neither of my files had — the input save's hash **bound to the output autosaves' hashes**, so a capture cannot be silently used against the wrong save. Rather than a second format in the engine, `tools/tcb_from_json.py` joins them: CB's JSON stays the capture of record, `.tcb` stays the engine's parser-free input, and one narrow script knows both. Fed CB's tool's own output over the L4 log, the adapter produces a replay **byte-identical** to the one from `aiorders_to_tcb.py`. Honest bound on that agreement: both decoders read the *same log*, so it checks the two decoders against each other and not the log. **And it caught a defect in CB's decoder, which is the same one I shipped and measured (§3):** its list-5 record maps the element's memory words straight onto the frame's wire order (`{systemId, ship, terraform, sciences, …}`). That is wrong by at least one position. The adapter overrides it to `?` and says why in its own docstring; **CB should drop the list-5 record from `RECORDS` rather than rely on the adapter to mask it**, because the JSON is the capture of record and a wrong typing in it will outlive this note. CB's `raw_words` are unaffected and remain right. ### 1.2 CB's real capture, run CB's own VM146 capture landed while this lane was closing and it goes straight through the adapter and the replayer unchanged: `verify/results/turncommands/cb-turn2to3.tcb`. It reproduces this lane's canonical result exactly — **108 → 62, closed 46, regressed 0**, `ModCount` 24 — from a completely separate run of the game on a different guest. It also brings two things L4's dump could not: * **The AI seeds, observed live:** netId **32 → `0x9dfad635`**, **496 → `0xb482b2d8`**, **512 → `0xfb81e993`**, all `pinned=false` (observed, not forced). These are the first seeds any capture has carried, and they are what makes this a turn record a future `game/ai` can be *re-derived* against rather than replayed from. * **The heap payloads**, which close two rows of the gap list below outright: * **list 8's route is `[272]`** — one hop to system 272. Lane L4 §5.3 recorded this as "`turn3-state.sav`'s waypoint says 272 but the capture does not prove it". It is proved now. * **list 10's counted vector is `[1728]`** — a **ship** id, and one that exists in the input save. `{systemId 288, fleetId 34, ships [1728]}` is now a typed reading with values, and "assign these ships to this fleet at this system" fits all three. Still not a proof of the command's *name* — the effect has never been observed — but the payload's type no longer leaves it open. Those two elements move from `incomplete` to `declined` in the replay report: the capture is now complete and the *subsystems* are what is missing. CB's capture also does the one thing neither of my files did and every future capture must: it records the input save's hash **bound to the output autosaves' hashes**, both of which match the published determinism oracle, with a `hooks=off` control run named. That binding is what makes a capture usable as a byte-match oracle at all. **The converter.** `tools/aiorders_to_tcb.py` turns lane L4's shim dump into a `.tcb` mechanically: it applies the per-list field mapping, undoes the one list whose writer runs backwards, reinterprets the words the record says are floats, turns a vector's begin/end pair into a length, and writes `?` where the window could not reach. Both existing captures are converted and checked in at `verify/results/turncommands/l4-turn{1to2,2to3}.tcb`. **Lane CB does not need to write an emitter:** the existing `aiorders=on` hook plus this converter already produces the file. If CB does emit `.tcb` directly, the seed and `name` records are the two things the current dump has no field for. --- ## 2. Apply order, and how it was verified The batch is a flat run of thirty steps: twenty-seven per-**list** steps, each looping over *every* player's block before the next step begins, with three per-**player** gate loops spliced in. ``` lists 6 11 20 19 17 18 5 23 24 gate loop A { group5 (free), research target (bump), research rate (bump) } 0x0088fdb0 lists 1 4 3 21 2 22 9 10 12 13 14 15 16 7 8 25 27 26 gate loop B { research boost (bump) } 0x008907b1 gate loop C { group 4 (bump) } 0x0089080a ``` So **one player's list-6 commands are applied before another player's list-11 commands**, the list sequence starts at 6 and is not 1..27, and the member offsets it walks are not ascending either. **How it was verified, and what the verification cannot show.** 1. **Address monotonicity, nine of thirty positions.** Six appliers are inlined into the batch and each writes `ModCount` in place, so a watchpoint run recovered their addresses; the three gate loop heads are known too. In schedule order those nine are `0x0088fdb0`, `0x0088fe0a`, `0x008902fe`, `0x008903b9`, `0x0089046c`, `0x008905c8`, `0x008907b1`, `0x008907bc`, `0x0089080a` — **strictly increasing**, at steps 9, 9, 18, 19, 20, 23, 28, 28, 29. That chain is an independent re-derivation. The other twenty-one lists' handlers are out of line and this lane has no record of their call-site addresses inside the batch, so their relative order is **inherited** from AI4's read of the `add edi, imm` chain, not re-derived. The test says nine. 2. **Bijection and non-sortedness**, asserted: every list exactly once, every located gate exactly once, and neither the list sequence nor the offset sequence ascending. That is what rules out the two obvious wrong implementations — `for (list = 1..27)` and walking the block in memory order. 3. **The civilian-ratios gate is absent from the schedule on purpose.** It has no located applier anywhere in the routine, so including it would be claiming a cost of zero for something whose cost is unknown. It is caught separately and makes the whole count report itself as a lower bound. **Stated plainly: apply order is unfalsifiable on every workload the campaign holds.** Both captures put every non-empty list on one player and every command on one system, so any permutation produces the same save and the same count. The order is implemented for the workload that will need it, and tested against the instruction stream rather than against an outcome. The workload that would make it falsifiable is **two players commanding the same object in one turn** — the cheapest is a two-human `/concurrent` game (lane G2's Tier 0), where both clients order fleets at one system. --- ## 3. A falsified prediction, and the finding it paid for **RB-P3 predicted `regressed 0`. The first run regressed two leaves**, and the cause is a real fact about the original. The converter mapped list 5's dumped words straight onto the wire order of the rates frame (`SRs, SRt, SRsc, SRtf, SRi, SRoh, SRnr`). The replay then wrote the AI's single non-zero slider into `SRt`, and `Sys[288 "Ke'Dolarra"]` came out with `SRt 1.0 / SRsc 0.0` against an oracle holding `SRt 0.0 / SRsc 1.0`. **The memory field order of `Game::StarSystem::OutputRates` is not its wire order.** What is known precisely: * the only non-zero word in every dumped element of list 5, on both turns, is at **memory index 2** of the element (index 0 is the system id, so it is the frame's **second** member); * the same command **on the wire**, in a save that carries issued orders, puts its only non-zero in **`SRsc`**, the frame's **third** member (lane Q, `SAVE_FORMAT.md` §11, cross-checked to the Planetary Budget slider pushed fully to Construction); * `turn2-state.sav` and `turn3-state.sav` both hold `Sys[288] Rts = {SRs 0, SRt 0, SRsc 1.0, SRtf 0, SRi 0, SRoh 0, SRnr 0}` — so the AI's command re-issues the state the save already holds, which is why a *correct* applier is invisible and an incorrect one is immediately visible. So **memory member 1 is wire member `SRsc`**: one correspondence pinned, six unread. One non-zero slider cannot determine a permutation of seven, and the converter no longer pretends otherwise — it emits the system id and seven `?`, and the replayer counts the command and declines it. Two experiments settle it, both cheap: * **one UI run** — push two *different* sliders to two *different* values on one system, End Turn with `aiorders=on`, and read the permutation straight off the element; * **cheaper, and no VM at all** — a save taken after issuing rates carries the same command on the **wire**, where every field is NAMED. `zuul-turn17-orders2.sav` has one. A `.tcb` converted from a save's own `TurnCommands_v5` block needs no memory mapping, and would also give the `.tcb` format a second, independent producer. **This is the highest-value next step on the capture side** and it is pure host work. Two notes on how this was caught, because they generalise. The plain closed/regressed measurement found it, before the control did — but the control (RB-P4, `--replay-count-only` vs a full replay must be byte-identical when every modelled handler is a no-op) would have found it too, and it is what now stands guard: on the canonical pair those two runs are byte-identical, which is the evidence that the handlers we *do* run write where they claim to rather than agreeing with the oracle by luck. And the reason a wrong write was visible at all is that the command re-issues existing state: **a command that re-states the board is the best possible test of an applier**, and the corpus is full of them. --- ## 4. What a *complete* replay needs that this lane does not have This is the real Rung B backlog. Each row is the reason one command in the reference turn is counted but not applied. | # | needed for | what is missing | shape | |---|---|---|---| | 1 | **list 3 build, list 1 design** | **ship construction.** No phase in this engine builds a ship. The command carries a queue ordinal, a design id and a system; the effect is `srb`/`sri`, `Maint`, the savings debit, a hull id and `ShipIDs`. Nine leaves on the canonical pair. | engine work; lane B6/E2 have the map | | 2 | **list 1, list 3, list 8, list 10, list 14** | **the client's id allocator.** The build order names design **18** and the fleet order names fleet **34** — objects that do not exist in the input save, allocated *client-side before submission*, while the server's own master counter (`NMnx`) issued 1712 and 1776 the same turn. **Two id spaces**, and the small one is part of the wire protocol. A reimplementation that allocates on apply produces a structurally correct save with every AI-created id wrong. | **a watchpoint**, not a week of reading: break on the write that produces 18 and 34 | | 3 | ~~**list 8 fleet move** — the route's hops~~ | **CLOSED by lane CB's deep capture (§1.2): the route is `[272]`, one hop.** What remains is row 2 — the route belongs to a fleet the input save does not contain. | done | | 4 | **list 5 system rates** | **the memory field order of the rates frame** (§3). One correspondence pinned, six unread. | one UI run, or one save-sourced capture | | 5 | **list 10** | **a name and a meaning.** The payload is fully read now (§1.2): `{systemId 288, fleetId 34, ships [1728]}`, and 1728 is a ship the *input* save already contains. "Assign these ships to this fleet at this system" fits every word. The **effect** has still never been observed, so it is still not named. | a hook on `0x0088bed0` | | 6 | **list 14 fleet task** | **what the two modes do.** Two elements per AI fleet order, modes 0 then 1; the interface deposits one, mode 0 only. The *cost* is settled; the *effect* is not read at all. | a hook on the inlined applier | | 7 | **list 23 population** | **the `Population` body**, 24 bytes behind a vftable that the capture window does not follow. Free in `ModCount`, so it has never been forced. | one dumper change | | 8 | **the research-target gate** | **the techId → tech-name map.** The wire carries an integer (144, 90, 288); the save carries a name. The client resolves it *off the command* and passes a `char*`. The ids are not `index * 16` and are not indices into anything we hold. Today the capture can carry the observed name and the replay TRANSCRIBES it — reported in its own column, because that is not a reimplementation. **Three data points now exist**: 144 → `IND_Waldo`, 90 → `DRV_PlsFiss`, 288 → `XNC_TrnsMorr2`. | a watchpoint at the gate payload, or a read of `0x006c8580` | | 9 | **the research-boost gate** | savings spent to advance research; both halves unmodelled. Never observed set on an AI turn. | engine work | | 10 | **the group-4 and group-5 gates** | no read semantics; group 5 is Hiver-only and no save carries it. | rule 6 — a manufactured Hiver workload | | 11 | **the civilian-ratios gate** | **no applier located anywhere in the batch.** Its `ModCount` cost is *unknown*, not zero, and any turn that sets it reports a lower bound. | an image-wide search for the consumer | | 12 | **the load-time batch** | the process applies a batch at LOAD as well as at End Turn (`seq=1`, `n=1`, the local client's block alone, with the rate gate SET). Whether it charges `ModCount` is **untested**, and this lane excludes it by construction. If it does charge, a save loaded and immediately re-saved reads one higher. | one run: load, save, compare | | 13 | **the whole thing, on any interesting board** | **two turns, one AI empire, 28 stars, no contact.** Lists 2, 4, 6, 9, 11–13, 15–22 and 24–27 have never been non-empty. Apply order is unfalsifiable (§2). Nothing here generalises past a very quiet game. | manufacture the workload | Two things NOT on that list, deliberately. `Summary.Checksum` moves whenever anything else does and its inputs are unread — no command replay will touch it until the rest is right. And nothing in either capture draws a generator word, so the RNG frame is untouched by the stream; the coordinator's note that the replay interval begins at `BeginProcessTurn` rather than `ProcessTurn` is already satisfied here — the batch is drained *before* the first phase runs, which is where the End-Turn dispatcher calls it (`ApplyAllTurnCommands` at `+0x00784904`, `ProcessTurn` at `+0x0078491c`, `OnAllCombatDone_Tail` at `+0x00784d07`) — but the **hive-registration draws lane L1 found inside `BeginProcessTurn` are a separate interval and this lane does not model them.** --- ## 5. What this lane did not do 1. **No VM time and no new instrument.** Every live number here is lane L4's, re-read through a converter. Nothing new ran under a hook. 2. **The CT111 shim cross-build was not run** — CT111 refused the key from this host. `src/game/ai` is in the cross build, so per rule 13 **the integrator must run it before pushing**. The two 32-bit-specific hazards were audited by hand and one was real and is fixed: `strtoul` plus a `> 0xffffffffUL` test is tautological where `unsigned long` is 32 bits, which would both accept out-of-range input silently and trip `-Wextra`/`-Werror`; both sites now use `strtoull`. `%zu` appears in the new code and also in `src/app/main.cpp` today, so it is not a new exposure. 3. **The apply order was not tested against an outcome** and cannot be on this corpus (§2). 4. **`ToTurnCommandBlock` reuses the existing, already-verified cost model** rather than counting a second time — deliberately, so there is one implementation of the arithmetic and not two that can agree with each other while both being wrong (rule 8). 5. **The `.tcb` format has exactly one producer.** Until the save-sourced converter of §3 exists, every capture comes through one script and a bug in it is invisible. 6. **Seeds are carried and never used.** No capture in the corpus has any.