# The command stream, captured and bound to the autosave it produced Lane CB, 2026-09-08. Guest **VM146** (`sots-re-win10-146`, 192.168.10.146), build `cb2-20260908T2251Z`, worktree `wip/cb` off `main` `4f25f1e`. Predictions committed before the instrument changed: `sots-engine` `docs/CB-predictions.md` (commit `4457249`, before `ai_orders.cpp` was touched; the seed addendum at `c3d95dd`, before the turn-1 runs). Paired with lane **RB**, which builds `sots_turn --turn-commands `. This document is the input side. Captures live in `verify/results/turncommands/`; raw logs in `verify/results/shim/aiorders/cb-*.txt`. --- ## 0. Lead **The canonical pair is captured and it binds.** `ref-turn2.sav` + one End Turn, with the block dump installed, wrote `(Autosave EndTurn).sav` = `bb4fd9ac89f41e3b…` (66,732 B) and `(Autosave).sav` = `978041acd168b56e…` (67,219 B) — **both byte-identical to the published determinism oracle and to this lane's own `hooks=off` control taken on the same guest through the same click path**. So the stream and the autosave come from the same run, and the instrument did not change the turn it recorded. `(Autosave).sav` is byte-identical to `turn3-state.sav`, which is the file lane RB's replay has to reproduce. `verify/results/turncommands/cb-turn2to3.json` carries, for that one run: * the **complete block set** — 8 block slots, 6 gates and all 27 lists each, with every element; * the **three heap payloads no previous capture could read** — list 8's route, list 10's counted vector, list 23's `Population` body; * the **three AI client seeds** (lane L1's per-process words), `0x9dfad635` / `0xb482b2d8` / `0xfb81e993` for net ids 32 / 496 / 512; * the **input save and both output autosaves**, with sizes, sha256s, and the oracle comparison. Two things this lane did not expect to produce and did: 1. **The creation turn made reproducible.** `turn1-state → turn2` is the workload three lanes ran and got three different files from. With the AI client seeds **pinned** to the values an earlier run observed, a *different process* reproduced that run's command block — including the research pick that varies — and its autosave **byte for byte**. Lane L1's mechanism confirmed from the other end; details in §4. 2. **An independent, empirical measurement of every element's size**, from two runs on two guests that wrote byte-identical autosaves. Any element word that differs between them cannot be a word the applier reads. The measured sizes agree with the record sizes read from the writer's instruction stream on all six lists — two methods, no disagreement. 3. **Two corrections to lane L4's list-23 reading**, and one to my own list-5 record — the latter found by lane RB while consuming this capture, which is the pairing working as intended. --- ## 1. What the canonical stream is `turn2-state.sav` (`ab4ac2d7…`, 66,739 B) → one End Turn → `turn3-state.sav` (`978041ac…`). Batch `seq=2`, `n=8` blocks at stride 0x1b4. Slots 0–3 carry player ids **16, 32, 496, 512**; slots 4–7 carry `playerId = 0` with uninitialised gate payloads (lane L4's P1 falsifier, reproduced). **Every one of the four real players has the research-rate gate set** and no other gate; no research-target gate is set on this turn by anybody. Only player 32 emits commands: | list | n | wire record, as captured | |---|---:|---| | 3 build | 1 | `{ordinal 2, designId 18, systemId 288, 0}` | | 5 system rates | 1 | `{systemId 288, 0, 1.0f, 0, 0, 0, 0, 0}` | | 8 fleet move | 1 | `{fleetId 34, nHops 1, [272]}` | | 10 | 1 | `{systemId 288, fleetId 34, count 1, [1728]}` | | 14 fleet task | 2 | `{34, 0, true}` and `{34, 1, true}` | | 23 population | 1 | `{systemId 288, vptr, 24-byte body}` | The three rows in bold type below are new; everything else reproduces lane L4 exactly, from a different guest and a different build, which is worth stating as the corroboration it is. **P3a held. The route is system 272.** Lane L4 could not read it and recorded that `turn3-state.sav`'s waypoint says 272 without being able to prove the command said so. It does. The waypoint in the output save and the route in the input command are the same id, so a replayer needs to carry only one. **P3b half-held, and the falsified half is the useful one.** List 10's counted vector holds one value, **1728** — a *master-counter* id (a multiple of 16), **not** a small client-allocated one in the 18/34 space. So the client allocates the fleet (34) and the design (18) out of its own small counter but names the ship out of the server's. Lane L4's open question — "the client id counter is not located, and 18 and 34 differ by 16, the master stride" — now has a boundary: **the two id spaces coexist inside a single element**, and whatever the client counter is, it is not what names ships. That is a sharper statement of the hole than the one L4 left, and it is still a hole. **P3c half-held, and the falsified half is a better answer.** List 23's body is 24 bytes, as predicted: `[0x00bd8d50, 1, 2, , 1000000000, 0]`. The first word is a vftable pointer in the same neighbourhood as the element's own, so this is **one 24-byte polymorphic object**, not six ints, and the `1000000000` is population-shaped. I predicted it would **differ between the turn-1 and turn-2 workloads**, on the strength of lane L4's `+1` / `−1` trailing. It does not. Captured on both turns, **five of the six words are identical** and the only one that moves is word 3, a heap pointer: ``` turn2 -> 3 [12422480, 1, 2, 0x3369f4c8, 1000000000, 0] turn1 -> 2 [12422480, 1, 2, 0x343ba3e8, 1000000000, 0] ``` So the `+1`/`−1` that made it look turn-dependent was the noise word outside the element (§2), and the payload itself is the same on every turn anyone has captured. It is still **not named** — this lane says what bytes it carries and nothing more (rule 6) — but a replayer is not going to be caught out by it on these two workloads. --- ## 2. The element sizes, measured rather than read Lane L4 dumped a fixed 12-word window; this lane's is configurable and was run at 32. A wide window runs off the end of a small heap node into the next one, and the words past the record are **neighbouring heap**, not zero. That is a nuisance and it is also an instrument: **Two runs of the same turn, on two different guests, both of which wrote byte-identical autosaves (`bb4fd9ac` / `978041ac`). Any element word that differs between them cannot be a word the applier reads.** | list | stable prefix, L4 (VM145) vs CB (VM146) | record size from the writer | agree? | |---|---:|---:|---| | 3 build | 4 words | 4 | yes | | 5 system rates | 8 words | 8 | yes | | 8 fleet move | 1 + vector | 1 + 3 | yes | | 10 | 2 + vector | 2 + 3 | yes | | 14 fleet task | 2 words + **one byte** | 2 + bool | yes | | 23 population | 1 + vptr + vector | 1 + 1 + 3 | yes | Two consequences worth having: **List 14's third field is one byte, and the other three bytes of its word are garbage.** The two runs read `0x00B5B601` and `0x01138601` — same low byte, different upper three, which look like the low bytes of a heap pointer. A replayer that compares the whole word will find a difference that is not there. **Correction to `ai-order-capture.md` §0: list 23's element has no trailing int.** Lane L4 recorded `Population{vptr, vector(24 B), −1}` on turn 2→3 and `{…, 1}` on turn 1→2, and read those as a field. This lane's turn-2 run reads **`0x0035765f`** in that slot — a value that also appears mid-element in list 8 in L4's own log, i.e. an allocator cookie. Two runs, byte-identical outputs, different value: it is not a field. The `+1`/`−1` L4 saw was heap noise that happened to look like a signed flag, which is exactly why a wide window is worth having and exactly why the campaign types elements offline instead of in the instrument. --- ## 3. The seeds Lane L1 found each AI client's generator is seeded with a fresh per-process 32-bit word. This lane records them next to the block, because **the seeds are part of the stream**: with them, the capture is a reproducible pair; without them it is the AI's answer with no record of its input. `Game::StrategyApp::RunAI` is entered **three times per process on a save load**, once per AI client, and the fourth stack argument is the seed. On the canonical run: | call | net id | seed | |---:|---:|---| | 1 | 32 | `0x9dfad635` | | 2 | 496 | `0xb482b2d8` | | 3 | 512 | `0xfb81e993` | The net ids are the same four-value set the block slots carry (16 is the human and gets no client), so a seed binds to a block by net id with no ambiguity. ### 3.1 One instrument defect, caught by its own coverage line The first attempt reported `MH_ERROR_ALREADY_CREATED`: the trace framework installs **its own** template hook on `RunAI` (that is L1's probe), and MinHook refuses a second detour on one target. The module printed ``` COVERAGE: aiseed hook NOT INSTALLED -- the capture will carry NO seeds, and an absent seed list means the instrument failed, not that the clients were unseeded ``` and that line is the only reason the run was not published with a silently empty seed list. It is method rule 1 doing exactly the job it was written for. The fix is a config line (`hook.Game::StrategyApp::RunAI=off`), and the detour count is unchanged: one instrument on that function, not two. --- ## 4. The creation turn, and the control that had to be built `turn1-state.sav` + one End Turn is the turn three AI players each **choose** a research target. Three lanes have run it and got three different autosaves; the campaign's oracle does not extend to it. This lane ran it three times. ### 4.1 Run C3, unpinned — and a fifth value for the tie set The block reproduces lane L4's turn-1 block exactly: player 32 emits lists `{1:1, 3:1, 5:1, 23:1}` with `{ordinal 1, designId 18, systemId 288, 0}`, and the three AI players set research-target gates. **The design's name is read directly for the first time** — `"Honor Lance"`, short-string optimised inside the list-1 element, recovered by the string scanner at word 2. The targets are **144 / 90 / 282** for players 32 / 496 / 512. The first two are lane L4's values. **282 is new.** L4 measured 288 for player 512; lane L5 named four outcomes by tech name. So the outcome set for that one decision has at least one more member than anyone has enumerated, and Rung C's "k = 6, all six nameable" claim is **not yet closed** — this run did not land inside the named set as an id, and nobody has mapped 282 to a name. That is the honest status. The autosave is `d59bb9f2fd0eb535…` (66,746 B) — which is exactly the file lane L5's own `hooks=off` run produced. Two lanes, two processes, same outcome, by chance. ### 4.2 Runs C5a and C5b, pinned — the creation turn IS reproducible C5a and C5b are C3 re-run with `aiseed=pin` set to the three seeds **C3 itself observed**. Each is a separate process; each drew its own three natural seeds, all different from C3's and from each other's, and each had them overwritten before the client constructor saw them. | | C3 (unpinned) | C5a (pinned) | C5b (pinned) | |---|---|---|---| | natural seeds drawn | `e70a4703` / `0c63ca36` / `372be4df` | `227235b9` / `762d125e` / `9d77d4e6` | `be649c5c` / `2200dbd2` / `d418f3f8` | | seeds **used** | its own | **C3's** | **C3's** | | research targets 32 / 496 / 512 | 144 / 90 / **282** | 144 / 90 / **282** | 144 / 90 / **282** | | all six gates, all four blocks | — | **identical** | **identical** | | element records (lists 1, 3, 5, 23) | — | **identical** | **identical** | | `(Autosave).sav` | `d59bb9f2fd0eb535…` | **`d59bb9f2fd0eb535…`** | **`d59bb9f2fd0eb535…`** | **The workload that three lanes could not reproduce is reproducible once the seeds are pinned — three processes, nine distinct natural seeds, one autosave.** Lane L1's mechanism is confirmed from the other end: the seed is *the* per-process input, and `game/ai` is a deterministic function of (save, per-client seed). Rung B's input is a **closed** record — save + seeds + block — not a log of something that happened once. Two runs agreeing on a `k > 1` workload is a `1/k` coincidence; three is `1/k²`, about 3 % at lane L4's k = 6. That is the reason the third run was worth the eight minutes. The single difference anywhere in the decoded stream is **word 3 of list 23's unnamed `Population` body** — `876301736`, `245`, `61179324` across the three. By the argument of §2 (runs whose autosaves are byte-identical cannot differ in a word the applier reads) that word is noise, now on three data points rather than two. Localising it is exactly what a control on a `k > 1` workload is *for*, and it did not exist before this lane could pin. --- ## 5. Rule 19, and the control that does not exist without pinning **On the canonical pair the control is real.** `hooks=off` on VM146 reproduced both oracle hashes (run C0), and the instrumented run reproduced them again (run C1). That pair is a valid rule-19 control because the workload is deterministic: there is one right answer and both runs gave it. **On `turn1-state` there is no such thing as a single-run control, and this lane will not pretend otherwise.** Three lanes have run `hooks=off` on that save and got three different files. With an outcome set of size k, an instrumented run agreeing with one un-instrumented run is a **1/k coincidence**; at lane L4's k = 6 that is a 17 % chance of a clean bill of health that means nothing. A control on that workload requires the seeds *pinned*, which removes the only per-process input. That is what the `aiseed=pin` mode exists for. --- ## 6. The capture format, for lane RB `verify/results/turncommands/.json`, written by `tools/turncommands_capture.py` from the run's `shim.aiorders.txt` plus the run's own saves. Version 1. ``` capture_version, lane, guest, build, captured_utc instrument { words, deep, seed, pins, detours } -- the run is self-describing workload, note binding { input { file, bytes, sha256 } outputs [ { file, bytes, sha256, path, oracle_sha256, matches_oracle } ] control_run } ai_seeds [ { call, netId, observed, used, pinned } ] list_size_self_check { mismatches, note } batches [ { seq, n, blocks_va, stride, blocks: [ { index, playerId, gates { researchRate, researchTarget, researchBoost, group4, group5, civilianRatios }, list_sizes [27], lists [ { list, member_offset, size, elements: [ { index, node_va, raw_words [...], <- GROUND TRUTH vectors [ { at_word, first_va, capacity, count, words, truncated } ], strings [ { at_word, sso, len, text } ], decoded { record, , wire [...], record_words } } ] } ] } ] } ] ``` Rules for consuming it, in order of how much they cost to get wrong: 1. **`raw_words` is the ground truth. `decoded` is a typing.** If they ever disagree, `raw_words` wins, and the fix belongs in the tool, not in the replayer. 2. **Use `decoded.wire`** — it is the values in the order the writer emits them, with the counted vectors already spliced in as `count` followed by elements. It is what a replayer feeds the applier. 3. **`record_words` says how much of `raw_words` is the element.** Everything past it is the next heap node. The window is deliberately wider than any record. 4. **A `vectors` entry at a word offset no record names is noise** and must be ignored. The follower is a shape test on a wide window and it fires on coincidences; every genuine payload in this capture sits at the word index its record predicts (list 8 `w1`, list 10 `w2`, list 23 `w2`), and the decoder reads only those. 5. **`binding.outputs` is what makes the capture checkable.** A capture with no outputs, or whose hashes are not the ones the run produced, cannot be used for a byte-match and the tool writes a `WARNING` into the file saying so. 6. **`instrument.deep` must be `1`.** With `deep=0` the heap payloads are *absent*, not empty, and the tool writes `WARNING_INCOMPLETE`. A fleet move with no route is not a fleet move with a route of length zero. 7. **List 3's element is stored backwards in memory** relative to its wire order. The tool handles it; anyone reading `raw_words` directly must. --- ## 7. Coverage, stated as loudly as the result (rule 15) * **Six of twenty-seven lists** have ever carried an element, in this capture or any other: 1, 3, 5, 8, 10, 14, 23 (seven, counting list 1 on the turn-1 workload). Lists 2, 4, 6, 7, 9, 11–22, 24–27 are empty in every block of every run. **A replayer's handling of those twenty is untested by this lane and must be labelled as such** — including list 7, colonize, which lane Q typed from a *human* orders save and which no AI capture has ever exercised. * **One of six gates** is set on the canonical pair (research rate). Two on the turn-1 workload (rate and target). Gates 4, 5 and 6 have never been observed set by anything. * **One AI player with anything to do**, on a 28-star board with no contact: no colonise, no invade, no raid, no diplomacy, no combat, no trade route, no spy program. * **The `Population` body and list 10's payload are carried, not named.** This lane reports their bytes and declines to say what they mean. * The element-size measurement rests on **two runs**. Two is enough to show a word is *unstable*; it is not enough to show a word is *stable*, and the stable prefixes above are therefore upper bounds on noise, not proofs of a record. They agree with the writer's own instruction stream, which is the second, independent leg. --- ### 7.1 Two planned runs that were not made, and why * **C2, the `ModCount` trap set** (`watch=on` alongside the block dump). Planned as a second, independent check on the stream — predicted 10 bumps on the canonical turn, attributed per EIP. **Not run.** Lane RB got there first from the other side: replaying this capture puts `/Sim/ModCount` on **24 exactly with zero residual** on the canonical pair, which is a stronger statement than a trap multiset would have been (it checks the *arithmetic over the whole stream*, not the *sites*). The trap run would still add one thing nothing else has — per-EIP attribution, which would say *which handler* pays, not just that the total is right — and it remains a cheap hour for whoever wants it. `shim.cfg.cbtrap` is built and staged. * **C4, a second unpinned turn-1 run.** Planned to give P5's "two streams differing in exactly one word". **Superseded.** The pinned pair answers the same question and answers it better: an unpinned pair shows that two runs *can* differ, while the pinned triple shows *what makes them differ* and removes it. P5 as written is unresolved and should stay on the board that way — this lane has one unpinned turn-1 stream, not two. ## 8. The guest, left as found (and one thing that was not preserved) VM146 stopped, `binkw32.dll` restored to lane L5's build (md5 `F8DFC4AE…`, verified equal to `C:\SOTS\shimdist-l5\binkw32.dll`), `shim.cfg` back to `hooks=off`, `SavedGames` at eight files so the Load-dialog row geometry is unchanged. **Not preserved:** the two autosaves that were in `SavedGames` when this lane took the guest. `cbdeploy.ps1` deletes them before every run — deliberately, so that a run which writes none is distinguishable from one that rewrites them, and so the dialog lists a constant six files. The autosaves now present are run **C5b's** outputs (`a3f9dc4b…` / `d59bb9f2…`), both of which are recorded byte-for-byte in `verify/results/saves/`. They are scratch files every lane's deploy script deletes anyway, and nothing in the campaign depends on them, but the deletion was not announced in advance and is recorded here rather than left for someone to notice. Left in place for the next lane: `C:\SOTS\shimdist-cb` (build `cb-final-20260908T2320Z` plus the four `shim.cfg.cb*` variants) and `C:\SOTS\cb\` (`cbdeploy.ps1`, `cbstep.ps1`, `cbgrab.ps1`, `cbrun.ps1`, and the five collected run directories). The WSL-side drivers are in `verify/harness/cb/`.