CB's VM146 capture goes through the JSON->tcb adapter and the replayer unchanged and reproduces this lane's canonical result from a separate run on a different guest: 108 -> 62, closed 46, regressed 0, ModCount 24. It carries the first AI seeds any capture has held (32/496/512), and its deep dump closes two gap rows: list 8's route is [272] and list 10's vector is [1728], a ship the input save already contains.
21 KiB
Replaying a recorded turn: what a command stream reproduces, and what it does not
Lane RB, 2026-09-08. Host work, no VM. Engine worktree wip/rb off main 4f25f1e; predictions
committed as sots-engine docs/RB-predictions.md at 1e6474b, before the module existed
(rule 2). Every number below was taken on a fresh build-host created with rm -rf (rule 24).
Consumes: ai-order-emission.md (AI4, the apply order and the cost table), ai-order-capture.md
(L4, the live dumps), turncommands-block.md (Q, the wire shape),
turn1-to-turn2-nondeterminism.md (L5). Produces the first end-to-end turn record the
campaign holds.
0. Lead: ModCount is reachable, and it is the only leaf a stream closes on the reference turn
/Sim/ModCount has been unreachable from a save for the whole campaign, because it counts a
thing a save does not contain. Our standalone wrote 14 where the original writes 24: the
two driver bumps and nothing else. Replaying the turn's recorded command blocks puts it on 24,
exactly, with no fitting and no fudge term.
Canonical pair turn2-state.sav -> turn3-state.sav, fresh build:
| leaves diverging | closed | regressed | |
|---|---|---|---|
| do nothing | 108 | — | — |
| standalone, no stream | 63 | 45 | 0 |
| standalone + recorded stream | 62 | 46 | 0 |
The one leaf the stream closes is /Sim/ModCount, and its twenty-four decompose with zero
residual:
2 turn drivers (S00, T00) -- ours already
4 research-rate gates, one per submitting block (16, 32, 496, 512)
1 list 5 system rates
1 list 3 build order
1 list 10 (unnamed)
2 list 14 fleet task, modes 0 and 1 against ONE fleet
1 list 8 fleet move
0 list 23 population -- the 17..27 half is free
--
12 delta, on a save that carried 12
Nothing else closes, and that is the honest result rather than a disappointment. Of the ten commands, three have a handler this engine holds and all three are no-ops on this workload — the AI re-issues the research rate and the planetary-budget sliders the save already carries. The other seven need subsystems we do not have (§4).
The secondary pair is where the stream does work, and it produces the campaign's first matched triple — an input save, the stream captured from the run that consumed it, and that run's own autosave:
| oracle | closed by the stream | still diverging |
|---|---|---|
the recording's own autosave (l4-turn1to2-instrumented-autosave.sav) |
7 | — |
the historical turn2-state.sav, from a different process |
6 | player 512's ResTNm |
Both at regressed 0. The seven are ModCount (0 → 12) plus ResRate and ResTNm on players
32, 496 and 512. Against the historical oracle only six close, and the one that does not is
exactly the leaf lanes L4 and L5 showed is decided per-process: the recording picked
XNC_TrnsMorr2, turn2-state.sav holds BIO_GnMod. That was written down as a prediction before
the run and it is the strongest thing in this lane — a replay that had closed it would have meant
something was copying the oracle instead of the capture.
1. The on-disk form: .tcb
Line-oriented, whitespace-separated, # comments, magic first. A shim can emit it with fprintf
and a lane can read it without a tool.
tcb 1
meta source l4-turn2to3-aiorders.txt
seed <netId> <32-bit word> -- one per AI client; absent means "not recorded"
block <idx> <playerId> -- the batch SLOT, ascending; playerId 0 = a slot no client wrote
gate <idx> rate|target|boost|group4|group5|civilian <payload...>
list <idx> <listNo> <count> -- required for every non-empty list
elem <idx> <listNo> <elemIdx> <field>... -- fields in WIRE order
A field is one token: iN int, fN float, b0/b1 bool, s:TEXT string, vN a counted vector
whose length is known and whose values are not, vN:a,b,c one that is fully read, and ? a scalar
the instrument could not reach.
Three parts of that are not decoration.
? and bare vN are the point. A dump reads a fixed window of each element and cannot follow
a pointer, so a route arrives as "one element, contents unknown". That is a different fact from
"no element" and from "an element of zeros": the command still costs its bump and its effect still
cannot be applied. Recording the ignorance is what lets the counter be right while the state is
honestly left alone. The reference turn has four such elements out of ten.
A declared count that disagrees with the elements present is an ERROR, not a warning. A capture that lost an element would otherwise produce a counter quietly one short, and nothing downstream could tell that from a turn that really issued one fewer command. Thirteen malformed captures are in the engine's test as rejection cases.
seed is carried even though nothing consumes it. Lane L1 established that each AI client
seeds one MT19937 with a single word at construction, so a decision is a function of (board,
seed) and a capture that records the seeds can be re-derived rather than replayed. The field
exists now so a capture taken today is still the right file when game/ai lands. A reference save
whose seeds were never logged — turn2-state.sav — is not reproducible by any process, the
original included, and its commands can only ever be replayed.
1.1 Two capture tools, one reader — and a defect one of them shares with my first attempt
Lane CB was building tools/turncommands_capture.py in parallel, emitting JSON. That is the
better capture of record and it should stay: raw element words as ground truth, heap vectors and
strings the deep dump followed, the AI seeds, the container self-check, and — the part neither of
my files had — the input save's hash bound to the output autosaves' hashes, so a capture cannot
be silently used against the wrong save.
Rather than a second format in the engine, tools/tcb_from_json.py joins them: CB's JSON stays the
capture of record, .tcb stays the engine's parser-free input, and one narrow script knows both.
Fed CB's tool's own output over the L4 log, the adapter produces a replay byte-identical to the
one from aiorders_to_tcb.py. Honest bound on that agreement: both decoders read the same log,
so it checks the two decoders against each other and not the log.
And it caught a defect in CB's decoder, which is the same one I shipped and measured (§3): its
list-5 record maps the element's memory words straight onto the frame's wire order
({systemId, ship, terraform, sciences, …}). That is wrong by at least one position. The adapter
overrides it to ? and says why in its own docstring; CB should drop the list-5 record from
RECORDS rather than rely on the adapter to mask it, because the JSON is the capture of record
and a wrong typing in it will outlive this note. CB's raw_words are unaffected and remain right.
1.2 CB's real capture, run
CB's own VM146 capture landed while this lane was closing and it goes straight through the adapter
and the replayer unchanged: verify/results/turncommands/cb-turn2to3.tcb. It reproduces this
lane's canonical result exactly — 108 → 62, closed 46, regressed 0, ModCount 24 — from a
completely separate run of the game on a different guest. It also brings two things L4's dump
could not:
-
The AI seeds, observed live: netId 32 →
0x9dfad635, 496 →0xb482b2d8, 512 →0xfb81e993, allpinned=false(observed, not forced). These are the first seeds any capture has carried, and they are what makes this a turn record a futuregame/aican be re-derived against rather than replayed from. -
The heap payloads, which close two rows of the gap list below outright:
- list 8's route is
[272]— one hop to system 272. Lane L4 §5.3 recorded this as "turn3-state.sav's waypoint says 272 but the capture does not prove it". It is proved now. - list 10's counted vector is
[1728]— a ship id, and one that exists in the input save.{systemId 288, fleetId 34, ships [1728]}is now a typed reading with values, and "assign these ships to this fleet at this system" fits all three. Still not a proof of the command's name — the effect has never been observed — but the payload's type no longer leaves it open.
Those two elements move from
incompletetodeclinedin the replay report: the capture is now complete and the subsystems are what is missing. - list 8's route is
CB's capture also does the one thing neither of my files did and every future capture must:
it records the input save's hash bound to the output autosaves' hashes, both of which match
the published determinism oracle, with a hooks=off control run named. That binding is what makes
a capture usable as a byte-match oracle at all.
The converter. tools/aiorders_to_tcb.py turns lane L4's shim dump into a .tcb mechanically:
it applies the per-list field mapping, undoes the one list whose writer runs backwards, reinterprets
the words the record says are floats, turns a vector's begin/end pair into a length, and writes ?
where the window could not reach. Both existing captures are converted and checked in at
verify/results/turncommands/l4-turn{1to2,2to3}.tcb. Lane CB does not need to write an emitter:
the existing aiorders=on hook plus this converter already produces the file. If CB does emit
.tcb directly, the seed and name records are the two things the current dump has no field for.
2. Apply order, and how it was verified
The batch is a flat run of thirty steps: twenty-seven per-list steps, each looping over every player's block before the next step begins, with three per-player gate loops spliced in.
lists 6 11 20 19 17 18 5 23 24
gate loop A { group5 (free), research target (bump), research rate (bump) } 0x0088fdb0
lists 1 4 3 21 2 22 9 10 12 13 14 15 16 7 8 25 27 26
gate loop B { research boost (bump) } 0x008907b1
gate loop C { group 4 (bump) } 0x0089080a
So one player's list-6 commands are applied before another player's list-11 commands, the list sequence starts at 6 and is not 1..27, and the member offsets it walks are not ascending either.
How it was verified, and what the verification cannot show.
- Address monotonicity, nine of thirty positions. Six appliers are inlined into the batch and
each writes
ModCountin place, so a watchpoint run recovered their addresses; the three gate loop heads are known too. In schedule order those nine are0x0088fdb0,0x0088fe0a,0x008902fe,0x008903b9,0x0089046c,0x008905c8,0x008907b1,0x008907bc,0x0089080a— strictly increasing, at steps 9, 9, 18, 19, 20, 23, 28, 28, 29. That chain is an independent re-derivation. The other twenty-one lists' handlers are out of line and this lane has no record of their call-site addresses inside the batch, so their relative order is inherited from AI4's read of theadd edi, immchain, not re-derived. The test says nine. - Bijection and non-sortedness, asserted: every list exactly once, every located gate exactly
once, and neither the list sequence nor the offset sequence ascending. That is what rules out the
two obvious wrong implementations —
for (list = 1..27)and walking the block in memory order. - The civilian-ratios gate is absent from the schedule on purpose. It has no located applier anywhere in the routine, so including it would be claiming a cost of zero for something whose cost is unknown. It is caught separately and makes the whole count report itself as a lower bound.
Stated plainly: apply order is unfalsifiable on every workload the campaign holds. Both captures
put every non-empty list on one player and every command on one system, so any permutation produces
the same save and the same count. The order is implemented for the workload that will need it, and
tested against the instruction stream rather than against an outcome. The workload that would make
it falsifiable is two players commanding the same object in one turn — the cheapest is a
two-human /concurrent game (lane G2's Tier 0), where both clients order fleets at one system.
3. A falsified prediction, and the finding it paid for
RB-P3 predicted regressed 0. The first run regressed two leaves, and the cause is a real fact
about the original.
The converter mapped list 5's dumped words straight onto the wire order of the rates frame
(SRs, SRt, SRsc, SRtf, SRi, SRoh, SRnr). The replay then wrote the AI's single non-zero slider
into SRt, and Sys[288 "Ke'Dolarra"] came out with SRt 1.0 / SRsc 0.0 against an oracle
holding SRt 0.0 / SRsc 1.0.
The memory field order of Game::StarSystem::OutputRates is not its wire order. What is known
precisely:
- the only non-zero word in every dumped element of list 5, on both turns, is at memory index 2 of the element (index 0 is the system id, so it is the frame's second member);
- the same command on the wire, in a save that carries issued orders, puts its only non-zero in
SRsc, the frame's third member (lane Q,SAVE_FORMAT.md§11, cross-checked to the Planetary Budget slider pushed fully to Construction); turn2-state.savandturn3-state.savboth holdSys[288] Rts = {SRs 0, SRt 0, SRsc 1.0, SRtf 0, SRi 0, SRoh 0, SRnr 0}— so the AI's command re-issues the state the save already holds, which is why a correct applier is invisible and an incorrect one is immediately visible.
So memory member 1 is wire member SRsc: one correspondence pinned, six unread. One non-zero
slider cannot determine a permutation of seven, and the converter no longer pretends otherwise — it
emits the system id and seven ?, and the replayer counts the command and declines it.
Two experiments settle it, both cheap:
- one UI run — push two different sliders to two different values on one system, End Turn
with
aiorders=on, and read the permutation straight off the element; - cheaper, and no VM at all — a save taken after issuing rates carries the same command on the
wire, where every field is NAMED.
zuul-turn17-orders2.savhas one. A.tcbconverted from a save's ownTurnCommands_v5block needs no memory mapping, and would also give the.tcbformat a second, independent producer. This is the highest-value next step on the capture side and it is pure host work.
Two notes on how this was caught, because they generalise. The plain closed/regressed measurement
found it, before the control did — but the control (RB-P4, --replay-count-only vs a full replay
must be byte-identical when every modelled handler is a no-op) would have found it too, and it is
what now stands guard: on the canonical pair those two runs are byte-identical, which is the
evidence that the handlers we do run write where they claim to rather than agreeing with the
oracle by luck. And the reason a wrong write was visible at all is that the command re-issues
existing state: a command that re-states the board is the best possible test of an applier, and
the corpus is full of them.
4. What a complete replay needs that this lane does not have
This is the real Rung B backlog. Each row is the reason one command in the reference turn is counted but not applied.
| # | needed for | what is missing | shape |
|---|---|---|---|
| 1 | list 3 build, list 1 design | ship construction. No phase in this engine builds a ship. The command carries a queue ordinal, a design id and a system; the effect is srb/sri, Maint, the savings debit, a hull id and ShipIDs. Nine leaves on the canonical pair. |
engine work; lane B6/E2 have the map |
| 2 | list 1, list 3, list 8, list 10, list 14 | the client's id allocator. The build order names design 18 and the fleet order names fleet 34 — objects that do not exist in the input save, allocated client-side before submission, while the server's own master counter (NMnx) issued 1712 and 1776 the same turn. Two id spaces, and the small one is part of the wire protocol. A reimplementation that allocates on apply produces a structurally correct save with every AI-created id wrong. |
a watchpoint, not a week of reading: break on the write that produces 18 and 34 |
| 3 | CLOSED by lane CB's deep capture (§1.2): the route is [272], one hop. What remains is row 2 — the route belongs to a fleet the input save does not contain. |
done | |
| 4 | list 5 system rates | the memory field order of the rates frame (§3). One correspondence pinned, six unread. | one UI run, or one save-sourced capture |
| 5 | list 10 | a name and a meaning. The payload is fully read now (§1.2): {systemId 288, fleetId 34, ships [1728]}, and 1728 is a ship the input save already contains. "Assign these ships to this fleet at this system" fits every word. The effect has still never been observed, so it is still not named. |
a hook on 0x0088bed0 |
| 6 | list 14 fleet task | what the two modes do. Two elements per AI fleet order, modes 0 then 1; the interface deposits one, mode 0 only. The cost is settled; the effect is not read at all. | a hook on the inlined applier |
| 7 | list 23 population | the Population body, 24 bytes behind a vftable that the capture window does not follow. Free in ModCount, so it has never been forced. |
one dumper change |
| 8 | the research-target gate | the techId → tech-name map. The wire carries an integer (144, 90, 288); the save carries a name. The client resolves it off the command and passes a char*. The ids are not index * 16 and are not indices into anything we hold. Today the capture can carry the observed name and the replay TRANSCRIBES it — reported in its own column, because that is not a reimplementation. Three data points now exist: 144 → IND_Waldo, 90 → DRV_PlsFiss, 288 → XNC_TrnsMorr2. |
a watchpoint at the gate payload, or a read of 0x006c8580 |
| 9 | the research-boost gate | savings spent to advance research; both halves unmodelled. Never observed set on an AI turn. | engine work |
| 10 | the group-4 and group-5 gates | no read semantics; group 5 is Hiver-only and no save carries it. | rule 6 — a manufactured Hiver workload |
| 11 | the civilian-ratios gate | no applier located anywhere in the batch. Its ModCount cost is unknown, not zero, and any turn that sets it reports a lower bound. |
an image-wide search for the consumer |
| 12 | the load-time batch | the process applies a batch at LOAD as well as at End Turn (seq=1, n=1, the local client's block alone, with the rate gate SET). Whether it charges ModCount is untested, and this lane excludes it by construction. If it does charge, a save loaded and immediately re-saved reads one higher. |
one run: load, save, compare |
| 13 | the whole thing, on any interesting board | two turns, one AI empire, 28 stars, no contact. Lists 2, 4, 6, 9, 11–13, 15–22 and 24–27 have never been non-empty. Apply order is unfalsifiable (§2). Nothing here generalises past a very quiet game. | manufacture the workload |
Two things NOT on that list, deliberately. Summary.Checksum moves whenever anything else does and
its inputs are unread — no command replay will touch it until the rest is right. And nothing in
either capture draws a generator word, so the RNG frame is untouched by the stream; the coordinator's
note that the replay interval begins at BeginProcessTurn rather than ProcessTurn is already
satisfied here — the batch is drained before the first phase runs, which is where the End-Turn
dispatcher calls it (ApplyAllTurnCommands at +0x00784904, ProcessTurn at +0x0078491c,
OnAllCombatDone_Tail at +0x00784d07) — but the hive-registration draws lane L1 found inside
BeginProcessTurn are a separate interval and this lane does not model them.
5. What this lane did not do
- No VM time and no new instrument. Every live number here is lane L4's, re-read through a converter. Nothing new ran under a hook.
- The CT111 shim cross-build was not run — CT111 refused the key from this host.
src/game/aiis in the cross build, so per rule 13 the integrator must run it before pushing. The two 32-bit-specific hazards were audited by hand and one was real and is fixed:strtoulplus a> 0xffffffffULtest is tautological whereunsigned longis 32 bits, which would both accept out-of-range input silently and trip-Wextra/-Werror; both sites now usestrtoull.%zuappears in the new code and also insrc/app/main.cpptoday, so it is not a new exposure. - The apply order was not tested against an outcome and cannot be on this corpus (§2).
ToTurnCommandBlockreuses the existing, already-verified cost model rather than counting a second time — deliberately, so there is one implementation of the arithmetic and not two that can agree with each other while both being wrong (rule 8).- The
.tcbformat has exactly one producer. Until the save-sourced converter of §3 exists, every capture comes through one script and a bug in it is invisible. - Seeds are carried and never used. No capture in the corpus has any.