# The AI's command block, read out of the running game Lane L4, 2026-09-08. Guest **VM145** (`sots-re-win10-145`, 192.168.10.145), build `l4-45bf085-dirty-20260908T2113Z`. Predictions committed before the module was written: `sots-engine` `docs/L4-predictions.md` (commit `45bf085`, before `src/shim/hooks/ai_orders.cpp` existed). Closes the live half of `ai-order-emission.md` (AI4), `ai-stepping-and-passes.md` (AI3), `ai-task-system.md` (AI2) and `ai-turn-logic.md` (AI1). **Everything in those four documents was static reading. This is the first time anything in `src/game/ai` has run under an instrument.** Raw logs: `verify/results/shim/aiorders/l4-turn{1to2,2to3}-aiorders.txt`. --- ## 0. Lead: what the real AI emitted, and what our model emits Two workloads, one End Turn each, every submitted block dumped at `StrategySim::ApplyTurnCommandBatch` 0x0088f9b0 where all of them are complete in memory. **Turn 1 → 2** (`turn1-state.sav`). Player 32, the only AI with an empire: | list | n | element (wire order) | |---|---:|---| | **1** new design | 1 | a `ShipDesignDef` object carrying the name string **`"Honor Lance"`** inline | | **3** build | 1 | `{ordinal 1, designId 18, systemId 288, 0}` | | **5** system rates | 1 | `{systemId 288, OutputRates{0, 1.0f, 0, 0, 0, 0, 0}}` | | **23** population | 1 | `{systemId 288, Population{vptr, vector(24 B), 1}}` | plus gates `rate = 0.8` and `target = techId 144`. **All twenty-three other lists empty.** **Turn 2 → 3** (`ref-turn2.sav`). Same player: | list | n | element (wire order) | |---|---:|---| | **3** build | 1 | `{ordinal 2, designId 18, systemId 288, 0}` | | **5** system rates | 1 | `{systemId 288, OutputRates{0, 1.0f, 0, …}}` | | **8** fleet move | 1 | `{fleetId 34, route[1]}` | | **10** | 1 | `{systemId 288, fleetId 34, counted[1]}` | | **14** fleet task | **2** | `{34, 0, true}` and `{34, 1, true}` | | **23** population | 1 | `{systemId 288, Population{vptr, vector(24 B), −1}}` | plus gate `rate = 0.8`; **no research target** on any player this turn. **Against our model.** `sots-engine`'s `src/game/ai/orders.h` reproduces **both blocks exactly** — list for list, element for element, and both turns land on the measured `ModCount` delta of 12. That is now a test, `tests/game_ai/test_live_blocks.cpp`, 44 checks, built from the dumped values and kept deliberately separate from `test_orders.cpp` (which is the record of what static reading predicted, and must not be fitted to this). **The model was right about the arithmetic and incomplete about the content.** Three things it did not have: 1. **A list-23 element, on every turn.** No save in eleven has ever carried an element in the free half of the table (lists 17–27), so that whole row of the cost model was a hypothesis in the rule-6 sense. It is now exercised twice, and **the counter still lands on 12** — the free half is free, measured, from the first workload that ever populated it. This is the single most valuable thing the capture produced, and nobody predicted it. 2. **The ids in the commands are client-allocated.** The build order names **design 18** before the server has issued it, and the fleet order names **fleet 34**, an object that does not exist in the input save. Our model treated ids as opaque; a reimplementation has to allocate them where the original does or every id in the resulting save is wrong. 3. **Build, rates and population all name system 288** — the AI's home. One decision, three commands. --- ## 1. Predictions, then outcomes `docs/L4-predictions.md` §1–§3, in order. Four held, two were falsified, and both falsifiers are worth more than the predictions were. ### P1 — the block set. **FALSIFIED, usefully.** Predicted `n == 4`, the four submitting players. Measured **`n == 8`**: the batch is sized to `playerCount`, and the four `Species == 4` monster factions occupy slots 4–7 with **`playerId == 0`, every gate clear and all twenty-seven lists empty**. AI3 §1.2 already read the mechanism — `ResumePlaying` does `clear(&S->+0x174)` then `resize(&S->+0x174, playerCount)` — and I did not join it to AI4's "the monster factions submit no block at all". Both are right: the *vector* has eight slots, the *submissions* are four. The untouched slots are not merely empty, they are **uninitialised**: the rate-gate payload reads as garbage floats (`8.97e-44`, `2.62e+33` on the two runs) with the gate bit clear. Consequence for a reimplementation: iterate `playerCount` slots and let the clear gates do the filtering; do not build a list of "submitting players". The four ids that *are* set are `16, 32, 496, 512` in save-player order, so `block+0x04` is the save player id, not an index. That half of P1 held. ### P2 — the turn-1 block. **HELD, plus one list nobody predicted.** Predicted lists {1, 3, 5} on player 32 and nothing on 16/496/512. Measured **{1, 3, 5, 23}** on 32 and nothing on the other three. The sharp falsifier was list 8/10/14 being non-empty on the turn the AI creates its first fleet — **they were all empty**, so AI4's P2 attribution of the twelve stands. Four research-rate gates and **three** research-target gates (`144`, `90`, `288`), the human's target gate clear. Exactly AI4's P2 and P5. ### P3 — the element values. **HELD for the build order; the design id is not in the window.** `designId == 18` and `systemId == 288` in list 3, and list 5 names the same system: both as predicted. The list-1 element is a polymorphic object whose first 48 bytes are a vftable pointer, a word, and a `std::string` holding `"Honor Lance"` (`_Mysize` 11, `_Myres` 15 — short-string optimisation, so the name is inline). **The id is past the dump window**, so "the design command carries its own id" is not proved from list 1 directly — but it is proved from list 3, which names design 18 in the *same block*, before the server has issued anything. ### P4 — the probes on turn 1. **Held on the control, wrong on one row.** | probe | predicted | measured (1→2) | measured (2→3) | |---|---|---:|---:| | `RunTaskList` | 6 = 3 AI × 2 passes | **6** (3+3) | **6** (3+3) | | `BuildTurnCommands` | 4 | 5 | 5 | | `RequestBuildForTask` | >0, both passes | 10 (5+5) | 8 (4+4) | | `AssignFleetsAndIssueOrders` | entered, no elements | 0 | 2 (1+1) | | `IssueRouteForFleets` | entered, no elements | 0 | 2 (1+1) | | `AITRaid::Execute` | 0 | **0** | **0** | | `StrategyClient::OrderList16` | 0 | **0** | **0** | | `AITAdvanceIdleShips::Execute` | entered | 6 (3+3) | 6 (3+3) | | `IsClaimedByAnotherTask` | >0 | **0** | 10 (9+1) | `RunTaskList == 6` is the headline: **three AI agents, each stepped once, two passes each**, which is AI3's P1 confirmed live at the agent level and AI4's "three AI players, not one" confirmed from a second instrument. `IsClaimedByAnotherTask == 0` on turn 1 falsifies my "called often" — on a board with no fleets it is never reached at all. Workload-dependent, and my prediction did not say so. `BuildTurnCommands == 5` where four blocks are submitted: three are the AI clients (one after each agent's pass 1), one fires before any AI has run (game setup / load), and one more at the end. The pass and agent columns on those two are stale globals and cannot attribute them, so I am **not** claiming which is the human's. ### P5 — the turn-2 block. **HELD, plus list 23 again.** Predicted {3:1, 5:1, 8:1, 10:1, 14:2}; measured exactly that, plus list 23. ### P6 — list 14 is two elements against one fleet, keyed on mode. **HELD, at the values.** `{34, 0, true}` and `{34, 1, true}`: same fleet, modes 0 then 1, and it is the same fleet id list 8's route names. AI2's P1 — inferred from a call site, then supported by two `ModCount` bumps — is now read off the element values. The interface's single element (`human-turn2-orders.sav`, `{1456, 0, true}`) is the same record with mode 0 only. ### P7 — the fleet id. **FALSIFIED, and this is the important one.** Predicted `F == 1744`, the fleet that exists at submit time, with the new fleet 34 assigned by the server on apply. Measured **`F == 34`**. Fleet 34 does not exist in `ref-turn2.sav`. It exists in `turn3-state.sav`, as `"Beta Fleet"`. So **the client allocates the object and its id before it submits, and ships the id in the command.** Design 18 is the same story from the other turn: the *other* new design that turn — a monster faction's, created server-side with no command block — took **1712** from the save's master id counter (`NMnx` 106 → 109), while the AI's took **18**. There are therefore **two id spaces**, and the small one is **client-allocated and part of the wire protocol**. For Rung B this is a hard constraint: a reimplementation that assigns ids on apply produces a structurally correct save with every AI-created id wrong. I do not know the client counter's rule. `18` and `34` differ by 16, which is the master counter's stride, so it looks like the same `id = index * 16` scheme running off a different, small base (index 1 and index 2 plus 2). **Unread**; it is the first thing the next lane should chase, and it is a watchpoint, not a week of reading. ### P8 — list 10's first word. **HALF-FALSIFIED, and the name is now supportable.** Predicted the first `i32` is the fleet. Measured `{systemId 288, fleetId 34, counted vector of 1}` — system first, fleet second. Lane Q's record `{i32, i32, counted i32}` is right; the reading is "at system 288, fleet 34, [one object]". The counted element's value is in the heap vector and the dump does not follow it (§5.3), so the payload is still not named. AI4 §4.5 declined to name list 10 on adjacency alone and was right to; it now has values, and "assign these ships to this fleet at this system" fits all three words, with the tail unread. ### P9 — AITRaid. **NOT SETTLED, and the probe says exactly why.** `StrategyClient::OrderList16` 0x007635f0 was entered **zero** times on both turns. That is a non-answer about pass 0 — and the companion probe says which non-answer: **`AITRaid::Execute` was also entered zero times, on both turns.** The task never ran. AI3 §2.4 stays open, and it stays open for a *stated* reason instead of an assumed one, which is the whole point of rule 20. The workload that would settle it needs `AITRaid` in a task list. Neither of the corpus's reachable turns has one, and the two boards differ only in whether the AI owns a fleet — so owning a fleet is not the trigger. ### Pass 0 writes nothing — **confirmed by element count, which is stronger than the entry count.** The three pass-1-gated emission exits were entered in **both** passes, in equal numbers: * turn 1: `RequestBuildForTask` 5 in pass 0 and 5 in pass 1 → **one** list-1 and **one** list-3 element in the block; * turn 2: the same, plus `AssignFleetsAndIssueOrders` and `IssueRouteForFleets` once per pass → **one** list-8, **one** list-10 and **two** list-14 elements. If pass 0 emitted, every count would double. AI3's P2 holds, measured from the output rather than inferred from the gate. The pass sweeps are otherwise symmetric: every probe's pass-0 count equals its pass-1 count, with one exception — `IsClaimedByAnotherTask` runs **9 times in pass 0 and once in pass 1**. The claim registry is already populated by the time the second sweep runs, so most candidates are filtered before the test is reached. That is consistent with AI3's two-tier quota model and is the only asymmetry in either run. --- ## 2. Which tasks actually fire (AI3 §5, live) Eight `Execute` bodies probed, covering the nine classes AI2 called planners plus `AITRaid` and `AITAdvanceIdleShips`. On both turns: | body | 1→2 | 2→3 | which agents | |---|---:|---:|---| | `AITAdvanceIdleShips::Execute` | 6 | 6 | all three, both passes | | `AITBuildDeepScanShips::Execute` | 4 | 4 | **only 496 and 512**, both passes | | `AITColonize`, `AITEscortGateInvade`, `AITInvade`, `AITNodeBore`, `AITBuildPoliceShips`, `AITRaid` | 0 | 0 | — | Two things follow, and the second is uncomfortable. **`AITAdvanceIdleShips` is in every agent's list and is entered on both passes**, exactly as its priority-0, pass-1-body shape predicts. It is a good control and it read non-zero on every run. **Player 32's task list contains none of the six named planner bodies.** Its build orders came from `RequestBuildForTask`, entered 5 times per pass on turn 1 and 4 times per pass on turn 2, under tasks whose `Execute` bodies were not in my probe set. So the answer to "which of the nine planner tasks fire on a real turn" is **one of them, `AITBuildDeepScanShips`, and only for the two AI players that own nothing** — and the one AI that actually plays is driven by tasks nobody has probed yet. The nine were the wrong nine to probe. AI3's §5 correction of AI2 stands on the call graph; this lane cannot add to it, and says so. The event ring gives the exact per-agent sequence for both turns (`aievent` lines in the raw logs). Turn 2→3, agent 0x335c1040 (player 32), pass 0, in order: ``` RunTaskList → Acquire → IsClaimed → IssueRouteForFleets → AssignFleetsAndIssueOrders → Acquire → IsClaimed ×2 → RequestBuild (×4 more of this pair) → AITAdvanceIdleShips ``` and pass 1 is the same walk with the claim tests gone. --- ## 3. Rule 19: the control, and what it cost to take it `ref-turn2.sav` + one End Turn, with **seventeen MinHook detours installed** (the batch dump plus sixteen entry probes): | file | measured | published oracle | |---|---|---| | `(Autosave EndTurn).sav` | 66,732 B `bb4fd9ac89f41e3b` | `bb4fd9ac89f41e3b` ✓ | | `(Autosave).sav` | 67,219 B `978041acd168b56e` | `978041acd168b56e` ✓ | **Byte-identical.** Two things at once: VM145, which is a ZFS clone nobody had checked, reproduces the reference guest exactly; and this lane's instrument is behaviour-neutral. Lane H's own entry probes were explicitly `probes=off` in every configuration here, because that set is the one measured to move an autosave by four bytes. ### 3.1 And a control that did **not** pass — the turn-1 workload is not reproducible `turn1-state.sav` + one End Turn does **not** reproduce `turn2-state.sav`, on this build, with or without instruments. Three runs, three different files: | run | `(Autosave).sav` | |---|---| | reference `turn2-state.sav` | 66,739 B `ab4ac2d7e2977260` | | `hooks=off` | 66,746 B `d59bb9f2fd0eb535` | | full instrument | 66,740 B `e43ec1d2b443c101` | Diffed field by field through the save reader, **exactly one field differs across all three**: ``` p512.ResTNm: ref = BIO_GnMod hooks=off = XNC_TrnsHum2 instrumented = XNC_TrnsMorr2 ``` `NMnx`, `ModCount`, every id list, every design, every fleet, every other player's target and rate: identical. And the command block shows it at the source — player 512's research-target gate carried techId **288** on the instrumented run, a different id on the others. So: **the research-target choice of an AI player that owns nothing is not reproducible run to run.** Players 16, 32 and 496 are stable across all three; only 512 moves. Three samples, one field. This is a correction to an assumption, not to a published claim: `determinism-oracle.md` verifies the oracle for the **`ref-turn2` workload** and says so; nobody had run the turn-1 workload three times. It is now on the record that the turn-1 workload is **not** an oracle, and any lane comparing bytes on it is comparing against noise. I did not find the mechanism. It is not the instrument (the `hooks=off` run diverges from the reference too, and by a *different* value). The shape — one player, one choice, stable everywhere else — fits a tie-break that reads a heap address or an uninitialised word, and that is a watchpoint on the selection site, not a reading exercise. --- ## 4. Corrections to earlier findings * **`ai-order-emission.md` §2 — "the four `Species == 4` factions submit no command block at all".** Right about the submissions, wrong about the block array: **the batch is `n = playerCount = 8`** and those four occupy slots with `playerId == 0` and uninitialised gate payloads. AI3 §1.2 had already read the `resize(playerCount)` that forces it. * **`ai-order-emission.md` §1 / §3 P1 — the 17..27 half of the cost table was read from the instruction stream only, with no workload.** It now has one: **list 23 carries one element on every AI turn measured**, and both turns still cost exactly 12. The free half is confirmed free. * **`ai-order-emission.md` §4.5 — list 10 "named only by position".** It now has values: `{systemId, fleetId, counted vector}`, with the system leading. Still not named; the vector's contents are unread. * **`turncommands-block.md` §3, list 3.** The wire record `{ordinal, designId, systemId, w}` is correct — but **the in-memory element is in the opposite order**, because that list's writer (`0x00822870`) emits `+0x14, +0x10, +0x0c, +0x08`, descending. It is the only one of the five writers checked that reverses; lists 5, 8, 10, 14 and 23 all write ascending. Anyone reading these elements out of memory needs that per-list, not as a rule. * **`ai-task-system.md` / `ai-stepping-and-passes.md` — "the AI's fleet order names the fleet it moves".** It names a fleet the input save does not contain (§1, P7). --- ## 5. What this lane did **not** do 1. **AI3 §2.4 is still open.** `AITRaid` never ran on either workload, so the list-16 pass-0 question is untouched — but now for a measured reason rather than an assumed one (§1, P9). 2. **AI3's P3 is untouched.** An entry counter cannot see the steal branch inside `IsClaimedByAnotherTask`; this lane measured only that it is called (10 times on turn 2, 0 on turn 1). Stated in advance in `docs/L4-predictions.md` §4. 3. **The dump reads 48 bytes per element and does not follow pointers.** Three payloads are therefore unread: the route in list 8 (one hop, id unknown — `turn3-state.sav`'s waypoint says 272 but the capture does not prove it), the counted vector in list 10 (one element), and the `Population` body in list 23 (24 bytes behind a vftable). Each is one more indirection in the dumper. 4. **The design id is not visible in the list-1 element** (§1, P3). It is inferred from list 3 naming design 18 in the same block. 5. **The client id counter is not located** (§1, P7). This is the largest remaining hole and it is a watchpoint: break on the write that produces 18 and 34. 6. **Which tasks drive player 32 is not known** (§2). Five build-shaped tasks per pass reached `RequestBuildForTask` and none of them is one of the eight `Execute` bodies probed. The probe set was chosen from AI3 §5's list of nine and that list is not where this AI's decisions come from. 7. **Two workloads, one AI empire, 28 stars, no contact.** Every count here is a count on a very quiet board: no colonise order, no invade, no raid, no diplomacy, no combat. Lists 2, 4, 6, 7, 9, 11–22 and 24–27 were empty on both turns and remain unexercised (rule 6). The distinct-state count for this lane is **two turns and one real AI player** — that is thin, and it is the caveat that matters most for anything generalised from here. 8. **The `+0x6c` `CivilianRatios` gate and the `+0x3c` Hiver gate were clear on all eight blocks on both turns**, as expected; nothing new about either. --- ## 6. The instrument, for the next lane `sots-engine` `src/shim/hooks/ai_orders.{h,cpp}`, configured by three keys: * `aiorders=on|off` — the block dump: **one** register-transparent entry stub on `StrategySim::ApplyTurnCommandBatch`, which receives `(blocks, n)` as stack arguments with every submitted block complete at a fixed 0x1b4 stride. Prints six gates, twenty-seven list lengths and 48 bytes per element per block. * `aiprobes=off|all|N` — sixteen entry counters, lane H's asm-stub pattern with its own table so lane H's set is untouched. `N` installs the first N, so the set bisects in one build. * `aiorders.out=`. Two design points worth keeping: **Row 0 is `RunTaskList`, and it is both the control and the pass recorder.** Its stub reads the `pass` stack argument before tail-jumping, so every later probe hit is attributed to a pass. That is what turned "pass 0 emits nothing" from an inference into a measurement. The global is stale once `RunTaskList` returns and the report says so; the `run` column in the event ring is what makes the staleness readable. **Every list is measured twice** — walked, and read from `_Mysize` — and a disagreement prints `MISMATCH`. Nothing printed it on 8 blocks × 27 lists × 2 runs, which is the evidence that the container layout is right rather than that the block is empty (rule 1). And a run whose `RunTaskList` count is zero prints `CONTROL ZERO` and says every other row is unmeasured, not absent — which is exactly what the load-time batch (`seq=1`, `n=1`, the local client's block alone) does print. ### 6.1 Lab notes * **`move X Y` then `click X Y` in the same click-helper batch is reliable; a bare `click` is not.** Roughly half of bare clicks were delivered at the *previous* cursor position, which reads as "the click did nothing" and then as "the next click did the previous thing". Two runs were nearly lost to it before the pattern was clear. * **Reset `SavedGames` to a fixed two-file set before every run.** With only `ref-turn2.sav` and `turn1-state.sav` present the Load dialog rows are always y=262 and y=291 and the click path never has to be re-derived. `C:\SOTS\ui\l4deploy.ps1` does it. * Startup to main menu on VM145 was 80–95 s. Verify by screenshot; never sleep and click. * VM145 left restored: `SavedGames` back to the 9-file pre-L4 set (autosaves byte-identical to the oracle), `binkw32.dll` and `shim.cfg` back to the W3 build and `w3mod` config, game stopped. `C:\SOTS\shimdist-l4`, `C:\SOTS\ui\l4\` and `C:\SOTS\ui\l4{deploy,click,grab,release}.ps1` left in place — they are a working template for the next lane.