sots-re/findings/subsystems/ai-order-capture.md

23 KiB
Raw Blame History

The AI's command block, read out of the running game

Lane L4, 2026-09-08. Guest VM145 (sots-re-win10-145, 192.168.10.145), build l4-45bf085-dirty-20260908T2113Z. Predictions committed before the module was written: sots-engine docs/L4-predictions.md (commit 45bf085, before src/shim/hooks/ai_orders.cpp existed).

Closes the live half of ai-order-emission.md (AI4), ai-stepping-and-passes.md (AI3), ai-task-system.md (AI2) and ai-turn-logic.md (AI1). Everything in those four documents was static reading. This is the first time anything in src/game/ai has run under an instrument.

Raw logs: verify/results/shim/aiorders/l4-turn{1to2,2to3}-aiorders.txt.


0. Lead: what the real AI emitted, and what our model emits

Two workloads, one End Turn each, every submitted block dumped at StrategySim::ApplyTurnCommandBatch 0x0088f9b0 where all of them are complete in memory.

Turn 1 → 2 (turn1-state.sav). Player 32, the only AI with an empire:

list n element (wire order)
1 new design 1 a ShipDesignDef object carrying the name string "Honor Lance" inline
3 build 1 {ordinal 1, designId 18, systemId 288, 0}
5 system rates 1 {systemId 288, OutputRates{0, 1.0f, 0, 0, 0, 0, 0}}
23 population 1 {systemId 288, Population{vptr, vector(24 B), 1}}

plus gates rate = 0.8 and target = techId 144. All twenty-three other lists empty.

Turn 2 → 3 (ref-turn2.sav). Same player:

list n element (wire order)
3 build 1 {ordinal 2, designId 18, systemId 288, 0}
5 system rates 1 {systemId 288, OutputRates{0, 1.0f, 0, …}}
8 fleet move 1 {fleetId 34, route[1]}
10 1 {systemId 288, fleetId 34, counted[1]}
14 fleet task 2 {34, 0, true} and {34, 1, true}
23 population 1 {systemId 288, Population{vptr, vector(24 B), −1}}

plus gate rate = 0.8; no research target on any player this turn.

Against our model. sots-engine's src/game/ai/orders.h reproduces both blocks exactly — list for list, element for element, and both turns land on the measured ModCount delta of 12. That is now a test, tests/game_ai/test_live_blocks.cpp, 44 checks, built from the dumped values and kept deliberately separate from test_orders.cpp (which is the record of what static reading predicted, and must not be fitted to this).

The model was right about the arithmetic and incomplete about the content. Three things it did not have:

  1. A list-23 element, on every turn. No save in eleven has ever carried an element in the free half of the table (lists 17–27), so that whole row of the cost model was a hypothesis in the rule-6 sense. It is now exercised twice, and the counter still lands on 12 — the free half is free, measured, from the first workload that ever populated it. This is the single most valuable thing the capture produced, and nobody predicted it.
  2. The ids in the commands are client-allocated. The build order names design 18 before the server has issued it, and the fleet order names fleet 34, an object that does not exist in the input save. Our model treated ids as opaque; a reimplementation has to allocate them where the original does or every id in the resulting save is wrong.
  3. Build, rates and population all name system 288 — the AI's home. One decision, three commands.

1. Predictions, then outcomes

docs/L4-predictions.md §1–§3, in order. Four held, two were falsified, and both falsifiers are worth more than the predictions were.

P1 — the block set. FALSIFIED, usefully.

Predicted n == 4, the four submitting players. Measured n == 8: the batch is sized to playerCount, and the four Species == 4 monster factions occupy slots 4–7 with playerId == 0, every gate clear and all twenty-seven lists empty.

AI3 §1.2 already read the mechanism — ResumePlaying does clear(&S->+0x174) then resize(&S->+0x174, playerCount) — and I did not join it to AI4's "the monster factions submit no block at all". Both are right: the vector has eight slots, the submissions are four. The untouched slots are not merely empty, they are uninitialised: the rate-gate payload reads as garbage floats (8.97e-44, 2.62e+33 on the two runs) with the gate bit clear.

Consequence for a reimplementation: iterate playerCount slots and let the clear gates do the filtering; do not build a list of "submitting players".

The four ids that are set are 16, 32, 496, 512 in save-player order, so block+0x04 is the save player id, not an index. That half of P1 held.

P2 — the turn-1 block. HELD, plus one list nobody predicted.

Predicted lists {1, 3, 5} on player 32 and nothing on 16/496/512. Measured {1, 3, 5, 23} on 32 and nothing on the other three. The sharp falsifier was list 8/10/14 being non-empty on the turn the AI creates its first fleet — they were all empty, so AI4's P2 attribution of the twelve stands.

Four research-rate gates and three research-target gates (144, 90, 288), the human's target gate clear. Exactly AI4's P2 and P5.

P3 — the element values. HELD for the build order; the design id is not in the window.

designId == 18 and systemId == 288 in list 3, and list 5 names the same system: both as predicted. The list-1 element is a polymorphic object whose first 48 bytes are a vftable pointer, a word, and a std::string holding "Honor Lance" (_Mysize 11, _Myres 15 — short-string optimisation, so the name is inline). The id is past the dump window, so "the design command carries its own id" is not proved from list 1 directly — but it is proved from list 3, which names design 18 in the same block, before the server has issued anything.

P4 — the probes on turn 1. Held on the control, wrong on one row.

probe predicted measured (1→2) measured (2→3)
RunTaskList 6 = 3 AI × 2 passes 6 (3+3) 6 (3+3)
BuildTurnCommands 4 5 5
RequestBuildForTask >0, both passes 10 (5+5) 8 (4+4)
AssignFleetsAndIssueOrders entered, no elements 0 2 (1+1)
IssueRouteForFleets entered, no elements 0 2 (1+1)
AITRaid::Execute 0 0 0
StrategyClient::OrderList16 0 0 0
AITAdvanceIdleShips::Execute entered 6 (3+3) 6 (3+3)
IsClaimedByAnotherTask >0 0 10 (9+1)

RunTaskList == 6 is the headline: three AI agents, each stepped once, two passes each, which is AI3's P1 confirmed live at the agent level and AI4's "three AI players, not one" confirmed from a second instrument.

IsClaimedByAnotherTask == 0 on turn 1 falsifies my "called often" — on a board with no fleets it is never reached at all. Workload-dependent, and my prediction did not say so.

BuildTurnCommands == 5 where four blocks are submitted: three are the AI clients (one after each agent's pass 1), one fires before any AI has run (game setup / load), and one more at the end. The pass and agent columns on those two are stale globals and cannot attribute them, so I am not claiming which is the human's.

P5 — the turn-2 block. HELD, plus list 23 again.

Predicted {3:1, 5:1, 8:1, 10:1, 14:2}; measured exactly that, plus list 23.

P6 — list 14 is two elements against one fleet, keyed on mode. HELD, at the values.

{34, 0, true} and {34, 1, true}: same fleet, modes 0 then 1, and it is the same fleet id list 8's route names. AI2's P1 — inferred from a call site, then supported by two ModCount bumps — is now read off the element values. The interface's single element (human-turn2-orders.sav, {1456, 0, true}) is the same record with mode 0 only.

P7 — the fleet id. FALSIFIED, and this is the important one.

Predicted F == 1744, the fleet that exists at submit time, with the new fleet 34 assigned by the server on apply. Measured F == 34.

Fleet 34 does not exist in ref-turn2.sav. It exists in turn3-state.sav, as "Beta Fleet". So the client allocates the object and its id before it submits, and ships the id in the command. Design 18 is the same story from the other turn: the other new design that turn — a monster faction's, created server-side with no command block — took 1712 from the save's master id counter (NMnx 106 → 109), while the AI's took 18.

There are therefore two id spaces, and the small one is client-allocated and part of the wire protocol. For Rung B this is a hard constraint: a reimplementation that assigns ids on apply produces a structurally correct save with every AI-created id wrong.

I do not know the client counter's rule. 18 and 34 differ by 16, which is the master counter's stride, so it looks like the same id = index * 16 scheme running off a different, small base (index 1 and index 2 plus 2). Unread; it is the first thing the next lane should chase, and it is a watchpoint, not a week of reading.

P8 — list 10's first word. HALF-FALSIFIED, and the name is now supportable.

Predicted the first i32 is the fleet. Measured {systemId 288, fleetId 34, counted vector of 1} — system first, fleet second. Lane Q's record {i32, i32, counted i32} is right; the reading is "at system 288, fleet 34, [one object]". The counted element's value is in the heap vector and the dump does not follow it (§5.3), so the payload is still not named. AI4 §4.5 declined to name list 10 on adjacency alone and was right to; it now has values, and "assign these ships to this fleet at this system" fits all three words, with the tail unread.

P9 — AITRaid. NOT SETTLED, and the probe says exactly why.

StrategyClient::OrderList16 0x007635f0 was entered zero times on both turns. That is a non-answer about pass 0 — and the companion probe says which non-answer: AITRaid::Execute was also entered zero times, on both turns. The task never ran. AI3 §2.4 stays open, and it stays open for a stated reason instead of an assumed one, which is the whole point of rule 20.

The workload that would settle it needs AITRaid in a task list. Neither of the corpus's reachable turns has one, and the two boards differ only in whether the AI owns a fleet — so owning a fleet is not the trigger.

Pass 0 writes nothing — confirmed by element count, which is stronger than the entry count.

The three pass-1-gated emission exits were entered in both passes, in equal numbers:

  • turn 1: RequestBuildForTask 5 in pass 0 and 5 in pass 1 → one list-1 and one list-3 element in the block;
  • turn 2: the same, plus AssignFleetsAndIssueOrders and IssueRouteForFleets once per pass → one list-8, one list-10 and two list-14 elements.

If pass 0 emitted, every count would double. AI3's P2 holds, measured from the output rather than inferred from the gate.

The pass sweeps are otherwise symmetric: every probe's pass-0 count equals its pass-1 count, with one exception — IsClaimedByAnotherTask runs 9 times in pass 0 and once in pass 1. The claim registry is already populated by the time the second sweep runs, so most candidates are filtered before the test is reached. That is consistent with AI3's two-tier quota model and is the only asymmetry in either run.


2. Which tasks actually fire (AI3 §5, live)

Eight Execute bodies probed, covering the nine classes AI2 called planners plus AITRaid and AITAdvanceIdleShips. On both turns:

body 1→2 2→3 which agents
AITAdvanceIdleShips::Execute 6 6 all three, both passes
AITBuildDeepScanShips::Execute 4 4 only 496 and 512, both passes
AITColonize, AITEscortGateInvade, AITInvade, AITNodeBore, AITBuildPoliceShips, AITRaid 0 0 —

Two things follow, and the second is uncomfortable.

AITAdvanceIdleShips is in every agent's list and is entered on both passes, exactly as its priority-0, pass-1-body shape predicts. It is a good control and it read non-zero on every run.

Player 32's task list contains none of the six named planner bodies. Its build orders came from RequestBuildForTask, entered 5 times per pass on turn 1 and 4 times per pass on turn 2, under tasks whose Execute bodies were not in my probe set. So the answer to "which of the nine planner tasks fire on a real turn" is one of them, AITBuildDeepScanShips, and only for the two AI players that own nothing — and the one AI that actually plays is driven by tasks nobody has probed yet. The nine were the wrong nine to probe. AI3's §5 correction of AI2 stands on the call graph; this lane cannot add to it, and says so.

The event ring gives the exact per-agent sequence for both turns (aievent lines in the raw logs). Turn 2→3, agent 0x335c1040 (player 32), pass 0, in order:

RunTaskList → Acquire → IsClaimed → IssueRouteForFleets → AssignFleetsAndIssueOrders
            → Acquire → IsClaimed ×2 → RequestBuild   (×4 more of this pair)
            → AITAdvanceIdleShips

and pass 1 is the same walk with the claim tests gone.


3. Rule 19: the control, and what it cost to take it

ref-turn2.sav + one End Turn, with seventeen MinHook detours installed (the batch dump plus sixteen entry probes):

file measured published oracle
(Autosave EndTurn).sav 66,732 B bb4fd9ac89f41e3b bb4fd9ac89f41e3b ✓
(Autosave).sav 67,219 B 978041acd168b56e 978041acd168b56e ✓

Byte-identical. Two things at once: VM145, which is a ZFS clone nobody had checked, reproduces the reference guest exactly; and this lane's instrument is behaviour-neutral. Lane H's own entry probes were explicitly probes=off in every configuration here, because that set is the one measured to move an autosave by four bytes.

3.1 And a control that did not pass — the turn-1 workload is not reproducible

Lane L5 got here first, on VM146, from the other end. turn1-to-turn2-nondeterminism.md is the owner of this result and it is the better-designed experiment: three runs including a pair with identical hooks that still disagreed, which rules out the instrument in a way my configurations cannot. What follows is an independent third-instrument corroboration and one thing it adds.

turn1-state.sav + one End Turn does not reproduce turn2-state.sav, on this build, with or without instruments. Three runs, three different files:

run (Autosave).sav
reference turn2-state.sav 66,739 B ab4ac2d7e2977260
hooks=off 66,746 B d59bb9f2fd0eb535
full instrument 66,740 B e43ec1d2b443c101

Diffed field by field through the save reader, exactly one field differs across all three:

p512.ResTNm:  ref = BIO_GnMod   hooks=off = XNC_TrnsHum2   instrumented = XNC_TrnsMorr2

NMnx, ModCount, every id list, every design, every fleet, every other player's target and rate: identical. And the command block shows it at the source — player 512's research-target gate carried techId 288 on the instrumented run, a different id on the others.

So: the research-target choice of an AI player that owns nothing is not reproducible run to run. Players 16, 32 and 496 are stable across all three; only 512 moves. Three samples, one field.

What this instrument adds to L5's result. L5 measured the divergence in the save. This lane sees it in the TurnCommands block: player 512's research-target gate carried techId 288 on the instrumented run and a different id on the others, in the block the client submits. So the divergent decision is made client-side, before submission, and reaches the save as a command like any other — it is not the sim diverging on identical input. That is consistent with L5's inference (a tie broken by per-process iteration order over a pointer-keyed container) and it removes the sim from the list of suspects, which their instrument could not do.

Second corroborating detail: across L5's three runs and this lane's two, the observed values are BIO_GnMod, XNC_TrnsLir2, XNC_TrnsHvr2, XNC_TrnsHum2, XNC_TrnsMorr2 — five distinct values in five runs, and four of the five are the same tech with a different species suffix. A selection walking a species- or player-keyed container and taking whichever arm it reaches first fits that shape exactly; a numeric roll over a flat tech list does not.

The right probe is the one L5 names: read the AI client's generator state after construction in two processes, and — from this side — put a write watchpoint on the block's research-target payload at block+0x10 to catch the writer with its call chain.


4. Corrections to earlier findings

  • ai-order-emission.md §2 — "the four Species == 4 factions submit no command block at all". Right about the submissions, wrong about the block array: the batch is n = playerCount = 8 and those four occupy slots with playerId == 0 and uninitialised gate payloads. AI3 §1.2 had already read the resize(playerCount) that forces it.
  • ai-order-emission.md §1 / §3 P1 — the 17..27 half of the cost table was read from the instruction stream only, with no workload. It now has one: list 23 carries one element on every AI turn measured, and both turns still cost exactly 12. The free half is confirmed free.
  • ai-order-emission.md §4.5 — list 10 "named only by position". It now has values: {systemId, fleetId, counted vector}, with the system leading. Still not named; the vector's contents are unread.
  • turncommands-block.md §3, list 3. The wire record {ordinal, designId, systemId, w} is correct — but the in-memory element is in the opposite order, because that list's writer (0x00822870) emits +0x14, +0x10, +0x0c, +0x08, descending. It is the only one of the five writers checked that reverses; lists 5, 8, 10, 14 and 23 all write ascending. Anyone reading these elements out of memory needs that per-list, not as a rule.
  • ai-task-system.md / ai-stepping-and-passes.md — "the AI's fleet order names the fleet it moves". It names a fleet the input save does not contain (§1, P7).

5. What this lane did not do

  1. AI3 §2.4 is still open. AITRaid never ran on either workload, so the list-16 pass-0 question is untouched — but now for a measured reason rather than an assumed one (§1, P9).
  2. AI3's P3 is untouched. An entry counter cannot see the steal branch inside IsClaimedByAnotherTask; this lane measured only that it is called (10 times on turn 2, 0 on turn 1). Stated in advance in docs/L4-predictions.md §4.
  3. The dump reads 48 bytes per element and does not follow pointers. Three payloads are therefore unread: the route in list 8 (one hop, id unknown — turn3-state.sav's waypoint says 272 but the capture does not prove it), the counted vector in list 10 (one element), and the Population body in list 23 (24 bytes behind a vftable). Each is one more indirection in the dumper.
  4. The design id is not visible in the list-1 element (§1, P3). It is inferred from list 3 naming design 18 in the same block.
  5. The client id counter is not located (§1, P7). This is the largest remaining hole and it is a watchpoint: break on the write that produces 18 and 34.
  6. Which tasks drive player 32 is not known (§2). Five build-shaped tasks per pass reached RequestBuildForTask and none of them is one of the eight Execute bodies probed. The probe set was chosen from AI3 §5's list of nine and that list is not where this AI's decisions come from.
  7. Two workloads, one AI empire, 28 stars, no contact. Every count here is a count on a very quiet board: no colonise order, no invade, no raid, no diplomacy, no combat. Lists 2, 4, 6, 7, 9, 11–22 and 24–27 were empty on both turns and remain unexercised (rule 6). The distinct-state count for this lane is two turns and one real AI player — that is thin, and it is the caveat that matters most for anything generalised from here.
  8. The +0x6c CivilianRatios gate and the +0x3c Hiver gate were clear on all eight blocks on both turns, as expected; nothing new about either.

6. The instrument, for the next lane

sots-engine src/shim/hooks/ai_orders.{h,cpp}, configured by three keys:

  • aiorders=on|off — the block dump: one register-transparent entry stub on StrategySim::ApplyTurnCommandBatch, which receives (blocks, n) as stack arguments with every submitted block complete at a fixed 0x1b4 stride. Prints six gates, twenty-seven list lengths and 48 bytes per element per block.
  • aiprobes=off|all|N — sixteen entry counters, lane H's asm-stub pattern with its own table so lane H's set is untouched. N installs the first N, so the set bisects in one build.
  • aiorders.out=<path>.

Two design points worth keeping:

Row 0 is RunTaskList, and it is both the control and the pass recorder. Its stub reads the pass stack argument before tail-jumping, so every later probe hit is attributed to a pass. That is what turned "pass 0 emits nothing" from an inference into a measurement. The global is stale once RunTaskList returns and the report says so; the run column in the event ring is what makes the staleness readable.

Every list is measured twice — walked, and read from _Mysize — and a disagreement prints MISMATCH. Nothing printed it on 8 blocks × 27 lists × 2 runs, which is the evidence that the container layout is right rather than that the block is empty (rule 1). And a run whose RunTaskList count is zero prints CONTROL ZERO and says every other row is unmeasured, not absent — which is exactly what the load-time batch (seq=1, n=1, the local client's block alone) does print.

6.1 Lab notes

  • move X Y then click X Y in the same click-helper batch is reliable; a bare click is not. Roughly half of bare clicks were delivered at the previous cursor position, which reads as "the click did nothing" and then as "the next click did the previous thing". Two runs were nearly lost to it before the pattern was clear.
  • Reset SavedGames to a fixed two-file set before every run. With only ref-turn2.sav and turn1-state.sav present the Load dialog rows are always y=262 and y=291 and the click path never has to be re-derived. C:\SOTS\ui\l4deploy.ps1 does it.
  • Startup to main menu on VM145 was 80–95 s. Verify by screenshot; never sleep and click.
  • VM145 left restored: SavedGames back to the 9-file pre-L4 set (autosaves byte-identical to the oracle), binkw32.dll and shim.cfg back to the W3 build and w3mod config, game stopped. C:\SOTS\shimdist-l4, C:\SOTS\ui\l4\ and C:\SOTS\ui\l4{deploy,click,grab,release}.ps1 left in place — they are a working template for the next lane.