Commit graph

103 commits

Author SHA1 Message Date
alex
c67374d2e3 CB: take lane L4's ai_orders instrument as the base for the command-stream capture
The capture lane needs the block dump L4 built; branching off main without it
would mean writing the same detour twice. Header regenerated from sots-re
(rule 14), not hand-resolved: 1,217 entries.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ARBgSooAfokKUy6wKUKEyZ
2026-09-08 18:28:18 -04:00
lane-l4
c842d44d62 L4: engine doc for the research-selection capture 2026-09-08 18:26:59 -04:00
alex
4f25f1e8c0 merge lane L1 (header regenerated from addresses.json + fragments, not hand-resolved) 2026-09-08 18:25:53 -04:00
alex
fa53e02b53 L1: probe the AI client seed across two processes -- it is fresh every time
Two hooks, Mars::RNG::Seed 0x0049fdf0 and StrategyApp::RunAI 0x008706f0, and a
config that turns everything else off. Two launches from the same save, load
only -- the AI clients are constructed on load, so no End Turn is needed.

Result: every AI client's generator seed differs between processes (net 32,
496 and 512 all move), while the record structure is byte-for-byte the same
shape and one Seed call with seed=0 produces an identical state in both runs.
So the turn1-state -> turn2 nondeterminism is a SEED effect, not the ordering
effect that was predicted, and lane AI1's 'every draw from the static generator
returns 0' is falsified by measurement.

The prediction said the opposite and is left in docs/L1-predictions.md with its
outcome underneath.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ARBgSooAfokKUy6wKUKEyZ
2026-09-08 18:22:14 -04:00
lane-l4
ec8d841dba L4: predictions for the research-selection tie set, before the instrument 2026-09-08 18:06:31 -04:00
lane-l4
c769394b57 L4: credit lane L5 for the turn-1 nondeterminism, and record what the block dump adds to it 2026-09-08 18:03:18 -04:00
lane-l4
4a5c218a35 L4: read the AI's command block out of the running game
New shim module src/shim/hooks/ai_orders.{h,cpp}: one register-transparent entry
stub on StrategySim::ApplyTurnCommandBatch dumps every submitted TurnCommands block
(six gates, 27 list lengths, element bytes) at the point where all of them are
complete in memory; sixteen entry probes, with RunTaskList's stub recording the
pass so every later hit is pass-attributed.

Two workloads on VM145, one End Turn each. The rule-19 control passed with all
seventeen detours installed: both autosaves byte-identical to the published oracle.

What the AI actually emits, and three things no reading had produced:
  - a list-23 element on EVERY turn, the first element ever observed in the free
    half of the cost table -- and both turns still cost the measured 12;
  - the ids in AI commands are client-allocated and travel in the command (design
    18, fleet 34; neither exists in the input save);
  - pass 0 emits nothing, measured from element counts rather than inferred.

tests/game_ai/test_live_blocks.cpp rebuilds both captured blocks through the public
OrderClient API and asserts the list profile, element values, gate counts and
ModCount total: 44 checks. Kept separate from test_orders.cpp, which stays the
record of what static reading predicted.

Gates: clean_room_check OK, host ctest 55/55, CT111 shim cross-build exit 0.
2026-09-08 18:01:19 -04:00
alex
b39bb290e3 L5: the interest literals verified live at a boundary, with a control that fails
`ComputeBudget`'s savings-interest term is now compared against the running game at
a treasury the corpus actually contains. Three runs on VM146 from turn1-state.sav:

  A (widened floats, as shipped)   3,895 calls, 0 diverged, 0 undeclared writes
  B (exact decimals, the control)  2,718 calls, 1,359 diverged

The game fills savingsInterest with 499 at a treasury of 50,000, and with 380 at
38,100 -- the exact decimals pay 500 and 381. Every divergence in B lands on a
treasury that is a multiple of 100 and no other state diverges at all, which is
exactly the arithmetic. G3's rule-23 reading is now measured, not inferred, and the
one-money error is shown to propagate into `available` and `researchMoney` too.

The control also settles why the earlier 4,437-call green run was green: slot 5 IS
diffed and the harness CAN see it, so that run simply presented no boundary state.
Coverage is therefore reported as distinct states, not calls: 5 distinct treasuries,
2 of them on the boundary.

Two further rule-23 constants found in the same routine by an operand-width sweep,
corrected, and honestly marked UNVERIFIED because no reference turn can see them:

  - the research-yield factor is a widened 0.85f while its two neighbours in the
    same product are exact doubles. Boundary: research money a multiple of 40,000;
    the run presented 9 distinct values and none is.
  - the three research modifiers are summed in single precision, not double.
    Boundary: two of the three non-zero; the corpus has shrm = TRM = 0.

Both are pinned by boundary cases in test_economy.cpp that fail with the decimals.

Also verified live, in the same run:
  - T31's difficulty-column recovery. The live ServerPlayer+0xf9 / NPC flags on all
    eight players are exactly what lane PL's save-only inversion claims, including
    the awkward system-owning player that is still ambiguous because it is an NPC.
  - BANKRUPTCY_PROTECTION_LIMIT_FACTOR reads 3.29999995 = (float)3.3. Its file image
    is zero because the loader fills it at run time, so lane PL-3 had to assume the
    value; it is now measured and the assumption was right.

Falsified, and recorded as such: the difficulty-mods record does NOT sit inline at
ServerPlayer+0x36c -- that field is a heap pointer on all eight players. The row IS
reachable from a ServerPlayer (which corrects the hook's standing coverage note),
but the fitted {3.0,1.5}/{1.0,1.0} pair remains unverified. The hook logs the
pointer and does not follow it.

The `verified` column stays 0, deliberately. Every phase this compare touches is
Partial for reasons upstream of it, and promoting one because part of it was checked
is the drift app_test_catalog exists to catch. What moved is models; see
docs/L5-live-verification.md for each one with its coverage.

Gates run separately: clean-room OK, host ctest 54/54, CT111 shim cross-build exit 0.
2026-09-08 17:48:42 -04:00
alex
2947e24ed9 L1: hook BeginProcessTurn and the three script-object writers a turn reaches
Lane SV recovered the script-object subsystem statically and predicted that
SVSOSwarmQueen::RegisterHives takes one RNG_NextInt per new hive inside
StrategyServer::BeginProcessTurn -- which runs inside lane Z's autosave bracket
and outside every one of its subtotals, so a draw there had never been
attributed by anything.

Five nested trace hooks, each declaring the strategic generator as a region and
each carrying a model evaluated at entry so the record can disagree with it:

  StrategyServer::BeginProcessTurn        the unhooked interval, plus a region
                                          over Frame so the increment is a fact
    SVSOSwarmQueen::OnTurnBegin           evt 0x13, vtable slot +0x60
      SVSOSwarmQueen::RegisterHives       predict_new_hives from the original's
                                          own two predicates; reads the LO/HI
                                          config pointers live
      SVSOSwarmQueen::TickHives           the NextQ slip and its three gates
    SVSOSlaversRefuel::UpdateDifficultyTier  tail phase 20; a 4-byte region over
                                          CDiff so a store and its ABSENCE are
                                          distinguishable (method rule 20)

Also watch.mode=snlv: one arming line moves slot 1 from the NVO map's _Mysize to
the target system's SnLv, and the arming sweep now prints SnLv with its decoded
per-player 2-bit levels for every system, which costs no debug register.

Measured on VM140: hive creation costs 2 words in BeginProcessTurn and the
residual outside the two turn drivers is 2, not 0; the next turn it is 0 again.
LO=20 HI=30. Spica's SnLv reads 0x200 with AFlags 0. The oracle reproduced
byte for byte with every one of these detours live.

Predictions and outcomes: docs/L1-predictions.md.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ARBgSooAfokKUy6wKUKEyZ
2026-09-08 17:43:10 -04:00
lane-l4
45bf085981 L4: predictions for the live AI order capture, before the module exists 2026-09-08 17:07:15 -04:00
alex
fe8ed45de4 merge lane SV: script-object event bus (CMake lists merged into their constructs, rule 22) 2026-09-08 16:48:49 -04:00
alex
2d93404c67 SV: the script-object event bus, and what it writes in a turn
`SvSctOb` is not written by direct calls. Every update goes through an event bus: a
driver notifies the root object with an integer id, the root fans the delivery out to
every child, and each delivery is two steps -- a generic handler that takes the id, then
one event-specific vtable slot that does not. The id -> slot map is a 33-entry jump table
in the image, so which class reacts to which event is recovered and exhaustive rather than
inferred from what the saves happen to show. Five of the 33 rows are not in slot order,
including two the tail sends.

Three handlers write the eight leaves that diverged:

  * the slavers' difficulty tier, on the tail's end-of-turn delivery -- a three-record
    stack table scanned against the frame, boundaries 1/50/100, stored only on a change,
    and at frame 100 and above the scan runs off the end and stores nothing, so the tier
    can never reach 2;
  * the refugees' one-shot latch, on the turn-begin delivery, with a design instantiation
    behind the same latch that nothing here can do;
  * the swarm queen's hives, also at turn begin, registered on the systems carrying the
    SWARM's scenario tag (the queen's constructor stores 3 for that and 10 for its own
    encounter id) and then ticked -- and the tick is the whole explanation of a target
    turn that reads 31 after one turn and 32 after the next. It is not re-rolled; it slips
    forward by one every turn the spawn gates stay shut.

New host phase H03 for the turn-begin delivery, run right after the frame counter where
the original sends it, and tail phase T20 implemented. Rules are pure in game/sim.

Measured on CT111, closed and regressed stated separately:

  default                 turn1->turn2  209 -> 126 (was 128)  closed 83, regressed 0
                          turn2->turn3  108 ->  67 (was  69)  closed 41, regressed 0
  --commit-blocked=H03    turn1->turn2  209 -> 124            closed 87, regressed 2
                          turn2->turn3  108 ->  67            closed 41, regressed 0

Registering a hive closes the four leaves that say which systems have hives and that they
have no queens, and opens two carrying a target turn known to be wrong: the original draws
it from the strategic generator inside the turn-begin step, outside both turn drivers, and
neither the two data-file constants nor the generator's position there is settled. That
trade is a flag, not a default.

The prediction in docs/SV-script-objects.md was committed before the build, and P5 was
wrong: it called the second pair a null control, and the second pair is where the slip
rule is tested EXACTLY -- two hives, two target turns, both landing on the oracle with no
draw and no fitting.

Gates run separately: clean-room OK, host ctest 51/51, CT111 shim cross-build clean.
2026-09-08 16:46:25 -04:00
alex
5d245e4c8b merge lane W3: TShn writer trapped and gate named (158/158); Player.Status predicate settled; AI4 prediction confirmed 2026-09-08 16:42:20 -04:00
alex
ca1b8c2cba docs: name 0x00743ec0 in the W3 predictions -- clean_room_check rejected the raw identifier
The prediction text is unchanged; only the decompiler-style identifier is replaced with the name
lane E3 had already recorded for it. Rule 13's gate did exactly its job on a doc that had been
committed before the build.
2026-09-08 16:39:23 -04:00
alex
9e914ef358 W3: watch.mode=tshn -- arm the NVO/TShn record and the trade+spy containers from the same S
Second mode in lane W2's watchpoint module. No second hook: only different arithmetic on the S the
ApplyAllTurnCommands detour already holds.

- picks the target system by predicate at arm time (AFlags == 0, NVO non-empty) and logs all 28
  systems, so the choice is auditable rather than a hard-coded pointer;
- probes both candidate ServerSystem bases and logs how many systems validate under each, which
  settled a documentation dispute (+0x274: 9, +0x26c: 0) by measurement;
- prints the trade-route and spy-program vector triples, which is the workload-confirmation
  instrument two earlier lanes lacked.

Rule 19 control passed: the armed run reproduced the determinism oracle byte for byte.
2026-09-08 16:38:55 -04:00
alex
26b041106f PL: decompose the /Sim/players residual; S00 PvSav snapshot; T31 recovers its difficulty column from the save; the bankruptcy protection factor is a widened float
Closed 2 on the reference pair and 4 on pair 2, regressed 0, with no operator
input. Measured, both pairs, never netted.

- S00 stamps PvSav from Sav on every live player, before any phase can move it.
  0 closed on pair 1 (a no-op there), 2 on pair 2.
- T31 recovers the per-player difficulty column by recomputing BnkEl from the
  colony state the input save was written from and comparing against the BnkEl
  the save carries. That removes the --ai-player flag as a blocker and turns the
  phase's self-check into a real one: it used to compare its POST-turn result
  against the PRE-turn stored value, so its 6-of-8 only ever covered the six
  players whose limit does not move. The load-time check passes for every live
  player of all eleven corpus saves. T31 Blocked -> Partial; BnkPr still needs
  the tuning constant.
- BANKRUPTCY_PROTECTION_LIMIT_FACTOR is multiplied in as fmul dword ptr, so it
  is a float32 in the image; the engine narrows it now. Zero leaves move on this
  corpus -- all seven of its BnkPr records land where the two constants agree --
  and three hand-written test expectations moved (rule 23).

docs/PL-players-residual.md carries the decomposition, the predictions written
before the build, where they were wrong, and the ranked remainder.
2026-09-08 16:34:40 -04:00
alex
c852965e0e PL: predictions for the players residual, written before the build (rule 2) 2026-09-08 16:19:34 -04:00
alex
bb81d3db2d docs: W3 predictions committed before the build -- TShn writer, trade/spy container confirmation, rule-19 control 2026-09-08 16:04:17 -04:00
alex
aabd8a3506 merge lane G3: civilian growth; app CMake list and includes union-resolved into their constructs (rule 22) 2026-09-08 15:54:32 -04:00
alex
ed6602e9ed S11 civilian growth; fix ComputeBudget's interest literals
Reference pair turn1->turn2: 81 leaves closed, 0 regressed (was 78/0).
Pair turn2->turn3: 39 closed, 0 regressed (was 36/0). With
--commit-blocked=T31 --ai-player 1: 83/0 and 41/0.

game/sim/colony: GrowCivilianPopulations models ServerSystem's civilian
growth sub-pass. The whole system's delta is clamped to 20,000,000 -- an
int64 column of the population-type table, built in the executable from
its own literals -- and on both reference pairs that clamp, not the growth
curve and not any carrying capacity, is what decides the value: the
uncapped delta is 7.5x it and the capacity headroom 25x it. So the pass
commits with no tuning table loaded, and says by how much each unmodelled
input would have to be wrong before it mattered.

The one input genuinely off the wire is the per-species civilian capacity
factor. It is handled by running the pass twice, once with the modelled
capacity and once with the system's own wire-known dcs limit, and
committing only when the two agree. Imperial growth is deliberately NOT
committed: it is a no-op on this corpus and would need a capacity the
corpus can bound from below but not from above.

game/sim/economy: both interest rates in ComputeBudget are WIDENED FLOAT
literals, (double)0.01f and (double)0.15f, and are then truncated -- so a
treasury of exactly 50,000 earns 499, not 500. This module used the exact
decimals, which left the human's savings one money high on the first
reference pair and exact on the second. Sixteen hand-computed test
expectations moved by one; they were derived from the model, not measured.
The live ComputeBudget compare (4,437 calls, 0 divergences) did not catch
this because it presented only 20 distinct states and none sat on a
rounding boundary.

game/sim/colony: ShipRepairCost, the last unmodelled input of the output
turn path. The demand is still 0 -- its two design fields are cached stats
the save does not carry -- but the zero is now evidenced rather than
silent: S13 reports the candidate set, and the independent colony keeps a
ten-ship fleet over a colony whose savings close exactly at zero demand.

Gates run as separate commands: clean-room OK, host ctest 49/49, and the
CT111 shim cross-build exit 0 (required: game/sim is compiled into the
shim). The host build and report were also re-run on CT111 and produced
identical numbers.

docs/G3-civilian-growth.md; notes repo
findings/subsystems/population-growth.md.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ARBgSooAfokKUy6wKUKEyZ
2026-09-08 15:52:38 -04:00
alex
5d01c7c7de merge lane EV: P11 event posting; +4/+3 leaves; NO_RESEARCH gate corrected; PostEvent turn is Frame not ModCount 2026-09-08 15:42:58 -04:00
alex
e30d43dd1b P11: post the no-research event into the save's turn bucket
The engine has modelled the event log since lane E, but the standalone never wrote
any of it into the save it produces. This wires the two together for the one event
on the reference pair the standalone can compute, and corrects the condition.

What a turn actually posts, measured over all eleven saves: two events on
turn1 -> turn2 and three on turn2 -> turn3, and only two players in the whole
corpus ever hold an event at all. Order is readable off the ids -- the build pass
posts before the research pass. See docs/EV-events.md section 1.

The condition had two of the original's three tests. The missing one is "no tech
finished on this turn or later", and it is what keeps the event off the turn a tech
lands; zuul-turn23 exercises it. The third input, whether the player holds a target
at the moment of the check, is not on the wire -- three of the four real players
acquire one during the turn, which is AI research selection -- so the operator's
--ai-player roster stands in for it as a stated hypothesis and the phase stays
blocked without it.

No prose in the engine: the record's text is resolved through a caller-supplied
lookup over the operator's own installed string table, and a run without a data
root posts nothing rather than writing a record it cannot fill.

Measured (CT111 host build, real data root), closed and regressed never netted:
  turn1 -> turn2  209 -> 127  closed 82 (+4 this lane), regressed 0
  turn2 -> turn3  108 ->  69  closed 39 (+3 this lane), regressed 0
The posted record agrees with the oracle on all eight of its fields. Controls: with
the roster withheld the phase over-fires and regresses 16 leaves on pair 1; with the
string table withheld it posts nothing and regresses none.

Gates run separately: tools/clean_room_check.sh OK; host ctest 50/50; CT111 shim
cross-build OK (exports 66 names identical to binkw32.dll, staged in
/srv/re-lab/shim/dist-ev); CT111 host ctest 50/50.
2026-09-08 15:39:42 -04:00
alex
cc77de3429 merge lane W2: watchpoints + multiplayer Tier 0 (CMake list and main.cpp union-resolved to keep both instruments; header regenerated) 2026-09-08 15:28:28 -04:00
alex
4065808a05 docs: EV prediction committed before the build -- P11 posts one event, closed 4/3, regressed 0 2026-09-08 15:23:16 -04:00
alex
db99971efb merge lane H probe module (header regenerated, not hand-resolved) 2026-09-08 15:00:42 -04:00
lane-b6
db3909bcd4 docs: B6 measured -- prediction held, closed 0 regressed 0 on both reference pairs
209 -> 158 (closed 51, regressed 0) and 108 -> 87 (closed 21, regressed 0), identical to the
pre-lane baseline, measured on CT111. Falsification hypotheses 1 and 2 refuted by measurement
(three empty BQ frames per reference save, every hbq false, the one command block empty);
hypothesis 3 refuted by the corpus oracle; hypothesis 4 stands as a labelled hypothesis.

construction.h also carries the ship/fleet birth shape as read but NOT implemented, with the
reason stated: the newborn hull copies four cached stat words the engine does not yet compute.
2026-09-08 14:54:21 -04:00
lane-b6
9f8d75d9cb docs: B6 prediction before the build -- the build queue closes nothing on either reference pair
Written before any code (earned rule 2). Both reference saves carry three BQ frames with
zero orders, every hbq false and the only TurnCommands_v5 block empty, so the pass has
nothing to advance; the archived destroyer is created inside the turn by the AI. Includes
the falsification list and names the oracle pair the campaign already owns.
2026-09-08 14:33:16 -04:00
alex
a26cd183f1 lane H: entry-probe module, ProcessTeamRecord hook, probes= config, and the outcome of the five predictions 2026-09-08 13:38:59 -04:00
alex
5e409cfa05 lane A2: S04, the alliance mask -- 80 player-records, 560 fields, 0 mismatches
The spine's fourth phase, read byte-for-byte and implemented:

    almem[i] = (1 << i) | (ALid != -1 ? AL : 0)

with i the player's POSITION IN THE PLAYER VECTOR, not its index field. Both
inputs are on the wire and so is the output, through the turn-record archive,
so the phase is checkable against bytes the original wrote:

    app_turn_record: 11 saves, 80 player-records, 560 fields, 0 mismatches
                     (was 480 fields over six fields; almem is the seventh)

The eight zero masks of the corpus's earliest archived turn are PREDICTED, not
excluded: the archiving phase also runs on load, and the load path does not run
the spine. BuildTurnRecord takes spineRan and models it, so all 80 records are
compared.

Three parts of the rule the corpus cannot separate -- the bit index, the OR,
and the ALid guard -- are pinned in app_alliance with the separating inputs no
save provides, and app_turn_record prints that it could not separate them.

Divergence, closed and regressed reported separately:

    turn1->turn2  default          209 -> 204   closed 5, regressed 0
    turn1->turn2  --commit-blocked 209 -> 189   closed 29, regressed 9  (was 17)
    turn2->turn3  default          108 -> 103   closed 5, regressed 0
    turn2->turn3  --commit-blocked 108 -> 106   closed 13, regressed 11 (was 19)

T36 stays blocked: nine leaves would still be wrong (inc x3, sav x3 behind the
budget; three census leaves behind the design catalogue). It now closes all 24
turnstats leaves on the reference pair, so it becomes a clean +24 once those
two land.

Prediction and falsification committed first in 49ae628.
Gates run separately: clean-room OK; host ctest 43/43. No src/shim touched.
2026-09-08 12:44:00 -04:00
alex
a01305fcf2 lane H: P2b-alt, a competing prediction from the callee bodies (still before any run) 2026-09-08 12:31:29 -04:00
alex
1932377b92 lane H: five probe predictions, committed before the build 2026-09-08 12:23:42 -04:00
alex
49ae628220 lane A2: prediction for the alliance mask and ModCount, written before the build 2026-09-08 12:23:06 -04:00
alex
0ebc222f45 lane N: the population -> base-output term, live-verified
Reads the whole colony output chain off the instruction stream (every range
disassembled to the next function start) and compares two of its functions
against the running game.

The population -> output law is linear and is carried by the executable:
output points per head are typeOutputModifier x 1.8 / 500000, and the
three-row population-type table is built in code rather than loaded, so the
imperial (1.0) and civilian (0.33f) modifiers are facts about the binary.

A system's total output is a SUM of three terms, not one multiplicative
chain. The station bonus scales only the imperial term and morale only the
civilian one, so OutputModifiers no longer carries either; they belong to
GroupOutputInputs. The function previously described as the base-output term
is the over-harvest RESOURCE demand, and it is corrected in place.

Live on VM140, both hooks in compare mode over two species and two workloads:
GroupOutput 13,105 calls / 0 divergences; ComputeTotalOutput 11,252 calls /
1 divergence of one ulp, in a value its caller rounds to an integer. Both
functions declare a whole-object Guard: 0 undeclared writes in 24,357 calls,
which is what makes the side-effect-free claim a measurement.

sim::Narrow forces the double rounding a 32-bit x87 build otherwise skips;
without it every civilian row came out one ulp low.

Also fixes ComputeBankruptcyLimits' elimination divisor, which was the
decimal -0.15 rather than the image's widened float -0.15000000596046448.
The two disagree for every maximum income divisible by 3 and for essentially
every empire above ~3,000,000.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ARBgSooAfokKUy6wKUKEyZ
2026-09-08 12:11:36 -04:00
lane Y
5cd301dcf5 Y: correct the divisor-defect numbers -- 6 of 25 corpus records, and the rate rises with magnitude 2026-09-08 11:39:36 -04:00
lane Y
655034fb51 Y: outcomes -- four predictions held, one was wrong in its framing and the correction is the result 2026-09-08 11:37:06 -04:00
lane Y
20e5112398 Y: predictions for the RNG model and the tail turn record, written before the build 2026-09-08 11:21:00 -04:00
alex
b48d860f8a merge lane Z: per-turn RNG ledger (header regenerated, CMakeLists union-resolved) 2026-09-08 10:58:10 -04:00
alex
39c01422f7 src/app: the standalone -- load a save, run a turn, write a save
`sots_turn` loads a save through the engine's own reader, walks the published
phase order of all three turn drivers, runs what we hold, prints what we do
not, and writes the result back through the engine's own writer.

The phase catalog carries all 32 + 12 + 37 phases whether or not they are
implemented, so an unimplemented phase is a named no-op that appears in the run
log rather than a silent absence. 14 of the 44 turn-driver phases are modelled,
7 commit anything, 2 of the 37 tail phases are modelled.

Modelled but NOT committed is a first-class state. A phase whose formula we hold
and whose inputs we do not is evaluated, reported, and left unwritten unless
--commit-blocked is passed. That distinction was earned: committing phase 31's
player-status restore regressed two leaves that had agreed with the oracle
before the turn, because the phase writes 1 and the file carries 4.

Measured against the game's own post-turn saves, leaves localised by
state_checksum.py with coverage proved by re-serialisation:

  turn1-state -> turn2-state   209 -> 204 diverging, closed 5, regressed 0
  turn2-state -> turn3-state   108 -> 103 diverging, closed 5, regressed 0

Two tests: app_catalog (the tables stay complete and nothing claims to be
verified against a live game) and app_turn (11 saves driven; an untouched load
re-serialises byte-identically, a turn leaves the file re-readable, and no
blocked or stub phase writes anything). Skips cleanly without SOTS_SAVES_DIR.

ctest 38/38, clean-room OK. src/shim untouched. docs/S-standalone.md has the
full gap list.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ARBgSooAfokKUy6wKUKEyZ
2026-09-08 10:35:45 -04:00
alex
2f4fcfcb21 Z: P9 outcome -- turn 64, one word, every clause held 2026-09-08 10:24:17 -04:00
alex
0b5679cf30 Z: P10 outcome -- wrong, and the falsification is the result 2026-09-08 10:15:06 -04:00
alex
d7e2080fb6 Z: P10 -- what Auto Resolve should cost, written before the click
The turn-54 End Turn stopped on a Von Neumann encounter at Gallandro. That is
the workload the finding says does not exist and lane J asks for: every
encounter measured so far had the no-battle flag set, so the combat resolver
has never run under an instrument.

Three falsifications, and the interesting one is the third: a non-zero cost
that does not appear under ApplyEncounterResult means combat proper is drawing,
which nothing hooks, and the bracket residual goes positive for the first time.
2026-09-08 10:12:09 -04:00
alex
dc43f93910 mars::rng: the seven draw entry points, with their word costs
A per-turn RNG budget is only as good as the entry-point table, and ours had
three of the seven. Adds the two that are modellable and documents the rest.

  float_range(lo, hi)     exactly one word.  Narrows TWICE -- the scaled product
                          is stored to a 4-byte float before lo is added, and the
                          sum is stored again.  Evaluating in double and narrowing
                          once disagrees on a measurable fraction of words, and the
                          test asserts the two models are distinguishable so the
                          shortcut cannot creep back.
  int_range_bell(lo, hi)  AT LEAST TWO words.  Triangular, not uniform: the span is
                          split into h/2 and h - h/2 (truncating toward zero) and
                          each half drawn inclusively, first half first.  The bounds
                          reach the draw as unsigned, so an inverted range yields a
                          huge first bound rather than an empty one; reproduced, not
                          corrected.

Documented but deliberately not modelled: a truncated-normal integer range built
on rejection sampling around a Box-Muller pair.  It costs TWO WORDS PER ATTEMPT
and the attempt count is unbounded, and predicting its stream position needs log,
sqrt and cos to agree bit for bit with the original CRT.  Nothing in the strategic
turn reaches it.  It is recorded so a ledger that meets it does not score its two
words as one draw.

Also recorded in docs/mars-rng.md, because each is a way a word budget goes wrong:

  * two calling conventions for one generator -- three entry points take the state
    block (the object plus four bytes) and four take the object itself, and one
    caller uses both within forty bytes of itself;
  * two different divisors in the same image, 1/(2^32 - 1) for the unit draw and
    2^-32 (with a +0.5 offset on the word) for the normal path;
  * the unit draw is inlined at twenty-eight sites across eleven functions, so any
    budget assembled by counting calls is a LOWER BOUND.  Exactly one of those
    eleven is reachable from the strategic turn driver.

Host ctest 36/36 and tools/clean_room_check.sh run as separate commands, both
clean.  No src/shim change, so no cross-build is implicated.
2026-09-08 09:59:05 -04:00
alex
fe39c82bea Z: P9 -- the first node-line draw lands on turn 64 and costs exactly 1 word
Written with the run in flight at turn 34. min_life is decrementing by exactly
1 per turn (43 down to 30), so the traffic term contributes nothing on this map
and the oldest mortal line expires 30 turns out.

Four ways it can be wrong, each with its own symptom. Every prediction this
lane has made so far compared 0 against 0; this one does not.
2026-09-08 09:50:21 -04:00
alex
6589a3985f Z: the run, and one falsified prediction of my own
Eight End Turns across two saves. A turn costs 18-22 generator words, all of
it inside StrategyServer::ProcessTurn; the tail costs 0; the residual outside
the two drivers is exactly 0 on every complete bracket. The generator does not
move between turns at all, so the interval a standalone has to reproduce is
closed at both ends.

Checked against the save files, not just against itself: the turn-6 autosave
pair gives 18 words read from the two Sim.RNG blobs, with twists == 0 -- so
that number never passes through a twist implementation and the agreement is
about the game rather than about two copies of one algorithm.

P6 was wrong. S+0x8 advances 12-14 times per turn, not twice; both drivers are
hooked so the other increments come from somewhere unidentified. Lane K's 'at
least twice' was right and its conclusion is strengthened.

Node-line decay still has not fired, and the hook now says how far away it is
rather than that it did not happen: 51 of 53 lines are permanent, the mortal
ones are dug ~1/turn by the Zuul, and each is ~40 turns from expiry.
2026-09-08 09:43:46 -04:00
alex
496a5124c9 Z: measure the strategic RNG, do not assume it
The turn's RNG cost has never been measured end to end. combat-done-tail.md
found two draw sites in OnAllCombatDone_Tail that nothing models and that both
run before the autosave, so a reimplementation that reproduces both ProcessTurn
functions exactly still diverges the first turn a node line expires.

RngLedger recovers an ABSOLUTE WORD POSITION from (mt[624], left) alone, by
indexing the forward-only chain of blocks the twist generates. Word deltas
between any two observations are then exact -- across twists, across NextInt
rejection loops, and across draws nobody hooked. That last point is not
theoretical: the image has four draw entry points, one of which (NextUInt
0x004f7670) appears in no previous lane's primitive set, plus inlined draws in
twelve functions. A primitive-counting hook would have undercounted silently.

Six nested trace hooks bracket one End Turn between the two autosaves and
attribute the words: the two turn drivers, the two tail phases that can draw,
and ProcessNodeSpaceTravel because it runs twice a turn. NodeLineDecay carries
a real model -- one word per expired node line under NodePath::RemainingLife --
so compare mode checks the count rather than reporting it.

Corrections from the instruction stream, both load-bearing:
  * StrategyHost::Autosave is ret 8, not ret 4, and returns the std::string* in
    EAX. A void-returning hook would have dropped it at both call sites.
  * node-line decay's 0x20000-fleet skip runs AFTER the Chance(0.5f) call, not
    before, so it cannot change the draw count -- combat-done-tail.md reads as
    if it gated the roll.

fpu.sample_turn releases StrategyServer::ProcessTurn, which the fpu sampler and
this ledger both want and MinHook grants to one of them. Default on: no
existing run changes behaviour.

Host ctest 37/37; shim cross-built on CT111; clean-room check OK.
2026-09-08 09:16:24 -04:00
alex
45a7a8c7b8 Z: prediction for the per-turn RNG ledger, written before the build
Six nested trace hooks bracket one End Turn between the two autosaves and
attribute every word the strategic generator consumes to a phase. Position is
recovered from (mt[624], left) alone via a forward-only block chain, so word
deltas are exact across twists and across NextInt rejection loops.

Eight predictions with their falsifications, including the two that matter:
the tail runs on a no-combat turn (lane K inferred it), and a quiet turn's
residual outside the two drivers is zero.
2026-09-08 09:02:32 -04:00
alex
c883a325ad lane T merge fixups: name the roll-succeeded branch instead of its FUN_ id (clean-room); mark describe_i32 maybe_unused so the shim cross-builds 2026-09-08 08:13:39 -04:00
alex
a7ca208b63 T: hook for Game::ServerPlayer::ProcessTurn, with the prediction committed first
Descriptor + pure adapter + host tests for the per-player turn driver. Not
deployed; the WIN32 half is unbuilt here (no cross-compiler on this host).

The declared boundary is narrower than the function on purpose. Phases 2, 3
and 6 -- the savings apply, the aid records and the research refund -- are pure
functions of ComputeBudget's 22 slots and ProcessResearch's overBudget, and
both live in the original's own stack frame. Reaching them would mean calling
ComputeBudget ourselves (it repairs ships in orbit, audit #6), reading the
nested B1/B3 hooks (audit #5, the self-fulfilling compare), or inferring them
from the Sav delta. So they are guarded, not checked, and the three formulas
are written and unit-tested but not wired into the verdict.

Declared: the phase-7 clear, the RebAI decay, the descending timed-bonus
sweep, plus roll_flags and rng as observations ours never writes. Guards over
the whole ServerPlayer and the TechTree header.

docs/T-turn-driver.md states, before any run: which regions must not diverge,
which checks are weak by construction on the reference save, what falsifies
the ResearchRollPending reading, and the save that would finally fire the
branch nobody has seen.

host ctest 36/36 (was 35/35); clean_room_check OK.
2026-09-08 08:08:47 -04:00
alex
8e45b43638 A: type the AIAgent custom-data blocks; named coverage 98.0% -> 99.9%
Game::StrategyAIAgent::Streamable and the ten shapes under it. The whole
writer is unconditional -- the branch the decompiler shows around lnat is an
inlined vector destructor whose operator delete is marked noreturn, and both
paths converge -- so the recovered sequence and a single record are the same
sequence, and all 36 items match with 0 wire-only and 0 shape-only.

CD blocks are now selected by the CDT id at the same ordinal, in both
directions; the one .TurnCommands_v5 block per save still falls to a Node.

Also: Sim's Attrib was not an empty frame, it was an AttribMap holding a count
of 0, and typing it closes those two items too. And StreamableEnum<T> writes a
frame containing one int, not a bare int, so SysMem/mts/nalat are arrays of
one-int frames -- byte-neutral, since all three have count 0 in every save,
but the previous typing was wrong.

Conformance 74 shapes/769 items -> 86/838, still 0 MISMATCH. Round trip
byte-identical on all four saves; ratchet 97.5 -> 99.8. Every container that
is empty in all four saves is named as such in the notes; the new unit test
populates each one, since nothing else exercises them.
2026-09-08 07:33:06 -04:00
alex
872e214d8e merge lane W: SvSctOb/DOpts/spies2 typed; named coverage 97.1 -> 98.0%; conformance 74 shapes 0 mismatch 2026-09-08 07:06:53 -04:00