Compare commits

...

4 commits

Author SHA1 Message Date
alex
f8f5581204 board: standalone scaffold runs on all 11 saves; 5 leaves closed, 0 regressed; ranked byte-match blockers 2026-09-08 10:39:39 -04:00
alex
c504729341 lane S2: the standalone scaffold, and the measured distance to the byte-match
tools/standalone_report.py drives sots-engine's sots_turn over each
consecutive-turn save pair and diffs the result against the game's own
post-turn save with state_checksum.py, which localises to named leaves and
proves its own coverage by re-serialisation.

  turn1-state -> turn2-state   baseline 209 diverging, after 204, closed 5
  turn2-state -> turn3-state   baseline 108 diverging, after 103, closed 5
  regressed 0 on both

`regressed` is reported next to `closed` and never netted off. It earned its
place immediately: committing the phase-31 player-status restore turned two
agreeing leaves into disagreeing ones, because the phase writes 1 and the file
carries 4.

The stable-system stand-in feeding the colony pass is a labelled hypothesis and
it survived a changed workload -- the same 3 ntdev leaves closed on both pairs,
six agreements, zero disagreements.

Two things deliberately NOT implemented: the TShn/ltis counters (18 leaves, a
`+1` would close them, but "+1 across one observed turn" is a hypothesis, not a
reading), and the RNG state write-back (an advanced-but-incomplete generator is
wrong in a different way from an untouched one).

dashboard.py gains section 6, reading verify/results/standalone/status.json:
phases modelled/committed per driver, baseline vs after, closed vs regressed,
the subsystem breakdown of what still differs, and the RNG gap. Sections 6-8
renumbered to 7-9; the delta footer tracks the two new counts.
DASHBOARD_README.md documents every number.

findings/control-flow/standalone-scaffold.md has the ranked blocker list.
2026-09-08 10:35:56 -04:00
alex
66fdf0f2df Z: P9 held on every clause -- the node-line draw landed on turn 64 and cost one word
Predicted at turn 34 with the run in flight, from min_life falling by exactly
1 per turn: the first phase-11 draw on turn 64, exactly 1 word, tail total 1,
bracket = ProcessTurn + 1. The game was played to turn 64 and every clause
held. predict_words, computed before the original ran, said 1 and the
measurement said 1 -- a real check of the model, against 63 preceding turns
where 0 matched 0 and checked nothing.

So the defect lane K warned about is no longer latent: on that turn a
reimplementation modelling ProcessTurn perfectly would have written an autosave
one word out of step.

And the instrument's thinnest part ran live on the same turn -- ProcessTurn
crossed a block boundary (left 11 -> 615, one twist, 20 words) and the bracket
still reconciled to residual 0.
2026-09-08 10:24:05 -04:00
alex
9fa1ee2600 Z: the first instrumented battle in this campaign costs zero RNG words
A Von Neumann encounter at Gallandro on turn 54 gave the workload the finding
said did not exist. Auto-resolved, with P10 committed before the click.

P10 predicted a non-zero tail cost and was wrong: res_no_battle flipped to 0
for the first time in 55 turns, the fleet was destroyed, and the generator
moved by zero. The bracket residual stayed 0, so combat proper drew nothing
either -- all 22 words were inside ProcessTurn, exactly as on a peaceful turn.

That is the strong form of lane J's static reading, and it means a
reimplementation can model a turn's RNG while modelling nothing about combat.
One auto-resolved encounter against an NPC is not combat in general, and 10.2
says so at length.
2026-09-08 10:14:53 -04:00
10 changed files with 2686 additions and 40 deletions

View file

@ -1,6 +1,6 @@
# SotS RE campaign — coverage dashboard
Generated 2026-09-08 14:07 UTC · `sots-re` @ c15a45d,2026-09-08 · `sots-engine` @ bcf4297,2026-09-08 (115 commits) · regenerate with `tools/dashboard.py`
Generated 2026-09-08 14:35 UTC · `sots-re` @ 66fdf0f,2026-09-08 · `sots-engine` @ bcf4297,2026-09-08 (115 commits) · regenerate with `tools/dashboard.py`
> **North star:** A functional reimplementation of the engine — behavior-equivalent, NOT byte-for-byte
@ -73,7 +73,39 @@ Board `engine:` rows: verified **18**, mapped 0, in flight 0 (of 18) — verifie
| P2-M3 Mars brace-block parser | ⬜ backlog | 0% | Mars::Script pull tokenizer (Open 0x008cd7d0, ReadToken 0x008cd2f0, Next 0x008cd3e0, SkipB |
| P2-M4 gobio VFS read | ⬜ backlog | 0% | choke point: bool __cdecl gobio::ReadFile(const char*, IBuffer**) 0x008d5140; FileSystemSe |
## 6. Verification ledger
## 6. Standalone (`src/app`) — distance to the byte-match
Turn-driver phases: **14/44** modelled (7 committed) `[███░░░░░░░] 32%`
| | verified | implemented | partial | blocked | stub |
|---|---:|---:|---:|---:|---:|
| turn drivers (44) | 0 | 2 | 5 | 7 | 30 |
| post-combat tail (37) | 0 | 1 | 0 | 1 | 35 |
Reference pair `turn1-state.sav` → `turn2-state.sav`, leaves localised by `state_checksum.py` (coverage proved by re-serialisation):
- baseline (a standalone that does nothing): **209** leaves diverge
- after one standalone turn: **204** leaves diverge — closed 5, regressed 0
- byte match: ❌ not yet `[░░░░░░░░░░] 2%`
Where the remaining divergence lives:
| Subsystem | Leaves |
|---|---:|
| `/Sim/players` | 82 |
| `/Sim/systems` | 80 |
| `/Sim/turnstats` | 24 |
| `/Sim/SvSctOb` | 8 |
| `/Sim/DesignIDs[]` | 1 |
| `/Sim/FleetIDs[]` | 1 |
| `/Sim/ModCount` | 1 |
| `/Sim/NMnx` | 1 |
Generator: 0 word(s) modelled per turn; unattributed per turn: 18-20 (lane Z, in flight). A byte-match is impossible until that closes — the generator state is saved state.
Detail: `verify/results/standalone/report.txt`.
## 7. Verification ledger
- ✅ Saves strict: 11/11 (strict exit 0, 0 errors, 0 warnings)
- ✅ Design rules: 127/127
@ -82,7 +114,7 @@ Board `engine:` rows: verified **18**, mapped 0, in flight 0 (of 18) — verifie
- ✅ M0 evidence present (`verify/results/shim/m0.log`)
- ✅ Determinism oracle: verified
## 7. Open questions
## 8. Open questions
Open **26** · resolved/parked 11 · backlog items: Now 4, Next 3, Later 2, Breadth queue 7, Parked 1, From the RE how-to 4, Behavioral slice 4
@ -94,12 +126,13 @@ Most recent open:
- SAVE_FORMAT tag corrections (fix Python reader + spec) — real on-disk tags: `otnF` (not `ontF`) in…
- Not traced end-to-end — `Species/_NPC/weapons/*.weapon` loading and the `.effect` dictionary entry…
## 8. Delta since previous dashboard
## 9. Delta since previous dashboard
- verified targets: 110 → 119 (+9) · mapped-or-better: 136 → 145 (+9)
- engine LOC: 35,544 → 35,582 (+38) · test files: 94 → 94 (+0) · checks: 3,155 → 3,170 (+15)
- addresses verified: 734 → 777 (+43) · recovered layouts: 384 → 384 (+0) · open questions: 26 → 26 (+0)
- verified targets: 119 → 119 (+0) · mapped-or-better: 145 → 145 (+0)
- engine LOC: 35,582 → 35,582 (+0) · test files: 94 → 94 (+0) · checks: 3,170 → 3,170 (+0)
- addresses verified: 777 → 777 (+0) · recovered layouts: 384 → 384 (+0) · open questions: 26 → 26 (+0)
- standalone leaves closed: n/a · leaves still diverging: n/a
---
warnings: board.md: unknown types subsystems; mars-rng.md: no oracle total row parsed; mars-stream.md: no oracle total row parsed; mars-vfs.md: no oracle total row parsed
<!-- dashboard-metrics {"verified": 119, "mapped_plus": 145, "targets": 172, "loc": 35582, "tests": 94, "checks": 3170, "addr_verified": 777, "addr_total": 801, "layouts": 384, "open_q": 26} -->
<!-- dashboard-metrics {"verified": 119, "mapped_plus": 145, "targets": 172, "loc": 35582, "tests": 94, "checks": 3170, "addr_verified": 777, "addr_total": 801, "layouts": 384, "open_q": 26, "sa_closed": 5, "sa_left": 204} -->

View file

@ -177,3 +177,9 @@ Status flow: `backlog → in-progress → mapped → verified` (or `blocked`).
| RNG entry points: SEVEN, not three | objects | verified | high | 100% | 2026-09-08 | Lane I: NextFloat, NextInt, Chance, plus NextUInt, FloatRange 0x0047d8a0 (1 word, NARROWS TWICE, in ProcessTurn's closure at depth 3 with two call sites), IntRangeBell 0x008e6d80 (triangular, >=2 words) and GaussianRange 0x008e6e30 (2 words PER ATTEMPT, UNBOUNDED, both draws inlined). AND IT SCALES BY 2^-32 WHERE NextFloat SCALES BY 1/(2^32-1) - TWO DIVISORS IN ONE IMAGE. GaussianRange documented but deliberately NOT modelled |
| RNG residual: 18-20 words STILL UNEXPLAINED | verify | backlog | — | 0% | 2026-09-08 | THE HONEST HEADLINE. Lane J predicted the two inlined functions would explain lane Z's per-turn gap. **NOT CONFIRMED.** Both new sources are gated and neither has been measured: 0x007aa240 contributes 0..(|contacts|x|detectors|) and whether its +0xfc gate passed is not knowable statically; ProbabilisticJump's second word is 0 unless a type-5 waypoint fails its arrival test. RESIDUAL: 18-20 words, essentially ALL of it. WHAT IS NOW PROVABLE IS THE NEGATIVE: there is NO TWENTY-THIRD MECHANISM - the complete draw-site inventory of the ProcessTurn closure (1,426 functions) is 22 sites (21 entry-point calls + 1 inlined), so the 18-20 words are distributed among exactly those. THE SEARCH SPACE CLOSES; THE COUNT DOES NOT. One extra bracket on EncounterDetect_ProcessTeamRecord with |contacts|/|detectors| in the argument record turns the formula into a one-turn test |
| CORRECTION: no Mars MT19937 variant | objects | verified | high | 100% | 2026-09-08 | Lane I correcting ITSELF (rule 11). It had written into Ghidra that 0xff3a58ad/0xffffdf8c are a Mars variant of MT19937. THEY ARE NOT - they are the textbook masks applied BEFORE the shift: (y & 0xff3a58ad) << 7 == (y << 7) & 0x9d2c5680, verified over 200k words. `mars::rng` in sots-engine was NEVER WRONG. Corrected in place in Ghidra and in the fragment. NOTE FOR FUTURE SCANS: a scan for the TEXTBOOK constants finds NOTHING in this image, which is exactly why rule 16's scan must use these pre-shift values |
| STANDALONE SCAFFOLD RUNS | engine | verified | high | 100% | 2026-09-08 | THE NORTH STAR MADE CONCRETE. `sots_turn SAVE --roundtrip --phases --out POST.sav --metric m.json`: loads a real .sav through mars::stream, PROVES THE FOUNDATION FIRST (re-serialises the untouched parse and byte-compares against the inflated stream; if that fails the run STOPS), walks the published phase order of ALL THREE drivers (host steps, StrategyServer::ProcessTurn's 32 with ServerPlayer::ProcessTurn's 12 nested at phase 13, then OnAllCombatDone_Tail's 37), PRINTS EVERY PHASE IT DOES NOT RUN, writes post-turn state back through write_save + gzip, emits a completion metric. Driven over all 11 saves: 11/11 load, run, re-serialise and re-read cleanly. Integrator-verified on CT111 with real saves: ctest 38/38, and the trace output labels hypotheses INLINE (e.g. "3 judged stable by the owned/not-abandoned stand-in (HYPOTHESIS -- the original asks a callee)") |
| standalone divergence vs the oracle | verify | mapped | high | 100% | 2026-09-08 | Leaves localised by state_checksum.py, coverage PROVED on every save in every comparison. turn1->turn2: baseline 209 diverge -> 204, CLOSED 5, REGRESSED 0. turn2->turn3: 108 -> 103, closed 5, regressed 0. Closed the same five on both pairs: /Summary/Turn, /Sim/Frame, and ntdev on Gamma Cephei / Ke'Dolarra / Koa'Vo. Remaining 204: /Sim/players 82, /Sim/systems 80, /Sim/turnstats 24, /Sim/SvSctOb 8, plus 10 singletons (ModCount, RNG, Checksum, NMnx, cmbtid, four id lists, NumFlts). BYTE MATCH: NO, AND IT CANNOT BE YET - 0 of the 18-20 per-turn generator words are modelled and the generator is saved state |
| standalone completion metric | meta | verified | high | 100% | 2026-09-08 | verify/results/standalone/status.json, read by tools/dashboard.py's new section 6 (regenerate with tools/standalone_report.py). **14 of 44 turn-driver phases modelled, 7 committed; 2 of 37 tail phases. VERIFIED 0 - DELIBERATELY**: in that table `verified` means "compared against the live game", lane S2 held no VM, and `app_catalog` FAILS THE BUILD if that ever drifts upward silently. Implemented+committed (9): H00 BeginProcessTurn, H01 SaveWriterInvariants, S00 ModCount bump, S11 SystemTurn (tuning-free subset), P07 clear timed-research accumulators, P08 rebellion output decay, P09 timed research bonuses (last->first, ORDER LOAD-BEARING), P10 ResearchRollPending consume, T00 tail ModCount bump. Blocked - evaluated, reported, NOT written (8): P01 P02 P03 P05 P06 all behind the unresolved population->base-output term, plus P11, S31, T31. Stub - named no-ops visible in output: 65 |
| THE FINDING THAT SHAPED THE DESIGN: report regressed, never net it off | meta | verified | high | 100% | 2026-09-08 | Lane S2 first implemented AND COMMITTED S31's player-status restore. The comparison tool immediately reported TWO REGRESSED LEAVES on turn2->turn3: two Player.Status words that AGREED with the oracle before the turn and DISAGREED after. The phase writes 1, the file carries 4, a load resets to 0 - so a writer between phase 31 and the autosave is unaccounted. S31 is now blocked, and `regressed` is reported NEXT TO `closed` in every run, NEVER NETTED OFF. Consequent design rule: a phase whose FORMULA we hold but whose INPUTS we do not is EVALUATED AND REPORTED, NOT WRITTEN, unless --commit-blocked. Same for the generator (--commit-rng) |
| standalone: two things deliberately NOT implemented | verify | backlog | — | 0% | 2026-09-08 | (1) TShn/ltis - 18 of the remaining 204 leaves, moving 1->2 on 8-10 systems on BOTH pairs. A `+1` closes them in ten lines. NOTHING NAMES THEIR WRITER, so that is a hypothesis not a reading; named in the docs as the cheapest measured target. (2) RNG write-back - an advanced-but-incomplete state is WRONG DIFFERENTLY from an untouched one. One hypothesis IS under test and survived a changed workload: `stable = owned && !abandoned && !destroyed` (the original asks a callee) closed the same 3 ntdev leaves on both pairs - six agreements, zero disagreements - and is labelled a hypothesis in code, log AND docs |
| RANKED blockers to the byte-match | meta | mapped | high | 100% | 2026-09-08 | Lane S2's ranking, which is now the project's critical path: (1) THE RNG LEDGER - lane Z; nothing in src/app can close it. (2) The population->base-output term - ONE FORMULA GATING 5 OF THE 44 PHASES. (3) The 37-phase post-combat tail, which is the driver THE AUTOSAVE IS WRITTEN FROM. (4) The `nve` visibility record (32 leaves, one mechanism x 8). (5) The event pipeline. (6) Summary.Checksum. (7) The Player.Status writer. (8) ModCount - the real turn advances it 12-44 times from writers spread across BOTH drivers |

View file

@ -0,0 +1,124 @@
# The standalone, and the measured distance to the byte-match
Lane S2, 2026-09-08. Engine branch `wip/standalone`; full documentation in
`sots-engine/docs/S-standalone.md`. This note records what the lane measured, what it declined
to implement, and the two numbers the campaign should track from here.
---
## 0. The headline
`sots_turn` loads a real save through the engine's own reader, walks the **published phase
order of all three turn drivers**, runs what we hold, prints what we do not, and writes a save
through the engine's own writer.
```
turn1-state.sav -> turn2-state.sav (a real End Turn)
baseline (a standalone that does nothing) 209 leaves diverge
after one standalone turn 204 leaves diverge
closed 5, regressed 0
turn2-state.sav -> turn3-state.sav
baseline 108 -> 103, closed 5, regressed 0
```
Phases: **14 of 44** turn-driver phases modelled, **7** committing anything; **2 of 37** of the
post-combat tail. Generator words modelled per turn: **0** of the 18–20 consumed.
Leaves are `verify/state-checksum/state_checksum.py`'s named leaves; every one of the six
saves involved reported `coverage: PROVED` on the same run, so the diff cannot be hiding
anything.
Regenerate with `tools/standalone_report.py`; outputs land in `verify/results/standalone/`
and the dashboard's new section 6 reads `status.json`.
## 1. What the scaffold is for
Everything the milestone needs already existed in pieces — a save reader with 100 % named
coverage, a verified budget roll-up, a verified research slice, a mechanism-verified movement
model, and byte-for-byte maps of both turn drivers. What did not exist was **a place to put
them and a number that says how far they get**. That is the whole of this lane.
The phase catalog (`src/app/phase_catalog.cpp`) is the roadmap: all 32 + 12 + 37 phases, each
carrying a status and a note. Running `sots_turn --phases` prints the turn, and every
unimplemented phase prints itself. There is no way for a phase to be silently absent, and
`app_catalog` fails the build if a table develops a gap or if anything claims to be `verified`
(which in that table means "compared against the live game" — lane S2 held no VM).
## 2. The finding that justifies the design: a committed phase can make things worse
The first version of this lane implemented `StrategyServer::ProcessTurn` phase 31's
player-status restore — `Status = 1` — and committed it. The comparison tool immediately
reported **two regressed leaves** on the `turn2 → turn3` pair: two `Player.Status` words that
**agreed** with the oracle before the turn and disagreed after it.
The phase writes 1. The post-turn file carries 4. A load resets it to 0. So a writer between
phase 31 and the autosave is unaccounted, and the input save happened to already carry the
right answer.
This is the same shape as the campaign's oldest lesson in a new place: running more code is
not the same as knowing more. The standalone therefore separates **modelled** from
**committed**, and a phase whose inputs are not modelled is evaluated, reported and *not
written* unless `--commit-blocked` is passed. `regressed` is reported next to `closed` in every
run, never netted off.
Five phases are blocked behind one unresolved formula (below); `S31` and `T31` are blocked
behind their own; `P11` is blocked behind the event-text table.
## 3. Two things this lane declined to implement
**The `TShn` / `ltis` counters.** 18 leaves of the remaining 204 are `TShn` and `ltis` moving
`1 -> 2` on 8–10 systems, on both turn pairs. They look exactly like per-turn counters and a
`+1` would close 18 leaves in ten lines of code. Nothing in the campaign names their writer, so
"+1 per turn across one observed turn" is a hypothesis, not a reading, and rule 6 says label it
as one. They are the **cheapest measured target on the board** and they are named here so the
next lane can close them properly rather than plausibly.
**The RNG state.** The generator advances during a real turn; the standalone leaves the blob
byte-identical by default. An advanced-but-incomplete state is wrong in a different way from an
untouched one, and the untouched one at least reports the truth. `--commit-rng` is there for
the day lane Z's ledger closes.
## 4. One hypothesis under test, and it survived a changed workload
`ProcessColonyTurn` takes `stable` as an input; in the original it is a callee's verdict. The
standalone stands in `owned && !abandoned && !destroyed`, labelled a hypothesis in the code and
printed as one in the run log.
It drives `ntdev`, which is a named leaf, so it is falsifiable. On `turn1 → turn2` it judged 3
of 28 systems stable and closed exactly the 3 `ntdev` leaves the oracle moved. On
`turn2 → turn3` — a different turn, a different set of orders — it closed the same 3 again.
**Six agreements, zero disagreements, across two workloads.** Not proof; recorded as such.
## 5. What is now measurably in the way
Ordered by what must be solved, not by size.
| # | blocker | cost in leaves on the reference pair | who can close it |
|---|---|---:|---|
| 1 | **the RNG ledger** — 18–20 words/turn, none modelled, and the generator is saved state | 1 leaf, and it makes a byte-match *arithmetically impossible* | lane Z (in flight); nothing in `src/app` |
| 2 | **the population → base-output term** — one unresolved formula that blocks `P01 P02 P05 P06 T31` | ~10 directly (`Sav`, `BnkPr`, `BnkEl` on 4 players), and it gates 5 of the 44 phases | a formula lane against the live game |
| 3 | **the post-combat tail**, 37 phases, none implemented, and the driver the autosave is written from | 24 (`turnstats`) + the bankruptcy limits + observed designs + player reports | its own milestone |
| 4 | **the per-player system-visibility record** (`nve`) | 32 — one mechanism, eight repetitions | spine phase 24 or tail phase 21 |
| 5 | **the event pipeline** — buckets, ids and localised text | ~20 across the players | needs the string table |
| 6 | **`Summary.Checksum`** — algorithm unknown | 1, and it is the last leaf to fall | — |
| 7 | **the `Player.Status` writer** — phase writes 1, file carries 4 | 4 | small, self-contained |
| 8 | **`ModCount`** — advances 12–44 times a turn from writers across both drivers; we model 2 | 1 | falls out of implementing the other phases |
## 6. Gates run, separately
```
tools/clean_room_check.sh -> clean-room check: OK (with src/app + tests/app staged)
ctest --preset host -> 100% tests passed out of 38 (was 36; +app_catalog, +app_turn)
app_turn with SOTS_SAVES_DIR set -> 11 saves driven, 0 failures
```
`src/shim/` was not touched, so no CT111 cross-build was required.
## 7. What the corpus cannot answer
The two pairs above are the only true End-Turn transitions we hold. The other nine saves are
single states: `app_turn` drives a turn over each of them and asserts the file survives, but
there is no oracle to diff against. A third and fourth consecutive-turn pair — especially on a
Zuul game, where the research and node-travel paths differ — would make every number in §0
sturdier for the cost of two End Turns on VM140.

View file

@ -305,19 +305,9 @@ would have said so by construction rather than by anyone noticing.
## 8. What is not settled, listed as loudly as the results
* **The combat resolver has still never run under an instrument.** Every encounter this workload produced
had the no-battle flag set, so `ApplyEncounterResult` was a no-op every time; its measured 0 words says
nothing whatever about combat's RNG cost, and the residual-0 result above holds only for turns with no
battle.
Lane J read all 7,641 bytes of it in parallel with this run (`combat-resolver.md`; and note **7,641**, not
the 7,499 Ghidra reports — method rule 17). Its conclusion pairs with this one exactly: **the resolver has
no unconditional draw.** All three sites in its subtree are conditional — a node-cannon `NextInt`, an
inlined `NextFloat` per back-engineering candidate, and a `NextInt` per successful roll of that. So lane
J's cheap first prediction is directly testable with this instrument: **a plain fleet battle with no node
cannon and no salvage should cost the same 18–22 words as a peaceful turn.** That is the next run this
hook family should do, and it needs a workload nobody has built yet: a save where two hostile fleets
actually meet.
* ~~The combat resolver has never run under an instrument.~~ **It has now, once — see §10. It cost 0
words.** What remains unsettled is everything a single auto-resolved encounter cannot speak for; §10.2
lists it.
* **Node-line expiry did not fire.** See §9 for the quantified distance rather than an absence.
* **A turn with a genuinely empty encounter vector was not observed** (§4). Both saves produce exactly one
sighting encounter on every turn. The tail-runs-every-turn claim is settled; the no-encounters variant is
@ -396,17 +386,105 @@ Three things follow, none of which was knowable before:
3. It explains why the campaign never noticed: no save in the corpus is within 40 turns of a decay event,
and the ones that could get there are the newest saves in it.
**Phase 11's draw therefore remains a path no save exercises — a hypothesis, and labelled one.** What *is*
now instruction-verified is the predicate that decides it (§6.1), and what is measured is the distance.
### 9.1 The prediction, and the turn it came true
### 9.1 The model is stated but not yet tested
Because `min_life` was falling by exactly 1 per turn — 43, 42, 41 … 30 at turn 34, with the traffic term
contributing nothing on this map — a numeric prediction became possible, and it was committed to
`sots-engine/docs/Z-tail-rng.md` §6 at turn 34 with the run still in flight:
`predict_words` is computed at hook entry, before the original runs, and recorded on every node-line-decay
record. It read **0** on every call and the measured delta was **0** on every call. That agreement is worth
exactly nothing as a test of the model — it is the "0 diverged while comparing nothing" shape this campaign
has already paid for — and it is reported that way rather than as a green tick.
> **P9. The first phase-11 draw happens on turn 64, and costs exactly 1 word.** Node-line decay records
> `np_min_life = 0`, `predict_words = 1`, and a measured `rng` delta of 1; the tail's total becomes 1
> instead of 0; the bracket total becomes `ProcessTurn + 1`.
Compare mode was **not** run on this hook for the same reason: with a prediction of 0 and a measurement of 0,
`ours` would advance the scratch generator by nothing, diff clean, and prove only that the harness works.
The model becomes checkable the first time a line expires, and the descriptor is ready for that day.
The game was played to turn 64. **Every clause held.**
| turn | node paths | mortal | min life | ≤5 | expired | `predict_words` | node-decay words | **tail words** | `ProcessTurn` | bracket | residual |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 62 | 64 | 13 | 2 | 1 | 0 | 0 | 0 | 0 | 18 | 18 | 0 |
| 63 | 64 | 13 | 1 | 1 | 0 | 0 | 0 | 0 | 18 | 18 | 0 |
| **64** | 64 | 13 | — | 0 | **1** | **1** | **1** | **1** | 20 | **21** | **0** |
**That is the first non-zero tail cost this campaign has ever recorded**, and it is exactly the draw lane K
found by reading 0x007ae095. The defect lane K warned about is no longer latent, no longer inferred and no
longer a hypothesis: **on turn 64 of this save, a reimplementation that models `ProcessTurn` perfectly and
stops would have written an autosave one generator word out of step, and every subsequent turn would
diverge.**
`predict_words` is computed at hook entry, *before* the original runs, from the same `RemainingLife`
predicate transcribed in §6.1. It said 1; the measurement said 1. That is a real check of the model — as
opposed to the 63 preceding turns, where it said 0 and the measurement said 0, which checked nothing and was
reported that way.
### 9.2 The twist path ran, live, on the same turn
The instrument's thinnest part (§8: "the block-chain machinery has never run live") was exercised on this
very turn. `ProcessTurn` entered with `left = 11` and left with `left = 615` — it **crossed a block
boundary**, the generator twisted, and the ledger reported `11 + (624 − 615) = 20` words. Sixty-three turns
of measurements had all sat inside a single block, so every earlier word count reduced to a subtraction; this
one did not, and the bracket still reconciled to a residual of 0.
### 9.3 What is still not settled about node lines
* **One expiry, on one map.** 51 of the 64 lines are permanent; the 13 mortal ones are Zuul-dug. A
non-Zuul game may never produce a mortal line at all.
* **The `0x20000`-fleet gate has never been exercised**, because it only matters when the roll *succeeds*
and no fleet was riding this line. Whether the roll succeeded here is not visible in a word count — the
draw costs 1 either way, which is the whole point of §6's correction.
* **Two expiries on one turn has never been observed** (`np_within5` read 1, never 2), so "one word per
expired line" is confirmed for *one* line and extrapolated for two.
---
## 10. A real battle, measured — and it costs nothing
The turn-54 End Turn of the long run stopped on an **Encounter at Gallandro**: the player's five ships
(3 DE Colonizer, 2 DE Armor) against a **Von Neumann**. That is the workload §8 said did not exist and lane
J's `combat-resolver.md` asked for — every encounter in 54 turns until this one had `res->+0x4` set, making
`ApplyEncounterResult` a whole-function no-op. **Auto Resolve** was chosen (the dialog's four options are
Fight Manually / Auto Resolve / Fight Manually If Opponent Does / Retreat), and the prediction was committed
to `sots-engine/docs/Z-tail-rng.md` §7 with the dialog still on screen and unclicked.
| turn | `res_no_battle` | `ApplyEncounterResult` words | tail words | `ProcessTurn` words | bracket | residual |
|---|---|---|---|---|---|---|
| 52 | 1 | 0 | 0 | 18 | 18 | 0 |
| 53 | 1 | 0 | 0 | 18 | 18 | 0 |
| 54 | 1 | 0 | 0 | 16 | 16 | 0 |
| **55** | **0** | **0** | **0** | 22 | **22** | **0** |
**P10 predicted a non-zero tail cost and was wrong.** The first battle this campaign has ever instrumented
moved the strategic generator by **zero words**, and the bracket residual stayed 0 — so combat proper
(`RunCombatRound` / the combat server, which run between `ProcessTurn` and the tail and are hooked by
nobody) drew nothing either. Every one of the turn's 22 words was inside `StrategyServer::ProcessTurn`, just
as on a peaceful turn.
That is the *strong* form of lane J's reading. Lane J established from the instruction stream that the
resolver has **no unconditional draw** — its three sites are a node-cannon `NextInt`, an inlined `NextFloat`
per back-engineering candidate, and a `NextInt` per successful roll of that. This run shows that on an
ordinary encounter **none of the three fires**, and lane J's own cheap prediction — *a plain fleet battle
should cost the same as a peaceful turn* — holds exactly.
### 10.1 Why this matters to the standalone
A reimplementation that models a strategic turn's RNG and **nothing about combat** reproduces the generator
correctly through a battle. Combat's effect on the save is entirely in the state it writes, not in the
generator it advances. That is a much cheaper milestone than "read the 7,641-byte resolver first", and it
was not knowable before this run: the honest prior was lane K's "draw counts are entirely combat-dependent
and unknown".
### 10.2 What one battle does not settle — and it is a lot
* **One encounter, auto-resolved.** `Auto Resolve` may not take the same path as a manually fought battle;
the tactical engine has its own `Mars::CombatSim` generator at `sim+0x108` (§7) which nothing here
watches. A manually fought battle is a different experiment and has still never been run.
* **The opponent was a Von Neumann**, an NPC pseudo-player, not a rival empire's war fleet. No node cannon
was present, so R1 could not fire; whether R2's salvage roll was skipped because no candidate had a
non-zero salvage slot, or because the arm was not reached at all, is not distinguishable from a word
count of 0.
* **A cost of 0 is the easiest number to produce by accident.** It is exactly what a hook that compared
nothing would report. The reasons to believe it here are that the same hook reported 18–22 for
`ProcessTurn` on the same turn, that `res_no_battle` flipped to 0 for the first time in 55 turns on
exactly the turn the battle happened, and that the player's fleet was destroyed — the battle demonstrably
occurred. It is still one observation.
* **No `EVENT_*` or state-side check was made.** These hooks declare the generator and nothing else, so
this says the battle was RNG-free and says nothing about whether it was *computed* correctly.

View file

@ -82,13 +82,34 @@ verified / all engine rows. In flight = `in-progress`.
Board rows whose Target starts with `P2-M`. Glyphs: ✅ verified or mapped, 🔄 in-progress,
⬜ backlog, ⛔ blocked. Notes are truncated to 90 characters.
## 6. Verification ledger
## 6. Standalone progress
Source: `verify/results/standalone/status.json`, written by `tools/standalone_report.py` (which
drives `sots-engine`'s `sots_turn` over each consecutive-turn save pair and diffs the result
against the game's own post-turn save with `verify/state-checksum/state_checksum.py`). The file
is optional: absent, unreadable, or carrying an unexpected `schema` -> the section renders
"Not measured" plus a footer warning, and the delta row shows `n/a`.
| Number | Source |
|---|---|
| turn-driver phases modelled / total | `phases.modelled` / `phases.total`. **Total is 44 = the 32 phases of `StrategyServer::ProcessTurn` + the 12 of `ServerPlayer::ProcessTurn`.** The host steps around the drivers are deliberately excluded from the denominator; the 37-phase post-combat tail is a separate row |
| committed | phases whose result is actually written to the save. A `blocked` phase is *modelled* (it runs and reports) but not *committed*, so `committed <= modelled` always |
| status breakdown | `verified` (compared against the live game) / `implemented` / `partial` / `blocked` (formula held, an input is not) / `stub` (named no-op) |
| baseline / after / closed / regressed | `reference.*`. Baseline = leaves that differ between the input save and the oracle, i.e. the distance a standalone that does nothing has to travel. `regressed` is reported next to `closed` and never netted off: a leaf that agreed before the turn and disagrees after it is a phase doing damage |
| byte match | `reference.byteMatch` — the milestone itself. The progress bar next to it is closed/baseline, not a claim about how much is left |
| subsystem table | `reference.subsystems`, the first 8 by leaf count |
| generator | `rng.wordsModelled` and the standing `rng.wordsPerTurnUnattributed` string |
The reference pair is the first entry of `PAIRS` in `standalone_report.py`; every pair it ran
is in `status.json` under `pairs`, and the readable form is `verify/results/standalone/report.txt`.
## 7. Verification ledger
One line per evidence source: saves strict and design rules (as in section 3), each parsed
oracle, `verify/harness/compare/` directory present, `verify/results/shim/m0.log` present,
and the board status of the `determinism oracle` row.
## 7. Open questions
## 8. Open questions
`campaign/open-questions.md`: every bullet beginning `- **`. A bullet is *resolved/parked* if
its bold text starts with `RESOLVED`, `Resolved` or `(parked)`; everything else is open. "Most
@ -96,12 +117,12 @@ recent" = the last five open bullets in file order (the file is append-ordered),
first 100 characters. Backlog counts are numbered or bulleted items under each `## ` heading
of `campaign/backlog.md`.
## 8. Delta
## 9. Delta
The previous `DASHBOARD.md` carries a machine-readable footer comment
`<!-- dashboard-metrics {...} -->`. It is read before the file is overwritten; the section
shows old → new (±) for verified targets, mapped-or-better, engine LOC, test files, checks,
addresses verified, recovered layouts and open questions. A dashboard without that comment
addresses verified, recovered layouts, open questions, and the standalone's closed / still-diverging leaf counts. A dashboard without that comment
(or the first run) reports "first run".
## Adding a number

View file

@ -225,6 +225,26 @@ def parse_design_rules():
return None, None
# ---------------------------------------------------------------- standalone
def parse_standalone(path):
"""verify/results/standalone/status.json, written by tools/standalone_report.py.
Absent or malformed -> None, so the section renders as 'not measured' rather than
breaking the dashboard."""
if not os.path.exists(path):
return None
try:
with open(path, encoding="utf-8") as f:
d = json.load(f)
except (OSError, ValueError) as e:
warn(f"standalone status.json unreadable: {e}")
return None
if d.get("schema") != "sots-standalone-status/1":
warn(f"standalone status.json: unexpected schema {d.get('schema')!r}")
return None
return d
# ---------------------------------------------------------------- engine
SRC_EXT = (".cpp", ".h", ".c")
@ -438,6 +458,53 @@ def render(engine, prev):
warn("board.md: no P2-M* rows")
L.append("")
# 6 standalone progress -- the north star, measured
sa = parse_standalone(rp("verify", "results", "standalone", "status.json"))
L += ["## 6. Standalone (`src/app`) — distance to the byte-match", ""]
sa_closed = sa_left = 0
if not sa:
L += ["Not measured. Build `sots-engine`'s host preset and run "
"`tools/standalone_report.py`.", ""]
else:
ph, tp = sa.get("phases") or {}, sa.get("tailPhases") or {}
ref = sa.get("reference") or {}
sa_closed = ref.get("closed") or 0
sa_left = ref.get("divergingAfterTurn") or 0
base = ref.get("baselineDiverging") or 0
if ph:
L += [f"Turn-driver phases: **{ph['modelled']}/{ph['total']}** modelled "
f"({ph['committed']} committed) {bar(ph['modelled'], ph['total'])}",
"",
f"| | verified | implemented | partial | blocked | stub |",
"|---|---:|---:|---:|---:|---:|",
f"| turn drivers (44) | {ph['verified']} | {ph['implemented']} | "
f"{ph['partial']} | {ph['blocked']} | {ph['stub']} |"]
if tp:
L.append(f"| post-combat tail ({tp['total']}) | {tp['verified']} | "
f"{tp['implemented']} | {tp['partial']} | {tp['blocked']} | "
f"{tp['stub']} |")
L.append("")
L += [f"Reference pair `{ref.get('input')}` → `{ref.get('oracle')}`, leaves localised by "
"`state_checksum.py` (coverage proved by re-serialisation):", "",
f"- baseline (a standalone that does nothing): **{base}** leaves diverge",
f"- after one standalone turn: **{sa_left}** leaves diverge "
f"— closed {sa_closed}, regressed {ref.get('regressed')}",
f"- byte match: {'✅ YES' if ref.get('byteMatch') else '❌ not yet'} "
f"{bar(sa_closed, base) if base else ''}", ""]
subs = ref.get("subsystems") or {}
if subs:
L += ["Where the remaining divergence lives:", "",
"| Subsystem | Leaves |", "|---|---:|"]
for k, v in list(subs.items())[:8]:
L.append(f"| `{k}` | {v} |")
L.append("")
rng = sa.get("rng") or {}
L += [f"Generator: {rng.get('wordsModelled')} word(s) modelled per turn; "
f"unattributed per turn: {rng.get('wordsPerTurnUnattributed')}. "
"A byte-match is impossible until that closes — the generator state is saved "
"state.", "",
"Detail: `verify/results/standalone/report.txt`.", ""]
# 7 verification ledger
def status_of(target_re):
for r in rows:
@ -446,7 +513,7 @@ def render(engine, prev):
return "not on board"
m0 = os.path.exists(rp("verify", "results", "shim", "m0.log"))
harness = os.path.isdir(rp("verify", "harness", "compare"))
L += ["## 6. Verification ledger", "",
L += ["## 7. Verification ledger", "",
f"- {'✅' if s_ok else '❌'} Saves strict: {nsaves}/{nsaves} ({s_detail})",
f"- {'✅' if d_ok == d_tot and d_ok else '❌'} Design rules: {d_ok}/{d_tot}",
"- " + (" · ".join(f"{'✅' if a == f else '❌'} oracle {nm} {a}/{f}" for nm, a, f in oracles) if oracles else "❌ oracle parsers: none parsed"),
@ -457,7 +524,7 @@ def render(engine, prev):
# 8 open questions
oq, closed = parse_questions(rp("campaign", "open-questions.md"))
bl = parse_backlog(rp("campaign", "backlog.md"))
L += ["## 7. Open questions", "",
L += ["## 8. Open questions", "",
f"Open **{len(oq)}** · resolved/parked {len(closed)} · backlog items: " +
(", ".join(f"{k} {v}" for k, v in bl.items()) if bl else "?"), "", "Most recent open:", ""]
for t in oq[-5:][::-1]:
@ -466,8 +533,9 @@ def render(engine, prev):
# 9 delta
cur = dict(verified=by["verified"], mapped_plus=mapped_plus, targets=total, loc=tot_loc, tests=tot_tf,
checks=tot_ck, addr_verified=a_ver, addr_total=a_total, layouts=len(layouts), open_q=len(oq))
L += ["## 8. Delta since previous dashboard", ""]
checks=tot_ck, addr_verified=a_ver, addr_total=a_total, layouts=len(layouts), open_q=len(oq),
sa_closed=sa_closed, sa_left=sa_left)
L += ["## 9. Delta since previous dashboard", ""]
if prev:
def d(k):
if k not in prev:
@ -476,7 +544,8 @@ def render(engine, prev):
return f"{prev[k]:,} → {cur[k]:,} ({diff:+,})"
L += [f"- verified targets: {d('verified')} · mapped-or-better: {d('mapped_plus')}",
f"- engine LOC: {d('loc')} · test files: {d('tests')} · checks: {d('checks')}",
f"- addresses verified: {d('addr_verified')} · recovered layouts: {d('layouts')} · open questions: {d('open_q')}", ""]
f"- addresses verified: {d('addr_verified')} · recovered layouts: {d('layouts')} · open questions: {d('open_q')}",
f"- standalone leaves closed: {d('sa_closed')} · leaves still diverging: {d('sa_left')}", ""]
else:
L += ["- first run (no previous `DASHBOARD.md` metrics found)", ""]

263
tools/standalone_report.py Normal file
View file

@ -0,0 +1,263 @@
#!/usr/bin/env python3
"""Measure the standalone against the oracle, and record the distance.
The milestone is: the standalone loads a save, runs one strategic turn, and writes an
autosave that byte-matches what the original produces from the same state. This tool
measures how far off that is, in the only currency the campaign trusts -- named leaves of
`verify/state-checksum/state_checksum.py`, whose coverage is proved by re-serialisation.
For each (before, after) pair of real saves it computes three numbers:
baseline leaves that differ between the INPUT save and the oracle's post-turn save.
This is the distance a standalone that does nothing has to travel.
result leaves that differ between OUR post-turn save and the oracle's.
closed baseline - result, and -- separately -- any leaf we made worse.
`closed` alone would be a comfortable number, so `regressed` is reported next to it: a leaf
that agreed with the oracle before the turn and disagrees after it is a phase doing damage,
and it is counted and named rather than netted off.
tools/standalone_report.py # every pair, write the JSON + text report
tools/standalone_report.py --print # also echo the report
tools/standalone_report.py --pair A.sav B.sav # one ad-hoc pair
tools/standalone_report.py --binary PATH # a sots_turn built elsewhere
Outputs (overwritten):
verify/results/standalone/status.json the completion metric, read by tools/dashboard.py
verify/results/standalone/report.txt the human divergence report
"""
import argparse
import datetime
import json
import os
import shutil
import subprocess
import sys
import tempfile
RE_ROOT = os.path.abspath(os.path.join(os.path.dirname(os.path.abspath(__file__)), ".."))
CK_DIR = os.path.join(RE_ROOT, "verify", "state-checksum")
SAVE_DIR = os.path.join(RE_ROOT, "verify", "results", "saves")
OUT_DIR = os.path.join(RE_ROOT, "verify", "results", "standalone")
sys.path.insert(0, CK_DIR)
import state_checksum as ck # noqa: E402
# The consecutive-turn pairs the corpus holds. A pair is (input, oracle): the oracle is the
# state the game itself produced by ending a turn on the input. Only the first family is a
# true End-Turn transition of one game; the others are listed so a regression on them is
# still visible, with their nature stated.
PAIRS = [
("turn1-state.sav", "turn2-state.sav", "real End Turn"),
("turn2-state.sav", "turn3-state.sav", "real End Turn"),
]
DEFAULT_BINARIES = [
os.path.expanduser("~/sots-engine-wt-standalone/build-host/src/app/sots_turn"),
os.path.expanduser("~/sots-engine/build-host/src/app/sots_turn"),
"/srv/re-lab/build/sots-engine-s2/src/app/sots_turn",
]
def find_binary(explicit):
if explicit:
return explicit if os.path.exists(explicit) else None
for p in DEFAULT_BINARIES:
if os.path.exists(p):
return p
return shutil.which("sots_turn")
def leaf_paths(a, b, limit=200000):
"""The set of leaf paths on which two checksummed saves differ."""
entries = ck.diff(a.root, b.root, limit=limit)
return {e.path: e for e in entries}
def run_pair(binary, src, oracle, note, workdir, keep_saves):
out_sav = os.path.join(workdir, "post-" + os.path.basename(src))
metric = os.path.join(workdir, "metric-" + os.path.basename(src) + ".json")
cmd = [binary, src, "--out", out_sav, "--metric", metric, "--roundtrip"]
proc = subprocess.run(cmd, capture_output=True, text=True)
row = {
"input": os.path.basename(src),
"oracle": os.path.basename(oracle),
"note": note,
"exit": proc.returncode,
"stdout": proc.stdout.strip().splitlines()[-30:],
}
if proc.returncode != 0:
row["error"] = proc.stderr.strip()[:2000]
return row, []
ck_in = ck.checksum_save(src)
ck_or = ck.checksum_save(oracle)
ck_ours = ck.checksum_save(out_sav)
base = leaf_paths(ck_in, ck_or)
ours = leaf_paths(ck_ours, ck_or)
closed = sorted(set(base) - set(ours))
regressed = sorted(set(ours) - set(base))
remaining = sorted(set(ours) & set(base))
row.update({
"coverage": {
"input": ck_in.coverage.get("ok"),
"oracle": ck_or.coverage.get("ok"),
"ours": ck_ours.coverage.get("ok"),
},
"roots": {"input": ck_in.digest, "oracle": ck_or.digest, "ours": ck_ours.digest},
"baselineDiverging": len(base),
"divergingAfterTurn": len(ours),
"closed": len(closed),
"regressed": len(regressed),
"closedPaths": closed,
"regressedPaths": [repr(ours[p]) for p in regressed],
"remainingSample": [repr(ours[p]) for p in remaining[:40]],
"byteMatch": ck_ours.digest == ck_or.digest,
})
if os.path.exists(metric):
with open(metric) as f:
row["standalone"] = json.load(f)
if keep_saves:
dst = os.path.join(OUT_DIR, os.path.basename(out_sav))
shutil.copyfile(out_sav, dst)
row["savedTo"] = os.path.relpath(dst, RE_ROOT)
return row, remaining
def subsystem_breakdown(remaining):
"""Group the remaining divergences by the subsystem they land in."""
buckets = {}
for p in remaining:
parts = p.strip("/").split("/")
key = "/" + "/".join(parts[:2]) if len(parts) > 1 else "/" + parts[0]
buckets[key] = buckets.get(key, 0) + 1
return dict(sorted(buckets.items(), key=lambda kv: -kv[1]))
def main(argv=None):
ap = argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("--binary", help="path to sots_turn")
ap.add_argument("--pair", nargs=2, metavar=("INPUT", "ORACLE"),
help="one ad-hoc (input, oracle) save pair")
ap.add_argument("--print", dest="echo", action="store_true", help="echo the report")
ap.add_argument("--no-write", action="store_true", help="render only, touch nothing")
ap.add_argument("--keep-saves", action="store_true",
help="copy each post-turn save into verify/results/standalone/")
args = ap.parse_args(argv)
binary = find_binary(args.binary)
if not binary:
print("standalone_report: sots_turn not found; build sots-engine's host preset first",
file=sys.stderr)
print(" looked in: " + ", ".join(DEFAULT_BINARIES), file=sys.stderr)
return 2
pairs = ([(args.pair[0], args.pair[1], "ad-hoc")] if args.pair
else [(os.path.join(SAVE_DIR, a), os.path.join(SAVE_DIR, b), n)
for a, b, n in PAIRS])
pairs = [(a, b, n) for a, b, n in pairs if os.path.exists(a) and os.path.exists(b)]
if not pairs:
print("standalone_report: no save pairs available, nothing to measure", file=sys.stderr)
return 0
if not args.no_write:
os.makedirs(OUT_DIR, exist_ok=True)
rows, all_remaining = [], []
with tempfile.TemporaryDirectory() as tmp:
for src, oracle, note in pairs:
row, remaining = run_pair(binary, src, oracle, note, tmp,
args.keep_saves and not args.no_write)
rows.append(row)
if row["input"] == os.path.basename(pairs[0][0]):
all_remaining = remaining
ok = [r for r in rows if r.get("exit") == 0]
ref = ok[0] if ok else {}
status = {
"schema": "sots-standalone-status/1",
"generated": datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ"),
"binary": binary,
"reference": {
"input": ref.get("input"),
"oracle": ref.get("oracle"),
"baselineDiverging": ref.get("baselineDiverging"),
"divergingAfterTurn": ref.get("divergingAfterTurn"),
"closed": ref.get("closed"),
"regressed": ref.get("regressed"),
"byteMatch": ref.get("byteMatch"),
"subsystems": subsystem_breakdown(all_remaining),
},
"phases": (ref.get("standalone") or {}).get("spine"),
"tailPhases": (ref.get("standalone") or {}).get("tail"),
"rng": {
"wordsModelled": ((ref.get("standalone") or {}).get("run") or {}).get("rngWords"),
"wordsPerTurnUnattributed": "18-20 (lane Z, in flight)",
},
"pairs": rows,
}
lines = []
w = lines.append
w("# standalone vs the oracle")
w("")
w(f"generated {status['generated']} binary {binary}")
w("")
ph = status["phases"] or {}
tp = status["tailPhases"] or {}
if ph:
w(f"phases: {ph['modelled']}/{ph['total']} of the two turn drivers modelled, "
f"{ph['committed']} committed "
f"(implemented {ph['implemented']}, partial {ph['partial']}, "
f"blocked {ph['blocked']}, stub {ph['stub']})")
if tp:
w(f" {tp['modelled']}/{tp['total']} of the post-combat tail modelled")
w("")
for r in rows:
w(f"## {r['input']} -> {r['oracle']} ({r['note']})")
if r.get("exit"):
w(f" FAILED, exit {r['exit']}: {r.get('error', '')[:400]}")
w("")
continue
w(f" baseline (do nothing) {r['baselineDiverging']:4d} leaves diverge")
w(f" after one standalone turn {r['divergingAfterTurn']:4d} leaves diverge")
w(f" closed {r['closed']}, regressed {r['regressed']}, "
f"byte match: {'YES' if r['byteMatch'] else 'no'}")
w(f" coverage proved on all three saves: {r['coverage']}")
if r["closedPaths"]:
w(" closed:")
for p in r["closedPaths"]:
w(f" + {p}")
if r["regressedPaths"]:
w(" REGRESSED (agreed before the turn, disagrees after):")
for p in r["regressedPaths"]:
w(f" - {p}")
w("")
if all_remaining:
w("## what still differs on the reference pair, by subsystem")
for k, v in subsystem_breakdown(all_remaining).items():
w(f" {v:4d} {k}")
w("")
w("## first 40 remaining, named")
for p in (ok[0]["remainingSample"] if ok else []):
w(f" {p}")
report = "\n".join(lines) + "\n"
if not args.no_write:
with open(os.path.join(OUT_DIR, "status.json"), "w") as f:
json.dump(status, f, indent=2)
f.write("\n")
with open(os.path.join(OUT_DIR, "report.txt"), "w") as f:
f.write(report)
print(f"wrote {os.path.relpath(OUT_DIR, RE_ROOT)}/status.json and report.txt")
if args.echo or args.no_write:
print(report)
return 0
if __name__ == "__main__":
sys.exit(main())

View file

@ -0,0 +1,88 @@
# standalone vs the oracle
generated 2026-09-08T14:30:05Z binary /home/alex/sots-engine-wt-standalone/build-host/src/app/sots_turn
phases: 14/44 of the two turn drivers modelled, 7 committed (implemented 2, partial 5, blocked 7, stub 30)
2/37 of the post-combat tail modelled
## turn1-state.sav -> turn2-state.sav (real End Turn)
baseline (do nothing) 209 leaves diverge
after one standalone turn 204 leaves diverge
closed 5, regressed 0, byte match: no
coverage proved on all three saves: {'input': True, 'oracle': True, 'ours': True}
closed:
+ /Sim/Frame
+ /Sim/systems/Sys[112 "Gamma Cephei"]/ntdev
+ /Sim/systems/Sys[288 "Ke'Dolarra"]/ntdev
+ /Sim/systems/Sys[304 "Koa’Vo"]/ntdev
+ /Summary/Turn
## turn2-state.sav -> turn3-state.sav (real End Turn)
baseline (do nothing) 108 leaves diverge
after one standalone turn 103 leaves diverge
closed 5, regressed 0, byte match: no
coverage proved on all three saves: {'input': True, 'oracle': True, 'ours': True}
closed:
+ /Sim/Frame
+ /Sim/systems/Sys[112 "Gamma Cephei"]/ntdev
+ /Sim/systems/Sys[288 "Ke'Dolarra"]/ntdev
+ /Sim/systems/Sys[304 "Koa’Vo"]/ntdev
+ /Summary/Turn
## what still differs on the reference pair, by subsystem
82 /Sim/players
80 /Sim/systems
24 /Sim/turnstats
8 /Sim/SvSctOb
1 /Sim/DesignIDs[]
1 /Sim/FleetIDs[]
1 /Sim/ModCount
1 /Sim/NMnx
1 /Sim/NumFlts
1 /Sim/RNG
1 /Sim/ShipIDs[]
1 /Sim/cmbtid
1 /Sim/fleets
1 /Summary/Checksum
## first 40 remaining, named
/Sim/DesignIDs[]: removed [], added [18, 1712] (41 -> 43 entries)
/Sim/FleetIDs[]: removed [], added [1744] (6 -> 7 entries)
/Sim/ModCount: 2 -> 12
/Sim/NMnx: 106 -> 109
/Sim/NumFlts: 6 -> 7
/Sim/RNG/.: '<raw 2503 B 9e6887688129a8f6>' -> '<raw 2503 B ef4d678696ed4c53>'
/Sim/ShipIDs[]: removed [], added [1728] (15 -> 16 entries)
/Sim/SvSctOb/EncObj[3]/CDiff: -1 -> 0
/Sim/SvSctOb/EncObj[5]/Hives/.: only-in-A
/Sim/SvSctOb/EncObj[5]/Hives/.[0]: only-in-B
/Sim/SvSctOb/EncObj[5]/Hives/.[1]: only-in-B
/Sim/SvSctOb/EncObj[5]/Hives/.[2]: only-in-B
/Sim/SvSctOb/EncObj[6]/did: only-in-B
/Sim/SvSctOb/EncObj[6]/didc: 0 -> 1
/Sim/SvSctOb/EncObj[6]/ini: False -> True
/Sim/cmbtid: 1 -> 2
/Sim/fleets/Flt[1744 "Alpha Fleet"]: only-in-B
/Sim/players/Player[16 "re"]/BnkEl: -1590613 -> -1594593
/Sim/players/Player[16 "re"]/BnkPr: -787353 -> -789323
/Sim/players/Player[16 "re"]/Events/EvNxID: 0 -> 2
/Sim/players/Player[16 "re"]/Events/Events/.: only-in-A
/Sim/players/Player[16 "re"]/Events/Events/.[0]: only-in-B
/Sim/players/Player[16 "re"]/Events/Events/.[EvTurn=2]: only-in-B
/Sim/players/Player[16 "re"]/Sav: 50000 -> 289688
/Sim/players/Player[16 "re"]/Status: 0 -> 4
/Sim/players/Player[32 "Fane Lao"]/BnkEl: -1811273 -> -1815833
/Sim/players/Player[32 "Fane Lao"]/BnkPr: -896580 -> -898837
/Sim/players/Player[32 "Fane Lao"]/Events/EvNxID: 0 -> 2
/Sim/players/Player[32 "Fane Lao"]/Events/Events/.: only-in-A
/Sim/players/Player[32 "Fane Lao"]/Events/Events/.[0]: only-in-B
/Sim/players/Player[32 "Fane Lao"]/Events/Events/.[EvTurn=2]: only-in-B
/Sim/players/Player[32 "Fane Lao"]/FNG/FNGNum: 0 -> 1
/Sim/players/Player[32 "Fane Lao"]/Maint: 0 -> 500
/Sim/players/Player[32 "Fane Lao"]/NumDes: 5 -> 6
/Sim/players/Player[32 "Fane Lao"]/PvSav: 50000 -> 38100
/Sim/players/Player[32 "Fane Lao"]/ResRate: 0.25 -> 0.800000011920929 [1.34218e+07 ulp]
/Sim/players/Player[32 "Fane Lao"]/ResTNm: '' -> 'IND_Waldo'
/Sim/players/Player[32 "Fane Lao"]/Sav: 50000 -> 92651
/Sim/players/Player[32 "Fane Lao"]/ShipRecs/srb[0]: 0 -> 1
/Sim/players/Player[32 "Fane Lao"]/ShipRecs/srb[3]: only-in-B

File diff suppressed because it is too large Load diff

Binary file not shown.