sots-engine/docs/Y-standalone-rng-and-tail.md

11 KiB
Raw Permalink Blame History

Lane Y — the predictions, written before the build

Committed before src/app was touched, per earned rule 2. Falsification sections included: how each prediction could be wrong, and the symptom of each way.

Two things are being predicted: (a) what the standalone's generator does on lane Z's two calibrated oracle pairs, and (b) whether a turn-record model built from the wire reproduces the record the game already wrote into every save it made.


P-Y1 — the trade-raid model costs exactly 16 words on an 8-player save

Lane Z measured ServerTradeManager::GenerateTradeRaidEncounters firing two Mars::RNG::Chance rolls for every entry of the server's player vector, and showed that no back-edge contains either site, so one word per player per site is a bound, not an observation. Both probabilities are strictly inside (0, 1), so Chance takes neither early-out and spends exactly one word.

The reference save's player vector holds 8 entries (Sim.NumPlrs; the wire's players array is that same vector).

The standalone, run with --commit-rng on z2-endturn.sav, consumes exactly 16 generator words, all of them attributed to the trade-raid phase, and the post-turn left is 395 − 16 = 379. mt[624] is unchanged (no block boundary is crossed).

Same claim on z-t6-endturn.sav: 16 words, left 375 − 16 = 359.

Falsification.

how it could be wrong symptom
the loop is not over the server's player vector a word count that is not 2 × 8
Chance is p >= r or narrows differently word count still 16 (the count does not depend on the comparison) — this prediction cannot test it, and says so
the NPC-raid gate is not open on this save 8 words, not 16
the refugee-raid gate is open 24 words
the engine's next_float disagrees with the original invisible here; only the count is being tested

P-Y2 — the residual against the oracle is 4 and 2, not 0

The oracle pairs carry a known cost: z2-endturn → z2-autosave is 20 words, z-t6-endturn → z-t6-autosave is 18. Lane Z's per-site ledger attributes the difference to four sites the standalone does not model:

  • 0x0050329d (NextFloat) and 0x007929a4 (NextInt), both under StrategyServer::DetectEncounters — one word each per turn on 3 of 3 measured turns;
  • TechTree::ProcessResearch+0x1c8 and ServerPlayer::RollResearchEvent+0x2f — 0–1 words each, both downstream of the research allocation, which is downstream of ComputeBudget, which is blocked on the population → base-output term.

The standalone lands 4 words short on the turn-5 pair and 2 words short on the turn-6 pair, and the shortfall is exactly the four sites above: 2 detection + 2 research on turn 5, 2 detection + 0 research on turn 6. The generator state therefore does not match the oracle on either pair, and the answer to "does it land on the oracle's mt[624]/left" is no, by a stated amount.

Falsification. A residual other than 4 and 2 means the site ledger and the file oracle disagree about this turn, which would be a real contradiction between two of lane Z's three instruments and would have to be reported as one. A residual of 0 would mean the trade-raid count is coincidentally absorbing the other sites.

Why the shortfall is not closed by guessing. Modelling "2 detection words per turn" would fit three observations with no derived rule behind it, and lane I bounds that site's cost at 0 .. |contacts| × |detectors|. It is left unmodelled and reported.


P-Y3 — six fields of the turn record are recoverable from the wire

StrategyServer::FinalizeTurnRecords (tail phase 36) fills a per-player record and archives it by turn. The archive is on the wire: it is /Sim/turnstats/nply[]/hist/stats[], one stats element per turn per player, and every save carries an element for its own Frame.

A record built from a save's own state reproduces that save's archive entry for its own turn, on every player of every save in the corpus, for these six fields:

record field rule
trn the frame counter
col the number of owned-system ids
sav the player's savings
inc savings minus previous-turn savings
tch the number of tech-tree states in the completed state
pop summed over owned systems: population plus the pending population bonus

and it does not reproduce almem, tdinc, bat, the acquired/lost system lists, or the per-hull-class ship census, because each of those is written by a phase or an input the standalone does not hold.

Provenance, stated plainly. The six rules were found by search over the corpus in Python before this was written, so this is not a blind prediction of the rules. What it is a prediction of is that a C++ transcription over the engine's own typed shapes agrees with them on all 112 player-records — a test of the transcription and of the shape typing, not of the rules. Two of the six (pop, inc) were ambiguous in the Python search and were separated by the corpus: pop = Σ(Pop + pbon) matched 112/112 while Σ(Pop + dcs) matched 101/112.

Falsification. Any player-record where the six computed fields differ from the stored ones. Particularly: a save where a player's owned-system list references a system id the save does not carry (the sum would silently drop a term).


P-Y4 — committing phase 36 would regress, so it is not committed

The record the standalone can build is the record for the input save's turn, not for the turn it is about to produce: sav after the turn needs ApplyNetToSavings, which is blocked, and almem needs the alliance-mask rebuild, which is a stub.

Committing phase 36 would add a new archive element whose sav, inc, almem, tdinc and class census are all wrong, turning one container-shaped divergence per player into several leaf-shaped ones. The divergence report would show a large regressed count and a small closed one. It is therefore evaluated, reported, and not written. The divergence baseline stays 5 closed / 0 regressed on both pairs.

Falsification. If committing it closed more than it broke, the block was too conservative. That is measurable with --commit-blocked and is reported rather than argued.


P-Y5 — the bankruptcy divisor in the engine is one ulp wrong

ComputeBankruptcyLimits divides by -0.15. The constant the original loads is the double -0.15000000596046448, which is (double)(float)-0.15f. The two differ, and the difference is visible on a large enough maximum income.

ftol(maxIncome / -0.15) and ftol(maxIncome / -0.15000000596046448) differ for some maxIncome reachable in a real game, and the smallest such value is small enough to be found by a linear scan.

This lane does not fix it: the file is another lane's. It is measured and reported.

Falsification. If no maxIncome in the int32 range separates the two constants, the correction is real but inert and should be recorded as such.


Outcomes

Written after the run, against the predictions above. Nothing above was edited.

P-Y1 — held exactly

16 words on both calibrated pairs, all attributed to the trade-raid phase, left 379 and 359. The player vector read 8 on both, as the wire says. The falsification table's other branches did not fire: not 8 (the NPC gate was open), not 24 (the refugee gate was not).

P-Y2 — held exactly

Residual 4 on the turn-4→5 pair and 2 on the turn-5→6 pair, matching the per-call-site split (2 detection + 2 research, and 2 detection + 0 research). The state block is byte-identical in both cases and only left differs, so the two instruments disagree about nothing.

The answer to the milestone question is therefore "no", by 4 words and by 2.

One thing this pair of numbers does not establish, and the prediction should have said so: the residual is a subtraction, so it confirms the sum of the modelled and unmodelled terms, not each term separately. The static bound is what carries the 16; this arithmetic corroborates it.

P-Y3 — held

480 fields over 80 player-records across 11 saves, 0 mismatches, 0 dangling owned-system ids. The same check runs inside every turn and is printed as a phase note.

P-Y4 — held, and the shape of the result is more useful than the claim

--commit-blocked closes 29 and regresses 17 on the first reference pair, and closes 13 and regresses 19 on the second. Every regressed leaf is in the unmodelled set: almem x8, inc x3, sav x3, three of the ship census. Nothing outside it moved.

So "committing makes it worse" is not what the measurement says: on the first pair the turnstats subsystem is net +7 and on the second it is net -9. The record is right in six fields and wrong in five, and which way the net falls depends on the save. The phase stays blocked, and it becomes a clean +24 the moment the alliance mask and the budget land.

P-Y5 — WRONG in its framing, and the correction is the result

The prediction said the two constants differ "for some maxIncome reachable in a real game" and that the smallest such value could be found by a linear scan. The scan found it immediately, and the framing -- inherited from the phrase "one ulp out on large empires" -- was the wrong picture entirely.

trunc(m / -0.15) and trunc(m / -0.15000000596046448) differ for every m divisible by 3: 1000 of the first 3000 integers, smallest m = 3 giving -20 against -19. It is not precision loss at scale. m / 0.15 is an exact integer whenever 3 divides m, and the two constants straddle it.

And the multiples of 3 are only the floor: the disagreement rate is 33.3% up to about 1.2 million, 66.7% by 1.36 million and 100% at and above 3 million.

Measured on the corpus by inverting the stored limit: 6 of 25 player-records would get a different stored value from the engine's current constant than the game wrote. The defect is live, it is a whole unit rather than an ulp, and it is a third of all inputs at small values and all of them at large ones -- the opposite shape from "one ulp on large empires", which would have been harmless.

The file belongs to another lane in flight, so it is reported and not fixed here.

An unpredicted result: the blocked input is recoverable from every save

The elimination limit is injective in the maximum-income sum it is computed from (slope about 6.67), so it inverts. Every save in the corpus therefore states, for every player that owns anything, the exact sum of its systems' maximum money outputs -- which is the term that blocks ComputeBudget, both savings phases, the bankruptcy limits and five fields of the turn record.

That was not predicted and is the most useful thing this lane found. The same inversion recovers the protection-limit factor as 3.3 from the saves, independently of the game data files.