sots-engine/docs/R-recapture.md
alex e50d5e5900 recapture: golden traces with guards on the live game
Five hooks recaptured; no source change needed. ProcessResearch's over-budget
event now diverges as one line instead of a 609 KB save diff (15 calls, 3
diverged, RNG 15/15). ComputeBudget holds at 0 over 4284 compares but only 20
distinct states and 13 always-zero slots. MoveFleet diverges on 8 of 45 calls by
1 ULP of position - B4's clean verdict was a one-sample verdict. First guarded
captures for OnTechResearched (RollResearchEvent draw finally fired and matched),
the colony turn and MoveFleet. Guards mapped SetResearched live and found an
undeclared observed-tech vector append.
2026-09-08 03:04:34 -04:00

8.2 KiB
Raw Permalink Blame History

R — golden-trace recapture with guards (live VM, 2026-09-08)

Every golden trace went stale when Coverage + guard regions landed (3f0721f+, docs/harness-audit.md). This lane recaptured all five on the game. Coverage first: two hooks that shipped as "verified" now report divergences, and both are real.

Build recap-7584bad-20260908T0615Z — the tree at 7584bad plus four src/shim/shim.cfg.recap* variants. No source change was needed to produce any of these results; the audit's machinery did all of it. Workload: the reference save → Launch → End Turn (the B3 and misc runs continue to five End Turns, because the reference turn contains no research completion). The End-Turn oracle hashes matched on every run's first turn, so neither trace nor compare mode perturbs the game.

hook mode calls compared diverged exit coverage
Game::TechTree::ProcessResearch compare, 1 turn 3 3 1 1 3 guarded, 0 undeclared
Game::TechTree::ProcessResearch compare, 5 turns 15 15 3 1 15 guarded, 13 undeclared in 2 calls
Game::ServerPlayer::ComputeBudget compare 4284 4284 0 0 4284 guarded, 0 undeclared
Game::ServerPlayer::OnTechResearched compare 2 2 0 — 2 guarded, 6 undeclared in 2 calls
Game::ServerSystem::ProcessTurn compare 140 140 0 — 140 guarded, 15 undeclared in 6 calls
Game::StrategyServer::MoveFleet compare 45 45 8 1 45 guarded, 42 undeclared in 15 calls
all five trace (golden) 3630 — — 0 3630 guarded, 9 undeclared in 4 calls

The last three share one log; its exit 1 is MoveFleet's.

1. The over-budget event: visible, on the predicted call

side.events.after.v.next_id [exact] orig={"t":"i32","v":4} ours={"t":"i32","v":3}

One divergent call out of three on the reference turn, and that is its only divergent field. It is bit-for-bit the signal docs/B3.md could previously only get by inflating and diffing a 609 KB autosave (EvNxID 4 → 3). Everything the milestone claimed still holds on the same call: the node's progress and its over-budget flag both reproduce, and the single RNG draw matches (mt[624] hash, stream position and remaining count all identical). The mechanism in docs/harness-audit.md §2 works exactly as designed, including the masking — the player guard stayed silent on that call because events is now a Result region covering the word that moved.

Across five turns the same field diverges on the two completion calls as well, by two ids each (a completion posts more than one event). The rest of those two calls' divergences are the SetResearched fields B3 declared out of scope in advance. RNG matched on 15 of 15.

2. What the guards mapped for free

On the two completion calls the player and tree_header guards printed what SetResearched touches, without a single new hook:

  • the player's construction-cost and output modifiers, and on the other call the population modifier — i.e. OnTechResearched's own declared regions, seen from the caller. A clean cross-check.
  • the current-research-target pointer being cleared.
  • the tree's completion-order counter — harness-audit row 9, confirmed live.
  • the player's observed-tech vector growing (all three vector words move, so it reallocates). This is not in any coverage note anywhere, it is a third list append in the same neighbourhood as the event list, and it is serialized state. New gap.

OnTechResearched's own guard reports the event counter on both of its calls, so B2's "ours posts no events" gap is now measured rather than assumed — the guard names the byte, where docs/B2.md gotcha 4 told readers a changed save hash was expected and therefore not a finding.

3. MoveFleet: 8 of 45 diverge at the last bit

Every divergence is a 1-ULP position component (the largest, 64 ULP, is at a near-zero result where the absolute error is the same ~1.2e-7 as the others). The error sits at half an ULP of the inputs, not of the result — the signature of one rounding too many or too few in the position update, not a formula error: the step length is right, every ship-range region matches, and x, y and z each diverge on some call.

docs/B4.md reported this hook clean. It had one moving call, and that fleet still matches bit-for-bit on four of its five moves here. The bug only appears once other fleets move. This is the thin-coverage failure mode caught by widening the workload, not by changing anything.

All 15 moving calls are the same straight-run waypoint type; types 2–5 never occurred, so the node-line step, node routes, gate teleports, the probabilistic jump and every RNG draw in this hook remain untested. The two arrivals in the run compared clean — which still means nothing, since no arrival machinery is declared.

4. ComputeBudget: verdict held, coverage is narrower than the number suggests

4284 compared, 0 diverged, exit 0, and the budget_object guard caught nothing in 4284 calls — including the over-budget slot the audit flagged as written but never checked (row 11). It never took a different value here, which is weaker than "it is never written".

Honest width of that result:

  • the 4284 calls hold only 20 distinct (player, output) states; 4278 of them are the strategy-map UI polling one player, and the End-Turn pass itself is ~8 calls.
  • 13 of the 22 slots are zero on every call. Five of the six declared input-boundary slots are among them, so they prove nothing twice over. Eight modelled slots ever carry a value; that is the real evidence.

5. ServerSystem::ProcessTurn: 140 clean calls, two moving counters

Over five turns, exactly two declared regions ever changed: the turns-developing counter (15 times) and the recon countdown (7). Decay, both bonus pools, the resource total, the halt flags, the battle countdown and the generator were byte-identical before and after on all 140 calls. A 0-divergence verdict here means "two integers agreed and nothing else moved". The generator not moving on any of 140 colony turns re-confirms B4's headline over five turns.

Its guard reports the AI home system's fleet vector growing every turn (the build queue emitting a ship — harness-audit row 14, observed live), the civilian population vector on the first turn only, and two spans no contract entry names. The morale-event append (row 3) never fired: still unmeasured, not cleared.

6. Traps for the next lane

  1. The reference turn contains no research completion. OnTechResearched produces an empty log on the documented one-turn recipe, and an empty log still passes. Reaching a completion took five End Turns.
  2. Only the first End Turn is reproducible. It matched the oracle byte for byte on all four runs, and its research calls reproduce docs/B3.md exactly. From the fourth turn on, the AI picks a different research target than B3 recorded while the point totals stay nearly identical. Treat anything past turn one as a run, not the run.
  3. ComputeBudget is called ~4280 times per session by the UI and ~8 times by the turn pass, so any "N compares" headline for it is dominated by UI polling.
  4. Guards cost nothing worth measuring: the whole five-hook trace run wrote 3630 records with guards on every call and the End-Turn oracle still hashed exactly.

7. Still owed

  • The count-only events region recommended by docs/E-events.md — it would turn audit row 1 from "known defect" into "checked". Not wired here: the captures predate that module and a good capture is not worth redoing.
  • Pin the rounding in the position update against the instruction stream.
  • Name the three unidentified spans (two on the system, one on the fleet) in ghidra/addresses.json.
  • Declare or at least note the observed-tech vector append.
  • Nothing here exercised replace mode on any hook, the Zuul double roll, a rebellion, an addiction sweep, a plague, gate traffic, waypoint types 2–5, a silent completion, or the budget expense/aid/debt tail.

Full per-call evidence, offsets and field names live in the RE repo under findings/subsystems/golden-trace-recapture.md; logs and reports under verify/traces/recap-* and verify/results/compare/recap-*.