sots-re/findings/subsystems/fpu-precision-sensitivity.md
alex 849c5069fc lane F: x87 precision sensitivity measured; STATE_CHECKSUM 3.5 closed
Seven End Turns from ref-turn2.sav on VM140, six control words, whole-state
checksum on every post-turn autosave.

53-bit and 64-bit x87 give byte-identical state across all 35,394 leaves, so
an x64/SSE port computing in IEEE double has NO double-rounding budget to
preserve and floats=bits is free. Two settings do move state, each reproduced
on a repeat run:

  0x007f (24-bit)   Sys[112 "Gamma Cephei"]/Pop2/PopG/PopC 540000000 -> 540000002
  0x1a7f (round-up) Flt[34 "Beta Fleet"]/Pos/.[0] and /Pos/.[2], 1 ULP each

So the port must hold intermediates at 53 bits and use round-to-nearest --
both SSE defaults, now measured rather than assumed, each with a named
regression witness.

The briefed triple was under-powered: 0x027f is 53-bit (it differs from 0x127f
only in bit 12, infinity control, ignored since the 387) and 0x137f is 64-bit,
not a rounding change. Run as written all three come back identical, and that
would have "proved" something false on both axes that matter.

Evidence the forced word actually held: read-back at each force site plus 38
independent in-pipeline hook samples per run spanning turn phases 4, 6 and 8,
all reading the forced value. Mars::Application::Run calls
_controlfp(0x50000,0x3070300) at 0x0089f606 every frame, which is 0x127f, so
forcing at StrategyClient::EndTurn is wiped before the turn runs;
StrategyServer::BeginProcessTurn is the point that works.

Also re-confirms the End-Turn determinism oracle on engine cef889e:
bb4fd9ac... / 978041ac... unchanged.

New tools: verify/fpu-cw/cw_census.py, verify/fpu-cw/trace_bitdiff.py (the
latter exists because under a forced 24-bit word the CRT's own %g rendering
degrades, so trace text is not a valid comparison surface).
2026-09-08 05:07:53 -04:00

12 KiB
Raw Permalink Blame History

x87 precision sensitivity of the turn pipeline — measured (2026-09-08, lane F)

Question (verify/state-checksum/STATE_CHECKSUM.md §3.5): does any value the turn pipeline produces actually depend on the x87 control word? Until now every save on this host was made at the game's own fpu_cw = 0x127f, so "the x64/SSE port must match bit-for-bit" was policy, not measurement.

It is now measured. Seven End-Turn runs from the same ref-turn2.sav, six distinct control-word settings, whole-state checksum on each post-turn autosave.

Verdict

forced fpu_cw precision rounding (Autosave).sav vs baseline
(stock, none forced) 53-bit nearest 978041ac… 67,219 B — (baseline)
0x027f 53-bit nearest 978041ac… 67,219 B identical
0x127f 53-bit nearest 978041ac… 67,219 B identical
0x137f 64-bit (extended) nearest 978041ac… 67,219 B identical
0x007f 24-bit (single) nearest ba2435be… 67,222 B 2 leaves
0x1a7f 53-bit up (+∞) e48e25fa… 67,219 B 2 leaves

The pre-turn (Autosave EndTurn).sav is bb4fd9ac… in all seven runs — the state at the moment End Turn was pressed is upstream of every forced value, which is the run-to-run control.

For the reimplementation:

  1. The x87's 64-bit intermediates are not load-bearing. 53-bit and 64-bit produce the same state, leaf for leaf, across 35,394 leaves. A port that computes in IEEE double reproduces this turn exactly, and there is no double-rounding budget to preserve — the strict floats=bits policy costs an x64/SSE port nothing on this axis. §3.5 closes.
  2. Narrowing intermediates to float does change state. At 24-bit the home system's civilian population lands two people higher. So the port must hold in double exactly where the original holds in an x87 register, and narrow exactly where the original stores to a dword — which is precisely the discipline lane M documented for MoveFleet (movefleet-position-rounding.md: five separate float32 narrowings, products and sums held at 53 bits in between). Getting a narrowing point wrong is not a rounding nicety; it moves saved integers.
  3. Round-to-nearest is required. Round-toward-+∞ moves fleet 34's position by 1 ULP in two of three components. SSE defaults to round-to-nearest, so this is satisfied for free — but it is now a measured requirement rather than an assumption.

The two named precision-sensitive witnesses

Everything the experiment found, in full — this is the complete list for this turn:

0x007f (24-bit) vs baseline:
  /Sim/systems/Sys[112 "Gamma Cephei"]/Pop2/PopG/PopC : 540000000 -> 540000002
  /Summary/Checksum                                   : -769976634 -> -769976632   (derived)

0x1a7f (round-up) vs baseline:
  /Sim/fleets/Flt[34 "Beta Fleet"]/Pos/.[0] : -10.563499450683594 -> -10.563498497009277  [1 ulp]
  /Sim/fleets/Flt[34 "Beta Fleet"]/Pos/.[2] :   1.564733624458313 ->   1.5647337436676025 [1 ulp]

Both reproduce exactly on a repeat run (run-007f-rep, run-1a7f-rep) — same root hash, same leaves. Summary/Checksum moves by the sum of the changed bytes, consistent with the additive checksum already characterised in determinism-oracle.md; treat it as derived, never as state.

The two witnesses are on different axes and different subsystems: the population integer is precision-sensitive but not rounding-sensitive, and the fleet position is rounding-sensitive but not precision-sensitive. That is a useful shape — it says an SSE port can go wrong in two independent ways, and each has a cheap regression witness on turn 2 of ref-turn2.sav.

The evidence that the control word actually held

This is what makes the result mean anything: an experiment where the setting silently reverted would produce identical saves and a confident, false "nothing depends on precision".

1. Read-back at the point of forcing. Each force site logs observed-before → requested → read-back-after. Every one reports OK, e.g. (run-007f/shim.log):

fpu: FORCE at StrategyServer::BeginProcessTurn: observed=0x127f 53bit-double/nearest
     -> requested=0x007f -> readback=0x007f 24bit-single/nearest  OK
fpu: sample at StrategyServer::ProcessTurn: cw=0x007f 24bit-single/nearest

StrategyServer::ProcessTurn is a sample-only hook — it never writes the word. Its reading is independent evidence that the value forced at the turn gate survived into the simulation.

2. 38 independent in-pipeline samples per run. Every template hook already snapshots the control word it finds. In each of the five forced runs, all 38 samples read the forced value and nothing else (verify/results/fpu-cw/cw-census.txt):

run-007f:  ServerSystem::ProcessTurn cw=0x007f x28   call_id 1110..1137
           StrategyServer::MoveFleet cw=0x007f x7    call_id 1103..1109
           TechTree::ProcessResearch cw=0x007f x3    call_id 1140..1144
           ALL SAMPLES: 0x007f x38

Those three hooks sit in turn phases 4 (movement), 6 (colony) and 8 (per-player research) — so the samples span the pipeline from its first mover to its last research pass, in call-id order, with no gap and no other value. The stock run reads 0x127f x38 in the same places.

3. The forcing is a single write per turn, never re-applied inside the pipeline. Re-forcing at each hook would have guaranteed the samples without proving anything.

Why the forcing point matters: the game re-arms the word every frame

Mars::Application::Run (0x0089f5b0) calls _controlfp(0x50000, 0x3070300) once per frame at 0x0089f606, between OnUpdate (vft+0x18) and OnTick (vft+0x1c). In MSVC's abstract encoding that is _PC_53 | _IC_AFFINE under the mask _MCW_DN|_MCW_IC|_MCW_PC|_MCW_RC, which is exactly the x87 word 0x127f. There is no literal 0x127f anywhere in the image (find-constant-uses → 0 hits); the value is synthesised by the CRT. _controlfp is also called from Mars::Application::Initialize (0x008a10c6) and from FUN_00451920 (0x00451929, one caller, a data reference).

Two consequences, both visible in the logs:

  • Forcing at StrategyClient::EndTurn does not survive. The client's End Turn only raises SETurnEndPending and sends SNMEndTurn; the server's OnMessage → BeginProcessTurn → ProcessTurn runs on a later frame, and the per-frame _controlfp wipes the setting in between. Every log shows FORCE at StrategyClient::EndTurn … readback=0x007f at t, then FORCE at StrategyServer::BeginProcessTurn: observed=0x127f about a second later. BeginProcessTurn is the forcing point that works, because BeginProcessTurn and ProcessTurn are called from the same OnMessage, inside one OnUpdate, with no _controlfp between them.
  • The per-tick sampler never logs a change, and that is expected, not evidence: _controlfp runs immediately before OnTick, so OnTick always observes a freshly re-armed 0x127f. The sampler's real contribution is the negative one — it proves the process never sits in a forced state outside the turn, so nothing leaks between runs.

This also retro-explains why every previous lane measured fpu_cw = 0x127f at every hook: the main loop guarantees it, frame by frame.

The lane brief's three values do not span three FPU modes

The experiment as specified in §3.5 and in the lane brief would have been under-powered, and this is worth recording so nobody re-derives it:

value actually is brief called it
0x027f 53-bit, nearest (MSVC CRT default) "24-bit mantissa (single precision)"
0x127f 53-bit, nearest (+ infinity-control bit) "53-bit — the game's own setting" ✓
0x137f 64-bit extended, nearest "53-bit but a different rounding mode"

0x027f and 0x127f differ only in bit 12 (infinity control), which the 387 and every later x87 ignore — they are the same arithmetic. So the briefed triple spans two modes (53-bit and 64-bit), not three, tests single precision not at all, and tests rounding not at all. Run literally, all three come back identical (they do — see the table) and the honest conclusion from them alone would have been "no precision sensitivity", which is false for rounding and unproven for single precision.

The control-word layout: bits 0–5 exception masks, bits 8–9 precision control (00=24-bit, 10=53-bit, 11=64-bit), bits 10–11 rounding control (00=nearest, 01=down, 10=up, 11=truncate), bit 12 infinity control (ignored). So the genuine probes are 0x007f (24-bit) and 0x1a7f (53-bit, round-up), and those are the two runs that found something.

Trap: the trace's own float rendering is precision-sensitive

Under a forced 24-bit control word, the shim's trace text is not a valid comparison surface. The emitter prints an f32 with %.9g, and the CRT's digit generation is itself x87 arithmetic, so the same float32 renders differently:

127f run:  "z":{"t":"f32","v":1.56473362}      both are float32 0x3fc84931
007f run:  "z":{"t":"f32","v":1.5647336}       — the value did not move, the printf did
127f run:  "suit":{"t":"f32","v":1.65671718}   both are float32 0x3fd40f4f
007f run:  "suit":{"t":"f32","v":1.65671721}

A text diff of the two traces reports dozens of these as divergences. verify/fpu-cw/trace_bitdiff.py re-quantises every float to its IEEE bit pattern before comparing, and then the 24-bit run shows zero differences in any hooked call — the population change happens in a byte no hook declares. This is the STATE_CHECKSUM.md §1 thesis with a fresh example: the per-function verdict was clean and the state had still moved; only the whole-state checksum saw it. Any future tracecmp-style comparison across control words must compare bits, not text.

Incidental: guard spans differ between byte-identical runs

ServerSystem::ProcessTurn's undeclared-write spans are not stable run to run even when the saves are byte-identical (trace-bitdiff.txt, 127f vs off): the same three writes at system+420..432 are reported once as one 12-byte span and once as three 3-byte spans, and once shifted by a byte. Harmless for a verdict (the set of touched bytes is the same) but it means guard-span text is not a stable comparison surface either. Not chased — outside this lane.

Method

  • VM140, ref-turn2.sav (ab4ac2d7…), click path per determinism-oracle.md. Shim build fpucw-cef889e-20260908T0803Z, hook set held identical across all seven runs — the only difference between shim.cfg.fpu* files is the fpu.force line.
  • SavedGames\ is reset to the same four files before every run, so the Load dialog rows never move and the click path is constant.
  • Forcing implemented in sots-engine/src/shim/fpu_force.cpp (branch wip/fpucw): register-transparent asm stubs, same pattern as the Application::Initialize hook, so the [unverified] prototypes of the turn-gate functions are never relied on. shim.cfg keys fpu.force=<cw>|off and fpu.sample_ticks=on|off.
  • Evidence: verify/results/fpu-cw/run-<tag>/ — both autosaves, shim.log, shim.cfg, gzipped trace; plus cw-census.txt, state-checksum-diffs.txt, trace-bitdiff.txt.
  • Tools: verify/fpu-cw/cw_census.py (per-hook control-word census), verify/fpu-cw/trace_bitdiff.py (bit-exact trace comparison).

Coverage and limits

  • The whole-state checksum proves coverage on every run: 609,080 bytes rebuilt byte-for-byte from 35,394 leaves. "Identical" here means no leaf anywhere moved, not "no declared region moved".
  • One turn, one save, one galaxy. Turn 2→3 of ref-turn2.sav: 28 systems, 3 owned, 7 MoveFleet calls of which one fleet actually moves, 3 research passes, no combat, no Zuul, no plague/rebellion/terraform. The two witnesses are what this turn exposes; a turn with combat or more movers could expose more. The result "53-bit == 64-bit" is the one that generalises best, because it held across all 35,394 leaves and all 38 in-pipeline samples.
  • Nothing here says the original is x87-free — only that its results on this turn do not move between 53-bit and 64-bit intermediates, which is the property an SSE port needs.