sots-re/verify/results/research-completion-abi-independent/result-run-735fcb8f4876c10285b03fad.md

7.6 KiB

Independent ABI verification result — scoped static pass

Session run-735fcb8f4876c10285b03fad; verifier research-abi-independent; model openai/gpt-5.6-sol.

Scope classification

This is an independent static reproduction plus archived-state check. It is not original-assisted runtime execution, partial or full live comparison, independent replacement, or integrated replay. No game, allocator, constructor, destructor, exception path, or RNG operation executed. Accordingly, this result supports the contract's independent-cross-check criterion only within its static scope; it does not establish live allocator safety or replacement acceptance.

Preconditions and falsifiers

Before execution I required exact paired HEAD/common-directory/source bindings, executable and tool identity, positive nonempty execution, empty stderr, complete terminal bytes, and no unexpected skip. Semantic falsifiers were collapsed 0x2c/0x74 strides, raw string-header transfer, missing long-string frees, description omission from equality, ordinary equality accepting unordered x87 coordinates, or archived event values existing only as counters rather than direct tree state.

Required states remain distinct: empty/short/long strings; equal/description-only-different events; finite/unordered coordinates; empty/spare/full capacity; new/existing tech; duplicate/nonduplicate; empty/nonempty nested vectors; normal destruction/unwind; and prune/get-or-create boundaries. The raw instructions expose these branch alternatives, but this run did not execute those live states.

Identities and positive execution

The engine pair is HEAD 7741d42fc5e4e761e6449bdaf0e4a61d00036a23, common Git directory /home/alex/sots-engine/.git, binding ccd8e02083e8d2e2b3e97976ace2273c8f924dfc02a39e919004eaf3544c50fd. The RE pair is HEAD 3bfde5a70d874a723e797a695bbd847fd82c0aa7, common directory /home/alex/sots-re/.git, binding 6696fd5201e144843617cbf6d78b41b5287ad5dcc9fa1e8aaa861d52b64e72e8. Their dirty inventories were pre-existing. Input dumps/sots.exe is 7,898,624 bytes, SHA-256 970b7de729956a53094c7eb98aba4270aee98e2fed5daf0d39e290013c90c841. /usr/bin/objdump is GNU 2.38, SHA-256 1eaaef2e7f57c4c7f69115c495e2466f5a8c8e5f3bc42221d092382f30f9d4cd.

The owned reproducer executed 22 windows. All returned zero with nonempty stdout and empty stderr; there were no skips. All six repaired ownership narrow/wide stream hashes exactly equal the analyst manifest. Narrow streams end in c2; wide streams establish five c2 04 00 encodings and one c2 08 00, with every preceding instruction row equal. Four originally complete-stop controls also match the analyst hashes. The independent dedup pair likewise matches its repaired archive hashes and widens c2 to c2 08 00. manifest.json records every argv, stream size/hash and result rather than using a count as proof.

Independent challenges and key ABI observations

  • Historical-stop ablation: all seven narrow/wide tests reproduce objdump's truncation behavior. Therefore decoded ret text from a narrow stream is not accepted as complete-byte provenance.
  • Held-out complete-stop boundary: stopping the ObservedTech constructor at 0x0085630c omits its ret; stopping at 0x0085630d adds exactly 85630c: c3 ret. This challenges the assumption that every historical stop suffered the ret imm16 effect and confirms the plain-ret control.
  • Description equality: the fresh dedup window passes both +0x08 description objects to 0x0046f8c0, performs caller cleanup, tests AL, and reaches the match return only on zero. The helper independently selects candidate inline/heap storage at capacity 0x10, calls the byte-and- length comparator, and normalizes nonzero to one. Description-only difference therefore continues scanning; equal descriptions are required.
  • Held-out NaN negative control: each coordinate executes fucompp; fnstsw ax; test ah,0x44; jp mismatch. Equal yields one tested status bit and does not jump; unordered yields two tested bits and jumps. Thus an otherwise identical event with NaN in any coordinate does not deduplicate, even for identical NaN payloads. This is static control-flow interpretation, not a live fixture.
  • Ownership boundary: ObservedTech append advances by 0x2c; PlayerEvent append/scan advances by 0x74. The ObservedTech copy calls string assignment once; PlayerEvent copy calls it three times. PlayerEvent destruction separately tests all three capacities against 0x10 and calls the bound delete thunk for each long string. Reallocation calls the allocator thunk, advances in 0x2c, invokes old-element destruction and then the delete thunk. These facts reject raw header copying.
  • Constructor/default held-out bytes: the fresh PlayerEvent constructor initializes three SSO strings, IDs/action/location to zero, and loads position words from 0x00af0dc8..d0. A separate section-byte capture gives ff ff 7f 7f three times, i.e. three 0x7f7fffff words.

Archived state and RNG, checked independently

The strict save reader parsed turn3-state.sav with zero resyncs, zero hint failures, zero best-effort fallbacks and empty stderr. Direct tree state (not a reported event counter) contains EvNxID=4, a turn-3 bucket with two elements, and event ID 3: description Research Over Budget, message Research for Waldo Units has gone overbudget., image EVENT_RESEARCH_OVERBUDGET, location 0, action 1, chain ID 0, and three integer position words 2139095039 (0x7f7fffff). These parser leaves are explicitly guessed, so this is archived-value corroboration rather than live ABI proof.

An independent whole-save checksum/audit rebuilt all 609,080 inflated bytes exactly (firstDiff: null), covering 35,394 leaves. Root digest is e9c161e311f8ef8fad6f1aa1903dcf3f under raw float-bit policy and no masks. Actual opaque /Sim/RNG state is present as one 2,503-byte leaf with digest 0978fdf34ff7962f76c2de810dc93e0a; this run did not infer RNG from an administrative draw counter and did not claim that the static helper windows consume it.

Verdict and residuals

Scoped pass for the independent static cross-check. No predeclared prediction failed in this session and no new surprise was observed. Prior failed provenance predictions remain historical and resolved by their Astra decisions; matching repaired bytes does not rewrite them as successes.

Residuals remain material: no live short/long or spare/full-capacity fixture, no same-bucket equal versus description-only-different or NaN runtime fixture, no allocation-failure execution, no live allocator-family compatibility test, no runtime event construction, no original-game differential, no independent replacement, and no integrated replay. The contract currently has no evidence records, so this measurement is not by itself a contract-level verdict or promotion.

Reproduction

python3 verify/results/research-completion-abi-independent/reproduce_run_735fcb8f4876c10285b03fad.py
python3 verify/save-reader/save_reader.py verify/results/saves/turn3-state.sav --dump --json --strict > verify/results/research-completion-abi-independent/run-735fcb8f4876c10285b03fad/turn3-save-reader.json 2> verify/results/research-completion-abi-independent/run-735fcb8f4876c10285b03fad/turn3-save-reader.stderr.txt
python3 verify/state-checksum/state_checksum.py verify/results/saves/turn3-state.sav --json > verify/results/research-completion-abi-independent/run-735fcb8f4876c10285b03fad/turn3-state-checksum.json 2> verify/results/research-completion-abi-independent/run-735fcb8f4876c10285b03fad/turn3-state-checksum.stderr.txt

The reproducer refuses to overwrite its output directory; use a fresh session path for another run.