sots-re/verify/results/research-completion-abi-independent/result-run-735fcb8f4876c10285b03fad.md

106 lines
7.6 KiB
Markdown

# Independent ABI verification result — scoped static pass
Session `run-735fcb8f4876c10285b03fad`; verifier `research-abi-independent`; model
`openai/gpt-5.6-sol`.
## Scope classification
This is an **independent static reproduction plus archived-state check**. It is not original-assisted
runtime execution, partial or full live comparison, independent replacement, or integrated replay.
No game, allocator, constructor, destructor, exception path, or RNG operation executed. Accordingly,
this result supports the contract's `independent-cross-check` criterion only within its static scope;
it does not establish live allocator safety or replacement acceptance.
## Preconditions and falsifiers
Before execution I required exact paired HEAD/common-directory/source bindings, executable and tool
identity, positive nonempty execution, empty stderr, complete terminal bytes, and no unexpected skip.
Semantic falsifiers were collapsed `0x2c`/`0x74` strides, raw string-header transfer, missing
long-string frees, description omission from equality, ordinary equality accepting unordered x87
coordinates, or archived event values existing only as counters rather than direct tree state.
Required states remain distinct: empty/short/long strings; equal/description-only-different events;
finite/unordered coordinates; empty/spare/full capacity; new/existing tech; duplicate/nonduplicate;
empty/nonempty nested vectors; normal destruction/unwind; and prune/get-or-create boundaries. The raw
instructions expose these branch alternatives, but this run did not execute those live states.
## Identities and positive execution
The engine pair is HEAD `7741d42fc5e4e761e6449bdaf0e4a61d00036a23`, common Git directory
`/home/alex/sots-engine/.git`, binding
`ccd8e02083e8d2e2b3e97976ace2273c8f924dfc02a39e919004eaf3544c50fd`. The RE pair is HEAD
`3bfde5a70d874a723e797a695bbd847fd82c0aa7`, common directory `/home/alex/sots-re/.git`, binding
`6696fd5201e144843617cbf6d78b41b5287ad5dcc9fa1e8aaa861d52b64e72e8`. Their dirty inventories
were pre-existing. Input `dumps/sots.exe` is 7,898,624 bytes, SHA-256
`970b7de729956a53094c7eb98aba4270aee98e2fed5daf0d39e290013c90c841`. `/usr/bin/objdump` is GNU
2.38, SHA-256 `1eaaef2e7f57c4c7f69115c495e2466f5a8c8e5f3bc42221d092382f30f9d4cd`.
The owned reproducer executed 22 windows. All returned zero with nonempty stdout and empty stderr;
there were no skips. All six repaired ownership narrow/wide stream hashes exactly equal the analyst
manifest. Narrow streams end in `c2`; wide streams establish five `c2 04 00` encodings and one
`c2 08 00`, with every preceding instruction row equal. Four originally complete-stop controls also
match the analyst hashes. The independent dedup pair likewise matches its repaired archive hashes and
widens `c2` to `c2 08 00`. `manifest.json` records every argv, stream size/hash and result rather than
using a count as proof.
## Independent challenges and key ABI observations
* **Historical-stop ablation:** all seven narrow/wide tests reproduce objdump's truncation behavior.
Therefore decoded `ret` text from a narrow stream is not accepted as complete-byte provenance.
* **Held-out complete-stop boundary:** stopping the ObservedTech constructor at `0x0085630c` omits
its `ret`; stopping at `0x0085630d` adds exactly `85630c: c3 ret`. This challenges the assumption
that every historical stop suffered the `ret imm16` effect and confirms the plain-ret control.
* **Description equality:** the fresh dedup window passes both `+0x08` description objects to
`0x0046f8c0`, performs caller cleanup, tests `AL`, and reaches the match return only on zero. The
helper independently selects candidate inline/heap storage at capacity `0x10`, calls the byte-and-
length comparator, and normalizes nonzero to one. Description-only difference therefore continues
scanning; equal descriptions are required.
* **Held-out NaN negative control:** each coordinate executes `fucompp; fnstsw ax; test ah,0x44; jp
mismatch`. Equal yields one tested status bit and does not jump; unordered yields two tested bits
and jumps. Thus an otherwise identical event with NaN in any coordinate does not deduplicate, even
for identical NaN payloads. This is static control-flow interpretation, not a live fixture.
* **Ownership boundary:** ObservedTech append advances by `0x2c`; PlayerEvent append/scan advances by
`0x74`. The ObservedTech copy calls string assignment once; PlayerEvent copy calls it three times.
PlayerEvent destruction separately tests all three capacities against `0x10` and calls the bound
delete thunk for each long string. Reallocation calls the allocator thunk, advances in `0x2c`,
invokes old-element destruction and then the delete thunk. These facts reject raw header copying.
* **Constructor/default held-out bytes:** the fresh PlayerEvent constructor initializes three SSO
strings, IDs/action/location to zero, and loads position words from `0x00af0dc8..d0`. A separate
section-byte capture gives `ff ff 7f 7f` three times, i.e. three `0x7f7fffff` words.
## Archived state and RNG, checked independently
The strict save reader parsed `turn3-state.sav` with zero resyncs, zero hint failures, zero best-effort
fallbacks and empty stderr. Direct tree state (not a reported event counter) contains `EvNxID=4`, a
turn-3 bucket with two elements, and event ID 3: description `Research Over Budget`, message
`Research for Waldo Units has gone overbudget.`, image `EVENT_RESEARCH_OVERBUDGET`, location 0,
action 1, chain ID 0, and three integer position words `2139095039` (`0x7f7fffff`). These parser
leaves are explicitly guessed, so this is archived-value corroboration rather than live ABI proof.
An independent whole-save checksum/audit rebuilt all 609,080 inflated bytes exactly (`firstDiff:
null`), covering 35,394 leaves. Root digest is `e9c161e311f8ef8fad6f1aa1903dcf3f` under raw float-bit
policy and no masks. Actual opaque `/Sim/RNG` state is present as one 2,503-byte leaf with digest
`0978fdf34ff7962f76c2de810dc93e0a`; this run did not infer RNG from an administrative draw counter
and did not claim that the static helper windows consume it.
## Verdict and residuals
**Scoped pass for the independent static cross-check.** No predeclared prediction failed in this
session and no new surprise was observed. Prior failed provenance predictions remain historical and
resolved by their Astra decisions; matching repaired bytes does not rewrite them as successes.
Residuals remain material: no live short/long or spare/full-capacity fixture, no same-bucket equal
versus description-only-different or NaN runtime fixture, no allocation-failure execution, no live
allocator-family compatibility test, no runtime event construction, no original-game differential,
no independent replacement, and no integrated replay. The contract currently has no evidence records,
so this measurement is not by itself a contract-level verdict or promotion.
## Reproduction
```sh
python3 verify/results/research-completion-abi-independent/reproduce_run_735fcb8f4876c10285b03fad.py
python3 verify/save-reader/save_reader.py verify/results/saves/turn3-state.sav --dump --json --strict > verify/results/research-completion-abi-independent/run-735fcb8f4876c10285b03fad/turn3-save-reader.json 2> verify/results/research-completion-abi-independent/run-735fcb8f4876c10285b03fad/turn3-save-reader.stderr.txt
python3 verify/state-checksum/state_checksum.py verify/results/saves/turn3-state.sav --json > verify/results/research-completion-abi-independent/run-735fcb8f4876c10285b03fad/turn3-state-checksum.json 2> verify/results/research-completion-abi-independent/run-735fcb8f4876c10285b03fad/turn3-state-checksum.stderr.txt
```
The reproducer refuses to overwrite its output directory; use a fresh session path for another run.