sots-re/verify/results/research-completion-abi-independent/result-run-7d85d45cb2196e07025e5096.md

5.7 KiB

Independent ABI evidence-array review — scoped pass

Session run-7d85d45cb2196e07025e5096; actor research-abi-independent; requested model openai/gpt-5.6-sol.

Classification and decision

PASS for the two-criterion, non-integrated evidence array in its stated static scope. This is an independent static reproduction plus archived-state inspection. It is not original-assisted game execution, partial/full live comparison, independent replacement, or integrated replay. It proves neither live allocator-family safety nor replacement acceptance.

The pre-experiment plan fixed identity/execution falsifiers and required branch/distinct-state exposures. None fired. Historical failed provenance predictions remain failures resolved by Astra; the repaired result does not rewrite that history.

Identity and evidence integrity

The assigned engine worktree remained at HEAD 7741d42fc5e4e761e6449bdaf0e4a61d00036a23, common Git directory /home/alex/sots-engine/.git, source binding ccd8e02083e8d2e2b3e97976ace2273c8f924dfc02a39e919004eaf3544c50fd. The assigned RE worktree remained at HEAD 3bfde5a70d874a723e797a695bbd847fd82c0aa7, common Git directory /home/alex/sots-re/.git, source binding 6696fd5201e144843617cbf6d78b41b5287ad5dcc9fa1e8aaa861d52b64e72e8. Both bindings exactly match both evidence records; dirty inventories were pre-existing and neither assigned worktree was edited.

dumps/sots.exe is 7,898,624 bytes, SHA-256 970b7de729956a53094c7eb98aba4270aee98e2fed5daf0d39e290013c90c841. /usr/bin/objdump SHA-256 is 1eaaef2e7f57c4c7f69115c495e2466f5a8c8e5f3bc42221d092382f30f9d4cd. The archived save SHA-256 is 978041acd168b56ed8eb3f5e42e78d5e70eae6e6517d75e659a5eb7ca3d60921. All declared evidence input/outcome hashes rehashed exactly, and campaign validation passed.

Positive execution and static branch exposure

The fresh verifier-owned reproducer ran 22 objdump windows. Every process exited zero with nonempty stdout and empty stderr; no window skipped. All 22 names, stream sizes, hashes, stderr sizes and return codes exactly equal the prior independent manifest. Seven narrow/wide pairs preserve every preceding instruction row and widen terminal c2 to five ownership c2 04 00, one string allocator c2 08 00, and one dedup c2 08 00 encoding.

Raw instructions independently retain distinct 0x2c ObservedTech and 0x74 PlayerEvent stepping. ObservedTech growth calls allocator thunk 0x00924fb6, deep-copy helper 0x0079a150, element destruction and delete thunk 0x00924faa. PlayerEvent copy calls string assignment 0x00425430 three times. Destructor 0x0061ae90 has three separate capacity-0x10 guards and three delete-thunk calls. Thus the evidence rejects raw string-header transfer.

The dedup window executes three fucompp; fnstsw; test ah,0x44; jp mismatch sequences, so unordered coordinates take mismatch rather than equality. It compares message/image, passes description to 0x0046f8c0, and reaches the matching-element return only for the helper's equality result. The helper/callee expose inline/heap selection and byte/length inequality. These are static branch facts, not live equal/description-different/NaN executions.

Held-out instrument challenge

A verifier-owned stdlib PE parser independently mapped virtual addresses through the image section table without objdump. Direct .text bytes match all seven repaired return encodings and the plain c3 return at 0x0085630c; the byte immediately before that return is 5d, providing a negative boundary control. Direct .data bytes at 0x00af0dc8 are three ffff7f7f words. This supports the repaired terminal-byte and constructor-default claims without relying solely on decoder formatting.

Actual archived state and RNG

Fresh strict save parsing exited zero with nonempty JSON and empty stderr: 37,847 items, 2,453 frames, zero resyncs, zero hint failures and zero best-effort fallback. Direct tree traversal found an EvNxID value 4, a turn-3 bucket count 2, and event ID 3 with description Research Over Budget, message Research for Waldo Units has gone overbudget., image EVENT_RESEARCH_OVERBUDGET, location 0, action 1, chain 0 and three integer position words 2139095039 (0x7f7fffff). Guessed leaves are archived corroboration, not live ABI proof.

The separate whole-save tool rebuilt all 609,080 inflated bytes with firstDiff: null; root digest is e9c161e311f8ef8fad6f1aa1903dcf3f. Actual /Sim/RNG is one 2,503-byte leaf with digest 0978fdf34ff7962f76c2de810dc93e0a. RNG state was inspected directly rather than inferred from a counter; no RNG operation was executed.

Residuals and reproduction

Unexecuted distinct runtime states remain empty/short/long strings, spare/full capacity, new/existing tech, duplicate/nonduplicate and description-only-different events, finite/NaN coordinates, empty/nonempty nested vectors, normal/unwind destruction, allocation failure, and prune/get-or-create boundaries. There is no live original-game differential or allocator-family compatibility check. Whole-save equality is only an audit of this one archived save, not general workload coverage.

Executed commands/results:

python3 verify/results/research-completion-abi-independent/reproduce_run_7d85d45cb2196e07025e5096.py
# exit 0; manifest: 22/22 positive, all boundaries/direct-PE/state checks true
python3 tools/campaign.py --state-root /home/alex/sots-re validate
# success: four contract IDs listed

Exact next action: record a scoped independent PASS verdict over the current non-integrated evidence array, then checkpoint this session; lead may subsequently decide whether to transition/package for integration, where a fresh verifier must attest the integrated evidence array.