26 KiB
Independent rollout review
Follow-up verdict — 2026-09-09
Independent execution: openai/gpt-6-astra, logical session review-followup-1;
actual harness session ID unavailable. Selected host/shim baseline passes the reviewed
evidence-integrity scope. Controls require one additional correction (R8 below).
Full asset/live-game and research-replacement acceptance remain unestablished.
This section supersedes the initial available-tree disposition for the source identities below;
the original findings are retained as history.
Independently reproduced measurements
- Tooling: 19/19, publishing: 8/8, config/source scanner: 7/7, controls: 36/36.
Commands:
PYTHONDONTWRITEBYTECODE=1 python3 -m unittest discover -s verify/DIR -p 'test_*.py' -v, with DIR respectivelytooling,publishing,config,campaign. Total 70 tests. Disposable fixture repositories alone perform Git writes/commits. PYTHONDONTWRITEBYTECODE=1 python3 tools/check_agent_config.py --resolved: PASS, actual OpenCode loader and exact model availability. No model execution inferred from availability.campaign/current.jsonselected gate SHA-2568e14e00ee3ce7478ddfdef8de12183451e858c52d78dc28cffbdee47f5d087d8and replay87f92c2b54625ecbca1f3c0a37e57c43d03a4aff488edf4ea8a09a842cd6279d: actual retained bytes match. Gate/replay pass shared structural and publication validators; selection is measurement-only.- Under
/home/alex/.local/share/sots-runs/rollout-host-20260909-c, independently rehashed host executable78b2562ea2c56351f9f6f0f24d271afcced741eadbc148047780e6dfd101940band shim DLL381c91aecf10a9093f651f58c884356cd88e0253e79cb3bab32753fff6612c7e. 524 engine files and 10 RE execution-tool files match both retained snapshots and current source in bytes and modes. All 43 corpus files match their manifest. Engine baseline7741d42fc5e4e761e6449bdaf0e4a61d00036a23; RE baseline3bfde5a70d874a723e797a695bbd847fd82c0aa7; dirty identities come from manifests, not HEAD alone. - Actual
ctest.xmlexactly matches embedded JUnit rows: 59 unique inventory identities, 52 passed, 7 allowed skips, zero failures. Positive final summaries are exactly[43]formars_stream_save,mars_stream_domains,app_turn,app_turn_record. Every required check is true; eight recorded gate commands exit zero; every JUnit output exists without truncation. Skips:game_config_replay,game_data_realdata,game_design_realdata,game_design_census,game_sim_smoke_real_save,mars_text_realdata,mars_vfs_realdata. - Retained replay in
/home/alex/.local/share/sots-runs/rollout-replay-20260909uses turn2 input and turn3 oracle. Recomputed file/inflated/typed-state hashes of all three saves and complete ordered diff list match: 62 state differences, all three equality surfaces false. Status correctly remains measured. This review recalculated saved outputs; it did not rerun the full engine suite or build anything. - Additional independent mutations of the selected manifest rejected missing binary, omitted
execution, duplicate execution, zero corpus count, omitted
outputCompleterequirement and falseoutputComplete. These were in-memory fixtures, leaving selected evidence untouched.
Reproduction recipe for artifact checks: load current pointers with dashboard.pointer, run
gate_valid/replay_valid; rehash both binary paths; compare gate.file_row for every source row
under manifest live roots and retained source/re-tooling; recompute each input_manifest;
parse actual XML with junit_statuses and compare rows, inventory, allowed skips and corpus
summary counts; use standalone_report.save_state and complete state_checksum.diff on retained
pair paths. Temporary driver /tmp/opencode/review_followup_checks.py is disposable; this recipe,
measurements and the content-addressed selected manifests are durable evidence.
Prior findings disposition
| Finding | Current disposition |
|---|---|
| R1 | Closed for inspected gate/package: exact unique execution partition, positive per-test corpus summaries, bound executable and output completeness; negatives reproduced. |
| R2 | Closed under lead's revised scope: reporter has no acceptance status/flag; equality is measured or failed, explicit inputs are hashed, inherited SOTS environment stripped, output retained. Equal-pair and require-match negatives pass. |
| R3 | Closed: --shim is now boolean and uses the recorded snapshot's MinGW toolchain file; actual selected DLL/hash and successful configure/build are retained. Full assets still unavailable. |
| R4 | Closed for approved controls implementation: current source bindings rehashed on evidence/verdict/promotion, immutable inputs/binaries and exact criterion outcomes required; same-HEAD candidate/integrated mutation tests reproduced. |
| R5 | Closed for actual selected host package: shared validator/publication accept the explicit 52/7 partition and expose limitations. |
| R6 | Closed: intact old recovery accepted, corrupt/missing state rejected, freshness retained at handoff/end, resolution-only Astra entry allowed with ordinary workers blocked; tests reproduced. |
| R7 | Closed for reviewed snapshot-copy defect: copy checks bytes/modes on both sides; current and snapshot inventories match. Pre-copy mutation, symlink and build-named-source tests pass. |
Engine accounting review confirms S13 no longer duplicates child P counters; replay writes survive the final fold; runtime-only MT word count measures rejection across twists without changing the explicit serialized state layout. Corpus app tests independently compare committed generator state. No new engine defect established. Snapshot clean-room scanner and scoped tracecmp path checks pass. Gate RE identity explicitly covers execution tooling only, not every campaign/control file.
Completed controls and actual runner evidence
Latest checkpoint campaign/runtime/checkpoints/controls-bootstrap-620eb9d25a6ad2f06b68353c.json
passes basis/identity/artifact validation with recovery freshness disabled. Both R4/R6 surprise
records are resolved by the named lead decisions; no open controls surprises existed at inspection.
Actual smoke campaign/runtime/runs/run-df1472c13f31db3a4d5361f0.json, SHA-256
65dd1c1286742cd4d2401af5a229d7d014f60dd329dadee0dee78c499f61d663, records successful
architecture-review/Astra execution: 24 parsed events, no error events, successful stop,
matching fresh session/actor/model checkpoint and intact artifacts. Actual session
ses_f77cf5062ffeTknrMfnbJMmdNh; emitted model unavailable, honestly recorded as such.
Canonical config/prompt hashes match. Independently rerunning the actual loader with runner overlay
in recorded cwd reproduces effective hash 4cb65613de34e41eabf780e84cad062f14cbfcdd68569dc874514c50fa7baa5a
and agent hash 530f4bfeb28c09d36e45565706247ad9822f7c177161e1ee0a71fcb981a76b92.
This proves the recorded architecture route, not an actual normal Terra worker quantum.
R8 — MEDIUM: launcher cannot start required final integrated verifier
Locations: tools/campaign.py:156-159 (ROLE_STATUS), tools/run_agent.py:87-89.
The verifier role permits only verification, while Campaign.verdict allows integration
and final acceptance requires a fresh verdict after integrated evidence changes.
Thus the prescribed explicit-role launcher cannot schedule that final verifier execution.
Executed reproduction: in a disposable verify/campaign/test_controls.py::Controls fixture,
run verification(), verdict(), transition to integration as lead, add evidence(integrated=True)
as lead, then call run_agent.check_launch(case.c, case.c.load('slice'), 'verifier', 'independent').
Observed ControlError: role cannot launch in this contract status. This is a launcher/lifecycle
mismatch, not a bypass of the acceptance check. Fixture cleaned up; canonical state untouched.
Required resolution/check: lead formally resolves the new review finding; permit the independent verifier route for final integrated-package review with existing independence/checkpoint/surprise guards, or define an equally explicit supported final-verifier launch path. Regression must allow both verification and integration verifier entry and still reject owner/verifier overlap, stale package and unresolved surprises. No implementation correction made by reviewer.
Exact remaining integration checks
- Resolve R8, implement the approved launcher policy, independently rerun its focused regression and refresh source-bound controls measurements after that source change.
- Perform the documented normal noninteractive worker smoke under its actual worker role/model, paired baseline worktrees and intact recovery checkpoint; verify effective config, successful stop/session, fresh end checkpoint and unavailable-versus-emitted model provenance. Existing actual smoke covers architecture-review only; fake-process tests cover worker runner logic.
- Attach the actual integrated controls executable/interpreter, fixture inputs and criterion
result package; obtain an independent verdict over that exact final binding through the supported
lifecycle.
controls-bootstrappresently remainsneeds-revisionwith no evidence array; this review is not a lifecycle promotion or machine verdict. - Full-profile owner assets/trace inputs and research pilot completion-bearing live controls, complete writes/events/allocations/RNG/runtime inputs and original differential remain required for their separate acceptance scopes. Host/shim success and divergent replay do not satisfy them.
Additional reviewed source SHA-256 (gate/engine/reporter identities are retained in selected gate):
| Path | SHA-256 |
|---|---|
tools/campaign.py |
ca4eb2c42e88c1222ec60d6b99c58fffcc72746411c499c4cb0246d899ebbd10 |
tools/run_agent.py |
0f5ccb53dcde59b64fe999768737b6cbd2dabcc1272e5605c1f1ecd86e5d5d2b |
campaign/contract.schema.json |
a2c76ec042ca097a57c3c05c1e519e392d498d387a31808c7d0f70482d43c201 |
verify/campaign/test_controls.py |
656e5938fa5107022b73b32cf6fc05b6ff6e3e9ab239432d72b304b3f6ea1801 |
tools/select_evidence.py |
a6e648a77bcf799cc774ff9d0de83011283049ed0f0c44f6e08857a8410e9463 |
tools/dashboard.py |
c84aeeea6fb7d5e45ccec8882bacc59182149f108edd9eb1585eab9ede99911c |
tools/check_agent_config.py |
93115532cb38105777a5c5c20a9f2a808043370e0d4f560a9c82fe8e800076ec |
No implementation edits, delegates, actual-repository staging/commits, engine builds or lab mutation.
Historical initial available-tree pass
Reviewer: openai/gpt-6-astra, independent review-worker execution, 2026-09-09. Disposition: changes required; rollout acceptance is not established. Workers were still implementing during this review. Findings below apply to the identified snapshots; fixes and final integrated-tree review are pending. No engine builds, lab operations, delegation, staging or commits were performed by this reviewer. Four selected Python tooling tests passed, but the adversarial probes below exposed gaps outside those tests.
Pilot delivered first:
campaign/contracts/research-replacement.json— proposed, validated by campaign CLI.campaign/pilots/research-replacement.md— concrete archived W1 workload, full transitive write boundary, runtime assets/original dependencies, executable acceptance requirements and blockers. No replacement implementation or acceptance claim.
Snapshot identities
Canonical RE HEAD 3bfde5a70d874a723e797a695bbd847fd82c0aa7; engine HEAD
7741d42fc5e4e761e6449bdaf0e4a61d00036a23. Both have concurrent uncommitted rollout work;
HEAD alone does not identify reviewed content. SHA-256 at this pass:
| File | SHA-256 |
|---|---|
tools/gate.py |
700f8fcf1a051a37322cd51e2e1bb774f35cd82e0abc622613dc2b79f37e9a42 |
tools/evidence.py |
8235dc7e3bc23e5a942fb6f80be0693a0d2c66057f24a783a41b512d0d31c137 |
tools/standalone_report.py |
6e4aae37ff5ca82f719d1994f1561cb21840d3cf7978f9a2e0a3f4b1019b71d8 |
tools/dashboard.py |
6d8e116f554f8dae222694e3e37daf7b03922187e5fd841d43ee4cd52c69a49a |
tools/campaign.py |
bd563f43b536841677e49ed9b942e5ed46ad3c53456a921592f281b4f2f3c9a2 |
tools/run_agent.py |
c8aeec109340958bf7850d2a91b4d6b6e53ac0534e97973f6c0869e692a9b7ba |
campaign/contract.schema.json |
ae4796b1f8e8336ddb63e60d774a81ff965461e4b884b5a2488e3865c6a6c5e0 |
opencode.json |
d24be7fcee57d5eb04c6b7189822a2ca055dd4a70daafbea67f09a120945d49f |
engine src/app/turn.cpp |
a626ddf80a340de50b70f09bf91c0180d255c533e65edce1532d807fcf5109b4 |
engine src/mars/rng/mt19937.cpp |
0c60775e9437b8eaad49f67a96001c81947c8d9e6d2472e28b2b490515bac95b |
engine tests/app/test_turn.cpp |
6d45b1d2d065528f14ae83d95314f2b6559e53cc23f1b9e2531c7c09b4b97eab |
Findings requiring correction
R1 — HIGH: gate accepts incomplete execution and missing binary
Locations: tools/gate.py:137-153,162-164.
Discovery is compared with the 59-name inventory, but actual JUnit cases are not required to
cover that inventory. Only four corpus names must appear in passed. There is no nonempty
binary requirement. corpusNonzero checks that stdout lacks "0 save"; empty output passes,
and CTest --output-on-failure normally suppresses successful test output anyway.
Executed reproduction: import gate; isolate fake engine/.git and one .sav in a temporary directory; mock source_manifest/copy_snapshot and command execution (no build); return all 59 names from ctest_names, write JUnit with only mars_stream_save, mars_stream_domains, app_turn, app_turn_record, return exit 0 and empty stdout for commands, produce no binary. Call main with explicit engine/corpus/new out. Observed:
gate incomplete JUnit/no output/no binary: 0 passed 4 of 59 binary= {} corpusNonzero= True
This tests the validator boundary, not actual CTest behavior. A truncated/misconfigured runner result must fail rather than rely on CTest usually producing complete output.
Required fix/check: exact JUnit identity partition (passed/skipped/failed), no duplicates or unknown/missing tests, explicit positive corpus counts from machine-readable results, required binary artifact existence/hash, and negative tests for each omission. Host allowances must be distinguished from required executed tests. Do not infer a positive count from absent text.
R2 — HIGH: reporter labels equality-only, unbound execution as accepted
Locations: tools/standalone_report.py:68-74,82-105; tools/evidence.py:12-16.
Reporter requires only schema/status/binary hash from provenance. It does not validate full
gate profile/checks/source/input completeness, required phase execution, missing runtime inputs,
or workload coverage. Caller may supply the same save as input and oracle. --accept promotes
three equal digests without any independent execution proof or contract binding.
Executed reproduction: use the real turn3 save as both input and oracle, a temporary dummy
binary with matching minimal {schema,status,profile:"host",binary:{sha256}} manifest, and mock
only subprocess.run to copy input to --out and return 0. Actual reader/reconstruction/digest
checks run. Observed reporter identical input/oracle, no execution: 0 accepted.
This is a validation-boundary test, not a claim that current sots_turn is a copy-only program.
--roundtrip checks conservation first but does not skip RunStrategicTurn (main.cpp:214-222,318).
Required fix/check: keep equality measurement usable, but scoped acceptance requires an explicit workload contract, full source-bound provenance, positive required phase/branch counts, known input/dependency completeness and independent verification/integration gates. No-op and empty/partial provenance fixtures must fail acceptance. Persist/hash all consumed external engine arguments (data/commands), verify files unchanged over execution, and retain output saves or a durable reproducible package: currently output files are removed with TemporaryDirectory.
R3 — HIGH: full gate's toolchain argument rejects valid files and accepts directories
Locations: tools/gate.py:100,107-108,154-156.
--shim must be a directory, then that directory is passed as CMAKE_TOOLCHAIN_FILE.
A normal existing toolchain .cmake file fails argument validation; a directory cannot supply
the intended toolchain file. Full gate is required for acceptance but this route is unusable
as a conventional CMake toolchain interface. Inspection finding; no build attempted.
Reproduction: supply a valid engine/corpus/data and existing toolchain file to --shim;
observe parser rejection at line 108. A directory passes validation then is the literal
-DCMAKE_TOOLCHAIN_FILE=<directory> in the configure command.
Required fix/check: define/document --toolchain file versus shim source/build artifact; resolve and hash it, validate configure command in a unit test. Bind produced shim binary as well as host binary. Full test inputs also include SOTS_M1_TRACE, SOTS_SAVES_JSON, SOTS_GOB_DIR; current gate inherits these without recording hashes and does not provide an explicit interface.
R4 — HIGH: campaign acceptance source binding is only baseline path+commit
Locations: campaign/contract.schema.json:26-32; tools/campaign.py:299-307,325-335,365-369;
tools/run_agent.py:68-84,173,227.
Runner records actual source manifests, but lifecycle evidence compares source only to the
contract's baseline {path,commit}. Neither evidence nor verifier verdict references the runner's
actual source digest or a candidate/integrated tree digest. The contract basis is a hash of
task metadata; it is unchanged when uncommitted implementation changes. Integrated status is a
lead-supplied boolean and axis coverage, with no connection to which integrated source was built.
Reproduction by data flow: create evidence for source A under a baseline, obtain a bound verdict, then change implementation bytes without changing HEAD or artifact/contract JSON. check_evidence and check_verdict have no source-content input to detect source B. Source_after in run records does not enter these checks. This is a stale-evidence gap even with honest actors; it is distinct from the documented limitation that editable role strings are not authentication.
Required fix/check: bind candidate/integrated source-manifest digests, build binary and immutable inputs through evidence and independent verdict; transition must validate that binding. Use per-criterion results or a typed acceptance package rather than treating any hashed artifact with the same axis label as execution proof. Add a test that changes source content at the same HEAD and rejects its old verifier/integration result.
Late-published README (lines 27-37) explicitly assigns dirty-source manifest evaluation to the independent reviewer, not the CLI. Thus this is an acceptance-robustness gap requiring a lead decision, not a claim that the controls author promised a machine interpretation of arbitrary criteria. The human procedure must at minimum compare the actual candidate/integrated manifests; baseline equality in CLI output cannot be presented as that check.
R5 — MEDIUM: publication rejects valid host-profile skips
Locations: tools/dashboard.py:93-101; tools/gate.py:146-150.
Gate's expected list is all discovered tests, including permitted host skips. Publisher demands
expected <= passed, contradicting host skip allowances. A correct host result cannot be
published as scoped host evidence.
Executed reproduction: gate_valid with host status passed, nonempty source/input/binary
identities, tests expected=[a,asset], passed=[a], skipped=[asset], failed=[] reports
required tests not executed. Same outcome follows for the current 59/52/7 shape.
Required fix/check: shared schema distinguishes required executions and approved skips; validate exact partition and profile policy. Host results must visibly publish limitations, not be promoted to full acceptance. Gate currently initializes limitations=[] and never fills it.
R6 — MEDIUM: recovery rejects old durable checkpoints and blocks resolver entry
Locations: tools/run_agent.py:87-99; tools/campaign.py:266-278.
check_launch reuses end-of-quantum freshness validation, requiring a checkpoint younger than 15 minutes even when starting recovery from a successfully completed earlier session. Next-day resume therefore needs a fabricated new checkpoint before an agent can read/recover the old one. Separately, open_surprises is checked unconditionally for every role, so even a resolver cannot launch to investigate an unresolved surprise (despite resolver status mapping to blocked/revision).
Reproduction: a structurally valid, matching checkpoint older than 900 seconds fails check_launch; a resolver/blocked contract with an open surprise fails before role-specific work. Inspection/data-flow finding; final controls tests were not yet published.
Required fix/check: distinguish valid durable recovery state from a fresh end-run checkpoint; verify identities/basis/artifact hashes on recovery, and require freshness only after this quantum starts. Permit an explicitly assigned resolver's bounded investigation while affected workers stay blocked, or document a separate operational resolver launch path that does not pre-resolve evidence.
R7 — MEDIUM: built snapshot is not itself verified against the source manifest
Locations: tools/gate.py:29-49,57-61,119-129,157-158.
Manifest hashes live source, then copy_snapshot copies paths without checking copied hashes;
end check hashes live source again. A concurrent A -> B -> A change can copy B while both
live manifests attest A. Filesystem modes and symlink identity/targets are not bound; every
path component starting with build is excluded, including possible real source directories.
RE tools/inventory are recorded once but then used live without a post-check.
Reproduction: hash source A, write B before copy_snapshot, restore A before after-manifest; before == after but destination is B. This needs no Git commit or modification of the gate.
Required fix/check: hash the copied snapshot and compare its exact file/mode/link manifest before building; reject unsupported escaping links/submodules and define explicit exclusions. Snapshot/hash tools and test inventory actually executed. Baseline-pinned isolated worktrees and leases remain necessary; a before/after check alone does not establish which bytes were built.
Engine / lead configuration observations
- Engine S13 no longer aggregates child P01..P12 counters; the existing final fold can count child operations once. CountingRandom measures rejection words with the monotone MT counter; seed/load reset it, and save_state's explicit layout excludes it. New app test replays reported words against persisted generator state; new MT test exercises rejection across a twist. No defect established in these reviewed changes. Integrated independent execution remains pending.
- New lead workflow/architecture clearly distinguish original assistance, partial/full/scoped evidence, unsupported inputs and roles. Explicit Astra/Terra/5.5 routing and 40-step bounds appear in config; runner checks model availability and aborts on known mismatch, with no fallback.
- Config permissions primarily guard architecture edit paths and deny worker task calls. They do not constrain arbitrary shell or enforce exact string-array contract scope, as workflow correctly acknowledges. Reviewer has not exercised OpenCode's effective merged configuration; runner overlay binds model/steps but only canonical config bytes are hashed, not expanded prompt file bytes or all effective configuration. Include these in final provenance/config validation.
- Controls lock is canonical and serialized; leases require matching token/owner and forbid automatic stealing. Active-run worktree reservations are checked under the same lock. No concurrency acceptance claim: independent contention/crash-recovery tests are still required.
Research evidence surprises sent to lead
verify/results/shim/cr/cr-R1.log:75-94 adds seven drawsites/twelve probes to the narrative's
six interceptors: 25 sites total. Lines 109-115 and raw trace args show runtime CW 0x127f,
not initialization 0x027f. Archive save hashes and R1's exact 16-leaf residual independently
reproduced. These corrections and runtime asset text path are detailed in the pilot and checkpoint.
Reproduction record and remaining review
Executed with PYTHONDONTWRITEBYTECODE=1:
python3 tools/campaign.py --state-root /home/alex/sots-re validate research-replacement: pass.- Selected unittest methods in
verify.tooling.test_tooling.ToolingTests: inventory_is_explicit_and_unique, passed_gate_required, zero_discovered_tests_cannot_match_inventory, stale_report_output_is_rejected_before_execution (each prefixedtest_): 4/4 pass. /tmp/opencode/review_rollout_probes.py: isolated boundary mocks described under R1/R2/R5; output recorded above. Temporary fixture files are disposable; this report preserves recipes/results.- Raw save diagnostic command in pilot: R1 exact-bit/unmasked 16 differences.
No full local worker suite: one existing tooling test creates a temporary Git commit, excluded under this review assignment. No reviewer builds while engine source was being changed. At last inspection controls README/test directory were not yet present; schema and initial tools were inspected, not their promised completed validation package.
Final refresh at 2026-09-09T21:30:34Z: campaign/README.md appeared and was read. Its schema/API
matches this pilot and documents manual criterion evaluation, baseline-versus-dirty-source limits,
and the current fresh-checkpoint/no-open-surprise launch rules. It does not close R4's automatic
source-drift gap or R6's recovery/resolver operational concern. Completed controls test execution
and integrated evidence review remain pending; repository-creating tests need explicit permission
for their temporary Git commits or a non-committing equivalent under this review assignment.
Next review: after owners fix/finish, resume from review-worker-state.md, check changed hashes, re-run these negative cases and finalized controls/publishing tests, then independently review one stable integrated source-bound gate/replay package. Until then Task B is an initial adversarial review with final integration review pending, not an approval of workers' future edits.