sots-re/campaign/rollout/independent-review.md

376 lines
26 KiB
Markdown

# Independent rollout review
## Follow-up verdict — 2026-09-09
Independent execution: openai/gpt-6-astra, logical session `review-followup-1`;
actual harness session ID unavailable. **Selected host/shim baseline passes the reviewed
evidence-integrity scope. Controls require one additional correction (R8 below).**
Full asset/live-game and research-replacement acceptance remain unestablished.
This section supersedes the initial available-tree disposition for the source identities below;
the original findings are retained as history.
### Independently reproduced measurements
- Tooling: **19/19**, publishing: **8/8**, config/source scanner: **7/7**, controls: **36/36**.
Commands: `PYTHONDONTWRITEBYTECODE=1 python3 -m unittest discover -s verify/DIR -p 'test_*.py' -v`,
with DIR respectively `tooling`, `publishing`, `config`, `campaign`. Total **70 tests**.
Disposable fixture repositories alone perform Git writes/commits.
- `PYTHONDONTWRITEBYTECODE=1 python3 tools/check_agent_config.py --resolved`: PASS, actual
OpenCode loader and exact model availability. No model execution inferred from availability.
- `campaign/current.json` selected gate SHA-256
`8e14e00ee3ce7478ddfdef8de12183451e858c52d78dc28cffbdee47f5d087d8` and replay
`87f92c2b54625ecbca1f3c0a37e57c43d03a4aff488edf4ea8a09a842cd6279d`: actual retained bytes match.
Gate/replay pass shared structural and publication validators; selection is measurement-only.
- Under `/home/alex/.local/share/sots-runs/rollout-host-20260909-c`, independently rehashed host
executable `78b2562ea2c56351f9f6f0f24d271afcced741eadbc148047780e6dfd101940b` and shim DLL
`381c91aecf10a9093f651f58c884356cd88e0253e79cb3bab32753fff6612c7e`.
**524 engine files and 10 RE execution-tool files** match both retained snapshots and current
source in bytes and modes. All **43** corpus files match their manifest. Engine baseline
`7741d42fc5e4e761e6449bdaf0e4a61d00036a23`; RE baseline
`3bfde5a70d874a723e797a695bbd847fd82c0aa7`; dirty identities come from manifests, not HEAD alone.
- Actual `ctest.xml` exactly matches embedded JUnit rows: **59 unique inventory identities,
52 passed, 7 allowed skips, zero failures**. Positive final summaries are exactly `[43]` for
`mars_stream_save`, `mars_stream_domains`, `app_turn`, `app_turn_record`. Every required check
is true; eight recorded gate commands exit zero; every JUnit output exists without truncation.
Skips: `game_config_replay`, `game_data_realdata`, `game_design_realdata`, `game_design_census`,
`game_sim_smoke_real_save`, `mars_text_realdata`, `mars_vfs_realdata`.
- Retained replay in `/home/alex/.local/share/sots-runs/rollout-replay-20260909` uses turn2 input
and turn3 oracle. Recomputed file/inflated/typed-state hashes of all three saves and complete
ordered diff list match: **62 state differences**, all three equality surfaces false.
Status correctly remains **measured**. This review recalculated saved outputs; it did not rerun
the full engine suite or build anything.
- Additional independent mutations of the selected manifest rejected missing binary, omitted
execution, duplicate execution, zero corpus count, omitted `outputComplete` requirement and
false `outputComplete`. These were in-memory fixtures, leaving selected evidence untouched.
Reproduction recipe for artifact checks: load current pointers with `dashboard.pointer`, run
`gate_valid`/`replay_valid`; rehash both binary paths; compare `gate.file_row` for every source row
under manifest live roots and retained `source`/`re-tooling`; recompute each `input_manifest`;
parse actual XML with `junit_statuses` and compare rows, inventory, allowed skips and corpus
summary counts; use `standalone_report.save_state` and complete `state_checksum.diff` on retained
pair paths. Temporary driver `/tmp/opencode/review_followup_checks.py` is disposable; this recipe,
measurements and the content-addressed selected manifests are durable evidence.
### Prior findings disposition
| Finding | Current disposition |
|---|---|
| R1 | Closed for inspected gate/package: exact unique execution partition, positive per-test corpus summaries, bound executable and output completeness; negatives reproduced. |
| R2 | Closed under lead's revised scope: reporter has no acceptance status/flag; equality is measured or failed, explicit inputs are hashed, inherited SOTS environment stripped, output retained. Equal-pair and require-match negatives pass. |
| R3 | Closed: `--shim` is now boolean and uses the recorded snapshot's MinGW toolchain file; actual selected DLL/hash and successful configure/build are retained. Full assets still unavailable. |
| R4 | Closed for approved controls implementation: current source bindings rehashed on evidence/verdict/promotion, immutable inputs/binaries and exact criterion outcomes required; same-HEAD candidate/integrated mutation tests reproduced. |
| R5 | Closed for actual selected host package: shared validator/publication accept the explicit 52/7 partition and expose limitations. |
| R6 | Closed: intact old recovery accepted, corrupt/missing state rejected, freshness retained at handoff/end, resolution-only Astra entry allowed with ordinary workers blocked; tests reproduced. |
| R7 | Closed for reviewed snapshot-copy defect: copy checks bytes/modes on both sides; current and snapshot inventories match. Pre-copy mutation, symlink and build-named-source tests pass. |
Engine accounting review confirms S13 no longer duplicates child P counters; replay writes survive
the final fold; runtime-only MT word count measures rejection across twists without changing the
explicit serialized state layout. Corpus app tests independently compare committed generator state.
No new engine defect established. Snapshot clean-room scanner and scoped tracecmp path checks pass.
Gate RE identity explicitly covers execution tooling only, not every campaign/control file.
### Completed controls and actual runner evidence
Latest checkpoint `campaign/runtime/checkpoints/controls-bootstrap-620eb9d25a6ad2f06b68353c.json`
passes basis/identity/artifact validation with recovery freshness disabled. Both R4/R6 surprise
records are resolved by the named lead decisions; no open controls surprises existed at inspection.
Actual smoke `campaign/runtime/runs/run-df1472c13f31db3a4d5361f0.json`, SHA-256
`65dd1c1286742cd4d2401af5a229d7d014f60dd329dadee0dee78c499f61d663`, records successful
**architecture-review/Astra** execution: 24 parsed events, no error events, successful stop,
matching fresh session/actor/model checkpoint and intact artifacts. Actual session
`ses_f77cf5062ffeTknrMfnbJMmdNh`; emitted model unavailable, honestly recorded as such.
Canonical config/prompt hashes match. Independently rerunning the actual loader with runner overlay
in recorded cwd reproduces effective hash `4cb65613de34e41eabf780e84cad062f14cbfcdd68569dc874514c50fa7baa5a`
and agent hash `530f4bfeb28c09d36e45565706247ad9822f7c177161e1ee0a71fcb981a76b92`.
This proves the recorded architecture route, not an actual normal Terra worker quantum.
### R8 — MEDIUM: launcher cannot start required final integrated verifier
**Locations:** `tools/campaign.py:156-159` (`ROLE_STATUS`), `tools/run_agent.py:87-89`.
The verifier role permits only `verification`, while `Campaign.verdict` allows `integration`
and final acceptance requires a fresh verdict after integrated evidence changes.
Thus the prescribed explicit-role launcher cannot schedule that final verifier execution.
**Executed reproduction:** in a disposable `verify/campaign/test_controls.py::Controls` fixture,
run `verification()`, `verdict()`, transition to `integration` as lead, add `evidence(integrated=True)`
as lead, then call `run_agent.check_launch(case.c, case.c.load('slice'), 'verifier', 'independent')`.
Observed `ControlError: role cannot launch in this contract status`. This is a launcher/lifecycle
mismatch, not a bypass of the acceptance check. Fixture cleaned up; canonical state untouched.
**Required resolution/check:** lead formally resolves the new review finding; permit the independent
verifier route for final integrated-package review with existing independence/checkpoint/surprise
guards, or define an equally explicit supported final-verifier launch path. Regression must allow
both verification and integration verifier entry and still reject owner/verifier overlap, stale
package and unresolved surprises. No implementation correction made by reviewer.
### Exact remaining integration checks
1. Resolve R8, implement the approved launcher policy, independently rerun its focused regression
and refresh source-bound controls measurements after that source change.
2. Perform the documented normal noninteractive worker smoke under its actual worker role/model,
paired baseline worktrees and intact recovery checkpoint; verify effective config, successful
stop/session, fresh end checkpoint and unavailable-versus-emitted model provenance. Existing
actual smoke covers architecture-review only; fake-process tests cover worker runner logic.
3. Attach the actual integrated controls executable/interpreter, fixture inputs and criterion
result package; obtain an independent verdict over that exact final binding through the supported
lifecycle. `controls-bootstrap` presently remains `needs-revision` with **no evidence array**;
this review is not a lifecycle promotion or machine verdict.
4. Full-profile owner assets/trace inputs and research pilot completion-bearing live controls,
complete writes/events/allocations/RNG/runtime inputs and original differential remain required
for their separate acceptance scopes. Host/shim success and divergent replay do not satisfy them.
Additional reviewed source SHA-256 (gate/engine/reporter identities are retained in selected gate):
| Path | SHA-256 |
|---|---|
| `tools/campaign.py` | `ca4eb2c42e88c1222ec60d6b99c58fffcc72746411c499c4cb0246d899ebbd10` |
| `tools/run_agent.py` | `0f5ccb53dcde59b64fe999768737b6cbd2dabcc1272e5605c1f1ecd86e5d5d2b` |
| `campaign/contract.schema.json` | `a2c76ec042ca097a57c3c05c1e519e392d498d387a31808c7d0f70482d43c201` |
| `verify/campaign/test_controls.py` | `656e5938fa5107022b73b32cf6fc05b6ff6e3e9ab239432d72b304b3f6ea1801` |
| `tools/select_evidence.py` | `a6e648a77bcf799cc774ff9d0de83011283049ed0f0c44f6e08857a8410e9463` |
| `tools/dashboard.py` | `c84aeeea6fb7d5e45ccec8882bacc59182149f108edd9eb1585eab9ede99911c` |
| `tools/check_agent_config.py` | `93115532cb38105777a5c5c20a9f2a808043370e0d4f560a9c82fe8e800076ec` |
No implementation edits, delegates, actual-repository staging/commits, engine builds or lab mutation.
## Historical initial available-tree pass
Reviewer: openai/gpt-6-astra, independent review-worker execution, 2026-09-09.
**Disposition: changes required; rollout acceptance is not established.** Workers were still
implementing during this review. Findings below apply to the identified snapshots; fixes and
final integrated-tree review are pending. No engine builds, lab operations, delegation, staging
or commits were performed by this reviewer. Four selected Python tooling tests passed, but the
adversarial probes below exposed gaps outside those tests.
Pilot delivered first:
- `campaign/contracts/research-replacement.json` — **proposed**, validated by campaign CLI.
- `campaign/pilots/research-replacement.md` — concrete archived W1 workload, full transitive
write boundary, runtime assets/original dependencies, executable acceptance requirements
and blockers. No replacement implementation or acceptance claim.
## Snapshot identities
Canonical RE HEAD `3bfde5a70d874a723e797a695bbd847fd82c0aa7`; engine HEAD
`7741d42fc5e4e761e6449bdaf0e4a61d00036a23`. Both have concurrent uncommitted rollout work;
HEAD alone does not identify reviewed content. SHA-256 at this pass:
| File | SHA-256 |
|---|---|
| `tools/gate.py` | `700f8fcf1a051a37322cd51e2e1bb774f35cd82e0abc622613dc2b79f37e9a42` |
| `tools/evidence.py` | `8235dc7e3bc23e5a942fb6f80be0693a0d2c66057f24a783a41b512d0d31c137` |
| `tools/standalone_report.py` | `6e4aae37ff5ca82f719d1994f1561cb21840d3cf7978f9a2e0a3f4b1019b71d8` |
| `tools/dashboard.py` | `6d8e116f554f8dae222694e3e37daf7b03922187e5fd841d43ee4cd52c69a49a` |
| `tools/campaign.py` | `bd563f43b536841677e49ed9b942e5ed46ad3c53456a921592f281b4f2f3c9a2` |
| `tools/run_agent.py` | `c8aeec109340958bf7850d2a91b4d6b6e53ac0534e97973f6c0869e692a9b7ba` |
| `campaign/contract.schema.json` | `ae4796b1f8e8336ddb63e60d774a81ff965461e4b884b5a2488e3865c6a6c5e0` |
| `opencode.json` | `d24be7fcee57d5eb04c6b7189822a2ca055dd4a70daafbea67f09a120945d49f` |
| engine `src/app/turn.cpp` | `a626ddf80a340de50b70f09bf91c0180d255c533e65edce1532d807fcf5109b4` |
| engine `src/mars/rng/mt19937.cpp` | `0c60775e9437b8eaad49f67a96001c81947c8d9e6d2472e28b2b490515bac95b` |
| engine `tests/app/test_turn.cpp` | `6d45b1d2d065528f14ae83d95314f2b6559e53cc23f1b9e2531c7c09b4b97eab` |
## Findings requiring correction
### R1 — HIGH: gate accepts incomplete execution and missing binary
**Locations:** `tools/gate.py:137-153,162-164`.
Discovery is compared with the 59-name inventory, but actual JUnit cases are not required to
cover that inventory. Only four corpus names must appear in `passed`. There is no nonempty
binary requirement. `corpusNonzero` checks that stdout lacks `"0 save"`; empty output passes,
and CTest `--output-on-failure` normally suppresses successful test output anyway.
**Executed reproduction:** import gate; isolate fake engine/.git and one .sav in a temporary
directory; mock source_manifest/copy_snapshot and command execution (no build); return all 59
names from ctest_names, write JUnit with only mars_stream_save, mars_stream_domains, app_turn,
app_turn_record, return exit 0 and empty stdout for commands, produce no binary. Call main with
explicit engine/corpus/new out. Observed:
```text
gate incomplete JUnit/no output/no binary: 0 passed 4 of 59 binary= {} corpusNonzero= True
```
This tests the validator boundary, not actual CTest behavior. A truncated/misconfigured runner
result must fail rather than rely on CTest usually producing complete output.
**Required fix/check:** exact JUnit identity partition (passed/skipped/failed), no duplicates or
unknown/missing tests, explicit positive corpus counts from machine-readable results, required
binary artifact existence/hash, and negative tests for each omission. Host allowances must be
distinguished from required executed tests. Do not infer a positive count from absent text.
### R2 — HIGH: reporter labels equality-only, unbound execution as accepted
**Locations:** `tools/standalone_report.py:68-74,82-105`; `tools/evidence.py:12-16`.
Reporter requires only schema/status/binary hash from provenance. It does not validate full
gate profile/checks/source/input completeness, required phase execution, missing runtime inputs,
or workload coverage. Caller may supply the same save as input and oracle. `--accept` promotes
three equal digests without any independent execution proof or contract binding.
**Executed reproduction:** use the real turn3 save as both input and oracle, a temporary dummy
binary with matching minimal `{schema,status,profile:"host",binary:{sha256}}` manifest, and mock
only subprocess.run to copy input to `--out` and return 0. Actual reader/reconstruction/digest
checks run. Observed `reporter identical input/oracle, no execution: 0 accepted`.
This is a validation-boundary test, not a claim that current sots_turn is a copy-only program.
`--roundtrip` checks conservation first but does not skip RunStrategicTurn (main.cpp:214-222,318).
**Required fix/check:** keep equality measurement usable, but scoped acceptance requires an
explicit workload contract, full source-bound provenance, positive required phase/branch counts,
known input/dependency completeness and independent verification/integration gates. No-op and
empty/partial provenance fixtures must fail acceptance. Persist/hash all consumed external
engine arguments (data/commands), verify files unchanged over execution, and retain output saves
or a durable reproducible package: currently output files are removed with TemporaryDirectory.
### R3 — HIGH: full gate's toolchain argument rejects valid files and accepts directories
**Locations:** `tools/gate.py:100,107-108,154-156`.
`--shim` must be a directory, then that directory is passed as `CMAKE_TOOLCHAIN_FILE`.
A normal existing toolchain `.cmake` file fails argument validation; a directory cannot supply
the intended toolchain file. Full gate is required for acceptance but this route is unusable
as a conventional CMake toolchain interface. Inspection finding; no build attempted.
**Reproduction:** supply a valid engine/corpus/data and existing toolchain file to --shim;
observe parser rejection at line 108. A directory passes validation then is the literal
`-DCMAKE_TOOLCHAIN_FILE=<directory>` in the configure command.
**Required fix/check:** define/document --toolchain file versus shim source/build artifact;
resolve and hash it, validate configure command in a unit test. Bind produced shim binary as
well as host binary. Full test inputs also include SOTS_M1_TRACE, SOTS_SAVES_JSON, SOTS_GOB_DIR;
current gate inherits these without recording hashes and does not provide an explicit interface.
### R4 — HIGH: campaign acceptance source binding is only baseline path+commit
**Locations:** `campaign/contract.schema.json:26-32`; `tools/campaign.py:299-307,325-335,365-369`;
`tools/run_agent.py:68-84,173,227`.
Runner records actual source manifests, but lifecycle evidence compares `source` only to the
contract's baseline `{path,commit}`. Neither evidence nor verifier verdict references the runner's
actual source digest or a candidate/integrated tree digest. The contract basis is a hash of
task metadata; it is unchanged when uncommitted implementation changes. Integrated status is a
lead-supplied boolean and axis coverage, with no connection to which integrated source was built.
**Reproduction by data flow:** create evidence for source A under a baseline, obtain a bound
verdict, then change implementation bytes without changing HEAD or artifact/contract JSON.
check_evidence and check_verdict have no source-content input to detect source B. Source_after
in run records does not enter these checks. This is a stale-evidence gap even with honest actors;
it is distinct from the documented limitation that editable role strings are not authentication.
**Required fix/check:** bind candidate/integrated source-manifest digests, build binary and
immutable inputs through evidence and independent verdict; transition must validate that binding.
Use per-criterion results or a typed acceptance package rather than treating any hashed artifact
with the same axis label as execution proof. Add a test that changes source content at the same
HEAD and rejects its old verifier/integration result.
Late-published README (lines 27-37) explicitly assigns dirty-source manifest evaluation to the
independent reviewer, not the CLI. Thus this is an acceptance-robustness gap requiring a lead
decision, not a claim that the controls author promised a machine interpretation of arbitrary
criteria. The human procedure must at minimum compare the actual candidate/integrated manifests;
baseline equality in CLI output cannot be presented as that check.
### R5 — MEDIUM: publication rejects valid host-profile skips
**Locations:** `tools/dashboard.py:93-101`; `tools/gate.py:146-150`.
Gate's expected list is all discovered tests, including permitted host skips. Publisher demands
`expected <= passed`, contradicting host skip allowances. A correct host result cannot be
published as scoped host evidence.
**Executed reproduction:** gate_valid with host status passed, nonempty source/input/binary
identities, tests expected=[a,asset], passed=[a], skipped=[asset], failed=[] reports
`required tests not executed`. Same outcome follows for the current 59/52/7 shape.
**Required fix/check:** shared schema distinguishes required executions and approved skips;
validate exact partition and profile policy. Host results must visibly publish limitations,
not be promoted to full acceptance. Gate currently initializes limitations=[] and never fills it.
### R6 — MEDIUM: recovery rejects old durable checkpoints and blocks resolver entry
**Locations:** `tools/run_agent.py:87-99`; `tools/campaign.py:266-278`.
check_launch reuses end-of-quantum freshness validation, requiring a checkpoint younger than
15 minutes even when starting recovery from a successfully completed earlier session. Next-day
resume therefore needs a fabricated new checkpoint before an agent can read/recover the old one.
Separately, open_surprises is checked unconditionally for every role, so even a resolver cannot
launch to investigate an unresolved surprise (despite resolver status mapping to blocked/revision).
**Reproduction:** a structurally valid, matching checkpoint older than 900 seconds fails
check_launch; a resolver/blocked contract with an open surprise fails before role-specific work.
Inspection/data-flow finding; final controls tests were not yet published.
**Required fix/check:** distinguish valid durable recovery state from a fresh end-run checkpoint;
verify identities/basis/artifact hashes on recovery, and require freshness only after this quantum
starts. Permit an explicitly assigned resolver's bounded investigation while affected workers stay
blocked, or document a separate operational resolver launch path that does not pre-resolve evidence.
### R7 — MEDIUM: built snapshot is not itself verified against the source manifest
**Locations:** `tools/gate.py:29-49,57-61,119-129,157-158`.
Manifest hashes live source, then copy_snapshot copies paths without checking copied hashes;
end check hashes live source again. A concurrent A -> B -> A change can copy B while both
live manifests attest A. Filesystem modes and symlink identity/targets are not bound; every
path component starting with `build` is excluded, including possible real source directories.
RE tools/inventory are recorded once but then used live without a post-check.
**Reproduction:** hash source A, write B before copy_snapshot, restore A before after-manifest;
before == after but destination is B. This needs no Git commit or modification of the gate.
**Required fix/check:** hash the copied snapshot and compare its exact file/mode/link manifest
before building; reject unsupported escaping links/submodules and define explicit exclusions.
Snapshot/hash tools and test inventory actually executed. Baseline-pinned isolated worktrees and
leases remain necessary; a before/after check alone does not establish which bytes were built.
## Engine / lead configuration observations
- Engine S13 no longer aggregates child P01..P12 counters; the existing final fold can count
child operations once. CountingRandom measures rejection words with the monotone MT counter;
seed/load reset it, and save_state's explicit layout excludes it. New app test replays reported
words against persisted generator state; new MT test exercises rejection across a twist.
No defect established in these reviewed changes. Integrated independent execution remains pending.
- New lead workflow/architecture clearly distinguish original assistance, partial/full/scoped
evidence, unsupported inputs and roles. Explicit Astra/Terra/5.5 routing and 40-step bounds
appear in config; runner checks model availability and aborts on known mismatch, with no fallback.
- Config permissions primarily guard architecture edit paths and deny worker task calls. They
do not constrain arbitrary shell or enforce exact string-array contract scope, as workflow
correctly acknowledges. Reviewer has not exercised OpenCode's effective merged configuration;
runner overlay binds model/steps but only canonical config bytes are hashed, not expanded prompt
file bytes or all effective configuration. Include these in final provenance/config validation.
- Controls lock is canonical and serialized; leases require matching token/owner and forbid
automatic stealing. Active-run worktree reservations are checked under the same lock. No
concurrency acceptance claim: independent contention/crash-recovery tests are still required.
## Research evidence surprises sent to lead
`verify/results/shim/cr/cr-R1.log:75-94` adds seven drawsites/twelve probes to the narrative's
six interceptors: 25 sites total. Lines 109-115 and raw trace args show runtime CW `0x127f`,
not initialization `0x027f`. Archive save hashes and R1's exact 16-leaf residual independently
reproduced. These corrections and runtime asset text path are detailed in the pilot and checkpoint.
## Reproduction record and remaining review
Executed with PYTHONDONTWRITEBYTECODE=1:
- `python3 tools/campaign.py --state-root /home/alex/sots-re validate research-replacement`: pass.
- Selected unittest methods in `verify.tooling.test_tooling.ToolingTests`: inventory_is_explicit_and_unique,
passed_gate_required, zero_discovered_tests_cannot_match_inventory,
stale_report_output_is_rejected_before_execution (each prefixed `test_`): **4/4 pass**.
- `/tmp/opencode/review_rollout_probes.py`: isolated boundary mocks described under R1/R2/R5;
output recorded above. Temporary fixture files are disposable; this report preserves recipes/results.
- Raw save diagnostic command in pilot: R1 exact-bit/unmasked **16 differences**.
No full local worker suite: one existing tooling test creates a temporary Git commit, excluded
under this review assignment. No reviewer builds while engine source was being changed.
At last inspection controls README/test directory were not yet present; schema and initial tools
were inspected, not their promised completed validation package.
Final refresh at 2026-09-09T21:30:34Z: `campaign/README.md` appeared and was read. Its schema/API
matches this pilot and documents manual criterion evaluation, baseline-versus-dirty-source limits,
and the current fresh-checkpoint/no-open-surprise launch rules. It does not close R4's automatic
source-drift gap or R6's recovery/resolver operational concern. Completed controls test execution
and integrated evidence review remain pending; repository-creating tests need explicit permission
for their temporary Git commits or a non-committing equivalent under this review assignment.
**Next review:** after owners fix/finish, resume from review-worker-state.md, check changed hashes,
re-run these negative cases and finalized controls/publishing tests, then independently review
one stable integrated source-bound gate/replay package. Until then Task B is an initial adversarial
review with final integration review pending, not an approval of workers' future edits.