376 lines
26 KiB
Markdown
376 lines
26 KiB
Markdown
# Independent rollout review
|
|
|
|
## Follow-up verdict — 2026-09-09
|
|
|
|
Independent execution: openai/gpt-6-astra, logical session `review-followup-1`;
|
|
actual harness session ID unavailable. **Selected host/shim baseline passes the reviewed
|
|
evidence-integrity scope. Controls require one additional correction (R8 below).**
|
|
Full asset/live-game and research-replacement acceptance remain unestablished.
|
|
This section supersedes the initial available-tree disposition for the source identities below;
|
|
the original findings are retained as history.
|
|
|
|
### Independently reproduced measurements
|
|
|
|
- Tooling: **19/19**, publishing: **8/8**, config/source scanner: **7/7**, controls: **36/36**.
|
|
Commands: `PYTHONDONTWRITEBYTECODE=1 python3 -m unittest discover -s verify/DIR -p 'test_*.py' -v`,
|
|
with DIR respectively `tooling`, `publishing`, `config`, `campaign`. Total **70 tests**.
|
|
Disposable fixture repositories alone perform Git writes/commits.
|
|
- `PYTHONDONTWRITEBYTECODE=1 python3 tools/check_agent_config.py --resolved`: PASS, actual
|
|
OpenCode loader and exact model availability. No model execution inferred from availability.
|
|
- `campaign/current.json` selected gate SHA-256
|
|
`8e14e00ee3ce7478ddfdef8de12183451e858c52d78dc28cffbdee47f5d087d8` and replay
|
|
`87f92c2b54625ecbca1f3c0a37e57c43d03a4aff488edf4ea8a09a842cd6279d`: actual retained bytes match.
|
|
Gate/replay pass shared structural and publication validators; selection is measurement-only.
|
|
- Under `/home/alex/.local/share/sots-runs/rollout-host-20260909-c`, independently rehashed host
|
|
executable `78b2562ea2c56351f9f6f0f24d271afcced741eadbc148047780e6dfd101940b` and shim DLL
|
|
`381c91aecf10a9093f651f58c884356cd88e0253e79cb3bab32753fff6612c7e`.
|
|
**524 engine files and 10 RE execution-tool files** match both retained snapshots and current
|
|
source in bytes and modes. All **43** corpus files match their manifest. Engine baseline
|
|
`7741d42fc5e4e761e6449bdaf0e4a61d00036a23`; RE baseline
|
|
`3bfde5a70d874a723e797a695bbd847fd82c0aa7`; dirty identities come from manifests, not HEAD alone.
|
|
- Actual `ctest.xml` exactly matches embedded JUnit rows: **59 unique inventory identities,
|
|
52 passed, 7 allowed skips, zero failures**. Positive final summaries are exactly `[43]` for
|
|
`mars_stream_save`, `mars_stream_domains`, `app_turn`, `app_turn_record`. Every required check
|
|
is true; eight recorded gate commands exit zero; every JUnit output exists without truncation.
|
|
Skips: `game_config_replay`, `game_data_realdata`, `game_design_realdata`, `game_design_census`,
|
|
`game_sim_smoke_real_save`, `mars_text_realdata`, `mars_vfs_realdata`.
|
|
- Retained replay in `/home/alex/.local/share/sots-runs/rollout-replay-20260909` uses turn2 input
|
|
and turn3 oracle. Recomputed file/inflated/typed-state hashes of all three saves and complete
|
|
ordered diff list match: **62 state differences**, all three equality surfaces false.
|
|
Status correctly remains **measured**. This review recalculated saved outputs; it did not rerun
|
|
the full engine suite or build anything.
|
|
- Additional independent mutations of the selected manifest rejected missing binary, omitted
|
|
execution, duplicate execution, zero corpus count, omitted `outputComplete` requirement and
|
|
false `outputComplete`. These were in-memory fixtures, leaving selected evidence untouched.
|
|
|
|
Reproduction recipe for artifact checks: load current pointers with `dashboard.pointer`, run
|
|
`gate_valid`/`replay_valid`; rehash both binary paths; compare `gate.file_row` for every source row
|
|
under manifest live roots and retained `source`/`re-tooling`; recompute each `input_manifest`;
|
|
parse actual XML with `junit_statuses` and compare rows, inventory, allowed skips and corpus
|
|
summary counts; use `standalone_report.save_state` and complete `state_checksum.diff` on retained
|
|
pair paths. Temporary driver `/tmp/opencode/review_followup_checks.py` is disposable; this recipe,
|
|
measurements and the content-addressed selected manifests are durable evidence.
|
|
|
|
### Prior findings disposition
|
|
|
|
| Finding | Current disposition |
|
|
|---|---|
|
|
| R1 | Closed for inspected gate/package: exact unique execution partition, positive per-test corpus summaries, bound executable and output completeness; negatives reproduced. |
|
|
| R2 | Closed under lead's revised scope: reporter has no acceptance status/flag; equality is measured or failed, explicit inputs are hashed, inherited SOTS environment stripped, output retained. Equal-pair and require-match negatives pass. |
|
|
| R3 | Closed: `--shim` is now boolean and uses the recorded snapshot's MinGW toolchain file; actual selected DLL/hash and successful configure/build are retained. Full assets still unavailable. |
|
|
| R4 | Closed for approved controls implementation: current source bindings rehashed on evidence/verdict/promotion, immutable inputs/binaries and exact criterion outcomes required; same-HEAD candidate/integrated mutation tests reproduced. |
|
|
| R5 | Closed for actual selected host package: shared validator/publication accept the explicit 52/7 partition and expose limitations. |
|
|
| R6 | Closed: intact old recovery accepted, corrupt/missing state rejected, freshness retained at handoff/end, resolution-only Astra entry allowed with ordinary workers blocked; tests reproduced. |
|
|
| R7 | Closed for reviewed snapshot-copy defect: copy checks bytes/modes on both sides; current and snapshot inventories match. Pre-copy mutation, symlink and build-named-source tests pass. |
|
|
|
|
Engine accounting review confirms S13 no longer duplicates child P counters; replay writes survive
|
|
the final fold; runtime-only MT word count measures rejection across twists without changing the
|
|
explicit serialized state layout. Corpus app tests independently compare committed generator state.
|
|
No new engine defect established. Snapshot clean-room scanner and scoped tracecmp path checks pass.
|
|
Gate RE identity explicitly covers execution tooling only, not every campaign/control file.
|
|
|
|
### Completed controls and actual runner evidence
|
|
|
|
Latest checkpoint `campaign/runtime/checkpoints/controls-bootstrap-620eb9d25a6ad2f06b68353c.json`
|
|
passes basis/identity/artifact validation with recovery freshness disabled. Both R4/R6 surprise
|
|
records are resolved by the named lead decisions; no open controls surprises existed at inspection.
|
|
Actual smoke `campaign/runtime/runs/run-df1472c13f31db3a4d5361f0.json`, SHA-256
|
|
`65dd1c1286742cd4d2401af5a229d7d014f60dd329dadee0dee78c499f61d663`, records successful
|
|
**architecture-review/Astra** execution: 24 parsed events, no error events, successful stop,
|
|
matching fresh session/actor/model checkpoint and intact artifacts. Actual session
|
|
`ses_f77cf5062ffeTknrMfnbJMmdNh`; emitted model unavailable, honestly recorded as such.
|
|
Canonical config/prompt hashes match. Independently rerunning the actual loader with runner overlay
|
|
in recorded cwd reproduces effective hash `4cb65613de34e41eabf780e84cad062f14cbfcdd68569dc874514c50fa7baa5a`
|
|
and agent hash `530f4bfeb28c09d36e45565706247ad9822f7c177161e1ee0a71fcb981a76b92`.
|
|
This proves the recorded architecture route, not an actual normal Terra worker quantum.
|
|
|
|
### R8 — MEDIUM: launcher cannot start required final integrated verifier
|
|
|
|
**Locations:** `tools/campaign.py:156-159` (`ROLE_STATUS`), `tools/run_agent.py:87-89`.
|
|
The verifier role permits only `verification`, while `Campaign.verdict` allows `integration`
|
|
and final acceptance requires a fresh verdict after integrated evidence changes.
|
|
Thus the prescribed explicit-role launcher cannot schedule that final verifier execution.
|
|
|
|
**Executed reproduction:** in a disposable `verify/campaign/test_controls.py::Controls` fixture,
|
|
run `verification()`, `verdict()`, transition to `integration` as lead, add `evidence(integrated=True)`
|
|
as lead, then call `run_agent.check_launch(case.c, case.c.load('slice'), 'verifier', 'independent')`.
|
|
Observed `ControlError: role cannot launch in this contract status`. This is a launcher/lifecycle
|
|
mismatch, not a bypass of the acceptance check. Fixture cleaned up; canonical state untouched.
|
|
|
|
**Required resolution/check:** lead formally resolves the new review finding; permit the independent
|
|
verifier route for final integrated-package review with existing independence/checkpoint/surprise
|
|
guards, or define an equally explicit supported final-verifier launch path. Regression must allow
|
|
both verification and integration verifier entry and still reject owner/verifier overlap, stale
|
|
package and unresolved surprises. No implementation correction made by reviewer.
|
|
|
|
### Exact remaining integration checks
|
|
|
|
1. Resolve R8, implement the approved launcher policy, independently rerun its focused regression
|
|
and refresh source-bound controls measurements after that source change.
|
|
2. Perform the documented normal noninteractive worker smoke under its actual worker role/model,
|
|
paired baseline worktrees and intact recovery checkpoint; verify effective config, successful
|
|
stop/session, fresh end checkpoint and unavailable-versus-emitted model provenance. Existing
|
|
actual smoke covers architecture-review only; fake-process tests cover worker runner logic.
|
|
3. Attach the actual integrated controls executable/interpreter, fixture inputs and criterion
|
|
result package; obtain an independent verdict over that exact final binding through the supported
|
|
lifecycle. `controls-bootstrap` presently remains `needs-revision` with **no evidence array**;
|
|
this review is not a lifecycle promotion or machine verdict.
|
|
4. Full-profile owner assets/trace inputs and research pilot completion-bearing live controls,
|
|
complete writes/events/allocations/RNG/runtime inputs and original differential remain required
|
|
for their separate acceptance scopes. Host/shim success and divergent replay do not satisfy them.
|
|
|
|
Additional reviewed source SHA-256 (gate/engine/reporter identities are retained in selected gate):
|
|
|
|
| Path | SHA-256 |
|
|
|---|---|
|
|
| `tools/campaign.py` | `ca4eb2c42e88c1222ec60d6b99c58fffcc72746411c499c4cb0246d899ebbd10` |
|
|
| `tools/run_agent.py` | `0f5ccb53dcde59b64fe999768737b6cbd2dabcc1272e5605c1f1ecd86e5d5d2b` |
|
|
| `campaign/contract.schema.json` | `a2c76ec042ca097a57c3c05c1e519e392d498d387a31808c7d0f70482d43c201` |
|
|
| `verify/campaign/test_controls.py` | `656e5938fa5107022b73b32cf6fc05b6ff6e3e9ab239432d72b304b3f6ea1801` |
|
|
| `tools/select_evidence.py` | `a6e648a77bcf799cc774ff9d0de83011283049ed0f0c44f6e08857a8410e9463` |
|
|
| `tools/dashboard.py` | `c84aeeea6fb7d5e45ccec8882bacc59182149f108edd9eb1585eab9ede99911c` |
|
|
| `tools/check_agent_config.py` | `93115532cb38105777a5c5c20a9f2a808043370e0d4f560a9c82fe8e800076ec` |
|
|
|
|
No implementation edits, delegates, actual-repository staging/commits, engine builds or lab mutation.
|
|
|
|
## Historical initial available-tree pass
|
|
|
|
Reviewer: openai/gpt-6-astra, independent review-worker execution, 2026-09-09.
|
|
**Disposition: changes required; rollout acceptance is not established.** Workers were still
|
|
implementing during this review. Findings below apply to the identified snapshots; fixes and
|
|
final integrated-tree review are pending. No engine builds, lab operations, delegation, staging
|
|
or commits were performed by this reviewer. Four selected Python tooling tests passed, but the
|
|
adversarial probes below exposed gaps outside those tests.
|
|
|
|
Pilot delivered first:
|
|
- `campaign/contracts/research-replacement.json` — **proposed**, validated by campaign CLI.
|
|
- `campaign/pilots/research-replacement.md` — concrete archived W1 workload, full transitive
|
|
write boundary, runtime assets/original dependencies, executable acceptance requirements
|
|
and blockers. No replacement implementation or acceptance claim.
|
|
|
|
## Snapshot identities
|
|
|
|
Canonical RE HEAD `3bfde5a70d874a723e797a695bbd847fd82c0aa7`; engine HEAD
|
|
`7741d42fc5e4e761e6449bdaf0e4a61d00036a23`. Both have concurrent uncommitted rollout work;
|
|
HEAD alone does not identify reviewed content. SHA-256 at this pass:
|
|
|
|
| File | SHA-256 |
|
|
|---|---|
|
|
| `tools/gate.py` | `700f8fcf1a051a37322cd51e2e1bb774f35cd82e0abc622613dc2b79f37e9a42` |
|
|
| `tools/evidence.py` | `8235dc7e3bc23e5a942fb6f80be0693a0d2c66057f24a783a41b512d0d31c137` |
|
|
| `tools/standalone_report.py` | `6e4aae37ff5ca82f719d1994f1561cb21840d3cf7978f9a2e0a3f4b1019b71d8` |
|
|
| `tools/dashboard.py` | `6d8e116f554f8dae222694e3e37daf7b03922187e5fd841d43ee4cd52c69a49a` |
|
|
| `tools/campaign.py` | `bd563f43b536841677e49ed9b942e5ed46ad3c53456a921592f281b4f2f3c9a2` |
|
|
| `tools/run_agent.py` | `c8aeec109340958bf7850d2a91b4d6b6e53ac0534e97973f6c0869e692a9b7ba` |
|
|
| `campaign/contract.schema.json` | `ae4796b1f8e8336ddb63e60d774a81ff965461e4b884b5a2488e3865c6a6c5e0` |
|
|
| `opencode.json` | `d24be7fcee57d5eb04c6b7189822a2ca055dd4a70daafbea67f09a120945d49f` |
|
|
| engine `src/app/turn.cpp` | `a626ddf80a340de50b70f09bf91c0180d255c533e65edce1532d807fcf5109b4` |
|
|
| engine `src/mars/rng/mt19937.cpp` | `0c60775e9437b8eaad49f67a96001c81947c8d9e6d2472e28b2b490515bac95b` |
|
|
| engine `tests/app/test_turn.cpp` | `6d45b1d2d065528f14ae83d95314f2b6559e53cc23f1b9e2531c7c09b4b97eab` |
|
|
|
|
## Findings requiring correction
|
|
|
|
### R1 — HIGH: gate accepts incomplete execution and missing binary
|
|
|
|
**Locations:** `tools/gate.py:137-153,162-164`.
|
|
|
|
Discovery is compared with the 59-name inventory, but actual JUnit cases are not required to
|
|
cover that inventory. Only four corpus names must appear in `passed`. There is no nonempty
|
|
binary requirement. `corpusNonzero` checks that stdout lacks `"0 save"`; empty output passes,
|
|
and CTest `--output-on-failure` normally suppresses successful test output anyway.
|
|
|
|
**Executed reproduction:** import gate; isolate fake engine/.git and one .sav in a temporary
|
|
directory; mock source_manifest/copy_snapshot and command execution (no build); return all 59
|
|
names from ctest_names, write JUnit with only mars_stream_save, mars_stream_domains, app_turn,
|
|
app_turn_record, return exit 0 and empty stdout for commands, produce no binary. Call main with
|
|
explicit engine/corpus/new out. Observed:
|
|
|
|
```text
|
|
gate incomplete JUnit/no output/no binary: 0 passed 4 of 59 binary= {} corpusNonzero= True
|
|
```
|
|
|
|
This tests the validator boundary, not actual CTest behavior. A truncated/misconfigured runner
|
|
result must fail rather than rely on CTest usually producing complete output.
|
|
|
|
**Required fix/check:** exact JUnit identity partition (passed/skipped/failed), no duplicates or
|
|
unknown/missing tests, explicit positive corpus counts from machine-readable results, required
|
|
binary artifact existence/hash, and negative tests for each omission. Host allowances must be
|
|
distinguished from required executed tests. Do not infer a positive count from absent text.
|
|
|
|
### R2 — HIGH: reporter labels equality-only, unbound execution as accepted
|
|
|
|
**Locations:** `tools/standalone_report.py:68-74,82-105`; `tools/evidence.py:12-16`.
|
|
|
|
Reporter requires only schema/status/binary hash from provenance. It does not validate full
|
|
gate profile/checks/source/input completeness, required phase execution, missing runtime inputs,
|
|
or workload coverage. Caller may supply the same save as input and oracle. `--accept` promotes
|
|
three equal digests without any independent execution proof or contract binding.
|
|
|
|
**Executed reproduction:** use the real turn3 save as both input and oracle, a temporary dummy
|
|
binary with matching minimal `{schema,status,profile:"host",binary:{sha256}}` manifest, and mock
|
|
only subprocess.run to copy input to `--out` and return 0. Actual reader/reconstruction/digest
|
|
checks run. Observed `reporter identical input/oracle, no execution: 0 accepted`.
|
|
|
|
This is a validation-boundary test, not a claim that current sots_turn is a copy-only program.
|
|
`--roundtrip` checks conservation first but does not skip RunStrategicTurn (main.cpp:214-222,318).
|
|
|
|
**Required fix/check:** keep equality measurement usable, but scoped acceptance requires an
|
|
explicit workload contract, full source-bound provenance, positive required phase/branch counts,
|
|
known input/dependency completeness and independent verification/integration gates. No-op and
|
|
empty/partial provenance fixtures must fail acceptance. Persist/hash all consumed external
|
|
engine arguments (data/commands), verify files unchanged over execution, and retain output saves
|
|
or a durable reproducible package: currently output files are removed with TemporaryDirectory.
|
|
|
|
### R3 — HIGH: full gate's toolchain argument rejects valid files and accepts directories
|
|
|
|
**Locations:** `tools/gate.py:100,107-108,154-156`.
|
|
|
|
`--shim` must be a directory, then that directory is passed as `CMAKE_TOOLCHAIN_FILE`.
|
|
A normal existing toolchain `.cmake` file fails argument validation; a directory cannot supply
|
|
the intended toolchain file. Full gate is required for acceptance but this route is unusable
|
|
as a conventional CMake toolchain interface. Inspection finding; no build attempted.
|
|
|
|
**Reproduction:** supply a valid engine/corpus/data and existing toolchain file to --shim;
|
|
observe parser rejection at line 108. A directory passes validation then is the literal
|
|
`-DCMAKE_TOOLCHAIN_FILE=<directory>` in the configure command.
|
|
|
|
**Required fix/check:** define/document --toolchain file versus shim source/build artifact;
|
|
resolve and hash it, validate configure command in a unit test. Bind produced shim binary as
|
|
well as host binary. Full test inputs also include SOTS_M1_TRACE, SOTS_SAVES_JSON, SOTS_GOB_DIR;
|
|
current gate inherits these without recording hashes and does not provide an explicit interface.
|
|
|
|
### R4 — HIGH: campaign acceptance source binding is only baseline path+commit
|
|
|
|
**Locations:** `campaign/contract.schema.json:26-32`; `tools/campaign.py:299-307,325-335,365-369`;
|
|
`tools/run_agent.py:68-84,173,227`.
|
|
|
|
Runner records actual source manifests, but lifecycle evidence compares `source` only to the
|
|
contract's baseline `{path,commit}`. Neither evidence nor verifier verdict references the runner's
|
|
actual source digest or a candidate/integrated tree digest. The contract basis is a hash of
|
|
task metadata; it is unchanged when uncommitted implementation changes. Integrated status is a
|
|
lead-supplied boolean and axis coverage, with no connection to which integrated source was built.
|
|
|
|
**Reproduction by data flow:** create evidence for source A under a baseline, obtain a bound
|
|
verdict, then change implementation bytes without changing HEAD or artifact/contract JSON.
|
|
check_evidence and check_verdict have no source-content input to detect source B. Source_after
|
|
in run records does not enter these checks. This is a stale-evidence gap even with honest actors;
|
|
it is distinct from the documented limitation that editable role strings are not authentication.
|
|
|
|
**Required fix/check:** bind candidate/integrated source-manifest digests, build binary and
|
|
immutable inputs through evidence and independent verdict; transition must validate that binding.
|
|
Use per-criterion results or a typed acceptance package rather than treating any hashed artifact
|
|
with the same axis label as execution proof. Add a test that changes source content at the same
|
|
HEAD and rejects its old verifier/integration result.
|
|
|
|
Late-published README (lines 27-37) explicitly assigns dirty-source manifest evaluation to the
|
|
independent reviewer, not the CLI. Thus this is an acceptance-robustness gap requiring a lead
|
|
decision, not a claim that the controls author promised a machine interpretation of arbitrary
|
|
criteria. The human procedure must at minimum compare the actual candidate/integrated manifests;
|
|
baseline equality in CLI output cannot be presented as that check.
|
|
|
|
### R5 — MEDIUM: publication rejects valid host-profile skips
|
|
|
|
**Locations:** `tools/dashboard.py:93-101`; `tools/gate.py:146-150`.
|
|
|
|
Gate's expected list is all discovered tests, including permitted host skips. Publisher demands
|
|
`expected <= passed`, contradicting host skip allowances. A correct host result cannot be
|
|
published as scoped host evidence.
|
|
|
|
**Executed reproduction:** gate_valid with host status passed, nonempty source/input/binary
|
|
identities, tests expected=[a,asset], passed=[a], skipped=[asset], failed=[] reports
|
|
`required tests not executed`. Same outcome follows for the current 59/52/7 shape.
|
|
|
|
**Required fix/check:** shared schema distinguishes required executions and approved skips;
|
|
validate exact partition and profile policy. Host results must visibly publish limitations,
|
|
not be promoted to full acceptance. Gate currently initializes limitations=[] and never fills it.
|
|
|
|
### R6 — MEDIUM: recovery rejects old durable checkpoints and blocks resolver entry
|
|
|
|
**Locations:** `tools/run_agent.py:87-99`; `tools/campaign.py:266-278`.
|
|
|
|
check_launch reuses end-of-quantum freshness validation, requiring a checkpoint younger than
|
|
15 minutes even when starting recovery from a successfully completed earlier session. Next-day
|
|
resume therefore needs a fabricated new checkpoint before an agent can read/recover the old one.
|
|
Separately, open_surprises is checked unconditionally for every role, so even a resolver cannot
|
|
launch to investigate an unresolved surprise (despite resolver status mapping to blocked/revision).
|
|
|
|
**Reproduction:** a structurally valid, matching checkpoint older than 900 seconds fails
|
|
check_launch; a resolver/blocked contract with an open surprise fails before role-specific work.
|
|
Inspection/data-flow finding; final controls tests were not yet published.
|
|
|
|
**Required fix/check:** distinguish valid durable recovery state from a fresh end-run checkpoint;
|
|
verify identities/basis/artifact hashes on recovery, and require freshness only after this quantum
|
|
starts. Permit an explicitly assigned resolver's bounded investigation while affected workers stay
|
|
blocked, or document a separate operational resolver launch path that does not pre-resolve evidence.
|
|
|
|
### R7 — MEDIUM: built snapshot is not itself verified against the source manifest
|
|
|
|
**Locations:** `tools/gate.py:29-49,57-61,119-129,157-158`.
|
|
|
|
Manifest hashes live source, then copy_snapshot copies paths without checking copied hashes;
|
|
end check hashes live source again. A concurrent A -> B -> A change can copy B while both
|
|
live manifests attest A. Filesystem modes and symlink identity/targets are not bound; every
|
|
path component starting with `build` is excluded, including possible real source directories.
|
|
RE tools/inventory are recorded once but then used live without a post-check.
|
|
|
|
**Reproduction:** hash source A, write B before copy_snapshot, restore A before after-manifest;
|
|
before == after but destination is B. This needs no Git commit or modification of the gate.
|
|
|
|
**Required fix/check:** hash the copied snapshot and compare its exact file/mode/link manifest
|
|
before building; reject unsupported escaping links/submodules and define explicit exclusions.
|
|
Snapshot/hash tools and test inventory actually executed. Baseline-pinned isolated worktrees and
|
|
leases remain necessary; a before/after check alone does not establish which bytes were built.
|
|
|
|
## Engine / lead configuration observations
|
|
|
|
- Engine S13 no longer aggregates child P01..P12 counters; the existing final fold can count
|
|
child operations once. CountingRandom measures rejection words with the monotone MT counter;
|
|
seed/load reset it, and save_state's explicit layout excludes it. New app test replays reported
|
|
words against persisted generator state; new MT test exercises rejection across a twist.
|
|
No defect established in these reviewed changes. Integrated independent execution remains pending.
|
|
- New lead workflow/architecture clearly distinguish original assistance, partial/full/scoped
|
|
evidence, unsupported inputs and roles. Explicit Astra/Terra/5.5 routing and 40-step bounds
|
|
appear in config; runner checks model availability and aborts on known mismatch, with no fallback.
|
|
- Config permissions primarily guard architecture edit paths and deny worker task calls. They
|
|
do not constrain arbitrary shell or enforce exact string-array contract scope, as workflow
|
|
correctly acknowledges. Reviewer has not exercised OpenCode's effective merged configuration;
|
|
runner overlay binds model/steps but only canonical config bytes are hashed, not expanded prompt
|
|
file bytes or all effective configuration. Include these in final provenance/config validation.
|
|
- Controls lock is canonical and serialized; leases require matching token/owner and forbid
|
|
automatic stealing. Active-run worktree reservations are checked under the same lock. No
|
|
concurrency acceptance claim: independent contention/crash-recovery tests are still required.
|
|
|
|
## Research evidence surprises sent to lead
|
|
|
|
`verify/results/shim/cr/cr-R1.log:75-94` adds seven drawsites/twelve probes to the narrative's
|
|
six interceptors: 25 sites total. Lines 109-115 and raw trace args show runtime CW `0x127f`,
|
|
not initialization `0x027f`. Archive save hashes and R1's exact 16-leaf residual independently
|
|
reproduced. These corrections and runtime asset text path are detailed in the pilot and checkpoint.
|
|
|
|
## Reproduction record and remaining review
|
|
|
|
Executed with PYTHONDONTWRITEBYTECODE=1:
|
|
- `python3 tools/campaign.py --state-root /home/alex/sots-re validate research-replacement`: pass.
|
|
- Selected unittest methods in `verify.tooling.test_tooling.ToolingTests`: inventory_is_explicit_and_unique,
|
|
passed_gate_required, zero_discovered_tests_cannot_match_inventory,
|
|
stale_report_output_is_rejected_before_execution (each prefixed `test_`): **4/4 pass**.
|
|
- `/tmp/opencode/review_rollout_probes.py`: isolated boundary mocks described under R1/R2/R5;
|
|
output recorded above. Temporary fixture files are disposable; this report preserves recipes/results.
|
|
- Raw save diagnostic command in pilot: R1 exact-bit/unmasked **16 differences**.
|
|
|
|
No full local worker suite: one existing tooling test creates a temporary Git commit, excluded
|
|
under this review assignment. No reviewer builds while engine source was being changed.
|
|
At last inspection controls README/test directory were not yet present; schema and initial tools
|
|
were inspected, not their promised completed validation package.
|
|
|
|
Final refresh at 2026-09-09T21:30:34Z: `campaign/README.md` appeared and was read. Its schema/API
|
|
matches this pilot and documents manual criterion evaluation, baseline-versus-dirty-source limits,
|
|
and the current fresh-checkpoint/no-open-surprise launch rules. It does not close R4's automatic
|
|
source-drift gap or R6's recovery/resolver operational concern. Completed controls test execution
|
|
and integrated evidence review remain pending; repository-creating tests need explicit permission
|
|
for their temporary Git commits or a non-committing equivalent under this review assignment.
|
|
|
|
**Next review:** after owners fix/finish, resume from review-worker-state.md, check changed hashes,
|
|
re-run these negative cases and finalized controls/publishing tests, then independently review
|
|
one stable integrated source-bound gate/replay package. Until then Task B is an initial adversarial
|
|
review with final integration review pending, not an approval of workers' future edits.
|