sots-re/campaign/rollout/review-worker-state.md

13 KiB

Independent reviewer checkpoint

CURRENT HANDOFF — follow-up review complete, 2026-09-09

Model openai/gpt-6-astra; logical session review-followup-1; actual harness ID unavailable. Owned review files only changed. No delegates, implementation edits, builds, lab operations, actual-repository staging/commits or held leases. Disposable fixture Git mutations authorized. Verdict and durable recipes/source hashes: campaign/rollout/independent-review.md, follow-up section. 70/70 Python tests passed: tooling 19, publishing 8, config 7, controls 36. Actual loader/model check passed. Selected host/shim package integrity reproduced: 524 engine+10 RE source rows, 43 saves, exact 59 JUnit identities (52 pass/7 permitted skips), four positive [43] summaries. Retained replay hashes/diffs reproduced: measured/nonmatching, 62 state differences. Prior R1-R7 closed within the documented revised scope; full asset/live acceptance remains absent. Latest controls checkpoint integrity passed. Actual architecture-review smoke events/checkpoint and effective config reproduced (run-df1472c13f31db3a4d5361f0); emitted model unavailable. New surprise R8, MEDIUM, independently reproduced: campaign.py:156-159 permits verifier only in verification; run_agent.py:87-89 rejects verifier in integration despite required final verdict refresh. Recorded for lead in owned review; canonical surprise mutation outside owned scope was not performed. Controls acceptance remains pending lead resolution and scoped correction. Remaining checks: R8 regression/review, actual normal worker-role smoke (existing actual smoke is architecture-review), final integrated controls evidence/verdict, separate full asset/live gates. ONE exact next action: lead records and resolves R8 from independent-review.md, then assigns the bounded launcher correction and independent integration-verifier regression before promotion.

Follow-up quantum 1 — 2026-09-09, pre-experiment

Model openai/gpt-6-astra; logical session review-followup-1 (harness ID unavailable). Recovered owned assignment/checkpoint, campaign README/policy and initial R1-R7 findings. Controls repair handoff now says complete; no resource leases or lab/build operations. Selected gate digest 8e14e00ee3ce7478ddfdef8de12183451e858c52d78dc28cffbdee47f5d087d8; selected replay digest 87f92c2b54625ecbca1f3c0a37e57c43d03a4aff488edf4ea8a09a842cd6279d. Read current gate/evidence/reporter code: positive per-test counts, exact JUnit partition, copy byte/mode checks, explicit shim boolean and measured-only reporter implemented. These are observations, not yet independent artifact validation or acceptance. Next exact action: run local tooling/publishing tests and read-only selected-package hash, JUnit, corpus and retained replay checks, then inspect completed controls and run their tests.

Follow-up checkpoint 2 — stable package reproduced

Selected gate/replay digests above independently rehashed. Both binaries match (host 78b2562ea2c56351f9f6f0f24d271afcced741eadbc148047780e6dfd101940b; shim 381c91aecf10a9093f651f58c884356cd88e0253e79cb3bab32753fff6612c7e). All 524 engine + 10 RE execution-tool files match current and retained snapshot bytes/modes; 43 corpus files match. Actual ctest.xml equals embedded rows, 59 unique expected identities, 52 pass/7 permitted skip/0 fail; all four corpus summaries [43]; no output truncation. All eight recorded gate commands exit 0. Recomputed retained input/oracle/output state and byte hashes plus all 62 state diffs match replay manifest (turn2 -> turn3, measured/nonmatching). Tooling 19/19, publishing 8/8, config 7/7 passed. Additional missing binary/execution, duplicate execution, zero corpus and missing/false outputComplete probes rejected. Recipe currently /tmp/opencode/review_followup_checks.py; durable results are this checkpoint and selected immutable manifests. Source review of campaign.py/run_agent.py performed; R4/R6 tests and canonical checkpoint validation not yet executed. No builds/lab/delegation. Next exact action: run verify/campaign tests and live resolved config check, then inspect canonical controls checkpoint/run records and persist final scoped verdict.

Follow-up checkpoint 3 — controls reproduction and final probe

Controls 36/36 passed (6.815s), live check_agent_config.py --resolved passed. Read complete controls tests, source-binding/lifecycle/recovery/runner code and formal R4/R6 decisions; both recorded surprises resolved. Canonical latest controls checkpoint is controls-bootstrap-620eb9d25a6ad2f06b68353c.json. Found actual lead-run smoke record run-df1472c13f31db3a4d5361f0.json: architecture-review/Astra, complete, 24 events, observed model explicitly unavailable. Need independently check event/checkpoint/config hashes. Potential new operational issue: campaign.py ROLE_STATUS allows verifier only in verification, but acceptance requires refreshed final-package verdict in integration. Runner check_launch will reject that independent integration session. Probe in disposable controls fixture prepared; affected launcher acceptance remains pending resolution if reproduced. No source edits/leases. Next exact action: run /tmp/opencode/review_controls_checks.py to validate actual controls/smoke artifacts and reproduce the integration-verifier launch guard, then persist scoped review verdict.

Updated: 2026-09-09T21:30:34Z (handoff checkpoint 6). Model: openai/gpt-6-astra. Session: review-worker initial rollout assignment (no external harness session ID supplied). Phase: pilot delivered; available-tree independent review delivered; integrated re-review pending. Owned files: research-replacement contract/pilot, independent-review, this checkpoint. No delegation, builds, lab I/O, staging or commits performed.

Progress / next action

  • Read architecture-decision, controls assignment, CR findings, archived compare JSON and R1 log.
  • Contract README/schema not published at first inspection; keep pilot proposed and wait for actual API.
  • Next: inspect raw compressed traces/save differences and engine callback/event interfaces; draft pilot, then review available worker code. Gate/report tooling currently modified; workers still implementing.

Surprise for lead: CR narrative understates instrumentation

verify/results/shim/cr/cr-R1.log:75-94 shows seven successful drawsite detours and twelve successful probes in addition to the six detours described in findings/subsystems/research-replace.md:317-332. Thus archived R1 has 25 installed interception sites, not six. Lines 109-115 also show runtime FPU CW 0x127f, whereas findings line 334 says 0x027f in every run (same reported precision/rounding, different full word). Do not inherit narrative as a complete instrumentation manifest. Discriminating check: inspect N/R0 logs and trace metadata; fresh controls must bind complete installed-site manifests and actual runtime FPU words to the same binary/config/assets.

Current blockers

  • Research replace still omits callback effects, actual ObservedTech element and full event records.
  • Archived compare is partial (six undeclared spans, eight unmodelled notes), not acceptance.
  • No current runnable pilot acceptance/dependency decision or controls schema yet observed.

Checkpoint 2 evidence

  • Rehashed all four CR saves and turn3 input; hashes match archive narrative. Re-ran exact-bit, unmasked state checksum: R1 has exactly 16 differences (five primary fields, event/otch records and counts, five derived leaves). Raw R1 trace confirms allocation 144/2898, RNG left 413, full runtime CW 4735 = 0x127f. No current acceptance inferred.
  • sots-engine/src/game/events/research_events.h:95-105 supplies runtime TextLookup; src/app/event_phase.cpp:10-14 already adapts caller-supplied strings. Therefore CR's claim that event text inherently requires original PostEvent is too strong. Runtime user assets are an existing architectural path; the live allocator/ABI adapter and policy remain unresolved.
  • Existing src/shim/hooks/tech_effects.cpp:354-365 DOES have player writeback and an original node-bore updater dependency. Missing is research-cascade integration and full callback effects, not the total absence of a tech-effect writeback implementation. Use ApplyTechCompletion, not the already-researched-guarded ApplyTechEffect (tech_effects.h:156-162).
  • Full callback writes can extend to systems, ships, recursive grants and rebellion objects; CR's five primary player fields are the observed workload subset, not a full write set.

Checkpoint 3

  • Pilot narrative written: campaign/pilots/research-replacement.md (proposed; no acceptance).
  • Controls interface now published in controls-worker-state.md and contract.schema.json; use its string arrays, acceptance {id,axis,criterion}, null checkpoint and full baseline IDs.
  • Current HEAD identities independently read: RE 3bfde5a70d874a723e797a695bbd847fd82c0aa7; engine 7741d42fc5e4e761e6449bdaf0e4a61d00036a23; concurrent changes are uncommitted.
  • Early gate/report inspection shows potential fail-open paths: gate corpusNonzero checks absence of text rather than positive execution; JUnit completeness unbound; reporter unconditional --roundtrip and --accept checks equality only. Record precise findings after current-code check.
  • Read new lead workflow and engine architecture: consistent with separating scoped evidence.
  • Next exact action: write schema-conforming proposed research-replacement.json, then run local tooling unit checks/adversarial reproductions outside repositories and persist review findings.

Checkpoint 4

  • Both pilot files written; JSON conforms structurally to newly published schema, validation next.
  • Available reviewer scope: gate/reporter/evidence, dashboard projection, engine accounting diffs, lead workflow/architecture/config, controls schema/models. Controls CLI/runner/README not yet all published at last read; final controls review remains pending rather than inferred from intent.
  • Engine accounting changes remove parent aggregation and use nonserialized MT word count; tests now compare persisted generator state. No engine build/run by reviewer.
  • Correction to preliminary concern: reporter's --roundtrip validates untouched serialization but engine then DOES execute RunStrategicTurn; it is not an early-return bypass.
  • Prepared /tmp/opencode/review_rollout_probes.py for isolated mock-boundary tests: incomplete JUnit/empty execution output/missing binary gate; equality-only reporter acceptance; host-skip projection mismatch. Uses real save reader for reporter pair; mocks child process/build only. No commits/staging in probes. Existing worker test_source_manifest_detects_mutation creates a temporary commit, so reviewer will omit that test under this assignment's no-commit instruction.
  • Next exact action: run PYTHONDONTWRITEBYTECODE=1 python3 /tmp/opencode/review_rollout_probes.py.

Checkpoint 5 — local results / escalation

  • Probe output: pilot schema subset PASS; gate returned 0/passed with 4 of 59 JUnit cases, empty stdout and binary {}; reporter returned 0/accepted for copy-only mocked process with real identical input/oracle checksums; dashboard rejected an allowed host skip.
  • Four selected tooling tests passed. Omitted temporary-commit test; no repository/engine tests or builds run. Full controls tests/README still absent at last check.
  • Actual controls CLI validation now PASS: PYTHONDONTWRITEBYTECODE=1 python3 tools/campaign.py --state-root /home/alex/sots-re validate research-replacement.
  • Read newly published campaign.py/run_agent.py: baseline evidence source identity is only path+HEAD; end-run source hashes are not tied into acceptance schema. Runner startup also applies 15-minute checkpoint freshness to resume and rejects open surprises even for resolver role. Review pending finalized controls docs/tests. No canonical runtime mutated by reviewer.
  • Next exact action: persist severity-ranked independent-review.md with precise reproduction, source SHA-256 identities, available-scope review and remaining integration review requirements.

Handoff checkpoint 6

  • Delivered all four owned files, untracked/uncommitted; no worker source edits.
  • campaign/contracts/research-replacement.json remains proposed. Actual campaign CLI validation passed; controls README now published/read and format matches. Pilot SHA-256: contract 1108d30bf7e28f30b2446014503fc3200a8b8198a8fab0b866b873a81bf7b8b9; narrative 3101b8959f28ba2c6ffbac02a82b5ae0e841e7d78bd1f40f05b9d1d2278490ad.
  • campaign/rollout/independent-review.md records seven findings (R1-R4 high, R5-R7 medium), exact code locations/source hashes, reproduced R1/R2/R5 boundary failures, test results, historical evidence corrections and specific follow-up gates. Owners need to review/fix; reviewer does not alter their files or certify future changes.
  • Controls README confirms human evaluation of dirty-source evidence and criteria. R4 is explicitly framed as source-binding robustness gap for lead decision, not a false claim of a schema promise.
  • Outstanding: worker fixes/final controls tests, stable integrated-tree gate/replay reproduction, runtime assets and executable pilot acceptance/dependency decisions. No held resource leases.
  • Next exact action: lead resumes this reviewer after worker completion to re-read changed file hashes and re-run R1/R2/R5 adversarial reproductions against the integrated rollout.