172 lines
13 KiB
Markdown
172 lines
13 KiB
Markdown
# Independent reviewer checkpoint
|
|
|
|
## CURRENT HANDOFF — follow-up review complete, 2026-09-09
|
|
|
|
Model openai/gpt-6-astra; logical session review-followup-1; actual harness ID unavailable.
|
|
Owned review files only changed. No delegates, implementation edits, builds, lab operations,
|
|
actual-repository staging/commits or held leases. Disposable fixture Git mutations authorized.
|
|
Verdict and durable recipes/source hashes: campaign/rollout/independent-review.md, follow-up section.
|
|
70/70 Python tests passed: tooling 19, publishing 8, config 7, controls 36. Actual loader/model
|
|
check passed. Selected host/shim package integrity reproduced: 524 engine+10 RE source rows,
|
|
43 saves, exact 59 JUnit identities (52 pass/7 permitted skips), four positive [43] summaries.
|
|
Retained replay hashes/diffs reproduced: measured/nonmatching, 62 state differences.
|
|
Prior R1-R7 closed within the documented revised scope; full asset/live acceptance remains absent.
|
|
Latest controls checkpoint integrity passed. Actual architecture-review smoke events/checkpoint
|
|
and effective config reproduced (run-df1472c13f31db3a4d5361f0); emitted model unavailable.
|
|
New surprise R8, MEDIUM, independently reproduced: campaign.py:156-159 permits verifier only in
|
|
verification; run_agent.py:87-89 rejects verifier in integration despite required final verdict
|
|
refresh. Recorded for lead in owned review; canonical surprise mutation outside owned scope was
|
|
not performed. Controls acceptance remains pending lead resolution and scoped correction.
|
|
Remaining checks: R8 regression/review, actual normal worker-role smoke (existing actual smoke is
|
|
architecture-review), final integrated controls evidence/verdict, separate full asset/live gates.
|
|
ONE exact next action: lead records and resolves R8 from independent-review.md, then assigns the
|
|
bounded launcher correction and independent integration-verifier regression before promotion.
|
|
|
|
## Follow-up quantum 1 — 2026-09-09, pre-experiment
|
|
|
|
Model openai/gpt-6-astra; logical session review-followup-1 (harness ID unavailable).
|
|
Recovered owned assignment/checkpoint, campaign README/policy and initial R1-R7 findings.
|
|
Controls repair handoff now says complete; no resource leases or lab/build operations.
|
|
Selected gate digest 8e14e00ee3ce7478ddfdef8de12183451e858c52d78dc28cffbdee47f5d087d8;
|
|
selected replay digest 87f92c2b54625ecbca1f3c0a37e57c43d03a4aff488edf4ea8a09a842cd6279d.
|
|
Read current gate/evidence/reporter code: positive per-test counts, exact JUnit partition,
|
|
copy byte/mode checks, explicit shim boolean and measured-only reporter implemented.
|
|
These are observations, not yet independent artifact validation or acceptance.
|
|
Next exact action: run local tooling/publishing tests and read-only selected-package hash,
|
|
JUnit, corpus and retained replay checks, then inspect completed controls and run their tests.
|
|
|
|
### Follow-up checkpoint 2 — stable package reproduced
|
|
|
|
Selected gate/replay digests above independently rehashed. Both binaries match (host
|
|
78b2562ea2c56351f9f6f0f24d271afcced741eadbc148047780e6dfd101940b;
|
|
shim 381c91aecf10a9093f651f58c884356cd88e0253e79cb3bab32753fff6612c7e).
|
|
All 524 engine + 10 RE execution-tool files match current and retained snapshot bytes/modes;
|
|
43 corpus files match. Actual ctest.xml equals embedded rows, 59 unique expected identities,
|
|
52 pass/7 permitted skip/0 fail; all four corpus summaries [43]; no output truncation.
|
|
All eight recorded gate commands exit 0. Recomputed retained input/oracle/output state and byte
|
|
hashes plus all 62 state diffs match replay manifest (turn2 -> turn3, measured/nonmatching).
|
|
Tooling 19/19, publishing 8/8, config 7/7 passed. Additional missing binary/execution,
|
|
duplicate execution, zero corpus and missing/false outputComplete probes rejected.
|
|
Recipe currently /tmp/opencode/review_followup_checks.py; durable results are this checkpoint
|
|
and selected immutable manifests. Source review of campaign.py/run_agent.py performed;
|
|
R4/R6 tests and canonical checkpoint validation not yet executed. No builds/lab/delegation.
|
|
Next exact action: run verify/campaign tests and live resolved config check, then inspect
|
|
canonical controls checkpoint/run records and persist final scoped verdict.
|
|
|
|
### Follow-up checkpoint 3 — controls reproduction and final probe
|
|
|
|
Controls 36/36 passed (6.815s), live check_agent_config.py --resolved passed. Read complete
|
|
controls tests, source-binding/lifecycle/recovery/runner code and formal R4/R6 decisions;
|
|
both recorded surprises resolved. Canonical latest controls checkpoint is
|
|
controls-bootstrap-620eb9d25a6ad2f06b68353c.json. Found actual lead-run smoke record
|
|
run-df1472c13f31db3a4d5361f0.json: architecture-review/Astra, complete, 24 events, observed
|
|
model explicitly unavailable. Need independently check event/checkpoint/config hashes.
|
|
Potential new operational issue: campaign.py ROLE_STATUS allows verifier only in verification,
|
|
but acceptance requires refreshed final-package verdict in integration. Runner check_launch
|
|
will reject that independent integration session. Probe in disposable controls fixture prepared;
|
|
affected launcher acceptance remains pending resolution if reproduced. No source edits/leases.
|
|
Next exact action: run /tmp/opencode/review_controls_checks.py to validate actual controls/smoke
|
|
artifacts and reproduce the integration-verifier launch guard, then persist scoped review verdict.
|
|
|
|
Updated: 2026-09-09T21:30:34Z (handoff checkpoint 6). Model: openai/gpt-6-astra.
|
|
Session: review-worker initial rollout assignment (no external harness session ID supplied).
|
|
Phase: pilot delivered; available-tree independent review delivered; integrated re-review pending.
|
|
Owned files: research-replacement contract/pilot, independent-review, this checkpoint.
|
|
No delegation, builds, lab I/O, staging or commits performed.
|
|
|
|
## Progress / next action
|
|
- Read architecture-decision, controls assignment, CR findings, archived compare JSON and R1 log.
|
|
- Contract README/schema not published at first inspection; keep pilot proposed and wait for actual API.
|
|
- Next: inspect raw compressed traces/save differences and engine callback/event interfaces; draft pilot,
|
|
then review available worker code. Gate/report tooling currently modified; workers still implementing.
|
|
|
|
## Surprise for lead: CR narrative understates instrumentation
|
|
`verify/results/shim/cr/cr-R1.log:75-94` shows seven successful drawsite detours and twelve
|
|
successful probes in addition to the six detours described in findings/subsystems/research-replace.md:317-332.
|
|
Thus archived R1 has 25 installed interception sites, not six. Lines 109-115 also show runtime
|
|
FPU CW 0x127f, whereas findings line 334 says 0x027f in every run (same reported precision/rounding,
|
|
different full word). Do not inherit narrative as a complete instrumentation manifest.
|
|
Discriminating check: inspect N/R0 logs and trace metadata; fresh controls must bind complete
|
|
installed-site manifests and actual runtime FPU words to the same binary/config/assets.
|
|
|
|
## Current blockers
|
|
- Research replace still omits callback effects, actual ObservedTech element and full event records.
|
|
- Archived compare is partial (six undeclared spans, eight unmodelled notes), not acceptance.
|
|
- No current runnable pilot acceptance/dependency decision or controls schema yet observed.
|
|
|
|
## Checkpoint 2 evidence
|
|
- Rehashed all four CR saves and turn3 input; hashes match archive narrative. Re-ran exact-bit,
|
|
unmasked state checksum: R1 has exactly 16 differences (five primary fields, event/otch records
|
|
and counts, five derived leaves). Raw R1 trace confirms allocation 144/2898, RNG left 413,
|
|
full runtime CW 4735 = 0x127f. No current acceptance inferred.
|
|
- `sots-engine/src/game/events/research_events.h:95-105` supplies runtime TextLookup;
|
|
`src/app/event_phase.cpp:10-14` already adapts caller-supplied strings. Therefore CR's claim
|
|
that event text inherently requires original PostEvent is too strong. Runtime user assets
|
|
are an existing architectural path; the live allocator/ABI adapter and policy remain unresolved.
|
|
- Existing `src/shim/hooks/tech_effects.cpp:354-365` DOES have player writeback and an original
|
|
node-bore updater dependency. Missing is research-cascade integration and full callback effects,
|
|
not the total absence of a tech-effect writeback implementation. Use ApplyTechCompletion, not
|
|
the already-researched-guarded ApplyTechEffect (tech_effects.h:156-162).
|
|
- Full callback writes can extend to systems, ships, recursive grants and rebellion objects;
|
|
CR's five primary player fields are the observed workload subset, not a full write set.
|
|
|
|
## Checkpoint 3
|
|
- Pilot narrative written: campaign/pilots/research-replacement.md (proposed; no acceptance).
|
|
- Controls interface now published in controls-worker-state.md and contract.schema.json;
|
|
use its string arrays, acceptance {id,axis,criterion}, null checkpoint and full baseline IDs.
|
|
- Current HEAD identities independently read: RE 3bfde5a70d874a723e797a695bbd847fd82c0aa7;
|
|
engine 7741d42fc5e4e761e6449bdaf0e4a61d00036a23; concurrent changes are uncommitted.
|
|
- Early gate/report inspection shows potential fail-open paths: gate corpusNonzero checks absence
|
|
of text rather than positive execution; JUnit completeness unbound; reporter unconditional
|
|
--roundtrip and --accept checks equality only. Record precise findings after current-code check.
|
|
- Read new lead workflow and engine architecture: consistent with separating scoped evidence.
|
|
- Next exact action: write schema-conforming proposed research-replacement.json, then run local
|
|
tooling unit checks/adversarial reproductions outside repositories and persist review findings.
|
|
|
|
## Checkpoint 4
|
|
- Both pilot files written; JSON conforms structurally to newly published schema, validation next.
|
|
- Available reviewer scope: gate/reporter/evidence, dashboard projection, engine accounting diffs,
|
|
lead workflow/architecture/config, controls schema/models. Controls CLI/runner/README not yet all
|
|
published at last read; final controls review remains pending rather than inferred from intent.
|
|
- Engine accounting changes remove parent aggregation and use nonserialized MT word count;
|
|
tests now compare persisted generator state. No engine build/run by reviewer.
|
|
- Correction to preliminary concern: reporter's --roundtrip validates untouched serialization but
|
|
engine then DOES execute RunStrategicTurn; it is not an early-return bypass.
|
|
- Prepared `/tmp/opencode/review_rollout_probes.py` for isolated mock-boundary tests: incomplete
|
|
JUnit/empty execution output/missing binary gate; equality-only reporter acceptance; host-skip
|
|
projection mismatch. Uses real save reader for reporter pair; mocks child process/build only.
|
|
No commits/staging in probes. Existing worker test_source_manifest_detects_mutation creates a
|
|
temporary commit, so reviewer will omit that test under this assignment's no-commit instruction.
|
|
- Next exact action: run PYTHONDONTWRITEBYTECODE=1 python3 /tmp/opencode/review_rollout_probes.py.
|
|
|
|
## Checkpoint 5 — local results / escalation
|
|
- Probe output: pilot schema subset PASS; gate returned 0/passed with 4 of 59 JUnit cases,
|
|
empty stdout and binary {}; reporter returned 0/accepted for copy-only mocked process with
|
|
real identical input/oracle checksums; dashboard rejected an allowed host skip.
|
|
- Four selected tooling tests passed. Omitted temporary-commit test; no repository/engine tests
|
|
or builds run. Full controls tests/README still absent at last check.
|
|
- Actual controls CLI validation now PASS:
|
|
`PYTHONDONTWRITEBYTECODE=1 python3 tools/campaign.py --state-root /home/alex/sots-re validate research-replacement`.
|
|
- Read newly published campaign.py/run_agent.py: baseline evidence source identity is only path+HEAD;
|
|
end-run source hashes are not tied into acceptance schema. Runner startup also applies 15-minute
|
|
checkpoint freshness to resume and rejects open surprises even for resolver role. Review pending
|
|
finalized controls docs/tests. No canonical runtime mutated by reviewer.
|
|
- Next exact action: persist severity-ranked independent-review.md with precise reproduction,
|
|
source SHA-256 identities, available-scope review and remaining integration review requirements.
|
|
|
|
## Handoff checkpoint 6
|
|
- Delivered all four owned files, untracked/uncommitted; no worker source edits.
|
|
- `campaign/contracts/research-replacement.json` remains proposed. Actual campaign CLI validation
|
|
passed; controls README now published/read and format matches. Pilot SHA-256:
|
|
contract `1108d30bf7e28f30b2446014503fc3200a8b8198a8fab0b866b873a81bf7b8b9`;
|
|
narrative `3101b8959f28ba2c6ffbac02a82b5ae0e841e7d78bd1f40f05b9d1d2278490ad`.
|
|
- `campaign/rollout/independent-review.md` records seven findings (R1-R4 high, R5-R7 medium),
|
|
exact code locations/source hashes, reproduced R1/R2/R5 boundary failures, test results,
|
|
historical evidence corrections and specific follow-up gates. Owners need to review/fix;
|
|
reviewer does not alter their files or certify future changes.
|
|
- Controls README confirms human evaluation of dirty-source evidence and criteria. R4 is explicitly
|
|
framed as source-binding robustness gap for lead decision, not a false claim of a schema promise.
|
|
- Outstanding: worker fixes/final controls tests, stable integrated-tree gate/replay reproduction,
|
|
runtime assets and executable pilot acceptance/dependency decisions. No held resource leases.
|
|
- Next exact action: lead resumes this reviewer after worker completion to re-read changed file
|
|
hashes and re-run R1/R2/R5 adversarial reproductions against the integrated rollout.
|