sots-re/campaign/rollout/lead-state.md

83 lines
6.1 KiB
Markdown

# Lead checkpoint
Authority/model: active session openai/gpt-6-astra. User authorized architecture rollout and later
added five-VM debloating/passwordless lab access. Extra architecture agents explicitly corrected
to GPT-6 Astra. No commits/pushes authorized or performed.
## Active assignments
- Controls, publishing and independent review/pilot: three explicit GPT-6 Astra CLI sessions.
- Gate, engine accounting, Windows housekeeping: explicit GPT-5.6 Terra bounded workers.
- Lead owns AGENTS/CLAUDE entry points, model prompts/config, engine architecture/README/contribution
policy and integration. Bootstrap uses exclusive files in canonical trees; normal runner uses
paired worktrees. Logs in /tmp/opencode/sots-*-worker.log are transport; durable checkpoints in RE.
## Changes / checks
- Added canonical AGENTS policy, six role prompts, project OpenCode configs with explicit model
selection, auto compaction and 40-step quanta. Removed six old global Claude re-* agent files.
- Replaced engine README's stale numerical claims; documented engine/adapters/phase contracts and
existing generated wire-schema channel. Project configuration loads: opencode agent list includes
all six SOTS roles. Need schema/model drift tests once controls registry is final.
- Engine worker reports 52 pass / 7 explicit skips over fresh 59-test inventory and 43 saves.
- Gate worker delivered first version but needs adversarial review fixes before acceptance:
corpusNonzero absence check, missing JUnit completeness, toolchain argument shape, incomplete
asset/input provenance, source-copy integrity and full-profile handling. Main has inspected these.
## Astra resolutions
1. Gate run during concurrent source mutation: invalid measurement, as flagged. No game-mechanism
inference. Wait for engine/tool source quiescence, fresh gate snapshot and compare exact statuses.
2. Research archive installs more hooks than narrative: archived 16-difference observation survives;
complete instrumentation claim is overturned. Pilot requires complete installed-site manifest,
runtime FPU value and neutral control under exact config/binary/assets. No replacement promotion.
3. Default VM policy: a console login or passive VNC viewer alone is not a running test. Housekeeping
may apply non-disruptive background policies when inventory shows no game/instrumentation/test
execution; reboot/logoff requires stronger freedom and access recovery. Actual active work blocks.
4. Independent R2: remove reporter acceptance authority. --require-match only enforces equality;
valid result is measured. Campaign acceptance requires independent verdict + complete integrated
criteria. R4 actual source-content binding and R6 durable recovery are mandatory control fixes.
5. Snapshot clean-room scanner used git commands although gate snapshots have no .git; errors
could lead to a false OK. Replaced with an actual-file scanner supporting canonical Git trees
and plain snapshots, requiring nonzero inspected files and failing read errors. This closes
a gate-integrity defect without changing scientific acceptance scope.
6. Integrated gate b built host+shim and passed 52 tests/7 named skips, but correctly FAILED
corpusNonzero: CTest JUnit truncated successful output at 1024 bytes, removing final counters.
Set explicit output limits, capture verbose output once and fail if any JUnit output truncates.
Keep b failed; run a fresh gate c. Prior run a hit harness timeout and has no accepted manifest.
7. Password located in provisioning ISO and VM146 actual reboot/autologon succeeded using protected
LSA storage. VM141 outage claim was based on .141 instead of confirmed .143; retry correct guest.
Prior housekeeping session ended at quantum limit with outstanding leases; reconcile cleanup
and actual access, explicitly release stale ownership, then GPT-5.5 lab worker finishes.
## Exact next action
Read independent review, repair bounded worker defects, validate controls/model config, run a
fresh integrated host gate and canonical standalone measurement once sources are stable. Verify
housekeeping guest-by-guest, resolve its access/autologon blockers without disabling authentication.
Publish actual evidence/limitations to current pointers, regenerate projections, review both diffs.
## Integration checkpoint — 2026-09-09T22:03Z
- Fresh gate c passed: host + MinGW shim, 59 identities = 52 passed / 7 explicit skips,
all four corpus summaries = 43. Source/copy/tool/input integrity and clean-room scan passed.
- Canonical reporter measured 62 residual state differences; file/inflated/state mismatch remains.
Gate/replay manifests retained under campaign/evidence and selected in campaign/current.json.
- Independent Astra reviewer rehashed both binaries, 524 engine + 10 RE tool source files,
all 43 corpus saves and reproduced the exact 62 diffs; 19 tooling/8 publishing/7 config tests pass.
- Main re-ran final 36 controls + 7 config tests and live config/model validation, all passed.
- REAL normal noninteractive paired-worktree Astra launch succeeded:
campaign/runtime/runs/run-df1472c13f31db3a4d5361f0.json, status complete, actual session and
fresh canonical checkpoint captured. No source edits; no --auto permission bypass on this smoke.
- Remaining: final independent controls disposition and formal scoped promotion, housekeeping
final per-VM login outcomes and cleanup. No original-game replacement or full-assets acceptance.
## Final checkpoint
Rollout complete; authoritative result `campaign/rollout/RESULT.md`. Controls contract accepted
after source-identical paired handoff, real Terra verification/integration runs and fresh independent
final verdict. All five Windows guest profiles compliant; four reboot/autologon proofs, VM140
existing console preserved, all five independently reachable by key SSH with re console active.
All VM/credential leases released; no agent processes remain. 71 Python tests; fresh host+shim
build; 52 CTest passes/7 named skips over 43 saves; canonical replay measured with 62 residuals.
Code remains uncommitted for review. Next: reviewed integration snapshot -> research pilot.