sots-re/campaign/rollout/lead-state.md

6.1 KiB

Lead checkpoint

Authority/model: active session openai/gpt-6-astra. User authorized architecture rollout and later added five-VM debloating/passwordless lab access. Extra architecture agents explicitly corrected to GPT-6 Astra. No commits/pushes authorized or performed.

Active assignments

  • Controls, publishing and independent review/pilot: three explicit GPT-6 Astra CLI sessions.
  • Gate, engine accounting, Windows housekeeping: explicit GPT-5.6 Terra bounded workers.
  • Lead owns AGENTS/CLAUDE entry points, model prompts/config, engine architecture/README/contribution policy and integration. Bootstrap uses exclusive files in canonical trees; normal runner uses paired worktrees. Logs in /tmp/opencode/sots-*-worker.log are transport; durable checkpoints in RE.

Changes / checks

  • Added canonical AGENTS policy, six role prompts, project OpenCode configs with explicit model selection, auto compaction and 40-step quanta. Removed six old global Claude re-* agent files.
  • Replaced engine README's stale numerical claims; documented engine/adapters/phase contracts and existing generated wire-schema channel. Project configuration loads: opencode agent list includes all six SOTS roles. Need schema/model drift tests once controls registry is final.
  • Engine worker reports 52 pass / 7 explicit skips over fresh 59-test inventory and 43 saves.
  • Gate worker delivered first version but needs adversarial review fixes before acceptance: corpusNonzero absence check, missing JUnit completeness, toolchain argument shape, incomplete asset/input provenance, source-copy integrity and full-profile handling. Main has inspected these.

Astra resolutions

  1. Gate run during concurrent source mutation: invalid measurement, as flagged. No game-mechanism inference. Wait for engine/tool source quiescence, fresh gate snapshot and compare exact statuses.
  2. Research archive installs more hooks than narrative: archived 16-difference observation survives; complete instrumentation claim is overturned. Pilot requires complete installed-site manifest, runtime FPU value and neutral control under exact config/binary/assets. No replacement promotion.
  3. Default VM policy: a console login or passive VNC viewer alone is not a running test. Housekeeping may apply non-disruptive background policies when inventory shows no game/instrumentation/test execution; reboot/logoff requires stronger freedom and access recovery. Actual active work blocks.
  4. Independent R2: remove reporter acceptance authority. --require-match only enforces equality; valid result is measured. Campaign acceptance requires independent verdict + complete integrated criteria. R4 actual source-content binding and R6 durable recovery are mandatory control fixes.
  5. Snapshot clean-room scanner used git commands although gate snapshots have no .git; errors could lead to a false OK. Replaced with an actual-file scanner supporting canonical Git trees and plain snapshots, requiring nonzero inspected files and failing read errors. This closes a gate-integrity defect without changing scientific acceptance scope.
  6. Integrated gate b built host+shim and passed 52 tests/7 named skips, but correctly FAILED corpusNonzero: CTest JUnit truncated successful output at 1024 bytes, removing final counters. Set explicit output limits, capture verbose output once and fail if any JUnit output truncates. Keep b failed; run a fresh gate c. Prior run a hit harness timeout and has no accepted manifest.
  7. Password located in provisioning ISO and VM146 actual reboot/autologon succeeded using protected LSA storage. VM141 outage claim was based on .141 instead of confirmed .143; retry correct guest. Prior housekeeping session ended at quantum limit with outstanding leases; reconcile cleanup and actual access, explicitly release stale ownership, then GPT-5.5 lab worker finishes.

Exact next action

Read independent review, repair bounded worker defects, validate controls/model config, run a fresh integrated host gate and canonical standalone measurement once sources are stable. Verify housekeeping guest-by-guest, resolve its access/autologon blockers without disabling authentication. Publish actual evidence/limitations to current pointers, regenerate projections, review both diffs.

Integration checkpoint — 2026-09-09T22:03Z

  • Fresh gate c passed: host + MinGW shim, 59 identities = 52 passed / 7 explicit skips, all four corpus summaries = 43. Source/copy/tool/input integrity and clean-room scan passed.
  • Canonical reporter measured 62 residual state differences; file/inflated/state mismatch remains. Gate/replay manifests retained under campaign/evidence and selected in campaign/current.json.
  • Independent Astra reviewer rehashed both binaries, 524 engine + 10 RE tool source files, all 43 corpus saves and reproduced the exact 62 diffs; 19 tooling/8 publishing/7 config tests pass.
  • Main re-ran final 36 controls + 7 config tests and live config/model validation, all passed.
  • REAL normal noninteractive paired-worktree Astra launch succeeded: campaign/runtime/runs/run-df1472c13f31db3a4d5361f0.json, status complete, actual session and fresh canonical checkpoint captured. No source edits; no --auto permission bypass on this smoke.
  • Remaining: final independent controls disposition and formal scoped promotion, housekeeping final per-VM login outcomes and cleanup. No original-game replacement or full-assets acceptance.

Final checkpoint

Rollout complete; authoritative result campaign/rollout/RESULT.md. Controls contract accepted after source-identical paired handoff, real Terra verification/integration runs and fresh independent final verdict. All five Windows guest profiles compliant; four reboot/autologon proofs, VM140 existing console preserved, all five independently reachable by key SSH with re console active. All VM/credential leases released; no agent processes remain. 71 Python tests; fresh host+shim build; 52 CTest passes/7 named skips over 43 saves; canonical replay measured with 62 residuals. Code remains uncommitted for review. Next: reviewed integration snapshot -> research pilot.