fly-reset-to-milestone does not check compatibility and a flysim that refuses every checkpoint does not start. fly-loop-reset --check applies 05-deploy's rule (identical, or an adapter-only difference named in FLY_ACCEPT_ADAPTERS); the ladder picks the highest restorable rung and falls back to a restart when there is none.
90 lines
4.9 KiB
Markdown
90 lines
4.9 KiB
Markdown
# Automated loop recovery
|
|
|
|
`fly-watchdog` remains report-only. The separate `fly-loop-recover.timer` reads its
|
|
`/run/fly/wd/loop.json` every five minutes and unsticks a confirmed trap on its own, climbing
|
|
a ladder one step per trap that outlives the previous step:
|
|
|
|
| level | step | cost |
|
|
|---|---|---|
|
|
| 0 | `systemctl restart flysim.service` | none: the restore keeps the rung and learned state, clears session macro ledgers |
|
|
| 1 | reset to the current rung's milestone archive | progress since the rung was first reached |
|
|
| 2+ | reset to the archive below the best rung (never lower) | one visible rung |
|
|
|
|
- **Confirmed** means two fresh suspected watchdog reports (at most 10 minutes old, `action: none`)
|
|
from separate probes about 5 minutes apart. A clear report in between starts the count again.
|
|
- **Outlives** means the reports that confirm it again are at least 20 minutes after the step, so
|
|
their 10-brain-minute window lies after it. The helper waits that long after every step.
|
|
- **Budget:** at most two milestone resets per 24 hours. With the budget spent the step is a
|
|
restart, at most every three hours, until a reset is free again.
|
|
- **Starting over:** the ladder returns to level 0 when the fly reaches a new best rung, or after
|
|
six hours with no suspected report.
|
|
|
|
A milestone step only uses an archive the running build can restore: `fly-loop-reset --check`
|
|
compares the archive's compatibility string with `flysim --print-compatibility` and accepts an
|
|
adapter-only difference that `FLY_ACCEPT_ADAPTERS` in `/etc/fly/fly.env` names (the rule
|
|
`05-deploy.sh` applies). A flysim that refuses every checkpoint refuses to start, so an archive
|
|
from an older adapter is skipped for the next one down unless the deploy named its adapter; with
|
|
none restorable the step is a restart. Keep `FLY_ACCEPT_ADAPTERS` in the release env file so a
|
|
deploy does not drop it.
|
|
|
|
A milestone step runs `/opt/fly/bin/fly-loop-reset <rung>` through sudo (the one line in
|
|
`config/fly-sudoers`): stop flysim, `fly-reset-to-milestone`, start flysim. flysim is started again
|
|
even when the reset fails. The reset copies both stores to `/srv/fly/state.reset-<UTC>` first, as
|
|
in the runbook, and needs no deploy when the running release wrote the archive or
|
|
`FLY_ACCEPT_ADAPTERS` already names its adapter. After each step the helper waits up to four
|
|
minutes for `/status` to report `running` (and the target rank for a reset); otherwise the step is
|
|
recorded as failed and the ladder still climbs.
|
|
|
|
## On stream
|
|
|
|
Every step is announced 60 seconds ahead (`FLY_LOOP_COUNTDOWN`) in
|
|
`/run/fly/wd/recovery-notice.json` (`FLY_RECOVERY_NOTICE`), which the stage's recovery splash
|
|
reads. It holds the contract below, written atomically, mode 0644:
|
|
|
|
```json
|
|
{"v": 1, "id": "1790629095-reset", "phase": "countdown|acting|done|failed",
|
|
"action": "restart|reset", "fromRung": 12, "fromLabel": "MT. MOON",
|
|
"toRung": 11, "toLabel": "BOULDER BADGE", "reason": "unrewarded",
|
|
"loop": ["GO OBJECTIVE", "GO WARP"], "stuckSeconds": 900,
|
|
"announcedAt": 1790629095, "executeAt": 1790629155, "updatedAt": 1790629160}
|
|
```
|
|
|
|
`toRung`/`toLabel` are present for a reset only. Consumers ignore a notice whose `updatedAt` is
|
|
more than 15 minutes old.
|
|
|
|
## Model confirmation
|
|
|
|
An OpenAI-compatible router can be asked before each step. The model may only answer
|
|
`{"stuck": true|false}` (the reply may wrap it in prose; the first such object counts). It cannot
|
|
choose commands, rungs or buttons. `false` delays the step by one probe, at most three times in a
|
|
row; then the step goes ahead. If no model answers (unset, rate-limited, down, malformed), the
|
|
watchdog's confirmation stands alone. A model can delay recovery by about 15 minutes, never deny
|
|
it.
|
|
|
|
Provision a root-managed `/etc/fly/loop-recovery.env`, mode `0640 root:fly`, outside this public
|
|
checkout:
|
|
|
|
```
|
|
FLY_LOOP_ROUTER_URL=<base URL ending in /v1>
|
|
FLY_LOOP_ROUTER_KEY=<router key>
|
|
FLY_LOOP_MODELS=<model>,<fallback model>,...
|
|
```
|
|
|
|
Models are tried in order within a 90-second budget. Prefer fast free-tier chat models, and put
|
|
providers with generous free limits first: free OpenRouter models share a small daily quota and
|
|
the router's circuit breaker can close the whole provider for a while. Check each model against
|
|
a stuck and a healthy report before listing it. `FLY_LOOP_MODEL` (one model) is still read when
|
|
`FLY_LOOP_MODELS` is unset.
|
|
|
|
## Operating it
|
|
|
|
- Decisions: `journalctl -u fly-loop-recover.service`. Every step, veto and ladder restart is also
|
|
appended to `/var/lib/fly-loop-recover/history.jsonl`.
|
|
- Ladder state: `/var/lib/fly-loop-recover/state.json` (the unit's `StateDirectory`, so it
|
|
survives a reboot). Deleting it starts the ladder over.
|
|
- Stop automatic recovery: `systemctl disable --now fly-loop-recover.timer`.
|
|
- Record automatic steps you find in the journal in the host claim log when you next claim the
|
|
container; the unit has no access to that log.
|
|
|
|
This is an unstick mechanism, not a macro bug fix. A recurring trap still needs the
|
|
checkpoint-based loop review in `docs/loop-review.md`.
|