flybrain/infra/docs/loop-recovery.md
acamilo 63ecc32b2f loop recovery: an escalation ladder that unsticks a trap on its own
Restart flysim, then reset to the current rung's milestone, then to the
archive below the best rung (never lower), one step per confirmed trap
that outlives the previous one. Two resets a day, three-hour restarts
once they are spent; the ladder starts over at a new best rung or after
six quiet hours. State and history live in the unit's StateDirectory so
a reboot does not forget where the ladder stood.

A router model list confirms each step; a 'not stuck' answer delays it
at most three probes and no answer leaves the watchdog to decide alone.
Each step is announced 60 s ahead in /run/fly/wd/recovery-notice.json
for the stage's recovery splash.

fly-loop-reset is the one new root surface (a sudoers line); 05-deploy
now converges config/fly-sudoers so a release can add it.
2026-09-28 21:23:33 +00:00

82 lines
4.4 KiB
Markdown

# Automated loop recovery
`fly-watchdog` remains report-only. The separate `fly-loop-recover.timer` reads its
`/run/fly/wd/loop.json` every five minutes and unsticks a confirmed trap on its own, climbing
a ladder one step per trap that outlives the previous step:
| level | step | cost |
|---|---|---|
| 0 | `systemctl restart flysim.service` | none: the restore keeps the rung and learned state, clears session macro ledgers |
| 1 | reset to the current rung's milestone archive | progress since the rung was first reached |
| 2+ | reset to the archive below the best rung (never lower) | one visible rung |
- **Confirmed** means two fresh suspected watchdog reports (at most 10 minutes old, `action: none`)
from separate probes about 5 minutes apart. A clear report in between starts the count again.
- **Outlives** means the reports that confirm it again are at least 20 minutes after the step, so
their 10-brain-minute window lies after it. The helper waits that long after every step.
- **Budget:** at most two milestone resets per 24 hours. With the budget spent the step is a
restart, at most every three hours, until a reset is free again.
- **Starting over:** the ladder returns to level 0 when the fly reaches a new best rung, or after
six hours with no suspected report.
A milestone step runs `/opt/fly/bin/fly-loop-reset <rung>` through sudo (the one line in
`config/fly-sudoers`): stop flysim, `fly-reset-to-milestone`, start flysim. flysim is started again
even when the reset fails. The reset copies both stores to `/srv/fly/state.reset-<UTC>` first, as
in the runbook, and needs no deploy when the running release wrote the archive or
`FLY_ACCEPT_ADAPTERS` already names its adapter. After each step the helper waits up to four
minutes for `/status` to report `running` (and the target rank for a reset); otherwise the step is
recorded as failed and the ladder still climbs.
## On stream
Every step is announced 60 seconds ahead (`FLY_LOOP_COUNTDOWN`) in
`/run/fly/wd/recovery-notice.json` (`FLY_RECOVERY_NOTICE`), which the stage's recovery splash
reads. It holds the contract below, written atomically, mode 0644:
```json
{"v": 1, "id": "1790629095-reset", "phase": "countdown|acting|done|failed",
"action": "restart|reset", "fromRung": 12, "fromLabel": "MT. MOON",
"toRung": 11, "toLabel": "BOULDER BADGE", "reason": "unrewarded",
"loop": ["GO OBJECTIVE", "GO WARP"], "stuckSeconds": 900,
"announcedAt": 1790629095, "executeAt": 1790629155, "updatedAt": 1790629160}
```
`toRung`/`toLabel` are present for a reset only. Consumers ignore a notice whose `updatedAt` is
more than 15 minutes old.
## Model confirmation
An OpenAI-compatible router can be asked before each step. The model may only answer
`{"stuck": true|false}` (the reply may wrap it in prose; the first such object counts). It cannot
choose commands, rungs or buttons. `false` delays the step by one probe, at most three times in a
row; then the step goes ahead. If no model answers (unset, rate-limited, down, malformed), the
watchdog's confirmation stands alone. A model can delay recovery by about 15 minutes, never deny
it.
Provision a root-managed `/etc/fly/loop-recovery.env`, mode `0640 root:fly`, outside this public
checkout:
```
FLY_LOOP_ROUTER_URL=<base URL ending in /v1>
FLY_LOOP_ROUTER_KEY=<router key>
FLY_LOOP_MODELS=<model>,<fallback model>,...
```
Models are tried in order within a 90-second budget. Prefer fast free-tier chat models, and put
providers with generous free limits first: free OpenRouter models share a small daily quota and
the router's circuit breaker can close the whole provider for a while. Check each model against
a stuck and a healthy report before listing it. `FLY_LOOP_MODEL` (one model) is still read when
`FLY_LOOP_MODELS` is unset.
## Operating it
- Decisions: `journalctl -u fly-loop-recover.service`. Every step, veto and ladder restart is also
appended to `/var/lib/fly-loop-recover/history.jsonl`.
- Ladder state: `/var/lib/fly-loop-recover/state.json` (the unit's `StateDirectory`, so it
survives a reboot). Deleting it starts the ladder over.
- Stop automatic recovery: `systemctl disable --now fly-loop-recover.timer`.
- Record automatic steps you find in the journal in the host claim log when you next claim the
container; the unit has no access to that log.
This is an unstick mechanism, not a macro bug fix. A recurring trap still needs the
checkpoint-based loop review in `docs/loop-review.md`.