Restart flysim, then reset to the current rung's milestone, then to the archive below the best rung (never lower), one step per confirmed trap that outlives the previous one. Two resets a day, three-hour restarts once they are spent; the ladder starts over at a new best rung or after six quiet hours. State and history live in the unit's StateDirectory so a reboot does not forget where the ladder stood. A router model list confirms each step; a 'not stuck' answer delays it at most three probes and no answer leaves the watchdog to decide alone. Each step is announced 60 s ahead in /run/fly/wd/recovery-notice.json for the stage's recovery splash. fly-loop-reset is the one new root surface (a sudoers line); 05-deploy now converges config/fly-sudoers so a release can add it.
82 lines
4.4 KiB
Markdown
82 lines
4.4 KiB
Markdown
# Automated loop recovery
|
|
|
|
`fly-watchdog` remains report-only. The separate `fly-loop-recover.timer` reads its
|
|
`/run/fly/wd/loop.json` every five minutes and unsticks a confirmed trap on its own, climbing
|
|
a ladder one step per trap that outlives the previous step:
|
|
|
|
| level | step | cost |
|
|
|---|---|---|
|
|
| 0 | `systemctl restart flysim.service` | none: the restore keeps the rung and learned state, clears session macro ledgers |
|
|
| 1 | reset to the current rung's milestone archive | progress since the rung was first reached |
|
|
| 2+ | reset to the archive below the best rung (never lower) | one visible rung |
|
|
|
|
- **Confirmed** means two fresh suspected watchdog reports (at most 10 minutes old, `action: none`)
|
|
from separate probes about 5 minutes apart. A clear report in between starts the count again.
|
|
- **Outlives** means the reports that confirm it again are at least 20 minutes after the step, so
|
|
their 10-brain-minute window lies after it. The helper waits that long after every step.
|
|
- **Budget:** at most two milestone resets per 24 hours. With the budget spent the step is a
|
|
restart, at most every three hours, until a reset is free again.
|
|
- **Starting over:** the ladder returns to level 0 when the fly reaches a new best rung, or after
|
|
six hours with no suspected report.
|
|
|
|
A milestone step runs `/opt/fly/bin/fly-loop-reset <rung>` through sudo (the one line in
|
|
`config/fly-sudoers`): stop flysim, `fly-reset-to-milestone`, start flysim. flysim is started again
|
|
even when the reset fails. The reset copies both stores to `/srv/fly/state.reset-<UTC>` first, as
|
|
in the runbook, and needs no deploy when the running release wrote the archive or
|
|
`FLY_ACCEPT_ADAPTERS` already names its adapter. After each step the helper waits up to four
|
|
minutes for `/status` to report `running` (and the target rank for a reset); otherwise the step is
|
|
recorded as failed and the ladder still climbs.
|
|
|
|
## On stream
|
|
|
|
Every step is announced 60 seconds ahead (`FLY_LOOP_COUNTDOWN`) in
|
|
`/run/fly/wd/recovery-notice.json` (`FLY_RECOVERY_NOTICE`), which the stage's recovery splash
|
|
reads. It holds the contract below, written atomically, mode 0644:
|
|
|
|
```json
|
|
{"v": 1, "id": "1790629095-reset", "phase": "countdown|acting|done|failed",
|
|
"action": "restart|reset", "fromRung": 12, "fromLabel": "MT. MOON",
|
|
"toRung": 11, "toLabel": "BOULDER BADGE", "reason": "unrewarded",
|
|
"loop": ["GO OBJECTIVE", "GO WARP"], "stuckSeconds": 900,
|
|
"announcedAt": 1790629095, "executeAt": 1790629155, "updatedAt": 1790629160}
|
|
```
|
|
|
|
`toRung`/`toLabel` are present for a reset only. Consumers ignore a notice whose `updatedAt` is
|
|
more than 15 minutes old.
|
|
|
|
## Model confirmation
|
|
|
|
An OpenAI-compatible router can be asked before each step. The model may only answer
|
|
`{"stuck": true|false}` (the reply may wrap it in prose; the first such object counts). It cannot
|
|
choose commands, rungs or buttons. `false` delays the step by one probe, at most three times in a
|
|
row; then the step goes ahead. If no model answers (unset, rate-limited, down, malformed), the
|
|
watchdog's confirmation stands alone. A model can delay recovery by about 15 minutes, never deny
|
|
it.
|
|
|
|
Provision a root-managed `/etc/fly/loop-recovery.env`, mode `0640 root:fly`, outside this public
|
|
checkout:
|
|
|
|
```
|
|
FLY_LOOP_ROUTER_URL=<base URL ending in /v1>
|
|
FLY_LOOP_ROUTER_KEY=<router key>
|
|
FLY_LOOP_MODELS=<model>,<fallback model>,...
|
|
```
|
|
|
|
Models are tried in order within a 90-second budget. Prefer fast free-tier chat models, and put
|
|
providers with generous free limits first: free OpenRouter models share a small daily quota and
|
|
the router's circuit breaker can close the whole provider for a while. Check each model against
|
|
a stuck and a healthy report before listing it. `FLY_LOOP_MODEL` (one model) is still read when
|
|
`FLY_LOOP_MODELS` is unset.
|
|
|
|
## Operating it
|
|
|
|
- Decisions: `journalctl -u fly-loop-recover.service`. Every step, veto and ladder restart is also
|
|
appended to `/var/lib/fly-loop-recover/history.jsonl`.
|
|
- Ladder state: `/var/lib/fly-loop-recover/state.json` (the unit's `StateDirectory`, so it
|
|
survives a reboot). Deleting it starts the ladder over.
|
|
- Stop automatic recovery: `systemctl disable --now fly-loop-recover.timer`.
|
|
- Record automatic steps you find in the journal in the host claim log when you next claim the
|
|
container; the unit has no access to that log.
|
|
|
|
This is an unstick mechanism, not a macro bug fix. A recurring trap still needs the
|
|
checkpoint-based loop review in `docs/loop-review.md`.
|