From 6655a1b1c6a939ed8a51513db074eb614d288915 Mon Sep 17 00:00:00 2001 From: dev Date: Tue, 22 Sep 2026 17:43:43 +0000 Subject: [PATCH 1/2] session: coherent all-participant checkpoint and recovery STATE-01 over the FLYSESS1 envelope CONTRACT-01 specified. fly-session gains a `state` module: the durable store with its generations, its rotation and the commit order of checkpoint-envelope-v1 section 5, where the store manifest rename is the durable commit point; a compatibility block whose comparison names the identity that differs rather than one opaque digest; and a bounded writer that owns its payload handles until the bytes are committed or the job fails. The writer's queue slot is taken before the first State.Capture, so a saturated writer refuses a capture rather than queueing it without bound, and the refusal is a BUSY the stepping session survives. Capture and durability are two events: a capture completes when an immutable capture exists, and only the store manifest rename moves the durable mark. A lost save reply is an outcome, and the resolution asks the store about the same checkpoint instead of saving again. Both worker roles implement State.Capture, State.StageRestore and State.ActivateRestore, with once-only restore tokens bound to checkpoint, scope, payload and incarnation. A restore selects a complete compatible generation, imports every payload as a fresh artifact, stages the group, validates the coordinator's own ledgers, and only then activates; a failure anywhere leaves the fence closed and records every participant that staged as one that must be replaced. The fence lifts at Failed -> Restoring(k) -> Paused(k) and nowhere else. The task and the action executor gain the capture/validate_restore/install_restore interfaces workers-v1 section 4 lists, and the ledger can re-derive the event identities it issued under another epoch, which is what lets a resumed run's behaviour trace be compared with an uninterrupted one. media: check_required_audio now takes the observation's provenance instead of exempting boundary 0. A chunk is the audio of an interval, and the observation ActivateRestore installs covers none. checkpoint-envelope-v1 section 3 gains a dated amendment adding `environment` to the manifest, the holder of the world's own payload, which the table named for every other participant; `helperState`, which that table already listed, joins the required-field set in Rust and TypeScript. The fixture was regenerated by the existing example; the schema set and contractDigest are unchanged. state-media-v1 section 5 gains a dated amendment for three readings this slice enforces: the State RPCs' compatibilityDigest is the participant's, not the manifest's composition-level block; a restored observation carries no audio chunk; and a participant that staged into an abandoned install must be replaced. --- .../checkpoint-envelope-v1.md | 13 + .../session-framework/state-media-v1.md | 22 + packages/session-types/src/checkpoint.ts | 9 +- .../examples/update_fixtures.rs | 1 + .../fixtures/checkpoint-envelope.json | 32 +- .../fly-session-types/src/checkpoint.rs | 7 + services/flysim/crates/fly-session/README.md | 50 +- .../flysim/crates/fly-session/src/agent.rs | 573 ++++++- services/flysim/crates/fly-session/src/cli.rs | 17 + .../flysim/crates/fly-session/src/clock.rs | 24 + .../crates/fly-session/src/coordinator.rs | 1489 +++++++++++++++- .../crates/fly-session/src/environment.rs | 517 +++++- .../flysim/crates/fly-session/src/harness.rs | 179 +- .../flysim/crates/fly-session/src/launcher.rs | 22 + services/flysim/crates/fly-session/src/lib.rs | 1 + .../flysim/crates/fly-session/src/media.rs | 144 +- .../flysim/crates/fly-session/src/state.rs | 1494 +++++++++++++++++ .../flysim/crates/fly-session/src/task.rs | 249 +++ .../flysim/crates/fly-session/src/types.rs | 80 +- .../flysim/crates/fly-session/src/worker.rs | 12 +- .../crates/fly-session/tests/processes.rs | 3 + .../flysim/crates/fly-session/tests/state.rs | 947 +++++++++++ 22 files changed, 5837 insertions(+), 48 deletions(-) create mode 100644 services/flysim/crates/fly-session/src/state.rs create mode 100644 services/flysim/crates/fly-session/tests/state.rs diff --git a/docs/design/session-framework/checkpoint-envelope-v1.md b/docs/design/session-framework/checkpoint-envelope-v1.md index c83bb70..7cb833f 100644 --- a/docs/design/session-framework/checkpoint-envelope-v1.md +++ b/docs/design/session-framework/checkpoint-envelope-v1.md @@ -105,6 +105,19 @@ names; a manifest missing any of them is not a complete checkpoint. | `helperState` | External-helper state required for exact resume, as payload names | | `payloads` | `[{name, byteLength, digest}]`, mirroring the payload table | +**Amendment, 2026-09-22 (STATE-01).** The table above names a holder for every payload except +the environment's own, although section 6's fixture has one (`world`) and a group install has +to map it by name like any other participant's. The manifest therefore also records: + +| Field | Contents | +| --- | --- | +| `environment` | `{workerId, payload}`: which worker the world belonged to and the payload name holding its state | + +The reference implementations' required-field set was also missing `helperState`, which this +section has listed from the start. Both are now in `REQUIRED_MANIFEST_FIELDS` in Rust and in +TypeScript, and the fixture was regenerated by the existing example. The schema set is +untouched, so `contractDigest` is unchanged. + `payloads` is redundant with the table on purpose: the table is what a reader needs to map bytes, and the manifest is what a store lists, compares and reports without opening the payload area. A reader checks that the two agree. diff --git a/docs/design/session-framework/state-media-v1.md b/docs/design/session-framework/state-media-v1.md index 2669786..0061761 100644 --- a/docs/design/session-framework/state-media-v1.md +++ b/docs/design/session-framework/state-media-v1.md @@ -190,6 +190,28 @@ restored time. It cannot advance gameplay to manufacture it. Capture/reconstruct covers render/inspection state and any pending sensor pipeline. Agent state agrees with it; do not replay reward or recalibrate merely to fill missing cached data. +**Amendment, 2026-09-22 (STATE-01).** Three readings of this section, made explicit because +they are now enforced: + +- `compatibilityDigest` on `CaptureResult` and `StageRestoreParams` is the **participant's** + capture compatibility digest of [worker interfaces](workers-v1.md) section 2 -- profile, + resolved seed, numerical model version and effective instance configuration for an agent; + backend, content, patch, controller and parser identity for an environment. It is not the + manifest's `compatibility` block of section 4, which is the composition's and which the + coordinator compares before anything is asked to stage. Both exist because they answer + different questions, and a restore that passed the second could still be handing an agent + another agent's brain. +- The observation `ActivateRestore` returns ran no transition, so it carries **no audio + chunk**, and one in it is refused. Section 2's chunk is the audio of an interval and this + observation covers none; MEDIA-01 implemented that rule as "boundary 0 carries no chunk", + which is true of the only such observation that slice could produce and false of this one. + The rule is about provenance, not about the boundary number. +- A participant that staged into a group install the coordinator then abandoned must be + **replaced** before another restore, exactly as one that activated must. It is holding a + validated replacement state that nothing installed, and [session RPC](ipc-v1.md) section 6 + already refuses to silently reattach such a participant to an active epoch. Without this the + group's second attempt meets its own leftovers and calls them a conflict. + If emulator validation requires mutation, stage a stopped replacement emulator. If that cannot provide externally atomic resume, advertise episode-restart, not exact-checkpoint. After all activation acknowledgments, install the coordinator's staged task/executor/admission state diff --git a/packages/session-types/src/checkpoint.ts b/packages/session-types/src/checkpoint.ts index 63aac19..2df48f7 100644 --- a/packages/session-types/src/checkpoint.ts +++ b/packages/session-types/src/checkpoint.ts @@ -223,7 +223,12 @@ export function decode(input: Uint8Array): Envelope { }; } -/** The manifest fields state-media-v1 section 4 requires. */ +/** + * The manifest fields state-media-v1 section 4 requires. + * + * `helperState` and `environment` join the list under the 2026-09-22 amendment to + * checkpoint-envelope-v1 section 3. + */ export const REQUIRED_MANIFEST_FIELDS = [ 'envelopeVersion', 'checkpointId', @@ -236,6 +241,8 @@ export const REQUIRED_MANIFEST_FIELDS = [ 'compatibility', 'agents', 'coordinator', + 'environment', + 'helperState', 'payloads', ] as const; diff --git a/services/flysim/crates/fly-session-types/examples/update_fixtures.rs b/services/flysim/crates/fly-session-types/examples/update_fixtures.rs index 238287e..d68b865 100644 --- a/services/flysim/crates/fly-session-types/examples/update_fixtures.rs +++ b/services/flysim/crates/fly-session-types/examples/update_fixtures.rs @@ -199,6 +199,7 @@ fn checkpoint_envelope() -> String { "admissionState": null, "eventWatermarks": {"lastEventId": "evt-1", "lastOrdinal": "7"}, }, + "environment": {"workerId": "arena", "payload": "world"}, "helperState": [], "payloads": payload_table(), }); diff --git a/services/flysim/crates/fly-session-types/fixtures/checkpoint-envelope.json b/services/flysim/crates/fly-session-types/fixtures/checkpoint-envelope.json index 6636bae..70453c3 100644 --- a/services/flysim/crates/fly-session-types/fixtures/checkpoint-envelope.json +++ b/services/flysim/crates/fly-session-types/fixtures/checkpoint-envelope.json @@ -63,6 +63,10 @@ "lastOrdinal": "7" } }, + "environment": { + "workerId": "arena", + "payload": "world" + }, "helperState": [], "payloads": [ { @@ -115,49 +119,49 @@ } ], "envelope": { - "base64": "RkxZU0VTUzEBAAAAIAAAAAAIAAAFAAAAIAgAAAAAAAB7ImFnZW50cyI6W3siYWdlbnRJZCI6ImZseS1hIiwiYnJhaW5UaWNrcyI6IjI1MzQiLCJkYXRhc2V0RGlnZXN0IjoiNmMwYWYxZjA3ODRlZjYzYTM5M2VlNzdkNjE0ZTgyNDZjNjI1MDUxMzYwZjNmMWE0ODgzODM3NGM1ZDM1NWI1MiIsIm1vZGVsVmVyc2lvbiI6ImxpZi0xbXMtZjY0LXYyIiwicGF5bG9hZCI6ImFnZW50LWZseS1hIiwicGxhc3RpY2l0eVZlcnNpb24iOiJmbHkta2MtbWJvbi1yc3RkcC12MiIsInByb2ZpbGVEaWdlc3QiOiIxOTAwZWFiNmMwMjg0ODNkNzEyNjU5OWVlNmY1MGRlMGQyNzkwN2I1YzY1ZmE5MDUyNDU4MGI0YjBmOTg1MmIwIiwicmVtYWluZGVyIjp7ImRlbm9taW5hdG9yIjoiMyIsIm51bWVyYXRvciI6IjEwMDAwMDAifSwic2VlZCI6LTE4NDk0NjA2M31dLCJjaGVja3BvaW50SWQiOiJja3B0LTEiLCJjb21wYXRpYmlsaXR5Ijp7ImJhY2tlbmREaWdlc3QiOiIxMGUwOGE0MTllODUwZWJhMWViYmExOGZkZDI4ZWI3ZWMxYjdlOGJhYTliY2MzYjk3M2UyYjg4OTFlYzcyNmJlIiwiY29udGVudERpZ2VzdCI6ImVkNzAwMmI0MzllOWFjODQ1ZjIyMzU3ZDgyMmJhYzE0NDQ3MzBmYmRiNjAxNmQzZWM5NDMyMjk3YjllYzlmNzMiLCJjb250cm9sbGVyRGlnZXN0IjoiYzE0NzIxMzViMTRjNzdjOGJlZjk4ZTczZjcwMjA4MzI1ZmEwZGNmMWU2YmQ2NjhhZTliMzFhOWNlYTI5NWZlNyIsInBhcnNlckRpZ2VzdCI6ImIxN2Q0NTEyMTE1MDkyOGYyMTQ2YWY0OWUxOTVlZmYxZWVmNWQ2NzMyNWJlMjczYTczM2ZiNzRhY2FkYWEzNDIiLCJwYXRjaERpZ2VzdCI6ImE0ODk1ZWI0NGFmYzMzNmZlY2JiYTZlNTIwY2Q2N2UxNzhkYWNlMDI3NjY1NWQxMDJmY2VmZmE4ZTVmNzA1NzAiLCJzdGF0ZUZvcm1hdElkIjoiZmx5c2Vzcy0xIn0sImNvbXBvc2l0aW9uRGlnZXN0IjoiNzMwZDcyNWM4YTU5ZDNhNzMwM2RlZjJiZWQwNDFhNTc3ZWRiNDI1NWFhYmQ0ODg5Y2UxMjkxODMxMWQ5NTJmMCIsImNvb3JkaW5hdG9yIjp7ImFkbWlzc2lvblN0YXRlIjpudWxsLCJldmVudFdhdGVybWFya3MiOnsibGFzdEV2ZW50SWQiOiJldnQtMSIsImxhc3RPcmRpbmFsIjoiNyJ9LCJleGVjdXRvclN0YXRlIjpbeyJhZ2VudElkIjoiZmx5LWEiLCJwYXlsb2FkIjoiZXhlY3V0b3ItZmx5LWEifV0sInByaW9ySW5zcGVjdGlvbiI6InByaW9yLWluc3BlY3Rpb24iLCJ0YXNrTGVkZ2VyIjoidGFzay1sZWRnZXIifSwiZW52ZWxvcGVWZXJzaW9uIjoxLCJlcGlzb2RlSWQiOiJlcGlzb2RlLTEiLCJoZWxwZXJTdGF0ZSI6W10sInBheWxvYWRzIjpbeyJieXRlTGVuZ3RoIjoiMTciLCJkaWdlc3QiOiIxMzIxZGZmYjBjZGM2ZjkwOTJjYmY3ZmEyYTVmYzY4YmJlZDEyYzk5M2Q1YWQzOTgyNjQwMTI4MTBjZTliZjkzIiwibmFtZSI6ImFnZW50LWZseS1hIn0seyJieXRlTGVuZ3RoIjoiMTQiLCJkaWdlc3QiOiIzYWVlNjBkZjdlMjllZmViYTdmNWY5OWZjNTg2NzY0N2IzNmFlYmZmMWQ1ZDNjODM4ZGJmZjMyMzEyMmU2NDYyIiwibmFtZSI6ImV4ZWN1dG9yLWZseS1hIn0seyJieXRlTGVuZ3RoIjoiMTEiLCJkaWdlc3QiOiI0MGIwMGVkMmJiYmE5MDFkNjgyMDVmZjcxYjA0YTQ0YjllZTUzYzUxY2IzMTA5YWEyY2VhYTQ0ZjFjNDU3MjdlIiwibmFtZSI6InRhc2stbGVkZ2VyIn0seyJieXRlTGVuZ3RoIjoiMTAiLCJkaWdlc3QiOiIyYzEzYjdiNGQ5YTk5MTY4MDFhYjkxOTFjMzE0ZjMxYjA0NWU5YjljNWI2NjlhNmMwNDc0ZjAyMTdlZjc1YmY1IiwibmFtZSI6InByaW9yLWluc3BlY3Rpb24ifSx7ImJ5dGVMZW5ndGgiOiI2NCIsImRpZ2VzdCI6ImY1YTVmZDQyZDE2YTIwMzAyNzk4ZWY2ZWQzMDk5NzliNDMwMDNkMjMyMGQ5ZjBlOGVhOTgzMWE5Mjc1OWZiNGIiLCJuYW1lIjoid29ybGQifV0sInBvcnRNYXAiOlt7ImFnZW50SWQiOiJmbHktYSIsInBvcnRJZCI6InBvcnQtMSJ9XSwic2NoZWR1bGVySWQiOiJsb2Nrc3RlcC12MSIsInNvdXJjZVNjb3BlIjp7ImVwb2NoIjoiZXBvY2gtMSIsInNlc3Npb25JZCI6ImRlbW8iLCJzdGVwIjoiNDIifSwid29ybGRUaW1lIjp7ImRlbm9taW5hdG9yIjoiMSIsIm51bWVyYXRvciI6IjcwMDAwMDAwMCJ9fWFnZW50LWZseS1hAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAABQCgAAAAAAABEAAAAAAAAAEyHf+wzcb5CSy/f6Kl/Gi77RLJk9WtOYJkASgQzpv5NleGVjdXRvci1mbHktYQAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAaAoAAAAAAAAOAAAAAAAAADruYN9+Ke/rp/X5n8WGdkezauv/HV08g42/8yMSLmRidGFzay1sZWRnZXIAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAHgKAAAAAAAACwAAAAAAAABAsA7Su7qQHWggX/cbBKRLnuU8UcsxCaos6qRPHEVyfnByaW9yLWluc3BlY3Rpb24AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACICgAAAAAAAAoAAAAAAAAALBO3tNmpkWgBq5GRwxTzGwRem5xbZppsBHTwIX73W/V3b3JsZAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAmAoAAAAAAABAAAAAAAAAAPWl/ULRaiAwJ5jvbtMJl5tDAD0jINnw6OqYMaknWftLYWdlbnQgc3RhdGUgYnl0ZXMAAAAAAAAAZXhlY3V0b3Igc3RhdGUAAHsicmFuayI6MTB9AAAAAAB7Im1hcCI6NDB9AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAgLAAAAAAAAq++fEx+FvDZho/eB4imbENN4HZrGNC2OCAsI7/gp9r5GTFlTRVNTRg==", - "byteLength": 2824, + "base64": "RkxZU0VTUzEBAAAAIAAAADUIAAAFAAAAWAgAAAAAAAB7ImFnZW50cyI6W3siYWdlbnRJZCI6ImZseS1hIiwiYnJhaW5UaWNrcyI6IjI1MzQiLCJkYXRhc2V0RGlnZXN0IjoiNmMwYWYxZjA3ODRlZjYzYTM5M2VlNzdkNjE0ZTgyNDZjNjI1MDUxMzYwZjNmMWE0ODgzODM3NGM1ZDM1NWI1MiIsIm1vZGVsVmVyc2lvbiI6ImxpZi0xbXMtZjY0LXYyIiwicGF5bG9hZCI6ImFnZW50LWZseS1hIiwicGxhc3RpY2l0eVZlcnNpb24iOiJmbHkta2MtbWJvbi1yc3RkcC12MiIsInByb2ZpbGVEaWdlc3QiOiIxOTAwZWFiNmMwMjg0ODNkNzEyNjU5OWVlNmY1MGRlMGQyNzkwN2I1YzY1ZmE5MDUyNDU4MGI0YjBmOTg1MmIwIiwicmVtYWluZGVyIjp7ImRlbm9taW5hdG9yIjoiMyIsIm51bWVyYXRvciI6IjEwMDAwMDAifSwic2VlZCI6LTE4NDk0NjA2M31dLCJjaGVja3BvaW50SWQiOiJja3B0LTEiLCJjb21wYXRpYmlsaXR5Ijp7ImJhY2tlbmREaWdlc3QiOiIxMGUwOGE0MTllODUwZWJhMWViYmExOGZkZDI4ZWI3ZWMxYjdlOGJhYTliY2MzYjk3M2UyYjg4OTFlYzcyNmJlIiwiY29udGVudERpZ2VzdCI6ImVkNzAwMmI0MzllOWFjODQ1ZjIyMzU3ZDgyMmJhYzE0NDQ3MzBmYmRiNjAxNmQzZWM5NDMyMjk3YjllYzlmNzMiLCJjb250cm9sbGVyRGlnZXN0IjoiYzE0NzIxMzViMTRjNzdjOGJlZjk4ZTczZjcwMjA4MzI1ZmEwZGNmMWU2YmQ2NjhhZTliMzFhOWNlYTI5NWZlNyIsInBhcnNlckRpZ2VzdCI6ImIxN2Q0NTEyMTE1MDkyOGYyMTQ2YWY0OWUxOTVlZmYxZWVmNWQ2NzMyNWJlMjczYTczM2ZiNzRhY2FkYWEzNDIiLCJwYXRjaERpZ2VzdCI6ImE0ODk1ZWI0NGFmYzMzNmZlY2JiYTZlNTIwY2Q2N2UxNzhkYWNlMDI3NjY1NWQxMDJmY2VmZmE4ZTVmNzA1NzAiLCJzdGF0ZUZvcm1hdElkIjoiZmx5c2Vzcy0xIn0sImNvbXBvc2l0aW9uRGlnZXN0IjoiNzMwZDcyNWM4YTU5ZDNhNzMwM2RlZjJiZWQwNDFhNTc3ZWRiNDI1NWFhYmQ0ODg5Y2UxMjkxODMxMWQ5NTJmMCIsImNvb3JkaW5hdG9yIjp7ImFkbWlzc2lvblN0YXRlIjpudWxsLCJldmVudFdhdGVybWFya3MiOnsibGFzdEV2ZW50SWQiOiJldnQtMSIsImxhc3RPcmRpbmFsIjoiNyJ9LCJleGVjdXRvclN0YXRlIjpbeyJhZ2VudElkIjoiZmx5LWEiLCJwYXlsb2FkIjoiZXhlY3V0b3ItZmx5LWEifV0sInByaW9ySW5zcGVjdGlvbiI6InByaW9yLWluc3BlY3Rpb24iLCJ0YXNrTGVkZ2VyIjoidGFzay1sZWRnZXIifSwiZW52ZWxvcGVWZXJzaW9uIjoxLCJlbnZpcm9ubWVudCI6eyJwYXlsb2FkIjoid29ybGQiLCJ3b3JrZXJJZCI6ImFyZW5hIn0sImVwaXNvZGVJZCI6ImVwaXNvZGUtMSIsImhlbHBlclN0YXRlIjpbXSwicGF5bG9hZHMiOlt7ImJ5dGVMZW5ndGgiOiIxNyIsImRpZ2VzdCI6IjEzMjFkZmZiMGNkYzZmOTA5MmNiZjdmYTJhNWZjNjhiYmVkMTJjOTkzZDVhZDM5ODI2NDAxMjgxMGNlOWJmOTMiLCJuYW1lIjoiYWdlbnQtZmx5LWEifSx7ImJ5dGVMZW5ndGgiOiIxNCIsImRpZ2VzdCI6IjNhZWU2MGRmN2UyOWVmZWJhN2Y1Zjk5ZmM1ODY3NjQ3YjM2YWViZmYxZDVkM2M4MzhkYmZmMzIzMTIyZTY0NjIiLCJuYW1lIjoiZXhlY3V0b3ItZmx5LWEifSx7ImJ5dGVMZW5ndGgiOiIxMSIsImRpZ2VzdCI6IjQwYjAwZWQyYmJiYTkwMWQ2ODIwNWZmNzFiMDRhNDRiOWVlNTNjNTFjYjMxMDlhYTJjZWFhNDRmMWM0NTcyN2UiLCJuYW1lIjoidGFzay1sZWRnZXIifSx7ImJ5dGVMZW5ndGgiOiIxMCIsImRpZ2VzdCI6IjJjMTNiN2I0ZDlhOTkxNjgwMWFiOTE5MWMzMTRmMzFiMDQ1ZTliOWM1YjY2OWE2YzA0NzRmMDIxN2VmNzViZjUiLCJuYW1lIjoicHJpb3ItaW5zcGVjdGlvbiJ9LHsiYnl0ZUxlbmd0aCI6IjY0IiwiZGlnZXN0IjoiZjVhNWZkNDJkMTZhMjAzMDI3OThlZjZlZDMwOTk3OWI0MzAwM2QyMzIwZDlmMGU4ZWE5ODMxYTkyNzU5ZmI0YiIsIm5hbWUiOiJ3b3JsZCJ9XSwicG9ydE1hcCI6W3siYWdlbnRJZCI6ImZseS1hIiwicG9ydElkIjoicG9ydC0xIn1dLCJzY2hlZHVsZXJJZCI6ImxvY2tzdGVwLXYxIiwic291cmNlU2NvcGUiOnsiZXBvY2giOiJlcG9jaC0xIiwic2Vzc2lvbklkIjoiZGVtbyIsInN0ZXAiOiI0MiJ9LCJ3b3JsZFRpbWUiOnsiZGVub21pbmF0b3IiOiIxIiwibnVtZXJhdG9yIjoiNzAwMDAwMDAwIn19AAAAYWdlbnQtZmx5LWEAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAIgKAAAAAAAAEQAAAAAAAAATId/7DNxvkJLL9/oqX8aLvtEsmT1a05gmQBKBDOm/k2V4ZWN1dG9yLWZseS1hAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACgCgAAAAAAAA4AAAAAAAAAOu5g334p7+un9fmfxYZ2R7Nq6/8dXTyDjb/zIxIuZGJ0YXNrLWxlZGdlcgAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAsAoAAAAAAAALAAAAAAAAAECwDtK7upAdaCBf9xsEpEue5TxRyzEJqizqpE8cRXJ+cHJpb3ItaW5zcGVjdGlvbgAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAMAKAAAAAAAACgAAAAAAAAAsE7e02amRaAGrkZHDFPMbBF6bnFtmmmwEdPAhfvdb9XdvcmxkAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAADQCgAAAAAAAEAAAAAAAAAA9aX9QtFqIDAnmO9u0wmXm0MAPSMg2fDo6pgxqSdZ+0thZ2VudCBzdGF0ZSBieXRlcwAAAAAAAABleGVjdXRvciBzdGF0ZQAAeyJyYW5rIjoxMH0AAAAAAHsibWFwIjo0MH0AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAQAsAAAAAAAAhBlh9AWLTtVKckmeIzNn4DO5Yhn2C1nUA1T1RfxMXlkZMWVNFU1NG", + "byteLength": 2880, "layout": { "headerBytes": 32, "manifestOffset": "32", - "manifestBytes": 2048, - "tableOffset": "2080", + "manifestBytes": 2101, + "tableOffset": "2136", "tableEntryBytes": 112, "entries": [ { "name": "agent-fly-a", - "offset": "2640", + "offset": "2696", "byteLength": "17", "digest": "1321dffb0cdc6f9092cbf7fa2a5fc68bbed12c993d5ad398264012810ce9bf93" }, { "name": "executor-fly-a", - "offset": "2664", + "offset": "2720", "byteLength": "14", "digest": "3aee60df7e29efeba7f5f99fc5867647b36aebff1d5d3c838dbff323122e6462" }, { "name": "task-ledger", - "offset": "2680", + "offset": "2736", "byteLength": "11", "digest": "40b00ed2bbba901d68205ff71b04a44b9ee53c51cb3109aa2ceaa44f1c45727e" }, { "name": "prior-inspection", - "offset": "2696", + "offset": "2752", "byteLength": "10", "digest": "2c13b7b4d9a9916801ab9191c314f31b045e9b9c5b669a6c0474f0217ef75bf5" }, { "name": "world", - "offset": "2712", + "offset": "2768", "byteLength": "64", "digest": "f5a5fd42d16a20302798ef6ed309979b43003d2320d9f0e8ea9831a92759fb4b" } ], - "footerOffset": "2776", + "footerOffset": "2832", "footerBytes": 48, - "totalBytes": "2824" + "totalBytes": "2880" } }, "corruption": [ @@ -178,17 +182,17 @@ }, { "name": "a flipped payload byte", - "offset": 2640, + "offset": 2696, "reason": "every payload carries its own digest" }, { "name": "a flipped footer digest byte", - "offset": 2784, + "offset": 2840, "reason": "the footer digest must match the contents" }, { "name": "a flipped footer magic byte", - "offset": 2816, + "offset": 2872, "reason": "a truncated file cannot look complete" } ] diff --git a/services/flysim/crates/fly-session-types/src/checkpoint.rs b/services/flysim/crates/fly-session-types/src/checkpoint.rs index 341f5c5..985fcb1 100644 --- a/services/flysim/crates/fly-session-types/src/checkpoint.rs +++ b/services/flysim/crates/fly-session-types/src/checkpoint.rs @@ -296,6 +296,11 @@ pub fn decode(bytes: &[u8]) -> Result { /// The manifest fields state-media-v1 section 4 requires, checked as a set: a manifest that /// omits one of them is not a complete checkpoint. +/// +/// `helperState` and `environment` join the list under the 2026-09-22 amendment to +/// checkpoint-envelope-v1 section 3: the first has been in that section's table from the +/// start and was missing here, and the second is the holder of the world's own payload, which +/// the table named for every other participant and not for the environment. pub const REQUIRED_MANIFEST_FIELDS: &[&str] = &[ "envelopeVersion", "checkpointId", @@ -308,6 +313,8 @@ pub const REQUIRED_MANIFEST_FIELDS: &[&str] = &[ "compatibility", "agents", "coordinator", + "environment", + "helperState", "payloads", ]; diff --git a/services/flysim/crates/fly-session/README.md b/services/flysim/crates/fly-session/README.md index 8e114e2..33f0594 100644 --- a/services/flysim/crates/fly-session/README.md +++ b/services/flysim/crates/fly-session/README.md @@ -44,6 +44,7 @@ Ready(k) ─ Prepare all agents concurrently ─────────── | `metrics` | Latency percentiles and the machine's core and memory counters | | `measure` | The execution-mode comparison of the guide's section 5 | | `cli` | The binary's subcommands: `agent`, `environment`, `measure` | +| `state` | The durable checkpoint store over `FLYSESS1`: compatibility, generations, the bounded writer | | `harness` | The runnable composition: router, the flies, one arena, one coordinator | ## Execution modes and the launcher @@ -178,6 +179,38 @@ harness.shutdown().await; event ids derived from epoch, source step, rule and ordinal. - **Executors.** The stateless identity executor only, as v1 specifies. +## Checkpoints and recovery + +The durable store is `state`, over the `FLYSESS1` layout the contract crate owns. + +- **One boundary, every participant.** `Coordinator::capture` runs at `Ready(k)` or + `Paused(k)` only. It takes its queue slot *before* the first `State.Capture`, so a saturated + writer refuses the capture rather than queueing it without bound, and the refusal is a + `BUSY` a stepping session survives rather than an epoch failure. +- **Capture and durability are two events.** `State.Capture` completes when an immutable + capture exists; `Coordinator::await_durable` completes when the store manifest rename has + happened, which is the durable commit point. Only the second moves the durable mark. A lost + save reply is `SaveOutcome::ReplyLost`, and `Coordinator::resolve_durable` then asks the + store about the *same* checkpoint instead of saving again. +- **The writer is bounded twice**, by outstanding captures and by queued bytes, and it owns + its payload handles until the bytes are committed or the job fails. A queued *replaceable* + capture is superseded by a later one, releasing its holds; a durable one never is. +- **The install is a group.** A restore selects a complete compatible generation, imports its + payloads as fresh artifacts, stages every participant, validates the coordinator's own + ledgers, and only then activates. A failure anywhere leaves the fence closed, and every + participant that got as far as staging is recorded as one that must be replaced before + another restore is attempted. +- **The fence lifts once.** `Failed -> Restoring(k) -> Paused(k)`, at the end of a complete + install and nowhere else. A fenced session takes no step, publishes nothing, captures + nothing and holds no artifact handle. +- **Nothing old crosses.** The fence drops every media handle; the restore imports fresh + artifacts; the environment re-renders its pending sensor pipeline from recorded + reconstruction inputs; and the new epoch's first audio chunk resumes the preserved sample + position and marks the discontinuity. +- **Epoch metadata in a trace.** `scope.epoch`, the batch id and every task event id are + derived from the epoch, so a resumed run's behaviour is compared through + `EpochRebase`, which rewrites exactly those and fails on anything it does not recognise. + ## Where this crate narrows or adds to the contract crate - **Required views.** `WorldObservation::validate_against` checks the views a result carries @@ -196,9 +229,9 @@ harness.shutdown().await; - **Fake workers.** There is no neural model and no emulator. What is modelled exactly is the ordering, the identity rules and the retry rules, not any numerical behaviour. -- **No state methods.** `State.Capture`, `State.StageRestore` and `State.ActivateRestore` are - STATE-01. The phase machine has their edges (`Capturing`, `Restoring`) and the workers do not - advertise them as implemented methods. +- **One environment, one task.** A checkpoint records the composition it was taken from, and a + restore refuses one taken under another backend, content, patch, controller or parser + identity. It does not migrate between compositions, and it does not try. - **No audience input.** The admitted pre-step stimulation list exists and is always empty. - **Pacing is coarse.** The pacing deadline rounds one step to whole nanoseconds for sleeping only; simulation time stays rational and that rounding never re-enters the accumulator. @@ -259,6 +292,9 @@ The three integration suites do not all run over both transports, and cannot: - `tests/processes.rs` runs over the Unix socket only, in all three execution modes. A participant in a process of its own has no in-memory transport to reach the router by, so the mode is the axis that suite varies and the transport is fixed. +- `tests/media.rs` and `tests/state.rs` run over both transports *and* in all three execution + modes: each acceptance body is written once and registered twice, by `both_transports!` in + the in-process composition and by `all_modes!` over the socket. - `tests/session.rs`: one world advance per complete batch; every agent Prepared before the advance; one task evaluation per transition; every agent committed before the next Prepare or @@ -275,6 +311,14 @@ The three integration suites do not all run over both transports, and cannot: allocation -- plus the sequential/reversed/parallel trace comparison across all three modes and the two process-mode section 4 rows: a router restart during a world advance, and an old worker's reply after a restart. +- `tests/state.rs`: the STATE-01 acceptance bullets -- an uninterrupted run and a resumed run + committing the same behaviour once the epoch metadata is rebased, a corrupt payload failing + the install as a group for every participant and for the coordinator's own ledger, a lost + save reply and an uncommitted store manifest both leaving the durable mark where it was, a + refused activation resuming no part of the world, the capture queue staying bounded under a + stalled writer, and old media and another parser's state failing to cross a recovery -- + plus the once-only restore token, the superseded replaceable capture, and the fence that + lifts only through a complete restore. - `tests/failures.rs`: a duplicate Prepare after a lost reply; a duplicate Commit; the same batch with altered controls; a lost Advance result; a cached artifact consumed by its first caller; one Commit failing after another succeeded; a replaced registration; a reply from diff --git a/services/flysim/crates/fly-session/src/agent.rs b/services/flysim/crates/fly-session/src/agent.rs index a52822f..22b5ac5 100644 --- a/services/flysim/crates/fly-session/src/agent.rs +++ b/services/flysim/crates/fly-session/src/agent.rs @@ -7,7 +7,7 @@ //! stimulation, then reinforces once, and executes no tick at all. Every mutating step bumps //! one counter, which is how a test proves a duplicate request changed nothing. -use std::collections::BTreeMap; +use std::collections::{BTreeMap, BTreeSet}; use serde_json::Value; @@ -193,6 +193,12 @@ pub struct AgentFaults { pub prepare_delay_ms: u64, /// Hold `Agent.Commit` open for this long. pub commit_delay_ms: u64, + /// Refuse `State.StageRestore`, so a group install meets one participant that will not + /// validate while the others already have. + pub fail_stage_restore: bool, + /// Refuse `State.ActivateRestore` after this worker has already staged, so a group meets + /// a failure halfway through activation. + pub fail_activate_restore: bool, } /// One fake agent worker's configuration. @@ -234,6 +240,10 @@ pub struct FakeAgentWorker { context: Option, context_digest: Option, prepared: Option<(DomainRequestId, PreparedDecision)>, + /// A validated replacement state that the live session cannot see yet. + staged: Option, + /// Restore tokens this worker has activated. A token activates once. + activated: BTreeSet, } impl FakeAgentWorker { @@ -248,10 +258,17 @@ impl FakeAgentWorker { context: None, context_digest: None, prepared: None, + staged: None, + activated: BTreeSet::new(), config, } } + /// True while a validated replacement state is staged and not yet activated. + pub fn has_staged_restore(&self) -> bool { + self.staged.is_some() + } + pub fn status(&self) -> StatusCell { self.status.clone() } @@ -657,7 +674,11 @@ impl WorkerEndpoint for FakeAgentWorker { } fn capabilities(&self) -> Vec { - vec![id("agent-step-v1"), id("pixel-observation-v1")] + vec![ + id("agent-step-v1"), + id("pixel-observation-v1"), + id(crate::state::CHECKPOINT_CAPABILITY), + ] } fn status_cell(&self) -> StatusCell { @@ -669,7 +690,14 @@ impl WorkerEndpoint for FakeAgentWorker { } fn methods(&self) -> Vec<&'static str> { - vec!["Agent.Initialize", "Agent.Prepare", "Agent.Commit"] + vec![ + "Agent.Initialize", + "Agent.Prepare", + "Agent.Commit", + "State.Capture", + "State.StageRestore", + "State.ActivateRestore", + ] } fn handle<'a>(&'a mut self, ctx: HandlerCtx<'a>) -> BoxFuture<'a, DomainResult> { @@ -678,6 +706,9 @@ impl WorkerEndpoint for FakeAgentWorker { "Agent.Initialize" => self.initialize(&ctx).await, "Agent.Prepare" => self.prepare(&ctx).await, "Agent.Commit" => self.commit(&ctx).await, + "State.Capture" => self.state_capture(&ctx).await, + "State.StageRestore" => self.state_stage_restore(&ctx).await, + "State.ActivateRestore" => self.state_activate_restore(&ctx).await, other => Err(DomainError::before( ErrorCode::Unsupported, format!("{other} is not an agent method"), @@ -690,7 +721,12 @@ impl WorkerEndpoint for FakeAgentWorker { /// The retention class table an agent endpoint follows, for a caller that wants it. pub fn agent_op_class(method: &str) -> Option { match method { - "Agent.Initialize" => Some(OpClass::Lifecycle), + // `ipc-v1` section 5: lifecycle *and capture* replies are retained until + // `Worker.Acknowledge`, which is also what lets a duplicate restore request replay + // its cached reply rather than staging or activating twice. + "Agent.Initialize" | "State.Capture" | "State.StageRestore" | "State.ActivateRestore" => { + Some(OpClass::Lifecycle) + } "Agent.Prepare" | "Agent.Commit" => Some(OpClass::StepMutation), _ => None, } @@ -712,3 +748,532 @@ pub fn synthetic_profile(agent_id: &Id, tick_duration: &RationalNs, warmup_ticks /// The per-agent contexts a bootstrap produced, keyed by agent id. pub type Contexts = BTreeMap; + +// ------------------------------------------------------------------------------------------- +// STATE-01: capture and restore + +/// The numerical model version this worker implements. It is part of a capture's +/// compatibility identity: the same profile and seed under another model is not the same +/// state (`workers-v1` section 2). +pub const MODEL_VERSION: &str = "fake-lcg-v1"; + +/// The plasticity rule version, for the same reason. +pub const PLASTICITY_VERSION: &str = "fake-reinforce-v1"; + +/// The version this payload layout is written and read under. +pub const AGENT_PAYLOAD_VERSION: u64 = 1; + +/// The dataset identity a synthetic agent resolves. +/// +/// There is no connectome dataset behind this worker, and a checkpoint says so with a stable +/// identity rather than omitting the field: "no dataset" has to be distinguishable from "the +/// dataset was not recorded". +pub fn dataset_digest() -> Digest { + digest_of_bytes(b"fly-session/no-dataset-v1") +} + +/// The capture compatibility digest of one agent (`workers-v1` section 2). +/// +/// The profile digest identifies the profile definition; this additionally covers the +/// resolved seed, the numerical model version and the plasticity rule, because two agents +/// with the same profile digest and different seeds hold state that is not interchangeable. +/// Every field it covers is one the checkpoint manifest already records in that agent's row, +/// so a restore derives the expected digest from the manifest rather than from the payload it +/// is about to validate. +pub fn agent_compatibility_digest( + agent_id: &Id, + profile_digest: &Digest, + dataset_digest: &Digest, + model_version: &str, + plasticity_version: &str, + seed: i32, +) -> Digest { + let value = serde_json::json!({ + "agentId": agent_id.as_str(), + "profileDigest": profile_digest.as_str(), + "datasetDigest": dataset_digest.as_str(), + "modelVersion": model_version, + "plasticityVersion": plasticity_version, + "seed": seed, + }); + digest_of(&value).expect("an agent compatibility block canonicalizes") +} + +impl FakeModel { + /// Every field of the model, so a resumed agent is this agent and not a fresh one. + fn capture(&self) -> Value { + serde_json::json!({ + "seed": self.seed, + "state": self.state.to_string(), + "mutations": self.mutations.to_string(), + "ticks": self.ticks.to_string(), + "stimulations": self.stimulations.to_string(), + "reinforcements": self.reinforcements.to_string(), + "learningEnabled": self.learning_enabled, + "learningUpdates": self.learning_updates.to_string(), + "learningChanged": self.learning_changed.to_string(), + "lastSignal": self.last_signal, + "inputValue": self.input_value.to_string(), + "inputInstalls": self.input_installs.to_string(), + }) + } + + fn restored(value: &Value) -> DomainResult { + let number = |key: &str| -> DomainResult { + value + .get(key) + .and_then(Value::as_str) + .ok_or_else(|| incompatible(format!("the agent payload has no {key}")))? + .parse::() + .map_err(|_| incompatible(format!("the agent payload's {key} is not a U64"))) + }; + let seed = value + .get("seed") + .and_then(Value::as_i64) + .and_then(|v| i32::try_from(v).ok()) + .ok_or_else(|| incompatible("the agent payload has no seed"))?; + let input_value = value + .get("inputValue") + .and_then(Value::as_str) + .ok_or_else(|| incompatible("the agent payload has no inputValue"))? + .parse::() + .map_err(|_| incompatible("the agent payload's inputValue is not an integer"))?; + let last_signal = value + .get("lastSignal") + .and_then(Value::as_f64) + .filter(|v| v.is_finite()) + .ok_or_else(|| incompatible("the agent payload's lastSignal is not finite"))?; + let learning_enabled = value + .get("learningEnabled") + .and_then(Value::as_bool) + .ok_or_else(|| incompatible("the agent payload has no learningEnabled"))?; + Ok(FakeModel { + seed, + state: number("state")?, + mutations: number("mutations")?, + ticks: number("ticks")?, + stimulations: number("stimulations")?, + reinforcements: number("reinforcements")?, + learning_enabled, + learning_updates: number("learningUpdates")?, + learning_changed: number("learningChanged")?, + last_signal, + input_value, + input_installs: number("inputInstalls")?, + }) + } +} + +fn incompatible(message: impl std::fmt::Display) -> DomainError { + DomainError::before(ErrorCode::IncompatibleState, message) +} + +/// One staged restore, held outside the live agent until it is activated. +struct StagedAgent { + token: Id, + checkpoint_id: Id, + scope: Scope, + model: FakeModel, + accumulator: TickAccumulator, + context: TypedValue, + profile: AssetRef, + committed_step: u64, +} + +impl FakeAgentWorker { + /// This worker's own compatibility identity, from its configuration and a resolved seed. + fn compatibility_digest(&self, profile: &AssetRef, seed: i32) -> Digest { + agent_compatibility_digest( + &self.config.agent_id, + &profile.digest, + &dataset_digest(), + MODEL_VERSION, + PLASTICITY_VERSION, + seed, + ) + } + + /// `State.Capture`: an immutable snapshot of this agent at its committed boundary. + /// + /// It is allowed at `Ready(k)` only. A Prepared agent holds half a transition, and there + /// is no coherent boundary to file that under. + async fn state_capture(&mut self, ctx: &HandlerCtx<'_>) -> DomainResult { + let scope = ctx.scope()?.clone(); + self.check_epoch(&scope)?; + let AgentPhase::Ready(k) = self.phase.clone() else { + return Err(DomainError::before( + ErrorCode::InvalidPhase, + format!( + "State.Capture needs a quiescent Ready(k); this worker is {:?}", + self.phase + ), + )); + }; + if scope.step != k { + return Err(DomainError::before( + if scope.step < k { ErrorCode::StaleStep } else { ErrorCode::FutureStep }, + "State.Capture names a boundary this worker is not at", + )); + } + let params: CaptureParams = ctx.params()?; + let profile = self.profile.clone().expect("initialized"); + let context = self.context.clone().expect("initialized"); + let accumulator = self.accumulator.as_ref().expect("initialized"); + let previous = self.status.state(); + self.status.set_state(WorkerState::Capturing); + let payload = serde_json::json!({ + "payloadVersion": AGENT_PAYLOAD_VERSION, + "kind": "agent", + "agentId": self.config.agent_id.as_str(), + "checkpointId": params.checkpoint_id.as_str(), + "sourceScope": scope.to_json(), + "committedStep": k.to_string(), + "profile": profile.to_json(), + "modelVersion": MODEL_VERSION, + "plasticityVersion": PLASTICITY_VERSION, + "datasetDigest": dataset_digest().as_str(), + "model": self.model.capture(), + "accumulator": { + "tickDuration": accumulator.tick_duration().to_json(), + "remainder": accumulator.remainder().to_json(), + "executedTicks": accumulator.executed_ticks().to_string(), + "warmupOffset": accumulator.warmup_offset().to_string(), + }, + "context": context.to_json(), + }); + let bytes = canonicalize(&payload) + .map_err(|e| DomainError::invalid(format!("State.Capture: {}", e.0)))? + .into_bytes(); + let digest = digest_of_bytes(&bytes); + let artifact = crate::state::seal_payload(ctx.client, &bytes, &digest).await?; + // Capture is a read of the model, not a mutation of it: nothing above changed a + // counter, and the worker goes back to the boundary it was already at. + self.status.set_state(previous); + let result = CaptureResult { + checkpoint_id: params.checkpoint_id, + boundary: k, + compatibility_digest: self.compatibility_digest(&profile, self.model.seed()), + payload: artifact.reference().clone(), + }; + Ok(HandlerReply::with_artifacts( + object(result.to_json()), + vec![(crate::state::PAYLOAD_ATTACHMENT.to_owned(), artifact)], + )) + } + + /// `State.StageRestore`: validate a replacement state into a staging slot. + /// + /// Nothing the live session can see changes here, and the worker keeps whatever state it + /// had. It is allowed on an uninitialized replacement or a quiescent worker only; a + /// failed one is neither, which is why a group that failed is replaced rather than + /// reused. + async fn state_stage_restore(&mut self, ctx: &HandlerCtx<'_>) -> DomainResult { + let scope = ctx.scope()?.clone(); + if scope.session_id != self.config.session_id { + return Err(DomainError::before( + ErrorCode::IdentityMismatch, + "this worker belongs to another session", + )); + } + match &self.phase { + AgentPhase::Uninitialized | AgentPhase::Ready(_) => {} + other => { + return Err(DomainError::before( + ErrorCode::InvalidPhase, + format!( + "State.StageRestore needs an uninitialized replacement or a quiescent \ +worker; this worker is {other:?}" + ), + )); + } + } + if let Some(epoch) = &self.epoch + && *epoch == scope.epoch + { + return Err(DomainError::before( + ErrorCode::StaleEpoch, + "State.StageRestore proposes the epoch this worker is already running", + )); + } + let params: StageRestoreParams = ctx.params()?; + if params.source_scope.step != scope.step { + return Err(DomainError::invalid( + "State.StageRestore's scope step must be the source boundary", + )); + } + let artifact = ctx.artifact(crate::state::PAYLOAD_ATTACHMENT)?; + if artifact.reference() != ¶ms.payload { + return Err(DomainError::before( + ErrorCode::BufferInvalid, + "the staged payload attachment is not the artifact the request names", + )); + } + let bytes = artifact.read_all().await.map_err(|e| { + DomainError::before( + ErrorCode::BufferInvalid, + format!("the staged payload could not be read: {}", e.message), + ) + })?; + let declared = params + .payload + .digest + .clone() + .ok_or_else(|| incompatible("a checkpoint payload must carry a content digest"))?; + let actual = digest_of_bytes(&bytes); + if actual != declared || bytes.len() as u64 != params.payload.byte_length { + return Err(incompatible( + "the staged payload is not the content the request declares", + )); + } + let value: Value = serde_json::from_slice(&bytes) + .map_err(|e| DomainError::invalid(format!("the staged payload is not JSON: {e}")))?; + let text = |key: &str| -> DomainResult { + value + .get(key) + .and_then(Value::as_str) + .map(str::to_owned) + .ok_or_else(|| incompatible(format!("the agent payload has no {key}"))) + }; + if value.get("payloadVersion").and_then(Value::as_u64) != Some(AGENT_PAYLOAD_VERSION) { + return Err(incompatible("the agent payload is another payload version")); + } + if text("kind")? != "agent" { + return Err(incompatible("this payload is not an agent's state")); + } + if text("agentId")? != self.config.agent_id { + return Err(DomainError::before( + ErrorCode::IdentityMismatch, + "the staged payload belongs to another agent", + )); + } + if text("checkpointId")? != params.checkpoint_id { + return Err(incompatible("the staged payload belongs to another checkpoint")); + } + if text("modelVersion")? != MODEL_VERSION || text("plasticityVersion")? != PLASTICITY_VERSION + { + return Err(incompatible( + "the staged payload was captured under another numerical model", + )); + } + let source_scope = Scope::from_json( + value + .get("sourceScope") + .ok_or_else(|| incompatible("the agent payload has no sourceScope"))?, + ) + .map_err(|e| incompatible(format!("the agent payload's sourceScope: {}", e.0)))?; + if source_scope != params.source_scope { + return Err(incompatible( + "the staged payload was captured at another source scope", + )); + } + let committed_step: u64 = text("committedStep")? + .parse() + .map_err(|_| incompatible("the agent payload's committedStep is not a U64"))?; + if committed_step != params.source_scope.step { + return Err(incompatible( + "the staged payload's committed step is not the source boundary", + )); + } + let profile = AssetRef::from_json( + value + .get("profile") + .ok_or_else(|| incompatible("the agent payload has no profile"))?, + ) + .map_err(|e| incompatible(format!("the agent payload's profile: {}", e.0)))?; + let model = FakeModel::restored( + value + .get("model") + .ok_or_else(|| incompatible("the agent payload has no model"))?, + )?; + // The compatibility digest is recomputed from this worker's own configuration and the + // identity the payload declares. A capture of the same profile under another seed, or + // of another agent's brain, fails here and never reaches activation. + let computed = self.compatibility_digest(&profile, model.seed()); + if computed != params.compatibility_digest { + return Err(incompatible(format!( + "the staged state's compatibility {computed} is not the {} the restore \ +requires", + params.compatibility_digest + ))); + } + let accumulator_value = value + .get("accumulator") + .ok_or_else(|| incompatible("the agent payload has no accumulator"))?; + let rational = |key: &str| -> DomainResult { + RationalNs::from_json( + accumulator_value + .get(key) + .ok_or_else(|| incompatible(format!("the accumulator has no {key}")))?, + ) + .map_err(|e| incompatible(format!("the accumulator's {key}: {}", e.0))) + }; + let counter = |key: &str| -> DomainResult { + accumulator_value + .get(key) + .and_then(Value::as_str) + .ok_or_else(|| incompatible(format!("the accumulator has no {key}")))? + .parse::() + .map_err(|_| incompatible(format!("the accumulator's {key} is not a U64"))) + }; + let tick_duration = rational("tickDuration")?; + if tick_duration != self.config.tick_duration { + return Err(incompatible( + "the staged state was captured at another model tick duration", + )); + } + let accumulator = TickAccumulator::restored( + tick_duration, + rational("remainder")?, + counter("executedTicks")?, + counter("warmupOffset")?, + ) + .map_err(incompatible)?; + let context = TypedValue::from_json( + value + .get("context") + .ok_or_else(|| incompatible("the agent payload has no context"))?, + ) + .map_err(|e| incompatible(format!("the agent payload's context: {}", e.0)))?; + FakeAgentWorker::available_actions(&context)?; + + if self.config.faults.fail_stage_restore { + // The row where a group validates three participants and the fourth does not. + // Nothing is staged here and nothing is staged anywhere else either: the + // coordinator abandons the whole install. + return Err(incompatible( + "injected staging refusal: this participant's replacement state does not \ +validate", + )); + } + // One staged restore at a time. A second proposal replaces nothing silently. + if let Some(staged) = &self.staged { + return Err(DomainError::before( + ErrorCode::Conflict, + format!( + "this worker already holds the staged restore {} for checkpoint {}", + staged.token, staged.checkpoint_id + ), + )); + } + let token = restore_token(¶ms.checkpoint_id, &scope, &actual, &self.config.incarnation_id); + if self.activated.contains(&token) { + return Err(DomainError::before( + ErrorCode::Conflict, + "this exact restore was already activated on this worker", + )); + } + self.staged = Some(StagedAgent { + token: token.clone(), + checkpoint_id: params.checkpoint_id.clone(), + scope: scope.clone(), + model, + accumulator, + context, + profile, + committed_step, + }); + self.status.set_state(WorkerState::StagedRestore); + let result = StageRestoreResult { + checkpoint_id: params.checkpoint_id, + restore_token: token, + }; + Ok(HandlerReply::from(&result)) + } + + /// `State.ActivateRestore`: install the staged state under its new scope, without a tick. + /// + /// The token activates once. A duplicate domain request replays the cached reply through + /// the shell's result cache; a fresh request naming an already activated token is a + /// conflict, which is what stops a second group from being resumed from the same bytes. + async fn state_activate_restore( + &mut self, + ctx: &HandlerCtx<'_>, + ) -> DomainResult { + let params: ActivateRestoreParams = ctx.params()?; + if self.activated.contains(¶ms.restore_token) { + return Err(DomainError::before( + ErrorCode::Conflict, + "this restore token has already been activated", + )); + } + let Some(staged) = self.staged.take() else { + return Err(DomainError::before( + ErrorCode::InvalidPhase, + "this worker holds no staged restore", + )); + }; + if staged.token != params.restore_token { + // Put it back: naming another token is not a reason to discard this one. + let token = staged.token.clone(); + self.staged = Some(staged); + return Err(DomainError::before( + ErrorCode::IdentityMismatch, + format!("this worker's staged restore is {token}, not {}", params.restore_token), + )); + } + if self.config.faults.fail_activate_restore { + let token = staged.token.clone(); + self.staged = Some(staged); + self.status.set_state(WorkerState::Failed); + return Err(DomainError::new( + ErrorCode::BackendFailure, + format!("injected activation failure; {token} stays staged and unresumed"), + MutationCertainty::None, + )); + } + self.status.set_state(WorkerState::Restoring); + let StagedAgent { + token, + checkpoint_id, + scope, + model, + accumulator, + context, + profile, + committed_step, + } = staged; + self.model = model; + self.accumulator = Some(accumulator); + self.context_digest = Some(context.digest()); + self.context = Some(context); + self.profile = Some(profile); + self.epoch = Some(scope.epoch.clone()); + self.prepared = None; + self.phase = AgentPhase::Ready(committed_step); + self.activated.insert(token); + self.status.set_state(WorkerState::Ready); + self.status.set_scope(Some(scope_at( + &scope.session_id, + &scope.epoch, + committed_step, + ))); + self.status.advance_to(self.model.mutations()); + let result = ActivateRestoreResult { + committed_step, + checkpoint_id, + // An agent returns a null observation; the environment returns the world's. + observation: None, + }; + result + .validate_for_role(Role::Agent) + .map_err(|e| DomainError::invalid(e.0))?; + Ok(HandlerReply::from(&result)) + } +} + +/// A restore token bound to the checkpoint, the proposed scope, the payload bytes and the +/// worker incarnation staging them. +/// +/// `state-media-v1` section 5 binds a token to scope, payload and checkpoint. Binding it to +/// the incarnation as well is what keeps a token minted by a worker that has since been +/// replaced from activating anything on its replacement. +pub fn restore_token(checkpoint_id: &Id, scope: &Scope, payload_digest: &Digest, incarnation: &Id) -> Id { + let digest = digest_of_bytes( + format!( + "fly-session/restore-token-v1\n{checkpoint_id}\n{}\n{}\n{}\n{payload_digest}\n{incarnation}\n", + scope.session_id, scope.epoch, scope.step + ) + .as_bytes(), + ); + parse_id(&format!("rt-{}", &digest[..32])).expect("a hex suffix is an Id") +} diff --git a/services/flysim/crates/fly-session/src/cli.rs b/services/flysim/crates/fly-session/src/cli.rs index 1f7bc68..bbd6bbf 100644 --- a/services/flysim/crates/fly-session/src/cli.rs +++ b/services/flysim/crates/fly-session/src/cli.rs @@ -46,8 +46,10 @@ Worker options (agent and environment): agent: --agent ID --port ID --tick-numerator N --tick-denominator N --warmup-ticks N [--prepare-delay-ms N] [--commit-delay-ms N] [--fail-commit-at-step N] + [--fail-stage-restore 0|1] [--fail-activate-restore 0|1] environment: --worker ID --ports p1,p2 --step-numerator N --step-denominator N [--advance-delay-ms N] [--omit-view-at-boundary N] + [--fail-stage-restore 0|1] [--fail-activate-restore 0|1] Measure options: --steps N transitions per run (default 200) @@ -145,6 +147,17 @@ impl Options { } } + /// A flag whose value is `0` or `1`. Anything else is an error naming it, so a + /// mistyped injection is a failed launch rather than a fault that never fires. + fn flag(&self, name: &str) -> Result { + match self.0.get(name) { + None => Ok(false), + Some(value) if value == "0" => Ok(false), + Some(value) if value == "1" => Ok(true), + Some(value) => Err(format!("--{name}: {value:?} is not 0 or 1")), + } + } + fn opt_u64(&self, name: &str) -> Result, String> { match self.0.get(name) { None => Ok(None), @@ -197,6 +210,8 @@ fn serve(role: &str, options: &Options) -> Result<(), String> { fail_commit_at_step: options.opt_u64(flags::FAIL_COMMIT_AT_STEP)?, prepare_delay_ms: options.u64(flags::PREPARE_DELAY_MS, 0)?, commit_delay_ms: options.u64(flags::COMMIT_DELAY_MS, 0)?, + fail_stage_restore: options.flag(flags::FAIL_STAGE_RESTORE)?, + fail_activate_restore: options.flag(flags::FAIL_ACTIVATE_RESTORE)?, }, client_id: client_id.clone(), service: service.clone(), @@ -219,6 +234,8 @@ fn serve(role: &str, options: &Options) -> Result<(), String> { omit_audio_at_boundary: options.opt_u64(flags::OMIT_AUDIO_AT_BOUNDARY)?, overlapping_audio_at_boundary: options .opt_u64(flags::OVERLAPPING_AUDIO_AT_BOUNDARY)?, + fail_stage_restore: options.flag(flags::FAIL_STAGE_RESTORE)?, + fail_activate_restore: options.flag(flags::FAIL_ACTIVATE_RESTORE)?, }, client_id: client_id.clone(), service: service.clone(), diff --git a/services/flysim/crates/fly-session/src/clock.rs b/services/flysim/crates/fly-session/src/clock.rs index 080c12a..aa50ff2 100644 --- a/services/flysim/crates/fly-session/src/clock.rs +++ b/services/flysim/crates/fly-session/src/clock.rs @@ -29,6 +29,30 @@ impl TickAccumulator { }) } + /// The exact accumulator a capture recorded. + /// + /// The remainder is restored, never rounded or reset: a resumed agent that started its + /// first interval from zero would drift away from the run it is supposed to continue. + pub fn restored( + tick_duration: RationalNs, + remainder: RationalNs, + executed_ticks: u64, + warmup_offset: u64, + ) -> Result { + let mut accumulator = TickAccumulator::new(tick_duration)?; + remainder.validate().map_err(|e| e.0)?; + if remainder >= tick_duration { + return Err("a captured remainder is not below one model tick".to_owned()); + } + if warmup_offset > executed_ticks { + return Err("a captured warm-up offset exceeds the executed tick count".to_owned()); + } + accumulator.remainder = remainder; + accumulator.executed_ticks = executed_ticks; + accumulator.warmup_offset = warmup_offset; + Ok(accumulator) + } + pub fn tick_duration(&self) -> RationalNs { self.tick_duration } diff --git a/services/flysim/crates/fly-session/src/coordinator.rs b/services/flysim/crates/fly-session/src/coordinator.rs index e7e4bb7..bde1183 100644 --- a/services/flysim/crates/fly-session/src/coordinator.rs +++ b/services/flysim/crates/fly-session/src/coordinator.rs @@ -141,6 +141,10 @@ pub struct Deadlines { pub resolve_attempts: u32, /// A separate, larger budget for `Worker.Hello` and the `Initialize` methods. pub boot: Duration, + /// A separate budget for the `State.*` methods, which `ipc-v1` section 6 gives one: + /// a capture serializes a participant and a restore validates and installs one, and + /// neither is a step whose latency the probe was chosen for. + pub capture: Duration, } /// How long the resolution waits between attempts. @@ -167,6 +171,7 @@ impl Default for Deadlines { // Over sixteen seconds of pauses: the budget above is what terminates. resolve_attempts: 8192, boot: Duration::from_secs(30), + capture: Duration::from_secs(30), } } } @@ -177,6 +182,8 @@ pub struct Topics { pub descriptor: String, pub snapshots: String, pub events: String, + /// Where the distinct captured/queued/committed/failed/superseded checkpoint events go. + pub checkpoints: String, } impl Topics { @@ -185,6 +192,7 @@ impl Topics { descriptor: format!("session.{session_id}.descriptor"), snapshots: format!("session.{session_id}.snapshots"), events: format!("session.{session_id}.events"), + checkpoints: format!("session.{session_id}.checkpoints"), } } } @@ -202,6 +210,11 @@ pub struct AgentSlot { pub tick_duration: RationalNs, pub warmup_ticks: u64, pub committed_step: u64, + /// The tick count and remainder this agent last reported, which are what the checkpoint + /// manifest records for it. They are metadata about the payload, never a substitute for + /// it: the agent's own capture is the state that is restored. + pub brain_ticks: u64, + pub remainder: RationalNs, context: TypedValue, context_digest: Digest, prepared: Option, @@ -226,6 +239,8 @@ impl AgentSlot { tick_duration: RationalNs::ZERO, warmup_ticks: 0, committed_step: 0, + brain_ticks: 0, + remainder: RationalNs::ZERO, context: TypedValue::new(crate::task::context_schema(), Value::Object(Map::new())) .expect("an empty context object is a valid typed value"), context_digest: digest_of_bytes(b""), @@ -300,6 +315,12 @@ pub struct Coordinator { /// Set when the epoch failed: every old handle, route and reply is invalid from here on /// and only a coherent restore may lift it. fenced: bool, + /// The durable store's bounded writer, when this composition has one. + writer: Option, + /// The last checkpoint whose saved acknowledgment arrived, and its boundary. + durable: Option<(Id, u64)>, + /// Participants that installed a restore a group install then abandoned. + tainted: BTreeSet, started: std::time::Instant, last_advance_request: Option, last_commit_requests: Vec, @@ -361,6 +382,9 @@ impl Coordinator { metrics: Metrics::default(), blame: None, fenced: false, + writer: None, + durable: None, + tainted: BTreeSet::new(), started: std::time::Instant::now(), last_advance_request: None, last_commit_requests: Vec::new(), @@ -489,8 +513,16 @@ impl Coordinator { let (from, to) = self.phases.fail(); self.trace.phase(from, to); self.fenced = true; + // Whatever replies this session still owed an acknowledgment for belong to + // participants of an invalid epoch. Carrying them across a recovery would send + // `Worker.Acknowledge` to a registration that is gone. + self.lifecycle_acks.clear(); + // Every handle this session held on the old epoch's media goes with the fence: a + // recovery imports fresh artifacts and never expects one of these back. self.views.clear(); self.pending_views.clear(); + self.audio.clear(); + self.pending_audio.clear(); match &participant { Some(who) => self.audit.push(format!("fail:{detail}:{who}")), None => self.audit.push(format!("fail:{detail}")), @@ -631,6 +663,10 @@ impl Coordinator { (self.topics.descriptor.clone(), flybus::Retained::Latest), (self.topics.snapshots.clone(), flybus::Retained::Latest), (self.topics.events.clone(), flybus::Retained::None), + // Checkpoint events are a stream of distinct facts, not a latest value: a + // "committed" that replaced a "queued" would erase the distinction the durable + // commit rules are built on. + (self.topics.checkpoints.clone(), flybus::Retained::None), ] { self.bus.declare_topic(&name, retained).await.map_err(|e| { let error = DomainError::new( @@ -716,6 +752,15 @@ impl Coordinator { if let Err(e) = media::check_required_views(&result.descriptor, &result.observation) { return Err(self.fail_now(e, "observation-0")); } + // O[0] ran no transition either, so it carries no chunk, and that is checked rather + // than assumed from its boundary number. + if let Err(e) = media::check_required_audio( + &result.descriptor, + &result.observation, + media::ObservationOrigin::Installed, + ) { + return Err(self.fail_now(e, "observation-0")); + } self.descriptor = Some(result.descriptor); self.observation = Some(result.observation); self.lifecycle_acks.push((worker, reply.request_id.clone())); @@ -815,6 +860,10 @@ impl Coordinator { self.agents[index].tick_duration = result.tick_duration; self.agents[index].warmup_ticks = result.warmup_ticks; self.agents[index].committed_step = 0; + // At boundary 0 the agent has executed exactly its warm-up, with no interval + // consumed, so the remainder is zero. + self.agents[index].brain_ticks = result.telemetry.brain_ticks; + self.agents[index].remainder = RationalNs::ZERO; self.lifecycle_acks.push((slot_worker, reply.request_id.clone())); let agent_id = self.agents[index].agent_id.clone(); self.audit.push(format!("agent.initialize:{agent_id}")); @@ -1013,6 +1062,8 @@ impl Coordinator { self.blame(Some(worker.worker_id.clone())); let deadline = if method.ends_with("Initialize") || method == "Worker.Hello" { self.deadlines.boot + } else if method.starts_with("State.") { + self.deadlines.capture } else { self.deadlines.probe }; @@ -1692,6 +1743,12 @@ impl Coordinator { method, )); } + if let Some(slot) = self.agents.iter_mut().find(|slot| slot.agent_id == agent_id) { + // What the agent will be at once this transition commits: Prepare is the only + // phase that advances the accumulator. + slot.brain_ticks = decision.brain_ticks; + slot.remainder = decision.remainder; + } self.audit.push(format!("prepared:{agent_id}@{k}")); self.stats.prepares += 1; self.blame(None); @@ -2110,7 +2167,11 @@ impl Coordinator { } // Audio has no sensory role here, but its chunks still cannot overlap or go backwards // inside an epoch, and a stale one must not reach presentation as current. - if let Err(e) = media::check_required_audio(descriptor, &result.observation) { + if let Err(e) = media::check_required_audio( + descriptor, + &result.observation, + media::ObservationOrigin::Transition, + ) { return Err(self.fail_now(e, "step-result")); } if let Err(e) = self.timelines.accept(descriptor, &result.observation) { @@ -2605,3 +2666,1429 @@ impl Coordinator { Ok(()) } } + +// ------------------------------------------------------------------------------------------- +// STATE-01: coherent all-participant capture and recovery + +/// A capture that exists and has been queued, whose durable outcome has not arrived yet. +/// +/// `State.Capture` completes when an immutable capture exists, not when a backend save was +/// requested, and only durable completion produces a saved acknowledgment. Those are two +/// events, so they are two calls. +#[derive(Debug)] +pub struct CaptureTicket { + pub checkpoint_id: Id, + pub boundary: u64, + receiver: tokio::sync::oneshot::Receiver, +} + +/// What one completed group restore installed. +#[derive(Clone, Debug, PartialEq, Eq)] +pub struct RestoreReport { + pub checkpoint_id: Id, + pub boundary: u64, + pub epoch: Id, + /// Every participant that staged, in the order they were asked. + pub staged: Vec, + /// Every participant that activated, in the order they were asked. + pub activated: Vec, + /// The once-only token each participant staged under, so a caller can prove a token + /// activates once rather than being told so. + pub tokens: Vec<(Id, Id)>, + /// The artifacts the payloads were imported as. None of them existed before this restore. + pub imported: Vec, +} + +/// The payload name the coordinator's own session record travels under. +impl Coordinator { + /// Attaches the durable store's bounded writer. Without one this session captures nothing + /// and says so, rather than pretending to. + pub fn attach_store(&mut self, writer: crate::state::CheckpointWriter) { + self.writer = Some(writer); + } + + pub fn writer(&self) -> Option<&crate::state::CheckpointWriter> { + self.writer.as_ref() + } + + /// Stops the checkpoint writer and waits for its task. + pub async fn shutdown_store(&mut self) { + if let Some(writer) = self.writer.take() { + writer.shutdown().await; + } + } + + /// The last checkpoint whose *saved acknowledgment* this coordinator received, and its + /// boundary. It moves on durable completion and on nothing else. + pub fn durable(&self) -> Option<(Id, u64)> { + self.durable.clone() + } + + /// Participants that installed a restore in a group install that then failed. + /// + /// They hold state no group ever resumed. A further restore is refused until each one has + /// been replaced, which is how "a failure during activation cannot resume half a world" + /// survives the next attempt as well as this one. + pub fn tainted(&self) -> Vec { + self.tainted.iter().cloned().collect() + } + + /// The live composition's compatibility identities. + pub fn compatibility(&self) -> Outcome { + match &self.descriptor { + Some(descriptor) => Ok(crate::state::Compatibility::of(descriptor)), + None => Err(SessionFailure { + error: DomainError::before( + ErrorCode::InvalidPhase, + "the session has no environment descriptor to take compatibility from", + ), + phase: self.phases.phase().label(), + detail: "compatibility".to_owned(), + participant: None, + }), + } + } + + fn agent_compatibility(&self, slot: &AgentSlot) -> Digest { + crate::agent::agent_compatibility_digest( + &slot.agent_id, + &slot.profile.digest, + &crate::agent::dataset_digest(), + crate::agent::MODEL_VERSION, + crate::agent::PLASTICITY_VERSION, + slot.seed, + ) + } + + /// The coordinator's own session record: what it must hold again to resume this boundary. + /// + /// `state-media-v1` section 4 puts the next decision state, the admission state and the + /// event watermarks in the coordinator's own payloads. This is that payload: the + /// composition has no audience input, so the admission record says so explicitly rather + /// than being absent. + fn coordinator_record(&self, boundary: u64) -> Value { + let (last_source_step, issued) = self.task.event_watermarks(); + json!({ + "payloadVersion": 1, + "kind": "coordinator", + "committedStep": boundary.to_string(), + "admission": { + "audienceInput": "none-configured", + "admitted": Value::Array(Vec::new()), + }, + "eventWatermarks": { + "lastSourceStep": last_source_step.to_string(), + "issued": issued.to_string(), + }, + "audioPositions": Value::Object( + self.timelines + .positions() + .into_iter() + .map(|(stream, sample)| (stream, Value::String(sample.to_string()))) + .collect(), + ), + "agents": Value::Array( + self.agents + .iter() + .map(|slot| json!({ + "agentId": slot.agent_id.as_str(), + "portId": slot.port_id.as_str(), + "profile": slot.profile.to_json(), + "seed": slot.seed, + "workerThreads": slot.worker_threads.to_string(), + "committedStep": slot.committed_step.to_string(), + "brainTicks": slot.brain_ticks.to_string(), + "remainder": slot.remainder.to_json(), + "context": slot.context.to_json(), + })) + .collect(), + ), + }) + } + + /// Seals one coordinator-owned payload as an immutable artifact. + async fn seal_own(&mut self, name: &str, value: &Value) -> Outcome { + let bytes = match canonicalize(value) { + Ok(text) => text.into_bytes(), + Err(e) => { + return Err(self.fail_now( + DomainError::invalid(format!("checkpoint payload {name}: {}", e.0)), + "capture", + )); + } + }; + let digest = digest_of_bytes(&bytes); + let artifact = match crate::state::seal_payload(&self.bus, &bytes, &digest).await { + Ok(artifact) => artifact, + Err(e) => return Err(self.fail_now(e, "capture")), + }; + Ok(crate::state::CapturedPayload { + name: name.to_owned(), + byte_length: bytes.len() as u64, + digest, + artifact, + }) + } + + /// Takes one coherent all-participant capture at the committed boundary and queues it. + /// + /// The queue slot is taken *first*: a saturated writer refuses before a single + /// `State.Capture` is sent, which is the only way "reject or defer before capture" can be + /// true rather than aspirational. + pub async fn capture(&mut self, checkpoint_id: &Id, replaceable: bool) -> Outcome { + if self.fenced { + return Err(SessionFailure { + error: DomainError::before( + ErrorCode::InvalidPhase, + "the epoch is fenced; a fenced session captures nothing", + ), + phase: self.phases.phase().label(), + detail: "fenced".to_owned(), + participant: None, + }); + } + // A capture that arrives while a transition is in flight is a race, not a bug in the + // transaction: the supervisor asking for one does not know where the session is. It + // is refused by name and the epoch is untouched, unlike a phase edge the machine + // itself takes. + let Some(boundary) = self.phases.phase().committed_boundary() else { + self.audit.push(format!("capture-refused:{checkpoint_id}")); + return Err(SessionFailure { + error: DomainError::before( + ErrorCode::InvalidPhase, + "a coherent checkpoint is taken at a committed boundary only", + ), + phase: self.phases.phase().label(), + detail: "capture".to_owned(), + participant: None, + }); + }; + let origin = self.phases.phase(); + if self.writer.is_none() { + return Err(SessionFailure { + error: DomainError::before( + ErrorCode::Unsupported, + "this session has no checkpoint store attached", + ), + phase: self.phases.phase().label(), + detail: "capture".to_owned(), + participant: None, + }); + } + // Before anything is captured. + let reservation = match self.writer.as_ref().expect("checked").reserve() { + Ok(reservation) => reservation, + Err(e) => { + // A refused capture is not a session failure: the boundary stands, the + // session keeps stepping and the caller is told the queue is full. + self.audit.push(format!("capture-refused:{checkpoint_id}")); + return Err(SessionFailure { + error: e, + phase: self.phases.phase().label(), + detail: "capture".to_owned(), + participant: None, + }); + } + }; + self.transition(Phase::Capturing(boundary))?; + let result = self + .capture_group(checkpoint_id, boundary, replaceable, reservation) + .await; + match result { + Ok(ticket) => { + self.transition(origin)?; + self.audit.push(format!("captured:{checkpoint_id}@{boundary}")); + Ok(ticket) + } + Err(e) => Err(e), + } + } + + async fn capture_group( + &mut self, + checkpoint_id: &Id, + boundary: u64, + replaceable: bool, + reservation: crate::state::Reservation, + ) -> Outcome { + let scope = self.scope(boundary); + let compatibility = self.compatibility()?; + let params = object(CaptureParams { checkpoint_id: checkpoint_id.clone() }.to_json()); + let want = vec![crate::state::PAYLOAD_ATTACHMENT.to_owned()]; + let mut payloads = Vec::new(); + let mut acknowledge: Vec<(WorkerRef, DomainRequestId)> = Vec::new(); + + // The world first, then the agents in sorted order: one boundary, every participant. + let environment = self.environment.clone(); + let reply = self + .call( + &environment, + "State.Capture", + Some(scope.clone()), + params.clone(), + &[], + &want, + ) + .await?; + let world: CaptureResult = reply.parse().map_err(|e| self.fail_now(e, "capture"))?; + let world_payload = self.accept_capture( + &reply, + &world, + checkpoint_id, + boundary, + &compatibility.digest(), + crate::state::WORLD_PAYLOAD, + )?; + payloads.push(world_payload); + acknowledge.push((environment, reply.request_id.clone())); + + let mut agent_rows = Vec::new(); + for index in 0..self.agents.len() { + let slot_worker = self.agents[index].worker.clone(); + let agent_id = self.agents[index].agent_id.clone(); + let expected = self.agent_compatibility(&self.agents[index]); + let reply = self + .call( + &slot_worker, + "State.Capture", + Some(scope.clone()), + params.clone(), + &[], + &want, + ) + .await?; + let captured: CaptureResult = reply.parse().map_err(|e| self.fail_now(e, "capture"))?; + let name = crate::state::agent_payload(&agent_id); + let payload = + self.accept_capture(&reply, &captured, checkpoint_id, boundary, &expected, &name)?; + payloads.push(payload); + acknowledge.push((slot_worker, reply.request_id.clone())); + let slot = &self.agents[index]; + agent_rows.push(crate::state::AgentEntry { + agent_id: agent_id.clone(), + profile_digest: slot.profile.digest.clone(), + dataset_digest: crate::agent::dataset_digest(), + model_version: crate::agent::MODEL_VERSION.to_owned(), + plasticity_version: crate::agent::PLASTICITY_VERSION.to_owned(), + seed: slot.seed, + brain_ticks: slot.brain_ticks, + remainder: slot.remainder, + payload: name, + }); + } + + // The coordinator's own ledgers, sealed the same way so the writer treats every + // payload alike. + let ledger = self.task.capture().map_err(|e| self.fail_now(e, "capture"))?; + payloads.push( + self.seal_own(crate::state::TASK_LEDGER_PAYLOAD, &ledger.to_json()) + .await?, + ); + let inspection = self + .observation + .as_ref() + .map(|observation| observation.inspection.to_json()) + .ok_or_else(|| { + DomainError::before(ErrorCode::InvalidPhase, "the session never bootstrapped") + }); + let inspection = match inspection { + Ok(value) => value, + Err(e) => return Err(self.fail_now(e, "capture")), + }; + payloads.push( + self.seal_own(crate::state::PRIOR_INSPECTION_PAYLOAD, &inspection) + .await?, + ); + let mut executor_rows = Vec::new(); + for agent_id in self.agent_ids() { + let state = { + let executor = self.executors.get(&agent_id).ok_or_else(|| { + DomainError::before( + ErrorCode::IdentityMismatch, + format!("{agent_id} has no configured action executor"), + ) + }); + match executor { + Ok(executor) => executor.capture().map_err(|e| (e, agent_id.clone())), + Err(e) => Err((e, agent_id.clone())), + } + }; + let state = match state { + Ok(state) => state, + Err((e, who)) => { + self.blame(Some(who)); + return Err(self.fail_now(e, "capture")); + } + }; + let name = crate::state::executor_payload(&agent_id); + payloads.push(self.seal_own(&name, &state.to_json()).await?); + executor_rows.push((agent_id, name)); + } + let record = self.coordinator_record(boundary); + payloads.push( + self.seal_own(crate::state::ADMISSION_PAYLOAD, &record) + .await?, + ); + + let (last_source_step, issued) = self.task.event_watermarks(); + let world_time = match self.observation.as_ref().map(|o| o.world_time) { + Some(world_time) => Ok(world_time), + None => Err(self.fail_now( + DomainError::before( + ErrorCode::InvalidPhase, + "the session has no observation, so it is at no world time to record", + ), + "capture", + )), + }; + let manifest = crate::state::CheckpointManifest { + checkpoint_id: checkpoint_id.clone(), + source_scope: scope.clone(), + episode_id: self.episode_id.clone(), + world_time: world_time?, + scheduler_id: "lockstep-v1".to_owned(), + composition_digest: self.composition_digest(), + port_map: self + .agents + .iter() + .map(|slot| (slot.port_id.clone(), slot.agent_id.clone())) + .collect(), + compatibility: compatibility.clone(), + agents: agent_rows, + coordinator: crate::state::CoordinatorEntry { + task_ledger: crate::state::TASK_LEDGER_PAYLOAD.to_owned(), + prior_inspection: crate::state::PRIOR_INSPECTION_PAYLOAD.to_owned(), + executor_state: executor_rows, + admission_state: crate::state::ADMISSION_PAYLOAD.to_owned(), + event_watermarks: crate::state::EventWatermarks { + last_source_step, + issued, + }, + }, + environment: crate::state::EnvironmentEntry { + worker_id: self.environment.worker_id.clone(), + payload: crate::state::WORLD_PAYLOAD.to_owned(), + }, + // No external helper takes part in this composition, and the manifest says so + // rather than leaving the field out. + helper_state: Vec::new(), + payloads: payloads + .iter() + .map(|p| (p.name.clone(), p.byte_length, p.digest.clone())) + .collect(), + }; + + let submission = crate::state::CaptureSubmission { + checkpoint_id: checkpoint_id.clone(), + boundary, + session_id: self.session_id.clone(), + epoch: self.epoch.clone(), + episode_id: self.episode_id.clone(), + compatibility_digest: compatibility.digest(), + manifest: manifest.to_json(), + payloads, + replaceable, + }; + // The writer takes its own ownership of every payload here. Only then are the + // workers' cached capture replies released. + let receiver = { + let writer = self.writer.as_ref().expect("checked"); + match writer.submit(reservation, submission).await { + Ok(receiver) => receiver, + Err(e) => return Err(self.fail_now(e, "capture")), + } + }; + self.publish_checkpoint_event("captured", checkpoint_id, boundary, None).await?; + for (worker, request_id) in acknowledge { + self.lifecycle_acks.push((worker, request_id)); + } + self.acknowledge_lifecycle().await?; + Ok(CaptureTicket { + checkpoint_id: checkpoint_id.clone(), + boundary, + receiver, + }) + } + + /// Checks one participant's capture and turns it into a payload the writer can own. + fn accept_capture( + &mut self, + reply: &DomainReply, + result: &CaptureResult, + checkpoint_id: &Id, + boundary: u64, + expected_compatibility: &Digest, + name: &str, + ) -> Outcome { + if result.checkpoint_id != *checkpoint_id { + return Err(self.fail_now( + DomainError::before( + ErrorCode::IdentityMismatch, + "a capture names another checkpoint", + ), + "capture", + )); + } + if result.boundary != boundary { + return Err(self.fail_now( + DomainError::before( + ErrorCode::IdentityMismatch, + "a capture names another boundary; one checkpoint is one boundary", + ), + "capture", + )); + } + if result.compatibility_digest != *expected_compatibility { + return Err(self.fail_now( + DomainError::before( + ErrorCode::IncompatibleState, + "a capture reports a compatibility identity the composition does not hold", + ), + "capture", + )); + } + let digest = match &result.payload.digest { + Some(digest) => digest.clone(), + None => { + return Err(self.fail_now( + DomainError::before( + ErrorCode::BufferInvalid, + "a checkpoint payload must carry a content digest", + ), + "capture", + )); + } + }; + let artifact = match reply.artifacts.get(crate::state::PAYLOAD_ATTACHMENT) { + Some(artifact) if artifact.reference() == &result.payload => artifact.clone(), + _ => { + return Err(self.fail_now( + DomainError::before( + ErrorCode::BufferInvalid, + "a capture arrived without a live owned handle on its payload", + ), + "capture", + )); + } + }; + Ok(crate::state::CapturedPayload { + name: name.to_owned(), + byte_length: result.payload.byte_length, + digest, + artifact, + }) + } + + /// Waits for one capture's durable outcome and moves the durable mark only on a commit. + /// + /// A lost reply is an outcome here, not a hang and not a save: the high-water mark stays + /// where it was until [`Coordinator::resolve_durable`] asks the store about the same + /// operation. + pub async fn await_durable(&mut self, ticket: CaptureTicket) -> Outcome { + let CaptureTicket { checkpoint_id, boundary, receiver } = ticket; + let budget = self.deadlines.capture; + let outcome = + crate::state::CheckpointWriter::wait(receiver, &checkpoint_id, budget).await; + if let crate::state::SaveOutcome::Committed { .. } = &outcome { + self.durable = Some((checkpoint_id.clone(), boundary)); + self.audit.push(format!("durable:{checkpoint_id}@{boundary}")); + } else { + self.audit.push(format!("not-durable:{checkpoint_id}@{boundary}")); + } + Ok(outcome) + } + + /// Resolves a save whose reply was lost, by asking the store's durable metadata about the + /// *same* checkpoint. It never saves again. + /// + /// `Some(boundary)` means the store manifest lists that generation, which is the durable + /// commit point; `None` means it does not, and an unreferenced generation file stays + /// unreferenced. + pub async fn resolve_durable(&mut self, checkpoint_id: &Id) -> Outcome> { + let Some(writer) = self.writer.as_ref() else { + return Err(self.fail_now( + DomainError::before( + ErrorCode::Unsupported, + "this session has no checkpoint store attached", + ), + "resolve-durable", + )); + }; + let wanted = checkpoint_id.clone(); + let found = writer + .with_store(move |store| store.lookup(&wanted).map(|g| g.boundary)) + .await; + match found { + Some(boundary) => { + self.durable = Some((checkpoint_id.clone(), boundary)); + self.audit.push(format!("durable:{checkpoint_id}@{boundary}")); + Ok(Some(boundary)) + } + None => { + self.audit.push(format!("not-durable:{checkpoint_id}")); + Ok(None) + } + } + } + + /// Captures and waits for the durable outcome, which is what an ordinary caller wants. + pub async fn checkpoint(&mut self, checkpoint_id: &Id) -> Outcome { + let ticket = self.capture(checkpoint_id, false).await?; + self.await_durable(ticket).await + } + + async fn publish_checkpoint_event( + &mut self, + event: &str, + checkpoint_id: &Id, + boundary: u64, + detail: Option<&str>, + ) -> Outcome<()> { + let payload = json!({ + "event": event, + "sessionId": self.session_id.as_str(), + "epoch": self.epoch.as_str(), + "checkpointId": checkpoint_id.as_str(), + "boundary": boundary.to_string(), + "detail": detail.map_or(Value::Null, |d| Value::String(d.to_owned())), + }); + let topic = self.topics.checkpoints.clone(); + self.publish(&topic, object(payload), Vec::new()).await + } + + // --------------------------------------------------------------------------------------- + // Recovery + + /// Points the coordinator at a replacement participant while the epoch is fenced. + /// + /// A replacement is the only way a fenced participant comes back: `step-v1` section 7's + /// incarnation row says every live participant of a failed epoch belongs to an invalid + /// one, so the reference is exchanged deliberately here and never repaired in place. + pub fn replace_participant(&mut self, worker_id: &Id, worker: WorkerRef) -> Outcome<()> { + if !self.fenced { + return Err(self.fail_now( + DomainError::before( + ErrorCode::InvalidPhase, + "participants are replaced while the epoch is fenced, not during play", + ), + "replace", + )); + } + if worker.worker_id != *worker_id { + return Err(self.fail_now( + DomainError::before( + ErrorCode::IdentityMismatch, + "the replacement reference names another worker", + ), + "replace", + )); + } + if self.environment.worker_id == *worker_id { + self.environment = worker; + self.tainted.remove(worker_id); + self.audit.push(format!("replaced:{worker_id}")); + return Ok(()); + } + match self.agents.iter_mut().find(|slot| slot.agent_id == *worker_id) { + Some(slot) => { + slot.worker = worker; + self.tainted.remove(worker_id); + self.audit.push(format!("replaced:{worker_id}")); + Ok(()) + } + None => Err(self.fail_now( + DomainError::before( + ErrorCode::IdentityMismatch, + format!("{worker_id} is not a participant of this composition"), + ), + "replace", + )), + } + } + + /// Installs a coherent checkpoint into a fresh epoch and lifts the fence. + /// + /// The whole of `state-media-v1` section 6 in order: select a complete compatible durable + /// checkpoint, import its payloads as *new* artifacts, stage every participant, activate + /// every participant, install the coordinator's own staged state, verify identity and + /// boundary, flush the old media and parser state, and establish `Paused(k)`. A failure + /// at any point leaves the fence exactly where it was. + pub async fn restore( + &mut self, + checkpoint_id: Option<&Id>, + new_epoch: &Id, + ) -> Outcome { + if self.phases.phase() != Phase::Failed { + return Err(self.fail_now( + DomainError::before( + ErrorCode::InvalidPhase, + "a coherent restore starts from a failed epoch", + ), + "restore", + )); + } + if *new_epoch == self.epoch { + return Err(self.fail_now( + DomainError::before( + ErrorCode::StaleEpoch, + "a restore installs a fresh epoch, never the one that failed", + ), + "restore", + )); + } + if !self.tainted.is_empty() { + let who: Vec<&str> = self.tainted.iter().map(String::as_str).collect(); + return Err(self.fail_now( + DomainError::before( + ErrorCode::InvalidPhase, + format!( + "{} installed a restore that no group resumed and must be replaced first", + who.join(", ") + ), + ), + "restore", + )); + } + let Some(writer) = self.writer.as_ref() else { + return Err(self.fail_now( + DomainError::before( + ErrorCode::Unsupported, + "this session has no checkpoint store attached", + ), + "restore", + )); + }; + let wanted = checkpoint_id.cloned(); + let read = writer + .with_store(move |store| { + let record = store.select(wanted.as_ref())?; + let envelope = store.read(&record)?; + Ok::<_, DomainError>((record, envelope)) + }) + .await; + let (record, envelope) = match read { + Ok(pair) => pair, + Err(e) => return Err(self.fail_now(e, "restore")), + }; + let manifest = match crate::state::CheckpointManifest::from_json(&envelope.manifest) { + Ok(manifest) => manifest, + Err(e) => { + return Err(self.fail_now( + DomainError::before(ErrorCode::IncompatibleState, e), + "restore", + )); + } + }; + if let Err(e) = self.check_restore_identity(&manifest, &record) { + return Err(self.fail_now(e, "restore")); + } + let boundary = manifest.source_scope.step; + self.transition(Phase::Restoring(boundary))?; + let outcome = self.restore_group(&envelope, &manifest, new_epoch, boundary).await; + match outcome { + Ok(report) => Ok(report), + Err(e) => Err(e), + } + } + + /// Marks every participant of an abandoned group install as one that must be replaced. + /// + /// A participant that staged holds a replacement state nothing installed; one that + /// activated holds installed state no group resumed. Neither is a participant this + /// session may reuse, and `ipc-v1` section 6's last paragraph is explicit that v1 does + /// not silently reattach one. + fn taint_group(&mut self, staged: &[(Id, WorkerRef, Id)]) { + for (who, _, _) in staged { + self.tainted.insert(who.clone()); + } + } + + /// Every identity a restore checks before a single participant is asked to stage. + fn check_restore_identity( + &self, + manifest: &crate::state::CheckpointManifest, + record: &crate::state::GenerationRecord, + ) -> DomainResult<()> { + if manifest.source_scope.session_id != self.session_id { + return Err(DomainError::before( + ErrorCode::IdentityMismatch, + "the checkpoint belongs to another session", + )); + } + if manifest.episode_id != self.episode_id { + return Err(DomainError::before( + ErrorCode::IdentityMismatch, + "the checkpoint belongs to another episode", + )); + } + if manifest.scheduler_id != "lockstep-v1" { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + format!( + "the checkpoint was scheduled by {}, not lockstep-v1", + manifest.scheduler_id + ), + )); + } + if record.compatibility_digest != manifest.compatibility.digest() { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + "the store manifest and the envelope disagree about compatibility", + )); + } + let live = match &self.descriptor { + Some(descriptor) => crate::state::Compatibility::of(descriptor), + None => { + return Err(DomainError::before( + ErrorCode::InvalidPhase, + "the session has no environment descriptor to compare compatibility with", + )); + } + }; + // Names the identity that differs -- backend, content, patch, controller, parser or + // state format -- rather than one opaque digest mismatch. + manifest + .compatibility + .compare(&live) + .map_err(|e| DomainError::before(ErrorCode::IncompatibleState, e))?; + let live_ports: Vec<(Id, Id)> = self + .agents + .iter() + .map(|slot| (slot.port_id.clone(), slot.agent_id.clone())) + .collect(); + if manifest.port_map != live_ports { + return Err(DomainError::before( + ErrorCode::IdentityMismatch, + "the checkpoint's port map is not this composition's", + )); + } + if manifest.environment.worker_id != self.environment.worker_id { + return Err(DomainError::before( + ErrorCode::IdentityMismatch, + "the checkpoint's world is not this composition's environment", + )); + } + let recorded: Vec = manifest.agents.iter().map(|a| a.agent_id.clone()).collect(); + if recorded != self.agent_ids() { + return Err(DomainError::before( + ErrorCode::IdentityMismatch, + "the checkpoint's agents are not this composition's", + )); + } + for (row, slot) in manifest.agents.iter().zip(&self.agents) { + if row.profile_digest != slot.profile.digest || row.seed != slot.seed { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + format!( + "{}'s checkpoint was taken under another profile or seed", + row.agent_id + ), + )); + } + } + Ok(()) + } + + #[allow(clippy::too_many_lines)] + async fn restore_group( + &mut self, + envelope: &fly_session_types::checkpoint::Envelope, + manifest: &crate::state::CheckpointManifest, + new_epoch: &Id, + boundary: u64, + ) -> Outcome { + let scope = scope_at(&self.session_id, new_epoch, boundary); + // Step 3: import the payloads as *new* artifacts. Nothing the fence dropped is asked + // to come back, and a router that restarted has none of the old roots anyway. + let mut imported: BTreeMap = BTreeMap::new(); + for (name, bytes) in &envelope.payloads { + let digest = digest_of_bytes(bytes); + let artifact = match crate::state::seal_payload(&self.bus, bytes, &digest).await { + Ok(artifact) => artifact, + Err(e) => return Err(self.fail_now(e, "restore")), + }; + imported.insert(name.clone(), (artifact, digest, bytes.len() as u64)); + } + let imported_ids: Vec = imported + .values() + .map(|(artifact, _, _)| artifact.reference().artifact_id.clone()) + .collect(); + + // Step 3, continued: stage every participant. A refusal anywhere leaves nothing + // staged that will ever be activated, because the whole install is abandoned. + let mut staged: Vec<(Id, WorkerRef, Id)> = Vec::new(); + let mut order: Vec<(Id, WorkerRef, String, Digest)> = Vec::new(); + order.push(( + self.environment.worker_id.clone(), + self.environment.clone(), + manifest.environment.payload.clone(), + manifest.compatibility.digest(), + )); + for index in 0..self.agents.len() { + let slot = &self.agents[index]; + let row = manifest + .agents + .iter() + .find(|row| row.agent_id == slot.agent_id) + .expect("the agent set was checked"); + order.push(( + slot.agent_id.clone(), + slot.worker.clone(), + row.payload.clone(), + crate::agent::agent_compatibility_digest( + &row.agent_id, + &row.profile_digest, + &row.dataset_digest, + &row.model_version, + &row.plasticity_version, + row.seed, + ), + )); + } + for (who, worker, payload_name, compatibility_digest) in &order { + let Some((artifact, _, _)) = imported.get(payload_name) else { + return Err(self.fail_now( + DomainError::before( + ErrorCode::IncompatibleState, + format!("the checkpoint has no payload {payload_name} for {who}"), + ), + "stage-restore", + )); + }; + let params = StageRestoreParams { + checkpoint_id: manifest.checkpoint_id.clone(), + source_scope: manifest.source_scope.clone(), + compatibility_digest: compatibility_digest.clone(), + payload: artifact.reference().clone(), + }; + let attachments = [(crate::state::PAYLOAD_ATTACHMENT, artifact)]; + let reply = match self + .call( + worker, + "State.StageRestore", + Some(scope.clone()), + object(params.to_json()), + &attachments, + &[], + ) + .await + { + Ok(reply) => reply, + Err(failure) => { + // Whoever already staged is holding a replacement state this group will + // never install. Nothing is resumed, and none of them is reused. + for (done, _, _) in &staged { + self.tainted.insert(done.clone()); + } + return Err(failure); + } + }; + let result: StageRestoreResult = match reply.parse() { + Ok(result) => result, + Err(e) => { + for (done, _, _) in &staged { + self.tainted.insert(done.clone()); + } + return Err(self.fail_now(e, "stage-restore")); + } + }; + if result.checkpoint_id != manifest.checkpoint_id { + for (done, _, _) in &staged { + self.tainted.insert(done.clone()); + } + return Err(self.fail_now( + DomainError::before( + ErrorCode::IdentityMismatch, + "a staged restore names another checkpoint", + ), + "stage-restore", + )); + } + self.lifecycle_acks.push((worker.clone(), reply.request_id.clone())); + self.audit.push(format!("staged:{who}")); + staged.push((who.clone(), worker.clone(), result.restore_token)); + } + + // `state-media-v1` section 5: activation happens "after every participant **and + // coordinator** state validates". The coordinator's own staged ledgers are checked + // here, before a single token is activated, so a checkpoint whose task ledger or + // executor state is unreadable installs nothing anywhere. + let payloads: BTreeMap> = envelope + .payloads + .iter() + .map(|(name, bytes)| (name.clone(), bytes.clone())) + .collect(); + let descriptor = match &self.descriptor { + Some(descriptor) => descriptor.clone(), + None => { + self.taint_group(&staged); + return Err(self.fail_now( + DomainError::before( + ErrorCode::InvalidPhase, + "the session has no environment descriptor to restore against", + ), + "restore", + )); + } + }; + if let Err(e) = Coordinator::validate_coordinator_state( + manifest, + &payloads, + &descriptor, + self.agents.len(), + &*self.task, + &self.executors, + ) { + self.taint_group(&staged); + return Err(self.fail_now(e, "restore")); + } + + // Step 3, last part: activate. Anything that activates before a failure holds state + // no group resumed, so it is recorded as tainted and must be replaced. + let mut activated: Vec = Vec::new(); + let mut restored_observation: Option<(WorldObservation, BTreeMap)> = + None; + for (who, worker, token) in &staged { + let params = ActivateRestoreParams { restore_token: token.clone() }; + let want = if *who == self.environment.worker_id { + self.media_names.clone() + } else { + Vec::new() + }; + let reply = match self + .call( + worker, + "State.ActivateRestore", + Some(scope.clone()), + object(params.to_json()), + &[], + &want, + ) + .await + { + Ok(reply) => reply, + Err(failure) => { + self.taint_group(&staged); + return Err(failure); + } + }; + let result: ActivateRestoreResult = match reply.parse() { + Ok(result) => result, + Err(e) => { + self.taint_group(&staged); + return Err(self.fail_now(e, "activate-restore")); + } + }; + let role = if *who == self.environment.worker_id { + Role::Environment + } else { + Role::Agent + }; + if let Err(e) = result.validate_for_role(role) { + self.taint_group(&staged); + return Err(self.fail_now(DomainError::invalid(e.0), "activate-restore")); + } + if result.committed_step != boundary || result.checkpoint_id != manifest.checkpoint_id { + self.taint_group(&staged); + return Err(self.fail_now( + DomainError::before( + ErrorCode::IdentityMismatch, + "an activation names another boundary or checkpoint", + ), + "activate-restore", + )); + } + if let Some(observation) = result.observation { + let (views, audio) = media::split_attachments(reply.artifacts); + if !audio.is_empty() { + self.taint_group(&staged); + return Err(self.fail_now( + DomainError::new( + ErrorCode::BufferInvalid, + "a restored observation carried an audio chunk; no interval was played", + MutationCertainty::Unknown, + ), + "activate-restore", + )); + } + restored_observation = Some((observation, views)); + } + self.lifecycle_acks.push((worker.clone(), reply.request_id.clone())); + self.audit.push(format!("activated:{who}")); + activated.push(who.clone()); + } + + // Step 4: verify identity and boundary, and flush the old media and parser state. + let Some((observation, views)) = restored_observation else { + self.taint_group(&staged); + return Err(self.fail_now( + DomainError::before( + ErrorCode::IncompatibleState, + "no participant returned the restored world observation", + ), + "activate-restore", + )); + }; + let install = self.install_restored( + manifest, + new_epoch, + boundary, + &descriptor, + observation, + views, + &payloads, + ); + if let Err(e) = install { + self.taint_group(&staged); + return Err(self.fail_now(e, "restore")); + } + + // Step 5: Paused(k), and only now is the fence lifted. + self.epoch = new_epoch.clone(); + self.transition(Phase::Paused(boundary))?; + self.fenced = false; + self.acknowledge_lifecycle().await?; + self.publish_checkpoint_event("restored", &manifest.checkpoint_id, boundary, None) + .await?; + self.publish_descriptor().await?; + self.publish_snapshot(boundary, &BTreeMap::new(), &[], &[]).await?; + self.audit.push(format!("restored:{}@{boundary}", manifest.checkpoint_id)); + Ok(RestoreReport { + checkpoint_id: manifest.checkpoint_id.clone(), + boundary, + epoch: new_epoch.clone(), + staged: staged.iter().map(|(who, _, _)| who.clone()).collect(), + tokens: staged + .iter() + .map(|(who, _, token)| (who.clone(), token.clone())) + .collect(), + activated, + imported: imported_ids, + }) + } + + /// Reads one of the checkpoint's payloads as JSON, or says which one is unreadable. + fn read_payload(payloads: &BTreeMap>, name: &str) -> DomainResult { + let bytes = payloads.get(name).ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + format!("the checkpoint has no payload {name}"), + ) + })?; + serde_json::from_slice(bytes).map_err(|e| { + DomainError::before( + ErrorCode::IncompatibleState, + format!("payload {name} is not JSON: {e}"), + ) + }) + } + + /// Validates every coordinator-owned payload, changing nothing. + /// + /// This runs after the group has staged and before anything activates, which is the order + /// `state-media-v1` section 5 sets. It is a separate pass from the install below on + /// purpose: a group that cannot be resumed coherently must not have resumed some of it. + fn validate_coordinator_state( + manifest: &crate::state::CheckpointManifest, + payloads: &BTreeMap>, + descriptor: &EnvironmentDescriptor, + agents: usize, + task: &dyn crate::task::Task, + executors: &BTreeMap>, + ) -> DomainResult<()> { + let ledger = TypedValue::from_json(&Coordinator::read_payload( + payloads, + &manifest.coordinator.task_ledger, + )?) + .map_err(|e| DomainError::before(ErrorCode::IncompatibleState, e.0))?; + task.validate_restore(&ledger)?; + for (agent_id, name) in &manifest.coordinator.executor_state { + let state = TypedValue::from_json(&Coordinator::read_payload(payloads, name)?) + .map_err(|e| DomainError::before(ErrorCode::IncompatibleState, e.0))?; + let executor = executors.get(agent_id).ok_or_else(|| { + DomainError::before( + ErrorCode::IdentityMismatch, + format!("{agent_id} has no configured action executor"), + ) + })?; + executor.validate_restore(&state)?; + } + let inspection = TypedValue::from_json(&Coordinator::read_payload( + payloads, + &manifest.coordinator.prior_inspection, + )?) + .map_err(|e| DomainError::before(ErrorCode::IncompatibleState, e.0))?; + if inspection.schema != descriptor.inspection_schema { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + "the recorded prior inspection is not the declared inspection schema", + )); + } + let record = + Coordinator::read_payload(payloads, &manifest.coordinator.admission_state)?; + let recorded = record + .get("agents") + .and_then(Value::as_array) + .ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "the coordinator record has no agent list", + ) + })?; + if recorded.len() != agents { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + "the coordinator record names another number of agents", + )); + } + if record.get("audioPositions").and_then(Value::as_object).is_none() { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + "the coordinator record has no audio positions", + )); + } + Ok(()) + } + + /// Installs the coordinator's own staged state, after every participant activated. + #[allow(clippy::too_many_arguments)] + fn install_restored( + &mut self, + manifest: &crate::state::CheckpointManifest, + new_epoch: &Id, + boundary: u64, + descriptor: &EnvironmentDescriptor, + observation: WorldObservation, + views: BTreeMap, + payloads: &BTreeMap>, + ) -> DomainResult<()> { + if observation.boundary != boundary { + return Err(DomainError::before( + ErrorCode::IdentityMismatch, + "the restored observation is not at the restored boundary", + )); + } + if observation.world_time != manifest.world_time { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + "the restored observation's world time is not the checkpoint's", + )); + } + observation + .validate_against(descriptor) + .map_err(|e| DomainError::new(ErrorCode::BufferInvalid, e, MutationCertainty::Unknown))?; + media::check_required_views(descriptor, &observation)?; + // An installed observation ran no transition, so it carries no chunk. Old epoch audio + // arriving as current is exactly what this refuses. + media::check_required_audio(descriptor, &observation, media::ObservationOrigin::Installed)?; + for view in &observation.sensory_views { + let name = media::view_attachment(&view.view_id); + match views.get(&name) { + Some(artifact) if artifact.reference() == &view.pixels => {} + _ => { + return Err(DomainError::new( + ErrorCode::BufferInvalid, + format!( + "the restored view {} arrived without a live owned handle", + view.view_id + ), + MutationCertainty::Unknown, + )); + } + } + } + + let read = |name: &str| Coordinator::read_payload(payloads, name); + + let ledger = TypedValue::from_json(&read(&manifest.coordinator.task_ledger)?) + .map_err(|e| DomainError::before(ErrorCode::IncompatibleState, e.0))?; + self.task.validate_restore(&ledger)?; + for (agent_id, name) in &manifest.coordinator.executor_state { + let state = TypedValue::from_json(&read(name)?) + .map_err(|e| DomainError::before(ErrorCode::IncompatibleState, e.0))?; + let executor = self.executors.get(agent_id).ok_or_else(|| { + DomainError::before( + ErrorCode::IdentityMismatch, + format!("{agent_id} has no configured action executor"), + ) + })?; + executor.validate_restore(&state)?; + } + let record = read(&manifest.coordinator.admission_state)?; + let inspection = TypedValue::from_json(&read(&manifest.coordinator.prior_inspection)?) + .map_err(|e| DomainError::before(ErrorCode::IncompatibleState, e.0))?; + if inspection.schema != descriptor.inspection_schema { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + "the recorded prior inspection is not the declared inspection schema", + )); + } + let agents = record + .get("agents") + .and_then(Value::as_array) + .ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "the coordinator record has no agent list", + ) + })? + .clone(); + if agents.len() != self.agents.len() { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + "the coordinator record names another number of agents", + )); + } + let positions = record + .get("audioPositions") + .and_then(Value::as_object) + .ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "the coordinator record has no audio positions", + ) + })?; + let mut audio_positions = BTreeMap::new(); + for (stream, value) in positions { + let sample: u64 = value + .as_str() + .ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "a recorded audio position is not a canonical U64", + ) + })? + .parse() + .map_err(|_| { + DomainError::before( + ErrorCode::IncompatibleState, + "a recorded audio position is not a canonical U64", + ) + })?; + audio_positions.insert(stream.clone(), sample); + } + // Everything above validated. From here the coordinator installs, in one pass. + self.task.install_restore(new_epoch, &ledger)?; + for (agent_id, name) in &manifest.coordinator.executor_state { + let state = TypedValue::from_json(&read(name)?) + .map_err(|e| DomainError::before(ErrorCode::IncompatibleState, e.0))?; + let executor = self + .executors + .get_mut(agent_id) + .expect("checked immediately above"); + executor.install_restore(&state)?; + } + for value in &agents { + let agent_id = value.get("agentId").and_then(Value::as_str).ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "a coordinator agent record has no agentId", + ) + })?; + let slot = self + .agents + .iter_mut() + .find(|slot| slot.agent_id == agent_id) + .ok_or_else(|| { + DomainError::before( + ErrorCode::IdentityMismatch, + format!("the coordinator record names {agent_id}, which is not configured"), + ) + })?; + let context = TypedValue::from_json(value.get("context").ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "a coordinator agent record has no decision context", + ) + })?) + .map_err(|e| DomainError::before(ErrorCode::IncompatibleState, e.0))?; + let committed: u64 = value + .get("committedStep") + .and_then(Value::as_str) + .ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "a coordinator agent record has no committedStep", + ) + })? + .parse() + .map_err(|_| { + DomainError::before( + ErrorCode::IncompatibleState, + "a recorded committed step is not a canonical U64", + ) + })?; + if committed != boundary { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + "a coordinator agent record is at another boundary", + )); + } + let brain_ticks: u64 = value + .get("brainTicks") + .and_then(Value::as_str) + .ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "a coordinator agent record has no brainTicks", + ) + })? + .parse() + .map_err(|_| { + DomainError::before( + ErrorCode::IncompatibleState, + "a recorded tick count is not a canonical U64", + ) + })?; + let remainder = RationalNs::from_json(value.get("remainder").ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "a coordinator agent record has no remainder", + ) + })?) + .map_err(|e| DomainError::before(ErrorCode::IncompatibleState, e.0))?; + slot.context_digest = context.digest(); + slot.context = context; + slot.committed_step = committed; + slot.brain_ticks = brain_ticks; + slot.remainder = remainder; + slot.prepared = None; + slot.prepare_request = None; + } + // Old media and old parser state are replaced, never carried: the previous epoch's + // handles were dropped by the fence and the timelines start again at their preserved + // sample positions with a discontinuity. + self.views = views; + self.pending_views.clear(); + self.audio.clear(); + self.pending_audio.clear(); + self.timelines = AudioTimelines::restored(descriptor, &audio_positions)?; + self.observation = Some(observation); + self.episode = None; + self.last_advance_request = None; + self.last_commit_requests.clear(); + self.pacing = Some(Pacing::new(descriptor.step_duration)); + Ok(()) + } + + /// The epoch-derived identities of this session's behaviour trace, mapped onto `to_epoch`. + /// + /// `step-v1` section 8 compares behaviour across runs. A resumed run runs in a new epoch, + /// and scope, batch identity and every task event identity are derived from it, so a + /// comparison either accounts for that or compares nothing. This is what "accounting for + /// new epoch metadata" is: an explicit, total rewrite of the epoch-derived fields, which + /// fails rather than passing anything through it does not recognise. + pub fn rebase(&self, to_epoch: &Id) -> Outcome { + let events = self.task.rebase_ids(to_epoch).map_err(|e| SessionFailure { + error: e, + phase: self.phases.phase().label(), + detail: "rebase".to_owned(), + participant: None, + })?; + Ok(EpochRebase { + from: self.epoch.clone(), + to: to_epoch.clone(), + events, + }) + } +} diff --git a/services/flysim/crates/fly-session/src/environment.rs b/services/flysim/crates/fly-session/src/environment.rs index 8a2947b..9c1bd85 100644 --- a/services/flysim/crates/fly-session/src/environment.rs +++ b/services/flysim/crates/fly-session/src/environment.rs @@ -13,6 +13,8 @@ use std::collections::BTreeSet; +use serde_json::Value; + use crate::media::{self, AudioSource, RenderCounter, ViewPipeline}; use crate::task::{controller_schema_ref, inspection, inspection_schema}; // `crate::types` is this crate's facade over the shared `fly-session-types` crate; the @@ -48,6 +50,12 @@ pub struct EnvironmentFaults { pub omit_audio_at_boundary: Option, /// Emit an audio chunk that starts before the previous chunk ended. pub overlapping_audio_at_boundary: Option, + /// Refuse `State.StageRestore`, so a group install meets a participant that will not + /// validate. + pub fail_stage_restore: bool, + /// Refuse `State.ActivateRestore` after staging, so a group meets a failure halfway + /// through activation. + pub fail_activate_restore: bool, } #[derive(Clone, Debug)] @@ -86,6 +94,10 @@ pub struct CounterEnvironment { audio: Option, /// The frame served at the previous boundary, kept only so a fault can serve it again. previous_view: Option<(ViewRef, flybus::Artifact)>, + /// A validated replacement world the live session cannot see yet. + staged: Option, + /// Restore tokens this world has activated. A token activates once. + activated: BTreeSet, } impl CounterEnvironment { @@ -104,10 +116,17 @@ impl CounterEnvironment { pipeline: None, audio: None, previous_view: None, + staged: None, + activated: BTreeSet::new(), config, } } + /// True while a validated replacement world is staged and not yet activated. + pub fn has_staged_restore(&self) -> bool { + self.staged.is_some() + } + pub fn status(&self) -> StatusCell { self.status.clone() } @@ -477,7 +496,7 @@ impl WorkerEndpoint for CounterEnvironment { vec![ id("world-step-v1"), id("pixel-observation-v1"), - id("checkpoint-v1"), + id(crate::state::CHECKPOINT_CAPABILITY), ] } @@ -490,7 +509,13 @@ impl WorkerEndpoint for CounterEnvironment { } fn methods(&self) -> Vec<&'static str> { - vec!["Environment.Initialize", "Environment.Advance"] + vec![ + "Environment.Initialize", + "Environment.Advance", + "State.Capture", + "State.StageRestore", + "State.ActivateRestore", + ] } fn handle<'a>(&'a mut self, ctx: HandlerCtx<'a>) -> BoxFuture<'a, DomainResult> { @@ -498,6 +523,9 @@ impl WorkerEndpoint for CounterEnvironment { match ctx.method { "Environment.Initialize" => self.initialize(&ctx).await, "Environment.Advance" => self.advance(&ctx).await, + "State.Capture" => self.state_capture(&ctx).await, + "State.StageRestore" => self.state_stage_restore(&ctx).await, + "State.ActivateRestore" => self.state_activate_restore(&ctx).await, other => Err(DomainError::before( ErrorCode::Unsupported, format!("{other} is not an environment method"), @@ -516,3 +544,488 @@ pub fn synthetic_asset(asset_id: &str, body: &str) -> AssetRef { format: id("fly-config-v1"), } } + +// ------------------------------------------------------------------------------------------- +// STATE-01: capture and restore + +/// The version this payload layout is written and read under. +pub const WORLD_PAYLOAD_VERSION: u64 = 1; + +fn incompatible(message: impl std::fmt::Display) -> DomainError { + DomainError::before(ErrorCode::IncompatibleState, message) +} + +/// One staged restore, held outside the live world until it is activated. +struct StagedWorld { + token: Id, + checkpoint_id: Id, + scope: Scope, + episode_id: Id, + descriptor: EnvironmentDescriptor, + boundary: u64, + counter: i64, + world_time: RationalNs, + advances: u64, + frames: Vec<(u64, i64)>, + audio_next_sample: u64, + audio_phase: u64, + audio_accumulator: u128, + audio_denominator: u128, +} + +impl CounterEnvironment { + /// `State.Capture`: the world at its committed boundary, including its pending sensor + /// pipeline. + /// + /// The pipeline is recorded as reconstruction inputs -- the producing boundary and the + /// world counter of every retained frame -- and never as an artifact identity: a + /// transient artifact belongs to the router that is running now, and a checkpoint outlives + /// it. + async fn state_capture(&mut self, ctx: &HandlerCtx<'_>) -> DomainResult { + let scope = ctx.scope()?.clone(); + let Some(descriptor) = self.descriptor.clone() else { + return Err(DomainError::before( + ErrorCode::InvalidPhase, + "this environment is uninitialized", + )); + }; + if scope.session_id != self.config.session_id { + return Err(DomainError::before( + ErrorCode::IdentityMismatch, + "this environment belongs to another session", + )); + } + match &self.epoch { + Some(epoch) if *epoch == scope.epoch => {} + _ => { + return Err(DomainError::before( + ErrorCode::StaleEpoch, + "State.Capture names an epoch this environment has left", + )); + } + } + if scope.step != self.boundary { + return Err(DomainError::before( + if scope.step < self.boundary { + ErrorCode::StaleStep + } else { + ErrorCode::FutureStep + }, + "State.Capture must name the boundary the world is at", + )); + } + let params: CaptureParams = ctx.params()?; + let pipeline = self + .pipeline + .as_ref() + .ok_or_else(|| DomainError::before(ErrorCode::InvalidPhase, "no view pipeline"))?; + let audio = self + .audio + .as_ref() + .ok_or_else(|| DomainError::before(ErrorCode::InvalidPhase, "no audio source"))?; + let (accumulator, denominator) = audio.accumulator(); + let previous = self.status.state(); + self.status.set_state(WorkerState::Capturing); + let payload = serde_json::json!({ + "payloadVersion": WORLD_PAYLOAD_VERSION, + "kind": "world", + "workerId": self.config.worker_id.as_str(), + "checkpointId": params.checkpoint_id.as_str(), + "sourceScope": scope.to_json(), + "episodeId": self.episode_id.clone().expect("initialized").as_str(), + "committedStep": self.boundary.to_string(), + "counter": self.counter.to_string(), + "worldTime": self.world_time.to_json(), + "advances": self.advances.to_string(), + "descriptor": descriptor.to_json(), + "pipeline": { + // The declared delay's whole queue, oldest first. + "frames": pipeline + .retained() + .into_iter() + .map(|(boundary, counter)| serde_json::json!({ + "boundary": boundary.to_string(), + "counter": counter.to_string(), + })) + .collect::>(), + }, + "audio": { + "nextSample": audio.next_sample().to_string(), + "phase": audio.phase().to_string(), + "accumulator": accumulator.to_string(), + "denominator": denominator.to_string(), + "chunks": audio.chunks().to_string(), + }, + }); + let bytes = canonicalize(&payload) + .map_err(|e| DomainError::invalid(format!("State.Capture: {}", e.0)))? + .into_bytes(); + let digest = digest_of_bytes(&bytes); + let artifact = crate::state::seal_payload(ctx.client, &bytes, &digest).await?; + // A capture reads the world; it does not advance it. + self.status.set_state(previous); + let result = CaptureResult { + checkpoint_id: params.checkpoint_id, + boundary: self.boundary, + compatibility_digest: crate::state::Compatibility::of(&descriptor).digest(), + payload: artifact.reference().clone(), + }; + Ok(HandlerReply::with_artifacts( + object(result.to_json()), + vec![(crate::state::PAYLOAD_ATTACHMENT.to_owned(), artifact)], + )) + } + + /// `State.StageRestore`: validate a replacement world into a staging slot. + async fn state_stage_restore(&mut self, ctx: &HandlerCtx<'_>) -> DomainResult { + let scope = ctx.scope()?.clone(); + if scope.session_id != self.config.session_id { + return Err(DomainError::before( + ErrorCode::IdentityMismatch, + "this environment belongs to another session", + )); + } + if let Some(epoch) = &self.epoch + && *epoch == scope.epoch + { + return Err(DomainError::before( + ErrorCode::StaleEpoch, + "State.StageRestore proposes the epoch this environment is already running", + )); + } + if self.descriptor.is_some() { + // A world that is already running a boundary is not a quiescent replacement: the + // group replaces it rather than restoring over a live one. + return Err(DomainError::before( + ErrorCode::InvalidPhase, + "State.StageRestore needs an uninitialized replacement environment", + )); + } + let params: StageRestoreParams = ctx.params()?; + if params.source_scope.step != scope.step { + return Err(DomainError::invalid( + "State.StageRestore's scope step must be the source boundary", + )); + } + let artifact = ctx.artifact(crate::state::PAYLOAD_ATTACHMENT)?; + if artifact.reference() != ¶ms.payload { + return Err(DomainError::before( + ErrorCode::BufferInvalid, + "the staged payload attachment is not the artifact the request names", + )); + } + let bytes = artifact.read_all().await.map_err(|e| { + DomainError::before( + ErrorCode::BufferInvalid, + format!("the staged payload could not be read: {}", e.message), + ) + })?; + let declared = params + .payload + .digest + .clone() + .ok_or_else(|| incompatible("a checkpoint payload must carry a content digest"))?; + let actual = digest_of_bytes(&bytes); + if actual != declared || bytes.len() as u64 != params.payload.byte_length { + return Err(incompatible( + "the staged payload is not the content the request declares", + )); + } + let value: Value = serde_json::from_slice(&bytes) + .map_err(|e| DomainError::invalid(format!("the staged payload is not JSON: {e}")))?; + let text = |key: &str| -> DomainResult { + value + .get(key) + .and_then(Value::as_str) + .map(str::to_owned) + .ok_or_else(|| incompatible(format!("the world payload has no {key}"))) + }; + let number = |key: &str| -> DomainResult { + text(key)? + .parse::() + .map_err(|_| incompatible(format!("the world payload's {key} is not a U64"))) + }; + if value.get("payloadVersion").and_then(Value::as_u64) != Some(WORLD_PAYLOAD_VERSION) { + return Err(incompatible("the world payload is another payload version")); + } + if text("kind")? != "world" { + return Err(incompatible("this payload is not a world's state")); + } + if text("workerId")? != self.config.worker_id { + return Err(DomainError::before( + ErrorCode::IdentityMismatch, + "the staged payload belongs to another world", + )); + } + if text("checkpointId")? != params.checkpoint_id { + return Err(incompatible("the staged payload belongs to another checkpoint")); + } + let source_scope = Scope::from_json( + value + .get("sourceScope") + .ok_or_else(|| incompatible("the world payload has no sourceScope"))?, + ) + .map_err(|e| incompatible(format!("the world payload's sourceScope: {}", e.0)))?; + if source_scope != params.source_scope { + return Err(incompatible( + "the staged payload was captured at another source scope", + )); + } + let committed_step = number("committedStep")?; + if committed_step != params.source_scope.step { + return Err(incompatible( + "the staged payload's committed step is not the source boundary", + )); + } + let descriptor = EnvironmentDescriptor::from_json( + value + .get("descriptor") + .ok_or_else(|| incompatible("the world payload has no descriptor"))?, + ) + .map_err(|e| incompatible(format!("the world payload's descriptor: {}", e.0)))?; + // The replacement builds the descriptor it would advertise and compares. A world + // started with other ports, another cadence or another declared render delay is a + // different backend, not this one resumed. + let live = self.build_descriptor()?; + if descriptor != live { + return Err(incompatible( + "the staged world was captured under another environment descriptor", + )); + } + let expected = crate::state::Compatibility::of(&descriptor).digest(); + if expected != params.compatibility_digest { + return Err(incompatible(format!( + "the staged world's compatibility {expected} is not the {} the restore requires", + params.compatibility_digest + ))); + } + let counter: i64 = text("counter")? + .parse() + .map_err(|_| incompatible("the world payload's counter is not an integer"))?; + let world_time = RationalNs::from_json( + value + .get("worldTime") + .ok_or_else(|| incompatible("the world payload has no worldTime"))?, + ) + .map_err(|e| incompatible(format!("the world payload's worldTime: {}", e.0)))?; + let pipeline_value = value + .get("pipeline") + .and_then(|p| p.get("frames")) + .and_then(Value::as_array) + .ok_or_else(|| incompatible("the world payload has no pipeline frames"))?; + let mut frames = Vec::with_capacity(pipeline_value.len()); + for frame in pipeline_value { + let boundary = frame + .get("boundary") + .and_then(Value::as_str) + .ok_or_else(|| incompatible("a captured frame has no boundary"))? + .parse::() + .map_err(|_| incompatible("a captured frame's boundary is not a U64"))?; + let frame_counter = frame + .get("counter") + .and_then(Value::as_str) + .ok_or_else(|| incompatible("a captured frame has no counter"))? + .parse::() + .map_err(|_| incompatible("a captured frame's counter is not an integer"))?; + frames.push((boundary, frame_counter)); + } + match frames.last() { + Some((boundary, _)) if *boundary == committed_step => {} + _ => { + return Err(incompatible( + "the captured pipeline does not end at the committed boundary", + )); + } + } + let audio_value = value + .get("audio") + .ok_or_else(|| incompatible("the world payload has no audio state"))?; + let audio_number = |key: &str| -> DomainResult { + audio_value + .get(key) + .and_then(Value::as_str) + .ok_or_else(|| incompatible(format!("the captured audio state has no {key}")))? + .parse::() + .map_err(|_| incompatible(format!("the captured audio {key} is not a number"))) + }; + let audio_next_sample = u64::try_from(audio_number("nextSample")?) + .map_err(|_| incompatible("the captured audio position is outside U64"))?; + let audio_phase = u64::try_from(audio_number("phase")?) + .map_err(|_| incompatible("the captured audio phase is outside U64"))?; + + if self.config.faults.fail_stage_restore { + return Err(incompatible( + "injected staging refusal: this participant's replacement state does not \ +validate", + )); + } + if let Some(staged) = &self.staged { + return Err(DomainError::before( + ErrorCode::Conflict, + format!( + "this environment already holds the staged restore {} for checkpoint {}", + staged.token, staged.checkpoint_id + ), + )); + } + let token = crate::agent::restore_token( + ¶ms.checkpoint_id, + &scope, + &actual, + &self.config.incarnation_id, + ); + if self.activated.contains(&token) { + return Err(DomainError::before( + ErrorCode::Conflict, + "this exact restore was already activated on this environment", + )); + } + self.staged = Some(StagedWorld { + token: token.clone(), + checkpoint_id: params.checkpoint_id.clone(), + scope, + episode_id: parse_id(&text("episodeId")?) + .map_err(|e| incompatible(format!("the world payload's episodeId {e}")))?, + descriptor, + boundary: committed_step, + counter, + world_time, + advances: number("advances")?, + frames, + audio_next_sample, + audio_phase, + audio_accumulator: audio_number("accumulator")?, + audio_denominator: audio_number("denominator")?, + }); + self.status.set_state(WorkerState::StagedRestore); + let result = StageRestoreResult { + checkpoint_id: params.checkpoint_id, + restore_token: token, + }; + Ok(HandlerReply::from(&result)) + } + + /// `State.ActivateRestore`: install the staged world and return its coherent observation. + /// + /// Nothing advances. The pipeline's frames are rendered again into fresh artifacts of the + /// current store, which is what "the durable store imports fresh immutable bus artifacts" + /// means on the producing side, and the observation carries no audio chunk because no + /// interval was played. + async fn state_activate_restore( + &mut self, + ctx: &HandlerCtx<'_>, + ) -> DomainResult { + let params: ActivateRestoreParams = ctx.params()?; + if self.activated.contains(¶ms.restore_token) { + return Err(DomainError::before( + ErrorCode::Conflict, + "this restore token has already been activated", + )); + } + let Some(staged) = self.staged.take() else { + return Err(DomainError::before( + ErrorCode::InvalidPhase, + "this environment holds no staged restore", + )); + }; + if staged.token != params.restore_token { + let token = staged.token.clone(); + self.staged = Some(staged); + return Err(DomainError::before( + ErrorCode::IdentityMismatch, + format!( + "this environment's staged restore is {token}, not {}", + params.restore_token + ), + )); + } + if self.config.faults.fail_activate_restore { + let token = staged.token.clone(); + self.staged = Some(staged); + self.status.set_state(WorkerState::Failed); + return Err(DomainError::new( + ErrorCode::BackendFailure, + format!("injected activation failure; {token} stays staged and unresumed"), + MutationCertainty::None, + )); + } + self.status.set_state(WorkerState::Restoring); + let mut pipeline = ViewPipeline::new( + CounterEnvironment::view_descriptor(self.config.observation_delay_steps), + self.config.renders.clone(), + ); + pipeline.restore(ctx.client, &staged.frames).await?; + let audio = AudioSource::restored_from( + CounterEnvironment::audio_descriptor(), + staged.audio_next_sample, + staged.audio_phase, + staged.audio_accumulator, + staged.audio_denominator, + )?; + self.epoch = Some(staged.scope.epoch.clone()); + self.episode_id = Some(staged.episode_id.clone()); + self.descriptor = Some(staged.descriptor.clone()); + self.boundary = staged.boundary; + self.counter = staged.counter; + self.world_time = staged.world_time; + self.advances = staged.advances; + // Batch ids are unique within an epoch, and this is a new one. Keeping the old set + // would refuse nothing extra: a request under the old epoch is already refused by its + // scope. + self.batches.clear(); + self.pipeline = Some(pipeline); + self.audio = Some(audio); + self.previous_view = None; + self.activated.insert(staged.token); + self.status.set_state(WorkerState::Ready); + self.status.set_scope(Some(scope_at( + &staged.scope.session_id, + &staged.scope.epoch, + staged.boundary, + ))); + + let (observation, attachments) = self.restored_observation()?; + let result = ActivateRestoreResult { + committed_step: staged.boundary, + checkpoint_id: staged.checkpoint_id, + observation: Some(observation), + }; + result + .validate_for_role(Role::Environment) + .map_err(|e| DomainError::invalid(e.0))?; + let mut reply = HandlerReply::from(&result); + reply.artifacts = attachments; + Ok(reply) + } + + /// The observation the restored world is already at: no render, no advance, no audio. + fn restored_observation( + &mut self, + ) -> DomainResult<(WorldObservation, Vec<(String, flybus::Artifact)>)> { + let boundary = self.boundary; + let counter = self.counter; + let pipeline = self + .pipeline + .as_ref() + .ok_or_else(|| DomainError::before(ErrorCode::InvalidPhase, "no view pipeline"))?; + let (view, artifact) = pipeline.at(boundary).ok_or_else(|| { + incompatible("the restored pipeline holds no frame for the restored boundary") + })?; + self.previous_view = Some((view.clone(), artifact.clone())); + let observation = WorldObservation { + boundary, + world_time: self.world_time, + engine_frame: Some(boundary.to_string()), + sensory_views: vec![view.clone()], + inspection: inspection(counter, boundary), + broadcast_views: vec![view.clone()], + // No interval was played, so there is no chunk. A chunk here would be an old + // epoch's audio offered as current. + audio: Vec::new(), + }; + Ok(( + observation, + vec![(media::view_attachment(&view.view_id), artifact)], + )) + } +} diff --git a/services/flysim/crates/fly-session/src/harness.rs b/services/flysim/crates/fly-session/src/harness.rs index c537a9b..53159d1 100644 --- a/services/flysim/crates/fly-session/src/harness.rs +++ b/services/flysim/crates/fly-session/src/harness.rs @@ -26,6 +26,7 @@ use crate::media::{RenderCounter, SensorLog}; use crate::launcher::{ AgentLaunch, EnvironmentLaunch, Launcher, ReapOutcome, SUPERVISOR_CLIENT, ThreadBudget, }; +use crate::state::{CheckpointStore, CheckpointWriter, StoreConfig, StoreFaults, WriterConfig, WriterFaults}; use crate::task::{ActionExecutor, CounterTask, IdentityExecutor, Terminal}; // `crate::types` is this crate's facade over the shared `fly-session-types` crate; the // glob keeps the contract's own names in sight instead of restating them. @@ -85,6 +86,14 @@ pub struct HarnessConfig { /// The threads reserved for the coordinator, its router and its store. pub coordinator_threads: usize, pub environment_threads: usize, + /// How many committed generations the durable checkpoint store keeps. + pub store: StoreConfig, + /// The durable write faults this composition injects. + pub store_faults: StoreFaults, + /// The checkpoint queue's bounds. + pub writer: WriterConfig, + /// The writer faults this composition injects. + pub writer_faults: WriterFaults, } impl Default for HarnessConfig { @@ -107,6 +116,10 @@ impl Default for HarnessConfig { thread_budget: None, coordinator_threads: 1, environment_threads: 1, + store: StoreConfig::default(), + store_faults: StoreFaults::default(), + writer: WriterConfig::default(), + writer_faults: WriterFaults::default(), } } } @@ -134,6 +147,14 @@ const ENV_SERVICE: &str = "env.arena"; const ENV_CLIENT: &str = "environment"; const ENV_WORKER: &str = "arena"; const COORDINATOR_CLIENT: &str = "coordinator"; +/// The checkpoint writer's own bus identity. It publishes checkpoint events and nothing else. +const WRITER_CLIENT: &str = "checkpoint-writer"; + +/// How many times one participant may be replaced in a composition. +/// +/// Each replacement connects under its own client id, so a restart is visibly a new +/// participant rather than a silent reattachment, and the policy has to name them all. +const MAX_GENERATIONS: u32 = 8; fn agent_service(agent_id: &Id) -> String { format!("agent.{agent_id}") @@ -172,6 +193,10 @@ pub struct SessionHarness { /// The supervisor. It owns every participant's lifetime and thread allocation. pub launcher: Launcher, observers: Mutex>, + /// Which generation of each participant is running: 1 is the one the composition started. + generations: BTreeMap, + /// Where the durable checkpoint store lives, for a test that reads the files themselves. + checkpoint_root: std::path::PathBuf, } impl SessionHarness { @@ -204,12 +229,16 @@ impl SessionHarness { g.call = vec![Pattern::prefix("agent."), Pattern::prefix("env.")]; }), ) + // The writer publishes the checkpoint events and never calls a participant. + .client(WRITER_CLIENT, grants(|g| g.publish = vec![Pattern::prefix("session.")])) .client(ENV_CLIENT, grants(|g| g.register = vec![Pattern::exact(ENV_SERVICE)])) - .client( - &format!("{ENV_CLIENT}-r2"), - grants(|g| g.register = vec![Pattern::exact(ENV_SERVICE)]), - ) .client("observer", grants(|g| g.subscribe = vec![Pattern::prefix("session.")])); + for generation in 2..=MAX_GENERATIONS { + policy = policy.client( + &format!("{ENV_CLIENT}-r{generation}"), + grants(|g| g.register = vec![Pattern::exact(ENV_SERVICE)]), + ); + } for spec in &config.agents { let service = agent_service(&spec.agent_id); policy = policy.client( @@ -218,10 +247,12 @@ impl SessionHarness { ); // A replacement worker connects under its own client id, so a restart is visibly a // new participant rather than a silent reattachment to the active epoch. - policy = policy.client( - &format!("{}-r2", agent_client(&spec.agent_id)), - grants(|g| g.register = vec![Pattern::exact(&service)]), - ); + for generation in 2..=MAX_GENERATIONS { + policy = policy.client( + &format!("{}-r{generation}", agent_client(&spec.agent_id)), + grants(|g| g.register = vec![Pattern::exact(&service)]), + ); + } } let mut router_config = RouterConfig::new(&store_root); router_config.policy = policy; @@ -299,6 +330,21 @@ impl SessionHarness { } let coordinator_client = launcher.connect(COORDINATOR_CLIENT).await?; + // The durable store lives beside the router's artifact store and never inside it: a + // committed generation is outside the bus's ephemeral collection. + let checkpoint_root = root.join("checkpoints"); + let mut store = CheckpointStore::open(&checkpoint_root, config.store).map_err(refusal)?; + *store.faults_mut() = config.store_faults.clone(); + let writer_client = launcher.connect(WRITER_CLIENT).await?; + let writer = CheckpointWriter::start( + store, + config.writer, + config.writer_faults.clone(), + Some(( + writer_client, + format!("session.{}.checkpoints", config.session_id), + )), + ); let executors: BTreeMap> = config .agents .iter() @@ -316,6 +362,8 @@ impl SessionHarness { Box::new(CounterTask::new(&config.epoch, config.terminal)), executors, ); + let mut coordinator = coordinator; + coordinator.attach_store(writer); Ok(SessionHarness { coordinator, @@ -326,9 +374,16 @@ impl SessionHarness { sensors, launcher, observers: Mutex::new(Vec::new()), + generations: BTreeMap::new(), + checkpoint_root, }) } + /// Where the durable checkpoint store's generations and store manifest live. + pub fn checkpoint_root(&self) -> &std::path::Path { + &self.checkpoint_root + } + pub fn router(&self) -> &Router { self.launcher.router() } @@ -373,10 +428,11 @@ impl SessionHarness { .find(|spec| spec.agent_id == *agent_id) .expect("a configured agent") .clone(); + let generation = self.next_generation(agent_id)?; self.launcher.kill(agent_id).await; let tick_duration = millis(self.config.tick_ms).expect("a positive tick"); - let incarnation_id = - parse_id(&format!("{agent_id}-inc-2")).expect("an agent id plus a suffix is an Id"); + let incarnation_id = parse_id(&format!("{agent_id}-inc-{generation}")) + .expect("an agent id plus a suffix is an Id"); self.launcher .launch_agent(AgentLaunch { session_id: self.config.session_id.clone(), @@ -390,7 +446,7 @@ impl SessionHarness { // predecessor wrote, so a restore's sensory input is visible beside it. sensors: self.sensors.get(agent_id).cloned().unwrap_or_default(), faults: spec.faults.clone(), - client_id: format!("{}-r2", agent_client(agent_id)), + client_id: format!("{}-r{generation}", agent_client(agent_id)), service: agent_service(agent_id), }) .await @@ -403,6 +459,102 @@ impl SessionHarness { }) } + /// Replaces the environment with a fresh, uninitialized incarnation, as a restore needs. + pub async fn restart_environment(&mut self) -> Result { + let worker_id = id(ENV_WORKER); + let generation = self.next_generation(&worker_id)?; + self.launcher.kill(&worker_id).await; + let step_duration = hz(self.config.step_hz).expect("a positive cadence"); + let incarnation_id = parse_id(&format!("arena-inc-{generation}")) + .expect("a worker id plus a suffix is an Id"); + self.launcher + .launch_environment(EnvironmentLaunch { + session_id: self.config.session_id.clone(), + worker_id: worker_id.clone(), + incarnation_id: incarnation_id.clone(), + step_duration, + ports: self.config.agents.iter().map(|a| a.port_id.clone()).collect(), + worker_threads: self.config.environment_threads, + observation_delay_steps: self.config.observation_delay_steps, + renders: self.renders.clone(), + faults: self.config.environment_faults.clone(), + client_id: format!("{ENV_CLIENT}-r{generation}"), + service: ENV_SERVICE.to_owned(), + }) + .await + .map_err(refusal)?; + let worker = self.launcher.worker(&worker_id).expect("just launched"); + Ok(Restarted { + service: worker.identity.service.clone(), + service_incarnation: worker.service_incarnation.clone(), + incarnation_id, + }) + } + + fn next_generation(&mut self, worker_id: &Id) -> Result { + let slot = self.generations.entry(worker_id.clone()).or_insert(1); + if *slot >= MAX_GENERATIONS { + return Err(flybus::BusError::new( + flybus::ErrorCode::QuotaExceeded, + format!( + "{worker_id} has used all {MAX_GENERATIONS} configured client identities; a composition declares how many replacements it allows" + ), + )); + } + *slot += 1; + Ok(*slot) + } + + /// Replaces every participant and points the fenced coordinator at the replacements. + /// + /// This is what a recovery does before it restores: the old participants belong to an + /// invalid epoch, and the references the coordinator pinned are exchanged deliberately. + pub async fn replace_all_participants(&mut self) -> Result<(), flybus::BusError> { + let environment = self.environment_id(); + self.restart_environment().await?; + let worker = self + .launcher + .worker(&environment) + .expect("just launched") + .worker_ref(); + self.coordinator + .replace_participant(&environment, worker) + .map_err(|e| refusal(e.error))?; + for agent_id in self.config.agents.iter().map(|a| a.agent_id.clone()).collect::>() { + self.restart_agent(&agent_id).await?; + let worker = self + .launcher + .worker(&agent_id) + .expect("just launched") + .worker_ref(); + self.coordinator + .replace_participant(&agent_id, worker) + .map_err(|e| refusal(e.error))?; + } + Ok(()) + } + + /// Changes one agent's injected faults, so the replacement the next restart launches is + /// a participant without them. + /// + /// A fault is launch configuration, so clearing one is a relaunch and not a live change: + /// the worker running now keeps whatever it was started with. + pub fn set_agent_faults(&mut self, agent_id: &Id, faults: AgentFaults) { + if let Some(spec) = self + .config + .agents + .iter_mut() + .find(|spec| spec.agent_id == *agent_id) + { + spec.faults = faults; + } + } + + /// Changes the environment's injected faults, with the same relaunch rule. + pub fn set_environment_faults(&mut self, faults: EnvironmentFaults) { + self.config.environment_faults = faults; + } + /// Ends one participant without asking it, as a crash would. pub async fn kill(&mut self, worker_id: &Id) -> ReapOutcome { self.launcher.kill(worker_id).await @@ -466,7 +618,10 @@ impl SessionHarness { /// Reaps every participant and closes the router. pub async fn shutdown(self) { - let SessionHarness { coordinator, mut launcher, observers, .. } = self; + let SessionHarness { mut coordinator, mut launcher, observers, .. } = self; + // The writer task owns artifact handles and a blocking store. Leaving it running + // would leave both behind. + coordinator.shutdown_store().await; drop(coordinator); launcher.reap_all(&id("shutdown")).await; for observer in observers.into_inner().expect("not poisoned") { diff --git a/services/flysim/crates/fly-session/src/launcher.rs b/services/flysim/crates/fly-session/src/launcher.rs index c80551e..43d6705 100644 --- a/services/flysim/crates/fly-session/src/launcher.rs +++ b/services/flysim/crates/fly-session/src/launcher.rs @@ -1231,6 +1231,8 @@ pub(crate) mod flags { pub const PREPARE_DELAY_MS: &str = "prepare-delay-ms"; pub const COMMIT_DELAY_MS: &str = "commit-delay-ms"; pub const FAIL_COMMIT_AT_STEP: &str = "fail-commit-at-step"; + pub const FAIL_STAGE_RESTORE: &str = "fail-stage-restore"; + pub const FAIL_ACTIVATE_RESTORE: &str = "fail-activate-restore"; pub const WORKER: &str = "worker"; pub const PORTS: &str = "ports"; @@ -1271,6 +1273,8 @@ pub(crate) mod flags { PREPARE_DELAY_MS, COMMIT_DELAY_MS, FAIL_COMMIT_AT_STEP, + FAIL_STAGE_RESTORE, + FAIL_ACTIVATE_RESTORE, ]; /// What only the environment is given, media options included. pub const ENVIRONMENT_ONLY: &[&str] = &[ @@ -1285,6 +1289,8 @@ pub(crate) mod flags { TRUNCATED_VIEW_AT_BOUNDARY, OMIT_AUDIO_AT_BOUNDARY, OVERLAPPING_AUDIO_AT_BOUNDARY, + FAIL_STAGE_RESTORE, + FAIL_ACTIVATE_RESTORE, ]; /// What a measurement run or one of its row children is given. pub const MEASURE: &[&str] = &[MODE, AGENTS, STEPS, WARMUP_STEPS, WORKER_THREADS, MODES]; @@ -1327,6 +1333,11 @@ impl Started { arg(flags::WARMUP_TICKS, spec.warmup_ticks), arg(flags::PREPARE_DELAY_MS, spec.faults.prepare_delay_ms), arg(flags::COMMIT_DELAY_MS, spec.faults.commit_delay_ms), + arg(flags::FAIL_STAGE_RESTORE, u64::from(spec.faults.fail_stage_restore)), + arg( + flags::FAIL_ACTIVATE_RESTORE, + u64::from(spec.faults.fail_activate_restore), + ), ]; if let Some(step) = spec.faults.fail_commit_at_step { args.push(arg(flags::FAIL_COMMIT_AT_STEP, step)); @@ -1345,6 +1356,11 @@ impl Started { // The media options a world in another process needs to be exactly this // world. Its render counter and its agents' sensor logs stay there. arg(flags::OBSERVATION_DELAY_STEPS, spec.observation_delay_steps), + arg(flags::FAIL_STAGE_RESTORE, u64::from(spec.faults.fail_stage_restore)), + arg( + flags::FAIL_ACTIVATE_RESTORE, + u64::from(spec.faults.fail_activate_restore), + ), ]; for (flag, boundary) in [ (flags::OMIT_VIEW_AT_BOUNDARY, spec.faults.omit_view_at_boundary), @@ -1458,6 +1474,8 @@ mod flag_tests { truncated_view_at_boundary: Some(3), omit_audio_at_boundary: Some(4), overlapping_audio_at_boundary: Some(5), + fail_stage_restore: true, + fail_activate_restore: true, } } @@ -1491,6 +1509,8 @@ mod flag_tests { fail_commit_at_step: Some(2), prepare_delay_ms: 1, commit_delay_ms: 2, + fail_stage_restore: true, + fail_activate_restore: true, }, client_id: "worker-fly-a".to_owned(), service: "agent.fly-a".to_owned(), @@ -1535,6 +1555,8 @@ mod flag_tests { flags::TRUNCATED_VIEW_AT_BOUNDARY, flags::OMIT_AUDIO_AT_BOUNDARY, flags::OVERLAPPING_AUDIO_AT_BOUNDARY, + flags::FAIL_STAGE_RESTORE, + flags::FAIL_ACTIVATE_RESTORE, ] { assert!(written.contains(&format!("--{flag}")), "--{flag} is not written"); } diff --git a/services/flysim/crates/fly-session/src/lib.rs b/services/flysim/crates/fly-session/src/lib.rs index 15722b5..e6d66c0 100644 --- a/services/flysim/crates/fly-session/src/lib.rs +++ b/services/flysim/crates/fly-session/src/lib.rs @@ -34,6 +34,7 @@ pub mod media; pub mod metrics; pub mod phase; pub mod rpc; +pub mod state; pub mod task; pub mod worker; diff --git a/services/flysim/crates/fly-session/src/media.rs b/services/flysim/crates/fly-session/src/media.rs index 870ead9..1d6c138 100644 --- a/services/flysim/crates/fly-session/src/media.rs +++ b/services/flysim/crates/fly-session/src/media.rs @@ -133,7 +133,14 @@ pub fn arena_frame(descriptor: &ViewDescriptor, counter: i64, boundary: u64) -> /// cannot be served an arbitrary stale image. pub struct ViewPipeline { descriptor: ViewDescriptor, - frames: VecDeque<(u64, flybus::Artifact)>, + /// Each retained frame: its producing boundary, the world counter it was rendered from + /// and the owned handle on its immutable bytes. + /// + /// The counter is kept because it is the whole of the reconstruction input: a checkpoint + /// records `(boundary, counter)` per retained frame and a restore re-renders them into + /// fresh artifacts of the current store, rather than persisting a transient artifact + /// identity that cannot survive a router restart. + frames: VecDeque<(u64, i64, flybus::Artifact)>, renders: RenderCounter, } @@ -160,7 +167,7 @@ impl ViewPipeline { let bytes = arena_frame(&self.descriptor, counter, boundary); let artifact = seal(client, FRAME_CONTENT_TYPE, &bytes).await?; self.renders.bump(); - self.frames.push_back((boundary, artifact)); + self.frames.push_back((boundary, counter, artifact)); // Keep exactly the frames a declared delay can still require. while self.frames.len() > self.descriptor.observation_delay_steps as usize + 1 { self.frames.pop_front(); @@ -168,6 +175,50 @@ impl ViewPipeline { Ok(()) } + /// The reconstruction inputs of every retained frame, oldest first. + /// + /// This is what a checkpoint records for the pending sensor pipeline: the producing + /// boundary and the world counter, never an artifact identity. + pub fn retained(&self) -> Vec<(u64, i64)> { + self.frames + .iter() + .map(|(boundary, counter, _)| (*boundary, *counter)) + .collect() + } + + /// Rebuilds the pipeline from recorded reconstruction inputs, into fresh artifacts. + /// + /// Every frame is rendered again in the current store, so nothing a fence dropped is + /// expected to come back and no old artifact identity crosses the recovery. + pub async fn restore( + &mut self, + client: &flybus::Client, + frames: &[(u64, i64)], + ) -> DomainResult<()> { + if frames.len() > self.descriptor.observation_delay_steps as usize + 1 { + return Err(media_error(format!( + "a captured pipeline of {} frames does not fit a declared delay of {}", + frames.len(), + self.descriptor.observation_delay_steps + ))); + } + for window in frames.windows(2) { + if window[1].0 != window[0].0 + 1 { + return Err(media_error( + "a captured pipeline's producing boundaries are not consecutive", + )); + } + } + self.frames.clear(); + for (boundary, counter) in frames { + let bytes = arena_frame(&self.descriptor, *counter, *boundary); + let artifact = seal(client, FRAME_CONTENT_TYPE, &bytes).await?; + self.renders.bump(); + self.frames.push_back((*boundary, *counter, artifact)); + } + Ok(()) + } + /// Seals a frame of the wrong length, which is what a broken backend produces. The /// reference it returns describes the artifact honestly, so the shape check is the thing /// under test rather than a lie in the payload. @@ -181,7 +232,7 @@ impl ViewPipeline { bytes.truncate(bytes.len() - self.descriptor.row_stride as usize); let artifact = seal(client, FRAME_CONTENT_TYPE, &bytes).await?; self.renders.bump(); - self.frames.push_back((boundary, artifact)); + self.frames.push_back((boundary, counter, artifact)); while self.frames.len() > self.descriptor.observation_delay_steps as usize + 2 { self.frames.pop_front(); } @@ -198,8 +249,8 @@ impl ViewPipeline { pub fn frame_produced_at(&self, produced: u64) -> Option<(ViewRef, flybus::Artifact)> { self.frames .iter() - .find(|(step, _)| *step == produced) - .map(|(step, artifact)| { + .find(|(step, _, _)| *step == produced) + .map(|(step, _, artifact)| { ( ViewRef { view_id: self.descriptor.view_id.clone(), @@ -257,6 +308,52 @@ impl AudioSource { source } + /// The exact state a capture recorded: sample position, waveform phase and the + /// unconsumed fraction of a frame. + /// + /// Restoring the position alone would restart the waveform and round the remainder away, + /// which is a resample the restore rules refuse. The first chunk of the new epoch marks + /// the discontinuity the recovery established. + pub fn restored_from( + descriptor: AudioDescriptor, + next_sample: u64, + phase: u64, + accumulator: u128, + denominator: u128, + ) -> DomainResult { + if denominator == 0 { + return Err(DomainError::invalid( + "audio: a captured accumulator denominator of zero", + )); + } + if accumulator >= denominator { + return Err(DomainError::invalid( + "audio: a captured accumulator is not below one whole frame", + )); + } + if phase >= descriptor.sample_rate { + return Err(DomainError::invalid( + "audio: a captured phase is not below the sample rate", + )); + } + let mut source = AudioSource::new(descriptor, next_sample); + source.discontinuous = true; + source.phase = phase; + source.accumulator = accumulator; + source.denominator = denominator; + Ok(source) + } + + /// The waveform phase, for a capture. + pub fn phase(&self) -> u64 { + self.phase + } + + /// The unconsumed fraction of a frame and the denominator it is over, for a capture. + pub fn accumulator(&self) -> (u128, u128) { + (self.accumulator, self.denominator) + } + pub fn descriptor(&self) -> &AudioDescriptor { &self.descriptor } @@ -439,18 +536,47 @@ pub fn check_required_views( Ok(()) } -/// Every declared audio stream produces exactly one chunk per transition. +/// Where an observation came from. +/// +/// `state-media-v1` section 2 makes a chunk the audio of an *interval*, so whether an +/// observation must carry one is a question about its provenance and not about its boundary +/// number. MEDIA-01 wrote the rule as "boundary 0 carries no chunk", which is true of the one +/// observation that slice could produce without a transition and false of the other one: +/// `State.ActivateRestore` installs a coherent observation at boundary `k` without advancing +/// gameplay, and it covers no interval either. Naming the provenance is the fix; exempting +/// the restored observation from the validator instead would have left "must a chunk exist" +/// unanswered exactly where a stale chunk would do the most damage. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub enum ObservationOrigin { + /// The observation a completed transition produced. Its interval has audio. + Transition, + /// An observation established at a boundary without running a transition: + /// `Environment.Initialize`'s `O[0]` and `State.ActivateRestore`'s restored observation. + /// It covers no interval, so it carries no chunk and one in it is refused. + Installed, +} + +/// Every declared audio stream produces exactly one chunk per transition, and none at all in +/// an observation that is not one. /// /// The contract states the shape and the ordering of chunks, not whether one has to exist, so /// this is MEDIA-01's choice and it is deliberate: a session that tolerates a silently missing /// chunk cannot tell "this world produced no audio for this interval" from "the chunk was -/// lost", and the second is the case the retention rules care about. Boundary 0 has no -/// preceding interval and so carries no chunk. +/// lost", and the second is the case the retention rules care about. The mirror of that, which +/// STATE-01 needs, is that an installed observation carrying a chunk is a stale chunk being +/// offered as current, and is refused for the same reason. pub fn check_required_audio( descriptor: &EnvironmentDescriptor, observation: &WorldObservation, + origin: ObservationOrigin, ) -> DomainResult<()> { - if observation.boundary == 0 { + if origin == ObservationOrigin::Installed { + if let Some(chunk) = observation.audio.first() { + return Err(media_error(format!( + "audio stream {} produced a chunk for an observation that ran no transition", + chunk.stream_id + ))); + } return Ok(()); } for stream in &descriptor.audio { diff --git a/services/flysim/crates/fly-session/src/state.rs b/services/flysim/crates/fly-session/src/state.rs new file mode 100644 index 0000000..93ba77b --- /dev/null +++ b/services/flysim/crates/fly-session/src/state.rs @@ -0,0 +1,1494 @@ +//! STATE-01: the coherent all-participant checkpoint store over the `FLYSESS1` envelope. +//! +//! [`fly_session_types::checkpoint`] owns the byte layout. This module owns everything +//! `checkpoint-envelope-v1` section 7 defers to this slice: generations, rotation, the store +//! manifest and its durable commit point, the bounded capture queue, the compatibility +//! comparison and the group fence. +//! +//! ```text +//! State.Capture ──> every participant, at one committed boundary +//! ──> one envelope: manifest + one payload per participant and per +//! coordinator-owned ledger +//! writer ──> temp, fsync, rename, fsync dir, then the store manifest the same way +//! ^^^^ the store manifest rename is the durable commit point +//! State.StageRestore ──> validated into replacement state, once-only token +//! State.ActivateRestore ──> installed under the new epoch, without a tick +//! ``` +//! +//! Nothing here is best-effort. A saturated queue is a named `BUSY` refusal taken *before* a +//! capture is requested; a lost save reply is an explicit outcome that leaves durable +//! metadata where it was; an incompatible checkpoint names the field that differs; and an +//! unreferenced generation is never a restore candidate. +//! +//! `FLYSIM01` (`crates/flybrain-core/src/envelope.rs`, read by the legacy flysim store) is a +//! different format with a different magic and a different reader, and nothing here touches +//! it. + +use std::collections::VecDeque; +use std::io::Write as _; +use std::path::{Path, PathBuf}; +use std::sync::Arc; +use std::sync::atomic::{AtomicBool, AtomicU64, Ordering}; + +use serde_json::{Map, Value, json}; + +use fly_session_types::checkpoint::{self, Envelope}; + +// `crate::types` is this crate's facade over the shared `fly-session-types` crate; the glob +// keeps the contract's own names in sight instead of restating them. +use crate::types::*; + +/// The state-format identity this slice writes and reads. A checkpoint recorded under any +/// other one is refused by name rather than attempted. +pub const STATE_FORMAT_ID: &str = "flysess-1"; + +/// The worker capability `state-media-v1` section 5 makes the State methods conditional on. +pub const CHECKPOINT_CAPABILITY: &str = "checkpoint-v1"; + +/// The store manifest's own version. It is not the envelope version. +pub const STORE_MANIFEST_VERSION: u32 = 1; + +/// The file the store manifest is committed to. Its rename is the durable commit point. +pub const STORE_MANIFEST_FILE: &str = "manifest.json"; + +/// The content type a checkpoint payload travels under as a bus artifact. +pub const PAYLOAD_CONTENT_TYPE: &str = "application/x-fly-checkpoint-payload"; + +/// The attachment name a checkpoint payload travels under, in both directions. +pub const PAYLOAD_ATTACHMENT: &str = "payload"; + +/// Seals one checkpoint payload as an immutable artifact against its content digest. +/// +/// `state-media-v1` section 1 makes a digest mandatory on checkpoint payloads, so this seals +/// with one rather than computing it afterwards: a store that wrote the wrong bytes finds out +/// here and not at the next restore. +pub async fn seal_payload( + client: &flybus::Client, + bytes: &[u8], + digest: &Digest, +) -> DomainResult { + let mut writer = client + .artifacts() + .allocate(bytes.len() as u64, PAYLOAD_CONTENT_TYPE) + .await + .map_err(|e| store_error(format!("payload allocate: {}", e.message)))?; + writer + .write_all(bytes) + .map_err(|e| store_error(format!("payload write: {e}")))?; + let artifact = writer + .seal_with_digest(Some(digest.clone())) + .await + .map_err(|e| store_error(format!("payload seal: {}", e.message)))?; + match &artifact.reference().digest { + Some(sealed) if sealed == digest => Ok(artifact), + _ => Err(store_error("a sealed checkpoint payload has no matching content digest")), + } +} + +fn store_error(what: impl std::fmt::Display) -> DomainError { + DomainError::new(ErrorCode::BackendFailure, what, MutationCertainty::Unknown) +} + +fn incompatible(what: impl std::fmt::Display) -> DomainError { + DomainError::before(ErrorCode::IncompatibleState, what) +} + +// ---------------------------------------------------------------------------------------------- +// Payload names + +/// The payload name one agent's captured state is filed under. +pub fn agent_payload(agent_id: &str) -> String { + format!("agent-{agent_id}") +} + +/// The payload name one agent's action-executor state is filed under. +pub fn executor_payload(agent_id: &str) -> String { + format!("executor-{agent_id}") +} + +/// The environment's payload name. +pub const WORLD_PAYLOAD: &str = "world"; +/// The task ledger's payload name. +pub const TASK_LEDGER_PAYLOAD: &str = "task-ledger"; +/// The prior world inspection's payload name. +pub const PRIOR_INSPECTION_PAYLOAD: &str = "prior-inspection"; +/// The coordinator's admission state and event watermarks. +pub const ADMISSION_PAYLOAD: &str = "coordinator-admission"; + +// ---------------------------------------------------------------------------------------------- +// Compatibility + +/// The compatibility identities `state-media-v1` section 4 requires a checkpoint to record. +/// +/// Each one is a separate field on purpose: a restore that fails says which identity differs, +/// instead of reporting one opaque digest mismatch. The synthetic composition maps them onto +/// the identities it actually has: +/// +/// | Field | Where it comes from | +/// | --- | --- | +/// | `backend_digest` | the environment descriptor's backend identity | +/// | `content_digest` | the environment descriptor's content identity | +/// | `patch_digest` | the environment descriptor's resolved configuration identity | +/// | `controller_digest` | every declared port's controller schema, in descriptor order | +/// | `parser_digest` | the inspection schema and the task schema that reads it | +/// | `state_format_id` | [`STATE_FORMAT_ID`] | +#[derive(Clone, Debug, PartialEq, Eq)] +pub struct Compatibility { + pub backend_digest: Digest, + pub content_digest: Digest, + pub patch_digest: Digest, + pub controller_digest: Digest, + pub parser_digest: Digest, + pub state_format_id: Id, +} + +impl Compatibility { + /// The compatibility of one world, from the descriptor it advertises. + /// + /// Every identity comes from the descriptor, which is what lets the environment compute + /// exactly this digest for its own capture while the coordinator computes it for the + /// composition. The task's own schema is deliberately not in here: the captured task + /// ledger carries its schema and refuses another one, so folding it in would put one + /// identity in two places. + pub fn of(descriptor: &EnvironmentDescriptor) -> Compatibility { + let controllers = Value::Array( + descriptor + .ports + .iter() + .map(|port| { + json!({ + "portId": port.port_id.as_str(), + "controls": port.controls.to_json(), + }) + }) + .collect(), + ); + let parser = json!({ + "inspectionSchema": descriptor.inspection_schema.to_json(), + }); + Compatibility { + backend_digest: descriptor.backend_digest.clone(), + content_digest: descriptor.content_digest.clone(), + patch_digest: descriptor.configuration_digest.clone(), + controller_digest: digest_of(&controllers) + .expect("a validated controller schema canonicalizes"), + parser_digest: digest_of(&parser).expect("a validated schema reference canonicalizes"), + state_format_id: id(STATE_FORMAT_ID), + } + } + + pub fn to_json(&self) -> Value { + json!({ + "backendDigest": self.backend_digest.as_str(), + "contentDigest": self.content_digest.as_str(), + "patchDigest": self.patch_digest.as_str(), + "controllerDigest": self.controller_digest.as_str(), + "parserDigest": self.parser_digest.as_str(), + "stateFormatId": self.state_format_id.as_str(), + }) + } + + pub fn from_json(value: &Value) -> Result { + let digest = |key: &str| -> Result { + let text = value + .get(key) + .and_then(Value::as_str) + .ok_or_else(|| format!("compatibility: {key} is missing or not a string"))?; + if !is_digest(text) { + return Err(format!("compatibility: {key} is not a digest")); + } + Ok(text.to_owned()) + }; + let state_format_id = value + .get("stateFormatId") + .and_then(Value::as_str) + .ok_or_else(|| "compatibility: stateFormatId is missing".to_owned())?; + Ok(Compatibility { + backend_digest: digest("backendDigest")?, + content_digest: digest("contentDigest")?, + patch_digest: digest("patchDigest")?, + controller_digest: digest("controllerDigest")?, + parser_digest: digest("parserDigest")?, + state_format_id: parse_id(state_format_id) + .map_err(|e| format!("compatibility: stateFormatId {e}"))?, + }) + } + + /// The single digest a participant echoes in its capture and its stage request. + pub fn digest(&self) -> Digest { + digest_of(&self.to_json()).expect("a compatibility block canonicalizes") + } + + /// Names the first identity that differs. There is no tolerance and no "close enough". + pub fn compare(&self, live: &Compatibility) -> Result<(), String> { + for (field, recorded, current) in [ + ("stateFormatId", &self.state_format_id, &live.state_format_id), + ("backend", &self.backend_digest, &live.backend_digest), + ("content", &self.content_digest, &live.content_digest), + ("patch", &self.patch_digest, &live.patch_digest), + ("controller", &self.controller_digest, &live.controller_digest), + ("parser", &self.parser_digest, &live.parser_digest), + ] { + if recorded != current { + return Err(format!( + "the checkpoint's {field} identity {recorded} is not this composition's {current}" + )); + } + } + Ok(()) + } +} + +// ---------------------------------------------------------------------------------------------- +// The store manifest + +/// One committed generation, as the store manifest lists it. +#[derive(Clone, Debug, PartialEq, Eq)] +pub struct GenerationRecord { + pub checkpoint_id: Id, + /// The generation's file name inside the store directory. + pub file: String, + pub session_id: Id, + pub epoch: Id, + pub episode_id: Id, + pub boundary: u64, + pub compatibility_digest: Digest, + /// The SHA-256 of the whole envelope file, so a store can compare without opening it. + pub envelope_digest: Digest, + pub byte_length: u64, +} + +impl GenerationRecord { + fn to_json(&self) -> Value { + json!({ + "checkpointId": self.checkpoint_id.as_str(), + "file": self.file.as_str(), + "sessionId": self.session_id.as_str(), + "epoch": self.epoch.as_str(), + "episodeId": self.episode_id.as_str(), + "boundary": self.boundary.to_string(), + "compatibilityDigest": self.compatibility_digest.as_str(), + "envelopeDigest": self.envelope_digest.as_str(), + "byteLength": self.byte_length.to_string(), + }) + } + + fn from_json(value: &Value) -> Result { + let text = |key: &str| -> Result { + value + .get(key) + .and_then(Value::as_str) + .map(str::to_owned) + .ok_or_else(|| format!("store manifest: a generation has no {key}")) + }; + let number = |key: &str| -> Result { + text(key)? + .parse::() + .map_err(|_| format!("store manifest: {key} is not a canonical U64")) + }; + Ok(GenerationRecord { + checkpoint_id: parse_id(&text("checkpointId")?)?, + file: text("file")?, + session_id: parse_id(&text("sessionId")?)?, + epoch: parse_id(&text("epoch")?)?, + episode_id: parse_id(&text("episodeId")?)?, + boundary: number("boundary")?, + compatibility_digest: text("compatibilityDigest")?, + envelope_digest: text("envelopeDigest")?, + byte_length: number("byteLength")?, + }) + } +} + +/// The store's durable metadata: which generations exist and how far durability has reached. +#[derive(Clone, Debug, Default, PartialEq, Eq)] +pub struct StoreManifest { + pub generations: Vec, + /// The newest committed checkpoint, which is the high-water mark. `None` until the first + /// durable commit; that is "nothing has been committed", not a default. + pub high_water: Option, +} + +impl StoreManifest { + fn to_json(&self) -> Value { + json!({ + "storeManifestVersion": STORE_MANIFEST_VERSION, + "stateFormatId": STATE_FORMAT_ID, + "highWater": self.high_water.as_ref().map_or(Value::Null, |h| h.as_str().into()), + "generations": Value::Array(self.generations.iter().map(GenerationRecord::to_json).collect()), + }) + } + + fn from_json(value: &Value) -> Result { + let version = value + .get("storeManifestVersion") + .and_then(Value::as_u64) + .ok_or_else(|| "store manifest: no storeManifestVersion".to_owned())?; + if version != u64::from(STORE_MANIFEST_VERSION) { + return Err(format!("store manifest: unsupported version {version}")); + } + match value.get("stateFormatId").and_then(Value::as_str) { + Some(STATE_FORMAT_ID) => {} + Some(other) => { + return Err(format!("store manifest: state format {other} is not {STATE_FORMAT_ID}")); + } + None => return Err("store manifest: no stateFormatId".to_owned()), + } + let generations = value + .get("generations") + .and_then(Value::as_array) + .ok_or_else(|| "store manifest: generations must be an array".to_owned())? + .iter() + .map(GenerationRecord::from_json) + .collect::, _>>()?; + let high_water = match value.get("highWater") { + Some(Value::Null) | None => None, + Some(Value::String(s)) => Some(parse_id(s)?), + Some(_) => return Err("store manifest: highWater is neither null nor an Id".to_owned()), + }; + if let Some(mark) = &high_water + && !generations.iter().any(|g| g.checkpoint_id == *mark) + { + return Err("store manifest: the high-water mark names no listed generation".to_owned()); + } + Ok(StoreManifest { generations, high_water }) + } +} + +// ---------------------------------------------------------------------------------------------- +// The store + +/// How many committed generations the store keeps. +#[derive(Clone, Copy, Debug)] +pub struct StoreConfig { + /// Generations retained after a commit. The oldest are dropped, and only once the + /// manifest that no longer references them is itself committed. + pub keep_generations: usize, +} + +impl Default for StoreConfig { + fn default() -> StoreConfig { + StoreConfig { keep_generations: 3 } + } +} + +/// Deliberate durable-write faults, for the failure rows this slice has to demonstrate. +#[derive(Clone, Debug, Default)] +pub struct StoreFaults { + /// Stop after the generation file has been renamed and before the store manifest is + /// committed. The generation is then an unreferenced file, which is never a candidate. + pub stop_before_manifest_commit: bool, +} + +/// The durable checkpoint store: generation files, one store manifest and the commit order of +/// `checkpoint-envelope-v1` section 5. +pub struct CheckpointStore { + root: PathBuf, + config: StoreConfig, + manifest: StoreManifest, + faults: StoreFaults, + commits: u64, + /// Generations the committed manifest no longer references and whose files could not be + /// removed. + /// + /// Rotation happens after the durable commit point, so a file that will not unlink is a + /// leaked file and never a lost checkpoint. It is recorded rather than swallowed, because + /// a store that keeps failing to rotate is filling a disk quietly. + unrotated: Vec, +} + +impl CheckpointStore { + /// Opens or creates a store at `root`, reading whatever it already committed. + /// + /// A directory with no manifest is an empty store: nothing has been committed there yet. + /// A manifest that cannot be read is a failure, not an empty store. + pub fn open(root: impl Into, config: StoreConfig) -> DomainResult { + let root: PathBuf = root.into(); + std::fs::create_dir_all(&root) + .map_err(|e| store_error(format!("checkpoint store {}: {e}", root.display())))?; + let path = root.join(STORE_MANIFEST_FILE); + let manifest = match std::fs::read(&path) { + Ok(bytes) => { + let value: Value = serde_json::from_slice(&bytes) + .map_err(|e| store_error(format!("store manifest: {e}")))?; + StoreManifest::from_json(&value).map_err(store_error)? + } + Err(e) if e.kind() == std::io::ErrorKind::NotFound => StoreManifest::default(), + Err(e) => return Err(store_error(format!("store manifest: {e}"))), + }; + Ok(CheckpointStore { + root, + config, + manifest, + faults: StoreFaults::default(), + commits: 0, + unrotated: Vec::new(), + }) + } + + pub fn root(&self) -> &Path { + &self.root + } + + /// Re-reads the store manifest from disk. + /// + /// The store manifest is the durable metadata, and this session is not necessarily the + /// only thing that has ever written it: a previous run, a repair or an operator may have + /// committed or removed a generation. Reading it again is how a store finds that out, + /// rather than trusting a copy it happens to be holding. + pub fn reload(&mut self) -> DomainResult<()> { + let reopened = CheckpointStore::open(self.root.clone(), self.config)?; + self.manifest = reopened.manifest; + Ok(()) + } + + pub fn manifest(&self) -> &StoreManifest { + &self.manifest + } + + pub fn faults_mut(&mut self) -> &mut StoreFaults { + &mut self.faults + } + + /// How many durable commits this store has completed. + pub fn commits(&self) -> u64 { + self.commits + } + + /// Generations the committed manifest dropped whose files are still on disk. + pub fn unrotated(&self) -> &[String] { + &self.unrotated + } + + /// The committed generation with this checkpoint id, if the manifest lists it. + /// + /// This is the query a coordinator uses to resolve a save whose reply it never saw: it + /// asks the durable metadata about the *same* operation rather than saving again. + pub fn lookup(&self, checkpoint_id: &Id) -> Option<&GenerationRecord> { + self.manifest + .generations + .iter() + .find(|g| g.checkpoint_id == *checkpoint_id) + } + + /// The newest committed generation, or an explicit refusal when nothing is committed. + pub fn high_water(&self) -> Option<&GenerationRecord> { + let mark = self.manifest.high_water.as_ref()?; + self.lookup(mark) + } + + /// Selects a restore candidate: the named generation, or the high-water one. + pub fn select(&self, checkpoint_id: Option<&Id>) -> DomainResult { + match checkpoint_id { + Some(wanted) => self.lookup(wanted).cloned().ok_or_else(|| { + incompatible(format!( + "the store has no committed generation {wanted}; an unreferenced \ +temporary is never a restore candidate" + )) + }), + None => self.high_water().cloned().ok_or_else(|| { + incompatible("the store has committed no checkpoint to restore from") + }), + } + } + + /// Reads one committed generation back and validates the whole envelope. + pub fn read(&self, record: &GenerationRecord) -> DomainResult { + let path = self.root.join(&record.file); + let bytes = std::fs::read(&path) + .map_err(|e| store_error(format!("generation {}: {e}", record.file)))?; + if bytes.len() as u64 != record.byte_length { + return Err(incompatible(format!( + "generation {} is {} bytes; the store manifest records {}", + record.file, + bytes.len(), + record.byte_length + ))); + } + if digest_of_bytes(&bytes) != record.envelope_digest { + return Err(incompatible(format!( + "generation {} does not match the digest the store manifest records", + record.file + ))); + } + let envelope = checkpoint::decode(&bytes) + .map_err(|e| incompatible(format!("generation {}: {}", record.file, e.0)))?; + checkpoint::validate_manifest(&envelope) + .map_err(|e| incompatible(format!("generation {}: {}", record.file, e.0)))?; + Ok(envelope) + } + + /// The durable commit sequence of `checkpoint-envelope-v1` section 5, in that order. + /// + /// Blocking by construction: it fsyncs. The writer runs it off the session's runtime. + fn commit(&mut self, record: GenerationRecord, bytes: &[u8]) -> DomainResult<()> { + let temporary = self.root.join(format!("tmp-{}.flysess", record.checkpoint_id)); + let final_path = self.root.join(&record.file); + // 1. write the envelope to a temporary generation file, 2. fsync it + { + let mut file = std::fs::File::create(&temporary) + .map_err(|e| store_error(format!("generation temporary: {e}")))?; + file.write_all(bytes) + .map_err(|e| store_error(format!("generation temporary: {e}")))?; + file.sync_all() + .map_err(|e| store_error(format!("generation fsync: {e}")))?; + } + // 3. rename it to its final generation name, 4. fsync the store directory + std::fs::rename(&temporary, &final_path) + .map_err(|e| store_error(format!("generation rename: {e}")))?; + sync_dir(&self.root)?; + if self.faults.stop_before_manifest_commit { + // The generation file exists and nothing references it. It is not a restore + // candidate and the high-water mark has not moved. + return Err(store_error( + "injected failure after the generation was renamed and before the store \ +manifest was committed", + )); + } + // 5. write the store manifest to its own temporary, fsync, rename, fsync the directory. + let mut next = self.manifest.clone(); + next.generations.retain(|g| g.checkpoint_id != record.checkpoint_id); + next.generations.push(record.clone()); + next.high_water = Some(record.checkpoint_id.clone()); + let dropped = if next.generations.len() > self.config.keep_generations { + let excess = next.generations.len() - self.config.keep_generations; + next.generations.drain(..excess).collect::>() + } else { + Vec::new() + }; + let text = canonicalize(&next.to_json()) + .map_err(|e| store_error(format!("store manifest: {}", e.0)))?; + let manifest_temporary = self.root.join("tmp-manifest.json"); + { + let mut file = std::fs::File::create(&manifest_temporary) + .map_err(|e| store_error(format!("store manifest temporary: {e}")))?; + file.write_all(text.as_bytes()) + .map_err(|e| store_error(format!("store manifest temporary: {e}")))?; + file.sync_all() + .map_err(|e| store_error(format!("store manifest fsync: {e}")))?; + } + std::fs::rename(&manifest_temporary, self.root.join(STORE_MANIFEST_FILE)) + .map_err(|e| store_error(format!("store manifest rename: {e}")))?; + sync_dir(&self.root)?; + // Past the durable commit point. Rotation removes only files the committed manifest + // no longer references. + self.manifest = next; + self.commits += 1; + for old in dropped { + if std::fs::remove_file(self.root.join(&old.file)).is_err() { + self.unrotated.push(old.file); + } + } + Ok(()) + } +} + +fn sync_dir(path: &Path) -> DomainResult<()> { + let dir = std::fs::File::open(path) + .map_err(|e| store_error(format!("store directory {}: {e}", path.display())))?; + dir.sync_all() + .map_err(|e| store_error(format!("store directory fsync: {e}"))) +} + +// ---------------------------------------------------------------------------------------------- +// The checkpoint manifest this slice writes and reads + +/// One agent's row in a checkpoint manifest. +#[derive(Clone, Debug, PartialEq)] +pub struct AgentEntry { + pub agent_id: Id, + pub profile_digest: Digest, + pub dataset_digest: Digest, + pub model_version: String, + pub plasticity_version: String, + pub seed: i32, + pub brain_ticks: u64, + pub remainder: RationalNs, + pub payload: String, +} + +impl AgentEntry { + fn to_json(&self) -> Value { + json!({ + "agentId": self.agent_id.as_str(), + "profileDigest": self.profile_digest.as_str(), + "datasetDigest": self.dataset_digest.as_str(), + "modelVersion": self.model_version.as_str(), + "plasticityVersion": self.plasticity_version.as_str(), + "seed": self.seed, + "brainTicks": self.brain_ticks.to_string(), + "remainder": self.remainder.to_json(), + "payload": self.payload.as_str(), + }) + } + + fn from_json(value: &Value) -> Result { + let text = |key: &str| -> Result { + value + .get(key) + .and_then(Value::as_str) + .map(str::to_owned) + .ok_or_else(|| format!("checkpoint manifest: an agent row has no {key}")) + }; + let seed = value + .get("seed") + .and_then(Value::as_i64) + .ok_or_else(|| "checkpoint manifest: an agent row has no seed".to_owned())?; + let seed = i32::try_from(seed) + .map_err(|_| "checkpoint manifest: a seed is outside i32".to_owned())?; + let remainder = RationalNs::from_json( + value + .get("remainder") + .ok_or_else(|| "checkpoint manifest: an agent row has no remainder".to_owned())?, + ) + .map_err(|e| format!("checkpoint manifest: remainder: {}", e.0))?; + Ok(AgentEntry { + agent_id: parse_id(&text("agentId")?)?, + profile_digest: text("profileDigest")?, + dataset_digest: text("datasetDigest")?, + model_version: text("modelVersion")?, + plasticity_version: text("plasticityVersion")?, + seed, + brain_ticks: text("brainTicks")? + .parse() + .map_err(|_| "checkpoint manifest: brainTicks is not a canonical U64".to_owned())?, + remainder, + payload: text("payload")?, + }) + } +} + +/// The coordinator-owned state a checkpoint records, each as a payload name. +#[derive(Clone, Debug, PartialEq, Eq)] +pub struct CoordinatorEntry { + pub task_ledger: String, + pub prior_inspection: String, + pub executor_state: Vec<(Id, String)>, + pub admission_state: String, + pub event_watermarks: EventWatermarks, +} + +/// The event identity a resumed epoch continues from. +#[derive(Clone, Debug, PartialEq, Eq)] +pub struct EventWatermarks { + /// The highest source step any recorded event belongs to. + pub last_source_step: u64, + /// How many events the ledger has issued. + pub issued: u64, +} + +impl EventWatermarks { + fn to_json(&self) -> Value { + json!({ + "lastSourceStep": self.last_source_step.to_string(), + "issued": self.issued.to_string(), + }) + } + + fn from_json(value: &Value) -> Result { + let number = |key: &str| -> Result { + value + .get(key) + .and_then(Value::as_str) + .ok_or_else(|| format!("checkpoint manifest: eventWatermarks has no {key}"))? + .parse() + .map_err(|_| format!("checkpoint manifest: {key} is not a canonical U64")) + }; + Ok(EventWatermarks { + last_source_step: number("lastSourceStep")?, + issued: number("issued")?, + }) + } +} + +impl CoordinatorEntry { + fn to_json(&self) -> Value { + json!({ + "taskLedger": self.task_ledger.as_str(), + "priorInspection": self.prior_inspection.as_str(), + "executorState": Value::Array( + self.executor_state + .iter() + .map(|(agent_id, payload)| json!({ + "agentId": agent_id.as_str(), + "payload": payload.as_str(), + })) + .collect(), + ), + "admissionState": self.admission_state.as_str(), + "eventWatermarks": self.event_watermarks.to_json(), + }) + } + + fn from_json(value: &Value) -> Result { + let text = |key: &str| -> Result { + value + .get(key) + .and_then(Value::as_str) + .map(str::to_owned) + .ok_or_else(|| format!("checkpoint manifest: coordinator has no {key}")) + }; + let executors = value + .get("executorState") + .and_then(Value::as_array) + .ok_or_else(|| "checkpoint manifest: executorState must be an array".to_owned())?; + let mut executor_state = Vec::with_capacity(executors.len()); + for entry in executors { + let agent_id = entry + .get("agentId") + .and_then(Value::as_str) + .ok_or_else(|| "checkpoint manifest: an executor row has no agentId".to_owned())?; + let payload = entry + .get("payload") + .and_then(Value::as_str) + .ok_or_else(|| "checkpoint manifest: an executor row has no payload".to_owned())?; + executor_state.push((parse_id(agent_id)?, payload.to_owned())); + } + let watermarks = value + .get("eventWatermarks") + .ok_or_else(|| "checkpoint manifest: coordinator has no eventWatermarks".to_owned())?; + Ok(CoordinatorEntry { + task_ledger: text("taskLedger")?, + prior_inspection: text("priorInspection")?, + executor_state, + admission_state: text("admissionState")?, + event_watermarks: EventWatermarks::from_json(watermarks)?, + }) + } +} + +/// The environment's row: which worker the world belonged to and which payload holds it. +/// +/// `checkpoint-envelope-v1` section 3 named a holder for every payload except the world's; +/// the 2026-09-22 amendment to that section adds this one. +#[derive(Clone, Debug, PartialEq, Eq)] +pub struct EnvironmentEntry { + pub worker_id: Id, + pub payload: String, +} + +impl EnvironmentEntry { + fn to_json(&self) -> Value { + json!({ + "workerId": self.worker_id.as_str(), + "payload": self.payload.as_str(), + }) + } + + fn from_json(value: &Value) -> Result { + let text = |key: &str| -> Result { + value + .get(key) + .and_then(Value::as_str) + .map(str::to_owned) + .ok_or_else(|| format!("checkpoint manifest: environment has no {key}")) + }; + Ok(EnvironmentEntry { + worker_id: parse_id(&text("workerId")?)?, + payload: text("payload")?, + }) + } +} + +/// A complete checkpoint manifest, in the field names `checkpoint-envelope-v1` section 3 sets. +#[derive(Clone, Debug, PartialEq)] +pub struct CheckpointManifest { + pub checkpoint_id: Id, + pub source_scope: Scope, + pub episode_id: Id, + pub world_time: RationalNs, + pub scheduler_id: String, + pub composition_digest: Digest, + pub port_map: Vec<(Id, Id)>, + pub compatibility: Compatibility, + pub agents: Vec, + pub coordinator: CoordinatorEntry, + pub environment: EnvironmentEntry, + /// External-helper state required for exact resume, as payload names. The synthetic + /// composition has no external helper, so it records an empty list rather than omitting + /// the field: "no helper" is a statement, not a missing one. + pub helper_state: Vec, + pub payloads: Vec<(String, u64, Digest)>, +} + +impl CheckpointManifest { + pub fn to_json(&self) -> Value { + json!({ + "envelopeVersion": checkpoint::VERSION, + "checkpointId": self.checkpoint_id.as_str(), + "sourceScope": self.source_scope.to_json(), + "episodeId": self.episode_id.as_str(), + "worldTime": self.world_time.to_json(), + "schedulerId": self.scheduler_id.as_str(), + "compositionDigest": self.composition_digest.as_str(), + "portMap": Value::Array( + self.port_map + .iter() + .map(|(port_id, agent_id)| json!({ + "portId": port_id.as_str(), + "agentId": agent_id.as_str(), + })) + .collect(), + ), + "compatibility": self.compatibility.to_json(), + "agents": Value::Array(self.agents.iter().map(AgentEntry::to_json).collect()), + "coordinator": self.coordinator.to_json(), + "environment": self.environment.to_json(), + "helperState": Value::Array( + self.helper_state.iter().map(|n| Value::String(n.clone())).collect(), + ), + "payloads": Value::Array( + self.payloads + .iter() + .map(|(name, length, digest)| json!({ + "name": name.as_str(), + "byteLength": length.to_string(), + "digest": digest.as_str(), + })) + .collect(), + ), + }) + } + + /// Reads one back, refusing an incomplete manifest rather than filling anything in. + pub fn from_json(value: &Value) -> Result { + for field in checkpoint::REQUIRED_MANIFEST_FIELDS { + if value.get(*field).is_none() { + return Err(format!("checkpoint manifest: missing {field:?}")); + } + } + let text = |key: &str| -> Result { + value + .get(key) + .and_then(Value::as_str) + .map(str::to_owned) + .ok_or_else(|| format!("checkpoint manifest: {key} is missing or not a string")) + }; + let source_scope = Scope::from_json(&value["sourceScope"]) + .map_err(|e| format!("checkpoint manifest: sourceScope: {}", e.0))?; + let world_time = RationalNs::from_json(&value["worldTime"]) + .map_err(|e| format!("checkpoint manifest: worldTime: {}", e.0))?; + let mut port_map = Vec::new(); + for entry in value["portMap"] + .as_array() + .ok_or_else(|| "checkpoint manifest: portMap must be an array".to_owned())? + { + let port_id = entry + .get("portId") + .and_then(Value::as_str) + .ok_or_else(|| "checkpoint manifest: a port map row has no portId".to_owned())?; + let agent_id = entry + .get("agentId") + .and_then(Value::as_str) + .ok_or_else(|| "checkpoint manifest: a port map row has no agentId".to_owned())?; + port_map.push((parse_id(port_id)?, parse_id(agent_id)?)); + } + let agents = value["agents"] + .as_array() + .ok_or_else(|| "checkpoint manifest: agents must be an array".to_owned())? + .iter() + .map(AgentEntry::from_json) + .collect::, _>>()?; + let mut helper_state = Vec::new(); + for entry in value["helperState"] + .as_array() + .ok_or_else(|| "checkpoint manifest: helperState must be an array".to_owned())? + { + helper_state.push( + entry + .as_str() + .ok_or_else(|| "checkpoint manifest: a helper state entry is not a payload name".to_owned())? + .to_owned(), + ); + } + let mut payloads = Vec::new(); + for entry in value["payloads"] + .as_array() + .ok_or_else(|| "checkpoint manifest: payloads must be an array".to_owned())? + { + let name = entry + .get("name") + .and_then(Value::as_str) + .ok_or_else(|| "checkpoint manifest: a payload row has no name".to_owned())?; + let length: u64 = entry + .get("byteLength") + .and_then(Value::as_str) + .ok_or_else(|| "checkpoint manifest: a payload row has no byteLength".to_owned())? + .parse() + .map_err(|_| "checkpoint manifest: byteLength is not a canonical U64".to_owned())?; + let digest = entry + .get("digest") + .and_then(Value::as_str) + .ok_or_else(|| "checkpoint manifest: a payload row has no digest".to_owned())?; + payloads.push((name.to_owned(), length, digest.to_owned())); + } + Ok(CheckpointManifest { + checkpoint_id: parse_id(&text("checkpointId")?)?, + source_scope, + episode_id: parse_id(&text("episodeId")?)?, + world_time, + scheduler_id: text("schedulerId")?, + composition_digest: text("compositionDigest")?, + port_map, + compatibility: Compatibility::from_json(&value["compatibility"])?, + agents, + coordinator: CoordinatorEntry::from_json(&value["coordinator"])?, + environment: EnvironmentEntry::from_json( + value + .get("environment") + .ok_or_else(|| "checkpoint manifest: missing \"environment\"".to_owned())?, + )?, + helper_state, + payloads, + }) + } + + /// The payload name this manifest files one participant's state under. + pub fn payload_of(&self, worker_id: &Id) -> Option<&str> { + if self.environment.worker_id == *worker_id { + return Some(self.environment.payload.as_str()); + } + self.agents + .iter() + .find(|a| a.agent_id == *worker_id) + .map(|a| a.payload.as_str()) + } +} + +// ---------------------------------------------------------------------------------------------- +// The bounded writer + +/// One participant's captured payload, with the owned handle the writer keeps until the bytes +/// are committed or the job fails. +pub struct CapturedPayload { + pub name: String, + pub artifact: flybus::Artifact, + pub byte_length: u64, + pub digest: Digest, +} + +/// What a coordinator hands the writer once every participant has captured. +pub struct CaptureSubmission { + pub checkpoint_id: Id, + pub boundary: u64, + pub session_id: Id, + pub epoch: Id, + pub episode_id: Id, + pub compatibility_digest: Digest, + pub manifest: Value, + pub payloads: Vec, + /// A hot checkpoint may replace a queued hot checkpoint, releasing its holds. A durable + /// one never is: the retention table coalesces only queued replaceable captures. + pub replaceable: bool, +} + +impl CaptureSubmission { + fn byte_length(&self) -> u64 { + self.payloads.iter().map(|p| p.byte_length).sum() + } +} + +/// How one durable save ended. Every variant is a statement; none of them is a default. +#[derive(Clone, Debug, PartialEq, Eq)] +pub enum SaveOutcome { + /// Past the store manifest rename. This is the only variant that is a saved + /// acknowledgment and the only one that moves a high-water mark. + Committed { checkpoint_id: Id, boundary: u64, file: String }, + /// The write failed and its owned captures were released under the retry policy. It + /// never reports false durability. + Failed { checkpoint_id: Id, reason: String }, + /// A later replaceable capture took this one's place in the queue before it was written. + Superseded { checkpoint_id: Id, by: Id }, + /// The writer's reply never arrived. The operation's outcome is unknown from here, so + /// durable metadata does not move; the caller resolves the *same* operation against the + /// store manifest instead of saving again. + ReplyLost { checkpoint_id: Id }, +} + +impl SaveOutcome { + pub fn checkpoint_id(&self) -> &Id { + match self { + SaveOutcome::Committed { checkpoint_id, .. } + | SaveOutcome::Failed { checkpoint_id, .. } + | SaveOutcome::Superseded { checkpoint_id, .. } + | SaveOutcome::ReplyLost { checkpoint_id } => checkpoint_id, + } + } + + /// The event name this outcome publishes under. + pub fn event(&self) -> &'static str { + match self { + SaveOutcome::Committed { .. } => "committed", + SaveOutcome::Failed { .. } => "failed", + SaveOutcome::Superseded { .. } => "superseded", + SaveOutcome::ReplyLost { .. } => "failed", + } + } +} + +/// What a failed write does with the ephemeral captures it owns. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub enum RetryPolicy { + /// Release the owned captures and report the failure. The default: a capture is cheap to + /// take again at the next committed boundary, and holding one is not free. + ReleaseAndReport, + /// Try the durable sequence again, up to `attempts` times in total, keeping the owned + /// captures until they are committed or the attempts are spent. + RetryThenRelease { attempts: u32 }, +} + +/// The writer's bounds. Both are finite and both refuse before a capture is requested. +#[derive(Clone, Copy, Debug)] +pub struct WriterConfig { + /// Outstanding coherent captures. `state-media-v1` section 3's initial session default + /// is two. + pub queue_capacity: usize, + /// The total payload bytes the queue may hold. + pub max_queued_bytes: u64, + pub retry: RetryPolicy, +} + +impl Default for WriterConfig { + fn default() -> WriterConfig { + WriterConfig { + queue_capacity: 2, + max_queued_bytes: 64 * 1024 * 1024, + retry: RetryPolicy::ReleaseAndReport, + } + } +} + +/// Deliberate writer faults, for the rows this slice has to demonstrate. +#[derive(Clone, Default)] +pub struct WriterFaults { + /// Hold every job until the gate is opened, so a test can fill the queue on purpose. + pub gate: Option>, + /// Drop this job's reply channel after the durable sequence ran, which is a lost save + /// reply. + pub drop_reply_for: Option, +} + +impl std::fmt::Debug for WriterFaults { + fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { + f.debug_struct("WriterFaults") + .field("gate", &self.gate.is_some()) + .field("drop_reply_for", &self.drop_reply_for) + .finish() + } +} + +/// The writer's own counters, so "bounded" is something a test reads rather than believes. +#[derive(Clone, Copy, Debug, Default, PartialEq, Eq)] +pub struct WriterStats { + pub queued: u64, + pub committed: u64, + pub failed: u64, + pub superseded: u64, + /// Submissions refused because the queue or its byte budget was full. + pub rejected: u64, + /// Checkpoint events the bus would not take. The durable outcome the caller receives is + /// the authority; a publication that failed is counted here rather than disappearing. + pub events_dropped: u64, + /// The deepest the queue ever got. + pub peak_queue: usize, + /// The most payload bytes the queue ever held. + pub peak_bytes: u64, +} + +struct Job { + submission: CaptureSubmission, + reply: tokio::sync::oneshot::Sender, + permit: tokio::sync::OwnedSemaphorePermit, +} + +struct WriterShared { + config: WriterConfig, + faults: WriterFaults, + permits: Arc, + queue: std::sync::Mutex>, + queued_bytes: AtomicU64, + stats: std::sync::Mutex, + wake: tokio::sync::Notify, + stop: AtomicBool, + store: Arc>, + events: Option<(flybus::Client, String)>, +} + +/// A queue slot, taken before a capture is requested so a saturated writer refuses early. +/// +/// `state-media-v1` section 3's durable row says to reject or defer *before capture* when +/// saturated. Holding the slot from before `State.Capture` until the outcome is delivered is +/// what makes that true rather than aspirational. +pub struct Reservation { + permit: tokio::sync::OwnedSemaphorePermit, +} + +/// The bounded checkpoint writer. +pub struct CheckpointWriter { + shared: Arc, + task: tokio::task::JoinHandle<()>, +} + +impl CheckpointWriter { + /// Starts the writer over `store`. `events` is the bus client and topic the distinct + /// captured/queued/committed/failed/superseded publications go to. + pub fn start( + store: CheckpointStore, + config: WriterConfig, + faults: WriterFaults, + events: Option<(flybus::Client, String)>, + ) -> CheckpointWriter { + let shared = Arc::new(WriterShared { + config, + faults, + permits: Arc::new(tokio::sync::Semaphore::new(config.queue_capacity)), + queue: std::sync::Mutex::new(VecDeque::new()), + queued_bytes: AtomicU64::new(0), + stats: std::sync::Mutex::new(WriterStats::default()), + wake: tokio::sync::Notify::new(), + stop: AtomicBool::new(false), + store: Arc::new(std::sync::Mutex::new(store)), + events, + }); + let task = tokio::spawn(run_writer(shared.clone())); + CheckpointWriter { shared, task } + } + + pub fn stats(&self) -> WriterStats { + *self.shared.stats.lock().expect("the writer stats are never poisoned") + } + + pub fn config(&self) -> WriterConfig { + self.shared.config + } + + /// How many captures are outstanding: queued or being written. + pub fn outstanding(&self) -> usize { + self.shared.config.queue_capacity - self.shared.permits.available_permits() + } + + /// Takes a queue slot, or refuses by name. Nothing is captured without one. + pub fn reserve(&self) -> DomainResult { + match self.shared.permits.clone().try_acquire_owned() { + Ok(permit) => Ok(Reservation { permit }), + Err(_) => { + self.shared + .stats + .lock() + .expect("the writer stats are never poisoned") + .rejected += 1; + Err(DomainError::before( + ErrorCode::Busy, + format!( + "the checkpoint queue already holds its {} outstanding captures; a \ +capture is refused before it is requested rather than queued without bound", + self.shared.config.queue_capacity + ), + )) + } + } + } + + /// Hands one complete capture to the writer and returns the channel its outcome arrives + /// on. The reservation becomes the job's slot. + pub async fn submit( + &self, + reservation: Reservation, + submission: CaptureSubmission, + ) -> DomainResult> { + let bytes = submission.byte_length(); + let budget = self.shared.config.max_queued_bytes; + let held = self.shared.queued_bytes.load(Ordering::SeqCst); + if held + bytes > budget { + self.shared + .stats + .lock() + .expect("the writer stats are never poisoned") + .rejected += 1; + return Err(DomainError::before( + ErrorCode::Busy, + format!( + "the checkpoint queue holds {held} of {budget} bytes and this capture adds \ +{bytes}; the byte budget is finite and refuses before it is exceeded" + ), + )); + } + let (reply, receiver) = tokio::sync::oneshot::channel(); + let checkpoint_id = submission.checkpoint_id.clone(); + let queued_event = submission.as_event("queued"); + let superseded = { + let mut queue = self.shared.queue.lock().expect("the writer queue is never poisoned"); + let replaced = if submission.replaceable { + queue + .iter() + .position(|job| job.submission.replaceable) + .map(|index| queue.remove(index).expect("just found")) + } else { + None + }; + queue.push_back(Job { submission, reply, permit: reservation.permit }); + self.shared.queued_bytes.fetch_add(bytes, Ordering::SeqCst); + let mut stats = self.shared.stats.lock().expect("the writer stats are never poisoned"); + stats.queued += 1; + stats.peak_queue = stats.peak_queue.max(queue.len()); + stats.peak_bytes = stats + .peak_bytes + .max(self.shared.queued_bytes.load(Ordering::SeqCst)); + replaced + }; + if let Some(old) = superseded { + self.shared + .queued_bytes + .fetch_sub(old.submission.byte_length(), Ordering::SeqCst); + self.shared + .stats + .lock() + .expect("the writer stats are never poisoned") + .superseded += 1; + let outcome = SaveOutcome::Superseded { + checkpoint_id: old.submission.checkpoint_id.clone(), + by: checkpoint_id.clone(), + }; + publish_event( + &self.shared.events, + &self.shared.stats, + &outcome_event(&old.submission, &outcome), + ) + .await; + // A caller that dropped its ticket is not waiting for this; the store's own + // metadata is the durable record either way. + let _ = old.reply.send(outcome); + // Dropping the job releases its owned captures and its queue slot. + drop(old.submission); + drop(old.permit); + } + publish_event(&self.shared.events, &self.shared.stats, &object(queued_event)).await; + self.shared.wake.notify_one(); + Ok(receiver) + } + + /// Waits for one save's outcome. A dropped reply channel is a lost save reply, which is + /// an outcome and not a hang. + pub async fn wait( + receiver: tokio::sync::oneshot::Receiver, + checkpoint_id: &Id, + budget: std::time::Duration, + ) -> SaveOutcome { + match tokio::time::timeout(budget, receiver).await { + Ok(Ok(outcome)) => outcome, + Ok(Err(_)) | Err(_) => SaveOutcome::ReplyLost { checkpoint_id: checkpoint_id.clone() }, + } + } + + /// Runs `f` against the store, which is how a caller resolves a save whose reply it lost. + /// + /// The store's own work is blocking -- it fsyncs -- so it is reached on a blocking thread + /// rather than from the session's runtime. + pub async fn with_store(&self, f: F) -> T + where + F: FnOnce(&mut CheckpointStore) -> T + Send + 'static, + T: Send + 'static, + { + let store = self.shared.store.clone(); + tokio::task::spawn_blocking(move || { + let mut store = store.lock().expect("the checkpoint store is never poisoned"); + f(&mut store) + }) + .await + .expect("the checkpoint store task is never cancelled") + } + + /// Stops the writer and waits for its task. A leftover writer task would hold artifact + /// handles the session has finished with. + pub async fn shutdown(self) { + self.shared.stop.store(true, Ordering::SeqCst); + self.shared.wake.notify_one(); + let _ = self.task.await; + let mut queue = self.shared.queue.lock().expect("the writer queue is never poisoned"); + queue.clear(); + } +} + +impl CaptureSubmission { + fn as_event(&self, event: &'static str) -> Value { + json!({ + "event": event, + "checkpointId": self.checkpoint_id.as_str(), + "sessionId": self.session_id.as_str(), + "epoch": self.epoch.as_str(), + "episodeId": self.episode_id.as_str(), + "boundary": self.boundary.to_string(), + "byteLength": self.byte_length().to_string(), + }) + } +} + +fn outcome_event(submission: &CaptureSubmission, outcome: &SaveOutcome) -> Map { + let mut payload = object(submission.as_event(outcome.event())); + match outcome { + SaveOutcome::Committed { file, .. } => { + payload.insert("generation".into(), file.as_str().into()); + payload.insert("durable".into(), true.into()); + } + SaveOutcome::Failed { reason, .. } => { + payload.insert("reason".into(), reason.as_str().into()); + payload.insert("durable".into(), false.into()); + } + SaveOutcome::Superseded { by, .. } => { + payload.insert("supersededBy".into(), by.as_str().into()); + payload.insert("durable".into(), false.into()); + } + SaveOutcome::ReplyLost { .. } => { + payload.insert("reason".into(), "the save reply was lost".into()); + payload.insert("durable".into(), false.into()); + } + } + payload +} + +async fn publish_event( + events: &Option<(flybus::Client, String)>, + stats: &std::sync::Mutex, + payload: &Map, +) { + let Some((client, topic)) = events else { return }; + // A checkpoint event is telemetry about the store, not the durable record. A publication + // that cannot be delivered never changes what was committed, so the outcome the caller + // receives stays the authority -- and the drop is counted rather than retried, because a + // retry here would be exactly the implicit best-effort policy these contracts refuse. + if client.publish(topic, payload.clone(), &[]).await.is_err() { + stats + .lock() + .expect("the writer stats are never poisoned") + .events_dropped += 1; + } +} + +async fn run_writer(shared: Arc) { + loop { + let job = { + let mut queue = shared.queue.lock().expect("the writer queue is never poisoned"); + queue.pop_front() + }; + let Some(job) = job else { + if shared.stop.load(Ordering::SeqCst) { + return; + } + shared.wake.notified().await; + continue; + }; + if let Some(gate) = &shared.faults.gate { + // A deliberately stalled writer. The queue in front of it stays bounded, which is + // the point of the stall. + if let Ok(permit) = gate.acquire().await { + permit.forget(); + } + } + let Job { submission, reply, permit } = job; + shared + .queued_bytes + .fetch_sub(submission.byte_length(), Ordering::SeqCst); + let outcome = write_one(&shared, &submission).await; + { + let mut stats = shared.stats.lock().expect("the writer stats are never poisoned"); + match &outcome { + SaveOutcome::Committed { .. } => stats.committed += 1, + SaveOutcome::Failed { .. } | SaveOutcome::ReplyLost { .. } => stats.failed += 1, + SaveOutcome::Superseded { .. } => stats.superseded += 1, + } + } + publish_event(&shared.events, &shared.stats, &outcome_event(&submission, &outcome)).await; + let lost = shared.faults.drop_reply_for.as_ref() == Some(&submission.checkpoint_id); + // The writer owned these handles until the bytes were committed or the job failed. + // The job is over, so they and its queue slot go before the outcome is delivered: + // a caller that reads the queue depth the moment its outcome arrives must not see a + // slot this job has finished with. + drop(submission); + drop(permit); + if lost { + // The bytes are committed and the acknowledgment is lost. Durable metadata is in + // the store manifest, which is exactly where the caller must look. + drop(reply); + } else { + // A caller that dropped its ticket is not waiting for this. + let _ = reply.send(outcome); + } + } +} + +async fn write_one(shared: &Arc, submission: &CaptureSubmission) -> SaveOutcome { + let mut payloads = Vec::with_capacity(submission.payloads.len()); + for payload in &submission.payloads { + let bytes = match payload.artifact.read_all().await { + Ok(bytes) => bytes, + Err(e) => { + return SaveOutcome::Failed { + checkpoint_id: submission.checkpoint_id.clone(), + reason: format!("payload {}: {}", payload.name, e.message), + }; + } + }; + if bytes.len() as u64 != payload.byte_length || digest_of_bytes(&bytes) != payload.digest { + return SaveOutcome::Failed { + checkpoint_id: submission.checkpoint_id.clone(), + reason: format!( + "payload {} is not the content its capture declared", + payload.name + ), + }; + } + payloads.push((payload.name.clone(), bytes)); + } + let bytes = match checkpoint::encode(&submission.manifest, &payloads) { + Ok(bytes) => bytes, + Err(e) => { + return SaveOutcome::Failed { + checkpoint_id: submission.checkpoint_id.clone(), + reason: format!("envelope: {}", e.0), + }; + } + }; + let record = GenerationRecord { + checkpoint_id: submission.checkpoint_id.clone(), + file: format!("{}.flysess", submission.checkpoint_id), + session_id: submission.session_id.clone(), + epoch: submission.epoch.clone(), + episode_id: submission.episode_id.clone(), + boundary: submission.boundary, + compatibility_digest: submission.compatibility_digest.clone(), + envelope_digest: digest_of_bytes(&bytes), + byte_length: bytes.len() as u64, + }; + let attempts = match shared.config.retry { + RetryPolicy::ReleaseAndReport => 1, + RetryPolicy::RetryThenRelease { attempts } => attempts.max(1), + }; + let mut last = String::new(); + for _ in 0..attempts { + let record = record.clone(); + let bytes = bytes.clone(); + let store = shared.store.clone(); + // The durable sequence fsyncs twice. It runs on a blocking thread so the session's + // runtime is never the thing waiting on a disk. + let result = tokio::task::spawn_blocking(move || { + let mut store = store.lock().expect("the checkpoint store is never poisoned"); + store.commit(record, &bytes) + }) + .await; + match result { + Ok(Ok(())) => { + return SaveOutcome::Committed { + checkpoint_id: submission.checkpoint_id.clone(), + boundary: submission.boundary, + file: format!("{}.flysess", submission.checkpoint_id), + }; + } + Ok(Err(e)) => last = e.message, + Err(e) => last = format!("the checkpoint writer stopped: {e}"), + } + } + SaveOutcome::Failed { + checkpoint_id: submission.checkpoint_id.clone(), + reason: last, + } +} diff --git a/services/flysim/crates/fly-session/src/task.rs b/services/flysim/crates/fly-session/src/task.rs index b48e909..2eebbef 100644 --- a/services/flysim/crates/fly-session/src/task.rs +++ b/services/flysim/crates/fly-session/src/task.rs @@ -39,6 +39,16 @@ pub fn episode_schema() -> SchemaRef { synthetic_schema("arena.episode.v1", 1) } +/// The schema of a captured task ledger. +pub fn ledger_schema() -> SchemaRef { + synthetic_schema("arena.ledger.v1", 1) +} + +/// The schema of a captured action-executor state. +pub fn executor_schema() -> SchemaRef { + synthetic_schema("arena.executor.v1", 1) +} + pub fn controller_schema_ref() -> SchemaRef { synthetic_schema("arena.controller.v1", 1) } @@ -86,6 +96,29 @@ pub trait Task: Send { /// How many times `evaluate_transition` has run. A transition must evaluate once. fn evaluations(&self) -> u64; + + /// The checkpointable ledger at a committed boundary (`workers-v1` section 4). + fn capture(&self) -> DomainResult; + + /// Validates a captured ledger without installing it, so a group install can fail before + /// anything is changed. + fn validate_restore(&self, state: &TypedValue) -> DomainResult<()>; + + /// Installs a validated ledger under `epoch`. Event identity is derived from the epoch, + /// so the new one is part of the install rather than something the ledger keeps from the + /// epoch it was captured in. + fn install_restore(&mut self, epoch: &Id, state: &TypedValue) -> DomainResult<()>; + + /// Every event identity this ledger has issued, mapped onto the identity it would have + /// under `to_epoch`. + /// + /// `workers-v1` section 4 derives an event id from the epoch, so a trace recorded in one + /// epoch cannot be compared with a trace recorded in another until these are rebased. + /// The ledger owns the derivation, so it is the only thing that can do it. + fn rebase_ids(&self, to_epoch: &Id) -> DomainResult>; + + /// How far event identity has reached: the highest source step and the number issued. + fn event_watermarks(&self) -> (u64, u64); } /// Translates one selected decision into a controller intent, with no port assignment. @@ -98,6 +131,15 @@ pub trait ActionExecutor: Send { progress: &TypedValue, clock: &RationalNs, ) -> DomainResult<(ControllerIntent, Vec)>; + + /// Per-executor state at a committed boundary (`workers-v1` section 4). + fn capture(&self) -> DomainResult; + + /// Validates a captured executor state without installing it. + fn validate_restore(&self, state: &TypedValue) -> DomainResult<()>; + + /// Installs a validated executor state. + fn install_restore(&mut self, state: &TypedValue) -> DomainResult<()>; } /// The only executor v1 supports: it passes a direct-control decision through unchanged. @@ -123,6 +165,37 @@ impl ActionExecutor for IdentityExecutor { .map_err(|e| DomainError::invalid(format!("decision: {e}")))?; Ok((intent, Vec::new())) } + + /// The identity executor is stateless, and says so rather than capturing nothing. + /// + /// An empty object would be indistinguishable from a stateful executor whose capture went + /// missing, so the capture names the executor it came from and a restore refuses any + /// other one. + fn capture(&self) -> DomainResult { + TypedValue::new(executor_schema(), json!({"executor": "identity-v1"})) + .map_err(|e| DomainError::invalid(e.0)) + } + + fn validate_restore(&self, state: &TypedValue) -> DomainResult<()> { + if state.schema != executor_schema() { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + "the captured executor state does not carry the executor schema", + )); + } + match state.value.get("executor").and_then(Value::as_str) { + Some("identity-v1") => Ok(()), + other => Err(DomainError::before( + ErrorCode::IncompatibleState, + format!("the captured executor is {other:?}, not the identity executor"), + )), + } + } + + fn install_restore(&mut self, state: &TypedValue) -> DomainResult<()> { + // Stateless: validation is the whole of the install, and it is not skipped. + self.validate_restore(state) + } } /// When the counter task asks for a terminal episode transition. @@ -146,6 +219,10 @@ pub struct CounterTask { evaluations: u64, total_reward: f64, counter: i64, + /// The highest source step any issued event belongs to, and how many were issued. These + /// are the event watermarks a checkpoint records and a resumed epoch continues from. + last_source_step: u64, + issued_events: u64, terminal: Terminal, } @@ -159,6 +236,8 @@ impl CounterTask { evaluations: 0, total_reward: 0.0, counter: 0, + last_source_step: 0, + issued_events: 0, terminal, } } @@ -239,6 +318,7 @@ impl Task for CounterTask { payload: TypedValue::new(event_schema(), json!({"counter": self.counter})) .expect("a synthetic typed value fits the contract"), }]; + self.issued_events += events.len() as u64; Ok(Bootstrap { contexts, progress: self.progress_value(), events }) } @@ -314,6 +394,8 @@ impl Task for CounterTask { )); } + self.last_source_step = self.last_source_step.max(source_step); + self.issued_events += events.len() as u64; let next_contexts = self .agents .iter() @@ -346,6 +428,173 @@ impl Task for CounterTask { fn evaluations(&self) -> u64 { self.evaluations } + + fn capture(&self) -> DomainResult { + TypedValue::new( + ledger_schema(), + json!({ + "epoch": self.epoch.as_str(), + "agents": self.agents.iter().map(String::as_str).collect::>(), + "bindings": self + .bindings + .iter() + .map(|b| json!({"portId": b.port_id.as_str(), "agentId": b.agent_id.as_str()})) + .collect::>(), + "transitions": self.transitions, + "evaluations": self.evaluations, + "totalReward": self.total_reward, + "counter": self.counter, + "lastSourceStep": self.last_source_step, + "issuedEvents": self.issued_events, + }), + ) + .map_err(|e| DomainError::invalid(e.0)) + } + + fn validate_restore(&self, state: &TypedValue) -> DomainResult<()> { + if state.schema != ledger_schema() { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + "the captured ledger does not carry this task's schema", + )); + } + for field in [ + "epoch", + "agents", + "bindings", + "transitions", + "evaluations", + "totalReward", + "counter", + "lastSourceStep", + "issuedEvents", + ] { + if state.value.get(field).is_none() { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + format!("the captured ledger has no {field}"), + )); + } + } + let bindings = state + .value + .get("bindings") + .and_then(Value::as_array) + .ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "the captured ledger's bindings are not a list", + ) + })?; + if bindings.len() != self.bindings.len() && !self.bindings.is_empty() { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + "the captured ledger binds another number of ports", + )); + } + Ok(()) + } + + fn install_restore(&mut self, epoch: &Id, state: &TypedValue) -> DomainResult<()> { + self.validate_restore(state)?; + let number = |key: &str| -> DomainResult { + state.value.get(key).and_then(Value::as_u64).ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + format!("the captured ledger's {key} is not a whole number"), + ) + }) + }; + let mut agents = Vec::new(); + for value in state.value["agents"].as_array().expect("validated") { + let agent = value.as_str().ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "the captured ledger names an agent that is not a string", + ) + })?; + agents.push(parse_id(agent).map_err(|e| { + DomainError::before(ErrorCode::IncompatibleState, format!("ledger: {e}")) + })?); + } + let mut bindings = Vec::new(); + for value in state.value["bindings"].as_array().expect("validated") { + let port_id = value.get("portId").and_then(Value::as_str).ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "the captured ledger has a binding with no portId", + ) + })?; + let agent_id = value.get("agentId").and_then(Value::as_str).ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "the captured ledger has a binding with no agentId", + ) + })?; + bindings.push(PortBinding { + port_id: parse_id(port_id).map_err(|e| { + DomainError::before(ErrorCode::IncompatibleState, format!("ledger: {e}")) + })?, + agent_id: parse_id(agent_id).map_err(|e| { + DomainError::before(ErrorCode::IncompatibleState, format!("ledger: {e}")) + })?, + }); + } + let counter = state.value.get("counter").and_then(Value::as_i64).ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "the captured ledger's counter is not an integer", + ) + })?; + let total_reward = state + .value + .get("totalReward") + .and_then(Value::as_f64) + .filter(|v| v.is_finite()) + .ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "the captured ledger's totalReward is not a finite number", + ) + })?; + // The epoch is the caller's, not the capture's: event identity belongs to the epoch + // the ledger is being installed into. + self.epoch = epoch.clone(); + self.agents = agents; + self.bindings = bindings; + self.transitions = number("transitions")?; + self.evaluations = number("evaluations")?; + self.total_reward = total_reward; + self.counter = counter; + self.last_source_step = number("lastSourceStep")?; + self.issued_events = number("issuedEvents")?; + Ok(()) + } + + fn rebase_ids(&self, to_epoch: &Id) -> DomainResult> { + let mut out = BTreeMap::new(); + out.insert( + event_id(&self.epoch, 0, "bootstrap", 0), + event_id(to_epoch, 0, "bootstrap", 0), + ); + // The counter task issues exactly one `counter-delta` event per bound port per + // evaluated transition, in descriptor port order, so every identity it has ever + // issued is re-derivable from its ledger without keeping a list of them. + let ports = self.bindings.len() as u32; + for source_step in 1..=self.last_source_step { + for ordinal in 0..ports { + out.insert( + event_id(&self.epoch, source_step, "counter-delta", ordinal), + event_id(to_epoch, source_step, "counter-delta", ordinal), + ); + } + } + Ok(out) + } + + fn event_watermarks(&self) -> (u64, u64) { + (self.last_source_step, self.issued_events) + } } /// The inspection value the counter environment publishes. diff --git a/services/flysim/crates/fly-session/src/types.rs b/services/flysim/crates/fly-session/src/types.rs index c52c276..07041c2 100644 --- a/services/flysim/crates/fly-session/src/types.rs +++ b/services/flysim/crates/fly-session/src/types.rs @@ -8,13 +8,18 @@ //! `Id` and `Digest` are type aliases, because the shared crate carries both as validated //! `String`s from `flybus::wire` rather than forking the encodings into newtypes. +use std::collections::BTreeMap; + use serde_json::{Map, Value}; pub use fly_session_types::ArtifactRef; pub use fly_session_types::canonical::{ self, OperationKey, body_digest, canonicalize, digest_of, sha256_hex, }; -pub use fly_session_types::media::{AudioDescriptor, AudioRef, ViewDescriptor, ViewRef}; +pub use fly_session_types::media::{ + ActivateRestoreParams, ActivateRestoreResult, AudioDescriptor, AudioRef, CaptureParams, + CaptureResult, StageRestoreParams, StageRestoreResult, ViewDescriptor, ViewRef, +}; pub use fly_session_types::rpc::{ ErrorCode, MutationCertainty, SessionRpcFailure, SessionRpcOutcome, SessionRpcRequest, SessionRpcSuccess, @@ -214,6 +219,57 @@ pub fn outcome_identity( } } +/// The epoch-derived identities of a behaviour trace, rewritten onto one reference epoch. +/// +/// `step-v1` section 8 compares committed behaviour across runs, excluding wall time, request +/// ids "and other explicitly operational metadata". A resumed run's epoch is neither: it is +/// behaviour metadata, and `scope.epoch`, the batch id and every task event id are derived +/// from it. Comparing the two runs therefore means rewriting exactly those three things and +/// nothing else, which is what this does -- and it **fails** on anything it does not +/// recognise instead of passing it through, so a field that silently stopped being rebased +/// would fail the comparison rather than weaken it. +#[derive(Clone, Debug, PartialEq, Eq)] +pub struct EpochRebase { + pub from: Id, + pub to: Id, + /// Every event identity the task issued under `from`, and the identity it has under `to`. + pub events: BTreeMap, +} + +impl EpochRebase { + /// Rewrites one behaviour record. An identity this rebase does not know is an error. + pub fn apply(&self, behaviour: &TraceBehaviour) -> Result { + if behaviour.scope.epoch != self.from { + return Err(format!( + "this behaviour was recorded in epoch {}, not {}", + behaviour.scope.epoch, self.from + )); + } + let mut out = behaviour.clone(); + out.scope = Scope::new(&behaviour.scope.session_id, &self.to, behaviour.scope.step) + .map_err(|e| e.0)?; + let prefix = format!("batch-{}-", self.from); + let suffix = behaviour + .batch_id + .strip_prefix(&prefix) + .ok_or_else(|| format!("the batch id {} is not derived from {}", behaviour.batch_id, self.from))?; + out.batch_id = parse_id(&format!("batch-{}-{suffix}", self.to))?; + let map = |ids: &[Id]| -> Result, String> { + ids.iter() + .map(|id| { + self.events + .get(id) + .cloned() + .ok_or_else(|| format!("no rebased identity for the event {id}")) + }) + .collect() + }; + out.outcome_ids = map(&behaviour.outcome_ids)?; + out.event_ids = map(&behaviour.event_ids)?; + Ok(out) + } +} + /// One session phase transition, recorded whether or not it ends a step. #[derive(Clone, Debug, PartialEq, Eq)] pub struct PhaseTransition { @@ -253,6 +309,28 @@ impl TraceLog { .collect() } + /// Every transition's behaviour, rebased onto one epoch and canonicalized. + /// + /// This is the comparison a resumed run is held to: the same strings as + /// [`TraceLog::behavior`], with the epoch metadata accounted for and nothing else changed. + /// A resumed run's log holds transitions from two epochs -- the ones before the checkpoint + /// and the ones after the restore -- so a transition already recorded in `rebase.to` is + /// kept as it stands and one recorded in `rebase.from` is rewritten. A transition in a + /// third epoch is an error; there is no pass-through case. + pub fn behavior_rebased(&self, rebase: &EpochRebase) -> Result, String> { + self.transitions + .iter() + .map(|t| { + let behaviour = if t.behaviour.scope.epoch == rebase.to { + t.behaviour.clone() + } else { + rebase.apply(&t.behaviour)? + }; + canonicalize(&behaviour.to_json()).map_err(|e| e.0) + }) + .collect() + } + /// The phase path, as `from -> to` strings. pub fn phase_path(&self) -> Vec { self.phases.iter().map(|p| format!("{} -> {}", p.from, p.to)).collect() diff --git a/services/flysim/crates/fly-session/src/worker.rs b/services/flysim/crates/fly-session/src/worker.rs index 327d66b..d717878 100644 --- a/services/flysim/crates/fly-session/src/worker.rs +++ b/services/flysim/crates/fly-session/src/worker.rs @@ -597,7 +597,17 @@ async fn execute( fn classify_default(method: &str) -> Option { match method { "Agent.Prepare" | "Agent.Commit" | "Environment.Advance" => Some(OpClass::StepMutation), - "Agent.Initialize" | "Environment.Initialize" => Some(OpClass::Lifecycle), + // `ipc-v1` section 5 retains lifecycle *and capture* replies until + // `Worker.Acknowledge`. The restore methods join them: their replies carry a + // once-only token and, for an environment, the restored observation's artifact, and a + // duplicate domain request must replay that reply rather than stage or activate a + // second time. They are not step mutations -- they carry no committed step of their + // own and are not keyed by one. + "Agent.Initialize" + | "Environment.Initialize" + | "State.Capture" + | "State.StageRestore" + | "State.ActivateRestore" => Some(OpClass::Lifecycle), "Worker.Hello" | "Worker.Status" | "Worker.Acknowledge" | "Worker.Shutdown" => { Some(OpClass::ReadOnly) } diff --git a/services/flysim/crates/fly-session/tests/processes.rs b/services/flysim/crates/fly-session/tests/processes.rs index d2dcd53..17dd0e2 100644 --- a/services/flysim/crates/fly-session/tests/processes.rs +++ b/services/flysim/crates/fly-session/tests/processes.rs @@ -121,6 +121,7 @@ async fn a_slow_participant_is_resolved_rather_than_failed(mode: ExecutionMode) resolve: Duration::from_secs(20), resolve_attempts: 4096, boot: Duration::from_secs(30), + capture: Duration::from_secs(30), }; within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap(); let reports = within("run", f.harness.coordinator.run(2)) @@ -191,6 +192,7 @@ async fn a_resolution_says_which_of_its_two_bounds_ended_it(mode: ExecutionMode) resolve: Duration::from_millis(300), resolve_attempts: 8192, boot: Duration::from_secs(30), + capture: Duration::from_secs(30), }; within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap(); let started = Instant::now(); @@ -221,6 +223,7 @@ async fn a_resolution_says_which_of_its_two_bounds_ended_it(mode: ExecutionMode) resolve: Duration::from_secs(600), resolve_attempts: 3, boot: Duration::from_secs(30), + capture: Duration::from_secs(30), }; within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap(); let failure = within("step", f.harness.coordinator.step()) diff --git a/services/flysim/crates/fly-session/tests/state.rs b/services/flysim/crates/fly-session/tests/state.rs new file mode 100644 index 0000000..d7b3a0f --- /dev/null +++ b/services/flysim/crates/fly-session/tests/state.rs @@ -0,0 +1,947 @@ +//! STATE-01 acceptance: one coherent all-participant checkpoint, and the recovery that +//! installs it into a fresh epoch. +//! +//! Every test here is one of the slice's acceptance bullets or one of the failure-injection +//! rows the implementation guide's section 4 assigns to it. Each of them runs over both +//! transports and in all three execution modes: the recovery path crosses a process boundary +//! in exactly the places the media path does, so a row that holds in one mode has to hold in +//! all of them. +//! +//! The byte layout itself is proved against the contract crate in +//! `fly-session-types/tests/checkpoint_envelope.rs`; these prove the session's use of it. + +mod common; + +use std::collections::BTreeSet; +use std::path::Path; +use std::sync::Arc; + +use serde_json::{Value, json}; + +use common::{Fixture, count, fixture, fly_a, fly_b, within}; +use fly_session::agent::AgentFaults; +use fly_session::harness::{ExecutionMode, HarnessConfig, Via}; +use fly_session::phase::Phase; +use fly_session::state::{SaveOutcome, WriterFaults}; +use fly_session::types::*; +use fly_session_types::checkpoint; + +both_transports!( + an_uninterrupted_run_and_a_resumed_run_produce_matching_traces, + corrupting_any_single_participants_payload_fails_the_install_as_a_group, + a_lost_save_reply_does_not_advance_durable_metadata, + a_failure_during_activation_cannot_resume_half_a_world, + the_checkpoint_queue_under_stress_stays_bounded, + old_media_cannot_cross_a_recovery, + a_checkpoint_taken_by_another_parser_is_refused_by_name, + a_group_where_one_participant_will_not_stage_resumes_nothing, + a_restore_token_activates_only_once, + the_fence_lifts_only_through_a_coherent_restore, + a_queued_replaceable_capture_is_superseded_rather_than_duplicated, + a_capture_is_refused_anywhere_but_a_committed_boundary, +); + +all_modes!( + matching_traces_in_every_mode, + a_group_install_fails_as_a_group_in_every_mode, + a_lost_save_reply_holds_durable_metadata_in_every_mode, + half_an_activation_resumes_nothing_in_every_mode, + the_queue_stays_bounded_in_every_mode, + old_media_cannot_cross_in_every_mode, + another_parser_is_refused_in_every_mode, + a_refused_stage_resumes_nothing_in_every_mode, + a_token_activates_once_in_every_mode, + the_fence_lifts_only_by_restore_in_every_mode, +); + +const BEFORE: u64 = 2; +const AFTER: u64 = 2; + +fn ckpt(n: u32) -> Id { + id(&format!("ckpt-{n}")) +} + +/// A fixture in one transport and one execution mode. +/// +/// `Via` is the composition's; a dedicated thread or a separate process reaches the router +/// over a socket whatever it says, which the launcher decides and this does not second-guess. +async fn fx(via: Via, mode: ExecutionMode, config: HarnessConfig) -> Fixture { + fixture(via, HarnessConfig { mode, ..config }).await +} + +/// Runs to a committed boundary and takes one durable checkpoint there. +async fn run_and_checkpoint(f: &mut Fixture, steps: u64, checkpoint_id: &Id) { + within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap(); + within("run", f.harness.coordinator.run(steps)).await.unwrap(); + let outcome = within("checkpoint", f.harness.coordinator.checkpoint(checkpoint_id)) + .await + .unwrap(); + match outcome { + SaveOutcome::Committed { boundary, .. } => assert_eq!(boundary, steps), + other => panic!("the checkpoint was not committed: {other:?}"), + } + assert_eq!( + f.harness.coordinator.durable(), + Some((checkpoint_id.clone(), steps)), + "a committed save is the only thing that moves the durable mark" + ); +} + +/// Fails the epoch the way a participant death does, and checks the fence closed. +async fn fail_the_epoch(f: &mut Fixture) { + f.harness.kill(&fly_a()).await; + let failure = within("step", f.harness.coordinator.step()) + .await + .expect_err("a dead participant fails the epoch"); + assert!( + failure.participant.is_some(), + "a diagnosed failure names its participant: {failure}" + ); + assert_eq!(f.harness.coordinator.phase(), Phase::Failed); + assert!(f.harness.coordinator.is_fenced()); + assert_eq!( + f.harness.coordinator.live_view_handles(), + 0, + "the fence drops every handle the old epoch held" + ); +} + +/// The committed generation's bytes, as they are on disk. +fn generation_bytes(root: &Path, checkpoint_id: &Id) -> Vec { + std::fs::read(root.join(format!("{checkpoint_id}.flysess"))).expect("a committed generation") +} + +/// Writes a generation file and makes the store manifest describe it, as a repair tool or a +/// previous run would have left it. +fn write_generation(root: &Path, checkpoint_id: &Id, bytes: &[u8]) { + std::fs::write(root.join(format!("{checkpoint_id}.flysess")), bytes).expect("writable"); + let path = root.join("manifest.json"); + let mut manifest: Value = + serde_json::from_slice(&std::fs::read(&path).expect("a store manifest")).expect("json"); + let generations = manifest["generations"].as_array_mut().expect("an array"); + for generation in generations.iter_mut() { + if generation["checkpointId"] == json!(checkpoint_id.as_str()) { + generation["byteLength"] = json!(bytes.len().to_string()); + generation["envelopeDigest"] = json!(digest_of_bytes(bytes)); + let envelope = checkpoint::decode(bytes).expect("a well formed envelope"); + let compatibility = fly_session::state::Compatibility::from_json( + &envelope.manifest["compatibility"], + ) + .expect("a compatibility block"); + generation["compatibilityDigest"] = json!(compatibility.digest()); + } + } + std::fs::write(&path, serde_json::to_vec(&manifest).expect("json")).expect("writable"); +} + +/// Re-encodes one committed generation after `edit` has changed its manifest or its payloads. +fn rewrite_generation( + root: &Path, + checkpoint_id: &Id, + edit: impl FnOnce(&mut Value, &mut Vec<(String, Vec)>), +) { + let bytes = generation_bytes(root, checkpoint_id); + let envelope = checkpoint::decode(&bytes).expect("a committed generation decodes"); + let mut manifest = envelope.manifest.clone(); + let mut payloads = envelope.payloads.clone(); + edit(&mut manifest, &mut payloads); + // The manifest's payload table mirrors the envelope's, so it is rebuilt from the bytes + // that are actually being written rather than left to disagree. + manifest["payloads"] = Value::Array( + payloads + .iter() + .map(|(name, bytes)| { + json!({ + "name": name, + "byteLength": bytes.len().to_string(), + "digest": digest_of_bytes(bytes), + }) + }) + .collect(), + ); + let rewritten = checkpoint::encode(&manifest, &payloads).expect("a valid envelope"); + write_generation(root, checkpoint_id, &rewritten); +} + +/// Makes the store read its durable metadata again, after a test has edited it. +async fn reload_store(f: &mut Fixture) { + f.harness + .coordinator + .writer() + .expect("a store is attached") + .with_store(|store| store.reload()) + .await + .expect("the store manifest is still readable"); +} + +/// The names of every payload the checkpoint holds, participants first. +fn payload_names(root: &Path, checkpoint_id: &Id) -> Vec { + let bytes = generation_bytes(root, checkpoint_id); + checkpoint::decode(&bytes) + .expect("decodes") + .payloads + .into_iter() + .map(|(name, _)| name) + .collect() +} + +/// Every artifact identity this session's committed boundary is holding. +fn live_artifact_ids(f: &Fixture) -> BTreeSet { + f.harness + .coordinator + .media_handles() + .into_iter() + .map(|(_, reference)| reference.artifact_id) + .collect() +} + +/// Asserts that no participant is running the proposed epoch: nothing was installed. +async fn nothing_is_installed(f: &mut Fixture, epoch: &Id) { + let ids: Vec = std::iter::once(f.harness.environment_id()) + .chain(f.harness.config.agents.iter().map(|a| a.agent_id.clone())) + .collect(); + for worker_id in ids { + let (_, launcher) = f.harness.parts(); + let Ok(status) = launcher.health_check(&worker_id).await else { + // A participant that is gone is certainly not running the new epoch. + continue; + }; + if let Some(scope) = status.current_scope { + assert_ne!( + scope.epoch, *epoch, + "{worker_id} is running the epoch the abandoned install proposed" + ); + } + } + assert_eq!(f.harness.coordinator.phase(), Phase::Failed); + assert!(f.harness.coordinator.is_fenced()); + let refused = within("step", f.harness.coordinator.step()) + .await + .expect_err("a fenced session takes no step"); + assert_eq!(refused.error.code, ErrorCode::InvalidPhase); +} + +// =============================================================================================== +// Acceptance: an uninterrupted run and a resumed run produce matching traces + +async fn an_uninterrupted_run_and_a_resumed_run_produce_matching_traces(via: Via) { + matching_traces(via, ExecutionMode::InProcess).await; +} + +async fn matching_traces_in_every_mode(mode: ExecutionMode) { + matching_traces(Via::Unix, mode).await; +} + +/// The reference run and the resumed run commit the same behaviour, once the epoch metadata +/// the restore necessarily changed is accounted for. +/// +/// `step-v1` section 8's split is the whole of the comparison: the behaviour is what must +/// match, the request ids and wall time are excluded because they are operational, and the +/// epoch is neither -- it is behaviour metadata, so it is rewritten explicitly and everything +/// else is compared byte for byte. +async fn matching_traces(via: Via, mode: ExecutionMode) { + let mut reference = fx(via, mode, HarnessConfig::default()).await; + within("bootstrap", reference.harness.coordinator.bootstrap()).await.unwrap(); + within("run", reference.harness.coordinator.run(BEFORE + AFTER)).await.unwrap(); + let expected = reference.harness.coordinator.trace.behavior(); + assert_eq!(expected.len() as u64, BEFORE + AFTER); + reference.shutdown().await; + + let mut f = fx(via, mode, HarnessConfig::default()).await; + let checkpoint_id = ckpt(1); + run_and_checkpoint(&mut f, BEFORE, &checkpoint_id).await; + fail_the_epoch(&mut f).await; + within("replace", f.harness.replace_all_participants()).await.unwrap(); + let report = within( + "restore", + f.harness.coordinator.restore(Some(&checkpoint_id), &id("e2")), + ) + .await + .unwrap(); + assert_eq!(report.boundary, BEFORE); + assert_eq!(report.epoch, id("e2")); + assert_eq!(report.activated.len(), report.staged.len()); + assert!(!f.harness.coordinator.is_fenced(), "a coherent restore lifts the fence"); + assert_eq!(f.harness.coordinator.phase(), Phase::Paused(BEFORE)); + + f.harness.coordinator.resume().unwrap(); + within("resume", f.harness.coordinator.run(AFTER)).await.unwrap(); + assert_eq!(f.harness.coordinator.phase(), Phase::Ready(BEFORE + AFTER)); + + // Without accounting for the epoch the two traces disagree, which is what makes the + // rebase a statement rather than a formality. + let raw = f.harness.coordinator.trace.behavior(); + assert_eq!(raw.len() as u64, BEFORE + AFTER); + assert_ne!(raw, expected, "the resumed run runs in a different epoch"); + + let rebase = f.harness.coordinator.rebase(&id("e1")).unwrap(); + let resumed = f.harness.coordinator.trace.behavior_rebased(&rebase).unwrap(); + assert_eq!( + resumed, expected, + "a resumed run commits the behaviour the uninterrupted run committed" + ); + f.shutdown().await; +} + +// =============================================================================================== +// Acceptance: corrupt any participant and installation fails as a group + +async fn corrupting_any_single_participants_payload_fails_the_install_as_a_group(via: Via) { + group_install_is_all_or_nothing(via, ExecutionMode::InProcess).await; +} + +async fn a_group_install_fails_as_a_group_in_every_mode(mode: ExecutionMode) { + group_install_is_all_or_nothing(Via::Unix, mode).await; +} + +/// One corrupted payload -- any participant's, and the coordinator's own ledger too -- fails +/// the whole install, and the group afterwards is exactly as it was. +async fn group_install_is_all_or_nothing(via: Via, mode: ExecutionMode) { + let mut f = fx(via, mode, HarnessConfig::default()).await; + let checkpoint_id = ckpt(1); + run_and_checkpoint(&mut f, BEFORE, &checkpoint_id).await; + let root = f.harness.checkpoint_root().to_path_buf(); + let good = generation_bytes(&root, &checkpoint_id); + let names = payload_names(&root, &checkpoint_id); + // Every participant's payload, plus one of the coordinator's own ledgers. + let mut corrupt: Vec = names + .iter() + .filter(|name| name.starts_with("agent-") || *name == "world") + .cloned() + .collect(); + corrupt.push("task-ledger".to_owned()); + assert!(corrupt.len() >= 3, "the composition has several payloads: {names:?}"); + fail_the_epoch(&mut f).await; + + let mut epoch = 1u32; + for target in &corrupt { + epoch += 1; + let proposed = id(&format!("e{epoch}")); + // Each round starts from the intact bytes, so exactly one payload is corrupt. + write_generation(&root, &checkpoint_id, &good); + // A well formed envelope whose digests all agree, so what fails is the participant + // reading its own bytes and not the envelope reader in front of it. + let name = target.clone(); + rewrite_generation(&root, &checkpoint_id, |_manifest, payloads| { + for (payload, bytes) in payloads.iter_mut() { + if *payload == name { + let last = bytes.len() - 1; + bytes[last] ^= 0xff; + } + } + }); + reload_store(&mut f).await; + within("replace", f.harness.replace_all_participants()).await.unwrap(); + let failure = within( + "restore", + f.harness.coordinator.restore(Some(&checkpoint_id), &proposed), + ) + .await + .expect_err(&format!("a corrupt {target} must fail the install")); + assert!( + matches!( + failure.error.code, + ErrorCode::IncompatibleState | ErrorCode::InvalidArgument + ), + "a corrupt {target} is an explicit refusal, not {failure}" + ); + nothing_is_installed(&mut f, &proposed).await; + let tainted = f.harness.coordinator.tainted(); + if target == "world" { + assert!( + tainted.is_empty(), + "the first participant asked refused, so nothing staged: {tainted:?}" + ); + } else { + assert!( + !tainted.is_empty(), + "a participant that staged into an abandoned install must be replaced" + ); + } + } + + // The same group, the same store, the intact bytes: the failures above installed nothing + // that stops this from working. + epoch += 1; + write_generation(&root, &checkpoint_id, &good); + reload_store(&mut f).await; + within("replace", f.harness.replace_all_participants()).await.unwrap(); + let report = within( + "restore", + f.harness + .coordinator + .restore(Some(&checkpoint_id), &id(&format!("e{epoch}"))), + ) + .await + .unwrap(); + assert_eq!(report.boundary, BEFORE); + f.harness.coordinator.resume().unwrap(); + within("resume", f.harness.coordinator.run(1)).await.unwrap(); + f.shutdown().await; +} + +// =============================================================================================== +// Acceptance: a lost save reply does not advance durable metadata + +async fn a_lost_save_reply_does_not_advance_durable_metadata(via: Via) { + a_lost_save_reply(via, ExecutionMode::InProcess).await; +} + +async fn a_lost_save_reply_holds_durable_metadata_in_every_mode(mode: ExecutionMode) { + a_lost_save_reply(Via::Unix, mode).await; +} + +/// Two ways a save can end without a saved acknowledgment, and neither moves the mark. +async fn a_lost_save_reply(via: Via, mode: ExecutionMode) { + let lost = ckpt(1); + let config = HarnessConfig { + writer_faults: WriterFaults { + drop_reply_for: Some(lost.clone()), + ..WriterFaults::default() + }, + ..HarnessConfig::default() + }; + let mut f = fx(via, mode, config).await; + within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap(); + within("run", f.harness.coordinator.run(1)).await.unwrap(); + + let ticket = within("capture", f.harness.coordinator.capture(&lost, false)) + .await + .unwrap(); + let outcome = within("durable", f.harness.coordinator.await_durable(ticket)) + .await + .unwrap(); + assert_eq!(outcome, SaveOutcome::ReplyLost { checkpoint_id: lost.clone() }); + assert_eq!( + f.harness.coordinator.durable(), + None, + "a lost save reply never moves the durable mark" + ); + assert_eq!(count(&f.harness.coordinator.audit, &format!("durable:{lost}@1")), 0); + + // The resolution asks the durable metadata about the *same* operation. It never saves + // again, and here the bytes did reach the store manifest. + let resolved = within("resolve", f.harness.coordinator.resolve_durable(&lost)) + .await + .unwrap(); + assert_eq!(resolved, Some(1)); + assert_eq!(f.harness.coordinator.durable(), Some((lost.clone(), 1))); + let commits = f + .harness + .coordinator + .writer() + .expect("a store") + .with_store(|store| store.commits()) + .await; + assert_eq!(commits, 1, "the resolution queried the store; it did not save again"); + f.shutdown().await; + + // The other half: the generation file is renamed and the store manifest is never + // committed. That generation is unreferenced, so it is not a candidate and the mark is + // still where it was. + let orphan = ckpt(2); + let mut config = HarnessConfig::default(); + config.store_faults.stop_before_manifest_commit = true; + let mut g = fx(via, mode, config).await; + within("bootstrap", g.harness.coordinator.bootstrap()).await.unwrap(); + within("run", g.harness.coordinator.run(1)).await.unwrap(); + let outcome = within("checkpoint", g.harness.coordinator.checkpoint(&orphan)) + .await + .unwrap(); + match outcome { + SaveOutcome::Failed { checkpoint_id, .. } => assert_eq!(checkpoint_id, orphan), + other => panic!("an uncommitted manifest is not a saved checkpoint: {other:?}"), + } + assert_eq!(g.harness.coordinator.durable(), None); + let root = g.harness.checkpoint_root().to_path_buf(); + assert!( + root.join(format!("{orphan}.flysess")).is_file(), + "the generation file was written and renamed" + ); + let resolved = within("resolve", g.harness.coordinator.resolve_durable(&orphan)) + .await + .unwrap(); + assert_eq!(resolved, None, "an unreferenced generation is never a restore candidate"); + assert_eq!(g.harness.coordinator.durable(), None); + g.shutdown().await; +} + +// =============================================================================================== +// Acceptance: a failure during activation cannot resume half a world + +async fn a_failure_during_activation_cannot_resume_half_a_world(via: Via) { + half_an_activation(via, ExecutionMode::InProcess).await; +} + +async fn half_an_activation_resumes_nothing_in_every_mode(mode: ExecutionMode) { + half_an_activation(Via::Unix, mode).await; +} + +/// The world and the first agent activate; the second refuses. Nothing plays, the fence +/// stays closed, and the group cannot be resumed until every participant is replaced. +async fn half_an_activation(via: Via, mode: ExecutionMode) { + let mut f = fx(via, mode, HarnessConfig::default()).await; + let checkpoint_id = ckpt(1); + run_and_checkpoint(&mut f, BEFORE, &checkpoint_id).await; + fail_the_epoch(&mut f).await; + + // The replacement for fly-b refuses to activate after it has staged. + f.harness.set_agent_faults( + &fly_b(), + AgentFaults { fail_activate_restore: true, ..AgentFaults::default() }, + ); + within("replace", f.harness.replace_all_participants()).await.unwrap(); + let failure = within( + "restore", + f.harness.coordinator.restore(Some(&checkpoint_id), &id("e2")), + ) + .await + .expect_err("a refused activation fails the install"); + assert_eq!(failure.participant.as_deref(), Some(fly_b().as_str())); + assert_eq!(f.harness.coordinator.phase(), Phase::Failed); + assert!(f.harness.coordinator.is_fenced(), "the group stays fenced"); + let refused = within("step", f.harness.coordinator.step()) + .await + .expect_err("half a world never runs"); + assert_eq!(refused.error.code, ErrorCode::InvalidPhase); + + // The participants that got as far as installing hold state no group resumed. Another + // restore over them is refused by name rather than attempted. + let tainted = f.harness.coordinator.tainted(); + assert!(tainted.contains(&f.harness.environment_id()), "{tainted:?}"); + assert!(tainted.contains(&fly_a()), "{tainted:?}"); + let refused = within( + "restore", + f.harness.coordinator.restore(Some(&checkpoint_id), &id("e3")), + ) + .await + .expect_err("the half-installed group is not restored over"); + assert_eq!(refused.error.code, ErrorCode::InvalidPhase); + assert!(refused.error.message.contains(&fly_a()), "{refused}"); + + // A fresh group, without the injected refusal, resumes the same boundary. + f.harness.set_agent_faults(&fly_b(), AgentFaults::default()); + within("replace", f.harness.replace_all_participants()).await.unwrap(); + assert!(f.harness.coordinator.tainted().is_empty()); + let report = within( + "restore", + f.harness.coordinator.restore(Some(&checkpoint_id), &id("e4")), + ) + .await + .unwrap(); + assert_eq!(report.boundary, BEFORE); + assert!(!f.harness.coordinator.is_fenced()); + f.harness.coordinator.resume().unwrap(); + within("resume", f.harness.coordinator.run(1)).await.unwrap(); + f.shutdown().await; +} + +// =============================================================================================== +// Acceptance: the checkpoint queue under stress stays bounded + +async fn the_checkpoint_queue_under_stress_stays_bounded(via: Via) { + queue_stays_bounded(via, ExecutionMode::InProcess).await; +} + +async fn the_queue_stays_bounded_in_every_mode(mode: ExecutionMode) { + queue_stays_bounded(Via::Unix, mode).await; +} + +/// With the writer stalled, the queue fills to its configured bound and the next capture is +/// refused *before* a single participant is asked for one. +async fn queue_stays_bounded(via: Via, mode: ExecutionMode) { + let gate = Arc::new(tokio::sync::Semaphore::new(0)); + let mut config = HarnessConfig::default(); + config.writer_faults.gate = Some(gate.clone()); + config.writer.queue_capacity = 2; + let mut f = fx(via, mode, config).await; + within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap(); + within("run", f.harness.coordinator.run(1)).await.unwrap(); + + let first = within("capture", f.harness.coordinator.capture(&ckpt(1), false)) + .await + .unwrap(); + let second = within("capture", f.harness.coordinator.capture(&ckpt(2), false)) + .await + .unwrap(); + assert_eq!(f.harness.coordinator.writer().expect("a store").outstanding(), 2); + + let environment = f.harness.environment_id(); + let before = within("status", f.harness.progress_of(&environment)).await.unwrap(); + let refused = within("capture", f.harness.coordinator.capture(&ckpt(3), false)) + .await + .expect_err("a full queue refuses"); + assert_eq!(refused.error.code, ErrorCode::Busy); + assert_eq!(refused.error.mutation, MutationCertainty::None); + let after = within("status", f.harness.progress_of(&environment)).await.unwrap(); + assert_eq!( + before, after, + "a refused capture asks no participant for one: rejection happens before capture" + ); + assert_eq!(count(&f.harness.coordinator.audit, "capture-refused:ckpt-3"), 1); + + // A refused capture is not a failed session: the world keeps stepping. + assert!(!f.harness.coordinator.is_fenced()); + within("run", f.harness.coordinator.run(1)).await.unwrap(); + assert_eq!(f.harness.coordinator.phase(), Phase::Ready(2)); + + let stats = f.harness.coordinator.writer().expect("a store").stats(); + assert_eq!(stats.rejected, 1); + assert!(stats.peak_queue <= 2, "the queue never exceeded its bound: {stats:?}"); + assert!(stats.peak_bytes <= 64 * 1024 * 1024); + + gate.add_permits(16); + for (ticket, expected) in [(first, ckpt(1)), (second, ckpt(2))] { + let outcome = within("durable", f.harness.coordinator.await_durable(ticket)) + .await + .unwrap(); + match outcome { + SaveOutcome::Committed { checkpoint_id, .. } => assert_eq!(checkpoint_id, expected), + other => panic!("an unstalled writer commits: {other:?}"), + } + } + let stats = f.harness.coordinator.writer().expect("a store").stats(); + assert_eq!(stats.committed, 2); + assert_eq!(f.harness.coordinator.writer().expect("a store").outstanding(), 0); + f.shutdown().await; +} + +/// A queued replaceable capture is replaced by the next one, releasing its holds, rather than +/// both being written. +async fn a_queued_replaceable_capture_is_superseded_rather_than_duplicated(via: Via) { + let gate = Arc::new(tokio::sync::Semaphore::new(0)); + let mut config = HarnessConfig::default(); + config.writer_faults.gate = Some(gate.clone()); + config.writer.queue_capacity = 3; + let mut f = fx(via, ExecutionMode::InProcess, config).await; + within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap(); + within("run", f.harness.coordinator.run(1)).await.unwrap(); + + // The first job is taken by the writer and stalls at the gate; the next two are queued. + let durable = within("capture", f.harness.coordinator.capture(&ckpt(1), false)) + .await + .unwrap(); + let hot = within("capture", f.harness.coordinator.capture(&ckpt(2), true)) + .await + .unwrap(); + let newer = within("capture", f.harness.coordinator.capture(&ckpt(3), true)) + .await + .unwrap(); + + let superseded = within("durable", f.harness.coordinator.await_durable(hot)) + .await + .unwrap(); + assert_eq!( + superseded, + SaveOutcome::Superseded { checkpoint_id: ckpt(2), by: ckpt(3) }, + "only a queued replaceable capture is coalesced" + ); + gate.add_permits(16); + for ticket in [durable, newer] { + let outcome = within("durable", f.harness.coordinator.await_durable(ticket)) + .await + .unwrap(); + assert!(matches!(outcome, SaveOutcome::Committed { .. }), "{outcome:?}"); + } + let stats = f.harness.coordinator.writer().expect("a store").stats(); + assert_eq!(stats.superseded, 1); + assert_eq!(stats.committed, 2, "the superseded capture was never written"); + f.shutdown().await; +} + +// =============================================================================================== +// Acceptance: old media and old parser data cannot cross a recovery + +async fn old_media_cannot_cross_a_recovery(via: Via) { + old_media_cannot_cross(via, ExecutionMode::InProcess).await; +} + +async fn old_media_cannot_cross_in_every_mode(mode: ExecutionMode) { + old_media_cannot_cross(Via::Unix, mode).await; +} + +/// Nothing the fence dropped comes back: the payloads are re-imported as fresh artifacts, the +/// restored view is a new object, and the new epoch's first audio chunk resumes the preserved +/// sample position and marks the discontinuity. +async fn old_media_cannot_cross(via: Via, mode: ExecutionMode) { + let mut f = fx(via, mode, HarnessConfig::default()).await; + let checkpoint_id = ckpt(1); + run_and_checkpoint(&mut f, BEFORE, &checkpoint_id).await; + let old_ids = live_artifact_ids(&f); + assert!(!old_ids.is_empty(), "the committed boundary holds media handles"); + let positions = f.harness.coordinator.audio_positions(); + assert!(positions.values().any(|sample| *sample > 0), "audio has played"); + + fail_the_epoch(&mut f).await; + within("replace", f.harness.replace_all_participants()).await.unwrap(); + let report = within( + "restore", + f.harness.coordinator.restore(Some(&checkpoint_id), &id("e2")), + ) + .await + .unwrap(); + for imported in &report.imported { + assert!( + !old_ids.contains(imported), + "a checkpoint payload was imported as an artifact the old epoch already had" + ); + } + let new_ids = live_artifact_ids(&f); + assert!(!new_ids.is_empty()); + assert!( + new_ids.is_disjoint(&old_ids), + "the restored boundary's media are fresh artifacts, not the old epoch's" + ); + assert_eq!( + f.harness.coordinator.audio_positions(), + positions, + "crash restore preserves the sample position" + ); + + f.harness.coordinator.resume().unwrap(); + within("resume", f.harness.coordinator.run(1)).await.unwrap(); + let chunk = f + .harness + .coordinator + .observation() + .expect("a restored world") + .audio + .first() + .cloned() + .expect("one chunk per transition"); + assert!( + chunk.discontinuity, + "the first chunk of a fresh epoch marks the discontinuity the recovery established" + ); + assert_eq!( + chunk.first_sample, + positions["arena"], + "and it starts where the checkpointed stream stopped" + ); + f.shutdown().await; +} + +async fn a_checkpoint_taken_by_another_parser_is_refused_by_name(via: Via) { + another_parser_is_refused(via, ExecutionMode::InProcess).await; +} + +async fn another_parser_is_refused_in_every_mode(mode: ExecutionMode) { + another_parser_is_refused(Via::Unix, mode).await; +} + +/// A checkpoint whose recorded parser identity is not this composition's is refused before a +/// single participant is asked to stage, and the refusal names the identity that differs. +async fn another_parser_is_refused(via: Via, mode: ExecutionMode) { + let mut f = fx(via, mode, HarnessConfig::default()).await; + let checkpoint_id = ckpt(1); + run_and_checkpoint(&mut f, 1, &checkpoint_id).await; + let root = f.harness.checkpoint_root().to_path_buf(); + fail_the_epoch(&mut f).await; + rewrite_generation(&root, &checkpoint_id, |manifest, _payloads| { + manifest["compatibility"]["parserDigest"] = + json!(digest_of_bytes(b"some other inspection schema")); + }); + reload_store(&mut f).await; + within("replace", f.harness.replace_all_participants()).await.unwrap(); + let failure = within( + "restore", + f.harness.coordinator.restore(Some(&checkpoint_id), &id("e2")), + ) + .await + .expect_err("another parser's state is not this composition's"); + assert_eq!(failure.error.code, ErrorCode::IncompatibleState); + assert!( + failure.error.message.contains("parser"), + "the refusal names the identity that differs: {failure}" + ); + assert!( + f.harness.coordinator.tainted().is_empty(), + "nothing was asked to stage, so nothing has to be replaced" + ); + nothing_is_installed(&mut f, &id("e2")).await; + f.shutdown().await; +} + +// =============================================================================================== +// Failure-injection rows + +async fn a_group_where_one_participant_will_not_stage_resumes_nothing(via: Via) { + a_refused_stage(via, ExecutionMode::InProcess).await; +} + +async fn a_refused_stage_resumes_nothing_in_every_mode(mode: ExecutionMode) { + a_refused_stage(Via::Unix, mode).await; +} + +/// The checklist row: the install validates the participants ahead of it and the last one +/// fails. Nothing is resumed. +async fn a_refused_stage(via: Via, mode: ExecutionMode) { + let mut f = fx(via, mode, HarnessConfig::default()).await; + let checkpoint_id = ckpt(1); + run_and_checkpoint(&mut f, 1, &checkpoint_id).await; + fail_the_epoch(&mut f).await; + f.harness.set_agent_faults( + &fly_b(), + AgentFaults { fail_stage_restore: true, ..AgentFaults::default() }, + ); + within("replace", f.harness.replace_all_participants()).await.unwrap(); + let failure = within( + "restore", + f.harness.coordinator.restore(Some(&checkpoint_id), &id("e2")), + ) + .await + .expect_err("a participant that will not validate stops the install"); + assert_eq!(failure.error.code, ErrorCode::IncompatibleState); + assert_eq!(failure.participant.as_deref(), Some(fly_b().as_str())); + assert_eq!( + count(&f.harness.coordinator.audit, &format!("activated:{}", fly_a())), + 0, + "no participant activates when one of them will not stage" + ); + nothing_is_installed(&mut f, &id("e2")).await; + f.shutdown().await; +} + +async fn a_restore_token_activates_only_once(via: Via) { + a_token_activates_once(via, ExecutionMode::InProcess).await; +} + +async fn a_token_activates_once_in_every_mode(mode: ExecutionMode) { + a_token_activates_once(Via::Unix, mode).await; +} + +/// A token is bound to its checkpoint, scope and payload and activates once. A fresh request +/// naming it again is a conflict, and the old epoch's scope is stale on the new participants. +async fn a_token_activates_once(via: Via, mode: ExecutionMode) { + let mut f = fx(via, mode, HarnessConfig::default()).await; + let checkpoint_id = ckpt(1); + run_and_checkpoint(&mut f, 1, &checkpoint_id).await; + fail_the_epoch(&mut f).await; + within("replace", f.harness.replace_all_participants()).await.unwrap(); + let report = within( + "restore", + f.harness.coordinator.restore(Some(&checkpoint_id), &id("e2")), + ) + .await + .unwrap(); + let (who, token) = report.tokens.first().cloned().expect("a staged token"); + let worker = if who == f.harness.environment_id() { + f.harness.coordinator.environment_ref().clone() + } else { + f.harness.coordinator.agent_ref(&who).expect("a participant").clone() + }; + let scope = scope_at("demo", "e2", 1); + let refused = within( + "activate", + f.harness.coordinator.probe_raw( + &worker, + "State.ActivateRestore", + Some(scope), + json!({"restoreToken": token.as_str()}), + ), + ) + .await + .expect_err("a token activates once"); + assert_eq!(refused.code, ErrorCode::Conflict); + + // And the epoch the restore left behind is stale on every replacement. + let stale = within( + "prepare", + f.harness.coordinator.probe_raw( + &worker, + if who == f.harness.environment_id() { "Environment.Advance" } else { "Agent.Prepare" }, + Some(scope_at("demo", "e1", 1)), + json!({}), + ), + ) + .await + .expect_err("the old epoch is gone"); + assert!( + matches!(stale.code, ErrorCode::StaleEpoch | ErrorCode::InvalidArgument), + "an old-epoch request is refused: {stale}" + ); + f.shutdown().await; +} + +async fn the_fence_lifts_only_through_a_coherent_restore(via: Via) { + the_fence_lifts_only_by_restore(via, ExecutionMode::InProcess).await; +} + +async fn the_fence_lifts_only_by_restore_in_every_mode(mode: ExecutionMode) { + the_fence_lifts_only_by_restore(Via::Unix, mode).await; +} + +/// Nothing but a coherent restore moves a fenced session, and the restore lands on +/// `Paused(k)` rather than straight back into play. +async fn the_fence_lifts_only_by_restore(via: Via, mode: ExecutionMode) { + let mut f = fx(via, mode, HarnessConfig::default()).await; + let checkpoint_id = ckpt(1); + run_and_checkpoint(&mut f, 1, &checkpoint_id).await; + fail_the_epoch(&mut f).await; + + for refused in [ + within("step", f.harness.coordinator.step()).await.err(), + within("bootstrap", f.harness.coordinator.bootstrap()).await.err(), + within("capture", f.harness.coordinator.capture(&ckpt(9), false)) + .await + .err(), + ] { + let refused = refused.expect("a fenced session refuses"); + assert_eq!(refused.error.code, ErrorCode::InvalidPhase); + } + assert!(f.harness.coordinator.resume().is_err(), "a fenced session does not resume"); + assert!(f.harness.coordinator.is_fenced()); + + // A restore into the epoch that failed is refused: recovery establishes a fresh one. + within("replace", f.harness.replace_all_participants()).await.unwrap(); + let same = within( + "restore", + f.harness.coordinator.restore(Some(&checkpoint_id), &id("e1")), + ) + .await + .expect_err("a restore installs a fresh epoch"); + assert_eq!(same.error.code, ErrorCode::StaleEpoch); + + let report = within( + "restore", + f.harness.coordinator.restore(Some(&checkpoint_id), &id("e2")), + ) + .await + .unwrap(); + assert_eq!(report.boundary, 1); + assert!(!f.harness.coordinator.is_fenced()); + assert_eq!(f.harness.coordinator.phase(), Phase::Paused(1)); + f.harness.coordinator.resume().unwrap(); + within("resume", f.harness.coordinator.run(1)).await.unwrap(); + assert_eq!(f.harness.coordinator.phase(), Phase::Ready(2)); + f.shutdown().await; +} + +/// A capture belongs to a committed boundary, and the phase machine is what says so. +async fn a_capture_is_refused_anywhere_but_a_committed_boundary(via: Via) { + let mut f = fx(via, ExecutionMode::InProcess, HarnessConfig::default()).await; + // Before bootstrap the session is Starting, which is not a boundary at all. + let refused = within("capture", f.harness.coordinator.capture(&ckpt(1), false)) + .await + .expect_err("Starting is not a committed boundary"); + assert_eq!(refused.error.code, ErrorCode::InvalidPhase); + f.shutdown().await; + + // From a pause, which is a committed boundary, it works, and it returns there. + let mut f = fx(via, ExecutionMode::InProcess, HarnessConfig::default()).await; + within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap(); + within("run", f.harness.coordinator.run(1)).await.unwrap(); + f.harness.coordinator.request_pause(); + within("run", f.harness.coordinator.run(1)).await.unwrap(); + assert_eq!(f.harness.coordinator.phase(), Phase::Paused(2)); + let outcome = within("checkpoint", f.harness.coordinator.checkpoint(&ckpt(1))) + .await + .unwrap(); + assert!(matches!(outcome, SaveOutcome::Committed { boundary: 2, .. }), "{outcome:?}"); + assert_eq!( + f.harness.coordinator.phase(), + Phase::Paused(2), + "a capture returns to the boundary it came from" + ); + f.shutdown().await; +} From 45621903be44096d98293da8b2a66cc6508efae2 Mon Sep 17 00:00:00 2001 From: dev Date: Tue, 22 Sep 2026 18:45:51 +0000 Subject: [PATCH 2/2] session: name the two ways a durable wait ends without an acknowledgment Review follow-up on the checkpoint store. A dropped reply channel and an expired caller budget were both reported as ReplyLost. They are different facts -- the first means the write is over and its outcome did not reach here, the second means the save is still going -- so they are now separate outcomes, and the durable wait has its own budget rather than borrowing the one that bounds a call to a participant. Both still leave durable metadata where it was, and for both the resolution asks the store about the same checkpoint. The writer's two bounds refuse at different moments and the comment claimed otherwise: the outstanding-capture bound is taken before a capture is requested, and the byte budget cannot be, because a capture's size is not known until it exists. The byte check, the decision and the change to the byte total are now one critical section, the peak is sampled after a superseded job's bytes are gone, and the writer's own bookkeeping is over a type that holds only the outcomes a writer can produce. The manifest's coordinator.eventWatermarks is {lastSourceStep, issued}; the fixture illustrated {lastEventId, lastOrdinal}, and the illustration is what changed, because an event id is derived from the epoch and cannot be compared across the restore that gives the session a new one. checkpoint-envelope-v1 section 3 also now says, under the same dated amendment, that a required-manifest-field change must bump envelopeVersion once production files exist: contractDigest is taken over the schema set and does not cover this manifest, so the envelope version is the only thing that can carry such a change. --- .../checkpoint-envelope-v1.md | 12 ++ .../examples/update_fixtures.rs | 2 +- .../fixtures/checkpoint-envelope.json | 32 ++-- services/flysim/crates/fly-session/README.md | 18 +- .../crates/fly-session/src/coordinator.rs | 16 +- .../flysim/crates/fly-session/src/state.rs | 168 ++++++++++++------ .../crates/fly-session/tests/processes.rs | 3 + .../flysim/crates/fly-session/tests/state.rs | 57 +++++- 8 files changed, 229 insertions(+), 79 deletions(-) diff --git a/docs/design/session-framework/checkpoint-envelope-v1.md b/docs/design/session-framework/checkpoint-envelope-v1.md index 7cb833f..d58ed98 100644 --- a/docs/design/session-framework/checkpoint-envelope-v1.md +++ b/docs/design/session-framework/checkpoint-envelope-v1.md @@ -118,6 +118,18 @@ section has listed from the start. Both are now in `REQUIRED_MANIFEST_FIELDS` in TypeScript, and the fixture was regenerated by the existing example. The schema set is untouched, so `contractDigest` is unchanged. +`coordinator.eventWatermarks` is `{lastSourceStep, issued}`. The fixture illustrated +`{lastEventId, lastOrdinal}`, and it is the illustration that changed: an event id is derived +from the epoch, so a watermark spelled as one cannot be compared across the restore that +gives the session a new epoch, while a source step and an issued count can. + +**A required-manifest-field change is compatibility-relevant and `contractDigest` does not +cover it.** The digest is taken over the schema set, and this manifest is not in it, so +`envelopeVersion` is the only thing that can carry such a change. It stays `1` here only +because no production `FLYSESS1` file exists yet: once one does, adding or removing a required +manifest field **must** bump `envelopeVersion`, because a reader of the older version would +otherwise accept a file it cannot completely read, or refuse one it could. + `payloads` is redundant with the table on purpose: the table is what a reader needs to map bytes, and the manifest is what a store lists, compares and reports without opening the payload area. A reader checks that the two agree. diff --git a/services/flysim/crates/fly-session-types/examples/update_fixtures.rs b/services/flysim/crates/fly-session-types/examples/update_fixtures.rs index d68b865..bad3693 100644 --- a/services/flysim/crates/fly-session-types/examples/update_fixtures.rs +++ b/services/flysim/crates/fly-session-types/examples/update_fixtures.rs @@ -197,7 +197,7 @@ fn checkpoint_envelope() -> String { "priorInspection": "prior-inspection", "executorState": [{"agentId": "fly-a", "payload": "executor-fly-a"}], "admissionState": null, - "eventWatermarks": {"lastEventId": "evt-1", "lastOrdinal": "7"}, + "eventWatermarks": {"lastSourceStep": "42", "issued": "7"}, }, "environment": {"workerId": "arena", "payload": "world"}, "helperState": [], diff --git a/services/flysim/crates/fly-session-types/fixtures/checkpoint-envelope.json b/services/flysim/crates/fly-session-types/fixtures/checkpoint-envelope.json index 70453c3..b726585 100644 --- a/services/flysim/crates/fly-session-types/fixtures/checkpoint-envelope.json +++ b/services/flysim/crates/fly-session-types/fixtures/checkpoint-envelope.json @@ -59,8 +59,8 @@ ], "admissionState": null, "eventWatermarks": { - "lastEventId": "evt-1", - "lastOrdinal": "7" + "lastSourceStep": "42", + "issued": "7" } }, "environment": { @@ -119,49 +119,49 @@ } ], "envelope": { - "base64": "RkxZU0VTUzEBAAAAIAAAADUIAAAFAAAAWAgAAAAAAAB7ImFnZW50cyI6W3siYWdlbnRJZCI6ImZseS1hIiwiYnJhaW5UaWNrcyI6IjI1MzQiLCJkYXRhc2V0RGlnZXN0IjoiNmMwYWYxZjA3ODRlZjYzYTM5M2VlNzdkNjE0ZTgyNDZjNjI1MDUxMzYwZjNmMWE0ODgzODM3NGM1ZDM1NWI1MiIsIm1vZGVsVmVyc2lvbiI6ImxpZi0xbXMtZjY0LXYyIiwicGF5bG9hZCI6ImFnZW50LWZseS1hIiwicGxhc3RpY2l0eVZlcnNpb24iOiJmbHkta2MtbWJvbi1yc3RkcC12MiIsInByb2ZpbGVEaWdlc3QiOiIxOTAwZWFiNmMwMjg0ODNkNzEyNjU5OWVlNmY1MGRlMGQyNzkwN2I1YzY1ZmE5MDUyNDU4MGI0YjBmOTg1MmIwIiwicmVtYWluZGVyIjp7ImRlbm9taW5hdG9yIjoiMyIsIm51bWVyYXRvciI6IjEwMDAwMDAifSwic2VlZCI6LTE4NDk0NjA2M31dLCJjaGVja3BvaW50SWQiOiJja3B0LTEiLCJjb21wYXRpYmlsaXR5Ijp7ImJhY2tlbmREaWdlc3QiOiIxMGUwOGE0MTllODUwZWJhMWViYmExOGZkZDI4ZWI3ZWMxYjdlOGJhYTliY2MzYjk3M2UyYjg4OTFlYzcyNmJlIiwiY29udGVudERpZ2VzdCI6ImVkNzAwMmI0MzllOWFjODQ1ZjIyMzU3ZDgyMmJhYzE0NDQ3MzBmYmRiNjAxNmQzZWM5NDMyMjk3YjllYzlmNzMiLCJjb250cm9sbGVyRGlnZXN0IjoiYzE0NzIxMzViMTRjNzdjOGJlZjk4ZTczZjcwMjA4MzI1ZmEwZGNmMWU2YmQ2NjhhZTliMzFhOWNlYTI5NWZlNyIsInBhcnNlckRpZ2VzdCI6ImIxN2Q0NTEyMTE1MDkyOGYyMTQ2YWY0OWUxOTVlZmYxZWVmNWQ2NzMyNWJlMjczYTczM2ZiNzRhY2FkYWEzNDIiLCJwYXRjaERpZ2VzdCI6ImE0ODk1ZWI0NGFmYzMzNmZlY2JiYTZlNTIwY2Q2N2UxNzhkYWNlMDI3NjY1NWQxMDJmY2VmZmE4ZTVmNzA1NzAiLCJzdGF0ZUZvcm1hdElkIjoiZmx5c2Vzcy0xIn0sImNvbXBvc2l0aW9uRGlnZXN0IjoiNzMwZDcyNWM4YTU5ZDNhNzMwM2RlZjJiZWQwNDFhNTc3ZWRiNDI1NWFhYmQ0ODg5Y2UxMjkxODMxMWQ5NTJmMCIsImNvb3JkaW5hdG9yIjp7ImFkbWlzc2lvblN0YXRlIjpudWxsLCJldmVudFdhdGVybWFya3MiOnsibGFzdEV2ZW50SWQiOiJldnQtMSIsImxhc3RPcmRpbmFsIjoiNyJ9LCJleGVjdXRvclN0YXRlIjpbeyJhZ2VudElkIjoiZmx5LWEiLCJwYXlsb2FkIjoiZXhlY3V0b3ItZmx5LWEifV0sInByaW9ySW5zcGVjdGlvbiI6InByaW9yLWluc3BlY3Rpb24iLCJ0YXNrTGVkZ2VyIjoidGFzay1sZWRnZXIifSwiZW52ZWxvcGVWZXJzaW9uIjoxLCJlbnZpcm9ubWVudCI6eyJwYXlsb2FkIjoid29ybGQiLCJ3b3JrZXJJZCI6ImFyZW5hIn0sImVwaXNvZGVJZCI6ImVwaXNvZGUtMSIsImhlbHBlclN0YXRlIjpbXSwicGF5bG9hZHMiOlt7ImJ5dGVMZW5ndGgiOiIxNyIsImRpZ2VzdCI6IjEzMjFkZmZiMGNkYzZmOTA5MmNiZjdmYTJhNWZjNjhiYmVkMTJjOTkzZDVhZDM5ODI2NDAxMjgxMGNlOWJmOTMiLCJuYW1lIjoiYWdlbnQtZmx5LWEifSx7ImJ5dGVMZW5ndGgiOiIxNCIsImRpZ2VzdCI6IjNhZWU2MGRmN2UyOWVmZWJhN2Y1Zjk5ZmM1ODY3NjQ3YjM2YWViZmYxZDVkM2M4MzhkYmZmMzIzMTIyZTY0NjIiLCJuYW1lIjoiZXhlY3V0b3ItZmx5LWEifSx7ImJ5dGVMZW5ndGgiOiIxMSIsImRpZ2VzdCI6IjQwYjAwZWQyYmJiYTkwMWQ2ODIwNWZmNzFiMDRhNDRiOWVlNTNjNTFjYjMxMDlhYTJjZWFhNDRmMWM0NTcyN2UiLCJuYW1lIjoidGFzay1sZWRnZXIifSx7ImJ5dGVMZW5ndGgiOiIxMCIsImRpZ2VzdCI6IjJjMTNiN2I0ZDlhOTkxNjgwMWFiOTE5MWMzMTRmMzFiMDQ1ZTliOWM1YjY2OWE2YzA0NzRmMDIxN2VmNzViZjUiLCJuYW1lIjoicHJpb3ItaW5zcGVjdGlvbiJ9LHsiYnl0ZUxlbmd0aCI6IjY0IiwiZGlnZXN0IjoiZjVhNWZkNDJkMTZhMjAzMDI3OThlZjZlZDMwOTk3OWI0MzAwM2QyMzIwZDlmMGU4ZWE5ODMxYTkyNzU5ZmI0YiIsIm5hbWUiOiJ3b3JsZCJ9XSwicG9ydE1hcCI6W3siYWdlbnRJZCI6ImZseS1hIiwicG9ydElkIjoicG9ydC0xIn1dLCJzY2hlZHVsZXJJZCI6ImxvY2tzdGVwLXYxIiwic291cmNlU2NvcGUiOnsiZXBvY2giOiJlcG9jaC0xIiwic2Vzc2lvbklkIjoiZGVtbyIsInN0ZXAiOiI0MiJ9LCJ3b3JsZFRpbWUiOnsiZGVub21pbmF0b3IiOiIxIiwibnVtZXJhdG9yIjoiNzAwMDAwMDAwIn19AAAAYWdlbnQtZmx5LWEAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAIgKAAAAAAAAEQAAAAAAAAATId/7DNxvkJLL9/oqX8aLvtEsmT1a05gmQBKBDOm/k2V4ZWN1dG9yLWZseS1hAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACgCgAAAAAAAA4AAAAAAAAAOu5g334p7+un9fmfxYZ2R7Nq6/8dXTyDjb/zIxIuZGJ0YXNrLWxlZGdlcgAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAsAoAAAAAAAALAAAAAAAAAECwDtK7upAdaCBf9xsEpEue5TxRyzEJqizqpE8cRXJ+cHJpb3ItaW5zcGVjdGlvbgAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAMAKAAAAAAAACgAAAAAAAAAsE7e02amRaAGrkZHDFPMbBF6bnFtmmmwEdPAhfvdb9XdvcmxkAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAADQCgAAAAAAAEAAAAAAAAAA9aX9QtFqIDAnmO9u0wmXm0MAPSMg2fDo6pgxqSdZ+0thZ2VudCBzdGF0ZSBieXRlcwAAAAAAAABleGVjdXRvciBzdGF0ZQAAeyJyYW5rIjoxMH0AAAAAAHsibWFwIjo0MH0AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAQAsAAAAAAAAhBlh9AWLTtVKckmeIzNn4DO5Yhn2C1nUA1T1RfxMXlkZMWVNFU1NG", - "byteLength": 2880, + "base64": "RkxZU0VTUzEBAAAAIAAAADAIAAAFAAAAUAgAAAAAAAB7ImFnZW50cyI6W3siYWdlbnRJZCI6ImZseS1hIiwiYnJhaW5UaWNrcyI6IjI1MzQiLCJkYXRhc2V0RGlnZXN0IjoiNmMwYWYxZjA3ODRlZjYzYTM5M2VlNzdkNjE0ZTgyNDZjNjI1MDUxMzYwZjNmMWE0ODgzODM3NGM1ZDM1NWI1MiIsIm1vZGVsVmVyc2lvbiI6ImxpZi0xbXMtZjY0LXYyIiwicGF5bG9hZCI6ImFnZW50LWZseS1hIiwicGxhc3RpY2l0eVZlcnNpb24iOiJmbHkta2MtbWJvbi1yc3RkcC12MiIsInByb2ZpbGVEaWdlc3QiOiIxOTAwZWFiNmMwMjg0ODNkNzEyNjU5OWVlNmY1MGRlMGQyNzkwN2I1YzY1ZmE5MDUyNDU4MGI0YjBmOTg1MmIwIiwicmVtYWluZGVyIjp7ImRlbm9taW5hdG9yIjoiMyIsIm51bWVyYXRvciI6IjEwMDAwMDAifSwic2VlZCI6LTE4NDk0NjA2M31dLCJjaGVja3BvaW50SWQiOiJja3B0LTEiLCJjb21wYXRpYmlsaXR5Ijp7ImJhY2tlbmREaWdlc3QiOiIxMGUwOGE0MTllODUwZWJhMWViYmExOGZkZDI4ZWI3ZWMxYjdlOGJhYTliY2MzYjk3M2UyYjg4OTFlYzcyNmJlIiwiY29udGVudERpZ2VzdCI6ImVkNzAwMmI0MzllOWFjODQ1ZjIyMzU3ZDgyMmJhYzE0NDQ3MzBmYmRiNjAxNmQzZWM5NDMyMjk3YjllYzlmNzMiLCJjb250cm9sbGVyRGlnZXN0IjoiYzE0NzIxMzViMTRjNzdjOGJlZjk4ZTczZjcwMjA4MzI1ZmEwZGNmMWU2YmQ2NjhhZTliMzFhOWNlYTI5NWZlNyIsInBhcnNlckRpZ2VzdCI6ImIxN2Q0NTEyMTE1MDkyOGYyMTQ2YWY0OWUxOTVlZmYxZWVmNWQ2NzMyNWJlMjczYTczM2ZiNzRhY2FkYWEzNDIiLCJwYXRjaERpZ2VzdCI6ImE0ODk1ZWI0NGFmYzMzNmZlY2JiYTZlNTIwY2Q2N2UxNzhkYWNlMDI3NjY1NWQxMDJmY2VmZmE4ZTVmNzA1NzAiLCJzdGF0ZUZvcm1hdElkIjoiZmx5c2Vzcy0xIn0sImNvbXBvc2l0aW9uRGlnZXN0IjoiNzMwZDcyNWM4YTU5ZDNhNzMwM2RlZjJiZWQwNDFhNTc3ZWRiNDI1NWFhYmQ0ODg5Y2UxMjkxODMxMWQ5NTJmMCIsImNvb3JkaW5hdG9yIjp7ImFkbWlzc2lvblN0YXRlIjpudWxsLCJldmVudFdhdGVybWFya3MiOnsiaXNzdWVkIjoiNyIsImxhc3RTb3VyY2VTdGVwIjoiNDIifSwiZXhlY3V0b3JTdGF0ZSI6W3siYWdlbnRJZCI6ImZseS1hIiwicGF5bG9hZCI6ImV4ZWN1dG9yLWZseS1hIn1dLCJwcmlvckluc3BlY3Rpb24iOiJwcmlvci1pbnNwZWN0aW9uIiwidGFza0xlZGdlciI6InRhc2stbGVkZ2VyIn0sImVudmVsb3BlVmVyc2lvbiI6MSwiZW52aXJvbm1lbnQiOnsicGF5bG9hZCI6IndvcmxkIiwid29ya2VySWQiOiJhcmVuYSJ9LCJlcGlzb2RlSWQiOiJlcGlzb2RlLTEiLCJoZWxwZXJTdGF0ZSI6W10sInBheWxvYWRzIjpbeyJieXRlTGVuZ3RoIjoiMTciLCJkaWdlc3QiOiIxMzIxZGZmYjBjZGM2ZjkwOTJjYmY3ZmEyYTVmYzY4YmJlZDEyYzk5M2Q1YWQzOTgyNjQwMTI4MTBjZTliZjkzIiwibmFtZSI6ImFnZW50LWZseS1hIn0seyJieXRlTGVuZ3RoIjoiMTQiLCJkaWdlc3QiOiIzYWVlNjBkZjdlMjllZmViYTdmNWY5OWZjNTg2NzY0N2IzNmFlYmZmMWQ1ZDNjODM4ZGJmZjMyMzEyMmU2NDYyIiwibmFtZSI6ImV4ZWN1dG9yLWZseS1hIn0seyJieXRlTGVuZ3RoIjoiMTEiLCJkaWdlc3QiOiI0MGIwMGVkMmJiYmE5MDFkNjgyMDVmZjcxYjA0YTQ0YjllZTUzYzUxY2IzMTA5YWEyY2VhYTQ0ZjFjNDU3MjdlIiwibmFtZSI6InRhc2stbGVkZ2VyIn0seyJieXRlTGVuZ3RoIjoiMTAiLCJkaWdlc3QiOiIyYzEzYjdiNGQ5YTk5MTY4MDFhYjkxOTFjMzE0ZjMxYjA0NWU5YjljNWI2NjlhNmMwNDc0ZjAyMTdlZjc1YmY1IiwibmFtZSI6InByaW9yLWluc3BlY3Rpb24ifSx7ImJ5dGVMZW5ndGgiOiI2NCIsImRpZ2VzdCI6ImY1YTVmZDQyZDE2YTIwMzAyNzk4ZWY2ZWQzMDk5NzliNDMwMDNkMjMyMGQ5ZjBlOGVhOTgzMWE5Mjc1OWZiNGIiLCJuYW1lIjoid29ybGQifV0sInBvcnRNYXAiOlt7ImFnZW50SWQiOiJmbHktYSIsInBvcnRJZCI6InBvcnQtMSJ9XSwic2NoZWR1bGVySWQiOiJsb2Nrc3RlcC12MSIsInNvdXJjZVNjb3BlIjp7ImVwb2NoIjoiZXBvY2gtMSIsInNlc3Npb25JZCI6ImRlbW8iLCJzdGVwIjoiNDIifSwid29ybGRUaW1lIjp7ImRlbm9taW5hdG9yIjoiMSIsIm51bWVyYXRvciI6IjcwMDAwMDAwMCJ9fWFnZW50LWZseS1hAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACACgAAAAAAABEAAAAAAAAAEyHf+wzcb5CSy/f6Kl/Gi77RLJk9WtOYJkASgQzpv5NleGVjdXRvci1mbHktYQAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAmAoAAAAAAAAOAAAAAAAAADruYN9+Ke/rp/X5n8WGdkezauv/HV08g42/8yMSLmRidGFzay1sZWRnZXIAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAKgKAAAAAAAACwAAAAAAAABAsA7Su7qQHWggX/cbBKRLnuU8UcsxCaos6qRPHEVyfnByaW9yLWluc3BlY3Rpb24AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAC4CgAAAAAAAAoAAAAAAAAALBO3tNmpkWgBq5GRwxTzGwRem5xbZppsBHTwIX73W/V3b3JsZAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAyAoAAAAAAABAAAAAAAAAAPWl/ULRaiAwJ5jvbtMJl5tDAD0jINnw6OqYMaknWftLYWdlbnQgc3RhdGUgYnl0ZXMAAAAAAAAAZXhlY3V0b3Igc3RhdGUAAHsicmFuayI6MTB9AAAAAAB7Im1hcCI6NDB9AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAADgLAAAAAAAAX4L9WkdX0MViD3h5YJgf7VQgwocWU4XJoKX6MUwl+8hGTFlTRVNTRg==", + "byteLength": 2872, "layout": { "headerBytes": 32, "manifestOffset": "32", - "manifestBytes": 2101, - "tableOffset": "2136", + "manifestBytes": 2096, + "tableOffset": "2128", "tableEntryBytes": 112, "entries": [ { "name": "agent-fly-a", - "offset": "2696", + "offset": "2688", "byteLength": "17", "digest": "1321dffb0cdc6f9092cbf7fa2a5fc68bbed12c993d5ad398264012810ce9bf93" }, { "name": "executor-fly-a", - "offset": "2720", + "offset": "2712", "byteLength": "14", "digest": "3aee60df7e29efeba7f5f99fc5867647b36aebff1d5d3c838dbff323122e6462" }, { "name": "task-ledger", - "offset": "2736", + "offset": "2728", "byteLength": "11", "digest": "40b00ed2bbba901d68205ff71b04a44b9ee53c51cb3109aa2ceaa44f1c45727e" }, { "name": "prior-inspection", - "offset": "2752", + "offset": "2744", "byteLength": "10", "digest": "2c13b7b4d9a9916801ab9191c314f31b045e9b9c5b669a6c0474f0217ef75bf5" }, { "name": "world", - "offset": "2768", + "offset": "2760", "byteLength": "64", "digest": "f5a5fd42d16a20302798ef6ed309979b43003d2320d9f0e8ea9831a92759fb4b" } ], - "footerOffset": "2832", + "footerOffset": "2824", "footerBytes": 48, - "totalBytes": "2880" + "totalBytes": "2872" } }, "corruption": [ @@ -182,17 +182,17 @@ }, { "name": "a flipped payload byte", - "offset": 2696, + "offset": 2688, "reason": "every payload carries its own digest" }, { "name": "a flipped footer digest byte", - "offset": 2840, + "offset": 2832, "reason": "the footer digest must match the contents" }, { "name": "a flipped footer magic byte", - "offset": 2872, + "offset": 2864, "reason": "a truncated file cannot look complete" } ] diff --git a/services/flysim/crates/fly-session/README.md b/services/flysim/crates/fly-session/README.md index 33f0594..76b8b72 100644 --- a/services/flysim/crates/fly-session/README.md +++ b/services/flysim/crates/fly-session/README.md @@ -189,12 +189,18 @@ The durable store is `state`, over the `FLYSESS1` layout the contract crate owns `BUSY` a stepping session survives rather than an epoch failure. - **Capture and durability are two events.** `State.Capture` completes when an immutable capture exists; `Coordinator::await_durable` completes when the store manifest rename has - happened, which is the durable commit point. Only the second moves the durable mark. A lost - save reply is `SaveOutcome::ReplyLost`, and `Coordinator::resolve_durable` then asks the - store about the *same* checkpoint instead of saving again. -- **The writer is bounded twice**, by outstanding captures and by queued bytes, and it owns - its payload handles until the bytes are committed or the job fails. A queued *replaceable* - capture is superseded by a later one, releasing its holds; a durable one never is. + happened, which is the durable commit point. Only the second moves the durable mark, and the + three ways it can end without one are told apart: `Failed` (the write stopped), + `ReplyLost` (the write finished and the acknowledgment did not arrive) and + `DeadlineExpired` (the caller's own budget ran out while the save was still going). + `Coordinator::resolve_durable` then asks the store about the *same* checkpoint instead of + saving again. +- **The writer is bounded twice**, and the two bounds refuse at different moments. The + outstanding-capture bound is taken before a capture is requested; the byte budget cannot be, + because a capture's size is not known until it exists, so it refuses at submit and releases + the payloads with the refusal. The writer owns its payload handles until the bytes are + committed or the job fails. A queued *replaceable* capture is superseded by a later one, + releasing its holds; a durable one never is. - **The install is a group.** A restore selects a complete compatible generation, imports its payloads as fresh artifacts, stages every participant, validates the coordinator's own ledgers, and only then activates. A failure anywhere leaves the fence closed, and every diff --git a/services/flysim/crates/fly-session/src/coordinator.rs b/services/flysim/crates/fly-session/src/coordinator.rs index bde1183..f01d72b 100644 --- a/services/flysim/crates/fly-session/src/coordinator.rs +++ b/services/flysim/crates/fly-session/src/coordinator.rs @@ -145,6 +145,12 @@ pub struct Deadlines { /// a capture serializes a participant and a restore validates and installs one, and /// neither is a step whose latency the probe was chosen for. pub capture: Duration, + /// How long a caller waits for a *durable* acknowledgment. + /// + /// It is not [`Deadlines::capture`]: that one bounds a call to a participant, and this + /// one bounds two `fsync`s, a queue the caller shares with other captures and a disk. + /// Reusing the call budget here would make a slow disk look like an unresponsive worker. + pub durable: Duration, } /// How long the resolution waits between attempts. @@ -172,6 +178,7 @@ impl Default for Deadlines { resolve_attempts: 8192, boot: Duration::from_secs(30), capture: Duration::from_secs(30), + durable: Duration::from_secs(60), } } } @@ -3186,14 +3193,17 @@ impl Coordinator { /// operation. pub async fn await_durable(&mut self, ticket: CaptureTicket) -> Outcome { let CaptureTicket { checkpoint_id, boundary, receiver } = ticket; - let budget = self.deadlines.capture; + let budget = self.deadlines.durable; let outcome = crate::state::CheckpointWriter::wait(receiver, &checkpoint_id, budget).await; - if let crate::state::SaveOutcome::Committed { .. } = &outcome { + if outcome.is_durable() { self.durable = Some((checkpoint_id.clone(), boundary)); self.audit.push(format!("durable:{checkpoint_id}@{boundary}")); } else { - self.audit.push(format!("not-durable:{checkpoint_id}@{boundary}")); + // Named rather than lumped together: a failed write, a superseded capture, a lost + // reply and an expired caller budget are four different things to have to explain. + self.audit + .push(format!("not-durable:{}:{checkpoint_id}@{boundary}", outcome.event())); } Ok(outcome) } diff --git a/services/flysim/crates/fly-session/src/state.rs b/services/flysim/crates/fly-session/src/state.rs index 93ba77b..aa42d3c 100644 --- a/services/flysim/crates/fly-session/src/state.rs +++ b/services/flysim/crates/fly-session/src/state.rs @@ -999,10 +999,18 @@ pub enum SaveOutcome { Failed { checkpoint_id: Id, reason: String }, /// A later replaceable capture took this one's place in the queue before it was written. Superseded { checkpoint_id: Id, by: Id }, - /// The writer's reply never arrived. The operation's outcome is unknown from here, so - /// durable metadata does not move; the caller resolves the *same* operation against the - /// store manifest instead of saving again. + /// The writer finished this job and its reply channel was gone before the outcome could + /// be delivered. The write is over and its result is unknown from here. ReplyLost { checkpoint_id: Id }, + /// The caller's own budget ran out while the job was still queued or being written. The + /// save is not over: it may commit after this is reported. + /// + /// It is a different fact from [`SaveOutcome::ReplyLost`] and is named separately because + /// diagnosing one as the other is exactly the implicit best-effort reading these + /// contracts refuse. Both leave durable metadata where it was, and for both the caller + /// resolves the *same* operation against the store manifest instead of saving again -- + /// but only one of them is a save that has already stopped. + DeadlineExpired { checkpoint_id: Id }, } impl SaveOutcome { @@ -1011,17 +1019,54 @@ impl SaveOutcome { SaveOutcome::Committed { checkpoint_id, .. } | SaveOutcome::Failed { checkpoint_id, .. } | SaveOutcome::Superseded { checkpoint_id, .. } - | SaveOutcome::ReplyLost { checkpoint_id } => checkpoint_id, + | SaveOutcome::ReplyLost { checkpoint_id } + | SaveOutcome::DeadlineExpired { checkpoint_id } => checkpoint_id, } } /// The event name this outcome publishes under. + /// + /// Only the three the writer itself produces are ever published; the two caller-side + /// outcomes are what a caller saw, not what the store did, and the store does not announce + /// them on its own topic. pub fn event(&self) -> &'static str { match self { SaveOutcome::Committed { .. } => "committed", SaveOutcome::Failed { .. } => "failed", SaveOutcome::Superseded { .. } => "superseded", - SaveOutcome::ReplyLost { .. } => "failed", + SaveOutcome::ReplyLost { .. } | SaveOutcome::DeadlineExpired { .. } => "failed", + } + } + + /// True only past the durable commit point. + pub fn is_durable(&self) -> bool { + matches!(self, SaveOutcome::Committed { .. }) + } +} + +/// What the writer itself can produce for one job. +/// +/// The two caller-side outcomes -- a lost reply and an expired caller deadline -- are not in +/// here, because the writer cannot observe either. Keeping them out is what stops the writer's +/// own bookkeeping from carrying arms that can never run. +#[derive(Clone, Debug, PartialEq, Eq)] +enum WriteOutcome { + Committed { boundary: u64, file: String }, + Failed { reason: String }, +} + +impl WriteOutcome { + fn into_save(self, checkpoint_id: &Id) -> SaveOutcome { + match self { + WriteOutcome::Committed { boundary, file } => SaveOutcome::Committed { + checkpoint_id: checkpoint_id.clone(), + boundary, + file, + }, + WriteOutcome::Failed { reason } => SaveOutcome::Failed { + checkpoint_id: checkpoint_id.clone(), + reason, + }, } } } @@ -1037,13 +1082,21 @@ pub enum RetryPolicy { RetryThenRelease { attempts: u32 }, } -/// The writer's bounds. Both are finite and both refuse before a capture is requested. +/// The writer's bounds. Both are finite, and they refuse at different moments. +/// +/// [`WriterConfig::queue_capacity`] is the one a capture is refused *before* it is requested: +/// [`CheckpointWriter::reserve`] takes its slot first, which is what the durable row of +/// `state-media-v1` section 3 means by rejecting before capture. The byte budget cannot work +/// that way, because how many bytes a capture is worth is not known until the participants +/// have produced it; it is checked at [`CheckpointWriter::submit`], so an oversized capture is +/// refused after it exists and before it is queued, and its payloads are released with the +/// refusal. Both are named `BUSY` refusals and neither fails the epoch. #[derive(Clone, Copy, Debug)] pub struct WriterConfig { /// Outstanding coherent captures. `state-media-v1` section 3's initial session default - /// is two. + /// is two. Refused before a capture is requested. pub queue_capacity: usize, - /// The total payload bytes the queue may hold. + /// The total payload bytes the queue may hold. Refused at submit, once the size is known. pub max_queued_bytes: u64, pub retry: RetryPolicy, } @@ -1198,48 +1251,57 @@ capture is refused before it is requested rather than queued without bound", ) -> DomainResult> { let bytes = submission.byte_length(); let budget = self.shared.config.max_queued_bytes; - let held = self.shared.queued_bytes.load(Ordering::SeqCst); - if held + bytes > budget { - self.shared - .stats - .lock() - .expect("the writer stats are never poisoned") - .rejected += 1; - return Err(DomainError::before( - ErrorCode::Busy, - format!( - "the checkpoint queue holds {held} of {budget} bytes and this capture adds \ -{bytes}; the byte budget is finite and refuses before it is exceeded" - ), - )); - } let (reply, receiver) = tokio::sync::oneshot::channel(); let checkpoint_id = submission.checkpoint_id.clone(); let queued_event = submission.as_event("queued"); + // Reading the byte total, deciding on it and changing it are one critical section. + // Only one caller submits today, so a split could not be observed -- but a bound that + // is only correct while nobody else is submitting is not a bound. let superseded = { let mut queue = self.shared.queue.lock().expect("the writer queue is never poisoned"); - let replaced = if submission.replaceable { - queue - .iter() - .position(|job| job.submission.replaceable) - .map(|index| queue.remove(index).expect("just found")) + let replaced_index = if submission.replaceable { + queue.iter().position(|job| job.submission.replaceable) } else { None }; + // A capture that will take a queued replaceable one's place frees its bytes, so + // the budget is decided against what the queue will hold and not what it holds. + let freed = replaced_index.map_or(0, |index| queue[index].submission.byte_length()); + let held = self.shared.queued_bytes.load(Ordering::SeqCst); + let after = held.saturating_sub(freed) + bytes; + if after > budget { + self.shared + .stats + .lock() + .expect("the writer stats are never poisoned") + .rejected += 1; + return Err(DomainError::before( + ErrorCode::Busy, + format!( + "the checkpoint queue would hold {after} of {budget} bytes; the byte \ +budget is finite and refuses before it is exceeded" + ), + )); + } + let replaced = replaced_index.map(|index| queue.remove(index).expect("just found")); + if let Some(old) = &replaced { + self.shared + .queued_bytes + .fetch_sub(old.submission.byte_length(), Ordering::SeqCst); + } queue.push_back(Job { submission, reply, permit: reservation.permit }); self.shared.queued_bytes.fetch_add(bytes, Ordering::SeqCst); let mut stats = self.shared.stats.lock().expect("the writer stats are never poisoned"); stats.queued += 1; stats.peak_queue = stats.peak_queue.max(queue.len()); + // Sampled after the superseded job's bytes are gone, so the peak is a total the + // queue really held. stats.peak_bytes = stats .peak_bytes .max(self.shared.queued_bytes.load(Ordering::SeqCst)); replaced }; if let Some(old) = superseded { - self.shared - .queued_bytes - .fetch_sub(old.submission.byte_length(), Ordering::SeqCst); self.shared .stats .lock() @@ -1267,8 +1329,12 @@ capture is refused before it is requested rather than queued without bound", Ok(receiver) } - /// Waits for one save's outcome. A dropped reply channel is a lost save reply, which is - /// an outcome and not a hang. + /// Waits for one save's outcome, within the caller's own budget. + /// + /// The two ways this ends without an outcome are different facts and are reported as + /// themselves: the channel closing means the writer finished and the reply did not reach + /// here, and the budget running out means the save is still going. Neither is a hang and + /// neither is a save. pub async fn wait( receiver: tokio::sync::oneshot::Receiver, checkpoint_id: &Id, @@ -1276,7 +1342,8 @@ capture is refused before it is requested rather than queued without bound", ) -> SaveOutcome { match tokio::time::timeout(budget, receiver).await { Ok(Ok(outcome)) => outcome, - Ok(Err(_)) | Err(_) => SaveOutcome::ReplyLost { checkpoint_id: checkpoint_id.clone() }, + Ok(Err(_)) => SaveOutcome::ReplyLost { checkpoint_id: checkpoint_id.clone() }, + Err(_) => SaveOutcome::DeadlineExpired { checkpoint_id: checkpoint_id.clone() }, } } @@ -1342,6 +1409,10 @@ fn outcome_event(submission: &CaptureSubmission, outcome: &SaveOutcome) -> Map { + payload.insert("reason".into(), "the caller's durable budget expired".into()); + payload.insert("durable".into(), false.into()); + } } payload } @@ -1388,15 +1459,15 @@ async fn run_writer(shared: Arc) { shared .queued_bytes .fetch_sub(submission.byte_length(), Ordering::SeqCst); - let outcome = write_one(&shared, &submission).await; + let written = write_one(&shared, &submission).await; { let mut stats = shared.stats.lock().expect("the writer stats are never poisoned"); - match &outcome { - SaveOutcome::Committed { .. } => stats.committed += 1, - SaveOutcome::Failed { .. } | SaveOutcome::ReplyLost { .. } => stats.failed += 1, - SaveOutcome::Superseded { .. } => stats.superseded += 1, + match &written { + WriteOutcome::Committed { .. } => stats.committed += 1, + WriteOutcome::Failed { .. } => stats.failed += 1, } } + let outcome = written.into_save(&submission.checkpoint_id); publish_event(&shared.events, &shared.stats, &outcome_event(&submission, &outcome)).await; let lost = shared.faults.drop_reply_for.as_ref() == Some(&submission.checkpoint_id); // The writer owned these handles until the bytes were committed or the job failed. @@ -1416,21 +1487,19 @@ async fn run_writer(shared: Arc) { } } -async fn write_one(shared: &Arc, submission: &CaptureSubmission) -> SaveOutcome { +async fn write_one(shared: &Arc, submission: &CaptureSubmission) -> WriteOutcome { let mut payloads = Vec::with_capacity(submission.payloads.len()); for payload in &submission.payloads { let bytes = match payload.artifact.read_all().await { Ok(bytes) => bytes, Err(e) => { - return SaveOutcome::Failed { - checkpoint_id: submission.checkpoint_id.clone(), + return WriteOutcome::Failed { reason: format!("payload {}: {}", payload.name, e.message), }; } }; if bytes.len() as u64 != payload.byte_length || digest_of_bytes(&bytes) != payload.digest { - return SaveOutcome::Failed { - checkpoint_id: submission.checkpoint_id.clone(), + return WriteOutcome::Failed { reason: format!( "payload {} is not the content its capture declared", payload.name @@ -1442,8 +1511,7 @@ async fn write_one(shared: &Arc, submission: &CaptureSubmission) - let bytes = match checkpoint::encode(&submission.manifest, &payloads) { Ok(bytes) => bytes, Err(e) => { - return SaveOutcome::Failed { - checkpoint_id: submission.checkpoint_id.clone(), + return WriteOutcome::Failed { reason: format!("envelope: {}", e.0), }; } @@ -1477,8 +1545,7 @@ async fn write_one(shared: &Arc, submission: &CaptureSubmission) - .await; match result { Ok(Ok(())) => { - return SaveOutcome::Committed { - checkpoint_id: submission.checkpoint_id.clone(), + return WriteOutcome::Committed { boundary: submission.boundary, file: format!("{}.flysess", submission.checkpoint_id), }; @@ -1487,8 +1554,5 @@ async fn write_one(shared: &Arc, submission: &CaptureSubmission) - Err(e) => last = format!("the checkpoint writer stopped: {e}"), } } - SaveOutcome::Failed { - checkpoint_id: submission.checkpoint_id.clone(), - reason: last, - } + WriteOutcome::Failed { reason: last } } diff --git a/services/flysim/crates/fly-session/tests/processes.rs b/services/flysim/crates/fly-session/tests/processes.rs index 17dd0e2..0eaa267 100644 --- a/services/flysim/crates/fly-session/tests/processes.rs +++ b/services/flysim/crates/fly-session/tests/processes.rs @@ -122,6 +122,7 @@ async fn a_slow_participant_is_resolved_rather_than_failed(mode: ExecutionMode) resolve_attempts: 4096, boot: Duration::from_secs(30), capture: Duration::from_secs(30), + durable: Duration::from_secs(60), }; within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap(); let reports = within("run", f.harness.coordinator.run(2)) @@ -193,6 +194,7 @@ async fn a_resolution_says_which_of_its_two_bounds_ended_it(mode: ExecutionMode) resolve_attempts: 8192, boot: Duration::from_secs(30), capture: Duration::from_secs(30), + durable: Duration::from_secs(60), }; within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap(); let started = Instant::now(); @@ -224,6 +226,7 @@ async fn a_resolution_says_which_of_its_two_bounds_ended_it(mode: ExecutionMode) resolve_attempts: 3, boot: Duration::from_secs(30), capture: Duration::from_secs(30), + durable: Duration::from_secs(60), }; within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap(); let failure = within("step", f.harness.coordinator.step()) diff --git a/services/flysim/crates/fly-session/tests/state.rs b/services/flysim/crates/fly-session/tests/state.rs index d7b3a0f..b287c04 100644 --- a/services/flysim/crates/fly-session/tests/state.rs +++ b/services/flysim/crates/fly-session/tests/state.rs @@ -15,6 +15,7 @@ mod common; use std::collections::BTreeSet; use std::path::Path; use std::sync::Arc; +use std::time::Duration; use serde_json::{Value, json}; @@ -391,7 +392,8 @@ async fn a_lost_save_reply_holds_durable_metadata_in_every_mode(mode: ExecutionM a_lost_save_reply(Via::Unix, mode).await; } -/// Two ways a save can end without a saved acknowledgment, and neither moves the mark. +/// Three ways a save can end without a saved acknowledgment. None moves the mark, and each +/// is reported as itself rather than as the others. async fn a_lost_save_reply(via: Via, mode: ExecutionMode) { let lost = ckpt(1); let config = HarnessConfig { @@ -464,6 +466,59 @@ async fn a_lost_save_reply(via: Via, mode: ExecutionMode) { assert_eq!(resolved, None, "an unreferenced generation is never a restore candidate"); assert_eq!(g.harness.coordinator.durable(), None); g.shutdown().await; + + // The third: the caller's own budget runs out while the save is still going. That is a + // different fact from a lost reply -- this save has not stopped -- and it is reported as + // itself, because diagnosing one as the other is the implicit reading these rules refuse. + let slow = ckpt(3); + let gate = Arc::new(tokio::sync::Semaphore::new(0)); + let mut config = HarnessConfig::default(); + config.writer_faults.gate = Some(gate.clone()); + let mut h = fx(via, mode, config).await; + within("bootstrap", h.harness.coordinator.bootstrap()).await.unwrap(); + within("run", h.harness.coordinator.run(1)).await.unwrap(); + // The durable budget is the caller's own and is not the call budget: this shortens the + // wait for an acknowledgment without shortening a single call to a participant. + h.harness.coordinator.deadlines.durable = Duration::from_millis(100); + let ticket = within("capture", h.harness.coordinator.capture(&slow, false)) + .await + .unwrap(); + let outcome = within("durable", h.harness.coordinator.await_durable(ticket)) + .await + .unwrap(); + assert_eq!( + outcome, + SaveOutcome::DeadlineExpired { checkpoint_id: slow.clone() }, + "an expired caller budget is not a lost reply" + ); + assert_ne!(outcome, SaveOutcome::ReplyLost { checkpoint_id: slow.clone() }); + assert_eq!(h.harness.coordinator.durable(), None); + assert_eq!( + count( + &h.harness.coordinator.audit, + &format!("not-durable:failed:{slow}@1") + ), + 1, + "the session records that this capture is not durable: {:?}", + h.harness.coordinator.audit + ); + + // It really was still going: once the writer is let past its gate the same operation + // commits, and resolving it is what moves the mark. + gate.add_permits(16); + let deadline = std::time::Instant::now() + Duration::from_secs(10); + let resolved = loop { + let found = within("resolve", h.harness.coordinator.resolve_durable(&slow)) + .await + .unwrap(); + if found.is_some() || std::time::Instant::now() >= deadline { + break found; + } + tokio::time::sleep(Duration::from_millis(10)).await; + }; + assert_eq!(resolved, Some(1), "the save the caller stopped waiting for still committed"); + assert_eq!(h.harness.coordinator.durable(), Some((slow, 1))); + h.shutdown().await; } // ===============================================================================================