diff --git a/docs/design/session-framework/checkpoint-envelope-v1.md b/docs/design/session-framework/checkpoint-envelope-v1.md index c83bb70..d58ed98 100644 --- a/docs/design/session-framework/checkpoint-envelope-v1.md +++ b/docs/design/session-framework/checkpoint-envelope-v1.md @@ -105,6 +105,31 @@ names; a manifest missing any of them is not a complete checkpoint. | `helperState` | External-helper state required for exact resume, as payload names | | `payloads` | `[{name, byteLength, digest}]`, mirroring the payload table | +**Amendment, 2026-09-22 (STATE-01).** The table above names a holder for every payload except +the environment's own, although section 6's fixture has one (`world`) and a group install has +to map it by name like any other participant's. The manifest therefore also records: + +| Field | Contents | +| --- | --- | +| `environment` | `{workerId, payload}`: which worker the world belonged to and the payload name holding its state | + +The reference implementations' required-field set was also missing `helperState`, which this +section has listed from the start. Both are now in `REQUIRED_MANIFEST_FIELDS` in Rust and in +TypeScript, and the fixture was regenerated by the existing example. The schema set is +untouched, so `contractDigest` is unchanged. + +`coordinator.eventWatermarks` is `{lastSourceStep, issued}`. The fixture illustrated +`{lastEventId, lastOrdinal}`, and it is the illustration that changed: an event id is derived +from the epoch, so a watermark spelled as one cannot be compared across the restore that +gives the session a new epoch, while a source step and an issued count can. + +**A required-manifest-field change is compatibility-relevant and `contractDigest` does not +cover it.** The digest is taken over the schema set, and this manifest is not in it, so +`envelopeVersion` is the only thing that can carry such a change. It stays `1` here only +because no production `FLYSESS1` file exists yet: once one does, adding or removing a required +manifest field **must** bump `envelopeVersion`, because a reader of the older version would +otherwise accept a file it cannot completely read, or refuse one it could. + `payloads` is redundant with the table on purpose: the table is what a reader needs to map bytes, and the manifest is what a store lists, compares and reports without opening the payload area. A reader checks that the two agree. diff --git a/docs/design/session-framework/state-media-v1.md b/docs/design/session-framework/state-media-v1.md index 3342ab8..0285287 100644 --- a/docs/design/session-framework/state-media-v1.md +++ b/docs/design/session-framework/state-media-v1.md @@ -205,6 +205,28 @@ restored time. It cannot advance gameplay to manufacture it. Capture/reconstruct covers render/inspection state and any pending sensor pipeline. Agent state agrees with it; do not replay reward or recalibrate merely to fill missing cached data. +**Amendment, 2026-09-22 (STATE-01).** Three readings of this section, made explicit because +they are now enforced: + +- `compatibilityDigest` on `CaptureResult` and `StageRestoreParams` is the **participant's** + capture compatibility digest of [worker interfaces](workers-v1.md) section 2 -- profile, + resolved seed, numerical model version and effective instance configuration for an agent; + backend, content, patch, controller and parser identity for an environment. It is not the + manifest's `compatibility` block of section 4, which is the composition's and which the + coordinator compares before anything is asked to stage. Both exist because they answer + different questions, and a restore that passed the second could still be handing an agent + another agent's brain. +- The observation `ActivateRestore` returns ran no transition, so it carries **no audio + chunk**, and one in it is refused. Section 2's chunk is the audio of an interval and this + observation covers none; MEDIA-01 implemented that rule as "boundary 0 carries no chunk", + which is true of the only such observation that slice could produce and false of this one. + The rule is about provenance, not about the boundary number. +- A participant that staged into a group install the coordinator then abandoned must be + **replaced** before another restore, exactly as one that activated must. It is holding a + validated replacement state that nothing installed, and [session RPC](ipc-v1.md) section 6 + already refuses to silently reattach such a participant to an active epoch. Without this the + group's second attempt meets its own leftovers and calls them a conflict. + If emulator validation requires mutation, stage a stopped replacement emulator. If that cannot provide externally atomic resume, advertise episode-restart, not exact-checkpoint. After all activation acknowledgments, install the coordinator's staged task/executor/admission state diff --git a/packages/session-types/src/checkpoint.ts b/packages/session-types/src/checkpoint.ts index 63aac19..2df48f7 100644 --- a/packages/session-types/src/checkpoint.ts +++ b/packages/session-types/src/checkpoint.ts @@ -223,7 +223,12 @@ export function decode(input: Uint8Array): Envelope { }; } -/** The manifest fields state-media-v1 section 4 requires. */ +/** + * The manifest fields state-media-v1 section 4 requires. + * + * `helperState` and `environment` join the list under the 2026-09-22 amendment to + * checkpoint-envelope-v1 section 3. + */ export const REQUIRED_MANIFEST_FIELDS = [ 'envelopeVersion', 'checkpointId', @@ -236,6 +241,8 @@ export const REQUIRED_MANIFEST_FIELDS = [ 'compatibility', 'agents', 'coordinator', + 'environment', + 'helperState', 'payloads', ] as const; diff --git a/services/flysim/crates/fly-session-types/examples/update_fixtures.rs b/services/flysim/crates/fly-session-types/examples/update_fixtures.rs index 238287e..bad3693 100644 --- a/services/flysim/crates/fly-session-types/examples/update_fixtures.rs +++ b/services/flysim/crates/fly-session-types/examples/update_fixtures.rs @@ -197,8 +197,9 @@ fn checkpoint_envelope() -> String { "priorInspection": "prior-inspection", "executorState": [{"agentId": "fly-a", "payload": "executor-fly-a"}], "admissionState": null, - "eventWatermarks": {"lastEventId": "evt-1", "lastOrdinal": "7"}, + "eventWatermarks": {"lastSourceStep": "42", "issued": "7"}, }, + "environment": {"workerId": "arena", "payload": "world"}, "helperState": [], "payloads": payload_table(), }); diff --git a/services/flysim/crates/fly-session-types/fixtures/checkpoint-envelope.json b/services/flysim/crates/fly-session-types/fixtures/checkpoint-envelope.json index 6636bae..b726585 100644 --- a/services/flysim/crates/fly-session-types/fixtures/checkpoint-envelope.json +++ b/services/flysim/crates/fly-session-types/fixtures/checkpoint-envelope.json @@ -59,10 +59,14 @@ ], "admissionState": null, "eventWatermarks": { - "lastEventId": "evt-1", - "lastOrdinal": "7" + "lastSourceStep": "42", + "issued": "7" } }, + "environment": { + "workerId": "arena", + "payload": "world" + }, "helperState": [], "payloads": [ { @@ -115,49 +119,49 @@ } ], "envelope": { - "base64": "RkxZU0VTUzEBAAAAIAAAAAAIAAAFAAAAIAgAAAAAAAB7ImFnZW50cyI6W3siYWdlbnRJZCI6ImZseS1hIiwiYnJhaW5UaWNrcyI6IjI1MzQiLCJkYXRhc2V0RGlnZXN0IjoiNmMwYWYxZjA3ODRlZjYzYTM5M2VlNzdkNjE0ZTgyNDZjNjI1MDUxMzYwZjNmMWE0ODgzODM3NGM1ZDM1NWI1MiIsIm1vZGVsVmVyc2lvbiI6ImxpZi0xbXMtZjY0LXYyIiwicGF5bG9hZCI6ImFnZW50LWZseS1hIiwicGxhc3RpY2l0eVZlcnNpb24iOiJmbHkta2MtbWJvbi1yc3RkcC12MiIsInByb2ZpbGVEaWdlc3QiOiIxOTAwZWFiNmMwMjg0ODNkNzEyNjU5OWVlNmY1MGRlMGQyNzkwN2I1YzY1ZmE5MDUyNDU4MGI0YjBmOTg1MmIwIiwicmVtYWluZGVyIjp7ImRlbm9taW5hdG9yIjoiMyIsIm51bWVyYXRvciI6IjEwMDAwMDAifSwic2VlZCI6LTE4NDk0NjA2M31dLCJjaGVja3BvaW50SWQiOiJja3B0LTEiLCJjb21wYXRpYmlsaXR5Ijp7ImJhY2tlbmREaWdlc3QiOiIxMGUwOGE0MTllODUwZWJhMWViYmExOGZkZDI4ZWI3ZWMxYjdlOGJhYTliY2MzYjk3M2UyYjg4OTFlYzcyNmJlIiwiY29udGVudERpZ2VzdCI6ImVkNzAwMmI0MzllOWFjODQ1ZjIyMzU3ZDgyMmJhYzE0NDQ3MzBmYmRiNjAxNmQzZWM5NDMyMjk3YjllYzlmNzMiLCJjb250cm9sbGVyRGlnZXN0IjoiYzE0NzIxMzViMTRjNzdjOGJlZjk4ZTczZjcwMjA4MzI1ZmEwZGNmMWU2YmQ2NjhhZTliMzFhOWNlYTI5NWZlNyIsInBhcnNlckRpZ2VzdCI6ImIxN2Q0NTEyMTE1MDkyOGYyMTQ2YWY0OWUxOTVlZmYxZWVmNWQ2NzMyNWJlMjczYTczM2ZiNzRhY2FkYWEzNDIiLCJwYXRjaERpZ2VzdCI6ImE0ODk1ZWI0NGFmYzMzNmZlY2JiYTZlNTIwY2Q2N2UxNzhkYWNlMDI3NjY1NWQxMDJmY2VmZmE4ZTVmNzA1NzAiLCJzdGF0ZUZvcm1hdElkIjoiZmx5c2Vzcy0xIn0sImNvbXBvc2l0aW9uRGlnZXN0IjoiNzMwZDcyNWM4YTU5ZDNhNzMwM2RlZjJiZWQwNDFhNTc3ZWRiNDI1NWFhYmQ0ODg5Y2UxMjkxODMxMWQ5NTJmMCIsImNvb3JkaW5hdG9yIjp7ImFkbWlzc2lvblN0YXRlIjpudWxsLCJldmVudFdhdGVybWFya3MiOnsibGFzdEV2ZW50SWQiOiJldnQtMSIsImxhc3RPcmRpbmFsIjoiNyJ9LCJleGVjdXRvclN0YXRlIjpbeyJhZ2VudElkIjoiZmx5LWEiLCJwYXlsb2FkIjoiZXhlY3V0b3ItZmx5LWEifV0sInByaW9ySW5zcGVjdGlvbiI6InByaW9yLWluc3BlY3Rpb24iLCJ0YXNrTGVkZ2VyIjoidGFzay1sZWRnZXIifSwiZW52ZWxvcGVWZXJzaW9uIjoxLCJlcGlzb2RlSWQiOiJlcGlzb2RlLTEiLCJoZWxwZXJTdGF0ZSI6W10sInBheWxvYWRzIjpbeyJieXRlTGVuZ3RoIjoiMTciLCJkaWdlc3QiOiIxMzIxZGZmYjBjZGM2ZjkwOTJjYmY3ZmEyYTVmYzY4YmJlZDEyYzk5M2Q1YWQzOTgyNjQwMTI4MTBjZTliZjkzIiwibmFtZSI6ImFnZW50LWZseS1hIn0seyJieXRlTGVuZ3RoIjoiMTQiLCJkaWdlc3QiOiIzYWVlNjBkZjdlMjllZmViYTdmNWY5OWZjNTg2NzY0N2IzNmFlYmZmMWQ1ZDNjODM4ZGJmZjMyMzEyMmU2NDYyIiwibmFtZSI6ImV4ZWN1dG9yLWZseS1hIn0seyJieXRlTGVuZ3RoIjoiMTEiLCJkaWdlc3QiOiI0MGIwMGVkMmJiYmE5MDFkNjgyMDVmZjcxYjA0YTQ0YjllZTUzYzUxY2IzMTA5YWEyY2VhYTQ0ZjFjNDU3MjdlIiwibmFtZSI6InRhc2stbGVkZ2VyIn0seyJieXRlTGVuZ3RoIjoiMTAiLCJkaWdlc3QiOiIyYzEzYjdiNGQ5YTk5MTY4MDFhYjkxOTFjMzE0ZjMxYjA0NWU5YjljNWI2NjlhNmMwNDc0ZjAyMTdlZjc1YmY1IiwibmFtZSI6InByaW9yLWluc3BlY3Rpb24ifSx7ImJ5dGVMZW5ndGgiOiI2NCIsImRpZ2VzdCI6ImY1YTVmZDQyZDE2YTIwMzAyNzk4ZWY2ZWQzMDk5NzliNDMwMDNkMjMyMGQ5ZjBlOGVhOTgzMWE5Mjc1OWZiNGIiLCJuYW1lIjoid29ybGQifV0sInBvcnRNYXAiOlt7ImFnZW50SWQiOiJmbHktYSIsInBvcnRJZCI6InBvcnQtMSJ9XSwic2NoZWR1bGVySWQiOiJsb2Nrc3RlcC12MSIsInNvdXJjZVNjb3BlIjp7ImVwb2NoIjoiZXBvY2gtMSIsInNlc3Npb25JZCI6ImRlbW8iLCJzdGVwIjoiNDIifSwid29ybGRUaW1lIjp7ImRlbm9taW5hdG9yIjoiMSIsIm51bWVyYXRvciI6IjcwMDAwMDAwMCJ9fWFnZW50LWZseS1hAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAABQCgAAAAAAABEAAAAAAAAAEyHf+wzcb5CSy/f6Kl/Gi77RLJk9WtOYJkASgQzpv5NleGVjdXRvci1mbHktYQAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAaAoAAAAAAAAOAAAAAAAAADruYN9+Ke/rp/X5n8WGdkezauv/HV08g42/8yMSLmRidGFzay1sZWRnZXIAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAHgKAAAAAAAACwAAAAAAAABAsA7Su7qQHWggX/cbBKRLnuU8UcsxCaos6qRPHEVyfnByaW9yLWluc3BlY3Rpb24AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACICgAAAAAAAAoAAAAAAAAALBO3tNmpkWgBq5GRwxTzGwRem5xbZppsBHTwIX73W/V3b3JsZAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAmAoAAAAAAABAAAAAAAAAAPWl/ULRaiAwJ5jvbtMJl5tDAD0jINnw6OqYMaknWftLYWdlbnQgc3RhdGUgYnl0ZXMAAAAAAAAAZXhlY3V0b3Igc3RhdGUAAHsicmFuayI6MTB9AAAAAAB7Im1hcCI6NDB9AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAgLAAAAAAAAq++fEx+FvDZho/eB4imbENN4HZrGNC2OCAsI7/gp9r5GTFlTRVNTRg==", - "byteLength": 2824, + "base64": "RkxZU0VTUzEBAAAAIAAAADAIAAAFAAAAUAgAAAAAAAB7ImFnZW50cyI6W3siYWdlbnRJZCI6ImZseS1hIiwiYnJhaW5UaWNrcyI6IjI1MzQiLCJkYXRhc2V0RGlnZXN0IjoiNmMwYWYxZjA3ODRlZjYzYTM5M2VlNzdkNjE0ZTgyNDZjNjI1MDUxMzYwZjNmMWE0ODgzODM3NGM1ZDM1NWI1MiIsIm1vZGVsVmVyc2lvbiI6ImxpZi0xbXMtZjY0LXYyIiwicGF5bG9hZCI6ImFnZW50LWZseS1hIiwicGxhc3RpY2l0eVZlcnNpb24iOiJmbHkta2MtbWJvbi1yc3RkcC12MiIsInByb2ZpbGVEaWdlc3QiOiIxOTAwZWFiNmMwMjg0ODNkNzEyNjU5OWVlNmY1MGRlMGQyNzkwN2I1YzY1ZmE5MDUyNDU4MGI0YjBmOTg1MmIwIiwicmVtYWluZGVyIjp7ImRlbm9taW5hdG9yIjoiMyIsIm51bWVyYXRvciI6IjEwMDAwMDAifSwic2VlZCI6LTE4NDk0NjA2M31dLCJjaGVja3BvaW50SWQiOiJja3B0LTEiLCJjb21wYXRpYmlsaXR5Ijp7ImJhY2tlbmREaWdlc3QiOiIxMGUwOGE0MTllODUwZWJhMWViYmExOGZkZDI4ZWI3ZWMxYjdlOGJhYTliY2MzYjk3M2UyYjg4OTFlYzcyNmJlIiwiY29udGVudERpZ2VzdCI6ImVkNzAwMmI0MzllOWFjODQ1ZjIyMzU3ZDgyMmJhYzE0NDQ3MzBmYmRiNjAxNmQzZWM5NDMyMjk3YjllYzlmNzMiLCJjb250cm9sbGVyRGlnZXN0IjoiYzE0NzIxMzViMTRjNzdjOGJlZjk4ZTczZjcwMjA4MzI1ZmEwZGNmMWU2YmQ2NjhhZTliMzFhOWNlYTI5NWZlNyIsInBhcnNlckRpZ2VzdCI6ImIxN2Q0NTEyMTE1MDkyOGYyMTQ2YWY0OWUxOTVlZmYxZWVmNWQ2NzMyNWJlMjczYTczM2ZiNzRhY2FkYWEzNDIiLCJwYXRjaERpZ2VzdCI6ImE0ODk1ZWI0NGFmYzMzNmZlY2JiYTZlNTIwY2Q2N2UxNzhkYWNlMDI3NjY1NWQxMDJmY2VmZmE4ZTVmNzA1NzAiLCJzdGF0ZUZvcm1hdElkIjoiZmx5c2Vzcy0xIn0sImNvbXBvc2l0aW9uRGlnZXN0IjoiNzMwZDcyNWM4YTU5ZDNhNzMwM2RlZjJiZWQwNDFhNTc3ZWRiNDI1NWFhYmQ0ODg5Y2UxMjkxODMxMWQ5NTJmMCIsImNvb3JkaW5hdG9yIjp7ImFkbWlzc2lvblN0YXRlIjpudWxsLCJldmVudFdhdGVybWFya3MiOnsiaXNzdWVkIjoiNyIsImxhc3RTb3VyY2VTdGVwIjoiNDIifSwiZXhlY3V0b3JTdGF0ZSI6W3siYWdlbnRJZCI6ImZseS1hIiwicGF5bG9hZCI6ImV4ZWN1dG9yLWZseS1hIn1dLCJwcmlvckluc3BlY3Rpb24iOiJwcmlvci1pbnNwZWN0aW9uIiwidGFza0xlZGdlciI6InRhc2stbGVkZ2VyIn0sImVudmVsb3BlVmVyc2lvbiI6MSwiZW52aXJvbm1lbnQiOnsicGF5bG9hZCI6IndvcmxkIiwid29ya2VySWQiOiJhcmVuYSJ9LCJlcGlzb2RlSWQiOiJlcGlzb2RlLTEiLCJoZWxwZXJTdGF0ZSI6W10sInBheWxvYWRzIjpbeyJieXRlTGVuZ3RoIjoiMTciLCJkaWdlc3QiOiIxMzIxZGZmYjBjZGM2ZjkwOTJjYmY3ZmEyYTVmYzY4YmJlZDEyYzk5M2Q1YWQzOTgyNjQwMTI4MTBjZTliZjkzIiwibmFtZSI6ImFnZW50LWZseS1hIn0seyJieXRlTGVuZ3RoIjoiMTQiLCJkaWdlc3QiOiIzYWVlNjBkZjdlMjllZmViYTdmNWY5OWZjNTg2NzY0N2IzNmFlYmZmMWQ1ZDNjODM4ZGJmZjMyMzEyMmU2NDYyIiwibmFtZSI6ImV4ZWN1dG9yLWZseS1hIn0seyJieXRlTGVuZ3RoIjoiMTEiLCJkaWdlc3QiOiI0MGIwMGVkMmJiYmE5MDFkNjgyMDVmZjcxYjA0YTQ0YjllZTUzYzUxY2IzMTA5YWEyY2VhYTQ0ZjFjNDU3MjdlIiwibmFtZSI6InRhc2stbGVkZ2VyIn0seyJieXRlTGVuZ3RoIjoiMTAiLCJkaWdlc3QiOiIyYzEzYjdiNGQ5YTk5MTY4MDFhYjkxOTFjMzE0ZjMxYjA0NWU5YjljNWI2NjlhNmMwNDc0ZjAyMTdlZjc1YmY1IiwibmFtZSI6InByaW9yLWluc3BlY3Rpb24ifSx7ImJ5dGVMZW5ndGgiOiI2NCIsImRpZ2VzdCI6ImY1YTVmZDQyZDE2YTIwMzAyNzk4ZWY2ZWQzMDk5NzliNDMwMDNkMjMyMGQ5ZjBlOGVhOTgzMWE5Mjc1OWZiNGIiLCJuYW1lIjoid29ybGQifV0sInBvcnRNYXAiOlt7ImFnZW50SWQiOiJmbHktYSIsInBvcnRJZCI6InBvcnQtMSJ9XSwic2NoZWR1bGVySWQiOiJsb2Nrc3RlcC12MSIsInNvdXJjZVNjb3BlIjp7ImVwb2NoIjoiZXBvY2gtMSIsInNlc3Npb25JZCI6ImRlbW8iLCJzdGVwIjoiNDIifSwid29ybGRUaW1lIjp7ImRlbm9taW5hdG9yIjoiMSIsIm51bWVyYXRvciI6IjcwMDAwMDAwMCJ9fWFnZW50LWZseS1hAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACACgAAAAAAABEAAAAAAAAAEyHf+wzcb5CSy/f6Kl/Gi77RLJk9WtOYJkASgQzpv5NleGVjdXRvci1mbHktYQAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAmAoAAAAAAAAOAAAAAAAAADruYN9+Ke/rp/X5n8WGdkezauv/HV08g42/8yMSLmRidGFzay1sZWRnZXIAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAKgKAAAAAAAACwAAAAAAAABAsA7Su7qQHWggX/cbBKRLnuU8UcsxCaos6qRPHEVyfnByaW9yLWluc3BlY3Rpb24AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAC4CgAAAAAAAAoAAAAAAAAALBO3tNmpkWgBq5GRwxTzGwRem5xbZppsBHTwIX73W/V3b3JsZAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAyAoAAAAAAABAAAAAAAAAAPWl/ULRaiAwJ5jvbtMJl5tDAD0jINnw6OqYMaknWftLYWdlbnQgc3RhdGUgYnl0ZXMAAAAAAAAAZXhlY3V0b3Igc3RhdGUAAHsicmFuayI6MTB9AAAAAAB7Im1hcCI6NDB9AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAADgLAAAAAAAAX4L9WkdX0MViD3h5YJgf7VQgwocWU4XJoKX6MUwl+8hGTFlTRVNTRg==", + "byteLength": 2872, "layout": { "headerBytes": 32, "manifestOffset": "32", - "manifestBytes": 2048, - "tableOffset": "2080", + "manifestBytes": 2096, + "tableOffset": "2128", "tableEntryBytes": 112, "entries": [ { "name": "agent-fly-a", - "offset": "2640", + "offset": "2688", "byteLength": "17", "digest": "1321dffb0cdc6f9092cbf7fa2a5fc68bbed12c993d5ad398264012810ce9bf93" }, { "name": "executor-fly-a", - "offset": "2664", + "offset": "2712", "byteLength": "14", "digest": "3aee60df7e29efeba7f5f99fc5867647b36aebff1d5d3c838dbff323122e6462" }, { "name": "task-ledger", - "offset": "2680", + "offset": "2728", "byteLength": "11", "digest": "40b00ed2bbba901d68205ff71b04a44b9ee53c51cb3109aa2ceaa44f1c45727e" }, { "name": "prior-inspection", - "offset": "2696", + "offset": "2744", "byteLength": "10", "digest": "2c13b7b4d9a9916801ab9191c314f31b045e9b9c5b669a6c0474f0217ef75bf5" }, { "name": "world", - "offset": "2712", + "offset": "2760", "byteLength": "64", "digest": "f5a5fd42d16a20302798ef6ed309979b43003d2320d9f0e8ea9831a92759fb4b" } ], - "footerOffset": "2776", + "footerOffset": "2824", "footerBytes": 48, - "totalBytes": "2824" + "totalBytes": "2872" } }, "corruption": [ @@ -178,17 +182,17 @@ }, { "name": "a flipped payload byte", - "offset": 2640, + "offset": 2688, "reason": "every payload carries its own digest" }, { "name": "a flipped footer digest byte", - "offset": 2784, + "offset": 2832, "reason": "the footer digest must match the contents" }, { "name": "a flipped footer magic byte", - "offset": 2816, + "offset": 2864, "reason": "a truncated file cannot look complete" } ] diff --git a/services/flysim/crates/fly-session-types/src/checkpoint.rs b/services/flysim/crates/fly-session-types/src/checkpoint.rs index 341f5c5..985fcb1 100644 --- a/services/flysim/crates/fly-session-types/src/checkpoint.rs +++ b/services/flysim/crates/fly-session-types/src/checkpoint.rs @@ -296,6 +296,11 @@ pub fn decode(bytes: &[u8]) -> Result { /// The manifest fields state-media-v1 section 4 requires, checked as a set: a manifest that /// omits one of them is not a complete checkpoint. +/// +/// `helperState` and `environment` join the list under the 2026-09-22 amendment to +/// checkpoint-envelope-v1 section 3: the first has been in that section's table from the +/// start and was missing here, and the second is the holder of the world's own payload, which +/// the table named for every other participant and not for the environment. pub const REQUIRED_MANIFEST_FIELDS: &[&str] = &[ "envelopeVersion", "checkpointId", @@ -308,6 +313,8 @@ pub const REQUIRED_MANIFEST_FIELDS: &[&str] = &[ "compatibility", "agents", "coordinator", + "environment", + "helperState", "payloads", ]; diff --git a/services/flysim/crates/fly-session/README.md b/services/flysim/crates/fly-session/README.md index c3d0ca9..9b8ee00 100644 --- a/services/flysim/crates/fly-session/README.md +++ b/services/flysim/crates/fly-session/README.md @@ -45,6 +45,7 @@ Ready(k) ─ Prepare all agents concurrently ─────────── | `metrics` | Latency percentiles and the machine's core and memory counters | | `measure` | The execution-mode comparison of the guide's section 5 | | `cli` | The binary's subcommands: `agent`, `environment`, `measure` | +| `state` | The durable checkpoint store over `FLYSESS1`: compatibility, generations, the bounded writer | | `harness` | The runnable composition: router, the flies, one arena, one coordinator | ## Execution modes and the launcher @@ -179,6 +180,44 @@ harness.shutdown().await; event ids derived from epoch, source step, rule and ordinal. - **Executors.** The stateless identity executor only, as v1 specifies. +## Checkpoints and recovery + +The durable store is `state`, over the `FLYSESS1` layout the contract crate owns. + +- **One boundary, every participant.** `Coordinator::capture` runs at `Ready(k)` or + `Paused(k)` only. It takes its queue slot *before* the first `State.Capture`, so a saturated + writer refuses the capture rather than queueing it without bound, and the refusal is a + `BUSY` a stepping session survives rather than an epoch failure. +- **Capture and durability are two events.** `State.Capture` completes when an immutable + capture exists; `Coordinator::await_durable` completes when the store manifest rename has + happened, which is the durable commit point. Only the second moves the durable mark, and the + three ways it can end without one are told apart: `Failed` (the write stopped), + `ReplyLost` (the write finished and the acknowledgment did not arrive) and + `DeadlineExpired` (the caller's own budget ran out while the save was still going). + `Coordinator::resolve_durable` then asks the store about the *same* checkpoint instead of + saving again. +- **The writer is bounded twice**, and the two bounds refuse at different moments. The + outstanding-capture bound is taken before a capture is requested; the byte budget cannot be, + because a capture's size is not known until it exists, so it refuses at submit and releases + the payloads with the refusal. The writer owns its payload handles until the bytes are + committed or the job fails. A queued *replaceable* capture is superseded by a later one, + releasing its holds; a durable one never is. +- **The install is a group.** A restore selects a complete compatible generation, imports its + payloads as fresh artifacts, stages every participant, validates the coordinator's own + ledgers, and only then activates. A failure anywhere leaves the fence closed, and every + participant that got as far as staging is recorded as one that must be replaced before + another restore is attempted. +- **The fence lifts once.** `Failed -> Restoring(k) -> Paused(k)`, at the end of a complete + install and nowhere else. A fenced session takes no step, publishes nothing, captures + nothing and holds no artifact handle. +- **Nothing old crosses.** The fence drops every media handle; the restore imports fresh + artifacts; the environment re-renders its pending sensor pipeline from recorded + reconstruction inputs; and the new epoch's first audio chunk resumes the preserved sample + position and marks the discontinuity. +- **Epoch metadata in a trace.** `scope.epoch`, the batch id and every task event id are + derived from the epoch, so a resumed run's behaviour is compared through + `EpochRebase`, which rewrites exactly those and fails on anything it does not recognise. + ## Where this crate narrows or adds to the contract crate - **Required views.** `WorldObservation::validate_against` checks the views a result carries @@ -220,9 +259,9 @@ them. `implementation.md` sequences those after this slice and together with eac - **Fake workers.** There is no neural model and no emulator. What is modelled exactly is the ordering, the identity rules and the retry rules, not any numerical behaviour. -- **No state methods.** `State.Capture`, `State.StageRestore` and `State.ActivateRestore` are - STATE-01. The phase machine has their edges (`Capturing`, `Restoring`) and the workers do not - advertise them as implemented methods. +- **One environment, one task.** A checkpoint records the composition it was taken from, and a + restore refuses one taken under another backend, content, patch, controller or parser + identity. It does not migrate between compositions, and it does not try. - **No audience input.** The admitted pre-step stimulation list exists and is always empty. - **One descriptor revision.** A revision changes when the composition does, and the only in-session path to that is a group restore into a fresh epoch, which is STATE-01's. The @@ -291,11 +330,13 @@ The three integration suites do not all run over both transports, and cannot: - `tests/processes.rs` runs over the Unix socket only, in all three execution modes. A participant in a process of its own has no in-memory transport to reach the router by, so the mode is the axis that suite varies and the transport is fixed. -- `tests/media.rs` and `tests/publishing.rs` run over both transports, and each also generates - a subset once per execution mode. The publication boundary lives in the coordinator, so - unlike the render counter and the sensor log it crosses no process boundary and stays fully - observable in all three modes; `the_publication_boundary_holds_in_every_execution_mode` - asserts that rather than assuming it. +- `tests/media.rs`, `tests/state.rs` and `tests/publishing.rs` run over both transports *and* + in the execution modes: each acceptance body is written once and registered twice, by + `both_transports!` in the in-process composition and by `all_modes!` over the socket. + `tests/publishing.rs` registers a subset that way rather than all of it, because the + publication boundary lives in the coordinator: unlike the render counter and the sensor log + it crosses no process boundary and stays fully observable in all three modes, which + `the_publication_boundary_holds_in_every_execution_mode` asserts rather than assumes. - `tests/session.rs`: one world advance per complete batch; every agent Prepared before the advance; one task evaluation per transition; every agent committed before the next Prepare or @@ -320,6 +361,14 @@ The three integration suites do not all run over both transports, and cannot: committed action being the transition that just ended, one snapshot carrying every agent, application-owned state and cues, a held event batch, and the read-only query service -- plus the first two generated once per execution mode by `all_modes!`. +- `tests/state.rs`: the STATE-01 acceptance bullets -- an uninterrupted run and a resumed run + committing the same behaviour once the epoch metadata is rebased, a corrupt payload failing + the install as a group for every participant and for the coordinator's own ledger, a lost + save reply and an uncommitted store manifest both leaving the durable mark where it was, a + refused activation resuming no part of the world, the capture queue staying bounded under a + stalled writer, and old media and another parser's state failing to cross a recovery -- + plus the once-only restore token, the superseded replaceable capture, and the fence that + lifts only through a complete restore. - `tests/failures.rs`: a duplicate Prepare after a lost reply; a duplicate Commit; the same batch with altered controls; a lost Advance result; a cached artifact consumed by its first caller; one Commit failing after another succeeded; a replaced registration; a reply from diff --git a/services/flysim/crates/fly-session/src/agent.rs b/services/flysim/crates/fly-session/src/agent.rs index aa59bd3..21ca1a5 100644 --- a/services/flysim/crates/fly-session/src/agent.rs +++ b/services/flysim/crates/fly-session/src/agent.rs @@ -7,7 +7,7 @@ //! stimulation, then reinforces once, and executes no tick at all. Every mutating step bumps //! one counter, which is how a test proves a duplicate request changed nothing. -use std::collections::BTreeMap; +use std::collections::{BTreeMap, BTreeSet}; use serde_json::Value; @@ -193,6 +193,12 @@ pub struct AgentFaults { pub prepare_delay_ms: u64, /// Hold `Agent.Commit` open for this long. pub commit_delay_ms: u64, + /// Refuse `State.StageRestore`, so a group install meets one participant that will not + /// validate while the others already have. + pub fail_stage_restore: bool, + /// Refuse `State.ActivateRestore` after this worker has already staged, so a group meets + /// a failure halfway through activation. + pub fail_activate_restore: bool, } /// One fake agent worker's configuration. @@ -238,6 +244,10 @@ pub struct FakeAgentWorker { context: Option, context_digest: Option, prepared: Option<(DomainRequestId, PreparedDecision)>, + /// A validated replacement state that the live session cannot see yet. + staged: Option, + /// Restore tokens this worker has activated. A token activates once. + activated: BTreeSet, } impl FakeAgentWorker { @@ -252,10 +262,17 @@ impl FakeAgentWorker { context: None, context_digest: None, prepared: None, + staged: None, + activated: BTreeSet::new(), config, } } + /// True while a validated replacement state is staged and not yet activated. + pub fn has_staged_restore(&self) -> bool { + self.staged.is_some() + } + pub fn status(&self) -> StatusCell { self.status.clone() } @@ -666,7 +683,11 @@ impl WorkerEndpoint for FakeAgentWorker { } fn capabilities(&self) -> Vec { - vec![id("agent-step-v1"), id("pixel-observation-v1")] + vec![ + id("agent-step-v1"), + id("pixel-observation-v1"), + id(crate::state::CHECKPOINT_CAPABILITY), + ] } fn status_cell(&self) -> StatusCell { @@ -678,7 +699,14 @@ impl WorkerEndpoint for FakeAgentWorker { } fn methods(&self) -> Vec<&'static str> { - vec!["Agent.Initialize", "Agent.Prepare", "Agent.Commit"] + vec![ + "Agent.Initialize", + "Agent.Prepare", + "Agent.Commit", + "State.Capture", + "State.StageRestore", + "State.ActivateRestore", + ] } fn handle<'a>(&'a mut self, ctx: HandlerCtx<'a>) -> BoxFuture<'a, DomainResult> { @@ -687,6 +715,9 @@ impl WorkerEndpoint for FakeAgentWorker { "Agent.Initialize" => self.initialize(&ctx).await, "Agent.Prepare" => self.prepare(&ctx).await, "Agent.Commit" => self.commit(&ctx).await, + "State.Capture" => self.state_capture(&ctx).await, + "State.StageRestore" => self.state_stage_restore(&ctx).await, + "State.ActivateRestore" => self.state_activate_restore(&ctx).await, other => Err(DomainError::before( ErrorCode::Unsupported, format!("{other} is not an agent method"), @@ -699,7 +730,12 @@ impl WorkerEndpoint for FakeAgentWorker { /// The retention class table an agent endpoint follows, for a caller that wants it. pub fn agent_op_class(method: &str) -> Option { match method { - "Agent.Initialize" => Some(OpClass::Lifecycle), + // `ipc-v1` section 5: lifecycle *and capture* replies are retained until + // `Worker.Acknowledge`, which is also what lets a duplicate restore request replay + // its cached reply rather than staging or activating twice. + "Agent.Initialize" | "State.Capture" | "State.StageRestore" | "State.ActivateRestore" => { + Some(OpClass::Lifecycle) + } "Agent.Prepare" | "Agent.Commit" => Some(OpClass::StepMutation), _ => None, } @@ -769,3 +805,532 @@ pub fn synthetic_profile(agent_id: &Id, tick_duration: &RationalNs, warmup_ticks /// The per-agent contexts a bootstrap produced, keyed by agent id. pub type Contexts = BTreeMap; + +// ------------------------------------------------------------------------------------------- +// STATE-01: capture and restore + +/// The numerical model version this worker implements. It is part of a capture's +/// compatibility identity: the same profile and seed under another model is not the same +/// state (`workers-v1` section 2). +pub const MODEL_VERSION: &str = "fake-lcg-v1"; + +/// The plasticity rule version, for the same reason. +pub const PLASTICITY_VERSION: &str = "fake-reinforce-v1"; + +/// The version this payload layout is written and read under. +pub const AGENT_PAYLOAD_VERSION: u64 = 1; + +/// The dataset identity a synthetic agent resolves. +/// +/// There is no connectome dataset behind this worker, and a checkpoint says so with a stable +/// identity rather than omitting the field: "no dataset" has to be distinguishable from "the +/// dataset was not recorded". +pub fn dataset_digest() -> Digest { + digest_of_bytes(b"fly-session/no-dataset-v1") +} + +/// The capture compatibility digest of one agent (`workers-v1` section 2). +/// +/// The profile digest identifies the profile definition; this additionally covers the +/// resolved seed, the numerical model version and the plasticity rule, because two agents +/// with the same profile digest and different seeds hold state that is not interchangeable. +/// Every field it covers is one the checkpoint manifest already records in that agent's row, +/// so a restore derives the expected digest from the manifest rather than from the payload it +/// is about to validate. +pub fn agent_compatibility_digest( + agent_id: &Id, + profile_digest: &Digest, + dataset_digest: &Digest, + model_version: &str, + plasticity_version: &str, + seed: i32, +) -> Digest { + let value = serde_json::json!({ + "agentId": agent_id.as_str(), + "profileDigest": profile_digest.as_str(), + "datasetDigest": dataset_digest.as_str(), + "modelVersion": model_version, + "plasticityVersion": plasticity_version, + "seed": seed, + }); + digest_of(&value).expect("an agent compatibility block canonicalizes") +} + +impl FakeModel { + /// Every field of the model, so a resumed agent is this agent and not a fresh one. + fn capture(&self) -> Value { + serde_json::json!({ + "seed": self.seed, + "state": self.state.to_string(), + "mutations": self.mutations.to_string(), + "ticks": self.ticks.to_string(), + "stimulations": self.stimulations.to_string(), + "reinforcements": self.reinforcements.to_string(), + "learningEnabled": self.learning_enabled, + "learningUpdates": self.learning_updates.to_string(), + "learningChanged": self.learning_changed.to_string(), + "lastSignal": self.last_signal, + "inputValue": self.input_value.to_string(), + "inputInstalls": self.input_installs.to_string(), + }) + } + + fn restored(value: &Value) -> DomainResult { + let number = |key: &str| -> DomainResult { + value + .get(key) + .and_then(Value::as_str) + .ok_or_else(|| incompatible(format!("the agent payload has no {key}")))? + .parse::() + .map_err(|_| incompatible(format!("the agent payload's {key} is not a U64"))) + }; + let seed = value + .get("seed") + .and_then(Value::as_i64) + .and_then(|v| i32::try_from(v).ok()) + .ok_or_else(|| incompatible("the agent payload has no seed"))?; + let input_value = value + .get("inputValue") + .and_then(Value::as_str) + .ok_or_else(|| incompatible("the agent payload has no inputValue"))? + .parse::() + .map_err(|_| incompatible("the agent payload's inputValue is not an integer"))?; + let last_signal = value + .get("lastSignal") + .and_then(Value::as_f64) + .filter(|v| v.is_finite()) + .ok_or_else(|| incompatible("the agent payload's lastSignal is not finite"))?; + let learning_enabled = value + .get("learningEnabled") + .and_then(Value::as_bool) + .ok_or_else(|| incompatible("the agent payload has no learningEnabled"))?; + Ok(FakeModel { + seed, + state: number("state")?, + mutations: number("mutations")?, + ticks: number("ticks")?, + stimulations: number("stimulations")?, + reinforcements: number("reinforcements")?, + learning_enabled, + learning_updates: number("learningUpdates")?, + learning_changed: number("learningChanged")?, + last_signal, + input_value, + input_installs: number("inputInstalls")?, + }) + } +} + +fn incompatible(message: impl std::fmt::Display) -> DomainError { + DomainError::before(ErrorCode::IncompatibleState, message) +} + +/// One staged restore, held outside the live agent until it is activated. +struct StagedAgent { + token: Id, + checkpoint_id: Id, + scope: Scope, + model: FakeModel, + accumulator: TickAccumulator, + context: TypedValue, + profile: AssetRef, + committed_step: u64, +} + +impl FakeAgentWorker { + /// This worker's own compatibility identity, from its configuration and a resolved seed. + fn compatibility_digest(&self, profile: &AssetRef, seed: i32) -> Digest { + agent_compatibility_digest( + &self.config.agent_id, + &profile.digest, + &dataset_digest(), + MODEL_VERSION, + PLASTICITY_VERSION, + seed, + ) + } + + /// `State.Capture`: an immutable snapshot of this agent at its committed boundary. + /// + /// It is allowed at `Ready(k)` only. A Prepared agent holds half a transition, and there + /// is no coherent boundary to file that under. + async fn state_capture(&mut self, ctx: &HandlerCtx<'_>) -> DomainResult { + let scope = ctx.scope()?.clone(); + self.check_epoch(&scope)?; + let AgentPhase::Ready(k) = self.phase.clone() else { + return Err(DomainError::before( + ErrorCode::InvalidPhase, + format!( + "State.Capture needs a quiescent Ready(k); this worker is {:?}", + self.phase + ), + )); + }; + if scope.step != k { + return Err(DomainError::before( + if scope.step < k { ErrorCode::StaleStep } else { ErrorCode::FutureStep }, + "State.Capture names a boundary this worker is not at", + )); + } + let params: CaptureParams = ctx.params()?; + let profile = self.profile.clone().expect("initialized"); + let context = self.context.clone().expect("initialized"); + let accumulator = self.accumulator.as_ref().expect("initialized"); + let previous = self.status.state(); + self.status.set_state(WorkerState::Capturing); + let payload = serde_json::json!({ + "payloadVersion": AGENT_PAYLOAD_VERSION, + "kind": "agent", + "agentId": self.config.agent_id.as_str(), + "checkpointId": params.checkpoint_id.as_str(), + "sourceScope": scope.to_json(), + "committedStep": k.to_string(), + "profile": profile.to_json(), + "modelVersion": MODEL_VERSION, + "plasticityVersion": PLASTICITY_VERSION, + "datasetDigest": dataset_digest().as_str(), + "model": self.model.capture(), + "accumulator": { + "tickDuration": accumulator.tick_duration().to_json(), + "remainder": accumulator.remainder().to_json(), + "executedTicks": accumulator.executed_ticks().to_string(), + "warmupOffset": accumulator.warmup_offset().to_string(), + }, + "context": context.to_json(), + }); + let bytes = canonicalize(&payload) + .map_err(|e| DomainError::invalid(format!("State.Capture: {}", e.0)))? + .into_bytes(); + let digest = digest_of_bytes(&bytes); + let artifact = crate::state::seal_payload(ctx.client, &bytes, &digest).await?; + // Capture is a read of the model, not a mutation of it: nothing above changed a + // counter, and the worker goes back to the boundary it was already at. + self.status.set_state(previous); + let result = CaptureResult { + checkpoint_id: params.checkpoint_id, + boundary: k, + compatibility_digest: self.compatibility_digest(&profile, self.model.seed()), + payload: artifact.reference().clone(), + }; + Ok(HandlerReply::with_artifacts( + object(result.to_json()), + vec![(crate::state::PAYLOAD_ATTACHMENT.to_owned(), artifact)], + )) + } + + /// `State.StageRestore`: validate a replacement state into a staging slot. + /// + /// Nothing the live session can see changes here, and the worker keeps whatever state it + /// had. It is allowed on an uninitialized replacement or a quiescent worker only; a + /// failed one is neither, which is why a group that failed is replaced rather than + /// reused. + async fn state_stage_restore(&mut self, ctx: &HandlerCtx<'_>) -> DomainResult { + let scope = ctx.scope()?.clone(); + if scope.session_id != self.config.session_id { + return Err(DomainError::before( + ErrorCode::IdentityMismatch, + "this worker belongs to another session", + )); + } + match &self.phase { + AgentPhase::Uninitialized | AgentPhase::Ready(_) => {} + other => { + return Err(DomainError::before( + ErrorCode::InvalidPhase, + format!( + "State.StageRestore needs an uninitialized replacement or a quiescent \ +worker; this worker is {other:?}" + ), + )); + } + } + if let Some(epoch) = &self.epoch + && *epoch == scope.epoch + { + return Err(DomainError::before( + ErrorCode::StaleEpoch, + "State.StageRestore proposes the epoch this worker is already running", + )); + } + let params: StageRestoreParams = ctx.params()?; + if params.source_scope.step != scope.step { + return Err(DomainError::invalid( + "State.StageRestore's scope step must be the source boundary", + )); + } + let artifact = ctx.artifact(crate::state::PAYLOAD_ATTACHMENT)?; + if artifact.reference() != ¶ms.payload { + return Err(DomainError::before( + ErrorCode::BufferInvalid, + "the staged payload attachment is not the artifact the request names", + )); + } + let bytes = artifact.read_all().await.map_err(|e| { + DomainError::before( + ErrorCode::BufferInvalid, + format!("the staged payload could not be read: {}", e.message), + ) + })?; + let declared = params + .payload + .digest + .clone() + .ok_or_else(|| incompatible("a checkpoint payload must carry a content digest"))?; + let actual = digest_of_bytes(&bytes); + if actual != declared || bytes.len() as u64 != params.payload.byte_length { + return Err(incompatible( + "the staged payload is not the content the request declares", + )); + } + let value: Value = serde_json::from_slice(&bytes) + .map_err(|e| DomainError::invalid(format!("the staged payload is not JSON: {e}")))?; + let text = |key: &str| -> DomainResult { + value + .get(key) + .and_then(Value::as_str) + .map(str::to_owned) + .ok_or_else(|| incompatible(format!("the agent payload has no {key}"))) + }; + if value.get("payloadVersion").and_then(Value::as_u64) != Some(AGENT_PAYLOAD_VERSION) { + return Err(incompatible("the agent payload is another payload version")); + } + if text("kind")? != "agent" { + return Err(incompatible("this payload is not an agent's state")); + } + if text("agentId")? != self.config.agent_id { + return Err(DomainError::before( + ErrorCode::IdentityMismatch, + "the staged payload belongs to another agent", + )); + } + if text("checkpointId")? != params.checkpoint_id { + return Err(incompatible("the staged payload belongs to another checkpoint")); + } + if text("modelVersion")? != MODEL_VERSION || text("plasticityVersion")? != PLASTICITY_VERSION + { + return Err(incompatible( + "the staged payload was captured under another numerical model", + )); + } + let source_scope = Scope::from_json( + value + .get("sourceScope") + .ok_or_else(|| incompatible("the agent payload has no sourceScope"))?, + ) + .map_err(|e| incompatible(format!("the agent payload's sourceScope: {}", e.0)))?; + if source_scope != params.source_scope { + return Err(incompatible( + "the staged payload was captured at another source scope", + )); + } + let committed_step: u64 = text("committedStep")? + .parse() + .map_err(|_| incompatible("the agent payload's committedStep is not a U64"))?; + if committed_step != params.source_scope.step { + return Err(incompatible( + "the staged payload's committed step is not the source boundary", + )); + } + let profile = AssetRef::from_json( + value + .get("profile") + .ok_or_else(|| incompatible("the agent payload has no profile"))?, + ) + .map_err(|e| incompatible(format!("the agent payload's profile: {}", e.0)))?; + let model = FakeModel::restored( + value + .get("model") + .ok_or_else(|| incompatible("the agent payload has no model"))?, + )?; + // The compatibility digest is recomputed from this worker's own configuration and the + // identity the payload declares. A capture of the same profile under another seed, or + // of another agent's brain, fails here and never reaches activation. + let computed = self.compatibility_digest(&profile, model.seed()); + if computed != params.compatibility_digest { + return Err(incompatible(format!( + "the staged state's compatibility {computed} is not the {} the restore \ +requires", + params.compatibility_digest + ))); + } + let accumulator_value = value + .get("accumulator") + .ok_or_else(|| incompatible("the agent payload has no accumulator"))?; + let rational = |key: &str| -> DomainResult { + RationalNs::from_json( + accumulator_value + .get(key) + .ok_or_else(|| incompatible(format!("the accumulator has no {key}")))?, + ) + .map_err(|e| incompatible(format!("the accumulator's {key}: {}", e.0))) + }; + let counter = |key: &str| -> DomainResult { + accumulator_value + .get(key) + .and_then(Value::as_str) + .ok_or_else(|| incompatible(format!("the accumulator has no {key}")))? + .parse::() + .map_err(|_| incompatible(format!("the accumulator's {key} is not a U64"))) + }; + let tick_duration = rational("tickDuration")?; + if tick_duration != self.config.tick_duration { + return Err(incompatible( + "the staged state was captured at another model tick duration", + )); + } + let accumulator = TickAccumulator::restored( + tick_duration, + rational("remainder")?, + counter("executedTicks")?, + counter("warmupOffset")?, + ) + .map_err(incompatible)?; + let context = TypedValue::from_json( + value + .get("context") + .ok_or_else(|| incompatible("the agent payload has no context"))?, + ) + .map_err(|e| incompatible(format!("the agent payload's context: {}", e.0)))?; + FakeAgentWorker::available_actions(&context)?; + + if self.config.faults.fail_stage_restore { + // The row where a group validates three participants and the fourth does not. + // Nothing is staged here and nothing is staged anywhere else either: the + // coordinator abandons the whole install. + return Err(incompatible( + "injected staging refusal: this participant's replacement state does not \ +validate", + )); + } + // One staged restore at a time. A second proposal replaces nothing silently. + if let Some(staged) = &self.staged { + return Err(DomainError::before( + ErrorCode::Conflict, + format!( + "this worker already holds the staged restore {} for checkpoint {}", + staged.token, staged.checkpoint_id + ), + )); + } + let token = restore_token(¶ms.checkpoint_id, &scope, &actual, &self.config.incarnation_id); + if self.activated.contains(&token) { + return Err(DomainError::before( + ErrorCode::Conflict, + "this exact restore was already activated on this worker", + )); + } + self.staged = Some(StagedAgent { + token: token.clone(), + checkpoint_id: params.checkpoint_id.clone(), + scope: scope.clone(), + model, + accumulator, + context, + profile, + committed_step, + }); + self.status.set_state(WorkerState::StagedRestore); + let result = StageRestoreResult { + checkpoint_id: params.checkpoint_id, + restore_token: token, + }; + Ok(HandlerReply::from(&result)) + } + + /// `State.ActivateRestore`: install the staged state under its new scope, without a tick. + /// + /// The token activates once. A duplicate domain request replays the cached reply through + /// the shell's result cache; a fresh request naming an already activated token is a + /// conflict, which is what stops a second group from being resumed from the same bytes. + async fn state_activate_restore( + &mut self, + ctx: &HandlerCtx<'_>, + ) -> DomainResult { + let params: ActivateRestoreParams = ctx.params()?; + if self.activated.contains(¶ms.restore_token) { + return Err(DomainError::before( + ErrorCode::Conflict, + "this restore token has already been activated", + )); + } + let Some(staged) = self.staged.take() else { + return Err(DomainError::before( + ErrorCode::InvalidPhase, + "this worker holds no staged restore", + )); + }; + if staged.token != params.restore_token { + // Put it back: naming another token is not a reason to discard this one. + let token = staged.token.clone(); + self.staged = Some(staged); + return Err(DomainError::before( + ErrorCode::IdentityMismatch, + format!("this worker's staged restore is {token}, not {}", params.restore_token), + )); + } + if self.config.faults.fail_activate_restore { + let token = staged.token.clone(); + self.staged = Some(staged); + self.status.set_state(WorkerState::Failed); + return Err(DomainError::new( + ErrorCode::BackendFailure, + format!("injected activation failure; {token} stays staged and unresumed"), + MutationCertainty::None, + )); + } + self.status.set_state(WorkerState::Restoring); + let StagedAgent { + token, + checkpoint_id, + scope, + model, + accumulator, + context, + profile, + committed_step, + } = staged; + self.model = model; + self.accumulator = Some(accumulator); + self.context_digest = Some(context.digest()); + self.context = Some(context); + self.profile = Some(profile); + self.epoch = Some(scope.epoch.clone()); + self.prepared = None; + self.phase = AgentPhase::Ready(committed_step); + self.activated.insert(token); + self.status.set_state(WorkerState::Ready); + self.status.set_scope(Some(scope_at( + &scope.session_id, + &scope.epoch, + committed_step, + ))); + self.status.advance_to(self.model.mutations()); + let result = ActivateRestoreResult { + committed_step, + checkpoint_id, + // An agent returns a null observation; the environment returns the world's. + observation: None, + }; + result + .validate_for_role(Role::Agent) + .map_err(|e| DomainError::invalid(e.0))?; + Ok(HandlerReply::from(&result)) + } +} + +/// A restore token bound to the checkpoint, the proposed scope, the payload bytes and the +/// worker incarnation staging them. +/// +/// `state-media-v1` section 5 binds a token to scope, payload and checkpoint. Binding it to +/// the incarnation as well is what keeps a token minted by a worker that has since been +/// replaced from activating anything on its replacement. +pub fn restore_token(checkpoint_id: &Id, scope: &Scope, payload_digest: &Digest, incarnation: &Id) -> Id { + let digest = digest_of_bytes( + format!( + "fly-session/restore-token-v1\n{checkpoint_id}\n{}\n{}\n{}\n{payload_digest}\n{incarnation}\n", + scope.session_id, scope.epoch, scope.step + ) + .as_bytes(), + ); + parse_id(&format!("rt-{}", &digest[..32])).expect("a hex suffix is an Id") +} diff --git a/services/flysim/crates/fly-session/src/cli.rs b/services/flysim/crates/fly-session/src/cli.rs index ac87642..603e375 100644 --- a/services/flysim/crates/fly-session/src/cli.rs +++ b/services/flysim/crates/fly-session/src/cli.rs @@ -46,8 +46,10 @@ Worker options (agent and environment): agent: --agent ID --port ID --tick-numerator N --tick-denominator N --warmup-ticks N [--prepare-delay-ms N] [--commit-delay-ms N] [--fail-commit-at-step N] + [--fail-stage-restore 0|1] [--fail-activate-restore 0|1] environment: --worker ID --ports p1,p2 --step-numerator N --step-denominator N [--advance-delay-ms N] [--omit-view-at-boundary N] + [--fail-stage-restore 0|1] [--fail-activate-restore 0|1] Measure options: --steps N transitions per run (default 200) @@ -145,6 +147,17 @@ impl Options { } } + /// A flag whose value is `0` or `1`. Anything else is an error naming it, so a + /// mistyped injection is a failed launch rather than a fault that never fires. + fn flag(&self, name: &str) -> Result { + match self.0.get(name) { + None => Ok(false), + Some(value) if value == "0" => Ok(false), + Some(value) if value == "1" => Ok(true), + Some(value) => Err(format!("--{name}: {value:?} is not 0 or 1")), + } + } + fn opt_u64(&self, name: &str) -> Result, String> { match self.0.get(name) { None => Ok(None), @@ -198,6 +211,8 @@ fn serve(role: &str, options: &Options) -> Result<(), String> { fail_commit_at_step: options.opt_u64(flags::FAIL_COMMIT_AT_STEP)?, prepare_delay_ms: options.u64(flags::PREPARE_DELAY_MS, 0)?, commit_delay_ms: options.u64(flags::COMMIT_DELAY_MS, 0)?, + fail_stage_restore: options.flag(flags::FAIL_STAGE_RESTORE)?, + fail_activate_restore: options.flag(flags::FAIL_ACTIVATE_RESTORE)?, }, client_id: client_id.clone(), service: service.clone(), @@ -220,6 +235,8 @@ fn serve(role: &str, options: &Options) -> Result<(), String> { omit_audio_at_boundary: options.opt_u64(flags::OMIT_AUDIO_AT_BOUNDARY)?, overlapping_audio_at_boundary: options .opt_u64(flags::OVERLAPPING_AUDIO_AT_BOUNDARY)?, + fail_stage_restore: options.flag(flags::FAIL_STAGE_RESTORE)?, + fail_activate_restore: options.flag(flags::FAIL_ACTIVATE_RESTORE)?, }, client_id: client_id.clone(), service: service.clone(), diff --git a/services/flysim/crates/fly-session/src/clock.rs b/services/flysim/crates/fly-session/src/clock.rs index 080c12a..aa50ff2 100644 --- a/services/flysim/crates/fly-session/src/clock.rs +++ b/services/flysim/crates/fly-session/src/clock.rs @@ -29,6 +29,30 @@ impl TickAccumulator { }) } + /// The exact accumulator a capture recorded. + /// + /// The remainder is restored, never rounded or reset: a resumed agent that started its + /// first interval from zero would drift away from the run it is supposed to continue. + pub fn restored( + tick_duration: RationalNs, + remainder: RationalNs, + executed_ticks: u64, + warmup_offset: u64, + ) -> Result { + let mut accumulator = TickAccumulator::new(tick_duration)?; + remainder.validate().map_err(|e| e.0)?; + if remainder >= tick_duration { + return Err("a captured remainder is not below one model tick".to_owned()); + } + if warmup_offset > executed_ticks { + return Err("a captured warm-up offset exceeds the executed tick count".to_owned()); + } + accumulator.remainder = remainder; + accumulator.executed_ticks = executed_ticks; + accumulator.warmup_offset = warmup_offset; + Ok(accumulator) + } + pub fn tick_duration(&self) -> RationalNs { self.tick_duration } diff --git a/services/flysim/crates/fly-session/src/coordinator.rs b/services/flysim/crates/fly-session/src/coordinator.rs index fb5ddaa..76db2e4 100644 --- a/services/flysim/crates/fly-session/src/coordinator.rs +++ b/services/flysim/crates/fly-session/src/coordinator.rs @@ -12,7 +12,7 @@ use std::collections::{BTreeMap, BTreeSet}; use std::time::{Duration, Instant}; -use serde_json::{Map, Value}; +use serde_json::{Map, Value, json}; use crate::clock::Pacing; use crate::media::{self, AudioTimelines}; @@ -150,6 +150,16 @@ pub struct Deadlines { pub resolve_attempts: u32, /// A separate, larger budget for `Worker.Hello` and the `Initialize` methods. pub boot: Duration, + /// A separate budget for the `State.*` methods, which `ipc-v1` section 6 gives one: + /// a capture serializes a participant and a restore validates and installs one, and + /// neither is a step whose latency the probe was chosen for. + pub capture: Duration, + /// How long a caller waits for a *durable* acknowledgment. + /// + /// It is not [`Deadlines::capture`]: that one bounds a call to a participant, and this + /// one bounds two `fsync`s, a queue the caller shares with other captures and a disk. + /// Reusing the call budget here would make a slow disk look like an unresponsive worker. + pub durable: Duration, } /// How long the resolution waits between attempts. @@ -176,6 +186,8 @@ impl Default for Deadlines { // Over sixteen seconds of pauses: the budget above is what terminates. resolve_attempts: 8192, boot: Duration::from_secs(30), + capture: Duration::from_secs(30), + durable: Duration::from_secs(60), } } } @@ -194,6 +206,8 @@ pub struct Topics { pub descriptor: String, pub snapshots: String, pub events: String, + /// Where the distinct captured/queued/committed/failed/superseded checkpoint events go. + pub checkpoints: String, } impl Topics { @@ -202,6 +216,7 @@ impl Topics { descriptor: format!("session.{session_id}.descriptor"), snapshots: format!("session.{session_id}.snapshots"), events: format!("session.{session_id}.events"), + checkpoints: format!("session.{session_id}.checkpoints"), } } } @@ -224,6 +239,11 @@ pub struct AgentSlot { pub graph: Option, /// The telemetry of the last committed boundary, which is what the snapshot publishes. pub telemetry: Option, + /// The tick count and remainder this agent last reported, which are what the checkpoint + /// manifest records for it. They are metadata about the payload, never a substitute for + /// it: the agent's own capture is the state that is restored. + pub brain_ticks: u64, + pub remainder: RationalNs, context: TypedValue, context_digest: Digest, prepared: Option, @@ -250,6 +270,8 @@ impl AgentSlot { committed_step: 0, graph: None, telemetry: None, + brain_ticks: 0, + remainder: RationalNs::ZERO, context: TypedValue::new(crate::task::context_schema(), Value::Object(Map::new())) .expect("an empty context object is a valid typed value"), context_digest: digest_of_bytes(b""), @@ -332,6 +354,12 @@ pub struct Coordinator { /// Set when the epoch failed: every old handle, route and reply is invalid from here on /// and only a coherent restore may lift it. fenced: bool, + /// The durable store's bounded writer, when this composition has one. + writer: Option, + /// The last checkpoint whose saved acknowledgment arrived, and its boundary. + durable: Option<(Id, u64)>, + /// Participants that installed a restore a group install then abandoned. + tainted: BTreeSet, started: std::time::Instant, last_advance_request: Option, last_commit_requests: Vec, @@ -408,6 +436,9 @@ impl Coordinator { metrics: Metrics::default(), blame: None, fenced: false, + writer: None, + durable: None, + tainted: BTreeSet::new(), started: std::time::Instant::now(), last_advance_request: None, last_commit_requests: Vec::new(), @@ -558,8 +589,16 @@ impl Coordinator { let (from, to) = self.phases.fail(); self.trace.phase(from, to); self.fenced = true; + // Whatever replies this session still owed an acknowledgment for belong to + // participants of an invalid epoch. Carrying them across a recovery would send + // `Worker.Acknowledge` to a registration that is gone. + self.lifecycle_acks.clear(); + // Every handle this session held on the old epoch's media goes with the fence: a + // recovery imports fresh artifacts and never expects one of these back. self.views.clear(); self.pending_views.clear(); + self.audio.clear(); + self.pending_audio.clear(); match &participant { Some(who) => self.audit.push(format!("fail:{detail}:{who}")), None => self.audit.push(format!("fail:{detail}")), @@ -697,6 +736,8 @@ impl Coordinator { /// Declares every framework topic under the delivery policy the publisher holds. async fn declare_topics(&mut self) -> Outcome<()> { + // Every framework topic, including the checkpoint stream, is declared in one place + // under the delivery policy the publisher holds. match self.publisher.declare().await { Ok(()) => Ok(()), Err(e) => Err(self.fail_now(e, "declare-topic")), @@ -775,6 +816,15 @@ impl Coordinator { if let Err(e) = media::check_required_views(&result.descriptor, &result.observation) { return Err(self.fail_now(e, "observation-0")); } + // O[0] ran no transition either, so it carries no chunk, and that is checked rather + // than assumed from its boundary number. + if let Err(e) = media::check_required_audio( + &result.descriptor, + &result.observation, + media::ObservationOrigin::Installed, + ) { + return Err(self.fail_now(e, "observation-0")); + } self.descriptor = Some(result.descriptor); self.observation = Some(result.observation); self.lifecycle_acks.push((worker, reply.request_id.clone())); @@ -878,6 +928,10 @@ impl Coordinator { self.agents[index].tick_duration = result.tick_duration; self.agents[index].warmup_ticks = result.warmup_ticks; self.agents[index].committed_step = 0; + // At boundary 0 the agent has executed exactly its warm-up, with no interval + // consumed, so the remainder is zero. + self.agents[index].brain_ticks = result.telemetry.brain_ticks; + self.agents[index].remainder = RationalNs::ZERO; self.lifecycle_acks.push((slot_worker, reply.request_id.clone())); let agent_id = self.agents[index].agent_id.clone(); self.audit.push(format!("agent.initialize:{agent_id}")); @@ -1076,6 +1130,8 @@ impl Coordinator { self.blame(Some(worker.worker_id.clone())); let deadline = if method.ends_with("Initialize") || method == "Worker.Hello" { self.deadlines.boot + } else if method.starts_with("State.") { + self.deadlines.capture } else { self.deadlines.probe }; @@ -1758,6 +1814,12 @@ impl Coordinator { method, )); } + if let Some(slot) = self.agents.iter_mut().find(|slot| slot.agent_id == agent_id) { + // What the agent will be at once this transition commits: Prepare is the only + // phase that advances the accumulator. + slot.brain_ticks = decision.brain_ticks; + slot.remainder = decision.remainder; + } self.audit.push(format!("prepared:{agent_id}@{k}")); self.stats.prepares += 1; self.blame(None); @@ -2176,7 +2238,11 @@ impl Coordinator { } // Audio has no sensory role here, but its chunks still cannot overlap or go backwards // inside an epoch, and a stale one must not reach presentation as current. - if let Err(e) = media::check_required_audio(descriptor, &result.observation) { + if let Err(e) = media::check_required_audio( + descriptor, + &result.observation, + media::ObservationOrigin::Transition, + ) { return Err(self.fail_now(e, "step-result")); } if let Err(e) = self.timelines.accept(descriptor, &result.observation) { @@ -2839,3 +2905,1432 @@ impl Coordinator { } } +// ------------------------------------------------------------------------------------------- +// STATE-01: coherent all-participant capture and recovery + +/// A capture that exists and has been queued, whose durable outcome has not arrived yet. +/// +/// `State.Capture` completes when an immutable capture exists, not when a backend save was +/// requested, and only durable completion produces a saved acknowledgment. Those are two +/// events, so they are two calls. +#[derive(Debug)] +pub struct CaptureTicket { + pub checkpoint_id: Id, + pub boundary: u64, + receiver: tokio::sync::oneshot::Receiver, +} + +/// What one completed group restore installed. +#[derive(Clone, Debug, PartialEq, Eq)] +pub struct RestoreReport { + pub checkpoint_id: Id, + pub boundary: u64, + pub epoch: Id, + /// Every participant that staged, in the order they were asked. + pub staged: Vec, + /// Every participant that activated, in the order they were asked. + pub activated: Vec, + /// The once-only token each participant staged under, so a caller can prove a token + /// activates once rather than being told so. + pub tokens: Vec<(Id, Id)>, + /// The artifacts the payloads were imported as. None of them existed before this restore. + pub imported: Vec, +} + +/// The payload name the coordinator's own session record travels under. +impl Coordinator { + /// Attaches the durable store's bounded writer. Without one this session captures nothing + /// and says so, rather than pretending to. + pub fn attach_store(&mut self, writer: crate::state::CheckpointWriter) { + self.writer = Some(writer); + } + + pub fn writer(&self) -> Option<&crate::state::CheckpointWriter> { + self.writer.as_ref() + } + + /// Stops the checkpoint writer and waits for its task. + pub async fn shutdown_store(&mut self) { + if let Some(writer) = self.writer.take() { + writer.shutdown().await; + } + } + + /// The last checkpoint whose *saved acknowledgment* this coordinator received, and its + /// boundary. It moves on durable completion and on nothing else. + pub fn durable(&self) -> Option<(Id, u64)> { + self.durable.clone() + } + + /// Participants that installed a restore in a group install that then failed. + /// + /// They hold state no group ever resumed. A further restore is refused until each one has + /// been replaced, which is how "a failure during activation cannot resume half a world" + /// survives the next attempt as well as this one. + pub fn tainted(&self) -> Vec { + self.tainted.iter().cloned().collect() + } + + /// The live composition's compatibility identities. + pub fn compatibility(&self) -> Outcome { + match &self.descriptor { + Some(descriptor) => Ok(crate::state::Compatibility::of(descriptor)), + None => Err(SessionFailure { + error: DomainError::before( + ErrorCode::InvalidPhase, + "the session has no environment descriptor to take compatibility from", + ), + phase: self.phases.phase().label(), + detail: "compatibility".to_owned(), + participant: None, + }), + } + } + + fn agent_compatibility(&self, slot: &AgentSlot) -> Digest { + crate::agent::agent_compatibility_digest( + &slot.agent_id, + &slot.profile.digest, + &crate::agent::dataset_digest(), + crate::agent::MODEL_VERSION, + crate::agent::PLASTICITY_VERSION, + slot.seed, + ) + } + + /// The coordinator's own session record: what it must hold again to resume this boundary. + /// + /// `state-media-v1` section 4 puts the next decision state, the admission state and the + /// event watermarks in the coordinator's own payloads. This is that payload: the + /// composition has no audience input, so the admission record says so explicitly rather + /// than being absent. + fn coordinator_record(&self, boundary: u64) -> Value { + let (last_source_step, issued) = self.task.event_watermarks(); + json!({ + "payloadVersion": 1, + "kind": "coordinator", + "committedStep": boundary.to_string(), + "admission": { + "audienceInput": "none-configured", + "admitted": Value::Array(Vec::new()), + }, + "eventWatermarks": { + "lastSourceStep": last_source_step.to_string(), + "issued": issued.to_string(), + }, + "audioPositions": Value::Object( + self.timelines + .positions() + .into_iter() + .map(|(stream, sample)| (stream, Value::String(sample.to_string()))) + .collect(), + ), + "agents": Value::Array( + self.agents + .iter() + .map(|slot| json!({ + "agentId": slot.agent_id.as_str(), + "portId": slot.port_id.as_str(), + "profile": slot.profile.to_json(), + "seed": slot.seed, + "workerThreads": slot.worker_threads.to_string(), + "committedStep": slot.committed_step.to_string(), + "brainTicks": slot.brain_ticks.to_string(), + "remainder": slot.remainder.to_json(), + "context": slot.context.to_json(), + })) + .collect(), + ), + }) + } + + /// Seals one coordinator-owned payload as an immutable artifact. + async fn seal_own(&mut self, name: &str, value: &Value) -> Outcome { + let bytes = match canonicalize(value) { + Ok(text) => text.into_bytes(), + Err(e) => { + return Err(self.fail_now( + DomainError::invalid(format!("checkpoint payload {name}: {}", e.0)), + "capture", + )); + } + }; + let digest = digest_of_bytes(&bytes); + let artifact = match crate::state::seal_payload(&self.bus, &bytes, &digest).await { + Ok(artifact) => artifact, + Err(e) => return Err(self.fail_now(e, "capture")), + }; + Ok(crate::state::CapturedPayload { + name: name.to_owned(), + byte_length: bytes.len() as u64, + digest, + artifact, + }) + } + + /// Takes one coherent all-participant capture at the committed boundary and queues it. + /// + /// The queue slot is taken *first*: a saturated writer refuses before a single + /// `State.Capture` is sent, which is the only way "reject or defer before capture" can be + /// true rather than aspirational. + pub async fn capture(&mut self, checkpoint_id: &Id, replaceable: bool) -> Outcome { + if self.fenced { + return Err(SessionFailure { + error: DomainError::before( + ErrorCode::InvalidPhase, + "the epoch is fenced; a fenced session captures nothing", + ), + phase: self.phases.phase().label(), + detail: "fenced".to_owned(), + participant: None, + }); + } + // A capture that arrives while a transition is in flight is a race, not a bug in the + // transaction: the supervisor asking for one does not know where the session is. It + // is refused by name and the epoch is untouched, unlike a phase edge the machine + // itself takes. + let Some(boundary) = self.phases.phase().committed_boundary() else { + self.audit.push(format!("capture-refused:{checkpoint_id}")); + return Err(SessionFailure { + error: DomainError::before( + ErrorCode::InvalidPhase, + "a coherent checkpoint is taken at a committed boundary only", + ), + phase: self.phases.phase().label(), + detail: "capture".to_owned(), + participant: None, + }); + }; + let origin = self.phases.phase(); + if self.writer.is_none() { + return Err(SessionFailure { + error: DomainError::before( + ErrorCode::Unsupported, + "this session has no checkpoint store attached", + ), + phase: self.phases.phase().label(), + detail: "capture".to_owned(), + participant: None, + }); + } + // Before anything is captured. + let reservation = match self.writer.as_ref().expect("checked").reserve() { + Ok(reservation) => reservation, + Err(e) => { + // A refused capture is not a session failure: the boundary stands, the + // session keeps stepping and the caller is told the queue is full. + self.audit.push(format!("capture-refused:{checkpoint_id}")); + return Err(SessionFailure { + error: e, + phase: self.phases.phase().label(), + detail: "capture".to_owned(), + participant: None, + }); + } + }; + self.transition(Phase::Capturing(boundary))?; + let result = self + .capture_group(checkpoint_id, boundary, replaceable, reservation) + .await; + match result { + Ok(ticket) => { + self.transition(origin)?; + self.audit.push(format!("captured:{checkpoint_id}@{boundary}")); + Ok(ticket) + } + Err(e) => Err(e), + } + } + + async fn capture_group( + &mut self, + checkpoint_id: &Id, + boundary: u64, + replaceable: bool, + reservation: crate::state::Reservation, + ) -> Outcome { + let scope = self.scope(boundary); + let compatibility = self.compatibility()?; + let params = object(CaptureParams { checkpoint_id: checkpoint_id.clone() }.to_json()); + let want = vec![crate::state::PAYLOAD_ATTACHMENT.to_owned()]; + let mut payloads = Vec::new(); + let mut acknowledge: Vec<(WorkerRef, DomainRequestId)> = Vec::new(); + + // The world first, then the agents in sorted order: one boundary, every participant. + let environment = self.environment.clone(); + let reply = self + .call( + &environment, + "State.Capture", + Some(scope.clone()), + params.clone(), + &[], + &want, + ) + .await?; + let world: CaptureResult = reply.parse().map_err(|e| self.fail_now(e, "capture"))?; + let world_payload = self.accept_capture( + &reply, + &world, + checkpoint_id, + boundary, + &compatibility.digest(), + crate::state::WORLD_PAYLOAD, + )?; + payloads.push(world_payload); + acknowledge.push((environment, reply.request_id.clone())); + + let mut agent_rows = Vec::new(); + for index in 0..self.agents.len() { + let slot_worker = self.agents[index].worker.clone(); + let agent_id = self.agents[index].agent_id.clone(); + let expected = self.agent_compatibility(&self.agents[index]); + let reply = self + .call( + &slot_worker, + "State.Capture", + Some(scope.clone()), + params.clone(), + &[], + &want, + ) + .await?; + let captured: CaptureResult = reply.parse().map_err(|e| self.fail_now(e, "capture"))?; + let name = crate::state::agent_payload(&agent_id); + let payload = + self.accept_capture(&reply, &captured, checkpoint_id, boundary, &expected, &name)?; + payloads.push(payload); + acknowledge.push((slot_worker, reply.request_id.clone())); + let slot = &self.agents[index]; + agent_rows.push(crate::state::AgentEntry { + agent_id: agent_id.clone(), + profile_digest: slot.profile.digest.clone(), + dataset_digest: crate::agent::dataset_digest(), + model_version: crate::agent::MODEL_VERSION.to_owned(), + plasticity_version: crate::agent::PLASTICITY_VERSION.to_owned(), + seed: slot.seed, + brain_ticks: slot.brain_ticks, + remainder: slot.remainder, + payload: name, + }); + } + + // The coordinator's own ledgers, sealed the same way so the writer treats every + // payload alike. + let ledger = self.task.capture().map_err(|e| self.fail_now(e, "capture"))?; + payloads.push( + self.seal_own(crate::state::TASK_LEDGER_PAYLOAD, &ledger.to_json()) + .await?, + ); + let inspection = self + .observation + .as_ref() + .map(|observation| observation.inspection.to_json()) + .ok_or_else(|| { + DomainError::before(ErrorCode::InvalidPhase, "the session never bootstrapped") + }); + let inspection = match inspection { + Ok(value) => value, + Err(e) => return Err(self.fail_now(e, "capture")), + }; + payloads.push( + self.seal_own(crate::state::PRIOR_INSPECTION_PAYLOAD, &inspection) + .await?, + ); + let mut executor_rows = Vec::new(); + for agent_id in self.agent_ids() { + let state = { + let executor = self.executors.get(&agent_id).ok_or_else(|| { + DomainError::before( + ErrorCode::IdentityMismatch, + format!("{agent_id} has no configured action executor"), + ) + }); + match executor { + Ok(executor) => executor.capture().map_err(|e| (e, agent_id.clone())), + Err(e) => Err((e, agent_id.clone())), + } + }; + let state = match state { + Ok(state) => state, + Err((e, who)) => { + self.blame(Some(who)); + return Err(self.fail_now(e, "capture")); + } + }; + let name = crate::state::executor_payload(&agent_id); + payloads.push(self.seal_own(&name, &state.to_json()).await?); + executor_rows.push((agent_id, name)); + } + let record = self.coordinator_record(boundary); + payloads.push( + self.seal_own(crate::state::ADMISSION_PAYLOAD, &record) + .await?, + ); + + let (last_source_step, issued) = self.task.event_watermarks(); + let world_time = match self.observation.as_ref().map(|o| o.world_time) { + Some(world_time) => Ok(world_time), + None => Err(self.fail_now( + DomainError::before( + ErrorCode::InvalidPhase, + "the session has no observation, so it is at no world time to record", + ), + "capture", + )), + }; + let manifest = crate::state::CheckpointManifest { + checkpoint_id: checkpoint_id.clone(), + source_scope: scope.clone(), + episode_id: self.episode_id.clone(), + world_time: world_time?, + scheduler_id: "lockstep-v1".to_owned(), + composition_digest: self.composition_digest(), + port_map: self + .agents + .iter() + .map(|slot| (slot.port_id.clone(), slot.agent_id.clone())) + .collect(), + compatibility: compatibility.clone(), + agents: agent_rows, + coordinator: crate::state::CoordinatorEntry { + task_ledger: crate::state::TASK_LEDGER_PAYLOAD.to_owned(), + prior_inspection: crate::state::PRIOR_INSPECTION_PAYLOAD.to_owned(), + executor_state: executor_rows, + admission_state: crate::state::ADMISSION_PAYLOAD.to_owned(), + event_watermarks: crate::state::EventWatermarks { + last_source_step, + issued, + }, + }, + environment: crate::state::EnvironmentEntry { + worker_id: self.environment.worker_id.clone(), + payload: crate::state::WORLD_PAYLOAD.to_owned(), + }, + // No external helper takes part in this composition, and the manifest says so + // rather than leaving the field out. + helper_state: Vec::new(), + payloads: payloads + .iter() + .map(|p| (p.name.clone(), p.byte_length, p.digest.clone())) + .collect(), + }; + + let submission = crate::state::CaptureSubmission { + checkpoint_id: checkpoint_id.clone(), + boundary, + session_id: self.session_id.clone(), + epoch: self.epoch.clone(), + episode_id: self.episode_id.clone(), + compatibility_digest: compatibility.digest(), + manifest: manifest.to_json(), + payloads, + replaceable, + }; + // The writer takes its own ownership of every payload here. Only then are the + // workers' cached capture replies released. + let receiver = { + let writer = self.writer.as_ref().expect("checked"); + match writer.submit(reservation, submission).await { + Ok(receiver) => receiver, + Err(e) => return Err(self.fail_now(e, "capture")), + } + }; + self.publish_checkpoint_event("captured", checkpoint_id, boundary, None).await?; + for (worker, request_id) in acknowledge { + self.lifecycle_acks.push((worker, request_id)); + } + self.acknowledge_lifecycle().await?; + Ok(CaptureTicket { + checkpoint_id: checkpoint_id.clone(), + boundary, + receiver, + }) + } + + /// Checks one participant's capture and turns it into a payload the writer can own. + fn accept_capture( + &mut self, + reply: &DomainReply, + result: &CaptureResult, + checkpoint_id: &Id, + boundary: u64, + expected_compatibility: &Digest, + name: &str, + ) -> Outcome { + if result.checkpoint_id != *checkpoint_id { + return Err(self.fail_now( + DomainError::before( + ErrorCode::IdentityMismatch, + "a capture names another checkpoint", + ), + "capture", + )); + } + if result.boundary != boundary { + return Err(self.fail_now( + DomainError::before( + ErrorCode::IdentityMismatch, + "a capture names another boundary; one checkpoint is one boundary", + ), + "capture", + )); + } + if result.compatibility_digest != *expected_compatibility { + return Err(self.fail_now( + DomainError::before( + ErrorCode::IncompatibleState, + "a capture reports a compatibility identity the composition does not hold", + ), + "capture", + )); + } + let digest = match &result.payload.digest { + Some(digest) => digest.clone(), + None => { + return Err(self.fail_now( + DomainError::before( + ErrorCode::BufferInvalid, + "a checkpoint payload must carry a content digest", + ), + "capture", + )); + } + }; + let artifact = match reply.artifacts.get(crate::state::PAYLOAD_ATTACHMENT) { + Some(artifact) if artifact.reference() == &result.payload => artifact.clone(), + _ => { + return Err(self.fail_now( + DomainError::before( + ErrorCode::BufferInvalid, + "a capture arrived without a live owned handle on its payload", + ), + "capture", + )); + } + }; + Ok(crate::state::CapturedPayload { + name: name.to_owned(), + byte_length: result.payload.byte_length, + digest, + artifact, + }) + } + + /// Waits for one capture's durable outcome and moves the durable mark only on a commit. + /// + /// A lost reply is an outcome here, not a hang and not a save: the high-water mark stays + /// where it was until [`Coordinator::resolve_durable`] asks the store about the same + /// operation. + pub async fn await_durable(&mut self, ticket: CaptureTicket) -> Outcome { + let CaptureTicket { checkpoint_id, boundary, receiver } = ticket; + let budget = self.deadlines.durable; + let outcome = + crate::state::CheckpointWriter::wait(receiver, &checkpoint_id, budget).await; + if outcome.is_durable() { + self.durable = Some((checkpoint_id.clone(), boundary)); + self.audit.push(format!("durable:{checkpoint_id}@{boundary}")); + } else { + // Named rather than lumped together: a failed write, a superseded capture, a lost + // reply and an expired caller budget are four different things to have to explain. + self.audit + .push(format!("not-durable:{}:{checkpoint_id}@{boundary}", outcome.event())); + } + Ok(outcome) + } + + /// Resolves a save whose reply was lost, by asking the store's durable metadata about the + /// *same* checkpoint. It never saves again. + /// + /// `Some(boundary)` means the store manifest lists that generation, which is the durable + /// commit point; `None` means it does not, and an unreferenced generation file stays + /// unreferenced. + pub async fn resolve_durable(&mut self, checkpoint_id: &Id) -> Outcome> { + let Some(writer) = self.writer.as_ref() else { + return Err(self.fail_now( + DomainError::before( + ErrorCode::Unsupported, + "this session has no checkpoint store attached", + ), + "resolve-durable", + )); + }; + let wanted = checkpoint_id.clone(); + let found = writer + .with_store(move |store| store.lookup(&wanted).map(|g| g.boundary)) + .await; + match found { + Some(boundary) => { + self.durable = Some((checkpoint_id.clone(), boundary)); + self.audit.push(format!("durable:{checkpoint_id}@{boundary}")); + Ok(Some(boundary)) + } + None => { + self.audit.push(format!("not-durable:{checkpoint_id}")); + Ok(None) + } + } + } + + /// Captures and waits for the durable outcome, which is what an ordinary caller wants. + pub async fn checkpoint(&mut self, checkpoint_id: &Id) -> Outcome { + let ticket = self.capture(checkpoint_id, false).await?; + self.await_durable(ticket).await + } + + async fn publish_checkpoint_event( + &mut self, + event: &str, + checkpoint_id: &Id, + boundary: u64, + detail: Option<&str>, + ) -> Outcome<()> { + let payload = json!({ + "event": event, + "sessionId": self.session_id.as_str(), + "epoch": self.epoch.as_str(), + "checkpointId": checkpoint_id.as_str(), + "boundary": boundary.to_string(), + "detail": detail.map_or(Value::Null, |d| Value::String(d.to_owned())), + }); + let outcome = self.publisher.publish_checkpoint(object(payload)).await; + self.settle(outcome, "checkpoint-event")?; + Ok(()) + } + + // --------------------------------------------------------------------------------------- + // Recovery + + /// Points the coordinator at a replacement participant while the epoch is fenced. + /// + /// A replacement is the only way a fenced participant comes back: `step-v1` section 7's + /// incarnation row says every live participant of a failed epoch belongs to an invalid + /// one, so the reference is exchanged deliberately here and never repaired in place. + pub fn replace_participant(&mut self, worker_id: &Id, worker: WorkerRef) -> Outcome<()> { + if !self.fenced { + return Err(self.fail_now( + DomainError::before( + ErrorCode::InvalidPhase, + "participants are replaced while the epoch is fenced, not during play", + ), + "replace", + )); + } + if worker.worker_id != *worker_id { + return Err(self.fail_now( + DomainError::before( + ErrorCode::IdentityMismatch, + "the replacement reference names another worker", + ), + "replace", + )); + } + if self.environment.worker_id == *worker_id { + self.environment = worker; + self.tainted.remove(worker_id); + self.audit.push(format!("replaced:{worker_id}")); + return Ok(()); + } + match self.agents.iter_mut().find(|slot| slot.agent_id == *worker_id) { + Some(slot) => { + slot.worker = worker; + self.tainted.remove(worker_id); + self.audit.push(format!("replaced:{worker_id}")); + Ok(()) + } + None => Err(self.fail_now( + DomainError::before( + ErrorCode::IdentityMismatch, + format!("{worker_id} is not a participant of this composition"), + ), + "replace", + )), + } + } + + /// Installs a coherent checkpoint into a fresh epoch and lifts the fence. + /// + /// The whole of `state-media-v1` section 6 in order: select a complete compatible durable + /// checkpoint, import its payloads as *new* artifacts, stage every participant, activate + /// every participant, install the coordinator's own staged state, verify identity and + /// boundary, flush the old media and parser state, and establish `Paused(k)`. A failure + /// at any point leaves the fence exactly where it was. + pub async fn restore( + &mut self, + checkpoint_id: Option<&Id>, + new_epoch: &Id, + ) -> Outcome { + if self.phases.phase() != Phase::Failed { + return Err(self.fail_now( + DomainError::before( + ErrorCode::InvalidPhase, + "a coherent restore starts from a failed epoch", + ), + "restore", + )); + } + if *new_epoch == self.epoch { + return Err(self.fail_now( + DomainError::before( + ErrorCode::StaleEpoch, + "a restore installs a fresh epoch, never the one that failed", + ), + "restore", + )); + } + if !self.tainted.is_empty() { + let who: Vec<&str> = self.tainted.iter().map(String::as_str).collect(); + return Err(self.fail_now( + DomainError::before( + ErrorCode::InvalidPhase, + format!( + "{} installed a restore that no group resumed and must be replaced first", + who.join(", ") + ), + ), + "restore", + )); + } + let Some(writer) = self.writer.as_ref() else { + return Err(self.fail_now( + DomainError::before( + ErrorCode::Unsupported, + "this session has no checkpoint store attached", + ), + "restore", + )); + }; + let wanted = checkpoint_id.cloned(); + let read = writer + .with_store(move |store| { + let record = store.select(wanted.as_ref())?; + let envelope = store.read(&record)?; + Ok::<_, DomainError>((record, envelope)) + }) + .await; + let (record, envelope) = match read { + Ok(pair) => pair, + Err(e) => return Err(self.fail_now(e, "restore")), + }; + let manifest = match crate::state::CheckpointManifest::from_json(&envelope.manifest) { + Ok(manifest) => manifest, + Err(e) => { + return Err(self.fail_now( + DomainError::before(ErrorCode::IncompatibleState, e), + "restore", + )); + } + }; + if let Err(e) = self.check_restore_identity(&manifest, &record) { + return Err(self.fail_now(e, "restore")); + } + let boundary = manifest.source_scope.step; + self.transition(Phase::Restoring(boundary))?; + let outcome = self.restore_group(&envelope, &manifest, new_epoch, boundary).await; + match outcome { + Ok(report) => Ok(report), + Err(e) => Err(e), + } + } + + /// Marks every participant of an abandoned group install as one that must be replaced. + /// + /// A participant that staged holds a replacement state nothing installed; one that + /// activated holds installed state no group resumed. Neither is a participant this + /// session may reuse, and `ipc-v1` section 6's last paragraph is explicit that v1 does + /// not silently reattach one. + fn taint_group(&mut self, staged: &[(Id, WorkerRef, Id)]) { + for (who, _, _) in staged { + self.tainted.insert(who.clone()); + } + } + + /// Every identity a restore checks before a single participant is asked to stage. + fn check_restore_identity( + &self, + manifest: &crate::state::CheckpointManifest, + record: &crate::state::GenerationRecord, + ) -> DomainResult<()> { + if manifest.source_scope.session_id != self.session_id { + return Err(DomainError::before( + ErrorCode::IdentityMismatch, + "the checkpoint belongs to another session", + )); + } + if manifest.episode_id != self.episode_id { + return Err(DomainError::before( + ErrorCode::IdentityMismatch, + "the checkpoint belongs to another episode", + )); + } + if manifest.scheduler_id != "lockstep-v1" { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + format!( + "the checkpoint was scheduled by {}, not lockstep-v1", + manifest.scheduler_id + ), + )); + } + if record.compatibility_digest != manifest.compatibility.digest() { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + "the store manifest and the envelope disagree about compatibility", + )); + } + let live = match &self.descriptor { + Some(descriptor) => crate::state::Compatibility::of(descriptor), + None => { + return Err(DomainError::before( + ErrorCode::InvalidPhase, + "the session has no environment descriptor to compare compatibility with", + )); + } + }; + // Names the identity that differs -- backend, content, patch, controller, parser or + // state format -- rather than one opaque digest mismatch. + manifest + .compatibility + .compare(&live) + .map_err(|e| DomainError::before(ErrorCode::IncompatibleState, e))?; + let live_ports: Vec<(Id, Id)> = self + .agents + .iter() + .map(|slot| (slot.port_id.clone(), slot.agent_id.clone())) + .collect(); + if manifest.port_map != live_ports { + return Err(DomainError::before( + ErrorCode::IdentityMismatch, + "the checkpoint's port map is not this composition's", + )); + } + if manifest.environment.worker_id != self.environment.worker_id { + return Err(DomainError::before( + ErrorCode::IdentityMismatch, + "the checkpoint's world is not this composition's environment", + )); + } + let recorded: Vec = manifest.agents.iter().map(|a| a.agent_id.clone()).collect(); + if recorded != self.agent_ids() { + return Err(DomainError::before( + ErrorCode::IdentityMismatch, + "the checkpoint's agents are not this composition's", + )); + } + for (row, slot) in manifest.agents.iter().zip(&self.agents) { + if row.profile_digest != slot.profile.digest || row.seed != slot.seed { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + format!( + "{}'s checkpoint was taken under another profile or seed", + row.agent_id + ), + )); + } + } + Ok(()) + } + + #[allow(clippy::too_many_lines)] + async fn restore_group( + &mut self, + envelope: &fly_session_types::checkpoint::Envelope, + manifest: &crate::state::CheckpointManifest, + new_epoch: &Id, + boundary: u64, + ) -> Outcome { + let scope = scope_at(&self.session_id, new_epoch, boundary); + // Step 3: import the payloads as *new* artifacts. Nothing the fence dropped is asked + // to come back, and a router that restarted has none of the old roots anyway. + let mut imported: BTreeMap = BTreeMap::new(); + for (name, bytes) in &envelope.payloads { + let digest = digest_of_bytes(bytes); + let artifact = match crate::state::seal_payload(&self.bus, bytes, &digest).await { + Ok(artifact) => artifact, + Err(e) => return Err(self.fail_now(e, "restore")), + }; + imported.insert(name.clone(), (artifact, digest, bytes.len() as u64)); + } + let imported_ids: Vec = imported + .values() + .map(|(artifact, _, _)| artifact.reference().artifact_id.clone()) + .collect(); + + // Step 3, continued: stage every participant. A refusal anywhere leaves nothing + // staged that will ever be activated, because the whole install is abandoned. + let mut staged: Vec<(Id, WorkerRef, Id)> = Vec::new(); + let mut order: Vec<(Id, WorkerRef, String, Digest)> = Vec::new(); + order.push(( + self.environment.worker_id.clone(), + self.environment.clone(), + manifest.environment.payload.clone(), + manifest.compatibility.digest(), + )); + for index in 0..self.agents.len() { + let slot = &self.agents[index]; + let row = manifest + .agents + .iter() + .find(|row| row.agent_id == slot.agent_id) + .expect("the agent set was checked"); + order.push(( + slot.agent_id.clone(), + slot.worker.clone(), + row.payload.clone(), + crate::agent::agent_compatibility_digest( + &row.agent_id, + &row.profile_digest, + &row.dataset_digest, + &row.model_version, + &row.plasticity_version, + row.seed, + ), + )); + } + for (who, worker, payload_name, compatibility_digest) in &order { + let Some((artifact, _, _)) = imported.get(payload_name) else { + return Err(self.fail_now( + DomainError::before( + ErrorCode::IncompatibleState, + format!("the checkpoint has no payload {payload_name} for {who}"), + ), + "stage-restore", + )); + }; + let params = StageRestoreParams { + checkpoint_id: manifest.checkpoint_id.clone(), + source_scope: manifest.source_scope.clone(), + compatibility_digest: compatibility_digest.clone(), + payload: artifact.reference().clone(), + }; + let attachments = [(crate::state::PAYLOAD_ATTACHMENT, artifact)]; + let reply = match self + .call( + worker, + "State.StageRestore", + Some(scope.clone()), + object(params.to_json()), + &attachments, + &[], + ) + .await + { + Ok(reply) => reply, + Err(failure) => { + // Whoever already staged is holding a replacement state this group will + // never install. Nothing is resumed, and none of them is reused. + for (done, _, _) in &staged { + self.tainted.insert(done.clone()); + } + return Err(failure); + } + }; + let result: StageRestoreResult = match reply.parse() { + Ok(result) => result, + Err(e) => { + for (done, _, _) in &staged { + self.tainted.insert(done.clone()); + } + return Err(self.fail_now(e, "stage-restore")); + } + }; + if result.checkpoint_id != manifest.checkpoint_id { + for (done, _, _) in &staged { + self.tainted.insert(done.clone()); + } + return Err(self.fail_now( + DomainError::before( + ErrorCode::IdentityMismatch, + "a staged restore names another checkpoint", + ), + "stage-restore", + )); + } + self.lifecycle_acks.push((worker.clone(), reply.request_id.clone())); + self.audit.push(format!("staged:{who}")); + staged.push((who.clone(), worker.clone(), result.restore_token)); + } + + // `state-media-v1` section 5: activation happens "after every participant **and + // coordinator** state validates". The coordinator's own staged ledgers are checked + // here, before a single token is activated, so a checkpoint whose task ledger or + // executor state is unreadable installs nothing anywhere. + let payloads: BTreeMap> = envelope + .payloads + .iter() + .map(|(name, bytes)| (name.clone(), bytes.clone())) + .collect(); + let descriptor = match &self.descriptor { + Some(descriptor) => descriptor.clone(), + None => { + self.taint_group(&staged); + return Err(self.fail_now( + DomainError::before( + ErrorCode::InvalidPhase, + "the session has no environment descriptor to restore against", + ), + "restore", + )); + } + }; + if let Err(e) = Coordinator::validate_coordinator_state( + manifest, + &payloads, + &descriptor, + self.agents.len(), + &*self.task, + &self.executors, + ) { + self.taint_group(&staged); + return Err(self.fail_now(e, "restore")); + } + + // Step 3, last part: activate. Anything that activates before a failure holds state + // no group resumed, so it is recorded as tainted and must be replaced. + let mut activated: Vec = Vec::new(); + let mut restored_observation: Option<(WorldObservation, BTreeMap)> = + None; + for (who, worker, token) in &staged { + let params = ActivateRestoreParams { restore_token: token.clone() }; + let want = if *who == self.environment.worker_id { + self.media_names.clone() + } else { + Vec::new() + }; + let reply = match self + .call( + worker, + "State.ActivateRestore", + Some(scope.clone()), + object(params.to_json()), + &[], + &want, + ) + .await + { + Ok(reply) => reply, + Err(failure) => { + self.taint_group(&staged); + return Err(failure); + } + }; + let result: ActivateRestoreResult = match reply.parse() { + Ok(result) => result, + Err(e) => { + self.taint_group(&staged); + return Err(self.fail_now(e, "activate-restore")); + } + }; + let role = if *who == self.environment.worker_id { + Role::Environment + } else { + Role::Agent + }; + if let Err(e) = result.validate_for_role(role) { + self.taint_group(&staged); + return Err(self.fail_now(DomainError::invalid(e.0), "activate-restore")); + } + if result.committed_step != boundary || result.checkpoint_id != manifest.checkpoint_id { + self.taint_group(&staged); + return Err(self.fail_now( + DomainError::before( + ErrorCode::IdentityMismatch, + "an activation names another boundary or checkpoint", + ), + "activate-restore", + )); + } + if let Some(observation) = result.observation { + let (views, audio) = media::split_attachments(reply.artifacts); + if !audio.is_empty() { + self.taint_group(&staged); + return Err(self.fail_now( + DomainError::new( + ErrorCode::BufferInvalid, + "a restored observation carried an audio chunk; no interval was played", + MutationCertainty::Unknown, + ), + "activate-restore", + )); + } + restored_observation = Some((observation, views)); + } + self.lifecycle_acks.push((worker.clone(), reply.request_id.clone())); + self.audit.push(format!("activated:{who}")); + activated.push(who.clone()); + } + + // Step 4: verify identity and boundary, and flush the old media and parser state. + let Some((observation, views)) = restored_observation else { + self.taint_group(&staged); + return Err(self.fail_now( + DomainError::before( + ErrorCode::IncompatibleState, + "no participant returned the restored world observation", + ), + "activate-restore", + )); + }; + let install = self.install_restored( + manifest, + new_epoch, + boundary, + &descriptor, + observation, + views, + &payloads, + ); + if let Err(e) = install { + self.taint_group(&staged); + return Err(self.fail_now(e, "restore")); + } + + // Step 5: Paused(k), and only now is the fence lifted. + self.epoch = new_epoch.clone(); + self.transition(Phase::Paused(boundary))?; + self.fenced = false; + self.acknowledge_lifecycle().await?; + self.publish_checkpoint_event("restored", &manifest.checkpoint_id, boundary, None) + .await?; + self.publish_descriptor().await?; + self.publish_snapshot(boundary, &BTreeMap::new(), &[], &[]).await?; + self.audit.push(format!("restored:{}@{boundary}", manifest.checkpoint_id)); + Ok(RestoreReport { + checkpoint_id: manifest.checkpoint_id.clone(), + boundary, + epoch: new_epoch.clone(), + staged: staged.iter().map(|(who, _, _)| who.clone()).collect(), + tokens: staged + .iter() + .map(|(who, _, token)| (who.clone(), token.clone())) + .collect(), + activated, + imported: imported_ids, + }) + } + + /// Reads one of the checkpoint's payloads as JSON, or says which one is unreadable. + fn read_payload(payloads: &BTreeMap>, name: &str) -> DomainResult { + let bytes = payloads.get(name).ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + format!("the checkpoint has no payload {name}"), + ) + })?; + serde_json::from_slice(bytes).map_err(|e| { + DomainError::before( + ErrorCode::IncompatibleState, + format!("payload {name} is not JSON: {e}"), + ) + }) + } + + /// Validates every coordinator-owned payload, changing nothing. + /// + /// This runs after the group has staged and before anything activates, which is the order + /// `state-media-v1` section 5 sets. It is a separate pass from the install below on + /// purpose: a group that cannot be resumed coherently must not have resumed some of it. + fn validate_coordinator_state( + manifest: &crate::state::CheckpointManifest, + payloads: &BTreeMap>, + descriptor: &EnvironmentDescriptor, + agents: usize, + task: &dyn crate::task::Task, + executors: &BTreeMap>, + ) -> DomainResult<()> { + let ledger = TypedValue::from_json(&Coordinator::read_payload( + payloads, + &manifest.coordinator.task_ledger, + )?) + .map_err(|e| DomainError::before(ErrorCode::IncompatibleState, e.0))?; + task.validate_restore(&ledger)?; + for (agent_id, name) in &manifest.coordinator.executor_state { + let state = TypedValue::from_json(&Coordinator::read_payload(payloads, name)?) + .map_err(|e| DomainError::before(ErrorCode::IncompatibleState, e.0))?; + let executor = executors.get(agent_id).ok_or_else(|| { + DomainError::before( + ErrorCode::IdentityMismatch, + format!("{agent_id} has no configured action executor"), + ) + })?; + executor.validate_restore(&state)?; + } + let inspection = TypedValue::from_json(&Coordinator::read_payload( + payloads, + &manifest.coordinator.prior_inspection, + )?) + .map_err(|e| DomainError::before(ErrorCode::IncompatibleState, e.0))?; + if inspection.schema != descriptor.inspection_schema { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + "the recorded prior inspection is not the declared inspection schema", + )); + } + let record = + Coordinator::read_payload(payloads, &manifest.coordinator.admission_state)?; + let recorded = record + .get("agents") + .and_then(Value::as_array) + .ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "the coordinator record has no agent list", + ) + })?; + if recorded.len() != agents { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + "the coordinator record names another number of agents", + )); + } + if record.get("audioPositions").and_then(Value::as_object).is_none() { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + "the coordinator record has no audio positions", + )); + } + Ok(()) + } + + /// Installs the coordinator's own staged state, after every participant activated. + #[allow(clippy::too_many_arguments)] + fn install_restored( + &mut self, + manifest: &crate::state::CheckpointManifest, + new_epoch: &Id, + boundary: u64, + descriptor: &EnvironmentDescriptor, + observation: WorldObservation, + views: BTreeMap, + payloads: &BTreeMap>, + ) -> DomainResult<()> { + if observation.boundary != boundary { + return Err(DomainError::before( + ErrorCode::IdentityMismatch, + "the restored observation is not at the restored boundary", + )); + } + if observation.world_time != manifest.world_time { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + "the restored observation's world time is not the checkpoint's", + )); + } + observation + .validate_against(descriptor) + .map_err(|e| DomainError::new(ErrorCode::BufferInvalid, e, MutationCertainty::Unknown))?; + media::check_required_views(descriptor, &observation)?; + // An installed observation ran no transition, so it carries no chunk. Old epoch audio + // arriving as current is exactly what this refuses. + media::check_required_audio(descriptor, &observation, media::ObservationOrigin::Installed)?; + for view in &observation.sensory_views { + let name = media::view_attachment(&view.view_id); + match views.get(&name) { + Some(artifact) if artifact.reference() == &view.pixels => {} + _ => { + return Err(DomainError::new( + ErrorCode::BufferInvalid, + format!( + "the restored view {} arrived without a live owned handle", + view.view_id + ), + MutationCertainty::Unknown, + )); + } + } + } + + let read = |name: &str| Coordinator::read_payload(payloads, name); + + let ledger = TypedValue::from_json(&read(&manifest.coordinator.task_ledger)?) + .map_err(|e| DomainError::before(ErrorCode::IncompatibleState, e.0))?; + self.task.validate_restore(&ledger)?; + for (agent_id, name) in &manifest.coordinator.executor_state { + let state = TypedValue::from_json(&read(name)?) + .map_err(|e| DomainError::before(ErrorCode::IncompatibleState, e.0))?; + let executor = self.executors.get(agent_id).ok_or_else(|| { + DomainError::before( + ErrorCode::IdentityMismatch, + format!("{agent_id} has no configured action executor"), + ) + })?; + executor.validate_restore(&state)?; + } + let record = read(&manifest.coordinator.admission_state)?; + let inspection = TypedValue::from_json(&read(&manifest.coordinator.prior_inspection)?) + .map_err(|e| DomainError::before(ErrorCode::IncompatibleState, e.0))?; + if inspection.schema != descriptor.inspection_schema { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + "the recorded prior inspection is not the declared inspection schema", + )); + } + let agents = record + .get("agents") + .and_then(Value::as_array) + .ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "the coordinator record has no agent list", + ) + })? + .clone(); + if agents.len() != self.agents.len() { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + "the coordinator record names another number of agents", + )); + } + let positions = record + .get("audioPositions") + .and_then(Value::as_object) + .ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "the coordinator record has no audio positions", + ) + })?; + let mut audio_positions = BTreeMap::new(); + for (stream, value) in positions { + let sample: u64 = value + .as_str() + .ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "a recorded audio position is not a canonical U64", + ) + })? + .parse() + .map_err(|_| { + DomainError::before( + ErrorCode::IncompatibleState, + "a recorded audio position is not a canonical U64", + ) + })?; + audio_positions.insert(stream.clone(), sample); + } + // Everything above validated. From here the coordinator installs, in one pass. + self.task.install_restore(new_epoch, &ledger)?; + for (agent_id, name) in &manifest.coordinator.executor_state { + let state = TypedValue::from_json(&read(name)?) + .map_err(|e| DomainError::before(ErrorCode::IncompatibleState, e.0))?; + let executor = self + .executors + .get_mut(agent_id) + .expect("checked immediately above"); + executor.install_restore(&state)?; + } + for value in &agents { + let agent_id = value.get("agentId").and_then(Value::as_str).ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "a coordinator agent record has no agentId", + ) + })?; + let slot = self + .agents + .iter_mut() + .find(|slot| slot.agent_id == agent_id) + .ok_or_else(|| { + DomainError::before( + ErrorCode::IdentityMismatch, + format!("the coordinator record names {agent_id}, which is not configured"), + ) + })?; + let context = TypedValue::from_json(value.get("context").ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "a coordinator agent record has no decision context", + ) + })?) + .map_err(|e| DomainError::before(ErrorCode::IncompatibleState, e.0))?; + let committed: u64 = value + .get("committedStep") + .and_then(Value::as_str) + .ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "a coordinator agent record has no committedStep", + ) + })? + .parse() + .map_err(|_| { + DomainError::before( + ErrorCode::IncompatibleState, + "a recorded committed step is not a canonical U64", + ) + })?; + if committed != boundary { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + "a coordinator agent record is at another boundary", + )); + } + let brain_ticks: u64 = value + .get("brainTicks") + .and_then(Value::as_str) + .ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "a coordinator agent record has no brainTicks", + ) + })? + .parse() + .map_err(|_| { + DomainError::before( + ErrorCode::IncompatibleState, + "a recorded tick count is not a canonical U64", + ) + })?; + let remainder = RationalNs::from_json(value.get("remainder").ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "a coordinator agent record has no remainder", + ) + })?) + .map_err(|e| DomainError::before(ErrorCode::IncompatibleState, e.0))?; + slot.context_digest = context.digest(); + slot.context = context; + slot.committed_step = committed; + slot.brain_ticks = brain_ticks; + slot.remainder = remainder; + slot.prepared = None; + slot.prepare_request = None; + } + // Old media and old parser state are replaced, never carried: the previous epoch's + // handles were dropped by the fence and the timelines start again at their preserved + // sample positions with a discontinuity. + self.views = views; + self.pending_views.clear(); + self.audio.clear(); + self.pending_audio.clear(); + self.timelines = AudioTimelines::restored(descriptor, &audio_positions)?; + self.observation = Some(observation); + self.episode = None; + self.last_advance_request = None; + self.last_commit_requests.clear(); + self.pacing = Some(Pacing::new(descriptor.step_duration)); + Ok(()) + } + + /// The epoch-derived identities of this session's behaviour trace, mapped onto `to_epoch`. + /// + /// `step-v1` section 8 compares behaviour across runs. A resumed run runs in a new epoch, + /// and scope, batch identity and every task event identity are derived from it, so a + /// comparison either accounts for that or compares nothing. This is what "accounting for + /// new epoch metadata" is: an explicit, total rewrite of the epoch-derived fields, which + /// fails rather than passing anything through it does not recognise. + pub fn rebase(&self, to_epoch: &Id) -> Outcome { + let events = self.task.rebase_ids(to_epoch).map_err(|e| SessionFailure { + error: e, + phase: self.phases.phase().label(), + detail: "rebase".to_owned(), + participant: None, + })?; + Ok(EpochRebase { + from: self.epoch.clone(), + to: to_epoch.clone(), + events, + }) + } +} diff --git a/services/flysim/crates/fly-session/src/environment.rs b/services/flysim/crates/fly-session/src/environment.rs index 8a2947b..9c1bd85 100644 --- a/services/flysim/crates/fly-session/src/environment.rs +++ b/services/flysim/crates/fly-session/src/environment.rs @@ -13,6 +13,8 @@ use std::collections::BTreeSet; +use serde_json::Value; + use crate::media::{self, AudioSource, RenderCounter, ViewPipeline}; use crate::task::{controller_schema_ref, inspection, inspection_schema}; // `crate::types` is this crate's facade over the shared `fly-session-types` crate; the @@ -48,6 +50,12 @@ pub struct EnvironmentFaults { pub omit_audio_at_boundary: Option, /// Emit an audio chunk that starts before the previous chunk ended. pub overlapping_audio_at_boundary: Option, + /// Refuse `State.StageRestore`, so a group install meets a participant that will not + /// validate. + pub fail_stage_restore: bool, + /// Refuse `State.ActivateRestore` after staging, so a group meets a failure halfway + /// through activation. + pub fail_activate_restore: bool, } #[derive(Clone, Debug)] @@ -86,6 +94,10 @@ pub struct CounterEnvironment { audio: Option, /// The frame served at the previous boundary, kept only so a fault can serve it again. previous_view: Option<(ViewRef, flybus::Artifact)>, + /// A validated replacement world the live session cannot see yet. + staged: Option, + /// Restore tokens this world has activated. A token activates once. + activated: BTreeSet, } impl CounterEnvironment { @@ -104,10 +116,17 @@ impl CounterEnvironment { pipeline: None, audio: None, previous_view: None, + staged: None, + activated: BTreeSet::new(), config, } } + /// True while a validated replacement world is staged and not yet activated. + pub fn has_staged_restore(&self) -> bool { + self.staged.is_some() + } + pub fn status(&self) -> StatusCell { self.status.clone() } @@ -477,7 +496,7 @@ impl WorkerEndpoint for CounterEnvironment { vec![ id("world-step-v1"), id("pixel-observation-v1"), - id("checkpoint-v1"), + id(crate::state::CHECKPOINT_CAPABILITY), ] } @@ -490,7 +509,13 @@ impl WorkerEndpoint for CounterEnvironment { } fn methods(&self) -> Vec<&'static str> { - vec!["Environment.Initialize", "Environment.Advance"] + vec![ + "Environment.Initialize", + "Environment.Advance", + "State.Capture", + "State.StageRestore", + "State.ActivateRestore", + ] } fn handle<'a>(&'a mut self, ctx: HandlerCtx<'a>) -> BoxFuture<'a, DomainResult> { @@ -498,6 +523,9 @@ impl WorkerEndpoint for CounterEnvironment { match ctx.method { "Environment.Initialize" => self.initialize(&ctx).await, "Environment.Advance" => self.advance(&ctx).await, + "State.Capture" => self.state_capture(&ctx).await, + "State.StageRestore" => self.state_stage_restore(&ctx).await, + "State.ActivateRestore" => self.state_activate_restore(&ctx).await, other => Err(DomainError::before( ErrorCode::Unsupported, format!("{other} is not an environment method"), @@ -516,3 +544,488 @@ pub fn synthetic_asset(asset_id: &str, body: &str) -> AssetRef { format: id("fly-config-v1"), } } + +// ------------------------------------------------------------------------------------------- +// STATE-01: capture and restore + +/// The version this payload layout is written and read under. +pub const WORLD_PAYLOAD_VERSION: u64 = 1; + +fn incompatible(message: impl std::fmt::Display) -> DomainError { + DomainError::before(ErrorCode::IncompatibleState, message) +} + +/// One staged restore, held outside the live world until it is activated. +struct StagedWorld { + token: Id, + checkpoint_id: Id, + scope: Scope, + episode_id: Id, + descriptor: EnvironmentDescriptor, + boundary: u64, + counter: i64, + world_time: RationalNs, + advances: u64, + frames: Vec<(u64, i64)>, + audio_next_sample: u64, + audio_phase: u64, + audio_accumulator: u128, + audio_denominator: u128, +} + +impl CounterEnvironment { + /// `State.Capture`: the world at its committed boundary, including its pending sensor + /// pipeline. + /// + /// The pipeline is recorded as reconstruction inputs -- the producing boundary and the + /// world counter of every retained frame -- and never as an artifact identity: a + /// transient artifact belongs to the router that is running now, and a checkpoint outlives + /// it. + async fn state_capture(&mut self, ctx: &HandlerCtx<'_>) -> DomainResult { + let scope = ctx.scope()?.clone(); + let Some(descriptor) = self.descriptor.clone() else { + return Err(DomainError::before( + ErrorCode::InvalidPhase, + "this environment is uninitialized", + )); + }; + if scope.session_id != self.config.session_id { + return Err(DomainError::before( + ErrorCode::IdentityMismatch, + "this environment belongs to another session", + )); + } + match &self.epoch { + Some(epoch) if *epoch == scope.epoch => {} + _ => { + return Err(DomainError::before( + ErrorCode::StaleEpoch, + "State.Capture names an epoch this environment has left", + )); + } + } + if scope.step != self.boundary { + return Err(DomainError::before( + if scope.step < self.boundary { + ErrorCode::StaleStep + } else { + ErrorCode::FutureStep + }, + "State.Capture must name the boundary the world is at", + )); + } + let params: CaptureParams = ctx.params()?; + let pipeline = self + .pipeline + .as_ref() + .ok_or_else(|| DomainError::before(ErrorCode::InvalidPhase, "no view pipeline"))?; + let audio = self + .audio + .as_ref() + .ok_or_else(|| DomainError::before(ErrorCode::InvalidPhase, "no audio source"))?; + let (accumulator, denominator) = audio.accumulator(); + let previous = self.status.state(); + self.status.set_state(WorkerState::Capturing); + let payload = serde_json::json!({ + "payloadVersion": WORLD_PAYLOAD_VERSION, + "kind": "world", + "workerId": self.config.worker_id.as_str(), + "checkpointId": params.checkpoint_id.as_str(), + "sourceScope": scope.to_json(), + "episodeId": self.episode_id.clone().expect("initialized").as_str(), + "committedStep": self.boundary.to_string(), + "counter": self.counter.to_string(), + "worldTime": self.world_time.to_json(), + "advances": self.advances.to_string(), + "descriptor": descriptor.to_json(), + "pipeline": { + // The declared delay's whole queue, oldest first. + "frames": pipeline + .retained() + .into_iter() + .map(|(boundary, counter)| serde_json::json!({ + "boundary": boundary.to_string(), + "counter": counter.to_string(), + })) + .collect::>(), + }, + "audio": { + "nextSample": audio.next_sample().to_string(), + "phase": audio.phase().to_string(), + "accumulator": accumulator.to_string(), + "denominator": denominator.to_string(), + "chunks": audio.chunks().to_string(), + }, + }); + let bytes = canonicalize(&payload) + .map_err(|e| DomainError::invalid(format!("State.Capture: {}", e.0)))? + .into_bytes(); + let digest = digest_of_bytes(&bytes); + let artifact = crate::state::seal_payload(ctx.client, &bytes, &digest).await?; + // A capture reads the world; it does not advance it. + self.status.set_state(previous); + let result = CaptureResult { + checkpoint_id: params.checkpoint_id, + boundary: self.boundary, + compatibility_digest: crate::state::Compatibility::of(&descriptor).digest(), + payload: artifact.reference().clone(), + }; + Ok(HandlerReply::with_artifacts( + object(result.to_json()), + vec![(crate::state::PAYLOAD_ATTACHMENT.to_owned(), artifact)], + )) + } + + /// `State.StageRestore`: validate a replacement world into a staging slot. + async fn state_stage_restore(&mut self, ctx: &HandlerCtx<'_>) -> DomainResult { + let scope = ctx.scope()?.clone(); + if scope.session_id != self.config.session_id { + return Err(DomainError::before( + ErrorCode::IdentityMismatch, + "this environment belongs to another session", + )); + } + if let Some(epoch) = &self.epoch + && *epoch == scope.epoch + { + return Err(DomainError::before( + ErrorCode::StaleEpoch, + "State.StageRestore proposes the epoch this environment is already running", + )); + } + if self.descriptor.is_some() { + // A world that is already running a boundary is not a quiescent replacement: the + // group replaces it rather than restoring over a live one. + return Err(DomainError::before( + ErrorCode::InvalidPhase, + "State.StageRestore needs an uninitialized replacement environment", + )); + } + let params: StageRestoreParams = ctx.params()?; + if params.source_scope.step != scope.step { + return Err(DomainError::invalid( + "State.StageRestore's scope step must be the source boundary", + )); + } + let artifact = ctx.artifact(crate::state::PAYLOAD_ATTACHMENT)?; + if artifact.reference() != ¶ms.payload { + return Err(DomainError::before( + ErrorCode::BufferInvalid, + "the staged payload attachment is not the artifact the request names", + )); + } + let bytes = artifact.read_all().await.map_err(|e| { + DomainError::before( + ErrorCode::BufferInvalid, + format!("the staged payload could not be read: {}", e.message), + ) + })?; + let declared = params + .payload + .digest + .clone() + .ok_or_else(|| incompatible("a checkpoint payload must carry a content digest"))?; + let actual = digest_of_bytes(&bytes); + if actual != declared || bytes.len() as u64 != params.payload.byte_length { + return Err(incompatible( + "the staged payload is not the content the request declares", + )); + } + let value: Value = serde_json::from_slice(&bytes) + .map_err(|e| DomainError::invalid(format!("the staged payload is not JSON: {e}")))?; + let text = |key: &str| -> DomainResult { + value + .get(key) + .and_then(Value::as_str) + .map(str::to_owned) + .ok_or_else(|| incompatible(format!("the world payload has no {key}"))) + }; + let number = |key: &str| -> DomainResult { + text(key)? + .parse::() + .map_err(|_| incompatible(format!("the world payload's {key} is not a U64"))) + }; + if value.get("payloadVersion").and_then(Value::as_u64) != Some(WORLD_PAYLOAD_VERSION) { + return Err(incompatible("the world payload is another payload version")); + } + if text("kind")? != "world" { + return Err(incompatible("this payload is not a world's state")); + } + if text("workerId")? != self.config.worker_id { + return Err(DomainError::before( + ErrorCode::IdentityMismatch, + "the staged payload belongs to another world", + )); + } + if text("checkpointId")? != params.checkpoint_id { + return Err(incompatible("the staged payload belongs to another checkpoint")); + } + let source_scope = Scope::from_json( + value + .get("sourceScope") + .ok_or_else(|| incompatible("the world payload has no sourceScope"))?, + ) + .map_err(|e| incompatible(format!("the world payload's sourceScope: {}", e.0)))?; + if source_scope != params.source_scope { + return Err(incompatible( + "the staged payload was captured at another source scope", + )); + } + let committed_step = number("committedStep")?; + if committed_step != params.source_scope.step { + return Err(incompatible( + "the staged payload's committed step is not the source boundary", + )); + } + let descriptor = EnvironmentDescriptor::from_json( + value + .get("descriptor") + .ok_or_else(|| incompatible("the world payload has no descriptor"))?, + ) + .map_err(|e| incompatible(format!("the world payload's descriptor: {}", e.0)))?; + // The replacement builds the descriptor it would advertise and compares. A world + // started with other ports, another cadence or another declared render delay is a + // different backend, not this one resumed. + let live = self.build_descriptor()?; + if descriptor != live { + return Err(incompatible( + "the staged world was captured under another environment descriptor", + )); + } + let expected = crate::state::Compatibility::of(&descriptor).digest(); + if expected != params.compatibility_digest { + return Err(incompatible(format!( + "the staged world's compatibility {expected} is not the {} the restore requires", + params.compatibility_digest + ))); + } + let counter: i64 = text("counter")? + .parse() + .map_err(|_| incompatible("the world payload's counter is not an integer"))?; + let world_time = RationalNs::from_json( + value + .get("worldTime") + .ok_or_else(|| incompatible("the world payload has no worldTime"))?, + ) + .map_err(|e| incompatible(format!("the world payload's worldTime: {}", e.0)))?; + let pipeline_value = value + .get("pipeline") + .and_then(|p| p.get("frames")) + .and_then(Value::as_array) + .ok_or_else(|| incompatible("the world payload has no pipeline frames"))?; + let mut frames = Vec::with_capacity(pipeline_value.len()); + for frame in pipeline_value { + let boundary = frame + .get("boundary") + .and_then(Value::as_str) + .ok_or_else(|| incompatible("a captured frame has no boundary"))? + .parse::() + .map_err(|_| incompatible("a captured frame's boundary is not a U64"))?; + let frame_counter = frame + .get("counter") + .and_then(Value::as_str) + .ok_or_else(|| incompatible("a captured frame has no counter"))? + .parse::() + .map_err(|_| incompatible("a captured frame's counter is not an integer"))?; + frames.push((boundary, frame_counter)); + } + match frames.last() { + Some((boundary, _)) if *boundary == committed_step => {} + _ => { + return Err(incompatible( + "the captured pipeline does not end at the committed boundary", + )); + } + } + let audio_value = value + .get("audio") + .ok_or_else(|| incompatible("the world payload has no audio state"))?; + let audio_number = |key: &str| -> DomainResult { + audio_value + .get(key) + .and_then(Value::as_str) + .ok_or_else(|| incompatible(format!("the captured audio state has no {key}")))? + .parse::() + .map_err(|_| incompatible(format!("the captured audio {key} is not a number"))) + }; + let audio_next_sample = u64::try_from(audio_number("nextSample")?) + .map_err(|_| incompatible("the captured audio position is outside U64"))?; + let audio_phase = u64::try_from(audio_number("phase")?) + .map_err(|_| incompatible("the captured audio phase is outside U64"))?; + + if self.config.faults.fail_stage_restore { + return Err(incompatible( + "injected staging refusal: this participant's replacement state does not \ +validate", + )); + } + if let Some(staged) = &self.staged { + return Err(DomainError::before( + ErrorCode::Conflict, + format!( + "this environment already holds the staged restore {} for checkpoint {}", + staged.token, staged.checkpoint_id + ), + )); + } + let token = crate::agent::restore_token( + ¶ms.checkpoint_id, + &scope, + &actual, + &self.config.incarnation_id, + ); + if self.activated.contains(&token) { + return Err(DomainError::before( + ErrorCode::Conflict, + "this exact restore was already activated on this environment", + )); + } + self.staged = Some(StagedWorld { + token: token.clone(), + checkpoint_id: params.checkpoint_id.clone(), + scope, + episode_id: parse_id(&text("episodeId")?) + .map_err(|e| incompatible(format!("the world payload's episodeId {e}")))?, + descriptor, + boundary: committed_step, + counter, + world_time, + advances: number("advances")?, + frames, + audio_next_sample, + audio_phase, + audio_accumulator: audio_number("accumulator")?, + audio_denominator: audio_number("denominator")?, + }); + self.status.set_state(WorkerState::StagedRestore); + let result = StageRestoreResult { + checkpoint_id: params.checkpoint_id, + restore_token: token, + }; + Ok(HandlerReply::from(&result)) + } + + /// `State.ActivateRestore`: install the staged world and return its coherent observation. + /// + /// Nothing advances. The pipeline's frames are rendered again into fresh artifacts of the + /// current store, which is what "the durable store imports fresh immutable bus artifacts" + /// means on the producing side, and the observation carries no audio chunk because no + /// interval was played. + async fn state_activate_restore( + &mut self, + ctx: &HandlerCtx<'_>, + ) -> DomainResult { + let params: ActivateRestoreParams = ctx.params()?; + if self.activated.contains(¶ms.restore_token) { + return Err(DomainError::before( + ErrorCode::Conflict, + "this restore token has already been activated", + )); + } + let Some(staged) = self.staged.take() else { + return Err(DomainError::before( + ErrorCode::InvalidPhase, + "this environment holds no staged restore", + )); + }; + if staged.token != params.restore_token { + let token = staged.token.clone(); + self.staged = Some(staged); + return Err(DomainError::before( + ErrorCode::IdentityMismatch, + format!( + "this environment's staged restore is {token}, not {}", + params.restore_token + ), + )); + } + if self.config.faults.fail_activate_restore { + let token = staged.token.clone(); + self.staged = Some(staged); + self.status.set_state(WorkerState::Failed); + return Err(DomainError::new( + ErrorCode::BackendFailure, + format!("injected activation failure; {token} stays staged and unresumed"), + MutationCertainty::None, + )); + } + self.status.set_state(WorkerState::Restoring); + let mut pipeline = ViewPipeline::new( + CounterEnvironment::view_descriptor(self.config.observation_delay_steps), + self.config.renders.clone(), + ); + pipeline.restore(ctx.client, &staged.frames).await?; + let audio = AudioSource::restored_from( + CounterEnvironment::audio_descriptor(), + staged.audio_next_sample, + staged.audio_phase, + staged.audio_accumulator, + staged.audio_denominator, + )?; + self.epoch = Some(staged.scope.epoch.clone()); + self.episode_id = Some(staged.episode_id.clone()); + self.descriptor = Some(staged.descriptor.clone()); + self.boundary = staged.boundary; + self.counter = staged.counter; + self.world_time = staged.world_time; + self.advances = staged.advances; + // Batch ids are unique within an epoch, and this is a new one. Keeping the old set + // would refuse nothing extra: a request under the old epoch is already refused by its + // scope. + self.batches.clear(); + self.pipeline = Some(pipeline); + self.audio = Some(audio); + self.previous_view = None; + self.activated.insert(staged.token); + self.status.set_state(WorkerState::Ready); + self.status.set_scope(Some(scope_at( + &staged.scope.session_id, + &staged.scope.epoch, + staged.boundary, + ))); + + let (observation, attachments) = self.restored_observation()?; + let result = ActivateRestoreResult { + committed_step: staged.boundary, + checkpoint_id: staged.checkpoint_id, + observation: Some(observation), + }; + result + .validate_for_role(Role::Environment) + .map_err(|e| DomainError::invalid(e.0))?; + let mut reply = HandlerReply::from(&result); + reply.artifacts = attachments; + Ok(reply) + } + + /// The observation the restored world is already at: no render, no advance, no audio. + fn restored_observation( + &mut self, + ) -> DomainResult<(WorldObservation, Vec<(String, flybus::Artifact)>)> { + let boundary = self.boundary; + let counter = self.counter; + let pipeline = self + .pipeline + .as_ref() + .ok_or_else(|| DomainError::before(ErrorCode::InvalidPhase, "no view pipeline"))?; + let (view, artifact) = pipeline.at(boundary).ok_or_else(|| { + incompatible("the restored pipeline holds no frame for the restored boundary") + })?; + self.previous_view = Some((view.clone(), artifact.clone())); + let observation = WorldObservation { + boundary, + world_time: self.world_time, + engine_frame: Some(boundary.to_string()), + sensory_views: vec![view.clone()], + inspection: inspection(counter, boundary), + broadcast_views: vec![view.clone()], + // No interval was played, so there is no chunk. A chunk here would be an old + // epoch's audio offered as current. + audio: Vec::new(), + }; + Ok(( + observation, + vec![(media::view_attachment(&view.view_id), artifact)], + )) + } +} diff --git a/services/flysim/crates/fly-session/src/harness.rs b/services/flysim/crates/fly-session/src/harness.rs index 938557c..334e6ad 100644 --- a/services/flysim/crates/fly-session/src/harness.rs +++ b/services/flysim/crates/fly-session/src/harness.rs @@ -26,6 +26,7 @@ use crate::media::{RenderCounter, SensorLog}; use crate::launcher::{ AgentLaunch, EnvironmentLaunch, Launcher, ReapOutcome, SUPERVISOR_CLIENT, ThreadBudget, }; +use crate::state::{CheckpointStore, CheckpointWriter, StoreConfig, StoreFaults, WriterConfig, WriterFaults}; use crate::task::{ActionExecutor, CounterTask, IdentityExecutor, Terminal}; // `crate::types` is this crate's facade over the shared `fly-session-types` crate; the // glob keeps the contract's own names in sight instead of restating them. @@ -90,6 +91,14 @@ pub struct HarnessConfig { /// The threads reserved for the coordinator, its router and its store. pub coordinator_threads: usize, pub environment_threads: usize, + /// How many committed generations the durable checkpoint store keeps. + pub store: StoreConfig, + /// The durable write faults this composition injects. + pub store_faults: StoreFaults, + /// The checkpoint queue's bounds. + pub writer: WriterConfig, + /// The writer faults this composition injects. + pub writer_faults: WriterFaults, } impl Default for HarnessConfig { @@ -112,6 +121,10 @@ impl Default for HarnessConfig { thread_budget: None, coordinator_threads: 1, environment_threads: 1, + store: StoreConfig::default(), + store_faults: StoreFaults::default(), + writer: WriterConfig::default(), + writer_faults: WriterFaults::default(), } } } @@ -139,6 +152,14 @@ const ENV_SERVICE: &str = "env.arena"; const ENV_CLIENT: &str = "environment"; const ENV_WORKER: &str = "arena"; const COORDINATOR_CLIENT: &str = "coordinator"; +/// The checkpoint writer's own bus identity. It publishes checkpoint events and nothing else. +const WRITER_CLIENT: &str = "checkpoint-writer"; + +/// How many times one participant may be replaced in a composition. +/// +/// Each replacement connects under its own client id, so a restart is visibly a new +/// participant rather than a silent reattachment, and the policy has to name them all. +const MAX_GENERATIONS: u32 = 8; fn agent_service(agent_id: &Id) -> String { format!("agent.{agent_id}") @@ -187,6 +208,10 @@ pub struct SessionHarness { observers: Mutex>, /// Which configured observer identity the next consumer takes. next_observer: std::sync::atomic::AtomicUsize, + /// Which generation of each participant is running: 1 is the one the composition started. + generations: BTreeMap, + /// Where the durable checkpoint store lives, for a test that reads the files themselves. + checkpoint_root: std::path::PathBuf, } impl SessionHarness { @@ -223,11 +248,9 @@ impl SessionHarness { g.call = vec![Pattern::prefix("agent."), Pattern::prefix("env.")]; }), ) + // The writer publishes the checkpoint events and never calls a participant. + .client(WRITER_CLIENT, grants(|g| g.publish = vec![Pattern::prefix("session.")])) .client(ENV_CLIENT, grants(|g| g.register = vec![Pattern::exact(ENV_SERVICE)])) - .client( - &format!("{ENV_CLIENT}-r2"), - grants(|g| g.register = vec![Pattern::exact(ENV_SERVICE)]), - ) // A presentation consumer subscribes and may call the repair service. It can // publish nothing, register nothing and reach no worker: "viewers/browser clients // never obtain worker control" (publishing-v1 section 7). A bus client id is one @@ -256,6 +279,14 @@ impl SessionHarness { g.manage_topics = vec![Pattern::prefix("session.")]; }), ); + // A replacement environment connects under its own client id, one per generation; + // this subsumes the single `-r2` identity the publication slice had configured. + for generation in 2..=MAX_GENERATIONS { + policy = policy.client( + &format!("{ENV_CLIENT}-r{generation}"), + grants(|g| g.register = vec![Pattern::exact(ENV_SERVICE)]), + ); + } for spec in &config.agents { let service = agent_service(&spec.agent_id); policy = policy.client( @@ -264,10 +295,12 @@ impl SessionHarness { ); // A replacement worker connects under its own client id, so a restart is visibly a // new participant rather than a silent reattachment to the active epoch. - policy = policy.client( - &format!("{}-r2", agent_client(&spec.agent_id)), - grants(|g| g.register = vec![Pattern::exact(&service)]), - ); + for generation in 2..=MAX_GENERATIONS { + policy = policy.client( + &format!("{}-r{generation}", agent_client(&spec.agent_id)), + grants(|g| g.register = vec![Pattern::exact(&service)]), + ); + } } let mut router_config = RouterConfig::new(&store_root); router_config.policy = policy; @@ -346,6 +379,21 @@ impl SessionHarness { } let coordinator_client = launcher.connect(COORDINATOR_CLIENT).await?; + // The durable store lives beside the router's artifact store and never inside it: a + // committed generation is outside the bus's ephemeral collection. + let checkpoint_root = root.join("checkpoints"); + let mut store = CheckpointStore::open(&checkpoint_root, config.store).map_err(refusal)?; + *store.faults_mut() = config.store_faults.clone(); + let writer_client = launcher.connect(WRITER_CLIENT).await?; + let writer = CheckpointWriter::start( + store, + config.writer, + config.writer_faults.clone(), + Some(( + writer_client, + format!("session.{}.checkpoints", config.session_id), + )), + ); let executors: BTreeMap> = config .agents .iter() @@ -363,6 +411,8 @@ impl SessionHarness { Box::new(CounterTask::new(&config.epoch, config.terminal)), executors, ); + let mut coordinator = coordinator; + coordinator.attach_store(writer); Ok(SessionHarness { coordinator, @@ -374,9 +424,16 @@ impl SessionHarness { launcher, observers: Mutex::new(Vec::new()), next_observer: std::sync::atomic::AtomicUsize::new(0), + generations: BTreeMap::new(), + checkpoint_root, }) } + /// Where the durable checkpoint store's generations and store manifest live. + pub fn checkpoint_root(&self) -> &std::path::Path { + &self.checkpoint_root + } + pub fn router(&self) -> &Router { self.launcher.router() } @@ -469,10 +526,11 @@ impl SessionHarness { .find(|spec| spec.agent_id == *agent_id) .expect("a configured agent") .clone(); + let generation = self.next_generation(agent_id)?; self.launcher.kill(agent_id).await; let tick_duration = millis(self.config.tick_ms).expect("a positive tick"); - let incarnation_id = - parse_id(&format!("{agent_id}-inc-2")).expect("an agent id plus a suffix is an Id"); + let incarnation_id = parse_id(&format!("{agent_id}-inc-{generation}")) + .expect("an agent id plus a suffix is an Id"); self.launcher .launch_agent(AgentLaunch { session_id: self.config.session_id.clone(), @@ -487,7 +545,7 @@ impl SessionHarness { // predecessor wrote, so a restore's sensory input is visible beside it. sensors: self.sensors.get(agent_id).cloned().unwrap_or_default(), faults: spec.faults.clone(), - client_id: format!("{}-r2", agent_client(agent_id)), + client_id: format!("{}-r{generation}", agent_client(agent_id)), service: agent_service(agent_id), }) .await @@ -500,6 +558,102 @@ impl SessionHarness { }) } + /// Replaces the environment with a fresh, uninitialized incarnation, as a restore needs. + pub async fn restart_environment(&mut self) -> Result { + let worker_id = id(ENV_WORKER); + let generation = self.next_generation(&worker_id)?; + self.launcher.kill(&worker_id).await; + let step_duration = hz(self.config.step_hz).expect("a positive cadence"); + let incarnation_id = parse_id(&format!("arena-inc-{generation}")) + .expect("a worker id plus a suffix is an Id"); + self.launcher + .launch_environment(EnvironmentLaunch { + session_id: self.config.session_id.clone(), + worker_id: worker_id.clone(), + incarnation_id: incarnation_id.clone(), + step_duration, + ports: self.config.agents.iter().map(|a| a.port_id.clone()).collect(), + worker_threads: self.config.environment_threads, + observation_delay_steps: self.config.observation_delay_steps, + renders: self.renders.clone(), + faults: self.config.environment_faults.clone(), + client_id: format!("{ENV_CLIENT}-r{generation}"), + service: ENV_SERVICE.to_owned(), + }) + .await + .map_err(refusal)?; + let worker = self.launcher.worker(&worker_id).expect("just launched"); + Ok(Restarted { + service: worker.identity.service.clone(), + service_incarnation: worker.service_incarnation.clone(), + incarnation_id, + }) + } + + fn next_generation(&mut self, worker_id: &Id) -> Result { + let slot = self.generations.entry(worker_id.clone()).or_insert(1); + if *slot >= MAX_GENERATIONS { + return Err(flybus::BusError::new( + flybus::ErrorCode::QuotaExceeded, + format!( + "{worker_id} has used all {MAX_GENERATIONS} configured client identities; a composition declares how many replacements it allows" + ), + )); + } + *slot += 1; + Ok(*slot) + } + + /// Replaces every participant and points the fenced coordinator at the replacements. + /// + /// This is what a recovery does before it restores: the old participants belong to an + /// invalid epoch, and the references the coordinator pinned are exchanged deliberately. + pub async fn replace_all_participants(&mut self) -> Result<(), flybus::BusError> { + let environment = self.environment_id(); + self.restart_environment().await?; + let worker = self + .launcher + .worker(&environment) + .expect("just launched") + .worker_ref(); + self.coordinator + .replace_participant(&environment, worker) + .map_err(|e| refusal(e.error))?; + for agent_id in self.config.agents.iter().map(|a| a.agent_id.clone()).collect::>() { + self.restart_agent(&agent_id).await?; + let worker = self + .launcher + .worker(&agent_id) + .expect("just launched") + .worker_ref(); + self.coordinator + .replace_participant(&agent_id, worker) + .map_err(|e| refusal(e.error))?; + } + Ok(()) + } + + /// Changes one agent's injected faults, so the replacement the next restart launches is + /// a participant without them. + /// + /// A fault is launch configuration, so clearing one is a relaunch and not a live change: + /// the worker running now keeps whatever it was started with. + pub fn set_agent_faults(&mut self, agent_id: &Id, faults: AgentFaults) { + if let Some(spec) = self + .config + .agents + .iter_mut() + .find(|spec| spec.agent_id == *agent_id) + { + spec.faults = faults; + } + } + + /// Changes the environment's injected faults, with the same relaunch rule. + pub fn set_environment_faults(&mut self, faults: EnvironmentFaults) { + self.config.environment_faults = faults; + } + /// Ends one participant without asking it, as a crash would. pub async fn kill(&mut self, worker_id: &Id) -> ReapOutcome { self.launcher.kill(worker_id).await @@ -563,7 +717,10 @@ impl SessionHarness { /// Reaps every participant and closes the router. pub async fn shutdown(self) { - let SessionHarness { coordinator, mut launcher, observers, .. } = self; + let SessionHarness { mut coordinator, mut launcher, observers, .. } = self; + // The writer task owns artifact handles and a blocking store. Leaving it running + // would leave both behind. + coordinator.shutdown_store().await; drop(coordinator); launcher.reap_all(&id("shutdown")).await; for observer in observers.into_inner().expect("not poisoned") { diff --git a/services/flysim/crates/fly-session/src/launcher.rs b/services/flysim/crates/fly-session/src/launcher.rs index 0f8469c..f778858 100644 --- a/services/flysim/crates/fly-session/src/launcher.rs +++ b/services/flysim/crates/fly-session/src/launcher.rs @@ -1235,6 +1235,8 @@ pub(crate) mod flags { pub const COMMIT_DELAY_MS: &str = "commit-delay-ms"; pub const FAIL_COMMIT_AT_STEP: &str = "fail-commit-at-step"; pub const GRAPH_VARIANT: &str = "graph-variant"; + pub const FAIL_STAGE_RESTORE: &str = "fail-stage-restore"; + pub const FAIL_ACTIVATE_RESTORE: &str = "fail-activate-restore"; pub const WORKER: &str = "worker"; pub const PORTS: &str = "ports"; @@ -1276,6 +1278,8 @@ pub(crate) mod flags { COMMIT_DELAY_MS, FAIL_COMMIT_AT_STEP, GRAPH_VARIANT, + FAIL_STAGE_RESTORE, + FAIL_ACTIVATE_RESTORE, ]; /// What only the environment is given, media options included. pub const ENVIRONMENT_ONLY: &[&str] = &[ @@ -1290,6 +1294,8 @@ pub(crate) mod flags { TRUNCATED_VIEW_AT_BOUNDARY, OMIT_AUDIO_AT_BOUNDARY, OVERLAPPING_AUDIO_AT_BOUNDARY, + FAIL_STAGE_RESTORE, + FAIL_ACTIVATE_RESTORE, ]; /// What a measurement run or one of its row children is given. pub const MEASURE: &[&str] = &[MODE, AGENTS, STEPS, WARMUP_STEPS, WORKER_THREADS, MODES]; @@ -1333,6 +1339,11 @@ impl Started { arg(flags::PREPARE_DELAY_MS, spec.faults.prepare_delay_ms), arg(flags::COMMIT_DELAY_MS, spec.faults.commit_delay_ms), arg(flags::GRAPH_VARIANT, spec.graph_variant), + arg(flags::FAIL_STAGE_RESTORE, u64::from(spec.faults.fail_stage_restore)), + arg( + flags::FAIL_ACTIVATE_RESTORE, + u64::from(spec.faults.fail_activate_restore), + ), ]; if let Some(step) = spec.faults.fail_commit_at_step { args.push(arg(flags::FAIL_COMMIT_AT_STEP, step)); @@ -1351,6 +1362,11 @@ impl Started { // The media options a world in another process needs to be exactly this // world. Its render counter and its agents' sensor logs stay there. arg(flags::OBSERVATION_DELAY_STEPS, spec.observation_delay_steps), + arg(flags::FAIL_STAGE_RESTORE, u64::from(spec.faults.fail_stage_restore)), + arg( + flags::FAIL_ACTIVATE_RESTORE, + u64::from(spec.faults.fail_activate_restore), + ), ]; for (flag, boundary) in [ (flags::OMIT_VIEW_AT_BOUNDARY, spec.faults.omit_view_at_boundary), @@ -1465,6 +1481,8 @@ mod flag_tests { truncated_view_at_boundary: Some(3), omit_audio_at_boundary: Some(4), overlapping_audio_at_boundary: Some(5), + fail_stage_restore: true, + fail_activate_restore: true, } } @@ -1499,6 +1517,8 @@ mod flag_tests { fail_commit_at_step: Some(2), prepare_delay_ms: 1, commit_delay_ms: 2, + fail_stage_restore: true, + fail_activate_restore: true, }, client_id: "worker-fly-a".to_owned(), service: "agent.fly-a".to_owned(), @@ -1543,6 +1563,8 @@ mod flag_tests { flags::TRUNCATED_VIEW_AT_BOUNDARY, flags::OMIT_AUDIO_AT_BOUNDARY, flags::OVERLAPPING_AUDIO_AT_BOUNDARY, + flags::FAIL_STAGE_RESTORE, + flags::FAIL_ACTIVATE_RESTORE, ] { assert!(written.contains(&format!("--{flag}")), "--{flag} is not written"); } diff --git a/services/flysim/crates/fly-session/src/lib.rs b/services/flysim/crates/fly-session/src/lib.rs index 312b53d..6849001 100644 --- a/services/flysim/crates/fly-session/src/lib.rs +++ b/services/flysim/crates/fly-session/src/lib.rs @@ -35,6 +35,7 @@ pub mod metrics; pub mod phase; pub mod publish; pub mod rpc; +pub mod state; pub mod task; pub mod worker; diff --git a/services/flysim/crates/fly-session/src/media.rs b/services/flysim/crates/fly-session/src/media.rs index 97af979..c73ed54 100644 --- a/services/flysim/crates/fly-session/src/media.rs +++ b/services/flysim/crates/fly-session/src/media.rs @@ -133,7 +133,14 @@ pub fn arena_frame(descriptor: &ViewDescriptor, counter: i64, boundary: u64) -> /// cannot be served an arbitrary stale image. pub struct ViewPipeline { descriptor: ViewDescriptor, - frames: VecDeque<(u64, flybus::Artifact)>, + /// Each retained frame: its producing boundary, the world counter it was rendered from + /// and the owned handle on its immutable bytes. + /// + /// The counter is kept because it is the whole of the reconstruction input: a checkpoint + /// records `(boundary, counter)` per retained frame and a restore re-renders them into + /// fresh artifacts of the current store, rather than persisting a transient artifact + /// identity that cannot survive a router restart. + frames: VecDeque<(u64, i64, flybus::Artifact)>, renders: RenderCounter, } @@ -160,7 +167,7 @@ impl ViewPipeline { let bytes = arena_frame(&self.descriptor, counter, boundary); let artifact = seal(client, FRAME_CONTENT_TYPE, &bytes).await?; self.renders.bump(); - self.frames.push_back((boundary, artifact)); + self.frames.push_back((boundary, counter, artifact)); // Keep exactly the frames a declared delay can still require. while self.frames.len() > self.descriptor.observation_delay_steps as usize + 1 { self.frames.pop_front(); @@ -168,6 +175,50 @@ impl ViewPipeline { Ok(()) } + /// The reconstruction inputs of every retained frame, oldest first. + /// + /// This is what a checkpoint records for the pending sensor pipeline: the producing + /// boundary and the world counter, never an artifact identity. + pub fn retained(&self) -> Vec<(u64, i64)> { + self.frames + .iter() + .map(|(boundary, counter, _)| (*boundary, *counter)) + .collect() + } + + /// Rebuilds the pipeline from recorded reconstruction inputs, into fresh artifacts. + /// + /// Every frame is rendered again in the current store, so nothing a fence dropped is + /// expected to come back and no old artifact identity crosses the recovery. + pub async fn restore( + &mut self, + client: &flybus::Client, + frames: &[(u64, i64)], + ) -> DomainResult<()> { + if frames.len() > self.descriptor.observation_delay_steps as usize + 1 { + return Err(media_error(format!( + "a captured pipeline of {} frames does not fit a declared delay of {}", + frames.len(), + self.descriptor.observation_delay_steps + ))); + } + for window in frames.windows(2) { + if window[1].0 != window[0].0 + 1 { + return Err(media_error( + "a captured pipeline's producing boundaries are not consecutive", + )); + } + } + self.frames.clear(); + for (boundary, counter) in frames { + let bytes = arena_frame(&self.descriptor, *counter, *boundary); + let artifact = seal(client, FRAME_CONTENT_TYPE, &bytes).await?; + self.renders.bump(); + self.frames.push_back((*boundary, *counter, artifact)); + } + Ok(()) + } + /// Seals a frame of the wrong length, which is what a broken backend produces. The /// reference it returns describes the artifact honestly, so the shape check is the thing /// under test rather than a lie in the payload. @@ -181,7 +232,7 @@ impl ViewPipeline { bytes.truncate(bytes.len() - self.descriptor.row_stride as usize); let artifact = seal(client, FRAME_CONTENT_TYPE, &bytes).await?; self.renders.bump(); - self.frames.push_back((boundary, artifact)); + self.frames.push_back((boundary, counter, artifact)); while self.frames.len() > self.descriptor.observation_delay_steps as usize + 2 { self.frames.pop_front(); } @@ -198,8 +249,8 @@ impl ViewPipeline { pub fn frame_produced_at(&self, produced: u64) -> Option<(ViewRef, flybus::Artifact)> { self.frames .iter() - .find(|(step, _)| *step == produced) - .map(|(step, artifact)| { + .find(|(step, _, _)| *step == produced) + .map(|(step, _, artifact)| { ( ViewRef { view_id: self.descriptor.view_id.clone(), @@ -257,6 +308,52 @@ impl AudioSource { source } + /// The exact state a capture recorded: sample position, waveform phase and the + /// unconsumed fraction of a frame. + /// + /// Restoring the position alone would restart the waveform and round the remainder away, + /// which is a resample the restore rules refuse. The first chunk of the new epoch marks + /// the discontinuity the recovery established. + pub fn restored_from( + descriptor: AudioDescriptor, + next_sample: u64, + phase: u64, + accumulator: u128, + denominator: u128, + ) -> DomainResult { + if denominator == 0 { + return Err(DomainError::invalid( + "audio: a captured accumulator denominator of zero", + )); + } + if accumulator >= denominator { + return Err(DomainError::invalid( + "audio: a captured accumulator is not below one whole frame", + )); + } + if phase >= descriptor.sample_rate { + return Err(DomainError::invalid( + "audio: a captured phase is not below the sample rate", + )); + } + let mut source = AudioSource::new(descriptor, next_sample); + source.discontinuous = true; + source.phase = phase; + source.accumulator = accumulator; + source.denominator = denominator; + Ok(source) + } + + /// The waveform phase, for a capture. + pub fn phase(&self) -> u64 { + self.phase + } + + /// The unconsumed fraction of a frame and the denominator it is over, for a capture. + pub fn accumulator(&self) -> (u128, u128) { + (self.accumulator, self.denominator) + } + pub fn descriptor(&self) -> &AudioDescriptor { &self.descriptor } @@ -451,18 +548,47 @@ pub fn check_required_views( Ok(()) } -/// Every declared audio stream produces exactly one chunk per transition. +/// Where an observation came from. +/// +/// `state-media-v1` section 2 makes a chunk the audio of an *interval*, so whether an +/// observation must carry one is a question about its provenance and not about its boundary +/// number. MEDIA-01 wrote the rule as "boundary 0 carries no chunk", which is true of the one +/// observation that slice could produce without a transition and false of the other one: +/// `State.ActivateRestore` installs a coherent observation at boundary `k` without advancing +/// gameplay, and it covers no interval either. Naming the provenance is the fix; exempting +/// the restored observation from the validator instead would have left "must a chunk exist" +/// unanswered exactly where a stale chunk would do the most damage. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub enum ObservationOrigin { + /// The observation a completed transition produced. Its interval has audio. + Transition, + /// An observation established at a boundary without running a transition: + /// `Environment.Initialize`'s `O[0]` and `State.ActivateRestore`'s restored observation. + /// It covers no interval, so it carries no chunk and one in it is refused. + Installed, +} + +/// Every declared audio stream produces exactly one chunk per transition, and none at all in +/// an observation that is not one. /// /// The contract states the shape and the ordering of chunks, not whether one has to exist, so /// this is MEDIA-01's choice and it is deliberate: a session that tolerates a silently missing /// chunk cannot tell "this world produced no audio for this interval" from "the chunk was -/// lost", and the second is the case the retention rules care about. Boundary 0 has no -/// preceding interval and so carries no chunk. +/// lost", and the second is the case the retention rules care about. The mirror of that, which +/// STATE-01 needs, is that an installed observation carrying a chunk is a stale chunk being +/// offered as current, and is refused for the same reason. pub fn check_required_audio( descriptor: &EnvironmentDescriptor, observation: &WorldObservation, + origin: ObservationOrigin, ) -> DomainResult<()> { - if observation.boundary == 0 { + if origin == ObservationOrigin::Installed { + if let Some(chunk) = observation.audio.first() { + return Err(media_error(format!( + "audio stream {} produced a chunk for an observation that ran no transition", + chunk.stream_id + ))); + } return Ok(()); } for stream in &descriptor.audio { diff --git a/services/flysim/crates/fly-session/src/publish.rs b/services/flysim/crates/fly-session/src/publish.rs index 485dff7..20f0200 100644 --- a/services/flysim/crates/fly-session/src/publish.rs +++ b/services/flysim/crates/fly-session/src/publish.rs @@ -500,6 +500,10 @@ pub struct Publisher { descriptor_topic: TopicPolicy, snapshot_topic: TopicPolicy, event_topic: TopicPolicy, + /// The STATE-01 checkpoint stream. A stream of distinct facts, so it is a bounded + /// delivery and never a latest value: a "committed" that replaced a "queued" would erase + /// the distinction the durable commit rules are built on. + checkpoint_topic: TopicPolicy, outbox: EventOutbox, state: SharedState, ledger: Ledger, @@ -520,6 +524,7 @@ impl Publisher { descriptor_topic: TopicPolicy::latest(&topics.descriptor), snapshot_topic: TopicPolicy::latest(&topics.snapshots), event_topic: TopicPolicy::bounded(&topics.events, EVENT_BATCH_DEPTH), + checkpoint_topic: TopicPolicy::bounded(&topics.checkpoints, EVENT_BATCH_DEPTH), outbox: EventOutbox::new(EVENT_BATCH_DEPTH), state: Arc::new(Mutex::new(PublishedState::default())), ledger: Ledger::default(), @@ -533,6 +538,7 @@ impl Publisher { self.descriptor_topic.clone(), self.snapshot_topic.clone(), self.event_topic.clone(), + self.checkpoint_topic.clone(), ] } @@ -675,6 +681,22 @@ impl Publisher { } } +impl Publisher { + /// Publishes one checkpoint fact on the checkpoint stream. + /// + /// It goes through the same named outcomes as everything else: a durable-commit fact that + /// an observer refuses is counted and does not fail the session, because the durable + /// acknowledgment is the store's, not the subscriber's -- "bus publish acceptance and + /// delivery consumption are not durable acknowledgments" (publishing-v1 section 6). + pub async fn publish_checkpoint(&mut self, payload: Map) -> PublicationOutcome { + let topic = self.checkpoint_topic.topic.clone(); + let outcome = + PublicationOutcome::from_bus(&topic, self.bus.publish(&topic, payload, &[]).await); + self.ledger.record(&outcome); + outcome + } +} + // ---------------------------------------------------------------------------------------------- // The query service: the repair path diff --git a/services/flysim/crates/fly-session/src/state.rs b/services/flysim/crates/fly-session/src/state.rs new file mode 100644 index 0000000..aa42d3c --- /dev/null +++ b/services/flysim/crates/fly-session/src/state.rs @@ -0,0 +1,1558 @@ +//! STATE-01: the coherent all-participant checkpoint store over the `FLYSESS1` envelope. +//! +//! [`fly_session_types::checkpoint`] owns the byte layout. This module owns everything +//! `checkpoint-envelope-v1` section 7 defers to this slice: generations, rotation, the store +//! manifest and its durable commit point, the bounded capture queue, the compatibility +//! comparison and the group fence. +//! +//! ```text +//! State.Capture ──> every participant, at one committed boundary +//! ──> one envelope: manifest + one payload per participant and per +//! coordinator-owned ledger +//! writer ──> temp, fsync, rename, fsync dir, then the store manifest the same way +//! ^^^^ the store manifest rename is the durable commit point +//! State.StageRestore ──> validated into replacement state, once-only token +//! State.ActivateRestore ──> installed under the new epoch, without a tick +//! ``` +//! +//! Nothing here is best-effort. A saturated queue is a named `BUSY` refusal taken *before* a +//! capture is requested; a lost save reply is an explicit outcome that leaves durable +//! metadata where it was; an incompatible checkpoint names the field that differs; and an +//! unreferenced generation is never a restore candidate. +//! +//! `FLYSIM01` (`crates/flybrain-core/src/envelope.rs`, read by the legacy flysim store) is a +//! different format with a different magic and a different reader, and nothing here touches +//! it. + +use std::collections::VecDeque; +use std::io::Write as _; +use std::path::{Path, PathBuf}; +use std::sync::Arc; +use std::sync::atomic::{AtomicBool, AtomicU64, Ordering}; + +use serde_json::{Map, Value, json}; + +use fly_session_types::checkpoint::{self, Envelope}; + +// `crate::types` is this crate's facade over the shared `fly-session-types` crate; the glob +// keeps the contract's own names in sight instead of restating them. +use crate::types::*; + +/// The state-format identity this slice writes and reads. A checkpoint recorded under any +/// other one is refused by name rather than attempted. +pub const STATE_FORMAT_ID: &str = "flysess-1"; + +/// The worker capability `state-media-v1` section 5 makes the State methods conditional on. +pub const CHECKPOINT_CAPABILITY: &str = "checkpoint-v1"; + +/// The store manifest's own version. It is not the envelope version. +pub const STORE_MANIFEST_VERSION: u32 = 1; + +/// The file the store manifest is committed to. Its rename is the durable commit point. +pub const STORE_MANIFEST_FILE: &str = "manifest.json"; + +/// The content type a checkpoint payload travels under as a bus artifact. +pub const PAYLOAD_CONTENT_TYPE: &str = "application/x-fly-checkpoint-payload"; + +/// The attachment name a checkpoint payload travels under, in both directions. +pub const PAYLOAD_ATTACHMENT: &str = "payload"; + +/// Seals one checkpoint payload as an immutable artifact against its content digest. +/// +/// `state-media-v1` section 1 makes a digest mandatory on checkpoint payloads, so this seals +/// with one rather than computing it afterwards: a store that wrote the wrong bytes finds out +/// here and not at the next restore. +pub async fn seal_payload( + client: &flybus::Client, + bytes: &[u8], + digest: &Digest, +) -> DomainResult { + let mut writer = client + .artifacts() + .allocate(bytes.len() as u64, PAYLOAD_CONTENT_TYPE) + .await + .map_err(|e| store_error(format!("payload allocate: {}", e.message)))?; + writer + .write_all(bytes) + .map_err(|e| store_error(format!("payload write: {e}")))?; + let artifact = writer + .seal_with_digest(Some(digest.clone())) + .await + .map_err(|e| store_error(format!("payload seal: {}", e.message)))?; + match &artifact.reference().digest { + Some(sealed) if sealed == digest => Ok(artifact), + _ => Err(store_error("a sealed checkpoint payload has no matching content digest")), + } +} + +fn store_error(what: impl std::fmt::Display) -> DomainError { + DomainError::new(ErrorCode::BackendFailure, what, MutationCertainty::Unknown) +} + +fn incompatible(what: impl std::fmt::Display) -> DomainError { + DomainError::before(ErrorCode::IncompatibleState, what) +} + +// ---------------------------------------------------------------------------------------------- +// Payload names + +/// The payload name one agent's captured state is filed under. +pub fn agent_payload(agent_id: &str) -> String { + format!("agent-{agent_id}") +} + +/// The payload name one agent's action-executor state is filed under. +pub fn executor_payload(agent_id: &str) -> String { + format!("executor-{agent_id}") +} + +/// The environment's payload name. +pub const WORLD_PAYLOAD: &str = "world"; +/// The task ledger's payload name. +pub const TASK_LEDGER_PAYLOAD: &str = "task-ledger"; +/// The prior world inspection's payload name. +pub const PRIOR_INSPECTION_PAYLOAD: &str = "prior-inspection"; +/// The coordinator's admission state and event watermarks. +pub const ADMISSION_PAYLOAD: &str = "coordinator-admission"; + +// ---------------------------------------------------------------------------------------------- +// Compatibility + +/// The compatibility identities `state-media-v1` section 4 requires a checkpoint to record. +/// +/// Each one is a separate field on purpose: a restore that fails says which identity differs, +/// instead of reporting one opaque digest mismatch. The synthetic composition maps them onto +/// the identities it actually has: +/// +/// | Field | Where it comes from | +/// | --- | --- | +/// | `backend_digest` | the environment descriptor's backend identity | +/// | `content_digest` | the environment descriptor's content identity | +/// | `patch_digest` | the environment descriptor's resolved configuration identity | +/// | `controller_digest` | every declared port's controller schema, in descriptor order | +/// | `parser_digest` | the inspection schema and the task schema that reads it | +/// | `state_format_id` | [`STATE_FORMAT_ID`] | +#[derive(Clone, Debug, PartialEq, Eq)] +pub struct Compatibility { + pub backend_digest: Digest, + pub content_digest: Digest, + pub patch_digest: Digest, + pub controller_digest: Digest, + pub parser_digest: Digest, + pub state_format_id: Id, +} + +impl Compatibility { + /// The compatibility of one world, from the descriptor it advertises. + /// + /// Every identity comes from the descriptor, which is what lets the environment compute + /// exactly this digest for its own capture while the coordinator computes it for the + /// composition. The task's own schema is deliberately not in here: the captured task + /// ledger carries its schema and refuses another one, so folding it in would put one + /// identity in two places. + pub fn of(descriptor: &EnvironmentDescriptor) -> Compatibility { + let controllers = Value::Array( + descriptor + .ports + .iter() + .map(|port| { + json!({ + "portId": port.port_id.as_str(), + "controls": port.controls.to_json(), + }) + }) + .collect(), + ); + let parser = json!({ + "inspectionSchema": descriptor.inspection_schema.to_json(), + }); + Compatibility { + backend_digest: descriptor.backend_digest.clone(), + content_digest: descriptor.content_digest.clone(), + patch_digest: descriptor.configuration_digest.clone(), + controller_digest: digest_of(&controllers) + .expect("a validated controller schema canonicalizes"), + parser_digest: digest_of(&parser).expect("a validated schema reference canonicalizes"), + state_format_id: id(STATE_FORMAT_ID), + } + } + + pub fn to_json(&self) -> Value { + json!({ + "backendDigest": self.backend_digest.as_str(), + "contentDigest": self.content_digest.as_str(), + "patchDigest": self.patch_digest.as_str(), + "controllerDigest": self.controller_digest.as_str(), + "parserDigest": self.parser_digest.as_str(), + "stateFormatId": self.state_format_id.as_str(), + }) + } + + pub fn from_json(value: &Value) -> Result { + let digest = |key: &str| -> Result { + let text = value + .get(key) + .and_then(Value::as_str) + .ok_or_else(|| format!("compatibility: {key} is missing or not a string"))?; + if !is_digest(text) { + return Err(format!("compatibility: {key} is not a digest")); + } + Ok(text.to_owned()) + }; + let state_format_id = value + .get("stateFormatId") + .and_then(Value::as_str) + .ok_or_else(|| "compatibility: stateFormatId is missing".to_owned())?; + Ok(Compatibility { + backend_digest: digest("backendDigest")?, + content_digest: digest("contentDigest")?, + patch_digest: digest("patchDigest")?, + controller_digest: digest("controllerDigest")?, + parser_digest: digest("parserDigest")?, + state_format_id: parse_id(state_format_id) + .map_err(|e| format!("compatibility: stateFormatId {e}"))?, + }) + } + + /// The single digest a participant echoes in its capture and its stage request. + pub fn digest(&self) -> Digest { + digest_of(&self.to_json()).expect("a compatibility block canonicalizes") + } + + /// Names the first identity that differs. There is no tolerance and no "close enough". + pub fn compare(&self, live: &Compatibility) -> Result<(), String> { + for (field, recorded, current) in [ + ("stateFormatId", &self.state_format_id, &live.state_format_id), + ("backend", &self.backend_digest, &live.backend_digest), + ("content", &self.content_digest, &live.content_digest), + ("patch", &self.patch_digest, &live.patch_digest), + ("controller", &self.controller_digest, &live.controller_digest), + ("parser", &self.parser_digest, &live.parser_digest), + ] { + if recorded != current { + return Err(format!( + "the checkpoint's {field} identity {recorded} is not this composition's {current}" + )); + } + } + Ok(()) + } +} + +// ---------------------------------------------------------------------------------------------- +// The store manifest + +/// One committed generation, as the store manifest lists it. +#[derive(Clone, Debug, PartialEq, Eq)] +pub struct GenerationRecord { + pub checkpoint_id: Id, + /// The generation's file name inside the store directory. + pub file: String, + pub session_id: Id, + pub epoch: Id, + pub episode_id: Id, + pub boundary: u64, + pub compatibility_digest: Digest, + /// The SHA-256 of the whole envelope file, so a store can compare without opening it. + pub envelope_digest: Digest, + pub byte_length: u64, +} + +impl GenerationRecord { + fn to_json(&self) -> Value { + json!({ + "checkpointId": self.checkpoint_id.as_str(), + "file": self.file.as_str(), + "sessionId": self.session_id.as_str(), + "epoch": self.epoch.as_str(), + "episodeId": self.episode_id.as_str(), + "boundary": self.boundary.to_string(), + "compatibilityDigest": self.compatibility_digest.as_str(), + "envelopeDigest": self.envelope_digest.as_str(), + "byteLength": self.byte_length.to_string(), + }) + } + + fn from_json(value: &Value) -> Result { + let text = |key: &str| -> Result { + value + .get(key) + .and_then(Value::as_str) + .map(str::to_owned) + .ok_or_else(|| format!("store manifest: a generation has no {key}")) + }; + let number = |key: &str| -> Result { + text(key)? + .parse::() + .map_err(|_| format!("store manifest: {key} is not a canonical U64")) + }; + Ok(GenerationRecord { + checkpoint_id: parse_id(&text("checkpointId")?)?, + file: text("file")?, + session_id: parse_id(&text("sessionId")?)?, + epoch: parse_id(&text("epoch")?)?, + episode_id: parse_id(&text("episodeId")?)?, + boundary: number("boundary")?, + compatibility_digest: text("compatibilityDigest")?, + envelope_digest: text("envelopeDigest")?, + byte_length: number("byteLength")?, + }) + } +} + +/// The store's durable metadata: which generations exist and how far durability has reached. +#[derive(Clone, Debug, Default, PartialEq, Eq)] +pub struct StoreManifest { + pub generations: Vec, + /// The newest committed checkpoint, which is the high-water mark. `None` until the first + /// durable commit; that is "nothing has been committed", not a default. + pub high_water: Option, +} + +impl StoreManifest { + fn to_json(&self) -> Value { + json!({ + "storeManifestVersion": STORE_MANIFEST_VERSION, + "stateFormatId": STATE_FORMAT_ID, + "highWater": self.high_water.as_ref().map_or(Value::Null, |h| h.as_str().into()), + "generations": Value::Array(self.generations.iter().map(GenerationRecord::to_json).collect()), + }) + } + + fn from_json(value: &Value) -> Result { + let version = value + .get("storeManifestVersion") + .and_then(Value::as_u64) + .ok_or_else(|| "store manifest: no storeManifestVersion".to_owned())?; + if version != u64::from(STORE_MANIFEST_VERSION) { + return Err(format!("store manifest: unsupported version {version}")); + } + match value.get("stateFormatId").and_then(Value::as_str) { + Some(STATE_FORMAT_ID) => {} + Some(other) => { + return Err(format!("store manifest: state format {other} is not {STATE_FORMAT_ID}")); + } + None => return Err("store manifest: no stateFormatId".to_owned()), + } + let generations = value + .get("generations") + .and_then(Value::as_array) + .ok_or_else(|| "store manifest: generations must be an array".to_owned())? + .iter() + .map(GenerationRecord::from_json) + .collect::, _>>()?; + let high_water = match value.get("highWater") { + Some(Value::Null) | None => None, + Some(Value::String(s)) => Some(parse_id(s)?), + Some(_) => return Err("store manifest: highWater is neither null nor an Id".to_owned()), + }; + if let Some(mark) = &high_water + && !generations.iter().any(|g| g.checkpoint_id == *mark) + { + return Err("store manifest: the high-water mark names no listed generation".to_owned()); + } + Ok(StoreManifest { generations, high_water }) + } +} + +// ---------------------------------------------------------------------------------------------- +// The store + +/// How many committed generations the store keeps. +#[derive(Clone, Copy, Debug)] +pub struct StoreConfig { + /// Generations retained after a commit. The oldest are dropped, and only once the + /// manifest that no longer references them is itself committed. + pub keep_generations: usize, +} + +impl Default for StoreConfig { + fn default() -> StoreConfig { + StoreConfig { keep_generations: 3 } + } +} + +/// Deliberate durable-write faults, for the failure rows this slice has to demonstrate. +#[derive(Clone, Debug, Default)] +pub struct StoreFaults { + /// Stop after the generation file has been renamed and before the store manifest is + /// committed. The generation is then an unreferenced file, which is never a candidate. + pub stop_before_manifest_commit: bool, +} + +/// The durable checkpoint store: generation files, one store manifest and the commit order of +/// `checkpoint-envelope-v1` section 5. +pub struct CheckpointStore { + root: PathBuf, + config: StoreConfig, + manifest: StoreManifest, + faults: StoreFaults, + commits: u64, + /// Generations the committed manifest no longer references and whose files could not be + /// removed. + /// + /// Rotation happens after the durable commit point, so a file that will not unlink is a + /// leaked file and never a lost checkpoint. It is recorded rather than swallowed, because + /// a store that keeps failing to rotate is filling a disk quietly. + unrotated: Vec, +} + +impl CheckpointStore { + /// Opens or creates a store at `root`, reading whatever it already committed. + /// + /// A directory with no manifest is an empty store: nothing has been committed there yet. + /// A manifest that cannot be read is a failure, not an empty store. + pub fn open(root: impl Into, config: StoreConfig) -> DomainResult { + let root: PathBuf = root.into(); + std::fs::create_dir_all(&root) + .map_err(|e| store_error(format!("checkpoint store {}: {e}", root.display())))?; + let path = root.join(STORE_MANIFEST_FILE); + let manifest = match std::fs::read(&path) { + Ok(bytes) => { + let value: Value = serde_json::from_slice(&bytes) + .map_err(|e| store_error(format!("store manifest: {e}")))?; + StoreManifest::from_json(&value).map_err(store_error)? + } + Err(e) if e.kind() == std::io::ErrorKind::NotFound => StoreManifest::default(), + Err(e) => return Err(store_error(format!("store manifest: {e}"))), + }; + Ok(CheckpointStore { + root, + config, + manifest, + faults: StoreFaults::default(), + commits: 0, + unrotated: Vec::new(), + }) + } + + pub fn root(&self) -> &Path { + &self.root + } + + /// Re-reads the store manifest from disk. + /// + /// The store manifest is the durable metadata, and this session is not necessarily the + /// only thing that has ever written it: a previous run, a repair or an operator may have + /// committed or removed a generation. Reading it again is how a store finds that out, + /// rather than trusting a copy it happens to be holding. + pub fn reload(&mut self) -> DomainResult<()> { + let reopened = CheckpointStore::open(self.root.clone(), self.config)?; + self.manifest = reopened.manifest; + Ok(()) + } + + pub fn manifest(&self) -> &StoreManifest { + &self.manifest + } + + pub fn faults_mut(&mut self) -> &mut StoreFaults { + &mut self.faults + } + + /// How many durable commits this store has completed. + pub fn commits(&self) -> u64 { + self.commits + } + + /// Generations the committed manifest dropped whose files are still on disk. + pub fn unrotated(&self) -> &[String] { + &self.unrotated + } + + /// The committed generation with this checkpoint id, if the manifest lists it. + /// + /// This is the query a coordinator uses to resolve a save whose reply it never saw: it + /// asks the durable metadata about the *same* operation rather than saving again. + pub fn lookup(&self, checkpoint_id: &Id) -> Option<&GenerationRecord> { + self.manifest + .generations + .iter() + .find(|g| g.checkpoint_id == *checkpoint_id) + } + + /// The newest committed generation, or an explicit refusal when nothing is committed. + pub fn high_water(&self) -> Option<&GenerationRecord> { + let mark = self.manifest.high_water.as_ref()?; + self.lookup(mark) + } + + /// Selects a restore candidate: the named generation, or the high-water one. + pub fn select(&self, checkpoint_id: Option<&Id>) -> DomainResult { + match checkpoint_id { + Some(wanted) => self.lookup(wanted).cloned().ok_or_else(|| { + incompatible(format!( + "the store has no committed generation {wanted}; an unreferenced \ +temporary is never a restore candidate" + )) + }), + None => self.high_water().cloned().ok_or_else(|| { + incompatible("the store has committed no checkpoint to restore from") + }), + } + } + + /// Reads one committed generation back and validates the whole envelope. + pub fn read(&self, record: &GenerationRecord) -> DomainResult { + let path = self.root.join(&record.file); + let bytes = std::fs::read(&path) + .map_err(|e| store_error(format!("generation {}: {e}", record.file)))?; + if bytes.len() as u64 != record.byte_length { + return Err(incompatible(format!( + "generation {} is {} bytes; the store manifest records {}", + record.file, + bytes.len(), + record.byte_length + ))); + } + if digest_of_bytes(&bytes) != record.envelope_digest { + return Err(incompatible(format!( + "generation {} does not match the digest the store manifest records", + record.file + ))); + } + let envelope = checkpoint::decode(&bytes) + .map_err(|e| incompatible(format!("generation {}: {}", record.file, e.0)))?; + checkpoint::validate_manifest(&envelope) + .map_err(|e| incompatible(format!("generation {}: {}", record.file, e.0)))?; + Ok(envelope) + } + + /// The durable commit sequence of `checkpoint-envelope-v1` section 5, in that order. + /// + /// Blocking by construction: it fsyncs. The writer runs it off the session's runtime. + fn commit(&mut self, record: GenerationRecord, bytes: &[u8]) -> DomainResult<()> { + let temporary = self.root.join(format!("tmp-{}.flysess", record.checkpoint_id)); + let final_path = self.root.join(&record.file); + // 1. write the envelope to a temporary generation file, 2. fsync it + { + let mut file = std::fs::File::create(&temporary) + .map_err(|e| store_error(format!("generation temporary: {e}")))?; + file.write_all(bytes) + .map_err(|e| store_error(format!("generation temporary: {e}")))?; + file.sync_all() + .map_err(|e| store_error(format!("generation fsync: {e}")))?; + } + // 3. rename it to its final generation name, 4. fsync the store directory + std::fs::rename(&temporary, &final_path) + .map_err(|e| store_error(format!("generation rename: {e}")))?; + sync_dir(&self.root)?; + if self.faults.stop_before_manifest_commit { + // The generation file exists and nothing references it. It is not a restore + // candidate and the high-water mark has not moved. + return Err(store_error( + "injected failure after the generation was renamed and before the store \ +manifest was committed", + )); + } + // 5. write the store manifest to its own temporary, fsync, rename, fsync the directory. + let mut next = self.manifest.clone(); + next.generations.retain(|g| g.checkpoint_id != record.checkpoint_id); + next.generations.push(record.clone()); + next.high_water = Some(record.checkpoint_id.clone()); + let dropped = if next.generations.len() > self.config.keep_generations { + let excess = next.generations.len() - self.config.keep_generations; + next.generations.drain(..excess).collect::>() + } else { + Vec::new() + }; + let text = canonicalize(&next.to_json()) + .map_err(|e| store_error(format!("store manifest: {}", e.0)))?; + let manifest_temporary = self.root.join("tmp-manifest.json"); + { + let mut file = std::fs::File::create(&manifest_temporary) + .map_err(|e| store_error(format!("store manifest temporary: {e}")))?; + file.write_all(text.as_bytes()) + .map_err(|e| store_error(format!("store manifest temporary: {e}")))?; + file.sync_all() + .map_err(|e| store_error(format!("store manifest fsync: {e}")))?; + } + std::fs::rename(&manifest_temporary, self.root.join(STORE_MANIFEST_FILE)) + .map_err(|e| store_error(format!("store manifest rename: {e}")))?; + sync_dir(&self.root)?; + // Past the durable commit point. Rotation removes only files the committed manifest + // no longer references. + self.manifest = next; + self.commits += 1; + for old in dropped { + if std::fs::remove_file(self.root.join(&old.file)).is_err() { + self.unrotated.push(old.file); + } + } + Ok(()) + } +} + +fn sync_dir(path: &Path) -> DomainResult<()> { + let dir = std::fs::File::open(path) + .map_err(|e| store_error(format!("store directory {}: {e}", path.display())))?; + dir.sync_all() + .map_err(|e| store_error(format!("store directory fsync: {e}"))) +} + +// ---------------------------------------------------------------------------------------------- +// The checkpoint manifest this slice writes and reads + +/// One agent's row in a checkpoint manifest. +#[derive(Clone, Debug, PartialEq)] +pub struct AgentEntry { + pub agent_id: Id, + pub profile_digest: Digest, + pub dataset_digest: Digest, + pub model_version: String, + pub plasticity_version: String, + pub seed: i32, + pub brain_ticks: u64, + pub remainder: RationalNs, + pub payload: String, +} + +impl AgentEntry { + fn to_json(&self) -> Value { + json!({ + "agentId": self.agent_id.as_str(), + "profileDigest": self.profile_digest.as_str(), + "datasetDigest": self.dataset_digest.as_str(), + "modelVersion": self.model_version.as_str(), + "plasticityVersion": self.plasticity_version.as_str(), + "seed": self.seed, + "brainTicks": self.brain_ticks.to_string(), + "remainder": self.remainder.to_json(), + "payload": self.payload.as_str(), + }) + } + + fn from_json(value: &Value) -> Result { + let text = |key: &str| -> Result { + value + .get(key) + .and_then(Value::as_str) + .map(str::to_owned) + .ok_or_else(|| format!("checkpoint manifest: an agent row has no {key}")) + }; + let seed = value + .get("seed") + .and_then(Value::as_i64) + .ok_or_else(|| "checkpoint manifest: an agent row has no seed".to_owned())?; + let seed = i32::try_from(seed) + .map_err(|_| "checkpoint manifest: a seed is outside i32".to_owned())?; + let remainder = RationalNs::from_json( + value + .get("remainder") + .ok_or_else(|| "checkpoint manifest: an agent row has no remainder".to_owned())?, + ) + .map_err(|e| format!("checkpoint manifest: remainder: {}", e.0))?; + Ok(AgentEntry { + agent_id: parse_id(&text("agentId")?)?, + profile_digest: text("profileDigest")?, + dataset_digest: text("datasetDigest")?, + model_version: text("modelVersion")?, + plasticity_version: text("plasticityVersion")?, + seed, + brain_ticks: text("brainTicks")? + .parse() + .map_err(|_| "checkpoint manifest: brainTicks is not a canonical U64".to_owned())?, + remainder, + payload: text("payload")?, + }) + } +} + +/// The coordinator-owned state a checkpoint records, each as a payload name. +#[derive(Clone, Debug, PartialEq, Eq)] +pub struct CoordinatorEntry { + pub task_ledger: String, + pub prior_inspection: String, + pub executor_state: Vec<(Id, String)>, + pub admission_state: String, + pub event_watermarks: EventWatermarks, +} + +/// The event identity a resumed epoch continues from. +#[derive(Clone, Debug, PartialEq, Eq)] +pub struct EventWatermarks { + /// The highest source step any recorded event belongs to. + pub last_source_step: u64, + /// How many events the ledger has issued. + pub issued: u64, +} + +impl EventWatermarks { + fn to_json(&self) -> Value { + json!({ + "lastSourceStep": self.last_source_step.to_string(), + "issued": self.issued.to_string(), + }) + } + + fn from_json(value: &Value) -> Result { + let number = |key: &str| -> Result { + value + .get(key) + .and_then(Value::as_str) + .ok_or_else(|| format!("checkpoint manifest: eventWatermarks has no {key}"))? + .parse() + .map_err(|_| format!("checkpoint manifest: {key} is not a canonical U64")) + }; + Ok(EventWatermarks { + last_source_step: number("lastSourceStep")?, + issued: number("issued")?, + }) + } +} + +impl CoordinatorEntry { + fn to_json(&self) -> Value { + json!({ + "taskLedger": self.task_ledger.as_str(), + "priorInspection": self.prior_inspection.as_str(), + "executorState": Value::Array( + self.executor_state + .iter() + .map(|(agent_id, payload)| json!({ + "agentId": agent_id.as_str(), + "payload": payload.as_str(), + })) + .collect(), + ), + "admissionState": self.admission_state.as_str(), + "eventWatermarks": self.event_watermarks.to_json(), + }) + } + + fn from_json(value: &Value) -> Result { + let text = |key: &str| -> Result { + value + .get(key) + .and_then(Value::as_str) + .map(str::to_owned) + .ok_or_else(|| format!("checkpoint manifest: coordinator has no {key}")) + }; + let executors = value + .get("executorState") + .and_then(Value::as_array) + .ok_or_else(|| "checkpoint manifest: executorState must be an array".to_owned())?; + let mut executor_state = Vec::with_capacity(executors.len()); + for entry in executors { + let agent_id = entry + .get("agentId") + .and_then(Value::as_str) + .ok_or_else(|| "checkpoint manifest: an executor row has no agentId".to_owned())?; + let payload = entry + .get("payload") + .and_then(Value::as_str) + .ok_or_else(|| "checkpoint manifest: an executor row has no payload".to_owned())?; + executor_state.push((parse_id(agent_id)?, payload.to_owned())); + } + let watermarks = value + .get("eventWatermarks") + .ok_or_else(|| "checkpoint manifest: coordinator has no eventWatermarks".to_owned())?; + Ok(CoordinatorEntry { + task_ledger: text("taskLedger")?, + prior_inspection: text("priorInspection")?, + executor_state, + admission_state: text("admissionState")?, + event_watermarks: EventWatermarks::from_json(watermarks)?, + }) + } +} + +/// The environment's row: which worker the world belonged to and which payload holds it. +/// +/// `checkpoint-envelope-v1` section 3 named a holder for every payload except the world's; +/// the 2026-09-22 amendment to that section adds this one. +#[derive(Clone, Debug, PartialEq, Eq)] +pub struct EnvironmentEntry { + pub worker_id: Id, + pub payload: String, +} + +impl EnvironmentEntry { + fn to_json(&self) -> Value { + json!({ + "workerId": self.worker_id.as_str(), + "payload": self.payload.as_str(), + }) + } + + fn from_json(value: &Value) -> Result { + let text = |key: &str| -> Result { + value + .get(key) + .and_then(Value::as_str) + .map(str::to_owned) + .ok_or_else(|| format!("checkpoint manifest: environment has no {key}")) + }; + Ok(EnvironmentEntry { + worker_id: parse_id(&text("workerId")?)?, + payload: text("payload")?, + }) + } +} + +/// A complete checkpoint manifest, in the field names `checkpoint-envelope-v1` section 3 sets. +#[derive(Clone, Debug, PartialEq)] +pub struct CheckpointManifest { + pub checkpoint_id: Id, + pub source_scope: Scope, + pub episode_id: Id, + pub world_time: RationalNs, + pub scheduler_id: String, + pub composition_digest: Digest, + pub port_map: Vec<(Id, Id)>, + pub compatibility: Compatibility, + pub agents: Vec, + pub coordinator: CoordinatorEntry, + pub environment: EnvironmentEntry, + /// External-helper state required for exact resume, as payload names. The synthetic + /// composition has no external helper, so it records an empty list rather than omitting + /// the field: "no helper" is a statement, not a missing one. + pub helper_state: Vec, + pub payloads: Vec<(String, u64, Digest)>, +} + +impl CheckpointManifest { + pub fn to_json(&self) -> Value { + json!({ + "envelopeVersion": checkpoint::VERSION, + "checkpointId": self.checkpoint_id.as_str(), + "sourceScope": self.source_scope.to_json(), + "episodeId": self.episode_id.as_str(), + "worldTime": self.world_time.to_json(), + "schedulerId": self.scheduler_id.as_str(), + "compositionDigest": self.composition_digest.as_str(), + "portMap": Value::Array( + self.port_map + .iter() + .map(|(port_id, agent_id)| json!({ + "portId": port_id.as_str(), + "agentId": agent_id.as_str(), + })) + .collect(), + ), + "compatibility": self.compatibility.to_json(), + "agents": Value::Array(self.agents.iter().map(AgentEntry::to_json).collect()), + "coordinator": self.coordinator.to_json(), + "environment": self.environment.to_json(), + "helperState": Value::Array( + self.helper_state.iter().map(|n| Value::String(n.clone())).collect(), + ), + "payloads": Value::Array( + self.payloads + .iter() + .map(|(name, length, digest)| json!({ + "name": name.as_str(), + "byteLength": length.to_string(), + "digest": digest.as_str(), + })) + .collect(), + ), + }) + } + + /// Reads one back, refusing an incomplete manifest rather than filling anything in. + pub fn from_json(value: &Value) -> Result { + for field in checkpoint::REQUIRED_MANIFEST_FIELDS { + if value.get(*field).is_none() { + return Err(format!("checkpoint manifest: missing {field:?}")); + } + } + let text = |key: &str| -> Result { + value + .get(key) + .and_then(Value::as_str) + .map(str::to_owned) + .ok_or_else(|| format!("checkpoint manifest: {key} is missing or not a string")) + }; + let source_scope = Scope::from_json(&value["sourceScope"]) + .map_err(|e| format!("checkpoint manifest: sourceScope: {}", e.0))?; + let world_time = RationalNs::from_json(&value["worldTime"]) + .map_err(|e| format!("checkpoint manifest: worldTime: {}", e.0))?; + let mut port_map = Vec::new(); + for entry in value["portMap"] + .as_array() + .ok_or_else(|| "checkpoint manifest: portMap must be an array".to_owned())? + { + let port_id = entry + .get("portId") + .and_then(Value::as_str) + .ok_or_else(|| "checkpoint manifest: a port map row has no portId".to_owned())?; + let agent_id = entry + .get("agentId") + .and_then(Value::as_str) + .ok_or_else(|| "checkpoint manifest: a port map row has no agentId".to_owned())?; + port_map.push((parse_id(port_id)?, parse_id(agent_id)?)); + } + let agents = value["agents"] + .as_array() + .ok_or_else(|| "checkpoint manifest: agents must be an array".to_owned())? + .iter() + .map(AgentEntry::from_json) + .collect::, _>>()?; + let mut helper_state = Vec::new(); + for entry in value["helperState"] + .as_array() + .ok_or_else(|| "checkpoint manifest: helperState must be an array".to_owned())? + { + helper_state.push( + entry + .as_str() + .ok_or_else(|| "checkpoint manifest: a helper state entry is not a payload name".to_owned())? + .to_owned(), + ); + } + let mut payloads = Vec::new(); + for entry in value["payloads"] + .as_array() + .ok_or_else(|| "checkpoint manifest: payloads must be an array".to_owned())? + { + let name = entry + .get("name") + .and_then(Value::as_str) + .ok_or_else(|| "checkpoint manifest: a payload row has no name".to_owned())?; + let length: u64 = entry + .get("byteLength") + .and_then(Value::as_str) + .ok_or_else(|| "checkpoint manifest: a payload row has no byteLength".to_owned())? + .parse() + .map_err(|_| "checkpoint manifest: byteLength is not a canonical U64".to_owned())?; + let digest = entry + .get("digest") + .and_then(Value::as_str) + .ok_or_else(|| "checkpoint manifest: a payload row has no digest".to_owned())?; + payloads.push((name.to_owned(), length, digest.to_owned())); + } + Ok(CheckpointManifest { + checkpoint_id: parse_id(&text("checkpointId")?)?, + source_scope, + episode_id: parse_id(&text("episodeId")?)?, + world_time, + scheduler_id: text("schedulerId")?, + composition_digest: text("compositionDigest")?, + port_map, + compatibility: Compatibility::from_json(&value["compatibility"])?, + agents, + coordinator: CoordinatorEntry::from_json(&value["coordinator"])?, + environment: EnvironmentEntry::from_json( + value + .get("environment") + .ok_or_else(|| "checkpoint manifest: missing \"environment\"".to_owned())?, + )?, + helper_state, + payloads, + }) + } + + /// The payload name this manifest files one participant's state under. + pub fn payload_of(&self, worker_id: &Id) -> Option<&str> { + if self.environment.worker_id == *worker_id { + return Some(self.environment.payload.as_str()); + } + self.agents + .iter() + .find(|a| a.agent_id == *worker_id) + .map(|a| a.payload.as_str()) + } +} + +// ---------------------------------------------------------------------------------------------- +// The bounded writer + +/// One participant's captured payload, with the owned handle the writer keeps until the bytes +/// are committed or the job fails. +pub struct CapturedPayload { + pub name: String, + pub artifact: flybus::Artifact, + pub byte_length: u64, + pub digest: Digest, +} + +/// What a coordinator hands the writer once every participant has captured. +pub struct CaptureSubmission { + pub checkpoint_id: Id, + pub boundary: u64, + pub session_id: Id, + pub epoch: Id, + pub episode_id: Id, + pub compatibility_digest: Digest, + pub manifest: Value, + pub payloads: Vec, + /// A hot checkpoint may replace a queued hot checkpoint, releasing its holds. A durable + /// one never is: the retention table coalesces only queued replaceable captures. + pub replaceable: bool, +} + +impl CaptureSubmission { + fn byte_length(&self) -> u64 { + self.payloads.iter().map(|p| p.byte_length).sum() + } +} + +/// How one durable save ended. Every variant is a statement; none of them is a default. +#[derive(Clone, Debug, PartialEq, Eq)] +pub enum SaveOutcome { + /// Past the store manifest rename. This is the only variant that is a saved + /// acknowledgment and the only one that moves a high-water mark. + Committed { checkpoint_id: Id, boundary: u64, file: String }, + /// The write failed and its owned captures were released under the retry policy. It + /// never reports false durability. + Failed { checkpoint_id: Id, reason: String }, + /// A later replaceable capture took this one's place in the queue before it was written. + Superseded { checkpoint_id: Id, by: Id }, + /// The writer finished this job and its reply channel was gone before the outcome could + /// be delivered. The write is over and its result is unknown from here. + ReplyLost { checkpoint_id: Id }, + /// The caller's own budget ran out while the job was still queued or being written. The + /// save is not over: it may commit after this is reported. + /// + /// It is a different fact from [`SaveOutcome::ReplyLost`] and is named separately because + /// diagnosing one as the other is exactly the implicit best-effort reading these + /// contracts refuse. Both leave durable metadata where it was, and for both the caller + /// resolves the *same* operation against the store manifest instead of saving again -- + /// but only one of them is a save that has already stopped. + DeadlineExpired { checkpoint_id: Id }, +} + +impl SaveOutcome { + pub fn checkpoint_id(&self) -> &Id { + match self { + SaveOutcome::Committed { checkpoint_id, .. } + | SaveOutcome::Failed { checkpoint_id, .. } + | SaveOutcome::Superseded { checkpoint_id, .. } + | SaveOutcome::ReplyLost { checkpoint_id } + | SaveOutcome::DeadlineExpired { checkpoint_id } => checkpoint_id, + } + } + + /// The event name this outcome publishes under. + /// + /// Only the three the writer itself produces are ever published; the two caller-side + /// outcomes are what a caller saw, not what the store did, and the store does not announce + /// them on its own topic. + pub fn event(&self) -> &'static str { + match self { + SaveOutcome::Committed { .. } => "committed", + SaveOutcome::Failed { .. } => "failed", + SaveOutcome::Superseded { .. } => "superseded", + SaveOutcome::ReplyLost { .. } | SaveOutcome::DeadlineExpired { .. } => "failed", + } + } + + /// True only past the durable commit point. + pub fn is_durable(&self) -> bool { + matches!(self, SaveOutcome::Committed { .. }) + } +} + +/// What the writer itself can produce for one job. +/// +/// The two caller-side outcomes -- a lost reply and an expired caller deadline -- are not in +/// here, because the writer cannot observe either. Keeping them out is what stops the writer's +/// own bookkeeping from carrying arms that can never run. +#[derive(Clone, Debug, PartialEq, Eq)] +enum WriteOutcome { + Committed { boundary: u64, file: String }, + Failed { reason: String }, +} + +impl WriteOutcome { + fn into_save(self, checkpoint_id: &Id) -> SaveOutcome { + match self { + WriteOutcome::Committed { boundary, file } => SaveOutcome::Committed { + checkpoint_id: checkpoint_id.clone(), + boundary, + file, + }, + WriteOutcome::Failed { reason } => SaveOutcome::Failed { + checkpoint_id: checkpoint_id.clone(), + reason, + }, + } + } +} + +/// What a failed write does with the ephemeral captures it owns. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub enum RetryPolicy { + /// Release the owned captures and report the failure. The default: a capture is cheap to + /// take again at the next committed boundary, and holding one is not free. + ReleaseAndReport, + /// Try the durable sequence again, up to `attempts` times in total, keeping the owned + /// captures until they are committed or the attempts are spent. + RetryThenRelease { attempts: u32 }, +} + +/// The writer's bounds. Both are finite, and they refuse at different moments. +/// +/// [`WriterConfig::queue_capacity`] is the one a capture is refused *before* it is requested: +/// [`CheckpointWriter::reserve`] takes its slot first, which is what the durable row of +/// `state-media-v1` section 3 means by rejecting before capture. The byte budget cannot work +/// that way, because how many bytes a capture is worth is not known until the participants +/// have produced it; it is checked at [`CheckpointWriter::submit`], so an oversized capture is +/// refused after it exists and before it is queued, and its payloads are released with the +/// refusal. Both are named `BUSY` refusals and neither fails the epoch. +#[derive(Clone, Copy, Debug)] +pub struct WriterConfig { + /// Outstanding coherent captures. `state-media-v1` section 3's initial session default + /// is two. Refused before a capture is requested. + pub queue_capacity: usize, + /// The total payload bytes the queue may hold. Refused at submit, once the size is known. + pub max_queued_bytes: u64, + pub retry: RetryPolicy, +} + +impl Default for WriterConfig { + fn default() -> WriterConfig { + WriterConfig { + queue_capacity: 2, + max_queued_bytes: 64 * 1024 * 1024, + retry: RetryPolicy::ReleaseAndReport, + } + } +} + +/// Deliberate writer faults, for the rows this slice has to demonstrate. +#[derive(Clone, Default)] +pub struct WriterFaults { + /// Hold every job until the gate is opened, so a test can fill the queue on purpose. + pub gate: Option>, + /// Drop this job's reply channel after the durable sequence ran, which is a lost save + /// reply. + pub drop_reply_for: Option, +} + +impl std::fmt::Debug for WriterFaults { + fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { + f.debug_struct("WriterFaults") + .field("gate", &self.gate.is_some()) + .field("drop_reply_for", &self.drop_reply_for) + .finish() + } +} + +/// The writer's own counters, so "bounded" is something a test reads rather than believes. +#[derive(Clone, Copy, Debug, Default, PartialEq, Eq)] +pub struct WriterStats { + pub queued: u64, + pub committed: u64, + pub failed: u64, + pub superseded: u64, + /// Submissions refused because the queue or its byte budget was full. + pub rejected: u64, + /// Checkpoint events the bus would not take. The durable outcome the caller receives is + /// the authority; a publication that failed is counted here rather than disappearing. + pub events_dropped: u64, + /// The deepest the queue ever got. + pub peak_queue: usize, + /// The most payload bytes the queue ever held. + pub peak_bytes: u64, +} + +struct Job { + submission: CaptureSubmission, + reply: tokio::sync::oneshot::Sender, + permit: tokio::sync::OwnedSemaphorePermit, +} + +struct WriterShared { + config: WriterConfig, + faults: WriterFaults, + permits: Arc, + queue: std::sync::Mutex>, + queued_bytes: AtomicU64, + stats: std::sync::Mutex, + wake: tokio::sync::Notify, + stop: AtomicBool, + store: Arc>, + events: Option<(flybus::Client, String)>, +} + +/// A queue slot, taken before a capture is requested so a saturated writer refuses early. +/// +/// `state-media-v1` section 3's durable row says to reject or defer *before capture* when +/// saturated. Holding the slot from before `State.Capture` until the outcome is delivered is +/// what makes that true rather than aspirational. +pub struct Reservation { + permit: tokio::sync::OwnedSemaphorePermit, +} + +/// The bounded checkpoint writer. +pub struct CheckpointWriter { + shared: Arc, + task: tokio::task::JoinHandle<()>, +} + +impl CheckpointWriter { + /// Starts the writer over `store`. `events` is the bus client and topic the distinct + /// captured/queued/committed/failed/superseded publications go to. + pub fn start( + store: CheckpointStore, + config: WriterConfig, + faults: WriterFaults, + events: Option<(flybus::Client, String)>, + ) -> CheckpointWriter { + let shared = Arc::new(WriterShared { + config, + faults, + permits: Arc::new(tokio::sync::Semaphore::new(config.queue_capacity)), + queue: std::sync::Mutex::new(VecDeque::new()), + queued_bytes: AtomicU64::new(0), + stats: std::sync::Mutex::new(WriterStats::default()), + wake: tokio::sync::Notify::new(), + stop: AtomicBool::new(false), + store: Arc::new(std::sync::Mutex::new(store)), + events, + }); + let task = tokio::spawn(run_writer(shared.clone())); + CheckpointWriter { shared, task } + } + + pub fn stats(&self) -> WriterStats { + *self.shared.stats.lock().expect("the writer stats are never poisoned") + } + + pub fn config(&self) -> WriterConfig { + self.shared.config + } + + /// How many captures are outstanding: queued or being written. + pub fn outstanding(&self) -> usize { + self.shared.config.queue_capacity - self.shared.permits.available_permits() + } + + /// Takes a queue slot, or refuses by name. Nothing is captured without one. + pub fn reserve(&self) -> DomainResult { + match self.shared.permits.clone().try_acquire_owned() { + Ok(permit) => Ok(Reservation { permit }), + Err(_) => { + self.shared + .stats + .lock() + .expect("the writer stats are never poisoned") + .rejected += 1; + Err(DomainError::before( + ErrorCode::Busy, + format!( + "the checkpoint queue already holds its {} outstanding captures; a \ +capture is refused before it is requested rather than queued without bound", + self.shared.config.queue_capacity + ), + )) + } + } + } + + /// Hands one complete capture to the writer and returns the channel its outcome arrives + /// on. The reservation becomes the job's slot. + pub async fn submit( + &self, + reservation: Reservation, + submission: CaptureSubmission, + ) -> DomainResult> { + let bytes = submission.byte_length(); + let budget = self.shared.config.max_queued_bytes; + let (reply, receiver) = tokio::sync::oneshot::channel(); + let checkpoint_id = submission.checkpoint_id.clone(); + let queued_event = submission.as_event("queued"); + // Reading the byte total, deciding on it and changing it are one critical section. + // Only one caller submits today, so a split could not be observed -- but a bound that + // is only correct while nobody else is submitting is not a bound. + let superseded = { + let mut queue = self.shared.queue.lock().expect("the writer queue is never poisoned"); + let replaced_index = if submission.replaceable { + queue.iter().position(|job| job.submission.replaceable) + } else { + None + }; + // A capture that will take a queued replaceable one's place frees its bytes, so + // the budget is decided against what the queue will hold and not what it holds. + let freed = replaced_index.map_or(0, |index| queue[index].submission.byte_length()); + let held = self.shared.queued_bytes.load(Ordering::SeqCst); + let after = held.saturating_sub(freed) + bytes; + if after > budget { + self.shared + .stats + .lock() + .expect("the writer stats are never poisoned") + .rejected += 1; + return Err(DomainError::before( + ErrorCode::Busy, + format!( + "the checkpoint queue would hold {after} of {budget} bytes; the byte \ +budget is finite and refuses before it is exceeded" + ), + )); + } + let replaced = replaced_index.map(|index| queue.remove(index).expect("just found")); + if let Some(old) = &replaced { + self.shared + .queued_bytes + .fetch_sub(old.submission.byte_length(), Ordering::SeqCst); + } + queue.push_back(Job { submission, reply, permit: reservation.permit }); + self.shared.queued_bytes.fetch_add(bytes, Ordering::SeqCst); + let mut stats = self.shared.stats.lock().expect("the writer stats are never poisoned"); + stats.queued += 1; + stats.peak_queue = stats.peak_queue.max(queue.len()); + // Sampled after the superseded job's bytes are gone, so the peak is a total the + // queue really held. + stats.peak_bytes = stats + .peak_bytes + .max(self.shared.queued_bytes.load(Ordering::SeqCst)); + replaced + }; + if let Some(old) = superseded { + self.shared + .stats + .lock() + .expect("the writer stats are never poisoned") + .superseded += 1; + let outcome = SaveOutcome::Superseded { + checkpoint_id: old.submission.checkpoint_id.clone(), + by: checkpoint_id.clone(), + }; + publish_event( + &self.shared.events, + &self.shared.stats, + &outcome_event(&old.submission, &outcome), + ) + .await; + // A caller that dropped its ticket is not waiting for this; the store's own + // metadata is the durable record either way. + let _ = old.reply.send(outcome); + // Dropping the job releases its owned captures and its queue slot. + drop(old.submission); + drop(old.permit); + } + publish_event(&self.shared.events, &self.shared.stats, &object(queued_event)).await; + self.shared.wake.notify_one(); + Ok(receiver) + } + + /// Waits for one save's outcome, within the caller's own budget. + /// + /// The two ways this ends without an outcome are different facts and are reported as + /// themselves: the channel closing means the writer finished and the reply did not reach + /// here, and the budget running out means the save is still going. Neither is a hang and + /// neither is a save. + pub async fn wait( + receiver: tokio::sync::oneshot::Receiver, + checkpoint_id: &Id, + budget: std::time::Duration, + ) -> SaveOutcome { + match tokio::time::timeout(budget, receiver).await { + Ok(Ok(outcome)) => outcome, + Ok(Err(_)) => SaveOutcome::ReplyLost { checkpoint_id: checkpoint_id.clone() }, + Err(_) => SaveOutcome::DeadlineExpired { checkpoint_id: checkpoint_id.clone() }, + } + } + + /// Runs `f` against the store, which is how a caller resolves a save whose reply it lost. + /// + /// The store's own work is blocking -- it fsyncs -- so it is reached on a blocking thread + /// rather than from the session's runtime. + pub async fn with_store(&self, f: F) -> T + where + F: FnOnce(&mut CheckpointStore) -> T + Send + 'static, + T: Send + 'static, + { + let store = self.shared.store.clone(); + tokio::task::spawn_blocking(move || { + let mut store = store.lock().expect("the checkpoint store is never poisoned"); + f(&mut store) + }) + .await + .expect("the checkpoint store task is never cancelled") + } + + /// Stops the writer and waits for its task. A leftover writer task would hold artifact + /// handles the session has finished with. + pub async fn shutdown(self) { + self.shared.stop.store(true, Ordering::SeqCst); + self.shared.wake.notify_one(); + let _ = self.task.await; + let mut queue = self.shared.queue.lock().expect("the writer queue is never poisoned"); + queue.clear(); + } +} + +impl CaptureSubmission { + fn as_event(&self, event: &'static str) -> Value { + json!({ + "event": event, + "checkpointId": self.checkpoint_id.as_str(), + "sessionId": self.session_id.as_str(), + "epoch": self.epoch.as_str(), + "episodeId": self.episode_id.as_str(), + "boundary": self.boundary.to_string(), + "byteLength": self.byte_length().to_string(), + }) + } +} + +fn outcome_event(submission: &CaptureSubmission, outcome: &SaveOutcome) -> Map { + let mut payload = object(submission.as_event(outcome.event())); + match outcome { + SaveOutcome::Committed { file, .. } => { + payload.insert("generation".into(), file.as_str().into()); + payload.insert("durable".into(), true.into()); + } + SaveOutcome::Failed { reason, .. } => { + payload.insert("reason".into(), reason.as_str().into()); + payload.insert("durable".into(), false.into()); + } + SaveOutcome::Superseded { by, .. } => { + payload.insert("supersededBy".into(), by.as_str().into()); + payload.insert("durable".into(), false.into()); + } + SaveOutcome::ReplyLost { .. } => { + payload.insert("reason".into(), "the save reply was lost".into()); + payload.insert("durable".into(), false.into()); + } + SaveOutcome::DeadlineExpired { .. } => { + payload.insert("reason".into(), "the caller's durable budget expired".into()); + payload.insert("durable".into(), false.into()); + } + } + payload +} + +async fn publish_event( + events: &Option<(flybus::Client, String)>, + stats: &std::sync::Mutex, + payload: &Map, +) { + let Some((client, topic)) = events else { return }; + // A checkpoint event is telemetry about the store, not the durable record. A publication + // that cannot be delivered never changes what was committed, so the outcome the caller + // receives stays the authority -- and the drop is counted rather than retried, because a + // retry here would be exactly the implicit best-effort policy these contracts refuse. + if client.publish(topic, payload.clone(), &[]).await.is_err() { + stats + .lock() + .expect("the writer stats are never poisoned") + .events_dropped += 1; + } +} + +async fn run_writer(shared: Arc) { + loop { + let job = { + let mut queue = shared.queue.lock().expect("the writer queue is never poisoned"); + queue.pop_front() + }; + let Some(job) = job else { + if shared.stop.load(Ordering::SeqCst) { + return; + } + shared.wake.notified().await; + continue; + }; + if let Some(gate) = &shared.faults.gate { + // A deliberately stalled writer. The queue in front of it stays bounded, which is + // the point of the stall. + if let Ok(permit) = gate.acquire().await { + permit.forget(); + } + } + let Job { submission, reply, permit } = job; + shared + .queued_bytes + .fetch_sub(submission.byte_length(), Ordering::SeqCst); + let written = write_one(&shared, &submission).await; + { + let mut stats = shared.stats.lock().expect("the writer stats are never poisoned"); + match &written { + WriteOutcome::Committed { .. } => stats.committed += 1, + WriteOutcome::Failed { .. } => stats.failed += 1, + } + } + let outcome = written.into_save(&submission.checkpoint_id); + publish_event(&shared.events, &shared.stats, &outcome_event(&submission, &outcome)).await; + let lost = shared.faults.drop_reply_for.as_ref() == Some(&submission.checkpoint_id); + // The writer owned these handles until the bytes were committed or the job failed. + // The job is over, so they and its queue slot go before the outcome is delivered: + // a caller that reads the queue depth the moment its outcome arrives must not see a + // slot this job has finished with. + drop(submission); + drop(permit); + if lost { + // The bytes are committed and the acknowledgment is lost. Durable metadata is in + // the store manifest, which is exactly where the caller must look. + drop(reply); + } else { + // A caller that dropped its ticket is not waiting for this. + let _ = reply.send(outcome); + } + } +} + +async fn write_one(shared: &Arc, submission: &CaptureSubmission) -> WriteOutcome { + let mut payloads = Vec::with_capacity(submission.payloads.len()); + for payload in &submission.payloads { + let bytes = match payload.artifact.read_all().await { + Ok(bytes) => bytes, + Err(e) => { + return WriteOutcome::Failed { + reason: format!("payload {}: {}", payload.name, e.message), + }; + } + }; + if bytes.len() as u64 != payload.byte_length || digest_of_bytes(&bytes) != payload.digest { + return WriteOutcome::Failed { + reason: format!( + "payload {} is not the content its capture declared", + payload.name + ), + }; + } + payloads.push((payload.name.clone(), bytes)); + } + let bytes = match checkpoint::encode(&submission.manifest, &payloads) { + Ok(bytes) => bytes, + Err(e) => { + return WriteOutcome::Failed { + reason: format!("envelope: {}", e.0), + }; + } + }; + let record = GenerationRecord { + checkpoint_id: submission.checkpoint_id.clone(), + file: format!("{}.flysess", submission.checkpoint_id), + session_id: submission.session_id.clone(), + epoch: submission.epoch.clone(), + episode_id: submission.episode_id.clone(), + boundary: submission.boundary, + compatibility_digest: submission.compatibility_digest.clone(), + envelope_digest: digest_of_bytes(&bytes), + byte_length: bytes.len() as u64, + }; + let attempts = match shared.config.retry { + RetryPolicy::ReleaseAndReport => 1, + RetryPolicy::RetryThenRelease { attempts } => attempts.max(1), + }; + let mut last = String::new(); + for _ in 0..attempts { + let record = record.clone(); + let bytes = bytes.clone(); + let store = shared.store.clone(); + // The durable sequence fsyncs twice. It runs on a blocking thread so the session's + // runtime is never the thing waiting on a disk. + let result = tokio::task::spawn_blocking(move || { + let mut store = store.lock().expect("the checkpoint store is never poisoned"); + store.commit(record, &bytes) + }) + .await; + match result { + Ok(Ok(())) => { + return WriteOutcome::Committed { + boundary: submission.boundary, + file: format!("{}.flysess", submission.checkpoint_id), + }; + } + Ok(Err(e)) => last = e.message, + Err(e) => last = format!("the checkpoint writer stopped: {e}"), + } + } + WriteOutcome::Failed { reason: last } +} diff --git a/services/flysim/crates/fly-session/src/task.rs b/services/flysim/crates/fly-session/src/task.rs index b48e909..2eebbef 100644 --- a/services/flysim/crates/fly-session/src/task.rs +++ b/services/flysim/crates/fly-session/src/task.rs @@ -39,6 +39,16 @@ pub fn episode_schema() -> SchemaRef { synthetic_schema("arena.episode.v1", 1) } +/// The schema of a captured task ledger. +pub fn ledger_schema() -> SchemaRef { + synthetic_schema("arena.ledger.v1", 1) +} + +/// The schema of a captured action-executor state. +pub fn executor_schema() -> SchemaRef { + synthetic_schema("arena.executor.v1", 1) +} + pub fn controller_schema_ref() -> SchemaRef { synthetic_schema("arena.controller.v1", 1) } @@ -86,6 +96,29 @@ pub trait Task: Send { /// How many times `evaluate_transition` has run. A transition must evaluate once. fn evaluations(&self) -> u64; + + /// The checkpointable ledger at a committed boundary (`workers-v1` section 4). + fn capture(&self) -> DomainResult; + + /// Validates a captured ledger without installing it, so a group install can fail before + /// anything is changed. + fn validate_restore(&self, state: &TypedValue) -> DomainResult<()>; + + /// Installs a validated ledger under `epoch`. Event identity is derived from the epoch, + /// so the new one is part of the install rather than something the ledger keeps from the + /// epoch it was captured in. + fn install_restore(&mut self, epoch: &Id, state: &TypedValue) -> DomainResult<()>; + + /// Every event identity this ledger has issued, mapped onto the identity it would have + /// under `to_epoch`. + /// + /// `workers-v1` section 4 derives an event id from the epoch, so a trace recorded in one + /// epoch cannot be compared with a trace recorded in another until these are rebased. + /// The ledger owns the derivation, so it is the only thing that can do it. + fn rebase_ids(&self, to_epoch: &Id) -> DomainResult>; + + /// How far event identity has reached: the highest source step and the number issued. + fn event_watermarks(&self) -> (u64, u64); } /// Translates one selected decision into a controller intent, with no port assignment. @@ -98,6 +131,15 @@ pub trait ActionExecutor: Send { progress: &TypedValue, clock: &RationalNs, ) -> DomainResult<(ControllerIntent, Vec)>; + + /// Per-executor state at a committed boundary (`workers-v1` section 4). + fn capture(&self) -> DomainResult; + + /// Validates a captured executor state without installing it. + fn validate_restore(&self, state: &TypedValue) -> DomainResult<()>; + + /// Installs a validated executor state. + fn install_restore(&mut self, state: &TypedValue) -> DomainResult<()>; } /// The only executor v1 supports: it passes a direct-control decision through unchanged. @@ -123,6 +165,37 @@ impl ActionExecutor for IdentityExecutor { .map_err(|e| DomainError::invalid(format!("decision: {e}")))?; Ok((intent, Vec::new())) } + + /// The identity executor is stateless, and says so rather than capturing nothing. + /// + /// An empty object would be indistinguishable from a stateful executor whose capture went + /// missing, so the capture names the executor it came from and a restore refuses any + /// other one. + fn capture(&self) -> DomainResult { + TypedValue::new(executor_schema(), json!({"executor": "identity-v1"})) + .map_err(|e| DomainError::invalid(e.0)) + } + + fn validate_restore(&self, state: &TypedValue) -> DomainResult<()> { + if state.schema != executor_schema() { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + "the captured executor state does not carry the executor schema", + )); + } + match state.value.get("executor").and_then(Value::as_str) { + Some("identity-v1") => Ok(()), + other => Err(DomainError::before( + ErrorCode::IncompatibleState, + format!("the captured executor is {other:?}, not the identity executor"), + )), + } + } + + fn install_restore(&mut self, state: &TypedValue) -> DomainResult<()> { + // Stateless: validation is the whole of the install, and it is not skipped. + self.validate_restore(state) + } } /// When the counter task asks for a terminal episode transition. @@ -146,6 +219,10 @@ pub struct CounterTask { evaluations: u64, total_reward: f64, counter: i64, + /// The highest source step any issued event belongs to, and how many were issued. These + /// are the event watermarks a checkpoint records and a resumed epoch continues from. + last_source_step: u64, + issued_events: u64, terminal: Terminal, } @@ -159,6 +236,8 @@ impl CounterTask { evaluations: 0, total_reward: 0.0, counter: 0, + last_source_step: 0, + issued_events: 0, terminal, } } @@ -239,6 +318,7 @@ impl Task for CounterTask { payload: TypedValue::new(event_schema(), json!({"counter": self.counter})) .expect("a synthetic typed value fits the contract"), }]; + self.issued_events += events.len() as u64; Ok(Bootstrap { contexts, progress: self.progress_value(), events }) } @@ -314,6 +394,8 @@ impl Task for CounterTask { )); } + self.last_source_step = self.last_source_step.max(source_step); + self.issued_events += events.len() as u64; let next_contexts = self .agents .iter() @@ -346,6 +428,173 @@ impl Task for CounterTask { fn evaluations(&self) -> u64 { self.evaluations } + + fn capture(&self) -> DomainResult { + TypedValue::new( + ledger_schema(), + json!({ + "epoch": self.epoch.as_str(), + "agents": self.agents.iter().map(String::as_str).collect::>(), + "bindings": self + .bindings + .iter() + .map(|b| json!({"portId": b.port_id.as_str(), "agentId": b.agent_id.as_str()})) + .collect::>(), + "transitions": self.transitions, + "evaluations": self.evaluations, + "totalReward": self.total_reward, + "counter": self.counter, + "lastSourceStep": self.last_source_step, + "issuedEvents": self.issued_events, + }), + ) + .map_err(|e| DomainError::invalid(e.0)) + } + + fn validate_restore(&self, state: &TypedValue) -> DomainResult<()> { + if state.schema != ledger_schema() { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + "the captured ledger does not carry this task's schema", + )); + } + for field in [ + "epoch", + "agents", + "bindings", + "transitions", + "evaluations", + "totalReward", + "counter", + "lastSourceStep", + "issuedEvents", + ] { + if state.value.get(field).is_none() { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + format!("the captured ledger has no {field}"), + )); + } + } + let bindings = state + .value + .get("bindings") + .and_then(Value::as_array) + .ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "the captured ledger's bindings are not a list", + ) + })?; + if bindings.len() != self.bindings.len() && !self.bindings.is_empty() { + return Err(DomainError::before( + ErrorCode::IncompatibleState, + "the captured ledger binds another number of ports", + )); + } + Ok(()) + } + + fn install_restore(&mut self, epoch: &Id, state: &TypedValue) -> DomainResult<()> { + self.validate_restore(state)?; + let number = |key: &str| -> DomainResult { + state.value.get(key).and_then(Value::as_u64).ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + format!("the captured ledger's {key} is not a whole number"), + ) + }) + }; + let mut agents = Vec::new(); + for value in state.value["agents"].as_array().expect("validated") { + let agent = value.as_str().ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "the captured ledger names an agent that is not a string", + ) + })?; + agents.push(parse_id(agent).map_err(|e| { + DomainError::before(ErrorCode::IncompatibleState, format!("ledger: {e}")) + })?); + } + let mut bindings = Vec::new(); + for value in state.value["bindings"].as_array().expect("validated") { + let port_id = value.get("portId").and_then(Value::as_str).ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "the captured ledger has a binding with no portId", + ) + })?; + let agent_id = value.get("agentId").and_then(Value::as_str).ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "the captured ledger has a binding with no agentId", + ) + })?; + bindings.push(PortBinding { + port_id: parse_id(port_id).map_err(|e| { + DomainError::before(ErrorCode::IncompatibleState, format!("ledger: {e}")) + })?, + agent_id: parse_id(agent_id).map_err(|e| { + DomainError::before(ErrorCode::IncompatibleState, format!("ledger: {e}")) + })?, + }); + } + let counter = state.value.get("counter").and_then(Value::as_i64).ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "the captured ledger's counter is not an integer", + ) + })?; + let total_reward = state + .value + .get("totalReward") + .and_then(Value::as_f64) + .filter(|v| v.is_finite()) + .ok_or_else(|| { + DomainError::before( + ErrorCode::IncompatibleState, + "the captured ledger's totalReward is not a finite number", + ) + })?; + // The epoch is the caller's, not the capture's: event identity belongs to the epoch + // the ledger is being installed into. + self.epoch = epoch.clone(); + self.agents = agents; + self.bindings = bindings; + self.transitions = number("transitions")?; + self.evaluations = number("evaluations")?; + self.total_reward = total_reward; + self.counter = counter; + self.last_source_step = number("lastSourceStep")?; + self.issued_events = number("issuedEvents")?; + Ok(()) + } + + fn rebase_ids(&self, to_epoch: &Id) -> DomainResult> { + let mut out = BTreeMap::new(); + out.insert( + event_id(&self.epoch, 0, "bootstrap", 0), + event_id(to_epoch, 0, "bootstrap", 0), + ); + // The counter task issues exactly one `counter-delta` event per bound port per + // evaluated transition, in descriptor port order, so every identity it has ever + // issued is re-derivable from its ledger without keeping a list of them. + let ports = self.bindings.len() as u32; + for source_step in 1..=self.last_source_step { + for ordinal in 0..ports { + out.insert( + event_id(&self.epoch, source_step, "counter-delta", ordinal), + event_id(to_epoch, source_step, "counter-delta", ordinal), + ); + } + } + Ok(out) + } + + fn event_watermarks(&self) -> (u64, u64) { + (self.last_source_step, self.issued_events) + } } /// The inspection value the counter environment publishes. diff --git a/services/flysim/crates/fly-session/src/types.rs b/services/flysim/crates/fly-session/src/types.rs index 4d1eec0..8648308 100644 --- a/services/flysim/crates/fly-session/src/types.rs +++ b/services/flysim/crates/fly-session/src/types.rs @@ -8,13 +8,18 @@ //! `Id` and `Digest` are type aliases, because the shared crate carries both as validated //! `String`s from `flybus::wire` rather than forking the encodings into newtypes. +use std::collections::BTreeMap; + use serde_json::{Map, Value}; pub use fly_session_types::ArtifactRef; pub use fly_session_types::canonical::{ self, OperationKey, body_digest, canonicalize, digest_of, sha256_hex, }; -pub use fly_session_types::media::{AudioDescriptor, AudioRef, ViewDescriptor, ViewRef}; +pub use fly_session_types::media::{ + ActivateRestoreParams, ActivateRestoreResult, AudioDescriptor, AudioRef, CaptureParams, + CaptureResult, StageRestoreParams, StageRestoreResult, ViewDescriptor, ViewRef, +}; pub use fly_session_types::rpc::{ ErrorCode, MutationCertainty, SessionRpcFailure, SessionRpcOutcome, SessionRpcRequest, SessionRpcSuccess, @@ -217,6 +222,57 @@ pub fn outcome_identity( } } +/// The epoch-derived identities of a behaviour trace, rewritten onto one reference epoch. +/// +/// `step-v1` section 8 compares committed behaviour across runs, excluding wall time, request +/// ids "and other explicitly operational metadata". A resumed run's epoch is neither: it is +/// behaviour metadata, and `scope.epoch`, the batch id and every task event id are derived +/// from it. Comparing the two runs therefore means rewriting exactly those three things and +/// nothing else, which is what this does -- and it **fails** on anything it does not +/// recognise instead of passing it through, so a field that silently stopped being rebased +/// would fail the comparison rather than weaken it. +#[derive(Clone, Debug, PartialEq, Eq)] +pub struct EpochRebase { + pub from: Id, + pub to: Id, + /// Every event identity the task issued under `from`, and the identity it has under `to`. + pub events: BTreeMap, +} + +impl EpochRebase { + /// Rewrites one behaviour record. An identity this rebase does not know is an error. + pub fn apply(&self, behaviour: &TraceBehaviour) -> Result { + if behaviour.scope.epoch != self.from { + return Err(format!( + "this behaviour was recorded in epoch {}, not {}", + behaviour.scope.epoch, self.from + )); + } + let mut out = behaviour.clone(); + out.scope = Scope::new(&behaviour.scope.session_id, &self.to, behaviour.scope.step) + .map_err(|e| e.0)?; + let prefix = format!("batch-{}-", self.from); + let suffix = behaviour + .batch_id + .strip_prefix(&prefix) + .ok_or_else(|| format!("the batch id {} is not derived from {}", behaviour.batch_id, self.from))?; + out.batch_id = parse_id(&format!("batch-{}-{suffix}", self.to))?; + let map = |ids: &[Id]| -> Result, String> { + ids.iter() + .map(|id| { + self.events + .get(id) + .cloned() + .ok_or_else(|| format!("no rebased identity for the event {id}")) + }) + .collect() + }; + out.outcome_ids = map(&behaviour.outcome_ids)?; + out.event_ids = map(&behaviour.event_ids)?; + Ok(out) + } +} + /// One session phase transition, recorded whether or not it ends a step. #[derive(Clone, Debug, PartialEq, Eq)] pub struct PhaseTransition { @@ -256,6 +312,28 @@ impl TraceLog { .collect() } + /// Every transition's behaviour, rebased onto one epoch and canonicalized. + /// + /// This is the comparison a resumed run is held to: the same strings as + /// [`TraceLog::behavior`], with the epoch metadata accounted for and nothing else changed. + /// A resumed run's log holds transitions from two epochs -- the ones before the checkpoint + /// and the ones after the restore -- so a transition already recorded in `rebase.to` is + /// kept as it stands and one recorded in `rebase.from` is rewritten. A transition in a + /// third epoch is an error; there is no pass-through case. + pub fn behavior_rebased(&self, rebase: &EpochRebase) -> Result, String> { + self.transitions + .iter() + .map(|t| { + let behaviour = if t.behaviour.scope.epoch == rebase.to { + t.behaviour.clone() + } else { + rebase.apply(&t.behaviour)? + }; + canonicalize(&behaviour.to_json()).map_err(|e| e.0) + }) + .collect() + } + /// The phase path, as `from -> to` strings. pub fn phase_path(&self) -> Vec { self.phases.iter().map(|p| format!("{} -> {}", p.from, p.to)).collect() diff --git a/services/flysim/crates/fly-session/src/worker.rs b/services/flysim/crates/fly-session/src/worker.rs index 327d66b..d717878 100644 --- a/services/flysim/crates/fly-session/src/worker.rs +++ b/services/flysim/crates/fly-session/src/worker.rs @@ -597,7 +597,17 @@ async fn execute( fn classify_default(method: &str) -> Option { match method { "Agent.Prepare" | "Agent.Commit" | "Environment.Advance" => Some(OpClass::StepMutation), - "Agent.Initialize" | "Environment.Initialize" => Some(OpClass::Lifecycle), + // `ipc-v1` section 5 retains lifecycle *and capture* replies until + // `Worker.Acknowledge`. The restore methods join them: their replies carry a + // once-only token and, for an environment, the restored observation's artifact, and a + // duplicate domain request must replay that reply rather than stage or activate a + // second time. They are not step mutations -- they carry no committed step of their + // own and are not keyed by one. + "Agent.Initialize" + | "Environment.Initialize" + | "State.Capture" + | "State.StageRestore" + | "State.ActivateRestore" => Some(OpClass::Lifecycle), "Worker.Hello" | "Worker.Status" | "Worker.Acknowledge" | "Worker.Shutdown" => { Some(OpClass::ReadOnly) } diff --git a/services/flysim/crates/fly-session/tests/processes.rs b/services/flysim/crates/fly-session/tests/processes.rs index d2dcd53..0eaa267 100644 --- a/services/flysim/crates/fly-session/tests/processes.rs +++ b/services/flysim/crates/fly-session/tests/processes.rs @@ -121,6 +121,8 @@ async fn a_slow_participant_is_resolved_rather_than_failed(mode: ExecutionMode) resolve: Duration::from_secs(20), resolve_attempts: 4096, boot: Duration::from_secs(30), + capture: Duration::from_secs(30), + durable: Duration::from_secs(60), }; within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap(); let reports = within("run", f.harness.coordinator.run(2)) @@ -191,6 +193,8 @@ async fn a_resolution_says_which_of_its_two_bounds_ended_it(mode: ExecutionMode) resolve: Duration::from_millis(300), resolve_attempts: 8192, boot: Duration::from_secs(30), + capture: Duration::from_secs(30), + durable: Duration::from_secs(60), }; within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap(); let started = Instant::now(); @@ -221,6 +225,8 @@ async fn a_resolution_says_which_of_its_two_bounds_ended_it(mode: ExecutionMode) resolve: Duration::from_secs(600), resolve_attempts: 3, boot: Duration::from_secs(30), + capture: Duration::from_secs(30), + durable: Duration::from_secs(60), }; within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap(); let failure = within("step", f.harness.coordinator.step()) diff --git a/services/flysim/crates/fly-session/tests/state.rs b/services/flysim/crates/fly-session/tests/state.rs new file mode 100644 index 0000000..b287c04 --- /dev/null +++ b/services/flysim/crates/fly-session/tests/state.rs @@ -0,0 +1,1002 @@ +//! STATE-01 acceptance: one coherent all-participant checkpoint, and the recovery that +//! installs it into a fresh epoch. +//! +//! Every test here is one of the slice's acceptance bullets or one of the failure-injection +//! rows the implementation guide's section 4 assigns to it. Each of them runs over both +//! transports and in all three execution modes: the recovery path crosses a process boundary +//! in exactly the places the media path does, so a row that holds in one mode has to hold in +//! all of them. +//! +//! The byte layout itself is proved against the contract crate in +//! `fly-session-types/tests/checkpoint_envelope.rs`; these prove the session's use of it. + +mod common; + +use std::collections::BTreeSet; +use std::path::Path; +use std::sync::Arc; +use std::time::Duration; + +use serde_json::{Value, json}; + +use common::{Fixture, count, fixture, fly_a, fly_b, within}; +use fly_session::agent::AgentFaults; +use fly_session::harness::{ExecutionMode, HarnessConfig, Via}; +use fly_session::phase::Phase; +use fly_session::state::{SaveOutcome, WriterFaults}; +use fly_session::types::*; +use fly_session_types::checkpoint; + +both_transports!( + an_uninterrupted_run_and_a_resumed_run_produce_matching_traces, + corrupting_any_single_participants_payload_fails_the_install_as_a_group, + a_lost_save_reply_does_not_advance_durable_metadata, + a_failure_during_activation_cannot_resume_half_a_world, + the_checkpoint_queue_under_stress_stays_bounded, + old_media_cannot_cross_a_recovery, + a_checkpoint_taken_by_another_parser_is_refused_by_name, + a_group_where_one_participant_will_not_stage_resumes_nothing, + a_restore_token_activates_only_once, + the_fence_lifts_only_through_a_coherent_restore, + a_queued_replaceable_capture_is_superseded_rather_than_duplicated, + a_capture_is_refused_anywhere_but_a_committed_boundary, +); + +all_modes!( + matching_traces_in_every_mode, + a_group_install_fails_as_a_group_in_every_mode, + a_lost_save_reply_holds_durable_metadata_in_every_mode, + half_an_activation_resumes_nothing_in_every_mode, + the_queue_stays_bounded_in_every_mode, + old_media_cannot_cross_in_every_mode, + another_parser_is_refused_in_every_mode, + a_refused_stage_resumes_nothing_in_every_mode, + a_token_activates_once_in_every_mode, + the_fence_lifts_only_by_restore_in_every_mode, +); + +const BEFORE: u64 = 2; +const AFTER: u64 = 2; + +fn ckpt(n: u32) -> Id { + id(&format!("ckpt-{n}")) +} + +/// A fixture in one transport and one execution mode. +/// +/// `Via` is the composition's; a dedicated thread or a separate process reaches the router +/// over a socket whatever it says, which the launcher decides and this does not second-guess. +async fn fx(via: Via, mode: ExecutionMode, config: HarnessConfig) -> Fixture { + fixture(via, HarnessConfig { mode, ..config }).await +} + +/// Runs to a committed boundary and takes one durable checkpoint there. +async fn run_and_checkpoint(f: &mut Fixture, steps: u64, checkpoint_id: &Id) { + within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap(); + within("run", f.harness.coordinator.run(steps)).await.unwrap(); + let outcome = within("checkpoint", f.harness.coordinator.checkpoint(checkpoint_id)) + .await + .unwrap(); + match outcome { + SaveOutcome::Committed { boundary, .. } => assert_eq!(boundary, steps), + other => panic!("the checkpoint was not committed: {other:?}"), + } + assert_eq!( + f.harness.coordinator.durable(), + Some((checkpoint_id.clone(), steps)), + "a committed save is the only thing that moves the durable mark" + ); +} + +/// Fails the epoch the way a participant death does, and checks the fence closed. +async fn fail_the_epoch(f: &mut Fixture) { + f.harness.kill(&fly_a()).await; + let failure = within("step", f.harness.coordinator.step()) + .await + .expect_err("a dead participant fails the epoch"); + assert!( + failure.participant.is_some(), + "a diagnosed failure names its participant: {failure}" + ); + assert_eq!(f.harness.coordinator.phase(), Phase::Failed); + assert!(f.harness.coordinator.is_fenced()); + assert_eq!( + f.harness.coordinator.live_view_handles(), + 0, + "the fence drops every handle the old epoch held" + ); +} + +/// The committed generation's bytes, as they are on disk. +fn generation_bytes(root: &Path, checkpoint_id: &Id) -> Vec { + std::fs::read(root.join(format!("{checkpoint_id}.flysess"))).expect("a committed generation") +} + +/// Writes a generation file and makes the store manifest describe it, as a repair tool or a +/// previous run would have left it. +fn write_generation(root: &Path, checkpoint_id: &Id, bytes: &[u8]) { + std::fs::write(root.join(format!("{checkpoint_id}.flysess")), bytes).expect("writable"); + let path = root.join("manifest.json"); + let mut manifest: Value = + serde_json::from_slice(&std::fs::read(&path).expect("a store manifest")).expect("json"); + let generations = manifest["generations"].as_array_mut().expect("an array"); + for generation in generations.iter_mut() { + if generation["checkpointId"] == json!(checkpoint_id.as_str()) { + generation["byteLength"] = json!(bytes.len().to_string()); + generation["envelopeDigest"] = json!(digest_of_bytes(bytes)); + let envelope = checkpoint::decode(bytes).expect("a well formed envelope"); + let compatibility = fly_session::state::Compatibility::from_json( + &envelope.manifest["compatibility"], + ) + .expect("a compatibility block"); + generation["compatibilityDigest"] = json!(compatibility.digest()); + } + } + std::fs::write(&path, serde_json::to_vec(&manifest).expect("json")).expect("writable"); +} + +/// Re-encodes one committed generation after `edit` has changed its manifest or its payloads. +fn rewrite_generation( + root: &Path, + checkpoint_id: &Id, + edit: impl FnOnce(&mut Value, &mut Vec<(String, Vec)>), +) { + let bytes = generation_bytes(root, checkpoint_id); + let envelope = checkpoint::decode(&bytes).expect("a committed generation decodes"); + let mut manifest = envelope.manifest.clone(); + let mut payloads = envelope.payloads.clone(); + edit(&mut manifest, &mut payloads); + // The manifest's payload table mirrors the envelope's, so it is rebuilt from the bytes + // that are actually being written rather than left to disagree. + manifest["payloads"] = Value::Array( + payloads + .iter() + .map(|(name, bytes)| { + json!({ + "name": name, + "byteLength": bytes.len().to_string(), + "digest": digest_of_bytes(bytes), + }) + }) + .collect(), + ); + let rewritten = checkpoint::encode(&manifest, &payloads).expect("a valid envelope"); + write_generation(root, checkpoint_id, &rewritten); +} + +/// Makes the store read its durable metadata again, after a test has edited it. +async fn reload_store(f: &mut Fixture) { + f.harness + .coordinator + .writer() + .expect("a store is attached") + .with_store(|store| store.reload()) + .await + .expect("the store manifest is still readable"); +} + +/// The names of every payload the checkpoint holds, participants first. +fn payload_names(root: &Path, checkpoint_id: &Id) -> Vec { + let bytes = generation_bytes(root, checkpoint_id); + checkpoint::decode(&bytes) + .expect("decodes") + .payloads + .into_iter() + .map(|(name, _)| name) + .collect() +} + +/// Every artifact identity this session's committed boundary is holding. +fn live_artifact_ids(f: &Fixture) -> BTreeSet { + f.harness + .coordinator + .media_handles() + .into_iter() + .map(|(_, reference)| reference.artifact_id) + .collect() +} + +/// Asserts that no participant is running the proposed epoch: nothing was installed. +async fn nothing_is_installed(f: &mut Fixture, epoch: &Id) { + let ids: Vec = std::iter::once(f.harness.environment_id()) + .chain(f.harness.config.agents.iter().map(|a| a.agent_id.clone())) + .collect(); + for worker_id in ids { + let (_, launcher) = f.harness.parts(); + let Ok(status) = launcher.health_check(&worker_id).await else { + // A participant that is gone is certainly not running the new epoch. + continue; + }; + if let Some(scope) = status.current_scope { + assert_ne!( + scope.epoch, *epoch, + "{worker_id} is running the epoch the abandoned install proposed" + ); + } + } + assert_eq!(f.harness.coordinator.phase(), Phase::Failed); + assert!(f.harness.coordinator.is_fenced()); + let refused = within("step", f.harness.coordinator.step()) + .await + .expect_err("a fenced session takes no step"); + assert_eq!(refused.error.code, ErrorCode::InvalidPhase); +} + +// =============================================================================================== +// Acceptance: an uninterrupted run and a resumed run produce matching traces + +async fn an_uninterrupted_run_and_a_resumed_run_produce_matching_traces(via: Via) { + matching_traces(via, ExecutionMode::InProcess).await; +} + +async fn matching_traces_in_every_mode(mode: ExecutionMode) { + matching_traces(Via::Unix, mode).await; +} + +/// The reference run and the resumed run commit the same behaviour, once the epoch metadata +/// the restore necessarily changed is accounted for. +/// +/// `step-v1` section 8's split is the whole of the comparison: the behaviour is what must +/// match, the request ids and wall time are excluded because they are operational, and the +/// epoch is neither -- it is behaviour metadata, so it is rewritten explicitly and everything +/// else is compared byte for byte. +async fn matching_traces(via: Via, mode: ExecutionMode) { + let mut reference = fx(via, mode, HarnessConfig::default()).await; + within("bootstrap", reference.harness.coordinator.bootstrap()).await.unwrap(); + within("run", reference.harness.coordinator.run(BEFORE + AFTER)).await.unwrap(); + let expected = reference.harness.coordinator.trace.behavior(); + assert_eq!(expected.len() as u64, BEFORE + AFTER); + reference.shutdown().await; + + let mut f = fx(via, mode, HarnessConfig::default()).await; + let checkpoint_id = ckpt(1); + run_and_checkpoint(&mut f, BEFORE, &checkpoint_id).await; + fail_the_epoch(&mut f).await; + within("replace", f.harness.replace_all_participants()).await.unwrap(); + let report = within( + "restore", + f.harness.coordinator.restore(Some(&checkpoint_id), &id("e2")), + ) + .await + .unwrap(); + assert_eq!(report.boundary, BEFORE); + assert_eq!(report.epoch, id("e2")); + assert_eq!(report.activated.len(), report.staged.len()); + assert!(!f.harness.coordinator.is_fenced(), "a coherent restore lifts the fence"); + assert_eq!(f.harness.coordinator.phase(), Phase::Paused(BEFORE)); + + f.harness.coordinator.resume().unwrap(); + within("resume", f.harness.coordinator.run(AFTER)).await.unwrap(); + assert_eq!(f.harness.coordinator.phase(), Phase::Ready(BEFORE + AFTER)); + + // Without accounting for the epoch the two traces disagree, which is what makes the + // rebase a statement rather than a formality. + let raw = f.harness.coordinator.trace.behavior(); + assert_eq!(raw.len() as u64, BEFORE + AFTER); + assert_ne!(raw, expected, "the resumed run runs in a different epoch"); + + let rebase = f.harness.coordinator.rebase(&id("e1")).unwrap(); + let resumed = f.harness.coordinator.trace.behavior_rebased(&rebase).unwrap(); + assert_eq!( + resumed, expected, + "a resumed run commits the behaviour the uninterrupted run committed" + ); + f.shutdown().await; +} + +// =============================================================================================== +// Acceptance: corrupt any participant and installation fails as a group + +async fn corrupting_any_single_participants_payload_fails_the_install_as_a_group(via: Via) { + group_install_is_all_or_nothing(via, ExecutionMode::InProcess).await; +} + +async fn a_group_install_fails_as_a_group_in_every_mode(mode: ExecutionMode) { + group_install_is_all_or_nothing(Via::Unix, mode).await; +} + +/// One corrupted payload -- any participant's, and the coordinator's own ledger too -- fails +/// the whole install, and the group afterwards is exactly as it was. +async fn group_install_is_all_or_nothing(via: Via, mode: ExecutionMode) { + let mut f = fx(via, mode, HarnessConfig::default()).await; + let checkpoint_id = ckpt(1); + run_and_checkpoint(&mut f, BEFORE, &checkpoint_id).await; + let root = f.harness.checkpoint_root().to_path_buf(); + let good = generation_bytes(&root, &checkpoint_id); + let names = payload_names(&root, &checkpoint_id); + // Every participant's payload, plus one of the coordinator's own ledgers. + let mut corrupt: Vec = names + .iter() + .filter(|name| name.starts_with("agent-") || *name == "world") + .cloned() + .collect(); + corrupt.push("task-ledger".to_owned()); + assert!(corrupt.len() >= 3, "the composition has several payloads: {names:?}"); + fail_the_epoch(&mut f).await; + + let mut epoch = 1u32; + for target in &corrupt { + epoch += 1; + let proposed = id(&format!("e{epoch}")); + // Each round starts from the intact bytes, so exactly one payload is corrupt. + write_generation(&root, &checkpoint_id, &good); + // A well formed envelope whose digests all agree, so what fails is the participant + // reading its own bytes and not the envelope reader in front of it. + let name = target.clone(); + rewrite_generation(&root, &checkpoint_id, |_manifest, payloads| { + for (payload, bytes) in payloads.iter_mut() { + if *payload == name { + let last = bytes.len() - 1; + bytes[last] ^= 0xff; + } + } + }); + reload_store(&mut f).await; + within("replace", f.harness.replace_all_participants()).await.unwrap(); + let failure = within( + "restore", + f.harness.coordinator.restore(Some(&checkpoint_id), &proposed), + ) + .await + .expect_err(&format!("a corrupt {target} must fail the install")); + assert!( + matches!( + failure.error.code, + ErrorCode::IncompatibleState | ErrorCode::InvalidArgument + ), + "a corrupt {target} is an explicit refusal, not {failure}" + ); + nothing_is_installed(&mut f, &proposed).await; + let tainted = f.harness.coordinator.tainted(); + if target == "world" { + assert!( + tainted.is_empty(), + "the first participant asked refused, so nothing staged: {tainted:?}" + ); + } else { + assert!( + !tainted.is_empty(), + "a participant that staged into an abandoned install must be replaced" + ); + } + } + + // The same group, the same store, the intact bytes: the failures above installed nothing + // that stops this from working. + epoch += 1; + write_generation(&root, &checkpoint_id, &good); + reload_store(&mut f).await; + within("replace", f.harness.replace_all_participants()).await.unwrap(); + let report = within( + "restore", + f.harness + .coordinator + .restore(Some(&checkpoint_id), &id(&format!("e{epoch}"))), + ) + .await + .unwrap(); + assert_eq!(report.boundary, BEFORE); + f.harness.coordinator.resume().unwrap(); + within("resume", f.harness.coordinator.run(1)).await.unwrap(); + f.shutdown().await; +} + +// =============================================================================================== +// Acceptance: a lost save reply does not advance durable metadata + +async fn a_lost_save_reply_does_not_advance_durable_metadata(via: Via) { + a_lost_save_reply(via, ExecutionMode::InProcess).await; +} + +async fn a_lost_save_reply_holds_durable_metadata_in_every_mode(mode: ExecutionMode) { + a_lost_save_reply(Via::Unix, mode).await; +} + +/// Three ways a save can end without a saved acknowledgment. None moves the mark, and each +/// is reported as itself rather than as the others. +async fn a_lost_save_reply(via: Via, mode: ExecutionMode) { + let lost = ckpt(1); + let config = HarnessConfig { + writer_faults: WriterFaults { + drop_reply_for: Some(lost.clone()), + ..WriterFaults::default() + }, + ..HarnessConfig::default() + }; + let mut f = fx(via, mode, config).await; + within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap(); + within("run", f.harness.coordinator.run(1)).await.unwrap(); + + let ticket = within("capture", f.harness.coordinator.capture(&lost, false)) + .await + .unwrap(); + let outcome = within("durable", f.harness.coordinator.await_durable(ticket)) + .await + .unwrap(); + assert_eq!(outcome, SaveOutcome::ReplyLost { checkpoint_id: lost.clone() }); + assert_eq!( + f.harness.coordinator.durable(), + None, + "a lost save reply never moves the durable mark" + ); + assert_eq!(count(&f.harness.coordinator.audit, &format!("durable:{lost}@1")), 0); + + // The resolution asks the durable metadata about the *same* operation. It never saves + // again, and here the bytes did reach the store manifest. + let resolved = within("resolve", f.harness.coordinator.resolve_durable(&lost)) + .await + .unwrap(); + assert_eq!(resolved, Some(1)); + assert_eq!(f.harness.coordinator.durable(), Some((lost.clone(), 1))); + let commits = f + .harness + .coordinator + .writer() + .expect("a store") + .with_store(|store| store.commits()) + .await; + assert_eq!(commits, 1, "the resolution queried the store; it did not save again"); + f.shutdown().await; + + // The other half: the generation file is renamed and the store manifest is never + // committed. That generation is unreferenced, so it is not a candidate and the mark is + // still where it was. + let orphan = ckpt(2); + let mut config = HarnessConfig::default(); + config.store_faults.stop_before_manifest_commit = true; + let mut g = fx(via, mode, config).await; + within("bootstrap", g.harness.coordinator.bootstrap()).await.unwrap(); + within("run", g.harness.coordinator.run(1)).await.unwrap(); + let outcome = within("checkpoint", g.harness.coordinator.checkpoint(&orphan)) + .await + .unwrap(); + match outcome { + SaveOutcome::Failed { checkpoint_id, .. } => assert_eq!(checkpoint_id, orphan), + other => panic!("an uncommitted manifest is not a saved checkpoint: {other:?}"), + } + assert_eq!(g.harness.coordinator.durable(), None); + let root = g.harness.checkpoint_root().to_path_buf(); + assert!( + root.join(format!("{orphan}.flysess")).is_file(), + "the generation file was written and renamed" + ); + let resolved = within("resolve", g.harness.coordinator.resolve_durable(&orphan)) + .await + .unwrap(); + assert_eq!(resolved, None, "an unreferenced generation is never a restore candidate"); + assert_eq!(g.harness.coordinator.durable(), None); + g.shutdown().await; + + // The third: the caller's own budget runs out while the save is still going. That is a + // different fact from a lost reply -- this save has not stopped -- and it is reported as + // itself, because diagnosing one as the other is the implicit reading these rules refuse. + let slow = ckpt(3); + let gate = Arc::new(tokio::sync::Semaphore::new(0)); + let mut config = HarnessConfig::default(); + config.writer_faults.gate = Some(gate.clone()); + let mut h = fx(via, mode, config).await; + within("bootstrap", h.harness.coordinator.bootstrap()).await.unwrap(); + within("run", h.harness.coordinator.run(1)).await.unwrap(); + // The durable budget is the caller's own and is not the call budget: this shortens the + // wait for an acknowledgment without shortening a single call to a participant. + h.harness.coordinator.deadlines.durable = Duration::from_millis(100); + let ticket = within("capture", h.harness.coordinator.capture(&slow, false)) + .await + .unwrap(); + let outcome = within("durable", h.harness.coordinator.await_durable(ticket)) + .await + .unwrap(); + assert_eq!( + outcome, + SaveOutcome::DeadlineExpired { checkpoint_id: slow.clone() }, + "an expired caller budget is not a lost reply" + ); + assert_ne!(outcome, SaveOutcome::ReplyLost { checkpoint_id: slow.clone() }); + assert_eq!(h.harness.coordinator.durable(), None); + assert_eq!( + count( + &h.harness.coordinator.audit, + &format!("not-durable:failed:{slow}@1") + ), + 1, + "the session records that this capture is not durable: {:?}", + h.harness.coordinator.audit + ); + + // It really was still going: once the writer is let past its gate the same operation + // commits, and resolving it is what moves the mark. + gate.add_permits(16); + let deadline = std::time::Instant::now() + Duration::from_secs(10); + let resolved = loop { + let found = within("resolve", h.harness.coordinator.resolve_durable(&slow)) + .await + .unwrap(); + if found.is_some() || std::time::Instant::now() >= deadline { + break found; + } + tokio::time::sleep(Duration::from_millis(10)).await; + }; + assert_eq!(resolved, Some(1), "the save the caller stopped waiting for still committed"); + assert_eq!(h.harness.coordinator.durable(), Some((slow, 1))); + h.shutdown().await; +} + +// =============================================================================================== +// Acceptance: a failure during activation cannot resume half a world + +async fn a_failure_during_activation_cannot_resume_half_a_world(via: Via) { + half_an_activation(via, ExecutionMode::InProcess).await; +} + +async fn half_an_activation_resumes_nothing_in_every_mode(mode: ExecutionMode) { + half_an_activation(Via::Unix, mode).await; +} + +/// The world and the first agent activate; the second refuses. Nothing plays, the fence +/// stays closed, and the group cannot be resumed until every participant is replaced. +async fn half_an_activation(via: Via, mode: ExecutionMode) { + let mut f = fx(via, mode, HarnessConfig::default()).await; + let checkpoint_id = ckpt(1); + run_and_checkpoint(&mut f, BEFORE, &checkpoint_id).await; + fail_the_epoch(&mut f).await; + + // The replacement for fly-b refuses to activate after it has staged. + f.harness.set_agent_faults( + &fly_b(), + AgentFaults { fail_activate_restore: true, ..AgentFaults::default() }, + ); + within("replace", f.harness.replace_all_participants()).await.unwrap(); + let failure = within( + "restore", + f.harness.coordinator.restore(Some(&checkpoint_id), &id("e2")), + ) + .await + .expect_err("a refused activation fails the install"); + assert_eq!(failure.participant.as_deref(), Some(fly_b().as_str())); + assert_eq!(f.harness.coordinator.phase(), Phase::Failed); + assert!(f.harness.coordinator.is_fenced(), "the group stays fenced"); + let refused = within("step", f.harness.coordinator.step()) + .await + .expect_err("half a world never runs"); + assert_eq!(refused.error.code, ErrorCode::InvalidPhase); + + // The participants that got as far as installing hold state no group resumed. Another + // restore over them is refused by name rather than attempted. + let tainted = f.harness.coordinator.tainted(); + assert!(tainted.contains(&f.harness.environment_id()), "{tainted:?}"); + assert!(tainted.contains(&fly_a()), "{tainted:?}"); + let refused = within( + "restore", + f.harness.coordinator.restore(Some(&checkpoint_id), &id("e3")), + ) + .await + .expect_err("the half-installed group is not restored over"); + assert_eq!(refused.error.code, ErrorCode::InvalidPhase); + assert!(refused.error.message.contains(&fly_a()), "{refused}"); + + // A fresh group, without the injected refusal, resumes the same boundary. + f.harness.set_agent_faults(&fly_b(), AgentFaults::default()); + within("replace", f.harness.replace_all_participants()).await.unwrap(); + assert!(f.harness.coordinator.tainted().is_empty()); + let report = within( + "restore", + f.harness.coordinator.restore(Some(&checkpoint_id), &id("e4")), + ) + .await + .unwrap(); + assert_eq!(report.boundary, BEFORE); + assert!(!f.harness.coordinator.is_fenced()); + f.harness.coordinator.resume().unwrap(); + within("resume", f.harness.coordinator.run(1)).await.unwrap(); + f.shutdown().await; +} + +// =============================================================================================== +// Acceptance: the checkpoint queue under stress stays bounded + +async fn the_checkpoint_queue_under_stress_stays_bounded(via: Via) { + queue_stays_bounded(via, ExecutionMode::InProcess).await; +} + +async fn the_queue_stays_bounded_in_every_mode(mode: ExecutionMode) { + queue_stays_bounded(Via::Unix, mode).await; +} + +/// With the writer stalled, the queue fills to its configured bound and the next capture is +/// refused *before* a single participant is asked for one. +async fn queue_stays_bounded(via: Via, mode: ExecutionMode) { + let gate = Arc::new(tokio::sync::Semaphore::new(0)); + let mut config = HarnessConfig::default(); + config.writer_faults.gate = Some(gate.clone()); + config.writer.queue_capacity = 2; + let mut f = fx(via, mode, config).await; + within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap(); + within("run", f.harness.coordinator.run(1)).await.unwrap(); + + let first = within("capture", f.harness.coordinator.capture(&ckpt(1), false)) + .await + .unwrap(); + let second = within("capture", f.harness.coordinator.capture(&ckpt(2), false)) + .await + .unwrap(); + assert_eq!(f.harness.coordinator.writer().expect("a store").outstanding(), 2); + + let environment = f.harness.environment_id(); + let before = within("status", f.harness.progress_of(&environment)).await.unwrap(); + let refused = within("capture", f.harness.coordinator.capture(&ckpt(3), false)) + .await + .expect_err("a full queue refuses"); + assert_eq!(refused.error.code, ErrorCode::Busy); + assert_eq!(refused.error.mutation, MutationCertainty::None); + let after = within("status", f.harness.progress_of(&environment)).await.unwrap(); + assert_eq!( + before, after, + "a refused capture asks no participant for one: rejection happens before capture" + ); + assert_eq!(count(&f.harness.coordinator.audit, "capture-refused:ckpt-3"), 1); + + // A refused capture is not a failed session: the world keeps stepping. + assert!(!f.harness.coordinator.is_fenced()); + within("run", f.harness.coordinator.run(1)).await.unwrap(); + assert_eq!(f.harness.coordinator.phase(), Phase::Ready(2)); + + let stats = f.harness.coordinator.writer().expect("a store").stats(); + assert_eq!(stats.rejected, 1); + assert!(stats.peak_queue <= 2, "the queue never exceeded its bound: {stats:?}"); + assert!(stats.peak_bytes <= 64 * 1024 * 1024); + + gate.add_permits(16); + for (ticket, expected) in [(first, ckpt(1)), (second, ckpt(2))] { + let outcome = within("durable", f.harness.coordinator.await_durable(ticket)) + .await + .unwrap(); + match outcome { + SaveOutcome::Committed { checkpoint_id, .. } => assert_eq!(checkpoint_id, expected), + other => panic!("an unstalled writer commits: {other:?}"), + } + } + let stats = f.harness.coordinator.writer().expect("a store").stats(); + assert_eq!(stats.committed, 2); + assert_eq!(f.harness.coordinator.writer().expect("a store").outstanding(), 0); + f.shutdown().await; +} + +/// A queued replaceable capture is replaced by the next one, releasing its holds, rather than +/// both being written. +async fn a_queued_replaceable_capture_is_superseded_rather_than_duplicated(via: Via) { + let gate = Arc::new(tokio::sync::Semaphore::new(0)); + let mut config = HarnessConfig::default(); + config.writer_faults.gate = Some(gate.clone()); + config.writer.queue_capacity = 3; + let mut f = fx(via, ExecutionMode::InProcess, config).await; + within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap(); + within("run", f.harness.coordinator.run(1)).await.unwrap(); + + // The first job is taken by the writer and stalls at the gate; the next two are queued. + let durable = within("capture", f.harness.coordinator.capture(&ckpt(1), false)) + .await + .unwrap(); + let hot = within("capture", f.harness.coordinator.capture(&ckpt(2), true)) + .await + .unwrap(); + let newer = within("capture", f.harness.coordinator.capture(&ckpt(3), true)) + .await + .unwrap(); + + let superseded = within("durable", f.harness.coordinator.await_durable(hot)) + .await + .unwrap(); + assert_eq!( + superseded, + SaveOutcome::Superseded { checkpoint_id: ckpt(2), by: ckpt(3) }, + "only a queued replaceable capture is coalesced" + ); + gate.add_permits(16); + for ticket in [durable, newer] { + let outcome = within("durable", f.harness.coordinator.await_durable(ticket)) + .await + .unwrap(); + assert!(matches!(outcome, SaveOutcome::Committed { .. }), "{outcome:?}"); + } + let stats = f.harness.coordinator.writer().expect("a store").stats(); + assert_eq!(stats.superseded, 1); + assert_eq!(stats.committed, 2, "the superseded capture was never written"); + f.shutdown().await; +} + +// =============================================================================================== +// Acceptance: old media and old parser data cannot cross a recovery + +async fn old_media_cannot_cross_a_recovery(via: Via) { + old_media_cannot_cross(via, ExecutionMode::InProcess).await; +} + +async fn old_media_cannot_cross_in_every_mode(mode: ExecutionMode) { + old_media_cannot_cross(Via::Unix, mode).await; +} + +/// Nothing the fence dropped comes back: the payloads are re-imported as fresh artifacts, the +/// restored view is a new object, and the new epoch's first audio chunk resumes the preserved +/// sample position and marks the discontinuity. +async fn old_media_cannot_cross(via: Via, mode: ExecutionMode) { + let mut f = fx(via, mode, HarnessConfig::default()).await; + let checkpoint_id = ckpt(1); + run_and_checkpoint(&mut f, BEFORE, &checkpoint_id).await; + let old_ids = live_artifact_ids(&f); + assert!(!old_ids.is_empty(), "the committed boundary holds media handles"); + let positions = f.harness.coordinator.audio_positions(); + assert!(positions.values().any(|sample| *sample > 0), "audio has played"); + + fail_the_epoch(&mut f).await; + within("replace", f.harness.replace_all_participants()).await.unwrap(); + let report = within( + "restore", + f.harness.coordinator.restore(Some(&checkpoint_id), &id("e2")), + ) + .await + .unwrap(); + for imported in &report.imported { + assert!( + !old_ids.contains(imported), + "a checkpoint payload was imported as an artifact the old epoch already had" + ); + } + let new_ids = live_artifact_ids(&f); + assert!(!new_ids.is_empty()); + assert!( + new_ids.is_disjoint(&old_ids), + "the restored boundary's media are fresh artifacts, not the old epoch's" + ); + assert_eq!( + f.harness.coordinator.audio_positions(), + positions, + "crash restore preserves the sample position" + ); + + f.harness.coordinator.resume().unwrap(); + within("resume", f.harness.coordinator.run(1)).await.unwrap(); + let chunk = f + .harness + .coordinator + .observation() + .expect("a restored world") + .audio + .first() + .cloned() + .expect("one chunk per transition"); + assert!( + chunk.discontinuity, + "the first chunk of a fresh epoch marks the discontinuity the recovery established" + ); + assert_eq!( + chunk.first_sample, + positions["arena"], + "and it starts where the checkpointed stream stopped" + ); + f.shutdown().await; +} + +async fn a_checkpoint_taken_by_another_parser_is_refused_by_name(via: Via) { + another_parser_is_refused(via, ExecutionMode::InProcess).await; +} + +async fn another_parser_is_refused_in_every_mode(mode: ExecutionMode) { + another_parser_is_refused(Via::Unix, mode).await; +} + +/// A checkpoint whose recorded parser identity is not this composition's is refused before a +/// single participant is asked to stage, and the refusal names the identity that differs. +async fn another_parser_is_refused(via: Via, mode: ExecutionMode) { + let mut f = fx(via, mode, HarnessConfig::default()).await; + let checkpoint_id = ckpt(1); + run_and_checkpoint(&mut f, 1, &checkpoint_id).await; + let root = f.harness.checkpoint_root().to_path_buf(); + fail_the_epoch(&mut f).await; + rewrite_generation(&root, &checkpoint_id, |manifest, _payloads| { + manifest["compatibility"]["parserDigest"] = + json!(digest_of_bytes(b"some other inspection schema")); + }); + reload_store(&mut f).await; + within("replace", f.harness.replace_all_participants()).await.unwrap(); + let failure = within( + "restore", + f.harness.coordinator.restore(Some(&checkpoint_id), &id("e2")), + ) + .await + .expect_err("another parser's state is not this composition's"); + assert_eq!(failure.error.code, ErrorCode::IncompatibleState); + assert!( + failure.error.message.contains("parser"), + "the refusal names the identity that differs: {failure}" + ); + assert!( + f.harness.coordinator.tainted().is_empty(), + "nothing was asked to stage, so nothing has to be replaced" + ); + nothing_is_installed(&mut f, &id("e2")).await; + f.shutdown().await; +} + +// =============================================================================================== +// Failure-injection rows + +async fn a_group_where_one_participant_will_not_stage_resumes_nothing(via: Via) { + a_refused_stage(via, ExecutionMode::InProcess).await; +} + +async fn a_refused_stage_resumes_nothing_in_every_mode(mode: ExecutionMode) { + a_refused_stage(Via::Unix, mode).await; +} + +/// The checklist row: the install validates the participants ahead of it and the last one +/// fails. Nothing is resumed. +async fn a_refused_stage(via: Via, mode: ExecutionMode) { + let mut f = fx(via, mode, HarnessConfig::default()).await; + let checkpoint_id = ckpt(1); + run_and_checkpoint(&mut f, 1, &checkpoint_id).await; + fail_the_epoch(&mut f).await; + f.harness.set_agent_faults( + &fly_b(), + AgentFaults { fail_stage_restore: true, ..AgentFaults::default() }, + ); + within("replace", f.harness.replace_all_participants()).await.unwrap(); + let failure = within( + "restore", + f.harness.coordinator.restore(Some(&checkpoint_id), &id("e2")), + ) + .await + .expect_err("a participant that will not validate stops the install"); + assert_eq!(failure.error.code, ErrorCode::IncompatibleState); + assert_eq!(failure.participant.as_deref(), Some(fly_b().as_str())); + assert_eq!( + count(&f.harness.coordinator.audit, &format!("activated:{}", fly_a())), + 0, + "no participant activates when one of them will not stage" + ); + nothing_is_installed(&mut f, &id("e2")).await; + f.shutdown().await; +} + +async fn a_restore_token_activates_only_once(via: Via) { + a_token_activates_once(via, ExecutionMode::InProcess).await; +} + +async fn a_token_activates_once_in_every_mode(mode: ExecutionMode) { + a_token_activates_once(Via::Unix, mode).await; +} + +/// A token is bound to its checkpoint, scope and payload and activates once. A fresh request +/// naming it again is a conflict, and the old epoch's scope is stale on the new participants. +async fn a_token_activates_once(via: Via, mode: ExecutionMode) { + let mut f = fx(via, mode, HarnessConfig::default()).await; + let checkpoint_id = ckpt(1); + run_and_checkpoint(&mut f, 1, &checkpoint_id).await; + fail_the_epoch(&mut f).await; + within("replace", f.harness.replace_all_participants()).await.unwrap(); + let report = within( + "restore", + f.harness.coordinator.restore(Some(&checkpoint_id), &id("e2")), + ) + .await + .unwrap(); + let (who, token) = report.tokens.first().cloned().expect("a staged token"); + let worker = if who == f.harness.environment_id() { + f.harness.coordinator.environment_ref().clone() + } else { + f.harness.coordinator.agent_ref(&who).expect("a participant").clone() + }; + let scope = scope_at("demo", "e2", 1); + let refused = within( + "activate", + f.harness.coordinator.probe_raw( + &worker, + "State.ActivateRestore", + Some(scope), + json!({"restoreToken": token.as_str()}), + ), + ) + .await + .expect_err("a token activates once"); + assert_eq!(refused.code, ErrorCode::Conflict); + + // And the epoch the restore left behind is stale on every replacement. + let stale = within( + "prepare", + f.harness.coordinator.probe_raw( + &worker, + if who == f.harness.environment_id() { "Environment.Advance" } else { "Agent.Prepare" }, + Some(scope_at("demo", "e1", 1)), + json!({}), + ), + ) + .await + .expect_err("the old epoch is gone"); + assert!( + matches!(stale.code, ErrorCode::StaleEpoch | ErrorCode::InvalidArgument), + "an old-epoch request is refused: {stale}" + ); + f.shutdown().await; +} + +async fn the_fence_lifts_only_through_a_coherent_restore(via: Via) { + the_fence_lifts_only_by_restore(via, ExecutionMode::InProcess).await; +} + +async fn the_fence_lifts_only_by_restore_in_every_mode(mode: ExecutionMode) { + the_fence_lifts_only_by_restore(Via::Unix, mode).await; +} + +/// Nothing but a coherent restore moves a fenced session, and the restore lands on +/// `Paused(k)` rather than straight back into play. +async fn the_fence_lifts_only_by_restore(via: Via, mode: ExecutionMode) { + let mut f = fx(via, mode, HarnessConfig::default()).await; + let checkpoint_id = ckpt(1); + run_and_checkpoint(&mut f, 1, &checkpoint_id).await; + fail_the_epoch(&mut f).await; + + for refused in [ + within("step", f.harness.coordinator.step()).await.err(), + within("bootstrap", f.harness.coordinator.bootstrap()).await.err(), + within("capture", f.harness.coordinator.capture(&ckpt(9), false)) + .await + .err(), + ] { + let refused = refused.expect("a fenced session refuses"); + assert_eq!(refused.error.code, ErrorCode::InvalidPhase); + } + assert!(f.harness.coordinator.resume().is_err(), "a fenced session does not resume"); + assert!(f.harness.coordinator.is_fenced()); + + // A restore into the epoch that failed is refused: recovery establishes a fresh one. + within("replace", f.harness.replace_all_participants()).await.unwrap(); + let same = within( + "restore", + f.harness.coordinator.restore(Some(&checkpoint_id), &id("e1")), + ) + .await + .expect_err("a restore installs a fresh epoch"); + assert_eq!(same.error.code, ErrorCode::StaleEpoch); + + let report = within( + "restore", + f.harness.coordinator.restore(Some(&checkpoint_id), &id("e2")), + ) + .await + .unwrap(); + assert_eq!(report.boundary, 1); + assert!(!f.harness.coordinator.is_fenced()); + assert_eq!(f.harness.coordinator.phase(), Phase::Paused(1)); + f.harness.coordinator.resume().unwrap(); + within("resume", f.harness.coordinator.run(1)).await.unwrap(); + assert_eq!(f.harness.coordinator.phase(), Phase::Ready(2)); + f.shutdown().await; +} + +/// A capture belongs to a committed boundary, and the phase machine is what says so. +async fn a_capture_is_refused_anywhere_but_a_committed_boundary(via: Via) { + let mut f = fx(via, ExecutionMode::InProcess, HarnessConfig::default()).await; + // Before bootstrap the session is Starting, which is not a boundary at all. + let refused = within("capture", f.harness.coordinator.capture(&ckpt(1), false)) + .await + .expect_err("Starting is not a committed boundary"); + assert_eq!(refused.error.code, ErrorCode::InvalidPhase); + f.shutdown().await; + + // From a pause, which is a committed boundary, it works, and it returns there. + let mut f = fx(via, ExecutionMode::InProcess, HarnessConfig::default()).await; + within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap(); + within("run", f.harness.coordinator.run(1)).await.unwrap(); + f.harness.coordinator.request_pause(); + within("run", f.harness.coordinator.run(1)).await.unwrap(); + assert_eq!(f.harness.coordinator.phase(), Phase::Paused(2)); + let outcome = within("checkpoint", f.harness.coordinator.checkpoint(&ckpt(1))) + .await + .unwrap(); + assert!(matches!(outcome, SaveOutcome::Committed { boundary: 2, .. }), "{outcome:?}"); + assert_eq!( + f.harness.coordinator.phase(), + Phase::Paused(2), + "a capture returns to the boundary it came from" + ); + f.shutdown().await; +}