session: coherent all-participant checkpoint and recovery
STATE-01 over the FLYSESS1 envelope CONTRACT-01 specified. fly-session gains a `state` module: the durable store with its generations, its rotation and the commit order of checkpoint-envelope-v1 section 5, where the store manifest rename is the durable commit point; a compatibility block whose comparison names the identity that differs rather than one opaque digest; and a bounded writer that owns its payload handles until the bytes are committed or the job fails. The writer's queue slot is taken before the first State.Capture, so a saturated writer refuses a capture rather than queueing it without bound, and the refusal is a BUSY the stepping session survives. Capture and durability are two events: a capture completes when an immutable capture exists, and only the store manifest rename moves the durable mark. A lost save reply is an outcome, and the resolution asks the store about the same checkpoint instead of saving again. Both worker roles implement State.Capture, State.StageRestore and State.ActivateRestore, with once-only restore tokens bound to checkpoint, scope, payload and incarnation. A restore selects a complete compatible generation, imports every payload as a fresh artifact, stages the group, validates the coordinator's own ledgers, and only then activates; a failure anywhere leaves the fence closed and records every participant that staged as one that must be replaced. The fence lifts at Failed -> Restoring(k) -> Paused(k) and nowhere else. The task and the action executor gain the capture/validate_restore/install_restore interfaces workers-v1 section 4 lists, and the ledger can re-derive the event identities it issued under another epoch, which is what lets a resumed run's behaviour trace be compared with an uninterrupted one. media: check_required_audio now takes the observation's provenance instead of exempting boundary 0. A chunk is the audio of an interval, and the observation ActivateRestore installs covers none. checkpoint-envelope-v1 section 3 gains a dated amendment adding `environment` to the manifest, the holder of the world's own payload, which the table named for every other participant; `helperState`, which that table already listed, joins the required-field set in Rust and TypeScript. The fixture was regenerated by the existing example; the schema set and contractDigest are unchanged. state-media-v1 section 5 gains a dated amendment for three readings this slice enforces: the State RPCs' compatibilityDigest is the participant's, not the manifest's composition-level block; a restored observation carries no audio chunk; and a participant that staged into an abandoned install must be replaced.
This commit is contained in:
parent
56db91cd9b
commit
6655a1b1c6
22 changed files with 5837 additions and 48 deletions
|
|
@ -105,6 +105,19 @@ names; a manifest missing any of them is not a complete checkpoint.
|
|||
| `helperState` | External-helper state required for exact resume, as payload names |
|
||||
| `payloads` | `[{name, byteLength, digest}]`, mirroring the payload table |
|
||||
|
||||
**Amendment, 2026-09-22 (STATE-01).** The table above names a holder for every payload except
|
||||
the environment's own, although section 6's fixture has one (`world`) and a group install has
|
||||
to map it by name like any other participant's. The manifest therefore also records:
|
||||
|
||||
| Field | Contents |
|
||||
| --- | --- |
|
||||
| `environment` | `{workerId, payload}`: which worker the world belonged to and the payload name holding its state |
|
||||
|
||||
The reference implementations' required-field set was also missing `helperState`, which this
|
||||
section has listed from the start. Both are now in `REQUIRED_MANIFEST_FIELDS` in Rust and in
|
||||
TypeScript, and the fixture was regenerated by the existing example. The schema set is
|
||||
untouched, so `contractDigest` is unchanged.
|
||||
|
||||
`payloads` is redundant with the table on purpose: the table is what a reader needs to map
|
||||
bytes, and the manifest is what a store lists, compares and reports without opening the
|
||||
payload area. A reader checks that the two agree.
|
||||
|
|
|
|||
|
|
@ -190,6 +190,28 @@ restored time. It cannot advance gameplay to manufacture it. Capture/reconstruct
|
|||
covers render/inspection state and any pending sensor pipeline. Agent state agrees with it;
|
||||
do not replay reward or recalibrate merely to fill missing cached data.
|
||||
|
||||
**Amendment, 2026-09-22 (STATE-01).** Three readings of this section, made explicit because
|
||||
they are now enforced:
|
||||
|
||||
- `compatibilityDigest` on `CaptureResult` and `StageRestoreParams` is the **participant's**
|
||||
capture compatibility digest of [worker interfaces](workers-v1.md) section 2 -- profile,
|
||||
resolved seed, numerical model version and effective instance configuration for an agent;
|
||||
backend, content, patch, controller and parser identity for an environment. It is not the
|
||||
manifest's `compatibility` block of section 4, which is the composition's and which the
|
||||
coordinator compares before anything is asked to stage. Both exist because they answer
|
||||
different questions, and a restore that passed the second could still be handing an agent
|
||||
another agent's brain.
|
||||
- The observation `ActivateRestore` returns ran no transition, so it carries **no audio
|
||||
chunk**, and one in it is refused. Section 2's chunk is the audio of an interval and this
|
||||
observation covers none; MEDIA-01 implemented that rule as "boundary 0 carries no chunk",
|
||||
which is true of the only such observation that slice could produce and false of this one.
|
||||
The rule is about provenance, not about the boundary number.
|
||||
- A participant that staged into a group install the coordinator then abandoned must be
|
||||
**replaced** before another restore, exactly as one that activated must. It is holding a
|
||||
validated replacement state that nothing installed, and [session RPC](ipc-v1.md) section 6
|
||||
already refuses to silently reattach such a participant to an active epoch. Without this the
|
||||
group's second attempt meets its own leftovers and calls them a conflict.
|
||||
|
||||
If emulator validation requires mutation, stage a stopped replacement emulator. If that cannot
|
||||
provide externally atomic resume, advertise episode-restart, not exact-checkpoint. After all
|
||||
activation acknowledgments, install the coordinator's staged task/executor/admission state
|
||||
|
|
|
|||
|
|
@ -223,7 +223,12 @@ export function decode(input: Uint8Array): Envelope {
|
|||
};
|
||||
}
|
||||
|
||||
/** The manifest fields state-media-v1 section 4 requires. */
|
||||
/**
|
||||
* The manifest fields state-media-v1 section 4 requires.
|
||||
*
|
||||
* `helperState` and `environment` join the list under the 2026-09-22 amendment to
|
||||
* checkpoint-envelope-v1 section 3.
|
||||
*/
|
||||
export const REQUIRED_MANIFEST_FIELDS = [
|
||||
'envelopeVersion',
|
||||
'checkpointId',
|
||||
|
|
@ -236,6 +241,8 @@ export const REQUIRED_MANIFEST_FIELDS = [
|
|||
'compatibility',
|
||||
'agents',
|
||||
'coordinator',
|
||||
'environment',
|
||||
'helperState',
|
||||
'payloads',
|
||||
] as const;
|
||||
|
||||
|
|
|
|||
|
|
@ -199,6 +199,7 @@ fn checkpoint_envelope() -> String {
|
|||
"admissionState": null,
|
||||
"eventWatermarks": {"lastEventId": "evt-1", "lastOrdinal": "7"},
|
||||
},
|
||||
"environment": {"workerId": "arena", "payload": "world"},
|
||||
"helperState": [],
|
||||
"payloads": payload_table(),
|
||||
});
|
||||
|
|
|
|||
|
|
@ -63,6 +63,10 @@
|
|||
"lastOrdinal": "7"
|
||||
}
|
||||
},
|
||||
"environment": {
|
||||
"workerId": "arena",
|
||||
"payload": "world"
|
||||
},
|
||||
"helperState": [],
|
||||
"payloads": [
|
||||
{
|
||||
|
|
@ -115,49 +119,49 @@
|
|||
}
|
||||
],
|
||||
"envelope": {
|
||||
"base64": "RkxZU0VTUzEBAAAAIAAAAAAIAAAFAAAAIAgAAAAAAAB7ImFnZW50cyI6W3siYWdlbnRJZCI6ImZseS1hIiwiYnJhaW5UaWNrcyI6IjI1MzQiLCJkYXRhc2V0RGlnZXN0IjoiNmMwYWYxZjA3ODRlZjYzYTM5M2VlNzdkNjE0ZTgyNDZjNjI1MDUxMzYwZjNmMWE0ODgzODM3NGM1ZDM1NWI1MiIsIm1vZGVsVmVyc2lvbiI6ImxpZi0xbXMtZjY0LXYyIiwicGF5bG9hZCI6ImFnZW50LWZseS1hIiwicGxhc3RpY2l0eVZlcnNpb24iOiJmbHkta2MtbWJvbi1yc3RkcC12MiIsInByb2ZpbGVEaWdlc3QiOiIxOTAwZWFiNmMwMjg0ODNkNzEyNjU5OWVlNmY1MGRlMGQyNzkwN2I1YzY1ZmE5MDUyNDU4MGI0YjBmOTg1MmIwIiwicmVtYWluZGVyIjp7ImRlbm9taW5hdG9yIjoiMyIsIm51bWVyYXRvciI6IjEwMDAwMDAifSwic2VlZCI6LTE4NDk0NjA2M31dLCJjaGVja3BvaW50SWQiOiJja3B0LTEiLCJjb21wYXRpYmlsaXR5Ijp7ImJhY2tlbmREaWdlc3QiOiIxMGUwOGE0MTllODUwZWJhMWViYmExOGZkZDI4ZWI3ZWMxYjdlOGJhYTliY2MzYjk3M2UyYjg4OTFlYzcyNmJlIiwiY29udGVudERpZ2VzdCI6ImVkNzAwMmI0MzllOWFjODQ1ZjIyMzU3ZDgyMmJhYzE0NDQ3MzBmYmRiNjAxNmQzZWM5NDMyMjk3YjllYzlmNzMiLCJjb250cm9sbGVyRGlnZXN0IjoiYzE0NzIxMzViMTRjNzdjOGJlZjk4ZTczZjcwMjA4MzI1ZmEwZGNmMWU2YmQ2NjhhZTliMzFhOWNlYTI5NWZlNyIsInBhcnNlckRpZ2VzdCI6ImIxN2Q0NTEyMTE1MDkyOGYyMTQ2YWY0OWUxOTVlZmYxZWVmNWQ2NzMyNWJlMjczYTczM2ZiNzRhY2FkYWEzNDIiLCJwYXRjaERpZ2VzdCI6ImE0ODk1ZWI0NGFmYzMzNmZlY2JiYTZlNTIwY2Q2N2UxNzhkYWNlMDI3NjY1NWQxMDJmY2VmZmE4ZTVmNzA1NzAiLCJzdGF0ZUZvcm1hdElkIjoiZmx5c2Vzcy0xIn0sImNvbXBvc2l0aW9uRGlnZXN0IjoiNzMwZDcyNWM4YTU5ZDNhNzMwM2RlZjJiZWQwNDFhNTc3ZWRiNDI1NWFhYmQ0ODg5Y2UxMjkxODMxMWQ5NTJmMCIsImNvb3JkaW5hdG9yIjp7ImFkbWlzc2lvblN0YXRlIjpudWxsLCJldmVudFdhdGVybWFya3MiOnsibGFzdEV2ZW50SWQiOiJldnQtMSIsImxhc3RPcmRpbmFsIjoiNyJ9LCJleGVjdXRvclN0YXRlIjpbeyJhZ2VudElkIjoiZmx5LWEiLCJwYXlsb2FkIjoiZXhlY3V0b3ItZmx5LWEifV0sInByaW9ySW5zcGVjdGlvbiI6InByaW9yLWluc3BlY3Rpb24iLCJ0YXNrTGVkZ2VyIjoidGFzay1sZWRnZXIifSwiZW52ZWxvcGVWZXJzaW9uIjoxLCJlcGlzb2RlSWQiOiJlcGlzb2RlLTEiLCJoZWxwZXJTdGF0ZSI6W10sInBheWxvYWRzIjpbeyJieXRlTGVuZ3RoIjoiMTciLCJkaWdlc3QiOiIxMzIxZGZmYjBjZGM2ZjkwOTJjYmY3ZmEyYTVmYzY4YmJlZDEyYzk5M2Q1YWQzOTgyNjQwMTI4MTBjZTliZjkzIiwibmFtZSI6ImFnZW50LWZseS1hIn0seyJieXRlTGVuZ3RoIjoiMTQiLCJkaWdlc3QiOiIzYWVlNjBkZjdlMjllZmViYTdmNWY5OWZjNTg2NzY0N2IzNmFlYmZmMWQ1ZDNjODM4ZGJmZjMyMzEyMmU2NDYyIiwibmFtZSI6ImV4ZWN1dG9yLWZseS1hIn0seyJieXRlTGVuZ3RoIjoiMTEiLCJkaWdlc3QiOiI0MGIwMGVkMmJiYmE5MDFkNjgyMDVmZjcxYjA0YTQ0YjllZTUzYzUxY2IzMTA5YWEyY2VhYTQ0ZjFjNDU3MjdlIiwibmFtZSI6InRhc2stbGVkZ2VyIn0seyJieXRlTGVuZ3RoIjoiMTAiLCJkaWdlc3QiOiIyYzEzYjdiNGQ5YTk5MTY4MDFhYjkxOTFjMzE0ZjMxYjA0NWU5YjljNWI2NjlhNmMwNDc0ZjAyMTdlZjc1YmY1IiwibmFtZSI6InByaW9yLWluc3BlY3Rpb24ifSx7ImJ5dGVMZW5ndGgiOiI2NCIsImRpZ2VzdCI6ImY1YTVmZDQyZDE2YTIwMzAyNzk4ZWY2ZWQzMDk5NzliNDMwMDNkMjMyMGQ5ZjBlOGVhOTgzMWE5Mjc1OWZiNGIiLCJuYW1lIjoid29ybGQifV0sInBvcnRNYXAiOlt7ImFnZW50SWQiOiJmbHktYSIsInBvcnRJZCI6InBvcnQtMSJ9XSwic2NoZWR1bGVySWQiOiJsb2Nrc3RlcC12MSIsInNvdXJjZVNjb3BlIjp7ImVwb2NoIjoiZXBvY2gtMSIsInNlc3Npb25JZCI6ImRlbW8iLCJzdGVwIjoiNDIifSwid29ybGRUaW1lIjp7ImRlbm9taW5hdG9yIjoiMSIsIm51bWVyYXRvciI6IjcwMDAwMDAwMCJ9fWFnZW50LWZseS1hAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAABQCgAAAAAAABEAAAAAAAAAEyHf+wzcb5CSy/f6Kl/Gi77RLJk9WtOYJkASgQzpv5NleGVjdXRvci1mbHktYQAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAaAoAAAAAAAAOAAAAAAAAADruYN9+Ke/rp/X5n8WGdkezauv/HV08g42/8yMSLmRidGFzay1sZWRnZXIAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAHgKAAAAAAAACwAAAAAAAABAsA7Su7qQHWggX/cbBKRLnuU8UcsxCaos6qRPHEVyfnByaW9yLWluc3BlY3Rpb24AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACICgAAAAAAAAoAAAAAAAAALBO3tNmpkWgBq5GRwxTzGwRem5xbZppsBHTwIX73W/V3b3JsZAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAmAoAAAAAAABAAAAAAAAAAPWl/ULRaiAwJ5jvbtMJl5tDAD0jINnw6OqYMaknWftLYWdlbnQgc3RhdGUgYnl0ZXMAAAAAAAAAZXhlY3V0b3Igc3RhdGUAAHsicmFuayI6MTB9AAAAAAB7Im1hcCI6NDB9AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAgLAAAAAAAAq++fEx+FvDZho/eB4imbENN4HZrGNC2OCAsI7/gp9r5GTFlTRVNTRg==",
|
||||
"byteLength": 2824,
|
||||
"base64": "RkxZU0VTUzEBAAAAIAAAADUIAAAFAAAAWAgAAAAAAAB7ImFnZW50cyI6W3siYWdlbnRJZCI6ImZseS1hIiwiYnJhaW5UaWNrcyI6IjI1MzQiLCJkYXRhc2V0RGlnZXN0IjoiNmMwYWYxZjA3ODRlZjYzYTM5M2VlNzdkNjE0ZTgyNDZjNjI1MDUxMzYwZjNmMWE0ODgzODM3NGM1ZDM1NWI1MiIsIm1vZGVsVmVyc2lvbiI6ImxpZi0xbXMtZjY0LXYyIiwicGF5bG9hZCI6ImFnZW50LWZseS1hIiwicGxhc3RpY2l0eVZlcnNpb24iOiJmbHkta2MtbWJvbi1yc3RkcC12MiIsInByb2ZpbGVEaWdlc3QiOiIxOTAwZWFiNmMwMjg0ODNkNzEyNjU5OWVlNmY1MGRlMGQyNzkwN2I1YzY1ZmE5MDUyNDU4MGI0YjBmOTg1MmIwIiwicmVtYWluZGVyIjp7ImRlbm9taW5hdG9yIjoiMyIsIm51bWVyYXRvciI6IjEwMDAwMDAifSwic2VlZCI6LTE4NDk0NjA2M31dLCJjaGVja3BvaW50SWQiOiJja3B0LTEiLCJjb21wYXRpYmlsaXR5Ijp7ImJhY2tlbmREaWdlc3QiOiIxMGUwOGE0MTllODUwZWJhMWViYmExOGZkZDI4ZWI3ZWMxYjdlOGJhYTliY2MzYjk3M2UyYjg4OTFlYzcyNmJlIiwiY29udGVudERpZ2VzdCI6ImVkNzAwMmI0MzllOWFjODQ1ZjIyMzU3ZDgyMmJhYzE0NDQ3MzBmYmRiNjAxNmQzZWM5NDMyMjk3YjllYzlmNzMiLCJjb250cm9sbGVyRGlnZXN0IjoiYzE0NzIxMzViMTRjNzdjOGJlZjk4ZTczZjcwMjA4MzI1ZmEwZGNmMWU2YmQ2NjhhZTliMzFhOWNlYTI5NWZlNyIsInBhcnNlckRpZ2VzdCI6ImIxN2Q0NTEyMTE1MDkyOGYyMTQ2YWY0OWUxOTVlZmYxZWVmNWQ2NzMyNWJlMjczYTczM2ZiNzRhY2FkYWEzNDIiLCJwYXRjaERpZ2VzdCI6ImE0ODk1ZWI0NGFmYzMzNmZlY2JiYTZlNTIwY2Q2N2UxNzhkYWNlMDI3NjY1NWQxMDJmY2VmZmE4ZTVmNzA1NzAiLCJzdGF0ZUZvcm1hdElkIjoiZmx5c2Vzcy0xIn0sImNvbXBvc2l0aW9uRGlnZXN0IjoiNzMwZDcyNWM4YTU5ZDNhNzMwM2RlZjJiZWQwNDFhNTc3ZWRiNDI1NWFhYmQ0ODg5Y2UxMjkxODMxMWQ5NTJmMCIsImNvb3JkaW5hdG9yIjp7ImFkbWlzc2lvblN0YXRlIjpudWxsLCJldmVudFdhdGVybWFya3MiOnsibGFzdEV2ZW50SWQiOiJldnQtMSIsImxhc3RPcmRpbmFsIjoiNyJ9LCJleGVjdXRvclN0YXRlIjpbeyJhZ2VudElkIjoiZmx5LWEiLCJwYXlsb2FkIjoiZXhlY3V0b3ItZmx5LWEifV0sInByaW9ySW5zcGVjdGlvbiI6InByaW9yLWluc3BlY3Rpb24iLCJ0YXNrTGVkZ2VyIjoidGFzay1sZWRnZXIifSwiZW52ZWxvcGVWZXJzaW9uIjoxLCJlbnZpcm9ubWVudCI6eyJwYXlsb2FkIjoid29ybGQiLCJ3b3JrZXJJZCI6ImFyZW5hIn0sImVwaXNvZGVJZCI6ImVwaXNvZGUtMSIsImhlbHBlclN0YXRlIjpbXSwicGF5bG9hZHMiOlt7ImJ5dGVMZW5ndGgiOiIxNyIsImRpZ2VzdCI6IjEzMjFkZmZiMGNkYzZmOTA5MmNiZjdmYTJhNWZjNjhiYmVkMTJjOTkzZDVhZDM5ODI2NDAxMjgxMGNlOWJmOTMiLCJuYW1lIjoiYWdlbnQtZmx5LWEifSx7ImJ5dGVMZW5ndGgiOiIxNCIsImRpZ2VzdCI6IjNhZWU2MGRmN2UyOWVmZWJhN2Y1Zjk5ZmM1ODY3NjQ3YjM2YWViZmYxZDVkM2M4MzhkYmZmMzIzMTIyZTY0NjIiLCJuYW1lIjoiZXhlY3V0b3ItZmx5LWEifSx7ImJ5dGVMZW5ndGgiOiIxMSIsImRpZ2VzdCI6IjQwYjAwZWQyYmJiYTkwMWQ2ODIwNWZmNzFiMDRhNDRiOWVlNTNjNTFjYjMxMDlhYTJjZWFhNDRmMWM0NTcyN2UiLCJuYW1lIjoidGFzay1sZWRnZXIifSx7ImJ5dGVMZW5ndGgiOiIxMCIsImRpZ2VzdCI6IjJjMTNiN2I0ZDlhOTkxNjgwMWFiOTE5MWMzMTRmMzFiMDQ1ZTliOWM1YjY2OWE2YzA0NzRmMDIxN2VmNzViZjUiLCJuYW1lIjoicHJpb3ItaW5zcGVjdGlvbiJ9LHsiYnl0ZUxlbmd0aCI6IjY0IiwiZGlnZXN0IjoiZjVhNWZkNDJkMTZhMjAzMDI3OThlZjZlZDMwOTk3OWI0MzAwM2QyMzIwZDlmMGU4ZWE5ODMxYTkyNzU5ZmI0YiIsIm5hbWUiOiJ3b3JsZCJ9XSwicG9ydE1hcCI6W3siYWdlbnRJZCI6ImZseS1hIiwicG9ydElkIjoicG9ydC0xIn1dLCJzY2hlZHVsZXJJZCI6ImxvY2tzdGVwLXYxIiwic291cmNlU2NvcGUiOnsiZXBvY2giOiJlcG9jaC0xIiwic2Vzc2lvbklkIjoiZGVtbyIsInN0ZXAiOiI0MiJ9LCJ3b3JsZFRpbWUiOnsiZGVub21pbmF0b3IiOiIxIiwibnVtZXJhdG9yIjoiNzAwMDAwMDAwIn19AAAAYWdlbnQtZmx5LWEAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAIgKAAAAAAAAEQAAAAAAAAATId/7DNxvkJLL9/oqX8aLvtEsmT1a05gmQBKBDOm/k2V4ZWN1dG9yLWZseS1hAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACgCgAAAAAAAA4AAAAAAAAAOu5g334p7+un9fmfxYZ2R7Nq6/8dXTyDjb/zIxIuZGJ0YXNrLWxlZGdlcgAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAsAoAAAAAAAALAAAAAAAAAECwDtK7upAdaCBf9xsEpEue5TxRyzEJqizqpE8cRXJ+cHJpb3ItaW5zcGVjdGlvbgAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAMAKAAAAAAAACgAAAAAAAAAsE7e02amRaAGrkZHDFPMbBF6bnFtmmmwEdPAhfvdb9XdvcmxkAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAADQCgAAAAAAAEAAAAAAAAAA9aX9QtFqIDAnmO9u0wmXm0MAPSMg2fDo6pgxqSdZ+0thZ2VudCBzdGF0ZSBieXRlcwAAAAAAAABleGVjdXRvciBzdGF0ZQAAeyJyYW5rIjoxMH0AAAAAAHsibWFwIjo0MH0AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAQAsAAAAAAAAhBlh9AWLTtVKckmeIzNn4DO5Yhn2C1nUA1T1RfxMXlkZMWVNFU1NG",
|
||||
"byteLength": 2880,
|
||||
"layout": {
|
||||
"headerBytes": 32,
|
||||
"manifestOffset": "32",
|
||||
"manifestBytes": 2048,
|
||||
"tableOffset": "2080",
|
||||
"manifestBytes": 2101,
|
||||
"tableOffset": "2136",
|
||||
"tableEntryBytes": 112,
|
||||
"entries": [
|
||||
{
|
||||
"name": "agent-fly-a",
|
||||
"offset": "2640",
|
||||
"offset": "2696",
|
||||
"byteLength": "17",
|
||||
"digest": "1321dffb0cdc6f9092cbf7fa2a5fc68bbed12c993d5ad398264012810ce9bf93"
|
||||
},
|
||||
{
|
||||
"name": "executor-fly-a",
|
||||
"offset": "2664",
|
||||
"offset": "2720",
|
||||
"byteLength": "14",
|
||||
"digest": "3aee60df7e29efeba7f5f99fc5867647b36aebff1d5d3c838dbff323122e6462"
|
||||
},
|
||||
{
|
||||
"name": "task-ledger",
|
||||
"offset": "2680",
|
||||
"offset": "2736",
|
||||
"byteLength": "11",
|
||||
"digest": "40b00ed2bbba901d68205ff71b04a44b9ee53c51cb3109aa2ceaa44f1c45727e"
|
||||
},
|
||||
{
|
||||
"name": "prior-inspection",
|
||||
"offset": "2696",
|
||||
"offset": "2752",
|
||||
"byteLength": "10",
|
||||
"digest": "2c13b7b4d9a9916801ab9191c314f31b045e9b9c5b669a6c0474f0217ef75bf5"
|
||||
},
|
||||
{
|
||||
"name": "world",
|
||||
"offset": "2712",
|
||||
"offset": "2768",
|
||||
"byteLength": "64",
|
||||
"digest": "f5a5fd42d16a20302798ef6ed309979b43003d2320d9f0e8ea9831a92759fb4b"
|
||||
}
|
||||
],
|
||||
"footerOffset": "2776",
|
||||
"footerOffset": "2832",
|
||||
"footerBytes": 48,
|
||||
"totalBytes": "2824"
|
||||
"totalBytes": "2880"
|
||||
}
|
||||
},
|
||||
"corruption": [
|
||||
|
|
@ -178,17 +182,17 @@
|
|||
},
|
||||
{
|
||||
"name": "a flipped payload byte",
|
||||
"offset": 2640,
|
||||
"offset": 2696,
|
||||
"reason": "every payload carries its own digest"
|
||||
},
|
||||
{
|
||||
"name": "a flipped footer digest byte",
|
||||
"offset": 2784,
|
||||
"offset": 2840,
|
||||
"reason": "the footer digest must match the contents"
|
||||
},
|
||||
{
|
||||
"name": "a flipped footer magic byte",
|
||||
"offset": 2816,
|
||||
"offset": 2872,
|
||||
"reason": "a truncated file cannot look complete"
|
||||
}
|
||||
]
|
||||
|
|
|
|||
|
|
@ -296,6 +296,11 @@ pub fn decode(bytes: &[u8]) -> Result<Envelope> {
|
|||
|
||||
/// The manifest fields state-media-v1 section 4 requires, checked as a set: a manifest that
|
||||
/// omits one of them is not a complete checkpoint.
|
||||
///
|
||||
/// `helperState` and `environment` join the list under the 2026-09-22 amendment to
|
||||
/// checkpoint-envelope-v1 section 3: the first has been in that section's table from the
|
||||
/// start and was missing here, and the second is the holder of the world's own payload, which
|
||||
/// the table named for every other participant and not for the environment.
|
||||
pub const REQUIRED_MANIFEST_FIELDS: &[&str] = &[
|
||||
"envelopeVersion",
|
||||
"checkpointId",
|
||||
|
|
@ -308,6 +313,8 @@ pub const REQUIRED_MANIFEST_FIELDS: &[&str] = &[
|
|||
"compatibility",
|
||||
"agents",
|
||||
"coordinator",
|
||||
"environment",
|
||||
"helperState",
|
||||
"payloads",
|
||||
];
|
||||
|
||||
|
|
|
|||
|
|
@ -44,6 +44,7 @@ Ready(k) ─ Prepare all agents concurrently ───────────
|
|||
| `metrics` | Latency percentiles and the machine's core and memory counters |
|
||||
| `measure` | The execution-mode comparison of the guide's section 5 |
|
||||
| `cli` | The binary's subcommands: `agent`, `environment`, `measure` |
|
||||
| `state` | The durable checkpoint store over `FLYSESS1`: compatibility, generations, the bounded writer |
|
||||
| `harness` | The runnable composition: router, the flies, one arena, one coordinator |
|
||||
|
||||
## Execution modes and the launcher
|
||||
|
|
@ -178,6 +179,38 @@ harness.shutdown().await;
|
|||
event ids derived from epoch, source step, rule and ordinal.
|
||||
- **Executors.** The stateless identity executor only, as v1 specifies.
|
||||
|
||||
## Checkpoints and recovery
|
||||
|
||||
The durable store is `state`, over the `FLYSESS1` layout the contract crate owns.
|
||||
|
||||
- **One boundary, every participant.** `Coordinator::capture` runs at `Ready(k)` or
|
||||
`Paused(k)` only. It takes its queue slot *before* the first `State.Capture`, so a saturated
|
||||
writer refuses the capture rather than queueing it without bound, and the refusal is a
|
||||
`BUSY` a stepping session survives rather than an epoch failure.
|
||||
- **Capture and durability are two events.** `State.Capture` completes when an immutable
|
||||
capture exists; `Coordinator::await_durable` completes when the store manifest rename has
|
||||
happened, which is the durable commit point. Only the second moves the durable mark. A lost
|
||||
save reply is `SaveOutcome::ReplyLost`, and `Coordinator::resolve_durable` then asks the
|
||||
store about the *same* checkpoint instead of saving again.
|
||||
- **The writer is bounded twice**, by outstanding captures and by queued bytes, and it owns
|
||||
its payload handles until the bytes are committed or the job fails. A queued *replaceable*
|
||||
capture is superseded by a later one, releasing its holds; a durable one never is.
|
||||
- **The install is a group.** A restore selects a complete compatible generation, imports its
|
||||
payloads as fresh artifacts, stages every participant, validates the coordinator's own
|
||||
ledgers, and only then activates. A failure anywhere leaves the fence closed, and every
|
||||
participant that got as far as staging is recorded as one that must be replaced before
|
||||
another restore is attempted.
|
||||
- **The fence lifts once.** `Failed -> Restoring(k) -> Paused(k)`, at the end of a complete
|
||||
install and nowhere else. A fenced session takes no step, publishes nothing, captures
|
||||
nothing and holds no artifact handle.
|
||||
- **Nothing old crosses.** The fence drops every media handle; the restore imports fresh
|
||||
artifacts; the environment re-renders its pending sensor pipeline from recorded
|
||||
reconstruction inputs; and the new epoch's first audio chunk resumes the preserved sample
|
||||
position and marks the discontinuity.
|
||||
- **Epoch metadata in a trace.** `scope.epoch`, the batch id and every task event id are
|
||||
derived from the epoch, so a resumed run's behaviour is compared through
|
||||
`EpochRebase`, which rewrites exactly those and fails on anything it does not recognise.
|
||||
|
||||
## Where this crate narrows or adds to the contract crate
|
||||
|
||||
- **Required views.** `WorldObservation::validate_against` checks the views a result carries
|
||||
|
|
@ -196,9 +229,9 @@ harness.shutdown().await;
|
|||
|
||||
- **Fake workers.** There is no neural model and no emulator. What is modelled exactly is the
|
||||
ordering, the identity rules and the retry rules, not any numerical behaviour.
|
||||
- **No state methods.** `State.Capture`, `State.StageRestore` and `State.ActivateRestore` are
|
||||
STATE-01. The phase machine has their edges (`Capturing`, `Restoring`) and the workers do not
|
||||
advertise them as implemented methods.
|
||||
- **One environment, one task.** A checkpoint records the composition it was taken from, and a
|
||||
restore refuses one taken under another backend, content, patch, controller or parser
|
||||
identity. It does not migrate between compositions, and it does not try.
|
||||
- **No audience input.** The admitted pre-step stimulation list exists and is always empty.
|
||||
- **Pacing is coarse.** The pacing deadline rounds one step to whole nanoseconds for sleeping
|
||||
only; simulation time stays rational and that rounding never re-enters the accumulator.
|
||||
|
|
@ -259,6 +292,9 @@ The three integration suites do not all run over both transports, and cannot:
|
|||
- `tests/processes.rs` runs over the Unix socket only, in all three execution modes. A
|
||||
participant in a process of its own has no in-memory transport to reach the router by, so
|
||||
the mode is the axis that suite varies and the transport is fixed.
|
||||
- `tests/media.rs` and `tests/state.rs` run over both transports *and* in all three execution
|
||||
modes: each acceptance body is written once and registered twice, by `both_transports!` in
|
||||
the in-process composition and by `all_modes!` over the socket.
|
||||
|
||||
- `tests/session.rs`: one world advance per complete batch; every agent Prepared before the
|
||||
advance; one task evaluation per transition; every agent committed before the next Prepare or
|
||||
|
|
@ -275,6 +311,14 @@ The three integration suites do not all run over both transports, and cannot:
|
|||
allocation -- plus the sequential/reversed/parallel trace comparison across all three modes
|
||||
and the two process-mode section 4 rows: a router restart during a world advance, and an old
|
||||
worker's reply after a restart.
|
||||
- `tests/state.rs`: the STATE-01 acceptance bullets -- an uninterrupted run and a resumed run
|
||||
committing the same behaviour once the epoch metadata is rebased, a corrupt payload failing
|
||||
the install as a group for every participant and for the coordinator's own ledger, a lost
|
||||
save reply and an uncommitted store manifest both leaving the durable mark where it was, a
|
||||
refused activation resuming no part of the world, the capture queue staying bounded under a
|
||||
stalled writer, and old media and another parser's state failing to cross a recovery --
|
||||
plus the once-only restore token, the superseded replaceable capture, and the fence that
|
||||
lifts only through a complete restore.
|
||||
- `tests/failures.rs`: a duplicate Prepare after a lost reply; a duplicate Commit; the same
|
||||
batch with altered controls; a lost Advance result; a cached artifact consumed by its first
|
||||
caller; one Commit failing after another succeeded; a replaced registration; a reply from
|
||||
|
|
|
|||
|
|
@ -7,7 +7,7 @@
|
|||
//! stimulation, then reinforces once, and executes no tick at all. Every mutating step bumps
|
||||
//! one counter, which is how a test proves a duplicate request changed nothing.
|
||||
|
||||
use std::collections::BTreeMap;
|
||||
use std::collections::{BTreeMap, BTreeSet};
|
||||
|
||||
use serde_json::Value;
|
||||
|
||||
|
|
@ -193,6 +193,12 @@ pub struct AgentFaults {
|
|||
pub prepare_delay_ms: u64,
|
||||
/// Hold `Agent.Commit` open for this long.
|
||||
pub commit_delay_ms: u64,
|
||||
/// Refuse `State.StageRestore`, so a group install meets one participant that will not
|
||||
/// validate while the others already have.
|
||||
pub fail_stage_restore: bool,
|
||||
/// Refuse `State.ActivateRestore` after this worker has already staged, so a group meets
|
||||
/// a failure halfway through activation.
|
||||
pub fail_activate_restore: bool,
|
||||
}
|
||||
|
||||
/// One fake agent worker's configuration.
|
||||
|
|
@ -234,6 +240,10 @@ pub struct FakeAgentWorker {
|
|||
context: Option<TypedValue>,
|
||||
context_digest: Option<Digest>,
|
||||
prepared: Option<(DomainRequestId, PreparedDecision)>,
|
||||
/// A validated replacement state that the live session cannot see yet.
|
||||
staged: Option<StagedAgent>,
|
||||
/// Restore tokens this worker has activated. A token activates once.
|
||||
activated: BTreeSet<Id>,
|
||||
}
|
||||
|
||||
impl FakeAgentWorker {
|
||||
|
|
@ -248,10 +258,17 @@ impl FakeAgentWorker {
|
|||
context: None,
|
||||
context_digest: None,
|
||||
prepared: None,
|
||||
staged: None,
|
||||
activated: BTreeSet::new(),
|
||||
config,
|
||||
}
|
||||
}
|
||||
|
||||
/// True while a validated replacement state is staged and not yet activated.
|
||||
pub fn has_staged_restore(&self) -> bool {
|
||||
self.staged.is_some()
|
||||
}
|
||||
|
||||
pub fn status(&self) -> StatusCell {
|
||||
self.status.clone()
|
||||
}
|
||||
|
|
@ -657,7 +674,11 @@ impl WorkerEndpoint for FakeAgentWorker {
|
|||
}
|
||||
|
||||
fn capabilities(&self) -> Vec<Id> {
|
||||
vec![id("agent-step-v1"), id("pixel-observation-v1")]
|
||||
vec![
|
||||
id("agent-step-v1"),
|
||||
id("pixel-observation-v1"),
|
||||
id(crate::state::CHECKPOINT_CAPABILITY),
|
||||
]
|
||||
}
|
||||
|
||||
fn status_cell(&self) -> StatusCell {
|
||||
|
|
@ -669,7 +690,14 @@ impl WorkerEndpoint for FakeAgentWorker {
|
|||
}
|
||||
|
||||
fn methods(&self) -> Vec<&'static str> {
|
||||
vec!["Agent.Initialize", "Agent.Prepare", "Agent.Commit"]
|
||||
vec![
|
||||
"Agent.Initialize",
|
||||
"Agent.Prepare",
|
||||
"Agent.Commit",
|
||||
"State.Capture",
|
||||
"State.StageRestore",
|
||||
"State.ActivateRestore",
|
||||
]
|
||||
}
|
||||
|
||||
fn handle<'a>(&'a mut self, ctx: HandlerCtx<'a>) -> BoxFuture<'a, DomainResult<HandlerReply>> {
|
||||
|
|
@ -678,6 +706,9 @@ impl WorkerEndpoint for FakeAgentWorker {
|
|||
"Agent.Initialize" => self.initialize(&ctx).await,
|
||||
"Agent.Prepare" => self.prepare(&ctx).await,
|
||||
"Agent.Commit" => self.commit(&ctx).await,
|
||||
"State.Capture" => self.state_capture(&ctx).await,
|
||||
"State.StageRestore" => self.state_stage_restore(&ctx).await,
|
||||
"State.ActivateRestore" => self.state_activate_restore(&ctx).await,
|
||||
other => Err(DomainError::before(
|
||||
ErrorCode::Unsupported,
|
||||
format!("{other} is not an agent method"),
|
||||
|
|
@ -690,7 +721,12 @@ impl WorkerEndpoint for FakeAgentWorker {
|
|||
/// The retention class table an agent endpoint follows, for a caller that wants it.
|
||||
pub fn agent_op_class(method: &str) -> Option<OpClass> {
|
||||
match method {
|
||||
"Agent.Initialize" => Some(OpClass::Lifecycle),
|
||||
// `ipc-v1` section 5: lifecycle *and capture* replies are retained until
|
||||
// `Worker.Acknowledge`, which is also what lets a duplicate restore request replay
|
||||
// its cached reply rather than staging or activating twice.
|
||||
"Agent.Initialize" | "State.Capture" | "State.StageRestore" | "State.ActivateRestore" => {
|
||||
Some(OpClass::Lifecycle)
|
||||
}
|
||||
"Agent.Prepare" | "Agent.Commit" => Some(OpClass::StepMutation),
|
||||
_ => None,
|
||||
}
|
||||
|
|
@ -712,3 +748,532 @@ pub fn synthetic_profile(agent_id: &Id, tick_duration: &RationalNs, warmup_ticks
|
|||
|
||||
/// The per-agent contexts a bootstrap produced, keyed by agent id.
|
||||
pub type Contexts = BTreeMap<Id, TypedValue>;
|
||||
|
||||
// -------------------------------------------------------------------------------------------
|
||||
// STATE-01: capture and restore
|
||||
|
||||
/// The numerical model version this worker implements. It is part of a capture's
|
||||
/// compatibility identity: the same profile and seed under another model is not the same
|
||||
/// state (`workers-v1` section 2).
|
||||
pub const MODEL_VERSION: &str = "fake-lcg-v1";
|
||||
|
||||
/// The plasticity rule version, for the same reason.
|
||||
pub const PLASTICITY_VERSION: &str = "fake-reinforce-v1";
|
||||
|
||||
/// The version this payload layout is written and read under.
|
||||
pub const AGENT_PAYLOAD_VERSION: u64 = 1;
|
||||
|
||||
/// The dataset identity a synthetic agent resolves.
|
||||
///
|
||||
/// There is no connectome dataset behind this worker, and a checkpoint says so with a stable
|
||||
/// identity rather than omitting the field: "no dataset" has to be distinguishable from "the
|
||||
/// dataset was not recorded".
|
||||
pub fn dataset_digest() -> Digest {
|
||||
digest_of_bytes(b"fly-session/no-dataset-v1")
|
||||
}
|
||||
|
||||
/// The capture compatibility digest of one agent (`workers-v1` section 2).
|
||||
///
|
||||
/// The profile digest identifies the profile definition; this additionally covers the
|
||||
/// resolved seed, the numerical model version and the plasticity rule, because two agents
|
||||
/// with the same profile digest and different seeds hold state that is not interchangeable.
|
||||
/// Every field it covers is one the checkpoint manifest already records in that agent's row,
|
||||
/// so a restore derives the expected digest from the manifest rather than from the payload it
|
||||
/// is about to validate.
|
||||
pub fn agent_compatibility_digest(
|
||||
agent_id: &Id,
|
||||
profile_digest: &Digest,
|
||||
dataset_digest: &Digest,
|
||||
model_version: &str,
|
||||
plasticity_version: &str,
|
||||
seed: i32,
|
||||
) -> Digest {
|
||||
let value = serde_json::json!({
|
||||
"agentId": agent_id.as_str(),
|
||||
"profileDigest": profile_digest.as_str(),
|
||||
"datasetDigest": dataset_digest.as_str(),
|
||||
"modelVersion": model_version,
|
||||
"plasticityVersion": plasticity_version,
|
||||
"seed": seed,
|
||||
});
|
||||
digest_of(&value).expect("an agent compatibility block canonicalizes")
|
||||
}
|
||||
|
||||
impl FakeModel {
|
||||
/// Every field of the model, so a resumed agent is this agent and not a fresh one.
|
||||
fn capture(&self) -> Value {
|
||||
serde_json::json!({
|
||||
"seed": self.seed,
|
||||
"state": self.state.to_string(),
|
||||
"mutations": self.mutations.to_string(),
|
||||
"ticks": self.ticks.to_string(),
|
||||
"stimulations": self.stimulations.to_string(),
|
||||
"reinforcements": self.reinforcements.to_string(),
|
||||
"learningEnabled": self.learning_enabled,
|
||||
"learningUpdates": self.learning_updates.to_string(),
|
||||
"learningChanged": self.learning_changed.to_string(),
|
||||
"lastSignal": self.last_signal,
|
||||
"inputValue": self.input_value.to_string(),
|
||||
"inputInstalls": self.input_installs.to_string(),
|
||||
})
|
||||
}
|
||||
|
||||
fn restored(value: &Value) -> DomainResult<FakeModel> {
|
||||
let number = |key: &str| -> DomainResult<u64> {
|
||||
value
|
||||
.get(key)
|
||||
.and_then(Value::as_str)
|
||||
.ok_or_else(|| incompatible(format!("the agent payload has no {key}")))?
|
||||
.parse::<u64>()
|
||||
.map_err(|_| incompatible(format!("the agent payload's {key} is not a U64")))
|
||||
};
|
||||
let seed = value
|
||||
.get("seed")
|
||||
.and_then(Value::as_i64)
|
||||
.and_then(|v| i32::try_from(v).ok())
|
||||
.ok_or_else(|| incompatible("the agent payload has no seed"))?;
|
||||
let input_value = value
|
||||
.get("inputValue")
|
||||
.and_then(Value::as_str)
|
||||
.ok_or_else(|| incompatible("the agent payload has no inputValue"))?
|
||||
.parse::<i64>()
|
||||
.map_err(|_| incompatible("the agent payload's inputValue is not an integer"))?;
|
||||
let last_signal = value
|
||||
.get("lastSignal")
|
||||
.and_then(Value::as_f64)
|
||||
.filter(|v| v.is_finite())
|
||||
.ok_or_else(|| incompatible("the agent payload's lastSignal is not finite"))?;
|
||||
let learning_enabled = value
|
||||
.get("learningEnabled")
|
||||
.and_then(Value::as_bool)
|
||||
.ok_or_else(|| incompatible("the agent payload has no learningEnabled"))?;
|
||||
Ok(FakeModel {
|
||||
seed,
|
||||
state: number("state")?,
|
||||
mutations: number("mutations")?,
|
||||
ticks: number("ticks")?,
|
||||
stimulations: number("stimulations")?,
|
||||
reinforcements: number("reinforcements")?,
|
||||
learning_enabled,
|
||||
learning_updates: number("learningUpdates")?,
|
||||
learning_changed: number("learningChanged")?,
|
||||
last_signal,
|
||||
input_value,
|
||||
input_installs: number("inputInstalls")?,
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
fn incompatible(message: impl std::fmt::Display) -> DomainError {
|
||||
DomainError::before(ErrorCode::IncompatibleState, message)
|
||||
}
|
||||
|
||||
/// One staged restore, held outside the live agent until it is activated.
|
||||
struct StagedAgent {
|
||||
token: Id,
|
||||
checkpoint_id: Id,
|
||||
scope: Scope,
|
||||
model: FakeModel,
|
||||
accumulator: TickAccumulator,
|
||||
context: TypedValue,
|
||||
profile: AssetRef,
|
||||
committed_step: u64,
|
||||
}
|
||||
|
||||
impl FakeAgentWorker {
|
||||
/// This worker's own compatibility identity, from its configuration and a resolved seed.
|
||||
fn compatibility_digest(&self, profile: &AssetRef, seed: i32) -> Digest {
|
||||
agent_compatibility_digest(
|
||||
&self.config.agent_id,
|
||||
&profile.digest,
|
||||
&dataset_digest(),
|
||||
MODEL_VERSION,
|
||||
PLASTICITY_VERSION,
|
||||
seed,
|
||||
)
|
||||
}
|
||||
|
||||
/// `State.Capture`: an immutable snapshot of this agent at its committed boundary.
|
||||
///
|
||||
/// It is allowed at `Ready(k)` only. A Prepared agent holds half a transition, and there
|
||||
/// is no coherent boundary to file that under.
|
||||
async fn state_capture(&mut self, ctx: &HandlerCtx<'_>) -> DomainResult<HandlerReply> {
|
||||
let scope = ctx.scope()?.clone();
|
||||
self.check_epoch(&scope)?;
|
||||
let AgentPhase::Ready(k) = self.phase.clone() else {
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::InvalidPhase,
|
||||
format!(
|
||||
"State.Capture needs a quiescent Ready(k); this worker is {:?}",
|
||||
self.phase
|
||||
),
|
||||
));
|
||||
};
|
||||
if scope.step != k {
|
||||
return Err(DomainError::before(
|
||||
if scope.step < k { ErrorCode::StaleStep } else { ErrorCode::FutureStep },
|
||||
"State.Capture names a boundary this worker is not at",
|
||||
));
|
||||
}
|
||||
let params: CaptureParams = ctx.params()?;
|
||||
let profile = self.profile.clone().expect("initialized");
|
||||
let context = self.context.clone().expect("initialized");
|
||||
let accumulator = self.accumulator.as_ref().expect("initialized");
|
||||
let previous = self.status.state();
|
||||
self.status.set_state(WorkerState::Capturing);
|
||||
let payload = serde_json::json!({
|
||||
"payloadVersion": AGENT_PAYLOAD_VERSION,
|
||||
"kind": "agent",
|
||||
"agentId": self.config.agent_id.as_str(),
|
||||
"checkpointId": params.checkpoint_id.as_str(),
|
||||
"sourceScope": scope.to_json(),
|
||||
"committedStep": k.to_string(),
|
||||
"profile": profile.to_json(),
|
||||
"modelVersion": MODEL_VERSION,
|
||||
"plasticityVersion": PLASTICITY_VERSION,
|
||||
"datasetDigest": dataset_digest().as_str(),
|
||||
"model": self.model.capture(),
|
||||
"accumulator": {
|
||||
"tickDuration": accumulator.tick_duration().to_json(),
|
||||
"remainder": accumulator.remainder().to_json(),
|
||||
"executedTicks": accumulator.executed_ticks().to_string(),
|
||||
"warmupOffset": accumulator.warmup_offset().to_string(),
|
||||
},
|
||||
"context": context.to_json(),
|
||||
});
|
||||
let bytes = canonicalize(&payload)
|
||||
.map_err(|e| DomainError::invalid(format!("State.Capture: {}", e.0)))?
|
||||
.into_bytes();
|
||||
let digest = digest_of_bytes(&bytes);
|
||||
let artifact = crate::state::seal_payload(ctx.client, &bytes, &digest).await?;
|
||||
// Capture is a read of the model, not a mutation of it: nothing above changed a
|
||||
// counter, and the worker goes back to the boundary it was already at.
|
||||
self.status.set_state(previous);
|
||||
let result = CaptureResult {
|
||||
checkpoint_id: params.checkpoint_id,
|
||||
boundary: k,
|
||||
compatibility_digest: self.compatibility_digest(&profile, self.model.seed()),
|
||||
payload: artifact.reference().clone(),
|
||||
};
|
||||
Ok(HandlerReply::with_artifacts(
|
||||
object(result.to_json()),
|
||||
vec![(crate::state::PAYLOAD_ATTACHMENT.to_owned(), artifact)],
|
||||
))
|
||||
}
|
||||
|
||||
/// `State.StageRestore`: validate a replacement state into a staging slot.
|
||||
///
|
||||
/// Nothing the live session can see changes here, and the worker keeps whatever state it
|
||||
/// had. It is allowed on an uninitialized replacement or a quiescent worker only; a
|
||||
/// failed one is neither, which is why a group that failed is replaced rather than
|
||||
/// reused.
|
||||
async fn state_stage_restore(&mut self, ctx: &HandlerCtx<'_>) -> DomainResult<HandlerReply> {
|
||||
let scope = ctx.scope()?.clone();
|
||||
if scope.session_id != self.config.session_id {
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::IdentityMismatch,
|
||||
"this worker belongs to another session",
|
||||
));
|
||||
}
|
||||
match &self.phase {
|
||||
AgentPhase::Uninitialized | AgentPhase::Ready(_) => {}
|
||||
other => {
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::InvalidPhase,
|
||||
format!(
|
||||
"State.StageRestore needs an uninitialized replacement or a quiescent \
|
||||
worker; this worker is {other:?}"
|
||||
),
|
||||
));
|
||||
}
|
||||
}
|
||||
if let Some(epoch) = &self.epoch
|
||||
&& *epoch == scope.epoch
|
||||
{
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::StaleEpoch,
|
||||
"State.StageRestore proposes the epoch this worker is already running",
|
||||
));
|
||||
}
|
||||
let params: StageRestoreParams = ctx.params()?;
|
||||
if params.source_scope.step != scope.step {
|
||||
return Err(DomainError::invalid(
|
||||
"State.StageRestore's scope step must be the source boundary",
|
||||
));
|
||||
}
|
||||
let artifact = ctx.artifact(crate::state::PAYLOAD_ATTACHMENT)?;
|
||||
if artifact.reference() != ¶ms.payload {
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::BufferInvalid,
|
||||
"the staged payload attachment is not the artifact the request names",
|
||||
));
|
||||
}
|
||||
let bytes = artifact.read_all().await.map_err(|e| {
|
||||
DomainError::before(
|
||||
ErrorCode::BufferInvalid,
|
||||
format!("the staged payload could not be read: {}", e.message),
|
||||
)
|
||||
})?;
|
||||
let declared = params
|
||||
.payload
|
||||
.digest
|
||||
.clone()
|
||||
.ok_or_else(|| incompatible("a checkpoint payload must carry a content digest"))?;
|
||||
let actual = digest_of_bytes(&bytes);
|
||||
if actual != declared || bytes.len() as u64 != params.payload.byte_length {
|
||||
return Err(incompatible(
|
||||
"the staged payload is not the content the request declares",
|
||||
));
|
||||
}
|
||||
let value: Value = serde_json::from_slice(&bytes)
|
||||
.map_err(|e| DomainError::invalid(format!("the staged payload is not JSON: {e}")))?;
|
||||
let text = |key: &str| -> DomainResult<String> {
|
||||
value
|
||||
.get(key)
|
||||
.and_then(Value::as_str)
|
||||
.map(str::to_owned)
|
||||
.ok_or_else(|| incompatible(format!("the agent payload has no {key}")))
|
||||
};
|
||||
if value.get("payloadVersion").and_then(Value::as_u64) != Some(AGENT_PAYLOAD_VERSION) {
|
||||
return Err(incompatible("the agent payload is another payload version"));
|
||||
}
|
||||
if text("kind")? != "agent" {
|
||||
return Err(incompatible("this payload is not an agent's state"));
|
||||
}
|
||||
if text("agentId")? != self.config.agent_id {
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::IdentityMismatch,
|
||||
"the staged payload belongs to another agent",
|
||||
));
|
||||
}
|
||||
if text("checkpointId")? != params.checkpoint_id {
|
||||
return Err(incompatible("the staged payload belongs to another checkpoint"));
|
||||
}
|
||||
if text("modelVersion")? != MODEL_VERSION || text("plasticityVersion")? != PLASTICITY_VERSION
|
||||
{
|
||||
return Err(incompatible(
|
||||
"the staged payload was captured under another numerical model",
|
||||
));
|
||||
}
|
||||
let source_scope = Scope::from_json(
|
||||
value
|
||||
.get("sourceScope")
|
||||
.ok_or_else(|| incompatible("the agent payload has no sourceScope"))?,
|
||||
)
|
||||
.map_err(|e| incompatible(format!("the agent payload's sourceScope: {}", e.0)))?;
|
||||
if source_scope != params.source_scope {
|
||||
return Err(incompatible(
|
||||
"the staged payload was captured at another source scope",
|
||||
));
|
||||
}
|
||||
let committed_step: u64 = text("committedStep")?
|
||||
.parse()
|
||||
.map_err(|_| incompatible("the agent payload's committedStep is not a U64"))?;
|
||||
if committed_step != params.source_scope.step {
|
||||
return Err(incompatible(
|
||||
"the staged payload's committed step is not the source boundary",
|
||||
));
|
||||
}
|
||||
let profile = AssetRef::from_json(
|
||||
value
|
||||
.get("profile")
|
||||
.ok_or_else(|| incompatible("the agent payload has no profile"))?,
|
||||
)
|
||||
.map_err(|e| incompatible(format!("the agent payload's profile: {}", e.0)))?;
|
||||
let model = FakeModel::restored(
|
||||
value
|
||||
.get("model")
|
||||
.ok_or_else(|| incompatible("the agent payload has no model"))?,
|
||||
)?;
|
||||
// The compatibility digest is recomputed from this worker's own configuration and the
|
||||
// identity the payload declares. A capture of the same profile under another seed, or
|
||||
// of another agent's brain, fails here and never reaches activation.
|
||||
let computed = self.compatibility_digest(&profile, model.seed());
|
||||
if computed != params.compatibility_digest {
|
||||
return Err(incompatible(format!(
|
||||
"the staged state's compatibility {computed} is not the {} the restore \
|
||||
requires",
|
||||
params.compatibility_digest
|
||||
)));
|
||||
}
|
||||
let accumulator_value = value
|
||||
.get("accumulator")
|
||||
.ok_or_else(|| incompatible("the agent payload has no accumulator"))?;
|
||||
let rational = |key: &str| -> DomainResult<RationalNs> {
|
||||
RationalNs::from_json(
|
||||
accumulator_value
|
||||
.get(key)
|
||||
.ok_or_else(|| incompatible(format!("the accumulator has no {key}")))?,
|
||||
)
|
||||
.map_err(|e| incompatible(format!("the accumulator's {key}: {}", e.0)))
|
||||
};
|
||||
let counter = |key: &str| -> DomainResult<u64> {
|
||||
accumulator_value
|
||||
.get(key)
|
||||
.and_then(Value::as_str)
|
||||
.ok_or_else(|| incompatible(format!("the accumulator has no {key}")))?
|
||||
.parse::<u64>()
|
||||
.map_err(|_| incompatible(format!("the accumulator's {key} is not a U64")))
|
||||
};
|
||||
let tick_duration = rational("tickDuration")?;
|
||||
if tick_duration != self.config.tick_duration {
|
||||
return Err(incompatible(
|
||||
"the staged state was captured at another model tick duration",
|
||||
));
|
||||
}
|
||||
let accumulator = TickAccumulator::restored(
|
||||
tick_duration,
|
||||
rational("remainder")?,
|
||||
counter("executedTicks")?,
|
||||
counter("warmupOffset")?,
|
||||
)
|
||||
.map_err(incompatible)?;
|
||||
let context = TypedValue::from_json(
|
||||
value
|
||||
.get("context")
|
||||
.ok_or_else(|| incompatible("the agent payload has no context"))?,
|
||||
)
|
||||
.map_err(|e| incompatible(format!("the agent payload's context: {}", e.0)))?;
|
||||
FakeAgentWorker::available_actions(&context)?;
|
||||
|
||||
if self.config.faults.fail_stage_restore {
|
||||
// The row where a group validates three participants and the fourth does not.
|
||||
// Nothing is staged here and nothing is staged anywhere else either: the
|
||||
// coordinator abandons the whole install.
|
||||
return Err(incompatible(
|
||||
"injected staging refusal: this participant's replacement state does not \
|
||||
validate",
|
||||
));
|
||||
}
|
||||
// One staged restore at a time. A second proposal replaces nothing silently.
|
||||
if let Some(staged) = &self.staged {
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::Conflict,
|
||||
format!(
|
||||
"this worker already holds the staged restore {} for checkpoint {}",
|
||||
staged.token, staged.checkpoint_id
|
||||
),
|
||||
));
|
||||
}
|
||||
let token = restore_token(¶ms.checkpoint_id, &scope, &actual, &self.config.incarnation_id);
|
||||
if self.activated.contains(&token) {
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::Conflict,
|
||||
"this exact restore was already activated on this worker",
|
||||
));
|
||||
}
|
||||
self.staged = Some(StagedAgent {
|
||||
token: token.clone(),
|
||||
checkpoint_id: params.checkpoint_id.clone(),
|
||||
scope: scope.clone(),
|
||||
model,
|
||||
accumulator,
|
||||
context,
|
||||
profile,
|
||||
committed_step,
|
||||
});
|
||||
self.status.set_state(WorkerState::StagedRestore);
|
||||
let result = StageRestoreResult {
|
||||
checkpoint_id: params.checkpoint_id,
|
||||
restore_token: token,
|
||||
};
|
||||
Ok(HandlerReply::from(&result))
|
||||
}
|
||||
|
||||
/// `State.ActivateRestore`: install the staged state under its new scope, without a tick.
|
||||
///
|
||||
/// The token activates once. A duplicate domain request replays the cached reply through
|
||||
/// the shell's result cache; a fresh request naming an already activated token is a
|
||||
/// conflict, which is what stops a second group from being resumed from the same bytes.
|
||||
async fn state_activate_restore(
|
||||
&mut self,
|
||||
ctx: &HandlerCtx<'_>,
|
||||
) -> DomainResult<HandlerReply> {
|
||||
let params: ActivateRestoreParams = ctx.params()?;
|
||||
if self.activated.contains(¶ms.restore_token) {
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::Conflict,
|
||||
"this restore token has already been activated",
|
||||
));
|
||||
}
|
||||
let Some(staged) = self.staged.take() else {
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::InvalidPhase,
|
||||
"this worker holds no staged restore",
|
||||
));
|
||||
};
|
||||
if staged.token != params.restore_token {
|
||||
// Put it back: naming another token is not a reason to discard this one.
|
||||
let token = staged.token.clone();
|
||||
self.staged = Some(staged);
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::IdentityMismatch,
|
||||
format!("this worker's staged restore is {token}, not {}", params.restore_token),
|
||||
));
|
||||
}
|
||||
if self.config.faults.fail_activate_restore {
|
||||
let token = staged.token.clone();
|
||||
self.staged = Some(staged);
|
||||
self.status.set_state(WorkerState::Failed);
|
||||
return Err(DomainError::new(
|
||||
ErrorCode::BackendFailure,
|
||||
format!("injected activation failure; {token} stays staged and unresumed"),
|
||||
MutationCertainty::None,
|
||||
));
|
||||
}
|
||||
self.status.set_state(WorkerState::Restoring);
|
||||
let StagedAgent {
|
||||
token,
|
||||
checkpoint_id,
|
||||
scope,
|
||||
model,
|
||||
accumulator,
|
||||
context,
|
||||
profile,
|
||||
committed_step,
|
||||
} = staged;
|
||||
self.model = model;
|
||||
self.accumulator = Some(accumulator);
|
||||
self.context_digest = Some(context.digest());
|
||||
self.context = Some(context);
|
||||
self.profile = Some(profile);
|
||||
self.epoch = Some(scope.epoch.clone());
|
||||
self.prepared = None;
|
||||
self.phase = AgentPhase::Ready(committed_step);
|
||||
self.activated.insert(token);
|
||||
self.status.set_state(WorkerState::Ready);
|
||||
self.status.set_scope(Some(scope_at(
|
||||
&scope.session_id,
|
||||
&scope.epoch,
|
||||
committed_step,
|
||||
)));
|
||||
self.status.advance_to(self.model.mutations());
|
||||
let result = ActivateRestoreResult {
|
||||
committed_step,
|
||||
checkpoint_id,
|
||||
// An agent returns a null observation; the environment returns the world's.
|
||||
observation: None,
|
||||
};
|
||||
result
|
||||
.validate_for_role(Role::Agent)
|
||||
.map_err(|e| DomainError::invalid(e.0))?;
|
||||
Ok(HandlerReply::from(&result))
|
||||
}
|
||||
}
|
||||
|
||||
/// A restore token bound to the checkpoint, the proposed scope, the payload bytes and the
|
||||
/// worker incarnation staging them.
|
||||
///
|
||||
/// `state-media-v1` section 5 binds a token to scope, payload and checkpoint. Binding it to
|
||||
/// the incarnation as well is what keeps a token minted by a worker that has since been
|
||||
/// replaced from activating anything on its replacement.
|
||||
pub fn restore_token(checkpoint_id: &Id, scope: &Scope, payload_digest: &Digest, incarnation: &Id) -> Id {
|
||||
let digest = digest_of_bytes(
|
||||
format!(
|
||||
"fly-session/restore-token-v1\n{checkpoint_id}\n{}\n{}\n{}\n{payload_digest}\n{incarnation}\n",
|
||||
scope.session_id, scope.epoch, scope.step
|
||||
)
|
||||
.as_bytes(),
|
||||
);
|
||||
parse_id(&format!("rt-{}", &digest[..32])).expect("a hex suffix is an Id")
|
||||
}
|
||||
|
|
|
|||
|
|
@ -46,8 +46,10 @@ Worker options (agent and environment):
|
|||
agent: --agent ID --port ID --tick-numerator N --tick-denominator N
|
||||
--warmup-ticks N [--prepare-delay-ms N] [--commit-delay-ms N]
|
||||
[--fail-commit-at-step N]
|
||||
[--fail-stage-restore 0|1] [--fail-activate-restore 0|1]
|
||||
environment: --worker ID --ports p1,p2 --step-numerator N --step-denominator N
|
||||
[--advance-delay-ms N] [--omit-view-at-boundary N]
|
||||
[--fail-stage-restore 0|1] [--fail-activate-restore 0|1]
|
||||
|
||||
Measure options:
|
||||
--steps N transitions per run (default 200)
|
||||
|
|
@ -145,6 +147,17 @@ impl Options {
|
|||
}
|
||||
}
|
||||
|
||||
/// A flag whose value is `0` or `1`. Anything else is an error naming it, so a
|
||||
/// mistyped injection is a failed launch rather than a fault that never fires.
|
||||
fn flag(&self, name: &str) -> Result<bool, String> {
|
||||
match self.0.get(name) {
|
||||
None => Ok(false),
|
||||
Some(value) if value == "0" => Ok(false),
|
||||
Some(value) if value == "1" => Ok(true),
|
||||
Some(value) => Err(format!("--{name}: {value:?} is not 0 or 1")),
|
||||
}
|
||||
}
|
||||
|
||||
fn opt_u64(&self, name: &str) -> Result<Option<u64>, String> {
|
||||
match self.0.get(name) {
|
||||
None => Ok(None),
|
||||
|
|
@ -197,6 +210,8 @@ fn serve(role: &str, options: &Options) -> Result<(), String> {
|
|||
fail_commit_at_step: options.opt_u64(flags::FAIL_COMMIT_AT_STEP)?,
|
||||
prepare_delay_ms: options.u64(flags::PREPARE_DELAY_MS, 0)?,
|
||||
commit_delay_ms: options.u64(flags::COMMIT_DELAY_MS, 0)?,
|
||||
fail_stage_restore: options.flag(flags::FAIL_STAGE_RESTORE)?,
|
||||
fail_activate_restore: options.flag(flags::FAIL_ACTIVATE_RESTORE)?,
|
||||
},
|
||||
client_id: client_id.clone(),
|
||||
service: service.clone(),
|
||||
|
|
@ -219,6 +234,8 @@ fn serve(role: &str, options: &Options) -> Result<(), String> {
|
|||
omit_audio_at_boundary: options.opt_u64(flags::OMIT_AUDIO_AT_BOUNDARY)?,
|
||||
overlapping_audio_at_boundary: options
|
||||
.opt_u64(flags::OVERLAPPING_AUDIO_AT_BOUNDARY)?,
|
||||
fail_stage_restore: options.flag(flags::FAIL_STAGE_RESTORE)?,
|
||||
fail_activate_restore: options.flag(flags::FAIL_ACTIVATE_RESTORE)?,
|
||||
},
|
||||
client_id: client_id.clone(),
|
||||
service: service.clone(),
|
||||
|
|
|
|||
|
|
@ -29,6 +29,30 @@ impl TickAccumulator {
|
|||
})
|
||||
}
|
||||
|
||||
/// The exact accumulator a capture recorded.
|
||||
///
|
||||
/// The remainder is restored, never rounded or reset: a resumed agent that started its
|
||||
/// first interval from zero would drift away from the run it is supposed to continue.
|
||||
pub fn restored(
|
||||
tick_duration: RationalNs,
|
||||
remainder: RationalNs,
|
||||
executed_ticks: u64,
|
||||
warmup_offset: u64,
|
||||
) -> Result<TickAccumulator, String> {
|
||||
let mut accumulator = TickAccumulator::new(tick_duration)?;
|
||||
remainder.validate().map_err(|e| e.0)?;
|
||||
if remainder >= tick_duration {
|
||||
return Err("a captured remainder is not below one model tick".to_owned());
|
||||
}
|
||||
if warmup_offset > executed_ticks {
|
||||
return Err("a captured warm-up offset exceeds the executed tick count".to_owned());
|
||||
}
|
||||
accumulator.remainder = remainder;
|
||||
accumulator.executed_ticks = executed_ticks;
|
||||
accumulator.warmup_offset = warmup_offset;
|
||||
Ok(accumulator)
|
||||
}
|
||||
|
||||
pub fn tick_duration(&self) -> RationalNs {
|
||||
self.tick_duration
|
||||
}
|
||||
|
|
|
|||
File diff suppressed because it is too large
Load diff
|
|
@ -13,6 +13,8 @@
|
|||
|
||||
use std::collections::BTreeSet;
|
||||
|
||||
use serde_json::Value;
|
||||
|
||||
use crate::media::{self, AudioSource, RenderCounter, ViewPipeline};
|
||||
use crate::task::{controller_schema_ref, inspection, inspection_schema};
|
||||
// `crate::types` is this crate's facade over the shared `fly-session-types` crate; the
|
||||
|
|
@ -48,6 +50,12 @@ pub struct EnvironmentFaults {
|
|||
pub omit_audio_at_boundary: Option<u64>,
|
||||
/// Emit an audio chunk that starts before the previous chunk ended.
|
||||
pub overlapping_audio_at_boundary: Option<u64>,
|
||||
/// Refuse `State.StageRestore`, so a group install meets a participant that will not
|
||||
/// validate.
|
||||
pub fail_stage_restore: bool,
|
||||
/// Refuse `State.ActivateRestore` after staging, so a group meets a failure halfway
|
||||
/// through activation.
|
||||
pub fail_activate_restore: bool,
|
||||
}
|
||||
|
||||
#[derive(Clone, Debug)]
|
||||
|
|
@ -86,6 +94,10 @@ pub struct CounterEnvironment {
|
|||
audio: Option<AudioSource>,
|
||||
/// The frame served at the previous boundary, kept only so a fault can serve it again.
|
||||
previous_view: Option<(ViewRef, flybus::Artifact)>,
|
||||
/// A validated replacement world the live session cannot see yet.
|
||||
staged: Option<StagedWorld>,
|
||||
/// Restore tokens this world has activated. A token activates once.
|
||||
activated: BTreeSet<Id>,
|
||||
}
|
||||
|
||||
impl CounterEnvironment {
|
||||
|
|
@ -104,10 +116,17 @@ impl CounterEnvironment {
|
|||
pipeline: None,
|
||||
audio: None,
|
||||
previous_view: None,
|
||||
staged: None,
|
||||
activated: BTreeSet::new(),
|
||||
config,
|
||||
}
|
||||
}
|
||||
|
||||
/// True while a validated replacement world is staged and not yet activated.
|
||||
pub fn has_staged_restore(&self) -> bool {
|
||||
self.staged.is_some()
|
||||
}
|
||||
|
||||
pub fn status(&self) -> StatusCell {
|
||||
self.status.clone()
|
||||
}
|
||||
|
|
@ -477,7 +496,7 @@ impl WorkerEndpoint for CounterEnvironment {
|
|||
vec![
|
||||
id("world-step-v1"),
|
||||
id("pixel-observation-v1"),
|
||||
id("checkpoint-v1"),
|
||||
id(crate::state::CHECKPOINT_CAPABILITY),
|
||||
]
|
||||
}
|
||||
|
||||
|
|
@ -490,7 +509,13 @@ impl WorkerEndpoint for CounterEnvironment {
|
|||
}
|
||||
|
||||
fn methods(&self) -> Vec<&'static str> {
|
||||
vec!["Environment.Initialize", "Environment.Advance"]
|
||||
vec![
|
||||
"Environment.Initialize",
|
||||
"Environment.Advance",
|
||||
"State.Capture",
|
||||
"State.StageRestore",
|
||||
"State.ActivateRestore",
|
||||
]
|
||||
}
|
||||
|
||||
fn handle<'a>(&'a mut self, ctx: HandlerCtx<'a>) -> BoxFuture<'a, DomainResult<HandlerReply>> {
|
||||
|
|
@ -498,6 +523,9 @@ impl WorkerEndpoint for CounterEnvironment {
|
|||
match ctx.method {
|
||||
"Environment.Initialize" => self.initialize(&ctx).await,
|
||||
"Environment.Advance" => self.advance(&ctx).await,
|
||||
"State.Capture" => self.state_capture(&ctx).await,
|
||||
"State.StageRestore" => self.state_stage_restore(&ctx).await,
|
||||
"State.ActivateRestore" => self.state_activate_restore(&ctx).await,
|
||||
other => Err(DomainError::before(
|
||||
ErrorCode::Unsupported,
|
||||
format!("{other} is not an environment method"),
|
||||
|
|
@ -516,3 +544,488 @@ pub fn synthetic_asset(asset_id: &str, body: &str) -> AssetRef {
|
|||
format: id("fly-config-v1"),
|
||||
}
|
||||
}
|
||||
|
||||
// -------------------------------------------------------------------------------------------
|
||||
// STATE-01: capture and restore
|
||||
|
||||
/// The version this payload layout is written and read under.
|
||||
pub const WORLD_PAYLOAD_VERSION: u64 = 1;
|
||||
|
||||
fn incompatible(message: impl std::fmt::Display) -> DomainError {
|
||||
DomainError::before(ErrorCode::IncompatibleState, message)
|
||||
}
|
||||
|
||||
/// One staged restore, held outside the live world until it is activated.
|
||||
struct StagedWorld {
|
||||
token: Id,
|
||||
checkpoint_id: Id,
|
||||
scope: Scope,
|
||||
episode_id: Id,
|
||||
descriptor: EnvironmentDescriptor,
|
||||
boundary: u64,
|
||||
counter: i64,
|
||||
world_time: RationalNs,
|
||||
advances: u64,
|
||||
frames: Vec<(u64, i64)>,
|
||||
audio_next_sample: u64,
|
||||
audio_phase: u64,
|
||||
audio_accumulator: u128,
|
||||
audio_denominator: u128,
|
||||
}
|
||||
|
||||
impl CounterEnvironment {
|
||||
/// `State.Capture`: the world at its committed boundary, including its pending sensor
|
||||
/// pipeline.
|
||||
///
|
||||
/// The pipeline is recorded as reconstruction inputs -- the producing boundary and the
|
||||
/// world counter of every retained frame -- and never as an artifact identity: a
|
||||
/// transient artifact belongs to the router that is running now, and a checkpoint outlives
|
||||
/// it.
|
||||
async fn state_capture(&mut self, ctx: &HandlerCtx<'_>) -> DomainResult<HandlerReply> {
|
||||
let scope = ctx.scope()?.clone();
|
||||
let Some(descriptor) = self.descriptor.clone() else {
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::InvalidPhase,
|
||||
"this environment is uninitialized",
|
||||
));
|
||||
};
|
||||
if scope.session_id != self.config.session_id {
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::IdentityMismatch,
|
||||
"this environment belongs to another session",
|
||||
));
|
||||
}
|
||||
match &self.epoch {
|
||||
Some(epoch) if *epoch == scope.epoch => {}
|
||||
_ => {
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::StaleEpoch,
|
||||
"State.Capture names an epoch this environment has left",
|
||||
));
|
||||
}
|
||||
}
|
||||
if scope.step != self.boundary {
|
||||
return Err(DomainError::before(
|
||||
if scope.step < self.boundary {
|
||||
ErrorCode::StaleStep
|
||||
} else {
|
||||
ErrorCode::FutureStep
|
||||
},
|
||||
"State.Capture must name the boundary the world is at",
|
||||
));
|
||||
}
|
||||
let params: CaptureParams = ctx.params()?;
|
||||
let pipeline = self
|
||||
.pipeline
|
||||
.as_ref()
|
||||
.ok_or_else(|| DomainError::before(ErrorCode::InvalidPhase, "no view pipeline"))?;
|
||||
let audio = self
|
||||
.audio
|
||||
.as_ref()
|
||||
.ok_or_else(|| DomainError::before(ErrorCode::InvalidPhase, "no audio source"))?;
|
||||
let (accumulator, denominator) = audio.accumulator();
|
||||
let previous = self.status.state();
|
||||
self.status.set_state(WorkerState::Capturing);
|
||||
let payload = serde_json::json!({
|
||||
"payloadVersion": WORLD_PAYLOAD_VERSION,
|
||||
"kind": "world",
|
||||
"workerId": self.config.worker_id.as_str(),
|
||||
"checkpointId": params.checkpoint_id.as_str(),
|
||||
"sourceScope": scope.to_json(),
|
||||
"episodeId": self.episode_id.clone().expect("initialized").as_str(),
|
||||
"committedStep": self.boundary.to_string(),
|
||||
"counter": self.counter.to_string(),
|
||||
"worldTime": self.world_time.to_json(),
|
||||
"advances": self.advances.to_string(),
|
||||
"descriptor": descriptor.to_json(),
|
||||
"pipeline": {
|
||||
// The declared delay's whole queue, oldest first.
|
||||
"frames": pipeline
|
||||
.retained()
|
||||
.into_iter()
|
||||
.map(|(boundary, counter)| serde_json::json!({
|
||||
"boundary": boundary.to_string(),
|
||||
"counter": counter.to_string(),
|
||||
}))
|
||||
.collect::<Vec<_>>(),
|
||||
},
|
||||
"audio": {
|
||||
"nextSample": audio.next_sample().to_string(),
|
||||
"phase": audio.phase().to_string(),
|
||||
"accumulator": accumulator.to_string(),
|
||||
"denominator": denominator.to_string(),
|
||||
"chunks": audio.chunks().to_string(),
|
||||
},
|
||||
});
|
||||
let bytes = canonicalize(&payload)
|
||||
.map_err(|e| DomainError::invalid(format!("State.Capture: {}", e.0)))?
|
||||
.into_bytes();
|
||||
let digest = digest_of_bytes(&bytes);
|
||||
let artifact = crate::state::seal_payload(ctx.client, &bytes, &digest).await?;
|
||||
// A capture reads the world; it does not advance it.
|
||||
self.status.set_state(previous);
|
||||
let result = CaptureResult {
|
||||
checkpoint_id: params.checkpoint_id,
|
||||
boundary: self.boundary,
|
||||
compatibility_digest: crate::state::Compatibility::of(&descriptor).digest(),
|
||||
payload: artifact.reference().clone(),
|
||||
};
|
||||
Ok(HandlerReply::with_artifacts(
|
||||
object(result.to_json()),
|
||||
vec![(crate::state::PAYLOAD_ATTACHMENT.to_owned(), artifact)],
|
||||
))
|
||||
}
|
||||
|
||||
/// `State.StageRestore`: validate a replacement world into a staging slot.
|
||||
async fn state_stage_restore(&mut self, ctx: &HandlerCtx<'_>) -> DomainResult<HandlerReply> {
|
||||
let scope = ctx.scope()?.clone();
|
||||
if scope.session_id != self.config.session_id {
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::IdentityMismatch,
|
||||
"this environment belongs to another session",
|
||||
));
|
||||
}
|
||||
if let Some(epoch) = &self.epoch
|
||||
&& *epoch == scope.epoch
|
||||
{
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::StaleEpoch,
|
||||
"State.StageRestore proposes the epoch this environment is already running",
|
||||
));
|
||||
}
|
||||
if self.descriptor.is_some() {
|
||||
// A world that is already running a boundary is not a quiescent replacement: the
|
||||
// group replaces it rather than restoring over a live one.
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::InvalidPhase,
|
||||
"State.StageRestore needs an uninitialized replacement environment",
|
||||
));
|
||||
}
|
||||
let params: StageRestoreParams = ctx.params()?;
|
||||
if params.source_scope.step != scope.step {
|
||||
return Err(DomainError::invalid(
|
||||
"State.StageRestore's scope step must be the source boundary",
|
||||
));
|
||||
}
|
||||
let artifact = ctx.artifact(crate::state::PAYLOAD_ATTACHMENT)?;
|
||||
if artifact.reference() != ¶ms.payload {
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::BufferInvalid,
|
||||
"the staged payload attachment is not the artifact the request names",
|
||||
));
|
||||
}
|
||||
let bytes = artifact.read_all().await.map_err(|e| {
|
||||
DomainError::before(
|
||||
ErrorCode::BufferInvalid,
|
||||
format!("the staged payload could not be read: {}", e.message),
|
||||
)
|
||||
})?;
|
||||
let declared = params
|
||||
.payload
|
||||
.digest
|
||||
.clone()
|
||||
.ok_or_else(|| incompatible("a checkpoint payload must carry a content digest"))?;
|
||||
let actual = digest_of_bytes(&bytes);
|
||||
if actual != declared || bytes.len() as u64 != params.payload.byte_length {
|
||||
return Err(incompatible(
|
||||
"the staged payload is not the content the request declares",
|
||||
));
|
||||
}
|
||||
let value: Value = serde_json::from_slice(&bytes)
|
||||
.map_err(|e| DomainError::invalid(format!("the staged payload is not JSON: {e}")))?;
|
||||
let text = |key: &str| -> DomainResult<String> {
|
||||
value
|
||||
.get(key)
|
||||
.and_then(Value::as_str)
|
||||
.map(str::to_owned)
|
||||
.ok_or_else(|| incompatible(format!("the world payload has no {key}")))
|
||||
};
|
||||
let number = |key: &str| -> DomainResult<u64> {
|
||||
text(key)?
|
||||
.parse::<u64>()
|
||||
.map_err(|_| incompatible(format!("the world payload's {key} is not a U64")))
|
||||
};
|
||||
if value.get("payloadVersion").and_then(Value::as_u64) != Some(WORLD_PAYLOAD_VERSION) {
|
||||
return Err(incompatible("the world payload is another payload version"));
|
||||
}
|
||||
if text("kind")? != "world" {
|
||||
return Err(incompatible("this payload is not a world's state"));
|
||||
}
|
||||
if text("workerId")? != self.config.worker_id {
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::IdentityMismatch,
|
||||
"the staged payload belongs to another world",
|
||||
));
|
||||
}
|
||||
if text("checkpointId")? != params.checkpoint_id {
|
||||
return Err(incompatible("the staged payload belongs to another checkpoint"));
|
||||
}
|
||||
let source_scope = Scope::from_json(
|
||||
value
|
||||
.get("sourceScope")
|
||||
.ok_or_else(|| incompatible("the world payload has no sourceScope"))?,
|
||||
)
|
||||
.map_err(|e| incompatible(format!("the world payload's sourceScope: {}", e.0)))?;
|
||||
if source_scope != params.source_scope {
|
||||
return Err(incompatible(
|
||||
"the staged payload was captured at another source scope",
|
||||
));
|
||||
}
|
||||
let committed_step = number("committedStep")?;
|
||||
if committed_step != params.source_scope.step {
|
||||
return Err(incompatible(
|
||||
"the staged payload's committed step is not the source boundary",
|
||||
));
|
||||
}
|
||||
let descriptor = EnvironmentDescriptor::from_json(
|
||||
value
|
||||
.get("descriptor")
|
||||
.ok_or_else(|| incompatible("the world payload has no descriptor"))?,
|
||||
)
|
||||
.map_err(|e| incompatible(format!("the world payload's descriptor: {}", e.0)))?;
|
||||
// The replacement builds the descriptor it would advertise and compares. A world
|
||||
// started with other ports, another cadence or another declared render delay is a
|
||||
// different backend, not this one resumed.
|
||||
let live = self.build_descriptor()?;
|
||||
if descriptor != live {
|
||||
return Err(incompatible(
|
||||
"the staged world was captured under another environment descriptor",
|
||||
));
|
||||
}
|
||||
let expected = crate::state::Compatibility::of(&descriptor).digest();
|
||||
if expected != params.compatibility_digest {
|
||||
return Err(incompatible(format!(
|
||||
"the staged world's compatibility {expected} is not the {} the restore requires",
|
||||
params.compatibility_digest
|
||||
)));
|
||||
}
|
||||
let counter: i64 = text("counter")?
|
||||
.parse()
|
||||
.map_err(|_| incompatible("the world payload's counter is not an integer"))?;
|
||||
let world_time = RationalNs::from_json(
|
||||
value
|
||||
.get("worldTime")
|
||||
.ok_or_else(|| incompatible("the world payload has no worldTime"))?,
|
||||
)
|
||||
.map_err(|e| incompatible(format!("the world payload's worldTime: {}", e.0)))?;
|
||||
let pipeline_value = value
|
||||
.get("pipeline")
|
||||
.and_then(|p| p.get("frames"))
|
||||
.and_then(Value::as_array)
|
||||
.ok_or_else(|| incompatible("the world payload has no pipeline frames"))?;
|
||||
let mut frames = Vec::with_capacity(pipeline_value.len());
|
||||
for frame in pipeline_value {
|
||||
let boundary = frame
|
||||
.get("boundary")
|
||||
.and_then(Value::as_str)
|
||||
.ok_or_else(|| incompatible("a captured frame has no boundary"))?
|
||||
.parse::<u64>()
|
||||
.map_err(|_| incompatible("a captured frame's boundary is not a U64"))?;
|
||||
let frame_counter = frame
|
||||
.get("counter")
|
||||
.and_then(Value::as_str)
|
||||
.ok_or_else(|| incompatible("a captured frame has no counter"))?
|
||||
.parse::<i64>()
|
||||
.map_err(|_| incompatible("a captured frame's counter is not an integer"))?;
|
||||
frames.push((boundary, frame_counter));
|
||||
}
|
||||
match frames.last() {
|
||||
Some((boundary, _)) if *boundary == committed_step => {}
|
||||
_ => {
|
||||
return Err(incompatible(
|
||||
"the captured pipeline does not end at the committed boundary",
|
||||
));
|
||||
}
|
||||
}
|
||||
let audio_value = value
|
||||
.get("audio")
|
||||
.ok_or_else(|| incompatible("the world payload has no audio state"))?;
|
||||
let audio_number = |key: &str| -> DomainResult<u128> {
|
||||
audio_value
|
||||
.get(key)
|
||||
.and_then(Value::as_str)
|
||||
.ok_or_else(|| incompatible(format!("the captured audio state has no {key}")))?
|
||||
.parse::<u128>()
|
||||
.map_err(|_| incompatible(format!("the captured audio {key} is not a number")))
|
||||
};
|
||||
let audio_next_sample = u64::try_from(audio_number("nextSample")?)
|
||||
.map_err(|_| incompatible("the captured audio position is outside U64"))?;
|
||||
let audio_phase = u64::try_from(audio_number("phase")?)
|
||||
.map_err(|_| incompatible("the captured audio phase is outside U64"))?;
|
||||
|
||||
if self.config.faults.fail_stage_restore {
|
||||
return Err(incompatible(
|
||||
"injected staging refusal: this participant's replacement state does not \
|
||||
validate",
|
||||
));
|
||||
}
|
||||
if let Some(staged) = &self.staged {
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::Conflict,
|
||||
format!(
|
||||
"this environment already holds the staged restore {} for checkpoint {}",
|
||||
staged.token, staged.checkpoint_id
|
||||
),
|
||||
));
|
||||
}
|
||||
let token = crate::agent::restore_token(
|
||||
¶ms.checkpoint_id,
|
||||
&scope,
|
||||
&actual,
|
||||
&self.config.incarnation_id,
|
||||
);
|
||||
if self.activated.contains(&token) {
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::Conflict,
|
||||
"this exact restore was already activated on this environment",
|
||||
));
|
||||
}
|
||||
self.staged = Some(StagedWorld {
|
||||
token: token.clone(),
|
||||
checkpoint_id: params.checkpoint_id.clone(),
|
||||
scope,
|
||||
episode_id: parse_id(&text("episodeId")?)
|
||||
.map_err(|e| incompatible(format!("the world payload's episodeId {e}")))?,
|
||||
descriptor,
|
||||
boundary: committed_step,
|
||||
counter,
|
||||
world_time,
|
||||
advances: number("advances")?,
|
||||
frames,
|
||||
audio_next_sample,
|
||||
audio_phase,
|
||||
audio_accumulator: audio_number("accumulator")?,
|
||||
audio_denominator: audio_number("denominator")?,
|
||||
});
|
||||
self.status.set_state(WorkerState::StagedRestore);
|
||||
let result = StageRestoreResult {
|
||||
checkpoint_id: params.checkpoint_id,
|
||||
restore_token: token,
|
||||
};
|
||||
Ok(HandlerReply::from(&result))
|
||||
}
|
||||
|
||||
/// `State.ActivateRestore`: install the staged world and return its coherent observation.
|
||||
///
|
||||
/// Nothing advances. The pipeline's frames are rendered again into fresh artifacts of the
|
||||
/// current store, which is what "the durable store imports fresh immutable bus artifacts"
|
||||
/// means on the producing side, and the observation carries no audio chunk because no
|
||||
/// interval was played.
|
||||
async fn state_activate_restore(
|
||||
&mut self,
|
||||
ctx: &HandlerCtx<'_>,
|
||||
) -> DomainResult<HandlerReply> {
|
||||
let params: ActivateRestoreParams = ctx.params()?;
|
||||
if self.activated.contains(¶ms.restore_token) {
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::Conflict,
|
||||
"this restore token has already been activated",
|
||||
));
|
||||
}
|
||||
let Some(staged) = self.staged.take() else {
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::InvalidPhase,
|
||||
"this environment holds no staged restore",
|
||||
));
|
||||
};
|
||||
if staged.token != params.restore_token {
|
||||
let token = staged.token.clone();
|
||||
self.staged = Some(staged);
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::IdentityMismatch,
|
||||
format!(
|
||||
"this environment's staged restore is {token}, not {}",
|
||||
params.restore_token
|
||||
),
|
||||
));
|
||||
}
|
||||
if self.config.faults.fail_activate_restore {
|
||||
let token = staged.token.clone();
|
||||
self.staged = Some(staged);
|
||||
self.status.set_state(WorkerState::Failed);
|
||||
return Err(DomainError::new(
|
||||
ErrorCode::BackendFailure,
|
||||
format!("injected activation failure; {token} stays staged and unresumed"),
|
||||
MutationCertainty::None,
|
||||
));
|
||||
}
|
||||
self.status.set_state(WorkerState::Restoring);
|
||||
let mut pipeline = ViewPipeline::new(
|
||||
CounterEnvironment::view_descriptor(self.config.observation_delay_steps),
|
||||
self.config.renders.clone(),
|
||||
);
|
||||
pipeline.restore(ctx.client, &staged.frames).await?;
|
||||
let audio = AudioSource::restored_from(
|
||||
CounterEnvironment::audio_descriptor(),
|
||||
staged.audio_next_sample,
|
||||
staged.audio_phase,
|
||||
staged.audio_accumulator,
|
||||
staged.audio_denominator,
|
||||
)?;
|
||||
self.epoch = Some(staged.scope.epoch.clone());
|
||||
self.episode_id = Some(staged.episode_id.clone());
|
||||
self.descriptor = Some(staged.descriptor.clone());
|
||||
self.boundary = staged.boundary;
|
||||
self.counter = staged.counter;
|
||||
self.world_time = staged.world_time;
|
||||
self.advances = staged.advances;
|
||||
// Batch ids are unique within an epoch, and this is a new one. Keeping the old set
|
||||
// would refuse nothing extra: a request under the old epoch is already refused by its
|
||||
// scope.
|
||||
self.batches.clear();
|
||||
self.pipeline = Some(pipeline);
|
||||
self.audio = Some(audio);
|
||||
self.previous_view = None;
|
||||
self.activated.insert(staged.token);
|
||||
self.status.set_state(WorkerState::Ready);
|
||||
self.status.set_scope(Some(scope_at(
|
||||
&staged.scope.session_id,
|
||||
&staged.scope.epoch,
|
||||
staged.boundary,
|
||||
)));
|
||||
|
||||
let (observation, attachments) = self.restored_observation()?;
|
||||
let result = ActivateRestoreResult {
|
||||
committed_step: staged.boundary,
|
||||
checkpoint_id: staged.checkpoint_id,
|
||||
observation: Some(observation),
|
||||
};
|
||||
result
|
||||
.validate_for_role(Role::Environment)
|
||||
.map_err(|e| DomainError::invalid(e.0))?;
|
||||
let mut reply = HandlerReply::from(&result);
|
||||
reply.artifacts = attachments;
|
||||
Ok(reply)
|
||||
}
|
||||
|
||||
/// The observation the restored world is already at: no render, no advance, no audio.
|
||||
fn restored_observation(
|
||||
&mut self,
|
||||
) -> DomainResult<(WorldObservation, Vec<(String, flybus::Artifact)>)> {
|
||||
let boundary = self.boundary;
|
||||
let counter = self.counter;
|
||||
let pipeline = self
|
||||
.pipeline
|
||||
.as_ref()
|
||||
.ok_or_else(|| DomainError::before(ErrorCode::InvalidPhase, "no view pipeline"))?;
|
||||
let (view, artifact) = pipeline.at(boundary).ok_or_else(|| {
|
||||
incompatible("the restored pipeline holds no frame for the restored boundary")
|
||||
})?;
|
||||
self.previous_view = Some((view.clone(), artifact.clone()));
|
||||
let observation = WorldObservation {
|
||||
boundary,
|
||||
world_time: self.world_time,
|
||||
engine_frame: Some(boundary.to_string()),
|
||||
sensory_views: vec![view.clone()],
|
||||
inspection: inspection(counter, boundary),
|
||||
broadcast_views: vec![view.clone()],
|
||||
// No interval was played, so there is no chunk. A chunk here would be an old
|
||||
// epoch's audio offered as current.
|
||||
audio: Vec::new(),
|
||||
};
|
||||
Ok((
|
||||
observation,
|
||||
vec![(media::view_attachment(&view.view_id), artifact)],
|
||||
))
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -26,6 +26,7 @@ use crate::media::{RenderCounter, SensorLog};
|
|||
use crate::launcher::{
|
||||
AgentLaunch, EnvironmentLaunch, Launcher, ReapOutcome, SUPERVISOR_CLIENT, ThreadBudget,
|
||||
};
|
||||
use crate::state::{CheckpointStore, CheckpointWriter, StoreConfig, StoreFaults, WriterConfig, WriterFaults};
|
||||
use crate::task::{ActionExecutor, CounterTask, IdentityExecutor, Terminal};
|
||||
// `crate::types` is this crate's facade over the shared `fly-session-types` crate; the
|
||||
// glob keeps the contract's own names in sight instead of restating them.
|
||||
|
|
@ -85,6 +86,14 @@ pub struct HarnessConfig {
|
|||
/// The threads reserved for the coordinator, its router and its store.
|
||||
pub coordinator_threads: usize,
|
||||
pub environment_threads: usize,
|
||||
/// How many committed generations the durable checkpoint store keeps.
|
||||
pub store: StoreConfig,
|
||||
/// The durable write faults this composition injects.
|
||||
pub store_faults: StoreFaults,
|
||||
/// The checkpoint queue's bounds.
|
||||
pub writer: WriterConfig,
|
||||
/// The writer faults this composition injects.
|
||||
pub writer_faults: WriterFaults,
|
||||
}
|
||||
|
||||
impl Default for HarnessConfig {
|
||||
|
|
@ -107,6 +116,10 @@ impl Default for HarnessConfig {
|
|||
thread_budget: None,
|
||||
coordinator_threads: 1,
|
||||
environment_threads: 1,
|
||||
store: StoreConfig::default(),
|
||||
store_faults: StoreFaults::default(),
|
||||
writer: WriterConfig::default(),
|
||||
writer_faults: WriterFaults::default(),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
|
@ -134,6 +147,14 @@ const ENV_SERVICE: &str = "env.arena";
|
|||
const ENV_CLIENT: &str = "environment";
|
||||
const ENV_WORKER: &str = "arena";
|
||||
const COORDINATOR_CLIENT: &str = "coordinator";
|
||||
/// The checkpoint writer's own bus identity. It publishes checkpoint events and nothing else.
|
||||
const WRITER_CLIENT: &str = "checkpoint-writer";
|
||||
|
||||
/// How many times one participant may be replaced in a composition.
|
||||
///
|
||||
/// Each replacement connects under its own client id, so a restart is visibly a new
|
||||
/// participant rather than a silent reattachment, and the policy has to name them all.
|
||||
const MAX_GENERATIONS: u32 = 8;
|
||||
|
||||
fn agent_service(agent_id: &Id) -> String {
|
||||
format!("agent.{agent_id}")
|
||||
|
|
@ -172,6 +193,10 @@ pub struct SessionHarness {
|
|||
/// The supervisor. It owns every participant's lifetime and thread allocation.
|
||||
pub launcher: Launcher,
|
||||
observers: Mutex<Vec<Client>>,
|
||||
/// Which generation of each participant is running: 1 is the one the composition started.
|
||||
generations: BTreeMap<Id, u32>,
|
||||
/// Where the durable checkpoint store lives, for a test that reads the files themselves.
|
||||
checkpoint_root: std::path::PathBuf,
|
||||
}
|
||||
|
||||
impl SessionHarness {
|
||||
|
|
@ -204,12 +229,16 @@ impl SessionHarness {
|
|||
g.call = vec![Pattern::prefix("agent."), Pattern::prefix("env.")];
|
||||
}),
|
||||
)
|
||||
// The writer publishes the checkpoint events and never calls a participant.
|
||||
.client(WRITER_CLIENT, grants(|g| g.publish = vec![Pattern::prefix("session.")]))
|
||||
.client(ENV_CLIENT, grants(|g| g.register = vec![Pattern::exact(ENV_SERVICE)]))
|
||||
.client(
|
||||
&format!("{ENV_CLIENT}-r2"),
|
||||
grants(|g| g.register = vec![Pattern::exact(ENV_SERVICE)]),
|
||||
)
|
||||
.client("observer", grants(|g| g.subscribe = vec![Pattern::prefix("session.")]));
|
||||
for generation in 2..=MAX_GENERATIONS {
|
||||
policy = policy.client(
|
||||
&format!("{ENV_CLIENT}-r{generation}"),
|
||||
grants(|g| g.register = vec![Pattern::exact(ENV_SERVICE)]),
|
||||
);
|
||||
}
|
||||
for spec in &config.agents {
|
||||
let service = agent_service(&spec.agent_id);
|
||||
policy = policy.client(
|
||||
|
|
@ -218,10 +247,12 @@ impl SessionHarness {
|
|||
);
|
||||
// A replacement worker connects under its own client id, so a restart is visibly a
|
||||
// new participant rather than a silent reattachment to the active epoch.
|
||||
policy = policy.client(
|
||||
&format!("{}-r2", agent_client(&spec.agent_id)),
|
||||
grants(|g| g.register = vec![Pattern::exact(&service)]),
|
||||
);
|
||||
for generation in 2..=MAX_GENERATIONS {
|
||||
policy = policy.client(
|
||||
&format!("{}-r{generation}", agent_client(&spec.agent_id)),
|
||||
grants(|g| g.register = vec![Pattern::exact(&service)]),
|
||||
);
|
||||
}
|
||||
}
|
||||
let mut router_config = RouterConfig::new(&store_root);
|
||||
router_config.policy = policy;
|
||||
|
|
@ -299,6 +330,21 @@ impl SessionHarness {
|
|||
}
|
||||
|
||||
let coordinator_client = launcher.connect(COORDINATOR_CLIENT).await?;
|
||||
// The durable store lives beside the router's artifact store and never inside it: a
|
||||
// committed generation is outside the bus's ephemeral collection.
|
||||
let checkpoint_root = root.join("checkpoints");
|
||||
let mut store = CheckpointStore::open(&checkpoint_root, config.store).map_err(refusal)?;
|
||||
*store.faults_mut() = config.store_faults.clone();
|
||||
let writer_client = launcher.connect(WRITER_CLIENT).await?;
|
||||
let writer = CheckpointWriter::start(
|
||||
store,
|
||||
config.writer,
|
||||
config.writer_faults.clone(),
|
||||
Some((
|
||||
writer_client,
|
||||
format!("session.{}.checkpoints", config.session_id),
|
||||
)),
|
||||
);
|
||||
let executors: BTreeMap<Id, Box<dyn ActionExecutor>> = config
|
||||
.agents
|
||||
.iter()
|
||||
|
|
@ -316,6 +362,8 @@ impl SessionHarness {
|
|||
Box::new(CounterTask::new(&config.epoch, config.terminal)),
|
||||
executors,
|
||||
);
|
||||
let mut coordinator = coordinator;
|
||||
coordinator.attach_store(writer);
|
||||
|
||||
Ok(SessionHarness {
|
||||
coordinator,
|
||||
|
|
@ -326,9 +374,16 @@ impl SessionHarness {
|
|||
sensors,
|
||||
launcher,
|
||||
observers: Mutex::new(Vec::new()),
|
||||
generations: BTreeMap::new(),
|
||||
checkpoint_root,
|
||||
})
|
||||
}
|
||||
|
||||
/// Where the durable checkpoint store's generations and store manifest live.
|
||||
pub fn checkpoint_root(&self) -> &std::path::Path {
|
||||
&self.checkpoint_root
|
||||
}
|
||||
|
||||
pub fn router(&self) -> &Router {
|
||||
self.launcher.router()
|
||||
}
|
||||
|
|
@ -373,10 +428,11 @@ impl SessionHarness {
|
|||
.find(|spec| spec.agent_id == *agent_id)
|
||||
.expect("a configured agent")
|
||||
.clone();
|
||||
let generation = self.next_generation(agent_id)?;
|
||||
self.launcher.kill(agent_id).await;
|
||||
let tick_duration = millis(self.config.tick_ms).expect("a positive tick");
|
||||
let incarnation_id =
|
||||
parse_id(&format!("{agent_id}-inc-2")).expect("an agent id plus a suffix is an Id");
|
||||
let incarnation_id = parse_id(&format!("{agent_id}-inc-{generation}"))
|
||||
.expect("an agent id plus a suffix is an Id");
|
||||
self.launcher
|
||||
.launch_agent(AgentLaunch {
|
||||
session_id: self.config.session_id.clone(),
|
||||
|
|
@ -390,7 +446,7 @@ impl SessionHarness {
|
|||
// predecessor wrote, so a restore's sensory input is visible beside it.
|
||||
sensors: self.sensors.get(agent_id).cloned().unwrap_or_default(),
|
||||
faults: spec.faults.clone(),
|
||||
client_id: format!("{}-r2", agent_client(agent_id)),
|
||||
client_id: format!("{}-r{generation}", agent_client(agent_id)),
|
||||
service: agent_service(agent_id),
|
||||
})
|
||||
.await
|
||||
|
|
@ -403,6 +459,102 @@ impl SessionHarness {
|
|||
})
|
||||
}
|
||||
|
||||
/// Replaces the environment with a fresh, uninitialized incarnation, as a restore needs.
|
||||
pub async fn restart_environment(&mut self) -> Result<Restarted, flybus::BusError> {
|
||||
let worker_id = id(ENV_WORKER);
|
||||
let generation = self.next_generation(&worker_id)?;
|
||||
self.launcher.kill(&worker_id).await;
|
||||
let step_duration = hz(self.config.step_hz).expect("a positive cadence");
|
||||
let incarnation_id = parse_id(&format!("arena-inc-{generation}"))
|
||||
.expect("a worker id plus a suffix is an Id");
|
||||
self.launcher
|
||||
.launch_environment(EnvironmentLaunch {
|
||||
session_id: self.config.session_id.clone(),
|
||||
worker_id: worker_id.clone(),
|
||||
incarnation_id: incarnation_id.clone(),
|
||||
step_duration,
|
||||
ports: self.config.agents.iter().map(|a| a.port_id.clone()).collect(),
|
||||
worker_threads: self.config.environment_threads,
|
||||
observation_delay_steps: self.config.observation_delay_steps,
|
||||
renders: self.renders.clone(),
|
||||
faults: self.config.environment_faults.clone(),
|
||||
client_id: format!("{ENV_CLIENT}-r{generation}"),
|
||||
service: ENV_SERVICE.to_owned(),
|
||||
})
|
||||
.await
|
||||
.map_err(refusal)?;
|
||||
let worker = self.launcher.worker(&worker_id).expect("just launched");
|
||||
Ok(Restarted {
|
||||
service: worker.identity.service.clone(),
|
||||
service_incarnation: worker.service_incarnation.clone(),
|
||||
incarnation_id,
|
||||
})
|
||||
}
|
||||
|
||||
fn next_generation(&mut self, worker_id: &Id) -> Result<u32, flybus::BusError> {
|
||||
let slot = self.generations.entry(worker_id.clone()).or_insert(1);
|
||||
if *slot >= MAX_GENERATIONS {
|
||||
return Err(flybus::BusError::new(
|
||||
flybus::ErrorCode::QuotaExceeded,
|
||||
format!(
|
||||
"{worker_id} has used all {MAX_GENERATIONS} configured client identities; a composition declares how many replacements it allows"
|
||||
),
|
||||
));
|
||||
}
|
||||
*slot += 1;
|
||||
Ok(*slot)
|
||||
}
|
||||
|
||||
/// Replaces every participant and points the fenced coordinator at the replacements.
|
||||
///
|
||||
/// This is what a recovery does before it restores: the old participants belong to an
|
||||
/// invalid epoch, and the references the coordinator pinned are exchanged deliberately.
|
||||
pub async fn replace_all_participants(&mut self) -> Result<(), flybus::BusError> {
|
||||
let environment = self.environment_id();
|
||||
self.restart_environment().await?;
|
||||
let worker = self
|
||||
.launcher
|
||||
.worker(&environment)
|
||||
.expect("just launched")
|
||||
.worker_ref();
|
||||
self.coordinator
|
||||
.replace_participant(&environment, worker)
|
||||
.map_err(|e| refusal(e.error))?;
|
||||
for agent_id in self.config.agents.iter().map(|a| a.agent_id.clone()).collect::<Vec<_>>() {
|
||||
self.restart_agent(&agent_id).await?;
|
||||
let worker = self
|
||||
.launcher
|
||||
.worker(&agent_id)
|
||||
.expect("just launched")
|
||||
.worker_ref();
|
||||
self.coordinator
|
||||
.replace_participant(&agent_id, worker)
|
||||
.map_err(|e| refusal(e.error))?;
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Changes one agent's injected faults, so the replacement the next restart launches is
|
||||
/// a participant without them.
|
||||
///
|
||||
/// A fault is launch configuration, so clearing one is a relaunch and not a live change:
|
||||
/// the worker running now keeps whatever it was started with.
|
||||
pub fn set_agent_faults(&mut self, agent_id: &Id, faults: AgentFaults) {
|
||||
if let Some(spec) = self
|
||||
.config
|
||||
.agents
|
||||
.iter_mut()
|
||||
.find(|spec| spec.agent_id == *agent_id)
|
||||
{
|
||||
spec.faults = faults;
|
||||
}
|
||||
}
|
||||
|
||||
/// Changes the environment's injected faults, with the same relaunch rule.
|
||||
pub fn set_environment_faults(&mut self, faults: EnvironmentFaults) {
|
||||
self.config.environment_faults = faults;
|
||||
}
|
||||
|
||||
/// Ends one participant without asking it, as a crash would.
|
||||
pub async fn kill(&mut self, worker_id: &Id) -> ReapOutcome {
|
||||
self.launcher.kill(worker_id).await
|
||||
|
|
@ -466,7 +618,10 @@ impl SessionHarness {
|
|||
|
||||
/// Reaps every participant and closes the router.
|
||||
pub async fn shutdown(self) {
|
||||
let SessionHarness { coordinator, mut launcher, observers, .. } = self;
|
||||
let SessionHarness { mut coordinator, mut launcher, observers, .. } = self;
|
||||
// The writer task owns artifact handles and a blocking store. Leaving it running
|
||||
// would leave both behind.
|
||||
coordinator.shutdown_store().await;
|
||||
drop(coordinator);
|
||||
launcher.reap_all(&id("shutdown")).await;
|
||||
for observer in observers.into_inner().expect("not poisoned") {
|
||||
|
|
|
|||
|
|
@ -1231,6 +1231,8 @@ pub(crate) mod flags {
|
|||
pub const PREPARE_DELAY_MS: &str = "prepare-delay-ms";
|
||||
pub const COMMIT_DELAY_MS: &str = "commit-delay-ms";
|
||||
pub const FAIL_COMMIT_AT_STEP: &str = "fail-commit-at-step";
|
||||
pub const FAIL_STAGE_RESTORE: &str = "fail-stage-restore";
|
||||
pub const FAIL_ACTIVATE_RESTORE: &str = "fail-activate-restore";
|
||||
|
||||
pub const WORKER: &str = "worker";
|
||||
pub const PORTS: &str = "ports";
|
||||
|
|
@ -1271,6 +1273,8 @@ pub(crate) mod flags {
|
|||
PREPARE_DELAY_MS,
|
||||
COMMIT_DELAY_MS,
|
||||
FAIL_COMMIT_AT_STEP,
|
||||
FAIL_STAGE_RESTORE,
|
||||
FAIL_ACTIVATE_RESTORE,
|
||||
];
|
||||
/// What only the environment is given, media options included.
|
||||
pub const ENVIRONMENT_ONLY: &[&str] = &[
|
||||
|
|
@ -1285,6 +1289,8 @@ pub(crate) mod flags {
|
|||
TRUNCATED_VIEW_AT_BOUNDARY,
|
||||
OMIT_AUDIO_AT_BOUNDARY,
|
||||
OVERLAPPING_AUDIO_AT_BOUNDARY,
|
||||
FAIL_STAGE_RESTORE,
|
||||
FAIL_ACTIVATE_RESTORE,
|
||||
];
|
||||
/// What a measurement run or one of its row children is given.
|
||||
pub const MEASURE: &[&str] = &[MODE, AGENTS, STEPS, WARMUP_STEPS, WORKER_THREADS, MODES];
|
||||
|
|
@ -1327,6 +1333,11 @@ impl Started {
|
|||
arg(flags::WARMUP_TICKS, spec.warmup_ticks),
|
||||
arg(flags::PREPARE_DELAY_MS, spec.faults.prepare_delay_ms),
|
||||
arg(flags::COMMIT_DELAY_MS, spec.faults.commit_delay_ms),
|
||||
arg(flags::FAIL_STAGE_RESTORE, u64::from(spec.faults.fail_stage_restore)),
|
||||
arg(
|
||||
flags::FAIL_ACTIVATE_RESTORE,
|
||||
u64::from(spec.faults.fail_activate_restore),
|
||||
),
|
||||
];
|
||||
if let Some(step) = spec.faults.fail_commit_at_step {
|
||||
args.push(arg(flags::FAIL_COMMIT_AT_STEP, step));
|
||||
|
|
@ -1345,6 +1356,11 @@ impl Started {
|
|||
// The media options a world in another process needs to be exactly this
|
||||
// world. Its render counter and its agents' sensor logs stay there.
|
||||
arg(flags::OBSERVATION_DELAY_STEPS, spec.observation_delay_steps),
|
||||
arg(flags::FAIL_STAGE_RESTORE, u64::from(spec.faults.fail_stage_restore)),
|
||||
arg(
|
||||
flags::FAIL_ACTIVATE_RESTORE,
|
||||
u64::from(spec.faults.fail_activate_restore),
|
||||
),
|
||||
];
|
||||
for (flag, boundary) in [
|
||||
(flags::OMIT_VIEW_AT_BOUNDARY, spec.faults.omit_view_at_boundary),
|
||||
|
|
@ -1458,6 +1474,8 @@ mod flag_tests {
|
|||
truncated_view_at_boundary: Some(3),
|
||||
omit_audio_at_boundary: Some(4),
|
||||
overlapping_audio_at_boundary: Some(5),
|
||||
fail_stage_restore: true,
|
||||
fail_activate_restore: true,
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -1491,6 +1509,8 @@ mod flag_tests {
|
|||
fail_commit_at_step: Some(2),
|
||||
prepare_delay_ms: 1,
|
||||
commit_delay_ms: 2,
|
||||
fail_stage_restore: true,
|
||||
fail_activate_restore: true,
|
||||
},
|
||||
client_id: "worker-fly-a".to_owned(),
|
||||
service: "agent.fly-a".to_owned(),
|
||||
|
|
@ -1535,6 +1555,8 @@ mod flag_tests {
|
|||
flags::TRUNCATED_VIEW_AT_BOUNDARY,
|
||||
flags::OMIT_AUDIO_AT_BOUNDARY,
|
||||
flags::OVERLAPPING_AUDIO_AT_BOUNDARY,
|
||||
flags::FAIL_STAGE_RESTORE,
|
||||
flags::FAIL_ACTIVATE_RESTORE,
|
||||
] {
|
||||
assert!(written.contains(&format!("--{flag}")), "--{flag} is not written");
|
||||
}
|
||||
|
|
|
|||
|
|
@ -34,6 +34,7 @@ pub mod media;
|
|||
pub mod metrics;
|
||||
pub mod phase;
|
||||
pub mod rpc;
|
||||
pub mod state;
|
||||
pub mod task;
|
||||
pub mod worker;
|
||||
|
||||
|
|
|
|||
|
|
@ -133,7 +133,14 @@ pub fn arena_frame(descriptor: &ViewDescriptor, counter: i64, boundary: u64) ->
|
|||
/// cannot be served an arbitrary stale image.
|
||||
pub struct ViewPipeline {
|
||||
descriptor: ViewDescriptor,
|
||||
frames: VecDeque<(u64, flybus::Artifact)>,
|
||||
/// Each retained frame: its producing boundary, the world counter it was rendered from
|
||||
/// and the owned handle on its immutable bytes.
|
||||
///
|
||||
/// The counter is kept because it is the whole of the reconstruction input: a checkpoint
|
||||
/// records `(boundary, counter)` per retained frame and a restore re-renders them into
|
||||
/// fresh artifacts of the current store, rather than persisting a transient artifact
|
||||
/// identity that cannot survive a router restart.
|
||||
frames: VecDeque<(u64, i64, flybus::Artifact)>,
|
||||
renders: RenderCounter,
|
||||
}
|
||||
|
||||
|
|
@ -160,7 +167,7 @@ impl ViewPipeline {
|
|||
let bytes = arena_frame(&self.descriptor, counter, boundary);
|
||||
let artifact = seal(client, FRAME_CONTENT_TYPE, &bytes).await?;
|
||||
self.renders.bump();
|
||||
self.frames.push_back((boundary, artifact));
|
||||
self.frames.push_back((boundary, counter, artifact));
|
||||
// Keep exactly the frames a declared delay can still require.
|
||||
while self.frames.len() > self.descriptor.observation_delay_steps as usize + 1 {
|
||||
self.frames.pop_front();
|
||||
|
|
@ -168,6 +175,50 @@ impl ViewPipeline {
|
|||
Ok(())
|
||||
}
|
||||
|
||||
/// The reconstruction inputs of every retained frame, oldest first.
|
||||
///
|
||||
/// This is what a checkpoint records for the pending sensor pipeline: the producing
|
||||
/// boundary and the world counter, never an artifact identity.
|
||||
pub fn retained(&self) -> Vec<(u64, i64)> {
|
||||
self.frames
|
||||
.iter()
|
||||
.map(|(boundary, counter, _)| (*boundary, *counter))
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// Rebuilds the pipeline from recorded reconstruction inputs, into fresh artifacts.
|
||||
///
|
||||
/// Every frame is rendered again in the current store, so nothing a fence dropped is
|
||||
/// expected to come back and no old artifact identity crosses the recovery.
|
||||
pub async fn restore(
|
||||
&mut self,
|
||||
client: &flybus::Client,
|
||||
frames: &[(u64, i64)],
|
||||
) -> DomainResult<()> {
|
||||
if frames.len() > self.descriptor.observation_delay_steps as usize + 1 {
|
||||
return Err(media_error(format!(
|
||||
"a captured pipeline of {} frames does not fit a declared delay of {}",
|
||||
frames.len(),
|
||||
self.descriptor.observation_delay_steps
|
||||
)));
|
||||
}
|
||||
for window in frames.windows(2) {
|
||||
if window[1].0 != window[0].0 + 1 {
|
||||
return Err(media_error(
|
||||
"a captured pipeline's producing boundaries are not consecutive",
|
||||
));
|
||||
}
|
||||
}
|
||||
self.frames.clear();
|
||||
for (boundary, counter) in frames {
|
||||
let bytes = arena_frame(&self.descriptor, *counter, *boundary);
|
||||
let artifact = seal(client, FRAME_CONTENT_TYPE, &bytes).await?;
|
||||
self.renders.bump();
|
||||
self.frames.push_back((*boundary, *counter, artifact));
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Seals a frame of the wrong length, which is what a broken backend produces. The
|
||||
/// reference it returns describes the artifact honestly, so the shape check is the thing
|
||||
/// under test rather than a lie in the payload.
|
||||
|
|
@ -181,7 +232,7 @@ impl ViewPipeline {
|
|||
bytes.truncate(bytes.len() - self.descriptor.row_stride as usize);
|
||||
let artifact = seal(client, FRAME_CONTENT_TYPE, &bytes).await?;
|
||||
self.renders.bump();
|
||||
self.frames.push_back((boundary, artifact));
|
||||
self.frames.push_back((boundary, counter, artifact));
|
||||
while self.frames.len() > self.descriptor.observation_delay_steps as usize + 2 {
|
||||
self.frames.pop_front();
|
||||
}
|
||||
|
|
@ -198,8 +249,8 @@ impl ViewPipeline {
|
|||
pub fn frame_produced_at(&self, produced: u64) -> Option<(ViewRef, flybus::Artifact)> {
|
||||
self.frames
|
||||
.iter()
|
||||
.find(|(step, _)| *step == produced)
|
||||
.map(|(step, artifact)| {
|
||||
.find(|(step, _, _)| *step == produced)
|
||||
.map(|(step, _, artifact)| {
|
||||
(
|
||||
ViewRef {
|
||||
view_id: self.descriptor.view_id.clone(),
|
||||
|
|
@ -257,6 +308,52 @@ impl AudioSource {
|
|||
source
|
||||
}
|
||||
|
||||
/// The exact state a capture recorded: sample position, waveform phase and the
|
||||
/// unconsumed fraction of a frame.
|
||||
///
|
||||
/// Restoring the position alone would restart the waveform and round the remainder away,
|
||||
/// which is a resample the restore rules refuse. The first chunk of the new epoch marks
|
||||
/// the discontinuity the recovery established.
|
||||
pub fn restored_from(
|
||||
descriptor: AudioDescriptor,
|
||||
next_sample: u64,
|
||||
phase: u64,
|
||||
accumulator: u128,
|
||||
denominator: u128,
|
||||
) -> DomainResult<AudioSource> {
|
||||
if denominator == 0 {
|
||||
return Err(DomainError::invalid(
|
||||
"audio: a captured accumulator denominator of zero",
|
||||
));
|
||||
}
|
||||
if accumulator >= denominator {
|
||||
return Err(DomainError::invalid(
|
||||
"audio: a captured accumulator is not below one whole frame",
|
||||
));
|
||||
}
|
||||
if phase >= descriptor.sample_rate {
|
||||
return Err(DomainError::invalid(
|
||||
"audio: a captured phase is not below the sample rate",
|
||||
));
|
||||
}
|
||||
let mut source = AudioSource::new(descriptor, next_sample);
|
||||
source.discontinuous = true;
|
||||
source.phase = phase;
|
||||
source.accumulator = accumulator;
|
||||
source.denominator = denominator;
|
||||
Ok(source)
|
||||
}
|
||||
|
||||
/// The waveform phase, for a capture.
|
||||
pub fn phase(&self) -> u64 {
|
||||
self.phase
|
||||
}
|
||||
|
||||
/// The unconsumed fraction of a frame and the denominator it is over, for a capture.
|
||||
pub fn accumulator(&self) -> (u128, u128) {
|
||||
(self.accumulator, self.denominator)
|
||||
}
|
||||
|
||||
pub fn descriptor(&self) -> &AudioDescriptor {
|
||||
&self.descriptor
|
||||
}
|
||||
|
|
@ -439,18 +536,47 @@ pub fn check_required_views(
|
|||
Ok(())
|
||||
}
|
||||
|
||||
/// Every declared audio stream produces exactly one chunk per transition.
|
||||
/// Where an observation came from.
|
||||
///
|
||||
/// `state-media-v1` section 2 makes a chunk the audio of an *interval*, so whether an
|
||||
/// observation must carry one is a question about its provenance and not about its boundary
|
||||
/// number. MEDIA-01 wrote the rule as "boundary 0 carries no chunk", which is true of the one
|
||||
/// observation that slice could produce without a transition and false of the other one:
|
||||
/// `State.ActivateRestore` installs a coherent observation at boundary `k` without advancing
|
||||
/// gameplay, and it covers no interval either. Naming the provenance is the fix; exempting
|
||||
/// the restored observation from the validator instead would have left "must a chunk exist"
|
||||
/// unanswered exactly where a stale chunk would do the most damage.
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
|
||||
pub enum ObservationOrigin {
|
||||
/// The observation a completed transition produced. Its interval has audio.
|
||||
Transition,
|
||||
/// An observation established at a boundary without running a transition:
|
||||
/// `Environment.Initialize`'s `O[0]` and `State.ActivateRestore`'s restored observation.
|
||||
/// It covers no interval, so it carries no chunk and one in it is refused.
|
||||
Installed,
|
||||
}
|
||||
|
||||
/// Every declared audio stream produces exactly one chunk per transition, and none at all in
|
||||
/// an observation that is not one.
|
||||
///
|
||||
/// The contract states the shape and the ordering of chunks, not whether one has to exist, so
|
||||
/// this is MEDIA-01's choice and it is deliberate: a session that tolerates a silently missing
|
||||
/// chunk cannot tell "this world produced no audio for this interval" from "the chunk was
|
||||
/// lost", and the second is the case the retention rules care about. Boundary 0 has no
|
||||
/// preceding interval and so carries no chunk.
|
||||
/// lost", and the second is the case the retention rules care about. The mirror of that, which
|
||||
/// STATE-01 needs, is that an installed observation carrying a chunk is a stale chunk being
|
||||
/// offered as current, and is refused for the same reason.
|
||||
pub fn check_required_audio(
|
||||
descriptor: &EnvironmentDescriptor,
|
||||
observation: &WorldObservation,
|
||||
origin: ObservationOrigin,
|
||||
) -> DomainResult<()> {
|
||||
if observation.boundary == 0 {
|
||||
if origin == ObservationOrigin::Installed {
|
||||
if let Some(chunk) = observation.audio.first() {
|
||||
return Err(media_error(format!(
|
||||
"audio stream {} produced a chunk for an observation that ran no transition",
|
||||
chunk.stream_id
|
||||
)));
|
||||
}
|
||||
return Ok(());
|
||||
}
|
||||
for stream in &descriptor.audio {
|
||||
|
|
|
|||
1494
services/flysim/crates/fly-session/src/state.rs
Normal file
1494
services/flysim/crates/fly-session/src/state.rs
Normal file
File diff suppressed because it is too large
Load diff
|
|
@ -39,6 +39,16 @@ pub fn episode_schema() -> SchemaRef {
|
|||
synthetic_schema("arena.episode.v1", 1)
|
||||
}
|
||||
|
||||
/// The schema of a captured task ledger.
|
||||
pub fn ledger_schema() -> SchemaRef {
|
||||
synthetic_schema("arena.ledger.v1", 1)
|
||||
}
|
||||
|
||||
/// The schema of a captured action-executor state.
|
||||
pub fn executor_schema() -> SchemaRef {
|
||||
synthetic_schema("arena.executor.v1", 1)
|
||||
}
|
||||
|
||||
pub fn controller_schema_ref() -> SchemaRef {
|
||||
synthetic_schema("arena.controller.v1", 1)
|
||||
}
|
||||
|
|
@ -86,6 +96,29 @@ pub trait Task: Send {
|
|||
|
||||
/// How many times `evaluate_transition` has run. A transition must evaluate once.
|
||||
fn evaluations(&self) -> u64;
|
||||
|
||||
/// The checkpointable ledger at a committed boundary (`workers-v1` section 4).
|
||||
fn capture(&self) -> DomainResult<TypedValue>;
|
||||
|
||||
/// Validates a captured ledger without installing it, so a group install can fail before
|
||||
/// anything is changed.
|
||||
fn validate_restore(&self, state: &TypedValue) -> DomainResult<()>;
|
||||
|
||||
/// Installs a validated ledger under `epoch`. Event identity is derived from the epoch,
|
||||
/// so the new one is part of the install rather than something the ledger keeps from the
|
||||
/// epoch it was captured in.
|
||||
fn install_restore(&mut self, epoch: &Id, state: &TypedValue) -> DomainResult<()>;
|
||||
|
||||
/// Every event identity this ledger has issued, mapped onto the identity it would have
|
||||
/// under `to_epoch`.
|
||||
///
|
||||
/// `workers-v1` section 4 derives an event id from the epoch, so a trace recorded in one
|
||||
/// epoch cannot be compared with a trace recorded in another until these are rebased.
|
||||
/// The ledger owns the derivation, so it is the only thing that can do it.
|
||||
fn rebase_ids(&self, to_epoch: &Id) -> DomainResult<BTreeMap<Id, Id>>;
|
||||
|
||||
/// How far event identity has reached: the highest source step and the number issued.
|
||||
fn event_watermarks(&self) -> (u64, u64);
|
||||
}
|
||||
|
||||
/// Translates one selected decision into a controller intent, with no port assignment.
|
||||
|
|
@ -98,6 +131,15 @@ pub trait ActionExecutor: Send {
|
|||
progress: &TypedValue,
|
||||
clock: &RationalNs,
|
||||
) -> DomainResult<(ControllerIntent, Vec<TaskEvent>)>;
|
||||
|
||||
/// Per-executor state at a committed boundary (`workers-v1` section 4).
|
||||
fn capture(&self) -> DomainResult<TypedValue>;
|
||||
|
||||
/// Validates a captured executor state without installing it.
|
||||
fn validate_restore(&self, state: &TypedValue) -> DomainResult<()>;
|
||||
|
||||
/// Installs a validated executor state.
|
||||
fn install_restore(&mut self, state: &TypedValue) -> DomainResult<()>;
|
||||
}
|
||||
|
||||
/// The only executor v1 supports: it passes a direct-control decision through unchanged.
|
||||
|
|
@ -123,6 +165,37 @@ impl ActionExecutor for IdentityExecutor {
|
|||
.map_err(|e| DomainError::invalid(format!("decision: {e}")))?;
|
||||
Ok((intent, Vec::new()))
|
||||
}
|
||||
|
||||
/// The identity executor is stateless, and says so rather than capturing nothing.
|
||||
///
|
||||
/// An empty object would be indistinguishable from a stateful executor whose capture went
|
||||
/// missing, so the capture names the executor it came from and a restore refuses any
|
||||
/// other one.
|
||||
fn capture(&self) -> DomainResult<TypedValue> {
|
||||
TypedValue::new(executor_schema(), json!({"executor": "identity-v1"}))
|
||||
.map_err(|e| DomainError::invalid(e.0))
|
||||
}
|
||||
|
||||
fn validate_restore(&self, state: &TypedValue) -> DomainResult<()> {
|
||||
if state.schema != executor_schema() {
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::IncompatibleState,
|
||||
"the captured executor state does not carry the executor schema",
|
||||
));
|
||||
}
|
||||
match state.value.get("executor").and_then(Value::as_str) {
|
||||
Some("identity-v1") => Ok(()),
|
||||
other => Err(DomainError::before(
|
||||
ErrorCode::IncompatibleState,
|
||||
format!("the captured executor is {other:?}, not the identity executor"),
|
||||
)),
|
||||
}
|
||||
}
|
||||
|
||||
fn install_restore(&mut self, state: &TypedValue) -> DomainResult<()> {
|
||||
// Stateless: validation is the whole of the install, and it is not skipped.
|
||||
self.validate_restore(state)
|
||||
}
|
||||
}
|
||||
|
||||
/// When the counter task asks for a terminal episode transition.
|
||||
|
|
@ -146,6 +219,10 @@ pub struct CounterTask {
|
|||
evaluations: u64,
|
||||
total_reward: f64,
|
||||
counter: i64,
|
||||
/// The highest source step any issued event belongs to, and how many were issued. These
|
||||
/// are the event watermarks a checkpoint records and a resumed epoch continues from.
|
||||
last_source_step: u64,
|
||||
issued_events: u64,
|
||||
terminal: Terminal,
|
||||
}
|
||||
|
||||
|
|
@ -159,6 +236,8 @@ impl CounterTask {
|
|||
evaluations: 0,
|
||||
total_reward: 0.0,
|
||||
counter: 0,
|
||||
last_source_step: 0,
|
||||
issued_events: 0,
|
||||
terminal,
|
||||
}
|
||||
}
|
||||
|
|
@ -239,6 +318,7 @@ impl Task for CounterTask {
|
|||
payload: TypedValue::new(event_schema(), json!({"counter": self.counter}))
|
||||
.expect("a synthetic typed value fits the contract"),
|
||||
}];
|
||||
self.issued_events += events.len() as u64;
|
||||
Ok(Bootstrap { contexts, progress: self.progress_value(), events })
|
||||
}
|
||||
|
||||
|
|
@ -314,6 +394,8 @@ impl Task for CounterTask {
|
|||
));
|
||||
}
|
||||
|
||||
self.last_source_step = self.last_source_step.max(source_step);
|
||||
self.issued_events += events.len() as u64;
|
||||
let next_contexts = self
|
||||
.agents
|
||||
.iter()
|
||||
|
|
@ -346,6 +428,173 @@ impl Task for CounterTask {
|
|||
fn evaluations(&self) -> u64 {
|
||||
self.evaluations
|
||||
}
|
||||
|
||||
fn capture(&self) -> DomainResult<TypedValue> {
|
||||
TypedValue::new(
|
||||
ledger_schema(),
|
||||
json!({
|
||||
"epoch": self.epoch.as_str(),
|
||||
"agents": self.agents.iter().map(String::as_str).collect::<Vec<_>>(),
|
||||
"bindings": self
|
||||
.bindings
|
||||
.iter()
|
||||
.map(|b| json!({"portId": b.port_id.as_str(), "agentId": b.agent_id.as_str()}))
|
||||
.collect::<Vec<_>>(),
|
||||
"transitions": self.transitions,
|
||||
"evaluations": self.evaluations,
|
||||
"totalReward": self.total_reward,
|
||||
"counter": self.counter,
|
||||
"lastSourceStep": self.last_source_step,
|
||||
"issuedEvents": self.issued_events,
|
||||
}),
|
||||
)
|
||||
.map_err(|e| DomainError::invalid(e.0))
|
||||
}
|
||||
|
||||
fn validate_restore(&self, state: &TypedValue) -> DomainResult<()> {
|
||||
if state.schema != ledger_schema() {
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::IncompatibleState,
|
||||
"the captured ledger does not carry this task's schema",
|
||||
));
|
||||
}
|
||||
for field in [
|
||||
"epoch",
|
||||
"agents",
|
||||
"bindings",
|
||||
"transitions",
|
||||
"evaluations",
|
||||
"totalReward",
|
||||
"counter",
|
||||
"lastSourceStep",
|
||||
"issuedEvents",
|
||||
] {
|
||||
if state.value.get(field).is_none() {
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::IncompatibleState,
|
||||
format!("the captured ledger has no {field}"),
|
||||
));
|
||||
}
|
||||
}
|
||||
let bindings = state
|
||||
.value
|
||||
.get("bindings")
|
||||
.and_then(Value::as_array)
|
||||
.ok_or_else(|| {
|
||||
DomainError::before(
|
||||
ErrorCode::IncompatibleState,
|
||||
"the captured ledger's bindings are not a list",
|
||||
)
|
||||
})?;
|
||||
if bindings.len() != self.bindings.len() && !self.bindings.is_empty() {
|
||||
return Err(DomainError::before(
|
||||
ErrorCode::IncompatibleState,
|
||||
"the captured ledger binds another number of ports",
|
||||
));
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
fn install_restore(&mut self, epoch: &Id, state: &TypedValue) -> DomainResult<()> {
|
||||
self.validate_restore(state)?;
|
||||
let number = |key: &str| -> DomainResult<u64> {
|
||||
state.value.get(key).and_then(Value::as_u64).ok_or_else(|| {
|
||||
DomainError::before(
|
||||
ErrorCode::IncompatibleState,
|
||||
format!("the captured ledger's {key} is not a whole number"),
|
||||
)
|
||||
})
|
||||
};
|
||||
let mut agents = Vec::new();
|
||||
for value in state.value["agents"].as_array().expect("validated") {
|
||||
let agent = value.as_str().ok_or_else(|| {
|
||||
DomainError::before(
|
||||
ErrorCode::IncompatibleState,
|
||||
"the captured ledger names an agent that is not a string",
|
||||
)
|
||||
})?;
|
||||
agents.push(parse_id(agent).map_err(|e| {
|
||||
DomainError::before(ErrorCode::IncompatibleState, format!("ledger: {e}"))
|
||||
})?);
|
||||
}
|
||||
let mut bindings = Vec::new();
|
||||
for value in state.value["bindings"].as_array().expect("validated") {
|
||||
let port_id = value.get("portId").and_then(Value::as_str).ok_or_else(|| {
|
||||
DomainError::before(
|
||||
ErrorCode::IncompatibleState,
|
||||
"the captured ledger has a binding with no portId",
|
||||
)
|
||||
})?;
|
||||
let agent_id = value.get("agentId").and_then(Value::as_str).ok_or_else(|| {
|
||||
DomainError::before(
|
||||
ErrorCode::IncompatibleState,
|
||||
"the captured ledger has a binding with no agentId",
|
||||
)
|
||||
})?;
|
||||
bindings.push(PortBinding {
|
||||
port_id: parse_id(port_id).map_err(|e| {
|
||||
DomainError::before(ErrorCode::IncompatibleState, format!("ledger: {e}"))
|
||||
})?,
|
||||
agent_id: parse_id(agent_id).map_err(|e| {
|
||||
DomainError::before(ErrorCode::IncompatibleState, format!("ledger: {e}"))
|
||||
})?,
|
||||
});
|
||||
}
|
||||
let counter = state.value.get("counter").and_then(Value::as_i64).ok_or_else(|| {
|
||||
DomainError::before(
|
||||
ErrorCode::IncompatibleState,
|
||||
"the captured ledger's counter is not an integer",
|
||||
)
|
||||
})?;
|
||||
let total_reward = state
|
||||
.value
|
||||
.get("totalReward")
|
||||
.and_then(Value::as_f64)
|
||||
.filter(|v| v.is_finite())
|
||||
.ok_or_else(|| {
|
||||
DomainError::before(
|
||||
ErrorCode::IncompatibleState,
|
||||
"the captured ledger's totalReward is not a finite number",
|
||||
)
|
||||
})?;
|
||||
// The epoch is the caller's, not the capture's: event identity belongs to the epoch
|
||||
// the ledger is being installed into.
|
||||
self.epoch = epoch.clone();
|
||||
self.agents = agents;
|
||||
self.bindings = bindings;
|
||||
self.transitions = number("transitions")?;
|
||||
self.evaluations = number("evaluations")?;
|
||||
self.total_reward = total_reward;
|
||||
self.counter = counter;
|
||||
self.last_source_step = number("lastSourceStep")?;
|
||||
self.issued_events = number("issuedEvents")?;
|
||||
Ok(())
|
||||
}
|
||||
|
||||
fn rebase_ids(&self, to_epoch: &Id) -> DomainResult<BTreeMap<Id, Id>> {
|
||||
let mut out = BTreeMap::new();
|
||||
out.insert(
|
||||
event_id(&self.epoch, 0, "bootstrap", 0),
|
||||
event_id(to_epoch, 0, "bootstrap", 0),
|
||||
);
|
||||
// The counter task issues exactly one `counter-delta` event per bound port per
|
||||
// evaluated transition, in descriptor port order, so every identity it has ever
|
||||
// issued is re-derivable from its ledger without keeping a list of them.
|
||||
let ports = self.bindings.len() as u32;
|
||||
for source_step in 1..=self.last_source_step {
|
||||
for ordinal in 0..ports {
|
||||
out.insert(
|
||||
event_id(&self.epoch, source_step, "counter-delta", ordinal),
|
||||
event_id(to_epoch, source_step, "counter-delta", ordinal),
|
||||
);
|
||||
}
|
||||
}
|
||||
Ok(out)
|
||||
}
|
||||
|
||||
fn event_watermarks(&self) -> (u64, u64) {
|
||||
(self.last_source_step, self.issued_events)
|
||||
}
|
||||
}
|
||||
|
||||
/// The inspection value the counter environment publishes.
|
||||
|
|
|
|||
|
|
@ -8,13 +8,18 @@
|
|||
//! `Id` and `Digest` are type aliases, because the shared crate carries both as validated
|
||||
//! `String`s from `flybus::wire` rather than forking the encodings into newtypes.
|
||||
|
||||
use std::collections::BTreeMap;
|
||||
|
||||
use serde_json::{Map, Value};
|
||||
|
||||
pub use fly_session_types::ArtifactRef;
|
||||
pub use fly_session_types::canonical::{
|
||||
self, OperationKey, body_digest, canonicalize, digest_of, sha256_hex,
|
||||
};
|
||||
pub use fly_session_types::media::{AudioDescriptor, AudioRef, ViewDescriptor, ViewRef};
|
||||
pub use fly_session_types::media::{
|
||||
ActivateRestoreParams, ActivateRestoreResult, AudioDescriptor, AudioRef, CaptureParams,
|
||||
CaptureResult, StageRestoreParams, StageRestoreResult, ViewDescriptor, ViewRef,
|
||||
};
|
||||
pub use fly_session_types::rpc::{
|
||||
ErrorCode, MutationCertainty, SessionRpcFailure, SessionRpcOutcome, SessionRpcRequest,
|
||||
SessionRpcSuccess,
|
||||
|
|
@ -214,6 +219,57 @@ pub fn outcome_identity(
|
|||
}
|
||||
}
|
||||
|
||||
/// The epoch-derived identities of a behaviour trace, rewritten onto one reference epoch.
|
||||
///
|
||||
/// `step-v1` section 8 compares committed behaviour across runs, excluding wall time, request
|
||||
/// ids "and other explicitly operational metadata". A resumed run's epoch is neither: it is
|
||||
/// behaviour metadata, and `scope.epoch`, the batch id and every task event id are derived
|
||||
/// from it. Comparing the two runs therefore means rewriting exactly those three things and
|
||||
/// nothing else, which is what this does -- and it **fails** on anything it does not
|
||||
/// recognise instead of passing it through, so a field that silently stopped being rebased
|
||||
/// would fail the comparison rather than weaken it.
|
||||
#[derive(Clone, Debug, PartialEq, Eq)]
|
||||
pub struct EpochRebase {
|
||||
pub from: Id,
|
||||
pub to: Id,
|
||||
/// Every event identity the task issued under `from`, and the identity it has under `to`.
|
||||
pub events: BTreeMap<Id, Id>,
|
||||
}
|
||||
|
||||
impl EpochRebase {
|
||||
/// Rewrites one behaviour record. An identity this rebase does not know is an error.
|
||||
pub fn apply(&self, behaviour: &TraceBehaviour) -> Result<TraceBehaviour, String> {
|
||||
if behaviour.scope.epoch != self.from {
|
||||
return Err(format!(
|
||||
"this behaviour was recorded in epoch {}, not {}",
|
||||
behaviour.scope.epoch, self.from
|
||||
));
|
||||
}
|
||||
let mut out = behaviour.clone();
|
||||
out.scope = Scope::new(&behaviour.scope.session_id, &self.to, behaviour.scope.step)
|
||||
.map_err(|e| e.0)?;
|
||||
let prefix = format!("batch-{}-", self.from);
|
||||
let suffix = behaviour
|
||||
.batch_id
|
||||
.strip_prefix(&prefix)
|
||||
.ok_or_else(|| format!("the batch id {} is not derived from {}", behaviour.batch_id, self.from))?;
|
||||
out.batch_id = parse_id(&format!("batch-{}-{suffix}", self.to))?;
|
||||
let map = |ids: &[Id]| -> Result<Vec<Id>, String> {
|
||||
ids.iter()
|
||||
.map(|id| {
|
||||
self.events
|
||||
.get(id)
|
||||
.cloned()
|
||||
.ok_or_else(|| format!("no rebased identity for the event {id}"))
|
||||
})
|
||||
.collect()
|
||||
};
|
||||
out.outcome_ids = map(&behaviour.outcome_ids)?;
|
||||
out.event_ids = map(&behaviour.event_ids)?;
|
||||
Ok(out)
|
||||
}
|
||||
}
|
||||
|
||||
/// One session phase transition, recorded whether or not it ends a step.
|
||||
#[derive(Clone, Debug, PartialEq, Eq)]
|
||||
pub struct PhaseTransition {
|
||||
|
|
@ -253,6 +309,28 @@ impl TraceLog {
|
|||
.collect()
|
||||
}
|
||||
|
||||
/// Every transition's behaviour, rebased onto one epoch and canonicalized.
|
||||
///
|
||||
/// This is the comparison a resumed run is held to: the same strings as
|
||||
/// [`TraceLog::behavior`], with the epoch metadata accounted for and nothing else changed.
|
||||
/// A resumed run's log holds transitions from two epochs -- the ones before the checkpoint
|
||||
/// and the ones after the restore -- so a transition already recorded in `rebase.to` is
|
||||
/// kept as it stands and one recorded in `rebase.from` is rewritten. A transition in a
|
||||
/// third epoch is an error; there is no pass-through case.
|
||||
pub fn behavior_rebased(&self, rebase: &EpochRebase) -> Result<Vec<String>, String> {
|
||||
self.transitions
|
||||
.iter()
|
||||
.map(|t| {
|
||||
let behaviour = if t.behaviour.scope.epoch == rebase.to {
|
||||
t.behaviour.clone()
|
||||
} else {
|
||||
rebase.apply(&t.behaviour)?
|
||||
};
|
||||
canonicalize(&behaviour.to_json()).map_err(|e| e.0)
|
||||
})
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// The phase path, as `from -> to` strings.
|
||||
pub fn phase_path(&self) -> Vec<String> {
|
||||
self.phases.iter().map(|p| format!("{} -> {}", p.from, p.to)).collect()
|
||||
|
|
|
|||
|
|
@ -597,7 +597,17 @@ async fn execute<E: WorkerEndpoint>(
|
|||
fn classify_default(method: &str) -> Option<OpClass> {
|
||||
match method {
|
||||
"Agent.Prepare" | "Agent.Commit" | "Environment.Advance" => Some(OpClass::StepMutation),
|
||||
"Agent.Initialize" | "Environment.Initialize" => Some(OpClass::Lifecycle),
|
||||
// `ipc-v1` section 5 retains lifecycle *and capture* replies until
|
||||
// `Worker.Acknowledge`. The restore methods join them: their replies carry a
|
||||
// once-only token and, for an environment, the restored observation's artifact, and a
|
||||
// duplicate domain request must replay that reply rather than stage or activate a
|
||||
// second time. They are not step mutations -- they carry no committed step of their
|
||||
// own and are not keyed by one.
|
||||
"Agent.Initialize"
|
||||
| "Environment.Initialize"
|
||||
| "State.Capture"
|
||||
| "State.StageRestore"
|
||||
| "State.ActivateRestore" => Some(OpClass::Lifecycle),
|
||||
"Worker.Hello" | "Worker.Status" | "Worker.Acknowledge" | "Worker.Shutdown" => {
|
||||
Some(OpClass::ReadOnly)
|
||||
}
|
||||
|
|
|
|||
|
|
@ -121,6 +121,7 @@ async fn a_slow_participant_is_resolved_rather_than_failed(mode: ExecutionMode)
|
|||
resolve: Duration::from_secs(20),
|
||||
resolve_attempts: 4096,
|
||||
boot: Duration::from_secs(30),
|
||||
capture: Duration::from_secs(30),
|
||||
};
|
||||
within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap();
|
||||
let reports = within("run", f.harness.coordinator.run(2))
|
||||
|
|
@ -191,6 +192,7 @@ async fn a_resolution_says_which_of_its_two_bounds_ended_it(mode: ExecutionMode)
|
|||
resolve: Duration::from_millis(300),
|
||||
resolve_attempts: 8192,
|
||||
boot: Duration::from_secs(30),
|
||||
capture: Duration::from_secs(30),
|
||||
};
|
||||
within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap();
|
||||
let started = Instant::now();
|
||||
|
|
@ -221,6 +223,7 @@ async fn a_resolution_says_which_of_its_two_bounds_ended_it(mode: ExecutionMode)
|
|||
resolve: Duration::from_secs(600),
|
||||
resolve_attempts: 3,
|
||||
boot: Duration::from_secs(30),
|
||||
capture: Duration::from_secs(30),
|
||||
};
|
||||
within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap();
|
||||
let failure = within("step", f.harness.coordinator.step())
|
||||
|
|
|
|||
947
services/flysim/crates/fly-session/tests/state.rs
Normal file
947
services/flysim/crates/fly-session/tests/state.rs
Normal file
|
|
@ -0,0 +1,947 @@
|
|||
//! STATE-01 acceptance: one coherent all-participant checkpoint, and the recovery that
|
||||
//! installs it into a fresh epoch.
|
||||
//!
|
||||
//! Every test here is one of the slice's acceptance bullets or one of the failure-injection
|
||||
//! rows the implementation guide's section 4 assigns to it. Each of them runs over both
|
||||
//! transports and in all three execution modes: the recovery path crosses a process boundary
|
||||
//! in exactly the places the media path does, so a row that holds in one mode has to hold in
|
||||
//! all of them.
|
||||
//!
|
||||
//! The byte layout itself is proved against the contract crate in
|
||||
//! `fly-session-types/tests/checkpoint_envelope.rs`; these prove the session's use of it.
|
||||
|
||||
mod common;
|
||||
|
||||
use std::collections::BTreeSet;
|
||||
use std::path::Path;
|
||||
use std::sync::Arc;
|
||||
|
||||
use serde_json::{Value, json};
|
||||
|
||||
use common::{Fixture, count, fixture, fly_a, fly_b, within};
|
||||
use fly_session::agent::AgentFaults;
|
||||
use fly_session::harness::{ExecutionMode, HarnessConfig, Via};
|
||||
use fly_session::phase::Phase;
|
||||
use fly_session::state::{SaveOutcome, WriterFaults};
|
||||
use fly_session::types::*;
|
||||
use fly_session_types::checkpoint;
|
||||
|
||||
both_transports!(
|
||||
an_uninterrupted_run_and_a_resumed_run_produce_matching_traces,
|
||||
corrupting_any_single_participants_payload_fails_the_install_as_a_group,
|
||||
a_lost_save_reply_does_not_advance_durable_metadata,
|
||||
a_failure_during_activation_cannot_resume_half_a_world,
|
||||
the_checkpoint_queue_under_stress_stays_bounded,
|
||||
old_media_cannot_cross_a_recovery,
|
||||
a_checkpoint_taken_by_another_parser_is_refused_by_name,
|
||||
a_group_where_one_participant_will_not_stage_resumes_nothing,
|
||||
a_restore_token_activates_only_once,
|
||||
the_fence_lifts_only_through_a_coherent_restore,
|
||||
a_queued_replaceable_capture_is_superseded_rather_than_duplicated,
|
||||
a_capture_is_refused_anywhere_but_a_committed_boundary,
|
||||
);
|
||||
|
||||
all_modes!(
|
||||
matching_traces_in_every_mode,
|
||||
a_group_install_fails_as_a_group_in_every_mode,
|
||||
a_lost_save_reply_holds_durable_metadata_in_every_mode,
|
||||
half_an_activation_resumes_nothing_in_every_mode,
|
||||
the_queue_stays_bounded_in_every_mode,
|
||||
old_media_cannot_cross_in_every_mode,
|
||||
another_parser_is_refused_in_every_mode,
|
||||
a_refused_stage_resumes_nothing_in_every_mode,
|
||||
a_token_activates_once_in_every_mode,
|
||||
the_fence_lifts_only_by_restore_in_every_mode,
|
||||
);
|
||||
|
||||
const BEFORE: u64 = 2;
|
||||
const AFTER: u64 = 2;
|
||||
|
||||
fn ckpt(n: u32) -> Id {
|
||||
id(&format!("ckpt-{n}"))
|
||||
}
|
||||
|
||||
/// A fixture in one transport and one execution mode.
|
||||
///
|
||||
/// `Via` is the composition's; a dedicated thread or a separate process reaches the router
|
||||
/// over a socket whatever it says, which the launcher decides and this does not second-guess.
|
||||
async fn fx(via: Via, mode: ExecutionMode, config: HarnessConfig) -> Fixture {
|
||||
fixture(via, HarnessConfig { mode, ..config }).await
|
||||
}
|
||||
|
||||
/// Runs to a committed boundary and takes one durable checkpoint there.
|
||||
async fn run_and_checkpoint(f: &mut Fixture, steps: u64, checkpoint_id: &Id) {
|
||||
within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap();
|
||||
within("run", f.harness.coordinator.run(steps)).await.unwrap();
|
||||
let outcome = within("checkpoint", f.harness.coordinator.checkpoint(checkpoint_id))
|
||||
.await
|
||||
.unwrap();
|
||||
match outcome {
|
||||
SaveOutcome::Committed { boundary, .. } => assert_eq!(boundary, steps),
|
||||
other => panic!("the checkpoint was not committed: {other:?}"),
|
||||
}
|
||||
assert_eq!(
|
||||
f.harness.coordinator.durable(),
|
||||
Some((checkpoint_id.clone(), steps)),
|
||||
"a committed save is the only thing that moves the durable mark"
|
||||
);
|
||||
}
|
||||
|
||||
/// Fails the epoch the way a participant death does, and checks the fence closed.
|
||||
async fn fail_the_epoch(f: &mut Fixture) {
|
||||
f.harness.kill(&fly_a()).await;
|
||||
let failure = within("step", f.harness.coordinator.step())
|
||||
.await
|
||||
.expect_err("a dead participant fails the epoch");
|
||||
assert!(
|
||||
failure.participant.is_some(),
|
||||
"a diagnosed failure names its participant: {failure}"
|
||||
);
|
||||
assert_eq!(f.harness.coordinator.phase(), Phase::Failed);
|
||||
assert!(f.harness.coordinator.is_fenced());
|
||||
assert_eq!(
|
||||
f.harness.coordinator.live_view_handles(),
|
||||
0,
|
||||
"the fence drops every handle the old epoch held"
|
||||
);
|
||||
}
|
||||
|
||||
/// The committed generation's bytes, as they are on disk.
|
||||
fn generation_bytes(root: &Path, checkpoint_id: &Id) -> Vec<u8> {
|
||||
std::fs::read(root.join(format!("{checkpoint_id}.flysess"))).expect("a committed generation")
|
||||
}
|
||||
|
||||
/// Writes a generation file and makes the store manifest describe it, as a repair tool or a
|
||||
/// previous run would have left it.
|
||||
fn write_generation(root: &Path, checkpoint_id: &Id, bytes: &[u8]) {
|
||||
std::fs::write(root.join(format!("{checkpoint_id}.flysess")), bytes).expect("writable");
|
||||
let path = root.join("manifest.json");
|
||||
let mut manifest: Value =
|
||||
serde_json::from_slice(&std::fs::read(&path).expect("a store manifest")).expect("json");
|
||||
let generations = manifest["generations"].as_array_mut().expect("an array");
|
||||
for generation in generations.iter_mut() {
|
||||
if generation["checkpointId"] == json!(checkpoint_id.as_str()) {
|
||||
generation["byteLength"] = json!(bytes.len().to_string());
|
||||
generation["envelopeDigest"] = json!(digest_of_bytes(bytes));
|
||||
let envelope = checkpoint::decode(bytes).expect("a well formed envelope");
|
||||
let compatibility = fly_session::state::Compatibility::from_json(
|
||||
&envelope.manifest["compatibility"],
|
||||
)
|
||||
.expect("a compatibility block");
|
||||
generation["compatibilityDigest"] = json!(compatibility.digest());
|
||||
}
|
||||
}
|
||||
std::fs::write(&path, serde_json::to_vec(&manifest).expect("json")).expect("writable");
|
||||
}
|
||||
|
||||
/// Re-encodes one committed generation after `edit` has changed its manifest or its payloads.
|
||||
fn rewrite_generation(
|
||||
root: &Path,
|
||||
checkpoint_id: &Id,
|
||||
edit: impl FnOnce(&mut Value, &mut Vec<(String, Vec<u8>)>),
|
||||
) {
|
||||
let bytes = generation_bytes(root, checkpoint_id);
|
||||
let envelope = checkpoint::decode(&bytes).expect("a committed generation decodes");
|
||||
let mut manifest = envelope.manifest.clone();
|
||||
let mut payloads = envelope.payloads.clone();
|
||||
edit(&mut manifest, &mut payloads);
|
||||
// The manifest's payload table mirrors the envelope's, so it is rebuilt from the bytes
|
||||
// that are actually being written rather than left to disagree.
|
||||
manifest["payloads"] = Value::Array(
|
||||
payloads
|
||||
.iter()
|
||||
.map(|(name, bytes)| {
|
||||
json!({
|
||||
"name": name,
|
||||
"byteLength": bytes.len().to_string(),
|
||||
"digest": digest_of_bytes(bytes),
|
||||
})
|
||||
})
|
||||
.collect(),
|
||||
);
|
||||
let rewritten = checkpoint::encode(&manifest, &payloads).expect("a valid envelope");
|
||||
write_generation(root, checkpoint_id, &rewritten);
|
||||
}
|
||||
|
||||
/// Makes the store read its durable metadata again, after a test has edited it.
|
||||
async fn reload_store(f: &mut Fixture) {
|
||||
f.harness
|
||||
.coordinator
|
||||
.writer()
|
||||
.expect("a store is attached")
|
||||
.with_store(|store| store.reload())
|
||||
.await
|
||||
.expect("the store manifest is still readable");
|
||||
}
|
||||
|
||||
/// The names of every payload the checkpoint holds, participants first.
|
||||
fn payload_names(root: &Path, checkpoint_id: &Id) -> Vec<String> {
|
||||
let bytes = generation_bytes(root, checkpoint_id);
|
||||
checkpoint::decode(&bytes)
|
||||
.expect("decodes")
|
||||
.payloads
|
||||
.into_iter()
|
||||
.map(|(name, _)| name)
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// Every artifact identity this session's committed boundary is holding.
|
||||
fn live_artifact_ids(f: &Fixture) -> BTreeSet<String> {
|
||||
f.harness
|
||||
.coordinator
|
||||
.media_handles()
|
||||
.into_iter()
|
||||
.map(|(_, reference)| reference.artifact_id)
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// Asserts that no participant is running the proposed epoch: nothing was installed.
|
||||
async fn nothing_is_installed(f: &mut Fixture, epoch: &Id) {
|
||||
let ids: Vec<Id> = std::iter::once(f.harness.environment_id())
|
||||
.chain(f.harness.config.agents.iter().map(|a| a.agent_id.clone()))
|
||||
.collect();
|
||||
for worker_id in ids {
|
||||
let (_, launcher) = f.harness.parts();
|
||||
let Ok(status) = launcher.health_check(&worker_id).await else {
|
||||
// A participant that is gone is certainly not running the new epoch.
|
||||
continue;
|
||||
};
|
||||
if let Some(scope) = status.current_scope {
|
||||
assert_ne!(
|
||||
scope.epoch, *epoch,
|
||||
"{worker_id} is running the epoch the abandoned install proposed"
|
||||
);
|
||||
}
|
||||
}
|
||||
assert_eq!(f.harness.coordinator.phase(), Phase::Failed);
|
||||
assert!(f.harness.coordinator.is_fenced());
|
||||
let refused = within("step", f.harness.coordinator.step())
|
||||
.await
|
||||
.expect_err("a fenced session takes no step");
|
||||
assert_eq!(refused.error.code, ErrorCode::InvalidPhase);
|
||||
}
|
||||
|
||||
// ===============================================================================================
|
||||
// Acceptance: an uninterrupted run and a resumed run produce matching traces
|
||||
|
||||
async fn an_uninterrupted_run_and_a_resumed_run_produce_matching_traces(via: Via) {
|
||||
matching_traces(via, ExecutionMode::InProcess).await;
|
||||
}
|
||||
|
||||
async fn matching_traces_in_every_mode(mode: ExecutionMode) {
|
||||
matching_traces(Via::Unix, mode).await;
|
||||
}
|
||||
|
||||
/// The reference run and the resumed run commit the same behaviour, once the epoch metadata
|
||||
/// the restore necessarily changed is accounted for.
|
||||
///
|
||||
/// `step-v1` section 8's split is the whole of the comparison: the behaviour is what must
|
||||
/// match, the request ids and wall time are excluded because they are operational, and the
|
||||
/// epoch is neither -- it is behaviour metadata, so it is rewritten explicitly and everything
|
||||
/// else is compared byte for byte.
|
||||
async fn matching_traces(via: Via, mode: ExecutionMode) {
|
||||
let mut reference = fx(via, mode, HarnessConfig::default()).await;
|
||||
within("bootstrap", reference.harness.coordinator.bootstrap()).await.unwrap();
|
||||
within("run", reference.harness.coordinator.run(BEFORE + AFTER)).await.unwrap();
|
||||
let expected = reference.harness.coordinator.trace.behavior();
|
||||
assert_eq!(expected.len() as u64, BEFORE + AFTER);
|
||||
reference.shutdown().await;
|
||||
|
||||
let mut f = fx(via, mode, HarnessConfig::default()).await;
|
||||
let checkpoint_id = ckpt(1);
|
||||
run_and_checkpoint(&mut f, BEFORE, &checkpoint_id).await;
|
||||
fail_the_epoch(&mut f).await;
|
||||
within("replace", f.harness.replace_all_participants()).await.unwrap();
|
||||
let report = within(
|
||||
"restore",
|
||||
f.harness.coordinator.restore(Some(&checkpoint_id), &id("e2")),
|
||||
)
|
||||
.await
|
||||
.unwrap();
|
||||
assert_eq!(report.boundary, BEFORE);
|
||||
assert_eq!(report.epoch, id("e2"));
|
||||
assert_eq!(report.activated.len(), report.staged.len());
|
||||
assert!(!f.harness.coordinator.is_fenced(), "a coherent restore lifts the fence");
|
||||
assert_eq!(f.harness.coordinator.phase(), Phase::Paused(BEFORE));
|
||||
|
||||
f.harness.coordinator.resume().unwrap();
|
||||
within("resume", f.harness.coordinator.run(AFTER)).await.unwrap();
|
||||
assert_eq!(f.harness.coordinator.phase(), Phase::Ready(BEFORE + AFTER));
|
||||
|
||||
// Without accounting for the epoch the two traces disagree, which is what makes the
|
||||
// rebase a statement rather than a formality.
|
||||
let raw = f.harness.coordinator.trace.behavior();
|
||||
assert_eq!(raw.len() as u64, BEFORE + AFTER);
|
||||
assert_ne!(raw, expected, "the resumed run runs in a different epoch");
|
||||
|
||||
let rebase = f.harness.coordinator.rebase(&id("e1")).unwrap();
|
||||
let resumed = f.harness.coordinator.trace.behavior_rebased(&rebase).unwrap();
|
||||
assert_eq!(
|
||||
resumed, expected,
|
||||
"a resumed run commits the behaviour the uninterrupted run committed"
|
||||
);
|
||||
f.shutdown().await;
|
||||
}
|
||||
|
||||
// ===============================================================================================
|
||||
// Acceptance: corrupt any participant and installation fails as a group
|
||||
|
||||
async fn corrupting_any_single_participants_payload_fails_the_install_as_a_group(via: Via) {
|
||||
group_install_is_all_or_nothing(via, ExecutionMode::InProcess).await;
|
||||
}
|
||||
|
||||
async fn a_group_install_fails_as_a_group_in_every_mode(mode: ExecutionMode) {
|
||||
group_install_is_all_or_nothing(Via::Unix, mode).await;
|
||||
}
|
||||
|
||||
/// One corrupted payload -- any participant's, and the coordinator's own ledger too -- fails
|
||||
/// the whole install, and the group afterwards is exactly as it was.
|
||||
async fn group_install_is_all_or_nothing(via: Via, mode: ExecutionMode) {
|
||||
let mut f = fx(via, mode, HarnessConfig::default()).await;
|
||||
let checkpoint_id = ckpt(1);
|
||||
run_and_checkpoint(&mut f, BEFORE, &checkpoint_id).await;
|
||||
let root = f.harness.checkpoint_root().to_path_buf();
|
||||
let good = generation_bytes(&root, &checkpoint_id);
|
||||
let names = payload_names(&root, &checkpoint_id);
|
||||
// Every participant's payload, plus one of the coordinator's own ledgers.
|
||||
let mut corrupt: Vec<String> = names
|
||||
.iter()
|
||||
.filter(|name| name.starts_with("agent-") || *name == "world")
|
||||
.cloned()
|
||||
.collect();
|
||||
corrupt.push("task-ledger".to_owned());
|
||||
assert!(corrupt.len() >= 3, "the composition has several payloads: {names:?}");
|
||||
fail_the_epoch(&mut f).await;
|
||||
|
||||
let mut epoch = 1u32;
|
||||
for target in &corrupt {
|
||||
epoch += 1;
|
||||
let proposed = id(&format!("e{epoch}"));
|
||||
// Each round starts from the intact bytes, so exactly one payload is corrupt.
|
||||
write_generation(&root, &checkpoint_id, &good);
|
||||
// A well formed envelope whose digests all agree, so what fails is the participant
|
||||
// reading its own bytes and not the envelope reader in front of it.
|
||||
let name = target.clone();
|
||||
rewrite_generation(&root, &checkpoint_id, |_manifest, payloads| {
|
||||
for (payload, bytes) in payloads.iter_mut() {
|
||||
if *payload == name {
|
||||
let last = bytes.len() - 1;
|
||||
bytes[last] ^= 0xff;
|
||||
}
|
||||
}
|
||||
});
|
||||
reload_store(&mut f).await;
|
||||
within("replace", f.harness.replace_all_participants()).await.unwrap();
|
||||
let failure = within(
|
||||
"restore",
|
||||
f.harness.coordinator.restore(Some(&checkpoint_id), &proposed),
|
||||
)
|
||||
.await
|
||||
.expect_err(&format!("a corrupt {target} must fail the install"));
|
||||
assert!(
|
||||
matches!(
|
||||
failure.error.code,
|
||||
ErrorCode::IncompatibleState | ErrorCode::InvalidArgument
|
||||
),
|
||||
"a corrupt {target} is an explicit refusal, not {failure}"
|
||||
);
|
||||
nothing_is_installed(&mut f, &proposed).await;
|
||||
let tainted = f.harness.coordinator.tainted();
|
||||
if target == "world" {
|
||||
assert!(
|
||||
tainted.is_empty(),
|
||||
"the first participant asked refused, so nothing staged: {tainted:?}"
|
||||
);
|
||||
} else {
|
||||
assert!(
|
||||
!tainted.is_empty(),
|
||||
"a participant that staged into an abandoned install must be replaced"
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
// The same group, the same store, the intact bytes: the failures above installed nothing
|
||||
// that stops this from working.
|
||||
epoch += 1;
|
||||
write_generation(&root, &checkpoint_id, &good);
|
||||
reload_store(&mut f).await;
|
||||
within("replace", f.harness.replace_all_participants()).await.unwrap();
|
||||
let report = within(
|
||||
"restore",
|
||||
f.harness
|
||||
.coordinator
|
||||
.restore(Some(&checkpoint_id), &id(&format!("e{epoch}"))),
|
||||
)
|
||||
.await
|
||||
.unwrap();
|
||||
assert_eq!(report.boundary, BEFORE);
|
||||
f.harness.coordinator.resume().unwrap();
|
||||
within("resume", f.harness.coordinator.run(1)).await.unwrap();
|
||||
f.shutdown().await;
|
||||
}
|
||||
|
||||
// ===============================================================================================
|
||||
// Acceptance: a lost save reply does not advance durable metadata
|
||||
|
||||
async fn a_lost_save_reply_does_not_advance_durable_metadata(via: Via) {
|
||||
a_lost_save_reply(via, ExecutionMode::InProcess).await;
|
||||
}
|
||||
|
||||
async fn a_lost_save_reply_holds_durable_metadata_in_every_mode(mode: ExecutionMode) {
|
||||
a_lost_save_reply(Via::Unix, mode).await;
|
||||
}
|
||||
|
||||
/// Two ways a save can end without a saved acknowledgment, and neither moves the mark.
|
||||
async fn a_lost_save_reply(via: Via, mode: ExecutionMode) {
|
||||
let lost = ckpt(1);
|
||||
let config = HarnessConfig {
|
||||
writer_faults: WriterFaults {
|
||||
drop_reply_for: Some(lost.clone()),
|
||||
..WriterFaults::default()
|
||||
},
|
||||
..HarnessConfig::default()
|
||||
};
|
||||
let mut f = fx(via, mode, config).await;
|
||||
within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap();
|
||||
within("run", f.harness.coordinator.run(1)).await.unwrap();
|
||||
|
||||
let ticket = within("capture", f.harness.coordinator.capture(&lost, false))
|
||||
.await
|
||||
.unwrap();
|
||||
let outcome = within("durable", f.harness.coordinator.await_durable(ticket))
|
||||
.await
|
||||
.unwrap();
|
||||
assert_eq!(outcome, SaveOutcome::ReplyLost { checkpoint_id: lost.clone() });
|
||||
assert_eq!(
|
||||
f.harness.coordinator.durable(),
|
||||
None,
|
||||
"a lost save reply never moves the durable mark"
|
||||
);
|
||||
assert_eq!(count(&f.harness.coordinator.audit, &format!("durable:{lost}@1")), 0);
|
||||
|
||||
// The resolution asks the durable metadata about the *same* operation. It never saves
|
||||
// again, and here the bytes did reach the store manifest.
|
||||
let resolved = within("resolve", f.harness.coordinator.resolve_durable(&lost))
|
||||
.await
|
||||
.unwrap();
|
||||
assert_eq!(resolved, Some(1));
|
||||
assert_eq!(f.harness.coordinator.durable(), Some((lost.clone(), 1)));
|
||||
let commits = f
|
||||
.harness
|
||||
.coordinator
|
||||
.writer()
|
||||
.expect("a store")
|
||||
.with_store(|store| store.commits())
|
||||
.await;
|
||||
assert_eq!(commits, 1, "the resolution queried the store; it did not save again");
|
||||
f.shutdown().await;
|
||||
|
||||
// The other half: the generation file is renamed and the store manifest is never
|
||||
// committed. That generation is unreferenced, so it is not a candidate and the mark is
|
||||
// still where it was.
|
||||
let orphan = ckpt(2);
|
||||
let mut config = HarnessConfig::default();
|
||||
config.store_faults.stop_before_manifest_commit = true;
|
||||
let mut g = fx(via, mode, config).await;
|
||||
within("bootstrap", g.harness.coordinator.bootstrap()).await.unwrap();
|
||||
within("run", g.harness.coordinator.run(1)).await.unwrap();
|
||||
let outcome = within("checkpoint", g.harness.coordinator.checkpoint(&orphan))
|
||||
.await
|
||||
.unwrap();
|
||||
match outcome {
|
||||
SaveOutcome::Failed { checkpoint_id, .. } => assert_eq!(checkpoint_id, orphan),
|
||||
other => panic!("an uncommitted manifest is not a saved checkpoint: {other:?}"),
|
||||
}
|
||||
assert_eq!(g.harness.coordinator.durable(), None);
|
||||
let root = g.harness.checkpoint_root().to_path_buf();
|
||||
assert!(
|
||||
root.join(format!("{orphan}.flysess")).is_file(),
|
||||
"the generation file was written and renamed"
|
||||
);
|
||||
let resolved = within("resolve", g.harness.coordinator.resolve_durable(&orphan))
|
||||
.await
|
||||
.unwrap();
|
||||
assert_eq!(resolved, None, "an unreferenced generation is never a restore candidate");
|
||||
assert_eq!(g.harness.coordinator.durable(), None);
|
||||
g.shutdown().await;
|
||||
}
|
||||
|
||||
// ===============================================================================================
|
||||
// Acceptance: a failure during activation cannot resume half a world
|
||||
|
||||
async fn a_failure_during_activation_cannot_resume_half_a_world(via: Via) {
|
||||
half_an_activation(via, ExecutionMode::InProcess).await;
|
||||
}
|
||||
|
||||
async fn half_an_activation_resumes_nothing_in_every_mode(mode: ExecutionMode) {
|
||||
half_an_activation(Via::Unix, mode).await;
|
||||
}
|
||||
|
||||
/// The world and the first agent activate; the second refuses. Nothing plays, the fence
|
||||
/// stays closed, and the group cannot be resumed until every participant is replaced.
|
||||
async fn half_an_activation(via: Via, mode: ExecutionMode) {
|
||||
let mut f = fx(via, mode, HarnessConfig::default()).await;
|
||||
let checkpoint_id = ckpt(1);
|
||||
run_and_checkpoint(&mut f, BEFORE, &checkpoint_id).await;
|
||||
fail_the_epoch(&mut f).await;
|
||||
|
||||
// The replacement for fly-b refuses to activate after it has staged.
|
||||
f.harness.set_agent_faults(
|
||||
&fly_b(),
|
||||
AgentFaults { fail_activate_restore: true, ..AgentFaults::default() },
|
||||
);
|
||||
within("replace", f.harness.replace_all_participants()).await.unwrap();
|
||||
let failure = within(
|
||||
"restore",
|
||||
f.harness.coordinator.restore(Some(&checkpoint_id), &id("e2")),
|
||||
)
|
||||
.await
|
||||
.expect_err("a refused activation fails the install");
|
||||
assert_eq!(failure.participant.as_deref(), Some(fly_b().as_str()));
|
||||
assert_eq!(f.harness.coordinator.phase(), Phase::Failed);
|
||||
assert!(f.harness.coordinator.is_fenced(), "the group stays fenced");
|
||||
let refused = within("step", f.harness.coordinator.step())
|
||||
.await
|
||||
.expect_err("half a world never runs");
|
||||
assert_eq!(refused.error.code, ErrorCode::InvalidPhase);
|
||||
|
||||
// The participants that got as far as installing hold state no group resumed. Another
|
||||
// restore over them is refused by name rather than attempted.
|
||||
let tainted = f.harness.coordinator.tainted();
|
||||
assert!(tainted.contains(&f.harness.environment_id()), "{tainted:?}");
|
||||
assert!(tainted.contains(&fly_a()), "{tainted:?}");
|
||||
let refused = within(
|
||||
"restore",
|
||||
f.harness.coordinator.restore(Some(&checkpoint_id), &id("e3")),
|
||||
)
|
||||
.await
|
||||
.expect_err("the half-installed group is not restored over");
|
||||
assert_eq!(refused.error.code, ErrorCode::InvalidPhase);
|
||||
assert!(refused.error.message.contains(&fly_a()), "{refused}");
|
||||
|
||||
// A fresh group, without the injected refusal, resumes the same boundary.
|
||||
f.harness.set_agent_faults(&fly_b(), AgentFaults::default());
|
||||
within("replace", f.harness.replace_all_participants()).await.unwrap();
|
||||
assert!(f.harness.coordinator.tainted().is_empty());
|
||||
let report = within(
|
||||
"restore",
|
||||
f.harness.coordinator.restore(Some(&checkpoint_id), &id("e4")),
|
||||
)
|
||||
.await
|
||||
.unwrap();
|
||||
assert_eq!(report.boundary, BEFORE);
|
||||
assert!(!f.harness.coordinator.is_fenced());
|
||||
f.harness.coordinator.resume().unwrap();
|
||||
within("resume", f.harness.coordinator.run(1)).await.unwrap();
|
||||
f.shutdown().await;
|
||||
}
|
||||
|
||||
// ===============================================================================================
|
||||
// Acceptance: the checkpoint queue under stress stays bounded
|
||||
|
||||
async fn the_checkpoint_queue_under_stress_stays_bounded(via: Via) {
|
||||
queue_stays_bounded(via, ExecutionMode::InProcess).await;
|
||||
}
|
||||
|
||||
async fn the_queue_stays_bounded_in_every_mode(mode: ExecutionMode) {
|
||||
queue_stays_bounded(Via::Unix, mode).await;
|
||||
}
|
||||
|
||||
/// With the writer stalled, the queue fills to its configured bound and the next capture is
|
||||
/// refused *before* a single participant is asked for one.
|
||||
async fn queue_stays_bounded(via: Via, mode: ExecutionMode) {
|
||||
let gate = Arc::new(tokio::sync::Semaphore::new(0));
|
||||
let mut config = HarnessConfig::default();
|
||||
config.writer_faults.gate = Some(gate.clone());
|
||||
config.writer.queue_capacity = 2;
|
||||
let mut f = fx(via, mode, config).await;
|
||||
within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap();
|
||||
within("run", f.harness.coordinator.run(1)).await.unwrap();
|
||||
|
||||
let first = within("capture", f.harness.coordinator.capture(&ckpt(1), false))
|
||||
.await
|
||||
.unwrap();
|
||||
let second = within("capture", f.harness.coordinator.capture(&ckpt(2), false))
|
||||
.await
|
||||
.unwrap();
|
||||
assert_eq!(f.harness.coordinator.writer().expect("a store").outstanding(), 2);
|
||||
|
||||
let environment = f.harness.environment_id();
|
||||
let before = within("status", f.harness.progress_of(&environment)).await.unwrap();
|
||||
let refused = within("capture", f.harness.coordinator.capture(&ckpt(3), false))
|
||||
.await
|
||||
.expect_err("a full queue refuses");
|
||||
assert_eq!(refused.error.code, ErrorCode::Busy);
|
||||
assert_eq!(refused.error.mutation, MutationCertainty::None);
|
||||
let after = within("status", f.harness.progress_of(&environment)).await.unwrap();
|
||||
assert_eq!(
|
||||
before, after,
|
||||
"a refused capture asks no participant for one: rejection happens before capture"
|
||||
);
|
||||
assert_eq!(count(&f.harness.coordinator.audit, "capture-refused:ckpt-3"), 1);
|
||||
|
||||
// A refused capture is not a failed session: the world keeps stepping.
|
||||
assert!(!f.harness.coordinator.is_fenced());
|
||||
within("run", f.harness.coordinator.run(1)).await.unwrap();
|
||||
assert_eq!(f.harness.coordinator.phase(), Phase::Ready(2));
|
||||
|
||||
let stats = f.harness.coordinator.writer().expect("a store").stats();
|
||||
assert_eq!(stats.rejected, 1);
|
||||
assert!(stats.peak_queue <= 2, "the queue never exceeded its bound: {stats:?}");
|
||||
assert!(stats.peak_bytes <= 64 * 1024 * 1024);
|
||||
|
||||
gate.add_permits(16);
|
||||
for (ticket, expected) in [(first, ckpt(1)), (second, ckpt(2))] {
|
||||
let outcome = within("durable", f.harness.coordinator.await_durable(ticket))
|
||||
.await
|
||||
.unwrap();
|
||||
match outcome {
|
||||
SaveOutcome::Committed { checkpoint_id, .. } => assert_eq!(checkpoint_id, expected),
|
||||
other => panic!("an unstalled writer commits: {other:?}"),
|
||||
}
|
||||
}
|
||||
let stats = f.harness.coordinator.writer().expect("a store").stats();
|
||||
assert_eq!(stats.committed, 2);
|
||||
assert_eq!(f.harness.coordinator.writer().expect("a store").outstanding(), 0);
|
||||
f.shutdown().await;
|
||||
}
|
||||
|
||||
/// A queued replaceable capture is replaced by the next one, releasing its holds, rather than
|
||||
/// both being written.
|
||||
async fn a_queued_replaceable_capture_is_superseded_rather_than_duplicated(via: Via) {
|
||||
let gate = Arc::new(tokio::sync::Semaphore::new(0));
|
||||
let mut config = HarnessConfig::default();
|
||||
config.writer_faults.gate = Some(gate.clone());
|
||||
config.writer.queue_capacity = 3;
|
||||
let mut f = fx(via, ExecutionMode::InProcess, config).await;
|
||||
within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap();
|
||||
within("run", f.harness.coordinator.run(1)).await.unwrap();
|
||||
|
||||
// The first job is taken by the writer and stalls at the gate; the next two are queued.
|
||||
let durable = within("capture", f.harness.coordinator.capture(&ckpt(1), false))
|
||||
.await
|
||||
.unwrap();
|
||||
let hot = within("capture", f.harness.coordinator.capture(&ckpt(2), true))
|
||||
.await
|
||||
.unwrap();
|
||||
let newer = within("capture", f.harness.coordinator.capture(&ckpt(3), true))
|
||||
.await
|
||||
.unwrap();
|
||||
|
||||
let superseded = within("durable", f.harness.coordinator.await_durable(hot))
|
||||
.await
|
||||
.unwrap();
|
||||
assert_eq!(
|
||||
superseded,
|
||||
SaveOutcome::Superseded { checkpoint_id: ckpt(2), by: ckpt(3) },
|
||||
"only a queued replaceable capture is coalesced"
|
||||
);
|
||||
gate.add_permits(16);
|
||||
for ticket in [durable, newer] {
|
||||
let outcome = within("durable", f.harness.coordinator.await_durable(ticket))
|
||||
.await
|
||||
.unwrap();
|
||||
assert!(matches!(outcome, SaveOutcome::Committed { .. }), "{outcome:?}");
|
||||
}
|
||||
let stats = f.harness.coordinator.writer().expect("a store").stats();
|
||||
assert_eq!(stats.superseded, 1);
|
||||
assert_eq!(stats.committed, 2, "the superseded capture was never written");
|
||||
f.shutdown().await;
|
||||
}
|
||||
|
||||
// ===============================================================================================
|
||||
// Acceptance: old media and old parser data cannot cross a recovery
|
||||
|
||||
async fn old_media_cannot_cross_a_recovery(via: Via) {
|
||||
old_media_cannot_cross(via, ExecutionMode::InProcess).await;
|
||||
}
|
||||
|
||||
async fn old_media_cannot_cross_in_every_mode(mode: ExecutionMode) {
|
||||
old_media_cannot_cross(Via::Unix, mode).await;
|
||||
}
|
||||
|
||||
/// Nothing the fence dropped comes back: the payloads are re-imported as fresh artifacts, the
|
||||
/// restored view is a new object, and the new epoch's first audio chunk resumes the preserved
|
||||
/// sample position and marks the discontinuity.
|
||||
async fn old_media_cannot_cross(via: Via, mode: ExecutionMode) {
|
||||
let mut f = fx(via, mode, HarnessConfig::default()).await;
|
||||
let checkpoint_id = ckpt(1);
|
||||
run_and_checkpoint(&mut f, BEFORE, &checkpoint_id).await;
|
||||
let old_ids = live_artifact_ids(&f);
|
||||
assert!(!old_ids.is_empty(), "the committed boundary holds media handles");
|
||||
let positions = f.harness.coordinator.audio_positions();
|
||||
assert!(positions.values().any(|sample| *sample > 0), "audio has played");
|
||||
|
||||
fail_the_epoch(&mut f).await;
|
||||
within("replace", f.harness.replace_all_participants()).await.unwrap();
|
||||
let report = within(
|
||||
"restore",
|
||||
f.harness.coordinator.restore(Some(&checkpoint_id), &id("e2")),
|
||||
)
|
||||
.await
|
||||
.unwrap();
|
||||
for imported in &report.imported {
|
||||
assert!(
|
||||
!old_ids.contains(imported),
|
||||
"a checkpoint payload was imported as an artifact the old epoch already had"
|
||||
);
|
||||
}
|
||||
let new_ids = live_artifact_ids(&f);
|
||||
assert!(!new_ids.is_empty());
|
||||
assert!(
|
||||
new_ids.is_disjoint(&old_ids),
|
||||
"the restored boundary's media are fresh artifacts, not the old epoch's"
|
||||
);
|
||||
assert_eq!(
|
||||
f.harness.coordinator.audio_positions(),
|
||||
positions,
|
||||
"crash restore preserves the sample position"
|
||||
);
|
||||
|
||||
f.harness.coordinator.resume().unwrap();
|
||||
within("resume", f.harness.coordinator.run(1)).await.unwrap();
|
||||
let chunk = f
|
||||
.harness
|
||||
.coordinator
|
||||
.observation()
|
||||
.expect("a restored world")
|
||||
.audio
|
||||
.first()
|
||||
.cloned()
|
||||
.expect("one chunk per transition");
|
||||
assert!(
|
||||
chunk.discontinuity,
|
||||
"the first chunk of a fresh epoch marks the discontinuity the recovery established"
|
||||
);
|
||||
assert_eq!(
|
||||
chunk.first_sample,
|
||||
positions["arena"],
|
||||
"and it starts where the checkpointed stream stopped"
|
||||
);
|
||||
f.shutdown().await;
|
||||
}
|
||||
|
||||
async fn a_checkpoint_taken_by_another_parser_is_refused_by_name(via: Via) {
|
||||
another_parser_is_refused(via, ExecutionMode::InProcess).await;
|
||||
}
|
||||
|
||||
async fn another_parser_is_refused_in_every_mode(mode: ExecutionMode) {
|
||||
another_parser_is_refused(Via::Unix, mode).await;
|
||||
}
|
||||
|
||||
/// A checkpoint whose recorded parser identity is not this composition's is refused before a
|
||||
/// single participant is asked to stage, and the refusal names the identity that differs.
|
||||
async fn another_parser_is_refused(via: Via, mode: ExecutionMode) {
|
||||
let mut f = fx(via, mode, HarnessConfig::default()).await;
|
||||
let checkpoint_id = ckpt(1);
|
||||
run_and_checkpoint(&mut f, 1, &checkpoint_id).await;
|
||||
let root = f.harness.checkpoint_root().to_path_buf();
|
||||
fail_the_epoch(&mut f).await;
|
||||
rewrite_generation(&root, &checkpoint_id, |manifest, _payloads| {
|
||||
manifest["compatibility"]["parserDigest"] =
|
||||
json!(digest_of_bytes(b"some other inspection schema"));
|
||||
});
|
||||
reload_store(&mut f).await;
|
||||
within("replace", f.harness.replace_all_participants()).await.unwrap();
|
||||
let failure = within(
|
||||
"restore",
|
||||
f.harness.coordinator.restore(Some(&checkpoint_id), &id("e2")),
|
||||
)
|
||||
.await
|
||||
.expect_err("another parser's state is not this composition's");
|
||||
assert_eq!(failure.error.code, ErrorCode::IncompatibleState);
|
||||
assert!(
|
||||
failure.error.message.contains("parser"),
|
||||
"the refusal names the identity that differs: {failure}"
|
||||
);
|
||||
assert!(
|
||||
f.harness.coordinator.tainted().is_empty(),
|
||||
"nothing was asked to stage, so nothing has to be replaced"
|
||||
);
|
||||
nothing_is_installed(&mut f, &id("e2")).await;
|
||||
f.shutdown().await;
|
||||
}
|
||||
|
||||
// ===============================================================================================
|
||||
// Failure-injection rows
|
||||
|
||||
async fn a_group_where_one_participant_will_not_stage_resumes_nothing(via: Via) {
|
||||
a_refused_stage(via, ExecutionMode::InProcess).await;
|
||||
}
|
||||
|
||||
async fn a_refused_stage_resumes_nothing_in_every_mode(mode: ExecutionMode) {
|
||||
a_refused_stage(Via::Unix, mode).await;
|
||||
}
|
||||
|
||||
/// The checklist row: the install validates the participants ahead of it and the last one
|
||||
/// fails. Nothing is resumed.
|
||||
async fn a_refused_stage(via: Via, mode: ExecutionMode) {
|
||||
let mut f = fx(via, mode, HarnessConfig::default()).await;
|
||||
let checkpoint_id = ckpt(1);
|
||||
run_and_checkpoint(&mut f, 1, &checkpoint_id).await;
|
||||
fail_the_epoch(&mut f).await;
|
||||
f.harness.set_agent_faults(
|
||||
&fly_b(),
|
||||
AgentFaults { fail_stage_restore: true, ..AgentFaults::default() },
|
||||
);
|
||||
within("replace", f.harness.replace_all_participants()).await.unwrap();
|
||||
let failure = within(
|
||||
"restore",
|
||||
f.harness.coordinator.restore(Some(&checkpoint_id), &id("e2")),
|
||||
)
|
||||
.await
|
||||
.expect_err("a participant that will not validate stops the install");
|
||||
assert_eq!(failure.error.code, ErrorCode::IncompatibleState);
|
||||
assert_eq!(failure.participant.as_deref(), Some(fly_b().as_str()));
|
||||
assert_eq!(
|
||||
count(&f.harness.coordinator.audit, &format!("activated:{}", fly_a())),
|
||||
0,
|
||||
"no participant activates when one of them will not stage"
|
||||
);
|
||||
nothing_is_installed(&mut f, &id("e2")).await;
|
||||
f.shutdown().await;
|
||||
}
|
||||
|
||||
async fn a_restore_token_activates_only_once(via: Via) {
|
||||
a_token_activates_once(via, ExecutionMode::InProcess).await;
|
||||
}
|
||||
|
||||
async fn a_token_activates_once_in_every_mode(mode: ExecutionMode) {
|
||||
a_token_activates_once(Via::Unix, mode).await;
|
||||
}
|
||||
|
||||
/// A token is bound to its checkpoint, scope and payload and activates once. A fresh request
|
||||
/// naming it again is a conflict, and the old epoch's scope is stale on the new participants.
|
||||
async fn a_token_activates_once(via: Via, mode: ExecutionMode) {
|
||||
let mut f = fx(via, mode, HarnessConfig::default()).await;
|
||||
let checkpoint_id = ckpt(1);
|
||||
run_and_checkpoint(&mut f, 1, &checkpoint_id).await;
|
||||
fail_the_epoch(&mut f).await;
|
||||
within("replace", f.harness.replace_all_participants()).await.unwrap();
|
||||
let report = within(
|
||||
"restore",
|
||||
f.harness.coordinator.restore(Some(&checkpoint_id), &id("e2")),
|
||||
)
|
||||
.await
|
||||
.unwrap();
|
||||
let (who, token) = report.tokens.first().cloned().expect("a staged token");
|
||||
let worker = if who == f.harness.environment_id() {
|
||||
f.harness.coordinator.environment_ref().clone()
|
||||
} else {
|
||||
f.harness.coordinator.agent_ref(&who).expect("a participant").clone()
|
||||
};
|
||||
let scope = scope_at("demo", "e2", 1);
|
||||
let refused = within(
|
||||
"activate",
|
||||
f.harness.coordinator.probe_raw(
|
||||
&worker,
|
||||
"State.ActivateRestore",
|
||||
Some(scope),
|
||||
json!({"restoreToken": token.as_str()}),
|
||||
),
|
||||
)
|
||||
.await
|
||||
.expect_err("a token activates once");
|
||||
assert_eq!(refused.code, ErrorCode::Conflict);
|
||||
|
||||
// And the epoch the restore left behind is stale on every replacement.
|
||||
let stale = within(
|
||||
"prepare",
|
||||
f.harness.coordinator.probe_raw(
|
||||
&worker,
|
||||
if who == f.harness.environment_id() { "Environment.Advance" } else { "Agent.Prepare" },
|
||||
Some(scope_at("demo", "e1", 1)),
|
||||
json!({}),
|
||||
),
|
||||
)
|
||||
.await
|
||||
.expect_err("the old epoch is gone");
|
||||
assert!(
|
||||
matches!(stale.code, ErrorCode::StaleEpoch | ErrorCode::InvalidArgument),
|
||||
"an old-epoch request is refused: {stale}"
|
||||
);
|
||||
f.shutdown().await;
|
||||
}
|
||||
|
||||
async fn the_fence_lifts_only_through_a_coherent_restore(via: Via) {
|
||||
the_fence_lifts_only_by_restore(via, ExecutionMode::InProcess).await;
|
||||
}
|
||||
|
||||
async fn the_fence_lifts_only_by_restore_in_every_mode(mode: ExecutionMode) {
|
||||
the_fence_lifts_only_by_restore(Via::Unix, mode).await;
|
||||
}
|
||||
|
||||
/// Nothing but a coherent restore moves a fenced session, and the restore lands on
|
||||
/// `Paused(k)` rather than straight back into play.
|
||||
async fn the_fence_lifts_only_by_restore(via: Via, mode: ExecutionMode) {
|
||||
let mut f = fx(via, mode, HarnessConfig::default()).await;
|
||||
let checkpoint_id = ckpt(1);
|
||||
run_and_checkpoint(&mut f, 1, &checkpoint_id).await;
|
||||
fail_the_epoch(&mut f).await;
|
||||
|
||||
for refused in [
|
||||
within("step", f.harness.coordinator.step()).await.err(),
|
||||
within("bootstrap", f.harness.coordinator.bootstrap()).await.err(),
|
||||
within("capture", f.harness.coordinator.capture(&ckpt(9), false))
|
||||
.await
|
||||
.err(),
|
||||
] {
|
||||
let refused = refused.expect("a fenced session refuses");
|
||||
assert_eq!(refused.error.code, ErrorCode::InvalidPhase);
|
||||
}
|
||||
assert!(f.harness.coordinator.resume().is_err(), "a fenced session does not resume");
|
||||
assert!(f.harness.coordinator.is_fenced());
|
||||
|
||||
// A restore into the epoch that failed is refused: recovery establishes a fresh one.
|
||||
within("replace", f.harness.replace_all_participants()).await.unwrap();
|
||||
let same = within(
|
||||
"restore",
|
||||
f.harness.coordinator.restore(Some(&checkpoint_id), &id("e1")),
|
||||
)
|
||||
.await
|
||||
.expect_err("a restore installs a fresh epoch");
|
||||
assert_eq!(same.error.code, ErrorCode::StaleEpoch);
|
||||
|
||||
let report = within(
|
||||
"restore",
|
||||
f.harness.coordinator.restore(Some(&checkpoint_id), &id("e2")),
|
||||
)
|
||||
.await
|
||||
.unwrap();
|
||||
assert_eq!(report.boundary, 1);
|
||||
assert!(!f.harness.coordinator.is_fenced());
|
||||
assert_eq!(f.harness.coordinator.phase(), Phase::Paused(1));
|
||||
f.harness.coordinator.resume().unwrap();
|
||||
within("resume", f.harness.coordinator.run(1)).await.unwrap();
|
||||
assert_eq!(f.harness.coordinator.phase(), Phase::Ready(2));
|
||||
f.shutdown().await;
|
||||
}
|
||||
|
||||
/// A capture belongs to a committed boundary, and the phase machine is what says so.
|
||||
async fn a_capture_is_refused_anywhere_but_a_committed_boundary(via: Via) {
|
||||
let mut f = fx(via, ExecutionMode::InProcess, HarnessConfig::default()).await;
|
||||
// Before bootstrap the session is Starting, which is not a boundary at all.
|
||||
let refused = within("capture", f.harness.coordinator.capture(&ckpt(1), false))
|
||||
.await
|
||||
.expect_err("Starting is not a committed boundary");
|
||||
assert_eq!(refused.error.code, ErrorCode::InvalidPhase);
|
||||
f.shutdown().await;
|
||||
|
||||
// From a pause, which is a committed boundary, it works, and it returns there.
|
||||
let mut f = fx(via, ExecutionMode::InProcess, HarnessConfig::default()).await;
|
||||
within("bootstrap", f.harness.coordinator.bootstrap()).await.unwrap();
|
||||
within("run", f.harness.coordinator.run(1)).await.unwrap();
|
||||
f.harness.coordinator.request_pause();
|
||||
within("run", f.harness.coordinator.run(1)).await.unwrap();
|
||||
assert_eq!(f.harness.coordinator.phase(), Phase::Paused(2));
|
||||
let outcome = within("checkpoint", f.harness.coordinator.checkpoint(&ckpt(1)))
|
||||
.await
|
||||
.unwrap();
|
||||
assert!(matches!(outcome, SaveOutcome::Committed { boundary: 2, .. }), "{outcome:?}");
|
||||
assert_eq!(
|
||||
f.harness.coordinator.phase(),
|
||||
Phase::Paused(2),
|
||||
"a capture returns to the boundary it came from"
|
||||
);
|
||||
f.shutdown().await;
|
||||
}
|
||||
Loading…
Add table
Reference in a new issue