flybrain/docs/design/session-framework/state-media-v1.md
dev 6655a1b1c6 session: coherent all-participant checkpoint and recovery
STATE-01 over the FLYSESS1 envelope CONTRACT-01 specified.

fly-session gains a `state` module: the durable store with its generations, its
rotation and the commit order of checkpoint-envelope-v1 section 5, where the store
manifest rename is the durable commit point; a compatibility block whose comparison
names the identity that differs rather than one opaque digest; and a bounded writer
that owns its payload handles until the bytes are committed or the job fails.

The writer's queue slot is taken before the first State.Capture, so a saturated
writer refuses a capture rather than queueing it without bound, and the refusal is a
BUSY the stepping session survives. Capture and durability are two events: a capture
completes when an immutable capture exists, and only the store manifest rename moves
the durable mark. A lost save reply is an outcome, and the resolution asks the store
about the same checkpoint instead of saving again.

Both worker roles implement State.Capture, State.StageRestore and
State.ActivateRestore, with once-only restore tokens bound to checkpoint, scope,
payload and incarnation. A restore selects a complete compatible generation, imports
every payload as a fresh artifact, stages the group, validates the coordinator's own
ledgers, and only then activates; a failure anywhere leaves the fence closed and
records every participant that staged as one that must be replaced. The fence lifts
at Failed -> Restoring(k) -> Paused(k) and nowhere else.

The task and the action executor gain the capture/validate_restore/install_restore
interfaces workers-v1 section 4 lists, and the ledger can re-derive the event
identities it issued under another epoch, which is what lets a resumed run's
behaviour trace be compared with an uninterrupted one.

media: check_required_audio now takes the observation's provenance instead of
exempting boundary 0. A chunk is the audio of an interval, and the observation
ActivateRestore installs covers none.

checkpoint-envelope-v1 section 3 gains a dated amendment adding `environment` to the
manifest, the holder of the world's own payload, which the table named for every
other participant; `helperState`, which that table already listed, joins the
required-field set in Rust and TypeScript. The fixture was regenerated by the
existing example; the schema set and contractDigest are unchanged.

state-media-v1 section 5 gains a dated amendment for three readings this slice
enforces: the State RPCs' compatibilityDigest is the participant's, not the
manifest's composition-level block; a restored observation carries no audio chunk;
and a participant that staged into an abandoned install must be replaced.
2026-09-22 17:43:43 +00:00

15 KiB
Raw Blame History

Session artifacts, native media and recovery

Status: draft 2, 2026-09-18. Flybus owns generic artifact storage, delivery ownership, retention and garbage collection. This document specifies the domain meaning of those artifacts: native observations, clock association and coherent session checkpoints. Read session RPC, step ordering and worker interfaces.

1. Use the bus ArtifactRef

Large payload fields use ArtifactRef from bus-v1, and every referenced artifact is listed in the surrounding bus attachments. There is no separate BinaryRef, buffer-region registry, coordinator lease endpoint, or Buffer.Release/Reclaim protocol. Profile/dataset release assets use AssetRef (a persistent content identity); transient bus ArtifactRefs are not those assets.

The environment publishes one immutable native image. The coordinator may forward the same owned handle to multiple agent Commit calls and publish it for presentation. Flybus creates destination ownership before releasing the source. It does not send another copy of the pixel bytes per recipient through its sockets.

An agent consumes/encodes pixels during Initialize/Commit and drops its handle when no longer used. A renderer may keep its extracted handle after dropping the message; the DeliveryGuard keeps the artifact alive until rendering has finished. A domain cached RPC result keeps its own handles so replay remains valid after the original recipient consumes its delivery.

Content digests are optional on transient live frames, mandatory on checkpoint payloads and persistent asset import. Ownership/index/byte-shape validation is always required. A digest does not replace epoch or observation-time identity.

2. Native observation types

interface ViewDescriptor {
  viewId: Id;
  width: number; height: number;
  format: "rgba8"; rowStride: number;
  pixelAspect: { numerator: number; denominator: number };
  observationDelaySteps: number;
}
interface ViewRef { viewId: Id; producedStep: U64; pixels: ArtifactRef }
interface AudioDescriptor {
  streamId: Id; sampleRate: number; channels: number;
  format: "f32le-interleaved";
}
interface AudioRef {
  streamId: Id; firstSample: U64; sampleFrames: number;
  samples: ArtifactRef; discontinuity: boolean;
}

Epoch is inherited from the domain observation; the bus treats it as opaque payload. View dimensions are integers 1..4096, rowStride exactly 4×width, no padded rows in v1. Pixel aspect numerator/denominator are positive integers <=65535; observationDelaySteps is integer 0..8. Pixels are top-left RGBA8 and artifact length equals rowStride×height. Other formats require a media-schema change, not special-case code inside the router.

Required sensory views have producedStep equal to max(0, observation.boundary - observationDelaySteps). Bootstrap may repeat O[0] until the declared pipeline delay fills. Beyond that, missing/extra-delay sensory input is a step failure, not an arbitrary latest frame. Observer publication may omit/coalesce frames while preserving each artifact's actual producing boundary.

Audio sampleRate is integer 8000..192000, channels 1..8, sampleFrames 0..192000 per chunk. Samples are finite f32; artifact length is sampleFrames×channels×4. firstSample identifies the sample position relative to the episode's configured audio origin, with intended PTS firstSample/sampleRate. Crash restore preserves sample position under a new epoch; first chunk marks discontinuity. Within an epoch, chunks cannot overlap or go backwards.

Amendment, 2026-09-22 (MEDIA-01). Two readings of the paragraphs above, made explicit because they are now enforced:

  • The bootstrap window is exactly the boundaries where max(0, boundary - observationDelaySteps) is zero, that is boundary <= observationDelaySteps. Inside it the repeated O[0] is the same artifact, not a fresh render of the same scene; outside it the producing boundary advances one per step, and a frame from any other boundary -- older or newer -- is a step failure. A producer therefore keeps a queue of observationDelaySteps + 1 frames and nothing more, so there is no older frame available to substitute.
  • Within an epoch, discontinuity marks a range the stream actually skipped. The first chunk after a restore marks it, and a later chunk may mark it when it starts past where the previous chunk ended; a chunk that continues the previous one exactly is continuous by construction and its flag is refused. Without that reading the restore rule is advisory, because a stream could set the flag on every chunk and satisfy it by accident. The requirement is one-directional: a fresh epoch's first chunk may mark a discontinuity, because section 6's recovery establishes a fresh timeline and publishes one.

The environment provides native game output. Sensor transformations belong to the agent profile. Resizing for viewers, overlays, composition, audio mixing/resampling, encoding, browser delivery and streaming belong to the application/presentation layer. No bus or generic session configuration assumes a 1080p show or Twitch output.

640×480 RGBA at 60 fps produces 73.728 MB/s of raw image data. Artifact fan-out references one stored object; reads and any staging/seal copy still consume memory bandwidth. This is reasonable to measure before introducing codecs or pooled GPU buffers. Native dimensions come from the backend, not a hardcoded GameCube or broadcast resolution.

3. Domain retention and backpressure

Use the same bus call/publish API for observations and artifacts. Bus ownership tracks bytes; the session decides which observations are required and when they have been used.

Use Rule
Required agent input Retain through encoding/Commit; no coalescing or overwrite
Step-result replay Retain in the endpoint's current/previous-step cache until domain eviction
Spectator snapshot Latest subscription, finite in-flight credits; release after actual use
Long rendering/storage job Explicit artifact hold with a finite byte/count budget
Hot checkpoint Coalesce only queued replaceable captures, releasing their holds
Durable checkpoint Acknowledge after durable commit; reject/defer before capture when saturated

Initial session defaults: two outstanding coherent captures and at most the bus-configured latest/in-flight frame credits per observer. Presentation audio can target 250 ms and cap at one second, but that is a presentation policy, not a bus or brain-clock requirement.

Budget cached step observations, active agent deliveries, retained latest and spectator holds together. The producer dropping its handle does not free cached/queued/in-use objects. A slow spectator exhausts its own credits; new latest messages replace its queued value. If it violates configured resource policy, disconnect/restart that observer instead of freeing live data or silently skipping simulation input. Global store exhaustion is an explicit fault or pause condition; the router cannot guess that a particular live object is disposable.

No coordinator tracks per-reader socket acknowledgments or calls a producer's reclaim method. The SDK and bus perform that bookkeeping. File-backed immutable mappings are safe after unlink; physical pages disappear when all OS mappings close. Pooled reuse is deferred until it provides equivalent safety. Bare handles in history are not persistent saved bytes.

4. Checkpoint identity and content

Exact-checkpoint sessions capture at Ready(k) or Paused(k), after all Agent.Commit replies, with no Advance in flight. Block next Prepare until all participants supply immutable captures.

The manifest records:

  • Envelope version, checkpoint ID and source session/epoch/step/episode/world time.
  • Coordinator scheduler/configuration identity and exact port-to-agent map.
  • Backend/content/patch/controller/parser/state-format compatibility.
  • Per-agent profile/dataset/model identities, seed, tick count/remainder and payload digests.
  • Task ledger, prior world inspection, per-agent executor state, next sensory/decision state or reproducible reconstruction inputs, admission state and event watermarks.
  • Payload names, lengths and hashes, including external-helper state required for exact resume.

Use a new envelope version; specify exact byte layout before production files. The historical letter-only chunk-name constraint is not silently widened, and FLYSIM01 remains separately readable. Persist payload bytes and durable content identity, not transient bus storeId, artifact IDs, ownership tokens, mappings or pointers.

A checkpoint writer owns the bus Artifact handles until bytes are committed or the job fails. It then drops them; durable files are outside Flybus's ephemeral GC. On restore, the durable store imports fresh immutable bus artifacts. sourceScope is provenance, while the new handles belong to the current router/store. Broker retention is never a substitute for a checkpoint.

5. State RPCs over the bus

Common to workers advertising checkpoint-v1:

interface CaptureParams { checkpointId: Id }
interface CaptureResult {
  checkpointId: Id; boundary: U64;
  compatibilityDigest: Digest;
  payload: ArtifactRef; // listed attachment; digest required
}
interface StageRestoreParams {
  checkpointId: Id; sourceScope: Scope;
  compatibilityDigest: Digest;
  payload: ArtifactRef; // newly imported owned attachment
}
interface StageRestoreResult { checkpointId: Id; restoreToken: Id }
interface ActivateRestoreParams { restoreToken: Id }
interface ActivateRestoreResult {
  committedStep: U64; checkpointId: Id;
  observation: WorldObservation | null; // environment required, agent null
}

State.Capture uses the committed scope. It completes after immutable capture exists, not when a backend save was requested. The cached reply retains its artifact until Worker.Acknowledge. The coordinator/writer obtains its own live ownership before acknowledging that cache.

State.StageRestore uses a proposed new epoch at source boundary k and is allowed only on an uninitialized replacement or a quiescent worker. Launcher configuration supplies the expected profile/backend identities; no implicit warm-up/reinitialization changes the saved brain. It validates into replacement state, without exposing mutations to the live session.

After every participant and coordinator state validates, State.ActivateRestore installs each staged token under that new scope without advancing a tick. Tokens are bound to scope/ payload/checkpoint and can activate only once; duplicate domain requests replay the cached reply, while a fresh request trying to reuse an activated token is a conflict.

The environment returns the coherent restored observation with fresh artifact references and restored time. It cannot advance gameplay to manufacture it. Capture/reconstruction therefore covers render/inspection state and any pending sensor pipeline. Agent state agrees with it; do not replay reward or recalibrate merely to fill missing cached data.

Amendment, 2026-09-22 (STATE-01). Three readings of this section, made explicit because they are now enforced:

  • compatibilityDigest on CaptureResult and StageRestoreParams is the participant's capture compatibility digest of worker interfaces section 2 -- profile, resolved seed, numerical model version and effective instance configuration for an agent; backend, content, patch, controller and parser identity for an environment. It is not the manifest's compatibility block of section 4, which is the composition's and which the coordinator compares before anything is asked to stage. Both exist because they answer different questions, and a restore that passed the second could still be handing an agent another agent's brain.
  • The observation ActivateRestore returns ran no transition, so it carries no audio chunk, and one in it is refused. Section 2's chunk is the audio of an interval and this observation covers none; MEDIA-01 implemented that rule as "boundary 0 carries no chunk", which is true of the only such observation that slice could produce and false of this one. The rule is about provenance, not about the boundary number.
  • A participant that staged into a group install the coordinator then abandoned must be replaced before another restore, exactly as one that activated must. It is holding a validated replacement state that nothing installed, and session RPC section 6 already refuses to silently reattach such a participant to an active epoch. Without this the group's second attempt meets its own leftovers and calls them a conflict.

If emulator validation requires mutation, stage a stopped replacement emulator. If that cannot provide externally atomic resume, advertise episode-restart, not exact-checkpoint. After all activation acknowledgments, install the coordinator's staged task/executor/admission state and establish Paused(new epoch,k). Failure during activation never permits half a group to run.

6. Durable commit, router failure and recovery

Write payload/envelope temporary generation, fsync, rename, fsync directory, then atomically write/fsync/rename the store manifest and fsync its directory. Manifest commit is the durable commit point. Unreferenced temporary generations are not automatic restore candidates.

Publish distinct captured/queued/committed/failed/superseded events over the same bus. Only durable completion produces a saved acknowledgment/high-water mark. Failed writes release owned ephemeral captures according to retry policy, without reporting false durability.

After participant/coordinator/router failure:

  1. Stop steps, abandon the epoch and fence old participants/routes.
  2. Connect to a live router and select a complete compatible durable checkpoint.
  3. Import its payloads as new artifacts; stage/activate every participant and coordinator.
  4. Verify identity/boundary, flush old media/parser queues and publish recovery/discontinuity.
  5. Establish Paused(k), then resume only after the group invariant holds.

A router restart loses ephemeral topics, queues, roots and correlations. Continuing with old handles is invalid even if some mapped bytes survived. Reconnect is not transparent mid-step recovery. The durable log records old/new epochs and abandoned step ranges; rollback can lose post-checkpoint work. Durable input replay requires a separate application/session journal policy, not an exactly-once claim about Flybus.

7. Episode reset

Reset differs from crash restore. The application selects a policy; the coordinator records the old episode's result/abort and creates a new epoch/episode at step zero. World initial state and retained/fresh brain components are explicit. Gain retention, eligibility/hold clearing, calibration and first sensory input are part of the policy, tested independently. Legacy Pokémon ratchet behavior remains in the legacy composition. Shared competitive worlds never restore one player's environment independently of the other players.