From d18bc043b35273d1b91f8131fc346daafa902e32 Mon Sep 17 00:00:00 2001 From: acamilo Date: Fri, 18 Sep 2026 09:40:03 -0400 Subject: [PATCH] docs(design): plan MaleCNS, modular sessions and Melee integration (cherry picked from commit 83090a9391aac4044cfc48958fef54c0869a0e7b) --- docs/design/malecns-modular-implementation.md | 247 +++++ docs/design/malecns-modular-sessions.md | 892 ++++++++++++++++++ docs/design/melee-framework-audit.md | 747 +++++++++++++++ 3 files changed, 1886 insertions(+) create mode 100644 docs/design/malecns-modular-implementation.md create mode 100644 docs/design/malecns-modular-sessions.md create mode 100644 docs/design/melee-framework-audit.md diff --git a/docs/design/malecns-modular-implementation.md b/docs/design/malecns-modular-implementation.md new file mode 100644 index 0000000..12b9de5 --- /dev/null +++ b/docs/design/malecns-modular-implementation.md @@ -0,0 +1,247 @@ +# Implementation backlog: MaleCNS and modular sessions + +Status: **planned, not started**. Written 2026-09-18. Companion to the +[design and code analysis](malecns-modular-sessions.md), based on `f7bc13a`. + +This is the execution order for that proposal. It does not authorize deployment or a live +stream. Reconcile the baseline with merged macro/shop/recovery work before implementation. +Existing feed/control contracts and TypeScript oracle rules remain binding. + +Concrete second-game audit: [Melee framework and emulator plan](melee-framework-audit.md). +Its MELEE-01/02 spikes specialize EMULATOR-01 below and can proceed alongside framework +extraction; they do not depend on importing MaleCNS first. + +## 1. Delivery strategy + +Deliver working vertical slices; keep the existing FAFB/Game Boy composition usable throughout. + +1. Establish a behavior baseline and explicit dataset/profile identities. +2. In parallel workstreams, characterize MaleCNS and extract the single-agent session runtime. +3. Demonstrate two isolated brains driving one ROM-free shared environment. +4. Expose that session through versioned feed/control contracts and a multi-agent broadcast. +5. Integrate a specifically chosen alternative emulator after its capability spike passes. +6. Finish physical package reorganization once the second consumer proves the boundaries. + +**First milestone:** a reproducible headless MaleCNS run and a reusable single-agent session. +**Second milestone:** two flies in a synthetic arena with coherent resume and a local broadcast. +**Third milestone:** two flies in the chosen fighting game, with documented task scaffolding. + +No elapsed-time estimate is committed before the import and emulator spikes establish their +unknowns. Split an item further when its contract and implementation cannot be reviewed together. + +## 2. Work queue + +Every item starts pending. Branch names are suggested implementation branches, not branches +already created. Builders use separate worktrees; the coordinator reviews contracts and results. + +### FOUNDATION-01 — Pin existing behavior + +- **Branch:** `test/session-baseline` +- **Depends on:** reconciliation with current main and related open work. +- Record effective legacy configuration, fingerprint, version strings and frame ordering. +- Add a ROM-free transition harness around the service's frame orchestration; capture brain + ticks, decoded actions, reward application and recovery effects with explicit clocks. +- Distinguish exact numerical replay from intentional macro/hold/transient reset on restore. +- **Done:** existing goldens pass; a trace fixture detects reordered vision/reward application; + legacy feed/API fixtures and compatibility identity are unchanged. + +### FOUNDATION-02 — Specify bundle and behavior identities + +- **Branch:** `feat/brain-profile-contract` +- **Depends on:** FOUNDATION-01. +- Specify dataset manifests, original-ID mapping, anatomical roles, sensory bindings, + readout bindings and composite behavior identity in a focused contract. +- Add profile resolution/validation around the existing core in TypeScript and Rust. +- Keep schema-1 fingerprinting and the legacy macro-role exception behind the legacy path. +- Add strict graph validation for new bundles, shared invalid fixtures and profile mismatch tests. +- **Done:** empty required populations, malformed CSR and incorrect profile restores fail; + legacy FAFB artifacts and default numerical version strings remain unchanged. + +### DATA-01 — Acquire and normalize MaleCNS + +- **Branch:** `feat/malecns-import` +- **Depends on:** FOUNDATION-02. +- Build a source-specific importer from official v1.0 tables with checksummed source locks. +- Reconcile the Codex versus neuPrint inventories, or explicitly select and document one. +- Preserve raw contact counts, transmitter evidence, original IDs and missing-data indicators. +- Emit deterministic graph bundles, weight≥1/weight≥5 comparison variants, and an exclusion/ + clipping/coverage report. Decide schema-1 feasibility from measured weight ranges. +- Add attribution and license records with actual artifacts; keep raw downloads out of git. +- **Done:** repeated builds match; endpoint/role/index invariants pass; both loaders agree; + no service startup or ordinary unit test needs a download. + +### DATA-02 — Characterize MaleCNS in the current kernel + +- **Branch:** `feat/malecns-baseline-profile` +- **Depends on:** DATA-01. +- Audit L1 geometry, hemisphere handling, KC/MBON/PAM mappings and brain-versus-VNC motor roles. +- Define a versioned fixed readout; keep task action partitions out of anatomical truth. +- Run learning-off first, then learning-on, using fixed sensory traces and multiple seeds. +- Generate TS reference goldens and compare Rust exactly; measure activity, saturation, + initialization, memory and per-phase latency for both graph thresholds. +- **Done:** publish a reproducible characterization report and profile choice. Stop task + integration if required mappings are missing or dynamics are unusable; any recalibration + becomes a named profile rather than an edit to the legacy model. + +### RUNTIME-01 — Separate environment execution from task interpretation + +- **Branch:** `refactor/environment-task-boundary` +- **Depends on:** FOUNDATION-02. +- Specify controller ports, digital/analog controls, rational cadence, media descriptors, + observation ownership and backend capabilities. +- Wrap binjgb as the first environment; retain Game Boy FFI/cache/state behavior. +- Keep Pokémon memory inspection, objective routing, macros and reward rules in its task. +- Preserve existing imports through a facade; avoid simultaneous directory moves. +- **Done:** existing single-agent action/reward traces match and a fake environment can be + driven through the same boundary without importing binjgb or task-specific addresses. + +### RUNTIME-02 — Extract the single-agent session + +- **Branch:** `refactor/session-runtime` +- **Depends on:** RUNTIME-01. +- Move deterministic agent/environment/task orchestration out of `Sim` into a library. +- Keep HTTP, WebSocket serialization, wall-clock publication and process supervision in flysim. +- Give session clock, action executor, task ledger and recovery state explicit owners. +- Wrap the existing composition with legacy ordering, checkpoint and reset semantics. +- **Done:** headless consumer runs a session without Twitch/browser; the legacy composition + passes its traces and restore tests; a slow snapshot consumer cannot stall simulation. + +### RUNTIME-03 — Add synchronized multi-agent sessions + +- **Branch:** `feat/multi-agent-arena` +- **Depends on:** RUNTIME-02. +- Implement a ROM-free two-player arena and per-port controller ownership. +- Evaluate both brains against one observation boundary; apply one complete action batch; + advance the world once. Start sequentially, then verify parallel execution equivalence. +- Isolate RNG, stimulation, decoder holds, gains, traces and rewards per agent; share only + immutable topology. Enforce a total worker budget and single-dispatcher pool ownership. +- Define participant failure, lateness and episode reset policies. +- **Done:** no cross-agent state leakage; swapping evaluation order leaves results unchanged; + one failed participant cannot accidentally advance a half-controlled match. + +### STATE-01 — Capture and resume whole sessions + +- **Branch:** `feat/session-checkpoints` +- **Depends on:** RUNTIME-03; specify the state contract during RUNTIME-02. +- Define the new envelope/manifest and preserve the `FLYSIM01` reader. +- Capture all agents, environment, task/executor/admission state and clock remainders at one + boundary. Bound off-thread write jobs and retain atomic manifest commit semantics. +- Validate all components before installing any restored state; define external-backend staging. +- **Done:** uninterrupted and resumed synthetic matches agree; corrupting any participant + refuses the generation without partial restore; crash-injection fallback tests pass. + +### WIRE-01 — Introduce session feed/control v2 + +- **Branch:** `feat/session-protocol-v2` +- **Depends on:** FOUNDATION-02, RUNTIME-02; use RUNTIME-03 fixtures for integration. +- Write binding contracts before consumer implementation: descriptors, scoped agents/events, + media IDs/timestamps, task progress, targeted stimulation and retry/idempotency behavior. +- Implement Rust/TS codecs, schemas and a fake server; preserve the legacy v1 surface. +- Specify descriptor reconnect behavior, asset/index identity, bounded message sizes and audio gaps. +- **Done:** cross-language fixtures pass for unequal neuron counts and shared/private views; + duplicate attachment kinds no longer collide; ambiguous targets and incompatible schemas fail. + +### PRESENTATION-01 — Compose multi-agent stage and bridge + +- **Branch:** `feat/multi-agent-broadcast` +- **Depends on:** WIRE-01, RUNTIME-03. +- Replace stage store/scaler singletons with session/agent instances and one paint scheduler. +- Resolve geometry from hashed descriptors, preserve the Game Boy presentation, and add a + shared-match layout with explicit audio ownership. +- Route bridge commands/redemptions to persistent session/agent identities; test lost responses, + retries and restart without applying an interaction twice or to a different agent. +- **Done:** local synthetic match broadcast works; two-agent PNGs receive operator review; + browser/fixture/legibility checks pass; bridge remains template-only and quiet-mode capable. + +### DATA-03 — Run MaleCNS through the complete application + +- **Branch:** `feat/malecns-session` +- **Depends on:** DATA-02 and descriptor-aware assets from WIRE-01/PRESENTATION-01. +- Expose explicit profile selection and create a fresh MaleCNS state namespace. +- Verify task/controller bindings, stimulation capability and displayed anatomy identity. +- A narrow single-agent descriptor extension may ship earlier only with matching v1 contract + and consumer updates; do not publish MaleCNS spikes as implicit FAFB indices. +- **Done:** local one-hour soak and restore drill pass; paired learning-off/on observations + are recorded without claiming improved play; FAFB remains available unchanged. + +### EMULATOR-01 — Establish the alternative backend's capabilities + +- **Branch:** `spike/fighting-game-backend` +- **Depends on:** RUNTIME-01; can proceed alongside later runtime work. +- Choose the exact game/version and emulator; Melee/Dolphin is a candidate, not a commitment. +- Prove pause/step, simultaneous ports, analog input, frame/audio capture, state inspection, + save/restore, process lifecycle and achievable cadence with synthetic controller traces. +- Prefer private IPC if embedding would leak emulator internals into the session library. +- **Done:** capability report includes pinned backend/content identity and reproducible results. + If bounded stepping or coherent restore fails, stop and revise the backend/requirements + before writing neural game logic. No game content enters repository fixtures. + +### EMULATOR-02 — Build the two-fly fighting-game slice + +- **Branch:** `feat/two-fly-fighting-game` +- **Depends on:** EMULATOR-01, STATE-01, PRESENTATION-01. +- Implement fixed controller mapping, match/round interpretation, positive attributed rewards, + observation policy and episode recovery. Display selected actions and actual controls. +- Validate one fly, two flies, round transitions, backend failure and resume in that order. +- Run side swaps and repeated seeds; compare learning-off and simple control baselines before + interpreting win rates. Separate show settings from controlled evaluation settings. +- **Done:** repeated local matches sustain declared cadence; restoration/failure policies work; + scaffold, interventions and limits are documented; reviewed match presentation is legible. + +### PACKAGE-01 — Finalize reusable packages and release compositions + +- **Branch:** `refactor/reusable-package-layout` +- **Depends on:** a useful second backend plus PRESENTATION-01. +- Extract proven crate/package boundaries from the design's module table; preserve facades. +- Move the Rust workspace only in a mechanical follow-up if it makes library consumption clearer. +- Update CI/build/vendor/golden paths, dataset/view asset packaging and compatibility preflight. +- Add minimal external-style Rust/TS consumers and a synthetic example composition. +- Reconcile current docs, stale template explanations and licensing/asset attribution. +- **Done:** both legacy and new compositions package successfully; incompatible state is + rejected before release selection; libraries run without importing broadcast services. + +## 3. Dependency map and first execution batch + +```text +FOUNDATION-01 → FOUNDATION-02 ┬→ DATA-01 → DATA-02 ───────────────→ DATA-03 + └→ RUNTIME-01 → RUNTIME-02 → RUNTIME-03 → STATE-01 + │ └→ WIRE-01 ───────┐ + └→ EMULATOR-01 PRESENTATION-01 + │ + STATE-01 + EMULATOR-01 + PRESENTATION-01 → EMULATOR-02 + second backend + presentation → PACKAGE-01 +``` + +The item dependency lists are authoritative; the diagram is a reading aid. + +When implementation begins, take **FOUNDATION-01 only** as the first build task. Then review +FOUNDATION-02's contract before assigning DATA-01 and RUNTIME-01 to independent worktrees. +Contract/schema authorship is serialized to avoid conflicting definitions. Deployment host +work remains serialized under the repository's claim protocol. + +## 4. Definition of done for every implementation branch + +- Scope and intentional behavior changes are stated; compatibility impact is explicit. +- Meaningful boundary tests cover the changed behavior; the TS oracle is not adjusted to + accommodate Rust output. Existing committed real-data goldens stay mandatory. +- `npm test`, `npm run typecheck`, `cargo test --workspace` (Rust workspace) and + `infra/tests/lint.sh` pass before merge. Visual changes also pass applicable Playwright + checks and PNG review. Optional full-MaleCNS/ROM runs record skips honestly. +- Performance-sensitive changes report representative activity, agent count, thread budget, + memory and tail latency. New experiments state what is modeled versus handwritten. +- Review the complete diff, merge with `--no-ff` when authorized, and update this queue with + commit, evidence and unresolved follow-ups. Rollback includes compatible state, not just code. + +## 5. Decisions needed before the relevant work starts + +| Decision | Deadline | Default recommendation | +| --- | --- | --- | +| MaleCNS inventory/filter policy | DATA-01 completion | Official versioned source; retain both threshold variants until measured | +| MaleCNS sensory/readout profile | DATA-02 | Audited L1 mapping with existing numerical model first | +| Exact fighting game and backend | EMULATOR-01 | Evaluate one concrete title/backend rather than supporting a console family at once | +| Number of flies and target resource budget | RUNTIME-03 performance gate | Two first; characterize four before promising it | +| Learning retention and sugar in matches | EMULATOR-02 task contract | Retention explicit; stimulation disabled in controlled comparisons | +| Package publication versus monorepo reuse | PACKAGE-01 | Monorepo libraries/examples first; public package publishing later | + +There is no need to resolve these now to plan another feature. This backlog is ready for +resumption at FOUNDATION-01. diff --git a/docs/design/malecns-modular-sessions.md b/docs/design/malecns-modular-sessions.md new file mode 100644 index 0000000..dcd82ac --- /dev/null +++ b/docs/design/malecns-modular-sessions.md @@ -0,0 +1,892 @@ +# MaleCNS and reusable streamed simulation sessions + +Status: **proposal, not an implemented contract**. Written 2026-09-18 against `f7bc13a` +on `main`. This document covers two related projects: adding MaleCNS v1.0 as another +connectome, and extracting reusable modules for other emulators, embodied environments, +and multiple flies. No dataset, neural semantics, deployed configuration, or wire contract +is changed by this document. + +The binding [feed](../feed-protocol.md) and [control](../control-api.md) contracts take +precedence. The TypeScript brain remains the oracle. Existing default versions +`lif-1ms-f64-v2` and `fly-kc-mbon-rstdp-v2` remain pinned. + +Reading map: sections 2–3 contain the code audit and MaleCNS analysis; sections 4–6 +define the proposed module/session boundaries; sections 7–8 give the extraction order, +implementation workstreams and acceptance gates; sections 9–10 record open questions and +sources. + +Execution queue: [implementation backlog](malecns-modular-implementation.md), with branch-sized +deliverables, dependencies and completion criteria. Start at FOUNDATION-01 when work resumes. + +Concrete follow-up: [Melee emulator and multi-fly framework audit](melee-framework-audit.md), +including source-checked Dolphin/libmelee integration options and full-stack performance gates. + +## 1. Recommendation + +1. **Add MaleCNS as a dataset/profile combination, not a replacement neural model.** + First run it through the existing LIF semantics with explicit, independently versioned + sensory, population, and readout mappings. Study different neuron dynamics separately. +2. **Make a session the unit of simulation ownership.** A session has one environment and + one or more independently stateful agents bound to its control ports. A shared match + advances once after all players have chosen actions from the same observation boundary. +3. **Extract along ownership and timing boundaries.** Separate anatomy, neural dynamics, + sensor encoding, action decoding, environment execution, task semantics, persistence, + observation transport, presentation, and audience interaction. Preserve current behavior + through a legacy composition while extracting these modules. +4. **Prove the design on a ROM-free two-player arena before a larger emulator.** Then build + a frame-stepped emulator integration. For “flies play Smash,” the first candidate should + be a specifically chosen title/backend, such as Melee with a pinned Dolphin integration; + “Smash” alone is not an emulator requirement. +5. **Retain a monorepo and one Rust workspace initially.** Reusable libraries do not require + a network of microservices, dynamic native plugins, or publishing unstable packages. + +The two tracks can progress independently after the identity/profile boundary is established. +MaleCNS does not require multiplayer; multiplayer does not require MaleCNS. The first useful +deliverables are a reproducible MaleCNS characterization run and a behavior-preserving +single-agent session API, not a wholesale rewrite. + +## 2. What the code actually does today + +Paths below are relative to the repository root. Rust paths beginning `core/`, `gb/`, or +`sim/` in this document abbreviate `services/flysim/crates/flybrain-core/`, +`services/flysim/crates/flybrain-gb/`, and `services/flysim/crates/flysim/` respectively. +These aliases refer to **current** paths, not proposed directories. + +| Boundary | Evidence inspected | Consequence | +| --- | --- | --- | +| Reference brain | `packages/brain/src/agent/agent.ts`; `core/src/agent.rs` | `NeuralAgent` composes network and decoder; image size and milliseconds/frame are configurable, but defaults are Game Boy-specific. It is already more reusable than the service. | +| Anatomy | `packages/brain/src/dataset/format.ts`; `core/src/dataset.rs` | CSR graph dimensions are dynamic. Weights are signed `i16`; roles and a two-dimensional visual-column table are part of the model input. Validation currently checks array lengths, not every graph invariant. | +| Dataset construction | `tools/build_flywire.py` | Five checksum-pinned Codex exports; stable indices from sorted root IDs; directed-pair aggregation; transmitter signs; clipping to ±32767; FAFB-specific role and L1-column extraction. Game macro populations are also generated here. | +| Neural dynamics | `core/src/lif.rs`; `packages/brain/src/model/lif.ts` | One-ms ticks, configurable gain/noise, role stimulation, image drive, and up to 64 tracked rate roles. It is not a general multimodal sensory API. | +| Learning | `core/src/plasticity.rs` | Strongest positive pre-role→post-role edges, default KC→MBON budget 16,384; gains/eligibility are per-agent. A caller supplies the reward scalar; PAM spikes do not generate it. | +| Parallelism | `core/src/lif.rs` (`SweepPlan`); `core/src/pool.rs` | Persistent deterministic within-brain pool. `broadcast` has one shared job slot and assumes one dispatcher at a time; cloning a plan is not permission to dispatch it concurrently. | +| Game interface | `gb/src/adapter.rs` | `GameAdapter` mixes reward detection, progress, decoder preset, recovery, ROM checks, and Pokémon-like tile/exit/objective queries. `MemoryReader::read8(u16)` is specifically a Game Boy-shaped interface. | +| Emulator | `gb/src/emulator.rs` | Concrete binjgb wrapper, 160×144 RGBA, eight-bit pad, one frame step, audio conversion, native save-state format. Explicit handles and `Send` allow multiple instances; `Sync` is deliberately absent. | +| Session ownership | `sim/src/simloop.rs` (`Sim`, `step_frame`) | One agent, emulator, adapter, ratchet, button mask, framebuffer, audio queue, sugar state, chat ring, and set of clocks. Application orchestration and game behavior share a large struct. | +| Actions | `sim/src/macros.rs`; `gb/src/macros.rs`; `gb/src/pokemon_red/macros/` | Neural channels select available macros, but the game-specific executor owns button sequences. The title screen uses raw input; scene handling, routing and targets are engineered behavior. | +| Recovery | `gb/src/recovery.rs`; `gb/src/ratchet.rs` | Game-only rewind retains brain clock, membrane, RNG, and gains; clears holds/eligibility and refreshes vision. A scalar progress ladder chooses a best save. This is not a general multiplayer reset policy. | +| Persistence | `sim/src/store.rs`; `gb/src/compatibility.rs` | Durable atomic envelope and manifest commit are reusable. Payload is one agent plus one emulator and ratchet. Compatibility names binjgb and `pokered`; native state identity includes size and target. | +| Public observation | `sim/src/snapshot.rs`; `packages/feed/src/{types,codec}.ts` | One flat brain/game snapshot; one attachment per kind; Game Boy buttons, fixed frame dimensions, Pokémon-shaped reward counters and a closed game-mode set. | +| Stage | `apps/stage/src/{App.tsx,feed/store.ts,feed/decode.ts}` | Good hot/paint/cold clock split, but mutable stores/scalers are singletons and the decoder checks 160×144 frames. Dataset URL is fixed to FAFB. | +| Game presentation | `apps/stage/src/games/` | A useful registry already exists, but config relabels v1 counters rather than declaring independent task schemas. | +| Twitch bridge | `services/bridge/src/{index,sim,commands,redemptions,templates}.ts` | Transport/client abstraction, test fakes, templates, rate limits, and redemption persistence are valuable. There is one sim URL and no agent target identity. | +| Packaging | `apps/stage/vite.config.ts`; `infra/build/package-release.sh`; `infra/05-deploy.sh` | Stage build copies FAFB artifacts; packaging defaults to FAFB; deploy preflights one compatibility string. Runtime/data/frontend are still one release composition. | + +### 2.1 Behaviors to preserve before extracting + +`Sim::step_frame` advances the brain from the previous visual input, decodes, applies +buttons, steps the emulator, sets the new visual input, samples rewards, stimulates per +reward event, reinforces their sum, updates macro availability, and observes the ratchet. +Control commands are drained before the frame step. This ordering is part of behavior. + +Do not replace this with `NeuralAgent::tick` merely because it looks like a convenient +wrapper: the service currently orchestrates substeps to sample reward from the frame just +produced. Moving reward or visual drive across that boundary changes trajectories. + +Other invariants: + +- Deterministic arithmetic and per-target propagation order, including across thread counts. +- Browser/network clients never stall the sim; snapshots are latest-value/drop-oldest. +- Checkpoint capture is coherent; encoding and storage happen off the sim thread. +- Failed restore does not silently reset a run. Existing legacy restore policies remain exact. +- No public control endpoint for button presses or game-memory writes. +- Sugar and learning reward are distinct mechanisms. Chat text never becomes neural input. +- The current positive-only reward doctrine remains the default for all shipped tasks. + +### 2.2 Existing compatibility gaps to handle deliberately + +The dataset fingerprint hashes metadata (including anatomical circuit roles) and six arrays. +`macro_*` roles are merged **after** hashing to preserve old checkpoints. Both language +implementations document that changing those populations could restore rates onto different +neurons without invalidating the checkpoint. The kernel parameter version also omits role +names; the service compatibility string does not fully identify the decoder/action mapping. + +Keep these historical behaviors in the legacy reader. For new profiles, add an explicit +behavior identity covering population bindings, sensor encoding, decoder configuration, +action executor, and reward catalog. Do not repair the old hash by changing it in place. + +Open macro/shop and recovery branches existed when this proposal was written. Before +implementation, rebase the inventory against their merged state and rerun characterization; +this proposal neither incorporates nor supersedes their unmerged changes. + +## 3. MaleCNS: anatomy, connectivity, and model are different things + +### 3.1 Dataset comparison and evidence limits + +| Property | FAFB v783 used here | MaleCNS v1.0 | +| --- | --- | --- | +| Specimen | Adult female | Adult male, independently imaged/reconstructed | +| Territory | Brain including optic lobes | Central brain, optic lobes, ventral nerve cord (VNC), intact neck connective | +| Local artifact | 139,255 neurons; 2,700,513 directed-pair edges | None imported in this repository | +| Available inventory counts | Codex lists 139,255 neurons | Codex lists 166,700; the Minecraft project's neuPrint `:Neuron` export reports 176,422. Selection rules must be reconciled before fixing a local count. | +| Input labels | Codex `root_id`, classification, consolidated types, column assignment | neuPrint/flat export `bodyId`/body IDs, class hierarchy, transmitter properties, sides, neuropils, cross-dataset type annotations | +| Added anatomical opportunity | Brain sensory→descending circuits | Brain↔VNC circuits, local motor circuitry, ascending feedback, additional sensory and motor populations | +| License evidence | Repository attribution: CC BY-NC 4.0 | Official MaleCNS download site: CC-BY; verify and retain the exact release license text when importing | + +MaleCNS is not FlyWire with extra neurons appended. IDs and dense array indices do not +correspond. Homologous cell types and registered anatomical spaces enable comparisons, not +automatic one-to-one neuron matching, state transfer, or graph concatenation. Male-specific +and sexually dimorphic circuits make a universal matching assumption especially misleading. + +The official MaleCNS site describes a finished, proofread and annotated CNS reconstruction. +This does not mean every synapse, cell type, or sensory column is equally certain. The +Minecraft derivative reports asymmetric visual-column coverage and missing soma positions; +our import must quantify coverage from its own pinned source. A connectome also does not +supply all synaptic physiology, electrical coupling, neuromodulation, body dynamics, or +behavioral competence. + +### 3.2 What “connections” means + +Separate these quantities in metadata, reports, and on-screen claims: + +1. Source neuron/segment inventory and the chosen included-neuron inventory. +2. Individual synaptic contacts/partner pairs (and separately pre-sites and post-sites). +3. Directed neuron-pair edges after aggregation. +4. Retained edges and retained synaptic weight after confidence/weight filtering. +5. Effective signed weights after the simulation's transmitter and clipping policy. + +The FAFB builder aggregates export rows by `(pre, post)` across rows, drops endpoints not +in its classification inventory, assigns a sign, and clips the summed magnitude. It adds +no explicit five-synapse threshold of its own. The source exports may already be filtered; +the builder cannot recover contacts absent upstream. Codex headline connection counts are +not necessarily the local aggregated graph's edge count. + +The Minecraft project's provenance reports ~25.9 million MaleCNS neuron-pair edges at +weight ≥1 and a bundled derivative of 6,287,749 edges at weight ≥5, representing +90,296,905 of 125,024,863 neuron-to-neuron synapses. It removes 40 autapses. Those are +**that project's reported query/build results**, not counts independently reproduced here. +Its five-contact cutoff retains roughly 24% of edges but 72% of synaptic weight. It is a +performance/modeling choice, not the definition of a complete CNS. + +The official bulk weight table includes **segments**, not only curated neurons. Loading +all rows as if they were all validated neurons would be a different experiment. Likewise, +neuPrint ROI-level adjacency rows must not be summed together with their already-aggregated +totals. Specify one authoritative edge representation and count every exclusion. + +### 3.3 Acquisition and reproducible construction + +Prefer official versioned bulk tables for the repeatable build, with a neuPrint query tool +for inspection and cross-checking. The official download page lists: + +- `body-annotations-male-cns-v1.0-minconf-0.5.feather` — curated annotations. +- `body-neurotransmitters-male-cns-v1.0.feather` — neuron-level transmitter information. +- `body-stats-male-cns-v1.0-minconf-0.5.feather` — segment statistics; large and broader than + the curated neuron list. +- `connectome-weights-male-cns-v1.0-minconf-0.5.feather` — full segment connection graph, + approximately 1.1 GB as listed by upstream. + +The much larger synaptic-point and partner tables are unnecessary for an initial point-neuron +simulation. Download them only for a question requiring synapse-level geometry. Neither +research downloads nor anatomy conversion belong in service startup or ordinary unit tests. + +Proposed build stages: + +```text +release manifest + checksummed source cache + → source-specific parser + → normalized neuron/edge tables + exclusion report + → selected anatomical graph + → model-specific signed-weight transform + → runtime CSR bundle + profile bindings + separate viewer bundle +``` + +Each stage records source release, database revision if queried, query/filter definitions, +source byte hashes, tool revision, and output hashes. Fetch time is provenance, not a random +input to the semantic graph hash. Sort body IDs numerically; encode original IDs as strings +in JSON so this common format also preserves FAFB IDs beyond JavaScript's safe integer range. +Preserve original annotations and cross-dataset type aliases separately from normalized roles. + +The first import report must reconcile the Codex/neuPrint count difference, or explicitly +choose and document one inventory without claiming equivalence. It must also report missing +IDs, unannotated neurons, empty required populations, missing geometry, unknown sides, +transmitter confidence/fallback counts, duplicate edges, autapses, clipped weights, and +retained contacts by region and threshold. + +Use deterministic serialization and gzip headers, following our existing reproducible +builder rather than copying the Minecraft artifact's timestamp-dependent container format. +Keep the source cache outside tracked artifacts; pin published runtime bundles by digest. + +When adding actual data, update `NOTICE`, `LICENSES.md` and bundle-local attribution/license +files in the same change. Keep FAFB-derived assets under their existing terms; a separately +licensed MaleCNS bundle does not relicense mixed fixtures, old goldens or viewer assets. + +### 3.4 Graph and sign policy + +Keep raw positive contact counts and transmitter evidence in the normalized data. Sign is a +model transform, not a measured property that should overwrite the source evidence. + +The legacy policy assigns GABA/GLUT negative and other/unknown transmitters positive; +conflicting per-edge transmitter rows become `MIXED`, then positive. Do not silently extend +that policy to histaminergic photoreceptor input. A MaleCNS policy must explicitly define +histamine, monoamines, mixed/unknown labels, confidence fallbacks, and whether transmitter +is chosen per neuron or per connection. None implies receptor-specific physiology. + +Recommended first profiles: + +- **`malecns-v1-lif-baseline`**: existing kernel, explicitly versioned sign policy and L1 + input mapping, fixed readout, plasticity initially disabled for characterization. +- **`malecns-v1-lif-learning`**: same anatomy/input/readout plus audited KC→MBON selection + and the existing reward rule. Learning is an experimental condition, not an assumed gain. +- Later **sensorimotor research profiles**: photoreceptors, mechanosensation, VNC outputs, + and possibly another neural model, each with its own identity and validation. + +Build weight≥1 and weight≥5 variants as **different graph identities** for the benchmark. +Do not select a production cutoff until activity, retained connectivity, memory and speed +have been measured. Preserve autapses by default in the new canonical graph; if an experiment +removes them, record the policy. The claim that point-neuron models have no use for autapses +is not a reason to discard observed connectivity silently. + +Schema-1 `i16` weights may be sufficient, but measure overflow rather than assume it. For +the first existing-kernel comparison, emit a schema-1-compatible runtime view only if its +quantization/clipping is explicitly reported. If wider weights are needed, implement an +additive format/loader and matching oracle path; do not reinterpret old `weights.binz`. + +Strengthen validation before constructing a network: CSR starts at zero, is monotone, ends +at edge count; targets/roles/visual indices are in range; required populations are present; +geometry is finite where marked valid; array lengths and index widths are representable; +declared hashes and artifact sizes match. Test the same invalid fixtures in both languages. + +### 3.5 Population and sensory mapping + +Existing role predicates cannot be copied unchanged. The current builder recognizes Codex +names such as `Kenyon_Cell`, `brain_motor_neuron`, `DAN` and `PAM*`; MaleCNS uses another +annotation vocabulary. Maintain a reviewed mapping table with source predicates, resulting +counts, hemisphere policy, aliases, and citations. Resolve each profile's required roles +at startup; an absent role must be an unsupported capability, not a silent empty population. + +In particular: + +- **Brain motor ≠ all motor.** Current macro pools combine 96 MBONs and 110 brain motor + neurons. Adding hundreds of VNC motor neurons under `motor` would silently change that + behavior. Use qualified roles such as `brain.motor`, `vnc.motor`, `brain.descending`, + `mb.kenyon`, `mb.output`, and `mb.pam` in new profiles, with legacy aliases only as needed. +- **Action groups are not anatomical facts.** `command_*` buckets use index modulo eight; + `macro_*` groups use a round-robin MBON/brain-motor pool. Move their construction into a + versioned task readout profile, outside the anatomy builder. Do not describe those groups + as natural “attack,” “jump,” or game-objective circuits. +- **Rate budgets are finite.** Both kernels currently cap tracked populations at 64. A + full CNS has many more interesting populations. Select a bounded control/telemetry set + initially; arbitrary bulk population analysis belongs in offline tooling. A larger mask + is a measured, oracle-tested change, not an unbounded string map in the tick loop. +- **L1 first, photoreceptors later.** Reuse the existing luminance projection only after + auditing MaleCNS L1 hex coordinates, both sides, and a declared hex→2D transform. Soma + coordinates are not visual-field coordinates. Missing columns remain explicitly missing; + do not synthesize them from array order or infer them from another specimen's neuron IDs. +- **Input profile and display view differ.** A game image may be resized/cropped for the + agent while the full frame is shown to viewers. Record crop, orientation, color transform, + and sampling geometry. Neither player may accidentally receive another player's private + view or adapter-only task observations. + +The existing `set_visual_frame` and `stimulate` API cannot represent a general collection of +odor/touch/proprioceptive inputs. A later sensory-drive interface needs explicit units, +target populations, additive/overriding rules, tick ordering, and deterministic noise streams. +It must be specified in TypeScript before a matching Rust implementation. Keep the old +image/stimulation path available byte-for-byte through its compatibility facade. + +### 3.6 What MaleCNS lets us investigate + +| Experiment | New capability | What still needs engineering/measurement | +| --- | --- | --- | +| Same game, another connectome | Compare datasets under a matched task interface | Cell-type mapping, gain/activity calibration, readout comparability, multiple seeds | +| Brain↔VNC control | Read actual descending, ascending and motor populations | Body/control mapping; gamepad commands are not muscles | +| Embodied fly arena | World smell/taste/touch/vision mapped into annotated sensory populations | Sensor transduction, proprioception and body dynamics; identify every reflex shortcut | +| Mixed-dataset two-fly match | FAFB and MaleCNS agents share one environment | Balanced observations, controller mapping, compute budgets, intervention rules | +| Circuit perturbation | Compare full graph with VNC feedback or defined pathways ablated | Separate graph identities; activity and behavioral controls; no biological claims from gameplay alone | + +The Minecraft project is a useful engineering comparison, not our validation oracle. It uses +a Shiu-style current-based LIF model with synaptic dynamics/delay, reports gain calibration, +and explicitly supplies odor-approach reflexes and higher-level looming drive where its +simulated pathways do not work. Its source graph, numerical model and embodiment differ +from ours simultaneously. Do not attribute its behavior solely to MaleCNS. + +### 3.7 Identity, restore, and performance + +Name the components independently: + +```text +anatomyId = source release + included inventory + graph/filter digest +modelId = numerical semantics + effective numeric configuration +sensorId = input encoding + anatomical binding digest +readoutId = population partition + decoder + action mapping digest +learningId = rule + selected-edge topology + reward-catalog identity +viewerId = positions/geometry + index mapping digest (presentation only) +``` + +The composite behavioral identity covers all behavior-affecting components; original source +IDs and display geometry cannot replace it. Same neuron count does not establish compatibility. +No FAFB neural checkpoint is restored into MaleCNS. A deliberately fresh MaleCNS brain may +start from a compatible game-only save, with a new run identity, clean calibration and reward +baselining; that is a new experiment, not continuation of the old fly. Gains are not mapped +between specimens by cell-type name. + +For rough capacity planning, the current CSR is `4(N+1) + 6E` bytes, excluding roles, +geometry, derived propagation structures and mutable state. At 176,422 neurons that is +about 38.4 MB for 6.29 M edges, or 156.1 MB for 25.9 M edges (decimal MB). This is roughly +2.3× or 9.6× our edge count, not a prediction of the same slowdown. Spike activity, fan-out, +plasticity selection, memory bandwidth and sharding determine runtime cost. Checkpoint +copies and renderer assets also need separate memory budgets. + +`LifNetwork` accepts shared `Arc` already. Reuse immutable anatomy across +same-profile agents; keep membrane, refractory state, RNG, rates, decoder holds, stimulation, +eligibility and learned gains private. Audit constructor-derived caches before moving them +into a shared topology object. Never share mutable gains merely because graphs match. + +Measure headless one-, two-, and four-agent runs with plasticity on/off, fixed inputs and +representative activity. Record resident/peak memory, initialization, state capture cost, +per-phase p50/p95/p99, spike distribution, and real-time factor. Existing CUDA code is an +optional backend requiring its own new-dataset equivalence/capacity gate, not assumed capacity. + +## 4. Reusable architecture + +### 4.1 Define the nouns first + +- **Dataset bundle:** immutable anatomical graph and source annotations. +- **Brain profile:** dataset plus numerical model, sensory/readout bindings and learning rule. +- **Agent:** one independently stateful brain, encoder, decoder and action executor. +- **Environment:** the world being advanced: one emulator instance, a linked-emulator group, + or an embodied simulator. Owns controller ports, world state, media, and native clock. +- **Task:** interpretation of environment state: rewards, progress, episode endings, allowed + macro actions and recovery policy. Pokémon is a task, not an environment API. +- **Session:** one environment plus agents, port assignments, scheduler, task state and clocks. +- **Broadcast:** presentation of one or several sessions, plus chat and audience interactions. + +An agent ID is not a Twitch username, controller port, array position or dataset ID. Session, +episode, agent, port, view, and event identities must be explicit and stable across restore. + +### 4.2 Dependency direction + +```text +source importers → dataset bundles + ↓ + neural core (TS oracle / Rust runtime) + ↓ + agent composition: sensors + readout + executor + ↓ +environment backend + task plugin → session runtime → observations/checkpoints + ↑ ↓ + command admission protocol adapters + ↑ ↓ + Twitch bridge stage / recorder +``` + +The neural core knows no emulator, task, network socket, chat, or UI. The environment knows +no neural populations or Twitch. The task may inspect backend-specific state through a +typed inspector, but neither inspection nor public presentation gives clients a memory-write +or controller-write API. Only the session commits agent-produced controls. + +### 4.3 Proposed modules and staged layout + +These are target responsibilities, not instructions to create every package immediately. +Start as modules; extract crates/packages once a second consumer demonstrates the boundary. +Keep the Rust workspace under `services/flysim` during semantic extraction so paths and +behavior do not change together. A later mechanical move can place reusable crates at the +root, updating CI/build/golden paths in one dedicated change. + +| Module / eventual location | Owns | Extraction source | +| --- | --- | --- | +| `packages/brain` | Reference numerical behavior and legacy public facade | Existing package; keep imports compatible | +| `crates/flybrain-core` | Rust numerical kernel, plasticity, generic population decoder | Existing `core/`; leave compatibility re-exports for presets | +| `crates/fly-dataset` | Manifest validation, artifact loading, source-ID/index mapping | `core/src/dataset.rs`; retain legacy fingerprint implementation | +| `tools/datasets/{fafb,malecns}` | Source-specific conversion to common bundles | Existing Python builder plus new importer; existing CLI wrapper remains | +| `crates/fly-session` | Agent ownership, clock coordination, action commit, event/reward routing | Orchestration extracted from `sim/src/simloop.rs` | +| `crates/fly-environment` | Backend capabilities, ports, observations, media, save-state interfaces | New small contract proven with binjgb and synthetic arena | +| `crates/fly-env-gb` | binjgb FFI, memory inspector and native save-state identity | `gb/src/{emulator,ffi}.rs`, build glue and vendor boundary | +| `crates/fly-task-pokemon`, `fly-task-platformer` | Audited reward rules, semantic state, macros, progress/recovery | `gb/src/{pokemon_red,platformer}/`; do not generalize tile routing into the core | +| `crates/fly-checkpoint` | Atomic storage and session envelope; legacy payload adapter | `sim/src/store.rs` plus core envelope helpers | +| `crates/fly-protocol` / `packages/feed` | Versioned wire schemas/codecs, legacy adapters, synthetic fixtures | `sim/src/snapshot.rs`, existing feed package; canonical schema with cross-language tests | +| `services/flysim` | Composition/config, HTTP/WS, process lifecycle, metrics | Thin host over reusable session library | +| `packages/stage-runtime` | Feed ingestion, per-session stores, paint loop, audio, fixture clock | Extract from `apps/stage/src/{feed,paint,audio,motion}` after multi-view prototype | +| `apps/stage` + presentation plugins | Layout, branding, task panels, audience-facing explanations | Existing page with legacy layout preserved | +| `services/bridge` + audience client module | Twitch transport/auth/redemptions; session-targeted interaction client | Existing bridge; extract provider-independent logic only when reused | +| `infra/` | Release composition, process supervision, capture, recordings | Existing tooling parameterized by session/broadcast manifest | + +Use static Rust composition or a small closed registry initially, with trait boundaries at +backend/task seams. Do not require stable native dynamic-plugin ABI. An out-of-process +emulator can implement the backend through private IPC; it is not a new public action API. +Keep backend-specific memory access private to its task implementation instead of widening +`read8(u16)` into a supposedly universal game-state abstraction. + +### 4.4 Environment and agent contracts + +Illustrative interfaces; concrete types must be written with tests during contract work: + +```rust +trait Environment { + fn descriptor(&self) -> &EnvironmentDescriptor; + fn observe(&mut self) -> Result; + fn advance(&mut self, actions: &ActionBatch) -> Result; + fn capture(&mut self) -> Result; + fn restore(&mut self, state: &EnvironmentCheckpoint) -> Result<()>; +} + +// One decision boundary, one action per configured controller port. +struct ActionBatch { + session_tick: u64, + ports: Vec, +} +``` + +`descriptor` declares rational step duration, controller schemas, views, audio streams, +save/restore availability, task inspection capabilities and determinism level. Capture and +restore return an explicit unsupported error when unavailable; configuration validates that +the chosen recovery policy can work. `WorldObservation` is a frame-boundary snapshot or +immutable handle, not an object allowing agents to advance the backend. + +Control schemas support digital buttons and bounded analog axes/triggers, with neutral +values, axis ranges, dead zones, and mutually exclusive directions where appropriate. +Preserve the Game Boy mask as one concrete codec. Analog controls need a fixed, versioned +decoder mapping; an 800-ms direction hold is not a sensible default for every fighting game. + +Separate three observation surfaces: + +1. **Agent sensory view:** pixels or declared synthetic senses the profile may consume. +2. **Task inspector:** audited state for reward/macro/episode logic. Access is part of the + disclosed scaffold, not implicitly available to the neural encoder. +3. **Broadcast view:** media and summaries for viewers, potentially richer than either player's + allowed sensory input. + +Readout produces semantic channel activations or continuous signals. A task-local action +executor translates these into port actions, optionally running a selected macro. It is +explicitly resettable/checkpointable and reports selected action versus actual controller +output. Pokémon pathfinding, dialog logic, and objective catalogs stay in its task plugin. + +Task outputs become scoped `RewardEvent`, `Progress`, `EpisodeEvent` and `ActionAvailability`. +Progress is a tagged value (`ladder`, `score`, `match`, `exploration`, or task extension), not +always a scalar rank. Rewards carry recipient agent/team, rule ID, event ID and observation +tick. A task cannot mutate neural state directly; the session routes accepted rewards once. + +An agent-facing API similarly separates `advance_brain(interval)`, `decide(observation, +availability)`, `encode_next(sensory_view)`, and `apply_outcome(rewards, stimulation)`. +The session owns their order. Agents cannot call `Environment::advance`, select another +port, or inspect another agent's mutable state. Start with a concrete LIF agent composition; +introduce a controller trait when synthetic controllers or a second neural model require it. +Test controllers implement the same decision surface but are identified as non-neural agents +in descriptors and experiment records. + +### 4.5 Example composition + +Illustrative configuration, not syntax supported by today's `flysim.toml`: + +```toml +[session] +id = "arena-demo" +environment = "synthetic-arena-v1" +task = "two-player-rounds-v1" +scheduler = "lockstep-v1" +master_seed = 1234 +recovery = "round-reset-keep-gains-v1" + +[[agents]] +id = "fly-a" +port = "player-1" +profile = "fafb-arena-baseline-v1" +sensory_view = "shared-camera" + +[[agents]] +id = "fly-b" +port = "player-2" +profile = "malecns-arena-baseline-v1" +sensory_view = "shared-camera" + +[broadcast] +layout = "shared-match-two-agents" +audio = "world" +audience_stimulation = false +``` + +Resolve profile IDs through a local, digest-pinned registry. Both agents may instead select +the same profile and share immutable topology while retaining independent state. The host +validates unique agent IDs and exclusive port ownership, view accessibility, profile/controller +compatibility, recovery capability and resource budget before starting. An independent-games +broadcast composes two such sessions; it does not misrepresent them as ports in one world. + +## 5. Multiple flies: concurrency is not multiplayer + +### 5.1 Three supported arrangements + +| Arrangement | Ownership / synchronization | First use | +| --- | --- | --- | +| Independent flies in independent games | One session/process each; optional broadcast composition | Parallel streams and experiments; existing deployment pattern generalizes easily | +| Several flies in one game | One environment, several ports, one session barrier | Local fighting/multiplayer games | +| Linked emulator instances | One composite environment owns all instances and link state | A later link-cable experiment; requires cycle-accurate link support, not two independent frame loops | + +For a shared arena there must not be one `Sim` loop per player, each calling `run_frame`. +That advances the world multiple times and gives an ordering advantage to one agent. + +### 5.2 Shared-world step semantics + +For decision boundary `t`, freeze observation `O[t]`, then: + +1. Admit queued audience/operator commands against stable session/agent identities; log their + effective tick. Chat remains presentation state. +2. Each agent advances its brain for the same environment interval, using its previously + encoded sensory input. Its clock remainder and RNG are private. +3. Each agent decodes and advances its action executor using the same `O[t]` task boundary. +4. Barrier: collect all port actions, validate ownership/ranges, and commit one complete batch. +5. Advance the environment **once** to obtain `O[t+1]` and timestamped media. +6. Encode the next sensory inputs, evaluate task events from the completed transition, apply + explicitly routed stimulation and rewards, and compute next action availability. +7. Apply any whole-session episode/recovery transition; capture coherent state and publish. + +Preserve the detailed legacy ordering inside the legacy single-agent composition. New +profiles identify their scheduling semantics explicitly rather than silently adopting a +different reward phase. Use rational environment time and integer substep accumulation for +new sessions; keep the legacy floating remainder arithmetic for old trajectories. The neural +clock may lead environment time by warm-up; persist that offset instead of pretending all +clocks start at zero. Rendering, physics and decision cadence may differ, but the backend +must define their relationship. + +Start with sequential agent evaluation for reproducibility. Then compare parallel agent +evaluation against the same action trace. Cap total worker budget: `agents × brain_threads` +can otherwise oversubscribe the machine. Use private pools for concurrent agents or serialize +dispatch into a pool; the current `WorkerPool` must not be concurrently reused through a +cloned `SweepPlan`. Sharing immutable graph buffers is independent of scheduling workers. + +If one agent is late, the default is to slow the **whole session** and report lag. A crashed +agent pauses/fails the match rather than silently becoming a neutral or scripted opponent. +A realtime external world that cannot pause requires a separate declared deadline/hold-last +policy, dropped-action telemetry and a different determinism claim. Do not hide that policy +inside the environment adapter or use wall-clock completion order as an action tie-breaker. + +### 5.3 Match state and learning + +Every agent has its own seed, calibration, gain vector, eligibility, reward totals and +stimulation cooldown. Derive seeds deterministically from a stored master seed and stable +agent ID; do not use thread scheduling or the default identical seed for every fly. + +For an initial fighting-game task: + +- Both agents receive the same shared camera unless the game has genuine private views. +- Controller-port swaps and seed repeats are part of evaluation; wins alone are confounded + by character, spawn, arena, action interface and side advantage. +- Award positive, explicitly attributed events such as a scored hit/round win; define + damage/self-damage/team attribution and duplicate detection before turning learning on. + A loss need not produce a negative reward; changing reward doctrine is a separate decision. +- Episode transitions may retain learned gains while clearing transient traces, or create + fresh brains for controlled trials. Record which policy was selected. +- Do not select a “best checkpoint” separately for each player in a shared world. There is + one world state. Tournament scores and historical results should not rewind with a match. +- Disable sugar for balanced evaluation. If enabled for a show, target a named agent under + a documented rule and log the intervention; do not call that an uncontrolled fair benchmark. + +### 5.4 Checkpoint and recovery semantics + +Distinguish three operations: + +| Operation | Restored/reset state | +| --- | --- | +| Crash resume | Coherent environment, every agent, scheduler remainders, task ledgers, pending actions/commands and executor state at one boundary | +| Task recovery | Policy-defined world rewind/reset and per-agent transient clearing; continued gains/brain clocks only when declared, as in the legacy ratchet | +| New episode | Task initial world state, explicit retained/fresh agent policy, new episode identity; session event sequence remains monotonic | + +Use a new session envelope version with a manifest mapping stable agent IDs to chunks and +listing graph/model/profile/backend/content/task/state-format identities. The current +envelope restricts chunk names to letters; do not simply append `agent/1/membrane` to it. +Specify a new container format or an explicit manifest-to-valid-chunk-name indirection. + +Validate every participant into staged state before mutating any live participant. A failed +environment import must not leave half the brains restored. For an external emulator, +restore a replacement stopped process when transactional in-place validation is impossible. +Capture at the action barrier with no backend step in flight. Reuse atomic payload write, +manifest commit, hot/durable tiers and off-thread serialization; bound outstanding snapshot +jobs so repeated copies cannot exhaust memory under slow storage. + +Exact replay requires action-executor and admission state, not just the neural envelope. +Legacy macros intentionally discard transient execution on restart; preserve that behavior +for v1 and label it as legacy continuation semantics, not exact session replay. New sessions +persist all behavior-affecting state or explicitly restart an episode under a documented rule. + +Keep `FLYSIM01` readable through a legacy adapter. Never silently rewrite a checkpoint on +read. Conversion is an explicit offline operation writing a new directory and run record. +Environment identity includes title/content digest, backend build/configuration, relevant +platform/native state format, and patch/symbol provenance; state size alone is not sufficient. + +## 6. Feed, stage and audience modules + +### 6.1 Feed v2 is required + +V1 is not a generic multi-agent format. Its attachment map forbids duplicate kinds, so two +`spikes` arrays cannot coexist; its button mask and 160×144 image are Game Boy-specific. +Platformer counters are currently folded onto names such as `pokedex` and `wildwin`. +Extending those conventions to fighting games would preserve syntax while losing meaning. + +Keep v1 stable for the legacy session. Introduce a negotiated v2 or a separate `/v2/feed` +endpoint, with a descriptor delivered before dependent snapshots and available on reconnect. +The contract change must update Rust, TypeScript, schema, fake service, fixture player, +stage and bridge together. Proposed shape: + +```text +SessionDescriptor + sessionId, protocol, descriptorRevision, environment/task identities, clock definition + agents[{agentId, profileId, datasetId, neuronCount, roles, controlPort}] + ports[{portId, controllerSchema}] + views[{viewId, dimensions, pixelFormat, sensory/broadcast use}] + audioStreams[{streamId, sampleRate, channels}] + assets[{id, contentHash, datasetIndexHash, localUrl, license/credit}] + +SessionSnapshot + descriptorRevision, seq, sessionTick, episodeId, environmentTime, wallTime, status + agents[{agentId, brainTime, rates, learning, selectedAction, actualControls, stimulation}] + progress: tagged task payload; events: scoped and sequenced + attachments[{id, kind, ownerId, byteLength, format, mediaTimestamp}] +``` + +Use bounded, schema-validated tagged payloads/namespaced task extensions, not arbitrary +unlimited JSON or executable server-supplied UI. Dataset identity and neuron index mapping +must accompany spike geometry: matching bitset length alone cannot establish alignment. +Include descriptor revision in every snapshot and reject stale/mismatched buffers. Large +monotonic IDs/times use a specified safe-integer bound or decimal strings across languages. + +Publish one shared camera/audio stream once, not once per agent. Independent sessions may +have independent streams. Timestamp audio/video to the session clock; define discontinuities +on reset, reconnect and lag. Latest-value video/telemetry may drop, while audio needs a +bounded timestamped buffer and explicit gap handling. Durable event IDs permit recovering +missed events; a drop-oldest snapshot feed is not an exactly-once event log. + +Bandwidth becomes an architectural concern: 640×480 RGBA at 30 Hz is ~36.9 MB/s before +framing; 1920×1080 is ~248.8 MB/s. Do not copy full-resolution raw frames per fly across +several sockets. Start with one bounded broadcast-resolution local stream and separate +agent input resolution; choose a compressed media transport only after measuring latency, +CPU/GPU cost and synchronization. A telemetry protocol need not become a video codec. + +### 6.2 Stage composition + +Extract `createSessionStore()` rather than adding `agent2` fields to the singleton `hot`. +Own rate scalers, button afterglow, ticker state, fixture clock and audio queues per session/ +agent. Keep one page paint scheduler; register surfaces against explicit view/agent IDs. +Geometry is loaded by descriptor/hash through an asset manifest, replacing the hardcoded +FAFB route in both `App.tsx` and `vite.config.ts`. + +Retain the existing Game Boy layout as a presentation plugin. Add composition primitives for +a shared match view with two agent summaries, or independent session tiles with one focused +audio source. A generic fallback shows status, media, controls and task labels without +inventing Pokémon counters. Unknown optional extensions can be omitted; unknown required +capabilities or mismatched descriptors must be visible rather than silently showing Pokémon. + +Keep UI runtime dependencies out of the numerical package. The current `@flybrain/brain` +view exports and optional Three.js peer can remain compatibility re-exports when viewer +geometry/helpers move to a dedicated view module. Controller labels belong in controller +schemas, not a UI import of the neural package's Game Boy preset. + +Review actual layout proposals as PNGs under `apps/stage/mockups/`, with existing legibility, +phone-scale and browser gates. This document proposes data/layout boundaries, not screen +copy or a replacement for visual sign-off. + +### 6.3 Audience interaction + +Keep Twitch authentication/EventSub and template-only replies in the bridge. Extract a +session-targeted interaction client with explicit `sessionId`, optional `agentId`, interaction +kind and idempotency key. A multi-agent request with no target is rejected unless a fixed, +declared target policy exists; presentation focus must never choose the recipient. + +Service admission owns per-session, per-agent and global limits. A profile advertises its +supported stimulation capability; an agent without PAM support returns unsupported rather +than pretending to accept “sugar.” `!stuck` becomes task-aware (ladder time versus round +time), while chat remains broadcast-scoped and independent of neural state. + +For redemptions, persist the resolved target and request identity before retrying. Distinguish +an HTTP timeout from a definite refusal: a timeout may occur after the sim accepted the +effect. V2 needs a durable or explicitly recoverable deduplication/status contract so bridge +restart cannot apply the same pulse twice or retarget a redemption to a new match. Do not +promise exactly-once behavior from the bridge intent log alone. Existing v1 remains as-is. + +### 6.4 Deployment and observability + +A deployable composition selects runtime binary, backend/task, agent profiles, dataset +bundles, view assets, session state namespace and broadcast layout. Generate service config +from that manifest plus the operator's external environment. Keep tokens and network-specific +values outside this repository. Package viewer artifacts separately from the full simulation +graph so the browser does not need every edge. + +Preserve existing process isolation: sim, bridge, browser, capture, local relay and Twitch +push can restart independently. One process/session remains the default for fault isolation; +a multi-agent match stays one coordinated session, not one independently supervised service +per port. Multi-session resource placement and broadcast composition are deployment concerns. + +Metrics distinguish session lag, environment step cost, per-agent step cost, barrier wait, +snapshot drops, audio discontinuities, checkpoint queue age and resource budget. Bound agent +labels to configured IDs; never label metrics by viewer name or arbitrary event text. +Health must distinguish paused, slow, disconnected backend and dead agent. A generic +watchdog cannot treat “no new Pokémon tiles” as a stall detector for all tasks. + +Deployment preflight checks every referenced profile/bundle and complete checkpoint identity +before selecting a release. Switching the binary back does not convert newer state: retain +the previous release's state namespace for rollback. Twitch publishing still requires the +operator's explicit approval for that run; implementation benchmarks use local sinks. + +## 7. Cleanup strategy: extract, then reorganize + +Prioritize coupling that prevents a second application, rather than renaming everything. + +1. **Characterize behavior and identify state owners.** Capture deterministic synthetic + observation→action→reward traces and legacy restore outcomes before moving code. +2. **Remove task construction from anatomy.** Introduce separately hashed readout bindings; + leave the current artifact and macro-role merge path frozen for legacy compatibility. +3. **Extract orchestration from transport.** A session step returns observations/events and + capture requests; it does not serialize HTTP/feed headers. `flysim` owns listener setup, + wall-clock publication and systemd integration. +4. **Split emulator from task.** Move binjgb behind an environment implementation, preserving + FFI/cache behavior. Move map/exit/objective queries into Pokémon task interfaces rather + than forcing every future game to implement them. +5. **Split task recovery from storage.** The ratchet decides a task transition; storage + commits an opaque coherent capture. Matches use round resets, not milestone archives. +6. **Version observation/control at the boundary.** Internal typed session snapshots become + v1 or v2 through adapters; core modules do not depend on wire enums. +7. **Make stage state instantiable and assets descriptor-driven.** Preserve hot/cold cadence + and fixture determinism while removing singletons and hardcoded dataset selection. +8. **Only then move directories/extract packages.** Keep re-exports/CLI wrappers during the + move, fix build/CI/vendor paths, and compile tiny consumers proving Rust/TS libraries can + be used without starting Twitch, a browser, or an emulator. +9. **Reconcile documentation.** Separate current contracts/reference from dated deployment + history. Audit contradictory “raw buttons only,” weighted-macro, throughput, token/setup, + and training-improvement claims against code. Update templates and scientific limitations + with actual profile capabilities, not a new generic claim of biological fidelity. + +Do not introduce a generic reward engine, universal memory address model, central plugin +marketplace, or per-neuron network transport. Those abstractions have no demonstrated second +consumer and would obscure the useful, small seams already present. + +## 8. Implementation plan and acceptance gates + +Each row is a reviewable change or small workstream, not one large feature branch. Follow +the repository's branch/worktree build-and-review workflow. Contract changes precede their +consumers. No estimate here assumes an emulator backend or biological mapping already works. + +| Phase | Deliverable and principal files | Dependencies | Acceptance / stop condition | +| --- | --- | --- | --- | +| P0 — baseline | Trace fixtures and boundary tests around `Sim::step_frame`, restore, dataset loader, stage singleton behavior; current-state documentation inventory | None; reconcile open branches first | Current TS/Rust goldens and legacy API/feed fixtures pinned; behavior ledger distinguishes intentional transient reset from exact replay | +| P1 — identities | Dataset/profile manifest, role resolver, behavior hash, synthetic fixtures; new strict validators in both languages | P0 | Missing/ambiguous roles fail; wrong profile refuses restore; FAFB legacy fingerprints/version strings unchanged | +| M1 — MaleCNS import | `tools/datasets/malecns`, source lock, inventory/exclusion report, schema-compatible baseline bundle, attribution | P1 | Counts reconcile to selected inventory; reproducible output hashes; graph invariants and both loaders agree; no runtime downloads | +| M2 — headless characterization | Profile-specific L1/sensor binding, role/readout audit, learning-off/on benches and new goldens | M1 | Stable/finite activity measured over repeated seeds; no silent empty populations; exact TS/Rust state agreement; unsupported inputs remain marked unsupported | +| M3 — task integration | Explicit MaleCNS profile selected by service config and matching stage assets, fresh state namespace | M2 and minimal descriptor/asset support from P5 | Fresh brain on audited game state; sugar capability verified; stage index/geometry identity correct; one-hour local soak plus restore drill; no automatic promotion over FAFB | +| P2 — environment/task split | Environment contract, binjgb wrapper, task-specific inspector; facade for existing `flybrain-gb` imports | P1 | Legacy action/reward traces identical; ROM-free tests pass; optional ROM-backed sample confirms stepping/audio/state behavior | +| P3 — session library | Single-agent session owns clocks, agent/task/executor state; service becomes transport/composition wrapper | P2 | Same legacy order and compatibility; renderer/bridge restart leaves run intact; fake backend runs without binjgb/ROM/Twitch | +| P4 — multi-agent + persistence | Two-agent synthetic arena, action barrier, isolated state, session envelope/recovery and failure behavior | P3 | One world step per batch; no port/order advantage; resume/parallel-vs-sequential equivalence; failure cannot partly commit a match | +| P5 — feed/control v2 | Descriptor, scoped snapshots/events/media, targeted stimulation and idempotency, TS/Rust schema fixtures/fake server | P1, P3; validate with P4 fixture | V1 still works for legacy; multi-agent attachments cannot collide; unknown target/profile rejected; retry/reconnect tests pass | +| P6 — modular stage/bridge | Instantiable stores, dataset assets, shared/independent views, target-aware commands/redemptions | P4–P5 | PNG review for two-agent layout; browser legibility/fixture tests; correct audio ownership and no agent cross-talk | +| E1 — new emulator spike | Select title/backend; private IPC or native wrapper; record frame, ports, media, inspection and restore capability matrix | P2, can run beside P4–P6 | Reliable bounded step + simultaneous controls, pinned content/backend identity, reproducible state round trip; stop before task implementation if unavailable | +| E2 — fighting-game vertical slice | Two flies, chosen game task, analog/digital readout, episode logic, attributed rewards, match view | E1, P4–P6 | Repeated local matches and side swaps; measured compute headroom; documented scaffold and interventions; win-rate claims require controls | +| P7 — packaging/reorg | Optional root Rust workspace move, library consumers, profile-based release/asset manifests, updated infra and docs | Useful second backend + P6 | Four merge suites and affected browser gates pass; old composition deploys locally; preflight rejects incompatible multi-agent state | + +M3 may use a small v1-compatible **additive descriptor extension**, if contracts and both +consumers are updated and the existing frame semantics stay unchanged. It must not publish +MaleCNS spikes under implicit FAFB geometry. Full multi-agent publishing still requires v2. + +### 8.1 A concrete first alternative-emulator spike + +Before committing to Melee/Dolphin or an N64 backend, establish: + +- An exact title/version and backend revision, available to the operator externally. +- A supported pause/advance boundary with all configured controller ports applied together. +- Whether rendering is required for stepping and whether frame capture is synchronous. +- Sample rate/channel metadata, audio latency and timestamps. +- Analog sticks/triggers and button semantics; neutral state on disconnect. +- Save-state completeness, version/platform constraints and reproducibility after restore. +- Supported task-state inspection (match/round/port state) without guessing memory offsets. +- Headless/runtime packaging, process lifecycle, resource cost and failure behavior. + +A library used for competitive tooling may expose controller and game-state APIs without +supporting arbitrary frame stepping or faithful visual capture. Verify capabilities instead +of assuming its name solves integration. Prefer a pinned private backend process if native +embedding would force emulator internals into our Rust runtime. Desktop keyboard automation +is unsuitable for simultaneous deterministic multi-port input. + +Start with synthetic constant/alternating controller traces and a test opponent before neural +control. Such traces are backend tests, not public “fly playing” footage. Then validate one +agent, two agents, episode boundaries, crashes and resume in that order. No copyrighted game +content is added to source control or test fixtures. + +### 8.2 Validation matrix + +**Numerical/format:** existing `core/tests/golden_{toy,real,agent,restore,platformer,versions}.rs` +remain gates. Add pinned MaleCNS subgraph fixtures and a full-artifact optional golden run, +with generated TS goldens and exact Rust comparisons. Include malformed CSR, missing roles, +wide IDs, hemisphere/coordinate errors, overflowing weights, and graph/profile mismatch. +The subgraph tests prove arithmetic/loader agreement, not full-CNS dynamics. + +**Scheduler:** synthetic backend asserts one advance per complete batch; swap agent evaluation +order, vary worker count, and inject a late/failing participant. Check equal observation +boundaries, deterministic seeds, no cross-agent gains/holds/stimulation, correct tick remainder, +and no reward twice at an episode boundary. Mixed profiles are allowed only if their clocks +and capabilities satisfy the same session contract. + +**Persistence:** kill/fault injection around capture/write/manifest commit; corrupt one agent +chunk, backend state or profile hash; verify all-or-none restore and fallback reporting. +Compare uninterrupted and resumed new-session action/state traces. For legacy runs compare +against the documented transient-reset behavior instead of demanding a newly invented one. + +**Protocol/UI/bridge:** cross-language v1/v2 fixtures; two agents with different neuron counts; +shared and private views; out-of-order descriptor/media, missing optional attachments, stale +snapshots and reconnect; replay seeking; duplicate redemption, lost HTTP response and bridge +restart; explicitly unsupported stimulation. Visual changes require PNG review and browser +tests, not prose approval of a hypothetical layout. + +**Performance/science:** benchmark full graph and thresholded graph with learning disabled +and enabled, fixed sensory traces, multiple seeds and side swaps. For game-performance +claims compare against random/readout baselines and learning-off, reporting scaffold, +recovery, intervention and episode policies. Measure sustained real-time factor and tail +latency under two/four flies plus actual browser/capture load; do not extrapolate a single +kernel throughput figure to an entire show. Initial target is ≥1.0× sustained at the declared +agent count with p99 step time within its cadence budget and no growing queues; select a +resource/headroom margin from the measured backend before release. + +Before each merge run the repository-required `npm test`, `npm run typecheck`, +`cargo test --workspace` from the Rust workspace, and `infra/tests/lint.sh`. Run affected +Playwright/PNG gates for stage changes. The existing committed FAFB real-data goldens remain +mandatory. New full-MaleCNS integration jobs and ROM-backed tests are explicit optional jobs +with recorded skips, never a hidden network/ROM dependency of normal CI. + +## 9. Decisions and open questions + +**Recommended decisions now:** preserve FAFB as the baseline; use official MaleCNS provenance; +freeze legacy arithmetic/identities; version profile behavior independently; make environment +and agent separate objects; use one session barrier for a shared match; retain static plugins +and process/session isolation; introduce v2 instead of stretching Game Boy fields indefinitely. + +**Questions answered by spikes rather than assumptions:** + +1. Which MaleCNS neuron inventory and confidence/threshold policy will be the published bundle? + Can we explain the differing inventories and quantify left/right sensory coverage? +2. Do existing LIF parameters give useful, stable activity on MaleCNS? If not, which explicit + profile calibration is justified, and does a different neural model warrant separate work? +3. Are verified L1 mappings adequate, or is the intended project really an embodied sensory + simulation requiring new encoders and VNC feedback? +4. Which Smash title/backend can satisfy deterministic stepping, simultaneous ports, media + capture and restoration at acceptable cost? +5. How many simultaneous brains fit the actual budget, with which mix of within-brain versus + between-brain workers? Is a compressed media path required? +6. Which match reset/learning/intervention policy defines the show, and which defines a + controlled comparison? They should be separate run configurations. + +Success is not just “another connectome loads” or “a second pad moves.” It is a new session +assembled from modules whose anatomy, numerical model, controller, task, recovery and +presentation assumptions are explicit, testable and reusable without changing the old fly. + +## 10. Sources and scope of the analysis + +Repository evidence is enumerated in section 2 and tied to the baseline commit above. Existing +reference documents: [dataset format](../dataset-format.md), [model](../model.md), +[plasticity](../plasticity.md), [readout](../readout.md), [limitations](../limitations.md), +[macros](macros.md), [architecture tour](../architecture-tour.md), and +[contribution/compatibility rules](../../CONTRIBUTING.md). Historical status notes are not +evidence that an unmeasured experiment succeeded. + +External sources inspected 2026-09-18: + +- [Official MaleCNS overview](https://www.janelia.org/project-team/flyem/male-cns-connectome): + anatomical coverage, collaboration, release dates and licensing statement. +- [Official MaleCNS downloads](https://male-cns.janelia.org/download): versioned bulk tables, + confidence cutoffs, segment versus neuron distinction, coordinate units and API guidance. +- [FlyWire overview](https://flywire.ai/): FAFB reconstruction provenance and brain coverage. +- [Codex dataset listing](https://codex.flywire.ai/): portal inventory counts, which are not + assumed to be identical to neuPrint query inventories or local runtime graphs. +- [Minecraft fly README](https://github.com/blendi-remade/fly-brain-minecraft/blob/main/README.md) + and [provenance](https://github.com/blendi-remade/fly-brain-minecraft/blob/main/PROVENANCE.md): + a separately authored MaleCNS derivative, thresholding and mapping decisions, and disclosed + sensory/motor limitations. These moving links are comparison material, not a locked data + dependency; M1 must acquire its own official source lock. + +This analysis reads the current implementation and upstream documentation. It does not +download/build the full MaleCNS dataset, independently validate the Minecraft benchmarks, +run a new emulator, or establish a performance/behavioral improvement. Those are explicit +deliverables with gates above. diff --git a/docs/design/melee-framework-audit.md b/docs/design/melee-framework-audit.md new file mode 100644 index 0000000..e1f5f51 --- /dev/null +++ b/docs/design/melee-framework-audit.md @@ -0,0 +1,747 @@ +# Melee: multi-fly runtime, emulator and broadcast audit + +Status: **research and proposed implementation plan**. Written 2026-09-18 against local +`f7bc13a`. Extends the [modular-session design](malecns-modular-sessions.md) and +[implementation backlog](malecns-modular-implementation.md) with a concrete second game. +Existing [feed](../feed-protocol.md) and [control](../control-api.md) contracts still win. +No emulator, game image, deployment host or live broadcast was run for this audit. + +## 1. Executive decision + +**Use Dolphin, initially a pinned mainline-based Slippi Dolphin build, as a separate backend +process. Use the maintained libmelee fork to accelerate the integration spike. Keep the +brains and session coordinator in Rust.** Prove its synchronization, rendered sensory input, +and recovery behavior before selecting a production build. Keep stock Dolphin plus a narrow +backend hook as the fallback if the Slippi path cannot satisfy those requirements cleanly. + +Do not port Melee to native code, embed Dolphin into the neural crate, run one emulator per +fighter, or assume “libmelee has `step()`” supplies the complete environment contract. + +Recommended first target: + +- Melee US 1.02, local two-player versus, one emulator and two independently stateful flies. +- Existing FAFB brains first. MaleCNS is an independent axis of experimentation, not a + prerequisite for solving the emulator and multiplayer problems. +- Fixed declared characters, stage, stock/time rules and input profiles; no netplay rollback. +- Pixel-based sensory input with an explicit resolution/aspect transform; state inspection + is for task measurement, match lifecycle and display, not a hidden fighting policy. +- Fixed GameCube controller readout with bounded analog values and short frame-based holds. +- Native-resolution rendering initially, local broadcast at 30 fps first; simulation/input + continue at the backend's approximately 60-Hz cadence. Promote to a 60-fps show only after + the full media path and encoder are verified. +- Two-fly synthetic arena remains the framework test case before game integration. + +The important scaling change is **one environment with many agents**, not simply a larger +ROM. Disc size is mostly a loading/storage concern. Runtime cost comes from PowerPC emulation, +graphics/audio, multiple neural simulations, synchronization and media copies. + +### Confidence labels used below + +- **Observed in source:** checked implementation or explicit upstream documentation. +- **Recommended:** proposed architecture/configuration, not implemented here. +- **Must measure:** cannot be established by reading code, including throughput, latency, + correct input-to-frame association and reliable recovery on the target platform. + +## 2. What the Melee decompilation gives us + +The supplied [doldecomp/melee](https://github.com/doldecomp/melee) repository is a matching +decompilation of **US 1.02**. Its README is `.github/README.md`, not the repository-root +`README.md`. The inspected revision is recorded in section 13. + +The README and `config/GALE01/config.yml` identify the matching `main.dol` SHA-1 as +`08e0bf20134dfcb260699671004527b2d6bb1a45`. That identifies the executable, **not the entire +disc image**. Our run manifest must separately identify externally supplied game content, +effective executable, modifications, emulator build and task interpretation. + +The decomp builds a GameCube executable, not a supported desktop port or emulator replacement. +The `dolphin` code within that source tree refers to Nintendo's SDK, not the Dolphin emulator +project. Rebuilding/relocating a DOL for instrumentation changes address identity; never use +stock symbol addresses on a shifted build. + +### 2.1 Useful inspection map + +| Upstream source | What was observed | How it helps our task | +| --- | --- | --- | +| `config/GALE01/{config.yml,symbols.txt}`; `docs/symbols.md` | Matching binary identity and named symbols with addresses, sections and attributes | Reproducible symbol/inspection manifest analogous to the Pokémon symbol generator | +| `src/melee/pl/player.h` and `player.c` | `StaticPlayer`, getters for stocks/damage/controller index, KO-by-player counters and self-destructs | Distinguish controller port, player slot and match attribution instead of assuming they are identical | +| `src/melee/ft/types.h` | `Fighter`, player/controller identifiers, buffered sticks/triggers/buttons, pressed/released edges, damage state, source-player field and move-instance information | Audit controls and potential reward attribution; source fields are hypotheses to validate against live transitions | +| `src/melee/gm/types.h` | Match frame/timer fields, `MatchEnd`, winner arrays and exit/results structures | Episode boundaries, timeout/results interpretation, one-time terminal rewards | +| `src/melee/gm/gmvsmelee.h` | Character/stage select, versus entry/exit, sudden-death and results transitions | An explicit lifecycle model instead of treating every screen as a playable frame | +| `src/melee/{cm,gr,mp,it}/` | Camera, stages, map/collision and item subsystems identified by upstream module structure | Follow-up inspection points for view geometry, hazards and projectiles; not all audited in this pass | + +Examples of concrete distinctions: + +- `StaticPlayer` has a controller index, a player ID and up to two sub-fighter entities. + Ice Climbers and transformations defeat “one visible fighter object = one agent.” +- The input struct tracks three-entry analog/button histories plus pressed/released buttons + and threshold timers. A constant button hold and a sequence of taps are different actions. +- Damage includes an annotated source-player number, but it must be checked for projectiles, + stale ownership, self-damage and indirect KOs before it becomes a reward source. +- The match structures include winner counts/arrays. A terminal result is not safely inferred + from whichever player's stock decrement happened to be sampled first. + +### 2.2 How to use it without making the framework Melee-specific + +Create a **task-local** inspection catalog: field name, binary/decomp revision, symbol or +pointer traversal, type/endianness, valid scenes, tested transitions and unsupported cases. +Prefer Slippi telemetry for fields it supplies reliably; use decomp-grounded inspection for +missing fields only after verification. Do not expose all emulator memory to every agent. + +If memory inspection is needed, decode big-endian integer/float fields and emulated 32-bit +pointers explicitly. Never cast emulated bytes to a native Rust/C struct whose layout, +pointer width or bitfield ordering is different. Sample a consistent backend boundary rather +than reading a moving process asynchronously. Accessors named in the decomp explain semantics; +they are not functions our host process can call in place of an adapter. + +Derive small constant/schema outputs and synthetic fixtures where appropriate. Do not vendor +the entire decomp, game executable, assets or save states merely to read stocks and damage. +The first working backend does not require rebuilding Melee. Custom hooks or patches are +separately identified scaffold and enter the run's behavior/content identity. + +## 3. Emulator options and recommendation + +All candidates below emulate GameCube software. The differentiator is the host integration, +not whether Melee can theoretically boot. + +| Candidate | Evidence / useful capability | Gap or cost | Decision | +| --- | --- | --- | --- | +| **Mainline-based Slippi Dolphin + maintained libmelee** | Structured game/port events, controller pipes, explicit blocking-input support; Linux rendered path available in the ecosystem | Pixel/audio access and coherent external save-state control still need integration; pin emulator, Gecko codes and parser together | **First spike and preferred initial integration** | +| **Stock Dolphin + narrow host hook** | Source has `Core::DoFrameStep`, CPU-thread coordination and `State::{SaveToBuffer,LoadFromBuffer}` | These are internal APIs, not a stable remote environment SDK; own a small patch and state parser/telemetry bridge | Fallback or eventual generic Dolphin backend if its maintenance cost is justified | +| **Felk Python-scripting Dolphin branch** | Inspected stubs expose controller/memory/save-state scripting and rendered-frame events | Historical branch; `await frameadvance()` is documented as waiting for a rendered-frame event, not proof of a paused one-step transaction | Research reference or temporary probe, not default production dependency | +| **Custom EXI/fast-forward Slippi-Ishiiruka** | Maintained libmelee README describes accelerated ML mode and EXI inputs | That documented fast path disables rendering; inspected libmelee rejects non-Null graphics for the EXI_AI build | Useful for explicitly state-driven offline research, not the pixel-fed live baseline | +| **Libretro Dolphin core** | Potential common frontend ABI | No core-specific synchronization/state/render benchmark was performed; another integration layer does not remove task semantics | Defer rather than introduce an unverified second dependency stack | +| **Native game port built from the decomp** | Source enables modding/research | Matching DOL compilation is not native execution; graphics, OS/SDK, timing and assets remain substantial work | Outside this project's first Melee phase | + +Use the maintained **`vladfi1/libmelee`**, not a floating install selected by an old tutorial. +`altf4/libmelee` says it is archived and points there. The maintained fork says it became +the PyPI `melee` source starting at 0.45.0; pin the actual chosen package/source revision and +its dependencies instead of assuming an unversioned `pip install melee` reproduces a run. + +The maintained README describes raw-state compatibility, but the inspected `Console.step()` +still invokes `__fixframeindexing` and `__fixiasa`. Therefore, confirm field semantics from +the installed source and observation fixtures rather than trusting README wording. Our +adapter identifies parser/normalization revision as part of task identity. + +### 3.1 What the blocking path actually does + +**Observed in source:** + +1. `libmelee.Console` defaults `blocking_input=False`. Setting it true writes + `Slippi/BlockingPipes` for its mainline backend. +2. `Console.step()` flushes its registered controllers and then dispatches game/menu events + until a frame boundary. It is not a method returning RGBA pixels or arbitrary game state. +3. Slippi's `EXI_DeviceSlippi.cpp` sets `g_need_input_for_frame` on game setup, menu frames + and frame bookends. +4. `Pipes.cpp::UpdateInput` checks the blocking setting and flag, waiting for commands through + `FLUSH`. Its Linux wait path uses `select`; the inspected Windows wait helper is not implemented. +5. `ControllerInterface::UpdateInput` updates devices and only then clears the flag. The + `FLUSH` handler explicitly avoids clearing it before the other devices have been read. + +That is strong evidence for trying a Linux multi-port blocking backend. It is **not yet a +measurement** that action batch `t` affects exactly our desired frame `t+1`, that menu/game +boundaries behave identically, or that rendering/audio are coherent with the telemetry. + +Keep at most **one batch outstanding**. The pipe implementation consumes buffered commands, +and a backlog of frame batches must not collapse into a latest-state input. Create an isolated +Dolphin user directory containing only the intended pipe devices: unused/abandoned devices +can participate in updates and leave blocking input waiting on a controller nobody drives. + +For two flies, stage both full controller states before calling the single owner's +`Console.step()`. Do not give each agent a `Console` loop. Verify which input sample corresponds +to returned pre/post-frame telemetry with distinguishable action pulses and deliberate delays +on each port. “Both bots sent commands” is weaker than “both commands landed on one frame.” + +### 3.2 What libmelee does not establish for us + +- No framebuffer-returning or whole-session save/load interface was found in the inspected + `Console` API. `DumpConfig` configures media dumping; a dump is not automatically a bounded, + timestamped sensory-frame transport. +- GameCube pad values are stateful. Unchanged buttons remain held; the backend must emit an + explicit complete state or compute trustworthy deltas, including release/neutral values. +- The library applies analog normalization (`fix_analog_stick`, `fix_analog_trigger`). Define + our canonical ranges and apply conversion exactly once; test round trips at neutral, + extremes, diagonals, dead zones and trigger-click thresholds. +- Rollback skipping and internal controller flushes are present. Use local offline matches + first; do not confuse filtered rollback frames with advancing brains through speculative time. +- Initial game events can flush neutral input internally. Record startup as lifecycle scaffold; + do not attribute that to a neural decision. +- Spectator transport keepalive, pipe blocking and process liveness are separate. A healthy + connection does not prove the match or renderer is advancing. + +## 4. Audit of our current system: keep, extract, replace + +Rust path aliases: `core/`, `gb/`, `sim/` mean the respective crates under +`services/flysim/crates/` named `flybrain-core`, `flybrain-gb`, and `flysim`. + +| Finding | Current evidence | Required change for Melee/multiple flies | Priority | +| --- | --- | --- | --- | +| Single world and brain bundled together | `sim/src/simloop.rs::Sim` owns one `NeuralAgent`, concrete `Emulator`, adapter and ratchet | Session owns one environment plus agent collection and explicit port map | Blocking | +| Direct Game Boy calls in frame loop | `step_frame`, `to_button_mask`, `set_buttons(u8)`, `run_frame`, fixed framebuffer copy | Backend interface with full action batch, rational cadence, observations and media capabilities | Blocking | +| Task trait carries Pokémon concepts | `gb/src/adapter.rs` includes `read8(u16)`, map/exit/objective hooks | Keep inspector and macros task-local; use generic reward/episode/progress outputs | Blocking | +| Timing defaults embed Game Boy | `core/src/agent.rs`, `sim/src/config.rs` | Session clock derived from backend, per-agent neural remainders; preserve old arithmetic in legacy facade | Blocking | +| Decoder tuned to walking through Pokémon maps | `packages/brain/src/readout/presets/gameboy.ts`: 800-ms directions, 85-ms A/B pulses with 480-ms cooldown | New fixed GameCube mapping; frame-scale controls, analog sticks/triggers, concurrent movement/action | Blocking | +| Neural code is already independently useful | `core/src/lif.rs`, `agent.rs`, `plasticity.rs`; TS counterparts | Reuse exact kernel and private per-agent state; do not replace neural semantics to integrate a game | Keep | +| Graph can be shared, pool cannot be concurrently dispatched | `Arc`, `SweepPlan`, `core/src/pool.rs` shared job slot | Immutable topology shared; distinct mutable state and controlled total scheduling budget | Blocking | +| CUDA exists but not an automatic service optimization | `core/src/lif/cuda.rs`; no `enable_cuda` call found in `sim/src/simloop.rs` | Explicit backend selection, equivalence/restore tests, profile first; don't promise GPU brains from graphics availability | Measured option | +| Snapshot header is a single Pokémon-shaped view | `sim/src/snapshot.rs`, `packages/feed/src/{types,codec}.ts` | Session descriptors, multiple agents/ports, task-specific progress, named attachments/media | Blocking for proper broadcast | +| Stage assumes Game Boy geometry and one fly | `apps/stage/src/{App.tsx,feed/store.ts,feed/decode.ts,lib/geometry.ts}` | Per-session/agent store instances, dynamic view aspect/dimensions, multi-agent match layout | Blocking for proper broadcast | +| Browser owns game audio | `apps/stage/src/audio/engine.ts`, 48-kHz feed, Pulse capture | Explicit audio producer and media-clock policy; do not play native Dolphin audio and forwarded PCM twice | Blocking | +| Capture already offers NVENC | `infra/bin/flycast-launch` | Reuse encode/relay/recording; measure new compositor/readback cost and revise 60-fps settings | Keep with changes | +| Encoder hardcodes H.264 level 4.1 | Both encoder functions in `flycast-launch` | 1080p60 needs a suitable level (normally 4.2 or automatic selection); changing only `FLY_FPS` is insufficient | Required for 1080p60 | +| Existing checkpoint payload is single-agent/binary-specific | `sim/src/store.rs`, `gb/src/compatibility.rs` | Coherent all-agent/world checkpoint plus backend/parser/patch/controller identity | Blocking for exact resume | +| Checkpoint writer queue is unbounded | `Sim::start_writer` uses `std::sync::mpsc::channel` | Bound/coalesce background work; larger emulator captures and several brain copies must not grow an unlimited queue | High | +| Existing “saved” event precedes durable commit | `checkpoint_with_reply` emits after enqueue; writer updates durable metrics on success | Distinguish capture/enqueue/commit/failure events; public status must not report a queued Melee save as durable | High | +| Recovery assumes a best progress ladder | `gb/src/{ratchet,recovery}.rs` | Match/episode reset policy; never rewind one player's world independently | Blocking | +| Health is mostly loop heartbeat | `sim/src/simloop.rs::Shared`, `sim/src/lib.rs` | Distinguish waiting at input barrier, intentional pause, backend timeout and deadlock; keep host supervision responsive | High | +| Deployment resource partitions reflect the old stack | `infra/units/flysim.service`, deploy cpuset construction | Budget Dolphin CPU/GPU plus N brains and media; measure and set new memory/process limits | Required before release | +| Bridge targets one sim | `services/bridge/src/{sim,commands,redemptions}.ts` | Explicit stable agent targeting and intervention policy; no viewer control-port endpoint | Before interactive show | + +Do not interpret dated CPU/GPU measurements in the repository as current free capacity. +The records identify bandwidth contention and graphics-sharing constraints, but this audit +does not inspect the host or establish that it can run two brains plus Dolphin in real time. + +## 5. A framework architecture that survives a third game + +### 5.1 Four runtime components, not one enormous adapter + +```text + private backend protocol +Rust session process <----------------------------------------------> backend helper + session clock / port ownership Python + libmelee initially + N agent states owns Dolphin process/user dir + sensory encoders │ + fixed readouts ├─ all controller pipes + task events / reward router ├─ telemetry/parser + episode policy / checkpoints └─ media + state hooks + │ │ + ├─ feed/control adapters Dolphin + └─ durable event/checkpoint store one local match + │ + stage/compositor → capture → local relay → optional Twitch push +``` + +The helper is an **internal backend implementation**, not an audience-accessible controller +service. The session remains the sole authority assigning actions to ports. Python never +simulates the neurons or selects actions. Keep it if measurements say its overhead is small; +replace its internals with Rust/native IPC only when an actual bottleneck or maintenance +requirement justifies it. + +An external emulator is a normal `Environment` implementation, not a special `if melee` +branch sprinkled throughout the session. For Game Boy, the same interface has an in-process +implementation. For an embodied world, it may be a native physics engine. Consumers do not +need to know which one owns the world. + +### 5.2 Framework contracts to extract + +| Contract | Owns | Must not know | +| --- | --- | --- | +| `Brain` / numerical core | Tick semantics, state, spikes, rates, learning updates | Game, process, controller labels or viewer | +| `SensorEncoder` | Declared observation→neural drive transform | Reward inspector state not declared as input | +| `Readout` | Fixed rate/signal→control-channel mapping | Opponent strategy, game addresses or pathfinding | +| `ActionExecutor` | Selected action→controller state; optional declared macro lifetime | Authority to invent a winning action when brain is silent | +| `Environment` | Native clock, port schema, action commit, observations/media, snapshot capabilities | Neural roles, Twitch or task reward weights | +| `Task` | Typed state interpretation, rewards, progress and episode outcomes | Direct neural mutation or direct controller writes | +| `EpisodePolicy` | Start/end/reset/recovery semantics | Hidden per-player rewind in a shared world | +| `Session` | Barrier, identity, agent isolation, routing, state capture and supervision | Melee memory offsets or Pokémon map IDs | +| `Presentation` | Descriptor-driven layout, media/audio and task panels | Emulator stepping or access to controller pipes | + +Use modules first, then crates/packages as second consumers appear. The existing monorepo +and Rust workspace can stay in place during extraction. Keep static compiled registries +initially; a stable dynamic-plugin ABI and public package registry are not prerequisites. + +**Framework acceptance test:** adding a synthetic third environment/task requires a backend +implementation, a task/profile and a composition manifest, not edits to the session loop, +neural core, protocol enums or generic stage store. A game-specific presentation plugin is +allowed. Wire extensions must be namespaced/schema-validated, not hardcoded into every panel. + +### 5.3 Minimal backend IPC + +Specify and test a small private protocol before implementing a network-shaped abstraction: + +```text +Hello → backend build/content/patch identity, capabilities, cadence, ports, views +Initialize(runConfig) → Ready(epoch, observationBoundary) +Advance(epoch, expectedBoundary, completePortBatch) + → StepResult(epoch, newBoundary, appliedBatchId, observation, mediaRefs) +Pause / Resume → explicit acknowledgment +Capture(epoch, boundary) → capture token + state digest + snapshot bytes/reference +Restore(captureToken) → new epoch + restored observation + success/failure +Shutdown → acknowledgment or bounded forced process termination +``` + +Only advertise `Capture/Restore` if implemented and tested. Otherwise expose an explicit +`restart_episode` capability and visible aborted-match policy; never claim exact resume. +Port batches use canonical buttons plus sticks in [-1,1] and triggers in [0,1], converted +once by the backend. Descriptor/schema versions pin conversions and active ports. + +Use a local framed socket for commands/small observations; use a bounded shared-memory ring +or equivalent for large media. Include epoch, frame/sample identity, dimensions, format and +generation on media references. Release/acquire ownership and slot lifetime prevent the +backend overwriting a sensory buffer while an encoder reads it. Reconnection invalidates old +handles; a stale frame must not be silently accepted because its byte length matches. + +One request in flight, explicit timeouts, bounded queues. After an uncertain `Advance` +response, do **not** resend blindly: the world may already have advanced. Resolve batch ID/ +boundary through an idempotent reply cache or fail/recover the session. A pipe transport with +no application acknowledgment needs validation against returned input telemetry, not a claim +of atomicity it cannot prove. + +Separate lifecycle I/O from the blocked advance operation so diagnostics/shutdown stay alive. +Do not issue a save operation scheduled on Dolphin's CPU thread while that same thread is +waiting forever for pipe input. Capture/pause needs a backend-owned quiescent point where +both the simulation and pending input consumption have known state. + +## 6. Multiple flies and the timing contract + +### 6.1 One match, one world clock + +Two flies in Melee normally means two controller ports in **one Dolphin instance**. Four +flies means four ports, not four copies of Melee joined through netplay. Several independent +matches are separate sessions/process trees and can share one broadcast director. + +For each committed boundary: + +1. Freeze each agent's allowed observation from the same world state. +2. Advance every brain by the environment's elapsed emulated time using its private remainder. +3. Decode independent actions; step any declared executors; assemble a complete port batch. +4. Commit the batch once and let the backend advance to the next acknowledged boundary. +5. Associate rendered sensory frames and telemetry with their actual producing boundary. +6. Route task events/rewards exactly once; encode the next input and publish a snapshot. + +Warm-up settles/calibrates brains with learning off while the environment is held at its +initial boundary. The next observation must not be from a game that ran freely through the +warm-up. Fixed offsets between brain and environment clocks are recorded. + +Use the selected backend's rational emulated cadence, not `GAMEBOY_MS_PER_FRAME` or an +unexamined exact 60. An approximately 60-Hz budget is about 16.7 ms, but logical game frame, +video interrupt, input poll, rendered presentation and Slippi frame bookend are distinct +events until the spike establishes their mapping. Session ticks are monotonic even when +Melee's signed frame number resets, starts before zero, or changes across menu scenes. + +### 6.2 Render latency and agent fairness + +Dual-threaded graphics can present frame `n` after telemetry for `n` is available. Label the +actual frame; do not attach “latest screenshot” to current state and assume equivalence. +Start with one declared fixed observation latency shared by all agents, and measure it. +If a pipeline deliberately adds one frame of latency, record that in the sensor profile. +The full-resolution spectator view may be delayed separately, provided overlays use the +matching presentation timestamps rather than future task data. + +Also test a potential pipeline deadlock: the helper waits for a rendered image while Dolphin +is waiting for the next controller flush needed to reach that presentation event. Fix the +backend's rendezvous or select a declared previous-frame sensory latency; do not unblock it +with an undisclosed neutral gameplay input. Game-state bookends alone do not prove the GPU +has completed a matching frame. + +Do not reduce brain integration from 1,000 to 500 ticks per emulated second to meet wall-clock +deadlines. That changes the model. A declared action-repeat interval can reduce decisions, +but normally still requires all neural ticks and correctly accumulated intermediate task +events; it does not halve the principal neural cost. Rendering every second game frame is +also a sensor change if the brain otherwise sees each frame, not merely a broadcast setting. + +On slow compute, the default is to slow the entire local match and report real-time factor. +Do not let one fly continue while the other misses turns. On a dead participant/backend, +pause or abort the match visibly; neutral fallback play is not silently substituted. + +### 6.3 No rollback netplay in the first release + +Slippi supports online play, but we do not need it to connect two local flies. Online rollback +would require rewinding **all** neural states, RNG, decoder/executor state, reward ledgers +and admission decisions at the same speculative boundary as the game, then replaying inputs. +Filtering repeated frames in libmelee is not that system. Keep offline local matches and +assert monotonic committed observations per epoch; classify unexpected rollback as an error +or explicit recovery transition rather than double-rewarding it. + +## 7. Controller, sensory and learning design + +### 7.1 A GameCube controller is not an eight-bit pad + +Support independent main and C sticks, analog shoulders, digital trigger clicks, face +buttons, start and D-pad. Movement and attack can overlap. Canonical neutral/release state +must be complete, so a missing command cannot accidentally leave attack or shield held. + +At about 60 Hz, the current 800-ms direction hold is roughly **48 game frames** and the +85-ms pulse roughly five. Reusing these values would dominate the fly's behavior regardless +of the dataset. Create a fixed Melee readout with explicit decisions in integer game frames, +bounded analog mappings, dead zones, tie handling and pulse/hold policies. Start with a small +declared set of stick magnitudes and directions if that makes validation easier; continuously +valued mappings can follow as a separate profile. + +Audit tap-jump, directional aerials/smash attacks, jump release, shields, simultaneous axes, +and conflicting inputs. Don't add state-conditioned auto-aim, automatic edge recovery or +combo execution under the label “controller mapping.” If later desired, publish those as +separate macro/assistance profiles with their own identity and comparison baseline. + +Start/system controls are lifecycle-sensitive. During an active match the profile may omit +pause entirely; initial match setup and between-match reset are disclosed episode scaffolding. +This is not permission for an API to press buttons. Specify whether setup uses an audited +initial state, internal deterministic menu setup, or a reset hook, and mark those frames as +non-neural setup with learning disabled. That expands the legacy “all presses” phrasing and +requires a deliberate task-policy/documentation decision before shipping it. + +### 7.2 What the fly sees + +For the initial pixel profile, both flies receive the same shared game camera with fixed +crop/aspect treatment, independent of spectator overlays. The legacy 160×144 input should +not stretch a 4:3 scene silently. Compare an aspect-preserving downsample/letterbox transform +with a separately versioned input-size profile; changing kernel retina dimensions currently +affects the numeric configuration identity. + +Current L1 projection samples luminance at 1,572 columns. Higher broadcast resolution does +not produce more sensory neurons, color recognition, motion estimation or knowledge of which +fighter the fly controls. Test whether each selected character remains visible across zoom +and stage movement; record the sparse sensory representation rather than assuming a human- +readable video is an adequate neural input. + +Three distinct modes must not be conflated: + +| Mode | Neural observation | Rendering implications | +| --- | --- | --- | +| Pixel baseline | Actual game image through fixed encoder | Needs real rendered frames even if no desktop GUI is shown | +| Structured-state experiment | Explicitly encoded positions, velocities, stocks, etc. | Can potentially use Null/fast-forward, but it is a new privileged-input model | +| Spectator-only rendering | Whatever the profile specifies; video for audience | May be independently compressed/delayed, never silently substituted for sensory input | + +“Headless” can mean no GUI while still rendering; “Null graphics” generally means no useful +pixel observation. The documented EXI fast-forward speed path cannot be advertised as the +performance of our pixel-fed broadcast. + +### 7.3 Reward and outcome attribution + +Implement a `MeleeTask` with typed per-player observations, match state, a ledger and positive +reward events. Its schema belongs to the task, not a generic `GameMode` enum. Start small: + +- Terminal match outcome, once, based on validated results/termination reason. +- Opponent damage and credited KOs only after verified ownership information is available. +- No reward for mere button activation, for losing a stock, or for scripted setup. + +Do not reward A for every increase in B's percent: self-damage, stage effects, reflected +projectiles, teams and another sub-fighter can invalidate that inference. The decomp's source- +player and KO tables guide inspection, but no runtime correctness is claimed until tested. +Unknown attribution produces a logged observation without a guessed reward. Keep fractional +damage until the rule deliberately quantizes; HUD damage and fighter damage may differ. + +Deduplication keys include epoch/match, producing frame and event identity. Handle multihits, +trades, simultaneous KOs, respawn percent reset, timeout, sudden death and disconnection as +separate cases. If source telemetry cannot distinguish a required case, narrow the first +ruleset or add a specific audited observation hook. + +Learning remains private per fly and synthetic reward modulation remains distinct from PAM +stimulation. Disable learning during kernel/controller/backend characterization; later compare +learning-on with learning-off, repeat seeds and swap sides/characters. Retaining gains between +rounds is a run policy. Competitive success is not guaranteed by increased model complexity. + +## 8. Performance plan: measure the actual critical path + +### 8.1 Budget equation + +For a lockstep pixel-fed match, approximate the critical wall-time interval as: + +```text +T_step = T_brains + T_readout/task + T_controller_IPC + + T_emulation_to_observation + T_required_render_readback + T_boundary_overhead + +T_brains ≈ sum(T_agent_i) [sequential evaluation] +T_brains ≳ max(T_agent_i) + barrier cost [parallel with sufficient independent resources] +``` + +The parallel estimate is a lower bound, not a promise: shared caches, memory bandwidth, +GPU contention and scheduling can make every brain slower. Media publication/encoding and +storage should be off the critical path, but their resource use and state capture still +affect it. Do not obtain a “60 fps” claim solely from Dolphin's display counter while the +brains advance fewer milliseconds or repeat stale observations. + +Proposed capacity gate: warm full-stack **unthrottled** throughput at least 1.2× the selected +game cadence for two flies, then a paced one-hour soak with no growing queues/lag and a +24-hour local endurance run before release. In the paced run, distinguish intentional wait +from compute time; report p50/p95/p99/max compute interval and deadline misses. The 1.2× +margin is a proposed engineering target, not a measured capability of current hardware. + +### 8.2 CPU/GPU strategy + +1. **Keep the first two brains on CPU.** Establish Dolphin JIT/render/media cost separately. + The repository's current service is CPU-composed even though a CUDA kernel exists. +2. **Allocate a total physical-core budget.** Compare sequential agents with modest within- + brain pools against agents running concurrently on disjoint core groups. Include Dolphin's + CPU/JIT thread, graphics worker, helper, browser, encoder and storage in the budget. Do not + launch four copies of the current per-brain pool size by default. +3. **Use native-resolution hardware rendering first.** Compare OpenGL/Vulkan on the chosen + build/platform; no blanket claim that one is faster. Measure render correctness and readback. + JIT is the performance baseline; an interpreter is a diagnostic baseline, not the live plan. +4. **Benchmark shader compilation and caches.** Report cold and warm starts separately. Choose + supported shader modes from measurements; a cache that hides startup hitches is not a + guarantee that a new stage/character will not compile something mid-match. +5. **Use NVENC where available, with a measured fallback policy.** Its encode engine does not + remove GPU rendering, memory allocation, color conversion or framebuffer-readback cost. + Automatic fallback to x264 can consume the cores the brains/emulator need; expose the + resulting degradation and test whether the declared session can still meet cadence. +6. **Only then test CUDA brains.** The current backend retains RNG/plasticity observation/rate + work on the host, uploads state inputs, and by default synchronizes membrane/refractory + state back each batch. It also allocates device graph/state per backend instance. Measure + two/four agents alongside Dolphin, browser graphics and NVENC; zero-copy shared graphs and + a GPU-wide scheduler are possible later work, not present features. + +For CUDA, preserve bit-exactness, gain-update ordering and checkpoint synchronization. +Batching across a future game action boundary is not valid just because it improves kernel +throughput. Keep the TypeScript oracle and existing version strings intact. + +The existing VirtualGL/Xvfb result demonstrates one Chromium rendering path, not that +Dolphin Vulkan/OpenGL works or is performant in the same container. Test the complete selected +graphics path. Reusing GPU passthrough requires no assumption of exclusive VRAM availability. +Resource availability must be measured in an approved, serialized deployment-host session. + +### 8.3 Media bandwidth and copies + +Uncompressed RGBA estimates, before copies/framing: + +| Image/cadence | Bytes per second | +| --- | ---: | +| Existing 160×144 at 30 fps | 2.76 MB/s | +| 640×480 at 30 fps | 36.86 MB/s | +| 640×480 at 60 fps | 73.73 MB/s | +| 1920×1080 at 60 fps | 497.66 MB/s | + +640×480 is a planning example, not an asserted fixed Dolphin framebuffer size. The backend +advertises actual dimensions/format/aspect. One shared camera is delivered once for the +match; two flies can sample one immutable image without duplicating its transport. If their +sensor transforms differ, encode separately against the same source frame. + +Prefer reducing/downsampling the sensory copy close to the renderer, ideally before GPU +readback, while preserving a separately timestamped spectator view. Benchmark against an +ordinary CPU path before adding device-buffer interop. A shared-memory ring removes socket +copies, not the GPU fence/readback itself. Slow viewers may drop frames; required sensory +frames may not disappear silently from the neural run. + +### 8.4 Benchmark ladder and decision records + +| Run | Configuration | Question / recorded output | +| --- | --- | --- | +| B0 | Synthetic two-port backend, no neurons | IPC latency, one-step semantics, barriers, timeouts, media-buffer ownership | +| B1 | Dolphin with fixed input traces, rendering/audio on, no brains | Cold/warm emulator cost, step timing, port alignment, render-to-state latency | +| B2 | Same run + helper/media extraction | Incremental parsing, copying, downsampling and audio cost | +| B3 | One FAFB brain | End-to-end reference and per-phase costs | +| B4 | Two FAFB brains, sequential vs parallel schedules | CPU/cache/bandwidth limits, balanced observation and input timing | +| B5 | B4 + actual stage, capture, relay, recording, checkpoints | Full critical path, A/V drift, encoder fallback and queue growth | +| B6 | Four brains and four active ports | Capacity characterization only until this independently passes the same gates | +| B7 | Matched MaleCNS and optional CUDA variants | Dataset and backend effects, measured independently before combined variants | + +Record backend/content/patch/profile digests, physical-core allocation, exact graphics settings, +sensor/broadcast resolutions, all clock rates, resident/peak memory, VRAM, thread usage, +real-time factor, latency distributions, audio under/overruns and dropped frames by purpose. +Store operator-specific machine details externally and publish only the portable methodology +and non-identifying results. No measurements were performed by this document-writing task. + +## 9. Broadcast architecture and audio ownership + +### 9.1 Two viable routes + +**Route A — stage receives game media.** Closest to the current architecture: backend emits +pixels/audio, stage composites game and overlays, ffmpeg captures the page. Start the local +prototype with bounded lower-resolution raw frames to validate semantics. If bandwidth and +copying dominate, add a compressed local media track (for example WebRTC) while telemetry +remains on the feed. Avoid encode→decode→encode unless its measured simplicity/latency tradeoff +is acceptable. Browser frame presentation timestamps must align overlays with displayed video. + +**Route B — compositor combines native game output and stage overlay.** Dolphin supplies its +rendered output to a compositor; the browser supplies a separate overlay surface. This can +avoid moving full-resolution game pixels through JavaScript, but requires an explicit shared +clock and a new capture composition. The fly's sensory image still needs a frame-identified +path from the backend. Capturing a desktop window on a wall clock is insufficient to establish +which image a brain used at a given game boundary. + +**Recommendation:** prototype Route A for the two-player local slice; benchmark Route B in +the media spike before committing to the long-run high-resolution pipeline. Preserve media +as a capability behind the environment interface so the choice does not change brain/task code. +The v2 protocol should be able to reference media streams, not mandate all video as WS RGBA. + +### 9.2 Audio and clock policy + +Today the page plays binjgb PCM and stream SFX into the Pulse sink. With Dolphin, select one +of these explicitly: + +- Dolphin PCM is captured/forwarded and played by the page, with native device output muted. +- Dolphin renders audio to the capture sink directly, and the page contributes only SFX. + +Do not run both. Declare sample format/rate, resampling location, timestamps, buffering and +discontinuity handling. Pause/reset/restore must flush or relabel buffered old-episode audio. +If wall time falls behind, measure pitch/time-stretch behavior; do not let “async resample” +hide minutes of simulation lag. Game timestamps, not arbitrary browser receipt time, define +the intended A/V relationship. + +Keep 30-fps broadcast and approximately 60-Hz gameplay as independent settings. For 1080p60, +the existing H.264 level 4.1 is too low for the normal macroblock-rate requirement; use a +compatible level such as 4.2 or encoder-selected level and validate the actual stream. Also +measure capture/compositor cadence, bitrate quality, encoder lookahead/latency, local recording +and audio synchronization. `FLY_FPS=60` alone is not a completed performance upgrade. + +### 9.3 Presentation changes + +The generic stage needs descriptor-driven game aspect ratio, two/four agent cards, per-port +button/stick indicators, per-agent learning/sugar state, shared match stocks/percent/results, +and scoped events. Neither “badges” nor “highest ladder rung” describes a match. + +Separate task data schema from layout. Keep one compositor clock and one selected world audio +stream; keep neural maps and rate scalers private per agent/dataset identity. Defer expensive +four-avatar/whole-connectome rendering until measured. All actual layout decisions require +PNG mockups and existing legibility/browser checks. This text does not approve a screen. + +## 10. Persistence and unattended operation + +### 10.1 Savestates are a capability, not a libmelee assumption + +Dolphin source provides buffer/file state operations, but `State.h` documents that operations +called off its CPU thread may be scheduled rather than executed immediately. An external +“save requested” is therefore not proof of a consistent capture at our agent boundary. +Slippi's internal rollback save-state commands likewise do not constitute an audited public +multi-agent checkpoint API. + +The selected backend must supply acknowledgment of the frozen boundary, resulting state +digest and completion. Save every brain, RNG, rate/calibration state, learning state, sensor/ +executor state, pending action identity, task ledger and environment together. Clock and +media epochs change on restore; discard pre-restore spectator/parser buffers. + +Re-create or explicitly reinitialize libmelee's parser caches and controller history after +restore. An emulator savestate does not include an external helper's `_frame`, previous +game state, normalization state or queued pipe data. Test how game-start metadata is supplied +when loading into mid-match; some telemetry protocols may need reseeding or restarting. + +When exact mid-match capture is unavailable, an initial prototype may visibly abort and +restart a match while retaining a declared brain checkpoint. Mark `resume=episode-restart` +in the descriptor. That is a deliberate narrower capability, not equivalent to crash resume. +It is not ready for a release that promises uninterrupted exact match continuation. + +### 10.2 Storage and health + +Bound pending captures and coalesce replaceable hot checkpoints. A durable request either +completes with a commit acknowledgment or fails explicitly; never drop it while reporting +success. Match-end result records are append-only and independent of world rewind. + +Keep backend/helper/brain health separate from “game frame did not advance.” An intentional +pause or waiting barrier is not a crash; a dead helper must not keep the session green by +merely refreshing an HTTP heartbeat. Export last completed boundary, in-flight request age, +barrier participant status and renderer progress. A hard stall has a timeout and explicit +match abort/recovery path, not repeated blind restarts of unrelated services. + +Manage Dolphin under the session's lifecycle or an explicitly coordinated systemd unit. If +Dolphin restarts, the session cannot keep sending frame `t+1` to a fresh match. Allocate unique +user directories, pipe names, telemetry ports and state namespaces per independent session. +Use deterministic configuration provisioning and checksums instead of reusing a developer's +desktop Dolphin settings or permitting auto-updates. + +The existing `flysim.service` memory ceiling and CPU partition were sized for a different +process graph. Set new cgroup/resource limits from measured high-water marks; account for +backend process, multiple brain copies and checkpoint transients. Release preflight checks +the complete backend/game/patch/parser/profile identity. Rollback retains compatible state +as well as the old executable. + +## 11. Staged implementation and go/no-go gates + +This specializes the existing backlog rather than replacing its foundation/session work. +Melee-specific spikes can start before the full framework reorganization is finished. + +| Item | Work and dependency | Evidence required before the next step | +| --- | --- | --- | +| **MELEE-01: backend selection spike** | Specializes EMULATOR-01. Pin mainline Slippi + maintained libmelee, content and Gecko codes; use isolated user config and two synthetic controllers | Boot/render/audio; block one then both ports; one batch/frame mapping; menu→match→results lifecycle; cleanup/restart. Choose this build or stock Dolphin + narrow hook based on results | +| **MELEE-02: capture/restore spike** | Alongside MELEE-01; prove sensory-frame identity, media export and save/load acknowledgment independently | Fixed pixel↔telemetry latency, bounded media storage, correct input after restore, parser/cache recovery. Explicit decision: exact resume or prototype-only episode restart | +| **MELEE-03: task observation audit** | Pin decomp; build field catalog, typed parser/inspector, lifecycle and synthetic event fixtures | Verified port/player/sub-fighter mapping; stocks/results; no guessed rewards; content/patch mismatches visibly disable unsupported semantic interpretation | +| **FRAMEWORK-01: generic backend + session** | Existing FOUNDATION-01/02 and RUNTIME-01/02; add the private IPC implementation behind `Environment` | Same legacy Game Boy traces; headless synthetic environment uses identical session API; no Melee branches in core loop | +| **FRAMEWORK-02: multi-agent and state** | Existing RUNTIME-03/STATE-01; integrate complete action batches, worker budget and chosen backend recovery capability | No cross-agent state leakage; one world step; changed evaluation order invariant; failed restore cannot partly install a match | +| **MELEE-04: fixed readout and sensory profile** | MELEE-01/02 + framework boundary; TS specification then Rust implementation for any new decoder semantics | Neutral/release, analog conversion, tap/hold/direction combinations, aspect-preserved neural input and recorded latency; no hidden combo/aim policy | +| **MELEE-05: first two-fly match** | MELEE-03/04 + FRAMEWORK-02; learning off, then audited positive rewards | Recorded action/observation timelines; paired side/seed trials; match terminal deduplication and visible reset/failure semantics | +| **MEDIA-01: full local show** | Existing WIRE-01/PRESENTATION-01; compare Route A/B, define audio owner and 30/60-fps profiles | PNG review, fake multi-agent fixtures, measured copies/latency/A/V drift; sustained media pipeline under checkpoint and shader-load events | +| **PERF-01: two-fly capacity gate** | Benchmark ladder B0–B5; optimize measured limiting phase | ≥1.2× warm unthrottled capacity target, one-hour paced soak and 24-hour local endurance; no queue/lag growth; documented CPU/GPU/memory envelope | +| **MELEE-06: expand carefully** | Passing two-fly slice | Four ports/teams, broader characters/stages, MaleCNS and CUDA are separate experiments, each with new tests and its own capacity result | +| **FRAMEWORK-03: finish packaging** | Existing PACKAGE-01 after useful second backend | Example third backend can be added without core/session/schema edits; isolated compositions package and preflight correctly | + +**Stop conditions:** no dependable step/input barrier; no identifiable pixel source for a +pixel-input claim; inability to attribute rewards under the claimed ruleset; unsupported +restore marketed as exact resume; or sustained capacity below the declared cadence. +Respond by changing the explicit supported scope, backend or resources—not by quietly skipping +neural ticks, adding a gameplay bot, hiding game stalls or reporting guessed measurements. + +### 11.1 Suggested first experiment script + +The first implementation should be a local measurement harness, not the final stream: + +1. Launch one pinned backend with known local content and two configured bot pads. +2. Enter a fixed local match through the declared setup procedure; record episode boundary. +3. Send distinct short left/right and A/jump pulse patterns on each port, including neutral + frames; log intended batches and observed raw/processed controller values. +4. Delay one port by a controlled wall-clock interval and verify no game boundary commits + until the complete batch is available. Repeat with port order reversed and four ports. +5. Capture a sequence of images/telemetry with frame identities; measure their association. +6. Save/restore at a known barrier if supported, replay the same actions, and compare task/ + input traces; verify helper state and buffered media are reset coherently. +7. Kill the helper/backend separately and verify bounded failure without accidental continued + play or permanent hangs. Test paused-state health independently. +8. Measure compute with no brains, one brain, two brains, then the full broadcast stack. + +Synthetic controller traces are test machinery, not footage presented as neural play. Keep +game content and environment-specific records outside source control; store portable metrics, +synthetic schemas and independently authored tests in the repository. + +### 11.2 Test matrix that catches Melee-specific failures + +- **Input:** two/four ports, inactive slots, delayed/missing flush, stale buffered commands, + full release, analog endpoints/deadzones, short taps, pressed versus held edges. +- **Identity:** controller↔player mapping, swapped ports, sub-fighters, transformations, + character/stage changes, wrong game revision, changed patch/parser normalization. +- **Events:** multi-hit, trade, self-damage, projectile ownership, stock reset, simultaneous + KO, timeout, sudden death, results re-entry, disconnect, restart after accepted reward. +- **Clocks/media:** game-frame reset, renderer lag, stale shared-memory generation, dropped + spectator frame versus required sensory frame, paused audio, mismatched overlay timestamps. +- **Recovery:** all-agent atomic validation, one corrupt state chunk, backend import failure, + helper parser not reinitialized, asynchronous save completion, hot-store coalescing and + durable-write failure. Test new exact resume separately from legacy transient-reset behavior. +- **Performance:** cold/warm shaders, high-activity matches, checkpoint capture bursts, CPU + encoder fallback, browser reconnect, two/four neural agents and measured GPU contention. + +All implementation merges retain repository-required TS tests/typecheck, Rust workspace +tests and infra lint; UI changes add Playwright and PNG review. Game-backed jobs are explicit +operator-provided tests. Normal CI uses synthetic observations/backends and existing goldens. + +## 12. Decisions to carry into implementation + +| Question | Recommended answer now | Still requires evidence/choice | +| --- | --- | --- | +| Which emulator? | Dolphin, first trying mainline-based Slippi + maintained libmelee | Exact build selected by synchronized-input/media/state spikes | +| Use the decomp to run the game natively? | No; use it to audit task/state/controller semantics | Custom instrumentation only for specifically missing observations | +| One emulator per fly? | No for one match; one per independent session | Four-port capability must be tested, not inferred from two ports | +| Which brain? | Two existing FAFB agents for integration baseline | MaleCNS comparison after mappings/dynamics pass their independent gates | +| CPU or GPU brain? | CPU baseline, share immutable graph | CUDA versus CPU benchmark under Dolphin + capture, not in isolation | +| Inputs to the brain? | Pixels with explicit fixed transform | Structured state is a distinct optional research profile | +| Start with macros? | Fixed controller mapping, no hidden aim/combo policy | Any later assist profile is separately disclosed and evaluated | +| How fast? | Backend-native gameplay/input cadence, 30-fps initial show | Full-stack two-agent capacity; optional 60-fps broadcast and four flies | +| How to resume? | Whole-session coherent state where supported | Episode-restart prototype if exact state interface is not yet available | +| How generic? | Concrete environment/task/agent/session contracts and composition examples | Extract public packages only after second/third consumers prove the seam | + +The first operator choices needed are the initial characters/stage/ruleset, desired show +cadence, and whether a visibly restarted match is acceptable during the prototype. They do +not block the synthetic framework work or source-level backend spike design. + +## 13. Sources and audit scope + +Local code evidence appears in section 4. Additional local files inspected include +`core/src/lif/cuda.rs`, `sim/src/pacing.rs`, `sim/src/simloop.rs::start_writer`, +`packages/brain/src/readout/presets/{gameboy,platformer}.ts`, +`apps/stage/src/audio/engine.ts`, `infra/bin/flycast-launch`, +`infra/units/flysim.service`, and the profiling/VirtualGL methods under `infra/docs/`. + +External source snapshots inspected on 2026-09-18 (pin actual dependencies again at spike start): + +| Repository/ref | Observed revision | Files used | +| --- | --- | --- | +| [doldecomp/melee](https://github.com/doldecomp/melee/tree/b9ec8a2eb48520753b2f8159ccc94d033fbf60ea) `master` | `b9ec8a2eb48520753b2f8159ccc94d033fbf60ea` | `.github/README.md`, `docs/symbols.md`, config, player/fighter/match headers and player implementation | +| [vladfi1/libmelee](https://github.com/vladfi1/libmelee/tree/bce21f09984b286e6d36bfd2939e4cd4691f94c2) `master` | `bce21f09984b286e6d36bfd2939e4cd4691f94c2` | README, `melee/console.py`, `melee/controller.py`, license metadata | +| [project-slippi/dolphin](https://github.com/project-slippi/dolphin/tree/41a7a3a110ed52999486ae1901c8fbb9a63d4f13) `slippi` | `41a7a3a110ed52999486ae1901c8fbb9a63d4f13` | Pipe backend, controller update loop, Slippi EXI events, `Core/State.h` | +| [dolphin-emu/dolphin](https://github.com/dolphin-emu/dolphin/tree/ee018d00e60b9eb727489908a8daec5c537f44a8) `master` | `ee018d00e60b9eb727489908a8daec5c537f44a8` | `Source/Core/Core/Core.h`, state/core module inventory | +| [Felk/dolphin](https://github.com/Felk/dolphin/tree/46b7eacd5c810c2d21ec5fe51ea1a9c61a7ceb3d) historical `scripting` branch | `46b7eacd5c810c2d21ec5fe51ea1a9c61a7ceb3d` | Scripting README, `python-stubs/dolphin/{event,savestate}.pyi` | +| [altf4/libmelee](https://github.com/altf4/libmelee/tree/1da979657122facd0750ea99cf6858255e198326) `main` | `1da979657122facd0750ea99cf6858255e198326` | Archive notice directing users to maintained fork | + +Dolphin files inspected carry GPL-2.0-or-later headers; libmelee repository metadata reports +LGPL-3.0. Pin and retain actual dependency licenses/notices when packaging. A separate process +is an architectural boundary, not an assertion that distribution obligations disappear. + +Source inspection supports the integration hypotheses and concrete constraints above. It +does not establish Dolphin throughput on the deployment hardware, verify any game-memory +field live, demonstrate a new neural behavior, or prove exact multi-port/frame/save semantics. +Those are the measured deliverables of MELEE-01/02 and the performance ladder.