Compare commits

...
Sign in to create a new pull request.

1 commit

Author SHA1 Message Date
acamilo
d18bc043b3 docs(design): plan MaleCNS, modular sessions and Melee integration
Some checks failed
ci / node 22 (test + typecheck) (push) Has been cancelled
ci / rust stable (cargo test --workspace --release) (push) Has been cancelled
ci / infra/tests/lint.sh (push) Has been cancelled
ci / playwright apps/stage (allowed to fail) (push) Has been cancelled
(cherry picked from commit 83090a9391aac4044cfc48958fef54c0869a0e7b)
2026-09-21 15:29:12 +00:00
3 changed files with 1886 additions and 0 deletions

View file

@ -0,0 +1,247 @@
# Implementation backlog: MaleCNS and modular sessions
Status: **planned, not started**. Written 2026-09-18. Companion to the
[design and code analysis](malecns-modular-sessions.md), based on `f7bc13a`.
This is the execution order for that proposal. It does not authorize deployment or a live
stream. Reconcile the baseline with merged macro/shop/recovery work before implementation.
Existing feed/control contracts and TypeScript oracle rules remain binding.
Concrete second-game audit: [Melee framework and emulator plan](melee-framework-audit.md).
Its MELEE-01/02 spikes specialize EMULATOR-01 below and can proceed alongside framework
extraction; they do not depend on importing MaleCNS first.
## 1. Delivery strategy
Deliver working vertical slices; keep the existing FAFB/Game Boy composition usable throughout.
1. Establish a behavior baseline and explicit dataset/profile identities.
2. In parallel workstreams, characterize MaleCNS and extract the single-agent session runtime.
3. Demonstrate two isolated brains driving one ROM-free shared environment.
4. Expose that session through versioned feed/control contracts and a multi-agent broadcast.
5. Integrate a specifically chosen alternative emulator after its capability spike passes.
6. Finish physical package reorganization once the second consumer proves the boundaries.
**First milestone:** a reproducible headless MaleCNS run and a reusable single-agent session.
**Second milestone:** two flies in a synthetic arena with coherent resume and a local broadcast.
**Third milestone:** two flies in the chosen fighting game, with documented task scaffolding.
No elapsed-time estimate is committed before the import and emulator spikes establish their
unknowns. Split an item further when its contract and implementation cannot be reviewed together.
## 2. Work queue
Every item starts pending. Branch names are suggested implementation branches, not branches
already created. Builders use separate worktrees; the coordinator reviews contracts and results.
### FOUNDATION-01 — Pin existing behavior
- **Branch:** `test/session-baseline`
- **Depends on:** reconciliation with current main and related open work.
- Record effective legacy configuration, fingerprint, version strings and frame ordering.
- Add a ROM-free transition harness around the service's frame orchestration; capture brain
ticks, decoded actions, reward application and recovery effects with explicit clocks.
- Distinguish exact numerical replay from intentional macro/hold/transient reset on restore.
- **Done:** existing goldens pass; a trace fixture detects reordered vision/reward application;
legacy feed/API fixtures and compatibility identity are unchanged.
### FOUNDATION-02 — Specify bundle and behavior identities
- **Branch:** `feat/brain-profile-contract`
- **Depends on:** FOUNDATION-01.
- Specify dataset manifests, original-ID mapping, anatomical roles, sensory bindings,
readout bindings and composite behavior identity in a focused contract.
- Add profile resolution/validation around the existing core in TypeScript and Rust.
- Keep schema-1 fingerprinting and the legacy macro-role exception behind the legacy path.
- Add strict graph validation for new bundles, shared invalid fixtures and profile mismatch tests.
- **Done:** empty required populations, malformed CSR and incorrect profile restores fail;
legacy FAFB artifacts and default numerical version strings remain unchanged.
### DATA-01 — Acquire and normalize MaleCNS
- **Branch:** `feat/malecns-import`
- **Depends on:** FOUNDATION-02.
- Build a source-specific importer from official v1.0 tables with checksummed source locks.
- Reconcile the Codex versus neuPrint inventories, or explicitly select and document one.
- Preserve raw contact counts, transmitter evidence, original IDs and missing-data indicators.
- Emit deterministic graph bundles, weight≥1/weight≥5 comparison variants, and an exclusion/
clipping/coverage report. Decide schema-1 feasibility from measured weight ranges.
- Add attribution and license records with actual artifacts; keep raw downloads out of git.
- **Done:** repeated builds match; endpoint/role/index invariants pass; both loaders agree;
no service startup or ordinary unit test needs a download.
### DATA-02 — Characterize MaleCNS in the current kernel
- **Branch:** `feat/malecns-baseline-profile`
- **Depends on:** DATA-01.
- Audit L1 geometry, hemisphere handling, KC/MBON/PAM mappings and brain-versus-VNC motor roles.
- Define a versioned fixed readout; keep task action partitions out of anatomical truth.
- Run learning-off first, then learning-on, using fixed sensory traces and multiple seeds.
- Generate TS reference goldens and compare Rust exactly; measure activity, saturation,
initialization, memory and per-phase latency for both graph thresholds.
- **Done:** publish a reproducible characterization report and profile choice. Stop task
integration if required mappings are missing or dynamics are unusable; any recalibration
becomes a named profile rather than an edit to the legacy model.
### RUNTIME-01 — Separate environment execution from task interpretation
- **Branch:** `refactor/environment-task-boundary`
- **Depends on:** FOUNDATION-02.
- Specify controller ports, digital/analog controls, rational cadence, media descriptors,
observation ownership and backend capabilities.
- Wrap binjgb as the first environment; retain Game Boy FFI/cache/state behavior.
- Keep Pokémon memory inspection, objective routing, macros and reward rules in its task.
- Preserve existing imports through a facade; avoid simultaneous directory moves.
- **Done:** existing single-agent action/reward traces match and a fake environment can be
driven through the same boundary without importing binjgb or task-specific addresses.
### RUNTIME-02 — Extract the single-agent session
- **Branch:** `refactor/session-runtime`
- **Depends on:** RUNTIME-01.
- Move deterministic agent/environment/task orchestration out of `Sim` into a library.
- Keep HTTP, WebSocket serialization, wall-clock publication and process supervision in flysim.
- Give session clock, action executor, task ledger and recovery state explicit owners.
- Wrap the existing composition with legacy ordering, checkpoint and reset semantics.
- **Done:** headless consumer runs a session without Twitch/browser; the legacy composition
passes its traces and restore tests; a slow snapshot consumer cannot stall simulation.
### RUNTIME-03 — Add synchronized multi-agent sessions
- **Branch:** `feat/multi-agent-arena`
- **Depends on:** RUNTIME-02.
- Implement a ROM-free two-player arena and per-port controller ownership.
- Evaluate both brains against one observation boundary; apply one complete action batch;
advance the world once. Start sequentially, then verify parallel execution equivalence.
- Isolate RNG, stimulation, decoder holds, gains, traces and rewards per agent; share only
immutable topology. Enforce a total worker budget and single-dispatcher pool ownership.
- Define participant failure, lateness and episode reset policies.
- **Done:** no cross-agent state leakage; swapping evaluation order leaves results unchanged;
one failed participant cannot accidentally advance a half-controlled match.
### STATE-01 — Capture and resume whole sessions
- **Branch:** `feat/session-checkpoints`
- **Depends on:** RUNTIME-03; specify the state contract during RUNTIME-02.
- Define the new envelope/manifest and preserve the `FLYSIM01` reader.
- Capture all agents, environment, task/executor/admission state and clock remainders at one
boundary. Bound off-thread write jobs and retain atomic manifest commit semantics.
- Validate all components before installing any restored state; define external-backend staging.
- **Done:** uninterrupted and resumed synthetic matches agree; corrupting any participant
refuses the generation without partial restore; crash-injection fallback tests pass.
### WIRE-01 — Introduce session feed/control v2
- **Branch:** `feat/session-protocol-v2`
- **Depends on:** FOUNDATION-02, RUNTIME-02; use RUNTIME-03 fixtures for integration.
- Write binding contracts before consumer implementation: descriptors, scoped agents/events,
media IDs/timestamps, task progress, targeted stimulation and retry/idempotency behavior.
- Implement Rust/TS codecs, schemas and a fake server; preserve the legacy v1 surface.
- Specify descriptor reconnect behavior, asset/index identity, bounded message sizes and audio gaps.
- **Done:** cross-language fixtures pass for unequal neuron counts and shared/private views;
duplicate attachment kinds no longer collide; ambiguous targets and incompatible schemas fail.
### PRESENTATION-01 — Compose multi-agent stage and bridge
- **Branch:** `feat/multi-agent-broadcast`
- **Depends on:** WIRE-01, RUNTIME-03.
- Replace stage store/scaler singletons with session/agent instances and one paint scheduler.
- Resolve geometry from hashed descriptors, preserve the Game Boy presentation, and add a
shared-match layout with explicit audio ownership.
- Route bridge commands/redemptions to persistent session/agent identities; test lost responses,
retries and restart without applying an interaction twice or to a different agent.
- **Done:** local synthetic match broadcast works; two-agent PNGs receive operator review;
browser/fixture/legibility checks pass; bridge remains template-only and quiet-mode capable.
### DATA-03 — Run MaleCNS through the complete application
- **Branch:** `feat/malecns-session`
- **Depends on:** DATA-02 and descriptor-aware assets from WIRE-01/PRESENTATION-01.
- Expose explicit profile selection and create a fresh MaleCNS state namespace.
- Verify task/controller bindings, stimulation capability and displayed anatomy identity.
- A narrow single-agent descriptor extension may ship earlier only with matching v1 contract
and consumer updates; do not publish MaleCNS spikes as implicit FAFB indices.
- **Done:** local one-hour soak and restore drill pass; paired learning-off/on observations
are recorded without claiming improved play; FAFB remains available unchanged.
### EMULATOR-01 — Establish the alternative backend's capabilities
- **Branch:** `spike/fighting-game-backend`
- **Depends on:** RUNTIME-01; can proceed alongside later runtime work.
- Choose the exact game/version and emulator; Melee/Dolphin is a candidate, not a commitment.
- Prove pause/step, simultaneous ports, analog input, frame/audio capture, state inspection,
save/restore, process lifecycle and achievable cadence with synthetic controller traces.
- Prefer private IPC if embedding would leak emulator internals into the session library.
- **Done:** capability report includes pinned backend/content identity and reproducible results.
If bounded stepping or coherent restore fails, stop and revise the backend/requirements
before writing neural game logic. No game content enters repository fixtures.
### EMULATOR-02 — Build the two-fly fighting-game slice
- **Branch:** `feat/two-fly-fighting-game`
- **Depends on:** EMULATOR-01, STATE-01, PRESENTATION-01.
- Implement fixed controller mapping, match/round interpretation, positive attributed rewards,
observation policy and episode recovery. Display selected actions and actual controls.
- Validate one fly, two flies, round transitions, backend failure and resume in that order.
- Run side swaps and repeated seeds; compare learning-off and simple control baselines before
interpreting win rates. Separate show settings from controlled evaluation settings.
- **Done:** repeated local matches sustain declared cadence; restoration/failure policies work;
scaffold, interventions and limits are documented; reviewed match presentation is legible.
### PACKAGE-01 — Finalize reusable packages and release compositions
- **Branch:** `refactor/reusable-package-layout`
- **Depends on:** a useful second backend plus PRESENTATION-01.
- Extract proven crate/package boundaries from the design's module table; preserve facades.
- Move the Rust workspace only in a mechanical follow-up if it makes library consumption clearer.
- Update CI/build/vendor/golden paths, dataset/view asset packaging and compatibility preflight.
- Add minimal external-style Rust/TS consumers and a synthetic example composition.
- Reconcile current docs, stale template explanations and licensing/asset attribution.
- **Done:** both legacy and new compositions package successfully; incompatible state is
rejected before release selection; libraries run without importing broadcast services.
## 3. Dependency map and first execution batch
```text
FOUNDATION-01 → FOUNDATION-02 ┬→ DATA-01 → DATA-02 ───────────────→ DATA-03
└→ RUNTIME-01 → RUNTIME-02 → RUNTIME-03 → STATE-01
│ └→ WIRE-01 ───────┐
└→ EMULATOR-01 PRESENTATION-01
│
STATE-01 + EMULATOR-01 + PRESENTATION-01 → EMULATOR-02
second backend + presentation → PACKAGE-01
```
The item dependency lists are authoritative; the diagram is a reading aid.
When implementation begins, take **FOUNDATION-01 only** as the first build task. Then review
FOUNDATION-02's contract before assigning DATA-01 and RUNTIME-01 to independent worktrees.
Contract/schema authorship is serialized to avoid conflicting definitions. Deployment host
work remains serialized under the repository's claim protocol.
## 4. Definition of done for every implementation branch
- Scope and intentional behavior changes are stated; compatibility impact is explicit.
- Meaningful boundary tests cover the changed behavior; the TS oracle is not adjusted to
accommodate Rust output. Existing committed real-data goldens stay mandatory.
- `npm test`, `npm run typecheck`, `cargo test --workspace` (Rust workspace) and
`infra/tests/lint.sh` pass before merge. Visual changes also pass applicable Playwright
checks and PNG review. Optional full-MaleCNS/ROM runs record skips honestly.
- Performance-sensitive changes report representative activity, agent count, thread budget,
memory and tail latency. New experiments state what is modeled versus handwritten.
- Review the complete diff, merge with `--no-ff` when authorized, and update this queue with
commit, evidence and unresolved follow-ups. Rollback includes compatible state, not just code.
## 5. Decisions needed before the relevant work starts
| Decision | Deadline | Default recommendation |
| --- | --- | --- |
| MaleCNS inventory/filter policy | DATA-01 completion | Official versioned source; retain both threshold variants until measured |
| MaleCNS sensory/readout profile | DATA-02 | Audited L1 mapping with existing numerical model first |
| Exact fighting game and backend | EMULATOR-01 | Evaluate one concrete title/backend rather than supporting a console family at once |
| Number of flies and target resource budget | RUNTIME-03 performance gate | Two first; characterize four before promising it |
| Learning retention and sugar in matches | EMULATOR-02 task contract | Retention explicit; stimulation disabled in controlled comparisons |
| Package publication versus monorepo reuse | PACKAGE-01 | Monorepo libraries/examples first; public package publishing later |
There is no need to resolve these now to plan another feature. This backlog is ready for
resumption at FOUNDATION-01.

View file

@ -0,0 +1,892 @@
# MaleCNS and reusable streamed simulation sessions
Status: **proposal, not an implemented contract**. Written 2026-09-18 against `f7bc13a`
on `main`. This document covers two related projects: adding MaleCNS v1.0 as another
connectome, and extracting reusable modules for other emulators, embodied environments,
and multiple flies. No dataset, neural semantics, deployed configuration, or wire contract
is changed by this document.
The binding [feed](../feed-protocol.md) and [control](../control-api.md) contracts take
precedence. The TypeScript brain remains the oracle. Existing default versions
`lif-1ms-f64-v2` and `fly-kc-mbon-rstdp-v2` remain pinned.
Reading map: sections 2–3 contain the code audit and MaleCNS analysis; sections 4–6
define the proposed module/session boundaries; sections 7–8 give the extraction order,
implementation workstreams and acceptance gates; sections 9–10 record open questions and
sources.
Execution queue: [implementation backlog](malecns-modular-implementation.md), with branch-sized
deliverables, dependencies and completion criteria. Start at FOUNDATION-01 when work resumes.
Concrete follow-up: [Melee emulator and multi-fly framework audit](melee-framework-audit.md),
including source-checked Dolphin/libmelee integration options and full-stack performance gates.
## 1. Recommendation
1. **Add MaleCNS as a dataset/profile combination, not a replacement neural model.**
First run it through the existing LIF semantics with explicit, independently versioned
sensory, population, and readout mappings. Study different neuron dynamics separately.
2. **Make a session the unit of simulation ownership.** A session has one environment and
one or more independently stateful agents bound to its control ports. A shared match
advances once after all players have chosen actions from the same observation boundary.
3. **Extract along ownership and timing boundaries.** Separate anatomy, neural dynamics,
sensor encoding, action decoding, environment execution, task semantics, persistence,
observation transport, presentation, and audience interaction. Preserve current behavior
through a legacy composition while extracting these modules.
4. **Prove the design on a ROM-free two-player arena before a larger emulator.** Then build
a frame-stepped emulator integration. For “flies play Smash,” the first candidate should
be a specifically chosen title/backend, such as Melee with a pinned Dolphin integration;
“Smash” alone is not an emulator requirement.
5. **Retain a monorepo and one Rust workspace initially.** Reusable libraries do not require
a network of microservices, dynamic native plugins, or publishing unstable packages.
The two tracks can progress independently after the identity/profile boundary is established.
MaleCNS does not require multiplayer; multiplayer does not require MaleCNS. The first useful
deliverables are a reproducible MaleCNS characterization run and a behavior-preserving
single-agent session API, not a wholesale rewrite.
## 2. What the code actually does today
Paths below are relative to the repository root. Rust paths beginning `core/`, `gb/`, or
`sim/` in this document abbreviate `services/flysim/crates/flybrain-core/`,
`services/flysim/crates/flybrain-gb/`, and `services/flysim/crates/flysim/` respectively.
These aliases refer to **current** paths, not proposed directories.
| Boundary | Evidence inspected | Consequence |
| --- | --- | --- |
| Reference brain | `packages/brain/src/agent/agent.ts`; `core/src/agent.rs` | `NeuralAgent` composes network and decoder; image size and milliseconds/frame are configurable, but defaults are Game Boy-specific. It is already more reusable than the service. |
| Anatomy | `packages/brain/src/dataset/format.ts`; `core/src/dataset.rs` | CSR graph dimensions are dynamic. Weights are signed `i16`; roles and a two-dimensional visual-column table are part of the model input. Validation currently checks array lengths, not every graph invariant. |
| Dataset construction | `tools/build_flywire.py` | Five checksum-pinned Codex exports; stable indices from sorted root IDs; directed-pair aggregation; transmitter signs; clipping to ±32767; FAFB-specific role and L1-column extraction. Game macro populations are also generated here. |
| Neural dynamics | `core/src/lif.rs`; `packages/brain/src/model/lif.ts` | One-ms ticks, configurable gain/noise, role stimulation, image drive, and up to 64 tracked rate roles. It is not a general multimodal sensory API. |
| Learning | `core/src/plasticity.rs` | Strongest positive pre-role→post-role edges, default KC→MBON budget 16,384; gains/eligibility are per-agent. A caller supplies the reward scalar; PAM spikes do not generate it. |
| Parallelism | `core/src/lif.rs` (`SweepPlan`); `core/src/pool.rs` | Persistent deterministic within-brain pool. `broadcast` has one shared job slot and assumes one dispatcher at a time; cloning a plan is not permission to dispatch it concurrently. |
| Game interface | `gb/src/adapter.rs` | `GameAdapter` mixes reward detection, progress, decoder preset, recovery, ROM checks, and Pokémon-like tile/exit/objective queries. `MemoryReader::read8(u16)` is specifically a Game Boy-shaped interface. |
| Emulator | `gb/src/emulator.rs` | Concrete binjgb wrapper, 160×144 RGBA, eight-bit pad, one frame step, audio conversion, native save-state format. Explicit handles and `Send` allow multiple instances; `Sync` is deliberately absent. |
| Session ownership | `sim/src/simloop.rs` (`Sim`, `step_frame`) | One agent, emulator, adapter, ratchet, button mask, framebuffer, audio queue, sugar state, chat ring, and set of clocks. Application orchestration and game behavior share a large struct. |
| Actions | `sim/src/macros.rs`; `gb/src/macros.rs`; `gb/src/pokemon_red/macros/` | Neural channels select available macros, but the game-specific executor owns button sequences. The title screen uses raw input; scene handling, routing and targets are engineered behavior. |
| Recovery | `gb/src/recovery.rs`; `gb/src/ratchet.rs` | Game-only rewind retains brain clock, membrane, RNG, and gains; clears holds/eligibility and refreshes vision. A scalar progress ladder chooses a best save. This is not a general multiplayer reset policy. |
| Persistence | `sim/src/store.rs`; `gb/src/compatibility.rs` | Durable atomic envelope and manifest commit are reusable. Payload is one agent plus one emulator and ratchet. Compatibility names binjgb and `pokered`; native state identity includes size and target. |
| Public observation | `sim/src/snapshot.rs`; `packages/feed/src/{types,codec}.ts` | One flat brain/game snapshot; one attachment per kind; Game Boy buttons, fixed frame dimensions, Pokémon-shaped reward counters and a closed game-mode set. |
| Stage | `apps/stage/src/{App.tsx,feed/store.ts,feed/decode.ts}` | Good hot/paint/cold clock split, but mutable stores/scalers are singletons and the decoder checks 160×144 frames. Dataset URL is fixed to FAFB. |
| Game presentation | `apps/stage/src/games/` | A useful registry already exists, but config relabels v1 counters rather than declaring independent task schemas. |
| Twitch bridge | `services/bridge/src/{index,sim,commands,redemptions,templates}.ts` | Transport/client abstraction, test fakes, templates, rate limits, and redemption persistence are valuable. There is one sim URL and no agent target identity. |
| Packaging | `apps/stage/vite.config.ts`; `infra/build/package-release.sh`; `infra/05-deploy.sh` | Stage build copies FAFB artifacts; packaging defaults to FAFB; deploy preflights one compatibility string. Runtime/data/frontend are still one release composition. |
### 2.1 Behaviors to preserve before extracting
`Sim::step_frame` advances the brain from the previous visual input, decodes, applies
buttons, steps the emulator, sets the new visual input, samples rewards, stimulates per
reward event, reinforces their sum, updates macro availability, and observes the ratchet.
Control commands are drained before the frame step. This ordering is part of behavior.
Do not replace this with `NeuralAgent::tick` merely because it looks like a convenient
wrapper: the service currently orchestrates substeps to sample reward from the frame just
produced. Moving reward or visual drive across that boundary changes trajectories.
Other invariants:
- Deterministic arithmetic and per-target propagation order, including across thread counts.
- Browser/network clients never stall the sim; snapshots are latest-value/drop-oldest.
- Checkpoint capture is coherent; encoding and storage happen off the sim thread.
- Failed restore does not silently reset a run. Existing legacy restore policies remain exact.
- No public control endpoint for button presses or game-memory writes.
- Sugar and learning reward are distinct mechanisms. Chat text never becomes neural input.
- The current positive-only reward doctrine remains the default for all shipped tasks.
### 2.2 Existing compatibility gaps to handle deliberately
The dataset fingerprint hashes metadata (including anatomical circuit roles) and six arrays.
`macro_*` roles are merged **after** hashing to preserve old checkpoints. Both language
implementations document that changing those populations could restore rates onto different
neurons without invalidating the checkpoint. The kernel parameter version also omits role
names; the service compatibility string does not fully identify the decoder/action mapping.
Keep these historical behaviors in the legacy reader. For new profiles, add an explicit
behavior identity covering population bindings, sensor encoding, decoder configuration,
action executor, and reward catalog. Do not repair the old hash by changing it in place.
Open macro/shop and recovery branches existed when this proposal was written. Before
implementation, rebase the inventory against their merged state and rerun characterization;
this proposal neither incorporates nor supersedes their unmerged changes.
## 3. MaleCNS: anatomy, connectivity, and model are different things
### 3.1 Dataset comparison and evidence limits
| Property | FAFB v783 used here | MaleCNS v1.0 |
| --- | --- | --- |
| Specimen | Adult female | Adult male, independently imaged/reconstructed |
| Territory | Brain including optic lobes | Central brain, optic lobes, ventral nerve cord (VNC), intact neck connective |
| Local artifact | 139,255 neurons; 2,700,513 directed-pair edges | None imported in this repository |
| Available inventory counts | Codex lists 139,255 neurons | Codex lists 166,700; the Minecraft project's neuPrint `:Neuron` export reports 176,422. Selection rules must be reconciled before fixing a local count. |
| Input labels | Codex `root_id`, classification, consolidated types, column assignment | neuPrint/flat export `bodyId`/body IDs, class hierarchy, transmitter properties, sides, neuropils, cross-dataset type annotations |
| Added anatomical opportunity | Brain sensory→descending circuits | Brain↔VNC circuits, local motor circuitry, ascending feedback, additional sensory and motor populations |
| License evidence | Repository attribution: CC BY-NC 4.0 | Official MaleCNS download site: CC-BY; verify and retain the exact release license text when importing |
MaleCNS is not FlyWire with extra neurons appended. IDs and dense array indices do not
correspond. Homologous cell types and registered anatomical spaces enable comparisons, not
automatic one-to-one neuron matching, state transfer, or graph concatenation. Male-specific
and sexually dimorphic circuits make a universal matching assumption especially misleading.
The official MaleCNS site describes a finished, proofread and annotated CNS reconstruction.
This does not mean every synapse, cell type, or sensory column is equally certain. The
Minecraft derivative reports asymmetric visual-column coverage and missing soma positions;
our import must quantify coverage from its own pinned source. A connectome also does not
supply all synaptic physiology, electrical coupling, neuromodulation, body dynamics, or
behavioral competence.
### 3.2 What “connections” means
Separate these quantities in metadata, reports, and on-screen claims:
1. Source neuron/segment inventory and the chosen included-neuron inventory.
2. Individual synaptic contacts/partner pairs (and separately pre-sites and post-sites).
3. Directed neuron-pair edges after aggregation.
4. Retained edges and retained synaptic weight after confidence/weight filtering.
5. Effective signed weights after the simulation's transmitter and clipping policy.
The FAFB builder aggregates export rows by `(pre, post)` across rows, drops endpoints not
in its classification inventory, assigns a sign, and clips the summed magnitude. It adds
no explicit five-synapse threshold of its own. The source exports may already be filtered;
the builder cannot recover contacts absent upstream. Codex headline connection counts are
not necessarily the local aggregated graph's edge count.
The Minecraft project's provenance reports ~25.9 million MaleCNS neuron-pair edges at
weight ≥1 and a bundled derivative of 6,287,749 edges at weight ≥5, representing
90,296,905 of 125,024,863 neuron-to-neuron synapses. It removes 40 autapses. Those are
**that project's reported query/build results**, not counts independently reproduced here.
Its five-contact cutoff retains roughly 24% of edges but 72% of synaptic weight. It is a
performance/modeling choice, not the definition of a complete CNS.
The official bulk weight table includes **segments**, not only curated neurons. Loading
all rows as if they were all validated neurons would be a different experiment. Likewise,
neuPrint ROI-level adjacency rows must not be summed together with their already-aggregated
totals. Specify one authoritative edge representation and count every exclusion.
### 3.3 Acquisition and reproducible construction
Prefer official versioned bulk tables for the repeatable build, with a neuPrint query tool
for inspection and cross-checking. The official download page lists:
- `body-annotations-male-cns-v1.0-minconf-0.5.feather` — curated annotations.
- `body-neurotransmitters-male-cns-v1.0.feather` — neuron-level transmitter information.
- `body-stats-male-cns-v1.0-minconf-0.5.feather` — segment statistics; large and broader than
the curated neuron list.
- `connectome-weights-male-cns-v1.0-minconf-0.5.feather` — full segment connection graph,
approximately 1.1 GB as listed by upstream.
The much larger synaptic-point and partner tables are unnecessary for an initial point-neuron
simulation. Download them only for a question requiring synapse-level geometry. Neither
research downloads nor anatomy conversion belong in service startup or ordinary unit tests.
Proposed build stages:
```text
release manifest + checksummed source cache
→ source-specific parser
→ normalized neuron/edge tables + exclusion report
→ selected anatomical graph
→ model-specific signed-weight transform
→ runtime CSR bundle + profile bindings + separate viewer bundle
```
Each stage records source release, database revision if queried, query/filter definitions,
source byte hashes, tool revision, and output hashes. Fetch time is provenance, not a random
input to the semantic graph hash. Sort body IDs numerically; encode original IDs as strings
in JSON so this common format also preserves FAFB IDs beyond JavaScript's safe integer range.
Preserve original annotations and cross-dataset type aliases separately from normalized roles.
The first import report must reconcile the Codex/neuPrint count difference, or explicitly
choose and document one inventory without claiming equivalence. It must also report missing
IDs, unannotated neurons, empty required populations, missing geometry, unknown sides,
transmitter confidence/fallback counts, duplicate edges, autapses, clipped weights, and
retained contacts by region and threshold.
Use deterministic serialization and gzip headers, following our existing reproducible
builder rather than copying the Minecraft artifact's timestamp-dependent container format.
Keep the source cache outside tracked artifacts; pin published runtime bundles by digest.
When adding actual data, update `NOTICE`, `LICENSES.md` and bundle-local attribution/license
files in the same change. Keep FAFB-derived assets under their existing terms; a separately
licensed MaleCNS bundle does not relicense mixed fixtures, old goldens or viewer assets.
### 3.4 Graph and sign policy
Keep raw positive contact counts and transmitter evidence in the normalized data. Sign is a
model transform, not a measured property that should overwrite the source evidence.
The legacy policy assigns GABA/GLUT negative and other/unknown transmitters positive;
conflicting per-edge transmitter rows become `MIXED`, then positive. Do not silently extend
that policy to histaminergic photoreceptor input. A MaleCNS policy must explicitly define
histamine, monoamines, mixed/unknown labels, confidence fallbacks, and whether transmitter
is chosen per neuron or per connection. None implies receptor-specific physiology.
Recommended first profiles:
- **`malecns-v1-lif-baseline`**: existing kernel, explicitly versioned sign policy and L1
input mapping, fixed readout, plasticity initially disabled for characterization.
- **`malecns-v1-lif-learning`**: same anatomy/input/readout plus audited KC→MBON selection
and the existing reward rule. Learning is an experimental condition, not an assumed gain.
- Later **sensorimotor research profiles**: photoreceptors, mechanosensation, VNC outputs,
and possibly another neural model, each with its own identity and validation.
Build weight≥1 and weight≥5 variants as **different graph identities** for the benchmark.
Do not select a production cutoff until activity, retained connectivity, memory and speed
have been measured. Preserve autapses by default in the new canonical graph; if an experiment
removes them, record the policy. The claim that point-neuron models have no use for autapses
is not a reason to discard observed connectivity silently.
Schema-1 `i16` weights may be sufficient, but measure overflow rather than assume it. For
the first existing-kernel comparison, emit a schema-1-compatible runtime view only if its
quantization/clipping is explicitly reported. If wider weights are needed, implement an
additive format/loader and matching oracle path; do not reinterpret old `weights.binz`.
Strengthen validation before constructing a network: CSR starts at zero, is monotone, ends
at edge count; targets/roles/visual indices are in range; required populations are present;
geometry is finite where marked valid; array lengths and index widths are representable;
declared hashes and artifact sizes match. Test the same invalid fixtures in both languages.
### 3.5 Population and sensory mapping
Existing role predicates cannot be copied unchanged. The current builder recognizes Codex
names such as `Kenyon_Cell`, `brain_motor_neuron`, `DAN` and `PAM*`; MaleCNS uses another
annotation vocabulary. Maintain a reviewed mapping table with source predicates, resulting
counts, hemisphere policy, aliases, and citations. Resolve each profile's required roles
at startup; an absent role must be an unsupported capability, not a silent empty population.
In particular:
- **Brain motor ≠ all motor.** Current macro pools combine 96 MBONs and 110 brain motor
neurons. Adding hundreds of VNC motor neurons under `motor` would silently change that
behavior. Use qualified roles such as `brain.motor`, `vnc.motor`, `brain.descending`,
`mb.kenyon`, `mb.output`, and `mb.pam` in new profiles, with legacy aliases only as needed.
- **Action groups are not anatomical facts.** `command_*` buckets use index modulo eight;
`macro_*` groups use a round-robin MBON/brain-motor pool. Move their construction into a
versioned task readout profile, outside the anatomy builder. Do not describe those groups
as natural “attack,” “jump,” or game-objective circuits.
- **Rate budgets are finite.** Both kernels currently cap tracked populations at 64. A
full CNS has many more interesting populations. Select a bounded control/telemetry set
initially; arbitrary bulk population analysis belongs in offline tooling. A larger mask
is a measured, oracle-tested change, not an unbounded string map in the tick loop.
- **L1 first, photoreceptors later.** Reuse the existing luminance projection only after
auditing MaleCNS L1 hex coordinates, both sides, and a declared hex→2D transform. Soma
coordinates are not visual-field coordinates. Missing columns remain explicitly missing;
do not synthesize them from array order or infer them from another specimen's neuron IDs.
- **Input profile and display view differ.** A game image may be resized/cropped for the
agent while the full frame is shown to viewers. Record crop, orientation, color transform,
and sampling geometry. Neither player may accidentally receive another player's private
view or adapter-only task observations.
The existing `set_visual_frame` and `stimulate` API cannot represent a general collection of
odor/touch/proprioceptive inputs. A later sensory-drive interface needs explicit units,
target populations, additive/overriding rules, tick ordering, and deterministic noise streams.
It must be specified in TypeScript before a matching Rust implementation. Keep the old
image/stimulation path available byte-for-byte through its compatibility facade.
### 3.6 What MaleCNS lets us investigate
| Experiment | New capability | What still needs engineering/measurement |
| --- | --- | --- |
| Same game, another connectome | Compare datasets under a matched task interface | Cell-type mapping, gain/activity calibration, readout comparability, multiple seeds |
| Brain↔VNC control | Read actual descending, ascending and motor populations | Body/control mapping; gamepad commands are not muscles |
| Embodied fly arena | World smell/taste/touch/vision mapped into annotated sensory populations | Sensor transduction, proprioception and body dynamics; identify every reflex shortcut |
| Mixed-dataset two-fly match | FAFB and MaleCNS agents share one environment | Balanced observations, controller mapping, compute budgets, intervention rules |
| Circuit perturbation | Compare full graph with VNC feedback or defined pathways ablated | Separate graph identities; activity and behavioral controls; no biological claims from gameplay alone |
The Minecraft project is a useful engineering comparison, not our validation oracle. It uses
a Shiu-style current-based LIF model with synaptic dynamics/delay, reports gain calibration,
and explicitly supplies odor-approach reflexes and higher-level looming drive where its
simulated pathways do not work. Its source graph, numerical model and embodiment differ
from ours simultaneously. Do not attribute its behavior solely to MaleCNS.
### 3.7 Identity, restore, and performance
Name the components independently:
```text
anatomyId = source release + included inventory + graph/filter digest
modelId = numerical semantics + effective numeric configuration
sensorId = input encoding + anatomical binding digest
readoutId = population partition + decoder + action mapping digest
learningId = rule + selected-edge topology + reward-catalog identity
viewerId = positions/geometry + index mapping digest (presentation only)
```
The composite behavioral identity covers all behavior-affecting components; original source
IDs and display geometry cannot replace it. Same neuron count does not establish compatibility.
No FAFB neural checkpoint is restored into MaleCNS. A deliberately fresh MaleCNS brain may
start from a compatible game-only save, with a new run identity, clean calibration and reward
baselining; that is a new experiment, not continuation of the old fly. Gains are not mapped
between specimens by cell-type name.
For rough capacity planning, the current CSR is `4(N+1) + 6E` bytes, excluding roles,
geometry, derived propagation structures and mutable state. At 176,422 neurons that is
about 38.4 MB for 6.29 M edges, or 156.1 MB for 25.9 M edges (decimal MB). This is roughly
2.3× or 9.6× our edge count, not a prediction of the same slowdown. Spike activity, fan-out,
plasticity selection, memory bandwidth and sharding determine runtime cost. Checkpoint
copies and renderer assets also need separate memory budgets.
`LifNetwork` accepts shared `Arc<BrainDataset>` already. Reuse immutable anatomy across
same-profile agents; keep membrane, refractory state, RNG, rates, decoder holds, stimulation,
eligibility and learned gains private. Audit constructor-derived caches before moving them
into a shared topology object. Never share mutable gains merely because graphs match.
Measure headless one-, two-, and four-agent runs with plasticity on/off, fixed inputs and
representative activity. Record resident/peak memory, initialization, state capture cost,
per-phase p50/p95/p99, spike distribution, and real-time factor. Existing CUDA code is an
optional backend requiring its own new-dataset equivalence/capacity gate, not assumed capacity.
## 4. Reusable architecture
### 4.1 Define the nouns first
- **Dataset bundle:** immutable anatomical graph and source annotations.
- **Brain profile:** dataset plus numerical model, sensory/readout bindings and learning rule.
- **Agent:** one independently stateful brain, encoder, decoder and action executor.
- **Environment:** the world being advanced: one emulator instance, a linked-emulator group,
or an embodied simulator. Owns controller ports, world state, media, and native clock.
- **Task:** interpretation of environment state: rewards, progress, episode endings, allowed
macro actions and recovery policy. Pokémon is a task, not an environment API.
- **Session:** one environment plus agents, port assignments, scheduler, task state and clocks.
- **Broadcast:** presentation of one or several sessions, plus chat and audience interactions.
An agent ID is not a Twitch username, controller port, array position or dataset ID. Session,
episode, agent, port, view, and event identities must be explicit and stable across restore.
### 4.2 Dependency direction
```text
source importers → dataset bundles
↓
neural core (TS oracle / Rust runtime)
↓
agent composition: sensors + readout + executor
↓
environment backend + task plugin → session runtime → observations/checkpoints
↑ ↓
command admission protocol adapters
↑ ↓
Twitch bridge stage / recorder
```
The neural core knows no emulator, task, network socket, chat, or UI. The environment knows
no neural populations or Twitch. The task may inspect backend-specific state through a
typed inspector, but neither inspection nor public presentation gives clients a memory-write
or controller-write API. Only the session commits agent-produced controls.
### 4.3 Proposed modules and staged layout
These are target responsibilities, not instructions to create every package immediately.
Start as modules; extract crates/packages once a second consumer demonstrates the boundary.
Keep the Rust workspace under `services/flysim` during semantic extraction so paths and
behavior do not change together. A later mechanical move can place reusable crates at the
root, updating CI/build/golden paths in one dedicated change.
| Module / eventual location | Owns | Extraction source |
| --- | --- | --- |
| `packages/brain` | Reference numerical behavior and legacy public facade | Existing package; keep imports compatible |
| `crates/flybrain-core` | Rust numerical kernel, plasticity, generic population decoder | Existing `core/`; leave compatibility re-exports for presets |
| `crates/fly-dataset` | Manifest validation, artifact loading, source-ID/index mapping | `core/src/dataset.rs`; retain legacy fingerprint implementation |
| `tools/datasets/{fafb,malecns}` | Source-specific conversion to common bundles | Existing Python builder plus new importer; existing CLI wrapper remains |
| `crates/fly-session` | Agent ownership, clock coordination, action commit, event/reward routing | Orchestration extracted from `sim/src/simloop.rs` |
| `crates/fly-environment` | Backend capabilities, ports, observations, media, save-state interfaces | New small contract proven with binjgb and synthetic arena |
| `crates/fly-env-gb` | binjgb FFI, memory inspector and native save-state identity | `gb/src/{emulator,ffi}.rs`, build glue and vendor boundary |
| `crates/fly-task-pokemon`, `fly-task-platformer` | Audited reward rules, semantic state, macros, progress/recovery | `gb/src/{pokemon_red,platformer}/`; do not generalize tile routing into the core |
| `crates/fly-checkpoint` | Atomic storage and session envelope; legacy payload adapter | `sim/src/store.rs` plus core envelope helpers |
| `crates/fly-protocol` / `packages/feed` | Versioned wire schemas/codecs, legacy adapters, synthetic fixtures | `sim/src/snapshot.rs`, existing feed package; canonical schema with cross-language tests |
| `services/flysim` | Composition/config, HTTP/WS, process lifecycle, metrics | Thin host over reusable session library |
| `packages/stage-runtime` | Feed ingestion, per-session stores, paint loop, audio, fixture clock | Extract from `apps/stage/src/{feed,paint,audio,motion}` after multi-view prototype |
| `apps/stage` + presentation plugins | Layout, branding, task panels, audience-facing explanations | Existing page with legacy layout preserved |
| `services/bridge` + audience client module | Twitch transport/auth/redemptions; session-targeted interaction client | Existing bridge; extract provider-independent logic only when reused |
| `infra/` | Release composition, process supervision, capture, recordings | Existing tooling parameterized by session/broadcast manifest |
Use static Rust composition or a small closed registry initially, with trait boundaries at
backend/task seams. Do not require stable native dynamic-plugin ABI. An out-of-process
emulator can implement the backend through private IPC; it is not a new public action API.
Keep backend-specific memory access private to its task implementation instead of widening
`read8(u16)` into a supposedly universal game-state abstraction.
### 4.4 Environment and agent contracts
Illustrative interfaces; concrete types must be written with tests during contract work:
```rust
trait Environment {
fn descriptor(&self) -> &EnvironmentDescriptor;
fn observe(&mut self) -> Result<WorldObservation>;
fn advance(&mut self, actions: &ActionBatch) -> Result<WorldStep>;
fn capture(&mut self) -> Result<EnvironmentCheckpoint>;
fn restore(&mut self, state: &EnvironmentCheckpoint) -> Result<()>;
}
// One decision boundary, one action per configured controller port.
struct ActionBatch {
session_tick: u64,
ports: Vec<PortAction>,
}
```
`descriptor` declares rational step duration, controller schemas, views, audio streams,
save/restore availability, task inspection capabilities and determinism level. Capture and
restore return an explicit unsupported error when unavailable; configuration validates that
the chosen recovery policy can work. `WorldObservation` is a frame-boundary snapshot or
immutable handle, not an object allowing agents to advance the backend.
Control schemas support digital buttons and bounded analog axes/triggers, with neutral
values, axis ranges, dead zones, and mutually exclusive directions where appropriate.
Preserve the Game Boy mask as one concrete codec. Analog controls need a fixed, versioned
decoder mapping; an 800-ms direction hold is not a sensible default for every fighting game.
Separate three observation surfaces:
1. **Agent sensory view:** pixels or declared synthetic senses the profile may consume.
2. **Task inspector:** audited state for reward/macro/episode logic. Access is part of the
disclosed scaffold, not implicitly available to the neural encoder.
3. **Broadcast view:** media and summaries for viewers, potentially richer than either player's
allowed sensory input.
Readout produces semantic channel activations or continuous signals. A task-local action
executor translates these into port actions, optionally running a selected macro. It is
explicitly resettable/checkpointable and reports selected action versus actual controller
output. Pokémon pathfinding, dialog logic, and objective catalogs stay in its task plugin.
Task outputs become scoped `RewardEvent`, `Progress`, `EpisodeEvent` and `ActionAvailability`.
Progress is a tagged value (`ladder`, `score`, `match`, `exploration`, or task extension), not
always a scalar rank. Rewards carry recipient agent/team, rule ID, event ID and observation
tick. A task cannot mutate neural state directly; the session routes accepted rewards once.
An agent-facing API similarly separates `advance_brain(interval)`, `decide(observation,
availability)`, `encode_next(sensory_view)`, and `apply_outcome(rewards, stimulation)`.
The session owns their order. Agents cannot call `Environment::advance`, select another
port, or inspect another agent's mutable state. Start with a concrete LIF agent composition;
introduce a controller trait when synthetic controllers or a second neural model require it.
Test controllers implement the same decision surface but are identified as non-neural agents
in descriptors and experiment records.
### 4.5 Example composition
Illustrative configuration, not syntax supported by today's `flysim.toml`:
```toml
[session]
id = "arena-demo"
environment = "synthetic-arena-v1"
task = "two-player-rounds-v1"
scheduler = "lockstep-v1"
master_seed = 1234
recovery = "round-reset-keep-gains-v1"
[[agents]]
id = "fly-a"
port = "player-1"
profile = "fafb-arena-baseline-v1"
sensory_view = "shared-camera"
[[agents]]
id = "fly-b"
port = "player-2"
profile = "malecns-arena-baseline-v1"
sensory_view = "shared-camera"
[broadcast]
layout = "shared-match-two-agents"
audio = "world"
audience_stimulation = false
```
Resolve profile IDs through a local, digest-pinned registry. Both agents may instead select
the same profile and share immutable topology while retaining independent state. The host
validates unique agent IDs and exclusive port ownership, view accessibility, profile/controller
compatibility, recovery capability and resource budget before starting. An independent-games
broadcast composes two such sessions; it does not misrepresent them as ports in one world.
## 5. Multiple flies: concurrency is not multiplayer
### 5.1 Three supported arrangements
| Arrangement | Ownership / synchronization | First use |
| --- | --- | --- |
| Independent flies in independent games | One session/process each; optional broadcast composition | Parallel streams and experiments; existing deployment pattern generalizes easily |
| Several flies in one game | One environment, several ports, one session barrier | Local fighting/multiplayer games |
| Linked emulator instances | One composite environment owns all instances and link state | A later link-cable experiment; requires cycle-accurate link support, not two independent frame loops |
For a shared arena there must not be one `Sim` loop per player, each calling `run_frame`.
That advances the world multiple times and gives an ordering advantage to one agent.
### 5.2 Shared-world step semantics
For decision boundary `t`, freeze observation `O[t]`, then:
1. Admit queued audience/operator commands against stable session/agent identities; log their
effective tick. Chat remains presentation state.
2. Each agent advances its brain for the same environment interval, using its previously
encoded sensory input. Its clock remainder and RNG are private.
3. Each agent decodes and advances its action executor using the same `O[t]` task boundary.
4. Barrier: collect all port actions, validate ownership/ranges, and commit one complete batch.
5. Advance the environment **once** to obtain `O[t+1]` and timestamped media.
6. Encode the next sensory inputs, evaluate task events from the completed transition, apply
explicitly routed stimulation and rewards, and compute next action availability.
7. Apply any whole-session episode/recovery transition; capture coherent state and publish.
Preserve the detailed legacy ordering inside the legacy single-agent composition. New
profiles identify their scheduling semantics explicitly rather than silently adopting a
different reward phase. Use rational environment time and integer substep accumulation for
new sessions; keep the legacy floating remainder arithmetic for old trajectories. The neural
clock may lead environment time by warm-up; persist that offset instead of pretending all
clocks start at zero. Rendering, physics and decision cadence may differ, but the backend
must define their relationship.
Start with sequential agent evaluation for reproducibility. Then compare parallel agent
evaluation against the same action trace. Cap total worker budget: `agents × brain_threads`
can otherwise oversubscribe the machine. Use private pools for concurrent agents or serialize
dispatch into a pool; the current `WorkerPool` must not be concurrently reused through a
cloned `SweepPlan`. Sharing immutable graph buffers is independent of scheduling workers.
If one agent is late, the default is to slow the **whole session** and report lag. A crashed
agent pauses/fails the match rather than silently becoming a neutral or scripted opponent.
A realtime external world that cannot pause requires a separate declared deadline/hold-last
policy, dropped-action telemetry and a different determinism claim. Do not hide that policy
inside the environment adapter or use wall-clock completion order as an action tie-breaker.
### 5.3 Match state and learning
Every agent has its own seed, calibration, gain vector, eligibility, reward totals and
stimulation cooldown. Derive seeds deterministically from a stored master seed and stable
agent ID; do not use thread scheduling or the default identical seed for every fly.
For an initial fighting-game task:
- Both agents receive the same shared camera unless the game has genuine private views.
- Controller-port swaps and seed repeats are part of evaluation; wins alone are confounded
by character, spawn, arena, action interface and side advantage.
- Award positive, explicitly attributed events such as a scored hit/round win; define
damage/self-damage/team attribution and duplicate detection before turning learning on.
A loss need not produce a negative reward; changing reward doctrine is a separate decision.
- Episode transitions may retain learned gains while clearing transient traces, or create
fresh brains for controlled trials. Record which policy was selected.
- Do not select a “best checkpoint” separately for each player in a shared world. There is
one world state. Tournament scores and historical results should not rewind with a match.
- Disable sugar for balanced evaluation. If enabled for a show, target a named agent under
a documented rule and log the intervention; do not call that an uncontrolled fair benchmark.
### 5.4 Checkpoint and recovery semantics
Distinguish three operations:
| Operation | Restored/reset state |
| --- | --- |
| Crash resume | Coherent environment, every agent, scheduler remainders, task ledgers, pending actions/commands and executor state at one boundary |
| Task recovery | Policy-defined world rewind/reset and per-agent transient clearing; continued gains/brain clocks only when declared, as in the legacy ratchet |
| New episode | Task initial world state, explicit retained/fresh agent policy, new episode identity; session event sequence remains monotonic |
Use a new session envelope version with a manifest mapping stable agent IDs to chunks and
listing graph/model/profile/backend/content/task/state-format identities. The current
envelope restricts chunk names to letters; do not simply append `agent/1/membrane` to it.
Specify a new container format or an explicit manifest-to-valid-chunk-name indirection.
Validate every participant into staged state before mutating any live participant. A failed
environment import must not leave half the brains restored. For an external emulator,
restore a replacement stopped process when transactional in-place validation is impossible.
Capture at the action barrier with no backend step in flight. Reuse atomic payload write,
manifest commit, hot/durable tiers and off-thread serialization; bound outstanding snapshot
jobs so repeated copies cannot exhaust memory under slow storage.
Exact replay requires action-executor and admission state, not just the neural envelope.
Legacy macros intentionally discard transient execution on restart; preserve that behavior
for v1 and label it as legacy continuation semantics, not exact session replay. New sessions
persist all behavior-affecting state or explicitly restart an episode under a documented rule.
Keep `FLYSIM01` readable through a legacy adapter. Never silently rewrite a checkpoint on
read. Conversion is an explicit offline operation writing a new directory and run record.
Environment identity includes title/content digest, backend build/configuration, relevant
platform/native state format, and patch/symbol provenance; state size alone is not sufficient.
## 6. Feed, stage and audience modules
### 6.1 Feed v2 is required
V1 is not a generic multi-agent format. Its attachment map forbids duplicate kinds, so two
`spikes` arrays cannot coexist; its button mask and 160×144 image are Game Boy-specific.
Platformer counters are currently folded onto names such as `pokedex` and `wildwin`.
Extending those conventions to fighting games would preserve syntax while losing meaning.
Keep v1 stable for the legacy session. Introduce a negotiated v2 or a separate `/v2/feed`
endpoint, with a descriptor delivered before dependent snapshots and available on reconnect.
The contract change must update Rust, TypeScript, schema, fake service, fixture player,
stage and bridge together. Proposed shape:
```text
SessionDescriptor
sessionId, protocol, descriptorRevision, environment/task identities, clock definition
agents[{agentId, profileId, datasetId, neuronCount, roles, controlPort}]
ports[{portId, controllerSchema}]
views[{viewId, dimensions, pixelFormat, sensory/broadcast use}]
audioStreams[{streamId, sampleRate, channels}]
assets[{id, contentHash, datasetIndexHash, localUrl, license/credit}]
SessionSnapshot
descriptorRevision, seq, sessionTick, episodeId, environmentTime, wallTime, status
agents[{agentId, brainTime, rates, learning, selectedAction, actualControls, stimulation}]
progress: tagged task payload; events: scoped and sequenced
attachments[{id, kind, ownerId, byteLength, format, mediaTimestamp}]
```
Use bounded, schema-validated tagged payloads/namespaced task extensions, not arbitrary
unlimited JSON or executable server-supplied UI. Dataset identity and neuron index mapping
must accompany spike geometry: matching bitset length alone cannot establish alignment.
Include descriptor revision in every snapshot and reject stale/mismatched buffers. Large
monotonic IDs/times use a specified safe-integer bound or decimal strings across languages.
Publish one shared camera/audio stream once, not once per agent. Independent sessions may
have independent streams. Timestamp audio/video to the session clock; define discontinuities
on reset, reconnect and lag. Latest-value video/telemetry may drop, while audio needs a
bounded timestamped buffer and explicit gap handling. Durable event IDs permit recovering
missed events; a drop-oldest snapshot feed is not an exactly-once event log.
Bandwidth becomes an architectural concern: 640×480 RGBA at 30 Hz is ~36.9 MB/s before
framing; 1920×1080 is ~248.8 MB/s. Do not copy full-resolution raw frames per fly across
several sockets. Start with one bounded broadcast-resolution local stream and separate
agent input resolution; choose a compressed media transport only after measuring latency,
CPU/GPU cost and synchronization. A telemetry protocol need not become a video codec.
### 6.2 Stage composition
Extract `createSessionStore()` rather than adding `agent2` fields to the singleton `hot`.
Own rate scalers, button afterglow, ticker state, fixture clock and audio queues per session/
agent. Keep one page paint scheduler; register surfaces against explicit view/agent IDs.
Geometry is loaded by descriptor/hash through an asset manifest, replacing the hardcoded
FAFB route in both `App.tsx` and `vite.config.ts`.
Retain the existing Game Boy layout as a presentation plugin. Add composition primitives for
a shared match view with two agent summaries, or independent session tiles with one focused
audio source. A generic fallback shows status, media, controls and task labels without
inventing Pokémon counters. Unknown optional extensions can be omitted; unknown required
capabilities or mismatched descriptors must be visible rather than silently showing Pokémon.
Keep UI runtime dependencies out of the numerical package. The current `@flybrain/brain`
view exports and optional Three.js peer can remain compatibility re-exports when viewer
geometry/helpers move to a dedicated view module. Controller labels belong in controller
schemas, not a UI import of the neural package's Game Boy preset.
Review actual layout proposals as PNGs under `apps/stage/mockups/`, with existing legibility,
phone-scale and browser gates. This document proposes data/layout boundaries, not screen
copy or a replacement for visual sign-off.
### 6.3 Audience interaction
Keep Twitch authentication/EventSub and template-only replies in the bridge. Extract a
session-targeted interaction client with explicit `sessionId`, optional `agentId`, interaction
kind and idempotency key. A multi-agent request with no target is rejected unless a fixed,
declared target policy exists; presentation focus must never choose the recipient.
Service admission owns per-session, per-agent and global limits. A profile advertises its
supported stimulation capability; an agent without PAM support returns unsupported rather
than pretending to accept “sugar.” `!stuck` becomes task-aware (ladder time versus round
time), while chat remains broadcast-scoped and independent of neural state.
For redemptions, persist the resolved target and request identity before retrying. Distinguish
an HTTP timeout from a definite refusal: a timeout may occur after the sim accepted the
effect. V2 needs a durable or explicitly recoverable deduplication/status contract so bridge
restart cannot apply the same pulse twice or retarget a redemption to a new match. Do not
promise exactly-once behavior from the bridge intent log alone. Existing v1 remains as-is.
### 6.4 Deployment and observability
A deployable composition selects runtime binary, backend/task, agent profiles, dataset
bundles, view assets, session state namespace and broadcast layout. Generate service config
from that manifest plus the operator's external environment. Keep tokens and network-specific
values outside this repository. Package viewer artifacts separately from the full simulation
graph so the browser does not need every edge.
Preserve existing process isolation: sim, bridge, browser, capture, local relay and Twitch
push can restart independently. One process/session remains the default for fault isolation;
a multi-agent match stays one coordinated session, not one independently supervised service
per port. Multi-session resource placement and broadcast composition are deployment concerns.
Metrics distinguish session lag, environment step cost, per-agent step cost, barrier wait,
snapshot drops, audio discontinuities, checkpoint queue age and resource budget. Bound agent
labels to configured IDs; never label metrics by viewer name or arbitrary event text.
Health must distinguish paused, slow, disconnected backend and dead agent. A generic
watchdog cannot treat “no new Pokémon tiles” as a stall detector for all tasks.
Deployment preflight checks every referenced profile/bundle and complete checkpoint identity
before selecting a release. Switching the binary back does not convert newer state: retain
the previous release's state namespace for rollback. Twitch publishing still requires the
operator's explicit approval for that run; implementation benchmarks use local sinks.
## 7. Cleanup strategy: extract, then reorganize
Prioritize coupling that prevents a second application, rather than renaming everything.
1. **Characterize behavior and identify state owners.** Capture deterministic synthetic
observation→action→reward traces and legacy restore outcomes before moving code.
2. **Remove task construction from anatomy.** Introduce separately hashed readout bindings;
leave the current artifact and macro-role merge path frozen for legacy compatibility.
3. **Extract orchestration from transport.** A session step returns observations/events and
capture requests; it does not serialize HTTP/feed headers. `flysim` owns listener setup,
wall-clock publication and systemd integration.
4. **Split emulator from task.** Move binjgb behind an environment implementation, preserving
FFI/cache behavior. Move map/exit/objective queries into Pokémon task interfaces rather
than forcing every future game to implement them.
5. **Split task recovery from storage.** The ratchet decides a task transition; storage
commits an opaque coherent capture. Matches use round resets, not milestone archives.
6. **Version observation/control at the boundary.** Internal typed session snapshots become
v1 or v2 through adapters; core modules do not depend on wire enums.
7. **Make stage state instantiable and assets descriptor-driven.** Preserve hot/cold cadence
and fixture determinism while removing singletons and hardcoded dataset selection.
8. **Only then move directories/extract packages.** Keep re-exports/CLI wrappers during the
move, fix build/CI/vendor paths, and compile tiny consumers proving Rust/TS libraries can
be used without starting Twitch, a browser, or an emulator.
9. **Reconcile documentation.** Separate current contracts/reference from dated deployment
history. Audit contradictory “raw buttons only,” weighted-macro, throughput, token/setup,
and training-improvement claims against code. Update templates and scientific limitations
with actual profile capabilities, not a new generic claim of biological fidelity.
Do not introduce a generic reward engine, universal memory address model, central plugin
marketplace, or per-neuron network transport. Those abstractions have no demonstrated second
consumer and would obscure the useful, small seams already present.
## 8. Implementation plan and acceptance gates
Each row is a reviewable change or small workstream, not one large feature branch. Follow
the repository's branch/worktree build-and-review workflow. Contract changes precede their
consumers. No estimate here assumes an emulator backend or biological mapping already works.
| Phase | Deliverable and principal files | Dependencies | Acceptance / stop condition |
| --- | --- | --- | --- |
| P0 — baseline | Trace fixtures and boundary tests around `Sim::step_frame`, restore, dataset loader, stage singleton behavior; current-state documentation inventory | None; reconcile open branches first | Current TS/Rust goldens and legacy API/feed fixtures pinned; behavior ledger distinguishes intentional transient reset from exact replay |
| P1 — identities | Dataset/profile manifest, role resolver, behavior hash, synthetic fixtures; new strict validators in both languages | P0 | Missing/ambiguous roles fail; wrong profile refuses restore; FAFB legacy fingerprints/version strings unchanged |
| M1 — MaleCNS import | `tools/datasets/malecns`, source lock, inventory/exclusion report, schema-compatible baseline bundle, attribution | P1 | Counts reconcile to selected inventory; reproducible output hashes; graph invariants and both loaders agree; no runtime downloads |
| M2 — headless characterization | Profile-specific L1/sensor binding, role/readout audit, learning-off/on benches and new goldens | M1 | Stable/finite activity measured over repeated seeds; no silent empty populations; exact TS/Rust state agreement; unsupported inputs remain marked unsupported |
| M3 — task integration | Explicit MaleCNS profile selected by service config and matching stage assets, fresh state namespace | M2 and minimal descriptor/asset support from P5 | Fresh brain on audited game state; sugar capability verified; stage index/geometry identity correct; one-hour local soak plus restore drill; no automatic promotion over FAFB |
| P2 — environment/task split | Environment contract, binjgb wrapper, task-specific inspector; facade for existing `flybrain-gb` imports | P1 | Legacy action/reward traces identical; ROM-free tests pass; optional ROM-backed sample confirms stepping/audio/state behavior |
| P3 — session library | Single-agent session owns clocks, agent/task/executor state; service becomes transport/composition wrapper | P2 | Same legacy order and compatibility; renderer/bridge restart leaves run intact; fake backend runs without binjgb/ROM/Twitch |
| P4 — multi-agent + persistence | Two-agent synthetic arena, action barrier, isolated state, session envelope/recovery and failure behavior | P3 | One world step per batch; no port/order advantage; resume/parallel-vs-sequential equivalence; failure cannot partly commit a match |
| P5 — feed/control v2 | Descriptor, scoped snapshots/events/media, targeted stimulation and idempotency, TS/Rust schema fixtures/fake server | P1, P3; validate with P4 fixture | V1 still works for legacy; multi-agent attachments cannot collide; unknown target/profile rejected; retry/reconnect tests pass |
| P6 — modular stage/bridge | Instantiable stores, dataset assets, shared/independent views, target-aware commands/redemptions | P4–P5 | PNG review for two-agent layout; browser legibility/fixture tests; correct audio ownership and no agent cross-talk |
| E1 — new emulator spike | Select title/backend; private IPC or native wrapper; record frame, ports, media, inspection and restore capability matrix | P2, can run beside P4–P6 | Reliable bounded step + simultaneous controls, pinned content/backend identity, reproducible state round trip; stop before task implementation if unavailable |
| E2 — fighting-game vertical slice | Two flies, chosen game task, analog/digital readout, episode logic, attributed rewards, match view | E1, P4–P6 | Repeated local matches and side swaps; measured compute headroom; documented scaffold and interventions; win-rate claims require controls |
| P7 — packaging/reorg | Optional root Rust workspace move, library consumers, profile-based release/asset manifests, updated infra and docs | Useful second backend + P6 | Four merge suites and affected browser gates pass; old composition deploys locally; preflight rejects incompatible multi-agent state |
M3 may use a small v1-compatible **additive descriptor extension**, if contracts and both
consumers are updated and the existing frame semantics stay unchanged. It must not publish
MaleCNS spikes under implicit FAFB geometry. Full multi-agent publishing still requires v2.
### 8.1 A concrete first alternative-emulator spike
Before committing to Melee/Dolphin or an N64 backend, establish:
- An exact title/version and backend revision, available to the operator externally.
- A supported pause/advance boundary with all configured controller ports applied together.
- Whether rendering is required for stepping and whether frame capture is synchronous.
- Sample rate/channel metadata, audio latency and timestamps.
- Analog sticks/triggers and button semantics; neutral state on disconnect.
- Save-state completeness, version/platform constraints and reproducibility after restore.
- Supported task-state inspection (match/round/port state) without guessing memory offsets.
- Headless/runtime packaging, process lifecycle, resource cost and failure behavior.
A library used for competitive tooling may expose controller and game-state APIs without
supporting arbitrary frame stepping or faithful visual capture. Verify capabilities instead
of assuming its name solves integration. Prefer a pinned private backend process if native
embedding would force emulator internals into our Rust runtime. Desktop keyboard automation
is unsuitable for simultaneous deterministic multi-port input.
Start with synthetic constant/alternating controller traces and a test opponent before neural
control. Such traces are backend tests, not public “fly playing” footage. Then validate one
agent, two agents, episode boundaries, crashes and resume in that order. No copyrighted game
content is added to source control or test fixtures.
### 8.2 Validation matrix
**Numerical/format:** existing `core/tests/golden_{toy,real,agent,restore,platformer,versions}.rs`
remain gates. Add pinned MaleCNS subgraph fixtures and a full-artifact optional golden run,
with generated TS goldens and exact Rust comparisons. Include malformed CSR, missing roles,
wide IDs, hemisphere/coordinate errors, overflowing weights, and graph/profile mismatch.
The subgraph tests prove arithmetic/loader agreement, not full-CNS dynamics.
**Scheduler:** synthetic backend asserts one advance per complete batch; swap agent evaluation
order, vary worker count, and inject a late/failing participant. Check equal observation
boundaries, deterministic seeds, no cross-agent gains/holds/stimulation, correct tick remainder,
and no reward twice at an episode boundary. Mixed profiles are allowed only if their clocks
and capabilities satisfy the same session contract.
**Persistence:** kill/fault injection around capture/write/manifest commit; corrupt one agent
chunk, backend state or profile hash; verify all-or-none restore and fallback reporting.
Compare uninterrupted and resumed new-session action/state traces. For legacy runs compare
against the documented transient-reset behavior instead of demanding a newly invented one.
**Protocol/UI/bridge:** cross-language v1/v2 fixtures; two agents with different neuron counts;
shared and private views; out-of-order descriptor/media, missing optional attachments, stale
snapshots and reconnect; replay seeking; duplicate redemption, lost HTTP response and bridge
restart; explicitly unsupported stimulation. Visual changes require PNG review and browser
tests, not prose approval of a hypothetical layout.
**Performance/science:** benchmark full graph and thresholded graph with learning disabled
and enabled, fixed sensory traces, multiple seeds and side swaps. For game-performance
claims compare against random/readout baselines and learning-off, reporting scaffold,
recovery, intervention and episode policies. Measure sustained real-time factor and tail
latency under two/four flies plus actual browser/capture load; do not extrapolate a single
kernel throughput figure to an entire show. Initial target is ≥1.0× sustained at the declared
agent count with p99 step time within its cadence budget and no growing queues; select a
resource/headroom margin from the measured backend before release.
Before each merge run the repository-required `npm test`, `npm run typecheck`,
`cargo test --workspace` from the Rust workspace, and `infra/tests/lint.sh`. Run affected
Playwright/PNG gates for stage changes. The existing committed FAFB real-data goldens remain
mandatory. New full-MaleCNS integration jobs and ROM-backed tests are explicit optional jobs
with recorded skips, never a hidden network/ROM dependency of normal CI.
## 9. Decisions and open questions
**Recommended decisions now:** preserve FAFB as the baseline; use official MaleCNS provenance;
freeze legacy arithmetic/identities; version profile behavior independently; make environment
and agent separate objects; use one session barrier for a shared match; retain static plugins
and process/session isolation; introduce v2 instead of stretching Game Boy fields indefinitely.
**Questions answered by spikes rather than assumptions:**
1. Which MaleCNS neuron inventory and confidence/threshold policy will be the published bundle?
Can we explain the differing inventories and quantify left/right sensory coverage?
2. Do existing LIF parameters give useful, stable activity on MaleCNS? If not, which explicit
profile calibration is justified, and does a different neural model warrant separate work?
3. Are verified L1 mappings adequate, or is the intended project really an embodied sensory
simulation requiring new encoders and VNC feedback?
4. Which Smash title/backend can satisfy deterministic stepping, simultaneous ports, media
capture and restoration at acceptable cost?
5. How many simultaneous brains fit the actual budget, with which mix of within-brain versus
between-brain workers? Is a compressed media path required?
6. Which match reset/learning/intervention policy defines the show, and which defines a
controlled comparison? They should be separate run configurations.
Success is not just “another connectome loads” or “a second pad moves.” It is a new session
assembled from modules whose anatomy, numerical model, controller, task, recovery and
presentation assumptions are explicit, testable and reusable without changing the old fly.
## 10. Sources and scope of the analysis
Repository evidence is enumerated in section 2 and tied to the baseline commit above. Existing
reference documents: [dataset format](../dataset-format.md), [model](../model.md),
[plasticity](../plasticity.md), [readout](../readout.md), [limitations](../limitations.md),
[macros](macros.md), [architecture tour](../architecture-tour.md), and
[contribution/compatibility rules](../../CONTRIBUTING.md). Historical status notes are not
evidence that an unmeasured experiment succeeded.
External sources inspected 2026-09-18:
- [Official MaleCNS overview](https://www.janelia.org/project-team/flyem/male-cns-connectome):
anatomical coverage, collaboration, release dates and licensing statement.
- [Official MaleCNS downloads](https://male-cns.janelia.org/download): versioned bulk tables,
confidence cutoffs, segment versus neuron distinction, coordinate units and API guidance.
- [FlyWire overview](https://flywire.ai/): FAFB reconstruction provenance and brain coverage.
- [Codex dataset listing](https://codex.flywire.ai/): portal inventory counts, which are not
assumed to be identical to neuPrint query inventories or local runtime graphs.
- [Minecraft fly README](https://github.com/blendi-remade/fly-brain-minecraft/blob/main/README.md)
and [provenance](https://github.com/blendi-remade/fly-brain-minecraft/blob/main/PROVENANCE.md):
a separately authored MaleCNS derivative, thresholding and mapping decisions, and disclosed
sensory/motor limitations. These moving links are comparison material, not a locked data
dependency; M1 must acquire its own official source lock.
This analysis reads the current implementation and upstream documentation. It does not
download/build the full MaleCNS dataset, independently validate the Minecraft benchmarks,
run a new emulator, or establish a performance/behavioral improvement. Those are explicit
deliverables with gates above.

View file

@ -0,0 +1,747 @@
# Melee: multi-fly runtime, emulator and broadcast audit
Status: **research and proposed implementation plan**. Written 2026-09-18 against local
`f7bc13a`. Extends the [modular-session design](malecns-modular-sessions.md) and
[implementation backlog](malecns-modular-implementation.md) with a concrete second game.
Existing [feed](../feed-protocol.md) and [control](../control-api.md) contracts still win.
No emulator, game image, deployment host or live broadcast was run for this audit.
## 1. Executive decision
**Use Dolphin, initially a pinned mainline-based Slippi Dolphin build, as a separate backend
process. Use the maintained libmelee fork to accelerate the integration spike. Keep the
brains and session coordinator in Rust.** Prove its synchronization, rendered sensory input,
and recovery behavior before selecting a production build. Keep stock Dolphin plus a narrow
backend hook as the fallback if the Slippi path cannot satisfy those requirements cleanly.
Do not port Melee to native code, embed Dolphin into the neural crate, run one emulator per
fighter, or assume “libmelee has `step()`” supplies the complete environment contract.
Recommended first target:
- Melee US 1.02, local two-player versus, one emulator and two independently stateful flies.
- Existing FAFB brains first. MaleCNS is an independent axis of experimentation, not a
prerequisite for solving the emulator and multiplayer problems.
- Fixed declared characters, stage, stock/time rules and input profiles; no netplay rollback.
- Pixel-based sensory input with an explicit resolution/aspect transform; state inspection
is for task measurement, match lifecycle and display, not a hidden fighting policy.
- Fixed GameCube controller readout with bounded analog values and short frame-based holds.
- Native-resolution rendering initially, local broadcast at 30 fps first; simulation/input
continue at the backend's approximately 60-Hz cadence. Promote to a 60-fps show only after
the full media path and encoder are verified.
- Two-fly synthetic arena remains the framework test case before game integration.
The important scaling change is **one environment with many agents**, not simply a larger
ROM. Disc size is mostly a loading/storage concern. Runtime cost comes from PowerPC emulation,
graphics/audio, multiple neural simulations, synchronization and media copies.
### Confidence labels used below
- **Observed in source:** checked implementation or explicit upstream documentation.
- **Recommended:** proposed architecture/configuration, not implemented here.
- **Must measure:** cannot be established by reading code, including throughput, latency,
correct input-to-frame association and reliable recovery on the target platform.
## 2. What the Melee decompilation gives us
The supplied [doldecomp/melee](https://github.com/doldecomp/melee) repository is a matching
decompilation of **US 1.02**. Its README is `.github/README.md`, not the repository-root
`README.md`. The inspected revision is recorded in section 13.
The README and `config/GALE01/config.yml` identify the matching `main.dol` SHA-1 as
`08e0bf20134dfcb260699671004527b2d6bb1a45`. That identifies the executable, **not the entire
disc image**. Our run manifest must separately identify externally supplied game content,
effective executable, modifications, emulator build and task interpretation.
The decomp builds a GameCube executable, not a supported desktop port or emulator replacement.
The `dolphin` code within that source tree refers to Nintendo's SDK, not the Dolphin emulator
project. Rebuilding/relocating a DOL for instrumentation changes address identity; never use
stock symbol addresses on a shifted build.
### 2.1 Useful inspection map
| Upstream source | What was observed | How it helps our task |
| --- | --- | --- |
| `config/GALE01/{config.yml,symbols.txt}`; `docs/symbols.md` | Matching binary identity and named symbols with addresses, sections and attributes | Reproducible symbol/inspection manifest analogous to the Pokémon symbol generator |
| `src/melee/pl/player.h` and `player.c` | `StaticPlayer`, getters for stocks/damage/controller index, KO-by-player counters and self-destructs | Distinguish controller port, player slot and match attribution instead of assuming they are identical |
| `src/melee/ft/types.h` | `Fighter`, player/controller identifiers, buffered sticks/triggers/buttons, pressed/released edges, damage state, source-player field and move-instance information | Audit controls and potential reward attribution; source fields are hypotheses to validate against live transitions |
| `src/melee/gm/types.h` | Match frame/timer fields, `MatchEnd`, winner arrays and exit/results structures | Episode boundaries, timeout/results interpretation, one-time terminal rewards |
| `src/melee/gm/gmvsmelee.h` | Character/stage select, versus entry/exit, sudden-death and results transitions | An explicit lifecycle model instead of treating every screen as a playable frame |
| `src/melee/{cm,gr,mp,it}/` | Camera, stages, map/collision and item subsystems identified by upstream module structure | Follow-up inspection points for view geometry, hazards and projectiles; not all audited in this pass |
Examples of concrete distinctions:
- `StaticPlayer` has a controller index, a player ID and up to two sub-fighter entities.
Ice Climbers and transformations defeat “one visible fighter object = one agent.”
- The input struct tracks three-entry analog/button histories plus pressed/released buttons
and threshold timers. A constant button hold and a sequence of taps are different actions.
- Damage includes an annotated source-player number, but it must be checked for projectiles,
stale ownership, self-damage and indirect KOs before it becomes a reward source.
- The match structures include winner counts/arrays. A terminal result is not safely inferred
from whichever player's stock decrement happened to be sampled first.
### 2.2 How to use it without making the framework Melee-specific
Create a **task-local** inspection catalog: field name, binary/decomp revision, symbol or
pointer traversal, type/endianness, valid scenes, tested transitions and unsupported cases.
Prefer Slippi telemetry for fields it supplies reliably; use decomp-grounded inspection for
missing fields only after verification. Do not expose all emulator memory to every agent.
If memory inspection is needed, decode big-endian integer/float fields and emulated 32-bit
pointers explicitly. Never cast emulated bytes to a native Rust/C struct whose layout,
pointer width or bitfield ordering is different. Sample a consistent backend boundary rather
than reading a moving process asynchronously. Accessors named in the decomp explain semantics;
they are not functions our host process can call in place of an adapter.
Derive small constant/schema outputs and synthetic fixtures where appropriate. Do not vendor
the entire decomp, game executable, assets or save states merely to read stocks and damage.
The first working backend does not require rebuilding Melee. Custom hooks or patches are
separately identified scaffold and enter the run's behavior/content identity.
## 3. Emulator options and recommendation
All candidates below emulate GameCube software. The differentiator is the host integration,
not whether Melee can theoretically boot.
| Candidate | Evidence / useful capability | Gap or cost | Decision |
| --- | --- | --- | --- |
| **Mainline-based Slippi Dolphin + maintained libmelee** | Structured game/port events, controller pipes, explicit blocking-input support; Linux rendered path available in the ecosystem | Pixel/audio access and coherent external save-state control still need integration; pin emulator, Gecko codes and parser together | **First spike and preferred initial integration** |
| **Stock Dolphin + narrow host hook** | Source has `Core::DoFrameStep`, CPU-thread coordination and `State::{SaveToBuffer,LoadFromBuffer}` | These are internal APIs, not a stable remote environment SDK; own a small patch and state parser/telemetry bridge | Fallback or eventual generic Dolphin backend if its maintenance cost is justified |
| **Felk Python-scripting Dolphin branch** | Inspected stubs expose controller/memory/save-state scripting and rendered-frame events | Historical branch; `await frameadvance()` is documented as waiting for a rendered-frame event, not proof of a paused one-step transaction | Research reference or temporary probe, not default production dependency |
| **Custom EXI/fast-forward Slippi-Ishiiruka** | Maintained libmelee README describes accelerated ML mode and EXI inputs | That documented fast path disables rendering; inspected libmelee rejects non-Null graphics for the EXI_AI build | Useful for explicitly state-driven offline research, not the pixel-fed live baseline |
| **Libretro Dolphin core** | Potential common frontend ABI | No core-specific synchronization/state/render benchmark was performed; another integration layer does not remove task semantics | Defer rather than introduce an unverified second dependency stack |
| **Native game port built from the decomp** | Source enables modding/research | Matching DOL compilation is not native execution; graphics, OS/SDK, timing and assets remain substantial work | Outside this project's first Melee phase |
Use the maintained **`vladfi1/libmelee`**, not a floating install selected by an old tutorial.
`altf4/libmelee` says it is archived and points there. The maintained fork says it became
the PyPI `melee` source starting at 0.45.0; pin the actual chosen package/source revision and
its dependencies instead of assuming an unversioned `pip install melee` reproduces a run.
The maintained README describes raw-state compatibility, but the inspected `Console.step()`
still invokes `__fixframeindexing` and `__fixiasa`. Therefore, confirm field semantics from
the installed source and observation fixtures rather than trusting README wording. Our
adapter identifies parser/normalization revision as part of task identity.
### 3.1 What the blocking path actually does
**Observed in source:**
1. `libmelee.Console` defaults `blocking_input=False`. Setting it true writes
`Slippi/BlockingPipes` for its mainline backend.
2. `Console.step()` flushes its registered controllers and then dispatches game/menu events
until a frame boundary. It is not a method returning RGBA pixels or arbitrary game state.
3. Slippi's `EXI_DeviceSlippi.cpp` sets `g_need_input_for_frame` on game setup, menu frames
and frame bookends.
4. `Pipes.cpp::UpdateInput` checks the blocking setting and flag, waiting for commands through
`FLUSH`. Its Linux wait path uses `select`; the inspected Windows wait helper is not implemented.
5. `ControllerInterface::UpdateInput` updates devices and only then clears the flag. The
`FLUSH` handler explicitly avoids clearing it before the other devices have been read.
That is strong evidence for trying a Linux multi-port blocking backend. It is **not yet a
measurement** that action batch `t` affects exactly our desired frame `t+1`, that menu/game
boundaries behave identically, or that rendering/audio are coherent with the telemetry.
Keep at most **one batch outstanding**. The pipe implementation consumes buffered commands,
and a backlog of frame batches must not collapse into a latest-state input. Create an isolated
Dolphin user directory containing only the intended pipe devices: unused/abandoned devices
can participate in updates and leave blocking input waiting on a controller nobody drives.
For two flies, stage both full controller states before calling the single owner's
`Console.step()`. Do not give each agent a `Console` loop. Verify which input sample corresponds
to returned pre/post-frame telemetry with distinguishable action pulses and deliberate delays
on each port. “Both bots sent commands” is weaker than “both commands landed on one frame.”
### 3.2 What libmelee does not establish for us
- No framebuffer-returning or whole-session save/load interface was found in the inspected
`Console` API. `DumpConfig` configures media dumping; a dump is not automatically a bounded,
timestamped sensory-frame transport.
- GameCube pad values are stateful. Unchanged buttons remain held; the backend must emit an
explicit complete state or compute trustworthy deltas, including release/neutral values.
- The library applies analog normalization (`fix_analog_stick`, `fix_analog_trigger`). Define
our canonical ranges and apply conversion exactly once; test round trips at neutral,
extremes, diagonals, dead zones and trigger-click thresholds.
- Rollback skipping and internal controller flushes are present. Use local offline matches
first; do not confuse filtered rollback frames with advancing brains through speculative time.
- Initial game events can flush neutral input internally. Record startup as lifecycle scaffold;
do not attribute that to a neural decision.
- Spectator transport keepalive, pipe blocking and process liveness are separate. A healthy
connection does not prove the match or renderer is advancing.
## 4. Audit of our current system: keep, extract, replace
Rust path aliases: `core/`, `gb/`, `sim/` mean the respective crates under
`services/flysim/crates/` named `flybrain-core`, `flybrain-gb`, and `flysim`.
| Finding | Current evidence | Required change for Melee/multiple flies | Priority |
| --- | --- | --- | --- |
| Single world and brain bundled together | `sim/src/simloop.rs::Sim` owns one `NeuralAgent`, concrete `Emulator`, adapter and ratchet | Session owns one environment plus agent collection and explicit port map | Blocking |
| Direct Game Boy calls in frame loop | `step_frame`, `to_button_mask`, `set_buttons(u8)`, `run_frame`, fixed framebuffer copy | Backend interface with full action batch, rational cadence, observations and media capabilities | Blocking |
| Task trait carries Pokémon concepts | `gb/src/adapter.rs` includes `read8(u16)`, map/exit/objective hooks | Keep inspector and macros task-local; use generic reward/episode/progress outputs | Blocking |
| Timing defaults embed Game Boy | `core/src/agent.rs`, `sim/src/config.rs` | Session clock derived from backend, per-agent neural remainders; preserve old arithmetic in legacy facade | Blocking |
| Decoder tuned to walking through Pokémon maps | `packages/brain/src/readout/presets/gameboy.ts`: 800-ms directions, 85-ms A/B pulses with 480-ms cooldown | New fixed GameCube mapping; frame-scale controls, analog sticks/triggers, concurrent movement/action | Blocking |
| Neural code is already independently useful | `core/src/lif.rs`, `agent.rs`, `plasticity.rs`; TS counterparts | Reuse exact kernel and private per-agent state; do not replace neural semantics to integrate a game | Keep |
| Graph can be shared, pool cannot be concurrently dispatched | `Arc<BrainDataset>`, `SweepPlan`, `core/src/pool.rs` shared job slot | Immutable topology shared; distinct mutable state and controlled total scheduling budget | Blocking |
| CUDA exists but not an automatic service optimization | `core/src/lif/cuda.rs`; no `enable_cuda` call found in `sim/src/simloop.rs` | Explicit backend selection, equivalence/restore tests, profile first; don't promise GPU brains from graphics availability | Measured option |
| Snapshot header is a single Pokémon-shaped view | `sim/src/snapshot.rs`, `packages/feed/src/{types,codec}.ts` | Session descriptors, multiple agents/ports, task-specific progress, named attachments/media | Blocking for proper broadcast |
| Stage assumes Game Boy geometry and one fly | `apps/stage/src/{App.tsx,feed/store.ts,feed/decode.ts,lib/geometry.ts}` | Per-session/agent store instances, dynamic view aspect/dimensions, multi-agent match layout | Blocking for proper broadcast |
| Browser owns game audio | `apps/stage/src/audio/engine.ts`, 48-kHz feed, Pulse capture | Explicit audio producer and media-clock policy; do not play native Dolphin audio and forwarded PCM twice | Blocking |
| Capture already offers NVENC | `infra/bin/flycast-launch` | Reuse encode/relay/recording; measure new compositor/readback cost and revise 60-fps settings | Keep with changes |
| Encoder hardcodes H.264 level 4.1 | Both encoder functions in `flycast-launch` | 1080p60 needs a suitable level (normally 4.2 or automatic selection); changing only `FLY_FPS` is insufficient | Required for 1080p60 |
| Existing checkpoint payload is single-agent/binary-specific | `sim/src/store.rs`, `gb/src/compatibility.rs` | Coherent all-agent/world checkpoint plus backend/parser/patch/controller identity | Blocking for exact resume |
| Checkpoint writer queue is unbounded | `Sim::start_writer` uses `std::sync::mpsc::channel` | Bound/coalesce background work; larger emulator captures and several brain copies must not grow an unlimited queue | High |
| Existing “saved” event precedes durable commit | `checkpoint_with_reply` emits after enqueue; writer updates durable metrics on success | Distinguish capture/enqueue/commit/failure events; public status must not report a queued Melee save as durable | High |
| Recovery assumes a best progress ladder | `gb/src/{ratchet,recovery}.rs` | Match/episode reset policy; never rewind one player's world independently | Blocking |
| Health is mostly loop heartbeat | `sim/src/simloop.rs::Shared`, `sim/src/lib.rs` | Distinguish waiting at input barrier, intentional pause, backend timeout and deadlock; keep host supervision responsive | High |
| Deployment resource partitions reflect the old stack | `infra/units/flysim.service`, deploy cpuset construction | Budget Dolphin CPU/GPU plus N brains and media; measure and set new memory/process limits | Required before release |
| Bridge targets one sim | `services/bridge/src/{sim,commands,redemptions}.ts` | Explicit stable agent targeting and intervention policy; no viewer control-port endpoint | Before interactive show |
Do not interpret dated CPU/GPU measurements in the repository as current free capacity.
The records identify bandwidth contention and graphics-sharing constraints, but this audit
does not inspect the host or establish that it can run two brains plus Dolphin in real time.
## 5. A framework architecture that survives a third game
### 5.1 Four runtime components, not one enormous adapter
```text
private backend protocol
Rust session process <----------------------------------------------> backend helper
session clock / port ownership Python + libmelee initially
N agent states owns Dolphin process/user dir
sensory encoders │
fixed readouts ├─ all controller pipes
task events / reward router ├─ telemetry/parser
episode policy / checkpoints └─ media + state hooks
│ │
├─ feed/control adapters Dolphin
└─ durable event/checkpoint store one local match
│
stage/compositor → capture → local relay → optional Twitch push
```
The helper is an **internal backend implementation**, not an audience-accessible controller
service. The session remains the sole authority assigning actions to ports. Python never
simulates the neurons or selects actions. Keep it if measurements say its overhead is small;
replace its internals with Rust/native IPC only when an actual bottleneck or maintenance
requirement justifies it.
An external emulator is a normal `Environment` implementation, not a special `if melee`
branch sprinkled throughout the session. For Game Boy, the same interface has an in-process
implementation. For an embodied world, it may be a native physics engine. Consumers do not
need to know which one owns the world.
### 5.2 Framework contracts to extract
| Contract | Owns | Must not know |
| --- | --- | --- |
| `Brain` / numerical core | Tick semantics, state, spikes, rates, learning updates | Game, process, controller labels or viewer |
| `SensorEncoder` | Declared observation→neural drive transform | Reward inspector state not declared as input |
| `Readout` | Fixed rate/signal→control-channel mapping | Opponent strategy, game addresses or pathfinding |
| `ActionExecutor` | Selected action→controller state; optional declared macro lifetime | Authority to invent a winning action when brain is silent |
| `Environment` | Native clock, port schema, action commit, observations/media, snapshot capabilities | Neural roles, Twitch or task reward weights |
| `Task` | Typed state interpretation, rewards, progress and episode outcomes | Direct neural mutation or direct controller writes |
| `EpisodePolicy` | Start/end/reset/recovery semantics | Hidden per-player rewind in a shared world |
| `Session` | Barrier, identity, agent isolation, routing, state capture and supervision | Melee memory offsets or Pokémon map IDs |
| `Presentation` | Descriptor-driven layout, media/audio and task panels | Emulator stepping or access to controller pipes |
Use modules first, then crates/packages as second consumers appear. The existing monorepo
and Rust workspace can stay in place during extraction. Keep static compiled registries
initially; a stable dynamic-plugin ABI and public package registry are not prerequisites.
**Framework acceptance test:** adding a synthetic third environment/task requires a backend
implementation, a task/profile and a composition manifest, not edits to the session loop,
neural core, protocol enums or generic stage store. A game-specific presentation plugin is
allowed. Wire extensions must be namespaced/schema-validated, not hardcoded into every panel.
### 5.3 Minimal backend IPC
Specify and test a small private protocol before implementing a network-shaped abstraction:
```text
Hello → backend build/content/patch identity, capabilities, cadence, ports, views
Initialize(runConfig) → Ready(epoch, observationBoundary)
Advance(epoch, expectedBoundary, completePortBatch)
→ StepResult(epoch, newBoundary, appliedBatchId, observation, mediaRefs)
Pause / Resume → explicit acknowledgment
Capture(epoch, boundary) → capture token + state digest + snapshot bytes/reference
Restore(captureToken) → new epoch + restored observation + success/failure
Shutdown → acknowledgment or bounded forced process termination
```
Only advertise `Capture/Restore` if implemented and tested. Otherwise expose an explicit
`restart_episode` capability and visible aborted-match policy; never claim exact resume.
Port batches use canonical buttons plus sticks in [-1,1] and triggers in [0,1], converted
once by the backend. Descriptor/schema versions pin conversions and active ports.
Use a local framed socket for commands/small observations; use a bounded shared-memory ring
or equivalent for large media. Include epoch, frame/sample identity, dimensions, format and
generation on media references. Release/acquire ownership and slot lifetime prevent the
backend overwriting a sensory buffer while an encoder reads it. Reconnection invalidates old
handles; a stale frame must not be silently accepted because its byte length matches.
One request in flight, explicit timeouts, bounded queues. After an uncertain `Advance`
response, do **not** resend blindly: the world may already have advanced. Resolve batch ID/
boundary through an idempotent reply cache or fail/recover the session. A pipe transport with
no application acknowledgment needs validation against returned input telemetry, not a claim
of atomicity it cannot prove.
Separate lifecycle I/O from the blocked advance operation so diagnostics/shutdown stay alive.
Do not issue a save operation scheduled on Dolphin's CPU thread while that same thread is
waiting forever for pipe input. Capture/pause needs a backend-owned quiescent point where
both the simulation and pending input consumption have known state.
## 6. Multiple flies and the timing contract
### 6.1 One match, one world clock
Two flies in Melee normally means two controller ports in **one Dolphin instance**. Four
flies means four ports, not four copies of Melee joined through netplay. Several independent
matches are separate sessions/process trees and can share one broadcast director.
For each committed boundary:
1. Freeze each agent's allowed observation from the same world state.
2. Advance every brain by the environment's elapsed emulated time using its private remainder.
3. Decode independent actions; step any declared executors; assemble a complete port batch.
4. Commit the batch once and let the backend advance to the next acknowledged boundary.
5. Associate rendered sensory frames and telemetry with their actual producing boundary.
6. Route task events/rewards exactly once; encode the next input and publish a snapshot.
Warm-up settles/calibrates brains with learning off while the environment is held at its
initial boundary. The next observation must not be from a game that ran freely through the
warm-up. Fixed offsets between brain and environment clocks are recorded.
Use the selected backend's rational emulated cadence, not `GAMEBOY_MS_PER_FRAME` or an
unexamined exact 60. An approximately 60-Hz budget is about 16.7 ms, but logical game frame,
video interrupt, input poll, rendered presentation and Slippi frame bookend are distinct
events until the spike establishes their mapping. Session ticks are monotonic even when
Melee's signed frame number resets, starts before zero, or changes across menu scenes.
### 6.2 Render latency and agent fairness
Dual-threaded graphics can present frame `n` after telemetry for `n` is available. Label the
actual frame; do not attach “latest screenshot” to current state and assume equivalence.
Start with one declared fixed observation latency shared by all agents, and measure it.
If a pipeline deliberately adds one frame of latency, record that in the sensor profile.
The full-resolution spectator view may be delayed separately, provided overlays use the
matching presentation timestamps rather than future task data.
Also test a potential pipeline deadlock: the helper waits for a rendered image while Dolphin
is waiting for the next controller flush needed to reach that presentation event. Fix the
backend's rendezvous or select a declared previous-frame sensory latency; do not unblock it
with an undisclosed neutral gameplay input. Game-state bookends alone do not prove the GPU
has completed a matching frame.
Do not reduce brain integration from 1,000 to 500 ticks per emulated second to meet wall-clock
deadlines. That changes the model. A declared action-repeat interval can reduce decisions,
but normally still requires all neural ticks and correctly accumulated intermediate task
events; it does not halve the principal neural cost. Rendering every second game frame is
also a sensor change if the brain otherwise sees each frame, not merely a broadcast setting.
On slow compute, the default is to slow the entire local match and report real-time factor.
Do not let one fly continue while the other misses turns. On a dead participant/backend,
pause or abort the match visibly; neutral fallback play is not silently substituted.
### 6.3 No rollback netplay in the first release
Slippi supports online play, but we do not need it to connect two local flies. Online rollback
would require rewinding **all** neural states, RNG, decoder/executor state, reward ledgers
and admission decisions at the same speculative boundary as the game, then replaying inputs.
Filtering repeated frames in libmelee is not that system. Keep offline local matches and
assert monotonic committed observations per epoch; classify unexpected rollback as an error
or explicit recovery transition rather than double-rewarding it.
## 7. Controller, sensory and learning design
### 7.1 A GameCube controller is not an eight-bit pad
Support independent main and C sticks, analog shoulders, digital trigger clicks, face
buttons, start and D-pad. Movement and attack can overlap. Canonical neutral/release state
must be complete, so a missing command cannot accidentally leave attack or shield held.
At about 60 Hz, the current 800-ms direction hold is roughly **48 game frames** and the
85-ms pulse roughly five. Reusing these values would dominate the fly's behavior regardless
of the dataset. Create a fixed Melee readout with explicit decisions in integer game frames,
bounded analog mappings, dead zones, tie handling and pulse/hold policies. Start with a small
declared set of stick magnitudes and directions if that makes validation easier; continuously
valued mappings can follow as a separate profile.
Audit tap-jump, directional aerials/smash attacks, jump release, shields, simultaneous axes,
and conflicting inputs. Don't add state-conditioned auto-aim, automatic edge recovery or
combo execution under the label “controller mapping.” If later desired, publish those as
separate macro/assistance profiles with their own identity and comparison baseline.
Start/system controls are lifecycle-sensitive. During an active match the profile may omit
pause entirely; initial match setup and between-match reset are disclosed episode scaffolding.
This is not permission for an API to press buttons. Specify whether setup uses an audited
initial state, internal deterministic menu setup, or a reset hook, and mark those frames as
non-neural setup with learning disabled. That expands the legacy “all presses” phrasing and
requires a deliberate task-policy/documentation decision before shipping it.
### 7.2 What the fly sees
For the initial pixel profile, both flies receive the same shared game camera with fixed
crop/aspect treatment, independent of spectator overlays. The legacy 160×144 input should
not stretch a 4:3 scene silently. Compare an aspect-preserving downsample/letterbox transform
with a separately versioned input-size profile; changing kernel retina dimensions currently
affects the numeric configuration identity.
Current L1 projection samples luminance at 1,572 columns. Higher broadcast resolution does
not produce more sensory neurons, color recognition, motion estimation or knowledge of which
fighter the fly controls. Test whether each selected character remains visible across zoom
and stage movement; record the sparse sensory representation rather than assuming a human-
readable video is an adequate neural input.
Three distinct modes must not be conflated:
| Mode | Neural observation | Rendering implications |
| --- | --- | --- |
| Pixel baseline | Actual game image through fixed encoder | Needs real rendered frames even if no desktop GUI is shown |
| Structured-state experiment | Explicitly encoded positions, velocities, stocks, etc. | Can potentially use Null/fast-forward, but it is a new privileged-input model |
| Spectator-only rendering | Whatever the profile specifies; video for audience | May be independently compressed/delayed, never silently substituted for sensory input |
“Headless” can mean no GUI while still rendering; “Null graphics” generally means no useful
pixel observation. The documented EXI fast-forward speed path cannot be advertised as the
performance of our pixel-fed broadcast.
### 7.3 Reward and outcome attribution
Implement a `MeleeTask` with typed per-player observations, match state, a ledger and positive
reward events. Its schema belongs to the task, not a generic `GameMode` enum. Start small:
- Terminal match outcome, once, based on validated results/termination reason.
- Opponent damage and credited KOs only after verified ownership information is available.
- No reward for mere button activation, for losing a stock, or for scripted setup.
Do not reward A for every increase in B's percent: self-damage, stage effects, reflected
projectiles, teams and another sub-fighter can invalidate that inference. The decomp's source-
player and KO tables guide inspection, but no runtime correctness is claimed until tested.
Unknown attribution produces a logged observation without a guessed reward. Keep fractional
damage until the rule deliberately quantizes; HUD damage and fighter damage may differ.
Deduplication keys include epoch/match, producing frame and event identity. Handle multihits,
trades, simultaneous KOs, respawn percent reset, timeout, sudden death and disconnection as
separate cases. If source telemetry cannot distinguish a required case, narrow the first
ruleset or add a specific audited observation hook.
Learning remains private per fly and synthetic reward modulation remains distinct from PAM
stimulation. Disable learning during kernel/controller/backend characterization; later compare
learning-on with learning-off, repeat seeds and swap sides/characters. Retaining gains between
rounds is a run policy. Competitive success is not guaranteed by increased model complexity.
## 8. Performance plan: measure the actual critical path
### 8.1 Budget equation
For a lockstep pixel-fed match, approximate the critical wall-time interval as:
```text
T_step = T_brains + T_readout/task + T_controller_IPC
+ T_emulation_to_observation + T_required_render_readback + T_boundary_overhead
T_brains ≈ sum(T_agent_i) [sequential evaluation]
T_brains ≳ max(T_agent_i) + barrier cost [parallel with sufficient independent resources]
```
The parallel estimate is a lower bound, not a promise: shared caches, memory bandwidth,
GPU contention and scheduling can make every brain slower. Media publication/encoding and
storage should be off the critical path, but their resource use and state capture still
affect it. Do not obtain a “60 fps” claim solely from Dolphin's display counter while the
brains advance fewer milliseconds or repeat stale observations.
Proposed capacity gate: warm full-stack **unthrottled** throughput at least 1.2× the selected
game cadence for two flies, then a paced one-hour soak with no growing queues/lag and a
24-hour local endurance run before release. In the paced run, distinguish intentional wait
from compute time; report p50/p95/p99/max compute interval and deadline misses. The 1.2×
margin is a proposed engineering target, not a measured capability of current hardware.
### 8.2 CPU/GPU strategy
1. **Keep the first two brains on CPU.** Establish Dolphin JIT/render/media cost separately.
The repository's current service is CPU-composed even though a CUDA kernel exists.
2. **Allocate a total physical-core budget.** Compare sequential agents with modest within-
brain pools against agents running concurrently on disjoint core groups. Include Dolphin's
CPU/JIT thread, graphics worker, helper, browser, encoder and storage in the budget. Do not
launch four copies of the current per-brain pool size by default.
3. **Use native-resolution hardware rendering first.** Compare OpenGL/Vulkan on the chosen
build/platform; no blanket claim that one is faster. Measure render correctness and readback.
JIT is the performance baseline; an interpreter is a diagnostic baseline, not the live plan.
4. **Benchmark shader compilation and caches.** Report cold and warm starts separately. Choose
supported shader modes from measurements; a cache that hides startup hitches is not a
guarantee that a new stage/character will not compile something mid-match.
5. **Use NVENC where available, with a measured fallback policy.** Its encode engine does not
remove GPU rendering, memory allocation, color conversion or framebuffer-readback cost.
Automatic fallback to x264 can consume the cores the brains/emulator need; expose the
resulting degradation and test whether the declared session can still meet cadence.
6. **Only then test CUDA brains.** The current backend retains RNG/plasticity observation/rate
work on the host, uploads state inputs, and by default synchronizes membrane/refractory
state back each batch. It also allocates device graph/state per backend instance. Measure
two/four agents alongside Dolphin, browser graphics and NVENC; zero-copy shared graphs and
a GPU-wide scheduler are possible later work, not present features.
For CUDA, preserve bit-exactness, gain-update ordering and checkpoint synchronization.
Batching across a future game action boundary is not valid just because it improves kernel
throughput. Keep the TypeScript oracle and existing version strings intact.
The existing VirtualGL/Xvfb result demonstrates one Chromium rendering path, not that
Dolphin Vulkan/OpenGL works or is performant in the same container. Test the complete selected
graphics path. Reusing GPU passthrough requires no assumption of exclusive VRAM availability.
Resource availability must be measured in an approved, serialized deployment-host session.
### 8.3 Media bandwidth and copies
Uncompressed RGBA estimates, before copies/framing:
| Image/cadence | Bytes per second |
| --- | ---: |
| Existing 160×144 at 30 fps | 2.76 MB/s |
| 640×480 at 30 fps | 36.86 MB/s |
| 640×480 at 60 fps | 73.73 MB/s |
| 1920×1080 at 60 fps | 497.66 MB/s |
640×480 is a planning example, not an asserted fixed Dolphin framebuffer size. The backend
advertises actual dimensions/format/aspect. One shared camera is delivered once for the
match; two flies can sample one immutable image without duplicating its transport. If their
sensor transforms differ, encode separately against the same source frame.
Prefer reducing/downsampling the sensory copy close to the renderer, ideally before GPU
readback, while preserving a separately timestamped spectator view. Benchmark against an
ordinary CPU path before adding device-buffer interop. A shared-memory ring removes socket
copies, not the GPU fence/readback itself. Slow viewers may drop frames; required sensory
frames may not disappear silently from the neural run.
### 8.4 Benchmark ladder and decision records
| Run | Configuration | Question / recorded output |
| --- | --- | --- |
| B0 | Synthetic two-port backend, no neurons | IPC latency, one-step semantics, barriers, timeouts, media-buffer ownership |
| B1 | Dolphin with fixed input traces, rendering/audio on, no brains | Cold/warm emulator cost, step timing, port alignment, render-to-state latency |
| B2 | Same run + helper/media extraction | Incremental parsing, copying, downsampling and audio cost |
| B3 | One FAFB brain | End-to-end reference and per-phase costs |
| B4 | Two FAFB brains, sequential vs parallel schedules | CPU/cache/bandwidth limits, balanced observation and input timing |
| B5 | B4 + actual stage, capture, relay, recording, checkpoints | Full critical path, A/V drift, encoder fallback and queue growth |
| B6 | Four brains and four active ports | Capacity characterization only until this independently passes the same gates |
| B7 | Matched MaleCNS and optional CUDA variants | Dataset and backend effects, measured independently before combined variants |
Record backend/content/patch/profile digests, physical-core allocation, exact graphics settings,
sensor/broadcast resolutions, all clock rates, resident/peak memory, VRAM, thread usage,
real-time factor, latency distributions, audio under/overruns and dropped frames by purpose.
Store operator-specific machine details externally and publish only the portable methodology
and non-identifying results. No measurements were performed by this document-writing task.
## 9. Broadcast architecture and audio ownership
### 9.1 Two viable routes
**Route A — stage receives game media.** Closest to the current architecture: backend emits
pixels/audio, stage composites game and overlays, ffmpeg captures the page. Start the local
prototype with bounded lower-resolution raw frames to validate semantics. If bandwidth and
copying dominate, add a compressed local media track (for example WebRTC) while telemetry
remains on the feed. Avoid encode→decode→encode unless its measured simplicity/latency tradeoff
is acceptable. Browser frame presentation timestamps must align overlays with displayed video.
**Route B — compositor combines native game output and stage overlay.** Dolphin supplies its
rendered output to a compositor; the browser supplies a separate overlay surface. This can
avoid moving full-resolution game pixels through JavaScript, but requires an explicit shared
clock and a new capture composition. The fly's sensory image still needs a frame-identified
path from the backend. Capturing a desktop window on a wall clock is insufficient to establish
which image a brain used at a given game boundary.
**Recommendation:** prototype Route A for the two-player local slice; benchmark Route B in
the media spike before committing to the long-run high-resolution pipeline. Preserve media
as a capability behind the environment interface so the choice does not change brain/task code.
The v2 protocol should be able to reference media streams, not mandate all video as WS RGBA.
### 9.2 Audio and clock policy
Today the page plays binjgb PCM and stream SFX into the Pulse sink. With Dolphin, select one
of these explicitly:
- Dolphin PCM is captured/forwarded and played by the page, with native device output muted.
- Dolphin renders audio to the capture sink directly, and the page contributes only SFX.
Do not run both. Declare sample format/rate, resampling location, timestamps, buffering and
discontinuity handling. Pause/reset/restore must flush or relabel buffered old-episode audio.
If wall time falls behind, measure pitch/time-stretch behavior; do not let “async resample”
hide minutes of simulation lag. Game timestamps, not arbitrary browser receipt time, define
the intended A/V relationship.
Keep 30-fps broadcast and approximately 60-Hz gameplay as independent settings. For 1080p60,
the existing H.264 level 4.1 is too low for the normal macroblock-rate requirement; use a
compatible level such as 4.2 or encoder-selected level and validate the actual stream. Also
measure capture/compositor cadence, bitrate quality, encoder lookahead/latency, local recording
and audio synchronization. `FLY_FPS=60` alone is not a completed performance upgrade.
### 9.3 Presentation changes
The generic stage needs descriptor-driven game aspect ratio, two/four agent cards, per-port
button/stick indicators, per-agent learning/sugar state, shared match stocks/percent/results,
and scoped events. Neither “badges” nor “highest ladder rung” describes a match.
Separate task data schema from layout. Keep one compositor clock and one selected world audio
stream; keep neural maps and rate scalers private per agent/dataset identity. Defer expensive
four-avatar/whole-connectome rendering until measured. All actual layout decisions require
PNG mockups and existing legibility/browser checks. This text does not approve a screen.
## 10. Persistence and unattended operation
### 10.1 Savestates are a capability, not a libmelee assumption
Dolphin source provides buffer/file state operations, but `State.h` documents that operations
called off its CPU thread may be scheduled rather than executed immediately. An external
“save requested” is therefore not proof of a consistent capture at our agent boundary.
Slippi's internal rollback save-state commands likewise do not constitute an audited public
multi-agent checkpoint API.
The selected backend must supply acknowledgment of the frozen boundary, resulting state
digest and completion. Save every brain, RNG, rate/calibration state, learning state, sensor/
executor state, pending action identity, task ledger and environment together. Clock and
media epochs change on restore; discard pre-restore spectator/parser buffers.
Re-create or explicitly reinitialize libmelee's parser caches and controller history after
restore. An emulator savestate does not include an external helper's `_frame`, previous
game state, normalization state or queued pipe data. Test how game-start metadata is supplied
when loading into mid-match; some telemetry protocols may need reseeding or restarting.
When exact mid-match capture is unavailable, an initial prototype may visibly abort and
restart a match while retaining a declared brain checkpoint. Mark `resume=episode-restart`
in the descriptor. That is a deliberate narrower capability, not equivalent to crash resume.
It is not ready for a release that promises uninterrupted exact match continuation.
### 10.2 Storage and health
Bound pending captures and coalesce replaceable hot checkpoints. A durable request either
completes with a commit acknowledgment or fails explicitly; never drop it while reporting
success. Match-end result records are append-only and independent of world rewind.
Keep backend/helper/brain health separate from “game frame did not advance.” An intentional
pause or waiting barrier is not a crash; a dead helper must not keep the session green by
merely refreshing an HTTP heartbeat. Export last completed boundary, in-flight request age,
barrier participant status and renderer progress. A hard stall has a timeout and explicit
match abort/recovery path, not repeated blind restarts of unrelated services.
Manage Dolphin under the session's lifecycle or an explicitly coordinated systemd unit. If
Dolphin restarts, the session cannot keep sending frame `t+1` to a fresh match. Allocate unique
user directories, pipe names, telemetry ports and state namespaces per independent session.
Use deterministic configuration provisioning and checksums instead of reusing a developer's
desktop Dolphin settings or permitting auto-updates.
The existing `flysim.service` memory ceiling and CPU partition were sized for a different
process graph. Set new cgroup/resource limits from measured high-water marks; account for
backend process, multiple brain copies and checkpoint transients. Release preflight checks
the complete backend/game/patch/parser/profile identity. Rollback retains compatible state
as well as the old executable.
## 11. Staged implementation and go/no-go gates
This specializes the existing backlog rather than replacing its foundation/session work.
Melee-specific spikes can start before the full framework reorganization is finished.
| Item | Work and dependency | Evidence required before the next step |
| --- | --- | --- |
| **MELEE-01: backend selection spike** | Specializes EMULATOR-01. Pin mainline Slippi + maintained libmelee, content and Gecko codes; use isolated user config and two synthetic controllers | Boot/render/audio; block one then both ports; one batch/frame mapping; menu→match→results lifecycle; cleanup/restart. Choose this build or stock Dolphin + narrow hook based on results |
| **MELEE-02: capture/restore spike** | Alongside MELEE-01; prove sensory-frame identity, media export and save/load acknowledgment independently | Fixed pixel↔telemetry latency, bounded media storage, correct input after restore, parser/cache recovery. Explicit decision: exact resume or prototype-only episode restart |
| **MELEE-03: task observation audit** | Pin decomp; build field catalog, typed parser/inspector, lifecycle and synthetic event fixtures | Verified port/player/sub-fighter mapping; stocks/results; no guessed rewards; content/patch mismatches visibly disable unsupported semantic interpretation |
| **FRAMEWORK-01: generic backend + session** | Existing FOUNDATION-01/02 and RUNTIME-01/02; add the private IPC implementation behind `Environment` | Same legacy Game Boy traces; headless synthetic environment uses identical session API; no Melee branches in core loop |
| **FRAMEWORK-02: multi-agent and state** | Existing RUNTIME-03/STATE-01; integrate complete action batches, worker budget and chosen backend recovery capability | No cross-agent state leakage; one world step; changed evaluation order invariant; failed restore cannot partly install a match |
| **MELEE-04: fixed readout and sensory profile** | MELEE-01/02 + framework boundary; TS specification then Rust implementation for any new decoder semantics | Neutral/release, analog conversion, tap/hold/direction combinations, aspect-preserved neural input and recorded latency; no hidden combo/aim policy |
| **MELEE-05: first two-fly match** | MELEE-03/04 + FRAMEWORK-02; learning off, then audited positive rewards | Recorded action/observation timelines; paired side/seed trials; match terminal deduplication and visible reset/failure semantics |
| **MEDIA-01: full local show** | Existing WIRE-01/PRESENTATION-01; compare Route A/B, define audio owner and 30/60-fps profiles | PNG review, fake multi-agent fixtures, measured copies/latency/A/V drift; sustained media pipeline under checkpoint and shader-load events |
| **PERF-01: two-fly capacity gate** | Benchmark ladder B0–B5; optimize measured limiting phase | ≥1.2× warm unthrottled capacity target, one-hour paced soak and 24-hour local endurance; no queue/lag growth; documented CPU/GPU/memory envelope |
| **MELEE-06: expand carefully** | Passing two-fly slice | Four ports/teams, broader characters/stages, MaleCNS and CUDA are separate experiments, each with new tests and its own capacity result |
| **FRAMEWORK-03: finish packaging** | Existing PACKAGE-01 after useful second backend | Example third backend can be added without core/session/schema edits; isolated compositions package and preflight correctly |
**Stop conditions:** no dependable step/input barrier; no identifiable pixel source for a
pixel-input claim; inability to attribute rewards under the claimed ruleset; unsupported
restore marketed as exact resume; or sustained capacity below the declared cadence.
Respond by changing the explicit supported scope, backend or resources—not by quietly skipping
neural ticks, adding a gameplay bot, hiding game stalls or reporting guessed measurements.
### 11.1 Suggested first experiment script
The first implementation should be a local measurement harness, not the final stream:
1. Launch one pinned backend with known local content and two configured bot pads.
2. Enter a fixed local match through the declared setup procedure; record episode boundary.
3. Send distinct short left/right and A/jump pulse patterns on each port, including neutral
frames; log intended batches and observed raw/processed controller values.
4. Delay one port by a controlled wall-clock interval and verify no game boundary commits
until the complete batch is available. Repeat with port order reversed and four ports.
5. Capture a sequence of images/telemetry with frame identities; measure their association.
6. Save/restore at a known barrier if supported, replay the same actions, and compare task/
input traces; verify helper state and buffered media are reset coherently.
7. Kill the helper/backend separately and verify bounded failure without accidental continued
play or permanent hangs. Test paused-state health independently.
8. Measure compute with no brains, one brain, two brains, then the full broadcast stack.
Synthetic controller traces are test machinery, not footage presented as neural play. Keep
game content and environment-specific records outside source control; store portable metrics,
synthetic schemas and independently authored tests in the repository.
### 11.2 Test matrix that catches Melee-specific failures
- **Input:** two/four ports, inactive slots, delayed/missing flush, stale buffered commands,
full release, analog endpoints/deadzones, short taps, pressed versus held edges.
- **Identity:** controller↔player mapping, swapped ports, sub-fighters, transformations,
character/stage changes, wrong game revision, changed patch/parser normalization.
- **Events:** multi-hit, trade, self-damage, projectile ownership, stock reset, simultaneous
KO, timeout, sudden death, results re-entry, disconnect, restart after accepted reward.
- **Clocks/media:** game-frame reset, renderer lag, stale shared-memory generation, dropped
spectator frame versus required sensory frame, paused audio, mismatched overlay timestamps.
- **Recovery:** all-agent atomic validation, one corrupt state chunk, backend import failure,
helper parser not reinitialized, asynchronous save completion, hot-store coalescing and
durable-write failure. Test new exact resume separately from legacy transient-reset behavior.
- **Performance:** cold/warm shaders, high-activity matches, checkpoint capture bursts, CPU
encoder fallback, browser reconnect, two/four neural agents and measured GPU contention.
All implementation merges retain repository-required TS tests/typecheck, Rust workspace
tests and infra lint; UI changes add Playwright and PNG review. Game-backed jobs are explicit
operator-provided tests. Normal CI uses synthetic observations/backends and existing goldens.
## 12. Decisions to carry into implementation
| Question | Recommended answer now | Still requires evidence/choice |
| --- | --- | --- |
| Which emulator? | Dolphin, first trying mainline-based Slippi + maintained libmelee | Exact build selected by synchronized-input/media/state spikes |
| Use the decomp to run the game natively? | No; use it to audit task/state/controller semantics | Custom instrumentation only for specifically missing observations |
| One emulator per fly? | No for one match; one per independent session | Four-port capability must be tested, not inferred from two ports |
| Which brain? | Two existing FAFB agents for integration baseline | MaleCNS comparison after mappings/dynamics pass their independent gates |
| CPU or GPU brain? | CPU baseline, share immutable graph | CUDA versus CPU benchmark under Dolphin + capture, not in isolation |
| Inputs to the brain? | Pixels with explicit fixed transform | Structured state is a distinct optional research profile |
| Start with macros? | Fixed controller mapping, no hidden aim/combo policy | Any later assist profile is separately disclosed and evaluated |
| How fast? | Backend-native gameplay/input cadence, 30-fps initial show | Full-stack two-agent capacity; optional 60-fps broadcast and four flies |
| How to resume? | Whole-session coherent state where supported | Episode-restart prototype if exact state interface is not yet available |
| How generic? | Concrete environment/task/agent/session contracts and composition examples | Extract public packages only after second/third consumers prove the seam |
The first operator choices needed are the initial characters/stage/ruleset, desired show
cadence, and whether a visibly restarted match is acceptable during the prototype. They do
not block the synthetic framework work or source-level backend spike design.
## 13. Sources and audit scope
Local code evidence appears in section 4. Additional local files inspected include
`core/src/lif/cuda.rs`, `sim/src/pacing.rs`, `sim/src/simloop.rs::start_writer`,
`packages/brain/src/readout/presets/{gameboy,platformer}.ts`,
`apps/stage/src/audio/engine.ts`, `infra/bin/flycast-launch`,
`infra/units/flysim.service`, and the profiling/VirtualGL methods under `infra/docs/`.
External source snapshots inspected on 2026-09-18 (pin actual dependencies again at spike start):
| Repository/ref | Observed revision | Files used |
| --- | --- | --- |
| [doldecomp/melee](https://github.com/doldecomp/melee/tree/b9ec8a2eb48520753b2f8159ccc94d033fbf60ea) `master` | `b9ec8a2eb48520753b2f8159ccc94d033fbf60ea` | `.github/README.md`, `docs/symbols.md`, config, player/fighter/match headers and player implementation |
| [vladfi1/libmelee](https://github.com/vladfi1/libmelee/tree/bce21f09984b286e6d36bfd2939e4cd4691f94c2) `master` | `bce21f09984b286e6d36bfd2939e4cd4691f94c2` | README, `melee/console.py`, `melee/controller.py`, license metadata |
| [project-slippi/dolphin](https://github.com/project-slippi/dolphin/tree/41a7a3a110ed52999486ae1901c8fbb9a63d4f13) `slippi` | `41a7a3a110ed52999486ae1901c8fbb9a63d4f13` | Pipe backend, controller update loop, Slippi EXI events, `Core/State.h` |
| [dolphin-emu/dolphin](https://github.com/dolphin-emu/dolphin/tree/ee018d00e60b9eb727489908a8daec5c537f44a8) `master` | `ee018d00e60b9eb727489908a8daec5c537f44a8` | `Source/Core/Core/Core.h`, state/core module inventory |
| [Felk/dolphin](https://github.com/Felk/dolphin/tree/46b7eacd5c810c2d21ec5fe51ea1a9c61a7ceb3d) historical `scripting` branch | `46b7eacd5c810c2d21ec5fe51ea1a9c61a7ceb3d` | Scripting README, `python-stubs/dolphin/{event,savestate}.pyi` |
| [altf4/libmelee](https://github.com/altf4/libmelee/tree/1da979657122facd0750ea99cf6858255e198326) `main` | `1da979657122facd0750ea99cf6858255e198326` | Archive notice directing users to maintained fork |
Dolphin files inspected carry GPL-2.0-or-later headers; libmelee repository metadata reports
LGPL-3.0. Pin and retain actual dependency licenses/notices when packaging. A separate process
is an architectural boundary, not an assertion that distribution obligations disappear.
Source inspection supports the integration hypotheses and concrete constraints above. It
does not establish Dolphin throughput on the deployment hardware, verify any game-memory
field live, demonstrate a new neural behavior, or prove exact multi-port/frame/save semantics.
Those are the measured deliverables of MELEE-01/02 and the performance ladder.