Commit graph

174 commits

Author SHA1 Message Date
flybrain
a6a8f797ff survey: the conversation probe, and every rectangle the cartridge drew
Row 56 is a gym's dialog for thirty brain minutes with YES, NO and NEXT dealt
in equal thirds, so the question is which box each press is answering and
whether the seam can see it at all. `state::drawn_boxes` asks the screen
instead of a pinned rectangle: every complete `TextBoxBorder` on the frame.
`FLY_PROBE_CATCH=dialog` walks the conversation one raw pulse at a time and
prints both readings of every frame side by side, with the pad the palette
deals for it, and tallies how the candidate readings separate the frames.

`FLY_PROBE_ANSWER` picks the button a frame with a box on it is answered with,
so the two arms of a choice can be walked separately.
2026-09-23 00:30:35 +00:00
acamilo
2de2dced0a tests: the service integration test expects the v6 adapter
Some checks are pending
ci / node 22 (test + typecheck) (push) Waiting to run
ci / rust stable (cargo test --workspace --release) (push) Waiting to run
ci / infra/tests/lint.sh (push) Waiting to run
ci / playwright apps/stage (allowed to fail) (push) Waiting to run
2026-09-22 23:00:14 +00:00
acamilo
cfb3c1c506 docs: v0.5.2 status
Some checks are pending
ci / node 22 (test + typecheck) (push) Waiting to run
ci / rust stable (cargo test --workspace --release) (push) Waiting to run
ci / infra/tests/lint.sh (push) Waiting to run
ci / playwright apps/stage (allowed to fail) (push) Waiting to run
2026-09-22 22:49:33 +00:00
acamilo
c24cbe0171 Merge fix/loop-row55: the shop pad is read from the drawn screen and offers only what can run 2026-09-22 22:49:30 +00:00
acamilo
6342adba39 docs: row 55 gates, as run on the merged tree 2026-09-22 22:48:11 +00:00
acamilo
e5fe6c46ac Merge main into fix/loop-row55: row 50's battle seam beside row 55's counter
Three conflicts, all of them two reviews appending in the same place, all
resolved by keeping both sides: `macros.md` (row 50 keeps section 12.18, row 55
becomes 12.19), `macros-traps.md` (both trap rows, both residual pairs, both arm
sections) and `scene_probe.rs` (both survey modes, `accept` and `shop`). Every
source file auto-merged.

The two rows compose on the cartridge: from the mart checkpoint the fly leaves
the counter, walks the town, and no `MOVE n` blocks any more, so the longest
chain of one macro blocked on an unchanged frame is 1 in any scene.
2026-09-22 22:38:10 +00:00
acamilo
549df6c323 docs: row 55, and the two claims about a mart the survey narrowed
`macros.md` section 12.18 is the review; `macros-wram.md` section 7.1 corrects
the address table, where `wListMenuID` was said to be cleared by every text
display and an item's position in the stock was said to be its cursor index;
`macros-traps.md` carries row 55, its two arms and its ROM run, and two new
residuals -- `wListScrollOffset` is not a pinned address, and a mart that has
never drawn a buy list reads `Unknown` rather than `Shop`.
2026-09-22 22:29:07 +00:00
acamilo
01c199fe18 probe: clippy's three readings of the counter survey 2026-09-22 22:29:07 +00:00
acamilo
351338038a docs: v0.5.1 status
Some checks are pending
ci / node 22 (test + typecheck) (push) Waiting to run
ci / rust stable (cargo test --workspace --release) (push) Waiting to run
ci / infra/tests/lint.sh (push) Waiting to run
ci / playwright apps/stage (allowed to fail) (push) Waiting to run
2026-09-22 22:13:09 +00:00
acamilo
349a1fd28c Merge fix/loop-row50: a battle turn is the fly's only while the move list is drawn 2026-09-22 22:13:08 +00:00
dev
e307b7dde4 docs: row 50's survey, its two trap hunts and its ROM run
The press survey, twenty brain minutes before and after on both checkpoints,
and the ROM-gated forest run.

The ethos check's 'fewer flagged windows, more distinct tiles' holds on the
tiles on both arms (296 to 430, 163 to 184) and does NOT hold on the windows
(33/73 to 70/73, 68/73 to 73/73). After the fix the fly spends three quarters
of each run inside battles it is actually fighting, and the hunt's tile rule
flags a fighting fly exactly as hard as a stuck one. Reported rather than
smoothed.
2026-09-22 22:11:01 +00:00
acamilo
01494a104e docs: the session framework slices that landed, and the two that are blocked
Some checks are pending
ci / node 22 (test + typecheck) (push) Waiting to run
ci / rust stable (cargo test --workspace --release) (push) Waiting to run
ci / infra/tests/lint.sh (push) Waiting to run
ci / playwright apps/stage (allowed to fail) (push) Waiting to run
2026-09-22 21:11:55 +00:00
acamilo
954b9f4db5 Merge fix/sf-resolution-bound-flake: a short acknowledgement is the contract, and two timing bets become claims 2026-09-22 21:11:49 +00:00
acamilo
b2afce1fce Merge feat/sf-publish-01: the publication boundary, committed snapshots and observer isolation over the same bus 2026-09-22 21:04:40 +00:00
dev
5e61c50728 session: an unreadable event batch is not the end of the stream
take_events kept returning Option<EventBatchView> and defaulting through the
question-mark operator, so a batch missing a field read as end of stream and the
ConsumerEvents enum added in the previous round described nothing. It returns
Batch or Unreadable now, and a test publishes a batch with no droppedBefore,
asserts it is reported as unreadable naming the field, and asserts the next real
batch still reads.

The checkpoint-envelope-v1 amendment cites the rule that lets a required
manifest field land with envelopeVersion still 1 while no production file
exists.
2026-09-22 21:03:59 +00:00
acamilo
50ab3d47ba session: hold an Acknowledge to the request it answered
Dropping the length check left nothing checking the reply against the
request at all: AcknowledgeResult::validate_against existed with no
caller, so a worker could acknowledge ids this session never asked
about. That is the other half of the rule the README states. A short
list is the worker reporting what it released and is accepted; an id
from outside the request is the worker reporting about someone else's
cache and is refused, named, before any mutation.

The bootstrap-path regression is now covered too. The earlier test calls
acknowledge_replies directly, which guards the check where it lives but
not where it lived, so a length check put back into acknowledge_lifecycle
left it green. The duplicate_lifecycle_acknowledge injection releases the
ids first, out of sight, so the call that method makes and checks is
already the second one -- the shape the section 6 resolution produces.
Verified by putting the old check back: three tests fail with it, none
without.

Also: last_resolution_attempts is cleared with last_resolution, so a
resolution ending before its first attempt no longer reports the
previous count; the guard half of the bound test asserts the fence like
the budget half; the attempts assertion checks a real bound rather than
u32::MAX; and two dead Instant bindings are gone.
2026-09-22 21:03:09 +00:00
dev
305a50d8df session: a graph identity does not cross a recovery
The index an agent attested to at Agent.Initialize joins its compatibility
identity and its checkpoint manifest row, so a replacement fly that built
another graph -- the same dataset, the same neuron count, another index --
cannot install a checkpoint taken under the first one. It was accepted before,
because agent_compatibility digested a dataset digest recomputed from a free
function rather than what the worker attested to, and the restored composition
was then published under its predecessor's indexDigest.

The refusal names what the worker is, not just that two digests differ. Dated
amendment to checkpoint-envelope-v1's agents row; no wire type and no schema
text change, so the contract digest is unchanged.
2026-09-22 20:36:51 +00:00
dev
6bf5687aca session: a snapshot publishes the telemetry of the transition that just ended
AgentSlot.telemetry was written only by the Agent.Initialize handler, so every
CommittedSnapshot carried warm-up telemetry labelled as boundary k while each
AgentCommitResult.telemetry was validated and dropped. The commit result is now
stored on the slot beside the committed step, and the media test asserts that
published telemetry advances across boundaries instead of merely being nonzero.

A refused snapshot is recorded and sequenced like a refused descriptor revision,
because the repair path exists for the consumer that did not receive it; the
sequence advances with the value rather than with the delivery, so two snapshots
can never share one. The query service counts an answer it could not deliver
rather than discarding the result, and an unreadable event batch is distinct
from the end of the stream.
2026-09-22 20:18:49 +00:00
acamilo
0cec73704c docs: v0.5.0 deployed, the run restarted from rung 8
Some checks are pending
ci / node 22 (test + typecheck) (push) Waiting to run
ci / rust stable (cargo test --workspace --release) (push) Waiting to run
ci / infra/tests/lint.sh (push) Waiting to run
ci / playwright apps/stage (allowed to fail) (push) Waiting to run
2026-09-22 20:00:43 +00:00
acamilo
270d4cb435 tests: the fly leaves the mart counter, and a blocked chain is measured
From the checkpoint the stream looped in: the fly is out of the shop scene and
off the map inside twenty brain minutes, no macro reports `blocked` more than
three times in a row in the shop with the map, the tile, the wallet and the scene
unchanged, and the shop never deals an empty pad.

The chain is recorded per scene as well as overall, because the battle scene has
one of its own and it belongs to another review.
2026-09-22 19:56:38 +00:00
acamilo
8523daf745 macros: a purchase the counter's cursor cannot reach is not on the pad
A mart's buy list scrolls. The cursor walks rows 0, 1, 2 and then the window
moves under it, so an item's position in the counter's stock is its cursor index
only for the first three entries; the offset that would name the rest is not a
pinned address. Pewter's counter carries seven items and the Antidote is its
fourth, so `BUY ANTIDOTE` there aimed a cursor step above the list's own max and
reported `blocked` on its first frame, having pressed nothing -- and a purchase
has no walk target, so nothing was recorded and the button came back on the next
hold.

One accessor answers "which index, if any" for both the pad and the plan, and the
clerk talking is refused through the seam rather than here.
2026-09-22 19:56:38 +00:00
acamilo
4a1df7c3be macros: the mart's list byte says the counter is open, not which screen is up
`wListMenuID` keeps PRICEDITEMLISTMENU for a whole mart visit: the clerk's text
is printed from inside the mart's own routine and never goes through the display
that clears it. So every frame of the Pewter mart read "the priced buy list",
including the "Here you are! Thank you!" box the stream was looking at, whose
leftover cursor bytes belong to a two-option box.

Which screen is up is now read from the figure the game draws, the same
construction the dialogue box's `waiting` test and the YES/NO prompt already
make: a full-width box waiting for a press is the clerk (`ShopScreen::Talking`),
the item window drawn is the buy list, and the item window blank is the counter
menu.
2026-09-22 19:56:38 +00:00
acamilo
da2cde176c probe: a counter survey that reads what the mart actually drew
`FLY_PROBE_CATCH=shop` drives a `BUY ...` from a checkpoint the stream looped in
and prints, per frame, what the seam reads at the counter, the halves of every
purchase's precondition, the outcome of the macro three times over, and what an
A press at the counter really opens. `screen_text` decodes the tile buffer,
because the question is which figure the game drew and the bytes that would name
it are not rewritten between two of the mart's screens.
2026-09-22 19:56:38 +00:00
acamilo
e3132362cf Merge main at 5512900: STATE-01, and the flybus coalescing fix
The flybus session_over_one_router failure my workspace runs were
counting is already fixed on main: the coalescing branch forces the
coalescing deterministically instead of asserting that a slow consumer
must skip. Merging before the runs so they measure a tree that exists.

One conflict, in the crate README, and it was two sections both newly
added at the same anchor rather than two versions of one thing. Both are
kept: STATE-01's checkpoint and recovery section, then the guidance on
which replies permit a subset, which stays immediately above the
contract-narrowing section where someone adding a check will meet it.

coordinator.rs and the process tests auto-merged. Checked rather than
assumed: the acknowledge equality check is still gone, acknowledge_replies
and last_resolution_attempts are present, STATE-01's capture, durable
and rebase surfaces are present, and both test changes survived.
2026-09-22 19:49:14 +00:00
acamilo
c1a972548b session: the death rows kill once the victim provably has the work
Both death rows killed after a fixed sleep, so under load the kill could
land before the call was dispatched. The bus then reports not-dispatched
and MutationCertainty::None, which is correct -- the participant never
received anything -- while the test demanded unknown. A full-workspace
run caught it: "left: None, right: None" at the certainty assertion.

kill_once_it_is_working polls the victim's own Worker.Status until it is
provably inside the operation before killing: the agent until it is
Preparing with an active request id, the world until it has recorded the
batch, which the arena does before its injected delay. Dispatch has then
demonstrably happened and unknown is the only correct certainty.

The wall-clock boundedness assertions are gone with them. The suite's
own `within` is the bound, and the victim is five seconds slow against
its twenty, so returning at all is the claim.
2026-09-22 19:48:10 +00:00
dev
7ce645f1dc session: a restored boundary is an installed one, and the revision follows the composition
A group restore re-establishes a committed boundary this epoch did not run a
transition into, so its snapshot carries no decision and no controls for any
agent, and a fresh epoch is a new compositionDigest, so the descriptor takes the
next revision rather than republishing revision 1 with different contents.

CommittedSnapshot's rule becomes: null at boundary 0 and at an installed
boundary, always together, and for every agent or none -- a snapshot where one
fly acted and another did not would be two boundaries in one value. Dated
amendment to publishing-v1 section 3, with the schema set, the fixtures and the
TypeScript package moved together and the digest regenerated.
2026-09-22 19:24:44 +00:00
dev
9501a5a17c macros: the survey numbers in the accessor are the corrected pulse's
The first pass of the press survey counted 187 refusals that were its own held
button: JoypadLowSensitivity acts on a key's edge, so a direction the fly was
already holding produced no press. With the pulse releasing first, the reading
is 264 honoured of 3,102 by the cursor bytes alone and 231 of 231 by the bytes
and the box.
2026-09-22 19:09:42 +00:00
acamilo
ef82a05a3a docs: which replies permit a subset, and which do not
The sweep behind this branch found one check demanding an exact match
where the contract permits a short answer, and four that were right to
demand one. The question that separates them belongs where the next
check gets written, not only in a run report: is the far side reporting
what it did, or being held to a requirement?

Worker.Acknowledge is the only reply of the first kind here, because
ipc-v1 section 5 makes it idempotent. The commit and batch checks are
the second kind and are named so nobody loosens them later in the name
of tolerance; they are what make a partial commit and an incomplete
batch fail.
2026-09-22 19:05:58 +00:00
acamilo
f6baeb5cfa session: a test for the acknowledge rule, not only for the flake
The short-list acknowledgment is a contract rule, so it gets a test that
says so rather than one that depends on a worker being slow.
acknowledge_replies carries the ipc-v1 section 5 sentence in its doc
comment and returns what the worker actually released;
acknowledge_lifecycle calls it, so bootstrap and the test exercise the
same path.

an_acknowledge_that_releases_nothing_is_not_a_failure drives the case
directly, once per execution mode: bootstrap releases every lifecycle
reply, the test asks for those ids again, the worker ignores them and
releases nothing, and the coordinator must accept the empty list, stay
unfenced, stay at its boundary and still play the next transition. It
fails if the equality check returns.
2026-09-22 19:03:18 +00:00
acamilo
4f2d4a5848 Merge main into feat/sf-publish-01 2026-09-22 19:03:01 +00:00
dev
e27306f171 Merge main into feat/sf-publish-01
# Conflicts:
#	services/flysim/crates/fly-session/README.md
#	services/flysim/crates/fly-session/src/coordinator.rs
#	services/flysim/crates/fly-session/src/harness.rs
#	services/flysim/crates/fly-session/src/launcher.rs
2026-09-22 19:02:57 +00:00
acamilo
55129002a4 tests: the placeholder fingerprint in the migration test no longer reads as a MAC address
Some checks are pending
ci / node 22 (test + typecheck) (push) Waiting to run
ci / rust stable (cargo test --workspace --release) (push) Waiting to run
ci / infra/tests/lint.sh (push) Waiting to run
ci / playwright apps/stage (allowed to fail) (push) Waiting to run
2026-09-22 18:59:17 +00:00
dev
c25c28efe3 test: hold the rung-nine forest run to a MOVE n that finishes, and a median battle
Row 50's two numbers, ROM-gated from the forest checkpoint. Before, on main:
MOVE n 940 starts and 838 blocked, 11 battles entered and 10 ended, worst 503
macros, median 48. After: 51 starts and 0 blocked, 13 entered and 13 ended,
worst 283, median 43, and the fly leaves the forest north through the gate.

The blocked share is the assertion the row is about; the median is what it
buys, and the worst battle is a tail rather than the run.
2026-09-22 18:57:18 +00:00
acamilo
083abe5f1f Merge feat/sf-state-01: a coherent all-participant checkpoint, group restore and a liftable fence 2026-09-22 18:57:18 +00:00
acamilo
2237e9a1a2 Merge main into feat/sf-publish-01 2026-09-22 18:54:26 +00:00
acamilo
6d67fa7ed2 session: a retried Acknowledge is not a failed epoch
The bound test failed two runs in thirty under load, and not on the
bound it was testing. Both failures were bootstrap: "a worker did not
acknowledge every lifecycle reply".

ipc-v1 section 5 says already released or unknown ids are ignored, and
the contract type already holds the acknowledged list to a subset of
the request. So the second Acknowledge of the same ids answers with an
empty list by design -- and the section 6 resolution produces exactly
that second Acknowledge whenever the first reply is slower than the
probe. Demanding the whole list back turned a safe, contract-sanctioned
retry into a failed epoch, which is a defect in the coordinator rather
than in the test: a slow lifecycle reply would do it to a real session
too.

The test made itself easy to hit by installing a fifty-millisecond
probe before bootstrap, so bootstrap's own lifecycle calls ran under a
budget meant for the step under test. It now bootstraps at ordinary
deadlines and tightens them afterwards.

The two bounds are also separated by construction rather than by clock.
Each half puts the bound it is not testing out of reach -- u32::MAX
attempts against a fifth of a second, three attempts against an hour --
so no scheduling delay can flip which one fires, and the silent
participant is ten minutes slow against a twenty-second test timeout,
so returning at all proves a bound ended it. The wall-clock assertion
is gone and the attempt count is asserted instead, which
last_resolution_attempts now records. One agent per composition, so the
participant the failure names is not a race either.
2026-09-22 18:47:42 +00:00
dev
f43fd4d9ff session: supportedStimuli is enforced, and the query method names are provisional
An undeclared stimulus kind is refused before the model is touched, proved by an
injection through Agent.Commit rather than by a unit call, so the declaration a
descriptor publishes is the thing the worker enforces.

The publishing-v1 section 2 amendment now says in its own words that
Session.GetDescriptor and Session.GetSnapshot are internal and provisional names,
which the later public v2 step may rename or supersede.
2026-09-22 18:47:10 +00:00
dev
388e112ad5 Merge main into feat/sf-state-01 2026-09-22 18:45:56 +00:00
dev
45621903be session: name the two ways a durable wait ends without an acknowledgment
Review follow-up on the checkpoint store.

A dropped reply channel and an expired caller budget were both reported as
ReplyLost. They are different facts -- the first means the write is over and its
outcome did not reach here, the second means the save is still going -- so they are
now separate outcomes, and the durable wait has its own budget rather than borrowing
the one that bounds a call to a participant. Both still leave durable metadata where
it was, and for both the resolution asks the store about the same checkpoint.

The writer's two bounds refuse at different moments and the comment claimed
otherwise: the outstanding-capture bound is taken before a capture is requested, and
the byte budget cannot be, because a capture's size is not known until it exists. The
byte check, the decision and the change to the byte total are now one critical
section, the peak is sampled after a superseded job's bytes are gone, and the writer's
own bookkeeping is over a type that holds only the outcomes a writer can produce.

The manifest's coordinator.eventWatermarks is {lastSourceStep, issued}; the fixture
illustrated {lastEventId, lastOrdinal}, and the illustration is what changed, because
an event id is derived from the epoch and cannot be compared across the restore that
gives the session a new one.

checkpoint-envelope-v1 section 3 also now says, under the same dated amendment, that
a required-manifest-field change must bump envelopeVersion once production files
exist: contractDigest is taken over the schema set and does not cover this manifest,
so the envelope version is the only thing that can carry such a change.
2026-09-22 18:45:51 +00:00
dev
6588897ab3 docs: a menu is up while its box is on screen (section 12.18, row 50)
macros.md gains 12.18 and macros-wram.md a section 10 for the accessor and the
survey that found it: press at every battle frame with a rollback pulse, and
ask every byte of WRAM and HRAM which of them separates a honoured press from
a refused one.

The half of row 50 that was wrong is where the fix is. The move list was not
drawn: MoveSelectionMenu's cursor bytes are never cleared and SelectMenuItem
decrements wCurrentMenuItem back into the one-based range on its way out, so a
turn's whole text and animation read as an open list with a placeable cursor.

The battle bag is the same trap on wListMenuID and is named rather than fixed,
because the figure that tells its list from the frame after it closes has not
been surveyed yet.
2026-09-22 18:43:05 +00:00
acamilo
55221a7c77 docs: v0.5.0 status 2026-09-22 18:41:17 +00:00
acamilo
a8afd8f448 Merge feat/catch-reward: a catch reward, adapter v6 with a v5 migration, and a rung restart 2026-09-22 18:41:15 +00:00
dev
e2bf4c0305 macros: the move list is the box on screen, not the cursor bytes it left behind
Row 50: MOVE n reported blocked 890 times in 1,431 macros on the cartridge,
every one of them on a frame the seam read as an open move list with a
placeable cursor.

MoveSelectionMenu writes wTopMenuItemY 12 and wTopMenuItemX 5 and nothing in
the game clears them, exactly as the two-option box's geometry outlives its
box. SelectMenuItem then decrements wCurrentMenuItem back to the 0-based slot
on its way out, which lands straight back inside the one-based range the
accessor reads. So the whole of a turn -- the text, the animation, the
enemy's reply -- read as the fly's own turn on an open list, the pad dealt
the four move buttons on it, and the cursor step pressed at a list nobody was
reading until its budget ran out.

The reading is the figure the menu draws, the same construction text_box's
waiting and yes_no_prompt already make: a box at (4, 12) fourteen wide with a
horizontal run over its top-left corner and the junction tile at (10, 12).
Surveyed with one rollback pulse per battle frame over 3,102 frames at the
rung-9 forest checkpoint: by the cursor bytes alone a real directional press
was honoured on 264 of them, and by the cursor bytes and the box on 231 of
231.

A frame whose list is not on screen is between turns, whose pad is the one
NEXT that advances text.
2026-09-22 18:37:36 +00:00
dev
2987b243f4 probe: press at every battle menu to see which reading says it is accepting input
Row 50 asks which WRAM reading tells a battle menu that is accepting input
from one that is only drawn, and the honest way to answer it is to press.
FLY_PROBE_CATCH=accept drives real battles, and on every battle frame it
exports the emulator state, issues one directional pulse, reads
wCurrentMenuItem and puts the state straight back -- so the run is not
perturbed by the measurement and every frame gets a ground truth.

Beside that it asks every byte of WRAM and HRAM whether its values on
accepting frames are disjoint from its values on refusing ones, so a
reading is found rather than nominated.

The pulse releases the buttons before it presses: JoypadLowSensitivity acts
on a key's edge, so a direction the fly is already holding reads as refused
for the measurement's reason and not the cartridge's.
2026-09-22 18:37:26 +00:00
acamilo
537cdd8ad9 Merge fix/flybus-conformance-rows: the conformance rows name the tests that now prove them
Some checks are pending
ci / node 22 (test + typecheck) (push) Waiting to run
ci / rust stable (cargo test --workspace --release) (push) Waiting to run
ci / infra/tests/lint.sh (push) Waiting to run
ci / playwright apps/stage (allowed to fail) (push) Waiting to run
2026-09-22 18:29:12 +00:00
dev
2cc8d0294d docs: the two sweep flakes are fixed, and one assertion that could not fail
bus-conformance.md still listed the example_demo root count and
session_over_one_router as not fixed here. Both are fixed on main now, so
the rows say what each test does instead: the example waits for the
producer's hold release before reading the counts, with its printed
output unchanged, and the integration renderer is held until the
publisher's twentieth receipt has returned, so its coalescing is forced
and the assertions are order, freshness, acceptance under a stalled
spectator and the replacement accounting, with the before and after
counts.

The delivery check in that test compared (step, sequence) against
(step, step), whose first element can never fail. It compares sequence
against step.
2026-09-22 18:28:54 +00:00
acamilo
882980823c Merge fix/flybus-coalescing-flake: the flybus suite asserts guarantees, not the machine's timing 2026-09-22 18:23:06 +00:00
acamilo
ed4cffcf59 Merge main into feat/sf-publish-01 2026-09-22 17:59:43 +00:00
dev
db30708c3f Merge main into feat/sf-state-01 2026-09-22 17:55:03 +00:00
acamilo
e76b3d1063 docs: v0.4.7 status 2026-09-22 17:46:03 +00:00