9.9 KiB
Earned rules
Rules this campaign paid for. Each one exists because it was violated first and something wrong got published, or nearly did. Lane briefs should reference this file rather than restating it.
Ordered by how much damage the violation caused.
1. A green verdict is not evidence. Coverage is.
tracecmp exit 0 means "the declared regions agreed". It does not mean the hook compared
anything. B4 found three hooks that each printed 0 diverged while comparing nothing:
describe_argsran beforeregions(), so every logged argument was one call stale;- the
StrategyServerhas two bases 4 bytes apart — reading the wrong one yields an empty player vector, zero declared regions, and a confident "1 call, 0 diverged"; - a Ghidra-base number was used as a raw struct offset (
0x64vs0x60), enumerating a vector's spare capacity.
Structural fixes now in place: compile-time-required Coverage on every descriptor (a hook without
one does not compile), guard regions that localise undeclared writes by offset rather than
reporting that a hash moved, replace-mode records, and tracecmp --strict-coverage. A hook claiming
"complete" while a guard caught an undeclared write counts as a divergence.
Read the trace, not the verdict. Every hook bug above was visible to a reader and invisible to the exit code.
2. Write the prediction down before the run
Commit the expected result before building, not before running. Lane U committed its prediction
before the build was even staged; lane P wrote its expected residual (next_id short by exactly 1
on every completion call) and lane V confirmed it live.
The strong form: a prediction that survives a changed workload tests the model, not the recording. Lane U's call 9 was not the call its prediction was written against — the AI had picked a different tech — and it unlocked three nodes at costs appearing in no earlier report. Numbers that never occurred cannot have been memorised.
Include a falsification section: how the model could be wrong, and the symptom of each way.
3. Static reading finds what behavioural comparison cannot
Both of these compared clean and were still wrong:
ApplyTechEffecthad an early return that would have made every real call a no-op;- the RNG divisor was wrong on 0.78% of draws and flipped zero decisions in 10⁶ trials.
Conversely, behavioural compare found the MoveFleet rounding that static reading had missed. They
are different instruments. Use both; trust neither alone.
4. Prefer the instruction stream to the decompiler for control flow
An inlined std::vector destructor looks exactly like a branch. The je skips only the
operator delete; both arms converge a few hundred bytes later. This has produced two wrong
published findings — turn-spine.md's "deferred end-of-turn tail" (there is no such branch) and a
near-miss on Streamable::Write.
Any decompiler if that wraps a whole tail is suspect until you have found the converge point in
the disassembly.
5. Size structs by enumeration, never by what the code touches
_Alval is std::allocator<char>, an empty class: it occupies a word and is never loaded or
stored, so it is invisible to touch-based analysis and undercounts by exactly 4. That produced a
std::string-is-0x18 scare against a correct campaign-wide 0x1c.
An enumeration — serializer, constructor, copy constructor, or container stride — can show
absence where a touch-scan cannot. Same trap is live for std::vector here:
{_Myfirst,_Mylast,_Myend,_Alval} = 0x10, allocator-last, the opposite of the shape the MSVC
textbooks describe.
Related: the on-disk primitive is not the memory kind. A member held as int16/int8 is
written by WriteInt and is four bytes on the wire.
6. A path no save exercises is a hypothesis
Label it as one. spies2, SysMem and mts were "closed" with count 0 in every save — no element
value has ever been observed. Lane A later found the element typing was wrong, and it was invisible
precisely because nothing exercised it. The flagging discipline is what made that checkable.
Saves can be manufactured. No species-5 save existed, so lane V built one and closed a row that had been disassembly-only for a day. Prefer building the workload over weakening the claim.
7. Round-trip byte-identity is not coverage
ar.any bodies round-trip trivially by copying bytes nobody understands. Measure the split between
items a field names and items a node merely carries. That reframing moved a "100% round-trip"
result to an honest 38%, and then to 99.9% by actual work.
8. Two independent-looking checks sharing a hidden assumption are one check
save_reader.py types items from its own kind catalog, not the schema passed to it — so on real
saves the catalog and the schema agreed with each other and were both wrong. Its failure mode is
silent agreement, not a desync.
9. Rankers rank; they never filter
A cohort/frequency ranker would have discarded the correct answer: the function that appends to
ObservedTech touches only two offsets on ServerPlayer, so every --min>=1 cohort filter drops
it. What closed the case was a plain query plus one call-graph lookup.
Displacement scan for recall, call graph for disambiguation. Neither alone sufficed. And report the false-positive rate honestly — that scanner's class-level precision is ~13% by function; its value is a 900× search-space cut, not accuracy.
10. Derive experiment parameters from the ISA, not from assumed mnemonics
An x87 sensitivity experiment was briefed as 0x027f / 0x127f / 0x137f. In fact 0x027f is 53-bit
(differing from 0x127f only in infinity control, ignored since the 387) and 0x137f is 64-bit
extended. Run literally it returns "all three identical" — true, and the conclusion drawn from it
would have been wrong on both axes, because it tests single precision not at all and rounding not at
all. The real probes are 0x007f and 0x1a7f.
Brief lanes to challenge the parameters, not just execute them. Lane F did, and lane G rejected its brief's premise outright and was right to.
11. Correcting the record beats defending it
Claims downgraded on evidence, by later lanes: "RNG matched 15/15" (workload luck, not a property);
B1's and B4's coverage after the harness audit; StreamableEnum typing; the deferred end-of-turn
tail; "the research roll draws exactly one NextFloat" (the branch it fires draws a second).
If you find an earlier finding wrong, say so plainly and correct it in place. Nothing here is anyone's reputation.
12. Fix the oracle openly, never quietly
Four defects were found in the reference save reader. The right move was not to patch them mid-campaign in silence: they were recorded, then fixed openly with a test per defect and byte-neutrality proven at item granularity (offset sequences identical across ~39,000 offsets per save; typed-value deltas balancing to the byte).
13. Run verification gates as separate commands
Never &&-chain them. Once, a clean-room failure short-circuited the chain, skipping build and test
while a separate push still ran — an unverified tree reached main.
On one merge both gates failed and both were real: a raw FUN_xxxxxxxx identifier had reached
docs/ in the public-capable engine repo, and the shim cross-build hit -Werror=unused-function
on a WIN32-only TU. Host ctest passed cleanly both times and would have hidden both.
If a lane cannot cross-build its own shim TU, it must say so, and the integrator must run the CT111 shim build before pushing.
14. Never hand-resolve a generated header
Always regenerate from ghidra/addresses.json plus the ghidra/addresses.d/<lane>.json fragments.
Hand-resolving a merge conflict there silently dropped a new entry and broke the build — twice.
Per-lane fragments exist because a shared single file caused three cross-lane sweeps in one day. Duplicate names across fragments are a hard error, never last-wins: two lanes disagreeing about an address is exactly what must not be papered over.
15. Report thin coverage as loudly as divergences
A green run that establishes twenty facts is not a strong result. ComputeBudget compared 4,284
calls with 0 divergences — but only 20 distinct states, with 4,278 of them the UI polling one
player, and 13 of 22 slots zero on every call, including 5 of 6 declared input-boundary slots.
The most useful sections of the best reports here have been the honest lists of what was not covered.
16. Inlined draws are invisible to call-graph sweeps
The combat resolver's NextFloat is inlined, so the only call-graph edge it leaves is
caller -> RNG_Twist, which reads as a bare Twist and is not one — it is the lazy twist inside
NextFloat. Lane K's "no NextFloat in that subtree" was wrong for exactly this reason.
An image-wide scan for the MT tempering immediates at instruction boundaries finds 14 game
functions with inlined draws that no RNG call-graph sweep can see, two of them inside
ProcessTurn's closure at depth 4.
Scan for the constants at instruction boundaries, not for calls. Any RNG accounting built from the call graph alone is a lower bound.
17. A truncated range makes a loop look like a sequence
Lane J read a function as straight-line because it dumped the range at Ghidra's reported size, and the outer back-edge fell just outside it. A delegated sweep made the same mistake independently. That changed the draw count of the very site it was there to measure.
This is the same defect lane X hit from the other side: clipping sweeps at fva + sizeInBytes lost
11% of functions to mid-instruction truncation, and sweeping to the next function start took
coverage 89% -> 100%. Ghidra's size is also simply wrong sometimes — it gave 7,499 bytes for a body
that is 7,641 and ends mid-instruction.
Never trust a function's end. Disassemble to the next function start and find the real boundary.