13 KiB
mars::parse — the engine's data-file readers
src/mars/parse/ reimplements the two text readers the original engine uses
for its catalogs: the brace-block key/value format (weapons, ship
sections, the tech tree, combat missions, badge/light defs, GUI scripts,
scenario and other block-form .txt) and the separate .effect format.
C++17, no dependencies, no exceptions across the API, bytes passed through
untouched (the data is cp1252).
Behaviour is specified by the original engine's Mars::Script tokenizer as
recovered in the RE repo (findings/subsystems/loader-prototypes.md §M3) and
mirrored by the RE repo's Python readers (sots-re/verify/parsers/mars_data.py,
effect_txt.py), which serve as the test oracle. The C++ is written from that
understanding, not transliterated.
src/mars/parse/
result.h Diagnostic {line, message}; Result<T>
value.h/.cpp classify / as_int / as_double / as_bool / iequals
script.h header-only pull tokenizer with the engine's semantics
(read_token / next / skip_block); shared with mars::text
blocks.h/.cpp brace-block reader: parse_blocks(text, Options) -> Result<Document>
effect.h/.cpp .effect reader: parse_effect(text) -> Result<EffectFile>
CMakeLists.txt static library `mars_parse` (include root: src/)
tests/mars_parse/
build_and_run.sh plain g++ build; unit tests; real-data oracle test
test_*.cpp hand-written samples, one test per rule
canon.h/.cpp canonical JSON rendering (matches oracle/dump.py)
oracle_check.cpp parses every real file, diffs against the oracle
oracle/dump.py runs the RE-repo readers, emits canonical JSON per file
1. The tokenizer — Script (script.h)
The original has no tree. Every loader (WeaponDef::ParseScript,
SectionDef::ParseScript, MasterTechTree::ParseTech, the GUI script
callbacks, GlobalConsts::LoadFile) drives one pull tokenizer with three
operations and interprets the steps itself:
| operation | what it does |
|---|---|
read_token(Raw&) |
one token. Ok = read, input remains; AtEnd = read, the token touches the end of input; NoInput = nothing to read |
next(ScriptToken&) |
reads a key; if it is } the step is Close. Otherwise reads a value; Open if the value is {, else Pair. Returns the status of the value read |
skip_block(depth) |
consumes tokens counting whole {/} tokens until depth 0 |
Token rules:
| rule | detail |
|---|---|
| whitespace | space, tab, CR, LF — nothing else (\v, \f, NBSP are token bytes) |
| bareword | a run of non-whitespace. Braces are not delimiters: { and } count only as whole tokens, so weapon{ is one word and abc"def" is one word |
| quoted | a token starting with ", ' or a backtick ends at the next occurrence of that same character; the quotes are stripped, no escape processing, whitespace and newlines inside are kept; an unterminated quote runs to end of input. Text glued after a closing quote starts the next token |
| comment | a token whose extracted text starts with // — bare or quoted — is a comment: the rest of the line is skipped and reading continues. 3//x is one word; "0 0 0"// junk is the value 0 0 0 then a comment |
| length | text is capped at 1023 bytes (the whole token is still consumed) |
| brace test | exact compare on the stripped text, so a quoted "{" / "}" acts as a brace |
Script is public so a per-consumer loader can be written against the same
step stream when a catalog needs the original's own key handling.
2. Brace-block trees — parse_blocks
parse_blocks runs the loaders' while (next() == Ok) loop and records the
steps: Pair → Entry::Kind::Pair (the key may be quoted — key_quoted),
Open → Entry::Kind::Block with a recursive body, Close → the block ends.
The file's top level is itself a body: catalog files hold one or more named
blocks, scenario .txt files mix top-level pairs and player{} blocks.
There is no "item" concept: a quoted string in key position is a key
(systemnames.txt lists become pairs of consecutive names), and a lone word
before } takes the } as its value. That is what the original sees.
End of input and other tolerances
Everything the original merely tolerates is recorded as a Diagnostic in
Document::warnings; with Options::strict each one is an error instead.
| condition | what the original does (and this reader) |
|---|---|
| input ends inside a block | the block simply ends (one warning per open block) |
} at top level |
ignored |
final KEY value whose value token touches the end of input (no trailing newline) |
the step returns AtEnd and the pair is dropped |
final KEY with no value; final NAME { touching the end |
dropped |
| unterminated quote | runs to end of input (and, touching it, that pair is dropped) |
{ in key position |
an ordinary key; kept |
} that closes a block and touches the end of input |
a plain close — no warning; the original stops there either way and the outcome is identical |
The shipped data exercises: 12 shipsections (11 unclosed outer blocks, Human
CrPropaganda with one } too many) and Data/Strategy/systemnames.txt,
whose hiver and liir lists have an odd number of names — the last name
pairs with the }, so tarkas/liir nest inside hiver and morrigi
inside liir, and two blocks are left open at the end. 27 catalog files end
with a } as their last byte (plain close). No brace file drops a pair.
Node model
struct Entry { // one step in a block, in file order
enum class Kind { Pair, Block } kind;
std::string key; // as written (case preserved, quotes stripped)
std::string value; // Pair: value text
bool quoted; // value was quoted
bool key_quoted; // key was quoted
std::unique_ptr<Node> block; // Block only
int line; // 1-based, counts '\n'
};
struct Node { // a named block; the root has name ""
std::string name;
std::vector<Entry> entries;
int line;
const Entry* first(key); std::vector<const Entry*> all(key); // pairs + blocks
const std::string* first_value(key); std::vector<std::string_view> values(key);
const Node* first_block(key); std::vector<const Node*> blocks(key);
bool has(key);
};
struct Document { Node root; std::vector<Diagnostic> warnings; };
Result<Document> parse_blocks(std::string_view, Options = {});
Design points, each tied to a data fact:
- Keys are matched case-insensitively (
first/all/… useiequals, the original uses_stricmp), the original spelling is kept. The data mixesRequires/requires,badge/Badge. - Repeated keys are not merged:
all("requires")returns them in file order. 970requires, 3,721bank, and the 19 banks that repeatturretsizeall rely on order being kept. - Pairs and blocks share the key namespace:
first/allsee both because 153 shipsections carryoption DRV_Xbesideoption { option A option B }.values()/blocks()filter by kind when a consumer wants only one. - Values are raw text.
50,.5,7e+8,TRUE,@WEAPON_Xand"8"all stay as written, withquotedrecording the difference. Conversion is explicit viavalue.h:classify()reports Int / Float / Bool / Text using the C-float shapes the data uses;as_int/as_double/as_boolconvert. (The original usesatoi/atofper key, so quoting never changes typing there; the oracle's canonical form types only barewords, and consumers wanting engine typing should convert regardless ofquoted.)
3. .effect format — grammar as implemented
Line-based, not brace-block (a different engine class, Mars::TextFileStream).
line 1: TXT magic (surrounding whitespace ignored)
KEY value [value ...] scalar entry; values are "quoted" or bare
KEY followed (after blank/comment lines) by
BEGIN ... END a nested group
- Lines split on LF; a trailing CR is dropped (three shipped files are CRLF).
//starts a comment unless inside double quotes; blank lines are skipped; indentation is cosmetic.- Tokens:
"…"(may hold spaces —NAME "New Emitter") or a run of non-whitespace. A"that never closes on its line is just part of a bare token. - A line with a single token is a group key and must be followed by
BEGIN.BEGINwith no pending key,ENDwith no open group, a group key followed by anything butBEGIN, a trailing group key, and unclosed groups at EOF are all errors. No leniency has been added: all 415 shipped files are balanced, and there is no evidence yet of what the engine would do otherwise. - Order is preserved and repeated keys are kept as separate entries because it
is semantic:
PARTICLEDATATYPE nprecedes the curves that belong to it andMODIFIERrepeats once per type.
struct EffectValue { std::string text; bool quoted; };
struct EffectEntry {
std::string key; // as written (UPPERCASE in the data)
std::vector<EffectValue> values; // scalar line: 1+ values
std::unique_ptr<EffectGroup> group; // group line
int line;
bool is_group() const; std::string_view value() const; // first value
};
struct EffectGroup { std::vector<EffectEntry> entries; first(key); all(key); };
struct EffectFile { EffectGroup root; };
Result<EffectFile> parse_effect(std::string_view);
4. Tests
tests/mars_parse/build_and_run.sh (plain g++ -std=c++17 -Wall -Wextra -Wpedantic -Werror):
- Unit tests — 51 cases with hand-written samples: the
Scriptstatus codes,next()step shapes,skip_block, the 1023-byte cap; nesting, same-line blocks, braces glued to words, the three quote characters, quotes inside words, quoted braces, the four-character whitespace set, comments bare / glued / quoted / at end of input, repeated keys in order, case-insensitive lookup, quoted keys and odd-count lists, lone word before},optionscalar+block,@TOKENverbatim, EOF-ends-blocks, stray},}as last byte, dropped final pair / key / block header, unterminated quote,{as key — each in lenient and strict mode where it applies — CRLF and line numbers, multi-line strings, high bytes, and every effect error path. - Real-data oracle test — runs when
SOTS_DATA_DIRpoints at an extractedsots.gobtree.oracle/dump.pyruns the RE-repo readers (SOTS_RE_PARSERS, default~/sots-re/verify/parsers; exit 3 → the C++ side runs count-only) and writes a canonical JSON per file;oracle_checkparses every file withmars_parse, renders the same canonical form (canon.cpp) and requires byte equality, plus equal warning counts.
The canonical form is the oracle's own dict shape (keys lower-cased and
sorted, repeats → lists, bare pair values typed), with floats rendered by both
sides as "\x01" + printf("%.17g") so formatting is shared, and cp1252 bytes
rendered one \u00xx per byte. A one-character change in an oracle file is
detected (verified).
Skips cleanly with SOTS_DATA_DIR unset. Nothing from the game is committed;
tests/mars_parse/build/ (which holds the oracle JSON) is git-ignored.
Oracle results (2026-09-07, engine-parity rules, owner's gob-extract)
| kind | files | parsed | oracle match | tolerances taken |
|---|---|---|---|---|
| shipsection | 875 | 875 | 875 | 12 |
| weapon | 207 | 207 | 207 | 0 |
| effect | 415 | 415 | 415 | — |
| tech (MasterTechList) | 1 | 1 | 1 | 0 |
| combat | 3 | 3 | 3 | 0 |
| script | 4 | 4 | 4 | 0 |
| def | 2 | 2 | 2 | 0 |
| block-form txt (Scenarios + 7 Data/Models files) | 24 | 24 | 24 | 1 (systemnames.txt) |
| total | 1,531 | 1,531 | 1,531 (100%) | 13 |
Against the previous (tree-grammar) reading, exactly one file's canonical
output changed: Data/Strategy/systemnames.txt (quoted names become pairs;
odd-count lists nest the following blocks). Every weapon, shipsection, tech,
combat, def, script and scenario file renders byte-identically, so nothing the
cross-links or catalogs consume moved. The remaining 64 of the Python suite's
1,595 files are CSV / flat key-value / manifest / positional tables — see
docs/mars-text.md.
5. Open questions
systemnames.txtconsumer. The generic step stream pairs its names up; the original's system-name loader "handles that format itself" (RE notes) and its exact handling of the odd-count lists is not recovered. A dedicated reader overScriptwill be needed for name generation parity.skip_blockgranularity. Assumed to count raw{/}tokens (so a}in value position inside a skipped block still closes it, unlikenext()). No shipped file distinguishes the two.- Malformed
.effectnesting stays an error here because no shipped file exercises it; the original may be more forgiving.