sots-engine/docs/mars-parse.md

13 KiB

mars::parse — the engine's data-file readers

src/mars/parse/ reimplements the two text readers the original engine uses for its catalogs: the brace-block key/value format (weapons, ship sections, the tech tree, combat missions, badge/light defs, GUI scripts, scenario and other block-form .txt) and the separate .effect format. C++17, no dependencies, no exceptions across the API, bytes passed through untouched (the data is cp1252).

Behaviour is specified by the original engine's Mars::Script tokenizer as recovered in the RE repo (findings/subsystems/loader-prototypes.md §M3) and mirrored by the RE repo's Python readers (sots-re/verify/parsers/mars_data.py, effect_txt.py), which serve as the test oracle. The C++ is written from that understanding, not transliterated.

src/mars/parse/
  result.h      Diagnostic {line, message}; Result<T>
  value.h/.cpp  classify / as_int / as_double / as_bool / iequals
  script.h      header-only pull tokenizer with the engine's semantics
                (read_token / next / skip_block); shared with mars::text
  blocks.h/.cpp brace-block reader: parse_blocks(text, Options) -> Result<Document>
  effect.h/.cpp .effect reader:     parse_effect(text)          -> Result<EffectFile>
  CMakeLists.txt  static library `mars_parse` (include root: src/)
tests/mars_parse/
  build_and_run.sh   plain g++ build; unit tests; real-data oracle test
  test_*.cpp         hand-written samples, one test per rule
  canon.h/.cpp       canonical JSON rendering (matches oracle/dump.py)
  oracle_check.cpp   parses every real file, diffs against the oracle
  oracle/dump.py     runs the RE-repo readers, emits canonical JSON per file

1. The tokenizer — Script (script.h)

The original has no tree. Every loader (WeaponDef::ParseScript, SectionDef::ParseScript, MasterTechTree::ParseTech, the GUI script callbacks, GlobalConsts::LoadFile) drives one pull tokenizer with three operations and interprets the steps itself:

operation what it does
read_token(Raw&) one token. Ok = read, input remains; AtEnd = read, the token touches the end of input; NoInput = nothing to read
next(ScriptToken&) reads a key; if it is } the step is Close. Otherwise reads a value; Open if the value is {, else Pair. Returns the status of the value read
skip_block(depth) consumes tokens counting whole {/} tokens until depth 0

Token rules:

rule detail
whitespace space, tab, CR, LF — nothing else (\v, \f, NBSP are token bytes)
bareword a run of non-whitespace. Braces are not delimiters: { and } count only as whole tokens, so weapon{ is one word and abc"def" is one word
quoted a token starting with ", ' or a backtick ends at the next occurrence of that same character; the quotes are stripped, no escape processing, whitespace and newlines inside are kept; an unterminated quote runs to end of input. Text glued after a closing quote starts the next token
comment a token whose extracted text starts with // — bare or quoted — is a comment: the rest of the line is skipped and reading continues. 3//x is one word; "0 0 0"// junk is the value 0 0 0 then a comment
length text is capped at 1023 bytes (the whole token is still consumed)
brace test exact compare on the stripped text, so a quoted "{" / "}" acts as a brace

Script is public so a per-consumer loader can be written against the same step stream when a catalog needs the original's own key handling.

2. Brace-block trees — parse_blocks

parse_blocks runs the loaders' while (next() == Ok) loop and records the steps: Pair → Entry::Kind::Pair (the key may be quoted — key_quoted), Open → Entry::Kind::Block with a recursive body, Close → the block ends. The file's top level is itself a body: catalog files hold one or more named blocks, scenario .txt files mix top-level pairs and player{} blocks.

There is no "item" concept: a quoted string in key position is a key (systemnames.txt lists become pairs of consecutive names), and a lone word before } takes the } as its value. That is what the original sees.

End of input and other tolerances

Everything the original merely tolerates is recorded as a Diagnostic in Document::warnings; with Options::strict each one is an error instead.

condition what the original does (and this reader)
input ends inside a block the block simply ends (one warning per open block)
} at top level ignored
final KEY value whose value token touches the end of input (no trailing newline) the step returns AtEnd and the pair is dropped
final KEY with no value; final NAME { touching the end dropped
unterminated quote runs to end of input (and, touching it, that pair is dropped)
{ in key position an ordinary key; kept
} that closes a block and touches the end of input a plain close — no warning; the original stops there either way and the outcome is identical

The shipped data exercises: 12 shipsections (11 unclosed outer blocks, Human CrPropaganda with one } too many) and Data/Strategy/systemnames.txt, whose hiver and liir lists have an odd number of names — the last name pairs with the }, so tarkas/liir nest inside hiver and morrigi inside liir, and two blocks are left open at the end. 27 catalog files end with a } as their last byte (plain close). No brace file drops a pair.

Node model

struct Entry {                      // one step in a block, in file order
    enum class Kind { Pair, Block } kind;
    std::string key;                // as written (case preserved, quotes stripped)
    std::string value;              // Pair: value text
    bool quoted;                    // value was quoted
    bool key_quoted;                // key was quoted
    std::unique_ptr<Node> block;    // Block only
    int line;                       // 1-based, counts '\n'
};
struct Node {                       // a named block; the root has name ""
    std::string name;
    std::vector<Entry> entries;
    int line;
    const Entry* first(key);  std::vector<const Entry*> all(key);      // pairs + blocks
    const std::string* first_value(key);  std::vector<std::string_view> values(key);
    const Node* first_block(key);        std::vector<const Node*> blocks(key);
    bool has(key);
};
struct Document { Node root; std::vector<Diagnostic> warnings; };
Result<Document> parse_blocks(std::string_view, Options = {});

Design points, each tied to a data fact:

  • Keys are matched case-insensitively (first/all/… use iequals, the original uses _stricmp), the original spelling is kept. The data mixes Requires/requires, badge/Badge.
  • Repeated keys are not merged: all("requires") returns them in file order. 970 requires, 3,721 bank, and the 19 banks that repeat turretsize all rely on order being kept.
  • Pairs and blocks share the key namespace: first/all see both because 153 shipsections carry option DRV_X beside option { option A option B }. values()/blocks() filter by kind when a consumer wants only one.
  • Values are raw text. 50, .5, 7e+8, TRUE, @WEAPON_X and "8" all stay as written, with quoted recording the difference. Conversion is explicit via value.h: classify() reports Int / Float / Bool / Text using the C-float shapes the data uses; as_int/as_double/as_bool convert. (The original uses atoi/atof per key, so quoting never changes typing there; the oracle's canonical form types only barewords, and consumers wanting engine typing should convert regardless of quoted.)

3. .effect format — grammar as implemented

Line-based, not brace-block (a different engine class, Mars::TextFileStream).

line 1:  TXT                          magic (surrounding whitespace ignored)
         KEY value [value ...]        scalar entry; values are "quoted" or bare
         KEY                          followed (after blank/comment lines) by
             BEGIN ... END            a nested group
  • Lines split on LF; a trailing CR is dropped (three shipped files are CRLF).
  • // starts a comment unless inside double quotes; blank lines are skipped; indentation is cosmetic.
  • Tokens: "…" (may hold spaces — NAME "New Emitter") or a run of non-whitespace. A " that never closes on its line is just part of a bare token.
  • A line with a single token is a group key and must be followed by BEGIN. BEGIN with no pending key, END with no open group, a group key followed by anything but BEGIN, a trailing group key, and unclosed groups at EOF are all errors. No leniency has been added: all 415 shipped files are balanced, and there is no evidence yet of what the engine would do otherwise.
  • Order is preserved and repeated keys are kept as separate entries because it is semantic: PARTICLEDATATYPE n precedes the curves that belong to it and MODIFIER repeats once per type.
struct EffectValue { std::string text; bool quoted; };
struct EffectEntry {
    std::string key;                       // as written (UPPERCASE in the data)
    std::vector<EffectValue> values;       // scalar line: 1+ values
    std::unique_ptr<EffectGroup> group;    // group line
    int line;
    bool is_group() const;  std::string_view value() const;   // first value
};
struct EffectGroup { std::vector<EffectEntry> entries; first(key); all(key); };
struct EffectFile  { EffectGroup root; };
Result<EffectFile> parse_effect(std::string_view);

4. Tests

tests/mars_parse/build_and_run.sh (plain g++ -std=c++17 -Wall -Wextra -Wpedantic -Werror):

  1. Unit tests — 51 cases with hand-written samples: the Script status codes, next() step shapes, skip_block, the 1023-byte cap; nesting, same-line blocks, braces glued to words, the three quote characters, quotes inside words, quoted braces, the four-character whitespace set, comments bare / glued / quoted / at end of input, repeated keys in order, case-insensitive lookup, quoted keys and odd-count lists, lone word before }, option scalar+block, @TOKEN verbatim, EOF-ends-blocks, stray }, } as last byte, dropped final pair / key / block header, unterminated quote, { as key — each in lenient and strict mode where it applies — CRLF and line numbers, multi-line strings, high bytes, and every effect error path.
  2. Real-data oracle test — runs when SOTS_DATA_DIR points at an extracted sots.gob tree. oracle/dump.py runs the RE-repo readers (SOTS_RE_PARSERS, default ~/sots-re/verify/parsers; exit 3 → the C++ side runs count-only) and writes a canonical JSON per file; oracle_check parses every file with mars_parse, renders the same canonical form (canon.cpp) and requires byte equality, plus equal warning counts.

The canonical form is the oracle's own dict shape (keys lower-cased and sorted, repeats → lists, bare pair values typed), with floats rendered by both sides as "\x01" + printf("%.17g") so formatting is shared, and cp1252 bytes rendered one \u00xx per byte. A one-character change in an oracle file is detected (verified).

Skips cleanly with SOTS_DATA_DIR unset. Nothing from the game is committed; tests/mars_parse/build/ (which holds the oracle JSON) is git-ignored.

Oracle results (2026-09-07, engine-parity rules, owner's gob-extract)

kind files parsed oracle match tolerances taken
shipsection 875 875 875 12
weapon 207 207 207 0
effect 415 415 415 —
tech (MasterTechList) 1 1 1 0
combat 3 3 3 0
script 4 4 4 0
def 2 2 2 0
block-form txt (Scenarios + 7 Data/Models files) 24 24 24 1 (systemnames.txt)
total 1,531 1,531 1,531 (100%) 13

Against the previous (tree-grammar) reading, exactly one file's canonical output changed: Data/Strategy/systemnames.txt (quoted names become pairs; odd-count lists nest the following blocks). Every weapon, shipsection, tech, combat, def, script and scenario file renders byte-identically, so nothing the cross-links or catalogs consume moved. The remaining 64 of the Python suite's 1,595 files are CSV / flat key-value / manifest / positional tables — see docs/mars-text.md.

5. Open questions

  • systemnames.txt consumer. The generic step stream pairs its names up; the original's system-name loader "handles that format itself" (RE notes) and its exact handling of the odd-count lists is not recovered. A dedicated reader over Script will be needed for name generation parity.
  • skip_block granularity. Assumed to count raw {/} tokens (so a } in value position inside a skipped block still closes it, unlike next()). No shipped file distinguishes the two.
  • Malformed .effect nesting stays an error here because no shipped file exercises it; the original may be more forgiving.