# `mars::text` — flat text readers `src/mars/text/` reimplements the three simple text formats the game reads next to the brace-block catalogs: flat `KEY value` tuning tables (and their positional-row cousins), the numbered id manifests, and the `#`/`//`-commented CSVs. Library target: `mars_text` (static, `src/mars/text/CMakeLists.txt`). Include as `mars/text/
.h`. C++17, no dependencies, no exceptions cross the API; every reader is total and reports oddities as `Problem`s. The behaviour spec is the RE repo's proven Python readers (`sots-re/verify/parsers/flat_kv.py`, `manifest.py`), which parse all 1,595 shipped data files. This module is checked against them file by file (below). ## API ```cpp namespace mars::text { // result.h struct Problem { enum class Kind {...}; Kind kind; int line; long long id; std::string message; }; template struct Result { T value; std::vector problems; bool ok() const; }; // value.h — lazily typed tokens enum class ValueKind { Int, Float, Bool, Bareword, String /*quoted*/ }; struct Token { std::string text; bool quoted; ValueKind kind() const; std::optional as_int() const; std::optional as_float() const; std::optional as_bool() const; }; struct Color { double r, g, b, a; int components; }; std::optional parse_color(std::string_view "r g b [a]"); // flat_kv.h struct KvEntry { int line; std::string key /*original spelling*/; std::vector values; }; class FlatKV { entries(); find(key) /*case-insensitive, last wins*/; has(key); get_int/get_float/get_bool/get_string/get_color(key); duplicates(); }; struct Row { int line; std::vector tokens; }; using Rows = std::vector; Result parse_flat_kv(std::string_view); Result parse_rows(std::string_view); std::string_view strip_comment(std::string_view line); std::vector split_tokens(std::string_view line); // manifest.h struct ManifestEntry { int line; int id; std::string name; }; class Manifest { entries(); deleted(); name_of(id); id_of(name) /*case-insensitive*/; is_deleted(id); }; Result parse_manifest(std::string_view); // csv.h struct CsvRow { int line; std::vector cells; }; struct Csv { std::optional> header; std::vector rows; }; Result parse_csv(std::string_view); // comments dropped, cells stripped Result> split_csv_records(std::string_view); // raw RFC-4180 split } ``` Input is the file's raw bytes (`std::string_view`). The files are cp1252 with mixed CRLF/LF; bytes ≥ 0x80 are passed through untouched and are never whitespace or case-folded. Whitespace is ASCII only (space, `\t`, `\r`, `\n`, `\v`, `\f`). ## Behaviour ### Flat `KEY value` tables (`parse_flat_kv`) and row tables (`parse_rows`) Per line: strip a trailing `//` comment (quote-aware: `//` inside `"..."` is text, and an unclosed `"` disables comment stripping for the rest of the line), trim, skip if empty, then tokenise: - `"..."` is one token (may be empty, may contain spaces/tabs/`//`); no escape processing, so `"Data\\x.txt"` keeps both backslashes. - Otherwise a run of non-whitespace is one bareword; a `"` that has no closing quote, or sits mid-word, is an ordinary character. `parse_flat_kv` takes the first token as the key and the remaining tokens as the value: none (key-only line), one (scalar) or several (a list — the reference reader returns a Python list here; none of the 20 shipped tables actually do this). `parse_rows` returns every line's token list. Typing is lazy: `Token::kind()` classifies a bareword exactly like the reference `coerce()`: `[+-]?\d+` → Int, `[+-]?(\d+\.\d*|\.\d+|\d+)([eE][+-]?\d+)?` → Float (`.5`, `-.8`, `1.`, `7e+8`), `true`/`false` any case → Bool, else Bareword (`0x10`, `inf`, `nan`, `1,5` are barewords). A quoted token is always a String — `"8"` never becomes a number. `as_float()` also accepts Int shapes; `as_int()` refuses values that do not fit in 64 bits. Keys: lookup is ASCII case-insensitive (`find("mars_default_color")` finds `MARS_DEFAULT_COLOR`); the entry keeps the original spelling and line. When a key repeats, `find()` returns the **last** line — the same "later line wins" default as the reference `parse_kv(on_dup='last')`. All entries stay in `entries()`; `duplicates()` lists repeated (case-folded) keys with their line numbers. The shipped tables contain no duplicates, exact or case-folded, so the case-folding in `duplicates()` is a deliberate superset of the reference's exact-key check. Colours are quoted `"r g b"` / `"r g b a"` strings (tabs inside are fine, as in `globals.txt`); `get_color` / `parse_color` require 3 or 4 numeric components. Problems: `UnbalancedQuote` (odd number of `"` on a line) is reported but the line is still parsed. None occur in shipped data. ### Id manifests (`parse_manifest`) Per trimmed line: blank → skip; a `// DELETED - ` marker anywhere on the line (word case-insensitive, whitespace optional) → tombstone in `deleted()`; other `//...` → comment; ` ` and nothing else → entry. Anything else is `Problem::Kind::Unrecognised` — including an entry followed by a trailing comment, matching the reference regex `^\s*(\d+)\s+(\S+)\s*$`. A repeated id is `DuplicateId` (the later entry is kept and wins lookups); an id that is both a tombstone and an entry is `DeletedAndAssigned`. `id_of()` folds case (`DEWar.SHIPSECTION` ↔ `dewar.shipsection`). ### CSV (`parse_csv`, `split_csv_records`) `split_csv_records` reproduces the reference (Python `csv`, excel dialect, non-strict) byte for byte on the shipped files: - `,` delimits; `\r\n`, `\n` and a lone `\r` each end a record; a blank line is an empty record `[]`; a trailing `,` yields a trailing empty cell; the last record needs no terminator. - A cell is quoted only if `"` is its **first byte**. Inside quotes `""` is a literal `"` and newline bytes are kept verbatim (`Strings.csv` has one such cell). Text after the closing quote is appended as-is (`"tail"junk` → `tailjunk`). - A `"` anywhere else is an ordinary byte: `EVENTMSG_DERELICT_UNRESOLVED"` keeps its quote, and ` "Effects/Blastoid1.effect"` (Asteroids.csv, quote after leading spaces) keeps both quotes. - End of input inside a quoted cell → `UnterminatedQuotedCell`, cell kept. `parse_csv` then drops comment rows (first cell, after leading whitespace, starts with `#` or `//`), empty rows and rows whose cells are all blank (`,,`), and strips every remaining cell. The header is the first `#` row with more than one cell that appears before any data row, with `#` and any `<`/`>` trimmed from each cell: `# ,` → `tech, human-pri`; `"# species","event"` → `species, event`; `#Key+A955,String,Size,Notes` → `Key+A955, ...`. A `#` row with a single cell is a plain comment and does not end the search. ## Oracle results `tests/mars_text/build_and_run.sh` compiles with `g++ -std=c++17 -Wall -Wextra -Werror`, runs the unit tests (165 checks, hand-written samples for every quirk above), the real-data facts test, and — when `SOTS_DATA_DIR` is set — dumps every file `mars_text` owns through both `dump_json` (C++) and `oracle/dump.py` (the Python readers) and compares them structurally, type-aware (`compare.py`). Run 2026-09-07 against the extracted `sots.gob` + `sots_local_en.gob` tree: | kind | files | agree | |---|---|---| | kv (`Data/**/*.txt`) | 20 | 20 | | rows (`_turrets.txt`, `_defaultweapons.txt`, `damfx*.txt`, `playercolors.txt`, `BadgeTable.txt`, `AvatarTable.txt`, `WeaponIconPlacements.txt`) | 8 | 8 | | manifest (`_weapons.txt`, 7 × `_shipsections.txt`) | 8 | 8 | | csv (AI tables, scenarios, `RealSpace.csv`, `SpriteTable.csv`, `Strings.csv`, `SpeechEvents.csv`, asteroids, music, sound_ui) | 28 | 28 | | **total** | **64** | **64 (100%)** | Compared per file: the full key → value map (exact keys, last wins, typed values), duplicate-key report, every row's typed tokens, manifest entries / tombstones / problems (kind, line, id), CSV header and every stripped cell. Real-data facts also asserted: `globals.txt` 364 keys, `MARS_DEFAULT_COLOR` = (255,177,39); `_turrets.txt` 42 × 8 tokens; `_weapons.txt` 123 ids, deleted 36/58/59; Human `_shipsections.txt` 145 ids with `dewar.shipsection` → 98; `Strings.csv` 5,722 raw records → 5,200 data rows → 5,196 distinct keys, exactly one multi-line raw cell; `aitechpri.csv` 0 rows with a 7-name schema. ## Deliberate divergences from the reference None affect shipped data (verified above); listed so nobody chases them: - Whitespace is ASCII-only. The Python readers operate on decoded `str`, so they would also treat cp1252 `0xA0` (NBSP) and bytes `0x1C`–`0x1F` as whitespace / line breaks. The relevant files contain no such bytes. - `duplicates()` folds key case; the reference compares exact keys. - Ints are 64-bit; the reference has unbounded ints. An overflowing bareword classifies as `Int` but `as_int()` returns nullopt. - The reference `csv` module rejects a NUL byte; this reader treats it as data. - `Problem` messages are our own wording; the oracle comparison matches on kind / line / id, not text. ## Open questions - The engine's own tolerance is unknown for the cases where the reference is strict (manifest entry with a trailing comment, a key that repeats). We copy the reference; both are absent from shipped data. - Multi-token `KEY a b c` values never occur in shipped tables; the API keeps them as a token list rather than guessing a type.