223 lines
12 KiB
Markdown
223 lines
12 KiB
Markdown
# `mars::text` — flat text readers
|
||
|
||
`src/mars/text/` reimplements the three simple text formats the game reads
|
||
next to the brace-block catalogs: flat `KEY value` tuning tables (and their
|
||
positional-row cousins), the numbered id manifests, and the `#`/`//`-commented
|
||
CSVs. Library target: `mars_text` (static, `src/mars/text/CMakeLists.txt`).
|
||
Include as `mars/text/<header>.h`. C++17, no dependencies, no exceptions cross
|
||
the API; every reader is total and reports oddities as `Problem`s.
|
||
|
||
The behaviour spec for the `KEY value` tables is the original engine's
|
||
`GlobalConsts` loader as recovered in the RE repo
|
||
(`findings/subsystems/loader-prototypes.md`, §M1 with the tokenizer of §M3);
|
||
the RE repo's Python readers (`sots-re/verify/parsers/flat_kv.py`,
|
||
`manifest.py`) implement the same rules and serve as the oracle this module
|
||
is checked against file by file (below).
|
||
|
||
## API
|
||
|
||
```cpp
|
||
namespace mars::text {
|
||
|
||
// result.h
|
||
struct Problem { enum class Kind {...}; Kind kind; int line; long long id; std::string message; };
|
||
template <class T> struct Result { T value; std::vector<Problem> problems; bool ok() const; };
|
||
|
||
// value.h — lazily typed tokens
|
||
enum class ValueKind { Int, Float, Bool, Bareword, String /*quoted*/ };
|
||
struct Token { std::string text; bool quoted; ValueKind kind() const;
|
||
std::optional<int64_t> as_int() const; std::optional<double> as_float() const;
|
||
std::optional<bool> as_bool() const; };
|
||
struct Color { double r, g, b, a; int components; };
|
||
std::optional<Color> parse_color(std::string_view "r g b [a]");
|
||
|
||
// flat_kv.h
|
||
struct KvEntry { int line; std::string key /*original spelling*/; Token value; };
|
||
class FlatKV { entries(); find(key) /*case-insensitive, FIRST wins*/; has(key);
|
||
get_int/get_float/get_bool/get_string/get_color(key); duplicates(); };
|
||
struct Row { int line; std::vector<Token> tokens; }; using Rows = std::vector<Row>;
|
||
Result<FlatKV> parse_flat_kv(std::string_view); // engine loader loop (mars/parse/script.h)
|
||
Result<Rows> parse_rows(std::string_view); // line-based
|
||
std::string_view strip_comment(std::string_view line); std::vector<Token> split_tokens(std::string_view line);
|
||
|
||
// manifest.h
|
||
struct ManifestEntry { int line; int id; std::string name; };
|
||
class Manifest { entries(); deleted(); name_of(id); id_of(name) /*case-insensitive*/; is_deleted(id); };
|
||
Result<Manifest> parse_manifest(std::string_view);
|
||
|
||
// csv.h
|
||
struct CsvRow { int line; std::vector<std::string> cells; };
|
||
struct Csv { std::optional<std::vector<std::string>> header; std::vector<CsvRow> rows; };
|
||
Result<Csv> parse_csv(std::string_view); // comments dropped, cells stripped
|
||
Result<std::vector<CsvRow>> split_csv_records(std::string_view); // raw RFC-4180 split
|
||
}
|
||
```
|
||
|
||
Input is the file's raw bytes (`std::string_view`). The files are cp1252 with
|
||
mixed CRLF/LF; bytes ≥ 0x80 are passed through untouched and are never
|
||
whitespace or case-folded.
|
||
|
||
## Behaviour
|
||
|
||
### Flat `KEY value` tables (`parse_flat_kv`)
|
||
|
||
This is the engine's `GlobalConsts::LoadFile` loop, verbatim in shape. The
|
||
file is **not** read by lines: it is stepped through with the same pull
|
||
tokenizer the brace-block catalogs use (`mars/parse/script.h`,
|
||
`Script::next()`), and each step is one of
|
||
|
||
| step | what the loader does |
|
||
|---|---|
|
||
| `KEY value` | one token each. `value` is consumed by the key's registered parser; a colour must therefore be quoted (`"48 29 2"`), and `LIST 1 2 three` is the two pairs `LIST=1`, `2=three` |
|
||
| `NAME {` | the block is skipped to its matching `}` (`Problem::Kind::SkippedBlock`) |
|
||
| `}` | ignored |
|
||
|
||
Tokenizer rules (script.h): whitespace is space/tab/CR/LF only; a bareword
|
||
runs to whitespace, so braces glued to a word are part of it; `"`, `'` and
|
||
backtick open a quoted token that ends at the same character, no escapes,
|
||
quotes stripped, an unterminated quote runs to end of input
|
||
(`UnterminatedQuote`); a token whose text starts with `//` is a comment to
|
||
end of line (so `3//x` is one word, `"0 0 0"// junk` is the value `0 0 0`
|
||
followed by a comment); tokens are capped at 1023 bytes.
|
||
|
||
Keys are matched **case-insensitively** and the **first occurrence wins**:
|
||
the loader erases a key from its expected-set once consumed, so a later
|
||
duplicate is logged "multiply defined" and ignored. `find()` returns the first
|
||
entry; every pair stays in `entries()` in file order; the later line is
|
||
reported as `DuplicateKey`; `duplicates()` lists repeated (case-folded) keys
|
||
with their line numbers. Unknown keys are the caller's business (the engine
|
||
ignores them); a registered key absent from the file keeps its default.
|
||
|
||
**End of input**: the loader stops at the first step that does not complete.
|
||
A final `KEY value` whose value token touches the end of the file — no
|
||
trailing newline or space — is therefore **dropped** (`DroppedTrailingPair`),
|
||
as is a final key with no value. Two shipped files lose a key this way:
|
||
`Data/Strategy/StrategyVars.txt` (`CIVILIAN_BURDEN_RATIO 0.5`) and
|
||
`Data/encounters.txt` (`HERALD_SPEECH_MAX_INVERVAL 45`) — the game runs on
|
||
those keys' compiled-in defaults.
|
||
|
||
Typing is lazy: `Token::kind()` classifies a bareword exactly like the
|
||
reference `coerce()`: `[+-]?\d+` → Int, `[+-]?(\d+\.\d*|\.\d+|\d+)([eE][+-]?\d+)?`
|
||
→ Float (`.5`, `-.8`, `1.`, `7e+8`), `true`/`false` any case → Bool, else
|
||
Bareword (`0x10`, `inf`, `nan`, `1,5` are barewords). A quoted token is always
|
||
a String — `"8"` never becomes a number. `as_float()` also accepts Int shapes;
|
||
`as_int()` refuses values that do not fit in 64 bits. (The engine types by the
|
||
key's registered parser — `sscanf` `%d`/`%f` — not by the text; see the RE
|
||
notes. `Token` keeps the text so a consumer can apply the registered parser.)
|
||
|
||
Colours are quoted `"r g b"` / `"r g b a"` strings (tabs inside are fine, as in
|
||
`globals.txt`); `get_color` / `parse_color` require 3 or 4 numeric components.
|
||
|
||
`KvEntry::line` counts `\n` only (the engine has no notion of lines; the
|
||
number is for diagnostics).
|
||
|
||
### Row tables (`parse_rows`)
|
||
|
||
Line-based (these files are read by other engine code whose reader has not
|
||
been recovered; the behaviour is the reference reader's). Per line: strip a
|
||
trailing `//` comment (quote-aware: `//` inside `"..."` is text, and an
|
||
unclosed `"` disables comment stripping for the rest of the line), trim, skip
|
||
if empty, then tokenise: `"..."` is one token (may be empty, may contain
|
||
spaces/tabs/`//`; no escapes), otherwise a run of non-whitespace is one
|
||
bareword; a `"` that has no closing quote, or sits mid-word, is an ordinary
|
||
character. Whitespace here is the C set (space, `\t`, `\r`, `\n`, `\v`, `\f`).
|
||
`UnbalancedQuote` (odd number of `"` on a line) is reported but the line is
|
||
still parsed. None occur in shipped data.
|
||
|
||
### Id manifests (`parse_manifest`)
|
||
|
||
Per trimmed line: blank → skip; a `// DELETED - <n>` marker anywhere on the
|
||
line (word case-insensitive, whitespace optional) → tombstone in `deleted()`;
|
||
other `//...` → comment; `<digits> <non-whitespace>` and nothing else → entry.
|
||
Anything else is `Problem::Kind::Unrecognised` — including an entry followed
|
||
by a trailing comment, matching the reference regex `^\s*(\d+)\s+(\S+)\s*$`.
|
||
A repeated id is `DuplicateId` (the later entry is kept and wins lookups); an
|
||
id that is both a tombstone and an entry is `DeletedAndAssigned`.
|
||
`id_of()` folds case (`DEWar.SHIPSECTION` ↔ `dewar.shipsection`).
|
||
|
||
### CSV (`parse_csv`, `split_csv_records`)
|
||
|
||
`split_csv_records` reproduces the reference (Python `csv`, excel dialect,
|
||
non-strict) byte for byte on the shipped files:
|
||
|
||
- `,` delimits; `\r\n`, `\n` and a lone `\r` each end a record; a blank line is
|
||
an empty record `[]`; a trailing `,` yields a trailing empty cell; the last
|
||
record needs no terminator.
|
||
- A cell is quoted only if `"` is its **first byte**. Inside quotes `""` is a
|
||
literal `"` and newline bytes are kept verbatim (`Strings.csv` has one such
|
||
cell). Text after the closing quote is appended as-is (`"tail"junk` →
|
||
`tailjunk`).
|
||
- A `"` anywhere else is an ordinary byte: `EVENTMSG_DERELICT_UNRESOLVED"`
|
||
keeps its quote, and ` "Effects/Blastoid1.effect"` (Asteroids.csv, quote
|
||
after leading spaces) keeps both quotes.
|
||
- End of input inside a quoted cell → `UnterminatedQuotedCell`, cell kept.
|
||
|
||
`parse_csv` then drops comment rows (first cell, after leading whitespace,
|
||
starts with `#` or `//`), empty rows and rows whose cells are all blank
|
||
(`,,`), and strips every remaining cell. The header is the first `#` row with
|
||
more than one cell that appears before any data row, with `#` and any `<`/`>`
|
||
trimmed from each cell: `# <tech>,<human-pri>` → `tech, human-pri`;
|
||
`"# species","event"` → `species, event`; `#Key+A955,String,Size,Notes` →
|
||
`Key+A955, ...`. A `#` row with a single cell is a plain comment and does not
|
||
end the search.
|
||
|
||
## Oracle results
|
||
|
||
`tests/mars_text/build_and_run.sh` compiles with `g++ -std=c++17 -Wall -Wextra
|
||
-Werror`, runs the unit tests (201 checks, hand-written samples for every
|
||
rule above, including a dedicated engine-step suite for the kv loader), the
|
||
real-data facts test, and — when `SOTS_DATA_DIR` is set — dumps every file
|
||
`mars_text` owns through both `dump_json` (C++) and `oracle/dump.py` (the
|
||
Python readers) and compares them structurally, type-aware (`compare.py`).
|
||
|
||
Run 2026-09-07 (engine-parity rules) against the extracted `sots.gob` +
|
||
`sots_local_en.gob` tree:
|
||
|
||
| kind | files | agree |
|
||
|---|---|---|
|
||
| kv (`Data/**/*.txt`) | 20 | 20 |
|
||
| rows (`_turrets.txt`, `_defaultweapons.txt`, `damfx*.txt`, `playercolors.txt`, `BadgeTable.txt`, `AvatarTable.txt`, `WeaponIconPlacements.txt`) | 8 | 8 |
|
||
| manifest (`_weapons.txt`, 7 × `_shipsections.txt`) | 8 | 8 |
|
||
| csv (AI tables, scenarios, `RealSpace.csv`, `SpriteTable.csv`, `Strings.csv`, `SpeechEvents.csv`, asteroids, music, sound_ui) | 28 | 28 |
|
||
| **total** | **64** | **64 (100%)** |
|
||
|
||
Compared per file: the full key → value map (case-insensitive keys under
|
||
their first spelling, first wins, typed values), duplicate-key report, every
|
||
row's typed tokens, manifest entries / tombstones / problems (kind, line, id),
|
||
CSV header and every stripped cell. Real-data facts also asserted:
|
||
`globals.txt` 364 keys, no duplicates, `MARS_DEFAULT_COLOR` = (255,177,39);
|
||
`StrategyVars.txt` 96 keys with `CIVILIAN_BURDEN_RATIO` dropped and
|
||
`encounters.txt` with `HERALD_SPEECH_MAX_INVERVAL` dropped (one
|
||
`DroppedTrailingPair` each); `_turrets.txt` 42 × 8 tokens; `_weapons.txt` 123
|
||
ids, deleted 36/58/59; Human `_shipsections.txt` 145 ids with
|
||
`dewar.shipsection` → 98; `Strings.csv` 5,722 raw records → 5,200 data rows →
|
||
5,196 distinct keys, exactly one multi-line raw cell; `aitechpri.csv` 0 rows
|
||
with a 7-name schema.
|
||
|
||
What changed when the kv reader moved from the earlier line-based, last-wins
|
||
reading to the engine's rules: no shipped kv file contains a duplicate key
|
||
(exact or case-folded), so first-wins flips no value; the only observable
|
||
difference is the two dropped trailing pairs above.
|
||
|
||
## Deliberate divergences from the reference
|
||
|
||
None affect shipped data (verified above); listed so nobody chases them:
|
||
|
||
- Row/manifest/CSV whitespace is ASCII-only. The Python readers operate on
|
||
decoded `str`, so they would also treat cp1252 `0xA0` (NBSP) and bytes
|
||
`0x1C`–`0x1F` as whitespace / line breaks. The relevant files contain no
|
||
such bytes. (The kv reader and the Python `Script` both use the engine's
|
||
exact four-character set.)
|
||
- Ints are 64-bit; the reference has unbounded ints. An overflowing bareword
|
||
classifies as `Int` but `as_int()` returns nullopt.
|
||
- The reference `csv` module rejects a NUL byte; this reader treats it as data.
|
||
- `Problem` messages are our own wording; the oracle comparison matches on
|
||
kind / line / id, not text.
|
||
|
||
## Open questions
|
||
|
||
- The `SkipBlock` used for `NAME {` inside a kv file is assumed to count raw
|
||
`{`/`}` tokens (`Script::skip_block`); no shipped kv file contains a block,
|
||
so this cannot be observed from data.
|
||
- The engine's own tolerance is unknown for a manifest entry with a trailing
|
||
comment; we copy the reference (absent from shipped data).
|