sots-engine/docs/mars-text.md

183 lines
9.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# `mars::text` — flat text readers
`src/mars/text/` reimplements the three simple text formats the game reads
next to the brace-block catalogs: flat `KEY value` tuning tables (and their
positional-row cousins), the numbered id manifests, and the `#`/`//`-commented
CSVs. Library target: `mars_text` (static, `src/mars/text/CMakeLists.txt`).
Include as `mars/text/<header>.h`. C++17, no dependencies, no exceptions cross
the API; every reader is total and reports oddities as `Problem`s.
The behaviour spec is the RE repo's proven Python readers
(`sots-re/verify/parsers/flat_kv.py`, `manifest.py`), which parse all 1,595
shipped data files. This module is checked against them file by file (below).
## API
```cpp
namespace mars::text {
// result.h
struct Problem { enum class Kind {...}; Kind kind; int line; long long id; std::string message; };
template <class T> struct Result { T value; std::vector<Problem> problems; bool ok() const; };
// value.h — lazily typed tokens
enum class ValueKind { Int, Float, Bool, Bareword, String /*quoted*/ };
struct Token { std::string text; bool quoted; ValueKind kind() const;
std::optional<int64_t> as_int() const; std::optional<double> as_float() const;
std::optional<bool> as_bool() const; };
struct Color { double r, g, b, a; int components; };
std::optional<Color> parse_color(std::string_view "r g b [a]");
// flat_kv.h
struct KvEntry { int line; std::string key /*original spelling*/; std::vector<Token> values; };
class FlatKV { entries(); find(key) /*case-insensitive, last wins*/; has(key);
get_int/get_float/get_bool/get_string/get_color(key); duplicates(); };
struct Row { int line; std::vector<Token> tokens; }; using Rows = std::vector<Row>;
Result<FlatKV> parse_flat_kv(std::string_view);
Result<Rows> parse_rows(std::string_view);
std::string_view strip_comment(std::string_view line); std::vector<Token> split_tokens(std::string_view line);
// manifest.h
struct ManifestEntry { int line; int id; std::string name; };
class Manifest { entries(); deleted(); name_of(id); id_of(name) /*case-insensitive*/; is_deleted(id); };
Result<Manifest> parse_manifest(std::string_view);
// csv.h
struct CsvRow { int line; std::vector<std::string> cells; };
struct Csv { std::optional<std::vector<std::string>> header; std::vector<CsvRow> rows; };
Result<Csv> parse_csv(std::string_view); // comments dropped, cells stripped
Result<std::vector<CsvRow>> split_csv_records(std::string_view); // raw RFC-4180 split
}
```
Input is the file's raw bytes (`std::string_view`). The files are cp1252 with
mixed CRLF/LF; bytes ≥ 0x80 are passed through untouched and are never
whitespace or case-folded. Whitespace is ASCII only (space, `\t`, `\r`, `\n`,
`\v`, `\f`).
## Behaviour
### Flat `KEY value` tables (`parse_flat_kv`) and row tables (`parse_rows`)
Per line: strip a trailing `//` comment (quote-aware: `//` inside `"..."` is
text, and an unclosed `"` disables comment stripping for the rest of the
line), trim, skip if empty, then tokenise:
- `"..."` is one token (may be empty, may contain spaces/tabs/`//`); no escape
processing, so `"Data\\x.txt"` keeps both backslashes.
- Otherwise a run of non-whitespace is one bareword; a `"` that has no
closing quote, or sits mid-word, is an ordinary character.
`parse_flat_kv` takes the first token as the key and the remaining tokens as
the value: none (key-only line), one (scalar) or several (a list — the
reference reader returns a Python list here; none of the 20 shipped tables
actually do this). `parse_rows` returns every line's token list.
Typing is lazy: `Token::kind()` classifies a bareword exactly like the
reference `coerce()`: `[+-]?\d+` → Int, `[+-]?(\d+\.\d*|\.\d+|\d+)([eE][+-]?\d+)?`
→ Float (`.5`, `-.8`, `1.`, `7e+8`), `true`/`false` any case → Bool, else
Bareword (`0x10`, `inf`, `nan`, `1,5` are barewords). A quoted token is always
a String — `"8"` never becomes a number. `as_float()` also accepts Int shapes;
`as_int()` refuses values that do not fit in 64 bits.
Keys: lookup is ASCII case-insensitive (`find("mars_default_color")` finds
`MARS_DEFAULT_COLOR`); the entry keeps the original spelling and line. When a
key repeats, `find()` returns the **last** line — the same "later line wins"
default as the reference `parse_kv(on_dup='last')`. All entries stay in
`entries()`; `duplicates()` lists repeated (case-folded) keys with their line
numbers. The shipped tables contain no duplicates, exact or case-folded, so
the case-folding in `duplicates()` is a deliberate superset of the reference's
exact-key check.
Colours are quoted `"r g b"` / `"r g b a"` strings (tabs inside are fine, as in
`globals.txt`); `get_color` / `parse_color` require 3 or 4 numeric components.
Problems: `UnbalancedQuote` (odd number of `"` on a line) is reported but the
line is still parsed. None occur in shipped data.
### Id manifests (`parse_manifest`)
Per trimmed line: blank → skip; a `// DELETED - <n>` marker anywhere on the
line (word case-insensitive, whitespace optional) → tombstone in `deleted()`;
other `//...` → comment; `<digits> <non-whitespace>` and nothing else → entry.
Anything else is `Problem::Kind::Unrecognised` — including an entry followed
by a trailing comment, matching the reference regex `^\s*(\d+)\s+(\S+)\s*$`.
A repeated id is `DuplicateId` (the later entry is kept and wins lookups); an
id that is both a tombstone and an entry is `DeletedAndAssigned`.
`id_of()` folds case (`DEWar.SHIPSECTION` ↔ `dewar.shipsection`).
### CSV (`parse_csv`, `split_csv_records`)
`split_csv_records` reproduces the reference (Python `csv`, excel dialect,
non-strict) byte for byte on the shipped files:
- `,` delimits; `\r\n`, `\n` and a lone `\r` each end a record; a blank line is
an empty record `[]`; a trailing `,` yields a trailing empty cell; the last
record needs no terminator.
- A cell is quoted only if `"` is its **first byte**. Inside quotes `""` is a
literal `"` and newline bytes are kept verbatim (`Strings.csv` has one such
cell). Text after the closing quote is appended as-is (`"tail"junk` →
`tailjunk`).
- A `"` anywhere else is an ordinary byte: `EVENTMSG_DERELICT_UNRESOLVED"`
keeps its quote, and ` "Effects/Blastoid1.effect"` (Asteroids.csv, quote
after leading spaces) keeps both quotes.
- End of input inside a quoted cell → `UnterminatedQuotedCell`, cell kept.
`parse_csv` then drops comment rows (first cell, after leading whitespace,
starts with `#` or `//`), empty rows and rows whose cells are all blank
(`,,`), and strips every remaining cell. The header is the first `#` row with
more than one cell that appears before any data row, with `#` and any `<`/`>`
trimmed from each cell: `# <tech>,<human-pri>` → `tech, human-pri`;
`"# species","event"` → `species, event`; `#Key+A955,String,Size,Notes` →
`Key+A955, ...`. A `#` row with a single cell is a plain comment and does not
end the search.
## Oracle results
`tests/mars_text/build_and_run.sh` compiles with `g++ -std=c++17 -Wall -Wextra
-Werror`, runs the unit tests (165 checks, hand-written samples for every
quirk above), the real-data facts test, and — when `SOTS_DATA_DIR` is set —
dumps every file `mars_text` owns through both `dump_json` (C++) and
`oracle/dump.py` (the Python readers) and compares them structurally,
type-aware (`compare.py`).
Run 2026-09-07 against the extracted `sots.gob` + `sots_local_en.gob` tree:
| kind | files | agree |
|---|---|---|
| kv (`Data/**/*.txt`) | 20 | 20 |
| rows (`_turrets.txt`, `_defaultweapons.txt`, `damfx*.txt`, `playercolors.txt`, `BadgeTable.txt`, `AvatarTable.txt`, `WeaponIconPlacements.txt`) | 8 | 8 |
| manifest (`_weapons.txt`, 7 × `_shipsections.txt`) | 8 | 8 |
| csv (AI tables, scenarios, `RealSpace.csv`, `SpriteTable.csv`, `Strings.csv`, `SpeechEvents.csv`, asteroids, music, sound_ui) | 28 | 28 |
| **total** | **64** | **64 (100%)** |
Compared per file: the full key → value map (exact keys, last wins, typed
values), duplicate-key report, every row's typed tokens, manifest entries /
tombstones / problems (kind, line, id), CSV header and every stripped cell.
Real-data facts also asserted: `globals.txt` 364 keys, `MARS_DEFAULT_COLOR` =
(255,177,39); `_turrets.txt` 42 × 8 tokens; `_weapons.txt` 123 ids, deleted
36/58/59; Human `_shipsections.txt` 145 ids with `dewar.shipsection` → 98;
`Strings.csv` 5,722 raw records → 5,200 data rows → 5,196 distinct keys,
exactly one multi-line raw cell; `aitechpri.csv` 0 rows with a 7-name schema.
## Deliberate divergences from the reference
None affect shipped data (verified above); listed so nobody chases them:
- Whitespace is ASCII-only. The Python readers operate on decoded `str`, so
they would also treat cp1252 `0xA0` (NBSP) and bytes `0x1C`–`0x1F` as
whitespace / line breaks. The relevant files contain no such bytes.
- `duplicates()` folds key case; the reference compares exact keys.
- Ints are 64-bit; the reference has unbounded ints. An overflowing bareword
classifies as `Int` but `as_int()` returns nullopt.
- The reference `csv` module rejects a NUL byte; this reader treats it as data.
- `Problem` messages are our own wording; the oracle comparison matches on
kind / line / id, not text.
## Open questions
- The engine's own tolerance is unknown for the cases where the reference is
strict (manifest entry with a trailing comment, a key that repeats). We copy
the reference; both are absent from shipped data.
- Multi-token `KEY a b c` values never occur in shipped tables; the API keeps
them as a token list rather than guessing a type.