.ngram binary format (v3)
A trained model serialized to disk: a portable binary encoding, gzipped. It is the cross-runtime contract — a fixed byte layout that Rust, PHP, and JS can all read with plain integer reads (no schema library, no reflection). This page summarizes the format faithfully; the normative specification is phony-core/FORMAT.md, and a port should read the two side by side.
The project is pre-release: there is no legacy format to support, and a reader rejects anything not starting with the magic.
Encoding conventions
| Token | Meaning |
|---|---|
uv | Unsigned LEB128 varint — little-endian base-128, 7 bits per byte, high bit = "more bytes follow". Small values take one byte. |
iv | Signed varint via zig-zag: (n << 1) ^ (n >> 31), then uv. |
u8 | A single raw byte (used only for token_type). |
str | A uv byte-length n, then n bytes of UTF-8 (no NUL terminator). |
dist | A weighted distribution: a uv count k, then k records of { id: uv, weight: uv }. |
lendist | A length distribution: a uv count k, then k records of { length: uv, weight: uv }. |
Every n-gram is interned once into a vocab and referred to by u32 id everywhere else, so the graph is compact.
Distribution ordering (normative)
Within a dist, edges are ordered by the target n-gram's lexicographic rank — which equals ascending id. A reader computes the cumulative weights from this order. This ordering is normative: sampling determinism across runtimes depends on it, combined with the portable SplitMix64 PRNG and cumulative-weight sampling.
Container
magic: 8 bytes = "PHNYNG03" (ASCII, uncompressed)
body: gzip streamThe 8-byte magic sits outside the gzip stream so a reader can dispatch without decompressing. Everything below is the decompressed body.
Body layout
order: uv
token_type: u8 (0 = char, 1 = word)
position_depth: uv
config: str (tokenizer config as JSON — provenance only)
metadata: str (metadata as JSON)
vocab_count: uv
vocab: vocab_count × str (sorted lexicographically; index = id)
first: dist (n-grams that start a word)
# graph, one record per id, in id order 0..vocab_count (id is implicit):
elements: vocab_count × { children: dist, last_children: dist }
positions_count: uv
positions: positions_count × { key: iv, d: dist }
paragraph_positions_count: uv
paragraph_positions: paragraph_positions_count × { key: iv, d: dist }
word_lengths: lendist
sentence_lengths: lendist
paragraph_lengths: lendist
originals_count: uv (number of unique training items)
originals: originals_count × { item: str, count: uv }Field notes
vocabis sorted lexicographically; a term's index is itsid. Every other structure refers to terms by that id.firstis the opener pool — the distribution over n-grams that start a word.elementsis the transition graph, one record per id in id order (the id is implicit from position).childrenis the normal continuation distribution;last_childrenthe distribution for word-final transitions.positions/paragraph_positionsare keyed positional openings (keyis a signed varint), used by the prose/texttoken type.word_lengths/sentence_lengths/paragraph_lengthsare learned length distributions.originalsare the training items, stored deduped as uniqueitem+count.real_wordsamples by count, reproducing the raw frequency distribution. This dedup is the single biggest size saving for prose models, and backsexclude_originals/real_word.
Rationale
- Why gzip. The vocab and originals are text and compress well. A future uncompressed variant will allow mmap zero-copy load (the graph arrays are already id-indexed for it).
- Varint weights. Edge weights and ids are varints, so a weight of
1(the overwhelmingly common case) is one byte — far smaller than fixed-widthu64. - Determinism. Sorted vocab, id-ordered ids, and ascending-id edges within every
dist, combined with the portable PRNG and cumulative-weight sampling, let any runtime reproduce identical output from a seed.
In memory
The same layout is the in-memory representation: a sorted vocab, a CSR graph (child_off / child_id / child_cum and the last_* arrays), and IdDist opener pools. The loader builds these directly from the bytes above, with no intermediate string-keyed structures.
See ngram-core for the engine that reads and samples this format, and the Model generator for how a model is invoked.