Skip to content

.ngram binary format (v3) ​

A trained model serialized to disk: a portable binary encoding, gzipped. It is the cross-runtime contract — a fixed byte layout that Rust, PHP, and JS can all read with plain integer reads (no schema library, no reflection). This page summarizes the format faithfully; the normative specification is phony-core/FORMAT.md, and a port should read the two side by side.

The project is pre-release: there is no legacy format to support, and a reader rejects anything not starting with the magic.

Encoding conventions ​

TokenMeaning
uvUnsigned LEB128 varint — little-endian base-128, 7 bits per byte, high bit = "more bytes follow". Small values take one byte.
ivSigned varint via zig-zag: (n << 1) ^ (n >> 31), then uv.
u8A single raw byte (used only for token_type).
strA uv byte-length n, then n bytes of UTF-8 (no NUL terminator).
distA weighted distribution: a uv count k, then k records of { id: uv, weight: uv }.
lendistA length distribution: a uv count k, then k records of { length: uv, weight: uv }.

Every n-gram is interned once into a vocab and referred to by u32 id everywhere else, so the graph is compact.

Distribution ordering (normative) ​

Within a dist, edges are ordered by the target n-gram's lexicographic rank — which equals ascending id. A reader computes the cumulative weights from this order. This ordering is normative: sampling determinism across runtimes depends on it, combined with the portable SplitMix64 PRNG and cumulative-weight sampling.

Container ​

magic:  8 bytes = "PHNYNG03"   (ASCII, uncompressed)
body:   gzip stream

The 8-byte magic sits outside the gzip stream so a reader can dispatch without decompressing. Everything below is the decompressed body.

Body layout ​

order:           uv
token_type:      u8         (0 = char, 1 = word)
position_depth:  uv
config:          str        (tokenizer config as JSON — provenance only)
metadata:        str        (metadata as JSON)

vocab_count:     uv
vocab:           vocab_count × str     (sorted lexicographically; index = id)

first:           dist                  (n-grams that start a word)

# graph, one record per id, in id order 0..vocab_count (id is implicit):
elements:        vocab_count × { children: dist, last_children: dist }

positions_count:           uv
positions:                 positions_count × { key: iv, d: dist }
paragraph_positions_count: uv
paragraph_positions:       paragraph_positions_count × { key: iv, d: dist }

word_lengths:      lendist
sentence_lengths:  lendist
paragraph_lengths: lendist

originals_count:   uv                  (number of unique training items)
originals:         originals_count × { item: str, count: uv }

Field notes ​

  • vocab is sorted lexicographically; a term's index is its id. Every other structure refers to terms by that id.
  • first is the opener pool — the distribution over n-grams that start a word.
  • elements is the transition graph, one record per id in id order (the id is implicit from position). children is the normal continuation distribution; last_children the distribution for word-final transitions.
  • positions / paragraph_positions are keyed positional openings (key is a signed varint), used by the prose/text token type.
  • word_lengths / sentence_lengths / paragraph_lengths are learned length distributions.
  • originals are the training items, stored deduped as unique item + count. real_word samples by count, reproducing the raw frequency distribution. This dedup is the single biggest size saving for prose models, and backs exclude_originals / real_word.

Rationale ​

  • Why gzip. The vocab and originals are text and compress well. A future uncompressed variant will allow mmap zero-copy load (the graph arrays are already id-indexed for it).
  • Varint weights. Edge weights and ids are varints, so a weight of 1 (the overwhelmingly common case) is one byte — far smaller than fixed-width u64.
  • Determinism. Sorted vocab, id-ordered ids, and ascending-id edges within every dist, combined with the portable PRNG and cumulative-weight sampling, let any runtime reproduce identical output from a seed.

In memory ​

The same layout is the in-memory representation: a sorted vocab, a CSR graph (child_off / child_id / child_cum and the last_* arrays), and IdDist opener pools. The loader builds these directly from the bytes above, with no intermediate string-keyed structures.

See ngram-core for the engine that reads and samples this format, and the Model generator for how a model is invoked.

Phony Cloud — Documentation & Specification