Skip to content

ngram-core API ​

The N-gram statistical-text engine — the Model generator. It learns from sample data and produces new, plausible values that were (usually) never in the training set: fast, deterministic, no LLM. You use it directly to train the .ngram models that the model generator then consumes.

Token types ​

A model is one of three kinds, chosen by what you train it on:

KindConstructorTokenTrains onProducesUse
chartrain_wordsa charactera word listnew wordsnames, usernames, codes
wordtrain_phrasesa worda phrase listnew phrasescompany/product names
texttrain_texta word + structureprosesentences/paragraphsbios, reviews

How the model works ​

Each n-gram node keeps two distributions: children (n-grams that continue a word) and last_children (n-grams that end one). Generation samples a target length from the learned word_lengths, then terminates on a real word-ending n-gram — so words get natural lengths and endings without rejection sampling, and the generated length matches the sampled target exactly (shorter only when the graph runs out of continuations).

train_text additionally learns a length hierarchy (word_lengths → sentence_lengths → paragraph_lengths) and positional n-grams (which openings tend to start vs end a sentence, and which sentences open vs close a paragraph), so generated prose has realistic structure.

In memory the graph is an id-interned CSR layout (a sorted vocab + contiguous u32/u64 arrays), so the walk needs no per-step hashing — see The .ngram format.

Training ​

Convenience constructors ​

rust
let model = NgramModel::train_words(&["Mehmet", "Ayşe", "Mustafa"], 3); // order-3 char model
let model = NgramModel::train_phrases(&["Acme Corp", "Globex Inc"], 2);
let model = NgramModel::train_text(corpus, 3);
// `*_with(items, order, &TokenizerConfig) -> Result<NgramModel, String>` variants
// take an explicit tokenizer; they error if the config holds an invalid regex
// (word filter or a separator). The default-tokenizer forms above stay infallible.
let model = NgramModel::train_words_with(&["İstanbul", "İzmir"], 3, &TokenizerConfig::turkish())?;

Streaming / incremental training ​

For large or remote data, feed batches — the counts form a commutative monoid, so any batch order (or merged workers) yields a byte-identical model:

rust
let mut b = ModelBuilder::words(3, &TokenizerConfig::default())?; // Result — invalid regex errors
b.feed_items(&["mehmet", "ayse"]);
b.feed_items(&["mustafa"]);
let model = b.finalize();
// also: ModelBuilder::{phrases, text} (both -> Result), .min_count(n), .merge(other), .feed_text(&str)

ModelBuilder::{words, phrases, text} return Result<ModelBuilder, String> — like the *_with constructors, they fail on an invalid tokenizer regex.

.min_count(n) (a builder option) prunes transitions seen fewer than n times at finalize — and drops any n-gram string whose every reference was pruned from the vocab, so rare (identifying) fragments leave the model entirely. It suppresses verbatim leakage and shrinks the model.

Choosing an order (and small corpora) ​

The n-gram order trades structure against novelty. A higher order sees more context, so output looks more like the training set — but with few samples that becomes memorisation: the graph has so few paths that generation just replays whole training items. A lower order mixes fragments more freely (more novel, less faithful).

Rules of thumb for char models (names, usernames, codes):

  • Order 3 is the default and the right starting point. For a small corpus (below ~1k items) stay at order 2–3; jumping to order 4–5 on a small list mostly reproduces inputs verbatim.
  • Reach for order 4–5 only with a large, varied corpus, where the extra context sharpens realism without collapsing to memorisation.
  • Turn on privacy guards when the corpus is small or sensitive (see below) — they're what keep a thin model from leaking its inputs.

The privacy levers stack: ModelBuilder::min_count(n) prunes rare (identifying) transitions at training time — set it to 2+ on a small corpus to drop hapax fragments and shrink the model; Constraints::exclude_originals guarantees no whole training item is ever emitted, at generation time. The tradeoff: exclude_originals narrows the reachable output space (each rejected candidate costs a retry), so on a tiny corpus it can exhaust retries and, once the privacy guards are satisfied, fall back to the first guard-passing candidate (or an empty string if every candidate tripped a guard — it never leaks past an active guard). Bigger or higher-min_count models leave more headroom.

Generation ​

All generation takes an explicit &mut SplitMix64 (the portable PRNG), so it's deterministic given a seed.

rust
let mut rng = SplitMix64::new(2026);
let c = Constraints::default();

model.word(&mut rng, &c);                       // a single word
model.words(&mut rng, 5, &c);                   // Vec<String>
model.real_word(&mut rng, Some("A"));           // an actual training item (optional prefix)
model.sentence(&mut rng, Some(8), ".", None);   // num_words, punctuation, starts_with
model.text(&mut rng, 200, "…", None);           // max_chars, suffix, starts_with
model.paragraph(&mut rng, Some(4), None);       // num_sentences, starts_with
model.poem(&mut rng, 4, 4, 8);                  // verses, stanza_length, max_words
model.acrostic(&mut rng, "PHONY", 8);           // initials, max_words
model.try_acrostic(&mut rng, "PHONY", 8)?;      // Err if an initial is unsatisfiable
ModeMethodOutput
wordword / wordsa novel word
sentencesentence / sentencesa sentence
texttexttruncated prose
paragraphparagrapha paragraph
poempoemstanzas
acrosticacrostic / try_acrostican acrostic — line k starts with initials[k]
real_wordreal_wordan actual training item (frequency-weighted)
manygenerate_manya Vec<String>; seeds its own RNG from a seed + count

If the model has no word starting with an initial, acrostic falls back to the bare (uppercased) initial itself — the acrostic property always holds — while try_acrostic returns an Err naming the unsatisfiable initial instead.

Constraints (per-call filters) ​

rust
struct Constraints {
    starts_with: Option<String>,
    ends_with: Option<String>,
    contains: Option<String>,
    min_length: Option<usize>,
    max_length: Option<usize>,
    exclude_originals: bool, // never emit a value that was in the training set
}

Constraints are enforced with up to 64 retries. On retry exhaustion, word returns the first candidate that at least passed the privacy guards (exclude_originals and the verbatim-run guard), even if a soft constraint (length / ends_with / contains) failed; if every candidate violated a privacy guard it returns an empty string — it never leaks training data past an active guard.

Sampling policy (set on the model) ​

rust
model.sampling = Sampling { temperature: 1.0, top_k: None, top_p: None, no_repeat: false };

temperature == 1.0 (the default) keeps the exact-integer, cross-runtime reproducible path; other values trade strict reproducibility for diversity. no_repeat resamples (bounded retries) when a word would repeat its predecessor, so sentences keep their sampled word count; if retries exhaust (e.g. a one-word vocabulary) the repeat is kept rather than shortening the sentence.

Anti-memorisation ​

rust
model.set_verbatim_guard(Some(8)); // reject output sharing a >8-char run with training

Determinism & keyed mapping ​

rust
model.generate_unique(seed, count, &c);  // distinct values (for unique/PK columns)
model.word_for(key, &c);                 // same key → same value (referential integrity)

word_for uses a portable FNV-1a hash, so a foreign key anonymised this way stays joinable.

Privacy & anonymisation ​

Phony's job is plausible data, not leaked data. The tools are layered and all opt-in:

ToolLayerEffect
ModelBuilder::min_count(n)trainingprune rare (identifying) transitions at the source
model.set_verbatim_guard(k)generationreject output sharing a > k-char verbatim run with training
Constraints::exclude_originalsgenerationnever reproduce a whole training item
model.word_for(key, …)generationdeterministic keyed mapping → referential integrity
model.generate_unique(…)generationdistinct values for unique / primary-key columns

Tokenization ​

N-grams are counted over tokens, and how text splits into tokens is configurable and stored in the model — so every runtime reproduces the same tokenization:

text ─▶ [paragraph split] ─▶ [sentence split] ─▶ words ─▶ [unit split] ─▶ units
rust
TokenizerConfig {
    word_separator, sentence_separator, paragraph_separator, // how text is split
    word_filter,                  // chars to drop from each word
    min_word_length, to_lowercase,
    locale,                       // provenance tag
};
// presets:
TokenizerConfig::turkish()       // keeps ç ğ ı ö ş ü, drops noise
TokenizerConfig::alphabetical()  // ASCII alphanumerics only
TokenizerConfig::for_locale("tr_TR")

Presets run as fast char predicates (not regex) on the hot path. Pass one to any *_with constructor or to ModelBuilder::{words, phrases, text} — both are fallible and reject a config whose word filter or separators aren't valid regexes.

Portable casing. When to_lowercase is set, lowercasing maps İ (U+0130) to a plain i — a single char, no combining dot — before Unicode lowercasing. This is part of the cross-runtime contract: every port must implement the same mapping so tokenization (and therefore models and output) stays byte-identical.

Inspecting a model ​

rust
model.order();          // the n-gram size
model.token_type();     // char | word | text
model.ngram_count();    // distinct n-grams
model.tokenizer();      // the stored TokenizerConfig
model.metadata();       // provenance
// capability probes the generation modes branch on:
model.has_positions();
model.has_sentence_lengths();
model.has_paragraph_lengths();
model.has_paragraph_positions();

The .ngram format ​

A .ngram file is a portable binary (magic PHNYNG03, gzipped): a documented little-endian, LEB128-varint layout that Rust, PHP, and JS can all read with plain integer reads. This is the cross-runtime contract.

rust
model.save("names.ngram")?;
let model = NgramModel::load("names.ngram")?;
let bytes = model.to_bytes()?;                  // in-memory
let model = NgramModel::from_bytes(&bytes)?;    // Err on anything malformed — never panics

from_bytes is hardened against corrupt or hostile input: bad magic, truncated bodies, non-canonical varints, out-of-range token ids, and absurd block counts all return Err (never a panic or a huge allocation), so untrusted .ngram files are safe to load.

PRNG ​

rust
let mut rng = SplitMix64::new(seed);
let raw = rng.next_u64();  // the raw output stream (pub — part of the portable contract)
let n = rng.below(100);    // uniform in [0, 100) — next_u64() % n

The PRNG, the format, and inverse-CDF sampling are all part of the portable spec so other runtimes reproduce identical output. The frozen known-answer vectors — PRNG streams, hashes, and a full train → serialize → generate pipeline — live in the conformance corpus at phony-core/conformance/vectors.json.

Phony Cloud — Documentation & Specification