Skip to content

Train your first model ​

In this tutorial you'll train your very first Phony model — a small n-gram model that has learned the shape of a set of names. You'll feed it 25 names, train in one command, and then poke at the result with three inspection commands. It takes about five minutes, and every command here is one you run yourself.

You don't need to know anything about n-grams yet — we'll explain that as we go.

Step 1 — Get the CLI ​

Phony's command-line tool is called phony. Build it once from the repository:

bash
cd phony-cli
cargo build

That produces the binary at target/debug/phony. To keep the commands below short, either put that binary on your PATH, or run it through Cargo — anywhere you see phony … you can instead type cargo run -- …. For example:

bash
cargo run -- --version
phony 0.0.1

The rest of this tutorial assumes you can type phony.

Step 2 — Prepare a tiny corpus ​

A model learns from examples. Ours will be a plain text file with one name per line. Create a file called names.txt:

text
Ada
Elif
Zeynep
Deniz
Mira
Ela
Lina
Nur
Aylin
Derin
Ceren
Sena
Yasemin
Melis
Ipek
Naz
Selin
Ece
Defne
Bade
Eylul
Irmak
Azra
Beren
Cansu

That's 25 names. Small on purpose — you can read the whole thing, and it trains instantly.

Step 3 — Train the model ​

One command:

bash
phony train names.txt -o names.ngram -n 3 -t char
✓ Trained char model (order 3) from 25 items
  → names.ngram (546 bytes)

That's it — you have a trained model in names.ngram. Let's unpack the flags:

  • names.txt — the input file (one item per line).
  • -o names.ngram — where to write the model. (--output also works.)
  • -n 3 — the n-gram order (--ngram-order). More on this in a moment.
  • -t char — the token type (--token-type). char means "learn the pattern of characters within a word" — exactly right for names and usernames. (The other choices are word for multi-word phrases and text for prose.)

What is an n-gram model, really? ​

Here's the whole idea in one sentence: an n-gram model records which short runs of characters tend to follow which.

With -n 3 (order 3), Phony looks at every run of 3 characters in your names and remembers what came next. It sees that after Ad comes a (from "Ada"), that names often start with a capital letter, that yn shows up (from "Zeynep"), and so on. It's not storing your names — it's storing the statistics of their shape: the little transitions that make a name feel Turkish rather than, say, Japanese.

Later, when you generate, Phony walks those transitions to build new names that follow the same patterns. A higher order (say -n 4) captures longer patterns and stays closer to the originals; a lower order (-n 2) is looser and more inventive. Order 3 is a good default.

You don't have to internalise this — just hold onto the intuition: the model learned the shape, not the list.

Step 4 — Look inside with info ​

Now inspect what you built:

bash
phony info names.ngram
Model: names.ngram
├── Token type    : Char
├── N-gram order  : 3
├── Trained items : 25
├── N-grams       : 57
├── Sentence stats: no
├── Paragraph stat: no
├── Sentence pos  : no
├── Paragraph pos : no
└── File size     : 546 bytes

Read that top-down: it's a char model of order 3, trained on 25 items, and it distilled them into 57 n-grams — the 57 distinct character-transitions it found worth remembering. The "sentence" and "paragraph" rows are all no because those only apply to text (prose) models; for names they're not used.

Step 5 — stats for the same numbers, plainly ​

stats shows the same core facts without the tree drawing — handy in scripts:

bash
phony stats names.ngram
Statistics: names.ngram
  Token type    : Char
  N-gram order  : 3
  Trained items : 25
  N-grams       : 57
  Sentence stats: no
  Paragraph stat: no
  Sentence pos  : no
  Paragraph pos : no

Step 6 — validate to check the file is sound ​

Finally, confirm the model file is well-formed and loadable:

bash
phony validate names.ngram
✓ Valid — Char model, order 3, 57 n-grams

validate is stricter than it looks: if a file is corrupt or isn't really a model, it says so and exits with a non-zero status (specifically exit code 4), so you can use it in CI. To see that in action, point it at a bogus file:

bash
echo "garbage" > bad.ngram
phony validate bad.ngram
✗ Invalid: not a .ngram file (bad magic)

("Bad magic" means the file doesn't start with the marker bytes every real .ngram begins with.)

What you learned ​

  • You built the phony CLI and ran your first commands.
  • You trained a char n-gram model from a one-name-per-line corpus with phony train, choosing the order (-n) and token type (-t).
  • An n-gram model stores the shape of your data — which character runs follow which — not the data itself.
  • You inspected it with info and stats, and confirmed it's sound with validate (which exits 4 on a bad file).

Curious how the model turns those statistics into novel values, and why it's all reproducible? That's the Model generator concept, and the deeper engine walkthrough.

Next ​

You have a model. Time to make it produce data.

→ Generate your first data

Phony Cloud — Documentation & Specification