Generate your first data
You have a trained model from the previous tutorial. Now you'll make Phony actually produce data. You'll write three small generator definitions — a logic generator, a list generator, and a model generator that uses your names.ngram — run each one, and watch the same seed reproduce the same values every time.
Everything here is a phony generate command against a little JSON file. Let's go.
What is a generator?
A generator is a named producer of values. You don't write code that loops and appends; you write a small JSON definition that describes the value you want, and Phony produces it — reproducibly, from a seed. Phony ships with a handful of built-in generators; the three you'll meet here are the ones you'll reach for most.
The definitions you write are PGDL — the Phony Generator Definition Language. It's just JSON.
Step 1 — A logic generator (a dice roll)
The simplest generator: pure arithmetic, no data, no model. Put this in dice.json:
{
"type": "logic",
"algorithm": "int_between",
"params": { "min": 1, "max": 6 }
}Read it aloud: "a logic generator using the int_between algorithm, between 1 and 6." Now run it:
phony generate dice.json --seed 42 --count 5[
2,
2,
3,
2,
2
]Five dice rolls. The two flags you'll use constantly:
--seed 42— the root seed. It's the single knob that decides the output.--count 5— how many values to produce. (-n 5is the short form.)
By default you get a pretty-printed JSON array.
Step 2 — See determinism for yourself
This is the property that makes Phony Phony. Run the exact same command again:
phony generate dice.json --seed 42 --count 5[
2,
2,
3,
2,
2
]Identical. Same seed, same definition → same data, every time, on any machine. Now change only the seed:
phony generate dice.json --seed 7 --count 5 -f jsonl2
4
6
2
4Different seed, different rolls. (-f jsonl switched the output to JSONL — one value per line — which is friendlier for piping into other tools. The default, -f json, is the pretty array you saw above.)
That reproducibility isn't luck; it's the core design. If you want the full story, see Determinism.
Step 3 — A list generator (pick from a set)
Logic invents numbers. A list generator picks from a set of values you supply — and you can weight the choices. Put this in plan.json:
{
"type": "list",
"source": "inline",
"values": [
{ "value": "free", "weight": 60 },
{ "value": "pro", "weight": 30 },
{ "value": "enterprise", "weight": 10 }
]
}source: "inline" means the values are right here in the file. The weights are relative: free is picked roughly six times as often as enterprise. Run it:
phony generate plan.json --seed 1 --count 6 -f jsonl"pro"
"free"
"free"
"pro"
"pro"
"enterprise"Six plans, weighted toward the cheaper tiers — and, being a generator, perfectly reproducible: run it again with --seed 1 and you'll get that same sequence.
Step 4 — A model generator (use your trained model)
Now the payoff from tutorial 1. A model generator produces novel values from a trained .ngram. But a model generator refers to its model by an asset name, not a file path — so first you register your model as an asset by dropping it in a tiny package.
A package is just a directory with a phony.json. Create one:
mkdir -p namespkg/models
cp names.ngram namespkg/models/names.ngramThen write namespkg/phony.json:
{
"name": "@me/names",
"version": "1.0.0",
"assets": {
"models": {
"@me/names:first_names": { "*": "models/names.ngram" }
}
}
}That says: "this package publishes a model asset named @me/names:first_names, and for any locale (*) it lives at models/names.ngram." (Packages are how Phony ships reusable data; you'll learn more in Packages.)
Now the generator definition, firstname.json, points at that asset name:
{
"type": "model",
"source": "@me/names:first_names",
"generation": { "mode": "word" }
}Run it, telling Phony to load the package with --package:
phony generate firstname.json --package namespkg --seed 3 --count 5 -f jsonl✓ Loaded @me/names v1.0.0 — 0 generator(s), 1 asset(s)
"Bade"
"Ada"
"Zeynep"
"Eylul"
"Mira"The first line (on stderr) is Phony confirming it loaded your package and found one asset. Then five names, generated from the shape your model learned. With a corpus this tiny the model often echoes names it saw — but feed it a few hundred and it starts inventing convincing new ones (you'll see exactly that in the next tutorial).
And, as always, it's reproducible — same --seed 3, same five names.
The envelope, gently
You wrote three definitions with different types, but they all share the same skeleton: a generator kind, some params, and a seed drives them. Under the hood every generator — built-in or third-party — compiles to one uniform envelope (use + params + constraints + modifiers). You never have to write that raw form; the friendly type shape you used is the envelope, spelled conveniently. The reward is that everything composes: any generator can stand in for any other. If you're curious, the Generators in depth page shows the envelope and every parameter.
What you learned
phony generate <file> --seed <n> --count <n>runs a PGDL definition.- Determinism: the same seed and definition always produce the same data; change the seed to get a different (but equally reproducible) run.
- Three everyday generators: logic (
int_between), list (weighted inline values), and model (novel values from your trained.ngram, registered as a package asset). -f jsongives a pretty array;-f jsonlgives one value per line.
Next
Single values are useful, but real records have several fields that must agree — a name and an email that matches it. That's composition.