Skip to content

Generate your first data ​

You have a trained model from the previous tutorial. Now you'll make Phony actually produce data. You'll write three small generator definitions — a logic generator, a list generator, and a model generator that uses your names.ngram — run each one, and watch the same seed reproduce the same values every time.

Everything here is a phony generate command against a little JSON file. Let's go.

What is a generator? ​

A generator is a named producer of values. You don't write code that loops and appends; you write a small JSON definition that describes the value you want, and Phony produces it — reproducibly, from a seed. Phony ships with a handful of built-in generators; the three you'll meet here are the ones you'll reach for most.

The definitions you write are PGDL — the Phony Generator Definition Language. It's just JSON.

Step 1 — A logic generator (a dice roll) ​

The simplest generator: pure arithmetic, no data, no model. Put this in dice.json:

json
{
  "type": "logic",
  "algorithm": "int_between",
  "params": { "min": 1, "max": 6 }
}

Read it aloud: "a logic generator using the int_between algorithm, between 1 and 6." Now run it:

bash
phony generate dice.json --seed 42 --count 5
json
[
  2,
  2,
  3,
  2,
  2
]

Five dice rolls. The two flags you'll use constantly:

  • --seed 42 — the root seed. It's the single knob that decides the output.
  • --count 5 — how many values to produce. (-n 5 is the short form.)

By default you get a pretty-printed JSON array.

Step 2 — See determinism for yourself ​

This is the property that makes Phony Phony. Run the exact same command again:

bash
phony generate dice.json --seed 42 --count 5
json
[
  2,
  2,
  3,
  2,
  2
]

Identical. Same seed, same definition → same data, every time, on any machine. Now change only the seed:

bash
phony generate dice.json --seed 7 --count 5 -f jsonl
2
4
6
2
4

Different seed, different rolls. (-f jsonl switched the output to JSONL — one value per line — which is friendlier for piping into other tools. The default, -f json, is the pretty array you saw above.)

That reproducibility isn't luck; it's the core design. If you want the full story, see Determinism.

Step 3 — A list generator (pick from a set) ​

Logic invents numbers. A list generator picks from a set of values you supply — and you can weight the choices. Put this in plan.json:

json
{
  "type": "list",
  "source": "inline",
  "values": [
    { "value": "free", "weight": 60 },
    { "value": "pro", "weight": 30 },
    { "value": "enterprise", "weight": 10 }
  ]
}

source: "inline" means the values are right here in the file. The weights are relative: free is picked roughly six times as often as enterprise. Run it:

bash
phony generate plan.json --seed 1 --count 6 -f jsonl
"pro"
"free"
"free"
"pro"
"pro"
"enterprise"

Six plans, weighted toward the cheaper tiers — and, being a generator, perfectly reproducible: run it again with --seed 1 and you'll get that same sequence.

Step 4 — A model generator (use your trained model) ​

Now the payoff from tutorial 1. A model generator produces novel values from a trained .ngram. But a model generator refers to its model by an asset name, not a file path — so first you register your model as an asset by dropping it in a tiny package.

A package is just a directory with a phony.json. Create one:

bash
mkdir -p namespkg/models
cp names.ngram namespkg/models/names.ngram

Then write namespkg/phony.json:

json
{
  "name": "@me/names",
  "version": "1.0.0",
  "assets": {
    "models": {
      "@me/names:first_names": { "*": "models/names.ngram" }
    }
  }
}

That says: "this package publishes a model asset named @me/names:first_names, and for any locale (*) it lives at models/names.ngram." (Packages are how Phony ships reusable data; you'll learn more in Packages.)

Now the generator definition, firstname.json, points at that asset name:

json
{
  "type": "model",
  "source": "@me/names:first_names",
  "generation": { "mode": "word" }
}

Run it, telling Phony to load the package with --package:

bash
phony generate firstname.json --package namespkg --seed 3 --count 5 -f jsonl
✓ Loaded @me/names v1.0.0 — 0 generator(s), 1 asset(s)
"Bade"
"Ada"
"Zeynep"
"Eylul"
"Mira"

The first line (on stderr) is Phony confirming it loaded your package and found one asset. Then five names, generated from the shape your model learned. With a corpus this tiny the model often echoes names it saw — but feed it a few hundred and it starts inventing convincing new ones (you'll see exactly that in the next tutorial).

And, as always, it's reproducible — same --seed 3, same five names.

The envelope, gently ​

You wrote three definitions with different types, but they all share the same skeleton: a generator kind, some params, and a seed drives them. Under the hood every generator — built-in or third-party — compiles to one uniform envelope (use + params + constraints + modifiers). You never have to write that raw form; the friendly type shape you used is the envelope, spelled conveniently. The reward is that everything composes: any generator can stand in for any other. If you're curious, the Generators in depth page shows the envelope and every parameter.

What you learned ​

  • phony generate <file> --seed <n> --count <n> runs a PGDL definition.
  • Determinism: the same seed and definition always produce the same data; change the seed to get a different (but equally reproducible) run.
  • Three everyday generators: logic (int_between), list (weighted inline values), and model (novel values from your trained .ngram, registered as a package asset).
  • -f json gives a pretty array; -f jsonl gives one value per line.

Next ​

Single values are useful, but real records have several fields that must agree — a name and an email that matches it. That's composition.

→ Build a coherent record

Phony Cloud — Documentation & Specification