Skip to content

Build a coherent record ​

Here's the real prize. So far each generator produced one value. But a real record has several fields that have to hang together — an email that actually matches the person's name, not a random address glued onto a random name. In this tutorial you'll build exactly that: a user profile where the email is derived from the name, so they can never disagree. Then you'll sprinkle in 1–3 tags.

This is a composition — a generator built out of other generators. Let's build it up piece by piece, then run the whole thing.

You'll reuse the names.ngram model from Train your first model, so keep it handy.

The idea: coherence comes from a let ​

A composition has a few moving parts:

  • let — named intermediate steps, computed in order. Each step binds a name to a value.
  • output — the final structured record, whose fields reference those names.

The trick to coherence: once a let step binds a name — say name — every later reference to name gets the same value. So if you compute the email fromname, and also put name in the output, they're guaranteed to match. Coherence isn't something you police; it falls out of referencing the same binding twice.

Step 1 — Set up the package ​

Compositions are named generators, so they live in a package (just like the model asset in the last tutorial). Create the package directory and copy your model in:

bash
mkdir -p profilepkg/models
cp names.ngram profilepkg/models/names.ngram

Step 2 — Write the composition ​

Create profilepkg/phony.json. Read it top to bottom — it mirrors the plan above:

json
{
  "name": "@me/profile",
  "version": "1.0.0",
  "assets": {
    "models": {
      "@me/profile:first_names": { "*": "models/names.ngram" }
    }
  },
  "generators": [
    {
      "name": "@me/profile:user",
      "let": {
        "name":  { "type": "model", "source": "@me/profile:first_names", "generation": { "mode": "word" } },
        "email": { "computed": "concat(lowercase(name), '@example.com')" },
        "tags":  {
          "repeat": {
            "count": "1-3",
            "of": { "type": "list", "source": "inline", "values": ["news", "beta", "vip", "sales"] },
            "unique": true
          }
        }
      },
      "output": {
        "name":  "{{ name }}",
        "email": "{{ email }}",
        "tags":  "{{ tags }}"
      }
    }
  ]
}

Three let steps, three kinds of step — this is the whole vocabulary of composition:

  1. name is a generator step: it draws a first name from your model (exactly the model generator from the last tutorial, now living inside a let).
  2. email is a computed step: { "computed": "…" } holds a PEL expression. Here it lowercases name and joins @example.com onto it. Because it reads name, it's tied to whatever name this row drew. (Note we use concat(...) to join strings — in PEL, + is for numbers only.)
  3. tags is a repeat step — covered next.

The output block assembles the final object. The {{ name }}-style placeholders pull each binding into a field.

Step 3 — Understand repeat (N items per record) ​

Sometimes one record needs several of something — a few tags, a handful of order lines. That's repeat. It's a let step shaped like this:

json
"tags": {
  "repeat": {
    "count": "1-3",
    "of": { "type": "list", "source": "inline", "values": ["news", "beta", "vip", "sales"] },
    "unique": true
  }
}
  • count: "1-3" — how many items, as an inclusive range drawn per record (so some users get one tag, some get three). It can also be a plain integer or a PEL expression.
  • of — the generator run once per item. Here, a list pick.
  • unique: true — de-duplicate within this record, so a user never gets news twice.

The result is an array, collected into the tags field.

Step 4 — Run it ​

Because a composition is a named generator, you run it with --use (its registered name) rather than by passing a file:

bash
phony generate --use @me/profile:user --package profilepkg --seed 5 --count 4
json
✓ Loaded @me/profile v1.0.0 — 1 generator(s), 1 asset(s)
[
  {
    "email": "bade@example.com",
    "name": "Bade",
    "tags": [
      "news"
    ]
  },
  {
    "email": "derin@example.com",
    "name": "Derin",
    "tags": [
      "vip"
    ]
  },
  {
    "email": "azra@example.com",
    "name": "Azra",
    "tags": [
      "beta",
      "news"
    ]
  },
  {
    "email": "selis@example.com",
    "name": "Selis",
    "tags": [
      "beta",
      "sales"
    ]
  }
]

Look at what you just built:

  • Every email matches its name. Bade → bade@example.com, Derin → derin@example.com. Not once by luck — by construction, because email was computed from the same name binding the output prints. That's coherence.
  • The tags vary — one to three each, never repeated within a record.
  • Look at the last row: Selis. That name isn't in your 25-name corpus — the model invented it, blending the shapes it learned (the -lis of "Melis", the Se- of "Sena"/"Selin"), and the email followed along automatically. This is the n-gram model doing exactly what tutorial 1 promised.

Step 5 — Confirm it's reproducible ​

Same seed, one value per line this time:

bash
phony generate --use @me/profile:user --package profilepkg --seed 5 --count 4 -f jsonl
✓ Loaded @me/profile v1.0.0 — 1 generator(s), 1 asset(s)
{"email":"bade@example.com","name":"Bade","tags":["news"]}
{"email":"derin@example.com","name":"Derin","tags":["vip"]}
{"email":"azra@example.com","name":"Azra","tags":["beta","news"]}
{"email":"selis@example.com","name":"Selis","tags":["beta","sales"]}

Same four users, same names, same emails, same tags — including how many tags each one got. The entire record, coherence and all, is a pure function of the seed.

What you learned ​

  • A composition builds one generator out of others, using let steps and an output record.
  • Coherence is free: derive a field from a prior let binding (the email from the name) and reference the same binding in the output — they can't disagree.
  • computed steps run PEL expressions over earlier steps (and remember: concat() for strings, since + is numeric-only).
  • repeat produces an array of N items per record, with an optional unique.
  • Your model earns its keep — inventing new, plausible names — and every field stays reproducible from the seed.

Where to go next ​

You've done the whole core loop: train a model, generate data, and compose coherent records. From here:

  • Concepts — the ideas behind determinism, generators, locales, and packages, explained once so the syntax clicks.
  • Generators in depth — every generator and every parameter, including variants, constraints, and the statistical and event-sequence generators.
  • Expressions (PEL) — the full function reference for computed, transform, and templates.
  • The gallery — worked examples to lift and adapt.

For the complete list of train and generate flags, see the CLI reference.

Phony Cloud — Documentation & Specification