Generators in depth
A generator is a named producer of values. Every generator — built-in or third-party — obeys one uniform envelope; the friendly per-type PGDL syntax compiles down to it. There are five core generators (Logic, List, Model, Statistical, Event Sequence) plus composition, the universal way to build a generator out of others.
The envelope
Internally, every generator invocation is:
{
"use": "@phony/core:logic.int_between",
"params": { "min": 18, "max": 80 },
"constraints": { },
"modifiers": { "unique": true, "nullable": 0.05, "transform": "<PEL>" }
}| Key | Meaning |
|---|---|
use | namespaced generator reference (built-ins are @phony/core:*) |
params | generator-specific inputs, validated against the generator's manifest |
constraints | generator-aware output filters (length, prefix, min/max) |
modifiers | cross-cutting, applied to any generator (see below) |
In PGDL you write the per-type shape ({ "type": "logic", "algorithm": "int_between", "params": {…} }) and the compiler maps type(+algorithm) → use, gathers the right keys into params, and lifts unique/nullable/transform into modifiers.
Whichever form you write, params are validated against the referenced generator's manifest on the generation path — every call to generate/ generate_one first checks for unknown keys, wrong types, and missing required params, and fails with a BadParams error before any value is drawn. A typo'd param never silently degrades into junk output.
Modifiers (work on every generator)
| Modifier | Type | Effect |
|---|---|---|
unique | bool | values are de-duplicated across a run of invocations (type-tagged keys, so "1" ≠ 1; nulls are exempt) |
nullable | number 0..1 | probability the value is null instead |
transform | string (PEL) | a PEL expression applied to the produced value, bound as value |
"score": {
"type": "logic",
"algorithm": "int_between",
"params": { "min": 0, "max": 100 },
"transform": "if(value >= 50, 'pass', 'fail')"
}Tiers
Built-ins are the Tier 0 primitive kernel (native, fast). Userland generators are Tier 1 (pure-data compositions = primitives + PEL — the 95% case) or Tier 2 (WASM, deferred). The envelope is the stable ABI that lets any of them plug in.
Custom generators (use)
The five core types are the built-ins. Because the registry is an open set, a PGDL schema can also reference any generator registered under its namespaced name with use instead of type:
"iban": { "use": "@acme/finance:iban", "params": { "country": "TR" } }params (plus constraints/unique/nullable/transform) work exactly as for a built-in — the generator validates params against its own manifest. Once declared, a custom generator is usable everywhere a built-in is: as a named generator, inside a composition body (e.g. {{ iban }}), or inline on a field. The generator must be registered in the runtime first (registry.register(...) in Rust, or by installing a package that provides it). An unknown use reference is an error at generation time.
Composition generators (Tier 1 — pure data)
A use reference can point at a built-in, a generator written in Rust, or a generator authored entirely as data — a composition. A composition is the 95% case for userland generators: primitives + PEL, no compiling. It is also how string composition and coherent multi-field records are expressed — what earlier drafts called the template and linked types (see the note below).
{
"name": "@acme/finance:gross",
"params": { "net": { "type": "number", "default": 100 },
"rate": { "type": "number", "default": 0.2 } },
"let": {
"tax": { "computed": "round(net * rate, 2)" }
},
"body": "{{ net + tax }}"
}| key | meaning |
|---|---|
name | the namespaced reference it registers under (what use points at) |
params | the composition's own inputs (each { type, default?, required? }) — its manifest |
let | named intermediate steps, evaluated in dependency order |
body | a template whose rendered value is the generator's output (string) |
variants | instead of body — weighted alternatives [{ pattern, weight }], one chosen per invocation (string); weight 0 = never chosen, all-zero weights are an error |
output | instead of body — a structured value whose string leaves are evaluated against the locals (returns an object/array, not a string) |
Each let step is one of three kinds: a generator def (any type/use shape), a { "computed": "<PEL>" } expression, or a { "repeat": … } step that produces an array (N items per parent, below). Input params and earlier let steps are in scope by name — in later steps, in the body, and in step params: a let step that uses another generator has its string params evaluated as PEL against the composition's locals, so values thread through to the sub-generator.
"let": { "tagged": { "use": "@acme/demo:tag", "params": { "word": "{{ first }}" } } }A whole-string {{ expr }} keeps its native type (number, boolean…); a string with embedded interpolation renders to text.
Compositions are isolated: their internal references resolve only against their own params and let steps, never the host schema — so the same composition behaves identically in any schema. They may still compose any other registered generator through a let step.
Once registered (register_composition(&mut registry, &def) in Rust, or loaded from an installed package) a composition is indistinguishable from a built-in at the call site — usable as a named generator, inline, or inside another composition.
template and linked folded into composition
Earlier drafts had two more types. Both are composition now:
template(string composition) → a composition with abody(and optional weightedvariants). The oldoperationspost-pipeline is just PEL function calls in the body ({{ slugify(name) }}).linked(coherent multi-field record) → either alistover an asset of objects (pick one whole record — its fields always agree), or a compositionoutputthat computes coherent fields over priorletsteps.
Repeating: N items per parent
A repeat let step invokes a sub-generator several times and collects the results into an array — the "N rows per parent" / Faker ->count(n) primitive (e.g. 3–5 line items for this order). It is a let step shaped { "repeat": { "count", "of", "unique"? } }:
{
"name": "@shop/demo:order",
"params": { "customer": { "type": "string", "default": "acme" } },
"let": {
"lines": {
"repeat": {
"count": "1-5",
"of": { "type": "logic", "algorithm": "int_between", "params": { "min": 1, "max": 9 } }
}
}
},
"output": { "customer": "{{ customer }}", "lines": "{{ lines }}" }
}| field | meaning |
|---|---|
count | how many items — a literal integer, an inclusive "a-b" range string drawn per invocation, or a PEL expression string over the locals ("n", "n * 2"); a fractional result floors, negatives clamp to 0 |
of | a sub-generator def (any type/use shape). Its string params are PEL templates evaluated against the parent's locals, so each item can reference the parent ({{ customer }}) |
unique | optional true — de-duplicate the items (bounded retries; errors if the item generator's domain is too small to fill count) |
Items are seeded deterministically: each item draws its own seed salted by the step name and the item ordinal, so the array is reproducible run-to-run yet the items vary independently. of may itself be a composition, so repeat nests.
Conditional routing & partitioning
Choosing a value (or a distribution) by another field is a composition pattern, not a separate generator. Select a value with a computed let and PEL if (e.g. a title from a gender); partition a distribution by drawing each candidate as a let and selecting with if; for many categories, key a map asset with list key/where so the partition lives in the data. See Generator capabilities & the platform boundary.
1. Logic — pure algorithm
No data, no locale, fully reproducible. use = @phony/core:logic.<algorithm>.
| algorithm | params | output |
|---|---|---|
uuid_v4 | — | random UUID |
uuid_v7 | — | time-sortable UUID v7: the first 48 bits carry a synthetic timestamp (the frozen reference instant in ms + the invocation index), the rest comes from the seed — ids sort by index, and stay a pure function of (seed, index, reference_time) |
ulid | — | Crockford-base32 ULID |
nanoid | length (default 21) | URL-safe id |
int_between | min, max, except | integer in [min, max], optionally excluding values |
digits | digits, strict, leading_zero | a fixed-digit integer (strict = exact width; leading_zero = zero-padded string) |
float_between | min, max, precision | float |
boolean | probability (default 0.5) | boolean |
datetime_between | start, end | ISO-8601 datetime |
date_between | start, end | ISO-8601 date |
time_between | start, end (HH:MM:SS) | time-of-day |
timestamp | start, end | Unix seconds (integer) |
sequence | start, step | start + index*step |
gaussian | mean, stddev | normal-distributed float |
exponential | lambda | exponential-distributed float |
start/end accept ISO-8601 or relative expressions (-1year, now, now+7days, today-30days) resolved against the reference instant.
"created_at": {
"type": "logic",
"algorithm": "datetime_between",
"params": { "start": "-2years", "end": "now" }
}Worked example — uuid_v7, whose first 48 bits are a synthetic timestamp (frozen reference instant + invocation index), so ids sort by index while staying a pure function of (seed, index, reference_time):
$ phony generate uuid_v7.json --seed 7 -n 3 --reference-time 1704067200 -f jsonl
"018cc251-f400-7614-a0ce-f74973cc9a72"
"018cc251-f401-76db-b8e2-e054ac49f61e"
"018cc251-f402-7f9a-b0ee-ae9833721222"Note the …f400 → …f401 → …f402 prefix marching with the index — monotonic ids without a wall clock.
2. List — finite valid set
Pick from a known-correct set. Data is inline or from a locale asset. Elements may be scalars, { value, weight } (weighted), or metadata objects (returned whole — so a list over an asset of objects yields one coherent record).
A weight of 0 means the element is never picked; if every weight is 0 there is nothing selectable, which is an error. (The same rule applies to weighted variants, statistical categorical values, and PEL's pick_weighted.)
"status": {
"type": "list",
"source": "inline",
"values": [
{ "value": "active", "weight": 70 },
{ "value": "churned", "weight": 30 }
]
},
"country": {
"type": "list",
"source": "geo.countries",
"locale_independent": true
}| key | meaning |
|---|---|
source | "inline", an asset name, or { "asset": "name" }; the asset may be an array or a map (for key) |
values | inline elements (when source is inline) |
key | when source is a map, pick from the source[key] bucket (grouped data) |
where | a PEL predicate over each element (item) + the caller's bindings; keep matches, then pick |
locale_independent | resolve the asset the same in every locale |
Object elements expose nested fields to templates: {{country.code}}. Picking from an asset whose elements are objects returns one whole record, so its fields always agree — a coherent city + country + postal with no "Ankara, Japan" (what earlier drafts called a linked source).
Filtered & hierarchical selection
where is a PEL predicate evaluated for each candidate, with item bound to it and any caller bindings in scope — so you can constrain a set and pick from the rest, including by a parent threaded in from a composition. (The predicate is compiled once and cached, not re-parsed per pick.)
// "a district of Adana" — filter a flat districts list
{ "type": "list", "source": "geo.districts", "where": "item.city == 'Adana'" }For pre-grouped data (each child stored once — no duplication), key indexes a map source directly:
// asset geo.districts = { "Adana": ["Seyhan", …], "Ankara": [ … ] }
{ "type": "list", "source": "geo.districts", "key": "Adana" }Compose them for a coherent il → ilçe → mahalle chain — each level filtered by the one above, with no (il, ilçe) duplication. A composition's output returns a structured object (the coherent tuple) instead of a rendered body string:
{
"name": "@tr/geo:adres",
"let": {
"il": { "type": "list", "source": "geo.cities" },
"ilce": { "type": "list", "source": "geo.districts", "where": "item.city == il.name" },
"mahalle": { "type": "list", "source": "geo.neighborhoods", "where": "item.district == ilce.name" }
},
"output": { "il": "{{ il.name }}", "ilce": "{{ ilce.name }}", "mahalle": "{{ mahalle.name }}" }
}3. Model — N-gram (locale-flavoured)
Generates novel values from a trained model (the Model generator). Requires a generation block.
"first_name": {
"type": "model",
"source": "person.first_names",
"generation": { "mode": "word", "params": { "starts_with": "A" } },
"constraints": { "min_length": 3, "max_length": 12 }
}| key | values |
|---|---|
source | the model asset name (a .ngram) |
generation.mode | word · sentence · text · paragraph · poem · acrostic · real_word |
generation.params | mode options: word_count, punctuation, max_chars, suffix, sentence_count, starts_with, … |
constraints | min_length · max_length · starts_with · ends_with · contains · exclude_originals |
params shapes the generation; constraints filter the output.
4. Statistical — distributions
Match real-world shape.
"order_status": {
"type": "statistical",
"mode": "categorical",
"values": [
{ "value": "completed", "weight": 70 },
{ "value": "pending", "weight": 30 }
]
},
"age": {
"type": "statistical",
"mode": "continuous",
"distribution": "normal",
"params": { "mean": 35, "stddev": 12 },
"constraints": { "min": 18, "max": 85 },
"differential_privacy": { "enabled": true, "epsilon": 1.0, "mechanism": "laplace" }
}| key | values |
|---|---|
mode | categorical (weighted set) · continuous (a distribution) · multivariate (correlated) |
distribution | normal lognormal exponential uniform poisson beta gamma |
params | distribution params (mean/stddev, mu/sigma, lambda, alpha/beta, shape/scale) |
constraints | min/max clamp (resampled) |
differential_privacy | { enabled, epsilon, mechanism: laplace|gaussian } |
Categorical weights follow the same rule as everywhere else: weight 0 means never drawn, and all-zero weights are a BadParams error.
multivariate draws correlated values jointly — variables (each { name, mean, stddev }; names must be non-empty and unique) following a correlations matrix (default identity), returned as an object:
"physique": {
"type": "statistical",
"mode": "multivariate",
"variables": [
{ "name": "height", "mean": 170, "stddev": 10 },
{ "name": "weight", "mean": 70, "stddev": 15 }
],
"correlations": [[1.0, 0.8], [0.8, 1.0]]
}For algebraic relationships (a field derived from another), use a composition's let/computed steps — each a PEL expression over the previous (bonus = round(salary * 0.15, 2)).
5. Event Sequence — chronological dates
A base event plus probabilistic, delayed follow-ups.
"timeline": {
"type": "event_sequence",
"events": [
{ "name": "created_at", "base": true, "range": { "start": "-1year", "end": "now" } },
{ "name": "paid_at", "after": "created_at", "delay": { "min": "0h", "max": "24h" }, "probability": 0.95 },
{ "name": "shipped_at", "after": "paid_at", "delay": { "min": "1d", "max": "3d" }, "probability": 0.90 }
]
}| event key | meaning |
|---|---|
base | the anchor event; range: { start, end } |
after | occurs after the named event |
delay | { min, max } durations (s/m/h/d/w) |
probability | chance the event occurs (else null, breaking the chain) |
The generator returns an object of event → ISO datetime | null; fields reference timeline.created_at, etc.
For exact parameter validation and design rationale, see Spec → Generator Types and the Generator × Asset matrix.