Skip to content

Generators in depth ​

A generator is a named producer of values. Every generator — built-in or third-party — obeys one uniform envelope; the friendly per-type PGDL syntax compiles down to it. There are five core generators (Logic, List, Model, Statistical, Event Sequence) plus composition, the universal way to build a generator out of others.

The envelope ​

Internally, every generator invocation is:

json
{
  "use":         "@phony/core:logic.int_between",
  "params":      { "min": 18, "max": 80 },
  "constraints": { },
  "modifiers":   { "unique": true, "nullable": 0.05, "transform": "<PEL>" }
}
KeyMeaning
usenamespaced generator reference (built-ins are @phony/core:*)
paramsgenerator-specific inputs, validated against the generator's manifest
constraintsgenerator-aware output filters (length, prefix, min/max)
modifierscross-cutting, applied to any generator (see below)

In PGDL you write the per-type shape ({ "type": "logic", "algorithm": "int_between", "params": {…} }) and the compiler maps type(+algorithm) → use, gathers the right keys into params, and lifts unique/nullable/transform into modifiers.

Whichever form you write, params are validated against the referenced generator's manifest on the generation path — every call to generate/ generate_one first checks for unknown keys, wrong types, and missing required params, and fails with a BadParams error before any value is drawn. A typo'd param never silently degrades into junk output.

Modifiers (work on every generator) ​

ModifierTypeEffect
uniqueboolvalues are de-duplicated across a run of invocations (type-tagged keys, so "1" ≠ 1; nulls are exempt)
nullablenumber 0..1probability the value is null instead
transformstring (PEL)a PEL expression applied to the produced value, bound as value
json
"score": {
  "type": "logic",
  "algorithm": "int_between",
  "params": { "min": 0, "max": 100 },
  "transform": "if(value >= 50, 'pass', 'fail')"
}

Tiers ​

Built-ins are the Tier 0 primitive kernel (native, fast). Userland generators are Tier 1 (pure-data compositions = primitives + PEL — the 95% case) or Tier 2 (WASM, deferred). The envelope is the stable ABI that lets any of them plug in.

Custom generators (use) ​

The five core types are the built-ins. Because the registry is an open set, a PGDL schema can also reference any generator registered under its namespaced name with use instead of type:

json
"iban": { "use": "@acme/finance:iban", "params": { "country": "TR" } }

params (plus constraints/unique/nullable/transform) work exactly as for a built-in — the generator validates params against its own manifest. Once declared, a custom generator is usable everywhere a built-in is: as a named generator, inside a composition body (e.g. {{ iban }}), or inline on a field. The generator must be registered in the runtime first (registry.register(...) in Rust, or by installing a package that provides it). An unknown use reference is an error at generation time.

Composition generators (Tier 1 — pure data) ​

A use reference can point at a built-in, a generator written in Rust, or a generator authored entirely as data — a composition. A composition is the 95% case for userland generators: primitives + PEL, no compiling. It is also how string composition and coherent multi-field records are expressed — what earlier drafts called the template and linked types (see the note below).

json
{
  "name": "@acme/finance:gross",
  "params": { "net": { "type": "number", "default": 100 },
              "rate": { "type": "number", "default": 0.2 } },
  "let": {
    "tax": { "computed": "round(net * rate, 2)" }
  },
  "body": "{{ net + tax }}"
}
keymeaning
namethe namespaced reference it registers under (what use points at)
paramsthe composition's own inputs (each { type, default?, required? }) — its manifest
letnamed intermediate steps, evaluated in dependency order
bodya template whose rendered value is the generator's output (string)
variantsinstead of body — weighted alternatives [{ pattern, weight }], one chosen per invocation (string); weight 0 = never chosen, all-zero weights are an error
outputinstead of body — a structured value whose string leaves are evaluated against the locals (returns an object/array, not a string)

Each let step is one of three kinds: a generator def (any type/use shape), a { "computed": "<PEL>" } expression, or a { "repeat": … } step that produces an array (N items per parent, below). Input params and earlier let steps are in scope by name — in later steps, in the body, and in step params: a let step that uses another generator has its string params evaluated as PEL against the composition's locals, so values thread through to the sub-generator.

json
"let": { "tagged": { "use": "@acme/demo:tag", "params": { "word": "{{ first }}" } } }

A whole-string {{ expr }} keeps its native type (number, boolean…); a string with embedded interpolation renders to text.

Compositions are isolated: their internal references resolve only against their own params and let steps, never the host schema — so the same composition behaves identically in any schema. They may still compose any other registered generator through a let step.

Once registered (register_composition(&mut registry, &def) in Rust, or loaded from an installed package) a composition is indistinguishable from a built-in at the call site — usable as a named generator, inline, or inside another composition.

template and linked folded into composition

Earlier drafts had two more types. Both are composition now:

  • template (string composition) → a composition with a body (and optional weighted variants). The old operations post-pipeline is just PEL function calls in the body ({{ slugify(name) }}).
  • linked (coherent multi-field record) → either a list over an asset of objects (pick one whole record — its fields always agree), or a composition output that computes coherent fields over prior let steps.

Repeating: N items per parent ​

A repeat let step invokes a sub-generator several times and collects the results into an array — the "N rows per parent" / Faker ->count(n) primitive (e.g. 3–5 line items for this order). It is a let step shaped { "repeat": { "count", "of", "unique"? } }:

json
{
  "name": "@shop/demo:order",
  "params": { "customer": { "type": "string", "default": "acme" } },
  "let": {
    "lines": {
      "repeat": {
        "count": "1-5",
        "of": { "type": "logic", "algorithm": "int_between", "params": { "min": 1, "max": 9 } }
      }
    }
  },
  "output": { "customer": "{{ customer }}", "lines": "{{ lines }}" }
}
fieldmeaning
counthow many items — a literal integer, an inclusive "a-b" range string drawn per invocation, or a PEL expression string over the locals ("n", "n * 2"); a fractional result floors, negatives clamp to 0
ofa sub-generator def (any type/use shape). Its string params are PEL templates evaluated against the parent's locals, so each item can reference the parent ({{ customer }})
uniqueoptional true — de-duplicate the items (bounded retries; errors if the item generator's domain is too small to fill count)

Items are seeded deterministically: each item draws its own seed salted by the step name and the item ordinal, so the array is reproducible run-to-run yet the items vary independently. of may itself be a composition, so repeat nests.

Conditional routing & partitioning ​

Choosing a value (or a distribution) by another field is a composition pattern, not a separate generator. Select a value with a computed let and PEL if (e.g. a title from a gender); partition a distribution by drawing each candidate as a let and selecting with if; for many categories, key a map asset with list key/where so the partition lives in the data. See Generator capabilities & the platform boundary.


1. Logic — pure algorithm ​

No data, no locale, fully reproducible. use = @phony/core:logic.<algorithm>.

algorithmparamsoutput
uuid_v4—random UUID
uuid_v7—time-sortable UUID v7: the first 48 bits carry a synthetic timestamp (the frozen reference instant in ms + the invocation index), the rest comes from the seed — ids sort by index, and stay a pure function of (seed, index, reference_time)
ulid—Crockford-base32 ULID
nanoidlength (default 21)URL-safe id
int_betweenmin, max, exceptinteger in [min, max], optionally excluding values
digitsdigits, strict, leading_zeroa fixed-digit integer (strict = exact width; leading_zero = zero-padded string)
float_betweenmin, max, precisionfloat
booleanprobability (default 0.5)boolean
datetime_betweenstart, endISO-8601 datetime
date_betweenstart, endISO-8601 date
time_betweenstart, end (HH:MM:SS)time-of-day
timestampstart, endUnix seconds (integer)
sequencestart, stepstart + index*step
gaussianmean, stddevnormal-distributed float
exponentiallambdaexponential-distributed float

start/end accept ISO-8601 or relative expressions (-1year, now, now+7days, today-30days) resolved against the reference instant.

json
"created_at": {
  "type": "logic",
  "algorithm": "datetime_between",
  "params": { "start": "-2years", "end": "now" }
}

Worked example — uuid_v7, whose first 48 bits are a synthetic timestamp (frozen reference instant + invocation index), so ids sort by index while staying a pure function of (seed, index, reference_time):

console
$ phony generate uuid_v7.json --seed 7 -n 3 --reference-time 1704067200 -f jsonl
"018cc251-f400-7614-a0ce-f74973cc9a72"
"018cc251-f401-76db-b8e2-e054ac49f61e"
"018cc251-f402-7f9a-b0ee-ae9833721222"

Note the …f400 → …f401 → …f402 prefix marching with the index — monotonic ids without a wall clock.

2. List — finite valid set ​

Pick from a known-correct set. Data is inline or from a locale asset. Elements may be scalars, { value, weight } (weighted), or metadata objects (returned whole — so a list over an asset of objects yields one coherent record).

A weight of 0 means the element is never picked; if every weight is 0 there is nothing selectable, which is an error. (The same rule applies to weighted variants, statistical categorical values, and PEL's pick_weighted.)

json
"status": {
  "type": "list",
  "source": "inline",
  "values": [
    { "value": "active",  "weight": 70 },
    { "value": "churned", "weight": 30 }
  ]
},
"country": {
  "type": "list",
  "source": "geo.countries",
  "locale_independent": true
}
keymeaning
source"inline", an asset name, or { "asset": "name" }; the asset may be an array or a map (for key)
valuesinline elements (when source is inline)
keywhen source is a map, pick from the source[key] bucket (grouped data)
wherea PEL predicate over each element (item) + the caller's bindings; keep matches, then pick
locale_independentresolve the asset the same in every locale

Object elements expose nested fields to templates: {{country.code}}. Picking from an asset whose elements are objects returns one whole record, so its fields always agree — a coherent city + country + postal with no "Ankara, Japan" (what earlier drafts called a linked source).

Filtered & hierarchical selection ​

where is a PEL predicate evaluated for each candidate, with item bound to it and any caller bindings in scope — so you can constrain a set and pick from the rest, including by a parent threaded in from a composition. (The predicate is compiled once and cached, not re-parsed per pick.)

json
// "a district of Adana" — filter a flat districts list
{ "type": "list", "source": "geo.districts", "where": "item.city == 'Adana'" }

For pre-grouped data (each child stored once — no duplication), key indexes a map source directly:

json
// asset geo.districts = { "Adana": ["Seyhan", …], "Ankara": [ … ] }
{ "type": "list", "source": "geo.districts", "key": "Adana" }

Compose them for a coherent il → ilçe → mahalle chain — each level filtered by the one above, with no (il, ilçe) duplication. A composition's output returns a structured object (the coherent tuple) instead of a rendered body string:

json
{
  "name": "@tr/geo:adres",
  "let": {
    "il":      { "type": "list", "source": "geo.cities" },
    "ilce":    { "type": "list", "source": "geo.districts",     "where": "item.city == il.name" },
    "mahalle": { "type": "list", "source": "geo.neighborhoods", "where": "item.district == ilce.name" }
  },
  "output": { "il": "{{ il.name }}", "ilce": "{{ ilce.name }}", "mahalle": "{{ mahalle.name }}" }
}

3. Model — N-gram (locale-flavoured) ​

Generates novel values from a trained model (the Model generator). Requires a generation block.

json
"first_name": {
  "type": "model",
  "source": "person.first_names",
  "generation": { "mode": "word", "params": { "starts_with": "A" } },
  "constraints": { "min_length": 3, "max_length": 12 }
}
keyvalues
sourcethe model asset name (a .ngram)
generation.modeword · sentence · text · paragraph · poem · acrostic · real_word
generation.paramsmode options: word_count, punctuation, max_chars, suffix, sentence_count, starts_with, …
constraintsmin_length · max_length · starts_with · ends_with · contains · exclude_originals

params shapes the generation; constraints filter the output.

4. Statistical — distributions ​

Match real-world shape.

json
"order_status": {
  "type": "statistical",
  "mode": "categorical",
  "values": [
    { "value": "completed", "weight": 70 },
    { "value": "pending",   "weight": 30 }
  ]
},
"age": {
  "type": "statistical",
  "mode": "continuous",
  "distribution": "normal",
  "params": { "mean": 35, "stddev": 12 },
  "constraints": { "min": 18, "max": 85 },
  "differential_privacy": { "enabled": true, "epsilon": 1.0, "mechanism": "laplace" }
}
keyvalues
modecategorical (weighted set) · continuous (a distribution) · multivariate (correlated)
distributionnormal lognormal exponential uniform poisson beta gamma
paramsdistribution params (mean/stddev, mu/sigma, lambda, alpha/beta, shape/scale)
constraintsmin/max clamp (resampled)
differential_privacy{ enabled, epsilon, mechanism: laplace|gaussian }

Categorical weights follow the same rule as everywhere else: weight 0 means never drawn, and all-zero weights are a BadParams error.

multivariate draws correlated values jointly — variables (each { name, mean, stddev }; names must be non-empty and unique) following a correlations matrix (default identity), returned as an object:

json
"physique": {
  "type": "statistical",
  "mode": "multivariate",
  "variables": [
    { "name": "height", "mean": 170, "stddev": 10 },
    { "name": "weight", "mean": 70,  "stddev": 15 }
  ],
  "correlations": [[1.0, 0.8], [0.8, 1.0]]
}

For algebraic relationships (a field derived from another), use a composition's let/computed steps — each a PEL expression over the previous (bonus = round(salary * 0.15, 2)).

5. Event Sequence — chronological dates ​

A base event plus probabilistic, delayed follow-ups.

json
"timeline": {
  "type": "event_sequence",
  "events": [
    { "name": "created_at", "base": true, "range": { "start": "-1year", "end": "now" } },
    { "name": "paid_at",    "after": "created_at", "delay": { "min": "0h", "max": "24h" }, "probability": 0.95 },
    { "name": "shipped_at", "after": "paid_at",    "delay": { "min": "1d", "max": "3d" }, "probability": 0.90 }
  ]
}
event keymeaning
basethe anchor event; range: { start, end }
afteroccurs after the named event
delay{ min, max } durations (s/m/h/d/w)
probabilitychance the event occurs (else null, breaking the chain)

The generator returns an object of event → ISO datetime | null; fields reference timeline.created_at, etc.


For exact parameter validation and design rationale, see Spec → Generator Types and the Generator × Asset matrix.

Phony Cloud — Documentation & Specification