Generator Types
Phony has five core generators — Logic, List, Model, Statistical, and Event Sequence — plus composition, the universal composer that combines them into strings and coherent multi-field records. Anything the core generators + composition + PEL can already express is a pattern, not a new type (see Generator Capabilities).
Implementation status
This page describes the current, built taxonomy. What remains deferred: the statistical algebraic mode is rejected by the engine (write it as a composition computed step); event_sequence per-event condition and non-uniform delay distribution are deferred (delays are uniform today). The former template and linked types are removed — the engine errors on them with a migration hint (see Folded types).
Overview: When to Use Each Type
| You need… | Use | Example |
|---|---|---|
| Values computed by an algorithm (IDs, numbers, dates) | Logic | uuid_v7, int_between |
| Values from a finite, valid set (must be correct) | List | HTTP codes, countries, enums |
| Values that feel natural for a locale | Model (N-gram) | Turkish names, company names |
| Values matching a real-world distribution or correlations | Statistical | ages ~ Normal(35, 12) |
| A chronologically valid series of timestamps | Event Sequence | created → paid → shipped |
| Values composed from other generators (strings or coherent records) | Composition | email, address, geo record |
Key insight:
- CORRECT data → List (pick from valid options) — HTTP 404 must be exactly 404
- REALISTIC data → Model (statistically similar) — a Turkish name should feel Turkish
- COMPUTED data → Logic (pure algorithm) — a UUID must be algorithmically valid
- DISTRIBUTED data → Statistical (matches a spec'd distribution)
- ORDERED data → Event Sequence (dates respect logical order)
- COMPOSED data → Composition (an email combines name + domain)
Type 1: Logic Generators
Pure algorithmic generation that doesn't require external data sources.
Characteristics
- No data files needed
- Deterministic given same seed
- Language/locale independent
- Fastest execution
Use Cases
| Generator | Output | Example |
|---|---|---|
uuid_v4 | UUID version 4 | 550e8400-e29b-41d4-a716-446655440000 |
uuid_v7 | UUID version 7 (time-sortable) | 018d6e5c-5c3a-7f1e-8b5a-2c4d6e8f0a1b |
ulid | ULID | 01ARZ3NDEKTSV4RRFFQ69G5FAV |
int_between | Random integer | 42 |
float_between | Random float | 3.14159 |
boolean | Random boolean | true |
datetime_between | Random datetime | 2024-03-15T14:30:00Z |
date_between | Random date | 2024-03-15 |
time_between | Random time | 14:30:00 |
timestamp | Unix timestamp | 1710512400 |
PGDL Syntax
{
"generators": {
"user_id": {
"type": "logic",
"algorithm": "uuid_v7"
},
"age": {
"type": "logic",
"algorithm": "int_between",
"params": { "min": 18, "max": 85 }
},
"price": {
"type": "logic",
"algorithm": "float_between",
"params": { "min": 0.01, "max": 9999.99, "precision": 2 }
},
"created_at": {
"type": "logic",
"algorithm": "datetime_between",
"params": { "start": "2023-01-01", "end": "now" }
},
"is_active": {
"type": "logic",
"algorithm": "boolean",
"params": { "probability": 0.85 }
}
}
}Available Algorithms
| Algorithm | Params | Description |
|---|---|---|
uuid_v4 | - | Random UUID |
uuid_v7 | - | Time-sortable UUID (synthetic, seed-derived timestamp) |
ulid | - | Universally Unique Lexicographically Sortable ID |
nanoid | length | Nano ID |
int_between | min, max, except | Random integer |
float_between | min, max, precision | Random float |
boolean | probability | Random boolean |
datetime_between | start, end | Random ISO-8601 datetime |
date_between | start, end | Random ISO-8601 date |
time_between | start, end | Random time |
timestamp | start, end | Unix timestamp |
sequence | start, step | Auto-incrementing |
gaussian | mean, stddev | Normal distribution |
exponential | lambda | Exponential distribution |
digits | digits, strict, leading_zero | Fixed-digit-count integer |
Type 2: List Generators
Selection from predefined, finite sets of valid values.
Characteristics
- Data must be correct (valid HTTP codes, real countries)
- Can be locale-specific or universal
- Fast O(1) random selection
- Supports weighted selection
Use Cases
| Category | Examples | Why List? |
|---|---|---|
| Standards | HTTP methods, status codes | Must be valid codes |
| Geography | Countries, currencies | ISO standards |
| Enums | Status values, categories | Application-specific |
| Lookup | Cities, districts | Real place names |
PGDL Syntax
{
"generators": {
"city": {
"type": "list",
"source": "lists/geo/cities.json"
},
"http_method": {
"type": "list",
"source": "inline",
"values": ["GET", "POST", "PUT", "DELETE", "PATCH"],
"locale_independent": true
},
"order_status": {
"type": "list",
"source": "inline",
"values": [
{ "value": "completed", "weight": 60 },
{ "value": "pending", "weight": 25 },
{ "value": "cancelled", "weight": 10 },
{ "value": "refunded", "weight": 5 }
]
},
"country": {
"type": "list",
"source": "lists/geo/countries.json"
}
}
}Note: List files returning objects (like countries with
{ name, code, phone_prefix }) allow accessing nested properties via{{country.code}}inside a composition.
List File Formats
Simple Array (JSON):
["İstanbul", "Ankara", "İzmir", "Bursa", "Antalya"]Weighted Array (JSON):
[
{ "value": "İstanbul", "weight": 40 },
{ "value": "Ankara", "weight": 20 },
{ "value": "İzmir", "weight": 15 },
{ "value": "Bursa", "weight": 10 },
{ "value": "Antalya", "weight": 15 }
]Object with Metadata (JSON):
[
{ "name": "Turkey", "code": "TR", "phone_prefix": "+90", "currency": "TRY" },
{ "name": "Germany", "code": "DE", "phone_prefix": "+49", "currency": "EUR" },
{ "name": "United States", "code": "US", "phone_prefix": "+1", "currency": "USD" }
]Accessing Nested Properties
A list over an object file returns one whole record; a composition let step draws it once, and every reference to its properties agrees:
{
"name": "@acme/geo:country_contact",
"let": {
"country": { "type": "list", "source": "lists/geo/countries.json" }
},
"output": {
"code": "{{ country.code }}",
"phone_prefix": "{{ country.phone_prefix }}",
"currency": "{{ country.currency }}"
}
}Type 3: Model Generators (N-gram)
Statistical generation that produces realistic, locale-specific output.
Deep Dive: For complete N-gram model architecture, training configuration, and generation algorithms, see N-gram Models.
Characteristics
- Trained on real data samples
- Produces statistically similar output (not identical)
- Locale-specific (Turkish names sound Turkish)
- Supports constraints (length, prefix)
Use Cases
| Domain | Why Model? |
|---|---|
| Person names | Should feel natural for the locale |
| Company names | Follow language patterns |
| Street names | Locale-specific naming conventions |
| Product names | Natural language patterns |
| Usernames | Realistic patterns |
PGDL Syntax
Model generators require a generation block specifying how to produce output:
{
"generators": {
"first_name": {
"type": "model",
"source": "models/person_names.ngram",
"generation": { "mode": "word" },
"constraints": { "min_length": 3, "max_length": 12 }
},
"username": {
"type": "model",
"source": "models/usernames.ngram",
"generation": { "mode": "word" },
"constraints": { "min_length": 4, "max_length": 16 }
},
"company_name": {
"type": "model",
"source": "models/company_names.ngram",
"generation": {
"mode": "word",
"params": { "starts_with": "A" }
}
},
"tagline": {
"type": "model",
"source": "models/slogans.ngram",
"generation": {
"mode": "sentence",
"params": {
"word_count": "{{number:3-8}}",
"punctuation": [".", "!"]
}
}
},
"bio": {
"type": "model",
"source": "models/text.ngram",
"generation": {
"mode": "text",
"params": {
"max_chars": 200,
"suffix": "..."
}
}
}
}
}Generation Modes
| Mode | Output | Use Case |
|---|---|---|
word | Single word | Names, usernames |
sentence | Single sentence | Taglines, mottos |
text | Truncated text | Bios, descriptions |
paragraph | Full paragraph | Articles, reviews |
poem | Poem with stanzas | Creative content |
acrostic | Acrostic poem | Hidden messages |
real_word | Pick from training | Known/real names |
Note: Generation parameters can reference other generators using
{{generator_name}}or inline random syntax like{{number:3-8}}. See N-gram Models for details.
Training Models
Models are trained using the phony CLI:
# Character mode (default) - for names, usernames
phony train names.txt -o models/names.ngram
# Word mode - for company names, product names
phony train companies.txt -o models/companies.ngram --token-type word
# Text mode - prose (sentence/paragraph generation)
phony train prose.txt -o models/prose.ngram --token-type text
# Full configuration example (all real flags)
phony train names.txt \
--output models/tr_TR/names.ngram \
--ngram-order 4 \
--token-type char \
--locale tr_TR \
--min-count 2 \
--min-word-length 3 \
--word-filter '[^a-zçğıöşü]+' \
--lowercaseRoadmap: structured inputs (
data.csv --column email,data.json --path "$.users[*].name") are not built yet — extract to a plain text file (one item per line) first.
Token Modes
| Mode | Description | Use Case |
|---|---|---|
char | Character-level N-grams | Names, usernames, made-up words |
word | Word-level N-grams | Company names, product names |
text | Char N-grams + learned word/sentence/paragraph lengths and positional openings | Prose (sentence/text/paragraph modes) |
Model File Structure
The .ngram file is a portable binary (magic PHNYNG03, gzipped body, varint-encoded, sorted vocab + CSR graph). See N-gram Models for the format summary and phony-core/FORMAT.md for the normative byte layout.
Type 4: Statistical Generators
Statistical generators produce data matching real-world distributions, not just random values.
Characteristics
- Preserves frequency distributions (categorical)
- Matches statistical distributions (continuous)
- Maintains correlations between columns (multivariate)
- Supports differential privacy
Use Cases
| Generator Mode | Output | Example |
|---|---|---|
categorical | Preserves value frequencies | 70% completed, 20% pending, 10% cancelled |
continuous | Follows distribution | Normal(mean=35, stddev=12) for ages |
multivariate | Preserves correlations | price correlates with sqft |
Status:
categorical,continuous, andmultivariateare built (the built multivariate shape isvariables: [{name, mean, stddev}, …]+ an N×Ncorrelationsmatrix). The formeralgebraicmode (math relationships liketotal = subtotal + tax - discount) is rejected by the engine — write it as a compositioncomputedstep.
PGDL Syntax
{
"generators": {
"order_status": {
"type": "statistical",
"mode": "categorical",
"values": [
{ "value": "completed", "weight": 70 },
{ "value": "pending", "weight": 20 },
{ "value": "cancelled", "weight": 10 }
]
},
"customer_age": {
"type": "statistical",
"mode": "continuous",
"distribution": "normal",
"params": { "mean": 35, "stddev": 12 },
"constraints": { "min": 18, "max": 85 }
},
"income": {
"type": "statistical",
"mode": "continuous",
"distribution": "lognormal",
"params": { "mu": 10.5, "sigma": 0.8 }
}
}
}Supported Distributions
| Distribution | Use Case | Parameters |
|---|---|---|
normal | Age, height, test scores | mean, stddev |
lognormal | Income, prices, durations | mu, sigma |
exponential | Wait times, failure rates | lambda |
uniform | Random selection | min, max |
poisson | Event counts per period | lambda |
beta | Probabilities, percentages | alpha, beta |
gamma | Wait times, rainfall | shape, scale |
Differential Privacy Option
Add mathematical privacy guarantees:
{
"generators": {
"salary": {
"type": "statistical",
"mode": "continuous",
"distribution": "normal",
"params": { "mean": 75000, "stddev": 25000 },
"differential_privacy": {
"enabled": true,
"epsilon": 1.0,
"mechanism": "laplace"
}
}
}
}Type 5: Event Sequence Generators
Event sequences generate chronologically valid date/time series.
Characteristics
- Dates respect logical order (order < ship < deliver)
- Configurable delays between events
- Probability-based optional events
- Supports business logic conditions
Use Cases
| Sequence | Events | Logic |
|---|---|---|
| Order lifecycle | created → paid → shipped → delivered | Each step after previous |
| User journey | signup → trial_start → trial_end → converted/churned | Branching paths |
| Employment | applied → interviewed → hired → onboarded → promoted | With probabilities |
PGDL Syntax
{
"generators": {
"order_timeline": {
"type": "event_sequence",
"events": [
{
"name": "created_at",
"base": true,
"range": { "start": "-1year", "end": "now" }
},
{
"name": "paid_at",
"after": "created_at",
"delay": { "min": "0h", "max": "24h" },
"probability": 0.95
},
{
"name": "shipped_at",
"after": "paid_at",
"delay": { "min": "1d", "max": "3d" },
"probability": 0.90
},
{
"name": "delivered_at",
"after": "shipped_at",
"delay": { "min": "1d", "max": "7d" },
"probability": 0.85
}
]
}
},
"entities": {
"Order": {
"fields": {
"created_at": { "generator": "order_timeline.created_at" },
"paid_at": { "generator": "order_timeline.paid_at" },
"shipped_at": { "generator": "order_timeline.shipped_at" },
"delivered_at": { "generator": "order_timeline.delivered_at" }
}
}
}
}Event Options
| Option | Description | Example |
|---|---|---|
base | The anchor event, generated first | true |
after | This event occurs after specified event | "created_at" |
delay | Time range between events | { "min": "1d", "max": "7d" } |
probability | Chance this event occurs (null if < 1.0) | 0.85 |
distribution | Delay distribution (deferred — uniform only today) | "exponential" |
condition | Only generate if condition met (deferred) | "status = 'shipped'" |
Composition — the universal composer
Composition is not a sixth data source; it is the envelope that combines the five core generators into strings and coherent multi-field records. It replaces what earlier drafts called the template and linked types — and covers conditional routing, partitioning, and derived (algebraic) columns as patterns.
A composition is a pure-data generator definition:
| key | meaning |
|---|---|
name | the namespaced reference it registers under (what use points at) |
params | the composition's own inputs (each { type, default?, required? }) — its manifest |
let | named intermediate steps, evaluated in dependency order |
body | a template whose rendered value is the generator's output (string) |
variants | instead of body — weighted alternatives [{ pattern, weight }], one chosen per invocation |
output | instead of body — a structured value whose string leaves are evaluated against the locals (returns an object/array) |
Each let step is either a generator def (any type/use shape) or a { "computed": "<PEL>" } expression; input params and earlier let steps are in scope by name. Compositions are isolated — their internal references resolve only against their own params and let steps, never the host schema.
Worked example
{
"name": "@acme/finance:gross",
"params": { "net": { "type": "number", "default": 100 },
"rate": { "type": "number", "default": 0.2 } },
"let": {
"tax": { "computed": "round(net * rate, 2)" }
},
"body": "{{ net + tax }}"
}And combining all the core types into one record:
{
"name": "@acme/demo:user_profile",
"let": {
"user_id": { "type": "logic", "algorithm": "uuid_v7" },
"country": { "type": "list", "source": "lists/geo/countries.json" },
"first_name": { "type": "model", "source": "models/tr_TR/first_names.ngram", "generation": { "mode": "word" } },
"last_name": { "type": "model", "source": "models/tr_TR/last_names.ngram", "generation": { "mode": "word" } }
},
"output": {
"id": "{{ user_id }}",
"name": "{{ first_name }} {{ last_name }}",
"country": "{{ country.name }} ({{ country.code }})",
"phone_prefix": "{{ country.phone_prefix }}"
}
}Every field that references country sees the same drawn record — coherence comes free from naming the value once in let.
Deep Dive: the full envelope semantics (PEL threading into step params, variant weighting rules, registration) are in the PGDL generators guide; composition patterns (conditional routing, partitioning, linking) are catalogued in Generator Capabilities.
Comparison Table
| Feature | Logic | List | Model | Statistical | Event Sequence | Composition |
|---|---|---|---|---|---|---|
| Data source | Algorithm | File/Inline | Trained model | Distribution spec | Event spec | Other generators |
| Locale-specific | No | Optional | Yes | No | No | Depends on parts |
| Output variety | Controlled | Finite | Infinite | Distribution-shaped | Structured | Controlled |
| Correctness | Guaranteed | Guaranteed | Statistical | Statistical | Order guaranteed | Depends on parts |
| Output shape | Scalar | Scalar or object | String | Number/category | Object (event → datetime) | String or object |
| Speed | Fastest | Very fast | Fast | Very fast | Fast | Depends on depth |
| Training needed | No | No | Yes | No | No | No |
| Typical use | IDs, numbers | Codes, enums | Names, text | Distributions, correlations | Timelines | Emails, addresses, coherent records |
Cross-Cutting Features: Consistency & Linking
Consistency
Built (generator level):
- Within one invocation — name a value once as a
letstep and reference it from many places; it is drawn once, so every reference agrees (see theuser_profileexample above). - Keyed generation — the deterministic
word_for(key)primitive: same key, same output, across invocations.
Target design (cloud platform): the consistency: {enabled, key, scope} block with table/entity/global scopes, and automatic PK→FK propagation with format-preserving encryption (FPE), operate across rows and tables — that is platform orchestration over the keyed primitive, not a generator feature. See Generator Capabilities and Advanced Concepts.
Linking
Coherent related columns ("Ankara" must come with "Turkey" and "06100") are a built composition pattern, not a type:
- Coherent by selection — a
listover an asset of objects picks one whole record; its fields always agree. - Coherent by construction — a composition
outputcomputes fields over sharedletsteps (e.g.bonus,tax,net_incomeall derived from one drawnsalary).
Target design (cloud platform): foreign-key linking across tables (referential integrity, cross-table aggregates) needs the cross-table row set and is deferred to the cloud layer — see Generator Capabilities and Advanced Concepts.
Folded types: template & linked (migration)
Earlier drafts had seven types. Two were folded into composition because they were redundant; the engine now errors with a migration hint when it sees them:
{"type": "template", …}→ "the 'template' type was folded into composition: author a composition with abodytemplate (optionally weightedvariants);operationsbecome PEL calls in the body"{"type": "linked", …}→ "the 'linked' type was folded: use alistover an asset to pick one coherent record, or a compositionoutputto compute coherent fields"
template → composition body / variants
Template was string composition: a pattern (or weighted variants) with placeholders and an operations post-pipeline.
// before (rejected by the engine)
{ "type": "template", "pattern": "{{product_name}}", "operations": ["slugify", "lowercase"] }
// after
{
"name": "@acme/web:slug",
"let": { "product_name": { "use": "@acme/catalog:product_name" } },
"body": "{{ lowercase(slugify(product_name)) }}"
}Weighted template variants carry over directly — composition variants are the same [{ "pattern": …, "weight": … }] shape.
linked → list over an object asset, or composition output/let
Linked drew multiple coherent columns from one selection (source) or from computed rules.
sourceform (pick one record, fields agree) → alistover the object asset inside a compositionlet, with anoutputexposing the fields — see Accessing Nested Properties above.rulesform (mathematically related fields) →letsteps +computedPEL, exposed viaoutput:
{
"name": "@acme/hr:financials",
"let": {
"salary": { "type": "statistical", "mode": "continuous", "distribution": "lognormal", "params": { "mu": 10, "sigma": 0.5 } },
"bonus": { "computed": "round(salary * uniform(0.05, 0.20), 2)" },
"tax": { "computed": "round((salary + bonus) * 0.25, 2)" }
},
"output": {
"salary": "{{ salary }}",
"bonus": "{{ bonus }}",
"tax": "{{ tax }}",
"net_income": "{{ salary + bonus - tax }}"
}
}