Skip to content

Generator Types ​

Phony has five core generators — Logic, List, Model, Statistical, and Event Sequence — plus composition, the universal composer that combines them into strings and coherent multi-field records. Anything the core generators + composition + PEL can already express is a pattern, not a new type (see Generator Capabilities).

Implementation status

This page describes the current, built taxonomy. What remains deferred: the statistical algebraic mode is rejected by the engine (write it as a composition computed step); event_sequence per-event condition and non-uniform delay distribution are deferred (delays are uniform today). The former template and linked types are removed — the engine errors on them with a migration hint (see Folded types).

Overview: When to Use Each Type ​

You need…UseExample
Values computed by an algorithm (IDs, numbers, dates)Logicuuid_v7, int_between
Values from a finite, valid set (must be correct)ListHTTP codes, countries, enums
Values that feel natural for a localeModel (N-gram)Turkish names, company names
Values matching a real-world distribution or correlationsStatisticalages ~ Normal(35, 12)
A chronologically valid series of timestampsEvent Sequencecreated → paid → shipped
Values composed from other generators (strings or coherent records)Compositionemail, address, geo record

Key insight:

  • CORRECT data → List (pick from valid options) — HTTP 404 must be exactly 404
  • REALISTIC data → Model (statistically similar) — a Turkish name should feel Turkish
  • COMPUTED data → Logic (pure algorithm) — a UUID must be algorithmically valid
  • DISTRIBUTED data → Statistical (matches a spec'd distribution)
  • ORDERED data → Event Sequence (dates respect logical order)
  • COMPOSED data → Composition (an email combines name + domain)

Type 1: Logic Generators ​

Pure algorithmic generation that doesn't require external data sources.

Characteristics ​

  • No data files needed
  • Deterministic given same seed
  • Language/locale independent
  • Fastest execution

Use Cases ​

GeneratorOutputExample
uuid_v4UUID version 4550e8400-e29b-41d4-a716-446655440000
uuid_v7UUID version 7 (time-sortable)018d6e5c-5c3a-7f1e-8b5a-2c4d6e8f0a1b
ulidULID01ARZ3NDEKTSV4RRFFQ69G5FAV
int_betweenRandom integer42
float_betweenRandom float3.14159
booleanRandom booleantrue
datetime_betweenRandom datetime2024-03-15T14:30:00Z
date_betweenRandom date2024-03-15
time_betweenRandom time14:30:00
timestampUnix timestamp1710512400

PGDL Syntax ​

json
{
  "generators": {
    "user_id": {
      "type": "logic",
      "algorithm": "uuid_v7"
    },
    "age": {
      "type": "logic",
      "algorithm": "int_between",
      "params": { "min": 18, "max": 85 }
    },
    "price": {
      "type": "logic",
      "algorithm": "float_between",
      "params": { "min": 0.01, "max": 9999.99, "precision": 2 }
    },
    "created_at": {
      "type": "logic",
      "algorithm": "datetime_between",
      "params": { "start": "2023-01-01", "end": "now" }
    },
    "is_active": {
      "type": "logic",
      "algorithm": "boolean",
      "params": { "probability": 0.85 }
    }
  }
}

Available Algorithms ​

AlgorithmParamsDescription
uuid_v4-Random UUID
uuid_v7-Time-sortable UUID (synthetic, seed-derived timestamp)
ulid-Universally Unique Lexicographically Sortable ID
nanoidlengthNano ID
int_betweenmin, max, exceptRandom integer
float_betweenmin, max, precisionRandom float
booleanprobabilityRandom boolean
datetime_betweenstart, endRandom ISO-8601 datetime
date_betweenstart, endRandom ISO-8601 date
time_betweenstart, endRandom time
timestampstart, endUnix timestamp
sequencestart, stepAuto-incrementing
gaussianmean, stddevNormal distribution
exponentiallambdaExponential distribution
digitsdigits, strict, leading_zeroFixed-digit-count integer

Type 2: List Generators ​

Selection from predefined, finite sets of valid values.

Characteristics ​

  • Data must be correct (valid HTTP codes, real countries)
  • Can be locale-specific or universal
  • Fast O(1) random selection
  • Supports weighted selection

Use Cases ​

CategoryExamplesWhy List?
StandardsHTTP methods, status codesMust be valid codes
GeographyCountries, currenciesISO standards
EnumsStatus values, categoriesApplication-specific
LookupCities, districtsReal place names

PGDL Syntax ​

json
{
  "generators": {
    "city": {
      "type": "list",
      "source": "lists/geo/cities.json"
    },
    "http_method": {
      "type": "list",
      "source": "inline",
      "values": ["GET", "POST", "PUT", "DELETE", "PATCH"],
      "locale_independent": true
    },
    "order_status": {
      "type": "list",
      "source": "inline",
      "values": [
        { "value": "completed", "weight": 60 },
        { "value": "pending", "weight": 25 },
        { "value": "cancelled", "weight": 10 },
        { "value": "refunded", "weight": 5 }
      ]
    },
    "country": {
      "type": "list",
      "source": "lists/geo/countries.json"
    }
  }
}

Note: List files returning objects (like countries with { name, code, phone_prefix }) allow accessing nested properties via {{country.code}} inside a composition.

List File Formats ​

Simple Array (JSON):

json
["İstanbul", "Ankara", "İzmir", "Bursa", "Antalya"]

Weighted Array (JSON):

json
[
  { "value": "İstanbul", "weight": 40 },
  { "value": "Ankara", "weight": 20 },
  { "value": "İzmir", "weight": 15 },
  { "value": "Bursa", "weight": 10 },
  { "value": "Antalya", "weight": 15 }
]

Object with Metadata (JSON):

json
[
  { "name": "Turkey", "code": "TR", "phone_prefix": "+90", "currency": "TRY" },
  { "name": "Germany", "code": "DE", "phone_prefix": "+49", "currency": "EUR" },
  { "name": "United States", "code": "US", "phone_prefix": "+1", "currency": "USD" }
]

Accessing Nested Properties ​

A list over an object file returns one whole record; a composition let step draws it once, and every reference to its properties agrees:

json
{
  "name": "@acme/geo:country_contact",
  "let": {
    "country": { "type": "list", "source": "lists/geo/countries.json" }
  },
  "output": {
    "code": "{{ country.code }}",
    "phone_prefix": "{{ country.phone_prefix }}",
    "currency": "{{ country.currency }}"
  }
}

Type 3: Model Generators (N-gram) ​

Statistical generation that produces realistic, locale-specific output.

Deep Dive: For complete N-gram model architecture, training configuration, and generation algorithms, see N-gram Models.

Characteristics ​

  • Trained on real data samples
  • Produces statistically similar output (not identical)
  • Locale-specific (Turkish names sound Turkish)
  • Supports constraints (length, prefix)

Use Cases ​

DomainWhy Model?
Person namesShould feel natural for the locale
Company namesFollow language patterns
Street namesLocale-specific naming conventions
Product namesNatural language patterns
UsernamesRealistic patterns

PGDL Syntax ​

Model generators require a generation block specifying how to produce output:

json
{
  "generators": {
    "first_name": {
      "type": "model",
      "source": "models/person_names.ngram",
      "generation": { "mode": "word" },
      "constraints": { "min_length": 3, "max_length": 12 }
    },
    "username": {
      "type": "model",
      "source": "models/usernames.ngram",
      "generation": { "mode": "word" },
      "constraints": { "min_length": 4, "max_length": 16 }
    },
    "company_name": {
      "type": "model",
      "source": "models/company_names.ngram",
      "generation": {
        "mode": "word",
        "params": { "starts_with": "A" }
      }
    },
    "tagline": {
      "type": "model",
      "source": "models/slogans.ngram",
      "generation": {
        "mode": "sentence",
        "params": {
          "word_count": "{{number:3-8}}",
          "punctuation": [".", "!"]
        }
      }
    },
    "bio": {
      "type": "model",
      "source": "models/text.ngram",
      "generation": {
        "mode": "text",
        "params": {
          "max_chars": 200,
          "suffix": "..."
        }
      }
    }
  }
}

Generation Modes ​

ModeOutputUse Case
wordSingle wordNames, usernames
sentenceSingle sentenceTaglines, mottos
textTruncated textBios, descriptions
paragraphFull paragraphArticles, reviews
poemPoem with stanzasCreative content
acrosticAcrostic poemHidden messages
real_wordPick from trainingKnown/real names

Note: Generation parameters can reference other generators using {{generator_name}} or inline random syntax like {{number:3-8}}. See N-gram Models for details.

Training Models ​

Models are trained using the phony CLI:

bash
# Character mode (default) - for names, usernames
phony train names.txt -o models/names.ngram

# Word mode - for company names, product names
phony train companies.txt -o models/companies.ngram --token-type word

# Text mode - prose (sentence/paragraph generation)
phony train prose.txt -o models/prose.ngram --token-type text

# Full configuration example (all real flags)
phony train names.txt \
  --output models/tr_TR/names.ngram \
  --ngram-order 4 \
  --token-type char \
  --locale tr_TR \
  --min-count 2 \
  --min-word-length 3 \
  --word-filter '[^a-zçğıöşü]+' \
  --lowercase

Roadmap: structured inputs (data.csv --column email, data.json --path "$.users[*].name") are not built yet — extract to a plain text file (one item per line) first.

Token Modes ​

ModeDescriptionUse Case
charCharacter-level N-gramsNames, usernames, made-up words
wordWord-level N-gramsCompany names, product names
textChar N-grams + learned word/sentence/paragraph lengths and positional openingsProse (sentence/text/paragraph modes)

Model File Structure ​

The .ngram file is a portable binary (magic PHNYNG03, gzipped body, varint-encoded, sorted vocab + CSR graph). See N-gram Models for the format summary and phony-core/FORMAT.md for the normative byte layout.


Type 4: Statistical Generators ​

Statistical generators produce data matching real-world distributions, not just random values.

Characteristics ​

  • Preserves frequency distributions (categorical)
  • Matches statistical distributions (continuous)
  • Maintains correlations between columns (multivariate)
  • Supports differential privacy

Use Cases ​

Generator ModeOutputExample
categoricalPreserves value frequencies70% completed, 20% pending, 10% cancelled
continuousFollows distributionNormal(mean=35, stddev=12) for ages
multivariatePreserves correlationsprice correlates with sqft

Status: categorical, continuous, and multivariate are built (the built multivariate shape is variables: [{name, mean, stddev}, …] + an N×N correlations matrix). The former algebraic mode (math relationships like total = subtotal + tax - discount) is rejected by the engine — write it as a composition computed step.

PGDL Syntax ​

json
{
  "generators": {
    "order_status": {
      "type": "statistical",
      "mode": "categorical",
      "values": [
        { "value": "completed", "weight": 70 },
        { "value": "pending", "weight": 20 },
        { "value": "cancelled", "weight": 10 }
      ]
    },
    "customer_age": {
      "type": "statistical",
      "mode": "continuous",
      "distribution": "normal",
      "params": { "mean": 35, "stddev": 12 },
      "constraints": { "min": 18, "max": 85 }
    },
    "income": {
      "type": "statistical",
      "mode": "continuous",
      "distribution": "lognormal",
      "params": { "mu": 10.5, "sigma": 0.8 }
    }
  }
}

Supported Distributions ​

DistributionUse CaseParameters
normalAge, height, test scoresmean, stddev
lognormalIncome, prices, durationsmu, sigma
exponentialWait times, failure rateslambda
uniformRandom selectionmin, max
poissonEvent counts per periodlambda
betaProbabilities, percentagesalpha, beta
gammaWait times, rainfallshape, scale

Differential Privacy Option ​

Add mathematical privacy guarantees:

json
{
  "generators": {
    "salary": {
      "type": "statistical",
      "mode": "continuous",
      "distribution": "normal",
      "params": { "mean": 75000, "stddev": 25000 },
      "differential_privacy": {
        "enabled": true,
        "epsilon": 1.0,
        "mechanism": "laplace"
      }
    }
  }
}

Type 5: Event Sequence Generators ​

Event sequences generate chronologically valid date/time series.

Characteristics ​

  • Dates respect logical order (order < ship < deliver)
  • Configurable delays between events
  • Probability-based optional events
  • Supports business logic conditions

Use Cases ​

SequenceEventsLogic
Order lifecyclecreated → paid → shipped → deliveredEach step after previous
User journeysignup → trial_start → trial_end → converted/churnedBranching paths
Employmentapplied → interviewed → hired → onboarded → promotedWith probabilities

PGDL Syntax ​

json
{
  "generators": {
    "order_timeline": {
      "type": "event_sequence",
      "events": [
        {
          "name": "created_at",
          "base": true,
          "range": { "start": "-1year", "end": "now" }
        },
        {
          "name": "paid_at",
          "after": "created_at",
          "delay": { "min": "0h", "max": "24h" },
          "probability": 0.95
        },
        {
          "name": "shipped_at",
          "after": "paid_at",
          "delay": { "min": "1d", "max": "3d" },
          "probability": 0.90
        },
        {
          "name": "delivered_at",
          "after": "shipped_at",
          "delay": { "min": "1d", "max": "7d" },
          "probability": 0.85
        }
      ]
    }
  },
  "entities": {
    "Order": {
      "fields": {
        "created_at": { "generator": "order_timeline.created_at" },
        "paid_at": { "generator": "order_timeline.paid_at" },
        "shipped_at": { "generator": "order_timeline.shipped_at" },
        "delivered_at": { "generator": "order_timeline.delivered_at" }
      }
    }
  }
}

Event Options ​

OptionDescriptionExample
baseThe anchor event, generated firsttrue
afterThis event occurs after specified event"created_at"
delayTime range between events{ "min": "1d", "max": "7d" }
probabilityChance this event occurs (null if < 1.0)0.85
distributionDelay distribution (deferred — uniform only today)"exponential"
conditionOnly generate if condition met (deferred)"status = 'shipped'"

Composition — the universal composer ​

Composition is not a sixth data source; it is the envelope that combines the five core generators into strings and coherent multi-field records. It replaces what earlier drafts called the template and linked types — and covers conditional routing, partitioning, and derived (algebraic) columns as patterns.

A composition is a pure-data generator definition:

keymeaning
namethe namespaced reference it registers under (what use points at)
paramsthe composition's own inputs (each { type, default?, required? }) — its manifest
letnamed intermediate steps, evaluated in dependency order
bodya template whose rendered value is the generator's output (string)
variantsinstead of body — weighted alternatives [{ pattern, weight }], one chosen per invocation
outputinstead of body — a structured value whose string leaves are evaluated against the locals (returns an object/array)

Each let step is either a generator def (any type/use shape) or a { "computed": "<PEL>" } expression; input params and earlier let steps are in scope by name. Compositions are isolated — their internal references resolve only against their own params and let steps, never the host schema.

Worked example ​

json
{
  "name": "@acme/finance:gross",
  "params": { "net": { "type": "number", "default": 100 },
              "rate": { "type": "number", "default": 0.2 } },
  "let": {
    "tax": { "computed": "round(net * rate, 2)" }
  },
  "body": "{{ net + tax }}"
}

And combining all the core types into one record:

json
{
  "name": "@acme/demo:user_profile",
  "let": {
    "user_id": { "type": "logic", "algorithm": "uuid_v7" },
    "country": { "type": "list", "source": "lists/geo/countries.json" },
    "first_name": { "type": "model", "source": "models/tr_TR/first_names.ngram", "generation": { "mode": "word" } },
    "last_name": { "type": "model", "source": "models/tr_TR/last_names.ngram", "generation": { "mode": "word" } }
  },
  "output": {
    "id": "{{ user_id }}",
    "name": "{{ first_name }} {{ last_name }}",
    "country": "{{ country.name }} ({{ country.code }})",
    "phone_prefix": "{{ country.phone_prefix }}"
  }
}

Every field that references country sees the same drawn record — coherence comes free from naming the value once in let.

Deep Dive: the full envelope semantics (PEL threading into step params, variant weighting rules, registration) are in the PGDL generators guide; composition patterns (conditional routing, partitioning, linking) are catalogued in Generator Capabilities.


Comparison Table ​

FeatureLogicListModelStatisticalEvent SequenceComposition
Data sourceAlgorithmFile/InlineTrained modelDistribution specEvent specOther generators
Locale-specificNoOptionalYesNoNoDepends on parts
Output varietyControlledFiniteInfiniteDistribution-shapedStructuredControlled
CorrectnessGuaranteedGuaranteedStatisticalStatisticalOrder guaranteedDepends on parts
Output shapeScalarScalar or objectStringNumber/categoryObject (event → datetime)String or object
SpeedFastestVery fastFastVery fastFastDepends on depth
Training neededNoNoYesNoNoNo
Typical useIDs, numbersCodes, enumsNames, textDistributions, correlationsTimelinesEmails, addresses, coherent records

Cross-Cutting Features: Consistency & Linking ​

Consistency ​

Built (generator level):

  • Within one invocation — name a value once as a let step and reference it from many places; it is drawn once, so every reference agrees (see the user_profile example above).
  • Keyed generation — the deterministic word_for(key) primitive: same key, same output, across invocations.

Target design (cloud platform): the consistency: {enabled, key, scope} block with table/entity/global scopes, and automatic PK→FK propagation with format-preserving encryption (FPE), operate across rows and tables — that is platform orchestration over the keyed primitive, not a generator feature. See Generator Capabilities and Advanced Concepts.

Linking ​

Coherent related columns ("Ankara" must come with "Turkey" and "06100") are a built composition pattern, not a type:

  • Coherent by selection — a list over an asset of objects picks one whole record; its fields always agree.
  • Coherent by construction — a composition output computes fields over shared let steps (e.g. bonus, tax, net_income all derived from one drawn salary).

Target design (cloud platform): foreign-key linking across tables (referential integrity, cross-table aggregates) needs the cross-table row set and is deferred to the cloud layer — see Generator Capabilities and Advanced Concepts.


Folded types: template & linked (migration) ​

Earlier drafts had seven types. Two were folded into composition because they were redundant; the engine now errors with a migration hint when it sees them:

  • {"type": "template", …} → "the 'template' type was folded into composition: author a composition with a body template (optionally weighted variants); operations become PEL calls in the body"
  • {"type": "linked", …} → "the 'linked' type was folded: use a list over an asset to pick one coherent record, or a composition output to compute coherent fields"

template → composition body / variants ​

Template was string composition: a pattern (or weighted variants) with placeholders and an operations post-pipeline.

json
// before (rejected by the engine)
{ "type": "template", "pattern": "{{product_name}}", "operations": ["slugify", "lowercase"] }

// after
{
  "name": "@acme/web:slug",
  "let": { "product_name": { "use": "@acme/catalog:product_name" } },
  "body": "{{ lowercase(slugify(product_name)) }}"
}

Weighted template variants carry over directly — composition variants are the same [{ "pattern": …, "weight": … }] shape.

linked → list over an object asset, or composition output/let ​

Linked drew multiple coherent columns from one selection (source) or from computed rules.

  • source form (pick one record, fields agree) → a list over the object asset inside a composition let, with an output exposing the fields — see Accessing Nested Properties above.
  • rules form (mathematically related fields) → let steps + computed PEL, exposed via output:
json
{
  "name": "@acme/hr:financials",
  "let": {
    "salary": { "type": "statistical", "mode": "continuous", "distribution": "lognormal", "params": { "mu": 10, "sigma": 0.5 } },
    "bonus": { "computed": "round(salary * uniform(0.05, 0.20), 2)" },
    "tax": { "computed": "round((salary + bonus) * 0.25, 2)" }
  },
  "output": {
    "salary": "{{ salary }}",
    "bonus": "{{ bonus }}",
    "tax": "{{ tax }}",
    "net_income": "{{ salary + bonus - tax }}"
  }
}

Phony Cloud — Documentation & Specification