Skip to content

Generator-level capabilities & the platform boundary ​

A generator is a pure function: (seed, index, assets, bindings) → value. It knows nothing about tables, rows, foreign keys, or the source database. Everything a generator needs to do — including conditional routing, partitioned distributions, and coherent multi-field records — is expressed with the core generators + composition, never with a new generator type.

This page draws the line between what belongs at the generator level (and how we already cover it) and what belongs above it (the cloud platform — masking existing production data, referential integrity across tables, orchestration), which is deliberately deferred.

Status

The composition patterns below are built. The platform-level items in the last sections are deferred to the cloud product; they are catalogued here so we don't mistake them for generator features.

Principle: no new type for what composition already does ​

We folded template and linked into composition because they were redundant. The same test gates every proposed generator: if the core generators + composition + PEL already express it, it is a pattern, not a new type. Conditional routing and partitioning fail that test — they are patterns.

Conditional routing & partitioning — composition patterns ​

A value chosen by a sibling field ​

gender → title — a computed let with PEL if. No sub-generator is drawn; the value is selected directly (lazy):

json
"let": {
  "gender": { "type": "list", "values": ["male", "female"] },
  "title":  { "computed": "if(gender == 'male', 'Mr', 'Ms')" }
}

A distribution chosen by a sibling field (partitioning) ​

salary by department — draw each candidate distribution as a let, then select with if. Eager (every branch is drawn, one is kept), which is fine for a handful of branches and deterministic either way (each let has its own derived seed):

json
"let": {
  "dept":         { "type": "list", "values": ["eng", "sales"] },
  "eng_salary":   { "type": "statistical", "mode": "continuous", "distribution": "normal", "params": { "mean": 160000, "stddev": 20000 } },
  "sales_salary": { "type": "statistical", "mode": "continuous", "distribution": "normal", "params": { "mean": 90000,  "stddev": 15000 } },
  "salary":       { "computed": "if(dept == 'eng', eng_salary, sales_salary)" }
}

Many categories — data, not branches ​

For a high-cardinality partition the branch is data, not code: key a map asset by the category and pick with list key (or where). No N-way if; the partition lives in the asset.

json
"let": {
  "country": { "type": "list", "source": "geo.countries" },
  "address": { "type": "list", "source": "address.formats", "key": "{{ country }}" }
}

Coherent multi-field records (linking) ​

A composition output emits an object whose fields are computed over shared let steps — coherent by construction (the former linked rules). A list over an asset of objects returns one whole record — coherent by selection (the former linked source).

N items per parent (repetition) ​

Producing an array of N sub-values — the Faker ->count(n) / "N line items per order" shape — is now expressible at the generator level with a repeatlet step: { "repeat": { "count", "of", "unique"? } }. count is a literal, an inclusive "a-b" range, or a PEL expression over the locals; of is a sub-generator whose params interpolate against the parent (so items can reference it); unique: true de-duplicates. Each item is seeded deterministically (salted by step name + ordinal), and of may itself be a composition, so repetition nests. See Generators → N items per parent.

Competitor map (Tonic) ​

Tonic is primarily a masking tool — many of its "generators" transform existing production values. Its generation-side features map onto our composition model; its masking-side features map onto our cloud platform.

Tonic featurePhony equivalentWhere
Conditional generatorcomposition let + ifgenerator (built)
Partitioning (distribution per category)eager let + if, or list key/wheregenerator (built)
Linking (coherent columns)composition output / list over an object assetgenerator (built)
Algebraic (derived column)composition computedgenerator (built)
Text compositioncomposition bodygenerator (built)
Categorical / Custom Categoricallist / statistical categoricalgenerator (built)
Sequential integerlogic.sequencegenerator (built)
Continuous / multivariatestatistical continuous / multivariategenerator (built)
Differential privacy (continuous)statistical differential_privacygenerator (built)
Uniquenessthe unique modifiergenerator (built)
Format-preserving encryption— (mask a real value, preserve format)platform (deferred)
Composite masking (JSON/XML/CSV/regex)— (mask part of an existing structured value)platform (deferred)
PK→FK propagationkeyed mapping (word_for) + cross-table applicationplatform (deferred)
Cross-table sum / aggregates— (needs the cross-table row set)platform (deferred)
Categorical differential privacy— (needs learned source frequencies)platform (deferred)
Privacy ranking / data-free taggenerator metadataplatform (deferred)

Deferred to the cloud platform (above the generator) ​

These are real, but they are not generator features — each needs an input value, the source-data distribution, or the cross-table row set:

  • Format-preserving encryption (FPE) — deterministically transform a real value while preserving its format (numeric stays numeric, length preserved). The core anonymisation primitive for DB-Sync; meaningless in pure generation, where there is no input value.
  • Composite / structured masking — apply a sub-generator to part of an existing JSON/XML/CSV/regex value.
  • PK/FK consistency under masking — apply one consistent (keyed) generator to a primary key and propagate it to every foreign key so joins survive. Our word_for(key) keyed mapping is the primitive; the propagation is platform orchestration.
  • Cross-table aggregates — e.g. an order total summed from its line items.
  • Categorical differential privacy & distribution fitting — DP noise and per-partition distributions learned from source data. In pure generation the distribution is author-specified, so there is nothing to privatise.
  • Privacy ranking & data-free classification — a per-generator privacy score and the "data-free ⇒ differentially private" classification, surfaced when masking.

Ergonomics since built ​

A multi-way switch PEL function — switch(subject, key: value, …, default: value) — is now built, so many-branch value selection reads better than nested if. Data-driven list key/where remains the right tool when the partition is high-cardinality.

Phony Cloud — Documentation & Specification