Generator-level capabilities & the platform boundary
A generator is a pure function: (seed, index, assets, bindings) → value. It knows nothing about tables, rows, foreign keys, or the source database. Everything a generator needs to do — including conditional routing, partitioned distributions, and coherent multi-field records — is expressed with the core generators + composition, never with a new generator type.
This page draws the line between what belongs at the generator level (and how we already cover it) and what belongs above it (the cloud platform — masking existing production data, referential integrity across tables, orchestration), which is deliberately deferred.
Status
The composition patterns below are built. The platform-level items in the last sections are deferred to the cloud product; they are catalogued here so we don't mistake them for generator features.
Principle: no new type for what composition already does
We folded template and linked into composition because they were redundant. The same test gates every proposed generator: if the core generators + composition + PEL already express it, it is a pattern, not a new type. Conditional routing and partitioning fail that test — they are patterns.
Conditional routing & partitioning — composition patterns
A value chosen by a sibling field
gender → title — a computed let with PEL if. No sub-generator is drawn; the value is selected directly (lazy):
"let": {
"gender": { "type": "list", "values": ["male", "female"] },
"title": { "computed": "if(gender == 'male', 'Mr', 'Ms')" }
}A distribution chosen by a sibling field (partitioning)
salary by department — draw each candidate distribution as a let, then select with if. Eager (every branch is drawn, one is kept), which is fine for a handful of branches and deterministic either way (each let has its own derived seed):
"let": {
"dept": { "type": "list", "values": ["eng", "sales"] },
"eng_salary": { "type": "statistical", "mode": "continuous", "distribution": "normal", "params": { "mean": 160000, "stddev": 20000 } },
"sales_salary": { "type": "statistical", "mode": "continuous", "distribution": "normal", "params": { "mean": 90000, "stddev": 15000 } },
"salary": { "computed": "if(dept == 'eng', eng_salary, sales_salary)" }
}Many categories — data, not branches
For a high-cardinality partition the branch is data, not code: key a map asset by the category and pick with list key (or where). No N-way if; the partition lives in the asset.
"let": {
"country": { "type": "list", "source": "geo.countries" },
"address": { "type": "list", "source": "address.formats", "key": "{{ country }}" }
}Coherent multi-field records (linking)
A composition output emits an object whose fields are computed over shared let steps — coherent by construction (the former linked rules). A list over an asset of objects returns one whole record — coherent by selection (the former linked source).
N items per parent (repetition)
Producing an array of N sub-values — the Faker ->count(n) / "N line items per order" shape — is now expressible at the generator level with a repeatlet step: { "repeat": { "count", "of", "unique"? } }. count is a literal, an inclusive "a-b" range, or a PEL expression over the locals; of is a sub-generator whose params interpolate against the parent (so items can reference it); unique: true de-duplicates. Each item is seeded deterministically (salted by step name + ordinal), and of may itself be a composition, so repetition nests. See Generators → N items per parent.
Competitor map (Tonic)
Tonic is primarily a masking tool — many of its "generators" transform existing production values. Its generation-side features map onto our composition model; its masking-side features map onto our cloud platform.
| Tonic feature | Phony equivalent | Where |
|---|---|---|
| Conditional generator | composition let + if | generator (built) |
| Partitioning (distribution per category) | eager let + if, or list key/where | generator (built) |
| Linking (coherent columns) | composition output / list over an object asset | generator (built) |
| Algebraic (derived column) | composition computed | generator (built) |
| Text composition | composition body | generator (built) |
| Categorical / Custom Categorical | list / statistical categorical | generator (built) |
| Sequential integer | logic.sequence | generator (built) |
| Continuous / multivariate | statistical continuous / multivariate | generator (built) |
| Differential privacy (continuous) | statistical differential_privacy | generator (built) |
| Uniqueness | the unique modifier | generator (built) |
| Format-preserving encryption | — (mask a real value, preserve format) | platform (deferred) |
| Composite masking (JSON/XML/CSV/regex) | — (mask part of an existing structured value) | platform (deferred) |
| PK→FK propagation | keyed mapping (word_for) + cross-table application | platform (deferred) |
| Cross-table sum / aggregates | — (needs the cross-table row set) | platform (deferred) |
| Categorical differential privacy | — (needs learned source frequencies) | platform (deferred) |
| Privacy ranking / data-free tag | generator metadata | platform (deferred) |
Deferred to the cloud platform (above the generator)
These are real, but they are not generator features — each needs an input value, the source-data distribution, or the cross-table row set:
- Format-preserving encryption (FPE) — deterministically transform a real value while preserving its format (numeric stays numeric, length preserved). The core anonymisation primitive for DB-Sync; meaningless in pure generation, where there is no input value.
- Composite / structured masking — apply a sub-generator to part of an existing JSON/XML/CSV/regex value.
- PK/FK consistency under masking — apply one consistent (keyed) generator to a primary key and propagate it to every foreign key so joins survive. Our
word_for(key)keyed mapping is the primitive; the propagation is platform orchestration. - Cross-table aggregates — e.g. an order total summed from its line items.
- Categorical differential privacy & distribution fitting — DP noise and per-partition distributions learned from source data. In pure generation the distribution is author-specified, so there is nothing to privatise.
- Privacy ranking & data-free classification — a per-generator privacy score and the "data-free ⇒ differentially private" classification, surfaced when masking.
Ergonomics since built
A multi-way switch PEL function — switch(subject, key: value, …, default: value) — is now built, so many-branch value selection reads better than nested if. Data-driven list key/where remains the right tool when the partition is high-cardinality.