Skip to content

Determinism ​

Phony data is a pure function of (schema, seed, reference-time). Same inputs → byte-identical output, every run, every machine, every language runtime. There is no hidden randomness and no wall-clock dependency.

Why it matters ​

  • Safe to commit. A test fixture that regenerates identically can live in your repo and CI — no flaky data.
  • Reviewable. A change to the schema produces a reviewable diff in the data.
  • Portable. The same schema in the CLI, an OSS library, or the Cloud yields the same rows — so you can develop against one and deploy against another.

How it works ​

The seed model: (root_seed, key, index). Every value Phony produces comes from a seed derived from three inputs by one portable recipe:

seed = mix64( root_seed  XOR  fnv1a64(key)  XOR  index × 0x9E3779B97F4A7C15 )
  • root_seed — the one number you choose. Change it and the whole dataset changes, coherently.
  • key — a name, not a position: a let step's name, or the JSON path of an output field ("a", "items.0.b"). It's hashed with FNV-1a/64.
  • index — the zero-based ordinal of the invocation (row 0, row 1, …), which sequence and cycle read directly. It's spread across the word by a multiply with the golden-ratio constant, then the whole thing runs through the SplitMix64 finalizer (mix64).

Each cell gets its own seed, so a PHP/JS/Python runtime can reproduce the exact same draws — and, crucially, the seeds are independent.

Positional independence. Because key is a name, the seed for one field never depends on its neighbours. Add a column to a schema, insert a let, or reorder fields, and every other value stays byte-identical — only the new field is new. This is what turns a schema change into a small, reviewable data diff instead of a total reshuffle, and it's why you can grow a committed fixture without churning the rows already in it.

Row consistency (coherence). Within one generator, references are memoized: a let binds a name to one value, and every reference to that name reads the same value. So email and full_name built from the same first agree — the record holds together. (Coherence is a property of the let, not of any table.)

Sub-stream salts. Independence goes all the way down. Each inline generator in a template draws a fresh sub-seed per occurrence, so two {{number}}s in one string differ. Each string leaf of an output is salted by its JSON path, so two sibling fields with the identical template still draw independently rather than echoing each other.

The frozen clock. "Now" never comes from the system clock — it comes from an injected reference instant. now(), today(), age(), and relative ranges like -1year or now+7days all resolve against it:

rust
let registry = Registry::with_builtins()
    .with_reference_time(1_704_067_200); // 2024-01-01T00:00:00Z

So a schema that says "created in the last year" produces the same dates tomorrow as it did today. (The CLI exposes this as --reference-time.)

Seeing it: same seed, same rows ​

A one-line generator, run twice, byte-for-byte:

console
$ phony generate --use "@demo:user" --package demo --seed 42 -n 4 -f jsonl
"ayse@example.com"
"deniz@example.com"
"zeynep@example.com"
"must@example.com"
$ phony generate --use "@demo:user" --package demo --seed 42 -n 4 -f jsonl   # again
"ayse@example.com"
"deniz@example.com"
"zeynep@example.com"
"must@example.com"

Same seed → the same four emails, on any machine, in any runtime. Change --seed and you get four different — but equally reproducible — emails.

Why it matters: agent-written tests ​

Determinism is what makes generated data safe to commit. Consider a test fixture an agent writes:

  • Non-deterministic tooling produces a fresh dataset each run, so the fixture either can't be asserted against (assertEquals on a random name is hopeless) or has to be regenerated and re-reviewed on every CI run — flaky by construction.
  • Phony, given (schema, seed, reference-time), regenerates the same rows forever. The agent asserts against concrete values, the fixture lives in the repo, the diff on a schema change is legible, and CI never flakes on the data.

One contract, many runtimes ​

Everything portable — the SplitMix64 PRNG, FNV-1a hashing, seed derivation, PEL semantics, generator output, and the whole n-gram train → serialize → generate pipeline — is pinned by a frozen conformance corpus (phony-core/conformance/vectors.json): a language port is correct exactly when it reproduces every vector byte-for-byte, and any engine drift fails CI. Even Unicode casing is part of the contract — lowercasing maps İ (U+0130, dotted capital I) to a plain i, no combining dot, in every runtime, so tokenization (and therefore every model and its output) never diverges between Rust and PHP.

What changes the output ​

Only three things: the schema, the seed, and the reference time. Change any one and you get a different — but equally valid and equally reproducible — dataset. Change none and you get the same data forever.

This cross-runtime reproducibility is the core promise that lets a portable .phony package behave identically everywhere. See the Execution Model for how the runtimes line up, and How generation works for the mechanics.

Phony Cloud — Documentation & Specification