Skip to content

Snapshots Architecture ​

Phony Snapshots provide point-in-time captures of synchronized databases, enabling instant rollback, environment provisioning, and data versioning. Like sync, snapshots are local-first: snapshot data lives in the customer's storage (local disk or their own S3/GCS bucket); Phony Cloud stores only snapshot metadata.

Implementation status

Everything on this page is target design — the snapshot engine and the control plane are not built yet.

Architecture Overview ​

The agent captures, stores, and restores snapshots inside the customer's infrastructure. The control plane holds the catalog: snapshot id, name, schema hash, sizes, chunk checksums, lineage (base → incrementals) — never the data.

┌─────────────────────────────────────────────────────────────────────────┐
│                    SNAPSHOTS ARCHITECTURE                                │
├─────────────────────────────────────────────────────────────────────────┤
│                                                                          │
│  CUSTOMER INFRASTRUCTURE (data plane)        PHONY CLOUD (control plane)│
│  ════════════════════════════════════        ═══════════════════════════│
│                                                                          │
│  ┌────────────────────────────────────┐                                 │
│  │         PHONY AGENT                │      ┌─────────────────────────┐│
│  │                                    │      │  SNAPSHOT CATALOG       ││
│  │  ┌─────────┐ ┌─────────┐ ┌───────┐│ meta │  (metadata only)        ││
│  │  │ Capture │ │ Storage │ │Restore││─────►│                         ││
│  │  │ Engine  │ │ Manager │ │Engine ││ out- │  • snapshot id, name    ││
│  │  │         │ │         │ │       ││ bound│  • schema hash          ││
│  │  │• Stream │ │• Local/ │ │• Par- ││ only │  • sizes, row counts    ││
│  │  │• Comp-  │ │  S3/GCS │ │  allel││      │  • chunk checksums      ││
│  │  │  ress   │ │  (YOURS)│ │• Ver- ││      │  • lineage (base→incr)  ││
│  │  │• Chunk  │ │• Encrypt│ │  ify  ││      │  • retention policy     ││
│  │  └────┬────┘ └────┬────┘ └───┬───┘│      └─────────────────────────┘│
│  │       └───────────┼──────────┘    │                                 │
│  └───────────────────┼───────────────┘       NEVER snapshot contents.  │
│                      ▼                                                  │
│  ┌────────────────────────────────────┐                                 │
│  │       CUSTOMER STORAGE             │                                 │
│  │                                    │                                 │
│  │  • Local disk (dev, single host)   │                                 │
│  │  • Customer S3 / GCS bucket        │                                 │
│  │  • Lifecycle/tiering = THEIR       │                                 │
│  │    bucket policy (we orchestrate)  │                                 │
│  └────────────────────────────────────┘                                 │
│                                                                          │
└─────────────────────────────────────────────────────────────────────────┘

Why snapshots matter in the agent era ​

A snapshot is a spawn point. Once an anonymized snapshot exists in your storage, the agent can materialize it into as many ephemeral environments as you want — a fresh database per CI run, per E2E suite, per coding-agent sandbox — and throw them away afterwards. Environments spawned from snapshots are never metered: only tracked snapshots are tiered (see product page). In a world where AI agents multiply test runs, this is deliberately the cheap, unlimited path.


Snapshot Types ​

Full Snapshot ​

Complete database capture:

┌─────────────────────────────────────────────────────────────────────────┐
│                    FULL SNAPSHOT                                         │
├─────────────────────────────────────────────────────────────────────────┤
│                                                                          │
│  Database                        Snapshot (in YOUR storage)             │
│  ════════                        ═══════════════════════════            │
│                                                                          │
│  ┌─────────────────┐            ┌─────────────────────────────────────┐│
│  │ users (500MB)   │──────────►│ snapshot_20260131_100000.tar.gz     ││
│  │ orders (2GB)    │   Stream   │                                     ││
│  │ products (100MB)│   + Gzip   │ Contents:                           ││
│  │ logs (5GB)      │            │ ├── manifest.json (metadata)        ││
│  └─────────────────┘            │ ├── schema.sql (DDL)                ││
│                                  │ ├── users.csv.gz (500MB → 80MB)    ││
│  Total: 7.6 GB                   │ ├── orders.csv.gz (2GB → 300MB)    ││
│                                  │ ├── products.csv.gz (100MB → 15MB) ││
│                                  │ └── logs.csv.gz (5GB → 600MB)      ││
│                                  │                                     ││
│                                  │ Compressed: ~1 GB                   ││
│                                  └─────────────────────────────────────┘│
│                                                                          │
└─────────────────────────────────────────────────────────────────────────┘

Incremental Snapshot ​

Only changes since last snapshot:

┌─────────────────────────────────────────────────────────────────────────┐
│                    INCREMENTAL SNAPSHOT                                  │
├─────────────────────────────────────────────────────────────────────────┤
│                                                                          │
│  Base Snapshot                Incremental                               │
│  (Day 1)                      (Day 2)                                   │
│  ═══════════                  ═══════════                               │
│                                                                          │
│  ┌─────────────────┐          ┌─────────────────┐                      │
│  │ Full DB dump    │          │ Only changes:   │                      │
│  │ 7.6 GB          │          │ • 50 new users  │                      │
│  │ Compressed: 1GB │          │ • 200 new orders│                      │
│  └─────────────────┘          │ • 10 updates    │                      │
│                               │                 │                      │
│                               │ Size: 5 MB      │                      │
│                               └─────────────────┘                      │
│                                                                          │
│  RESTORE PROCESS (runs in the agent):                                   │
│  ════════════════════════════════════                                   │
│  1. Restore base snapshot                                               │
│  2. Apply incremental changes in order                                  │
│  3. Verify checksums (against catalog metadata)                         │
│                                                                          │
└─────────────────────────────────────────────────────────────────────────┘

The base → incremental lineage graph is the piece the control plane tracks — it is what lets the dashboard offer "restore to Tuesday 15:30" without ever touching the data.


Storage Architecture ​

Bring your own storage ​

The Storage Manager writes to a backend the customer configures on the agent:

BackendUse case
Local diskDevelopment, single-host setups, air-gapped
Customer S3 bucketTeams; durable, shared across agents
Customer GCS bucketSame, on GCP

Phony never holds the bucket credentials in the cloud — they are agent-local configuration, like database credentials.

Chunked Storage ​

Large snapshots are split into chunks for parallel upload/download. Chunk paths point at the customer's bucket; the checksums live in the catalog so restores can verify integrity:

json
{
  "snapshot_id": "snap_abc123",
  "chunks": [
    {
      "index": 0,
      "table": "users",
      "offset": 0,
      "size": 104857600,
      "checksum": "sha256:abc123...",
      "path": "s3://customer-bucket/phony/snapshots/snap_abc123/chunk_000.gz"
    },
    {
      "index": 1,
      "table": "users",
      "offset": 104857600,
      "size": 104857600,
      "checksum": "sha256:def456...",
      "path": "s3://customer-bucket/phony/snapshots/snap_abc123/chunk_001.gz"
    }
  ]
}

Encryption ​

Snapshots are encrypted at rest in the customer's storage, with keys the customer controls:

┌─────────────────────────────────────────────────────────────────────────┐
│                    ENCRYPTION MODEL                                      │
├─────────────────────────────────────────────────────────────────────────┤
│                                                                          │
│  1. ENVELOPE ENCRYPTION (agent-side)                                    │
│  ═══════════════════════════════════                                    │
│                                                                          │
│  ┌─────────────────┐    ┌─────────────────┐    ┌─────────────────┐    │
│  │  Master Key     │    │  Data Key       │    │  Snapshot       │    │
│  │  (CUSTOMER's    │───►│  (per snapshot) │───►│  Chunks         │    │
│  │  KMS or local   │    │                 │    │                 │    │
│  │  keyfile)       │    │  Encrypted with │    │  Encrypted with │    │
│  │  Never leaves   │    │  master key     │    │  data key       │    │
│  │  customer infra │    │                 │    │                 │    │
│  └─────────────────┘    └─────────────────┘    └─────────────────┘    │
│                                                                          │
│  2. KEY HIERARCHY                                                       │
│  ═════════════════                                                      │
│                                                                          │
│  Customer Master Key (their KMS / keyfile)                              │
│       │                                                                 │
│       ├── Snapshot Key (snap_001) ─► Encrypted chunks                  │
│       ├── Snapshot Key (snap_002) ─► Encrypted chunks                  │
│       └── Snapshot Key (snap_003) ─► Encrypted chunks                  │
│                                                                          │
│  BENEFIT: Rotate the master key without re-encrypting all snapshots.   │
│  Alternatively: rely on bucket-native SSE (SSE-KMS) — it's your bucket. │
│                                                                          │
└─────────────────────────────────────────────────────────────────────────┘

Phony Cloud stores no keys — even a control-plane breach exposes only the catalog.

Replication and tiering ​

Cross-region replication and cold-storage tiering are your bucket policy, not Phony infrastructure: enable S3 Cross-Region Replication or GCS dual-region on your bucket, attach lifecycle rules (Standard → IA → Glacier) as you see fit. Phony's role is orchestration — retention policies in the control plane tell the agent which snapshots to keep, prune, or verify, and the catalog records where each one lives.


Lifecycle Management ​

Retention Policies ​

Retention is defined in the control plane and enforced by the agent against your storage:

json
{
  "retention": {
    "policy": "tiered",
    "rules": [
      { "age": "< 7 days", "keep": "all" },
      { "age": "7-30 days", "keep": "daily" },
      { "age": "30-90 days", "keep": "weekly" },
      { "age": "> 90 days", "keep": "monthly" }
    ],
    "minimum_retention": "30 days",
    "legal_hold": false
  }
}

Storage-class transitions (Standard-IA, Glacier, Deep Archive) are configured as lifecycle rules on your own bucket; the agent tolerates restore-latency from cold tiers and surfaces it in job status.


Restore Operations ​

All restore work happens in the agent: it reads chunks from your storage, decrypts and decompresses them locally, and loads the target database. The control plane supplies the catalog entry and records the outcome.

Full Restore ​

┌─────────────────────────────────────────────────────────────────────────┐
│                    FULL RESTORE PROCESS (in the agent)                   │
├─────────────────────────────────────────────────────────────────────────┤
│                                                                          │
│  1. PREPARE                                                             │
│  ══════════                                                             │
│  • Fetch catalog entry (chunk list + checksums) from control plane      │
│  • Validate snapshot integrity (checksums)                              │
│  • Verify target database connection                                    │
│  • Check available disk space                                           │
│  • Acquire exclusive lock on target                                     │
│                                                                          │
│  2. RESTORE SCHEMA                                                      │
│  ═════════════════                                                      │
│  • Drop existing tables (if overwrite=true)                            │
│  • Execute DDL from schema.sql                                         │
│  • Disable foreign key constraints                                      │
│  • Disable triggers                                                     │
│                                                                          │
│  3. RESTORE DATA                                                        │
│  ════════════════                                                       │
│  • Download chunks from YOUR storage in parallel (8 concurrent)         │
│  • Decompress and decrypt (locally)                                     │
│  • Bulk load into tables (COPY command)                                │
│  • Order: FK dependencies respected                                    │
│                                                                          │
│  4. FINALIZE                                                            │
│  ═══════════                                                            │
│  • Enable foreign key constraints                                       │
│  • Enable triggers                                                      │
│  • Rebuild indexes                                                      │
│  • Update statistics (ANALYZE)                                         │
│  • Verify row counts; report result to control plane                    │
│                                                                          │
└─────────────────────────────────────────────────────────────────────────┘

Table-Level Restore ​

Restore specific tables only:

bash
phony snapshot restore snap_abc123 \
  --target staging \
  --tables users,orders \
  --mode merge  # or 'replace'

Point-in-Time Restore ​

Using incremental snapshots:

bash
phony snapshot restore \
  --target staging \
  --point-in-time "2026-01-30T15:30:00Z"

# The agent asks the control plane's lineage graph for:
# 1. Base snapshot before timestamp
# 2. Incrementals to apply
# then restores locally to the exact point.

Ephemeral Environments ​

The spawn-point pattern — the reason snapshots exist in the agent era:

bash
# CI: fresh database per run, destroyed after
phony env spawn --from snap_abc123 --ttl 2h
# → postgres://localhost:54xx/phony_env_9f3a  (agent-local)

# Coding-agent sandbox: each agent gets its own copy
phony env spawn --from snap_abc123 --name agent-sandbox-42

Spawned environments are agent-local materializations of a snapshot. They are unlimited at every tier and (optionally) invisible to the control plane beyond an aggregate count.


Comparison Modes ​

Diff and drift detection run in the data plane (the agent has the data); the resulting summaries are reported up and rendered in the dashboard.

Diff Between Snapshots ​

bash
phony snapshot diff snap_001 snap_002

# Output (this summary — not the rows — goes to the control plane):
# Table: users
#   Added: 50 rows
#   Modified: 12 rows
#   Deleted: 3 rows
#
# Table: orders
#   Added: 200 rows
#   Modified: 0 rows
#   Deleted: 0 rows

Schema Drift Detection ​

bash
phony snapshot schema-diff snap_001 current

# Output:
# Table: users
#   + Column: phone_verified (boolean)
#   ~ Column: email (varchar(100) → varchar(255))
#
# Table: audit_logs
#   + New table

Schema drift is pure metadata, so it also feeds the control plane's schema-hash tracking: the dashboard can flag "production schema changed since your last snapshot" without any data access.


Automation ​

Scheduled Snapshots ​

Schedules live in the control plane; the agent executes them:

json
{
  "schedule": {
    "full": {
      "cron": "0 2 * * 0",
      "description": "Weekly full snapshot on Sunday 2 AM"
    },
    "incremental": {
      "cron": "0 2 * * 1-6",
      "description": "Daily incremental Mon-Sat 2 AM"
    }
  },
  "triggers": {
    "before_sync": true,
    "before_migration": true
  }
}

Event-Driven Snapshots ​

Webhooks are dispatched by the control plane from snapshot metadata events:

json
{
  "webhooks": {
    "on_create": "https://yourapp.com/webhook/snapshot-created",
    "on_restore_start": "https://yourapp.com/webhook/restore-started",
    "on_restore_complete": "https://yourapp.com/webhook/restore-completed"
  }
}

Performance ​

Bounded by your hardware and your storage bandwidth — there is no Phony-side hop in the data path.

Operation1 GB10 GB100 GB1 TB
Create (full)~1 min~5 min~30 min~4 hours
Create (incremental)~10 sec~30 sec~2 min~15 min
Restore (warm storage)~2 min~10 min~1 hour~8 hours
Restore (cold tier, e.g. Glacier)+5 min+5 min+10 min+30 min

Optimization factors:

  • Parallel chunk processing (up to 16 workers)
  • Compression ratio (typically 6-10x for text data)
  • Bandwidth between agent and your storage backend
  • Target database write speed

Tier Limits ​

Environments spawned from snapshots are unlimited at every tier; only tracked snapshots are tiered. Storage cost is yours (it's your bucket), which is why size is not the axis Phony meters.

FeatureFREESTARTERTEAMBUSINESS
Tracked snapshots31050Unlimited
Retention orchestration7 days30 days90 daysCustom
Incremental-✓✓✓
Snapshot diff--✓✓
Point-in-time---✓
Ephemeral environmentsUnlimitedUnlimitedUnlimitedUnlimited

Phony Cloud — Documentation & Specification