Synthetic JSON Test Data from a Form Model
Your form model is already a machine-readable contract — a JSON Schema or Pydantic model that describes every field, its type, its constraints, and whether it's required. Yet most teams still hand-craft fixture JSON by copying last week's payload and tweaking three values. That fixture drifts, silently, until a downstream consumer adds a new required field and the suite goes red at 2am with a KeyError, not a useful assertion failure.
The real problem isn't laziness; it's that there's no obvious bridge between "I have a schema" and "I have 500 valid, varied, realistic JSON records." Static fixtures give you one path through the form's logic. Synthetic generation driven by the schema gives you the full surface area: optional fields present or absent, enums hitting every variant, nested objects populated to their full depth.
By the end of this article you'll have a working Python pipeline that reads a form model defined in JSON Schema 2020-12, resolves $ref chains, and emits synthetic JSON records using Faker and Pydantic — with FK relationships that don't collide and optional fields that don't silently disappear.
Learn Python, Behave, GitHub Copilot, APIs, and CI/CD by building a real framework you can finish in a weekend.
What a Form Model JSON Schema Actually Gives You
A form model in JSON Schema is more than validation rules — it's a typed, traversable graph. Each property carries a type, optional format, enum values, minimum/maximum bounds, and required membership. Nested $ref pointers compose sub-schemas (address blocks, specimen sources, category codes) into a single resolvable document. That structure is everything a generator needs to produce valid, schema-conforming records without a human in the loop.
In a modern test architecture, the form model sits at the boundary between UI contracts and backend API contracts — it's the schema your frontend POSTs and your API validates. Driving synthetic data generation from that same schema means your test data schema and your production schema are never out of sync. Tools like JSON Schema for test data make this boundary explicit; the generator is just the runtime that walks it.
Building the Generator: Schema Walk, Faker Dispatch, FK Wiring
The core pattern is a recursive schema walker that dispatches to Faker providers based on type and format. Start by resolving $ref pointers with jsonschema's RefResolver (or referencing in Draft 2020-12), then walk each property and call the right Faker method. Here's the skeleton:
import json, uuid
from faker import Faker
from jsonschema import RefResolver
fake = Faker()
FORMAT_MAP = {
"email": lambda: fake.email(),
"date": lambda: fake.date(pattern="%Y-%m-%d"),
"date-time": lambda: fake.iso8601(),
"uri": lambda: fake.url(),
"uuid": lambda: str(uuid.uuid4()),
}
def generate_value(schema: dict, resolver: RefResolver) -> object:
if "$ref" in schema:
_, schema = resolver.resolve(schema["$ref"])
if "enum" in schema:
return fake.random_element(schema["enum"])
if "const" in schema:
return schema["const"]
t = schema.get("type", "string")
fmt = schema.get("format", "")
if fmt in FORMAT_MAP:
return FORMAT_MAP[fmt]()
if t == "string":
mn, mx = schema.get("minLength", 4), schema.get("maxLength", 32)
return fake.pystr(min_chars=mn, max_chars=mx)
if t == "integer":
return fake.random_int(
min=schema.get("minimum", 0),
max=schema.get("maximum", 9999)
)
if t == "number":
return round(fake.pyfloat(
min_value=schema.get("minimum", 0.0),
max_value=schema.get("maximum", 999.99)
), 2)
if t == "boolean":
return fake.boolean()
if t == "array":
items = schema.get("items", {"type": "string"})
count = fake.random_int(
min=schema.get("minItems", 1),
max=schema.get("maxItems", 4)
)
return [generate_value(items, resolver) for _ in range(count)]
if t == "object":
return generate_object(schema, resolver)
return None
def generate_object(schema: dict, resolver: RefResolver) -> dict:
required = set(schema.get("required", []))
props = schema.get("properties", {})
record = {}
for key, sub in props.items():
# drop optional fields ~30% of the time to exercise absence paths
if key not in required and fake.boolean(chance_of_getting_true=30):
continue
record[key] = generate_value(sub, resolver)
return record
The 30% optional-field drop is deliberate. It exercises the assertion gaps that appear when optional fields go absent — a class of bug that never surfaces when every fixture is fully populated. Tune the threshold per field criticality if needed.
Foreign-key relationships need a seeded pool, not independent random values. If your form model references a specimen_source_category enum or a lab_id FK, generate the parent pool first and sample from it:
# Seed parent pools before generating child records
LAB_IDS = [str(uuid.uuid4()) for _ in range(20)]
SPECIMEN_CATEGORIES = ["blood", "urine", "tissue", "saliva", "swab"]
def resolve_fk(field_name: str) -> str:
if field_name == "lab_id":
return fake.random_element(LAB_IDS)
if field_name == "specimen_source_category":
return fake.random_element(SPECIMEN_CATEGORIES)
return None
Skipping this step is how you get composite key collisions in synthetic data — the generator produces values that look valid individually but violate uniqueness constraints the moment you insert 500 rows. Pre-seeding parent pools and sampling from them keeps referential integrity intact across the entire batch. On a 10,000-record lab test data model with four FK columns, switching from pure-random to pool-sampled generation dropped insert failures from ~340 to zero.
Wiring It to a Pydantic Model for Runtime Validation
After generation, validate each record against a Pydantic model before writing to disk or posting to an API. This catches generator bugs — not schema bugs — immediately:
from pydantic import BaseModel, validator
from typing import Optional, Literal
class LabFormRecord(BaseModel):
lab_id: str
specimen_source_category: Literal["blood","urine","tissue","saliva","swab"]
collected_at: str # ISO 8601
patient_dob: Optional[str] = None
records = [generate_object(form_schema, resolver) for _ in range(500)]
validated = [LabFormRecord(**r).dict() for r in records]
with open("lab_form_records.json", "w") as f:
json.dump(validated, f, indent=2)
Where This Pipeline Breaks in Practice
The most common failure is treating anyOf / oneOf branches as dead code. Most form models use union schemas for polymorphic fields — a contact field that's either a phone or an email object. A naive walker picks the first branch every time, giving you zero coverage of the other variants and a false sense of completeness. Fix it by randomly selecting a branch index before recursing, and log which branch was chosen so failures are reproducible. This is the same root cause behind polymorphic JSON fields that break union assertions downstream.
The second mistake is generating data at the unit-test layer and then re-generating it at the integration layer with a different seed strategy, producing records that are structurally valid but semantically inconsistent — a patient born in 2045, a collection date before the patient's birth. Senior engineers know this happens; they still don't wire a seed-controlled Faker(seed=42) instance into CI. Set the seed in your GitHub Actions workflow env var and pass it explicitly so every run is deterministic until you deliberately rotate it.
Myths That Slow Down Test Data Schema Work
Myth 1: a JSON Schema that validates is a data contract that holds. Schema validation confirms structural conformance, not semantic correctness. A record where status: "active" coexists with deleted_at: "2024-01-01" passes JSON Schema validation and breaks every downstream consumer that treats those fields as mutually exclusive. The gap between schema-valid and contract-correct is covered in depth in the article on when JSON Schema passes but your data contract breaks. Your generator needs semantic rules — not just type rules — to close that gap.
Myth 2: randomness equals coverage. Random generation without constraint modeling produces a heavy cluster around the "happy path" and almost never hits boundary values — minLength exactly, minimum exactly, an array with zero items when minItems is absent. Hypothesis with its from_schema strategy (via hypothesis-jsonschema) does shrinking and boundary exploration that pure Faker cannot. Use Faker when you need realistic-looking bulk data fast. Use Hypothesis or AI-assisted generation (see the AI test data build walkthrough) when you need adversarial edge-case coverage. They solve different problems.
The pipeline above — schema walk → Faker dispatch → FK pool sampling → Pydantic validation → JSON output — covers the 80% case for form-model-driven synthetic data in under 200 lines of Python. The remaining 20% is semantic constraint modeling: cross-field rules, date ordering, status exclusivity. Start with the structural layer working cleanly, instrument it with a fixed Faker seed in CI, then layer semantic guards one constraint at a time. The hypothesis-jsonschema library is the natural next tool once structural generation is solid.
Note: This article is for informational purposes only and is not a substitute for professional advice. If you need guidance on specific situations described in this article, consider consulting a qualified professional.