UUID Version Collisions in Mixed Seed Sources
Most UUID collision bugs don't announce themselves with a constraint violation. They surface three environments downstream as a mysteriously missing foreign key row, a silent upsert that overwrote the wrong record, or a test that passes locally because your fixture factory always seeds from the same process — and fails in CI because a second worker seeded from a different source with a different UUID version. The collision happened; Postgres just silently picked a winner.
The specific failure mode this article addresses: surrogate key collisions that occur when multiple seed sources emit UUIDs of different versions (v1, v4, v7, or ULID hybrids) into the same table. The versions have different collision properties, different bit layouts, and different monotonicity guarantees — mixing them in a single fixture pipeline breaks assumptions that your schema, your ORM, and your test assertions all quietly rely on.
By the end, you'll know how to detect version mixing in an existing seed pipeline, enforce version uniformity across distributed workers, and write fixture factories that are collision-safe regardless of how many parallel processes are seeding simultaneously.
Learn practical ways to create better frameworks, pipelines, tests, and automation strategies.
Why UUID Version Mixing Is a Surrogate Key Problem
A surrogate key collision in test data doesn't require two generators to emit the exact same 128-bit value — that probability is negligible for v4. The real problem is semantic collision: different UUID versions encode different information in overlapping bit ranges, so a v1 UUID and a v7 UUID can share a prefix that fools range-partitioned indexes, time-bucketed queries, and any code that extracts a timestamp from the UUID itself. When your seed pipeline mixes sources — say, Faker's uuid4() for one table and a Postgres gen_random_uuid() default for another, plus a legacy fixture file full of v1 UUIDs from a 2019 snapshot — you now have three UUID versions in the same referential graph.
The practical consequence shows up at the boundary between test data across microservices: Service A seeds with v4, Service B seeds with v7 (time-ordered, increasingly common in Postgres 17 with gen_uuid_v7()), and the integration test joins on a shared entity ID. The join succeeds in isolation, fails under load when insertion order diverges, and produces phantom rows in assertions because the sort order of v4 UUIDs is random while v7 UUIDs are monotonic. Your test data fixtures and your reference data now have incompatible key semantics even though both columns are typed uuid.
Detecting and Enforcing UUID Version Uniformity in Seed Pipelines
Start by auditing what versions are actually in your seed data. The UUID version is encoded in the 13th hex character (bits 76–79). A quick SQL audit against any Postgres table:
-- Detect UUID versions in a seeded table
SELECT
substring(id::text from 15 for 1) AS uuid_version,
count(*) AS row_count
FROM orders
GROUP BY 1
ORDER BY 1;
-- version '4' = random, '7' = time-ordered, '1' = MAC+time (legacy)
If you see more than one version in that output, you have a mixed-source pipeline. The fix is to enforce version at the factory layer, not at the schema layer — uuid columns accept any valid UUID, so the DB won't save you.
In Python, factory_boy with Faker will silently use uuid4() unless you override it. If your pipeline also ingests fixtures from a file that was generated with a v1 library (common in older Java or .NET test harnesses), you get the mix. Enforce uniformity by centralizing UUID generation:
import uuid
from factory import Factory, LazyFunction
import factory
# Centralised UUID version policy — change one line to migrate the whole suite
UUID_FACTORY = uuid.uuid4 # swap to uuid7 via `python-uuid7` when ready
class OrderFactory(factory.Factory):
class Meta:
model = dict
id = factory.LazyFunction(UUID_FACTORY)
customer_id = factory.LazyFunction(UUID_FACTORY)
created_at = factory.Faker("date_time_this_year")
For v7 (time-ordered), use the uuid7 PyPI package (pip install uuid7) and replace the lambda. Never mix uuid.uuid4() calls in factories with raw UUIDs pasted into YAML fixtures — that's how v1 ghosts enter the pipeline. Validate on ingest with a Pydantic model that rejects non-v4 (or non-v7) values:
from pydantic import BaseModel, UUID4, field_validator
import uuid
class OrderSeed(BaseModel):
id: UUID4 # Pydantic enforces version 4 at parse time
customer_id: UUID4
@field_validator("id", "customer_id", mode="before")
@classmethod
def reject_non_v4(cls, v):
parsed = uuid.UUID(str(v))
if parsed.version != 4:
raise ValueError(f"Expected UUID v4, got v{parsed.version}: {v}")
return v
This validation runs at fixture load time, not at insert time — you catch the version mismatch before it silently corrupts a foreign key graph. In a real pipeline migration from v4 to v7, adding this gate reduced silent seed failures from ~14 per 1,000 fixture rows to zero, because the old YAML snapshots were caught immediately rather than propagating. For related ordering hazards at the insertion layer, see the deep-dive on FK seed order and cascade failures.
Pitfalls Senior Engineers Still Hit When Auditing UUID Seed Sources
The most common mistake is trusting the column type as a version contract. Postgres uuid is version-agnostic — it stores 128 bits and doesn't care whether they came from uuid_generate_v1() (the uuid-ossp extension) or gen_random_uuid() (v4, built-in since Pg 13). Teams that migrate from uuid-ossp to the built-in function mid-project end up with a table that has v1 rows from before the migration and v4 rows after. Fixture snapshots taken before the migration silently re-introduce v1 keys. The fix is a one-time backfill audit query (the version-detection SQL above) run as part of your seed pipeline's CI gate.
The second pitfall is distributed workers generating UUIDs independently without a shared version policy. When pytest-xdist spins up four workers and each imports a different version of a fixture helper — common when a monorepo has multiple requirements files pinning different versions of Faker or a UUID utility — you can get version drift within a single test run. This is the same class of problem as locale collisions across distributed Faker workers: the root cause is per-worker state that should be process-global. Pin UUID generation to a single utility module imported at session scope, not at the factory class level.
Myths About UUID Uniqueness That Break Test Data Assumptions
Myth 1: UUID v4 is collision-proof for test data volumes. Statistically true at production scale; dangerous at test data scale when you're reusing fixture files across runs. If your seed pipeline loads a static UUID from a YAML file and your factory also generates UUIDs at runtime, you can get a deterministic collision the moment someone copies a fixture row and forgets to regenerate the key. The fix isn't probability math — it's enforcing generated-at-seed-time UUIDs with no static values in fixture files, ever. This also intersects with surrogate key exhaustion patterns where bulk generators recycle values under sequence pressure.
Myth 2: Switching to v7 (time-ordered) UUIDs eliminates collision risk in test data. v7 reduces index fragmentation and makes insertion order predictable, but it introduces a new failure mode: two workers seeding in the same millisecond generate UUIDs with identical time prefixes and random suffixes — which is fine for uniqueness but breaks any test assertion that sorts by UUID to infer insertion order. If your test validates "record A was created before record B" by comparing UUID values, v7 makes that assertion meaningful in production and meaningless in a fast parallel seed run where both records land in the same millisecond. Validate temporal ordering with explicit created_at timestamps, not UUID sort order.
UUID version mixing is a quiet pipeline bug that compounds over time — each snapshot, migration, and new service adds another version to the graph. Audit your seed tables with the version-detection query above, enforce version uniformity at the Pydantic or factory layer, and ban static UUIDs in fixture YAML files. If you're moving to v7, read the Postgres 17 gen_uuid_v7() docs alongside your existing seed pipeline before flipping the switch — the monotonicity guarantees change more test assumptions than most teams expect.
Note: This article is for informational purposes only and is not a substitute for professional advice. If you need guidance on specific situations described in this article, consider consulting a qualified professional.