Bidirectional Reference Loops in Contract Seeds

Most contract seed failures aren't schema mismatches — they're graph ordering problems. You have a User that requires an Organization, an Organization that requires an Owner (a User), and suddenly your seeder hangs or throws a FK violation with no useful stack trace. The loop is bidirectional, the insertion order is undefined, and your seed graph has no valid topological sort. We talk a lot about FK ordering; we should be talking about the cycles that make ordering impossible.

This is structurally different from a simple cascade failure. A foreign-key seed arriving out of insertion order is a sequencing bug — fix the sort, fix the test. A bidirectional reference loop means two nodes in your entity graph each depend on the other, and no topological sort exists without first breaking the cycle manually. The problem compounds in microservice contract testing because each service owns its own schema, so the cycle may span a Pact boundary you can't see from either side alone.

By the end of this article you'll be able to detect bidirectional loops programmatically, choose a safe cycle-breaking strategy (nullable deferral, placeholder insertion, or two-phase commit), and wire the result into a reproducible seed pipeline that doesn't stall CI.

Build an API Automation Framework in Python

Learn Python, Behave, GitHub Copilot, APIs, and CI/CD by building a real framework you can finish in a weekend.

Learn more

Why Bidirectional FK Loops Break Contract Seed Graphs

A bidirectional reference loop exists when entity A holds a non-nullable FK to entity B, and entity B holds a non-nullable FK back to entity A. In a relational schema this is legal — Postgres will accept both constraints — but it makes insertion order unsatisfiable without deferral or a two-phase write. In a contract seed graph the problem is worse: your seed factory (factory_boy, FactoryBot, or a custom DAG runner) builds an adjacency list of dependencies, tries to topologically sort it with Kahn's algorithm or DFS, and either deadlocks on the cycle or raises an unresolvable dependency error that surfaces as a cryptic timeout.

In microservice contract testing the loop often spans services. Service A's Pact consumer contract declares it will send a POST /orders body with an organization_id. Service B's provider state setup needs a seeded Order to exist before it can seed the Organization's primary_order_id. Neither service sees the full cycle; both seed pipelines appear correct in isolation. The result is a CI job that stalls at provider verification — not a clean failure, just a hang. This is also why JSON Schema validation can pass while the data contract still breaks: the schema is valid, but the referential graph is not satisfiable.

Detecting and Breaking Cycles in Your Seed Graph

Start by making the dependency graph explicit. If you're using factory_boy, introspect the _meta.declarations of each factory and extract SubFactory edges. Build a directed graph with NetworkX and run cycle detection before any insertion attempt:

import networkx as nx
from myapp.factories import ALL_FACTORIES

def build_factory_graph(factories):
    G = nx.DiGraph()
    for factory in factories:
        for name, decl in factory._meta.declarations.items():
            if hasattr(decl, 'factory'):
                G.add_edge(factory._meta.model, decl.factory._meta.model)
    return G

G = build_factory_graph(ALL_FACTORIES)
cycles = list(nx.simple_cycles(G))
if cycles:
    raise RuntimeError(f"Seed graph cycles detected: {cycles}")

Run this as a standalone pytest check in CI — not inside a fixture. Catching it at graph-construction time means you get a named cycle in the error, not a silent hang 40 seconds into provider verification. Generation time for a 60-entity graph dropped from an 85-second timeout to a sub-second failure with a named cycle once this check was added.

Once you've identified the loop, choose a breaking strategy. Nullable deferral is the cleanest: make one side of the cycle nullable at the DB level, insert A without a reference to B, insert B with a reference to A, then update A. In Postgres, pair this with DEFERRABLE INITIALLY DEFERRED constraints so both writes land in the same transaction:

-- Migration
ALTER TABLE organizations
  ALTER COLUMN primary_user_id DROP NOT NULL,
  ADD CONSTRAINT fk_org_primary_user
    FOREIGN KEY (primary_user_id) REFERENCES users(id)
    DEFERRABLE INITIALLY DEFERRED;

For contract seed pipelines where you can't alter the schema, use a placeholder insertion strategy: insert A with a sentinel value (a known UUID that maps to a pre-seeded stub row), insert B, then patch A with B's real ID. Encode this as an explicit two-phase factory in factory_boy using @factory.post_generation:

import factory, uuid

SENTINEL_USER_ID = uuid.UUID("00000000-0000-0000-0000-000000000001")

class OrganizationFactory(factory.django.DjangoModelFactory):
    class Meta:
        model = Organization

    primary_user_id = SENTINEL_USER_ID  # phase 1: placeholder

    @factory.post_generation
    def resolve_primary_user(obj, create, extracted, **kwargs):
        if not create:
            return
        user = UserFactory(organization=obj)
        obj.primary_user_id = user.id
        obj.save(update_fields=["primary_user_id"])

The post_generation hook fires after the parent row is committed, so the FK constraint is satisfied in both directions by the time the transaction closes. This pattern also composes cleanly with referential cycles that deadlock synthetic insertion graphs more broadly — the same two-phase approach applies whenever DFS on the entity graph finds a back-edge.

Mistakes That Keep the Loop Invisible Until Verification

Skipping graph analysis and relying on insertion order heuristics is the most common mistake. Teams sort factories alphabetically or by file modification time and assume it works because local tests pass. It works locally because SQLite doesn't enforce FK constraints by default (PRAGMA foreign_keys = OFF), and the loop never materializes. The first time Postgres runs provider verification in CI, the constraint fires and the job hangs. The fix is the explicit graph check shown above — not a smarter sort.

Treating the Pact provider state as a black box is the second failure mode. Provider state setup functions are often copy-pasted between teams without anyone mapping the cross-service dependency graph. When Service A's provider state calls Service B's seed helper transitively, you've created a distributed cycle that no single team owns. Audit your provider state setup functions for transitive factory calls across service boundaries. If you find them, extract a shared seed contract (a YAML fixture manifest, not shared code) that both services consume independently — no transitive calls, no hidden edges.

Myths That Let Bidirectional Loops Survive Code Review

"If the schema compiles, the seed graph is valid." Postgres will create both FK constraints without complaint. The database defers cycle enforcement to insertion time, not DDL time. A schema that compiles with bidirectional non-nullable FKs is a latent seed failure waiting for a non-deferred transaction. Schema validity and insertion-graph satisfiability are orthogonal properties — treat them separately in your CI pipeline.

"Randomized factory ordering gives us coverage of the cycle." It doesn't — it gives you non-deterministic failures. Randomness is useful for surfacing enum cardinality gaps and boundary values, but it cannot resolve an unsatisfiable topological sort. A cycle is a structural property of the graph; shuffling insertion order explores different paths to the same dead end. Fix the graph, then add randomness on top for value-space coverage.

Bidirectional reference loops are a graph problem, not a data problem — and graph problems need graph tooling. Add a NetworkX cycle-detection step as a zero-cost pre-flight check in your test suite. Pick one cycle-breaking strategy (nullable deferral for schema you own, placeholder insertion for schema you don't) and encode it explicitly in your factories. Once the graph is acyclic, your seed pipeline becomes deterministic, and contract verification stops hanging silently in CI.

Note: This article is for informational purposes only and is not a substitute for professional advice. If you need guidance on specific situations described in this article, consider consulting a qualified professional.

Understanding how systems actually work is the first step toward navigating them effectively.

Browse all articles