LLMs That Fake IDs: Plausible but Invalid
An LLM seeded with locale="de_DE" will happily produce a German Personalausweis number that passes a visual sniff test and fails every real checksum validator you throw at it. The model learned the shape of the format — character class, length, prefix — but not the invariants: Luhn digits, modulo-97 check characters, or authority-specific prefix tables. Your downstream service rejects it silently, the test passes green, and you've validated nothing.
This is the core tension in AI-generated test data: LLMs are excellent at surface plausibility and poor at algorithmic constraints. When you add locale seeding to the mix, the problem compounds — the model now interpolates across locale-specific formats it may have seen only sparsely in training data, producing hybrids that look like no real jurisdiction's ID at all.
By the end of this article you'll know how to detect invalid-but-plausible IDs coming out of an LLM pipeline, how to enforce format contracts at generation time, and which scenarios actually warrant LLM generation versus a library like Faker or Mimesis.
Deliver on your own schedule and get paid for the time you choose to work.
Why Locale-Seeded LLMs Produce Structurally Broken IDs
Most national ID formats embed algorithmic integrity: a check digit computed from the preceding digits, a prefix drawn from a finite authority table, or a date component that must be a valid calendar date in a specific encoding. These constraints are not learnable from examples alone — the model would need to infer the algorithm from a corpus of valid IDs, and that corpus is deliberately sparse in training data for privacy reasons. What the model does learn is the regex skeleton: [A-Z]{1,2}\d{6,9} for a UK NI number, or a 10-digit block for a Brazilian CPF. It fills that skeleton with plausible-looking noise.
Locale seeding makes this worse because it shifts the model's prior toward a particular country's format without giving it the constraint rules. A prompt like "generate 50 test users with valid Brazilian CPF numbers, locale pt_BR" produces digit strings that match the 11-digit format and even include the standard XXX.XXX.XXX-XX punctuation — but the two check digits are random, not computed. This is a different failure mode from the semantic drift you get when LLMs paraphrase enum values; there the shape is wrong, here the shape is right and the math is wrong.
Enforcing ID Validity at Generation Time
The practical fix is a generate-then-validate-then-replace loop: let the LLM produce the surrounding record (name, address, locale-appropriate phone), but delegate every algorithmically-constrained field to a library that knows the invariants. Faker 24.x and Mimesis 12.x both ship locale-aware providers for CPF, BSN, PESEL, NRIC, and others. The LLM sets context; the library sets the ID.
import json
from faker import Faker
from faker.providers.person.pt_BR import Provider as BRProvider
fake = Faker("pt_BR")
fake.add_provider(BRProvider)
def enrich_llm_record(llm_record: dict) -> dict:
"""Replace LLM-generated ID fields with library-validated equivalents."""
record = llm_record.copy()
record["cpf"] = fake.cpf() # Faker computes both check digits
record["rg"] = fake.rg() # State-issued, format-correct
return record
# LLM output arrives as JSON; parse and patch in bulk
with open("llm_output.json") as f:
records = json.load(f)
clean = [enrich_llm_record(r) for r in records]
This keeps the LLM's value — coherent names, realistic addresses, plausible occupation/income combos — while stripping out the fields it can't compute correctly. Generation of 10,000 enriched records drops from ~40 seconds of LLM round-trips to under 3 seconds once the ID fields are handled locally.
When you genuinely need the LLM to produce the ID (e.g., you're testing a parser that must handle locale-specific edge cases), add a post-generation validation gate using stdnum (Python) or a purpose-built validator:
from stdnum.br import cpf as cpf_validator
from stdnum.nl import bsn as bsn_validator
VALIDATORS = {
"pt_BR": {"cpf": cpf_validator.is_valid},
"nl_NL": {"bsn": bsn_validator.is_valid},
}
def validate_ids(record: dict, locale: str) -> list[str]:
errors = []
for field, fn in VALIDATORS.get(locale, {}).items():
if field in record and not fn(record[field]):
errors.append(f"{field}={record[field]!r} failed {locale} validation")
return errors
Wire this into a Pydantic model with a @validator or @field_validator (v2) so the schema layer rejects bad IDs before they reach your database fixtures. Pair it with a JSON Schema 2020-12 pattern constraint for the structural check and the validator for the algorithmic check — the two layers catch different failure classes. If you're building locale-aware seed factories more broadly, the same principle applies to date fields; see how locale seeds corrupt Faker date arithmetic for a related class of silent failures.
Where Senior Engineers Still Get Burned
The first mistake is trusting format-only regex validation. A regex that matches \d{3}\.\d{3}\.\d{3}-\d{2} will pass every LLM-generated CPF regardless of check digit correctness. Teams add the regex to their JSON Schema, see it pass, and ship test data that will fail any service performing real CPF validation. The fix is one extra call to stdnum or equivalent — but it requires knowing the regex isn't enough, which isn't obvious until something downstream breaks.
The second mistake is mixing locales in a single LLM prompt without explicit per-record locale anchoring. Asking for "50 users across de_DE, fr_FR, and pt_BR" often produces German-formatted names with Brazilian ID numbers, or French addresses with Dutch BSN fields. The model interpolates across the locale examples in its context window. Solve it by batching per locale and passing a strict system prompt that names the exact fields and their format contracts — or, better, use the generate-then-replace pattern above and remove the ambiguity entirely. This is the same cross-locale contamination problem documented for locale collisions in distributed Faker workers, just occurring inside the model's context instead of across processes.
Myths About LLMs and ID Format Fidelity
Myth 1: A more capable model (GPT-4o, Claude 3.5 Sonnet) will get the check digits right. It won't — reliably. These models can compute a Luhn digit if you ask them to reason step-by-step, but at bulk generation scale they revert to pattern completion. Benchmarks on CPF generation show even frontier models producing ~30–40% invalid check digits when generating more than a handful of records in a single pass. Capability doesn't substitute for deterministic computation. Myth 2: Prompting the model with a few valid examples is enough to teach it the algorithm. Few-shot examples improve surface formatting but don't reliably transfer the underlying arithmetic. The model is doing token prediction, not algorithm execution. If you need the model to produce valid IDs, make it call a tool (function calling / tool use) that invokes stdnum or Faker server-side — that's the only reliable path.
Myth 3: Invalid IDs are only a problem if your service validates them. They're also a data-quality problem in test analytics, coverage reporting, and AI training pipelines built on synthetic data. An invalid CPF stored in a test database poisons any downstream model trained on that fixture. The blast radius extends beyond the immediate test suite, which is why fixing it at generation time — not at assertion time — is the right architectural choice.
The pattern is simple: use LLMs for what they're good at (coherent, contextually rich records) and route every algorithmically-constrained field through a library that computes the invariants correctly. Add a stdnum-backed Pydantic validator as a generation-time gate, not a test-time assertion. If you're auditing an existing fixture set, a JQ one-liner piped through a Python validator script will surface bad IDs in minutes — fix the generator, not the data.
Note: This article is for informational purposes only and is not a substitute for professional advice. If you need guidance on specific situations described in this article, consider consulting a qualified professional.