GUIDES · METHODOLOGY

How to create a synthetic population: the methodological guide.

A synthetic population is not a collection of LLM-generated personas. It is a statistical object built under constraints, with controlled properties and measured deviations. This guide walks through the six steps of a rigorous build.

DEFINITION

What is a synthetic population?

A synthetic population is a set of modeled individuals, from a few thousand to several million, that reproduces the statistical properties of a real population without containing any personal data. Each individual carries demographic, economic, behavioral, territorial and cultural attributes, consistent with one another and consistent at the scale of the whole.

The distinction from neighboring objects is structural. A marketing persona is a composite portrait: an embodied average, with no internal variance. A panel is a sample of real people: valuable, but limited in size, granularity and availability. A synthetic population combines the scale a panel cannot reach with the internal diversity a persona flattens. It is rebuilt from anonymous aggregates and segments: no personal data enters the system, which makes it natively compliant with the RGPD.

THE LLM-ONLY TRAP

Why not simply ask an LLM for profiles?

It is the most widespread method, and the most misleading. An LLM left to itself has a structural drift: it averages. For every profile requested, it produces the most probable individual and flattens the internal variance of the groups. Rare combinations and minority voices disappear: precisely what the decision needs to see. This drift is invisible to the eye, because each individual taken in isolation looks plausible. It is the population as a whole that is false.

Constraint-based generation mechanically does the opposite. Among all the distributions compatible with the imposed statistical references, it retains the most spread-out one, the one that adds no stereotype to the constraints. Real diversity is maintained at the reference frequencies, and the measurement of deviations allows a third party to verify it. It is the difference between a dressed-up intuition and a scientific object.

STEP 1

Define the decision perimeter.

A synthetic population is built for a decision, not in the abstract. The first step is to formulate the decision to be illuminated, then to derive the relevant populations from it: who lives the impact of this decision, on what territory, over what horizon. A price increase does not mobilize the same populations as a regulatory reform or a product launch. This framing determines the dimensions to model and conditions the quality of everything that follows.

STEP 2

Assemble the calibration data.

Calibration draws on two families of sources. Public data first: fine-grained census data, such as INSEE data at the IRIS level for French territorial populations, sector data published by regulatory authorities, national and European reference surveys. Proprietary data second, when the organization has it: internal typologies, behavioral histories, in-house segmentations, integrated after strict anonymization. This dual calibration produces the raw material of the next step: the statistical references the population will have to satisfy.

STEP 3

Formalize the statistical constraints.

The references translate into constraints on three levels. Marginal distributions bear on a single attribute: 52% women. Binary constraints cross two attributes: the proportion of women aged 35 to 44. Ternary constraints cross three: the proportion of male executives aged 45 to 54 living in a given region. These multidimensional constraints are what make the sociological reality of a population. A generation that only respects the marginals produces accurate averages and false cross-sections, with nothing to flag it. A real simulation typically mobilizes 30 to 60 crossed attributes, with numerous ternary constraints.

STEP 4

Generate under constraints.

Simultaneously satisfying dozens to hundreds of multidimensional constraints at the scale of millions of individuals is an NP-hard problem: exact formulations explode combinatorially. The methodological answer is Maximum Entropy Relaxation, inherited from the principle formalized by Jaynes in 1957 and from Régin's work on cardinality constraints: among all the distributions that satisfy the constraints, retain the one with maximum entropy, the least biased, which is unique and optimal in the sense of information theory. The problem reformulates into convex optimization, solvable with convergence guarantees.

This approach is publicly benchmarked against generalized raking, the reference method of statistical surveys, on NPORS datasets of 4 to 40 attributes with ternary constraints. Beyond 28 simultaneous attributes, its advantage is structural: that is the regime of real simulations. The results, the code and the methodology are published on ArXiv and reproducible by third parties.

STEP 5

Structure into granular typologies.

A statistically accurate population remains unreadable without structure. The next step divides it into granular typologies, typically 10 to 30 per simulation, built by crossing the dimensions relevant to the decision: relationship to the subject, personal trajectory, territorial anchoring, media exposure, cultural capital. These typologies are validated by cross-referencing with the public qualitative studies on the subject, to guarantee that they correspond to sociologically relevant divisions and not to statistical artifacts. This is the granularity at which the decision plays out: general averages hide the 4% of a population that concentrates 78% of a measure's risk.

STEP 6

Validate, measure, document.

A serious synthetic population ships with its proofs. Deviations from the target distributions, marginal, binary and ternary, are measured and documented in a calibration report. Internal variance is compared against survey references. The results produced on the population must be traceable down to the individual interviews that ground them, and therefore contestable line by line by independent third parties. A population whose properties are neither measured nor published is not a methodological object: it is a dressed-up hypothesis.

ACKNOWLEDGED LIMITS

What a synthetic population does not do.

Three limits delineate the use. It does not mechanically predict the future: it projects probabilistic trajectories that reduce uncertainty without eliminating it. It does not replace real data: when observed behavioral data exists and is relevant, it remains the reference that the simulation complements. It does not dispense with human decision-making: it illuminates the decision-maker's reasoning, it does not substitute for it. Naming these limits is a prerequisite of scientific rigor.

FAQ

Frequently asked questions.

How many individuals should be generated?

The scale is calibrated to the decision: from a few hundred thousand individuals for a territory or a segment to ten million for a multinational population. The stake is not raw volume but the preservation of granularity: minority typologies must remain represented at their real frequencies.

Does a synthetic population contain personal data?

No, by construction. It is generated from anonymous aggregates and from public or anonymized proprietary statistics. No personal data enters the system, which makes it natively compatible with the RGPD.

How is it different from a classic panel?

A panel interviews real people, in limited numbers and over long timelines. A synthetic population reaches a scale and granularity inaccessible to panels, with immediate availability. The two approaches are complementary: structural studies keep their value in panels, iterative explorations move to synthetic.

Can we use our own segmentations?

Yes. An organization's proprietary segments are formalized into instantiation instructions: the generated population then simultaneously respects the public references and the in-house segmentation, which remains the organization's property.

How long does the creation take?

On our system, generating a calibrated population and interviewing it is a matter of minutes: from the initial brief to the first investigation file, expect 20 to 30 minutes depending on the complexity of the case.

Une décision à prendre, une population de synthèse qui y répond, un éclairage

Create your first population.

The methodology described in this guide is industrialized in our system: describe the population you want to understand, in natural language, and meet the synthetic humans who compose it.

Contact our team →

GOING FURTHER