GUIDES · SYNTHETIC POPULATIONS
Build a synthetic population.
A synthetic population is not a collection of personas.
It is a set of fictional individuals generated under constraints to reproduce the diversity, distributions and relationships that structure a real population.
The goal is not to create a few plausible profiles. It is to build a collective object whose properties can be measured, controlled and documented.
Build diversity. Preserve variance. Measure what you built.
THE DISTINCTION
Persona, panel, synthetic population: three different objects.
PERSONA
An individual profile built to represent a type of user.
PANEL
A group of real people recruited to be questioned.
SYNTHETIC POPULATION
A set of fictional individuals generated to respect controlled collective properties.
A persona describes an individual. A panel observes real individuals. A synthetic population builds a representative system.
THE PROBLEM
An LLM alone tends toward the probable individual.
A language model can produce a plausible character. It is far less naturally suited to building a population whose distributions and relationships must remain controlled.
LLM alone
- Probable profile
- Convergence toward the average
- Reduced variance
- Under-represented minorities
Constraint-based generation
- Imposed distributions
- Controlled relationships
- Preserved diversity
- Measurable population
The plausibility of an individual does not guarantee the coherence of a population.
THE METHOD
Six steps to build a synthetic population.
DEFINE
Define the population relevant to the decision.
Territory, market, customers, citizens, employees or another human group.
CALIBRATE
Identify the distributions to reproduce.
Public data, sector data and proprietary data when available.
CONSTRAIN
Define the statistical relationships to preserve.
Marginals, binary relationships and more complex relationships.
GENERATE
Instantiate individuals under constraints.
Each individual must remain plausible while respecting the properties of the whole.
STRUCTURE
Make readable typologies emerge.
Turn millions of individuals into interpretable groups without losing internal diversity.
VALIDATE
Measure the properties of the resulting population.
Compare the generated population with the expected constraints and document deviations.
THE CONSTRAINTS
Preserve more than an average.
A population is not defined only by simple proportions. Relationships between variables matter as much as the variables themselves.
MARGINAL
52% women
Control the distribution of one variable.
BINARY
Women aged 35–44
Control the relationship between two variables.
TERNARY
Male executives aged 45–54 in one region
Control finer relationships across several dimensions.
The more relationships are preserved, the more the population retains its internal structure.
THE ENGINE
Maximum Entropy Relaxation.
When all constraints cannot be satisfied simultaneously, the system searches for a solution that respects the available constraints as closely as possible without artificially compressing diversity.
The principle is to preserve the maximum amount of information compatible with the known constraints.
- Observed constraints
- Compatibility
- Controlled relaxation
- Preserved diversity
Do not force a perfect population. Build the population most coherent with what is known.
See the methodological foundations →
STRUCTURE
From millions of individuals to readable groups.
A population can contain millions of individuals. It must still remain interpretable.
MILLIONS OF INDIVIDUALS
VARIANCE
GROUPING
TYPOLOGIES
Each typology groups similar individuals without removing their individual differences.
Structuring is not reducing. It is making complexity readable.
VALIDATE
A population comes with its evidence.
A synthetic population is not rigorous because it looks plausible. It is rigorous because its properties are measured and documented.
- 01 Marginals
- 02 Binary relationships
- 03 Ternary relationships
- 04 Variance
- 05 Deviations
- 06 Traceability
EXPECTED
What the population was supposed to reproduce.
OBSERVED
What the generated population actually reproduces.
DEVIATION
The difference between the two.
A population whose properties are neither measured nor documented is not a rigorous methodological object.
STATED LIMITATIONS
What a synthetic population is not.
A COPY OF REALITY
A synthetic population reproduces properties and explores plausible reactions. It is not a perfect duplicate of a real population.
A REPLACEMENT FOR REAL DATA
When real data exists, it remains essential for calibrating, validating and challenging the population.
A DATABASE OF PEOPLE
The simulated individuals are fictional. Simulation does not require personal data about the simulated people.
If real personal data is used to calibrate a project, it remains subject to the applicable rules.
Frequently asked questions.
How many individuals are needed?
The size depends on the level of detail required, the segments to represent and the relationships to preserve. A larger population is useful only if it provides a more relevant structure.
Is a synthetic population representative?
It is representative of the constraints and distributions used to build it, within the limits documented by its calibration report.
Can proprietary data be used?
Yes. It can complement public data to calibrate a specific population, subject to the rules applicable to the data concerned.
Do synthetic individuals correspond to real people?
No. They are fictional and do not correspond to identified natural persons.
How do you know whether the population is high quality?
By comparing its measured properties with the expected constraints and documenting the deviations.
YOUR POPULATION
Build the population your decision needs.
Tell us about the context, the available data and the decision you want to test.
Talk about your project →