Science · Foundations
Build a population, not a collection of personas.
We start from reference data, transformed into statistical constraints. A population distribution is generated under these constraints, then calibrated before being structured into queryable synthetic individuals.
Data → constraints → population.
- 01Reference data
Relevant public data (INSEE at IRIS level, sector data, national surveys) and anonymized proprietary data where available.
- 02Constraints
Marginal and cross-distributions that the population must satisfy.
- 03Maximum Entropy Relaxation
The constraints are recast as a convex optimization problem and satisfied in expectation rather than exactly.
- 04Population
An exponential-family distribution over complete individual configurations, then calibrated by measuring deviations from target distributions.
- 05Synthetic individuals
Complete, mutually coherent attribute configurations relevant to the decision being studied.
- 06Interviews
Structured qualitative interviews with contextual follow-up on points of friction.
Maximum Entropy Relaxation.
Formulation
p(x) = exp( Σₖ λₖ · fₖ(x) ) / Z(λ)
- p(x)probability of a complete individual configuration x
- fₖ(x)statistical constraint k (marginal, binary or ternary), satisfied in expectation
- λₖLagrange multiplier associated with constraint k, obtained through dual convex optimization
- Z(λ)normalization constant of the exponential family
Among the distributions compatible with the constraints, retain the one that adds the least additional structure.
SourceFrançois Pachet, Jean-Daniel Zucker — Maximum Entropy Relaxation of Multi-Way Cardinality Constraints for Synthetic Population Generation, ArXiv (preprint, not peer reviewed) — 2026 — ArXiv:2603.22558 View
Principle inherited from Jaynes (1957); cardinality constraints inherited from Régin (1996).
Matching each variable does not mean matching their cross-distributions.
Conceptual illustration. Marginal sampling can violate entire combinations of attributes without anything detecting it. The selected constraints also apply to cross-distributions.
A population is not an average.
The population retains diversity consistent with the selected constraints: each synthetic individual is a complete configuration of attributes, queried in that specific configuration rather than as an average profile. Reading typologies are an analytical tool; the simulation continues to operate at the individual level.
Conceptual illustration. No real data are shown.
From the population to queryable individuals.
- 01Question
Interview protocol calibrated for the decision being studied.
- 02Answer
The synthetic individual responds from their own configuration.
- 03Contextual follow-up
Identified points of friction are explored rather than bypassed.
- 04Synthesis
Insights are delivered with links to the interviews that produced them.
Population structure
- Reference data
- Statistical constraints
- Calibration and deviation measurement
The LLM is an interaction layer
- Contextualized interview dialogue
- Linguistic register of the individual being interviewed
- Structured delivery of responses
The language model is the execution layer, not the population: statistical properties are controlled and measured upstream.
Results that can be challenged line by line.
- 01Result delivered to the client.
- 02Individual interviews that produced it.
- 03Population and constraints applied.
- 04Reference data and calibration report.
What the method does not claim
- The synthetic population does not predict a real person.
- A synthetic individual has not lived the real-world experience.
- Without reliable reference data, the result remains exploratory.
- A simulation does not replace a regulatory obligation.
Your next decision
Which decision do you want to explore?
Describe your need. We can point you to the right level of support.