SCIENCE · METHODOLOGICAL FOUNDATIONS
Our synthetic populations are not LLM-simulated and eyeball-calibrated. They are generated under multi-dimensional statistical constraints, with explicitly controlled properties and deviations from the references measured and published.
WHERE WE COME FROM SCIENTIFICALLY
We do not invent a methodology out of nothing. We articulate four established scientific traditions in an unprecedented combination. Each brings a specific building block. Together, they produce what none could produce alone: decision intelligence through synthetic populations.
Our mathematical core. Imagine All The People combines constraint-satisfaction methods, inherited from Régin's work (1996) on the Global Cardinality Constraint, with probabilistic modeling methods, inherited from the Maximum Entropy principle formalized by Jaynes (1957). The goal is to build synthetic populations coherent with a set of reference statistics: marginal distributions, for example 52% women, but also cross-distributions over several attributes, such as the proportion of women aged 35 to 44, or of those women living in the Paris region. This approach does not claim to perfectly reconstruct a real population from partial data. What it does allow is to explicitly control the statistical properties deemed important, and to ensure that large-scale generation respects the general averages while producing coherent, minimally biased combinations of attributes.
Our calibration data. We inherit the methodologies for building statistically coherent populations: INSEE at IRIS level for territorial populations, sector data published by regulatory authorities, national and European reference surveys, long-term sector panels. These public data form the constraints our synthetic populations must satisfy. Our contribution is not to challenge these sources: it is to combine them rigorously at a scale physical panels cannot reach.
Our depth of interviewing. We inherit the methodologies of the in-depth qualitative interview: those of social scientists who probe deep motivations through contextualized follow-ups, of clinicians who distinguish the stated from the symptom, of ethnographers who enter the registers and codes of the populations they study. This tradition, which produced the finest insights in the human sciences, gives us the qualitative depth our dynamic agents are calibrated on.
Our execution layer. We inherit recent advances in large language models: contextualized dialogue, fine understanding of linguistic registers, structured restitution, massive parallelization of processing. This tradition, more recent but dense with progress, makes it possible to execute qualitative interviews at scale on our synthetic populations. It is our execution layer: not our scientific core. The distinction matters: generic AI produces average answers; our system produces in-depth qualitative interviews on populations with explicitly controlled statistical properties.
THE MATHEMATICAL CORE
Generating a synthetic population coherent at the scale of several million individuals that simultaneously satisfies dozens to hundreds of multi-dimensional statistical constraints is no trivial problem. It is NP-hard in theoretical computer science. Exact formulations explode combinatorially as the number and arity of constraints grow. A population generated by simple sampling from marginal distributions respects the general averages but systematically violates the multi-dimensional correlations: 45-to-54-year-old male executives living in Occitanie end up over- or under-represented relative to statistical reality, with nothing in the system detecting it.
We solve this problem with Maximum Entropy Relaxation, a founding principle inherited from statistical physics and formalized by Jaynes in 1957. The idea is mathematically simple and deep: among all the distributions that satisfy the known constraints, select the one that makes the fewest additional assumptions: the maximum-entropy distribution. This distribution is maximally unbiased given the imposed constraints. It is unique. It is mathematically optimal in the precise sense of information theory.
Our approach recasts the NP-hard problem of exact constraint satisfaction as a convex optimization problem solvable by standard numerical methods. The multi-dimensional constraints are satisfied in expectation rather than exactly, which yields a standard exponential-family distribution over the complete configurations of the population. This formulation dualizes into a convex optimization problem over the Lagrange multipliers: solvable by L-BFGS with convergence guarantees.
The statistical properties of our populations are explicitly controlled and verified by measurement: deviations from the target distributions, marginal, binary and ternary, are documented for every simulation in a calibration report appended to the inquiry dossier, and published on standardized benchmarks. This is what sets us apart from players who generate populations by simply prompting an LLM. An LLM left to its own devices has a structural drift: it averages. For each profile it produces the most probable individual and crushes the internal variance of groups: rare combinations and minority voices disappear, precisely what the decision needs to see. This drift is invisible to the eye, because each individual looks plausible; it is the population as a whole that is wrong. Maximum Entropy mechanically does the opposite: among all the distributions compatible with the references, it retains the most spread out, the one that adds no stereotype to the constraints. Generation under constraints keeps real diversity at the reference frequencies, and the deviation measurement lets a third party verify it.
This mathematical core is presented in detail in the article Maximum Entropy Relaxation of Multi-Way Cardinality Constraints for Synthetic Population Generation, published on ArXiv in April 2026. All the scientific publications our team contributes to are available from the Scientific papers page.
HOW WE WORK IN PRACTICE
Each synthetic population is calibrated on the public data relevant to the decision at hand: INSEE at IRIS level for territorial populations, sector data published by regulatory authorities, national and European reference surveys. When the client holds relevant proprietary data, it is added to the calibration after strict anonymization to refine the structuring of the constraints. This dual calibration produces the multi-dimensional statistical constraints our mathematical core will satisfy.
Once the constraints are defined, our system generates the synthetic population by Maximum Entropy Relaxation, then measures its deviations from the target distributions, unit marginal, binary and ternary, documented in the calibration report. Our populations simultaneously respect the broad demographic averages and the fine correlations between attributes: not just the general proportions, but the cross-distributions that structure sociological reality.
On this statistically controlled base, we structure the population into granular typologies, typically 10 to 30 per simulation depending on the case's complexity, built by crossing several dimensions relevant to the decision at hand. This structuring is validated by cross-reference with the public qualitative studies on the subject, ensuring that the typologies retained correspond to sociologically meaningful cuts and not statistical artifacts.
Our dynamic agents do not administer closed questionnaires. They conduct qualitative interviews structured by protocols calibrated for the decision at hand, with a capacity for contextualized follow-up that digs into friction points identified in real time. The number of follow-ups per interview, typically 7 to 12 with peaks of 20 to 30 on the most complex cases, is an explicit methodological parameter, adapted to the qualitative depth each simulation requires.
Every insight delivered to the client is traceable to the individual interviews that produced it. Every projection rests on the reasoning paths that built it. This full traceability is not a feature: it is a methodological requirement. It guarantees that our results can be challenged line by line by independent third parties, which is the condition of existence of a serious scientific approach.
HOW WE POSITION OURSELVES
| Classic panels and quantitative sociology | LLM simulation without a mathematical core | Our discipline | |
|---|---|---|---|
| Scale of interviewing | Panels of a few thousand individuals interviewed | Populations generated without coherence control at scale | Coherent synthetic populations of up to several million |
| Multi-dimensional statistical coherence | Marginal distributions respected by sampling | Drift toward the average: no property controlled or measured | Properties controlled and deviations measured by Maximum Entropy Relaxation |
| Qualitative depth | Limited to structured questionnaires | Average answers without contextualized follow-up | In-depth interviews with 7 to 12 follow-ups per person |
| Temporal dimension | Successive snapshots via repeated fieldwork | Instant output without temporal depth | Trajectories projected over several months or years |
| Scientific publication of the method | Sector methodology documented in the classic journals | Absent or limited to commercial publications | ArXiv and Frontiers in AI publications, academic partnerships |
HEALTHCARE
HEALTHCARE
This case illustrates our methodological approach in its most readable form. We generated by Maximum Entropy a synthetic population of 2.4 million parents of children aged 12 to 24 months, simultaneously satisfying the marginal, binary and ternary distributions calibrated on INSEE and Santé publique France data. This population was structured into 18 vaccine-hesitancy typologies built by crossing six dimensions: personal care experience, institutional trust, scientific cultural capital, family environment, media exposure, territorial political context. Our dynamic agents conducted several hundred thousand synthetic interviews, with 8 follow-ups per person on average and full traceability of every insight down to the interviews that produced it.
Read the full case →WHAT WE DO NOT DO
Every serious scientific approach begins by circumscribing its domain of application. Players who claim to explain everything, predict everything, solve everything are structurally less credible than those who acknowledge their limits. We identify three, structural to our discipline.
They project possible trajectories based on the dynamics identified in the interviews and the modeled cascades. These projections are probabilistic, not deterministic. They reduce uncertainty before a decision: they do not eliminate it. An unexpected exogenous event (an economic crisis, a geopolitical shock, a disruptive media event) can invalidate a projection however carefully it was built.
When real behavioral data exists and is relevant to the decision, it remains the reference: our simulations complement it, they do not replace it. Our synthetic populations deliver their full value when real data does not yet exist (a new product, a new territory, a new configuration) or is not enough (behaviors under unobserved scenarios).
We deliver insights, compared scenarios, identified tipping points. We do not make the decision in the decision-maker's place. Our deliverables are exploration instruments that accompany the executive's reasoning: they do not substitute for it. This distinction is essential: an organization that delegated its decisions to a simulation system, whatever it may be, would commit a major methodological and political error. Our tools sharpen the decision-maker's lucidity: they do not replace their responsibility.
Our methodological controls and validation protocols are detailed on the Validation & calibration page. The scientific benchmarks comparing our approach with classic methods are published on the Benchmarks page. The scientific publications our team contributes to are available from the Scientific papers page. Co-founder François Pachet and our ecosystem of academic partnerships are presented on the dedicated page. You can also contact our team directly to discuss the methodological parameters of a simulation adapted to your decision.
See the benchmarks →