Science · Foundations

Build a population, not a collection of personas.

We start from reference data, transformed into statistical constraints. A population distribution is generated under these constraints, then calibrated before being structured into queryable synthetic individuals.

DATACONSTRAINTSPOPULATIONReference aggregatesDistributions to satisfySynthetic individuals
01The methodological chain

Data → constraints → population.

  1. 01Reference data

    Relevant public data (INSEE at IRIS level, sector data, national surveys) and anonymized proprietary data where available.

  2. 02Constraints

    Marginal and cross-distributions that the population must satisfy.

  3. 03Maximum Entropy Relaxation

    The constraints are recast as a convex optimization problem and satisfied in expectation rather than exactly.

  4. 04Population

    An exponential-family distribution over complete individual configurations, then calibrated by measuring deviations from target distributions.

  5. 05Synthetic individuals

    Complete, mutually coherent attribute configurations relevant to the decision being studied.

  6. 06Interviews

    Structured qualitative interviews with contextual follow-up on points of friction.

02Constrained generation

Maximum Entropy Relaxation.

Formulation

p(x) = exp( Σₖ λₖ · fₖ(x) ) / Z(λ)

  • p(x)probability of a complete individual configuration x
  • fₖ(x)statistical constraint k (marginal, binary or ternary), satisfied in expectation
  • λₖLagrange multiplier associated with constraint k, obtained through dual convex optimization
  • Z(λ)normalization constant of the exponential family

Among the distributions compatible with the constraints, retain the one that adds the least additional structure.

SourceFrançois Pachet, Jean-Daniel Zucker — Maximum Entropy Relaxation of Multi-Way Cardinality Constraints for Synthetic Population Generation, ArXiv (preprint, not peer reviewed) — 2026ArXiv:2603.22558 View
Principle inherited from Jaynes (1957); cardinality constraints inherited from Régin (1996).

03Margins and cross-distributions

Matching each variable does not mean matching their cross-distributions.

UNIVARIATE MARGINSSEXAGEINCOMEEach variable matched separately.CROSS-DISTRIBUTIONSSEX × AGEAGE × INCOMESEX × AGE × INCOMEAttribute combinations are themselves constrained.

Conceptual illustration. Marginal sampling can violate entire combinations of attributes without anything detecting it. The selected constraints also apply to cross-distributions.

04From distributions to individuals

A population is not an average.

AVERAGED POPULATIONMinority voices disappear.HETEROGENEOUS POPULATIONRare combinations remain present at their reference frequencies.

The population retains diversity consistent with the selected constraints: each synthetic individual is a complete configuration of attributes, queried in that specific configuration rather than as an average profile. Reading typologies are an analytical tool; the simulation continues to operate at the individual level.

Conceptual illustration. No real data are shown.

05Querying the population

From the population to queryable individuals.

  1. 01Question

    Interview protocol calibrated for the decision being studied.

  2. 02Answer

    The synthetic individual responds from their own configuration.

  3. 03Contextual follow-up

    Identified points of friction are explored rather than bypassed.

  4. 04Synthesis

    Insights are delivered with links to the interviews that produced them.

Population structure

  • Reference data
  • Statistical constraints
  • Calibration and deviation measurement

The LLM is an interaction layer

  • Contextualized interview dialogue
  • Linguistic register of the individual being interviewed
  • Structured delivery of responses

The language model is the execution layer, not the population: statistical properties are controlled and measured upstream.

06Traceability & limitations

Results that can be challenged line by line.

  1. 01Result delivered to the client.
  2. 02Individual interviews that produced it.
  3. 03Population and constraints applied.
  4. 04Reference data and calibration report.

What the method does not claim

  • The synthetic population does not predict a real person.
  • A synthetic individual has not lived the real-world experience.
  • Without reliable reference data, the result remains exploratory.
  • A simulation does not replace a regulatory obligation.

Your next decision

Which decision do you want to explore?

Describe your need. We can point you to the right level of support.

What if you tested
your next decision?

State your decision. See the future it produces.

Explore the product