FRAUD PREVENTION · PREDICTIVE AI

How can predictive AI for fraud detection be designed without discriminating against legitimate customers ?

A European health insurer — 4,8 million insured members across seven countries — is preparing to deploy a predictive AI system for healthcare reimbursements. Annual fraud is estimated at €82 million, increasing by 8 % per year. The model has already been technically validated, with projected precision of 78 %.

What remains to be decided is not the model's performance: it is the trigger thresholds, degree of automation, information provided to the insured member, appeal procedure and governance of the system. In other words, how a score becomes a decision.

At home, an insured woman looks on her phone at a healthcare reimbursement statement, with a paper form and an open envelope on the table.
DECISION
Deploy, or not deploy, predictive AI for fraud detection
POPULATION
2,2 million insured members, 17 profiles of attitudes toward controls
WHAT IS TESTED
11 predictive-system configurations
HORIZON
24 months of operational and reputational impact

THE PROBLEM

A model can perform well
and still produce
and still produce a bad decision.

A predictive system learns regularities present in historical data. Some do describe fraudulent schemes. Others describe a chronic condition, a complex care pathway, a change in circumstances, an age, a territory, or a way of filling in forms.

The model does not distinguish these two families of regularities. It treats them as signal. And when a threshold turns that signal into a control, the issue ceases to be statistical: it becomes a decision made about a person.

The useful question is therefore not “ does the model detect more fraud ? ”, but: whom does it identify as suspicious, and what happens next to that person?

WHAT IS THE SAME FOR EVERYONE

  • a European health insurer, 4,8 million insured members, 7 countries
  • €82 million in estimated annual fraud
  • a model already technically validated, with 78 % projected precision
  • manual controls overwhelmed by volume
  • a rollout planned in three waves over 18 months
  • approval expected from the ethics committee

WHAT DIFFERS FOR EACH GROUP

  • frequency and nature of reimbursed care
  • presence of a chronic condition with an atypical care pathway
  • stability of professional and family circumstances
  • age and volume of healthcare use
  • previous experience of an administrative control
  • initial trust in the insurer and institutions

THE FALSE POSITIVE

Detecting more
is not
does not mean deciding better.

A legitimate person incorrectly classified as high risk does not experience a statistical error. They experience a control, a delay, a request for documentation, and the feeling of being suspected.

  1. 01PREDICTIONa risk score is assigned to the claim
  2. 02CONTROLthe score crosses a threshold and triggers an investigation
  3. 03FRICTIONsupporting documents to gather, reimbursement pending
  4. 04INTERPRETATION“why me?” — the question comes before the explanation
  5. 05TRUST / DISTRUSTthis is where judgment of the insurer is formed

This sequence is an analytical mechanism, not a result. It is a reminder that a model does not merely produce a score: it triggers human consequences, and those consequences are what the simulation examines.

BIAS AND CONSEQUENCE

A small difference
in a score
can produce a large difference
in someone's life.

A statistical bias and a human consequence are not measured at the same point. Between them lie a threshold, an automation rule and a procedure. Ethical analysis must cover the entire chain, not the model alone.

  1. HISTORICAL DATASeveral hundred behavioral and transactional variables drawn from real care pathways.
  2. PREDICTIVE SIGNALThe model learns regularities. Some describe fraud. Others describe a chronic illness, a move, an age.
  3. RISK CLASSIFICATIONA score is assigned to a claim. At this stage, nothing has yet happened to the insured member.
  4. CONTROL DECISIONThis is where the score becomes an action: request for supporting documents, reimbursement delay, investigation.
  5. CONSEQUENCE FOR THE INSURED MEMBERFriction, perceived suspicion, suspended reimbursement, and sometimes abandonment of the process.

ONE DECISION, MULTIPLE OBJECTIVES

These objectives
cannot all be maximized
at the same time.

FRAUD CAPTURE

€82 million in estimated annual fraud, increasing by 8 % per year, driven by increasingly sophisticated schemes.

FALSE POSITIVES

Legitimate insured members incorrectly subjected to controls, whose cost is visible neither in the model's precision nor in the savings achieved.

EQUITY ACROSS POPULATIONS

The overall error rate says nothing about how errors are distributed. That distribution is the ethical issue.

AUDITABILITY AND COMPLIANCE

A system that cannot be explained, traced and audited cannot be deployed, regardless of its performance.

WHAT IS TESTED

The same model.
Eleven decision architectures.

The protocol does not test “ an algorithm ”. It tests eleven ways of combining a prediction, a threshold, a human decision, information and an appeal. These six dimensions vary from one configuration to another.

  1. 01TRIGGER THRESHOLDA single threshold applied to the entire population, or differentiated thresholds by insured-member type.
  2. 02DEGREE OF AUTOMATIONAutomated resolution of the flag, or a human case manager systematically involved in the decision.
  3. 03INFORMATION PROVIDED TO THE INSURED MEMBERNo information, contextual information about a standard control, or transparent information about the system and the reason.
  4. 04APPEAL PROCEDURENo appeal, or a formal human appeal available within fifteen days with a reasoned explanation.
  5. 05ETHICAL GOVERNANCEEthics committee established after the first crises, or before deployment, with a published charter and external audit.
  6. 06MODEL REASSESSMENTModel deployed once, or quarterly reassessment of bias indicators and retraining on recent data.

5 200 synthetic insured members were interviewed individually, distributed proportionally across the 17 types, and exposed to the 11 configurations combined with different information and appeal arrangements, with dynamic follow-ups at tipping points: perception of the control, feeling of unfair suspicion, satisfaction with the appeal.

“ Insured member ”
does not describe
any specific risk.

SIMULATED POPULATION

The reconstructed population covers 2,2 million European health-insurance members, calibrated on public data — INSEE, DREES, open insurance-sector data — and on the insurer's proprietary data: member types, healthcare-use history, documented fraud history, demographic structure. No personally identifiable data entered the system.

It is structured into 17 types of relationship with insurance and controls, combining age, family structure, chronic or acute health status, sensitivity to controls, previous experience of administrative interactions and level of trust in institutions.

  • RARE CHRONIC CONDITIONS

    atypical care pathway due to medical necessity, high exposure to being flagged, limited room for explanation

  • TRANSITIONAL LIFE SITUATIONS

    recent professional or family change, significant transactional variations read as anomalies

  • SENIOR INSURED MEMBERS

    high and varied healthcare use, multiple providers, strong sensitivity to the feeling of being suspected

  • SIMPLE CARE PATHWAYS

    median and regular healthcare use, low probability of being flagged, little experience of controls

  • INSURED MEMBERS WITH HIGH INSTITUTIONAL TRUST

    control interpreted as a normal procedure as long as it is explained and time-bounded

  • PREVIOUS EXPERIENCE OF AN ERROR

    past administrative dispute, immediate interpretation of the control as repeated injustice

  • ADMINISTRATIVE VULNERABILITY

    difficulty gathering supporting documents, high cost of pursuing an appeal, risk of abandonment

  • INSURED MEMBERS NOT CONTROLLED BUT INFORMED

    never flagged, aware that the system exists, yet still forming a judgment of the insurer

These configurations reveal part of the population's heterogeneity. The simulation operates on synthetic individuals, not a handful of persona types.

REACTIONS

What is at stake
is not the model's performance.

  1. 01

    ERRORS ARE NOT RANDOMLY DISTRIBUTED

    False positives are concentrated in three types: rare chronic conditions with atypical pathways, people in transitional situations, and senior insured members with high healthcare use. These types represent 8 % of the insured population and 62 % of the false positives generated by the model. The overall error rate says nothing about this concentration.

  2. 02

    TRANSPARENCY COSTS LESS THAN OPACITY

    An unannounced control produces 34 % satisfaction when resolved favorably and 12 % when the claim is rejected. An explained control — nature of the system, reason for the flag, right of appeal — produces 72 % and 51 %. Opacity is read as illegitimate suspicion; explanation, as respect owed to the insured member.

  3. 03

    CONTESTABILITY MATTERS MORE THAN PRECISION

    At identical precision, the same model produces 24 % average satisfaction without a human appeal procedure and 68 % with a formal appeal available within fifteen days: a 44 points difference obtained without changing the model. The algorithm cannot be the final decision-making authority.

  4. 04

    DIFFERENTIATION IS NOT A COMPROMISE

    Differentiated thresholds by type — more permissive for the three sensitive types, stricter elsewhere — reduce false positives by 62 % without degrading fraud capture, with 34 % lower human-processing costs and 78 % lower projected reputational risk. Ethical optimization and operational optimization point in the same direction here.

Imagine All The People does not ask a model whether predictive AI is acceptable. Eleven decision architectures are tested against a heterogeneous synthetic population to observe who is controlled, who manages to justify their case, who abandons the process, and what each person then concludes about their insurer.

At a dining table, an insured woman sorts bills, prescriptions and forms into a folder to support a reimbursement claim.
The cost of a false positive is not visible in the model's precision. It is visible here, in supporting documents and weeks of waiting.

CONDITION FOR DEPLOYMENT

To be deployable,
the system had to
be auditable.

WHAT THE SYSTEM HAD TO DEMONSTRATE

Full traceability of the configurations tested, validation by the insurer's ethics committee before deployment, involvement of legal and compliance teams, annual external audit by an independent firm, and public reporting on bias and false-positive indicators.

WHY THIS CHANGES THE DESIGN

Explainability is not a convenience feature here. It conditions authorization to deploy. A higher-performing but non-auditable configuration was not an available option.

  1. METHOD
  2. TRACEABILITY
  3. AUDIT
  4. VALIDATION

RESULT

Five representative architectures,
assessed across four dimensions.

ARCHITECTUREDETECTIONFALSE POSITIVESEQUITY ACROSS POPULATIONSAUDITABILITY
01Single threshold, automated decision, no information provided to the insured memberhighhighlowlow
02Single threshold, automated decision, contextual informationhighhighlowmedium
03Single threshold, human case manager involved, formal appealhighmediummediummedium
04Differentiated thresholds by type, automated decisionhighlowmediummedium
05Differentiated thresholds, transparent information, human appeal, prior governance, continuous reassessmenthighlowhighhigh

Comparative reading based on simulated reactions across five architectures representative of the eleven configurations tested. The false-positive column is read in the opposite direction from the others: a high level signals a large volume of unjustified controls. A cautious configuration may detect less; an aggressive configuration controls more legitimate insured members.

THE MOST ROBUST

Differentiated thresholds by type transparent information human appeal within 15 days prior ethical governance continuous model reassessment

THE MOST FRAGILE

Single threshold automated decision no information provided to the insured member

The ethics committee selected the leading configuration, reinforced with a permanent committee including representatives of user associations, a published charter and an annual external audit. At 18 months after full deployment, annualized fraud savings reach €62 million versus €54 million projected, the false-positive rate among sensitive types is 4,2 %, insured-member satisfaction after a control is 71 %, and human appeal reverses the algorithmic decision in 34 % of cases where it is used. No significant media or advocacy challenge emerged over the period.

TAKEAWAY

The question was not
who looks like a fraudster.
It was: who would be controlled
because of that resemblance.

The model was good. That is what made the decision difficult. Its overall precision said nothing about how its errors were distributed: 8 % of the population accounted for 62 % of false positives, and that population was precisely the one whose care pathways are atypical by necessity.

The decisive adjustment therefore concerned not the algorithm but the architecture around it: differentiated thresholds, reasoned explanation, human appeal, governance established before deployment. With model precision unchanged, these parameters shifted insured-member satisfaction by 44 points.

TO KEEP
Predictive detection capability, and the operational gain it produces.
TO LIMIT
Direct automation of the control decision, which turns a score into an action.
TO INTRODUCE
Reasoned information, human appeal within fifteen days, ethics committee before deployment, external audit.
TO MONITOR
The distribution of errors by type, and model drift over time.

POSSIBLE FUTURES

The same model.
Three ways to use it.

A

AUTOMATISER

The predictive score directly triggers the control, with a single threshold for the whole population

  • maximum speed and operational efficiency
  • fraud capture in line with technical expectations
  • false positives concentrated among the most atypical types
  • high reputational and regulatory exposure

B

ASSISTER

The score becomes one signal among others, validated by a human case manager

  • the control decision becomes a human decision again
  • substantially higher insured-member satisfaction at identical model precision
  • higher operational processing cost
  • human biases and heterogeneous practices persist

C

GOUVERNER

Differentiated thresholds, transparent information, formal appeal, external audit, quarterly reassessment

  • errors cease to concentrate on the same populations
  • the system becomes explainable, traceable and contestable
  • trust is preserved in the large majority of crisis scenarios tested
  • more complex architecture and ongoing governance requirements

WHAT THE SIMULATION DOES NOT DO

It does not judge
predictive AI.

This case does not conclude that these systems should be deployed, nor that they should be rejected. Predictive detection delivers a real operational gain against costly and growing fraud. The model itself is not the issue.

What is simulated are the consequences of how its prediction is transformed into a decision: who is affected, with what errors, what asymmetries, and what reputational and regulatory consequences. The simulation makes these architectures comparable. It does not rank them on behalf of the institution.

METHOD

Before recommending,
we tested reactions.

  1. DECISION
  2. POPULATION
  3. SYSTEM CONFIGURATIONS
  4. SIMULATION OF CONSEQUENCES
  5. COMPARISON
  6. DECISION

Operational and reputational trajectories were projected over 24 months for each configuration. This case adds a capability to the library: testing not only the reaction to a decision, but how an automated system itself produces decisions about populations.

A real case.
An unnamed insurer.

This case is based on a simulation conducted for a European health insurer. The company is not named, no personally identifiable data entered the system, and detailed results remain the client's property. The comparisons published here are qualitative and constitute neither an algorithmic audit nor an assessment of regulatory compliance.

Your next decision

Which decision do you want to explore?

Describe your need. We can point you to the right level of support.

What if you tested
your next decision?

State your decision. See the future it produces.

Explore the product