Latest Trending Discover Timelines Categories
←All explainers

Technology explainer

Why Does AI Fail When It's Trained on One Population and Used on Another?

AI can fail across populations when inputs, disease prevalence, labels, devices, or workflows differ from training. External validation, subgroup metrics, calibration, prospective testing, and drift monitoring are needed before transfer.

Short answer: AI learns statistical relationships from a particular data-generating environment. When the people, animals, devices, hospitals, prevalence, labels, or workflows change, those relationships may no longer hold. This dataset shift can reduce accuracy, distort confidence, and concentrate errors in the population that was least represented during development.

Training performance is conditional

A model does not learn a universal rule merely because its training score is high. It learns a mapping from inputs to labels under the conditions present in its development data. The result may rely on genuine biology, but also on device noise, clinical habits, background prevalence, image preparation, documentation style, or other shortcuts.

If those conditions change, the input may still look valid to the software while carrying a different meaning. This makes transfer failures especially dangerous: the model can remain confident even as its error rate rises.

Three main forms of shift

Shift What changes Example
Covariate shift The distribution of inputs changes A scanner, microphone, age group, skin tone, species, or hospital differs from training.
Label or prevalence shift The frequency of outcomes changes A specialist clinic sees far more disease than primary care.
Concept shift The relationship between input and outcome changes A treatment, diagnostic definition, pathogen variant, or workflow changes what a signal predicts.

Several types can occur together. “Out of distribution” is therefore not one binary condition but a family of mismatches.

Why population differences matter

Age, anatomy, physiology, genetics, comorbidities, language, environment, and access to care can alter both the signal and the label. A pulse pattern learned mainly from adults may not transfer to infants. A skin-image model can struggle on pigmentation or conditions underrepresented in its training set. A model trained on humans may encounter different heart rates, chest anatomy, acoustic patterns, and diseases in other species.

The population also affects predictive value. Even if sensitivity and specificity stayed fixed, a positive alert is less likely to indicate disease where the condition is rare. A threshold selected in a referral hospital can therefore generate too many false positives in routine screening.

Shortcuts that fail outside the source site

Machine learning rewards any feature that predicts the label, not only the causal feature clinicians intended. A model may learn hospital-specific text, scanner borders, portable-machine markers, recording-room noise, or which tests doctors tend to order. These correlations can look powerful in a random split because the same environment appears in training and testing.

Testing on a later time period and an independent site is more demanding. It asks whether performance survives changes in patients, equipment, staff, and practice rather than merely recognizing the source dataset.

How to tell whether a model will transfer

  1. Define the intended-use population. State ages, species, care setting, geography, devices, languages, and exclusions.
  2. Compare datasets before deployment. Check prevalence, missingness, measurement ranges, acquisition protocols, and subgroup coverage.
  3. Perform external validation. Use untouched data from different sites and a later period, ideally collected independently.
  4. Report subgroup results. Overall averages can hide poor sensitivity or calibration in smaller groups.
  5. Recalibrate when justified. Adjusting probabilities or thresholds may correct prevalence differences, but cannot repair a model that learned the wrong features.
  6. Run prospectively and monitor. Measure real workflow effects, drift, false alerts, overrides, harms, and outcomes after launch.

Accuracy is not enough

Measure Question it answers
Sensitivity Of the true cases, how many did the system detect?
Specificity Of those without the condition, how many did it leave unflagged?
Positive predictive value When it alerts here, how often is it correct?
Calibration Do predicted probabilities match observed risk?
Subgroup performance Are errors distributed acceptably across intended users?

The acceptable balance depends on consequences. Missing a dangerous condition and triggering an unnecessary referral are not equivalent harms. Thresholds must reflect the actual role of the tool and the availability of confirmatory testing.

A concrete transfer failure

An AI-enabled digital stethoscope developed largely from human heart sounds was evaluated in dogs and cats at a veterinary teaching hospital. It missed most feline murmurs and incorrectly flagged healthy dogs for an irregular rhythm; experienced clinicians performed better. The device produced recordings, but its learned classifier did not automatically become a veterinary expert. Read An AI Stethoscope Missed Most Cat Heart Murmurs.

This case illustrates both domain and prevalence differences. Species vary in heart rate, anatomy, sound patterns, and common disease, while the clinical setting changes the case mix. Retraining on representative veterinary data may help, but it must be followed by independent validation for each claimed species and use.

Can adaptation solve the problem?

More representative data, transfer learning, domain adaptation, recalibration, and robust feature design can improve portability. Yet adaptation creates a new model that needs fresh evaluation. Fine-tuning on a small local sample may overfit, reduce performance elsewhere, or erase safeguards.

A safe system also needs uncertainty handling. It should reject unusable inputs, flag conditions beyond its scope, preserve a human fallback, and log version and device information so failures can be investigated.

The practical rule

Validate the model where, when, and on whom it will be used. If any of those change, treat performance as an open question. Deployment should have explicit boundaries, local testing, subgroup analysis, drift monitoring, and a procedure to suspend or roll back the system when the evidence no longer supports it.

First appeared in

An AI Stethoscope Missed Most Cat Heart Murmurs

A new version of NewTqnia is ready.