03 · What You Need to Know
Overfitting Produces Performance That Looks Better Than It Travels
What overfitting means in prediction research
A prediction model attempts to learn relationships between predictors and an outcome that will generalize to future individuals from the target population. During development, however, the model sees only a particular sample.
If the model is sufficiently flexible relative to the information in that sample, it can learn not only reproducible signal but also sample-specific noise. Apparent performance calculated from the development data then becomes optimistic: it makes the model look better than it is likely to perform in new observations.
Model fit
How well the model describes or predicts observations used during its development.
Generalization
How well the model predicts outcomes for relevant new individuals who were not used to construct it.
Start with how much information was available
Prediction-model sample size should be considered relative to the modelling task, not merely as a raw participant count. For binary outcomes, the number of participants experiencing the outcome can be especially important. For time-to-event models, the number of events likewise constrains available information.
A dataset containing 10,000 people may still provide limited information if the outcome occurs only 80 times and researchers consider dozens or hundreds of candidate predictor parameters.
PROBAST specifically asks whether there was a reasonable number of participants with the outcome relative to the number of candidate predictor parameters and whether overfitting and optimism were appropriately addressed. Empirical reviews using PROBAST have repeatedly identified insufficient events and inadequate handling of overfitting as common problems in prediction-model research.
Do not rely mechanically on the old events-per-variable rule
You may encounter the traditional recommendation of at least 10 outcome events per predictor variable in logistic or survival models. It can serve as a rough warning sign in some appraisals, but it is not a universal law and should not be treated as a definitive sample-size calculation.
Modern prediction-model sample-size methods consider additional features such as anticipated model fit, outcome frequency, total sample size, and the number of predictor parameters. The number of parameters also matters more than simply counting named predictors because categorical variables, nonlinear terms, splines, and interactions can consume multiple parameters.
The useful appraisal question is therefore not merely "Did they achieve 10 events per variable?" It is whether the development dataset contained enough information for the complexity of the model being estimated.
Count candidate predictors, not only predictors in the final model
A paper may present a final model containing six predictors and appear reassuringly simple. But perhaps the researchers started with 80 candidate variables, tried several transformations, examined interactions, tested alternative selection procedures, and retained the combination that performed best.
That search process is part of model development and contributes to overfitting.
Look in the Methods for the full set of candidate predictors and how they were selected. If the paper reports only the final variables without explaining the search that produced them, you cannot judge model complexity adequately.
Univariable screening is a warning sign
Some studies test each potential predictor individually and include only predictors with a statistically significant univariable association in the multivariable model. PROBAST specifically asks whether predictor selection based on univariable analysis was avoided.
A variable's value for multivariable prediction is not determined by whether its individual P-value crosses a threshold. Screening in this manner can create unstable selection and discard predictors that contribute useful information jointly with others.
Likewise, automated stepwise procedures can produce unstable models, particularly when information is limited relative to the number of candidate parameters.
Apparent performance is usually optimistic
If researchers develop a model and then calculate its performance using the exact same observations, the assessment is not an independent test. The model has already been optimized, directly or indirectly, for those data.
This apparent performance commonly overestimates performance in new individuals.
A particularly concerning paper might report excellent discrimination but provide no internal validation, no optimism correction, and no external validation. The performance estimate may describe how well the model learned its development sample rather than how well it predicts future cases.
Internal validation should reproduce the whole development process
Internal validation uses the development data to estimate how much performance may deteriorate in new observations. Resampling approaches such as bootstrapping are commonly used.
TRIPOD guidance emphasizes that all aspects of model building should be incorporated into internal validation, including predictor selection, penalization, transformations, and interaction testing. Validating only the final fitted equation ignores uncertainty introduced by the development process and can yield overoptimistic results.
This point is easy to miss. If researchers select variables using the entire dataset and only afterward cross-validate the already selected model, information from the validation observations has indirectly influenced model construction.
A simple train-test split is not automatically strong validation
Researchers sometimes randomly divide one dataset into a training set and a test set. This keeps the test observations out of model fitting and can provide an honest evaluation if the split is maintained properly.
However, splitting can be statistically inefficient, especially when the overall dataset is not large. The model is developed using only part of the available information, while performance is evaluated on a test sample that may itself be relatively small and imprecise.
Resampling methods such as bootstrapping or appropriately implemented cross-validation often use available development data more efficiently for internal validation. The best strategy depends on the modelling context and available information.
Cross-validation can still leak information
The phrase "we used cross-validation" should not end your appraisal.
Ask what happened inside each resampling iteration. Feature selection, preprocessing, normalization, imputation procedures that learn from data, hyperparameter tuning, and other data-driven operations should be structured so that validation observations do not influence the model being evaluated.
If researchers preprocess or select features using the complete dataset before cross-validation, information can leak from validation observations into model development. Reported performance may then remain optimistic.
Machine learning does not eliminate overfitting
Highly flexible machine-learning algorithms can model complex nonlinearities and interactions, but that flexibility can increase the capacity to fit noise when information is insufficient or tuning is poorly validated.
A systematic review of supervised machine-learning prediction models found high overall risk of bias in 88% of developed models assessed, with common problems involving inadequate sample size, missing-data handling, and failure to account adequately for overfitting and optimism.
The label "machine learning" therefore provides no assurance of generalizability. The same questions about information, model development, validation, and performance remain essential.
Discrimination alone is not enough
Prediction papers frequently emphasize discrimination, often using the area under the receiver operating characteristic curve or C-statistic. Discrimination measures how well a model distinguishes individuals who experience the outcome from those who do not.
Calibration asks a different question: do predicted probabilities agree adequately with observed outcome frequencies?
A model can discriminate reasonably well while producing systematically inaccurate probabilities. PROBAST therefore considers whether relevant performance measures, including both discrimination and calibration where appropriate, were evaluated. Empirical work has found calibration frequently omitted from prediction-model reports.
| Performance concept |
Question |
| Discrimination |
Can the model distinguish individuals with different outcomes? |
| Calibration |
Do predicted probabilities correspond adequately to observed outcome frequencies? |
| Overall performance |
How accurate are predictions when considering the prediction errors more broadly? |
| Clinical or decision utility |
Would using the model improve decisions compared with relevant alternatives? |
External validation answers a different question
External validation evaluates the developed model in data that were not used to construct it and that are meaningfully separate from the development sample. It provides evidence about how well the model performs beyond its original dataset.
Performance commonly changes when a model is evaluated in new populations or settings because case mix, predictor distributions, measurement procedures, outcome frequency, and other features differ.
External validation therefore adds evidence about transportability. It does not retroactively make poor development practices acceptable, and one favorable external validation does not prove universal performance.
A large drop between development and validation performance is informative
Suppose a paper reports an apparent C-statistic of 0.94 during development but optimism-corrected internal performance of 0.79. That difference suggests the original apparent performance substantially overstated what the model is likely to achieve in new data.
A similar drop during external validation can reveal limited generalizability, although population differences and measurement changes can also contribute.
Do not focus only on whether validation performance remains above an arbitrary threshold. Examine how much performance changes and whether calibration remains acceptable for the intended use.
Penalization and shrinkage can reduce overfitting
Methods such as ridge regression, lasso, elastic net, and other forms of penalization can constrain model complexity. Shrinkage can reduce overly extreme estimated coefficients. TRIPOD guidance discusses penalized approaches and bootstrap-derived shrinkage as strategies for reducing overfitting.
These techniques help, but their mere presence does not guarantee a good model. Penalty parameters may themselves require tuning, and tuning procedures need appropriate validation. A sophisticated method can still be implemented poorly. Statistics has no immunity-by-acronym clause.
Overfitting is not the only reason a model fails externally
A model can perform poorly in a new dataset because of overfitting, but other explanations include differences in population, outcome definition, predictor measurement, healthcare or educational practice, prevalence, temporal changes, and data quality.
BMJ guidance on external validation emphasizes evaluating whether participant selection, predictors, and outcome definitions in the validation data are appropriate for the intended model use.
Therefore, poor external performance is evidence that the model does not transport well to that validation setting. It does not identify overfitting as the sole cause.
Watch Out
An impressive accuracy, AUC, or C-statistic calculated on the same observations used to develop the model may be substantially optimistic. Always determine whether the reported performance is apparent, internally validated and optimism-corrected, or externally validated.