Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

How Do You Recognize Overfitting in a Prediction Model Paper?

An overfitted prediction model learns features of its development data that do not generalize well to new individuals. Warning signs include excessive model complexity, inadequate information, weak validation, and performance reported only on the data used to build the model.

393
Recognizing Overfitting in Prediction Models Guide 393 of 899
01 · The Question

Is the Model Predicting the Outcome or Remembering Its Development Data?

A prediction-model paper reports an area under the curve of 0.94. Accuracy is impressive. Perhaps the authors describe their machine-learning algorithm as outperforming existing approaches. Should you trust those numbers?

Not until you know how the model was developed and evaluated.

A flexible model can fit the dataset used to create it extraordinarily well, including idiosyncratic patterns and random noise that will not recur in future individuals. This is overfitting. The model appears excellent where it was developed but performs worse when asked to predict genuinely new observations.

You usually cannot diagnose overfitting from one performance statistic alone. Instead, look for conditions that make overfitting likely and for validation procedures designed to reveal and correct the resulting optimism.

02 · The Short Answer

Look for Too Much Model Complexity Relative to the Available Information

In Brief

Suspect overfitting when a prediction model is developed using too little information for its complexity, explores many candidate predictors or modelling choices, and then reports performance mainly or entirely on the same data used to construct the model. Stronger papers use appropriate internal validation to quantify optimism, report both discrimination and calibration, and ideally evaluate performance in independent data.

No single sample-size ratio or validation technique proves that overfitting is absent. Appraise the entire modelling process, including predictor selection, missing-data handling, model tuning, validation, and whether every data-driven development step was repeated inside the validation procedure.

03 · What You Need to Know

Overfitting Produces Performance That Looks Better Than It Travels

What overfitting means in prediction research

A prediction model attempts to learn relationships between predictors and an outcome that will generalize to future individuals from the target population. During development, however, the model sees only a particular sample.

If the model is sufficiently flexible relative to the information in that sample, it can learn not only reproducible signal but also sample-specific noise. Apparent performance calculated from the development data then becomes optimistic: it makes the model look better than it is likely to perform in new observations.

Model fit How well the model describes or predicts observations used during its development.
Generalization How well the model predicts outcomes for relevant new individuals who were not used to construct it.

Start with how much information was available

Prediction-model sample size should be considered relative to the modelling task, not merely as a raw participant count. For binary outcomes, the number of participants experiencing the outcome can be especially important. For time-to-event models, the number of events likewise constrains available information.

A dataset containing 10,000 people may still provide limited information if the outcome occurs only 80 times and researchers consider dozens or hundreds of candidate predictor parameters.

PROBAST specifically asks whether there was a reasonable number of participants with the outcome relative to the number of candidate predictor parameters and whether overfitting and optimism were appropriately addressed. Empirical reviews using PROBAST have repeatedly identified insufficient events and inadequate handling of overfitting as common problems in prediction-model research.

Do not rely mechanically on the old events-per-variable rule

You may encounter the traditional recommendation of at least 10 outcome events per predictor variable in logistic or survival models. It can serve as a rough warning sign in some appraisals, but it is not a universal law and should not be treated as a definitive sample-size calculation.

Modern prediction-model sample-size methods consider additional features such as anticipated model fit, outcome frequency, total sample size, and the number of predictor parameters. The number of parameters also matters more than simply counting named predictors because categorical variables, nonlinear terms, splines, and interactions can consume multiple parameters.

The useful appraisal question is therefore not merely "Did they achieve 10 events per variable?" It is whether the development dataset contained enough information for the complexity of the model being estimated.

Count candidate predictors, not only predictors in the final model

A paper may present a final model containing six predictors and appear reassuringly simple. But perhaps the researchers started with 80 candidate variables, tried several transformations, examined interactions, tested alternative selection procedures, and retained the combination that performed best.

That search process is part of model development and contributes to overfitting.

Look in the Methods for the full set of candidate predictors and how they were selected. If the paper reports only the final variables without explaining the search that produced them, you cannot judge model complexity adequately.

Univariable screening is a warning sign

Some studies test each potential predictor individually and include only predictors with a statistically significant univariable association in the multivariable model. PROBAST specifically asks whether predictor selection based on univariable analysis was avoided.

A variable's value for multivariable prediction is not determined by whether its individual P-value crosses a threshold. Screening in this manner can create unstable selection and discard predictors that contribute useful information jointly with others.

Likewise, automated stepwise procedures can produce unstable models, particularly when information is limited relative to the number of candidate parameters.

Apparent performance is usually optimistic

If researchers develop a model and then calculate its performance using the exact same observations, the assessment is not an independent test. The model has already been optimized, directly or indirectly, for those data.

This apparent performance commonly overestimates performance in new individuals.

A particularly concerning paper might report excellent discrimination but provide no internal validation, no optimism correction, and no external validation. The performance estimate may describe how well the model learned its development sample rather than how well it predicts future cases.

Internal validation should reproduce the whole development process

Internal validation uses the development data to estimate how much performance may deteriorate in new observations. Resampling approaches such as bootstrapping are commonly used.

TRIPOD guidance emphasizes that all aspects of model building should be incorporated into internal validation, including predictor selection, penalization, transformations, and interaction testing. Validating only the final fitted equation ignores uncertainty introduced by the development process and can yield overoptimistic results.

This point is easy to miss. If researchers select variables using the entire dataset and only afterward cross-validate the already selected model, information from the validation observations has indirectly influenced model construction.

A simple train-test split is not automatically strong validation

Researchers sometimes randomly divide one dataset into a training set and a test set. This keeps the test observations out of model fitting and can provide an honest evaluation if the split is maintained properly.

However, splitting can be statistically inefficient, especially when the overall dataset is not large. The model is developed using only part of the available information, while performance is evaluated on a test sample that may itself be relatively small and imprecise.

Resampling methods such as bootstrapping or appropriately implemented cross-validation often use available development data more efficiently for internal validation. The best strategy depends on the modelling context and available information.

Cross-validation can still leak information

The phrase "we used cross-validation" should not end your appraisal.

Ask what happened inside each resampling iteration. Feature selection, preprocessing, normalization, imputation procedures that learn from data, hyperparameter tuning, and other data-driven operations should be structured so that validation observations do not influence the model being evaluated.

If researchers preprocess or select features using the complete dataset before cross-validation, information can leak from validation observations into model development. Reported performance may then remain optimistic.

Machine learning does not eliminate overfitting

Highly flexible machine-learning algorithms can model complex nonlinearities and interactions, but that flexibility can increase the capacity to fit noise when information is insufficient or tuning is poorly validated.

A systematic review of supervised machine-learning prediction models found high overall risk of bias in 88% of developed models assessed, with common problems involving inadequate sample size, missing-data handling, and failure to account adequately for overfitting and optimism.

The label "machine learning" therefore provides no assurance of generalizability. The same questions about information, model development, validation, and performance remain essential.

Discrimination alone is not enough

Prediction papers frequently emphasize discrimination, often using the area under the receiver operating characteristic curve or C-statistic. Discrimination measures how well a model distinguishes individuals who experience the outcome from those who do not.

Calibration asks a different question: do predicted probabilities agree adequately with observed outcome frequencies?

A model can discriminate reasonably well while producing systematically inaccurate probabilities. PROBAST therefore considers whether relevant performance measures, including both discrimination and calibration where appropriate, were evaluated. Empirical work has found calibration frequently omitted from prediction-model reports.

Performance concept Question
Discrimination Can the model distinguish individuals with different outcomes?
Calibration Do predicted probabilities correspond adequately to observed outcome frequencies?
Overall performance How accurate are predictions when considering the prediction errors more broadly?
Clinical or decision utility Would using the model improve decisions compared with relevant alternatives?

External validation answers a different question

External validation evaluates the developed model in data that were not used to construct it and that are meaningfully separate from the development sample. It provides evidence about how well the model performs beyond its original dataset.

Performance commonly changes when a model is evaluated in new populations or settings because case mix, predictor distributions, measurement procedures, outcome frequency, and other features differ.

External validation therefore adds evidence about transportability. It does not retroactively make poor development practices acceptable, and one favorable external validation does not prove universal performance.

A large drop between development and validation performance is informative

Suppose a paper reports an apparent C-statistic of 0.94 during development but optimism-corrected internal performance of 0.79. That difference suggests the original apparent performance substantially overstated what the model is likely to achieve in new data.

A similar drop during external validation can reveal limited generalizability, although population differences and measurement changes can also contribute.

Do not focus only on whether validation performance remains above an arbitrary threshold. Examine how much performance changes and whether calibration remains acceptable for the intended use.

Penalization and shrinkage can reduce overfitting

Methods such as ridge regression, lasso, elastic net, and other forms of penalization can constrain model complexity. Shrinkage can reduce overly extreme estimated coefficients. TRIPOD guidance discusses penalized approaches and bootstrap-derived shrinkage as strategies for reducing overfitting.

These techniques help, but their mere presence does not guarantee a good model. Penalty parameters may themselves require tuning, and tuning procedures need appropriate validation. A sophisticated method can still be implemented poorly. Statistics has no immunity-by-acronym clause.

Overfitting is not the only reason a model fails externally

A model can perform poorly in a new dataset because of overfitting, but other explanations include differences in population, outcome definition, predictor measurement, healthcare or educational practice, prevalence, temporal changes, and data quality.

BMJ guidance on external validation emphasizes evaluating whether participant selection, predictors, and outcome definitions in the validation data are appropriate for the intended model use.

Therefore, poor external performance is evidence that the model does not transport well to that validation setting. It does not identify overfitting as the sole cause.

Watch Out

An impressive accuracy, AUC, or C-statistic calculated on the same observations used to develop the model may be substantially optimistic. Always determine whether the reported performance is apparent, internally validated and optimism-corrected, or externally validated.

04 · A Practical Example

When a Prediction Model Looks Almost Too Good

Hypothetical Example

A model predicts student dropout with 94% accuracy

Suppose researchers develop a dropout-prediction model using data from 600 students, of whom 55 dropped out. They begin with 70 candidate predictors, use automated variable selection, and report a final model containing eight predictors. The model achieves an AUC of 0.94 on the complete development dataset.

Look beyond the final eight predictors. The modelling process considered 70 candidate predictors. The complexity of the search is therefore much greater than the final equation alone suggests.
Check the outcome information. Only 55 dropout events are available. That amount of outcome information may be limited relative to the candidate predictor parameters and modelling flexibility.
Ask where the AUC came from. The reported 0.94 was calculated using the same dataset used for predictor selection and model estimation. It is apparent performance and is vulnerable to optimism.
Look for internal validation. Suppose no bootstrap validation, cross-validation, or optimism correction is reported. You therefore have little evidence about how much the development performance would deteriorate in new students.
Look beyond discrimination. Suppose calibration is not reported. Even if the model ranks high-risk students effectively, you cannot determine whether predicted probabilities such as 20%, 50%, or 80% correspond adequately to observed dropout risk.
Interpret the performance cautiously. The paper demonstrates strong apparent fit to the development data. It has not yet demonstrated that an AUC of 0.94 will generalize to new students.

If appropriately implemented internal validation reduced the AUC to 0.77, that corrected estimate would provide a more realistic indication of expected performance than the original 0.94. Independent external validation would then address how well the finalized model performs in genuinely separate data.

05 · What Researchers Often Get Wrong

Common Mistakes When Appraising Prediction Models

Misconception

A Very High AUC Means the Model Is Excellent

Not necessarily. Determine whether the AUC was calculated on development data, optimism-corrected internal validation, or independent validation. Also examine calibration and whether the model is useful for its intended decision.

Misconception

Only the Number of Predictors in the Final Model Matters

The entire modelling search matters. Candidate predictors, transformations, interactions, selection procedures, and tuning decisions can all contribute to overfitting even if the final model appears simple.

Misconception

Cross-Validation Automatically Prevents Overfitting

Cross-validation must be implemented correctly. Data-driven preprocessing, feature selection, and tuning performed before the validation split can leak information and produce optimistic estimates.

Misconception

Ten Events per Predictor Guarantees a Reliable Model

No single events-per-variable threshold guarantees adequate development. Sample-size requirements depend on outcome frequency, model complexity, anticipated performance, predictor parameters, and other design characteristics.

Misconception

Machine Learning Solves the Overfitting Problem

Flexible algorithms can themselves overfit. Appropriate tuning, validation, adequate information, transparent development, and evaluation of relevant performance measures remain necessary.

Misconception

External Validation Makes the Development Study Methodologically Sound

External validation provides important evidence about performance in new data, but it does not erase biased predictor selection, inappropriate missing-data handling, leakage, or other development problems.

06 · What This Means for You

Trace How the Model Was Built Before Trusting How Well It Performed

When appraising a prediction model, the headline performance statistic should be near the end of your evaluation rather than the beginning. First determine how much information the researchers had, how many modelling decisions they made, and whether performance was assessed in a way that accounts for those decisions.

A simple decision framework

If performance is reported only on the model-development data
Treat it as apparent performance that may substantially overestimate generalization.
If many candidate predictor parameters were explored with limited outcome information
Increase concern about overfitting and model instability.
If appropriate resampling validation repeats the complete model-development process
Give greater weight to the optimism-corrected performance than to the apparent development performance.
If the model performs well in genuinely independent data
Confidence in its ability to generalize to settings resembling that validation population can increase.
If only discrimination is reported
Look for calibration and performance measures relevant to the model's intended use before judging its usefulness.

A prediction model should ultimately be judged by its ability to make useful predictions for relevant new individuals, not by how impressively it describes the sample from which it learned.

07 · A Quick Checklist

How to Check a Prediction Model Paper for Overfitting

Before trusting the reported model performance, check:
How many participants and outcome events were available for model development.
How many candidate predictor parameters, transformations, and interactions were considered, not merely how many predictors survived into the final model.
Whether predictor selection relied on univariable significance testing or unstable automated procedures.
How missing data, continuous predictors, outliers, and other preprocessing decisions were handled.
Whether internal validation was performed and whether the entire model-building process was repeated within it.
Whether performance estimates were corrected for optimism rather than reported only on the development data.
Whether both discrimination and calibration are reported where appropriate.
Whether the finalized model has been evaluated in independent data relevant to the intended population and setting.
08 · Frequently Asked Questions

Questions About Overfitting in Prediction Models

What is overfitting in simple terms?

Overfitting occurs when a model learns sample-specific noise as though it were reproducible signal. It then performs better on the data used to develop it than it is likely to perform on relevant new observations.

Can a simple regression model overfit?

Yes. Overfitting is not unique to machine learning. Regression models can overfit when their complexity and data-driven development decisions are too great for the amount of available information.

Does cross-validation eliminate overfitting?

No. Properly implemented cross-validation can estimate generalization performance and help tune models, but leakage, inappropriate preprocessing, repeated tuning against validation results, or inadequate sample size can still produce optimistic conclusions.

What is optimism in a prediction model?

Optimism is the extent to which performance estimated from the development process is better than the performance expected in comparable new observations. Internal validation methods can estimate and correct for this optimism.

What is the difference between internal and external validation?

Internal validation uses the development dataset, usually through resampling or carefully structured data splitting, to estimate optimism and model stability. External validation evaluates the finalized model using separate data that were not used for model development.

Is a high AUC enough to show that a model is useful?

No. AUC measures discrimination, not calibration or decision usefulness. A model can rank individuals effectively while producing inaccurate risk estimates or offering little benefit for the decision it is intended to support.

Does poor external validation prove that the original model was overfitted?

Not necessarily. Overfitting is one possible explanation, but differences in populations, predictor measurement, outcome definitions, case mix, prevalence, time period, or setting can also reduce external performance.

09 · The Bottom Line

A Prediction Model Must Perform Beyond the Data That Created It

The Bottom Line

Recognize possible overfitting by examining whether model complexity and data-driven searching were excessive relative to the available information, and whether reported performance was properly validated rather than calculated only on the development data. An impressive development AUC or accuracy can be substantially optimistic.

Give more weight to appropriately optimism-corrected internal validation and performance in independent data, while also examining calibration and the intended use of the model. The real test of a prediction model is not how well it remembers its development sample but how reliably it predicts relevant new cases.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes