Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Could Confounding Explain the Relationship You Want to Study?

An observed relationship between X and Y may partly reflect other causes that influence both variables. Identifying confounding requires causal reasoning about how the data were generated, not simply controlling for every available variable.

671
Could Confounding Explain the Relationship? Guide 671 of 760
01 · The Question

Could Something Else Be Driving Both Variables?

You observe that students who use a learning platform more frequently earn higher grades. Perhaps the platform improves learning. But students with stronger motivation might also be more likely to use the platform and to study consistently, attend class, and complete assessments successfully.

If so, some of the observed relationship between platform use and grades may reflect differences in motivation rather than the effect of platform use itself.

This is the central concern of confounding. When you want to interpret an association causally, you need to ask whether common causes of the proposed exposure and outcome could create, exaggerate, reduce, or otherwise distort the relationship you observe.

02 · The Short Answer

Confounding Can Make an Association Look Like a Causal Effect

In Brief

Confounding can occur when causes of the exposure or treatment also influence the outcome, so the observed relationship between X and Y does not represent only the causal effect of X on Y.

Identifying confounding is fundamentally a causal problem. A variable should not be treated as a confounder merely because it correlates with X and Y, changes a regression coefficient, or happens to be available in the dataset. You need a defensible account of the causal structure.

03 · What You Need to Know

Confounding Is About How the Exposure and Outcome Came to Be Associated

Think About Common Causes, Not Just Third Variables

A simple way to begin is with three variables. Suppose X is your exposure, Y is your outcome, and C is another variable:

C → X
C → Y
X → Y

If C helps determine both X and Y, comparing people with different values of X may also compare people who differ systematically in C. Part of the observed X–Y association may therefore be attributable to that difference.

For example, students who voluntarily attend supplemental tutorials may obtain higher examination scores. Prior academic preparation could influence both tutorial attendance and later performance. Simply comparing attendees with non-attendees may therefore mix the possible effect of tutorials with pre-existing differences between the groups.

Confounding Is Not the Same as Correlation With Both Variables

A familiar rule says that a confounder must be associated with the exposure, associated with the outcome, and not lie on the causal pathway between them. Although this can be a useful introductory heuristic, modern causal inference cautions against identifying adjustment variables solely through observed statistical associations.

Why? Because variables can be statistically associated with both X and Y for several causal reasons. Some should be adjusted for. Some should not. Others can actually create bias if conditioned on.

Watch Out

Do not select confounders simply by testing which variables are statistically significant or by checking whether adding them changes the exposure coefficient. Confounding concerns the causal structure that generated the data, not merely patterns discovered in the sample.

A Confounder Is Different From a Mediator

Confounder A variable involved in a noncausal path between the exposure and outcome that may need to be addressed to estimate the causal effect of interest.
Mediator A variable through which some of the causal effect of the exposure on the outcome operates.

Suppose study time influences academic achievement partly because it improves mastery of course material:

Study time → mastery → achievement

If your goal is to estimate the total effect of study time on achievement, controlling for mastery may remove part of the very effect you want to estimate. Calling every variable related to the outcome a confounder therefore creates conceptual trouble rather quickly.

Some Variables Can Create Bias When You Control for Them

A collider is a variable caused by two other variables. Consider a simplified structure:

X → C ← Y

Here, C is a common effect rather than a common cause. Conditioning on C can induce an association between X and Y even when the relevant path would otherwise be closed.

This is one reason “control for everything” is not a safe default. Adjustment requires deciding which variables are needed for the causal effect being estimated and which should be left alone.

Directed Acyclic Graphs Can Make Your Assumptions Visible

A directed acyclic graph, usually called a DAG, represents assumptions about causal relationships using variables connected by arrows. Researchers can use DAGs to reason about open and closed paths between an exposure and outcome and to identify candidate adjustment sets under those assumptions.

A DAG does not learn the correct causal structure for you. Its arrows must be justified using substantive knowledge, previous evidence, temporal ordering, and defensible assumptions. Its value is that those assumptions become explicit enough to inspect and challenge.

Randomization Addresses Confounding Differently

In a well-conducted randomized experiment, treatment assignment is generated by a random mechanism. This is designed to make treatment groups exchangeable with respect to baseline causes of the outcome, apart from random imbalances.

Observational exposures usually lack that mechanism. People select, receive, or encounter exposures for reasons that may themselves affect the outcome. Understanding why some participants received or experienced X while others did not is therefore often central to identifying potential confounding.

Measured Confounding Is Not the Only Concern

Statistical adjustment can only directly use variables represented adequately in the data. Important common causes may be unmeasured, measured poorly, or represented by imperfect proxies.

This distinction matters because saying “we controlled for confounders” can sound stronger than the design warrants. A more defensible claim specifies which variables were adjusted for and why those variables were considered sufficient or useful under the assumed causal structure.

Confounders Can Be Measured Badly

Even when an important confounder is included, measurement error can leave residual confounding. A crude measure of socioeconomic circumstances, prior achievement, disease severity, motivation, or another complex construct may fail to capture enough of the relevant variation to achieve adequate adjustment.

This connects confounding directly to whether measurement error could distort the pattern you expect. Having a variable in the dataset is not equivalent to having measured it sufficiently well for the intended adjustment.

Confounding Can Strengthen, Weaken, or Reverse an Association

Confounding does not always create an impressive positive relationship from nothing. Depending on the causal structure and associations involved, it can make an effect appear stronger or weaker than it is, conceal an effect, or contribute to an observed association whose direction differs from the causal effect of interest.

You therefore cannot safely assume that “controlling for confounding” will always make an estimate smaller. The direction of the distortion depends on the underlying structure.

Confounding Is Only One Alternative Explanation

Even excellent control of confounding does not resolve every threat to causal interpretation. The proposed outcome might influence the exposure, creating ambiguity about the direction of causation. Selection processes and measurement problems can also generate misleading patterns.

It is therefore useful to treat confounding as one member of a broader set of plausible explanations for an expected research finding.

04 · A Practical Example

Does Optional Tutorial Attendance Improve Academic Performance?

Hypothetical Example

Tutorial attendance and examination scores

A university offers optional weekly tutorials. Students who attend regularly average higher final examination scores than students who rarely or never attend. The researcher wants to conclude that tutorial attendance improves achievement.

Observed pattern Students with greater tutorial attendance have higher examination scores.
Proposed explanation Tutorial participation improves understanding, which subsequently improves examination performance.
Potential confounding Academic motivation may influence both voluntary tutorial attendance and behaviors that improve examination performance. Prior achievement, available study time, or other pre-exposure characteristics could also matter.
Design implication Before collecting data, the researcher identifies plausible common causes using substantive knowledge and the assumed causal structure, then measures important pre-exposure variables adequately rather than simply collecting every convenient demographic characteristic.
Interpretation Adjustment may make the estimated tutorial effect more credible under the relevant assumptions, but an observational analysis cannot automatically guarantee that all confounding has been eliminated.

The important move occurs before the regression model is fitted. The researcher asks why some students attend tutorials and others do not, and whether those same reasons could also affect examination performance.

05 · What Researchers Often Get Wrong

Common Mistakes When Thinking About Confounding

Misconception

Any Variable Associated With X and Y Is a Confounder

Statistical association alone does not determine causal role. A variable may be a common cause, mediator, collider, consequence of another variable, proxy, or something else entirely. The causal structure matters.

Misconception

Controlling for More Variables Always Reduces Bias

No. Unnecessary adjustment can reduce precision, adjustment for mediators can remove part of an effect of interest, and conditioning on colliders can introduce bias. More covariates do not automatically produce a more causally credible model.

Misconception

A Variable Is a Confounder Only if It Is Statistically Significant

Significance testing is not a reliable procedure for identifying confounders. Whether adjustment is needed depends on the causal structure and estimand, not whether a coefficient crosses a conventional significance threshold in one sample.

Misconception

If the Adjusted and Unadjusted Estimates Are Similar, There Is No Confounding

A small change after adjustment does not prove the absence of confounding. The chosen adjustment set may omit important common causes, variables may be measured poorly, or different sources of bias may operate in opposing directions.

Misconception

Including Demographic Variables Means Confounding Has Been Handled

Age, sex, socioeconomic indicators, and other demographics may sometimes be relevant, but their routine inclusion does not establish adequate confounding control. The appropriate variables depend on the specific exposure, outcome, population, timing, and causal question.

Misconception

Confounding Makes an Observational Study Worthless

No. Observational studies can answer many important questions, and causal analyses can be informative under defensible assumptions and appropriate designs. The challenge is to make those assumptions explicit and avoid claiming that statistical adjustment mechanically eliminates all alternative explanations.

06 · What This Means for You

Identify Confounding Before Choosing Your Adjustment Variables

Begin with the causal question, not the regression model. Define the exposure, outcome, population, relevant time points, and causal effect you want to estimate. Then ask what caused participants to have different values of the exposure and which of those causes could also affect the outcome.

A simple decision framework

If a variable plausibly causes both X and Y
Consider whether it belongs in the adjustment strategy under your causal assumptions and measure it adequately before the outcome where possible.
If a variable occurs after X and lies on the pathway from X to Y
Do not automatically treat it as a confounder. Determine whether your estimand concerns the total effect or a more specific causal effect.
If a variable is a common effect of X and Y or their causes
Be cautious about conditioning on it because collider bias may be introduced.
If an important common cause cannot be measured adequately
Consider design alternatives, sensitivity analyses, or a more limited causal interpretation.
If you are unsure what role a variable plays
Return to substantive theory and the assumed causal structure rather than allowing software or significance tests to make the decision.

This exercise may expose a deeper issue: perhaps the proposed study cannot adequately separate your preferred explanation from a serious rival. That does not necessarily end the project, but it should influence whether the research idea can distinguish among the explanations that matter most.

07 · A Quick Checklist

Check for Confounding Before Finalizing the Design

Before interpreting X as a cause of Y, check:
Define the exposure, outcome, population, time frame, and causal effect you want to estimate.
Ask why some participants have different values of X and whether those causes could also influence Y.
Use substantive knowledge and causal reasoning rather than significance tests alone to identify adjustment variables.
Distinguish potential confounders from mediators and colliders before adjusting for them.
Measure important pre-exposure common causes with sufficient quality for the intended analysis.
Consider drawing a DAG to make causal assumptions and possible biasing paths explicit.
Assess whether important common causes remain unmeasured or poorly measured.
Match the strength of the causal conclusion to the confounding assumptions the study can reasonably defend.
08 · Frequently Asked Questions

Questions About Confounding in Research

What is confounding in simple terms?

Confounding occurs when the groups being compared differ in causes of the outcome for reasons connected with the exposure. As a result, the observed exposure-outcome association does not represent only the causal effect of the exposure.

What is a confounding variable?

The term is commonly used for a variable involved in confounding, often a common cause of the exposure and outcome. In causal inference, however, it is usually more useful to reason about the causal structure and an appropriate adjustment set than to classify variables using simple association-based rules.

Is a mediator a confounder?

Not simply because it is related to both exposure and outcome. A mediator lies on a causal pathway through which the exposure affects the outcome. Whether it should be adjusted for depends on the causal effect the researcher is trying to estimate.

Can regression eliminate confounding?

Regression can be used to adjust for measured variables under appropriate assumptions, but fitting an adjusted model does not itself guarantee control of confounding. The relevant variables must be identified appropriately and measured sufficiently well, and the analytical model must suit the causal question.

What is residual confounding?

Residual confounding is confounding that remains after adjustment. It may arise because an important confounder was omitted, measured inaccurately, categorized too crudely, or otherwise inadequately represented in the analysis.

Can randomization eliminate confounding?

Random assignment is designed to break systematic links between baseline causes of the outcome and treatment assignment, which provides a strong basis for causal comparison. Random imbalances can still occur, and post-randomization problems such as nonadherence, attrition, or missing data can complicate the analysis.

Can confounding reverse the apparent direction of an effect?

Yes, under some causal and numerical structures. Confounding can strengthen, weaken, conceal, or potentially reverse an observed association relative to the causal effect of interest. Its direction should therefore be reasoned about rather than assumed.

09 · The Bottom Line

Do Not Let the Regression Model Decide What Counts as a Confounder

The Bottom Line

Confounding could explain or distort the relationship you want to study when causes of the exposure also affect the outcome, making exposed and unexposed groups systematically different in ways relevant to that outcome.

Addressing it requires more than adding control variables. Define the causal question, reason about how the exposure and outcome arise, identify an appropriate adjustment strategy, and measure the relevant variables well. Statistical adjustment is most defensible when it follows causal reasoning rather than replacing it.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes