01 · The Question
Could Something Else Be Driving Both Variables?
You observe that students who use a learning platform more frequently earn higher grades. Perhaps the platform improves learning. But students with stronger motivation might also be more likely to use the platform and to study consistently, attend class, and complete assessments successfully.
If so, some of the observed relationship between platform use and grades may reflect differences in motivation rather than the effect of platform use itself.
This is the central concern of confounding. When you want to interpret an association causally, you need to ask whether common causes of the proposed exposure and outcome could create, exaggerate, reduce, or otherwise distort the relationship you observe.
02 · The Short Answer
Confounding Can Make an Association Look Like a Causal Effect
In Brief
Confounding can occur when causes of the exposure or treatment also influence the outcome, so the observed relationship between X and Y does not represent only the causal effect of X on Y.
Identifying confounding is fundamentally a causal problem. A variable should not be treated as a confounder merely because it correlates with X and Y, changes a regression coefficient, or happens to be available in the dataset. You need a defensible account of the causal structure.
03 · What You Need to Know
Confounding Is About How the Exposure and Outcome Came to Be Associated
Think About Common Causes, Not Just Third Variables
A simple way to begin is with three variables. Suppose X is your exposure, Y is your outcome, and C is another variable:
C → X
C → Y
X → Y
If C helps determine both X and Y, comparing people with different values of X may also compare people who differ systematically in C. Part of the observed X–Y association may therefore be attributable to that difference.
For example, students who voluntarily attend supplemental tutorials may obtain higher examination scores. Prior academic preparation could influence both tutorial attendance and later performance. Simply comparing attendees with non-attendees may therefore mix the possible effect of tutorials with pre-existing differences between the groups.
Confounding Is Not the Same as Correlation With Both Variables
A familiar rule says that a confounder must be associated with the exposure, associated with the outcome, and not lie on the causal pathway between them. Although this can be a useful introductory heuristic, modern causal inference cautions against identifying adjustment variables solely through observed statistical associations.
Why? Because variables can be statistically associated with both X and Y for several causal reasons. Some should be adjusted for. Some should not. Others can actually create bias if conditioned on.
Watch Out
Do not select confounders simply by testing which variables are statistically significant or by checking whether adding them changes the exposure coefficient. Confounding concerns the causal structure that generated the data, not merely patterns discovered in the sample.
A Confounder Is Different From a Mediator
Confounder
A variable involved in a noncausal path between the exposure and outcome that may need to be addressed to estimate the causal effect of interest.
Mediator
A variable through which some of the causal effect of the exposure on the outcome operates.
Suppose study time influences academic achievement partly because it improves mastery of course material:
Study time → mastery → achievement
If your goal is to estimate the total effect of study time on achievement, controlling for mastery may remove part of the very effect you want to estimate. Calling every variable related to the outcome a confounder therefore creates conceptual trouble rather quickly.
Some Variables Can Create Bias When You Control for Them
A collider is a variable caused by two other variables. Consider a simplified structure:
X → C ← Y
Here, C is a common effect rather than a common cause. Conditioning on C can induce an association between X and Y even when the relevant path would otherwise be closed.
This is one reason “control for everything” is not a safe default. Adjustment requires deciding which variables are needed for the causal effect being estimated and which should be left alone.
Directed Acyclic Graphs Can Make Your Assumptions Visible
A directed acyclic graph, usually called a DAG, represents assumptions about causal relationships using variables connected by arrows. Researchers can use DAGs to reason about open and closed paths between an exposure and outcome and to identify candidate adjustment sets under those assumptions.
A DAG does not learn the correct causal structure for you. Its arrows must be justified using substantive knowledge, previous evidence, temporal ordering, and defensible assumptions. Its value is that those assumptions become explicit enough to inspect and challenge.
Randomization Addresses Confounding Differently
In a well-conducted randomized experiment, treatment assignment is generated by a random mechanism. This is designed to make treatment groups exchangeable with respect to baseline causes of the outcome, apart from random imbalances.
Observational exposures usually lack that mechanism. People select, receive, or encounter exposures for reasons that may themselves affect the outcome. Understanding why some participants received or experienced X while others did not is therefore often central to identifying potential confounding.
Measured Confounding Is Not the Only Concern
Statistical adjustment can only directly use variables represented adequately in the data. Important common causes may be unmeasured, measured poorly, or represented by imperfect proxies.
This distinction matters because saying “we controlled for confounders” can sound stronger than the design warrants. A more defensible claim specifies which variables were adjusted for and why those variables were considered sufficient or useful under the assumed causal structure.
Confounders Can Be Measured Badly
Even when an important confounder is included, measurement error can leave residual confounding. A crude measure of socioeconomic circumstances, prior achievement, disease severity, motivation, or another complex construct may fail to capture enough of the relevant variation to achieve adequate adjustment.
This connects confounding directly to whether measurement error could distort the pattern you expect . Having a variable in the dataset is not equivalent to having measured it sufficiently well for the intended adjustment.
Confounding Can Strengthen, Weaken, or Reverse an Association
Confounding does not always create an impressive positive relationship from nothing. Depending on the causal structure and associations involved, it can make an effect appear stronger or weaker than it is, conceal an effect, or contribute to an observed association whose direction differs from the causal effect of interest.
You therefore cannot safely assume that “controlling for confounding” will always make an estimate smaller. The direction of the distortion depends on the underlying structure.
Confounding Is Only One Alternative Explanation
Even excellent control of confounding does not resolve every threat to causal interpretation. The proposed outcome might influence the exposure, creating ambiguity about the direction of causation . Selection processes and measurement problems can also generate misleading patterns.
It is therefore useful to treat confounding as one member of a broader set of plausible explanations for an expected research finding .
04 · A Practical Example
Does Optional Tutorial Attendance Improve Academic Performance?
Hypothetical Example
Tutorial attendance and examination scores
A university offers optional weekly tutorials. Students who attend regularly average higher final examination scores than students who rarely or never attend. The researcher wants to conclude that tutorial attendance improves achievement.
Observed pattern Students with greater tutorial attendance have higher examination scores.
Proposed explanation Tutorial participation improves understanding, which subsequently improves examination performance.
Potential confounding Academic motivation may influence both voluntary tutorial attendance and behaviors that improve examination performance. Prior achievement, available study time, or other pre-exposure characteristics could also matter.
Design implication Before collecting data, the researcher identifies plausible common causes using substantive knowledge and the assumed causal structure, then measures important pre-exposure variables adequately rather than simply collecting every convenient demographic characteristic.
Interpretation Adjustment may make the estimated tutorial effect more credible under the relevant assumptions, but an observational analysis cannot automatically guarantee that all confounding has been eliminated.
The important move occurs before the regression model is fitted. The researcher asks why some students attend tutorials and others do not, and whether those same reasons could also affect examination performance.
06 · What This Means for You
Identify Confounding Before Choosing Your Adjustment Variables
Begin with the causal question, not the regression model. Define the exposure, outcome, population, relevant time points, and causal effect you want to estimate. Then ask what caused participants to have different values of the exposure and which of those causes could also affect the outcome.
A simple decision framework
If a variable plausibly causes both X and Y
Consider whether it belongs in the adjustment strategy under your causal assumptions and measure it adequately before the outcome where possible.
If a variable occurs after X and lies on the pathway from X to Y
Do not automatically treat it as a confounder. Determine whether your estimand concerns the total effect or a more specific causal effect.
If a variable is a common effect of X and Y or their causes
Be cautious about conditioning on it because collider bias may be introduced.
If an important common cause cannot be measured adequately
Consider design alternatives, sensitivity analyses, or a more limited causal interpretation.
If you are unsure what role a variable plays
Return to substantive theory and the assumed causal structure rather than allowing software or significance tests to make the decision.
This exercise may expose a deeper issue: perhaps the proposed study cannot adequately separate your preferred explanation from a serious rival. That does not necessarily end the project, but it should influence whether the research idea can distinguish among the explanations that matter most .
07 · A Quick Checklist
Check for Confounding Before Finalizing the Design
Before interpreting X as a cause of Y, check:
Define the exposure, outcome, population, time frame, and causal effect you want to estimate.
Ask why some participants have different values of X and whether those causes could also influence Y.
Use substantive knowledge and causal reasoning rather than significance tests alone to identify adjustment variables.
Distinguish potential confounders from mediators and colliders before adjusting for them.
Measure important pre-exposure common causes with sufficient quality for the intended analysis.
Consider drawing a DAG to make causal assumptions and possible biasing paths explicit.
Assess whether important common causes remain unmeasured or poorly measured.
Match the strength of the causal conclusion to the confounding assumptions the study can reasonably defend.
09 · The Bottom Line
Do Not Let the Regression Model Decide What Counts as a Confounder
The Bottom Line
Confounding could explain or distort the relationship you want to study when causes of the exposure also affect the outcome, making exposed and unexposed groups systematically different in ways relevant to that outcome.
Addressing it requires more than adding control variables. Define the causal question, reason about how the exposure and outcome arise, identify an appropriate adjustment strategy, and measure the relevant variables well. Statistical adjustment is most defensible when it follows causal reasoning rather than replacing it.
11 · Cite this Guide
How to Cite This Guide
This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.
Recommended (Field Guide)
APA
MLA
Chicago
Copy Citation