03 · What You Need to Know
The Most Impressive Result Is Not Necessarily the Best Summary of the Evidence
Selection Can Change the Apparent Strength of a Study
Imagine a study with ten relevant analyses. Nine produce small or uncertain effects. One produces a large statistically significant effect.
If the manuscript presents only that one analysis, readers encounter a different evidential object from the complete study.
The reported estimate may be calculated correctly. The statistical test may be performed correctly. The problem is that selection has removed the context needed to judge how exceptional the finding was among the analyses conducted.
This is the central danger of selective reporting: the selected evidence can be accurate while the resulting impression is inaccurate.
Start With the Analyses That Had Evidential Priority Before You Saw the Results
When results compete for attention, the study's design can help determine which deserve greatest interpretive weight.
A prespecified primary outcome generally has a different role from a post hoc subgroup analysis. A planned confirmatory model differs from one selected after trying several specifications. A hypothesis formulated before the results differs from one generated by them.
These distinctions do not make secondary or exploratory findings worthless. They prevent favorable later findings from quietly replacing the questions the study was originally designed to answer.
For medical journal reporting, ICMJE recommends providing data on all primary and secondary outcomes identified in the Methods and linking conclusions to study goals without making claims inadequately supported by the data.
Different disciplines use different reporting standards, but the underlying logic is broadly useful: determine evidential priority from the design and research question, not simply from which result is most attractive after analysis.
Do Not Count Significant Results and Call the Majority the Truth
The complete evidence is not necessarily determined by counting how many analyses are significant.
Suppose two primary outcomes show no clear effect while six exploratory variants of a secondary measure are significant. A simple six-versus-two tally would be misleading because those tests do not necessarily have equal status or independence.
Likewise, one large, precise primary effect may reasonably carry more evidential weight than several noisy peripheral analyses.
“Complete results” means evaluating the relevant evidence according to its methodological role, quality, precision, multiplicity, and relationship to the research question. It does not mean democratic voting among p-values.
Ask Why the Selected Result Looks Better
Sometimes a selected result is stronger because it genuinely provides a better analysis.
Perhaps the primary model violated assumptions and a revised model is demonstrably more appropriate. Perhaps one measure has substantially better validity. Perhaps an effect is expected theoretically only in a particular population.
Those are methodological arguments that can be evaluated.
Other times, the result looks better because researchers searched until they found it.
The distinction is crucial. If the reason for preferring the result is essentially “this specification gives the clearest significance,” the selection itself becomes part of the evidence readers need to know about.
Analytical Instability Is a Result
Suppose your main association is statistically significant under one model but becomes small or uncertain when reasonable alternative specifications are used.
That does not mean you must average the models mechanically or abandon the study.
It means the conclusion depends on analytical choices.
That dependence is scientifically relevant. Hiding it converts a fragile finding into an apparently robust one.
This is why analytical flexibility becomes questionable when only the specification producing the preferred conclusion receives visibility.
A Null Primary Result Does Not Become Less Primary When a Secondary Result Is Significant
This situation is particularly tempting.
The primary analysis disappoints. A secondary outcome, subgroup, or alternative analysis looks excellent. Suddenly the paper has a possible success story.
You can report and discuss that finding. But the favorable result should not retrospectively replace the study's primary evidential question.
Research on spin demonstrates how easily this happens. Boutron and colleagues examined randomized trials with statistically nonsignificant primary outcomes and documented strategies that emphasized statistically significant secondary outcomes, subgroup analyses, within-group comparisons, or other favorable findings while distracting readers from the primary result.
In their sample, spin appeared in 37.5% of abstract Results sections and 58.3% of abstract Conclusions. These figures describe a specific sample of randomized trials published in 2006, not research generally. They illustrate, however, how selective emphasis can change the interpretation of otherwise accurately reported findings.
Complete Reporting Does Not Require Equal Prominence
A common objection is practical: “If I give every result equal attention, the paper will become unreadable.”
You do not need to.
A primary result may receive several paragraphs. A relevant secondary result may receive a sentence and a table entry. Detailed sensitivity analyses may sit in supplementary material. Exploratory results can be grouped.
The objective is not equal word counts. It is an accurate evidential hierarchy.
Selective emphasis
Gives more attention to findings because they are central, informative, methodologically credible, or substantively important while preserving relevant contrary evidence.
Selective reporting
Uses visibility or omission in a result-dependent way that leaves readers with a materially distorted impression of the evidence.
Sometimes the Right Response Is to Narrow the Claim
Suppose you hoped to conclude that an intervention improves academic performance broadly. The complete results show improvement only in engagement and one assessment measure, with no clear changes in retention or overall grades.
You do not necessarily need to conclude that the intervention “does not work.”
You may need a narrower claim:
“The intervention showed evidence of improving engagement and one assessment outcome, but the study did not demonstrate broad improvement across academic outcomes.”
The complete results have not destroyed the contribution. They have defined its boundaries.
Sometimes the Right Response Is to Increase Uncertainty
Suppose several reasonable analyses produce effect estimates in the same direction but with substantially different magnitudes and precision.
The central conclusion might survive, but with less certainty:
“The analyses generally suggest a positive association, although its magnitude was sensitive to model specification.”
Again, the paper remains publishable. The sentence simply communicates what the evidence actually supports.
Sometimes the Complete Results Change the Conclusion Entirely
There will also be cases where the selected finding cannot reasonably carry the paper's central claim.
If the primary analysis is null, numerous related analyses are null, and one post hoc subgroup among many produces a significant result, the complete evidence may not support concluding that the intervention is effective.
The subgroup finding can remain an exploratory observation worth investigating. What changes is the headline claim.
Do Not Treat the Abstract as a Place to Restore the Better Story
Researchers sometimes report mixed evidence responsibly in the Results and Discussion, then write an abstract that emphasizes only the favorable subset.
That is especially consequential because many readers will never reach the full paper.
The abstract should therefore reflect the same evidential hierarchy as the article. Highlighting the most interesting result is reasonable only when the central evidence remains visible.
Selective Reporting Can Affect the Wider Evidence Base
The consequences extend beyond one paper.
Chan and colleagues compared trial protocols with publications and found that statistically significant efficacy and harm outcomes had higher odds of being fully reported than nonsignificant outcomes. They also found that 62% of the trials examined had at least one primary outcome that was changed, introduced, or omitted between protocol and publication.
Those findings come from a specific sample of randomized trials and should not be generalized mechanically across disciplines. They illustrate why selective reporting matters cumulatively: if favorable results repeatedly have greater visibility, later readers and evidence syntheses receive a systematically altered research record.
Watch Out
If the study looks impressive only after you decide which results the reader is allowed to see, the selection process is doing part of the evidential work. That is a reason to revise the claim, not merely the Results section.
04 · A Practical Example
When One Excellent Result Sits Inside a Much Less Exciting Study
Hypothetical Example
An Intervention With One Striking Finding
A researcher evaluates a new teaching intervention. The prespecified primary outcome is final examination performance. Secondary outcomes include retention, attendance, engagement, satisfaction, and academic confidence.
Primary outcome
The examination-score difference is small and statistically inconclusive.
Secondary outcomes
Retention, attendance, and confidence show little clear difference. Engagement is somewhat higher.
Striking result
Satisfaction is substantially higher in the intervention group and produces p <.001.
Selected story
The manuscript focuses on satisfaction and concludes that the intervention is highly effective.
Complete story
The intervention substantially improved satisfaction, may have improved engagement, but did not provide clear evidence of broader academic benefits in the outcomes examined.
The complete story is less impressive if the hoped-for conclusion was “this intervention improves academic achievement.”
It may nevertheless be more scientifically useful. Perhaps satisfaction matters for implementation or student experience. That contribution can stand on its own without being inflated into evidence of academic effectiveness.
The strongest result does not need to disappear. It simply needs to remain the result it actually is.