03 · What You Need to Know
Sensitivity and Precision Answer Different Questions
Sensitivity asks: How much of the relevant evidence did you find?
Sensitivity, also called recall in information retrieval, concerns the proportion of relevant records successfully retrieved.
High sensitivity is therefore about coverage of the relevant literature, not the total number of records retrieved. A search can retrieve 20,000 records and still have poor sensitivity if important relevant studies are absent.
In practice, the complete number of relevant records is usually unknown. Researchers therefore use benchmark studies, relative recall, citation searching, reference checking, and missed-record analysis to evaluate whether a search appears sufficiently sensitive.
Precision asks: How much of what you found is actually relevant?
Precision concerns the composition of the retrieved result set.
Higher precision usually means less screening. If two searches retrieve the same relevant studies but one returns 2,000 records and the other returns 10,000, the smaller set is operationally preferable because researchers have fewer irrelevant records to assess.
But precision alone cannot tell you whether the search is good. A strategy retrieving ten highly relevant articles and missing another 90 would look wonderfully tidy while performing terribly if comprehensive evidence identification were the objective.
The same relevant records appear in both formulas
Sensitivity and precision are closely related because both include relevant records retrieved in the numerator. Their denominators differ.
| Measure |
Question it answers |
Main concern |
| Sensitivity |
What proportion of the relevant evidence did the search retrieve? |
Missed relevant records |
| Precision |
What proportion of retrieved records are relevant? |
Irrelevant retrieval and screening burden |
That difference explains why the measures can move in opposite directions.
Why increasing sensitivity often reduces precision
Suppose you are searching for artificial intelligence in higher education. You begin with highly specific terminology and retrieve a relatively clean set of records.
Then you add alternative names, acronyms, spelling variants, broader terminology, and additional subject headings so that studies using less predictable language can also be found.
Some newly retrieved records are relevant. Many others may not be.
Sensitivity can therefore increase because you capture relevant records previously missed, while precision falls because the broader strategy also admits more irrelevant records. Cochrane explicitly notes this relationship: increasing search comprehensiveness or sensitivity will reduce precision and usually retrieve more non-relevant reports.
This is not necessarily a flaw. It can be the expected cost of protecting against missed evidence.
Why increasing precision can reduce sensitivity
The reverse happens when you tighten a search.
You might add another concept with AND, restrict searching to titles, remove broad synonyms, apply a population filter, impose a study-design filter, or require a particular outcome.
These changes can eliminate irrelevant records and improve precision. Yet relevant records that fail the additional requirement disappear as well.
This is why adding more search terms can make a search worse even when the resulting set looks much cleaner.
The trade-off is real, but it is not always perfectly symmetrical
It is tempting to imagine sensitivity and precision as a simple seesaw: every gain in one must produce an equivalent loss in the other. Real searches are more interesting than that.
A badly chosen ambiguous synonym may generate thousands of irrelevant records while retrieving no unique relevant studies. Removing it improves precision without harming sensitivity.
Likewise, discovering an important missing synonym may retrieve several relevant studies with only modest additional noise, improving sensitivity at relatively little precision cost.
Search development should therefore look first for these relatively efficient improvements. The difficult trade-off begins when further gains in one measure genuinely require sacrificing the other.
For systematic reviews, sensitivity generally receives priority
Current Cochrane guidance states that systematic-review searches should aim to be as extensive as possible so that as many relevant studies as possible are included. For Cochrane intervention reviews specifically, the recommended objective is to maximize sensitivity while striving for reasonable precision.
The reasoning is methodological. Screening an irrelevant record costs time. Missing a relevant study can alter the evidence base from which conclusions are drawn.
The two costs are therefore not necessarily equivalent.
This does not mean precision is irrelevant. A strategy producing hundreds of thousands of avoidable false positives can make screening impractical and may indicate conceptual or technical problems. The phrase reasonable precision matters.
Low precision can be an acceptable consequence of a sensitive search
Researchers sometimes become alarmed when only a small percentage of search results prove relevant. For comprehensive evidence synthesis, that alone does not demonstrate poor search quality.
Cochrane notes that the high yield and relatively low precision associated with systematic-review searching may be less daunting than it initially appears because titles and abstracts can often be screened comparatively quickly. Its current guidance gives a conservative estimate of approximately 60 to 120 abstracts per hour, although actual rates will vary by review and reviewer.
The implication is not that screening burden should be ignored. Rather, screening additional irrelevant records may sometimes be preferable to excluding evidence before reviewers have an opportunity to assess it.
Precision matters more as screening burden becomes consequential
Suppose two strategies retrieve all 50 benchmark studies. Search A retrieves 3,000 total records. Search B retrieves 30,000.
If the additional 27,000 records contribute no unique relevant evidence, Search B gains nothing from its lower precision. The extra retrieval simply consumes resources.
A sensitive search is not an excuse for uncontrolled noise. Researchers should still diagnose ambiguous terminology, overly broad truncation, unnecessary synonyms, poor Boolean structure, inappropriate fields, and other avoidable causes of irrelevant retrieval.
The practical task is to determine when the search is precise enough without narrowing it beyond what the evidence-identification objective permits.
Search purpose changes the balance
Not every literature search is a systematic review.
| Search purpose |
Typical priority |
Reason |
| Comprehensive systematic review |
Sensitivity usually receives greater emphasis |
Missing eligible evidence can undermine completeness |
| Rapid evidence synthesis |
Explicit compromise may be necessary |
Time and resource constraints can require narrower methods |
| Scoping or mapping work |
Depends on the objective |
Broad conceptual coverage may matter more than a narrowly focused result set |
| Background reading |
Precision may receive more emphasis |
The researcher may need representative useful literature rather than exhaustive identification |
| Finding a few papers on a familiar topic |
Precision often matters more |
Retrieving every relevant publication may provide little additional practical value |
Calling one search "better" without knowing its purpose can therefore be misleading. A strategy suitable for quickly locating five strong background papers may be unacceptable for a systematic review, while a systematic-review strategy may feel unnecessarily broad for ordinary literature exploration.
The consequences of missing studies also matter
Not every missed record has the same potential importance.
If omitted studies differ systematically from retrieved studies, sensitivity problems can affect the conclusions drawn from the evidence base. This is one reason comprehensive evidence synthesis treats missed studies seriously.
Conversely, the cost of irrelevant retrieval is primarily operational: more records need to be screened, deduplicated, managed, and assessed.
That asymmetry helps explain why comprehensive reviews often tolerate low precision in pursuit of sensitivity. It does not establish that every additional record is worth retrieving. The marginal benefit of broader searching still needs to be considered.
Filters make the trade-off particularly visible
Search filters can dramatically improve precision. A methodological filter might remove thousands of records with ineligible study designs. An age filter may remove populations outside the review. An outcome block can sharply focus retrieval.
But every filter should be evaluated for what it loses as well as what it saves. Cochrane recommends considering validated, highly sensitive filters for randomized trials in appropriate databases and cautions that filters must suit the source and purpose.
This is why testing whether a filter is too restrictive is essentially an exercise in examining sensitivity and precision together.
Peer review can improve the balance
A search can have poor sensitivity because concepts or synonyms are missing. It can have poor precision because terms are unnecessarily broad. It can suffer both because Boolean logic, field codes, or filters are incorrect.
PRESS provides a structured framework for peer review of electronic search strategies. Its retained domains include translation of the research question, Boolean and proximity operators, subject headings, text-word searching, spelling and syntax, and limits and filters. Evidence reviewed during development of PRESS suggested that structured peer review can identify search errors and improve term selection.
Peer review does not eliminate the sensitivity-precision trade-off, but it can remove avoidable problems before researchers begin making genuine compromises between the two.
Do not optimize either metric without examining the records involved
A numerical improvement can conceal an undesirable methodological change.
Suppose a revision increases precision from 5% to 15%. Excellent, perhaps. But if benchmark sensitivity falls from 98% to 80%, the gain may be unacceptable for a comprehensive review.
Now suppose another revision raises precision from 5% to 10% while benchmark sensitivity remains 98%. That is a much more attractive improvement because avoidable noise has been removed without detectable loss among the benchmark records.
Numbers become useful when interpreted alongside the actual records gained and lost.
Watch Out
Do not choose a search simply because it has the highest sensitivity or highest precision. A search strategy exists to serve a research purpose, and both metrics need to be interpreted against that purpose and the consequences of retrieval errors.