Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What Should You Do When Different Databases Produce Dramatically Different Numbers of Results?

Different databases rarely produce identical result counts, even for searches intended to represent the same concepts. Learn how to distinguish legitimate database differences from problems in search translation.

140
When Databases Produce Different Result Counts Guide 140 of 899
01 · The Question

Why Can the Same Literature Search Produce Very Different Result Counts?

You translate a search into several databases and expect broadly comparable results. One database retrieves 2,000 records, another 7,000, and a third only 900. The discrepancy can look suspicious enough to suggest that one of the searches must be wrong.

Sometimes it is. A mistranslated field code, missing subject heading, unsupported proximity operator, or misplaced Boolean expression can substantially alter retrieval.

But large differences can also be legitimate. Databases do not contain identical collections of records, use identical controlled vocabularies, index documents in the same way, or necessarily expose the same searchable fields. The useful question is therefore not whether the counts match. It is whether each search accurately represents the intended concepts within that particular database.

02 · The Short Answer

Different Counts Are Expected, but Large Differences Should Be Explained

In Brief

Different databases can legitimately produce dramatically different result counts because they differ in content coverage, indexing, controlled vocabulary, searchable fields, record structure, and search functionality; however, unusually large discrepancies should prompt a line-by-line check of how the strategy was translated.

Do not try to make the numbers match. Compare the logic and retrieval behavior of corresponding search concepts, investigate where the counts diverge, and determine whether the difference reflects the databases themselves or an error in the translated strategy.

03 · What You Need to Know

Equivalent Search Logic Does Not Mean Identical Retrieval

Databases contain different literature

The most fundamental reason for different result counts is coverage. Bibliographic databases do not index exactly the same journals, conference proceedings, document types, geographic literature, or publication periods.

Cochrane notes that MEDLINE and Embase overlap, but the degree of overlap varies substantially by topic. It also identifies subject-specific and regional databases as potentially important because they may contain literature not indexed elsewhere.

Consequently, even perfectly translated searches can retrieve different numbers of records because they are searching different collections.

Database difference The strategies are appropriately translated, but retrieval differs because the databases contain or index literature differently.
Translation problem The intended search logic has changed during adaptation because syntax, fields, vocabulary, or operators were translated incorrectly.

Controlled vocabularies are not interchangeable

Databases may use different indexing vocabularies. MEDLINE uses Medical Subject Headings (MeSH), while Embase uses Emtree. Cochrane explicitly notes that the controlled-vocabulary terms and indexing approaches for MEDLINE and Embase are not identical.

A MeSH descriptor therefore cannot simply be pasted into an Embase strategy and assumed to have the same meaning or retrieval behavior. The corresponding Emtree concept may occupy a different position in its hierarchy, have different narrower terms, or be indexed differently.

For each database, identify the appropriate controlled-vocabulary term independently and inspect its scope and hierarchy.

Explosion can produce different conceptual breadth

Many controlled vocabularies are hierarchical. Exploding a subject heading generally retrieves records indexed with that heading together with relevant narrower terms beneath it, according to the particular system.

Two apparently corresponding headings can therefore retrieve quite different sets if their hierarchical structures differ.

Cochrane recommends using appropriate controlled vocabulary, including exploded terms where relevant, but this requires database-specific judgment. Matching the visible heading names is not enough.

The same free-text expression may search different fields

A search for a word can mean different things depending on the platform. One interface may search title and abstract by default. Another may include author keywords, subject fields, or additional metadata. A third may apply automatic query processing.

This can produce major count differences even when the text typed into the search box looks identical.

Check the documentation for the actual fields being searched. For a controlled comparison, explicitly specify comparable fields where appropriate rather than relying on defaults whose behavior differs between platforms.

Proximity operators are particularly easy to mistranslate

Databases use different syntax for requiring words to appear near one another. The operators themselves differ, and so can their treatment of word order, maximum distance, punctuation, and field boundaries.

A proximity expression copied literally from another platform may fail, behave differently, or be interpreted as ordinary text.

Translate the intended relationship, such as "these terms should occur within three words of each other in either order," rather than translating only the visible characters of the original query.

Truncation and wildcard behavior can differ too

A truncation symbol may represent any number of characters in one system but behave differently in another. Wildcards can represent optional or single characters, and some platforms impose restrictions on where or how they may be used.

These details matter because a small syntax difference can produce a large retrieval difference. A term intended to capture a controlled family of word variants may become either much narrower or much broader after translation.

Phrase searching may not behave identically

Quotation marks and other phrase-searching mechanisms are not universally interpreted in the same way. Some systems process phrases differently, and some database interfaces apply automatic transformations to unquoted terms.

If a count changes dramatically around a multiword expression, test the expression independently and inspect how the database reports or interprets the query.

Filters and limits may not be equivalent

Applying an apparently similar filter in two databases does not guarantee equivalent selection. Publication types, study designs, age groups, document types, languages, and other metadata may be indexed differently.

Cochrane cautions that search filters should be used with attention to the database and context. It also notes that filters appropriate for one resource may be unnecessary or inappropriate in another. For example, randomized-trial and human filters used for bibliographic databases should not simply be applied to CENTRAL.

When counts diverge sharply, temporarily remove filters and limits. If the discrepancy contracts, investigate how those restrictions operate in each database.

Compare concept blocks rather than only final totals

The final result count tells you that two searches differ, but not why.

A much more useful diagnostic approach is to compare corresponding concept blocks separately:

Search component Database A Database B Question to investigate
Population concept 42,000 47,000 Are coverage and terminology reasonably comparable?
Intervention concept 18,000 71,000 Which subject heading or free-text term causes the divergence?
Study-design filter 3,500,000 2,100,000 Are equivalent validated filters being used?
Final combination 1,240 4,860 Does the intervention difference explain most of the final gap?

These counts are hypothetical, but the diagnostic principle is important. Find the first point at which the strategies diverge substantially, then investigate that component.

More results do not mean better coverage

A database retrieving three times as many records has not necessarily found three times as much relevant evidence. It may contain broader subject coverage, more duplicate-like records, different document types, or a noisier translation of one concept.

Likewise, the smaller result set is not automatically more precise or more accurate.

Judge the searches through relevant retrieval, known-paper testing, inspection of unique records, and the appropriateness of the translated logic rather than by treating result count as a quality score.

Watch Out

Do not "correct" a database search merely because its result count differs from another database. Forcing counts to converge can distort a correctly translated strategy. Explain the difference before deciding that anything needs fixing.

04 · A Practical Example

Trace a Large Count Difference to the Responsible Search Line

Hypothetical Example

One database retrieves four times as many records

A review team translates a health-related search from one bibliographic database into another. The first search returns approximately 2,400 records. The second returns almost 10,000.

Do not compare only the totals The researchers run the population and intervention blocks separately in both databases.
Locate the divergence The population blocks are reasonably similar, but the intervention block retrieves nearly five times as many records in the second database.
Inspect individual lines Most free-text terms behave plausibly. One controlled-vocabulary heading, however, retrieves a much broader set than the apparently corresponding heading in the first database.
Inspect the hierarchy The second database's heading has different narrower concepts beneath it, and the search is exploding the term. Several of those narrower concepts extend beyond the intended intervention.
Revise based on meaning The team adjusts the controlled-vocabulary component to represent the intended concept in that database and then retests known relevant records and the resulting literature.

The problem was not that 10,000 was inherently "too many." The count discrepancy exposed a meaningful difference in how one database represented the concept.

05 · What Researchers Often Get Wrong

Common Mistakes When Database Counts Differ

Misconception

"Equivalent searches should produce similar numbers."

Not necessarily. Databases differ in coverage, indexing, controlled vocabulary, metadata, and retrieval functionality. Equivalent conceptual strategies can therefore produce substantially different counts.

Misconception

"The database with more results is more comprehensive."

A larger result set can contain more unique relevant literature, but count alone cannot establish comprehensiveness. It can also reflect broader coverage of irrelevant subject areas or a less precise search translation.

Misconception

"I should copy the same search string into every database."

Search strategies need database-specific translation. Controlled vocabularies, fields, proximity syntax, truncation, wildcards, phrase searching, and filters may all differ.

Misconception

"If the counts differ, I should adjust the smaller search until they match."

Matching counts is not a methodological objective. Change the strategy only when investigation reveals that a concept has been represented inadequately or translated incorrectly.

Misconception

"The same subject-heading label must retrieve the same concept."

Controlled vocabularies have different structures and indexing practices. Similar labels can have different scopes, hierarchies, and retrieval behavior.

06 · What This Means for You

Troubleshoot the Translation, Not the Number

When database counts differ dramatically, work from the inside out. Compare concept blocks, then individual terms, subject headings, fields, proximity expressions, and filters. The first substantial divergence usually tells you where to investigate.

A simple diagnostic framework

If corresponding concept blocks behave similarly but final totals differ moderately
Database coverage and indexing differences may plausibly explain much of the variation.
If one concept block differs dramatically
Inspect its individual free-text terms, subject headings, fields, and Boolean construction.
If a controlled-vocabulary line causes the discrepancy
Compare scope, hierarchy, explosion, and indexing rather than assuming the headings are direct equivalents.
If a free-text line behaves unexpectedly
Check fields, phrase processing, truncation, wildcards, automatic query processing, and proximity syntax.
If the strategy works in PubMed but the translated version behaves strangely elsewhere
Investigate the platform-specific problem rather than treating PubMed syntax as universal, particularly when a PubMed search fails in another database.

If the differences remain difficult to explain after line-by-line testing, external review may be worthwhile. Search translation is one of the points at which a research librarian or information specialist can identify database-specific problems that are easy to overlook.

07 · A Quick Checklist

What to Compare When Database Result Counts Differ

Before deciding that one search is wrong, check:
Confirm that each database actually covers the relevant disciplines, journals, publication periods, and document types.
Compare corresponding concept blocks separately rather than comparing only the final result totals.
Map controlled-vocabulary terms independently in each database and inspect their scope and hierarchy.
Verify whether subject headings are exploded and what narrower concepts the explosion includes.
Check which fields each free-text expression actually searches.
Translate proximity, truncation, wildcard, and phrase syntax according to the destination platform.
Check whether filters and limits are genuinely comparable and appropriate in both resources.
Test known relevant papers and inspect unique relevant retrieval rather than judging quality from counts alone.
08 · Frequently Asked Questions

Questions About Different Database Result Counts

Should two databases ever produce exactly the same number of results?

They can, but there is generally no reason to expect it. Different coverage, indexing, metadata, vocabulary, and search functionality normally produce at least some variation.

How different is too different?

There is no universal percentage or ratio. Treat the magnitude as a diagnostic signal. A large discrepancy deserves investigation, but whether it is problematic depends on what causes it.

Does the database with the largest result set contain the most relevant studies?

Not necessarily. It may contain unique relevant literature, but a larger result set can also contain substantially more irrelevant material. Compare relevant and unique retrieval rather than raw totals.

Can I use exactly the same keywords in every database?

Many free-text concepts can be carried across databases, but field codes, phrase behavior, truncation, wildcards, and proximity syntax may require adaptation. Controlled-vocabulary terms generally require database-specific mapping.

Should I remove terms from the larger search to make the counts closer?

No. Remove or modify a term only when its retrieval is inappropriate for the concept. Similarity of counts is not evidence that two strategies are better aligned.

Can different database coverage explain a very large difference?

Yes. Subject focus, journal coverage, document types, regional coverage, and time span can all contribute. Nevertheless, very large discrepancies are still worth checking for translation errors before attributing everything to coverage.

What if both searches retrieve my known relevant papers?

That is useful evidence but not complete validation. Known papers can reveal obvious failures, yet both searches may still differ in their ability to retrieve relevant literature you do not already know about.

09 · The Bottom Line

Explain the Difference Rather Than Trying to Eliminate It

The Bottom Line

When databases produce dramatically different result counts, do not try to make the numbers match; compare corresponding search components and determine whether the difference comes from legitimate variation in coverage and indexing or from a problem in search translation.

A strong multi-database search preserves the same conceptual intent while adapting to each resource's vocabulary, fields, syntax, and functionality. Different counts are therefore expected. What requires attention is an unexplained difference that reveals the concepts are no longer being searched equivalently.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes