Why don't we just use Google Scholar?
This is a question that often crops up, especially with experienced researchers who know what they are looking for. ARCOM CM Abstracts is not competing with Google Scholar. Each provides a different, equally important service and neither competes with the other. They are complementary.
What Google Scholar does well
Google Scholar is the largest scholarly index in existence (Gusenbauer, 2019). It is free, requires no institutional login, and searches full text rather than metadata alone. It reports citation counts, links alternative versions of the same work, and maintains author profiles. For construction management, its coverage of anything with a web presence is close to complete.
If you know that a paper exists, Google Scholar will almost certainly find it. This has been demonstrated directly. Gehanno, Rollin and Darmoni (2013) tested Google Scholar against the 738 original studies included in 29 systematic reviews published in the Cochrane Database of Systematic Reviews or JAMA in 2009, and retrieved all 738, concluding that had those reviews used Scholar alone, no reference would have been missed. A comparable test by Giustini and Kamel Boulos (2013) recovered around 95 per cent, though only through repeated iterative searching with each citation handled individually. On the measure of raw coverage, no curated database competes. This is not in dispute here.
Being indexed is not the same as being found
An item can be indexed and yet remain unreachable by any query that could plausibly be constructed.
A Google search query itself cannot be stated with precision. Google Scholar has no nested Boolean logic and no proximity operators, and misinterpretation of Boolean queries can cause relevant results to be omitted (Gusenbauer and Gauster, 2025). Field restrictions, that are so useful in crafting search enquiries, extend little beyond author and date. A query cannot exceed 256 characters (Gusenbauer and Haddaway, 2020).
Ranking, then, is the sole arbiter of what is seen. It involves weighing citation counts and recency, so a query returns what has been cited, not what has been written; and exclusively settling for high-ranking results degrades exploratory and systematic searching in a way that it does not degrade lookup (Gusenbauer, 2021). The method behind that weighting is undisclosed and changes without notice (Gusenbauer and Haddaway, 2021).
The result of a search cannot then be precisely captured. There is no version date or snapshot date for the index, which is recrawled continuously, and repeating the same complex query does not reliably return the same number of hits (Bramer, 2016). Most awkwardly for anyone attempting a systematic search, a query reporting several thousand hits will list no more than the first thousand, so the retrieved set cannot be listed even once. A set that cannot be listed cannot be compared with anything.
The deeper problem is the difficulty of claiming evidence of absence. Google publishes no list of indexed sources and no inclusion criteria beyond a general statement of intent. Coverage is inferred through testing and never declared, and the extent to which any given type of record is covered remains unknown (Gusenbauer, 2024). Silence in the results may mean nothing was written, or that the crawler was excluded, or that the ranking buried it far below where any user would look.
Google Scholar includes any document that resembles a research paper, and constructs its metadata from crawler and parser output rather than from the records publishers already supply (Jacsó, 2012). That includes predatory journal output, duplicated preprint versions, student assignments posted on departmental servers and, more recently, machine-generated text. There is no editorial gate because there is no editor; Gusenbauer and Haddaway (2020) count the absence of curation among their grounds for judging Scholar unsuitable as a principal search system.
In the interests of balance, it should also be pointed out that critiques of Google Scholar are contested. Nash (2026), for example, argues that the criticisms levelled at Google Scholar are unrelated to what PRISMA guidelines actually require, and that Scholar should be treated as a primary database like any other. That is not the position taken here, but the disagreement is worthy of consideration.
The business model
A service of this size, free at the point of use, invites the question of how it is paid for. Google Scholar carries no advertising and charges nothing, to users or to institutions, and there is no public API to license. Nothing visible from outside the company suggests that it generates revenue.
It has been maintained since 2004 by a small team, and it exists at Google's discretion. Its function within the company is strategic rather than commercial. It retains a valuable population of searchers within Google's habits and it feeds the wider index. Whatever strategic value it holds is a matter of internal judgement, and such judgements may be revised from time to time without explanation.
A major risk with this kind of business model is discontinuation, and there is precedent among other services once useful to the academic community. Google closed down Reader, Knol, Code and Answers. This is not a critique restricted to Google. Microsoft retired Academic in 2021, and that service included a lot more functionality, such as a documented API, persistent identifiers and was supported by institutional backing.
Of course, nothing indicates that Scholar is about to be withdrawn; that is not the issue. The point here is that, when it comes to the operation of Google Scholar, there is no service commitment, no governance structure, no accountable body and nobody to write to. A service that has not provided an undertaking cannot break one.
What AI changes
Google released Scholar Labs in late 2025. It applies a language model to the full text of papers and returns an assembled answer, setting aside citation counts and journal standing as ordering signals. Elicit, Consensus and similar tools work along the same lines. Two consequences follow for anyone trying to survey a field.
The first is silent omission. A ranked list displays its own incompleteness. For example, you can scroll to, say, result 47, notice that the quality is declining, and make your own decision to stop scrolling. A generated answer shows nothing of what the engine declined to include. Errors of exclusion become invisible. There is an easy test that anyone can try. Ask the same question in two different formulations and the returned set will differ, with nothing in either answer to indicate that the other answer exists.
The second is contamination. Machine-written text is now entering the literature that crawlers collect. Haider et al. (2024) identified 139 papers in Google Scholar suspected of undisclosed generative AI authorship, of which 19 appeared in indexed journals and 89 in non-indexed ones, and warned that this opens the way to strategic manipulation of the evidence base. An index admitting anything shaped like a paper will admit these, and the synthesis layer will then draw on them.
Artificial intelligence does not diminish the case for a bounded, described collection. On the contrary, the case is strengthened by it. A model can be told exactly what a catalogue of 30,000+ records contains, how those records were selected, and what falls outside the scope. Nobody can supply the equivalent statement for the set of documents Google happens to have crawled.
The direction of travel
The widely used tools for online discovery are gradually shifting from listing results to assembling an answer for the researcher. When an interface ceases to show its working, the composition of the underlying collection becomes the last remaining control on output quality. Whatever cannot be seen in the process must be trusted in the source. At present, in the case of Google Scholar, that composition is undisclosed.
The same questions, asked of this catalogue
Scope. The boundary for the CM Abstracts is published. Journals were selected through cross-citation analysis rather than reputation or convenience, and the list is explicit. Doctoral theses come from named repositories in named countries; items deposited are controlled by universities. Therefore, what lies outside the collection is evident, so the absence of specific pieces of research can be interpreted by the user.
Indexing. A controlled vocabulary of roughly 7,000 terms is applied to titles and abstracts. Topics record what is being studied, and Subjects record how it is being studied. Author-supplied keywords remain visible and searchable but play no part in classification, because uncontrolled terms generate unreliable matches.
Funding. This is through an annual service agreement with a learned society. The sum is small and the arrangement is contractual. ARCOM can end it, and would have to make such a decision in the open.
Ownership. ARCOM holds the data and the right to the domain, which is registered on ARCOM's behalf. It can move either at any point. The platform is EPrints, which is open source. Records export in standard interchange formats. Nothing here is locked in.
Function. This is a metadata catalogue, not a repository. It does not store full text and it does not competes with any publisher. Most visits end in a click through to a publisher or an institutional archive, if they do not lead to the export of a list of references. These routes have been intentionally designed into the system.
Machine readability. Structured metadata, consistent vocabulary, stable identifiers and a declared boundary make the collection legible to software as well as to people. Whichever tools researchers adopt, a described corpus produces better results than an undescribed one. Indeed, the website logs show that Google Scholar is constantly driving enquiries to these pages. What is described here, therefore, shapes what is found elsewhere, and that is a responsibility we accept and take seriously. A primary aim of this service is to help shape perceptions of the field.
In practice
Use Google Scholar; most of us do, daily. It remains the quickest route to a known item and the widest net for a first pass at an unfamiliar topic.
Use this catalogue when the question concerns the field rather than the paper. What has been written on governance in construction projects? How has the treatment of trust in contracts developed since 1995? Which doctoral theses have addressed labour productivity in West Africa? Questions of that kind require a defined collection with consistent description, and no general web index can supply one.
References
Bramer, W.M. (2016) Variation in number of hits for complex searches in Google Scholar. Journal of the Medical Library Association, 104(2), 143–145. https://doi.org/10.3163/1536-5050.104.2.009
Gehanno, J.-F., Rollin, L. and Darmoni, S. (2013) Is the coverage of Google Scholar enough to be used alone for systematic reviews. BMC Medical Informatics and Decision Making, 13(1), 7. https://doi.org/10.1186/1472-6947-13-7
Giustini, D. and Kamel Boulos, M.N. (2013) Google Scholar is not enough to be used alone for systematic reviews. Online Journal of Public Health Informatics, 5(2), 214. https://doi.org/10.5210/ojphi.v5i2.4623
Gusenbauer, M. (2019) Google Scholar to overshadow them all? Comparing the sizes of 12 academic search engines and bibliographic databases. Scientometrics, 118(1), 177–214. https://doi.org/10.1007/s11192-018-2958-5
Gusenbauer, M. (2021) The age of abundant scholarly information and its synthesis: a time when ‘just google it’ is no longer enough. Research Synthesis Methods, 12(6), 684–691. https://doi.org/10.1002/jrsm.1520
Gusenbauer, M. (2024) Searchsmart.org: guiding researchers to the best databases and search systems for systematic reviews and beyond. Research Synthesis Methods, 15(6), 1200–1213. https://doi.org/10.1002/jrsm.1746
Gusenbauer, M. and Gauster, S.P. (2025) How to search for literature in systematic reviews and meta-analyses: a comprehensive step-by-step guide. Technological Forecasting and Social Change, 212, 123833. https://doi.org/10.1016/j.techfore.2024.123833
Gusenbauer, M. and Haddaway, N.R. (2020) Which academic search systems are suitable for systematic reviews or meta-analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resources. Research Synthesis Methods, 11(2), 181–217. https://doi.org/10.1002/jrsm.1378
Gusenbauer, M. and Haddaway, N.R. (2021) What every researcher should know about searching: clarified concepts, search advice, and an agenda to improve finding in academia. Research Synthesis Methods, 12(2), 136–147. https://doi.org/10.1002/jrsm.1457
Haider, J., Söderström, K.R., Ekström, B. and Rödl, M. (2024) GPT-fabricated scientific papers on Google Scholar: key features, spread, and implications for preempting evidence manipulation. Harvard Kennedy School Misinformation Review, 5(5). https://doi.org/10.37016/mr-2020-156
Jacsó, P. (2012) Using Google Scholar for journal impact factors and the h-index in nationwide publishing assessments in academia: siren songs and air-raid sirens. Online Information Review, 36(3), 462–478. https://doi.org/10.1108/14684521211241503
Nash, C. (2026) Reassessing the role of Google Scholar in PRISMA-informed systematic reviews through a critical analysis of influential research identifying its limitations and empirical evidence. Publications, 14(2), 29. https://doi.org/10.3390/publications14020029
Last updated 18 August 2026