Skip to main content
Have a personal or library account? Click to login
‘GenAI’ Literature Search Tools and Scholarly Diversity: An Algorithmic Ethnographical Analysis Cover

‘GenAI’ Literature Search Tools and Scholarly Diversity: An Algorithmic Ethnographical Analysis

Open Access
|Jun 2026

Full Article

Introduction

This study arose following a conversation between two of the authors of this paper during doctoral supervision in June 2024. The doctoral researcher is undertaking a part-time structured PhD in E-Research and Technology Enhanced Learning. Having already successfully completed a structured set of modules over two years with the support of tutors and peers in their international cohort, they are now engaged in a period of intensive individual research with the support of their supervisor. Supervision meetings take place monthly via videoconferencing and every aspect of the doctoral research is discussed in depth.

The supervisor questioned the doctoral researcher’s decision to use only ‘Generative AI’ (‘GenAI’) tools to search for literature on their chosen topic and suggested it would be appropriate to use some well-known, ‘traditional’ literature search databases instead or at least in addition. The doctoral researcher mentioned that an effort was made to search for papers within the ‘traditional’ databases, but the results were less relevant to the research topic and the searches more time-consuming compared to ‘GenAI’ tools, which make the process more streamlined. In the absence of university guidance to researchers on the use of GenAI tools for research activities at this time, the focus of the conversation then turned to the ethical implications of researchers using such tools and whether the ‘GenAI’ literature search tools might be more inherently biased towards research conducted by male researchers in the USA and Europe than the ‘traditional’ literature search databases. We decided to test this notion with the help of academic colleagues and devised the study reported in this paper.

While ethical concerns about Generative AI (GenAI) in academia, such as training data limitations, language inequities and systemic biases have been well investigated (see for example, Anis & French, 2023; Dhingra et al., 2023; Yan, Sha et al., 2024), empirical research that examines ‘GenAI’ literature search tools remains critically scarce. Current research focuses on efficiency gains (Kumar & Gunn, 2024; Whitfield & Hofmann, 2023) or output quality (Bjelobaba et al., 2024). However, it does not attempt to quantitatively analyse how these tools contribute to scholarly visibility across gender and geographic lines. This area is particularly urgent given evidence that doctoral students using ‘GenAI’ search tools report zero ethical considerations about representation (Kumar and Gunn, 2024). Hence, without investigation, we risk automating the marginalisation of Global Majority, female, and early-career scholars under the guise of algorithmic neutrality, a phenomenon Kay et al. (2025) term as algorithmic epistemic injustice.

We should highlight from the outset that so-called ‘traditional’ academic databases are far from neutral. The choice of which database to use for a systematic literature review can have a marked effect on the results, with one study finding less than 6% commonality in articles retrieved from the same search performed across three ‘traditional’ databases (Wanyama et al., 2022). ‘Traditional’ databases already encode epistemic injustices through opaque “relevance” rankings (Jordan & Tsai, 2024). These rankings influence academics’ perceptions. An eye tracking experiment has demonstrated that higher ranking results are more often trusted, even when those ranked lower are more relevant to the search topic (Pan et al, 2007).

Researchers should consider power dynamics, cultural competence, and the ways in which they are culturally situated in terms of their own power and privilege when conducting research (Fassinger & Morrow, 2013). In this paper, we operationally define social justice as equitable representation of marginalised voices, measured through gender parity, geographic diversity, and reduced reliance on prestige metrics, such as journal impact factors (JIF) and citations that favour dominant groups in academic research and scholarly work.

We question herewith whether ‘GenAI’ literature search tools replicate or redress social biases around gender and geography in the field of educational research, responding to the call for empirical studies that examine such biases in depth (Li et al., 2025). Our inquiry is grounded in epistemic injustice theory (Fricker, 2008; Dotson, 2014), which examines how power imbalances distort whose knowledge is deemed credible. When ‘GenAI’ literature tools algorithmically rank scholarly work, they risk amplifying three forms of injustice: testimonial injustice, hermeneutical injustice and algorithmic epistemic injustice, which are explained in detail later in the paper. Our work is guided by the question: How do ‘GenAI’ literature search tools represent the gender and geographic diversity of scholarly work in comparison to ‘traditional’ literature search databases? As universities worldwide race to develop institutional AI research policies, our work is significant for higher education. Higher education institutions might use the findings to influence policy and guide their decisions about subscriptions to ‘GenAI’ literature search tools.

The ‘GenAI’ literature tools examined in this paper are the following: Consensus, Elicit, Research Rabbit, Scopus AI and Semantic Scholar. Consensus AI is a search engine that uses AI to extract insights from scientific papers, providing evidence-based answers to research questions (Consensus, n.d.). Elicit uses AI algorithms to automate literature reviews, summarising key findings and identifying relevant studies based on user queries (Elicit Help Center, n.d.), whilst Research Rabbit discovers and organises academic papers through AI-driven recommendations (ResearchRabbit, 2025), creating visual networks of related work. Scopus AI, powered by Elsevier’s Scopus database, uses GenAI to summarise research trends, suggest relevant articles, and answer questions about scholarly content (Elsevier, 2025). Semantic Scholar claims to employ advanced Natural Language Processing to analyse and rank academic papers, offering smart citations and highlighting influential research (Semantic Scholar, n.d.). Together, these tools claim to streamline the research process by enhancing discovery, summarisation, and analysis, saving researchers time and improving the efficiency of literature reviews and knowledge synthesis. Each tool integrates GenAI uniquely, catering to different needs in academic and scientific exploration. These ‘GenAI’ literature search tools were used to generate data but were not used for the following narrative literature review, which used ‘traditional’ databases.

Literature

Whilst there has been a great deal of recent interest in the topic of GenAI for academic research purposes generally, the literature published to date in the specific area of ‘GenAI’ literature search tools is somewhat limited. The quality of outputs was a concern in a study to examine the role of GenAI tools as knowledge brokers in mathematics education (Rycroft-Smith & Macey, 2025). In examining whether four different GenAI tools provide accurate information when promoted to summarise research on the topic of Venn diagrams, two literature search-specific tools, Consensus and Elicit, produced accurate, existing references, whilst the other ‘GenAI’ tools tested, Claude and ChatGPT 4.0, hallucinated references (Rycroft-Smith & Macey, 2025).

Much of the other extant literature in this area relates to users’ perceptions. A perception often held about ‘GenAI’ literature search tools is that they increase efficiency (Kumar & Gunn, 2024; Whitfield & Hofmann, 2023) but evidence to support this suggestion is inconsistent. This claim of enhanced efficiency is supported by the results of a study of ChatGPT-assisted systematic reviews (Spillias et al., 2024) but another empirical study in which groups of researchers were asked to find medical literature in response to a specific question found that groups given access to iris.ai did not produce significantly more results in the same timeframe than the control group who used Google Scholar, Web of Science, and PubMed (Schoeb et al., 2020). In education, however, enhanced efficiency is not necessarily always a positive outcome. This is because it is only achieved with a comparable loss of individual agency in the process (White, 2025).

A study analysing the perceptions of 26 doctoral researchers in Educational Technology at one university in the USA following an assignment in which they were encouraged to explore a range of ‘GenAI’ literature search tools found that the participants perceived the tools’ benefits to be facilitating discovery and making connections between papers, increasing efficiency, and making the literature review process less intimidating (Kumar & Gunn, 2024). The participants reflected, however, that the tools should not replace ‘traditional’ literature search databases but be supplementary, as they had concerns about the quality of outputs, the tools’ inferior usability or features, and concerns around privacy. These concerns reflect those identified around the use of GenAI more generally within academia, along with whether all students have the digital literacy necessary to use such tools well and whether critical engagement with ideas will be diminished by their use (Bozkurt, 2025; Mozelius & Humble, 2024; Whitfield & Hofmann, 2023).

It is notable that the concerns identified by the participants in Kumar and Gunn’s study (2024) did not include any ethical issues or concerns around the social justice implications of using GenAI. Moreover, whilst the authors themselves identify a need to understand how such tools might be used ethically, they do not expand on what they mean by ‘ethically’. Similarly, whilst an examination of five ‘GenAI’ literature search tools, including Research Rabbit, Elicit and Consensus, identified concerns around the quality of outputs in relation to the given prompt (Bjelobaba et al., 2024), the same study does not comment on the biases associated with such tools. It does, however, identify such issues in relation to some of the other types of GenAI tools that it examines.

Social justice concerns regarding the use of GenAI for academic purposes include the associated excessive energy consumption (Selwyn, 2022), limitations around the training data used (Anis & French, 2023) and the predominance of English language-based tools, which make them far more useful to those in some societies over others and amplify existing global digital disparity (Yan, Greiff, et al., 2024; Yan, Sha, et al., 2024). However, there are also suggestions that GenAI can enhance equitability in academia through helping researchers overcome the difficulties associated with marginalised backgrounds, speaking English as a second language and a lack of access to resources that privileged academics take for granted (Anis & French, 2023; Bjelobaba et al., 2024).

Whilst some academics highlight the potential harm associated with the systemic biases around gender, race and social class that Large Language Models and GenAI tools are capable of perpetuating (Dhingra et al., 2023; Mei et al., 2023; Yan, Greiff, et al., 2024) and others call for transparent reporting of their use within research and academic writing (Bozkurt, 2024; Tlili et al., 2025), we have yet to consider these concerns in relation to ‘GenAI’ literature search tools.

Theoretical Framework

Ideas around epistemic injustice frame this study theoretically and particularly those relating to scholarly diversity around the axes of gender and geography. Feminist philosopher Miranda Fricker identifies two kinds of epistemic injustice (Fricker, 2008), both of which are apparent in the world of academic publishing. The first of these is testimonial injustice, when the author’s credibility is diminished by the readers’ prejudice. Readers might like to think that we judge each piece of academic work on its own merits but inevitably, we have preconceived ideas about every journal article as we open it, based on the author’s name, institutional affiliation, their country of origin, and the status of the journal in which the work is published, for example. Gender bias is also an issue in academia. Despite evidence to suggest that female authors in the field of economics write with greater clarity than their male counterparts, their articles spend up to six months longer under peer review (Hengel, 2022). Journal articles with a female lead author are less likely to be published in prestigious journals and, when they are, they are less often cited than those with male lead authors (Bendels et al., 2018).

The second type of epistemic injustice is hermeneutical injustice (Fricker, 2008), when an unfair distribution of resources means that authors are disadvantaged when trying to understand their own social experience as compared with that of a wider society. This is also apparent in the world of academic writing and publishing, where journals based in Global Majority countries face challenges compared to their richer and more prestigious counterparts in North America and Europe (Bol et al., 2023). Authors with the opportunity to study within elite institutions have privileged access to teaching and resources that enable them to succeed when submitting their work to the journals most likely to be described as ‘high quality’ (Wellmon & Piper, 2017). Their institutions can afford subscriptions to such journals. Their teachers have published in such journals themselves and can provide advice based on their own interactions with reviewers and editors. Those with access to such resources are more likely to have their work accepted by top journals and when they do, have institutions able to pay the necessary fees to make their work open access. The figures published by one medical journal, for example, show stark regional disparities, with acceptance rates ranging from 24.5% for papers whose authors are based in North America to 0% for papers whose authors are from Africa and the Middle East (Onken et al., 2021).

A further type of epistemic injustice is theorised by Kristie Dotson, which she terms ‘third order exclusion’ (Dotson, 2014). This is when there are epistemic barriers to making the author’s ideas intelligible to the audience. The author’s writing may be dismissed as nonsense, dangerous or ridiculed by the dominant knowers who in this case might be the community of academic journal editors and reviewers. Dominant epistemic landscapes are characterised by the affective practices of the community of knowers, and it is their ways of thinking that undermine the credibility of oppressed individuals or groups as producers of knowledge, leading the oppressed to doubt their ability as knowledge producers (Hernandez, 2023). This conceptualisation of epistemic injustice has long been reflected in the suggestion that dominant groups only listen to the oppressed if they couch their ideas in ways that make those in power comfortable (Hill Collins, 2000).

Researchers build on the ideas of Fricker and others to propose a further type of epistemic injustice, which is particularly relevant to our work in this paper: algorithmic epistemic injustice (Byrnes & Spear, 2023). These authors are concerned with the epistemic injustice associated with the rapidly increasing use of machine learning in a medical context, where individuals’ knowledge, and particularly that of patients, is devalued by the authority afforded to automated systems. Their argument can also be applied within education and especially in the context of GenAI use (Kay et al, 2025). Kay et al. (2025) identify multiple ways in which GenAI negatively impacts human capacity to understand and trust oppressed and marginalised groups. For example, GenAI, they suggest, can amplify the voices of those who already dominate, thereby increasing testimonial injustice. GenAI can also heighten hermeneutical injustice through widening the information divide between privileged and marginalised groups and distorting or erasing the understanding of the experiences of those on the margins.

‘Traditional’ academic databases, such as Web of Science, have long played a role in making explicit, distributing and evaluating academic information, determining which research is deemed valuable and whose work is funded and disseminated (Fischer et al., 2020). Whilst GenAI has only recently been recognised as an influential stakeholder in the world of publishing (Bozkurt, 2024), all academic search engines have opaque, poorly understood algorithms that rank some writers’ work as ‘relevant’ over the work of others (Jordan & Tsai, 2024). In this study, we are concerned to determine whether the use of ‘GenAI’ literature search tools amplifies or reduces this effect in the context of literature reviews for the purposes of educational research.

Methods

In approaching the research design for the study, we faced the common paradox which faces all researchers wishing to understand algorithmic systems. Systems based on opaque algorithms and artificial intelligence are a classic example of a ‘black box problem’, when systems are so opaque, that they can only be understood through their inputs and outputs, rather than knowing how they function directly (Castelvecchi, 2016; von Eschenbach, 2021). This makes the need for research to understand their implications more pressing but presents a significant challenge for those attempting to conduct such research effectively.

While generative AI is relatively new as a broad social phenomenon, the issue of how to address this challenge and ways to conceptualise and study algorithms has been an issue in internet research prior to this. One of the authors’ previous experiences of scholarship in Internet Studies and digital methods provided a link between these fields and the present study. The critical study of algorithms – algorithms being broadly defined, but particularly in response to search engines and social media feeds – positions algorithms as a form of culture, and as such calls for the use of ethnographic methods (Gillespie, 2016; Seaver, 2017). Although the application of ethnographic methods to understand what is obscured by algorithms has previously been more frequently applied to social media, the principles remain relevant and adopting a stance to “follow the medium” and “consider the Internet not so much as an object of study, rather as a source of new methods and languages for understanding contemporary society” (Caliandro, 2018, p.553) remains fitting. As generative AI becomes increasingly embedded in academic culture and practices, there is likely to be further reconceptualising of the definition of algorithms and thinking about ways to approach ethnographic research (e.g. Cellard, 2022).For the present study, we adopted a methodological approach which can be characterised as aligning with ‘algorithmic ethnography’ (Christin, 2020a) to address these challenges and begin to understand the ways in which algorithms are mediating academic and scholarly culture in this context. Algorithmic ethnography is defined as “the ethnographic study of the computational systems enabling and shaping online interactions”, and “entails paying close attention to the role of algorithms in structuring the back and front end of the digital platforms that increasingly mediate digital exchanges” (Christin, 2020a, p.109). The definition of ‘algorithm’ is notoriously slippery and multiple (Seaver, 2017); in the context of this project, mindful of this, we adopt the position similar to Christin (2020a) that algorithms are technical mediators between individuals and their interaction with content (in this case, scholarly knowledge) online.

In adopting an algorithmic ethnographic stance, it is important to be attentive to the data infrastructures, the role of algorithmic sorting, and metrics deployed by platforms (Christin, 2020a). Algorithmic ethnography is also fitting given the focus and research question guiding the study, which are aligned with a stance of understanding and interpreting the effects of the technology in practice, rather than from a technical standpoint. For this project, we applied algorithmic ethnography by running a number of searches for a defined sample of topics, across the same range of academic literature search platforms. We then sought to identify trends across the platforms, drawing upon descriptive statistics and our interpretation of the observations and data. As such, through these processes we sought to enact the key strategies set out by Christin (2020b). The strategy of ‘algorithmic comparison’ was applied through the research design by seeking to compare features across several different platforms, while ‘algorithmic triangulation’ acknowledges the role that the algorithms themselves play in the study as a source of data.

The first step involved planning and identifying the object of focus for our inquiry, in terms of both the platforms we would include in the sample, and the searches we would run across the platforms. We sought to balance the sample of platforms to include a range of ‘traditional’ academic literature searches, alongside a range of tools which also provide a way to search the academic literature but state they use generative artificial intelligence to assist in the process. Through collaborative discussion, ten platforms were identified, as shown in Table 1. Except for Scopus AI, to which one of the authors had access via their institution, all the ‘GenAI’ literature search tools are payment free versions.

Table 1

Platforms included in the sample for searches.

‘TRADITIONAL’ DATABASES‘GENAI’ TOOLS
Academic Search UltimateConsensus
ERICElicit
Google ScholarResearch Rabbit
ScopusScopus AI
Web of ScienceSemantic Scholar

As the authors are all researchers with a range of experiences in educational technology and related fields, we each nominated potential topics to search for. We selected three topics initially, focusing on ones which were relatively short and well-defined terms, to avoid ambiguity. The three terms included ‘breakout rooms’, ‘learning management systems in higher education’, and ‘microcredentials’. As we engaged with the analytical process, we opted to add a fourth search term – ‘citizen science’ – to help verify some of the emergent trends.

Data collection took place between October 2024 and April 2025. All data were collected before the decision that ERIC, a US Department of Education funded ‘traditional’ database, would no longer be updated (Alonso, 2025). For each of the search terms, a search was conducted on each platform, and the first 20 journal article results were logged in a spreadsheet. Most platforms presented results according to ‘relevance ranking’ or similar by default (Jordan & Tsai, 2024), apart from Scopus which displayed results according to date by default. The ordering of results in Scopus was toggled to ordering ‘by relevance’, for comparability. By focusing on the order in which articles are presented by different platforms, we put into practice a focus on algorithmic sorting (Christin, 2020a) and applied the strategy of ‘algorithmic comparison’, a “similarity-and-difference approach to identify the instruments’ unique features” (Christin, 2020b, p.897).

Once the articles were logged, the spreadsheet was also populated with metadata and information relating to each one. Given the focus of the study, and to allow potential sources of bias to be viewed from multiple perspectives, the following data was added: (i) publication date; (ii) number of female first authors; (iv) number of countries represented by first authors and the number of these which are ‘Global Majority’ countries; (v) journal impact factor according to Web of Science (also indicative of whether the journal is indexed by Web of Science); and (vi) number of citations, according to Google Scholar.

In examining these variables, we were aware of the potential for inaccuracies in our data and the potential for our choice of variables to unhelpfully reinforce existing discourses. In examining the number of female first authors we were dependent on the preferred pronouns indicated by the authors in their biographies in online institutional profiles or social media, so there may have been instances where individual authors were misgendered. Whilst we did not restrict analysis to a binary gender framework, we identified no lead authors with non-binary pronouns in our sample.

We categorised the countries represented according to how many are Global Majority countries, broadly defined as those outside of the USA, Europe, Australia and New Zealand. Whilst we acknowledge that there could be much debate over this imperfect, working definition, we also recognize that most publications come from authors affiliated to institutions within the USA, Europe, Australia and New Zealand and that failing to include this factor would have been an omission in research of this kind.

Once collected, the data were analysed using a descriptive statistics approach (Marshall & Jonker, 2010), with an emphasis on displaying the data and understanding its characteristics and distribution, rather than applying statistical tests.

Findings

In this section, the data will be presented according to each of the factors which were examined, in turn. The focus here is upon presenting the data and identifying trends, with a full discussion bringing together the trends more generally in relation to the research question in the subsequent section.

Gender

Figure 1 shows the number of female first authors within the first 20 search results, for each platform, colour-coded according to the search topic.

Figure 1

Number of female first authors in the first 20 papers returned from each platform included in the sample. Colour of dots denotes the search term, and green dots show the mean average for each platform. Note that in instances where fewer than five data points are shown, this is due to overlapped data points of the same value – please refer to the colour coding.

Figure 1 suggests that the proportion of female first authors shows variation by platform, but also varies substantially according to each topic, so any platform differences are less clear. Focusing on the mean average per platform (as shown by green dots), there is an overall tendency towards balanced gender representation. Despite this tendency toward balance, there are platforms which represent the extremes of the distribution, with Scopus AI being furthest below the average (higher proportion of males), and Research Rabbit and Semantic Scholar being above average. Among the conventional academic databases, Academic Search Ultimate is the furthest below average, whilst Scopus identifies more female first authors than any other platform studied. There is also considerable variation between the results for different research topics, with the topics of citizen science and breakout rooms generating far more results with female first authors across all platforms and the topics of microcredentials and learning management systems generating far more results with male first authors.

Location

We examined the geographical range of scholars represented by the sampled papers by looking at the countries given by first authors’ affiliations. Location is a known bias within academic publishing (Bol et al., 2023), and the visibility of results within academic database searches (Czerniewicz & Wiens, 2013; Czerniewicz et al., 2017). The number of countries represented by first-author affiliations in the first 20 results are shown in Figure 2.

Figure 2

Number of countries represented by first authors in the first 20 papers returned from each platform included in the sample. Colour of dots denotes the search term, and green dots show the mean average for each platform. Note that in instances where fewer than five data points are shown, this is due to overlapped data points of the same value – please refer to the colour coding.

As was the case for gender, Figure 2 shows that the number of countries represented varies by both platform and topic. When considering diversity in terms of the number of countries, it may be notable that the three lowest mean average (less diverse) are ‘traditional’ literature search databases (Academic Search Ultimate, Google Scholar and ERIC), while two of the three highest are ‘GenAI-based’ tools (ScopusAI, Research Rabbit). The number of Global Majority countries in the first 20 results are shown in Table 2. Notably, the total for ‘GenAI-based’ tools is approximately ten percent higher than for ‘traditional’ literature search databases.

Table 2

Number of papers in the first 20 results per query, per platform, with first authors based in Global Majority countries.

PLATFORMBREAKOUT ROOMSMICROCREDENTIALSLEARNING MANAGEMENT SYSTEMS IN HIGHER EDUCATIONCITIZEN SCIENCETOTAL
Academic Search Ultimate4112228
ERIC4417327
Google Scholar9315023
Scopus6113318
Web of Science5012119
Total115
Consensus9011121
Elicit6414024
Research Rabbit9213125
Scopus AI6513327
Semantic Scholar9316129
Total126

Date

While publication date is unlikely to be a direct source of bias, it is possible that it may indirectly reflect a role of biased information such as citation counts, which accrue over time, and are known to be a source of bias (Oza, 2023). The distribution of the 20 sampled papers per platform according to date are shown in Figure 3. The distributions vary slightly according to query, but all apart from ‘Citizen Science’, which typically returns older results, are similar. Consensus, Elicit, ScopusAI and Google Scholar are notably more likely to include older results. The role of citation counts will also be addressed specifically in a subsequent section.

Figure 3

Distribution of publication date (year) for the first 20 papers returned from each platform included in the sample, colour coded according to search topic.

Web of Science Indexing

For each of the papers included in the sample, we recorded whether it is in a journal which is indexed by Web of Science. This is a notable characteristic, given the focus of our inquiry, for two reasons. First, there are known biases towards whether journals are indexed by Web of Science according to the country or region they are based in (Mills et al., 2023). Second, only Web of Science-indexed journals are given Journal Impact Factor (JIF) metrics, both being products of the parent company Clarivate. While we consider journal impact factors separately, our examination of whether or not a paper was Web of Science indexed provided a way of including and considering papers which are not indexed and do not have a JIF.

The number of papers published in Web of Science-indexed journals in the first 20 results per query and platform is shown in Figure 4. Here, a clear divide appears between major, publisher-linked platforms and others (both ‘GenAI’ tools and ‘traditional’ literature search databases). Web of Science (by definition), Scopus, Scopus AI and Academic Seach Ultimate are all relatively high in terms of this measure. The other platforms are notably lower, which may suggest the inclusion of potentially more diverse sources.

Figure 4

Number of papers in the first 20 papers returned from each platform included in the sample which are published in journals indexed by Web of Science. Colour of dots denotes the search term, and green dots show the mean average for each platform. Note that in instances where fewer than five data points are shown, this is due to overlapped data points of the same value – please refer to the colour coding.

Citation Counts

The number of citations a paper has accrued is a highly influential metric within academic publishing and evaluation of academic performance (Kulczycki, 2023). However, the choice of whose work to cite – and whose work not to cite – is an act which often serves to reinforce power imbalances in academia (Ray et al., 2024). As a widely available, numeric form of data, citation counts may be relatively easy to incorporate in computational ways of ranking search results. For the sampled papers, Google Scholar was used as a source of citation count data. While citation counts may be provided by the platforms themselves, Google Scholar was used as the source of citation data as it includes a wider range of sources compared to Web of Science, for example, which gives citation counts but only for indexed sources (e.g. Gerasimov et al., 2024). The range of citation counts for the sampled papers per platform, according to each topic, are shown in Figures 5, 6, 7, 8.

Figure 5

Citation counts according to Google Scholar for the first 20 papers returned from each platform included in the sample on the topic of breakout rooms. Three outliers – articles with citations >300 – are not shown (one each from Academic Search Ultimate, Elicit, and Scopus AI).

Figure 6

Citation counts according to Google Scholar for the first 20 papers returned from each platform included in the sample on the topic of citizen science.

Figure 7

Citation counts according to Google Scholar for the first 20 papers returned from each platform included in the sample on the topic of learning management systems in higher education. Three outliers – articles with citations >600 – are not shown (one from Consensus, and two from Scopus AI).

Figure 8

Citation counts according to Google Scholar for the first 20 papers returned from each platform included in the sample on the topic of microcredentials.

Figures 5, 6, 7, 8 show that the range of citation counts varies considerably according to each topic. As such, it is useful to consider the relative position of platforms. Table 3 shows the platforms per topic, ranked by median, from the largest median citation value (rank 1) to the smallest (rank 10). Within Table 3, there is not a clear distinction between the ‘GenAI’ tools and ‘traditional’ literature search databases; the sample of both types shows some which appear to be more influenced by citations, and others which are not. Looking at the ‘total’ ranking column, a lower aggregate figure here indicates that a platform is more often highly ranked, reflecting a more highly cited sample of papers. Within ‘traditional’ literature search databases, Google Scholar and Scopus stand out as being more often highly ranked, while this is true of Elicit and Consensus within the ‘GenAI’ literature search tools.

Table 3

Ranking by median number of citations for the sampled papers per query.

PLATFORMBREAKOUT ROOMSMICROCREDENTIALSLEARNING MANAGEMENT SYSTEMS IN HIGHER EDUCATIONCITIZEN SCIENCETOTAL
Academic Search Ultimate58.5101033.5
ERIC868931
Google Scholar714113
Scopus133714
Web of Science647825
Consensus451313
Elicit22228
Research Rabbit9.58.554.527.5
Scopus AI379625
Semantic Scholar9.51064.530

Journal Impact Factor

We also examined journal impact factor (JIF), as, alongside citations, this is a key metric within academic publishing and would potentially readily lend itself to computational methods of defining relevance. JIF is linked to the earlier discussions about the binary variable of whether or not a journal is indexed by Web of Science, as not all journals are assigned a JIF by Clarivate.

Boxplots showing the range of JIFs – for the papers within the sample which have them – are shown in Figures 9, 10, 11, 12.

Figure 9

Journal impact factors, for papers published in journals which have them (note differing ‘n‘), within the first 20 papers returned from each platform included in the sample on the topic of breakout rooms.

Figure 10

Journal impact factors, for papers published in journals which have them (note differing ‘n‘), within the first 20 papers returned from each platform included in the sample on the topic of citizen science.

Figure 11

Journal impact factors, for papers published in journals which have them (note differing ‘n‘), within the first 20 papers returned from each platform included in the sample on the topic of learning management systems in higher education.

Figure 12

Journal impact factors, for papers published in journals which have them (note differing ‘n’), within the first 20 papers returned from each platform included in the sample on the topic of microcredentials.

Similar to the citation data, the range of JIF within the sample varies considerably according to the search topic. Likewise, we consider the relative positions of the median JIF per platform, per query, as shown in Table 4. With the exception of Scopus AI, which tends to use high-JIF sources, the ‘GenAI’ tools may tend to use lower-JIF sources. However, in the ‘traditional’ literature search databases group, it is also notable that the two least JIF-linked platforms are Scopus and Web of Science.

Table 4

Ranking by median JIF for the sampled papers per query.

PLATFORMBREAKOUT ROOMSMICROCREDENTIALSLEARNING MANAGEMENT SYSTEMS IN HIGHER EDUCATIONCITIZEN SCIENCETOTAL
Academic Search Ultimate224.5917.5
ERIC23.531018.5
Google Scholar23.58316.4
Scopus5.597728.5
Web of Science97.54.5829
Consensus5.5102623.5
Elicit469423
Research Rabbit9561.521.5
Scopus AI11158
Semantic Scholar97.5101.528

Discussion

We set out to explore whether ‘GenAI’ literature search tools have any influential social biases around gender and geography in academic research. Using algorithmic ethnography, we analysed 800 search results retrieved through five ‘GenAI’ literature search tools and five ‘traditional’ literature search databases and following we provide a critical discussion of the findings and offer some practical implications for researchers.

‘GenAI’ Literature Search Tools: Some Patterns Over Clear Bias

To return to our overall question of how ‘GenAI’ literature search tools represent the diversity of scholarly work around gender and geography in comparison to ‘traditional’ literature search databases, we found little evidence to support the initial hypothesis that ‘GenAI’ literature search tools might be inherently biased towards research conducted by male researchers in Global Minority countries. In terms of gender, there is no overall discernible trend. The overall tendency towards a balanced gender representation contrasts with the extent of gender biases in academic publishing overall, which tend to favour male authors and reflect a range of sources of bias in the sector (Larivière et al., 2013). However, given that our focus is on education, within the social sciences, female academics may be more likely to be represented than in some other disciplines (Goyanes et al., 2025), which may account for the overall balance here. There is, however, considerable variation between platforms. Among the ‘GenAI’ tools, Research Rabbit and Semantic Scholar identified the highest number of articles with female first authors, whilst Scopus AI identified the lowest number. Among the ‘traditional’ literature search databases, Academic Search Ultimate is the furthest below average, whilst Scopus identifies more female first authors than any other platform studied. There is also considerable variation between the results for different research topics.

In terms of geographic diversity, the ‘GenAI-based’ searches were more likely to represent a higher number of countries than the ‘traditional’ database searches, and the proportion of papers identified from authors within institutions in Global Majority countries was approximately ten percent higher for the ‘GenAI-based’ searches. We have not identified any factors that might account for this difference. Whilst the ‘GenAI’ platforms, particularly Consensus, Elicit, Scopus AI and Scopus AI were notably more likely to include older results, a finding which may have indicated an algorithmic bias toward sources with higher numbers of citations, there is no clear distinction between the different types of platforms in terms of citation counts. Overall, there is no evidence within our study that ‘GenAI’ literature search tools display inherent algorithmic epistemic injustice (Byrnes & Spear, 2023) when compared with the ‘traditional’ literature search databases. Rather, the picture is more nuanced, with individual platforms in both groups displaying different biases when responding to different queries. Whilst the increased diversity associated with the ‘GenAI’ platforms may seem positive, this difference may be linked to systemic publishing inequities or rooted in hidden patterns of large-scale data extraction and exploitation.

Examination of the number of papers published in Web of Science-indexed journals show a clear divide appears between major, publisher-linked platforms and others, in both ‘GenAI’ and ‘traditional’ groups. Whilst Web of Science (by definition), Scopus, Scopus AI and Academic Seach Ultimate are all relatively high in terms of this measure, other platforms are notably lower, which may suggest the inclusion of potentially more diverse sources. Except for Scopus AI, the ‘GenAI’ tools may tend to use lower-JIF sources. When compiling our data, we noted our own perceptions of the platforms we used for the study, including our tendency to describe the results identified by the ‘traditional’ database searches and Scopus AI as ‘higher quality papers’. These perceptions may well be influenced by our ideas around the value of the major, publisher-linked platforms and journal impact factors. As such, these perceptions also reflect our own contribution to testimonial injustice (Fricker, 2008), when the author’s credibility is diminished by the readers’ prejudice. Our perceptions of what constitutes a ‘high quality’ publication are also reflective of the hermeneutical injustice (Fricker, 2008) and ‘third order exclusion’ (Dotson, 2014) apparent within the institutions within which we study and work and within in the wider world of academic writing and publishing.

Additional Observations on Accessibility and Cost

The research process revealed significant social justice concerns regarding accessibility and cost barriers associated with ‘GenAI’ literature search tools. A primary issue involves geographical restrictions, as certain ‘GenAI’ tools remain unavailable in specific countries, exacerbating global research inequalities (Castillo-Segura et al., 2023). This challenge became particularly evident when one co-author based in China encountered accessibility limitations, requiring a virtual private network to access some tools due to national internet regulations.

Furthermore, many advanced ‘GenAI’ tools operate through costly subscription models, creating financial barriers for underfunded researchers and institutions. For instance, Scopus AI remains exclusively available through institutional subscriptions, with only one co-author having access, thereby restricting researchers from less-resourced institutions. Similarly, while Elicit and Consensus offer enhanced features, their monthly upgrade fees (currently $12 and $8.99, respectively) present substantial cost barriers. Although Semantic Scholar and Research Rabbit provide free access, preliminary observations suggest that these tools deliver comparatively lower functionality than non-free alternatives, such as Scopus AI. While existing literature has not thoroughly examined cost-related barriers for these specific tools, scholars universally acknowledge that financial obstacles to research technologies disproportionately affect researchers in low- and middle-income countries and those lacking institutional support, thereby widening global research disparities (Cox & Abbott, 2021; Gabriel, 2024; Khisro & Fenlon, 2025). These accessibility and affordability challenges ultimately limit equitable participation in GenAI-enhanced academic research, particularly for marginalised scholar communities.

Towards a Better Integration: Some Practical Implications

Based on the findings, researchers require an appreciation of the variability of platforms, both ‘GenAI’ literature search tools and ‘traditional’ literature search databases, and the extent of their individual potential to enhance or limit epistemic diversity. Just as previous research uncovers the potential for some literature to be prioritised over others if only a single ‘traditional’ database is used (Heck et al., 2024; Wanyama et al., 2022), the same may be true for ‘GenAI’ review tools. Researchers might use multiple search tools strategically to achieve a higher diversity in results and a higher coverage. Combining geographically inclusive ‘GenAI’ tools like Semantic Scholar with discipline-specific databases, such as ERIC, could mitigate single-tool biases. Combining ‘GenAI’ tools with ‘traditional’ literature search databases might also enhance the geographic diversity of results. It is important to be aware that the diversity of results obtained from individual search tools is likely to vary for different search topics. Researchers need skills to critically evaluate literature search results using a social justice lens to enhance the rigor and trustworthiness of their research.

Higher education institutions must go beyond simple GenAI adoption policies to take account of algorithmic considerations when adopting and providing access to literature search platforms, both ‘GenAI’ literature search tools and ‘traditional’ literature search databases. Based on diversity of results according to search topic, they might consider specific tools for certain disciplines. Additionally, given that literature searching requires not only tool literacy but also the skills to critically evaluate search results, institutions should devise professional development programmes for staff and students to learn how to use the platforms and to learn about their algorithmic advantages and limitations from social justice perspectives. This is an essential part of preparing students for a future that not only requires technical skill but the development of a critical perspective regarding the use of GenAI (Bozkurt et al., 2024; Gupta et al., 2024). As institutions rush to integrate GenAI tools for research, we warn here against conflating diversity metrics with true equity. Surfacing Global Majority scholarship matters little if those works remain algorithmically buried beneath high-JIF publication, so institutions should esteem diverse knowledge generation and epistemologies and encourage scholars to rethink their own contribution to epistemic justice in terms of whose work they read and cite. We affirm Bozkurt’s call for the cultivation of the capacity for epistemic resilience, so that humans learn how to ‘ask the questions that machines cannot’ (Bozkurt, 2025).

Conclusion

We conclude that whilst there is there is no overall discernible difference in how the ‘GenAI’ literature search tools represent the diversity of scholarly work in terms of gender compared with ‘traditional’ literature search databases, there is a notable, unexpected positive difference in terms of geographical diversity, where the proportion of papers identified from authors within institutions in Global Majority countries was approximately ten percent higher for the ‘GenAI’-based tools.

We acknowledge some limitations that need to be considered when interpreting the results of this study. The sample was limited to 10 platforms and four search terms within a single discipline. Except for Scopus AI, the ‘GenAI’ literature search tools used were free versions. The use of premium versions of these tools might generate different results. In addition, the binary categorisation of so called ‘GenAI’ tools versus ‘traditional’ literature search databases may oversimplify the continuum of algorithmic sophistication, as even conventional databases employ AI-like ranking systems, e.g. Google Scholar ranking. Also, the snapshot in time and the ways in which tools are in continuous development or decline might not indicate how algorithmic updates change the results over time.

It is not just authors’ geography and gender that determine whose work is most valued in academia. Future research could consider additional factors that might be associated with epistemic injustice in searching for literature, including university ranking. The location of the publisher and those with editorial oversight of a journal article also determine a paper’s influence (Czerniewicz et al., 2017). Further research should also include the open access status of papers in the dataset, not from a perspective of bias as such, but to examine whether any evidence could be found to support the possibility that open access publications are more open to exploitation by GenAI (Decker, 2025; Pollock & Michael, 2024). Furthermore, a valuable area for future research would be to consider whether there is evidence emerging about if and how ‘GenAI’ tools are in turn shaping scholarly norms.

Data Accessibility Statement

The datasets generated and analysed during the current study are available in the Lancaster University repository, https://doi.org/10.17635/lancaster/researchdata/732.

Sustainable Development Goals (SDGs)

This study is linked to the following SDG(s): Quality education (SDG 4), Gender equality (SDG 5), Reduced inequalities (SDG 10).

Author Contributions (CRediT)

Kathy Chandler: Conceptualization, Data curation, Investigation, Methodology, Project administration, Supervision, Writing – original draft, Writing – review & editing. Katy Jordan: Conceptualization, Data curation, Formal Analysis, Investigation, Methodology, Project administration, Resources, Software, Supervision, Validation, Visualization, Writing – original draft. Ishaq Al-Naabi: Conceptualization, Data curation, Investigation, Methodology, Project administration, Writing – original draft. Panagiota Tzanni: Conceptualization, Data curation, Investigation, Methodology, Resources, Writing – original draft. Leone Gately: Conceptualization, Data curation, Writing – review & editing. All authors have read and agreed to the published version of the manuscript.

Language: English
Page range: 291 - 309
Submitted on: Jul 31, 2025
Accepted on: Mar 5, 2026
Published on: Jun 2, 2026
Published by: International Council for Open and Distance Education (ICDE)
In partnership with: Paradigm Publishing Services

© 2026 Kathy M. Chandler, Katy Jordan, Ishaq Al-Naabi, Panagiota Tzanni, Leone Gately, published by International Council for Open and Distance Education (ICDE)
This work is licensed under the Creative Commons Attribution 4.0 License.