Menglin Li (Miles) is a third-year undergraduate student majoring in English at City University of Macau. His research interests lie in applied linguistics, with a particular focus on language assessment and human–AI collaboration in educational contexts.
Email: milesmenglinli18@gmail.com
Title: "Is Human Judgment Becoming an Epistemic Anchor? A Bibliometric Study of Trust, Validity, and Human-AI Co-Assessment in Language Testing."
Author: Menglin Li
Abstract: The integration of artificial intelligence (AI) into language testing has profoundly reshaped assessment research and practice over the past two decades. While early work focused primarily on automated scoring accuracy and efficiency, recent scholarship increasingly foregrounds concerns of validity, fairness, explainability, and stakeholder trust. This study examines whether human judgment is structurally re-emerging as an epistemic anchor within AI-mediated language assessment systems.
Drawing on 270 publications indexed in the Web of Science (2004–2026), this study employs bibliometric methods including co-word analysis, bibliographic coupling, co-occurrence network analysis, and thematic mapping using Bibliometrix (R) and VOSviewer. Findings reveal three major phases of intellectual development: (1) an optimization-oriented phase centered on automated essay scoring and natural language processing; (2) a validation-oriented phase characterized by debates on bias, reliability, and fairness; and (3) a generative AI phase marked by rapid growth in large language model research and trust-related discourse.
Network centrality measures indicate that although automated scoring remains a dominant structural node, constructs such as “validity,” “rater,” “fairness,” and “explainability” increasingly function as bridging elements across thematic clusters. This shift suggests a structural rebalancing rather than technological displacement. AI systems are increasingly framed as epistemically incomplete without human interpretive mediation.
The study proposes that human judgment operates as an epistemic anchor—stabilizing interpretive authority, legitimizing algorithmic outputs, and mediating trust within human–AI co-assessment ecosystems. These findings contribute to emerging theoretical discussions on collaborative epistemology in language testing and underscore the necessity of human-centered governance in AI-based assessment systems.