Chatbot-mediated assessment in education: A systematic review of validity, feedback quality, and learner engagement
Enoch Oluwatomiwo Oladunmoye, Olaoye Adekunle Olufemi, Alakeme Nestor Johnson
Published May 28, 2026
Pages 454-467
Chatbot in educational assessment has grown at a rapid rate since 2019, with the number of Scopus-indexed publications growing by 381, 53 articles in 2019 to 255 in 2023. Although this has increased, such critical questions as the psychometric validity, quality of automated feedback and interest of the learner are not well synthesized. These are more acute in the African contexts, in which access and effectiveness are influenced by the infrastructural constraints and digital inequalities. This is a systematic review of chatbot-based assessment (CMA) evidence, emphasising three areas: validity and reliability, the quality of feedback, and learner engagement. It also critically compares the results in African and global contexts. Based on PRISMA 2020, a search in Scopus, Web of Science, ERIC, and PubMed resulted in 1,847 records. Following screening and inclusion criteria, 63 empirical studies published within 2018 were retained after screening and application of inclusion criteria. The Mixed Methods Appraisal Tool (MMAT) was used to evaluate the quality of the study, and the results were analyzed via thematic synthesis and through narrative methods. Findings show that out of 60 studies (n = 60), 71.4% (n = 45) studies found statistically significant improvements in learning outcomes linked to chatbot-mediated formative assessment. The quality of feedback was found to be multidimensional and included immediacy, specificity and corrective utility with effect sizes of d = 0.42 to 0.78. Nevertheless, psychometric rigor is low: Just 38% of the studies reported construct validity evidence and 22% included reliability coefficients. The representation of Africans was considerably low (7.9%, n = 5), and mostly restricted to Nigeria, Ghana, and South Africa, which points to gaps in geographical researches. In general, CMA shows good prospects to increase formative feedback and learner engagement on a large scale. Nevertheless, there are still ongoing issues of validity and fair access, especially in sub-Saharan Africa, where AI preparedness is 44.6/100 on average. The future research needs to focus on validity-driven design, African studies that are context-sensitive, and longitudinal studies on engagement outcomes.
Assessment
Chatbot
AI
Psychometric
Education
Africa
Enoch Oluwatomiwo Oladunmoye, Olaoye Adekunle Olufemi, Alakeme Nestor Johnson.
"Chatbot-mediated assessment in education: A systematic review of validity, feedback quality, and learner engagement."
African Multidisciplinary Journals of Development
, vol. 14
, no. 2
, 2026
, pp. 454-467