AI-based educational assessment used to inform consequential decisions about learner progression
Source article: Developing validity arguments for artificial intelligence-based assessment: Balancing affordances and threats
Abstract: Background Artificial intelligence (AI) is increasingly used to generate, score and interpret educational assessment, yet these applications are being adopted in a largely unregulated environment. This creates a paradox: Whereas AI systems used in clinical care are subject to formal scrutiny for safety, performance and monitoring, AI systems used to inform consequential decisions about learner progression and future clinical practice are not. Existing validity frameworks remain useful but may not fully account f…
Contested: both sides are scored from claims and sources, not community votes.
Hospital Universitari Doctor Peset, València 02 by 19Tarrestnom65. CC BY-SA 4.0 · https://creativecommons.org/licenses/by-sa/4.0
A 2026 conceptual review in Medical Education examined how artificial intelligence used to generate, score and interpret assessments affects validity. Using Kane's four inferences, the authors mapped threats such as prompt instability and domain shift and noted that AI assessment is advancing without formal scrutiny comparable to clinical AI.
The analysis matters because validity failures could lead to unfair progression decisions, subgroup inequities and erosion of academic integrity in medical training. The authors propose treating validity as continuous governance with auditability and multi-site evaluation, but the paper does not provide new empirical data on whether such governance improves outcomes.
- Review applied Kane's four inferences: scoring, generalisation, extrapolation and implications to AI-based assessment.
- Scoring threats identified include construct contamination, prompt instability and limited explainability.
- Generalisation threats include domain shift, rater culture differences and temporal drift across settings and time.
- Implications threats include subgroup inequities, automation bias, deskilling, weak accountability and erosion of academic integrity.
AI systems can generate, score and interpret educational assessments that inform learner progression, with design and governance determining whether cross-cutting mechanisms function as affordances.
AI-based assessment introduces distinct validity threats across scoring, generalisation, extrapolation and implications, including contamination, instability, inequities, automation bias and deskilling when used for consequential learner progression decisions.
The rundown
The authors reviewed validity for AI-based assessment using Kane's framework, noting AI is being adopted in a largely unregulated environment compared to clinical AI, creating risks for decisions about progression.
They argue defensible use requires explicit validity arguments, auditability, multi-site evaluation and continuous consequence monitoring as ongoing governance of a sociotechnical system rather than one-time validation.
Findings derive from a conceptual review synthesizing literature rather than new empirical multi-site validation of an AI assessment system.
Sources
- Peer-reviewedMedical Education2026-09-08
How should this claim be treated?
ace
The debate