At a glance: Calibrating AI in health professions assessment means matching the role and autonomy of AI to the purpose of the assessment, the consequences of error, the quality of available evidence, and the stage of the assessment cycle. AI may support tasks across design, delivery, interpretation, feedback, and decision-making, but the level of human oversight should increase as uncertainty and consequences increase
“Are you using AI in assessment?” sounds like a yes-or-no question. It’s not. Assessment spans five stages and six kinds of evidence, and AI’s role shifts at every one. A Dean comfortable with AI drafting distractors for a low-stakes quiz shouldn’t necessarily be comfortable with AI drafting the narrative deciding whether a resident progresses.
Assessment Is a Cycle, Not an Exam
Assessment, at its core, is the process of finding, gathering, and interpreting evidence about what a learner knows and can do.Assessment can serve different purposes across a continuum, from supporting learning and development to informing high-stakes judgments about progression, certification, or readiness for practice. Every method, a quiz, an OSCE station, a licensing exam, sits somewhere along that spectrum, and separately, along a second spectrum entirely: how much is at stake.
That second axis, low-stakes versus high-stakes, describes consequences rather than purpose. A 2% credit quiz and a licensing board exam can use the same question format yet sit at opposite ends of the stakes spectrum. Both axes matter because together they determine how much AI autonomy is defensible for a given task.
Underneath both axes sits a five-stage cycle every assessment moves through: Design, Deliver, Interpret, Feedback, and Decide. AI’s role differs by stage. At the Design stage, AI may often be used with greater flexibility because outputs can be reviewed before they affect learners. However, design decisions still shape validity, fairness, and downstream consequences, so human oversight remains important. At Interpret, evidence increasingly includes predictive analytics that can flag a struggling learner weeks before a final exam would, ground Lecturio has covered in depth around shifting from activity metrics to competency-focused analytics. By the time an institution reaches Decide, human accountability should be at its highest because progression, remediation, and certification decisions carry consequences that should not be delegated to an automated system.
This isn’t just an internal framework. The recent AMEE Guide on AI in health professions education assessment makes a similar case: institutions need a structured, theory-grounded path through AI’s role in assessment, not a single blanket policy, because navigating the uncertainty AI introduces requires deliberate design rather than ad hoc rules.
Six Ways Health Professions Education Gathers Evidence
Any assessment method falls into one of six categories of evidence. AI touches all six, but not equally, and not without new risks alongside the gains.
Selected Response
Selected-response assessment is one of the more established areas of AI-assisted assessment development, particularly for drafting stems, generating distractors, tagging content, and supporting localization. However, item quality still depends on expert review, blueprint alignment, and post-administration psychometric analysis. This is also the territory where discrimination indices and distractor quality matter most, ground Lecturio has already covered around guided, faculty-reviewed item generation, so we won’t re-litigate the psychometrics here.
Constructed Response
Written and open-ended responses reveal something selected-response items can’t: a learner’s own reasoning and written communication. AI’s role here is rubric development, preliminary scoring, and drafting individualized comments. Written and open-ended responses can reveal aspects of reasoning and communication that selected-response formats may not capture. AI may support rubric development, preliminary scoring, and draft feedback, but these uses introduce questions around construct representation, rubric quality, bias, and consistency. Faculty review remains especially important when feedback or scores carry meaningful consequences.
Dialogic Assessment
Oral and dialogic assessment is receiving renewed attention in the context of generative AI, partly because it allows educators to explore reasoning, justification, and adaptation in real time. AI may support prompt development, transcription, formative rehearsal, and post-hoc analysis, but automated scoring introduces risks related to accent, language, cultural communication norms, and construct validity.
Clinical Reasoning and Simulation
Digital clinical cases and branching simulations can generate rich process data about how learners gather information, prioritize hypotheses, and revise decisions. AI can support feedback on some of these reasoning processes, but it should not replace opportunities for learners to commit to and explain their own reasoning before assistance is introduced. This is also where over-reliance risk is highest: learners who lean on AI before building their own reasoning skills risk never building it at all. Lecturio covers this, and how Socratic-style AI tutoring guards against it, in a separate piece on scaling 24/7 student support.
Performance-Based Assessment
OSCEs are widely used to assess clinical, procedural, and communication performance under standardized conditions. AI’s emerging roles include station development, simulated-patient preparation, transcription, and rubric-aligned analysis of recordings, although automated performance scoring remains an early area of research. A 2026 preprint found that a properly configured multimodal AI system (using synchronized multi-camera video) reported higher reliability on selected physical-examination scoring tasks than the human raters in that study, though the authors note this required infrastructure not yet common at most institutions, and as a preprint, the findings haven’t been peer-reviewed. Treat this as an early but concrete signal, not yet a production-ready standard.
Workplace and Longitudinal Assessment
Postgraduate and residency training rely on narrative feedback and 360-degree observation, qualitative data that’s slow to synthesize and hard to standardize across supervisors, especially non-academic clinical educators. AI may help organize narrative comments, map text to competency frameworks, and identify patterns in feedback quality over time. These tools may also inform faculty-development efforts by highlighting where feedback is consistently vague or insufficiently specific. Recent work applying natural language processing to workplace-based assessment narratives found this useful for monitoring feedback quality trends over time, though its authors are clear it isn’t a substitute for expert review. The risks worth naming are bias accumulation across observations, and an early negative label becoming self-fulfilling if unchecked. Narrative data deserves particular care for this reason: AI synthesis can lend a veneer of objectivity to what is still a collection of subjective judgments, and that synthesis is only as trustworthy as the quality, context, and potential bias of the observations feeding it.
Every Opportunity Has a Corresponding Risk
Every advantage AI offers in assessment—speed, scalability, consistency, synthesis—can create a corresponding risk when introduced without appropriate oversight. In practice, it collapses into one column, because every advantage AI offers, scalability, speed, consistency, has a matching risk that shows up when governance is missing. This is the same principle behind Lecturio’s faculty-in-the-loop governance model: opportunity and risk are two sides of the same coin.
Automated sampling can also create structural errors even when individual items are technically sound. An item-selection algorithm may overrepresent a topic, competency, or content area unless blueprint constraints and review processes are built into the workflow. Nothing about the individual questions was wrong; the failure was an ungoverned randomization process letting one topic dominate a supposedly representative sample, a structural blind spot nobody was watching for.
The fix isn’t avoiding automation. It’s building oversight in from the start. The AMEE Guide frames this same challenge as an assessment uncertainty problem that institutions need to navigate deliberately, not one that resolves itself by hoping the tool behaves.
The De-Skilling Question: Why Human Judgment Still Leads at “Decide”
Design-stage AI use carries the lowest stakes in the cycle, because students aren’t yet exposed to anything. Decide carries the highest human responsibility, because a call about whether a learner remediates or progresses has consequences AI isn’t yet positioned to own.
The evidence for why that boundary matters is growing. A 2026 systematic review in Advances in Medical Education and Practice found AI-supported learning may be associated with efficiency and basic knowledge gains, but is less consistently supportive of higher-order clinical reasoning, and flagged deskilling and ‘upskilling inhibition’ as risks worth attention, while noting the existing evidence base on this question remains thin.
A 2026 Nature Medicine perspective goes further, naming this risk “never-skilling”: trainees who rely on AI during their earliest clinical years may never develop the foundational reasoning skills independent practice requires, distinct from deskilling in a clinician who already has that foundation. The authors propose a three-phase approach: establish baseline competency, build critical calibration, then integrate AI under supervision. These concerns strengthen the case for keeping consequential progression and remediation decisions under explicit human accountability.
Calibrating AI Autonomy: A Stakes-and-Uncertainty Matrix
The practical version of this comes down to two questions for any task: how much is at stake, and how much uncertainty is in the AI’s output?
| Lower uncertainty / well-validated use | Higher uncertainty / limited validation | |
| Lower stakes | AI may automate or draft with monitoring, e.g. item tagging, formative question generation, routine feedback support | AI may assist, with sampling and human review of outputs |
| Higher stakes | AI may standardize or support processes, but consequential judgments remain accountable to humans | AI should augment expert review only and should not determine progression, remediation, or certification independently |
Rolling this out takes four honest questions before deploying AI anywhere:
- Purpose: What learning need does this serve?
- Construct: What competency are we assessing?
- Evidence: Has this AI been validated for our learners’ context?
- Consequence: What happens if the AI is wrong?
- Governance: Who’s accountable?
From Fragmented AI Use to Stakes-Calibrated Governance
| Dimension | Poorly governed AI use | Stakes-calibrated governance |
| Scope of AI use | Adoption driven by tool availability rather than educational need | AI used selectively where there is a clear assessment purpose and evidence of value |
| Autonomy | Similar level of automation regardless of consequences | Human oversight increases as stakes, uncertainty, and consequences increase |
| Construct protection | Efficiency prioritized over what the assessment is intended to measure | AI use is checked for whether it alters or contaminates the construct being assessed |
| Oral and performance assessment | AI scoring adopted without adequate bias or validity evaluation | AI supports transcription or analysis while scoring processes are validated and reviewed |
| Workplace assessment | Narrative data synthesized without examining source quality or bias | AI-assisted synthesis retains traceability to observations and expert review |
| Accountability | Responsibility is unclear when AI contributes to decisions | Named individuals or committees remain accountable for consequential judgments |
| Transparency | Learners may not understand where AI is involved | AI use, review processes, and decision responsibility can be explained clearly to learners |
Responsible AI use in assessment does not require choosing between innovation and institutional control. It requires matching the role of AI to the purpose of the assessment, the construct being measured, the quality of available evidence, and the consequences of error. As stakes and uncertainty increase, so should human oversight and accountability. Schedule a Demo with the Lecturio team today.
Frequently Asked Questions
How does AI change the assessment cycle for health professions educators?
AI can support different stages of the assessment cycle, including drafting items and rubrics, organizing evidence, generating preliminary feedback, and synthesizing patterns across data. Its autonomy should narrow as consequences and uncertainty increase, particularly when assessment information contributes to remediation, progression, certification, or other high-stakes decisions.
Is it safe to let AI score oral exams or OSCEs?
AI may support transcription, prompt standardization, or rubric-aligned analysis in oral and performance assessments, but automated scoring remains an emerging area. Before high-stakes use, institutions need evidence of validity, reliability, fairness, and performance across relevant learner groups, alongside human review.
What’s the difference between formative and summative assessment when AI is involved?
Formative and summative describe the purpose of assessment rather than its format. Lower-stakes formative uses may permit greater AI involvement because errors are easier to detect and correct, but autonomy should still depend on the quality of the evidence, the construct being assessed, and the consequences of an incorrect output.
Can AI cause de-skilling or “never-skilling” in clinical reasoning?
Current research treats this as a credible, precautionary concern rather than a settled finding: while direct evidence from medical training is still limited, learners who rely on AI before building foundational reasoning skills may fail to develop them at all, distinct from deskilling in a clinician who already has that foundation.
Who should review AI-assisted workplace-based assessments before they reach a learner’s file?
A named faculty reviewer, assessment lead, or appropriately constituted competence committee should retain responsibility for the final interpretation. AI may help organize and synthesize narrative evidence, but consequential competency judgments should remain accountable to qualified human decision-makers.