From Authoring Items to Interpreting Evidence: A Dean’s Framework for Calibrating AI Across the Assessment Cycle

From Authoring Items to Interpreting Evidence: A Dean’s Framework for Calibrating AI Across the Assessment Cycle

Last update: September 3, 2026

  ·  

Author: Goran Stevanovski, MD

 | 

 | 

Assessment isn't one exam format. It's five stages and six kinds of evidence, and AI touches all of them differently. Drawing from a recent Lecturio global webinar, this piece gives Deans a practical framework for deciding exactly how much AI autonomy belongs at each stage, from low-stakes drafting to high-stakes progression decisions.
Lecturio thumbnail featuring the headline “From Authoring Items to Interpreting Evidence:” and the subheading “A Dean’s Framework for Calibrating AI Across the Assessment Cycle.” A white card on the right presents a 2x2 “Calibrating AI Autonomy” matrix comparing validation and stakes, with recommendations ranging from “Automate, Monitor” to “Augment Only, Never Decide.”

TABLE OF CONTENTS

At a glance: Calibrating AI in health professions assessment means matching the role and autonomy of AI to the purpose of the assessment, the consequences of error, the quality of available evidence, and the stage of the assessment cycle. AI may support tasks across design, delivery, interpretation, feedback, and decision-making, but the level of human oversight should increase as uncertainty and consequences increase


“Are you using AI in assessment?” sounds like a yes-or-no question. It’s not. Assessment spans five stages and six kinds of evidence, and AI’s role shifts at every one. A Dean comfortable with AI drafting distractors for a low-stakes quiz shouldn’t necessarily be comfortable with AI drafting the narrative deciding whether a resident progresses.

Assessment Is a Cycle, Not an Exam

Assessment, at its core, is the process of finding, gathering, and interpreting evidence about what a learner knows and can do.Assessment can serve different purposes across a continuum, from supporting learning and development to informing high-stakes judgments about progression, certification, or readiness for practice. Every method, a quiz, an OSCE station, a licensing exam, sits somewhere along that spectrum, and separately, along a second spectrum entirely: how much is at stake.

That second axis, low-stakes versus high-stakes, describes consequences rather than purpose. A 2% credit quiz and a licensing board exam can use the same question format yet sit at opposite ends of the stakes spectrum. Both axes matter because together they determine how much AI autonomy is defensible for a given task.

Underneath both axes sits a five-stage cycle every assessment moves through: Design, Deliver, Interpret, Feedback, and Decide. AI’s role differs by stage. At the Design stage, AI may often be used with greater flexibility because outputs can be reviewed before they affect learners. However, design decisions still shape validity, fairness, and downstream consequences, so human oversight remains important. At Interpret, evidence increasingly includes predictive analytics that can flag a struggling learner weeks before a final exam would, ground Lecturio has covered in depth around shifting from activity metrics to competency-focused analytics. By the time an institution reaches Decide, human accountability should be at its highest because progression, remediation, and certification decisions carry consequences that should not be delegated to an automated system.

This isn’t just an internal framework. The recent AMEE Guide on AI in health professions education assessment makes a similar case: institutions need a structured, theory-grounded path through AI’s role in assessment, not a single blanket policy, because navigating the uncertainty AI introduces requires deliberate design rather than ad hoc rules. 

Six Ways Health Professions Education Gathers Evidence

Any assessment method falls into one of six categories of evidence. AI touches all six, but not equally, and not without new risks alongside the gains.

Selected Response

Selected-response assessment is one of the more established areas of AI-assisted assessment development, particularly for drafting stems, generating distractors, tagging content, and supporting localization. However, item quality still depends on expert review, blueprint alignment, and post-administration psychometric analysis. This is also the territory where discrimination indices and distractor quality matter most, ground Lecturio has already covered around guided, faculty-reviewed item generation, so we won’t re-litigate the psychometrics here.

Constructed Response

Written and open-ended responses reveal something selected-response items can’t: a learner’s own reasoning and written communication. AI’s role here is rubric development, preliminary scoring, and drafting individualized comments. Written and open-ended responses can reveal aspects of reasoning and communication that selected-response formats may not capture. AI may support rubric development, preliminary scoring, and draft feedback, but these uses introduce questions around construct representation, rubric quality, bias, and consistency. Faculty review remains especially important when feedback or scores carry meaningful consequences.

Dialogic Assessment

Oral and dialogic assessment is receiving renewed attention in the context of generative AI, partly because it allows educators to explore reasoning, justification, and adaptation in real time. AI may support prompt development, transcription, formative rehearsal, and post-hoc analysis, but automated scoring introduces risks related to accent, language, cultural communication norms, and construct validity.

Clinical Reasoning and Simulation

Digital clinical cases and branching simulations can generate rich process data about how learners gather information, prioritize hypotheses, and revise decisions. AI can support feedback on some of these reasoning processes, but it should not replace opportunities for learners to commit to and explain their own reasoning before assistance is introduced. This is also where over-reliance risk is highest: learners who lean on AI before building their own reasoning skills risk never building it at all. Lecturio covers this, and how Socratic-style AI tutoring guards against it, in a separate piece on scaling 24/7 student support.

Performance-Based Assessment

OSCEs are widely used to assess clinical, procedural, and communication performance under standardized conditions. AI’s emerging roles include station development, simulated-patient preparation, transcription, and rubric-aligned analysis of recordings, although automated performance scoring remains an early area of research. A 2026 preprint found that a properly configured multimodal AI system (using synchronized multi-camera video) reported higher reliability on selected physical-examination scoring tasks than the human raters in that study, though the authors note this required infrastructure not yet common at most institutions, and as a preprint, the findings haven’t been peer-reviewed. Treat this as an early but concrete signal, not yet a production-ready standard. 

Workplace and Longitudinal Assessment

Postgraduate and residency training rely on narrative feedback and 360-degree observation, qualitative data that’s slow to synthesize and hard to standardize across supervisors, especially non-academic clinical educators. AI may help organize narrative comments, map text to competency frameworks, and identify patterns in feedback quality over time. These tools may also inform faculty-development efforts by highlighting where feedback is consistently vague or insufficiently specific. Recent work applying natural language processing to workplace-based assessment narratives found this useful for monitoring feedback quality trends over time, though its authors are clear it isn’t a substitute for expert review. The risks worth naming are bias accumulation across observations, and an early negative label becoming self-fulfilling if unchecked. Narrative data deserves particular care for this reason: AI synthesis can lend a veneer of objectivity to what is still a collection of subjective judgments, and that synthesis is only as trustworthy as the quality, context, and potential bias of the observations feeding it. 

Every Opportunity Has a Corresponding Risk

Every advantage AI offers in assessment—speed, scalability, consistency, synthesis—can create a corresponding risk when introduced without appropriate oversight. In practice, it collapses into one column, because every advantage AI offers, scalability, speed, consistency, has a matching risk that shows up when governance is missing. This is the same principle behind Lecturio’s faculty-in-the-loop governance model: opportunity and risk are two sides of the same coin.

Automated sampling can also create structural errors even when individual items are technically sound. An item-selection algorithm may overrepresent a topic, competency, or content area unless blueprint constraints and review processes are built into the workflow. Nothing about the individual questions was wrong; the failure was an ungoverned randomization process letting one topic dominate a supposedly representative sample, a structural blind spot nobody was watching for.

The fix isn’t avoiding automation. It’s building oversight in from the start. The AMEE Guide frames this same challenge as an assessment uncertainty problem that institutions need to navigate deliberately, not one that resolves itself by hoping the tool behaves. 

The De-Skilling Question: Why Human Judgment Still Leads at “Decide”

Design-stage AI use carries the lowest stakes in the cycle, because students aren’t yet exposed to anything. Decide carries the highest human responsibility, because a call about whether a learner remediates or progresses has consequences AI isn’t yet positioned to own.

The evidence for why that boundary matters is growing. A 2026 systematic review in Advances in Medical Education and Practice found AI-supported learning may be associated with efficiency and basic knowledge gains, but is less consistently supportive of higher-order clinical reasoning, and flagged deskilling and ‘upskilling inhibition’ as risks worth attention, while noting the existing evidence base on this question remains thin.

A 2026 Nature Medicine perspective goes further, naming this risk “never-skilling”: trainees who rely on AI during their earliest clinical years may never develop the foundational reasoning skills independent practice requires, distinct from deskilling in a clinician who already has that foundation. The authors propose a three-phase approach: establish baseline competency, build critical calibration, then integrate AI under supervision. These concerns strengthen the case for keeping consequential progression and remediation decisions under explicit human accountability.

Calibrating AI Autonomy: A Stakes-and-Uncertainty Matrix

The practical version of this comes down to two questions for any task: how much is at stake, and how much uncertainty is in the AI’s output?

Lower uncertainty / well-validated useHigher uncertainty / limited validation
Lower stakesAI may automate or draft with monitoring, e.g. item tagging, formative question generation, routine feedback supportAI may assist, with sampling and human review of outputs
Higher stakesAI may standardize or support processes, but consequential judgments remain accountable to humansAI should augment expert review only and should not determine progression, remediation, or certification independently

Rolling this out takes four honest questions before deploying AI anywhere: 

  • Purpose: What learning need does this serve? 
  • Construct: What competency are we assessing? 
  • Evidence: Has this AI been validated for our learners’ context? 
  • Consequence: What happens if the AI is wrong? 
  • Governance: Who’s accountable?

From Fragmented AI Use to Stakes-Calibrated Governance

DimensionPoorly governed AI useStakes-calibrated governance
Scope of AI useAdoption driven by tool availability rather than educational needAI used selectively where there is a clear assessment purpose and evidence of value
AutonomySimilar level of automation regardless of consequencesHuman oversight increases as stakes, uncertainty, and consequences increase
Construct protectionEfficiency prioritized over what the assessment is intended to measureAI use is checked for whether it alters or contaminates the construct being assessed
Oral and performance assessmentAI scoring adopted without adequate bias or validity evaluationAI supports transcription or analysis while scoring processes are validated and reviewed
Workplace assessmentNarrative data synthesized without examining source quality or biasAI-assisted synthesis retains traceability to observations and expert review
AccountabilityResponsibility is unclear when AI contributes to decisionsNamed individuals or committees remain accountable for consequential judgments
TransparencyLearners may not understand where AI is involvedAI use, review processes, and decision responsibility can be explained clearly to learners

Responsible AI use in assessment does not require choosing between innovation and institutional control. It requires matching the role of AI to the purpose of the assessment, the construct being measured, the quality of available evidence, and the consequences of error. As stakes and uncertainty increase, so should human oversight and accountability. Schedule a Demo with the Lecturio team today.


Frequently Asked Questions

How does AI change the assessment cycle for health professions educators?

AI can support different stages of the assessment cycle, including drafting items and rubrics, organizing evidence, generating preliminary feedback, and synthesizing patterns across data. Its autonomy should narrow as consequences and uncertainty increase, particularly when assessment information contributes to remediation, progression, certification, or other high-stakes decisions.

Is it safe to let AI score oral exams or OSCEs?

AI may support transcription, prompt standardization, or rubric-aligned analysis in oral and performance assessments, but automated scoring remains an emerging area. Before high-stakes use, institutions need evidence of validity, reliability, fairness, and performance across relevant learner groups, alongside human review.

What’s the difference between formative and summative assessment when AI is involved?

Formative and summative describe the purpose of assessment rather than its format. Lower-stakes formative uses may permit greater AI involvement because errors are easier to detect and correct, but autonomy should still depend on the quality of the evidence, the construct being assessed, and the consequences of an incorrect output.

Can AI cause de-skilling or “never-skilling” in clinical reasoning?

Current research treats this as a credible, precautionary concern rather than a settled finding: while direct evidence from medical training is still limited, learners who rely on AI before building foundational reasoning skills may fail to develop them at all, distinct from deskilling in a clinician who already has that foundation. 

Who should review AI-assisted workplace-based assessments before they reach a learner’s file?

A named faculty reviewer, assessment lead, or appropriately constituted competence committee should retain responsibility for the final interpretation. AI may help organize and synthesize narrative evidence, but consequential competency judgments should remain accountable to qualified human decision-makers.

Share this page:

Speak to us

Learn how Lecturio can help you

Authors

References

  1. Chen, J.-W., Tu, H.-L., Chang, C.-H., Hsu, W.-C., Wang, P.-C., Liao, C.-H., & Chen, M. (2025). Automated evaluation of reflection and feedback quality in workplace-based assessments by using natural language processing: Cross-sectional competency-based medical education study. JMIR Medical Education, 11, e81718. https://doi.org/10.2196/81718
  2. Eachempati, P., Komattil, R., & Arakala, A. (2025). Should oral examination be reimagined in the era of AI? Advances in Physiology Education, 49(1), 208-209. https://doi.org/10.1152/advan.00191.2024
  3. Kang, S., Holcomb, M. J., Shakur, A. H., Hein, D., Ngo, H.-T., Schuler, H., Jarrett, P. C., Dalton, T. O., & Jamieson, A. R. (2026). Automated assessment of OSCE physical exams using multimodal AI [Preprint]. medRxiv. https://doi.org/10.64898/2026.01.09.26343786
  4. Ke, Y., Jin, L., Ong, J. C. L., Thirunavukarasu, A. J., Car, J., Cheung, C. Y., Tham, Y. C., Ting, D. S. W., Ong, M. E. H., Compton, S., Narayan, A., Keane, P. A., Wong, T. Y., Bates, D. W., Tan, P., & Liu, N. (2026). AI-induced never-skilling in medical education. Nature Medicine, 32, 1997-2006. https://doi.org/10.1038/s41591-026-04438-y
  5. Masters, K., MacNeil, H., Benjamin, J., Carver, T., & Nemethy, K. (2025). Artificial intelligence in health professions education assessment: AMEE Guide No. 178. Medical Teacher, 47(9). https://doi.org/10.1080/0142159X.2024.2445037
  6. Turney, J., Young, T. M., Chauhan, D. R., Beeharry, R., & Mahmud, M. (2026). AI use for medical students: Impact on clinical skill acquisition and retention: A systematic review. Advances in Medical Education and Practice. Advance online publication. https://doi.org/10.2147/AMEP.S583763

User Reviews