Teaching Internal Medicine and Family Medicine Residents to Reason With Generative AI: A Multi-Site Randomized Controlled Trial (TEACH-AI)
Teaching Internal Medicine and Family Medicine Residents to Reason With Generative AI: A Multi-Site Randomized Controlled Trial (TEACH-AI)
The purpose of the TEACH-AI study is to assess whether a brief, structured workshop on artificial intelligence can improve the performance of medicine doctors in training (i.e. residents) in their diagnostic and management reasoning.
In this multi-site randomized controlled trial, internal medicine and family medicine residents are assigned either to receive an in-person workshop on safe, effective LLM use before a standardized AI-assisted assessment, or to complete the same assessment before receiving the workshop. Residents will review clinical cases that are fully synthetic, no protected health information is used, using a password protected LLM interface.
Large language models (LLMs) have rapidly entered routine use in medical education and clinical practice. Prior randomized trials have shown that, although LLMs can outperform individual clinicians on some reasoning benchmarks, providing physicians with access to an LLM does not necessarily improve their diagnostic or management performance, and LLMs alone may perform better than physician-plus-LLM teams. These findings suggest that human-AI collaboration is fraught, non-trivial and may require explicit training.
TEACH-AI is a pragmatic, multi-site, two-arm randomized controlled trial embedded within protected residency didactic time. The primary objective is to determine whether a single in-person workshop on the basics of LLMs and best practices of prompting/verification strategies improves residents' performance on an AI-assisted assessment, compared with residents who complete the simulation before receiving the workshop. All participants will ultimately receive the same workshop and the same assessment (either workshop-first vs assessment-first).
The simulation consists of multiple fully synthetic vignettes delivered via a password protected LLM interface. Residents interact freely with the LLM using natural-language prompts and then submit structured final responses regarding aspects such as leading diagnosis, differential, management plan, and/or justification. The platform will record prompts, model outputs, final answers, and timing. Vignette scoring combines correctness of the final diagnosis or management plan. Scoring is conducted by blinded faculty using standardized rubrics and then any discrepancies will be resolved through multiple rounds of discussions.
The trial will enroll up to 200 residents across four ACGME-accredited programs (internal medicine at Beth Israel Deaconess Medical Center, Stanford University, and Cambridge Health Alliance; family medicine at AdventHealth Orlando).
Inclusion Criteria:
Exclusion Criteria:
jkoshy@bidmc.harvard.edu617-754-4677
Palo Alto, California 94305, United States
kkeet@stanford.edu650-498-4559
raj.mehta.md@adventhealth.com407-845-8383
jkoshy@bidmc.harvard.edu617-754-4677
617-665-1000