Evaluation of the Feasibility and Effectiveness of Generative AI-Assisted Multidisciplinary Decision Support in Medical Intensive Care: A Pilot Randomized Controlled Trial
Evaluation of the Feasibility and Effectiveness of Generative AI-Assisted Multidisciplinary Decision Support in Medical Intensive Care: A Pilot Randomized Controlled Trial
This pilot study evaluated the feasibility and usefulness of generative artificial intelligence (AI) as a clinical decision-support tool for physicians working in a medical intensive care unit. Participating physicians were assigned by work period to either use a generative AI system in addition to usual clinical information resources or to use usual resources without generative AI. The assigned condition was then switched so that participants experienced both approaches. During the AI-assisted periods, physicians used de-identified clinical information and considered the AI-generated responses as reference information. All final clinical decisions remained the responsibility of the treating physicians. The study assessed acceptability, usability, satisfaction, perceived decision support, workload, confidence, and learning experience through repeated questionnaires.
This was a single-center, open-label, pilot cluster-randomized crossover study involving physicians working in a medical intensive care unit. Each participating physician was observed during a scheduled one-month rotation in the medical intensive care unit. At the beginning of each monthly rotation, participating physicians were divided into two clusters. The clusters were randomized to begin with either the ChatGPT-assisted condition or the control condition. After approximately two weeks, each cluster crossed over to the alternate condition for the remainder of the one-month rotation. This design allowed participating physicians to experience both study conditions within the same rotation.
During the AI-assisted condition, physicians were encouraged to use ChatGPT (OpenAI) as a reference tool to support clinical information review and decision-making. Only non-identifiable clinical information was permitted to be entered into ChatGPT. Patient names, medical record numbers, contact information, and other information that could directly identify an individual patient were not entered. Physicians summarized clinically relevant information in their own words and considered the responses generated by ChatGPT when planning patient management. The Situation-Background-Assessment-Recommendation framework was recommended as an optional structure for organizing clinical information, but its use was not mandatory. Physicians were otherwise free to formulate their queries and interact with ChatGPT according to their clinical needs. A suggested prompt encouraged ChatGPT to present multiple management options, together with their rationale, potential benefits and risks, relevant supporting evidence, and areas of uncertainty. ChatGPT did not make or implement clinical decisions.
During the control condition, physicians used usual information resources, including discussions with other clinicians, multidisciplinary rounds, consultations, textbooks, clinical practice guidelines, PubMed, and other established clinical reference services, without using ChatGPT or other generative AI tools for study-related clinical decision support.
All diagnostic and treatment decisions were made independently by the treating physicians. Repeated questionnaires assessed satisfaction, decision-making experience, confidence, perceived efficiency, workload, educational value, and other aspects of clinical decision support. At study completion, participants also evaluated usability, satisfaction, perceived learning, reliance on ChatGPT, intention for future use, and the extent to which ChatGPT-generated suggestions were reflected in their clinical plans.
Inclusion Criteria:
Exclusion Criteria: