AI & Prehospital ECG Analysis: A Randomized Controlled Trial of a Smartphone Large Language Model for Occlusion Myocardial Infarction Detection by Prehospital Providers
AI & Prehospital ECG Analysis: A Randomized Controlled Trial of a Smartphone Large Language Model for Occlusion Myocardial Infarction Detection by Prehospital Providers
Prehospital providers interpret 12-lead electrocardiograms (ECGs) under time pressure and without immediate expert support. Missed acute coronary occlusion - occlusion myocardial infarction (OMI) - delays reperfusion, while false positive interpretations trigger unnecessary catheterization laboratory activations. Multimodal large language models (LLMs) available on any smartphone can now analyze a photographed ECG, and prehospital providers have begun using them spontaneously. No randomized trial has evaluated whether this practice improves diagnostic performance.
This randomized controlled trial compares the diagnostic performance of prehospital providers interpreting ECG clinical vignettes with and without mandatory assistance from a single, version-locked smartphone large language model. Participants - paramedics, emergency medical technicians, nurses and physicians practicing in prehospital care in French-speaking Switzerland - are randomized 1:1 on a dedicated digital platform and answer 14 clinical vignettes presented in individually randomized order. Each vignette is built around a real, anonymized 12-lead ECG obtained during routine clinical care.
The primary outcome is the proportion of vignettes for which the participant correctly identifies the presence or absence of an OMI. Secondary outcomes are sensitivity, specificity, and the accuracy of the prehospital priority decision level.
Design and setting. Two-arm parallel-group randomized controlled trial conducted on a dedicated digital platform. Phase 1 takes place during a single in-person session at the Swiss French-speaking prehospital clinical research conference (Morat, Switzerland, 2 September 2026). Phase 2 extends recruitment to prehospital emergency services across French-speaking Switzerland (September to November 2026) using an identical standardized protocol.
Randomization. Individual 1:1 allocation performed automatically by the platform using permuted blocks of variable size (4 to 6), without stratification, after the demographic questionnaire and before the first vignette. Allocation cannot be changed once assigned. The presentation order of the 14 vignettes is independently randomized for each participant, which neutralizes position effects and prevents copying between neighbouring participants.
Intervention. Participants allocated to the intervention arm must consult the study-imposed large language model for every vignette before submitting their answer; the platform locks the submit button until use of the tool is confirmed. A single model (OpenAI GPT-4o) in an API version locked for the entire study is used by all participants, with a standardized, non-modifiable prompt identical for every participant and every vignette. Participants cannot add text, ask follow-up questions, or provide additional clinical context. Exposure to the tool is mandatory, but adherence to its interpretation is not: participants remain free to base their final answer on their own clinical reasoning. Control participants interpret the same ECGs unaided, with smartphones turned face down and out of reach.
Reference standard. For each vignette, the expected answers were defined a priori by the study cardiologist and locked, with any subsequent modification time-stamped in the platform audit log. The reference standard is the answer expected of a prehospital provider at the point of care, anchored on coronary angiography wherever angiography is discriminant. For non-ischaemic mimics, the expected answer depends on whether the acute presentation allows the condition to be distinguished from a coronary occlusion: it does not for the Takotsubo case included in the study, whose expected answer is therefore "yes", whereas acute pericarditis, being usually recognizable, has an expected answer of "no". This pre-specified departure from a purely angiographic standard is reported as such, and the primary analysis is repeated in a sensitivity analysis in which all mimics are classified as non-OMI.
Blinding. Participants and investigators cannot be masked. The statistician conducting the primary analysis receives coded group labels only, and the allocation key is released after database lock and approval of the statistical analysis plan.
Statistical analysis. Mixed-effects logistic regression with correct OMI identification at vignette level as the dependent variable, arm as the main fixed effect, grouped ECG category as a fixed covariate, and random intercepts for participant and for vignette. Intention-to-treat, alpha 0.05 two-sided. The target is 130 evaluable participants (65 per arm), giving 80% power to detect an absolute increase from 0.65 to 0.80 with an intraclass correlation up to approximately 0.43. Allowing for approximately 10% of sessions to be abandoned before completion, planned enrollment is 144 (72 per arm).
ECG material. All tracings originate from routine clinical care at the Geneva University Hospitals, were selected by a cardiologist to cover predefined electrocardiographic categories, and were anonymized at source before transmission to the research team. No patient is enrolled, followed or contacted, and the research team holds no key allowing re-identification.
Pilot. A technical pilot involving a small number of prehospital providers was conducted before the study start date to test the platform. Pilot sessions took place outside the standardized study conditions and are identified in the database by a dedicated centre code; their data are excluded from all analyses, as pre-specified in the protocol.
Inclusion Criteria:
Exclusion Criteria:
laurent.bourgeois@edu.ge.ch+41 22 388 34 04