This retrospective, non-interventional study evaluates the prognostic performance of open-weight Large Language Models (LLMs) in the setting of a German academic emergency department. Using a full census of all consecutive emergency department cases at University Hospital Cologne between 01 January 2023 and 31 December 2025 (approximately 100,000 cases), the study assesses whether LLMs can make reliable prognostic predictions (e.g., hospital admission, imaging, diagnosis, placement) based on the initial history, vital signs, and triage category. In addition, it quantifies how strongly automated anonymization and perturbation procedures affect the models' diagnostic accuracy. This is an Investigator-Initiated Trial (IIT) with no intervention on patients.
Inclusion Criteria:
Exclusion Criteria:
Original data will be used for the analysis
Anonymized/perturbed data will be used for the analysis
Cologne, 50937, Germany
Artificial Intelligence as a Decision Making Tool in Emergency Department
Application of Large Language Models in Emergency Neurology
Large Language Models Versus Anesthesiologists for ASA Physical Status Classification
Point-of-Care AI Assistance and Critical Care Outcomes: A Randomized Trial
Diagnostic Accuracy of GPT-4o and Claude 4.6 Sonnet in Turkish ED Anamnesis Notes
Development of a Natural Language Processing Tool to Enable Clinical Research in Emergency Medicine
Evaluation of AI Large Models for Diagnosis and Treatment in Real-World Cases: Multicenter Retrospective Study
A Large Language Model in Outpatient Care