A Prospective Multicenter Clinical-Performance Study of Federated Machine Learning for Automated Interpretation of Point-of-Care Cardiac Ultrasound
A Prospective Multicenter Clinical-Performance Study of Federated Machine Learning for Automated Interpretation of Point-of-Care Cardiac Ultrasound
This prospective, multicenter study will evaluate a federated machine-learning system designed to analyze focused cardiac point-of-care ultrasound examinations. Federated learning allows participating clinical sites to contribute to model development while keeping raw ultrasound images and directly identifiable patient information within each site's controlled computing environment. Encrypted model updates, rather than patient images, will be transmitted for secure aggregation.
The prospective validation cohort will include approximately 3,000 adults undergoing clinically indicated focused cardiac ultrasound. Model performance will be compared with an expert interpretation of a comprehensive transthoracic echocardiogram performed within 24 hours. The primary objective is to determine how accurately the model identifies reduced left ventricular systolic function, defined as a left ventricular ejection fraction below 40%.
During the initial validation period, the investigational software will operate in silent mode. Its results will not be displayed to treating clinicians and will not be used to diagnose participants, select treatment, or replace standard clinical interpretation.
The study will also evaluate image-quality classification, cardiac-view recognition, performance across clinical sites and ultrasound systems, model calibration, processing time, cybersecurity, privacy resilience, and performance across demographic and clinical subgroups. Long-term monitoring will assess whether model performance changes as clinical populations, ultrasound equipment, acquisition practices, and software environments evolve during the 2026-2037 study period.
Federated Learning for Point-of-Care Cardiac Ultrasound (FL-POCUS) is a prospective, multicenter clinical-performance study of a federated machine-learning system for focused cardiac point-of-care ultrasound. The study is intended to determine whether a diagnostic model can be developed and validated across multiple clinical environments without routinely transferring raw ultrasound images or directly identifiable participant information to a central training repository.
Participating institutions may use previously collected, locally governed ultrasound examinations during the federated model-development stage. Each institution will operate a local computing node using a common model architecture, data specification, and quality-control framework. Local model updates will be encrypted and transmitted to a secure aggregation service. The aggregated parameters will then be redistributed to participating sites for subsequent training rounds. Training events, software versions, data-quality findings, and model changes will be documented in an auditable version-control system.
Privacy protections will include access controls, secure aggregation, data-minimization procedures, cybersecurity testing, and evaluation for membership-inference and model-inversion risk. Raw ultrasound images, protected health information, consent records, participant identifiers, and authentication credentials will not be placed on a public blockchain. Any distributed ledger used by the study will be limited to document hashes, version identifiers, authorized attestations, and non-sensitive audit records.
After federated development is complete, the model will be version-locked before prospective clinical validation. Approximately 3,000 adult participants will be enrolled across at least six geographically and technically diverse clinical sites. Eligible participants will be undergoing a clinically indicated focused cardiac point-of-care ultrasound examination and will have an eligible comprehensive transthoracic echocardiogram available within 24 hours of the index examination.
The locked model will operate in silent mode. Investigational outputs will not be displayed to treating clinicians and will not influence diagnosis, treatment, patient disposition, or the decision to obtain additional testing. All clinical decisions will continue to be made through the participating institution's standard care processes.
The model will evaluate standard focused cardiac views, including parasternal long-axis, parasternal short-axis, apical four-chamber, and subcostal views when available. Investigational functions may include cardiac-view classification, image-quality assessment, identification of technically limited examinations, and detection of reduced left ventricular systolic function.
The primary outcome is the area under the receiver-operating-characteristic curve for detecting a reference-standard left ventricular ejection fraction below 40%. The proposed performance criterion is an area under the curve of at least 0.85, with the lower bound of the two-sided 95% confidence interval exceeding 0.80.
The reference standard will be established using comprehensive transthoracic echocardiography. Two qualified echocardiography readers, masked to the investigational model result, will independently review eligible reference examinations. Disagreements affecting the prespecified ejection-fraction category will be resolved by a third senior reader.
Secondary evaluations will include sensitivity, specificity, positive and negative predictive values, detection of severe systolic dysfunction, cardiac-view classification accuracy, agreement with expert image-quality assessments, nondiagnostic examination rate, calibration, processing time, site-level heterogeneity, and performance by ultrasound manufacturer and transducer type.
Prespecified subgroup analyses will examine performance by age, sex, race, ethnicity, body mass index category, clinical environment, cardiac rhythm, acquisition experience, study site, and ultrasound system. Subgroup results will be reported even when the model satisfies the overall primary performance criterion.
All attempted examinations will remain in the primary intention-to-diagnose analysis, including technically limited studies and examinations with incomplete cardiac views. A model output that cannot produce a valid diagnostic classification will be counted as a test failure in the primary analysis. An evaluable-case analysis may be performed as a secondary analysis.
The study includes a long-term performance-surveillance period extending through September 30, 2037. This period will evaluate model drift, calibration changes, equipment transitions, software updates, cybersecurity events, evolving acquisition practices, and changes in the enrolled population. Prospective data used for final validation will remain separated from model-development data unless a separately governed amendment authorizes a new model version. Any updated model will receive a new version designation and must undergo independent validation before clinical use.
Reportable study events will include unauthorized data disclosure, attempted reconstruction of participant information, incorrect linkage between examinations and participants, inadvertent display of investigational results, validation-data leakage into training, material subgroup-performance disparities, cybersecurity incidents, and significant protocol deviations.
The study will not authorize automated diagnosis or autonomous clinical management. Any later investigation in which model results are displayed to clinicians or used to influence care will require a separately approved protocol, updated risk assessment, applicable regulatory review, and independent Institutional Review Board authorization.
Inclusion Criteria:
Exclusion Criteria: