BEACON (Benchmarking AI for Clinical Oncology decisioNmaking) is a prospective, multicentre, comparative, blinded, non-interventional benchmark evaluating the treatment recommendations of five frontier large language models (LLMs) against the recommendations of multidisciplinary tumour boards (RCP) in oncology treatment planning. One hundred standardised synthetic cases (20 per localisation, across breast, lung, urological, digestive and gynaecological cancers) are submitted as identical structured input to two independent tumour boards per localisation and to five frontier LLMs. Each recommendation - human or model - is decomposed into five predefined decision domains (intent, surgery, radiotherapy, systemic therapy, work-up and biomarkers) and scored 0/1/2 for concordance against a two-tier reference: the consensus of the two tumour boards, complemented by an a priori locked guideline matrix (ESMO, NCCN). The primary endpoint is domain-level concordance between LLM and RCP consensus, expressed as a linearly weighted Cohen's kappa. A co-primary safety endpoint captures the proportion of recommendations carrying serious harm potential, because concordance alone can conceal dangerous errors. Because expert boards may disagree with one another on identical cases, model performance is always interpreted against the human consensus. BEACON is designed as reusable, openly licensed, pre-registered infrastructure: all synthetic cases, evaluation rubrics, the locked guideline matrix, scoring algorithms and verbatim prompts are released for full reproducibility.
Inclusion Criteria:
Exclusion Criteria:
jean-emmanuel.bibault@aphp.fr
jerome.lambert@u-paris.fr0142499742 ext. +33
Agreement Between Large Language Model-Generated Treatment Recommendations With Guideline-Based and Tumor Board Decisions in Gastrointestinal Cancer
Large Language Models to Aid Gynecological Oncology Treatment
Large Language Models Assist in Tumor MDT
Evaluating AI and Human Expert Decisions in Colorectal Cancer
Preliminary Evaluation of a Large Language Model-Based Tool for Complex Surgical Decision Support in Lung Cancer
Use and Acceptance of Large Language Models for Cancer Shared Decision-Making
Impact of COMORBIDities After Radical Cystectomy Using a Predictive Method With Artificial Intelligence
Concordance Between Large Language Model and Multidisciplinary Team Recommendations in Rectal Cancer