{"entity":{"id":"idea-data-living-llm-oncology-benchmark","kind":"idea","name":"A monthly-updated benchmark for AI answers to oncology questions with citation accuracy","aka":[],"tldr":"Test the large language models doctors and patients are already using against a continually refreshed set of cancer questions, scoring not just correct answers but whether the sources they cite are real and support the claim.","summary":"Clinicians and patients use general-purpose language models for oncology questions; evaluations are static, quickly outdated and rarely check citations. The proposal is a living benchmark: new questions each month drawn from recent practice changes, expert-graded answers, and scoring of citation validity and support, with public leaderboards and per-cancer breakdowns, run by an independent academic consortium.","asOf":"2026-09-08","links":[{"label":"Bottleneck evidence (Knowledge reaches practice too slowly): Morris, Wooding & Grant, The answer is 17 years, what is the question (JRSM 2011)","url":"https://doi.org/10.1258/jrsm.2011.110180"}],"tags":[],"related":[],"cancers":[],"sections":["ai-computation"],"technologies":[],"targets":[],"drugs":[],"companies":[],"institutions":[],"pathways":[],"terms":[],"trials":[],"people":[],"bottlenecks":["b-knowledge-diffusion","b-ai-validation","b-misinformation"],"keyPapers":["paper-morris-j-r-soc-med"],"journals":[],"dependsOn":[],"notes":[],"hypothesis":"Public, living evaluation will drive measurable improvement in citation accuracy and currency of oncology answers across models within a year, and will identify failure modes (outdated standards, hallucinated trials) that static benchmarks miss.","rationale":"Public benchmarks have driven progress in every area of machine learning; medical question benchmarks exist but are static and do not test currency, which is the key oncology failure.","test":"Run the benchmark monthly for a year on the major models; publish trends; check whether model releases show improvement on the citation and currency metrics.","maturity":"early-clinical","actor":"research","cost":"small","horizonYears":1},"route":"/ideas/idea-data-living-llm-oncology-benchmark/","neighbours":{"section":[{"id":"ai-computation","kind":"section","name":"AI & Computation","route":"/fronts/ai-computation/"}],"bottleneck":[{"id":"b-ai-validation","kind":"bottleneck","name":"AI that is built but not validated or deployed","route":"/bottlenecks/b-ai-validation/"},{"id":"b-knowledge-diffusion","kind":"bottleneck","name":"Knowledge reaches practice too slowly","route":"/bottlenecks/b-knowledge-diffusion/"},{"id":"b-misinformation","kind":"bottleneck","name":"Misinformation and unproven therapies","route":"/bottlenecks/b-misinformation/"}],"paper":[{"id":"paper-morris-j-r-soc-med","kind":"paper","name":"The answer is 17 years, what is the question: understanding time lags in translational research","route":"/key-papers/paper-morris-j-r-soc-med/"}]}}