{"entity":{"id":"idea-bio1-model-predictivity-benchmark","kind":"idea","name":"Score every model system on how well it predicted real trial results","aka":[],"tldr":"No one keeps score of which laboratory models actually predicted what happened in patients. A public scoreboard would show which models to trust.","summary":"Protein structure prediction improved rapidly once CASP created a blinded, periodic benchmark. An oncology equivalent would take drugs with known but embargoed clinical outcomes, ask model owners (organoids, PDX, chips, in silico) to submit blinded predictions of response rate or ranking, and publish accuracy by model class. Over time this creates evidence for which systems merit regulatory and investment weight.","asOf":"2026-09-08","links":[{"label":"Bottleneck evidence (Lab models that fail to predict what happens in patients): Wong, Siah & Lo, Estimation of clinical trial success rates (Biostatistics 2019)","url":"https://doi.org/10.1093/biostatistics/kxx069"}],"tags":[],"related":[],"cancers":[],"sections":[],"technologies":["organoids","pdx-models","ai-drug-design"],"targets":[],"drugs":[],"companies":[],"institutions":["broad-institute"],"pathways":[],"terms":[],"trials":[],"people":[],"bottlenecks":["b-preclinical-models","b-ai-validation","b-reproducibility"],"keyPapers":["paper-wong-biostatistics"],"journals":[],"dependsOn":[],"notes":[],"hypothesis":"Blinded benchmarking reveals large and reproducible differences between model classes in predicting clinical response rates, and participation improves accuracy across rounds.","rationale":"Community benchmarks with held-out truth transformed structural biology and machine learning; oncology model validation is currently self-reported and non-comparable.","test":"Run a first round with ten agents whose phase 2 results are complete but unpublished or paywalled, and publish accuracy metrics per submitted model class.","maturity":"speculative","actor":"data","cost":"small","horizonYears":3},"route":"/ideas/idea-bio1-model-predictivity-benchmark/","neighbours":{"technology":[{"id":"ai-drug-design","kind":"technology","name":"AI-driven drug & target discovery","route":"/technologies/ai-drug-design/"},{"id":"organoids","kind":"technology","name":"Patient-derived organoids","route":"/technologies/organoids/"},{"id":"pdx-models","kind":"technology","name":"Patient-derived xenografts","route":"/technologies/pdx-models/"}],"institution":[{"id":"broad-institute","kind":"institution","name":"Broad Institute of MIT and Harvard","route":"/institutions/broad-institute/"}],"bottleneck":[{"id":"b-ai-validation","kind":"bottleneck","name":"AI that is built but not validated or deployed","route":"/bottlenecks/b-ai-validation/"},{"id":"b-preclinical-models","kind":"bottleneck","name":"Lab models that fail to predict what happens in patients","route":"/bottlenecks/b-preclinical-models/"},{"id":"b-reproducibility","kind":"bottleneck","name":"Preclinical results do not reproduce","route":"/bottlenecks/b-reproducibility/"}],"paper":[{"id":"paper-wong-biostatistics","kind":"paper","name":"Estimation of clinical trial success rates and related parameters","route":"/key-papers/paper-wong-biostatistics/"}]}}