Merlin is a model trained on 15,000 CT scans with their reports that can find and describe hundreds of findings.
Merlin is a 3D vision-language model for abdominal CT from Stanford whose image encoder is aligned with both the free-text radiology report and structured electronic health record codes. The 2024 arXiv paper describes training on 15,000 CT scans, amounting to 6M images and 6M EHR codes, and shows zero-shot classification of hundreds of findings plus report generation. It is aimed at radiology research groups exploring report-supervised learning, a route that avoids hand labelling every finding. Its limits are that the data come from a single institution and cover the abdomen only, so generalisation to other scanners, populations and body regions is untested; for oncology the relevance is in finding and describing lesions rather than staging. For a newcomer: Merlin learned to read abdominal CT scans by studying the reports radiologists wrote about them.
3D image encoder aligned with report text and structured codes.
Query for this technology: (TITLE:"Merlin" OR ABSTRACT:"Merlin" OR TITLE:"Stanford abdominal CT vision-language model" OR ABSTRACT:"Stanford abdominal CT vision-language model") AND (cancer OR tumor OR tumour OR oncology OR carcinoma OR lymphoma OR leukemia OR leukaemia OR myeloma OR sarcoma OR melanoma OR glioma). Results are unfiltered search hits about Merlin (Stanford abdominal CT vision-language model), not a curated reading list.
Shares AI in oncology roadmap: pattern readers → foundation models → agents in the workflow, AI in radiology, CT (computed tomography) and the tags foundation-model, radiology.
Shares AI in oncology roadmap: pattern readers → foundation models → agents in the workflow, AI in radiology and the tags foundation-model, radiology.
Shares Pathology & radiology foundation models, AI in radiology and the tags foundation-model, radiology.
Shares AI in radiology and the tags foundation-model, radiology.
Shares Stanford Health Care / Stanford Cancer Institute, Pathology & radiology foundation models, AI in oncology roadmap: pattern readers → foundation models → agents in the workflow and the tag foundation-model.
Shares Pathology & radiology foundation models, AI in oncology roadmap: pattern readers → foundation models → agents in the workflow and the tag foundation-model.
Shares Pathology & radiology foundation models, AI in oncology roadmap: pattern readers → foundation models → agents in the workflow and the tag foundation-model.
Shares Stanford Health Care / Stanford Cancer Institute and the tag foundation-model.
Open-source projects that implement or serve this technology, from OnCo's own catalogue: licence and last activity as the repository reported them on the day of the fetch. Listing is not endorsement; check the licence before reuse and the validation before clinical use.
Stanford's vision-language foundation model for abdominal CT trained with radiology reports and diagnosis codes.