Stanford's Merlin puts vision-language AI on full 3D CT scans
The Nature study tested Merlin across 752 tasks and released code, weights and a de-identified abdominal CT dataset.
By Ryan Merket · Published
Primary source: Nature
Why it matters
Merlin shows how radiology AI is moving from single-purpose classifiers toward reusable 3D foundation models trained on routine hospital data, with open code and a gated dataset that make the work easier to test than most clinical AI papers.

Louis Blankemeier, Ashwin Kumar and Akshay S. Chaudhari's Stanford-led team published Merlin, a 3D vision-language foundation model for computed tomography, in a March 4th Nature paper that takes aim at one of radiology AI's practical gaps: most medical vision-language models have been built around 2D images and shorter text, while CT interpretation is volumetric, text-heavy and tied to patient history.
Merlin was trained on paired abdominal CT scans, diagnosis codes and radiology reports, using more than 6 million images from 15,331 CT scans, more than 1.8 million diagnosis codes and more than 6 million report tokens in the training set. The researchers evaluated the model on 6 task types and 752 individual tasks, including zero-shot findings classification, phenotype classification, image-report retrieval, 5-year chronic disease prediction, radiology report generation and 3D organ segmentation.
The paper's authors include researchers from Stanford's Center for Artificial Intelligence in Medicine and Imaging, Stanford Radiology, UC Berkeley, University of Wisconsin-Madison, Hospital Israelita Albert Einstein, University Hospital Zurich, Chang Gung Memorial Hospital and Google. Blankemeier and Kumar contributed equally, according to the paper, and Chaudhari served as corresponding author and principal investigator.
Merlin's core bet is that hospitals already sit on the supervision signal needed to train useful medical imaging models. Instead of relying on manually labeled CT scans for each downstream task, Merlin learns from CT volumes paired with structured electronic health record codes and the unstructured reports radiologists already write. That matters because CT studies can contain hundreds of axial slices, and abdominal CT interpretation often requires checking many organs, vessels and incidental findings in a single exam.
The team validated Merlin internally on 5,137 CT scans and externally on 44,098 CT scans from three independent clinical sites and two public datasets. In the Nature paper, Merlin outperformed 2D vision-language baselines, CT foundation models and off-the-shelf radiology models across the reported benchmark suite. Its zero-shot findings classification covered 30 abdominal CT findings, while phenotype classification covered 692 phenotypes derived from grouped diagnosis codes.
The strongest technical claim is practical rather than theatrical: the model can process an entire 3D CT volume at once. That separates Merlin from approaches that flatten CT into 2D slices or aggregate slice-level predictions after the fact. The authors argue that full-volume processing better fits anatomy, where structures change across all three spatial dimensions and where clinically meaningful findings may span multiple slices.
The report-generation results show both the promise and the boundary of the work. Merlin outperformed RadFM on the paper's reported report-generation metrics, including RadGraph-F1, BERTScore, ROUGE-2 and BLEU. In a qualitative example, however, the authors showed Merlin missing cholelithiasis that appeared in the human report and generating an invented series or slice location. That is the part of the paper that should keep buyers, founders and hospital AI committees from treating benchmark wins as clinical readiness.
The chronic disease prediction experiments are the bigger commercial clue. Merlin was fine-tuned to predict whether patients without a baseline diagnosis would develop one of six chronic diseases within five years: chronic kidney disease, osteoporosis, cardiovascular disease, ischemic heart disease, hypertension and diabetes. Using all downstream labels, Merlin reported an average AUROC of 0.757; using 10% of labels, it reported 0.708. The pitch embedded in those numbers is that CT scans ordered for one reason may contain risk signals for diseases that were not the original indication.
The researchers also released the Merlin GitHub repository, a Hugging Face model page, a PyPI package and the Merlin Abdominal CT Dataset. The dataset contains 25,494 abdominal CT scans paired with radiology reports, covering 18,317 unique patients, according to the paper's data availability statement. Access requires a data-use agreement, and the scans were de-identified and hosted through Stanford AIMI.
That release is a real contribution because medical imaging AI is often bottlenecked by private datasets, uneven labels and non-reproducible benchmarks. It also has limits. The external clinical datasets used for evaluation are not public because of patient privacy and data-use agreements. That means outside groups can inspect the code and use the released dataset, but they cannot fully reproduce every external validation claim from the paper.
Merlin also lands in a field where academic models and commercial interests increasingly overlap. The paper's competing-interests section discloses that Blankemeier has received consulting compensation from Google, is a co-founder and employee of Cognita Imaging and owns equity in Radiology Partners. Zhihong Chen is also listed as a co-founder and employee of Cognita Imaging with equity interest in Radiology Partners. Chaudhari, unrelated to this work, is listed as a co-founder receiving salary support from Cognita Imaging, with equity interests in Radiology Partners, Subtle Medical, Brain Key and LVIS Corp., and consulting work for several healthcare and imaging-related entities. The paper says all of their work on Merlin was performed as part of Stanford University.
The disclosure does not undercut the paper. It frames the market around it. A 3D CT foundation model that can classify findings, retrieve comparable reports, help draft reports, segment organs and estimate future disease risk sits close to several active business models in radiology AI: workflow assistance, coding support, opportunistic screening and imaging-derived risk stratification.
The authors are careful on deployment language. They write that Merlin may assist with interpretation and reduce radiologist burden, and the work remains a research model rather than a cleared clinical product. That distinction matters because radiology already has hundreds of AI-enabled devices cleared by the Food and Drug Administration, and many are narrow tools designed around specific findings or workflows. Merlin points toward a broader backbone model that health systems could adapt internally, especially because the paper reports training on a single NVIDIA A6000 GPU over about 160 hours.
For AI founders, the lesson is less about one Stanford model than about where the defensibility is moving. General vision-language AI has raced ahead on web images and captions. Clinical imaging rewards groups that can lawfully assemble multimodal hospital data, align it with reports and EHR codes, test it across sites and publish enough of the stack for others to verify. Merlin is one of the clearest public examples of that play in CT.