Instead of learning from internet data, this AI learned directly from years of routine hospital brain scans, delivering more accurate diagnoses, better radiology reports, and stronger clinical triage than leading frontier AI models in real-world testing.
Study: Health system learning enables generalist neuroimaging models. Image credit: Andrey_Popov/Shutterstock.com
Computed tomography (CT) and magnetic resonance imaging (MRI) scans have largely been excluded from the public datasets that train many artificial intelligence (AI) models, since these images often carry identifiable facial features. A recent study published in Nature Medicine explored whether training an AI model directly within a hospital, using the raw scans and workflows generated during routine patient care, could yield a high-performance neuroimaging AI model.
Medical Scans And Public Datasets
Modern AI models used in health and life science research are trained on large public datasets consisting of de-identified patient data. However, medical imaging rarely appears in these datasets. Hospitals generally cannot openly share most MRI and CT scans because facial features embedded in the images could identify patients, leaving general-purpose models with little real exposure to how disease actually looks inside the body.
Existing efforts to adapt AI for medicine usually rely on a narrow set of curated scans built around a single disease, such as Alzheimer's, and on paired radiology reports for those scans. Both approaches limit how much data a model can draw on and how well its skills can be applied to new images. However, general AI systems trained on the varied reality of real-world clinical imaging could be used in the future for tasks such as clinical interpretation and patient triage.
NeuroVFM Learned From Five Million Clinical Scan Volumes
In the present study, the researchers assembled 566,915 MRI and CT studies, totaling 5.24 million three-dimensional volumes, collected over two decades of routine care at Michigan Medicine to build a visual foundation model called NeuroVFM. They organized these scans into a diagnostic ontology covering 82 CT and 74 MRI conditions, spanning tumors, strokes, trauma, and congenital abnormalities.
The training of the model relied on a method called Volumetric Joint-Embedding Predictive Architecture (Vol-JEPA), which hides portions of a three-dimensional (3D) scan and trains the model to predict what the hidden regions should look like in an abstract, mathematical representation rather than reconstructing actual pixels. This allowed the system to learn anatomy and disease patterns without needing labels, annotations, or written radiology reports during self-supervised pretraining.
The team compared NeuroVFM against five other baseline approaches, with some systems trained on the same hospital data using different learning strategies, and others trained on massive public internet datasets. The performance of the models was tested on more than 21,000 CT and 29,000 MRI studies that were not used for training, spanning 156 diagnostic categories.
The researchers also checked eight public neuroimaging datasets covering conditions such as Alzheimer's disease, autism, and brain hemorrhage, to see whether NeuroVFM's learning held up beyond the in-house datasets. Finally, they paired the frozen NeuroVFM encoder with an open-source language model to generate written findings from scans. They tested this combination against GPT-5 and Claude Sonnet 4.5 on 300 clinician-reviewed studies and, separately, in a week-long prospective trial across the health system, where generated reports were used to flag studies needing urgent review.
Prospective Testing Demonstrated Better Clinical Triage Accuracy
The study found that a model trained directly on ordinary hospital imaging outperformed both proprietary AI systems and other specialized medical models at reading brain and head scans. NeuroVFM outperformed the five competing approaches across nearly all 156 diagnostic categories, achieving average accuracy scores of roughly 92.68 and 92.49 out of 100 on CT and MRI, respectively.
It also needed far fewer labeled examples of rare conditions to match the accuracy of its competing models, in some cases requiring less than half as many. Furthermore, on public datasets it had never encountered, including cohorts built around Alzheimer's disease and autism, the model held up well, suggesting that its learning reflected genuine anatomical knowledge rather than features specific to Michigan Medicine's own patient population.
The most interesting results came from real-world testing. When paired with an open-source language model to write radiology findings, NeuroVFM produced reports that clinicians preferred over GPT-5's more than twice as often. These reports also showed a significantly lower rate of factual errors and far fewer hallucinated findings.
Moreover, during the week-long trial, which spanned more than 1,100 real patient scans across the health system, the combined system triaged urgent and non-urgent cases with 92.6% accuracy, compared with 71.2% for a GPT-5-based system.
The researchers were also transparent about the model's limitations. During the prospective workflow, NeuroVFM-generated reports failed to identify critical findings in 21 of 155 patients with a genuine critical finding. While the study clarified that the errors were due to the model missing radiographic findings rather than mis-triage, the authors stressed that the model remains a research tool rather than an approved diagnostic device.
Hospital-Based AI Training Shows Broad Clinical Potential
Overall, the study highlighted that a generalist medical AI model need not depend on internet-scale data or hand-labeled datasets. The researchers demonstrated that using a hospital's own imaging and workflows for training enabled the NeuroVFM-based system to outperform GPT-5 and Claude Sonnet 4.5 on the neuroimaging report generation and triage tasks evaluated in the study.
Although broader validation across health systems is still required for NeuroVFM, the researchers believe that this generalist foundation model could serve as a blueprint that other hospitals could adapt.
Download your PDF copy now!
Journal Reference
Kondepudi, A., Rao, A., Zhao, C., Lyu, Y., Harake, S., Banerjee, S., Ogle, J., Joshi, R., Meissner, A.-K., Hou, X., Jiang, C., Chowdury, A., Srinivasan, A., Athey, B., Gulani, V., Pandey, A., Lee, H., & Hollon, T. (2026). Health system learning enables generalist neuroimaging models. Nature Medicine.
DOI:10.1038/s41591-026-04497-1.https://www.nature.com/articles/s41591-026-04497-1