Hospital-Trained AI Outperforms GPT-5 At Reading Brain Scans

Instead of learning from internet data, this AI learned directly from years of routine hospital brain scans, delivering more accurate diagnoses, better radiology reports, and stronger clinical triage than leading frontier AI models in real-world testing. 

Diverse Medical Team Analyzing Brain Scans On Computer Screens In Clinic.Study: Health system learning enables generalist neuroimaging models. Image credit: Andrey_Popov/Shutterstock.com

    Computed tomography (CT) and magnetic resonance imaging (MRI) scans have largely been excluded from the public datasets that train many artificial intelligence (AI) models, since these images often carry identifiable facial features. A recent study published in Nature Medicine explored whether training an AI model directly within a hospital, using the raw scans and workflows generated during routine patient care, could yield a high-performance neuroimaging AI model.

    Medical Scans And Public Datasets

    Modern AI models used in health and life science research are trained on large public datasets consisting of de-identified patient data. However, medical imaging rarely appears in these datasets. Hospitals generally cannot openly share most MRI and CT scans because facial features embedded in the images could identify patients, leaving general-purpose models with little real exposure to how disease actually looks inside the body.

    Existing efforts to adapt AI for medicine usually rely on a narrow set of curated scans built around a single disease, such as Alzheimer's, and on paired radiology reports for those scans. Both approaches limit how much data a model can draw on and how well its skills can be applied to new images. However, general AI systems trained on the varied reality of real-world clinical imaging could be used in the future for tasks such as clinical interpretation and patient triage.

    NeuroVFM Learned From Five Million Clinical Scan Volumes

    In the present study, the researchers assembled 566,915 MRI and CT studies, totaling 5.24 million three-dimensional volumes, collected over two decades of routine care at Michigan Medicine to build a visual foundation model called NeuroVFM. They organized these scans into a diagnostic ontology covering 82 CT and 74 MRI conditions, spanning tumors, strokes, trauma, and congenital abnormalities.

    The training of the model relied on a method called Volumetric Joint-Embedding Predictive Architecture (Vol-JEPA), which hides portions of a three-dimensional (3D) scan and trains the model to predict what the hidden regions should look like in an abstract, mathematical representation rather than reconstructing actual pixels. This allowed the system to learn anatomy and disease patterns without needing labels, annotations, or written radiology reports during self-supervised pretraining.

    The team compared NeuroVFM against five other baseline approaches, with some systems trained on the same hospital data using different learning strategies, and others trained on massive public internet datasets. The performance of the models was tested on more than 21,000 CT and 29,000 MRI studies that were not used for training, spanning 156 diagnostic categories.

    The researchers also checked eight public neuroimaging datasets covering conditions such as Alzheimer's disease, autism, and brain hemorrhage, to see whether NeuroVFM's learning held up beyond the in-house datasets. Finally, they paired the frozen NeuroVFM encoder with an open-source language model to generate written findings from scans. They tested this combination against GPT-5 and Claude Sonnet 4.5 on 300 clinician-reviewed studies and, separately, in a week-long prospective trial across the health system, where generated reports were used to flag studies needing urgent review.

    Prospective Testing Demonstrated Better Clinical Triage Accuracy

    The study found that a model trained directly on ordinary hospital imaging outperformed both proprietary AI systems and other specialized medical models at reading brain and head scans. NeuroVFM outperformed the five competing approaches across nearly all 156 diagnostic categories, achieving average accuracy scores of roughly 92.68 and 92.49 out of 100 on CT and MRI, respectively.

    It also needed far fewer labeled examples of rare conditions to match the accuracy of its competing models, in some cases requiring less than half as many. Furthermore, on public datasets it had never encountered, including cohorts built around Alzheimer's disease and autism, the model held up well, suggesting that its learning reflected genuine anatomical knowledge rather than features specific to Michigan Medicine's own patient population.

    The most interesting results came from real-world testing. When paired with an open-source language model to write radiology findings, NeuroVFM produced reports that clinicians preferred over GPT-5's more than twice as often. These reports also showed a significantly lower rate of factual errors and far fewer hallucinated findings.

    Moreover, during the week-long trial, which spanned more than 1,100 real patient scans across the health system, the combined system triaged urgent and non-urgent cases with 92.6% accuracy, compared with 71.2% for a GPT-5-based system.

    The researchers were also transparent about the model's limitations. During the prospective workflow, NeuroVFM-generated reports failed to identify critical findings in 21 of 155 patients with a genuine critical finding. While the study clarified that the errors were due to the model missing radiographic findings rather than mis-triage, the authors stressed that the model remains a research tool rather than an approved diagnostic device.

    Hospital-Based AI Training Shows Broad Clinical Potential

    Overall, the study highlighted that a generalist medical AI model need not depend on internet-scale data or hand-labeled datasets. The researchers demonstrated that using a hospital's own imaging and workflows for training enabled the NeuroVFM-based system to outperform GPT-5 and Claude Sonnet 4.5 on the neuroimaging report generation and triage tasks evaluated in the study.

    Although broader validation across health systems is still required for NeuroVFM, the researchers believe that this generalist foundation model could serve as a blueprint that other hospitals could adapt.

    Download your PDF copy now!

    Journal Reference

    Kondepudi, A., Rao, A., Zhao, C., Lyu, Y., Harake, S., Banerjee, S., Ogle, J., Joshi, R., Meissner, A.-K., Hou, X., Jiang, C., Chowdury, A., Srinivasan, A., Athey, B., Gulani, V., Pandey, A., Lee, H., & Hollon, T. (2026). Health system learning enables generalist neuroimaging models. Nature Medicine.
    DOI:10.1038/s41591-026-04497-1.https://www.nature.com/articles/s41591-026-04497-1

    Dr. Chinta Sidharthan

    Written by

    Dr. Chinta Sidharthan

    Chinta Sidharthan is a writer based in Bangalore, India. Her academic background is in evolutionary biology and genetics, and she has extensive experience in scientific research, teaching, science writing, and herpetology. Chinta holds a Ph.D. in evolutionary biology from the Indian Institute of Science and is passionate about science education, writing, animals, wildlife, and conservation. For her doctoral research, she explored the origins and diversification of blindsnakes in India, as a part of which she did extensive fieldwork in the jungles of southern India. She has received the Canadian Governor General’s bronze medal and Bangalore University gold medal for academic excellence and published her research in high-impact journals.

    Citations

    Please use one of the following formats to cite this article in your essay, paper or report:

    • APA

      Sidharthan, Chinta. (2026, July 20). Hospital-Trained AI Outperforms GPT-5 At Reading Brain Scans. AZoLifeSciences. Retrieved on July 20, 2026 from https://www.azolifesciences.com/news/20260720/Hospital-Trained-AI-Outperforms-GPT-5-At-Reading-Brain-Scans.aspx.

    • MLA

      Sidharthan, Chinta. "Hospital-Trained AI Outperforms GPT-5 At Reading Brain Scans". AZoLifeSciences. 20 July 2026. <https://www.azolifesciences.com/news/20260720/Hospital-Trained-AI-Outperforms-GPT-5-At-Reading-Brain-Scans.aspx>.

    • Chicago

      Sidharthan, Chinta. "Hospital-Trained AI Outperforms GPT-5 At Reading Brain Scans". AZoLifeSciences. https://www.azolifesciences.com/news/20260720/Hospital-Trained-AI-Outperforms-GPT-5-At-Reading-Brain-Scans.aspx. (accessed July 20, 2026).

    • Harvard

      Sidharthan, Chinta. 2026. Hospital-Trained AI Outperforms GPT-5 At Reading Brain Scans. AZoLifeSciences, viewed 20 July 2026, https://www.azolifesciences.com/news/20260720/Hospital-Trained-AI-Outperforms-GPT-5-At-Reading-Brain-Scans.aspx.

    Comments

    The opinions expressed here are the views of the writer and do not necessarily reflect the views and opinions of AZoLifeSciences.
    Post a new comment
    Post

    While we only use edited and approved content for Azthena answers, it may on occasions provide incorrect responses. Please confirm any data provided with the related suppliers or authors. We do not provide medical advice, if you search for medical information you must always consult a medical professional before acting on any information provided.

    Your questions, but not your email details will be shared with OpenAI and retained for 30 days in accordance with their privacy principles.

    Please do not ask questions that use sensitive or confidential information.

    Read the full Terms & Conditions.

    You might also like...
    Researchers Link Prenatal Fructose Exposure to Reduced Adult Neurogenesis