How AI Predicts Molecular Biomarkers From H&E Slides

Transforming Pathology Through Innovation
What Are Pathology Foundation Models
From Tissue Image to Biomarker
Challenges and Clinical Opportunities
References and Further Reading


Artificial intelligence is enabling routine H&E-stained pathology slides to predict molecular biomarkers, cancer subtype, prognosis, and treatment response by identifying tissue patterns that correlate with underlying molecular changes. Advances in whole-slide imaging and pathology foundation models are expanding the role of computational pathology while highlighting the need for rigorous clinical validation before widespread adoption.

Image credit: Komsan Loonprom/Shutterstock.com

Artificial intelligence (AI) is transforming digital pathology by identifying nuanced cellular and tissue patterns in routine hematoxylin and eosin (H&E) slides that are associated with molecular biomarkers, clinical outcomes, and treatment response. Traditionally, identifying genetic mutations and biomarker profiles required expensive, time-consuming tissue processing and molecular testing.

Today, deep learning algorithms can learn morphological and spatial patterns associated with these molecular features directly from digitized routine slides, potentially accelerating screening and helping prioritize confirmatory testing. However, these systems remain predictive tools rather than substitutes for validated immunohistochemical, genomic, or other molecular assays.1,3

Save This Article as a Free PDF for Future Reference. Click Here

Transforming Pathology Through Innovation

Pathological examination plays a central role in diagnosing disease, determining cancer stage, monitoring treatment, and guiding therapy. The process typically begins with tissue staining, a critical step for visualizing specific structures or biomarkers necessary for accurate diagnosis. H&E is the most widely used stain; hematoxylin renders nuclei blue-purple and eosin stains cytoplasm and extracellular matrix pink.1,2

Traditional pathological staining is a complex, multi-step procedure that encompasses tissue collection, processing, sectioning, deparaffinization, and staining. Each stage is essential for generating high-quality samples. However, the workflow is typically labor-intensive and time-consuming, and ancillary staining or molecular testing can further extend turnaround time. Frequently, multiple stains are required for a single diagnosis, further increasing complexity and costs. These challenges can contribute to diagnostic delays and affect patient care.2

Recent advances in digital pathology and computational modeling, particularly deep learning, are transforming diagnostic pathology. Virtual staining techniques harness AI to digitally convert images of tissue stained with one dye into representations that mimic staining with another. These innovations offer the potential to increase efficiency, reduce costs, and minimize variability. However, virtual stains are generated images and may introduce artifacts or biologically inaccurate details; current systems therefore require rigorous validation and are not positioned to replace conventional staining across routine clinical practice.2

What Are Pathology Foundation Models

Pathology foundation models are large-scale AI systems trained on diverse pathology image datasets, often using self-supervision. Unlike traditional AI models that require detailed pathologist-annotated data for each specific task, these models can learn from large datasets with minimal annotation. This reduces the workload for pathologists and enables a single model to adapt to a broad range of applications, such as disease diagnosis, prognosis prediction, cancer grading, biomarker and molecular marker prediction, and immunohistochemical scoring. They usually convert image regions into reusable numerical representations, or embeddings, that can be adapted to downstream tasks using smaller labeled datasets.3

A significant advantage of these models is their ability to analyze routine slides, particularly H&E-stained images, and estimate features that ordinarily require additional specialized tests. By identifying complex patterns in thousands of images, pathology foundation models can predict the probability of molecular markers, such as gene mutations or protein expression, classify cancer subtypes, and assess prognosis. Such predictions should be interpreted as probabilistic associations and generally require confirmatory clinical testing.3

Pathology foundation models such as Virchow, CHIEF, UNI, and CONCH have been developed from large collections of digitized pathology images, with some models also incorporating paired text or other clinical data. These models have shown strong performance in tumor subtyping, grading, and biomarker prediction. Virchow, for example, used a vision-transformer architecture trained with self-supervised learning on approximately 1.5 million H&E slides and was evaluated for pan-cancer detection, rare cancer recognition, cell identification, and digital biomarker prediction.3,4

Workflow showing how pathology foundation models are trained on multimodal data, validated, and adapted for diagnostic and biomarker prediction tasks.
Pathology foundation models are trained using multimodal data, including images, text, and multi-omics information, then validated and adapted for diverse downstream applications such as cancer classification, biomarker discovery, outcome prediction, precision diagnosis, and clinical decision support. Image credit: Cheng, C.H et al (2026).

From Tissue Image to Biomarker

The rapid adoption of whole slide imaging (WSI), coupled with advances in deep learning algorithms, has significantly accelerated the momentum behind AI-driven computational pathology.5 A major breakthrough has been the automated, quantitative analysis of tissue biomarkers in digital images, enabling pathologists to obtain objective and reproducible measurements from histological slides.

Such biomarkers include quantifiable features like cell counts, nuclear morphology (size and shape), mitotic figures, staining intensity, and tissue architectural patterns. Because these features can be consistently measured by computational tools, they help minimize subjectivity and promote more reliable, reproducible diagnoses.

In colorectal and other cancers, AI models can learn tissue patterns associated with molecular biomarkers such as microsatellite instability (MSI), although MSI ordinarily requires confirmation by immunohistochemical or molecular testing. Foundation models excel at prognostic modeling by learning complex tissue patterns linked to patient outcomes. For example, CHIEF, or Clinical Histopathology Imaging Evaluation Foundation, is a general-purpose weakly supervised pathology foundation model developed for cancer detection, tumor-origin identification, molecular-profile prediction, and prognosis estimation. CHIEF was developed using 60,530 whole-slide images spanning 19 anatomical sites and validated on 19,491 slides from 32 independent datasets.1,6

AI tools have also been instrumental in the diagnosis of lung and breast cancer. In breast cancer diagnosis, these tools can assist with detecting metastatic lesions in lymph nodes, grading tumors, and quantifying immunohistochemical biomarkers.7

While distinguishing lung squamous cell carcinoma from adenocarcinoma is usually straightforward, diagnosis becomes challenging with poorly differentiated tumors or limited biopsy material. In one study, a deep learning model trained on 579 H&E-stained whole-slide images from transbronchial lung biopsies classified adenocarcinoma, squamous cell carcinoma, small-cell lung cancer, and non-neoplastic tissue. The model achieved an area under the receiver-operating-characteristic curve of 0.99 in an independent set of 83 diagnostically challenging biopsy cases.8

Challenges and Clinical Opportunities

AI models for predicting molecular biomarkers from H&E slides face several core challenges. The complexity and heterogeneity of WSIs, stemming from variations in tissue preparation, staining, and scanning, make it difficult to develop models that generalize well across different clinical settings. Additionally, the creation of large, annotated datasets is hampered by labor-intensive labeling and strict privacy regulations, which complicate data sharing and collaborative efforts. These factors collectively limit the scalability and robustness of AI solutions in real-world environments. Molecular labels can also be noisy or incomplete because different assays, sampling sites, and positivity thresholds may produce discordant ground truth.1,3

Another significant obstacle is the lack of standardization in data formats, model outputs, and integration pathways. Without standardized and interpretable outputs, AI-driven results may be difficult to validate and integrate with existing pathology systems, limiting their practical utility in healthcare. Models that perform well on internal test sets may lose accuracy when applied to slides from other laboratories, populations, scanners, or staining protocols, a problem known as domain shift.1,7

External validation must therefore include independent institutions, clinically representative case mixes, relevant specimen types, and comparisons with accepted reference assays. Prospective studies should also assess whether the model improves turnaround time, diagnostic consistency, testing efficiency, treatment selection, or patient outcomes when embedded in an actual pathology workflow.7

Despite these challenges, several promising opportunities are emerging. Advances in multimodal approaches, integrating histology, genomics, and radiomics, offer the potential for more accurate and comprehensive biomarker prediction. Techniques such as federated and swarm learning allow the use of diverse datasets across institutions while maintaining data privacy, supporting the development of more generalizable models. Multimodal foundation models may additionally combine pathology images with reports, sequencing results, and clinical variables, but missing data and differences between institutions remain important limitations.1

The ongoing improvements in model interpretability, including the use of visual explainability tools, are beginning to address clinician concerns and regulatory requirements, fostering greater trust and transparency in AI-driven pathology. Heat maps and attention visualizations can identify regions that influenced a prediction, but they do not necessarily prove that a model used biologically valid reasoning. A successful clinical adoption will hinge on developing lightweight, resource-efficient AI models that can be widely deployed, including in low-resource settings. It will also require quality-control procedures, bias assessment, cybersecurity safeguards, human oversight, regulatory evaluation, and ongoing monitoring after deployment.1,7

For the foreseeable future, the most plausible clinical role is a tiered workflow in which AI screens routine H&E slides, flags patients with a high or low predicted probability of a biomarker, and supports decisions about which cases should undergo definitive molecular testing. This approach could conserve tissue and laboratory resources while retaining established assays as the clinical reference standard.3

References and Further Reading

  1. Zhang XM, et al. Artificial intelligence in digital pathology diagnosis and analysis: technologies, challenges, and future prospects. DOI: Unable to verify from the citation provided.
  2. Lin W, et al. Virtual staining for pathology: Challenges, limitations and perspectives. DOI: Unable to verify from the citation provided.
  3. Ochi M, Komura D, Ishikawa S. Pathology Foundation Models. DOI: Unable to verify from the citation provided.
  4. Vorontsov E, et al. A foundation model for clinical-grade computational pathology and rare cancers detection. DOI: Unable to verify from the citation provided.
  5. Masjoodi S, et al. Whole Slide Imaging (WSI) in Pathology. DOI: Unable to verify from the citation provided.
  6. Wang X, et al. A pathology foundation model for cancer diagnosis and prognosis prediction. DOI: Unable to verify from the citation provided.
  7. Cheng CH, Wong CC. The role of artificial intelligence-based foundation models and "copilots" in cancer pathology. DOI: Unable to verify from the citation provided.
  8. Kanavati F, et al. A deep learning model for the classification of indeterminate lung carcinoma in biopsy whole slide images. DOI:10.1038/s41598-020-64051-x, https://www.nature.com/articles/s41598-020-64051-x

Last Updated: Jul 31, 2026

Dr. Priyom Bose

Written by

Dr. Priyom Bose

Priyom holds a Ph.D. in Plant Biology and Biotechnology from the University of Madras, India. She is an active researcher and an experienced science writer. Priyom has also co-authored several original research articles that have been published in reputed peer-reviewed journals. She is also an avid reader and an amateur photographer.

Citations

Please use one of the following formats to cite this article in your essay, paper or report:

  • APA

    Bose, Priyom. (2026, July 31). How AI Predicts Molecular Biomarkers From H&E Slides. AZoLifeSciences. Retrieved on July 31, 2026 from https://www.azolifesciences.com/article/How-AI-Is-Predicting-Molecular-Biomarkers-From-HE-Slides.aspx.

  • MLA

    Bose, Priyom. "How AI Predicts Molecular Biomarkers From H&E Slides". AZoLifeSciences. 31 July 2026. <https://www.azolifesciences.com/article/How-AI-Is-Predicting-Molecular-Biomarkers-From-HE-Slides.aspx>.

  • Chicago

    Bose, Priyom. "How AI Predicts Molecular Biomarkers From H&E Slides". AZoLifeSciences. https://www.azolifesciences.com/article/How-AI-Is-Predicting-Molecular-Biomarkers-From-HE-Slides.aspx. (accessed July 31, 2026).

  • Harvard

    Bose, Priyom. 2026. How AI Predicts Molecular Biomarkers From H&E Slides. AZoLifeSciences, viewed 31 July 2026, https://www.azolifesciences.com/article/How-AI-Is-Predicting-Molecular-Biomarkers-From-HE-Slides.aspx.

Comments

The opinions expressed here are the views of the writer and do not necessarily reflect the views and opinions of AZoLifeSciences.
Post a new comment
Post

While we only use edited and approved content for Azthena answers, it may on occasions provide incorrect responses. Please confirm any data provided with the related suppliers or authors. We do not provide medical advice, if you search for medical information you must always consult a medical professional before acting on any information provided.

Your questions, but not your email details will be shared with OpenAI and retained for 30 days in accordance with their privacy principles.

Please do not ask questions that use sensitive or confidential information.

Read the full Terms & Conditions.