Endogenous metabolites participate in nearly every fundamental biological process, serving not only as metabolic intermediates but also as signaling molecules that influence physiology and disease. Advances in metabolomics and chemical proteomics have uncovered unexpected functions and protein partners for many familiar metabolites.
However, systematically identifying functional targets remains difficult because metabolite-protein binding affinities vary widely, biological effects depend on disease and cellular context, and experimental methods cannot always connect physical binding with a specific phenotype. Existing computational approaches often emphasize chemical structures or omics data while underusing information about diseases, phenotypes, and cellular locations. These limitations create a need for computational strategies that integrate multiple types of biomedical evidence to prioritize biologically meaningful metabolite targets.
A study (DOI: 10.48130/targetome-0026-0024) published in Targetome on 17 June 2026 by Hao Zhang's team, Shanghai University of Traditional Chinese Medicine, reports that integrating multidimensional biological fingerprints with an attention mechanism can accurately prioritize disease-related protein targets and reveal previously unverified metabolite-protein interactions.
To construct DeepETD, the researchers assembled 15,872 metabolite-protein pairs from the Human Metabolome Database, the IUPHAR/BPS Guide to PHARMACOLOGY, ChEMBL, and BindingDB. Pairs with half-maximal inhibitory concentration (IC50) values of 100 nmol·L−1 or below were treated as positive samples, producing 2,565 positive and 13,307 negative examples. The team mined PubMed abstracts and incorporated Disease Ontology and Human Phenotype Ontology data to generate fingerprints describing each metabolite and protein through associated diseases, cellular phenotypes, and subcellular locations. An attention layer assigned different weights to these features before a deep neural network estimated interaction probabilities.
The dataset was divided 80:20 for training and validation, while ten-fold cross-validation, ablation tests, and comparisons with XGBoost, CatBoost, DrugBAN, TransformerCPI, and MGNDTI assessed performance. DeepETD produced training and validation AUC-ROC values of 0.9771 and 0.8470, respectively, with corresponding accuracies of 92.74% and 82.49%. Ten-fold cross-validation yielded a mean AUC of 0.8484 and mean accuracy of 0.8344, indicating stable performance. Removing the attention mechanism significantly reduced accuracy, while eliminating disease-related features caused the greatest decline among the three fingerprint dimensions. In global ranking tests, the top 10 predictions reached a mean precision of 0.92 and represented a 5.73-fold enrichment over the validation set's baseline positive rate. The model also ranked known targets of dopamine and estradiol highly within Parkinson's disease and breast cancer contexts.
The researchers then applied DeepETD to 3,382 metabolites across ten disease contexts, generating 33,820 records for the Endogenous Metabolites Target Discovery Database (EMTDD). Microscale thermophoresis experiments tested predicted targets for testosterone and leukotriene B4. The assays supported direct binding between testosterone and NR2F6, as well as between leukotriene B4 and PRMT2 and ACAT1; molecular docking provided compatible binding models. Not every tested prediction was experimentally confirmed, underscoring the need for further validation.
Overall, DeepETD offers a computational complement to chemical proteomics by transforming dispersed biomedical knowledge into disease-aware target predictions. Its combination of interpretable feature weighting, competitive predictive performance, experimental testing, and an open database may help researchers prioritize targets before undertaking resource-intensive laboratory studies. The approach nevertheless depends on the availability and quality of existing biological annotations, and its fixed affinity threshold and imbalanced training data may introduce uncertainty.
Source:
Journal reference:
Xu, Z., et al. (2026) DeepETD: a novel deep-learning based model for endogenous metabolite target discovery. Targetome. DOI: 10.48130/targetome-0026-0024. https://www.maxapress.com/article/doi/10.48130/targetome-0026-0024