The human genome contains millions of genetic differences, but determining which ones disrupt biological function - and how - remains a fundamental challenge. Dr. Katja Luck and her team at the Institute of Molecular Biology (IMB) in Mainz have developed a new approach to tackle this problem in intrinsically disordered protein regions, which are particularly difficult to study. Combining protein motif analysis with structural modelling, they identified more than 1,000 genetic variants that are likely to have deleterious effects. Their findings, published in Nature Structural & Molecular Biology, offer new leads for understanding the molecular basis of many genetic diseases.
An uncharacterized variant in an IDR. Image Credit: Institute of Molecular Biology gGmbH
As genome sequencing becomes more common in both research and medicine, researchers are uncovering vast numbers of inter-individual genetic variants in humans. The biggest challenge now lies in identifying which of the many thousands of variants can disrupt protein function in ways that lead to disease. This information is especially important for understanding the mechanisms of rare genetic diseases, where patients often have many uncharacterized genetic variants and it is unclear which contribute to the disease and which are harmless.
More Than a Third of Missense Mutations Occur in Intrinsically Disordered Protein Regions
Katja Luck and her research team set out to address this critical gap between DNA sequence and protein function using a machine learning method that combines protein sequence motif searches and structural modelling of the corresponding regions with AlphaFold. Specifically, the team focused on protein interactions involving intrinsically disordered regions (IDRs), which are flexible stretches in proteins that do not fold into stable three-dimensional shapes, but are nonetheless essential for protein binding and function. IDRs contain 37% of all uncharacterized missense mutations, but they are very difficult to study because most existing tools to predict the effect of missense mutations work best on folded protein regions. As a result, many genetic variants in IDRs remain uncharacterized, creating a significant bottleneck in the study of genetic diseases.
Uncovering Which Genetic Variants Disrupt Protein Binding
For their study, they analyzed more than 50,000 published human protein-protein interactions and identified 1,300 interactions involving IDRs, which they structurally annotated. Katja and her team then identified missense variants falling within these interaction regions that are likely to disrupt protein binding, and hence have functional consequences. In total, Katja and her team identified 1,187 potentially pathogenic genetic variants in IDRs. Importantly, they were able to confirm several of these deleterious effects experimentally, including some variants that leading computational tools had predicted to be benign. This showed that even flexible, poorly structured regions of proteins can contain crucial interaction sites whose disruption may contribute directly to disease.
Most patients that undergo whole genome or exome sequencing remain without a genetic diagnosis because we cannot predict well which of the identified mutations are likely disease-causing. This hinders selection and development of therapies and prohibits a better understanding of the underlying disease mechanisms. Our study is an important step forward in developing tools for clinicians to close this gap.”
Dr. Katja Luck, Group Leader, IMB Mainz
From Uncertain Genetic Variants to Testable Disease Mechanisms
By shedding light on this long-overlooked part of the proteome, Katja and her team have improved our understanding of how human genetic variation connects to functional molecular mechanisms. These findings will be able to help researchers and clinicians better identify disease-causing mutations, particularly in rare disorders, and pave the way towards developing a cure.