Imagine doctors had the potential to understand exactly what cells caused a patient's cancer, or whether pathogens contributed to the disease. They could then use the information to tailor a treatment plan to the patient's specific cancer. But answering such questions would mean wading through data from thousands of experiments locked in massive databases around the globe. Moreover, the search would take several days at the very least.
Now, researchers at the Berlin Institute of Medical Systems Biology of the Max Delbrück Center (MDC-BIMSB) present a search engine that radically simplifies such tasks: "Malva." It is the first platform that can quickly sort through massive single-cell data using sequence information only, explains Daniel León-Periñán, first author of the study in "Nature." León-Periñán is a doctoral student in the Systems Biology of Gene Regulatory Elements lab of Dr. Nikolaus Rajewsky, Director of MDC-BIMSB.
Like Google did for the internet 30 years ago, Malva allows scientists and AI tools to search across millions of cells in seconds - without downloading huge files or needing a reference genome, and without deep computational expertise. Malva transforms static transcriptomic atlases into dynamic resources, which will further our understanding of RNA biology. It will also be potentially transformative in helping researchers understand how health slides into disease, or how and which cells respond to specific medical treatments."
Nikolaus Rajewsky, Senior Author of the paper, Max Delbrück Center
The Need for a Tool to Search RNA Data
Single-cell RNA sequencing gives researchers a remarkably detailed view of what is happening inside individual cells at any given point in time. Over the past decade, researchers around the world have amassed terabytes of data. But anyone wishing to mine it to learn more about a DNA or RNA sequence of interest faces multiple hurdles. They would need to download and reprocess petabytes of raw files, which no single lab has the capacity to do, and figure out how to standardize data from different sources.
What's more, because of the way the data is indexed, information about RNA isoforms - multiple RNA variants coded by the same gene - is extremely limited.
To simplify the task and to expand the types of questions the data can answer, León-Periñán and co-first author Dr. Nikos Karaiskos, also from the Rajewsky lab, have reprocessed data from public repositories and made it searchable by nucleotide sequence. Malva also indexes spatial data, so researchers can also find where in a tissue section a particular RNA is located.
Malva is designed to expand continuously. Every time new single-cell RNA data becomes available in the literature, it gets downloaded to a server located at the Max Delbrück Center, processed and added to existing data.
Broad Application
There are innumerable use-case scenarios, says Karaiskos. "They can range from the very simple, like: 'In what cell type is this particular gene expressed?' to much more complex."
Other platforms can also be used to answer simple questions, Karaiskos adds. But they can't, for example, answer questions about RNA biology. This is because RNA isoforms are not each indexed to a reference gene individually, but rather treated as a single entity. This makes it impossible to distinguish whether any particular RNA isoform is expressed in any specific cell type, he says. "Having the flexibility to search by RNA sequence in Malva gives us the ability to answer questions from this data that were previously not answerable."
Moreover, because these platforms map their data to a reference genome, which is a composite of a few individual human DNA samples, anyone looking for information about RNA produced from rare or unique gene variants may not find much. Malva, on the other hand, enables access to a much more diverse RNA dataset.
The Malva platform also includes sequence information from multiple species, including bacteria, viruses and fungi. Thus, researchers can study how these microorganisms affect human cells and cause disease.
The platform is currently freely available to scientists. The three researchers are in the early phases of launching a start-up to make Malva commercially available. A patent on the technology is pending.
Source:
Journal reference:
León-Periñán, et al. (2026) Ultrafast and reference-free sequence discovery in single-cell data. Nature. DOI: 10.1038/s41586-026-10975-w. https://www.nature.com/articles/s41586-026-10975-w