Max Planck Institute Unveils New Tools to Identify Pathogenic Variants

Researchers at the Max Planck Institute for Molecular Genetics have unveiled new bioinformatics tools called Dicast and TandemTwister in studies published in NAR Genomics & Bioinformatics and Genome Biology, designed to help identify pathogenic structural variants and tandem repeats in human genomes to improve genetic diagnostics.

With modern sequencing technologies able to locate disease causes precisely within the 3 billion base pairs of the human genome, only about 30 to 40 percent of patients receive a clear molecular diagnosis, according to Martin Vingron’s laboratory at the Max Planck Institute for Molecular Genetics. Historically, diagnostic focus centered primarily on point mutations, leaving complex structural variants and repetitive sequences largely unaddressed.

Machine Learning Targets Complex Structural Variants

Structural variants affect entire segments of the human genome, leaving sections missing, duplicated, or displaced. Because these genetic rearrangements are often larger than the short DNA reads generated by standard sequencing technologies, detecting them reliably has remained a persistent challenge for researchers and clinicians alike.

To overcome this hurdle, researchers in Martin Vingron’s laboratory turned to advanced machine learning algorithms. The core idea is that this method allows us to learn the patterns behind real structural variants and thus correctly classify new variants, explained Nico Alavi, the first author of the study published in Genome Biology.

Their newly developed tool, designated as Dicast, successfully filters out false-positive artifacts that previously required laborious manual checks. When tested against patient data, the algorithm detected all pathogenic structural variants while effectively discarding a substantial volume of false positives.

Genotyping Tandem Repeats With TandemTwister

Beyond large structural rearrangements, researchers focused on tandem repeats—short genetic motifs repeated consecutively directly one after another. These elements vary widely across healthy individuals and disease states alike, playing crucial roles in paternity testing, forensics, and particularly neurological conditions driven by replication errors.

In a second paper published in NAR Genomics & Bioinformatics, first authors Lion Ward Al Raei and Maryam Ghareghani detailed an algorithm designed to count basic repetition frequencies rapidly and precisely from sequencing data. The accompanying software package, titled TandemTwister, also provides advanced visualization features to help laboratories identify pathological repeats in patient datasets.

In the second paper, we have developed an algorithm that can quickly and precisely count, based on sequencing data, how often a basic motif is repeated. We also provide a tool that visualizes this data.

Martin Vingron, Max Planck Institute for Molecular Genetics

Bridging Basic Research and Clinical Diagnostics

Both bioinformatics tools are already available for use in basic research environments and clinical diagnostic pipelines. By streamlining the detection of structural variants and tandem repeats, the software aims to close the diagnostic gap for rare genetic disorders where traditional short-circuit screening falls short.

To transition these research breakthroughs into practical medical applications, members of the research team have established a startup venture called Lucid Genomics. Current sequencing technologies, combined with our specialized analysis algorithms, promise to further improve genetic diagnostics, Martin Vingron noted, highlighting how computational pattern recognition continues to reshape genomic medicine.

Pathogen Discovery: Bioinformatics Tools by Balaji Chattopadhyay