When the Genetic Map Gets It Wrong
Imagine using Google Maps to find your way somewhere, only to discover that the road you are driving on does not appear on the map. The navigation system tries to redirect you, warns that you have gone off course, and suggests another route. But there is nothing wrong with the road; the map is incomplete. A similar problem can arise when computers analyse the human genome. They compare a person’s DNA sequence with a reference sequence to identify genetic differences that might affect their health. If the reference fails to capture the natural diversity of human DNA, however, a normal genetic variant can look like a defect, or even a potentially serious mutation.
“In the natural history of a genetically influenced illness, an individual is born with a genetic predisposition (genotype), which may lead to a biological onset of the disease, followed by observable signs and symptoms (the phenotype) that prompt a diagnosis. Traditional curative medicine operates on a “phenotype-first” model, working backward from the patient’s symptoms to identify the underlying cause.” Nagy et al., 2026
A study published in GeroScience, involving researchers from Semmelweis University, draws attention to this problem and raises questions about how much we can rely on automated genome analysis when the reference used for comparison does not adequately represent human genetic diversity.
The results were presented in the study “Beyond the linear genome: how reference bias threatens preventive medicine and geroscience.” The authors are Gyula Richárd Nagy, Gyöngyi Munkácsy, Ankita Murmu, Giuliana Longo, Luis Izquierdo López, Vincenzo Cirigliano, and Balázs Győrffy.
The human genome contains approximately three billion base pairs. Its genetic code is written using four nucleotides: adenine, thymine, cytosine and guanine. Although humans share almost all of this sequence, their genomes differ at millions of positions. Most of these differences are part of normal human variation. Some have no known effect on health, while others can influence how genes function or affect a person’s risk of developing certain diseases. To identify these differences, computers compare an individual’s DNA sequence with a reference genome. One of the most widely used is GRCh38. But GRCh38 is not an average human genome. It is a reference assembly built from genetic data obtained from multiple sources, and it does not fully represent the range of genetic variation found across the world’s populations.
“The current standard, GRCh38, is not a statistical consensus of the wild type but rather a mosaic assembly containing reference minor alleles (RMAs) at clinically relevant loci. When a healthy individual carries the common, functional major allele at a locus where the reference genome harbors a rare or nonfunctional variant, standard bioinformatics pipelines—designed to detect deviations from the reference—mathematically define the healthy state as a variant.” Nagy et al., 2026
Return to the Google Maps analogy. If a map shows only one of several possible routes, any alternative may appear to be a deviation. Genome analysis can run into the same difficulty. When a patient’s DNA differs from the reference sequence, the software may interpret that difference as a genetic variant. In some cases, it introduces an artificial gap to make the two sequences align. That gap can then be misidentified as a deletion — a loss of DNA that may be flagged as potentially harmful. The result can look alarming, even when the patient carries a normal genetic variant. In other words, the software may detect the difference correctly but interpret it incorrectly.
Twenty healthy people, the same false alarm
In the study, researchers compared the whole genomes of 20 healthy Hungarians with the GRCh38 reference. Automated analysis flagged the same genetic change as a serious variant in all 20 participants. Further expert review found that the finding was not a genuine, medically relevant variant. In clinical practice, suspicious findings should undergo additional assessment before they are considered significant. That process, however, takes time and requires specialist expertise. If whole-genome sequencing were extended to millions of healthy people, reviewing every suspicious result could become a major bottleneck. Preventive genomics aims to identify increased disease risk before symptoms appear. That approach is useful only if the results are reliable enough to avoid unnecessary concern and prevent doctors from being led towards the wrong conclusions.
“This observational case series included 20 healthy adult volunteers seeking preventive genome sequencing. Approximately 6 mL of patient blood in EDTA vacutainers was collected from each participant. The Hungarian National Center for Public Health and Medicine approved the study (7177-5/2023/EÜIG) following the Declaration of Helsinki. Participants provided written informed consent and were uncompensated.” Nagy et al., 2026
Beyond a single map: a network of possible routes
One potential solution is the human pangenome, which represents genetic diversity differently from a conventional linear reference. If GRCh38 is like a map showing one main route, a pangenome is more like a map that includes several possible paths between the same points. It links shared sections of human genomes with alternative sequences found in different individuals and populations. This approach is known as a graph-based pangenome. Rather than forcing every DNA sequence to fit a single reference, it allows the software to follow alternative paths that reflect genuine human variation. If a patient carries a common variant that is not represented in the conventional reference, a pangenome could allow the software to recognise it as a valid alternative rather than treating it as a missing or altered sequence. This does not mean a pangenome will solve every problem in genetic analysis. It could, however, reduce errors that arise at the earliest stage, when DNA sequences are aligned against the reference. Pangenome models already exist, but their routine use in clinical laboratories will require better computing infrastructure, standardised procedures and extensive validation.
Why it matters for preventive medicine
Genetic testing is already used to diagnose inherited disorders, assess risk in families with known genetic predispositions and help guide treatment for certain types of cancer. In the future, genome analysis could help identify people who would benefit from earlier screening or closer monitoring for particular diseases. But genetic predisposition is not a prediction of what will inevitably happen. Health is also shaped by age, environment, lifestyle and many other factors. Distinguishing genuine genetic risks from errors introduced during computational analysis will therefore be essential.
“Future studies should focus on benchmarking experimental pangenome aligners against these advanced diagnostic tools. In conclusion, as medicine moves toward genotype-first screening, relying on a linear reference assembly risks hiding true clinical signals beneath systemic reference-biased noise. While transitioning to graph-based pangenome references represents a highly promising trajectory for the field, its routine implementation currently faces substantial practical barriers. These include the need for upgraded computational infrastructure, bioinformatic standardization, and rigorous regulatory validation prior to clinical adoption.” Nagy et al., 2026
There is also a possible problem in the opposite direction. Under certain circumstances, a genuinely harmful variant could go undetected if it matches a sequence represented as standard in the reference genome. The researchers highlight this possibility but stress that further work is needed to establish how often such a scenario actually occurs. The study has its limitations. It included only 20 people of European ancestry, so the findings cannot establish how frequently similar errors occur in other populations. Nor do they show that every computational pipeline will produce the same results.
Still, the underlying point is clear: a reliable genetic analysis depends not only on reading DNA accurately, but also on having a reference that reflects the diversity of the people being tested. A good navigation system should not treat every unfamiliar road as a wrong turn. Genetic analysis should follow the same principle: a difference from the reference is not, by itself, evidence of disease.
Image: Dr. Gyula Richárd Nagy, a clinical geneticist and associate professor in the Department of Obstetrics and Gynecology at Semmelweis University, Budapest, Hungary

