The human genome: Finished in 2003 but really finished in 2022?

A promise made in 1988

In 1988, a group of American scientists wrote down a goal. They wanted to read the complete DNA sequence of a human being. The word “complete” was there from the very first day.

The Human Genome Project began in 1990. In 2003 it was declared finished. There were speeches, press conferences and headlines. But the job was not really done. About eight per cent of the genome was still missing. Scientists knew the missing pieces were there. They simply could not read them.

Those last pieces were only filled in during the 2020s. Today scientists can read a genome from the tip of one chromosome to the tip of the other. This is called telomere-to-telomere sequencing, or T2T for short. Telomeres are the protective caps at the ends of chromosomes, so the name just means “from tip to tip”.

Figure 1. How the human reference genome changed between 1990 and 2026.

Why the machines kept failing

A sequencing machine cannot read a whole chromosome in one go. It chops the DNA into small pieces, reads each piece, and then a computer looks for overlaps and puts the pieces back in order. It is like rebuilding a shredded book.

This works well when every page is different. It fails badly when pages repeat. And human DNA repeats a lot. Near the middle of each chromosome sits the centromere, where a short block of about 171 letters is copied over and over for millions of letters. The old machines read pieces of only 500 to 800 letters, and later machines read as few as 100. In a repeat, a piece that short fits in ten thousand places at once. That is the same as fitting nowhere.

So those regions were left blank. In the reference file they appeared as long rows of the letter N, meaning “something is here, but we cannot read it”. For twenty years, no plant or animal genome was finished without such holes. Most human genomes were never rebuilt at all. They were simply compared against the old reference, which misses roughly fifteen per cent of the genome and quietly measures everyone against one person’s DNA.

The fix was surprisingly simple

The answer was to read longer pieces.

Nanopore sequencing pulls a single strand of DNA through a tiny hole and reads it as it passes. It can read millions of letters in one continuous piece. In 2018, it read straight through a human centromere for the first time. Around the same time, a second technology called PacBio HiFi arrived. Its reads are shorter but extremely accurate.

Neither one was enough on its own. Long reads made mistakes. Accurate reads were too short. Together, with new software to combine them, they worked. A group called the Telomere-to-Telomere Consortium formed in 2018 and finished the X chromosome by 2020. In 2022 it published the first complete human genome, and the Y chromosome followed in 2023.

Figure 2. How a complete genome is built today. Accuracy and length are combined, then the two parental copies are separated.

What was hiding in the gaps

The missing eight per cent was not junk. Much of it turned out to be recent duplications, where the genome has copied a gene and kept both versions. These copies are hard to tell apart, which is exactly why the old methods lost them. Once the full sequence was available, the amount of duplicated DNA that scientists could see roughly doubled.

This matters because duplicated regions change fast and often carry genes linked to diet, immunity, development and brain function. They are also where several important medical genes sit, including genes that control how the body breaks down common drugs.

From one genome to anyone’s genome

The first complete genome came from an unusual cell line with only one set of chromosomes. Real people have two, one from each parent, and telling them apart is hard. Newer software solved this, either by using DNA from the parents or by using a laboratory method called Hi-C that shows which pieces sit close together inside the cell.

That opened the door to the Human Pangenome Project, which is building a collection of complete genomes from people of many different ancestries instead of relying on a single reference. Complete genomes for apes, monkeys and common laboratory animals have followed. The goal has shifted three times in thirty years: first to finish the human genome, then to finish many, and now to be able to finish anyone’s.

One honest warning

It is worth being careful with the word “complete”. Even now, only a small number of published genomes meet the highest quality standard. Errors can still hide inside long repeats, where there is nothing unique to check against. Good practice is for researchers to explain exactly how a genome was built and checked, rather than treating “T2T” as a badge.

What this could mean in a hospital

Today, genetic testing usually means running one panel for one suspected condition, then another, then another. Each test compares the patient to a standard reference.

A complete personal genome flips this around. If the patient’s own genome has been read properly, a doctor can simply ask questions of it. Are there any unexpected repeat expansions? Which version of this drug-processing gene does this person carry? One sequence replaces a stack of separate tests, and the answers do not depend on how similar the patient happens to be to the old reference.

Figure 3. A possible future path from patient sample to clinical decision.

Reading is not the same as understanding

Here is the catch. We can now read the whole book, but we still cannot read most of it well.

When a change in DNA breaks a protein, scientists can usually predict the effect, because the code for building proteins is well understood. But most of our DNA does not code for proteins. It switches genes on and off, controls timing, and shapes how cells develop. Changes in those regions are much harder to interpret.

This is where artificial intelligence enters. AI models for protein structure only worked because decades of careful protein data existed first. Doing the same for whole genomes will need complete genomes paired with matching measurements from the same person or cell line, covering many ancestries, tissues and stages of life. That collection is only now being built.

Finishing the human genome took thirty-eight years. Understanding it will probably take longer. But for the first time, nothing is missing from the page.

Timeline: the Human Genome Project

YearWhat happened
1979A computer method for piecing short DNA reads together is described
1988A US report sets the goal: read the complete human DNA sequence
1990The Human Genome Project officially starts
2000The shotgun method is proven on the fruit fly genome
2001Two draft human genomes are published, one public and one private
2003The genome is declared complete, but repeats are left out
2003–2019New genomes are compared to the old reference instead of being rebuilt

Timeline: the telomere-to-telomere era

YearWhat happened
2018Nanopore long reads read through a human centromere for the first time
2018The Telomere-to-Telomere Consortium is formed
2020The first complete human chromosome, the X, is finished
2021Quality standards for error-free genome assembly are agreed
2022T2T-CHM13, the first complete human genome, is published
2022The Human Pangenome Project is announced
2023The Y chromosome is finished; software makes two-copy genomes possible
2024Faster software scales the method to many species
2025Complete ape genomes and a South Asian reference genome are released
2026A benchmark human genome and a complete macaque genome are published

Where complete genomes are useful

  • Diagnosis. One complete sequence can answer many questions at once, instead of running a separate test for each suspected condition.
  • Repeat diseases. Conditions caused by long stretches of repeated DNA can be measured directly rather than estimated.
  • Medicines. Genes that control how the body handles drugs sit in duplicated regions that were previously hard to read.
  • Cancer. Comparing a tumour to the patient’s own genome, rather than a standard reference, gives cleaner results.
  • Fairness in research. A collection of genomes from many ancestries reduces the bias built into a single reference.
  • Animals and conservation. Complete genomes allow a species to be recorded in full, including those close to extinction.

References

  1. Venter, J.C. et al. (2001). The sequence of the human genome. Science 291, 1304–1351.
  2. Zhang, S., Xu, N., Lu, Y., Nie, Y., Li, Z., de Gennaro, L., La Torraca, A., Fu, L., Zhang, Z., Chen, J., et al. (2026). Complete subtelomeric architectures in a complete rhesus macaque reference genome. Cell, 189, 4909–4921.e15. https://doi.org/10.1016/j.cell.2026.02.018
  3. Avsec, Ž., Latysheva, N., Cheng, J., Novati, G., Taylor, K. R., Ward, T., Bycroft, C., Nicolaisen, L., Arvaniti, E., Pan, J., et al. (2026). Advancing regulatory variant effect prediction with AlphaGenome. Nature, 649, 1206–1218. https://doi.org/10.1038/s41586-025-10014-0
  4. Phillippy, A. M., Mao, Y., Kang, Y., Šikić, M., & Miga, K. H. (2026). Filling the holes in whole genomes: A vision for personalized genomics from telomere to telomere. Cell, 189, 4825–4828. https://doi.org/10.1016/j.cell.2026.07.019

Photo of author

Dr. Jawahar

Dr. Jawahar is a plant biotechnologist specializing in stress physiology, molecular biology, tissue culture, and metabolic engineering. His research focuses on understanding the molecular mechanisms underlying salinity and drought tolerance, particularly the roles of osmolytes, abscisic acid (ABA) signaling, and stress-responsive genes. He has also contributed significantly to enhancing the production of valuable plant secondary metabolites, including colchicine, through in vitro culture and biotechnological approaches. Dr. Jawahar has authored numerous research articles, reviews, and book chapters published in leading journals and international publishers, including PLOS ONE, Environmental and Experimental Botany, Physiologia Plantarum, and Industrial Crops and Products. His research interests include functional genomics, metabolomics, crop improvement, and sustainable agricultural biotechnology.

Follow on X

LinkedIn

WhatsApp

Telegram