The Reference Genome: Mostly One Man From Buffalo

Every raw DNA file is measured against a reference genome that is about 70 percent one anonymous donor. Who he was, why reference does not mean normal, and what the complete and pangenome references change.

Every position in your raw file has a reference allele - the letter the human reference genome carries at that spot. It is tempting to read the reference as an average, a consensus, or a picture of a healthy human. It is none of those things. About 70 percent of it is the DNA of one man who answered a newspaper advertisement in Buffalo, New York, in 1997, and the rest is a patchwork of a few dozen other people. Knowing that changes how you read the words “reference” and “variant” in every genetic report you will ever see.

How the reference was built

The Human Genome Project did not sequence a person. It sequenced libraries: collections of DNA fragments, each carried in a bacterial artificial chromosome, that could be mapped and read piece by piece. Roswell Park Cancer Institute in Buffalo recruited volunteers through an advertisement in The Buffalo News in March 1997 and built libraries from a handful of them. One library, RPCI-11, from a male donor, was of unusually good quality and became the project’s workhorse. It supplied about two thirds of the original draft and around 70 percent of the current reference. The remaining 30 percent came mostly from about ten further samples, with small contributions from more than fifty others.

The draft was announced in 2001, a mostly complete version in 2003, and the two builds consumer services use - GRCh37 in 2009 and GRCh38 in 2013 - are refinements of the same mosaic. If your file lists positions on GRCh37, as almost all consumer files do, the coordinates and the reference letters at every one of them come mostly from RP11. Our article on genome builds explains why positions differ between the two.

Who he was

Nobody knows, by design. Samples were anonymised before the libraries were made, and the donor himself would not know he was chosen. His genome, though, is public, and it can be studied like anyone else’s. Analyses of the reference’s own variation suggest that he had mixed African and European ancestry, most consistent with an African American - a useful corrective to the assumption that the reference is a European genome.

In 2011 Chen and Butte ran the reference through the same disease-risk pipeline used for real patients and found that it carried more than 3,500 variants associated with over a hundred conditions, including an elevated genetic risk of type 1 diabetes. Of course it did. It is a person’s genome, and every person’s genome looks like that.

Reference does not mean normal

Three consequences follow, and each one shows up in how raw files are read.

The reference is not the common allele. At around two million positions the reference carries the minor allele - the rarer of the two - because the donor happened to. Our explainer on reading the two letters separates “reference versus variant” from “major versus minor”; this is why the distinction matters.

The reference is not the working allele. The best example is blood type. The one-base deletion at rs8176719 that disables the ABO enzyme and produces type O is what the reference genome has. The A and B forms of the gene are, relative to the reference, an insertion. The reference genome is blood group O, and at this position “variant” means “has a functioning enzyme”. Our guide to blood type in raw DNA shows how to read the marker.

The reference is not a person who ever lived. It is haploid, with one letter at every position where you have two, and it switches between donors along each chromosome. Nobody has ever had this genome. So when a VCF reports 0/0, it means “same as this mosaic here”, not “normal”, and when it reports an ALT allele it means “different from this mosaic”, not “mutant”. The VCF format guide covers how those columns are written.

Gaps, bias, and the newer references

GRCh38 still had gaps - centromeres, the ends of chromosomes, and long stretches of repeated sequence that the original technology could not resolve, about 8 percent of the genome in total. In 2022 the Telomere-to-Telomere consortium closed them, adding nearly 200 million bases. They did it with a different source, a cell line called CHM13 derived from a complete hydatidiform mole - a rare growth that carries two identical copies of a single paternal genome, so there is no second haplotype to confuse the assembly.

A single genome of any origin still carries the deeper problem, which is that everything measured against one sequence sees other people’s variation less well. Chips are designed from the reference, sequencing reads are aligned to it, and variants the reference donors lacked are harder to detect. The effect is largest for populations least represented in the reference, and is one of several reasons that genetic tools work better for people of European ancestry. The Human Pangenome Reference Consortium’s 2023 draft addressed this directly: 47 people from diverse ancestries, both copies of each chromosome, assembled into a graph rather than a line, adding 119 million bases of variable sequence that the linear reference could not represent.

What changes for your file

For now, very little. Consumer services still report on GRCh37, the letters in your file are the same whatever reference you compare them with, and imputation already uses panels of thousands of genomes rather than the reference alone. When services move to GRCh38 or beyond, positions will shift and liftover will matter again. Meanwhile the useful habit is a small one: read “reference” as “what one anonymous volunteer from Buffalo had here”, and read “variant” as “different from him”. Neither word is a judgement about you.

This article is educational only.

References

  • International Human Genome Sequencing Consortium. Initial sequencing and analysis of the human genome. Nature. 2001. PubMed 11237011
  • Osoegawa K, et al. A bacterial artificial chromosome library for sequencing the complete human genome. Genome Research. 2001. PubMed 11230172
  • Ballouz S, Dobin A, Gillis JA. Is it time to change the reference genome? Genome Biology. 2019. PubMed 31399121
  • Chen R, Butte AJ. The reference human genome demonstrates high risk of type 1 diabetes and other disorders. Pacific Symposium on Biocomputing. 2011. PubMed 21121051
  • Nurk S, et al. The complete sequence of a human genome. Science. 2022. PubMed 35357919
  • Liao WW, et al. A draft human pangenome reference. Nature. 2023. PubMed 37165242
  • Genome Reference Consortium. Human genome overview, NCBI.
  • dbSNP entry for rs8176719, NCBI.

Further reading