HPRC2: A human pangenome reference with near-complete coverage of common genetic variation
HPRC2: A human pangenome reference with near-complete coverage of common genetic variation
Lucas, J. K.; Hebbar, P.; Liao, W.-W.; Macias-Velasco, J. F.; Novak, A. M.; Asri, M.; Balacco, J. R.; Blair, A. P.; Ebler, J.; Gardner, J. M. V.; Geleta, M.; Groza, C.; Guarracino, A.; Heringer, P.; Hickey, G.; Lu, S.; Marin, M. G.; Markovic, C.; Mastoras, M.; Mayoud, C.; McNulty, B.; Menendez, J. M.; Minkina, A.; Mohanty, S. K.; Monlong, J.; Munson, K. M.; Oshima, K. K.; Porubsky, D.; Ranallo-Benavidez, T. R.; Seligmann, W. E.; Shemirani, R.; Violich, I.; Yoo, D.; Zhuo, X.; Albracht, D.; Alexandrov, I. A.; Allen, J.; Alsheikh-Ali, A. A.; Andrews, C.; Antipov, D.; Antonacci-Fulton, L.; Arguell
AbstractA pangenome reference overcomes the inherent limitation of any individual reference genome by integrating the variation present in a population. We present the Human Pangenome Reference Consortium's (HPRC) Release 2 (HPRC2), an openly available, second phase pangenome that is an approximately fivefold expansion in genome number over HPRC Release 1 (HPRC1) and measurable improvement in genome completeness, contiguity, and accuracy. Selecting samples with a principled algorithm prioritising common variant coverage, HPRC2 contributes 460 haplotypes that together capture over 99% of common variation observed in the All of Us Research Program v8 cohort. Combining high-coverage long and ultra-long reads with modern assemblers and polishers, we produce thousands of telomere-to-telomere (T2T) chromosomes, and relative to HPRC1 halve the number of structurally unreliable regions as well as individual base errors per haplotype. We complement the assemblies with whole genome multiple alignments and gene annotations, and derive formal pangenome coordinate systems for addressing off-reference variation, demonstrating that individual human genomes contain more than one hundred thousand variants not succinctly described with respect to existing reference genomes. We also present the first matched long-read backed pantranscriptome and panepigenome at this scale, provide continuous local-ancestry estimates spanning every genome, and outline a host of new tools and applications that leverage the pangenome resource for improved genomics analysis.