WGS
Catalog entries using this tag (links open the entry card on its page):
- 1000 Genomes — High-coverage on GRCh38 — Projects
- All of Us — Flagship genomic data release (WGS + WES + array) — Projects
- All of Us — Population-scale long-read sequencing — Projects
- CKB — Haplotype-resolved reference panel (WGS) — Projects
- gnomAD v3 — Genome-only release — Projects
- MVP — Large-scale disease-specific GWAS & WGS — Projects
- TOPMed — Freeze 5, first major WGS release — Projects
- TOPMed — Freeze 8, multi-ancestry expansion — Projects
- TOPMed — WGS freeze expansion & multi-omics integration — Projects
- UK Biobank — Telomere length GWAS (WGS) — Projects
- UK Biobank — WGS full 500k release — Projects
- UK Biobank — WGS-derived CNV PheWAS — Projects
- UK Biobank — WGS/WES phasing (SHAPEIT5) — Projects
- b37 — References
- b38 — References
- CHM13 — References
- CPC — References
- GRCh37.p13 — References
- GRCh38.p14 — References
- GRCh39 (indefinitely postponed) — References
- hg19 — References
- hg38 — References
- HPRC first draft pangenome — References
- hs37d5 — References
- humanG1Kv37 — References
Entries
1000 Genomes — High-coverage on GRCh38
STAGE_PERIOD
2019–
DESCRIPTION
High-coverage whole-genome sequencing of a subset of Phase 3 samples on GRCh38; improves rare-variant discovery, phasing, and structural-variant catalogs while staying aligned with the 1000 Genomes sample framework.
URL
All of Us — Flagship genomic data release (WGS + WES + array)
PUBMED_LINK
STAGE_PERIOD
2024
DESCRIPTION
Landmark genomic data release of whole-genome sequences (WGS), whole-exome sequences (WES), and genotyping array data from over 245,000 diverse participants. Demonstrated 2.1× increase in discovery power for rare variants by including diverse populations. Data accessible via the Researcher Workbench.
URL
TITLE
Genomic data in the All of Us Research Program
All of Us — Population-scale long-read sequencing
PUBMED_LINK
STAGE_PERIOD
2025
DESCRIPTION
Population-scale long-read whole-genome sequencing in the All of Us cohort, enabling comprehensive detection of structural variants, repeat expansions, and complex genomic regions not accessible with short-read sequencing.
URL
TITLE
Population-scale Long-read Sequencing in the All of Us Research Program
CKB — Haplotype-resolved reference panel (WGS)
PUBMED_LINK
STAGE_PERIOD
2023
DESCRIPTION
High-resolution haplotype-resolved reference panel from ~5,000 CKB whole-genome sequences. One of the most comprehensive East Asian imputation references for improved GWAS fine-mapping.
URL
TITLE
A high-resolution haplotype-resolved Reference panel constructed from the China Kadoorie Biobank Study
gnomAD v3 — Genome-only release
STAGE_PERIOD
2020–2021
DESCRIPTION
gnomAD v3 released 71,702 whole genomes aligned to GRCh38, providing coverage of both coding and non-coding regions. Enabled genome-wide constraint metrics and non-coding variant interpretation. The transition to GRCh38 improved compatibility with modern genomic analyses.
URL
MVP — Large-scale disease-specific GWAS & WGS
STAGE_PERIOD
2023–2025
DESCRIPTION
Expanded disease-specific GWAS across hundreds of traits leveraging the deep EHR phenotyping in the VA system. Whole-genome sequencing of a subset of participants for comprehensive variant discovery. MVP data contributed to multi-biobank meta-analyses with FinnGen and UK Biobank spanning thousands of phenotypes.
URL
TOPMed — Freeze 5, first major WGS release
STAGE_PERIOD
2017–2018
DESCRIPTION
First major whole-genome sequencing data freeze (Freeze 5) covering ~14k genomes from diverse ancestral populations including African American, Hispanic/Latino, Asian, and European individuals. Established the TOPMed variant discovery and quality control pipeline.
URL
TOPMed — Freeze 8, multi-ancestry expansion
STAGE_PERIOD
2019–2020
DESCRIPTION
Freeze 8 expanded WGS to ~62k participants, substantially increasing representation of African American, Hispanic/Latino, Asian, and other ancestries. Enabled deep variant discovery with >400M variants including many rare and population-specific alleles. Became the gold-standard imputation reference panel for multi-ethnic studies.
URL
TOPMed — WGS freeze expansion & multi-omics integration
STAGE_PERIOD
2022–2025
DESCRIPTION
Continued expansion of TOPMed WGS to over 100k participants across >50 NHLBI cohort studies. Integration of whole-genome sequencing with epigenomics, transcriptomics, proteomics, and metabolomics data. Open access to sequencing data and variant annotations through the BRAVEO browser and dbGaP, serving as a critical resource for the global genetics community.
URL
UK Biobank — Telomere length GWAS (WGS)
PUBMED_LINK
STAGE_PERIOD
2023.11
DESCRIPTION
Genetic architecture of telomere length in 462,666 UK Biobank whole-genome sequences, identifying both common and rare variant associations.
URL
TITLE
Genetic architecture of telomere length in 462,666 UK Biobank whole-genome sequences
UK Biobank — WGS full 500k release
PUBMED_LINK
STAGE_PERIOD
2023.11
DESCRIPTION
WGS released for all 500,000 participants — the largest whole-genome sequencing dataset ever released for medical research. Sequencing by deCODE Genetics & Wellcome Sanger Institute.
URL
TITLE
World's biggest set of human genome sequences opens to scientists
UK Biobank — WGS-derived CNV PheWAS
PUBMED_LINK
STAGE_PERIOD
2023.11
DESCRIPTION
Phenome-wide analysis of copy number variants in 470,727 UK Biobank genomes, including CNV-pQTL and phenome-wide associations.
URL
TITLE
Phenome-wide analysis of copy number variants in 470,727 UK Biobank genomes
UK Biobank — WGS/WES phasing (SHAPEIT5)
PUBMED_LINK
STAGE_PERIOD
2023.11
DESCRIPTION
Accurate rare variant phasing of WGS and WES data using SHAPEIT5, enabling haplotype-based analyses.
URL
TITLE
Accurate rare variant phasing of whole-genome and whole-exome sequencing data in the UK Biobank
b37
FULL NAME
Broad Institute Homo_sapiens_assembly19 (b37)
DESCRIPTION
GRCh37-compatible reference FASTA used across Broad Institute and 1000 Genomes workflows: chromosomes 1-22, X, Y, MT, plus GL/NC unlocalized and unplaced contigs (as in the distributed assembly19 package). Coordinate system matches the 1KG/b37 ecosystem used by many GWAS imputation and joint-calling pipelines.
URL
KEYWORDS
GRCh37; 1000 Genomes; Broad; b37; reference FASTA
Main citation
Broad Institute / 1000 Genomes Project. Homo_sapiens_assembly19.fasta (b37). https://data.broadinstitute.org/snowman/hg19/
b38
FULL NAME
Broad Institute Homo_sapiens_assembly38 (b38)
DESCRIPTION
GRCh38-based reference FASTA distributed with GATK and Broad pipelines (Homo_sapiens_assembly38), including primary chromosomes and standard alternate contigs (hs38d5 decoy is distributed separately). Default reference for many germline short-variant and joint-genotyping workflows on cloud and HPC.
URL
KEYWORDS
GRCh38; GATK; Broad; b38; reference FASTA
Main citation
Broad Institute. Homo_sapiens_assembly38.fasta (GATK GRCh38 reference bundle). https://storage.googleapis.com/genomics-public-data/references/hg38/v0/
CHM13
PUBMED_LINK
FULL NAME
T2T-CHM13 v1.1 complete hydatidiform mole assembly
DESCRIPTION
Telomere-to-telomere (T2T) assembly of the CHM13 hydatidiform mole cell line, providing the first gap-resolved maps of centromeres and the full Y (from a composite). Use as a complement to GRCh38 for studying repetitive and structurally variable loci; chromosome naming and coordinates differ from GRC primary assemblies; use liftover and T2T-specific tooling where appropriate.
URL
KEYWORDS
T2T; telomere-to-telomere; complete genome; CHM13; GRCh38 alternative
TITLE
The complete sequence of a human genome.
Main citation
Nurk S, Koren S, Rhie A, Rautiainen M, ...&, Phillippy AM. (2022) The complete sequence of a human genome. Science, 376 (6588) 44-53. doi:10.1126/science.abj6987. PMID 35357919
ABSTRACT
Since its initial release in 2000, the human reference genome has covered only the euchromatic fraction of the genome, leaving important heterochromatic regions unfinished. Addressing the remaining 8% of the genome, the Telomere-to-Telomere (T2T) Consortium presents a complete 3.055 billion-base pair sequence of a human genome, T2T-CHM13, that includes gapless assemblies for all chromosomes except Y, corrects errors in the prior references, and introduces nearly 200 million base pairs of sequence containing 1956 gene predictions, 99 of which are predicted to be protein coding. The completed regions include all centromeric satellite arrays, recent segmental duplications, and the short arms of all five acrocentric chromosomes, unlocking these complex regions of the genome to variational and functional studies.
DOI
10.1126/science.abj6987
CPC
PUBMED_LINK
FULL NAME
Chinese Pangenome Consortium (phase I core)
DESCRIPTION
Phase I data from the Chinese Pangenome Consortium: 116 high-quality haplotype-phased de novo assemblies from 58 core samples across 36 minority Chinese ethnic groups (high-fidelity long-read coverage). Adds substantial novel sequence and variant discovery relative to GRCh38 and supports population-specific reference panels for Asian-ancestry genomics.
URL
KEYWORDS
pangenome; Chinese populations; long-read; haplotype; GRCh38
TITLE
A pangenome reference of 36 Chinese populations.
Main citation
Gao Y, Yang X, Chen H, Tan X, ...&, Xu S. (2023) A pangenome reference of 36 Chinese populations. Nature, 619 (7968) 112-121. doi:10.1038/s41586-023-06173-7. PMID 37316654
ABSTRACT
Human genomics is witnessing an ongoing paradigm shift from a single reference sequence to a pangenome form, but populations of Asian ancestry are underrepresented. Here we present data from the first phase of the Chinese Pangenome Consortium, including a collection of 116 high-quality and haplotype-phased de novo assemblies based on 58 core samples representing 36 minority Chinese ethnic groups. With an average 30.65× high-fidelity long-read sequence coverage, an average contiguity N50 of more than 35.63 megabases and an average total size of 3.01 gigabases, the CPC core assemblies add 189 million base pairs of euchromatic polymorphic sequences and 1,367 protein-coding gene duplications to GRCh38. We identified 15.9 million small variants and 78,072 structural variants, of which 5.9 million small variants and 34,223 structural variants were not reported in a recently released pangenome reference1. The Chinese Pangenome Consortium data demonstrate a remarkable increase in the discovery of novel and missing sequences when individuals are included from underrepresented minority ethnic groups. The missing reference sequences were enriched with archaic-derived alleles and genes that confer essential functions related to keratinization, response to ultraviolet radiation, DNA repair, immunological responses and lifespan, implying great potential for shedding new light on human evolution and recovering missing heritability in complex disease mapping.
DOI
10.1038/s41586-023-06173-7
GRCh37.p13
FULL NAME
Genome Reference Consortium Human Build 37 patch release 13
DESCRIPTION
NCBI/GRC human assembly build 37, patch 13 (GCF_000001405.25): the authoritative GRCh37 patch-level reference used for stable accessioning and alignment. Distinct from UCSC hg19/Broad b37 contig naming; always verify chromosome naming and inclusion of ALT/patch contigs when mixing resources.
URL
KEYWORDS
GRCh37; GRC; NCBI; reference assembly; patch 13
Main citation
Genome Reference Consortium. Human genome assembly GRCh37.p13 (GCF_000001405.25). National Center for Biotechnology Information.
GRCh38.p14
FULL NAME
Genome Reference Consortium Human Build 38 patch release 14
DESCRIPTION
NCBI/GRC human assembly build 38, patch 14 (GCF_000001405.40): current GRC primary human reference on the GRCh38 line, including cumulative sequence fixes and scaffold updates through p14. Use this accession when you need the exact GRC patch level that matches NCBI/RefSeq alignment products.
URL
KEYWORDS
GRCh38; GRC; NCBI; reference assembly; patch 14
Main citation
Genome Reference Consortium. Human genome assembly GRCh38.p14 (GCF_000001405.40). National Center for Biotechnology Information. https://www.ncbi.nlm.nih.gov/assembly/GCF_000001405.40
GRCh39 (indefinitely postponed) (GRCh39)
FULL NAME
Genome Reference Consortium Human Build 39 (not pursued)
DESCRIPTION
The Genome Reference Consortium announced that work toward a distinct GRCh39 assembly line was indefinitely postponed; human reference updates continue on the GRCh38 series (patches) and complementary resources such as T2T-CHM13 and pangenome references. Check the GRC human page for current guidance and patch releases.
URL
KEYWORDS
GRC; GRCh39; reference assembly; postponed
Main citation
Genome Reference Consortium. Human genome reference updates (GRCh39 indefinitely postponed; continued GRCh38 patches). https://www.ncbi.nlm.nih.gov/grc/human
hg19
FULL NAME
UCSC hg19 (GRCh37) reference bundle
DESCRIPTION
UCSC Genome Browser distribution of the GRCh37-era human reference (hg19): chromosomes chr1-22, chrX, chrY, chrM, plus unlocalized and unplaced contigs, alternate loci (e.g. chr6_apd_hap1), and related patches as packaged for the browser. Widely used in legacy pipelines and liftOver chains to/from hg38.
URL
KEYWORDS
GRCh37; UCSC; reference genome; FASTA; legacy assembly
Main citation
UCSC Genome Browser. Human reference assembly hg19 (GRCh37-aligned). https://hgdownload.soe.ucsc.edu/goldenPath/hg19/
hg38
FULL NAME
UCSC hg38 (GRCh38) reference bundle
DESCRIPTION
UCSC Genome Browser distribution of the human reference aligned to GRCh38 (primary assembly plus standard patches and decoys as packaged in the browser bigZips downloads). Chromosome names use the chr1-chrM convention; coordinates match the corresponding GRC assembly for the same patch level when sequences are identical.
URL
KEYWORDS
GRCh38; UCSC; reference genome; FASTA; primary assembly
Main citation
UCSC Genome Browser. Human reference assembly hg38 (GRCh38-aligned). https://hgdownload.soe.ucsc.edu/goldenPath/hg38/
HPRC first draft pangenome (HPRC draft)
PUBMED_LINK
FULL NAME
Human Pangenome Reference Consortium first-draft pangenome
DESCRIPTION
First-draft human pangenome from the HPRC: 47 phased diploid assemblies from diverse samples, aligned and summarized relative to GRCh38. Adds substantial euchromatic polymorphic sequence and duplicated gene content versus a single linear reference; intended for pangenome-aware alignment, variant calling, and downstream graph-based genomics (see HPRC data portal and companion software).
URL
KEYWORDS
HPRC; pangenome; graph genome; haplotypes; GRCh38
TITLE
A draft human pangenome reference.
Main citation
Liao WW, Asri M, Ebler J, Doerr D, ...&, Paten B. (2023) A draft human pangenome reference. Nature, 617 (7960) 312-324. doi:10.1038/s41586-023-05896-x. PMID 37165242
ABSTRACT
Here the Human Pangenome Reference Consortium presents a first draft of the human pangenome reference. The pangenome contains 47 phased, diploid assemblies from a cohort of genetically diverse individuals1. These assemblies cover more than 99% of the expected sequence in each genome and are more than 99% accurate at the structural and base pair levels. Based on alignments of the assemblies, we generate a draft pangenome that captures known variants and haplotypes and reveals new alleles at structurally complex loci. We also add 119 million base pairs of euchromatic polymorphic sequences and 1,115 gene duplications relative to the existing reference GRCh38. Roughly 90 million of the additional base pairs are derived from structural variation. Using our draft pangenome to analyse short-read data reduced small variant discovery errors by 34% and increased the number of structural variants detected per haplotype by 104% compared with GRCh38-based workflows, which enabled the typing of the vast majority of structural variant alleles per sample.
DOI
10.1038/s41586-023-05896-x
hs37d5
FULL NAME
1000 Genomes GRCh37 + decoy (hs37d5)
DESCRIPTION
GRCh37 (b37-style) primary chromosomes and contigs plus the hs37d5 decoy sequence set (HuRef/BAC/Fosmid/NA12878-derived sequences) to reduce spurious alignments in short-read mapping. Standard reference for Phase 3-era 1000 Genomes alignment and many imputation and low-pass WGS workflows that target the 1KG coordinate system.
URL
KEYWORDS
GRCh37; decoy; 1000 Genomes; alignment; hs37d5
Main citation
1000 Genomes Project / Broad Institute. hs37d5 reference (GRCh37 plus decoy sequences). https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/technical/reference/phase2_reference_assembly_sequence/
humanG1Kv37
FULL NAME
1000 Genomes human_g1k_v37 reference
DESCRIPTION
GRCh37-based reference FASTA distributed by the 1000 Genomes Project (human_g1k_v37): chromosomes 1-22, X, Y, MT, plus GL unlocalized/unplaced contigs, without separate haplotype scaffolds or EBV. Commonly used as the Phase 1/III alignment reference when harmonizing with public 1KG VCFs and phase panels.
URL
KEYWORDS
GRCh37; 1000 Genomes; reference FASTA; human_g1k_v37
Main citation
1000 Genomes Project. human_g1k_v37 reference (GRCh37). https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/technical/reference/