The Genome Aggregation Database (gnomAD) Marina

Published  . 0 views
↓ Download
The Genome Aggregation Database (gnomAD) Marina
1 / 1
The Genome Aggregation Database (gnomAD) Marina - slide 1 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 2 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 3 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 4 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 5 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 6 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 7 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 8 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 9 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 10 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 11 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 12 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 13 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 14 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 15 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 16 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 17 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 18 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 19 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 20 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 21 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 22 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 23 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 24 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 25 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 26 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 27 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 28 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 29 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 30 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 31 of 32 The Genome Aggregation Database (gnomAD) Marina - slide 32 of 32
Description: The Genome Aggregation Database (gnomAD) Marina DiStefano, Ph.D., FACMG Analysis of 800,000 diverse sequenced humans in gnomAD improves clinical interpretation and provides insight into gene function gnomAD v4 Slides from Heidi Rehm The

Related Topics

Download Presentation

"The Genome Aggregation Database (gnomAD) Marina" is the property of its rightful owner. Permission is granted to download and print the materials on this website for personal, non-commercial use only, and to display it on your personal computer provided you do not modify the materials and that you retain all copyright notices contained in the materials. By downloading content from our website, you accept the terms of this agreement.

Presentation Transcript

slide1. The Genome Aggregation Database (gnomAD) Marina DiStefano, Ph.D., FACMG<br>
slide2. Analysis of >800,000 diverse sequenced humans in gnomAD improves clinical interpretation and provides insight into gene function gnomAD v4 Slides from Heidi Rehm<br>
slide3. The growth of gnomAD v2/3 = >195,000

v4 = 807,162 GRCh37 GRCh38<br>
slide4. Contributing projects
1000 Genomes
1958 Birth Cohort
African American Coronary Artery Calcification project (AACAC)
ALSGEN
Alzheimer's Disease Sequencing Project (ADSP)
Atrial Fibrillation Genetics Consortium (AFGen)
Duke Catheterization Genetics (CATHGEN)
Bangladesh Risk of Acute Vascular Events (BRAVE) Study
​​BIOcd-plus
BioHeart
BioMe Biobank
BioVU
BipEx
Bulgarian Trios
CCDG IBD sequencing project
COPD-Gene
Crohn's & Colitis Foundation (CCFA) Genetics Initiative
ENGAGE-TIMI
Estonian Genome Center, University of Tartu (EGCUT)
Finland-United States Investigation of NIDDM Genetics (FUSION)
Finnish Migraine Study
Finnish Twin Cohort Study
FINN-ADGEN
FINRISK
Framingham Heart Study
Gene Discoveries in Subjects with Crohn’s Disease of African Descent
Genetics of Cardiometabolic Health in the Amish
Genizon Biobank
Génome Québec - Genizon Biobank
Genomic Psychiatry Cohort
GoT2D
Genotype-Tissue Expression Project (GTEx)
Health2000
Human Genome Diversity Project
Inflammatory Bowel Disease:
1000IBD project
Helsinki University Hospital Finland
IBD Genomic Medicine Consortium (iGenoMed)
IBD: REMIND
IBD: Understanding the determinants of health outcomes
Inflammatory Bowel Disease Sequencing Study
NIDDK IBD Genetics Consortium
Quebec IBD Genetics Consortium
University of Miami IBD Collaborative
IMAGINE
International Genome Sample Resource (IGSR)
Jackson Heart Study
Jewish Genome Project - funded by Bonei Olam
Kuopio Alzheimer Study
LifeLines Cohort
Lung Tissue Research Consortium (LTRC)
Material and Information Resources for Inflammatory And Digestive Diseases Biobank
McLean Program for Neuropsychiatric Research, Psychotic Disorders Division
MESTA
METabolic Syndrome In Men (METSIM)
Mass General Brigham biobank
Molecular Genetics of Cognitive Disorders in Northern Finland
Multi-Ethnic Study of Atherosclerosis (MESA)
Myocardial Infarction Genetics Consortium (MIGen):
Leicester Exome Seq
North German MI Study
Ottawa Genomics Heart Study
Pakistan Risk of Myocardial Infarction Study (PROMIS)
Precocious Coronary Artery Disease Study (PROCARDIS)
Registre Gironi del COR (REGICOR)
South German MI Study
Variation in Recovery: Role of Gender on Outcomes of Young AMI Patients (VIRGO)
National Institute of Mental Health (NIMH) Controls
NHGRI CCDG
NHLBI-GO Exome Sequencing Project (ESP)
NHLBI TOPMed
NeuroDev
Nurses' Health Study
Osaka University Graduate School of Medicine
PEGASUS
Population Architecture Using Genomics and Epidemiology (PAGE) Consortium
PRISM
Pritzker Neuropsychiatric Disorders Research Consortium
Schizophrenia Exome Sequencing Meta-Analysis (SCHEMA)
SCHEMA - Japan
SCHEMA - Spain
Schizophrenia Trios from Taiwan
Sequencing Initiative Suomi (SiSu)
SHARE
SIGMA-T2D
SubPopulations and InteRmediate Outcome Measures In COPD Study (SPIROMICS)
SUPER Study – “A Finnish study of hereditary mechanisms of psychosis disorders”
Swedish Schizophrenia & Bipolar Studies
T2D-GENES
BioMe
GoDARTS
Framingham Heart Study
T2D-SEARCH
The Cancer Genome Atlas (TCGA)
The Fund for Resources for Psychiatric Research
The Genetics of Atrial Fibrillation
The Genetics of Cardiovascular Disease: Atrial Fibrillation and Atrioventricular Block
The Vanderbilt Atrial Fibrillation Ablation Registry (VAFAR)
TheWellcomeTrust Case Control Consortium
THL Biobank consent in accordance with the Finnish Biobank Act
UCSF atrial fibrillation cohort
UKIBDGC - Pharmacogenetic
UK BioBank
Whole Genome Sequencing in Psychiatric Disorders (WGSPD)
Women's Health Initiative (WHI) Where do gnomAD samples come from? 308 data contributors >100 studies >25 countries* Australia, Bangladesh, Belgium, Canada, China, England, Finland, France, Germany, Israel, Italy, Japan, Kenya, Korea, Lithuania, Mexico, Netherlands, Pakistan, Scotland, Singapore, Spain, Sweden, United Arab Emirates, USA, Wales
*Based on country of the study's institutional review board (IRB) https://gnomad.broadinstitute.org/about
https://gnomad.broadinstitute.org/stats 126K from v2 exomes

76K from v3 genomes

417K exomes from UKBB

188K exomes from many new sources
(sequenced at Broad) TCGA removed<br>
slide5. Breakdown of gnomAD cohort phenotypes *This category includes: GTEx, 1KG, UKBB, and the Qatar Genome Project, as well as the FinnGen and MGB biobank samples when no phenotype was specified
^ includes diseases like Crohn's disease, irritable bowel syndrome, interstitial cystitis, ulcerative colitis
** Neurodevelopmental controls are unaffected parents of children with confirmed or suspected de novo cause of their neurodevelopmental disorder https://gnomad.broadinstitute.org/stats<br>
slide6. Reminder gnomAD removes cohorts recruited for severe pediatric disease but it is NOT a database of universal controls
It represents the general population (cases and controls for common disease, biobanks containing all individuals, etc)
As such, there are individuals with disease in gnomAD<br>
slide7. Age and Sex Distribution 50.3% XX
49.7% XY<br>
slide8. >910 Million Variants in v4 Short variants
Total SNVs: 786,500,648
Total InDels: 122,583,462
Variant type counts
Synonymous: 9,643,254
Missense: 16,412,219
Nonsense: 726,924
Frameshift: 1,186,588
Canonical splice site: 542,514 Structural variants
1,199,117 genome SVs
627,947 Deletions
258,882 Duplications
711 CNVs
296,184 Insertions
2,185 Inversions
13,116 Complex
92 Canonical reciprocal translocations
66,903 rare (<1% site frequency) exome CNVs
30,877 Deletions
36,026 Duplications

Average number of CNVs/SVs per person
1 rare (<1% SF) coding CNV per individual
11,844 SVs per genome 61% of high quality variants in v4 exomes were not detected in prior releases https://gnomad.broadinstitute.org/stats<br>
slide9. 99.8% of possible synonymous CpG substitutions observed
97.1% missense
83.5% predicted loss-of-function (pLoF)

Note: Methylated CpGs are highly mutable Nearly saturating discovery of CpG transitions Jeremy Guez Highly methylated CpGs<br>
slide10. 26.0% of all synonymous variants observed
18.3% missense
11.6% pLoF Not yet saturating all observable variation Jeremy Guez All variants (CpG + Non-CpG) Take home: We still need lots more data!

What’s coming:
~450K WGS form AoU
Federated gnomAD<br>
slide11. 2.9x more diverse individuals than in v2

~169,000 inferred non-European samples gnomAD v4 diversity<br>
slide12. How many variants with AF > 0.1% are we discovering per genetic ancestry group in v4 compared to v2? Konrad Karczewski Number of individuals Variants with AF > 0.1% Variant discovery per genetic ancestry group<br>
slide13. Number of individuals Variants with AF > 0.1% gnomAD v4 gnomAD v4 Variant discovery per genetic ancestry group Largest genetic ancestry group in gnomAD v4 is European

But the proportion of variants with AF > 0.1% from European samples has barely increased

Number of variants per sample is lowest for European genetic ancestry group Konrad Karczewski<br>
slide14. Easing the burden of rare disease analysis Variants with AF > 0.01% All of the variants shown are variants that would be classified as VUS using only European frequencies

Plot shows non-synonymous variants with AF < 0.01% in European samples

Increase in representation in v4 means the number of variants that move from VUS to LB has increased

> 200k more variants can now be considered LB Konrad Karczewski<br>
slide15. Gene Constraint: Genes intolerant of variation are more likely responsible for disease phenotypes Comparing observed rare (AF < 0.1%) variants to expectation generated from mutational model rebuilt using gnomAD v4
Model is selection-neutral Kaitlin Samocha Synonymous r = 0.983 r = 0.975 r = 0.816 Missense pLoF<br>
slide16. Karczewski et al. (2020) 86% in v4 ExAC v2 v4 Genes powered to detect constraint 72% of genes in gnomAD v2 were well-powered to detect constraint against predicted loss-of-function variants (observe >= 10 pLoF)<br>
slide17. LOEUF scores in v4 LOEUF = loss-of-function observed/expected upper bound fraction (LOEUF) Kaitlin Samocha Percentage of gene list (%) LOEUF decile (%) Autosomal recessive Haploinsufficient Olfactory genes More constrained Less constrained<br>
slide18. Recommended criteria to identify novel candidate genes to be assessed, reported, and shared
A gene in which predicted loss of function (pLOF) (nonsense, frameshift, essential splice site, whole or partial gene deletion or other structural variant that disrupts the coding region of the gene) or missense variant is observed to be de novo in a proband and the gene is strongly constrained for pLOF and/or missense variation. For proband-only analyses, strongly constrained genes with heterozygous pLOF variants absent from population databases42 should be considered as well. See footnotea and Table 1 for further discussion of constraint. Posted February 11, 2024. doi: https://doi.org/10.1101/2024.02.05.579012<br>
slide19. Gene constraint on gnomAD v4 https://gnomad.broadinstitute.org/gene/ENSG00000171316?dataset=gnomad_r4<br>
slide20. Amer J Hum Genet, 2023.110:1496-1508. Not all pLOF variants are LOF, especially in population databases<br>
slide21. gnomAD FAQs<br>
slide22. Are all samples from versions 2 and 3 included in version 4, and is the method for variant identification consistent across these versions? Not all exome samples from v2 are in v4
Mostly due to slight shifts in QC or removing individuals who were related to a sample in the genome data
Most v3 genome samples are still present
Due to minor updates in HGDP/1KG subset<br>
slide23. How far away (in terms of # of individuals in the dataset) are we from finer scale constraint measures such as missense constraint at the codon level? We would need 40-100x our sample size in v2 to be able to determine significant deviations on a per codon level
Previously calculated >5 million individuals given current approaches<br>
slide24. Sometimes total filtering allele frequency is not provided, what are the reasons?
If filtering allele frequency is not provided is it OK to use highest allele frequency of 5 general continental populations such as Admixed American, African, East Asian, European (Non-Finnish), South Asian? Reasons FAF is not provided
variant is only present in a group we exclude from the calculations (e.g. Amish (ami), Ashkenazi Jewish (asj), European Finnish (fin), Middle Eastern (mid), and "Remaining Individuals" (rmi))
singletons
If FAF is not provided
Context-dependent since this depends on the associated disease<br>
slide25. Will allele frequencies be categorized by subethnic groups, such as Korean or Japanese, as seen in version 2? If so, when can we expect this breakdown? This is of interest, but not on our task list in the near future
Lower priority than fixing bugs and adding features not yet present in v4 (e.g. pext)<br>
slide26. Could you comment about phenotypes and if they were in most cases known not to contain severe paediatric-onset disorders? Historically removed cohorts ascertained for severe pediatric disease
Contains UKBB and other biobanks, with no filtering on phenotype/EHR<br>
slide27. Were any genome and exome data sourced from the same individual? Checked relatedness between individuals in exome and genome
Removed related individuals from exome data in v4 to prevent overestimation<br>
slide28. Will there be a non-cancer cohort in gnomAD v4? When will it be released? No, we do not feel that the prevalence of any disease is high enough to warrant subsets
increased inclusion of biobank samples
V4 not very different from v2 non-cancer subset (only removed cohorts recruited for cancer)<br>
slide29. Can you clarify how the structural variant datasets from v2 and v4 relate to one another?
What is the overlap between the v2 and v4 SV datasets (are all v2 samples included in v4?)? Might different algorithms/QC/etc account for the large frequency differences observed (higher false positive rates?)? e.g. Looking at a particular X chromosome SV, the frequency in v2 (DUP_X_52893) is rather high and only in males (16 males at >0.5% and no females), but the same region in v4 (DUP_CHRX_19591362) is more consistent with expectations (Mendelian inheritance), other population databases, and our internal experience (5 females and 2 males, and 100x less frequent than in v2).
Differences in v2 and v4 mostly due to:
Hg38 instead of Hg19
More samples in v4 versus v2 (some overlapping, ~15% of v4 is from v2)
Same GATK-SV pipeline was used for discovery
New SV refinement pipeline to improve precision
V2 trained filtering model based on trio family structures; aimed at minimizing mendelian violation rate
V4 trained and integrated multiple machine learning models; utilized inheritance info and long-read Pac-bio calls in matched samples<br>
slide30. Will MANE Plus Clinical transcripts be incorporated and displayed? No plans to include as of yet<br>
slide31. Is there a timeline that can be viewed by the public for when each feature is released? We make blog posts after major features are released
Changelog for minor and major feature<br>
slide32. Acknowledgments Samantha Baxter Katherine Chao Sinéad Chapman Mark Daly Phil Darnowsky Julia Goodrich Riley Grant Qin He Stephen Jahl Konrad Karczewski Kristen Laricchia Wenhan Lu Daniel MacArthur Ben Neale Anne O'Donnell-Luria Daniel Marten Heidi Rehm Kaitlin Samocha Matt Solomonson Christine Stevens Mike Talkowski Chris Vittal Ben Weisburd Mike Wilson Lauren Witzgall Core gnomAD team Data contributors Mitochondrial
Sarah E. Calvo
Nicole Lake
Eric Banks
David Benjamin
James Emery
Kiran Garimella
Laura Gauthier
Andrea Haessly
Monkol Lek
Vamsi K. Mootha
Sebastian Schoenherr
Megan Shand Structural Variants
Ryan Collins
Harrison Brand
Jack Fu
Eric Banks
Ted Brookings
Laura Gauthier
Chelsea Lowther
Tom Lyons
Sam Novod
Ted Sharpe
Mark Walker
Harold Wang
Christopher Whelan
Xuefang Zhao Data Generation
Eric Banks
Sam Bryant
Louis Bergelson
Kristian Cibulskis
Miguel Covarrubias
Laura Gauthier
Trevyn Langsford
Christopher Llanwarne
Ruchi Munshi
Sam Novod
Nikelle Petrillo
David Roazen
Valentin Ruano-Rubio
Megan Shand
Jonn Smith
Kathleen Tibbetts
Charlotte Tolonen Production and Analysis
Ryan Collins
Siwei Chen
Laura Gauthier
Sanna Gudmundsson
Dan King
Zan Koenig
Alicia Martin
Patrick Schultz
Eleanor Seaby
Moriel Singer-Berk
James Ware
Nicola Whiffin
Mary Yohannes
Jeremy Guez Broad Genomics Platform
Stacey Gabriel
Kristen Connolly
Steven Ferriera Ethics
Stacey Donnelly
Kelly Flannagan
Namrata Gupta
Emily Lipscomb
Andrea Saltzman
Molly Schleicher gnomAD SAB
Shawneequa Callier
Adam Frankish
Cecilia Lindgren
Andrés Moreno
Nicky Mulder

Yuki Okada
Tina Peseran
Sharon Plon
Aaron Quinlan
Heiko Runz Alumni
Jessica Alföldi (operations)
Irina Armean (production and analysis)
Beryl Cummings (production and analysis)
Eleina England (production and analysis)
Emily Evangelista (production and analysis)
Yossi Farjoun (data generation)
Laurent Francioli (production and analysis)
Jeff Gentry (data generation)
Thibault Jeandet (data generation)
Diane Kaplan (data generation)
Monkol Lek (production and analysis)
Eric Minikel (analysis)
William Phu (production and analysis)
Tim Poterba (production/Hail)
Dan Rhodes (production and analysis)
Nareh Sahakian (data generation)
Cotton Seed (production/Hail)
Rachel Son (production and analysis)
Jose Soto (data generation contributor)
Kat Tarasova (operations)
Grace Tiao (product owner, production and analysis)
Gordon Wade (data generation, structural variants)
Arcturus Wang (production/Hail)
Qingbo Wang (production and analysis)
Nick Watts (browser)<br>