Skip to content

GWAS

Catalog entries using this tag (links open the entry card on its page):

Entries

Causal ML for scGenomics (Causal ML sc)

AI GWAS Causal ML Single Cell Machine Learning Nat Genet
PUBMED_LINK
40164735
FULL NAME
Causal Machine Learning for Single-Cell Genomics
DESCRIPTION
A Perspective from Nature Genetics delineating the application of causal machine learning to single-cell genomics. Discusses causal models, challenges in inferring causative roles of genes from single-cell omics data combined with perturbation screens, and the potential for integrating causal ML with GWAS to understand disease mechanisms at single-cell resolution.
TITLE
Causal machine learning for single-cell genomics.
ABSTRACT
Advances in single-cell '-omics' allow unprecedented insights into the transcriptional profiles of individual cells and, when combined with large-scale perturbation screens, enable measuring of the effect of targeted perturbations on the whole transcriptome. In this Perspective, we delineate the application of causal machine learning to single-cell genomics and its associated challenges, presenting the causal model most commonly applied to single-cell biology.
DOI
10.1038/s41588-025-02124-2

DeepNull

AI GWAS Deep Learning Covariate Adjustment Statistical Power Nat Commun
PUBMED_LINK
35017556
FULL NAME
DeepNull - Deep Learning for Non-linear Covariate Adjustment in GWAS
DESCRIPTION
DeepNull is a method that identifies and adjusts for non-linear and interactive covariate effects in GWAS using a deep neural network. It maintains tight control of type I error while increasing statistical power by up to 20% in the presence of non-linear covariate effects. Published in Nature Communications.
TITLE
DeepNull models non-linear covariate effects to improve phenotypic prediction and association power.
ABSTRACT
Genome-wide association studies (GWASs) examine the association between genotype and phenotype while adjusting for a set of covariates. Although the covariates may have non-linear or interactive effects, due to the challenge of specifying the model, GWAS often neglect such terms. Here we introduce DeepNull, a method that identifies and adjusts for non-linear and interactive covariate effects using a deep neural network, maintaining tight control of the type I error while increasing statistical power by up to 20%.
DOI
10.1038/s41467-021-27930-0

DL for PRS Survey (DL PRS Survey)

AI GWAS Polygenic Risk Score Deep Learning Survey Review Brief Bioinform
PUBMED_LINK
40802796
FULL NAME
A Survey on Deep Learning for Polygenic Risk Scores
DESCRIPTION
A comprehensive survey of deep learning approaches for polygenic risk scores (PRS). Reviews how neural networks can model non-linear relationships between genetic variants and disease risk, going beyond traditional linear PRS methods, and assesses their performance across different traits and architectures. Published in Briefings in Bioinformatics.
TITLE
A survey on deep learning for polygenic risk scores.
ABSTRACT
Polygenic risk scores (PRS) combine the effects of multiple genetic variants to predict an individual's genetic predisposition to a disease. PRS typically rely on linear models, which assume that all genetic variants act independently. There is growing interest in applying deep learning neural networks to model PRS given their ability to model non-linear relationships. We conducted a survey of the literature to investigate how neural networks are being applied to PRS.
DOI
10.1093/bib/bbaf373

DL Thoracic Aorta GWAS (DL Aorta GWAS)

AI GWAS Deep Learning Medical Imaging UK Biobank Nat Genet
PUBMED_LINK
34837083
FULL NAME
Deep Learning Enables Genetic Analysis of the Human Thoracic Aorta
DESCRIPTION
Applied a pretrained CNN (transferred from natural image recognition, e.g. ResNet/Inception-like architecture) to 4.6 million cardiac MRI images from UK Biobank, trained on only 116 manually annotated samples to regress ascending and descending thoracic aorta dimensions. GWAS identified 82 ascending and 47 descending aorta loci. Demonstrates transfer learning from natural images to medical imaging for rapid biobank-scale phenotyping.
KEYWORDS
deep learning, CNN, ImageNet transfer learning, cardiac MRI, thoracic aorta, image regression, UK Biobank
TITLE
Deep learning enables genetic analysis of the human thoracic aorta.
ABSTRACT
Enlargement or aneurysm of the aorta predisposes to dissection, an important cause of sudden death. We trained a deep learning model to evaluate the dimensions of the ascending and descending thoracic aorta in 4.6 million cardiac magnetic resonance images from the UK Biobank. We then conducted genome-wide association studies in 39,688 individuals, identifying 82 loci associated with ascending and 47 with descending thoracic aortic diameter. Transcriptome-wide analyses, rare-variant burden tests and human aortic single nucleus RNA sequencing prioritized genes including FBN1 and MFAP5.
DOI
10.1038/s41588-021-00962-4

GWANN

AI GWAS Neural Network Alzheimer's Disease Gene-level Association Brief Bioinform
PUBMED_LINK
39775791
FULL NAME
GWANN - Genome-Wide Association Neural Networks
DESCRIPTION
GWANN (Genome-Wide Association Neural Networks) is a novel approach that uses neural networks to perform gene-level association studies. Applied to Alzheimer's disease in UK Biobank, GWANN identifies genes linked to family history of AD by aggregating SNP-level information at the gene level through neural network architectures. Published in Briefings in Bioinformatics.
TITLE
Genome-wide association neural networks identify genes linked to family history of Alzheimer's disease.
ABSTRACT
Augmenting traditional genome-wide association studies (GWAS) with advanced machine learning algorithms can allow the detection of novel signals in available cohorts. We introduce "genome-wide association neural networks (GWANN)", a novel approach that uses neural networks (NNs) to perform a gene-level association study with family history of Alzheimer's disease (AD) in UK Biobank.
DOI
10.1093/bib/bbae704

Haas ME (ML Liver Fat GWAS)

AI GWAS Imaging Machine Learning Liver Fat Abdominal MRI UK Biobank
PUBMED_LINK
34957434
FULL NAME
Machine Learning Enables New Insights into Genetic Contributions to Liver Fat Accumulation
DESCRIPTION
Developed an abdominal MRI-based machine-learning regression model (gradient-boosted regression on raw MRI signal intensities) to accurately estimate liver fat from UK Biobank abdominal MRI scans (correlation 0.97-0.99 with ground truth). Trained on 4,511 participants with gold-standard MRI biomarker measurements and applied to 32,192 additional individuals. GWAS identified 8 associated variants (5 novel: MTARC1, ADH1B, TRIB1, GPAM, MAST3) and a polygenic score strongly associated with future chronic liver disease risk (HR>1.32 per SD, p<9e-17).
KEYWORDS
MRI signal regression, liver fat quantification, abdominal MRI, hepatic steatosis, gradient boosting, UK Biobank
TITLE
Machine learning enables new insights into genetic contributions to liver fat accumulation.
Main citation
Haas ME, Pirruccello JP, Friedman SN, Wang M, ...&, Khera AV. (2021) Machine learning enables new insights into genetic contributions to liver fat accumulation. Cell Genom, 1 (3). doi:10.1016/j.xgen.2021.100066. PMID 34957434
ABSTRACT
Excess liver fat, called hepatic steatosis, is a leading risk factor for end-stage liver disease and cardiometabolic diseases but often remains undiagnosed in clinical practice because of the need for direct imaging assessments. We developed an abdominal MRI-based machine-learning algorithm to accurately estimate liver fat from a truth dataset of 4,511 middle-aged UK Biobank participants, enabling quantification in 32,192 additional individuals. A genome-wide association study of common genetic variants and liver fat replicated three known associations and identified five newly associated variants.
DOI
10.1016/j.xgen.2021.100066

iGWAS

AI GWAS Imaging Deep Learning Self-Supervised Learning Retinal Fundus Phenotyping Contrastive Learning
PUBMED_LINK
38728357
FULL NAME
Image-Based Genome-Wide Association of Self-Supervised Deep Phenotyping of Retina Fundus Images
DESCRIPTION
iGWAS uses self-supervised contrastive learning (SimCLR-style framework with a CNN encoder backbone) to extract a 128-dimensional phenotype vector directly from retinal fundus images without any manual labels. Trained on 40,000 EyePACS images via instance discrimination, then applied to 130,329 UK Biobank fundus images. GWAS on these 128 learned phenotypes identified 14 genome-wide significant loci. First demonstration of unsupervised deep phenotyping for image-based GWAS — discovering genetic associations without predefined human annotations.
KEYWORDS
self-supervised contrastive learning, SimCLR, CNN, retinal fundus, deep phenotyping, image-based GWAS, UK Biobank
TITLE
iGWAS: Image-Based Genome-Wide Association of Self-Supervised Deep Phenotyping of Retina Fundus Images.
Main citation
Xie Z, Zhang T, Kim S, Sun J, Forouzandeh P, Chen R, Zhi D. (2024) iGWAS: Image-Based Genome-Wide Association of Self-Supervised Deep Phenotyping of Retina Fundus Images. PLOS Genetics, 20(5):e1011273. doi:10.1371/journal.pgen.1011273. PMID 38728357
ABSTRACT
Existing imaging genetics studies have been mostly limited in scope by using imaging-derived phenotypes defined by human experts. Here, leveraging new breakthroughs in self-supervised deep representation learning, we propose a new approach, image-based genome-wide association study (iGWAS), for identifying genetic factors associated with phenotypes discovered from medical images using contrastive learning. Using retinal fundus photos, our model extracts a 128-dimensional vector representing features of the retina as phenotypes. We identified 14 loci with genome-wide significance.
DOI
10.1371/journal.pgen.1011273

Imaging Genomics Review

AI GWAS Imaging Imaging-derived phenotypes Review IDP MRI Deep Learning Mendelian Randomization Nat Rev Genet
PUBMED_LINK
42409965
FULL NAME
Genetic analysis of imaging-derived phenotypes
DESCRIPTION
A comprehensive review of imaging-derived phenotypes (IDPs) for genetic analysis. Covers MRI, CT, X-ray, OCT, and other imaging modalities; traditional (FSL, FreeSurfer) and deep learning (U-Net, nnU-Net, self-supervised contrastive learning) pipelines for IDP extraction; biobank-scale cohorts (UK Biobank, All of Us); GWAS of brain, cardiac, retinal, abdominal, and skeletal IDPs; Mendelian randomization for causal inference linking organ structure to disease; and emerging challenges in multi-organ phenomics, longitudinal imaging, and clinical translation of IDP-based PRS.
URL
https://www.nature.com/articles/s41576-026-00989-5
TITLE
Genetic analysis of imaging-derived phenotypes.
Main citation
Bian Y, Akey JM. (2026) Genetic analysis of imaging-derived phenotypes. Nature Reviews Genetics. doi:10.1038/s41576-026-00989-5. PMID 42409965
ABSTRACT
Imaging-derived phenotypes (IDPs) developed from medical imaging data, such as magnetic resonance imaging, computed tomography and X-ray scans, are traits that provide quantitative information on anatomical and functional properties of organs and tissues. IDPs are powerful tools for identifying biomarkers and studying disease mechanisms. When coupled with genetic data, IDPs can be analysed as heritable phenotypes using modern gene mapping methods to uncover genotype-phenotype relationships. The field of imaging genomics is rapidly maturing, with the emergence of high-quality imaging datasets collected in biobank-scale cohorts and sophisticated computational methods for extracting IDPs from imaging data, including tools that leverage machine learning. Here we review common imaging modalities and analytical approaches for developing IDPs, discuss biological insights gleaned from the large-scale genetic analysis of imaging traits and highlight emerging areas and remaining challenges that must be overcome to realize the full potential of IDPs for genetic analysis.
DOI
10.1038/s41576-026-00989-5

InsightGWAS (Migraine) (InsightGWAS)

AI GWAS Transformer Deep Learning Migraine Transfer Learning Nat Commun
PUBMED_LINK
41372126
FULL NAME
InsightGWAS - Transformer-based Deep Learning Enhances Migraine GWAS
DESCRIPTION
InsightGWAS is a Transformer-based model that enhances genetic discovery for migraine GWAS by integrating functional annotations and leveraging transfer learning from GWAS datasets of major depressive disorder. It identified 293 previously unreported loci from 53,109 cases and 230,876 controls. Published in Nature Communications.
TITLE
Transformer-based deep learning enhances discovery in migraine GWAS.
ABSTRACT
Migraine is a complex neurological disorder with substantial heritability, yet genome-wide association studies (GWAS) have explained only a fraction of its genetic component. We developed InsightGWAS, a Transformer-based model, to enhance genetic discovery for migraine by integrating functional annotations and leveraging transfer learning from GWAS datasets of major depressive disorder (MDD), identifying 293 previously unreported loci.
DOI
10.1038/s41467-025-65991-7

Khurshid S (DL LV Mass GWAS)

AI GWAS Imaging Deep Learning Cardiac MRI Left Ventricular Mass UK Biobank
PUBMED_LINK
36944631
FULL NAME
Clinical and Genetic Associations of Deep Learning-Derived Cardiac Magnetic Resonance-Based Left Ventricular Mass
DESCRIPTION
Applied a CNN-based segmentation model (U-Net style architecture) to automatically segment left ventricular myocardium from cardiac MRI in 43,230 UK Biobank participants. The segmented contours were used to compute left ventricular mass indexed to body surface area (LVMI), enabling GWAS that identified 12 associations (11 novel) implicating genes associated with cardiac contractility and cardiomyopathy. The LVMI polygenic risk score validated in independent Mass General Brigham cohort.
KEYWORDS
deep learning, cardiac MRI segmentation, U-Net, left ventricular mass, CNN, cardiomyopathy, UK Biobank
TITLE
Clinical and genetic associations of deep learning-derived cardiac magnetic resonance-based left ventricular mass.
Main citation
Khurshid S, Lazarte J, Pirruccello JP, ...&, Lubitz SA. (2023) Clinical and genetic associations of deep learning-derived cardiac magnetic resonance-based left ventricular mass. Nat Commun, 14 (1) 1558. doi:10.1038/s41467-023-37173-w. PMID 36944631
ABSTRACT
Left ventricular mass is a risk marker for cardiovascular events, and may indicate an underlying cardiomyopathy. Cardiac magnetic resonance is the gold-standard for left ventricular mass estimation, but is challenging to obtain at scale. Here, we use deep learning to enable genome-wide association study of cardiac magnetic resonance-derived left ventricular mass indexed to body surface area within 43,230 UK Biobank participants. We identify 12 genome-wide associations (1 known at TTN and 11 novel for left ventricular mass).
DOI
10.1038/s41467-023-37173-w

Liu Y (DL Organ MRI GWAS)

AI GWAS Imaging Deep Learning Abdominal MRI Organ Traits UK Biobank
PUBMED_LINK
34128465
FULL NAME
Genetic Architecture of 11 Organ Traits Derived from Abdominal MRI Using Deep Learning
DESCRIPTION
Applied a U-Net-based CNN segmentation pipeline to over 38,000 abdominal MRI scans from UK Biobank. The deep learning model automatically segmented 7 organs/tissues (liver, pancreas, kidneys, spleen, lungs, visceral adipose tissue, subcutaneous adipose tissue) and quantified their volume, fat content (via signal intensity), and iron content (via T2* mapping). GWAS on these 11 DL-derived traits identified 93 independent genome-wide significant associations (heritability 8-44%), including 4 novel liver trait associations.
KEYWORDS
deep learning, abdominal MRI segmentation, U-Net, organ volume quantification, liver fat, pancreas iron, UK Biobank
TITLE
Genetic architecture of 11 organ traits derived from abdominal MRI using deep learning.
Main citation
Liu Y, Basty N, Whitcher B, Bell JD, ...&, Cule M. (2021) Genetic architecture of 11 organ traits derived from abdominal MRI using deep learning. Elife, 10. doi:10.7554/eLife.65554. PMID 34128465
ABSTRACT
Cardiometabolic diseases are an increasing global health burden. While socioeconomic, environmental, behavioural, and genetic risk factors have been identified, a better understanding of the underlying mechanisms is required to develop more effective interventions. Magnetic resonance imaging (MRI) has been used to assess organ health, but biobank-scale studies are still in their infancy. Using over 38,000 abdominal MRI scans in the UK Biobank, we used deep learning to quantify volume, fat, and iron in seven organs and tissues, and demonstrate that imaging-derived phenotypes reflect health status. We identify 93 independent genome-wide significant associations.
DOI
10.7554/eLife.65554

MILTON

AI GWAS Machine Learning Disease Prediction UK Biobank Multi-omics Nat Genet
PUBMED_LINK
39261665
FULL NAME
MILTON - Machine Learning with Phenotype Associations for Disease Prediction
DESCRIPTION
MILTON is an ensemble machine learning framework that utilizes biomarkers and multi-omics data to predict 3,213 diseases in the UK Biobank. It predicts incident disease cases undiagnosed at time of recruitment and demonstrates utility in augmenting genetic association discovery by empowering case-control GWAS with predicted phenotypes. Published in Nature Genetics.
TITLE
Disease prediction with multi-omics and biomarkers empowers case-control genetic discoveries in the UK Biobank.
ABSTRACT
The emergence of biobank-level datasets offers new opportunities to discover novel biomarkers and develop predictive algorithms for human disease. Here, we present an ensemble machine-learning framework (machine learning with phenotype associations, MILTON) utilizing a range of biomarkers to predict 3,213 diseases in the UK Biobank. MILTON predicts incident disease cases undiagnosed at time of recruitment, largely outperforming available polygenic risk scores, and augments genetic association discovery.
DOI
10.1038/s41588-024-01898-1

MixEHR-SAGE

AI GWAS Topic Modeling PheWAS EHR Phenotyping UK Biobank Brief Bioinform
PUBMED_LINK
41627341
FULL NAME
MixEHR-SAGE - Multi-modal Topic Modeling for PheWAS and GWAS
DESCRIPTION
MixEHR-SAGE is a PheCode-guided multi-modal topic model that integrates diagnoses, procedures, and medications from EHR to enhance phenotyping for GWAS. By combining expert-informed priors with probabilistic inference, it identifies over 1000 interpretable phenotype topics from UK Biobank data and improves disease incidence prediction and GWAS discovery. Published in Briefings in Bioinformatics.
TITLE
PheCode-guided multi-modal topic modeling of electronic health records improves disease incidence prediction and GWAS discovery from UK Biobank.
ABSTRACT
Phenome-wide association studies rely on disease definitions derived from diagnostic codes, often failing to leverage the full richness of electronic health records (EHR). We present MixEHR-SAGE, a PheCode-guided multi-modal topic model that integrates diagnoses, procedures, and medications to enhance phenotyping from large-scale EHRs. Applied to 350,000 individuals with high-quality genetic data, MixEHR-SAGE-derived risk scores accurately predicted disease incidence and improved GWAS discovery.
DOI
10.1093/bib/bbag030

Ning C (DL LVRWT GWAS)

AI GWAS Imaging Deep Learning Cardiac MRI Left Ventricular Wall Hypertrophic Cardiomyopathy UK Biobank
PUBMED_LINK
38036550
FULL NAME
Genome-Wide Association Analysis of Left Ventricular Imaging-Derived Phenotypes Identifies 72 Risk Loci
DESCRIPTION
Built a CNN-based deep learning algorithm for automated segmentation of left ventricular myocardium from cardiac MRI, enabling precise calculation of 12 regional wall thickness (LVRWT) measurements in 42,194 UK Biobank participants. GWAS of these 12 CNN-derived LVRWT traits identified 72 significant genetic loci involved in heart development and contraction pathways. Mendelian randomization confirmed causal relationships with hypertrophic cardiomyopathy. The PRS of inferoseptal LVRWT enabled identification of high-risk individuals.
KEYWORDS
deep learning, cardiac MRI, CNN, left ventricular wall thickness segmentation, hypertrophic cardiomyopathy, UK Biobank
TITLE
Genome-wide association analysis of left ventricular imaging-derived phenotypes identifies 72 risk loci and yields genetic insights into hypertrophic cardiomyopathy.
Main citation
Ning C, Fan L, Jin M, ...&, Miao X. (2023) Genome-wide association analysis of left ventricular imaging-derived phenotypes identifies 72 risk loci and yields genetic insights into hypertrophic cardiomyopathy. Nat Commun, 14 (1) 7900. doi:10.1038/s41467-023-43771-5. PMID 38036550
ABSTRACT
Left ventricular regional wall thickness (LVRWT) is an independent predictor of morbidity and mortality in cardiovascular diseases (CVDs). To identify specific genetic influences on individual LVRWT, we established a novel deep learning algorithm to calculate 12 LVRWTs accurately in 42,194 individuals from the UK Biobank with cardiac magnetic resonance (CMR) imaging. Genome-wide association studies of CMR-derived 12 LVRWTs identified 72 significant genetic loci associated with at least one LVRWT phenotype.
DOI
10.1038/s41467-023-43771-5

PoPS

AI GWAS Gene Prioritization Machine Learning Polygenic Nat Genet
PUBMED_LINK
37443254
FULL NAME
PoPS - Polygenic Priority Score for Gene Prioritization
DESCRIPTION
PoPS (Polygenic Priority Score) is a method that learns trait-relevant gene features, such as cell-type-specific expression, to prioritize genes at GWAS loci. It leverages polygenic enrichments across multiple gene features to predict causal genes underlying complex traits and diseases. Published in Nature Genetics.
URL
https://github.com/FinucaneLab/pops
TITLE
Leveraging polygenic enrichments of gene features to predict genes underlying complex traits and diseases.
ABSTRACT
Genome-wide association studies (GWASs) are a valuable tool for understanding the biology of complex human traits and diseases, but associated variants rarely point directly to causal genes. In the present study, we introduce a new method, polygenic priority score (PoPS), that learns trait-relevant gene features, such as cell-type-specific expression, to prioritize genes at GWAS loci. PoPS and the closest gene individually outperform other gene prioritization methods.
DOI
10.1038/s41588-023-01443-6

Quickdraws

AI GWAS Variational Inference Mixed Model GPU Nat Genet
PUBMED_LINK
39789286
FULL NAME
Quickdraws - Scalable Variational Inference for Mixed-Model GWAS
DESCRIPTION
Quickdraws is a method that increases association power in quantitative and binary traits for GWAS without sacrificing computational efficiency, leveraging a spike-and-slab prior on variant effects, stochastic variational inference, and graphics processing unit acceleration. Published in Nature Genetics.
TITLE
A scalable variational inference approach for increased mixed-model association power.
ABSTRACT
The rapid growth of modern biobanks is creating new opportunities for large-scale genome-wide association studies (GWASs) and the analysis of complex traits. However, performing GWASs on millions of samples often leads to trade-offs between computational efficiency and statistical power, reducing the benefits of large-scale data collection efforts. We developed Quickdraws, a method that increases association power in quantitative and binary traits without sacrificing computational efficiency, leveraging a spike-and-slab prior on variant effects, stochastic variational inference and graphics processing unit acceleration.
DOI
10.1038/s41588-024-02044-7

SynSurr

AI GWAS Machine Learning Phenotype Imputation Synthetic Surrogates Nat Genet
PUBMED_LINK
38872030
FULL NAME
SynSurr - Synthetic Surrogates for GWAS of Missing Phenotypes
DESCRIPTION
SynSurr (Synthetic Surrogate analysis) is a method that makes GWAS on imputed phenotypes robust to imputation errors. Rather than replacing missing values, SynSurr jointly analyzes the observed and imputed data to provide calibrated association statistics, improving power for genome-wide association studies of partially missing phenotypes in population biobanks. Published in Nature Genetics.
TITLE
Synthetic surrogates improve power for genome-wide association studies of partially missing phenotypes in population biobanks.
ABSTRACT
Within population biobanks, incomplete measurement of certain traits limits the power for genetic discovery. Machine learning is increasingly used to impute the missing values from the available data. However, performing GWAS on imputed traits can introduce spurious associations. Here we introduce SynSurr analysis, which makes GWAS on imputed phenotypes robust to imputation errors by jointly analyzing observed and imputed data.
DOI
10.1038/s41588-024-01793-9

transferGWAS

AI GWAS Imaging Transfer Learning Deep Learning Retinal Fundus Representation Learning
PUBMED_LINK
35640976
FULL NAME
transferGWAS: GWAS of Images Using Deep Transfer Learning
DESCRIPTION
transferGWAS performs GWAS directly on full medical images using deep transfer learning: (1) a pretrained CNN (ResNet-based architecture, pretrained on ImageNet) extracts feature embeddings from raw images; (2) these learned representations are used as quantitative phenotypes for genetic association testing. Applied to UK Biobank retinal fundus images, identified 60 genomic regions including 7 novel candidate loci for eye-related traits. First demonstration of direct GWAS on whole images without predefined phenotype engineering.
URL
https://github.com/mkirchler/transferGWAS/
KEYWORDS
deep transfer learning, pretrained CNN, ResNet, retinal fundus, whole-image GWAS, representation learning, UK Biobank
TITLE
transferGWAS: GWAS of images using deep transfer learning.
Main citation
Kirchler M, Konigorski S, Norden M, Meltendorf C, Kloft M, Schurmann C, Lippert C. (2022) transferGWAS: GWAS of images using deep transfer learning. Bioinformatics, 38(14):3621-3628. doi:10.1093/bioinformatics/btac369. PMID 35640976
ABSTRACT
MOTIVATION: Medical images can provide rich information about diseases and their biology. However, investigating their association with genetic variation requires non-standard methods. We propose transferGWAS, a novel approach to perform genome-wide association studies directly on full medical images. First, we learn semantically meaningful representations of the images based on a transfer learning task, during which a deep neural network is trained on independent but similar data. Then, we perform genetic association tests with these representations. RESULTS: We validate the type I error rates and power of transferGWAS in simulation studies of synthetic images. Then we apply transferGWAS in a genome-wide association study of retinal fundus images from the UK Biobank. This first-of-a-kind GWAS of full imaging data yielded 60 genomic regions associated with retinal fundus images, of which 7 are novel candidate loci for eye-related traits and diseases.
DOI
10.1093/bioinformatics/btac369

IMRP-GxE

GWAS G×E Mendelian randomization Summary statistics Tool
PUBMED_LINK
38649715
FULL NAME
Mendelian randomization-based genome-wide screening for gene–environment interactions
DESCRIPTION
Screens for combined gene–environment interaction and environmental mediation by testing departure of marginal GWAS effects from GWIS main effects using an MR-style statistic (IMRP), applicable to summary statistics from separate GWAS and interaction meta-analyses.
URL
https://www.nature.com/articles/s41467-024-47806-3 ,https://github.com/XiaofengZhuCase/IMRP23 ,https://zenodo.org/records/10815731
KEYWORDS
GWAS, GWIS, G×E, Mendelian randomization, IMRP, summary statistics
TITLE
An approach to identify gene-environment interactions and reveal new biological insight in complex traits.
Main citation
Zhu X, Yang Y, Lorincz-Comi N, Li G, Bentley AR, de Vries PS, ...&, Aschard H. (2024) An approach to identify gene-environment interactions and reveal new biological insight in complex traits. Nat Commun, 15 (1) 3385. doi:10.1038/s41467-024-47806-3. PMID 38649715
ABSTRACT
There is a long-standing debate about the magnitude of the contribution of gene-environment interactions to phenotypic variations of complex traits owing to the low statistical power and few reported interactions to date. To address this issue, the Gene-Lifestyle Interactions Working Group within the Cohorts for Heart and Aging Research in Genetic Epidemiology Consortium has been spearheading efforts to investigate G×E in large and diverse samples through meta-analysis. Here, we present a powerful new approach to screen for interactions across the genome, an approach that shares substantial similarity to the Mendelian randomization framework. We identify and confirm 5 loci (6 independent signals) interacted with either cigarette smoking or alcohol consumption for serum lipids, and empirically demonstrate that interaction and mediation are the major contributors to genetic effect size heterogeneity across populations. The estimated lower bound of the interaction and environmentally mediated heritability is significant (P < 0.02) for low-density lipoprotein cholesterol and triglycerides in Cross-Population data. Our study improves the understanding of the genetic architecture and environmental contributions to complex traits.
DOI
10.1038/s41467-024-47806-3
ARROW_SUMMARY
Inputs: GWAS + GWIS summary stats (per-SNP β/SE; LD-pruned instruments; sample-overlap ρ if needed) → IMRP θ → T_MR-GxE (G×E + mediation)

PP-GWAS

GWAS Privacy-preserving GWAS Tool Summary statistics
PUBMED_LINK
41365878
DESCRIPTION
Privacy-preserving framework for multi-site GWAS on quantitative traits using a distributed linear mixed model and randomized encoding so servers never see raw genotypes or phenotypes—only obfuscated intermediates—while improving speed versus several cryptographic baselines.
URL
https://github.com/mdppml/PP-GWAS ,https://doi.org/10.1038/s41467-025-66771-z
KEYWORDS
Privacy-preserving GWAS, multi-site, quantitative traits, federated analysis
TITLE
PP-GWAS: Privacy Preserving Multi-Site Genome-wide Association Studies.
Main citation
Swaminathan A, Hannemann A, Ünal AB, Pfeifer N, ...&, Akgün M. (2025) PP-GWAS: Privacy Preserving Multi-Site Genome-wide Association Studies. Nat Commun, 16 (1) 11030. doi:10.1038/s41467-025-66771-z. PMID 41365878
ABSTRACT
Genome-wide association studies help uncover genetic influences on complex traits and diseases. Importantly, multi-site data collaborations enhance the statistical power of these studies but pose challenges due to the sensitivity of genomic data. Existing privacy-preserving approaches to performing multi-site genome-wide association studies rely on computationally expensive cryptographic techniques, which limit applicability. To address this, we present PP-GWAS, a privacy-preserving algorithm that improves efficiency and scalability while maintaining data privacy. Our method leverages randomized encoding within a distributed framework to perform stacked ridge regression on a linear mixed model, enabling robust analysis of quantitative phenotypes. We show experimentally using real-world and synthetic data that our approach achieves twice the computational speed of comparable methods while reducing resource consumption.
DOI
10.1038/s41467-025-66771-z

REGENIE

GWAS
PUBMED_LINK
34017140
DESCRIPTION
regenie is a C++ program for whole genome regression modelling of large genome-wide association studies. It is developed and supported by a team of scientists at the Regeneron Genetics Center.
URL
https://github.com/rgcgithub/regenie
KEYWORDS
whole genome regression
TITLE
Computationally efficient whole-genome regression for quantitative and binary traits.
Main citation
Mbatchou J, Barnard L, Backman J, Marcketta A, ...&, Marchini J. (2021) Computationally efficient whole-genome regression for quantitative and binary traits. Nat Genet, 53 (7) 1097-1103. doi:10.1038/s41588-021-00870-7. PMID 34017140
ABSTRACT
Genome-wide association analysis of cohorts with thousands of phenotypes is computationally expensive, particularly when accounting for sample relatedness or population structure. Here we present a novel machine-learning method called REGENIE for fitting a whole-genome regression model for quantitative and binary phenotypes that is substantially faster than alternatives in multi-trait analyses while maintaining statistical efficiency. The method naturally accommodates parallel analysis of multiple phenotypes and requires only local segments of the genotype matrix to be loaded in memory, in contrast to existing alternatives, which must load genome-wide matrices into memory. This results in substantial savings in compute time and memory usage. We introduce a fast, approximate Firth logistic regression test for unbalanced case-control phenotypes. The method is ideally suited to take advantage of distributed computing frameworks. We demonstrate the accuracy and computational benefits of this approach using the UK Biobank dataset with up to 407,746 individuals.
DOI
10.1038/s41588-021-00870-7

seismic

GWAS Single cell scRNA-seq Gene prioritization Tool
PUBMED_LINK
41034207
FULL NAME
Single-cell Expression Integration System for Mapping genetically Implicated Cell types
DESCRIPTION
R framework that links GWAS signals to single-cell-defined cell types via a cell-type gene specificity score (expression magnitude and consistency) and regression on gene-level association statistics, with influential-gene follow-up for interpretability.
URL
https://github.com/ylaboratory/seismic ,https://ylaboratory.github.io/seismic/ ,https://doi.org/10.1038/s41467-025-63753-z
KEYWORDS
GWAS, scRNA-seq, cell type, MAGMA, post-GWAS interpretation
TITLE
Disentangling associations between complex traits and cell types with seismic.
Main citation
Lai Q, Dannenfelser R, Roussarie JP, Yao V. (2025) Disentangling associations between complex traits and cell types with seismic. Nat Commun, 16 (1) 8744. doi:10.1038/s41467-025-63753-z. PMID 41034207
ABSTRACT
Integrating single-cell RNA sequencing with Genome-Wide Association Studies (GWAS) can uncover cell types involved in complex traits and disease. However, current methods often lack scalability, interpretability, and robustness. We present seismic, a framework that computes a novel specificity score capturing both expression magnitude and consistency across cell types and introduces influential gene analysis, an approach to identify genes driving each cell type-trait association. Across over 1000 cell-type characterizations at different granularities and 28 polygenic traits, seismic corroborates known associations and uncovers trait-relevant cell groups not apparent through other methodologies. In Parkinson's and Alzheimer's, seismic unveils both cell- and brain-region-specific differences in pathology. Analyzing a pathology-based Alzheimer's GWAS with seismic enables the identification of vulnerable neuron populations and molecular pathways implicated in their neurodegeneration. In general, seismic is a computationally efficient, powerful, and interpretable approach for mapping the relationships between polygenic traits and cell-type-specific expression, offering new insights into disease mechanisms.
DOI
10.1038/s41467-025-63753-z

BBJ — Cross-population atlas of 220 phenotypes

GWAS Cross-population Trans-ancestry
PUBMED_LINK
34594039
STAGE_PERIOD
2021
DESCRIPTION
Cross-population GWAS atlas integrating BBJ, UK Biobank, and FinnGen for 220 phenotypes. Identified thousands of novel loci through trans-ancestry meta-analysis, improving fine-mapping resolution with diverse ancestry. BBJ was critical for East Asian representation.
URL
https://biobankjp.org/
TITLE
A cross-population atlas of genetic associations for 220 human phenotypes

BBJ — Quantitative traits GWAS

GWAS Quantitative Traits Cross-population
PUBMED_LINK
29403010
STAGE_PERIOD
2018
DESCRIPTION
Large-scale GWAS of 58 quantitative traits in up to 148,000 Japanese individuals. Identified 585 trait-associated loci, 118 of which were novel. Demonstrated cell-type-specific gene expression mediation and genetic correlations between Japanese and European populations.
URL
https://biobankjp.org/
TITLE
Genetic analysis of quantitative traits in the Japanese population links cell types to complex human diseases

CKB — Lung cancer GWAS & polygenic risk score

GWAS PRS Lung Cancer
PUBMED_LINK
31326317
STAGE_PERIOD
2019
DESCRIPTION
Large-scale GWAS of lung cancer in Chinese populations identifying novel risk loci, plus development of a polygenic risk score for lung cancer risk stratification in the CKB prospective cohort.
URL
https://www.ckbiobank.org/
TITLE
Identification of risk loci and a polygenic risk score for lung cancer: a large-scale prospective cohort study in Chinese

TOPMed — Flagship publications & cross-trait analyses

GWAS Rare Variant Fine-mapping Blood Pressure Lipids
STAGE_PERIOD
2020–2022
DESCRIPTION
Series of landmark papers from TOPMed studies spanning blood pressure, lipid, pulmonary function, and glycemic trait GWAS leveraging the unique WGS data for rare variant association testing. Demonstrated value of deep WGS in diverse populations for fine-mapping and identifying causal variants. Multi-ancestry meta-analyses improved fine-mapping resolution at loci with divergent LD patterns.
URL
https://topmed.nhlbi.nih.gov/

AD gut microbiome host genetics (Chinese cohort)

GWAS
PUBMED_LINK
41782023
DESCRIPTION
Joint host whole-genome sequencing and gut microbiome profiling in 252 Chinese individuals with graded cognitive disability. Microbiome GWAS for latent enterosignature (Anaerostipes-enriched) abundance; integrates AD polygenic risk and brain cell-type expression context. Open-access report in Microbiome.
URL
https://link.springer.com/article/10.1186/s40168-026-02342-8
Main citation
Liu J, Cao J, Jia L, et al. (2026) Impacts of host genetics on gut microbiome composition in Alzheimer's disease. Microbiome. doi:10.1186/s40168-026-02342-8. PMID 41782023
MAIN ANCESTRY
EAS
METAGENOME
Gut bacteria

Cole JB-32193382

GWAS
PUBMED_LINK
32193382
DESCRIPTION
GWAS of food-frequency questionnaire traits in UK Biobank; hundreds of associated loci. Summary statistics for ~143 heritable dietary measures (BOLT-LMM); large tarball via Knowledge Portal.
URL
https://www.nature.com/articles/s41467-020-15193-0 ,https://www.kp4cd.org/node/351
Main citation
Cole JB, Florez JC, Hirschhorn JN. (2020) Comprehensive genomic analysis of dietary habits in UK Biobank identifies hundreds of genetic associations. Nat Commun, 11 (1) 1467. doi:10.1038/s41467-020-15193-0. PMID 32193382
DOI
10.1038/s41467-020-15193-0
MAIN ANCESTRY
EUR

Doherty A-30531941

GWAS
PUBMED_LINK
30531941
DESCRIPTION
GWAS of UK Biobank wrist accelerometer phenotypes (91,105 participants); 14 loci. Summary statistics (unadjusted and sex/BMI-adjusted) deposited in Oxford University Research Archive.
URL
https://www.nature.com/articles/s41467-018-07743-4 ,https://ora.ox.ac.uk/objects/uuid:ff479f44-bf35-48b9-9e67-e690a2937b22
Main citation
Doherty A, Smith-Byrne K, Ferreira T, et al. (2018) GWAS identifies 14 loci for device-measured physical activity and sleep duration. Nat Commun, 9 (1) 5257. doi:10.1038/s41467-018-07743-4. PMID 30531941
DOI
10.1038/s41467-018-07743-4
MAIN ANCESTRY
EUR

German gut microbiome GWAS (ABO / FUT2)

GWAS
PUBMED_LINK
33462482
DESCRIPTION
Genome-wide association analysis in 8,956 German individuals links host variation to single-taxon and overall gut microbiome composition; replicates ABO histo-blood group and FUT2 secretor associations with Bacteroides and Faecalibacterium. Includes Mendelian randomization for IBD-relevant microbial effects.
URL
https://www.nature.com/articles/s41588-020-00747-1
Main citation
Rühlemann MC, Hermes BM, Bang C, et al. (2021) Genome-wide association study in 8,956 German individuals identifies influence of ABO histo-blood groups on gut microbiome. Nat Genet, 53:147–155. doi:10.1038/s41588-020-00747-1. PMID 33462482
MAIN ANCESTRY
EUR
METAGENOME
Gut bacteria

Gut microbial structural variation GWAS (Dutch meta-analysis)

GWAS
PUBMED_LINK
38172637
DESCRIPTION
Meta-analysis of genome-wide associations between host genotypes and gut microbial structural variants (deletion dSVs and variable vSVs) from metagenomes in four Dutch cohorts (n = 9,015), with replication in Tanzania (n = 279). After frequency filters, GWAS used 3,552 common SV phenotypes aggregated across 49 bacterial species with sufficient metagenomic coverage. Highlights ABO/FUT2–GalNAc pathways and Faecalibacterium prausnitzii SVs.
URL
https://www.nature.com/articles/s41586-023-06893-w
Main citation
Zhernakova DV, Wang D, Liu L, et al. (2024) Host genetic regulation of human gut microbial structural variation. Nature, 625:813–821. doi:10.1038/s41586-023-06893-w. PMID 38172637
MAIN ANCESTRY
EUR
METAGENOME
Gut bacteria (metagenomic SVs)

Host genetics & gut microbiome (Nat Genet perspective)

GWAS
PUBMED_LINK
35115688
DESCRIPTION
Perspective on microbial GWAS (mbGWAS): state of the art, heterogeneity of microbiome assays, power, and directions for genetic analysis of the gut microbiome.
URL
https://www.nature.com/articles/s41588-021-00983-z
Main citation
Sanna S, Kurilshikov A, van der Graaf A, et al. (2022) Challenges and future directions for studying effects of host genetics on the gut microbiome. Nat Genet, 54:100–106. doi:10.1038/s41588-021-00983-z. PMID 35115688

Japanese gut microbiome–host genetics (Cell Reports)

GWAS
PUBMED_LINK
37935197
DESCRIPTION
Japanese shotgun metagenome and SNP-array GWAS (7,213,469 post-imputation variants, MAF >1%, Minimac4 Rsq >0.7): microbial traits included species, gene orthologs, and pathways. Two study waves—dataset 1 (gut microbiome, plasma metabolome, genotype): n = 300; dataset 2 (gut microbiome, genotype): n = 224—were combined by fixed-effect meta-analysis (total n = 524; Figure S1 / Table S1). Species GWAS used 423 microbial species (study-wide threshold 5×10⁻⁸/423); headline hits include PDE1C–Bacteroides intestinalis and TGIF2 / TGIF2-RAB5IF–B. acidifaciens, plus microbial gene ortholog associations with blood group A conditioned on East Asian FUT2 secretor status. Metabolome arm in dataset 1 supports microbiome–plasma integration. Public data resource.
URL
https://www.cell.com/cell-reports/fulltext/S2211-1247%2823%2901336-0
Main citation
Tomofuji Y, Kishikawa T, Sonehara K, et al. (2023) Analysis of gut microbiome, host genetics, and plasma metabolites reveals gut microbiome-host interactions in the Japanese population. Cell Rep, 42(11):113324. doi:10.1016/j.celrep.2023.113324. PMID 37935197
MAIN ANCESTRY
EAS
METAGENOME
Gut (shotgun metagenome)

Matoba N-31959922

GWAS
PUBMED_LINK
31959922
DESCRIPTION
Genome-wide association studies for 13 dietary habits in Japanese individuals (n = 58,610–165,084) from BioBank Japan; nine genome-wide significant loci reported.
URL
https://www.nature.com/articles/s41562-019-0805-1
Main citation
Matoba N, Akiyama M, Ishigaki K, et al. (2020) GWAS of 165,084 Japanese individuals identified nine loci associated with dietary habits. Nat Hum Behav, 4 (3) 308-316. doi:10.1038/s41562-019-0805-1. PMID 31959922
DOI
10.1038/s41562-019-0805-1
MAIN ANCESTRY
EAS

Merino J-34426670

GWAS
PUBMED_LINK
34426670
DESCRIPTION
Multi-trait GWAS meta-analysis of dietary intake combining UK Biobank and CHARGE; functional and brain-related follow-up. Summary statistics via UK Biobank showcase, dbGaP, and Type 2 Diabetes Knowledge Portal (see paper data availability).
URL
https://www.nature.com/articles/s41562-021-01182-w ,https://www.kp4cd.org/dataset_downloads/t2d
Main citation
Merino J, Dashti HS, Sarnowski C, et al. (2022) Genetic analysis of dietary intake identifies new loci and functional links with metabolic traits. Nat Hum Behav, 6 (1) 155-163. doi:10.1038/s41562-021-01182-w. PMID 34426670
DOI
10.1038/s41562-021-01182-w
MAIN ANCESTRY
EUR

MiBioGen gut microbiome GWAS (multi-cohort)

GWAS
PUBMED_LINK
33462485
DESCRIPTION
MiBioGen consortium meta-analysis: genome-wide host genotypes with 16S fecal microbiome data across 24 cohorts (n = 18,340). N_MICROBES = 410 genus-level groups in the MiBioGen framework (paper: nine genera detected in >95% of samples). Identifies 31 genome-wide significant loci affecting microbial taxa (e.g. lactase LCT, ABO, fucosyltransferase cluster).
URL
https://pmc.ncbi.nlm.nih.gov/articles/PMC8515199/ ,https://www.nature.com/articles/s41588-020-00763-1
Main citation
Kurilshikov A, Medina-Gomez C, Bacigalupe R, et al. (2021) Large-scale association analyses identify host factors influencing human gut microbiome composition. Nat Genet, 53:156–165. doi:10.1038/s41588-020-00763-1. PMID 33462485
MAIN ANCESTRY
Multi-ancestry
METAGENOME
Gut bacteria (16S)

Namba S-41606330

GWAS
PUBMED_LINK
41606330
DESCRIPTION
Cross-population atlas of gene–environment interactions (discovery n ≈ 440k EUR + Japanese; replication n ≈ 540k). Genome-wide G×E summary statistics released at NBDC Human Database (hum0197.v26.374-traits.v1) and GWAS Catalog (GCST90681837–GCST90690020).
URL
https://www.nature.com/articles/s41586-025-10054-6 ,https://humandbs.dbcls.jp/en ,https://www.ebi.ac.uk/gwas
Main citation
Namba S, Sonehara K, Koyanagi YN, et al. (2026) A cross-population compendium of gene-environment interactions. Nature, 651:688-697. doi:10.1038/s41586-025-10054-6. PMID 41606330
DOI
10.1038/s41586-025-10054-6
MAIN ANCESTRY
Multi-ancestry

Oral microbiota GWAS (ADDITION-PRO)

GWAS
PUBMED_LINK
38926497
DESCRIPTION
16S rRNA amplicon-based GWAS of salivary microbiota traits in unrelated Danish adults from the ADDITION-PRO cohort (n = 610). Identifies host SNPs associated with oral bacterial abundance and beta diversity; several variants link to metabolic traits. Oral (not gut) microbiota — listed here for host–microbiome genetics.
URL
https://www.nature.com/articles/s41598-024-65538-8
Main citation
Stankevic E, Kern T, Borisevich D, et al. (2024) Genome-wide association study identifies host genetic variants influencing oral microbiota diversity and metabolic health. Sci Rep, 14:14738. doi:10.1038/s41598-024-65538-8. PMID 38926497
MAIN ANCESTRY
EUR
METAGENOME
Oral (salivary)

Swedish gut metagenome GWAS (SCAPIS / HUNT replication)

GWAS
PUBMED_LINK
41688638
DESCRIPTION
Harmonized shotgun metagenome GWAS in 16,017 adults from four Swedish studies, with replication in 12,652 participants from the Norwegian HUNT study. Reports loci including OR51E1–OR51E2 (microbial richness), LCT, ABO, FUT2, MUC12, CORO7–HMOX2, SLC5A11, FOXP1, FUT3–FUT6, and species-level associations.
URL
https://www.nature.com/articles/s41588-026-02512-2
Main citation
Dekkers KF, Pertiwi K, Baldanzi G, et al. (2026) Genome-wide association analyses highlight the role of the intestinal molecular environment in human gut microbiota variation. Nat Genet, 58:540–549. doi:10.1038/s41588-026-02512-2. PMID 41688638
MAIN ANCESTRY
EUR
METAGENOME
Gut (shotgun metagenome)

UK Biobank family-based GWAS (Guan et al.)

GWAS UK Biobank
PUBMED_LINK
40065166
DESCRIPTION
Family-based GWAS (FGWAS) summary statistics estimating direct genetic effects (DGEs) with unified, robust, Young et al. Mendelian-imputation, and sib-difference estimators implemented in snipar. Unified estimator adds singletons via linear parental imputation (largest DGE effective sample size in homogeneous samples); robust estimator avoids allele-frequency imputation bias under strong structure or admixture. UK Biobank application (n up to ~408k unified White British; ~52k robust with ≥1 genotyped first-degree relative).
URL
https://www.nature.com/articles/s41588-025-02118-0 ,https://thessgac.com/ ,https://github.com/AlexTISYoung/snipar ,https://doi.org/10.5281/zenodo.14270274
Main citation
Guan J, Tan T, Nehzati SM, Bennett M, ...&, Young AS. (2025) Family-based genome-wide association study designs for increased power and robustness. Nat Genet, 57 (4) 1044-1052. doi:10.1038/s41588-025-02118-0. PMID 40065166
DOI
10.1038/s41588-025-02118-0
RELATED_BIOBANK
UK Biobank
MAIN ANCESTRY
EUR (unified / Young et al. in White British); multi-ancestry robust estimator

UK Biobank sleep traits (self-report)

GWAS
PUBMED_LINK
30846698
DESCRIPTION
Self-reported sleep-related GWAS summary statistics hosted on the Knowledge Portal (downloads per trait with READMEs). Landmark papers include Dashti et al. on sleep duration and daytime napping, Wang et al. on daytime sleepiness, and Lane et al. on chronotype and sleep disturbance (see portal publication list).
URL
https://www.kp4cd.org/node/235 ,https://www.nature.com/articles/s41467-019-08917-4
Main citation
Dashti HS, Jones SE, Wood AR, et al. (2019) Genome-wide association study identifies genetic loci for self-reported habitual sleep duration supported by accelerometer-derived estimates. Nat Commun, 10 (1) 1100. doi:10.1038/s41467-019-08917-4. PMID 30846698
DOI
10.1038/s41467-019-08917-4
MAIN ANCESTRY
EUR

Wang Z-36071172

GWAS
PUBMED_LINK
36071172
DESCRIPTION
Genome-wide association meta-analysis of physical activity and sedentary behaviour traits across ancestries; 99 associated loci. Supplementary tables and open materials via Nature Genetics and WashU Digital Commons.
URL
https://www.nature.com/articles/s41588-022-01165-1 ,https://digitalcommons.wustl.edu/oa_4/339/
Main citation
Wang Z, Emmerich A, Pillon NJ, et al. (2022) Genome-wide association analyses of physical activity and sedentary behavior provide insights into underlying mechanisms and roles in disease prevention. Nat Genet, 54 (9) 1332-1344. doi:10.1038/s41588-022-01165-1. PMID 36071172
DOI
10.1038/s41588-022-01165-1
MAIN ANCESTRY
Multi-ancestry