Multi-omics
Catalog entries using this tag (links open the entry card on its page):
- MILTON — AI
- scGPT — AI
- OmiGA — GWAS Tools
- Host transcriptome–gut microbiome integration (CRC / IBD / IBS) — Summary statistics
Entries
MILTON
PUBMED_LINK
FULL NAME
MILTON - Machine Learning with Phenotype Associations for Disease Prediction
DESCRIPTION
MILTON is an ensemble machine learning framework that utilizes biomarkers and multi-omics data to predict 3,213 diseases in the UK Biobank. It predicts incident disease cases undiagnosed at time of recruitment and demonstrates utility in augmenting genetic association discovery by empowering case-control GWAS with predicted phenotypes. Published in Nature Genetics.
TITLE
Disease prediction with multi-omics and biomarkers empowers case-control genetic discoveries in the UK Biobank.
ABSTRACT
The emergence of biobank-level datasets offers new opportunities to discover novel biomarkers and develop predictive algorithms for human disease. Here, we present an ensemble machine-learning framework (machine learning with phenotype associations, MILTON) utilizing a range of biomarkers to predict 3,213 diseases in the UK Biobank. MILTON predicts incident disease cases undiagnosed at time of recruitment, largely outperforming available polygenic risk scores, and augments genetic association discovery.
DOI
10.1038/s41588-024-01898-1
scGPT
PUBMED_LINK
FULL NAME
scGPT — Foundation Model for Single-Cell Multi-Omics Using Generative AI
DESCRIPTION
scGPT is a generative pretrained transformer foundation model for single-cell biology, pretrained on over 33 million human cells from 51 organs across 441 studies. Uses a GPT architecture adapted for gene expression data with a specialized attention mask. Outperforms traditional methods on cell type annotation, multi-batch integration, multi-omic integration, perturbation response prediction, and gene network inference. Represents a foundational AI model for cellular biology analogous to GPT for natural language.
URL
TITLE
scGPT: toward building a foundation model for single-cell multi-omics using generative AI.
Main citation
Cui H, Wang C, Maan H, Pang K, Luo F, Duan N, Wang B. (2024) scGPT: toward building a foundation model for single-cell multi-omics using generative AI. Nature Methods, 21(8):1470-1480. doi:10.1038/s41592-024-02201-0. PMID 38840054
ABSTRACT
Generative pretrained models have achieved remarkable success in various domains such as language and computer vision. Using burgeoning single-cell sequencing data, we have constructed a foundation model for single-cell biology, scGPT, based on a generative pretrained transformer across a repository of over 33 million cells. Our findings illustrate that scGPT effectively distills critical biological insights concerning genes and cells. Through further adaptation of transfer learning, scGPT can be optimized to achieve superior performance across diverse downstream applications including cell type annotation, multi-batch integration, multi-omic integration, perturbation response prediction and gene network inference.
DOI
10.1038/s41592-024-02201-0
OmiGA
PUBMED_LINK
DESCRIPTION
Toolkit for molecular QTL (molQTL) mapping using linear mixed models that handle complex relatedness, aimed at high-throughput omics phenotypes with strong performance for discovery, fine mapping, and trait–molQTL colocalization versus common linear-mapper pipelines.
URL
KEYWORDS
molQTL, xQTL, LMM, relatedness, colocalization, fine mapping
TITLE
OmiGA for ultra-efficient molecular quantitative trait loci mapping.
Main citation
Teng J, Zhang W, Gong W, Chen J, ...&, Zhang Z. (2026) OmiGA for ultra-efficient molecular quantitative trait loci mapping. Nat Commun, 17 (1) . doi:10.1038/s41467-026-68978-0. PMID 41680153
ABSTRACT
Molecular quantitative trait loci (molQTL) mapping is one of the most popular approaches to systematically characterize functional impacts of genomic variants, leading to advanced understanding of the regulatory mechanisms underpinning complex traits and diseases. However, when applied to high-throughput molecular phenotypes, the existing molQTL mapping tools often implement simple linear models, overlooking complex inter-individual relatedness, leading to false positives and insufficient statistical power. Here, we introduce OmiGA, an ultra-efficient omics genetic analysis toolkit, for molQTL mapping based on linear mixed model in populations with complex relatedness. Both computational simulations and real data analyses demonstrate that OmiGA outperforms the existing popular tools regarding molQTL discovery power, fine mapping of causal variants, colocalization of molQTL and trait associations, and computational efficiency. In summary, we recommend OmiGA for molQTL mapping in populations with complex relatedness, for example, those in the Farm animal Genotype-Tissue Expression project and family-based molQTL studies in humans.
DOI
10.1038/s41467-026-68978-0
Host transcriptome–gut microbiome integration (CRC / IBD / IBS)
PUBMED_LINK
DESCRIPTION
Paired colonic mucosal RNA-seq and 16S gut microbiome data (208 sample pairs) across colorectal cancer, IBD, and IBS; machine-learning integration (sparse CCA, lasso) maps shared and disease-specific host gene–microbe associations—not SNP-based GWAS.
URL
Main citation
Priya S, Burns MB, Ward T, et al. (2022) Identification of shared and disease-specific host gene–microbiome associations across human diseases using multi-omic integration. Nat Microbiol, 7:780–795. doi:10.1038/s41564-022-01121-z. PMID 35577971
TRANSCRIPTOME
Colonic mucosa RNA-seq
METAGENOME
Gut mucosa (16S)