PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “multi-omic imputation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

9 recordsLinked to original sources

OmicsPred as a centralised resource for genetic prediction of multi-omic traits.

Genetic prediction of multi-omic data has emerged as a cost-effective alternative to direct omics profiling, particularly useful for identifying molecular features associated with disease susceptibility. However, despite its popularity, multi-omic imputation models are fragmented across studies, hindering findability, accessibility, interoperability and re-use. To address this, we developed OmicsPred (https://www.omicspred.org), a centralised platform for the deposition and dissemination of genetic prediction models of multi-omic traits. OmicsPred unifies the most commonly used molecular imputation models (e.g. from PredictDB) and other published studies totalling 3,339,469 prediction models spanning transcriptomic, proteomic, and metabolomic traits (as of May 2026). Each model is accompanied by metadata describing score development and predictive performance, and distributed in formats compatible with popular analytic tools, such as PGS Catalog Calculator and MetaXcan. To demonstrate the utility of the resource for systematic target discovery, we perform a multi-omic phenome-wide association analysis in Million Veterans Program data.

Journal Article↗

Penalized likelihood optimization for censored missing value imputation in proteomics.

Label-free bottom-up proteomics using mass spectrometry and liquid chromatography has long been established as one of the most popular high-throughput analysis workflows for proteome characterization. However, it produces data hindered by complex and heterogeneous missing values, which imputation has long remained problematic. To cope with this, we introduce Pirat, an algorithm that harnesses this challenge using an original likelihood maximization strategy. Notably, it models the instrument limit by learning a global censoring mechanism from the data available. Moreover, it estimates the covariance matrix between enzymatic cleavage products (ie peptides or precursor ions), while offering a natural way to integrate complementary transcriptomic information when multi-omic assays are available. Our benchmarking on several datasets covering a variety of experimental designs (number of samples, acquisition mode, missingness patterns, etc.) and using a variety of metrics (differential analysis ground truth or imputation errors) shows that Pirat outperforms all pre-existing imputation methods. Beyond the interest of Pirat as an imputation tool, these results pinpoint the need for a paradigm change in proteomics imputation, as most pre-existing strategies could be boosted by incorporating similar models to account for the instrument censorship or for the correlation structures, either grounded to the analytical pipeline or arising from a multi-omic approach.

Proteomics↗

Protocol to perform integrative analysis of high-dimensional single-cell multimodal data using an interpretable deep learning technique.

The advent of single-cell multi-omics sequencing technology makes it possible for researchers to leverage multiple modalities for individual cells. Here, we present a protocol to perform integrative analysis of high-dimensional single-cell multimodal data using an interpretable deep learning technique called moETM. We describe steps for data preprocessing, multi-omics integration, inclusion of prior pathway knowledge, and cross-omics imputation. As a demonstration, we used the single-cell multi-omics data collected from bone marrow mononuclear cells (GSE194122) as in our original study. For complete details on the use and execution of this protocol, please refer to Zhou et al.1.

Deep Learning↗

A pan-cancer multi-omic SuperLearner for regulated cell death survival topologies.

INTRODUCTION: Regulated cell death (RCD) pathways influence tumor progression and immune modulation. We previously constructed a signature database mapping 25 RCD forms across seven multi-omic layers and 33 tumor types (CancerRCDShiny). Despite their ability to identify risk populations, translating these signatures into personalized clinical workflows requires a shift from cohort stratification to individualized risk mapping by modeling patient risk (survival topologies) to capture the non-linear dynamics of RCD signatures. METHODS: We engineered a pan-cancer multi-omic SuperLearner pipeline across 33 cancer types. Phase I performed zero-leakage harmonization and groupwise imputation to prevent cross-cohort amalgamation. Phase II deployed Elastic Net-regularized Cox regression as a CANARY diagnostic to map proportional hazards failures. Strata with a 35% missingness barrier entered Phase III, deploying a Quadripartite ensemble: Random Survival Forests, XGBoost, Survival-Boruta, and Multi-Task Logistic Regression, fused within an Elastic Net Multi-View Meta-Learner (MVL), with post-hoc TreeSHAP and LIME interpretability. RESULTS: The CANARY diagnostic demonstrated the structural invalidity of pan-cancer geometric proportional hazards. Across 96 admissible strata, Phase III executed algorithmic displacement: continuous multi-omic topologies suppressed static genomic mutations and copy number variations (85.7% vs. 0.0% apex retention). The MVL stabilized predictions against extreme variance; LIME surrogate validations (R 2&#x202f;<&#x202f;0.10) confirmed the systematic failure of linear interpretative proxies. N-dimensional TreeSHAP interaction mapping exposed synergistic and antagonistic rescue trajectories defining individualized Survival Topologies, which were invisible to additive models. The architecture was deployed as CancerRCDPredictor, a digital molecular tumor board with integrated LLM capabilities. The MVL SuperLearner achieved a median C-index of 0.749 (IQR: 0.722-0.836) across 96 modelable strata, with 95% bootstrap confidence intervals confirming precision (median width: 0.052) and permutation significance in 93.8% of strata (p&#x202f;<&#x202f;0.001). External CPTAC validation across ten cancer types demonstrated significant cross-cohort generalizability in clear cell renal carcinoma (KIRC; C-index 0.675, p&#x202f;=&#x202f;0.017) and modest performance across the remaining adequately powered cancers (median 0.582), underscoring the need for larger multi-institutional validation cohorts. CONCLUSION: This pan-cancer multi-omic SuperLearner bypasses linear topological failures, advancing beyond generalized stratification to establish a deterministically mapped architecture for predicting RCD-related survival topologies. Through the CancerRCDPredictor interface, multi-omic insights translate into individualized survival topology exploration, providing a foundation for future precision oncology validation.

SuperLearner↗

Alterations in ether lipid metabolism in obesity revealed by systems genomics of multi-omics datasets.

Ratios between two metabolites are sensitive indicators of metabolic changes. Lipidomic profiling studies have revealed that plasma ether lipids, a class of glycero- and glycerophospho-lipids with reported health benefits, are negatively associated with obesity. Here, we utilized lipid ratios as surrogate markers of lipid metabolism to explore the processes underlying the inverse relationship between ether lipid metabolism and obesity. Plasma lipidomics data from two independent human cohorts (n&#x2009;=&#x2009;10,339 and n&#x2009;=&#x2009;4,492) were integrated to assess the associations between 82 lipid ratios and obesity-related markers in males and females. Results were externally validated using mouse transcriptomics data from the Hybrid Mouse Diversity Panel (n&#x2009;=&#x2009;152-227 across 74 strains). Genome-wide association studies using imputed genotypes from a population cohort (n&#x2009;=&#x2009;4,492) were performed to examine the genetic architecture of the ratios. Findings showed that waist circumference (WC), body mass index, and waist-hip ratio were inversely associated with total plasmalogens relative to total phospholipids in both sexes. Ratios comprising product-substrate pairs positioned either side of enzymes involved in plasmalogen synthesis and degradation showed positive and negative associations with WC, respectively. Branched-chain fatty acids negatively correlated with WC, while omega-6 polyunsaturated fatty acids exhibited differing associations depending on their position within the pathway. Mouse transcriptomics corroborated these results. Genomics data showed strong associations between ratios containing choline-plasmalogens and single-nucleotide polymorphisms in the transmembrane protein 229B (TMEM229B) gene region. This work demonstrates the utility of lipid ratios in understanding lipid metabolism. By applying the ratios to multi-omic datasets, we identified alterations in enzymatic activity and genetic variants likely affecting ether lipid synthesis in obesity that could not have been obtained from lipidomics data alone. Additionally, we characterized a potential role for TMEM229B, offering new perspectives on ether lipid metabolism and regulation.

Humans↗

PLNMFG: Pseudo-label guided non-negative matrix factorization model with graph constraint for single-cell multi-omics data clustering.

The development of single-cell multi-omics sequencing technologies has enabled the simultaneous analysis of multi-omics data within the same cell. Accurate clustering of these cells is crucial for downstream analyses of complex biological functions. Despite significant advances in multi-omics integration approaches, current methodologies exhibit two major limitations. First, they inadequately incorporate prior biological knowledge from various omic layers. Second, these methods often conduct independent dimensionality reduction on individual omic datasets, thereby failing to capture the intrinsic complementary information and potentially overlooking crucial cross-platform interactions. Motivated by these, this study investigates a non-negative matrix factorization model called PLNMFG, which integrates the unified latent representation learning that retains the features between and within omics and the cluster structure learning that retains the intrinsic structure of the data into one joint framework. Specially, PLNMFG performs adaptive imputation to handle dropout events and uses prior pseudo-labels as constraints during the process of collective non-negative matrix factorization, as a result, a more robust latent representation that preserves the double similarity information is obtained. Graph Laplacian constraint is applied during clustering which further preserves structure characteristic of multi-omics data. In addition, the weight of each omic is adaptively learned based on the omic contribution. A series of experiments on 8 benchmark datasets show that our model performs well in terms of clustering accuracy and computational efficiency.

Single-Cell Analysis↗

National genomic projects in Asia and Africa: a review.

National genome projects (NGPs) are increasingly shaping precision medicine by improving representation of population-specific genetic diversity. This review compiles findings from NGPs across Asia and Africa, regions that remain underrepresented in global genomic databases despite their extensive demographic and genetic diversity. A total of 53 studies from 24 countries were identified to understand (1) the genomic approach utilized, (2) novel findings that have emerged, and (3) strategies for improving research in these regions. The NGPs implement population-based variome databases (20 NGPs), linear reference genome assemblies (8 NGPs), and graph-based pangenome assemblies (1 NGP). Novel variants ranged between 0.28% (China) and 19.6% (Iran), whereas rare variants accounted for up to 88.9% of the detected variants in the Chinese population. Each NGP documents its country's evolutionary and migration history, which impacts disease frequency and pharmacogenomic variants. Clinically, NGPs revealed strong population stratification in disease-associated and pharmacogenomic variants. For example, the&#xa0;GJB2&#xa0;rs72474224 hearing-loss variant ranged from 13% in Vietnam and 12% in Hong Kong to 0.0894% in Turkey, while the&#xa0;VKORC1&#xa0;rs9923231 pharmacogenomic variant reached 89.2% in Taiwan but was 20%-25% in European-related Russian subpopulations. These findings demonstrate that clinically relevant allele frequencies, pathogenicity assessments, and drug-response markers differ substantially across ancestries. This review highlights ongoing efforts and strategies to enhance the representativeness of genomic data through NGPs in Asia and Africa. We also suggest future directions for national projects, including integrating family-based studies, multi-omic data, and standardized pipelines to accelerate discovery and support the equitable implementation of precision medicine.

Humans↗

Multi-omics analysis identifies key genes and functional loci affecting teat number in American Large White and Landrace pigs and their application in optimizing genomic selection models.

BACKGROUND: Teat number is a crucial economic trait in pigs. It directly affects the ability of sows to lactate, which in turn influences the survival and health of piglets. The teat number of French Large White pigs is close to 16, while the teat number of American Large White and Landrace pigs is about 14. In order to improve the teat number of American Landrace and Large White pigs through molecular approaches and precise breeding techniques, we genotyped 2,131 American Landrace and 4,564 American Large White with teat number phenotype using a 50&#xa0;K SNP chip. Then, the SNP-chip data was imputed to the level of whole-genome sequencing (iWGS). Based on iWGS data, we conducted GWAS to identify novel, significant SNPs associated with teat number and to incorporate them into genomic selection. RESULTS: In Landrace pigs, significant SNPs for TTN mapped to SSC2, SSC7, SSC8, and SSC14; the SSC8 and SSC14 effects are novel. LTN mapped to SSC7, RTN to SSC7 and SSC8. The lead SSC7 SNP explained 2.60% of TTN phenotypic variance. In Large White pigs, significant SNPs were detected on SSC7 and SSC10 for TTN; SSC7, SSC10, and SSC12 for LTN; and SSC7 and SSC10 for RTN. The most significant locus on SSC7 accounted for 2.99% of the phenotypic variance in TTN. Additionally, a multi-population meta-analysis detected significant novel SNPs for LTN on SSC1 and SSC8. By utilizing Bayesian fine mapping, the most precise QTL confidence interval on SSC7 for both TTN and RTN in Large White pigs was reduced to 40&#xa0;kb. By integrating functional gene annotation with RNA-seq and ATAC-seq data from Erhualian and Bamaxiang pigs mammary placodes at embryonic day 26, we prioritized PTPN13, TRPV3, ZDHHC13, and BRD2 as novel candidate genes for teat number. We then incorporated the significant SNPs to GBLUP and benchmarked genomic-selection accuracy. In both breeds, fitting the top SNP as fixed maximized prediction for TTN and RTN, whereas treating all significant loci as an additional random effect optimized LTN. CONCLUSIONS: Our findings provide a theoretical basis for dissecting new key genes affecting teat number and for advancing molecular breeding of teat number in pigs.

Animals↗

Pan-genomics and multi-omics for deciphering genetic variation and accelerating genetic improvement in ruminant livestock.

Livestock reference genomes have transformed the discovery of variants associated with production, reproduction, health, and environmental adaptation. Nevertheless, a single linear reference represents only one mosaic haplotype and incompletely captures sequence diversity within a species, particularly structural variants, copy-number changes, repeat-rich regions, and breed-specific sequences. Pangenomes address this limitation by integrating multiple high-quality assemblies or population-scale variants into a unified sequence or graph representation. Concurrently, multi-omics approaches connect genomic variation with transcriptomic, epigenomic, manuscriptproteomic, metabolomic, and microbiome responses, thereby improving biological interpretation of genotype-phenotype relationships. This review synthesizes recent progress in livestock pangenomics and multi-omics, with emphasis on cattle, goats, sheep, water buffalo, and chickens. It describes advances in long-read and haplotype-resolved sequencing, graph construction, structural-variant discovery and genotyping, functional annotation, and integrative analysis. Recent pangenome studies have uncovered substantial non-reference sequence, reduced reference bias, identified breed- and population-specific structural variants, and resolved candidate variants underlying pigmentation, body size, tail morphology, cashmere production, altitude adaptation, and other economically relevant traits. However, translation into routine breeding remains constrained by uneven population representation, inconsistent structural-variant definitions, limited functional annotation, computational demands, and insufficient validation across environments. Future progress will depend on diverse near-complete assemblies, graph-aware imputation and genomic prediction, long-read transcriptomics, single-cell and spatial omics, rigorous causal validation, and open, interoperable resources. Together, these developments can support more accurate, resilient, and biologically informed livestock improvement. Importantly, current dairy-cattle evidence indicates that pangenome-derived structural variants can substantially improve variant discovery and functional interpretation while yielding only marginal average gains in routine genomic prediction, favoring targeted augmentation rather than wholesale replacement of established SNP-based evaluations.

Animals↗