PubMed HealthSearch

SEARCH · PubMed Health

Results for “high-dimensional omics data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

LimROTS: a hybrid method integrating empirical Bayes and reproducibility-optimized statistics for robust differential expression analysis.

MOTIVATION: Differential expression analysis plays a vital role in omics research enabling precise identification of features that associate with different phenotypes. This process is critical for uncovering biological differences between conditions, such as disease versus healthy states. In proteomics, several statistical methods have been used, ranging from simple t-tests to more advanced methods like DEqMS, limma and ROTS. However, a flexible method for reproducibility-optimized statistics tailored for clinical omics data has been lacking. RESULTS: In this study, we developed LimROTS, a hybrid method that integrates a linear regression model and the empirical Bayes approach with reproducibility optimized statistics, to create a novel moderated ranking statistic, for robust and flexible analysis of proteomics data. We validated its performance using twenty-one proteomics gold standard spike-in datasets with different protein mixtures, MS instruments, and techniques for benchmarking. This hybrid approach improves accuracy and reproducibility of complex proteomics data, making LimROTS a powerful tool for high-dimensional omics data analysis. AVAILABILITY AND IMPLEMENTATION: LimROTS has been implemented as an R/Bioconductor package, available at https://doi.org/doi:10.18129/B9.bioc.LimROTS. Additionally, the code used in this study is available in GitHub repository https://github.com/AliYoussef96/LimROTSmanuscript.

Bayes Theorem

Model-based multifacet clustering with high-dimensional omics applications.

High-dimensional omics data often contain intricate and multifaceted information, resulting in the coexistence of multiple plausible sample partitions based on different subsets of selected features. Conventional clustering methods typically yield only one clustering solution, limiting their capacity to fully capture all facets of cluster structures in high-dimensional data. To address this challenge, we propose a model-based multifacet clustering (MFClust) method based on a mixture of Gaussian mixture models, where the former mixture achieves facet assignment for gene features and the latter mixture determines cluster assignment of samples. We demonstrate superior facet and cluster assignment accuracy of MFClust through simulation studies. The proposed method is applied to three transcriptomic applications from postmortem brain and lung disease studies. The result captures multifacet clustering structures associated with critical clinical variables and provides intriguing biological insights for further hypothesis generation and discovery.

Humans

MO-GCAN: multi-omics integration based on graph convolutional and attention networks.

MOTIVATION: Cancer subtypes play a critical role in disease progression, prognosis, and treatment, making their detection essential for tailoring precision medicine. Studies have shown that multi-omics integration outperforms single-omics approaches in cancer subtyping tasks. However, due to the high-dimensionality of multi-omics data, many existing studies either fail to capture the correlation between true labels and learned features, or lack sufficient capacity to model complex biological representations. These limitations hinder the full potential of leveraging the rich and complementary information embedded in multi-omics datasets. RESULT: We propose a framework that leverages supervised feature learning and classification based on a graph-based learning approach with attention mechanism for cancer subtyping. More specifically, we train graph convolutional network models on each omics dataset to extract latent representations, which are then concatenated to form a comprehensive multi-omics feature embedding. We further develop sample fusion network based on the omics-specific graphs, incorporating the derived features and feeding them into a graph attention model for subtype classification. This two-stage multi-omics framework is applied to eight cancer types, with performance evaluated in terms of test accuracy, training time, macro-averaged precision, recall, and F-score. Experimental results show that the proposed method outperforms state-of-the-art approaches across various cancer types. Additionally, we provide empirical evidence supporting the hypothesis that retaining a limited number of high-confidence edges and utilizing enriched embeddings from intermediate graph neural network layers can improve predictive performance. AVAILABILITY AND IMPLEMENTATION: Data and the code are available at https://github.com/YD-00/MO-GCAN-Updated.git.

Neoplasms

Protocol to perform integrative analysis of high-dimensional single-cell multimodal data using an interpretable deep learning technique.

The advent of single-cell multi-omics sequencing technology makes it possible for researchers to leverage multiple modalities for individual cells. Here, we present a protocol to perform integrative analysis of high-dimensional single-cell multimodal data using an interpretable deep learning technique called moETM. We describe steps for data preprocessing, multi-omics integration, inclusion of prior pathway knowledge, and cross-omics imputation. As a demonstration, we used the single-cell multi-omics data collected from bone marrow mononuclear cells (GSE194122) as in our original study. For complete details on the use and execution of this protocol, please refer to Zhou et al.1.

Deep Learning

A novel and robust feature selection method with FDR control for omics-wide association analysis.

Omics-wide association analysis is a very important tool for medicine and human health study. However, the modern omics data sets collected often exhibit the high-dimensionality, unknown distribution response, unknown distribution features and unknown complex association relationships between the response and its explanatory features. Reliable association analysis results depend on an accurate modeling for such data sets. Most of the existing association analysis methods rely on the specific model assumptions and lack effective false discovery rate (FDR) control. To address these limitations, the paper firstly applies a single index model for omics data. The model shows robust performance in allowing the relationships between the response variable and linear combination of covariates to be connected by any unknown monotonic link function, and both the random error and the covariates can follow any unknown distribution. Then based on this model, the paper combines rank-based approach and symmetrized data aggregation approach to develop a novel and robust feature selection method for achieving fine-mapping of risk features while controlling the false positive rate of selection. The theoretical results support the proposed method and the analysis results of simulated data show the new method possesses effective and robust performance for all the scenarios. The new method is also used to analyze the two real datasets and identifies some risk features unreported by the existing finds.

Humans

A framework for block-wise missing data in multi-omics.

High-throughput technologies have generated vast amounts of omic data. It is a consensus that the integration of diverse omics sources improves predictive models and biomarker discovery. However, managing multiple omics data poses challenges such as data heterogeneity, noise, high-dimensionality and missing data, especially in block-wise patterns. This study addresses the challenges of high dimensionality and block-wise missing data through a regularization and constrained-based approach. The methodology is implemented in the R package bwm for binary and continuous response variables, and applied to breast cancer and exposome multi-omics datasets, achieving strong performance even in scenarios with missing data present in all omics. In binary classification task, our proposed model achieves accuracy in the range of 86% to 92%, and F1 in the range of 68% to 79%. And, in regression task the correlation between true and predicted responses is in the range of 72% to 76%. However, there is a slight decline in performance metrics as the percentage of missing data increases. In scenarios where block-wise missing data affects multiple omics, the model performance actually surpasses that of scenarios where missing data is present in only one omics. One possible explanation for this might be that the other scenarios introduce a greater diversity of observation profiles, leading to a more robust model. Depending on the specific omics being studied, there is greater consistency in feature selection when comparing block-wise missing data scenarios.

Humans

Composition-on-composition regression analysis for multi-omics integration of metagenomic data.

MOTIVATION: Compositional data are frequently encountered in many disciplines, such as in next-generation sequencing experiments widely used in biomedical studies. Regression analysis with compositional data as either responses or predictors has been well studied. However, when both responses and predictors are compositional, the inventory of analysis tools is surprisingly limited, especially in the high-dimensional setting. Among the few existing methods, most of them rely on a log-ratio transformation to move compositional data from the simplex to real numbers. Yet, a serious weakness of these methods is their failure to handle the substantial fraction of zeroes observed in data collected from next-generation sequencing experiments. RESULTS: To investigate associations between two high-dimensional multi-omics compositions, we propose a composition-on-composition (COC) regression analysis method which does not require log-ratio transformations and hence can handle zeroes in the data. To account for high dimensionality, we estimate regression coefficients using a penalized estimation equation approach. Finally, inference procedures for COC regression are also proposed. Superior performance of COC is demonstrated through both comprehensive numerical simulations and case studies. AVAILABILITY AND IMPLEMENTATION: Source R codes to implement COC method is available at https://github.com/nrios4/COC.

Regression Analysis

[Applications and Challenges of Deep Learning in Human Genome Research].

In recent years, the advent of high-throughput omics technologies has fueled an explosive growth in human genomic data. Uncovering the latent functions within this vast data has become a significant challenge in functional genomics research. While traditional statistical methods have proved successful for analyzing smaller-scale datasets in the past, they exhibit clear limitations in analytical efficiency and integrating multi-dimensional data, struggling to meet the escalating demands of contemporary genomic analysis. The introduction of deep learning (DL) technologies offers a novel paradigm for this field. This review systematically examines the advances in applying deep learning to human genomics research. Studies demonstrate that when ample labeled data is available, discriminative DL computational methods-such as Convolutional Neural Networks (CNNs) and Long Short-Term Memory networks (LSTMs)-achieve high accuracy and efficiency in genomic variant discovery tasks. Furthermore, generative DL methods, particularly Large Language Models (LLMs) leveraging self-supervised pre-training strategies, effectively integrate complex genomic information and exhibit superior performance in functional genomic sequence annotation and gene regulation studies. This review also explores the application of LLMs in multi-omics data integration and prediction. Looking ahead, the continued accumulation of long-read sequencing and high-dimensional data is expected to enable DL technologies to integrate increasingly complex and heterogeneous genomic information, playing an increasingly crucial role in human genomics research.

Deep Learning

AI-integrated digital breeding for crop improvement.

Crop breeding increasingly depends on the effective integration and interpretation of large, heterogeneous datasets spanning genomic, phenotypic, multi-omics, and environmental layers. Conventional breeding approaches are often insufficient to capture the complex relationships among these data or to support timely selection decisions. Digital breeding can help address this limitation by complementing field experimentation, mixed models, and genomic prediction with the integration of biological data and computational prediction throughout the breeding process. In particular, the rapid advancement of artificial intelligence (AI) has improved the analysis of high-dimensional datasets and broadened its application to trait prediction, selection, and breeding design. Here, we review recent developments in AI-enabled digital breeding, encompassing genomic, phenomic, and multi-omics data generation and analysis, predictive modeling, explainable and generative AI, and data-driven breeding decision support. We further discuss emerging AI applications, their current contributions to crop research and breeding, and the major considerations affecting their reliable and practical implementation. Collectively, this review provides a structured understanding of the roles of AI across the digital breeding process and offers guidance for future methodological development and practical application in crop improvement.

artificial intelligence

Artificial Intelligence for Natural Products Discovery and Development.

Natural products (NPs) remain a cornerstone of modern drug discovery, offering stereochemical complexity and diverse bioactivities that precisely modulate therapeutic targets, refined through billions of years of evolution. However, their research has long been hindered by inefficient, empirical workflows, high resource consumption, structural complexity, and the "multicomponent, multi-target" nature of their mechanisms. The exponential growth of genomic, metabolomic, and spectral data has overwhelmed conventional analytical methods, exposing critical bottlenecks in handling high-dimensional, heterogeneous datasets that exceed human interpretive capacity. Artificial intelligence (AI) is emerging as a transformative paradigm to address these challenges, integrating multi-omics and chemical data to shift NP research from fragmented empiricism toward mechanism-driven, precision-oriented development. By leveraging deep learning architectures- including graph neural networks, Transformers, and diffusion-based generative models-AI enables systematic decoding of NP biosynthesis, automated structure elucidation, rational target identification, knowledge extraction from vast unstructured scientific literature, and de novo molecular design. This review comprehensively surveys recent advances in AI applications across the full NP discovery and development pipeline, encompassing genome mining, structure-based and ligand-based virtual screening, multimodal structural characterization, lead optimization, and biosynthetic pathway engineering. We further examine the emerging roles of protein-centric, molecule- centric, and multimodal foundation models, as well as large language models, in bridging genotype-to-chemotype gaps and unlocking unstructured scientific knowledge. Finally, we discuss critical challenges including data scarcity, representational limitations for complex stereochemistry, physical plausibility in generative models, and the urgent need for experimental validation, while outlining future directions toward autonomous experimentation, closed-loop optimization, and human-AI collaborative discovery.

Artificial intelligence

Multiomics approaches to cardiovascular disease: technological innovations and clinical translation.

Cardiovascular diseases (CVDs) remain the leading cause of global morbidity and mortality, reflecting a persistent gap between clinical phenotyping and the molecular mechanisms that govern disease initiation, progression, and interindividual variability. Recent advances in emerging technologies have fundamentally reshaped cardiovascular physiology by enabling high-resolution, cross-layer profiling of the heart and vasculature across genomic, epigenomic, transcriptomic, proteomic, metabolomic, lipidomic, glycomic, and fluxomic layers, increasingly at single-cell and spatial resolution. These approaches reveal CVD as a coordinated, multilayered process driven by dynamic interactions among cell types, regulatory programs, and metabolic states, rather than isolated gene-level defects. In this review, we synthesize how emerging multiomic, computational, and functional genomic technologies are redefining the study of cardiovascular disease across molecular, cellular, and tissue levels. We highlight recent innovations in single-cell and spatial atlases, long-read sequencing, proteomics and metabolomics, integrative data modeling, and functional omics approaches, including genome-scale perturbation screens and single-cell perturbation frameworks. These platforms enable mechanistic dissection of regulatory circuits, distinguish primary disease drivers from secondary adaptations, and directly assess therapeutic reversibility, advancing the field beyond associative biomarker discovery toward mechanism-guided target prioritization. We further discuss key methodological and translational challenges accompanying high-dimensional cardiovascular data, including preanalytical variability, control selection, temporal misalignment across molecular layers, population diversity, and reference bias. By integrating technological innovation with computational rigor and functional validation, this review frames emerging omics-enabled strategies as a unified, physiologically grounded framework for translating molecular insight into clinically meaningful cardiovascular phenotypes and advancing precision cardiovascular medicine.

Humans

CCNA2 orchestrates the PI3K/AKT signaling axis to propel prostate cancer metastasis.

BACKGROUND: Prostate cancer (PCa) remains one of the most common malignancies in men, posing a persistent global burden in terms of both public health and socioeconomic costs. Although early detection is essential for improving patient outcomes, existing clinical tools, including prostate-specific antigen (PSA) screening, digital rectal examination, and transrectal ultrasound-guided biopsy, are hampered by suboptimal specificity and positive predictive value, resulting in frequent overdiagnosis and overtreatment of indolent lesions while missing a subset of aggressive tumors at an early stage. In this context, the rapid advancement of high-throughput omics technologies, coupled with sophisticated machine learning (ML) algorithms, provides a powerful computational framework to dissect high-dimensional genomic data, uncover latent gene expression signatures, and identify candidate biomarkers with superior discriminative performance over conventional clinicopathological parameters. Therefore, in this study, we sought to screen for crucial ML-based biomarkers associated with PCa, with a particular focus on systematically assessing the diagnostic and prognostic value of CCNA2. Leveraging large-scale transcriptomic cohorts from public repositories, we employed an ensemble of ML approaches to prioritize candidate genes and subsequently evaluated the diagnostic performance of CCNA2 through receiver operating characteristic curve analysis, as well as its prognostic utility via Kaplan-Meier survival estimation and multivariate Cox proportional hazards modeling. Our findings are anticipated to elucidate the molecular landscape of PCa and offer a promising biomarker candidate for early detection and risk stratification. METHODS: This study integrated single-cell RNA sequencing, bulk transcriptomic data from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) repositories, immunofluorescence, and multiple ML algorithms with in vitro functional assays to evaluate CCNA2 expression, clinical relevance, and biological behavior in PCa. RESULTS: CCNA2 was linked to metastasis and poor prognosis. High CCNA2 expression significantly correlated with adverse survival outcomes, and knockdown of CCNA2 suppressed proliferation, migration, and invasion in PCa cell lines. Mechanistically, CCNA2 modulated the PI3K/AKT signaling pathway. An ML-based diagnostic model incorporating CCNA2 demonstrated high predictive accuracy across multiple validation cohorts. CONCLUSIONS: CCNA2 serves as a promising prognostic biomarker and therapeutic target in prostate adenocarcinoma, driving tumor progression potentially via the PI3K/AKT axis.

CCNA2

PEARL: integrative multi-omics classification and omics feature discovery via deep graph learning.

MOTIVATION: Integrating multi-omics data provides valuable insights into biological processes by capturing information across multiple molecular layers, enabling a comprehensive understanding of complex diseases and driving advancements in precision medicine. However, existing computational methods for multi-omics integration face significant challenges, such as low reliability and poor generalizability, due to the high dimensionality and low sample size nature of omics data. RESULTS: To address these challenges, we present PEARL (Pearson-Enhanced spectrAl gRaph convoLutional networks), a novel deep graph learning method for biomedical classification and functional important omics features identification. PEARL leverages a simple yet effective learning architecture to achieve superior and robust performance in high-dimensional, low-sample-size multi-omics settings. Our results demonstrate that PEARL significantly outperforms existing state-of-the-art methods on both synthetic and real biomedical datasets. Furthermore, applied to Alzheimer's disease (AD) brain multi-omics data, features prioritized by PEARL lead to functionally important genes that demonstrate significant enrichment in AD-related pathways. These findings highlight PEARL's practical utility in biomedical research and its potential to enhance biological interpretability in multi-omics studies. AVAILABILITY AND IMPLEMENTATION: The source code of our computational framework is available at https://github.com/zqq121017/PEARL.

Multiomics

NExON-Bayes: a Bayesian approach to network estimation informed by ordinal covariates.

MOTIVATION: In heterogeneous disease settings, accounting for intrinsic sample variability is crucial for obtaining reliable and interpretable omic network estimates. However, most graphical model analyses of biomedical data assume homogeneous conditional dependence structures, potentially leading to misleading conclusions. To address this, we propose a joint Gaussian graphical model that leverages sample-level ordinal covariates (e.g. disease stage) to account for heterogeneity and improve the estimation of partial correlation structures. RESULTS: Our modelling framework, called NExON-Bayes, extends the graphical spike-and-slab framework to account for ordinal covariates, jointly estimating their relevance to the graph structure and leveraging them to improve the accuracy of network estimation. To scale to high-dimensional omic settings, we develop an efficient variational inference algorithm tailored to our model. Through simulations, we demonstrate that our method outperforms the vanilla graphical spike-and-slab (with no covariate information), as well as other state-of-the-art network approaches which exploit covariate information. Applying our method to reverse phase protein array data from patients diagnosed with stage I, II or III breast carcinoma, we estimate the behaviour of proteomic networks as cancer progresses. Our model provides insights not only through inspection of the estimated proteomic networks, but also of the estimated ordinal covariate dependencies of key groups of proteins within those networks, offering a comprehensive understanding of how biological pathways shift across disease stages. AVAILABILITY AND IMPLEMENTATION: A user-friendly R package for NExON-Bayes with tutorials is available on Github at github.com/jf687/NExON, and archived at https://doi.org/10.5281/zenodo.20312938. The source of the dataset used is cited in the relevant section.

Bayes Theorem

Next-Generation Disease Profiling by Integrating Histopathology with Spatial Multi-Omics Data.

The field of pathology has experienced several transformative changes in recent years with the advent of digital pathology and spatial multi-omics. These technologies have enhanced every aspect of pathology practice, from streamlining daily workflows to generating high-fidelity multi-omics data that provide pathologists with novel tools to refine disease profiling and clinical diagnosis. Each layer of multimodal data (genomic, metabolomic, proteomic, or transcriptomic) has uncovered a distinct facet of disease pathologies, and combined with machine learning/artificial intelligence-based data analysis and pattern recognition models, has provided holistic understanding of regulatory mechanisms underpinning them. However, high-dimensional data have far exceeded the volume, scale, and complexity of immunostaining methods implemented by pathologists and, thus, have generated significant challenges related to deconvolution, interpretation, and clinical translation. Furthermore, these multimodal studies have predominantly relied on computational methods to process data and extract disease-relevant insights, thus raising questions around relevance or role of a pathologist in this new era of multi-omics. This review will provide a perspective on the evolving fields of molecular histopathology and spatial -omics, leveraging them to approach disease profiling, and redefining the role of a pathologist during this process.

Humans

Emerging multidimensional biomarker system for cardiovascular-kidney-metabolic syndrome: from multi-omics integration to clinical artificial intelligence.

Cardiovascular-kidney-metabolic (CKM) syndrome is an emerging clinical entity that highlights the complex, bidirectional interplay among cardiovascular disease, chronic kidney disease, and metabolic disorders, representing a substantial and growing global health burden. This conceptualization marks a paradigm shift from viewing these conditions in isolation to understanding them as an interconnected disease continuum. Traditional biomarkers face significant limitations in the early detection, risk stratification, and precise management of CKM, necessitating a transition towards an integrated framework that captures its multisystem nature. This review systematically outlines an emerging multidimensional biomarker system encompassing key pathological axes such as metabolism, immuno-inflammation, oxidative stress, and biological aging, offering refined risk assessment beyond conventional metrics. The development of this system is propelled by revolutionary platforms, including accessible sampling techniques (e.g., dried blood spots), advanced in vitro models (e.g., multi-organ-on-a-chip), and multi-omics technologies. These platforms not only facilitate a deeper dissection of the heterogeneous origins and inter-organ crosstalk in CKM but also accelerate the discovery and validation of novel biomarkers. Concurrently, artificial intelligence serves as a pivotal tool for clinical translation, effectively integrating high-dimensional data to transform complex molecular profiles into actionable clinical insights. By enabling the construction of dynamic risk prediction and decision-support systems, this review charts a pathway toward proactive, individualized, and precise prevention and management of CKM syndrome.

Humans

engGNN: a dual-graph neural network for omics-based disease classification and feature selection.

Omics data, such as transcriptomics, proteomics, and metabolomics, provide critical insights into disease mechanisms and clinical outcomes. However, their high dimensionality, small sample sizes, and intricate biological networks pose major challenges for reliable prediction and meaningful interpretation. Graph neural networks offer a promising way to integrate prior knowledge by encoding feature relationships as graphs. Yet, existing methods typically rely solely on either an externally curated feature graph or a data-driven generated graph, which limits their ability to capture complementary information. To address this, we propose the external and generated Graph Neural Network (engGNN), a dual-graph framework that jointly leverages both external biological networks and data-driven generated graphs. Specifically, engGNN constructs a biologically informed undirected feature graph from established network databases and complements it with a directed feature graph derived from tree-ensemble models. This dual-graph design produces more comprehensive representations, thereby improving predictive performance and interpretability. Through extensive simulation studies and real-world applications to three independent gene expression datasets, engGNN consistently demonstrates strong classification performance compared with competitive baselines. Beyond classification, engGNN provides feature- and source-level interpretability, enabling biologically meaningful analyses such as pathway enrichment analysis. Taken together, these results highlight engGNN as a robust, flexible, and interpretable framework for disease classification and biomarker discovery in high-dimensional omics contexts.

Graph Neural Networks

T-SMmOTE: tweaked synthetic majority minority oversampling technique for data scarcity issue in multi omics studies.

MOTIVATION: Multiomics data offer a rich data mine for modeling complex as well as day-to-day diseases, but their practical deployment is constrained by the limited sample availability. To this end, generating synthetic samples is a viable remedy. Extant schemes operating along this line, however, are mostly limited to augmenting the minority class in imbalanced datasets and often produce synthetic samples that lack sufficient diversity and fail to faithfully capture the underlying data distribution. As a result, the full potential of synthetic augmentation in multi-omics learning remains underexplored. The aim is to address the data scarcity problem in multi-omics domain. We propose a synthetic oversampling framework, which is dedicated to addressing overall data scarcity in multi-omics datasets and the lack of diversity in synthetic samples. Contrary to conventional methods that restrict augmentation to minority classes and rely on interpolation of two neighbors, our method generates diverse yet distribution-aligned synthetic samples by interpolating three neighbors and extends this augmentation paradigm to the majority class. The framework first balances the dataset by generating synthetic minority samples, and subsequently augments the balanced dataset by oversampling both majority and minority classes. RESULTS: Empirical evaluation on multi-omics data obtained from three heterogeneous health scenarios-inflammatory bowel disease, multi-organ dysfunction syndrome, and colorectal cancer-substantiates the utility of the proposed scheme in improving the predictive performance. The models trained on T-SMmOTE-augmented data achieve higher Matthews correlation coefficient values, along with improvedscores for both majority and minority classes. Notably, oversampling of the majority class improves the cognition of the minority class as well. We also explore the consistency of the class distributions between the original and augmented class-specific datasets. These findings confirm the capability of our scheme to learn from small, high-dimensional multi-omics datasets and highlight its potential for non-invasive disease detection. AVAILABILITY AND IMPLEMENTATION: https://github.com/payelu/TSMm.

Journal Article