PubMed HealthSearch

SEARCH · PubMed Health

Results for “Integrative omics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

moiraine: an R package to construct reproducible pipelines for the application and comparison of multi-omics integration methods.

MOTIVATION: In the past decades, many statistical methods for integrating multi-omics data have been developed. They have been implemented into software tools, which differ widely in their programming choices, such as the format required for data input, or the format of the generated integration results. This lack of standards renders cumbersome and time-intensive the application and comparison of different integration tools to the same multi-omics dataset. RESULTS: We have developed the moiraine R package for constructing reproducible multi-omics integration pipelines, which enables users to apply one or more statistical methods for multi-omics integration to their own multi-omics dataset. moiraine facilitates the preprocessing of the omics datasets and automates their formatting for the integration step. It simplifies the interpretation and evaluation of the integration results through the construction of visualizations in which metadata about samples and features can easily be included. Crucially, it enables the comparison of results obtained with different integration tools, allowing users to assess the robustness of their results. AVAILABILITY AND IMPLEMENTATION: The moiraine R package is publicly available at https://github.com/Plant-Food-Research-Open/moiraine; an archival snapshot of the package is available on Zenodo at https://doi.org/10.5281/zenodo.17172718. A detailed tutorial is available at https://plant-food-research-open.github.io/moiraine-manual/.

Software

ARISE: RNA-anchored shared-edge topology and hierarchical fusion for spatial multi-omics integration.

MOTIVATION: Spatial multi-omics technologies jointly profile transcriptomes, proteins and chromatin accessibility in situ, enabling integrative analysis of tissue organization across molecular layers. However, most existing graph-based integration methods rely on independently constructed modality-specific k-nearest-neighbor graphs. When auxiliary modalities are sparse or noisy, these graphs can become topologically discordant, propagate spurious edges, weaken cross-modal alignment, and reduce spatial domain resolution. RESULTS: We present Anchored RNA for Integrated Spatial Embedding (ARISE), an RNA expression anchored framework for spatial multi-omics integration. ARISE defines a shared-edge topology by intersecting RNA feature-similarity and spatial-proximity graphs, encodes auxiliary modalities on this common scaffold, and integrates them through inside-out hierarchical fusion. We further show theoretically that graph intersection minimizes false-positive edges within a broad class of k-of-r graph fusion rules, providing a principled basis for topology anchoring. Across various spatial multi-omics benchmarks spanning simulated and real datasets in bi-modal and tri-modal settings, ARISE improves spatial domain identification, cross-modal consistency, and preservation of tissue structure relative to existing methods. Furthermore, the learned representation supports biologically meaningful downstream analyses, including marker-based domain annotation, pathway enrichment, and cis-regulatory inference, indicating that ARISE yields a robust and interpretable framework for spatial multi-omics integration. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/XiangxiangWang-code/ARISE. The archived version used in this study is available at https://doi.org/10.6084/m9.figshare.32686137.v2.

Multiomics

Multi-omics integration uncovers epigenetic control of metabolic reprogramming in triple-negative breast cancer.

Triple-negative breast cancer (TNBC) is an aggressive subtype characterized by the absence of estrogen, progesterone, and HER2 receptors, limiting effective targeted therapies. Increasing evidence suggests that metabolic reprogramming, a hallmark of TNBC progression, is driven by underlying epigenetic mechanisms such as DNA methylation. The represented study performed an integrative analysis of transcriptomic (RNA-seq) and methylome data to uncover the metabolic-epigenetic interplay in TNBC. Differential gene expression analysis using DESeq2 revealed significant dysregulation of key metabolic genes, including upregulation of genes encoding glycolytic and serine biosynthesis enzymes and downregulation of metabolic tumor suppressors. Genome-wide methylation profiling identified extensive cytosine-phosphate-guanine (CpG) hypermethylation events associated with transcriptional repression, particularly in promoter regions. Integrative analysis pinpointed a subset of metabolism-related genes exhibiting both differential expression and methylation, such as FBP1, RASSF1A, and PHGDH. Pathway enrichment analysis highlighted aberrations in glycolysis/gluconeogenesis, fatty acid metabolism, and one-carbon pathways (adjusted p&#x2009;<&#x2009;0.01). Importantly, TNBC patients with hypermethylated metabolic gene signatures displayed significantly shorter overall survival (log-rank p&#x2009;<&#x2009;0.05). These findings reveal that DNA methylation-driven metabolic dysregulation contributes to TNBC aggressiveness and may provide novel biomarkers and therapeutic targets at the metabolic-epigenetic interface.

Humans

Livestock Multi-Omics Integration: A Systematic Framework From Statistical Association to Causal Interpretation.

Livestock multi-omics integration is key to unraveling complex trait regulation, yet systematic, livestock-specific strategies remain scarce. This review traces the progression from single-omics accumulation to multi-dimensional integration, highlighting how large-scale genomic, epigenomic, and transcriptomic projects lay the foundation for functional dissection. We identify core impediments: extreme species diversity, marked data heterogeneity, limited sample sizes, and a pervasive reduction of multi-omics data to simplistic differential screens, resulting in low translational efficiency. We critically appraise four common pitfalls-overinterpreting correlation as causation, relegating proteomics to corroborating transcriptomics, incomplete microbiome-host integration lacking environmental context, and systematic neglect of metabolic fluxomics-and show how exposomics and fluxomics add necessary causal and dynamic dimensions. To address these, we propose a livestock-adapted three-tier analytical framework: (1) statistical association of cross-omics covariation patterns; (2) machine learning-driven feature mining and integrative modeling; and (3) causal interpretation encompassing Mendelian randomization, prior-knowledge-guided network inference, and physical causal evidence via fluxomics and metabolic control analysis. We further discuss how multimodal sequencing (single-cell, spatial, temporal) and generative AI can fundamentally mitigate heterogeneity and strengthen causal evidence. Finally, we outline future priorities in database standardization, livestock-specific benchmarking, and translational pipelines, charting a path from correlation-centric reporting to mechanistic causality and precision breeding.

Animals

Multi-omics integration uncovers adaptive responses of stomach and pyloric ceca to artificial feed in mandarin fish (Siniperca chuatsi).

The mandarin fish, as an obligate piscivore, is highly dependent on live bait, which restricts its intensive aquaculture. Although domestication has enabled it to partially accept formulated diets, the tissue-specific molecular adaptation mechanisms of its digestive tract to artificial feed remain unclear. In this study, we conducted an integrated analysis of mandarin fish fed with live bait or artificial diet for three weeks, combining growth performance evaluation, gastric histology, and paired transcriptomic and metabolomic analyses of the stomach and pyloric ceca. AD feeding significantly improved growth performance, while histological examination revealed marked hyperplasia of the gastric mucosa and disorganized fold structures. Transcriptomic analysis identified 5065 and 3381 differentially expressed genes in the stomach and pyloric ceca, respectively. In the stomach, the artificial diet induced a glutathione-dependent antioxidant response, accompanied by glycolytic reprogramming and coordinated upregulation of genes in the extracellular matrix (ECM)-receptor interaction signaling pathway, including those encoding collagen, laminin, and integrin. In the pyloric ceca, the tricarboxylic acid (TCA) cycle and oxidative phosphorylation were broadly suppressed, whereas glycosaminoglycan degradation and lysosomal pathways were activated. Metabolomic analysis showed that gastric metabolites were enriched in vascular and inflammatory mediator pathways, while metabolites in the pyloric ceca were enriched in peroxisome proliferator-activated receptor (PPAR) signaling, sphingolipid signaling, and steroid hormone biosynthesis pathways. Following artificial diet feeding, integrated multi-omics analysis of the stomach revealed significant enrichment of pathways such as phospholipase D signaling, sphingolipid signaling, and arachidonic acid metabolism, accompanied by the accumulation of key metabolites including sphingosine-1-phosphate, 20-hydroxyeicosatetraenoic acid, and cellobiose. Integrated analysis of the pyloric ceca identified significantly altered pathways, including sphingolipid metabolism, alpha-linolenic acid metabolism, and glutathione metabolism, along with elevated levels of sphingosine-1-phosphate, sphingosine galactoside, and 9-hydroxy-12-oxo-10,15-octadecadienoic acid, as well as decreased glutathionylspermidine. These findings systematically unveil the tissue-specific molecular adaptation characteristics of the mandarin fish digestive tract in response to artificial feed, providing an important basis for understanding the molecular mechanisms of dietary adaptation in carnivorous fish and for optimizing artificial feed formulations.

Animals

AI-Driven Multi-Omics Integration of Synthetic Colon Adenocarcinoma for Cluster-Guided PROTAC Candidate Design Targeting KRASG12D.

Colorectal cancer is a leading cause of cancer death, yet its molecular heterogeneity remains poorly translated into individualized treatment. We present a reproducible artificial intelligence (AI) framework that integrates multi-omics benchmarking, sample-level drug prioritization, E3 ubiquitin ligase selection, and shape-anchored Proteolysis Targeting Chimera (PROTAC) design for KRASG12D in colon adenocarcinoma (COAD). A controlled synthetic benchmark comprising 425 tumor and 41 simulated normal profiles, parameterized to match The Cancer Genome Atlas (TCGA) distributions, was used for pipeline verification. Among sixteen methods, the Balanced Latent Integration with Stability Selection (BLISS) model achieved the highest silhouette width (0.86) and competitive agreement (Adjusted Rand Index, ARI, 0.90). The pipeline was validated on real data: a TCGA COAD cohort (186 tumors) with independent Consensus Molecular Subtype (CMS) labels and a CPTAC cohort (104 tumors). Integration modestly recovered CMS (ARI 0.28), and stage, not molecular cluster, drove survival (log-rank p = 0.005 versus 0.81). Sample-level prioritization differed from cluster-level ranking in 82.6% of profiles, below chance (p < 0.0001), without indicating efficacy. Candidate NOVEL00489 showed a good MM-GBSA estimate, matching the reference ASP3082. Compounds are computational candidates requiring experimental validation. This establishes a transparent benchmark for in silico degrader generation in precision oncology.

Humans

Multi-Omics Integration Identifies a Five-Gene Metabolic Signature With Experimental Validation in Clear Cell Renal Cell Carcinoma.

BACKGROUND: Clear cell renal cell carcinoma (ccRCC) is hallmarked by profound metabolic reprogramming; however, its intricate crosstalk with the tumor immune microenvironment (TIME) and its clinical ramifications remain inadequately elucidated. This study aims to systematically decipher the metabolic-immune interplay in ccRCC through multi-omics integration, with the goal of identifying robust prognostic biomarkers and actionable therapeutic vulnerabilities. AIMS: This study aims to systematically decipher the metabolic-immune interplay in clear cell renal cell carcinoma (ccRCC) through multi&#x2011;omics integration, and to identify robust prognostic biomarkers and actionable therapeutic vulnerabilities that can inform precision risk stratification and individualized treatment strategies. METHODS: We integrated bulk transcriptomic, genomic, and clinical data from multiple ccRCC cohorts. Differential expression and functional enrichment analyses were performed to characterize metabolic pathway alterations. Mendelian randomization (MR) was employed to infer causal relationships between metabolic disorders and ccRCC risk. A machine learning-based prognostic framework, incorporating SHAP (SHapley Additive exPlanations) for feature interpretability, was constructed and rigorously validated. TIME heterogeneity was dissected using deconvolution algorithms, while drug sensitivity, tumor mutation burden (TMB), and TIDE scores were utilized to assess therapeutic responses and immune evasion. Candidate gene function was evaluated through in&#xa0;vitro gain- and loss-of-function assays, with expression validated via TCGA, HPA, western blot, and qRT-PCR. RESULTS: Enrichment analysis identified coordinated dysregulation in lipid metabolism, energy homeostasis, and hypoxia response pathways. MR analysis confirmed lipid metabolism disorders as a causal risk factor for ccRCC. Our machine-learning model, centered on five core SHAP-identified features (SUCLA2, ACAT1, PC, SUCLG1, and HMGCS2), demonstrated superior predictive accuracy over conventional clinical staging. Immune profiling unveiled dichotomous TIME states: the low-risk group retained active immune surveillance, whereas the high-risk group was enriched with immunosuppressive subsets. Drug sensitivity screening pinpointed LY2109761 and carmustine as high-risk-specific candidate agents. Furthermore, TMB and TIDE analyses stratified high-risk patients displaying genomic instability and immune evasion phenotypes. Functionally, SUCLA2 knockdown significantly enhanced ccRCC cell proliferation and invasion, while its overexpression suppressed these malignant phenotypes, corroborating its tumor-suppressive role. Expression patterns of the hub genes were consistently validated across multi-level datasets and experimental assays. CONCLUSION: This study establishes a precision oncology framework for ccRCC by functionally linking metabolic biomarkers, immunophenotypes, and stratified therapeutic strategies. Importantly, we identify SUCLA2 as a potential functional tumor suppressor and a promising target for further mechanistic and translational investigation.

Humans

Research progress and application prospects of multi-omics integration strategies in precision risk stratification of type 1 diabetes mellitus.

Type 1 diabetes (T1D) is a chronic metabolic disease mediated by autoimmunity. Its pathogenesis involves complex interactions between genetic susceptibility and environmental factors. Conventional T1D risk stratification primarily relies on genetic markers, islet autoantibodies, and glycemic indicators. Although these biomarkers remain indispensable in current clinical practice, they are often insufficient when used alone to accurately identify ultra-early high-risk individuals, predict disease progression rates, or support individualized preventive strategies. Consequently, more comprehensive molecular approaches are needed to improve precision risk stratification. In recent years, the rapid development of multi-omics technologies has provided new strategies for precise risk stratification of T1D. This narrative review critically evaluates how multi-omics integration strategies can improve precision risk stratification throughout the T1D disease continuum by integrating complementary molecular information from genomics, transcriptomics, proteomics, metabolomics, epigenomics, and the microbiome. Particular emphasis is placed on stage-specific biomarker discovery, multi-omics data integration frameworks, artificial intelligence-assisted prediction models, biomarker validation, and the opportunities and challenges associated with clinical translation. Current evidence suggests that integrated multi-omics approaches have the potential to improve risk prediction accuracy, distinguish heterogeneous disease trajectories, identify individuals at imminent risk of progression, and provide biologically informed targets for precision intervention. However, important challenges remain, including data harmonization, external validation, model interpretability, cost-effectiveness, and integration into routine clinical screening programs. Future research should prioritize prospective multicenter cohorts, standardized analytical pipelines, externally validated prediction models, and clinically interpretable multi-omics frameworks to facilitate the translation of precision risk stratification into routine T1D prevention and management.

Humans

Clustering individuals using INMTD: a novel versatile multi-view embedding framework integrating omics and imaging data.

MOTIVATION: Combining omics and images can lead to a more comprehensive clustering of individuals than classic single-view approaches. Among the various approaches for multi-view clustering, nonnegative matrix tri-factorization (NMTF) and nonnegative Tucker decomposition (NTD) are advantageous in learning low-rank embeddings with promising interpretability. Besides, there is a need to handle unwanted drivers of clusterings (i.e. confounders). RESULTS: In this work, we introduce a novel multi-view clustering method based on NMTF and NTD, named INMTD, which integrates omics and 3D imaging data to derive unconfounded subgroups of individuals. According to the adjusted Rand index, INMTD outperformed other clustering methods on a synthetic dataset with known clusters. In the application to real-life facial-genomic data, INMTD generated biologically relevant embeddings for individuals, genetics, and facial morphology. By removing confounded embedding vectors, we derived an unconfounded clustering with better internal and external quality; the genetic and facial annotations of each derived subgroup highlighted distinctive characteristics. In conclusion, INMTD can effectively integrate omics data and 3D images for unconfounded clustering with biologically meaningful interpretation. AVAILABILITY AND IMPLEMENTATION: INMTD is freely available at https://github.com/ZuqiLi/INMTD.

Cluster Analysis

MO-GCAN: multi-omics integration based on graph convolutional and attention networks.

MOTIVATION: Cancer subtypes play a critical role in disease progression, prognosis, and treatment, making their detection essential for tailoring precision medicine. Studies have shown that multi-omics integration outperforms single-omics approaches in cancer subtyping tasks. However, due to the high-dimensionality of multi-omics data, many existing studies either fail to capture the correlation between true labels and learned features, or lack sufficient capacity to model complex biological representations. These limitations hinder the full potential of leveraging the rich and complementary information embedded in multi-omics datasets. RESULT: We propose a framework that leverages supervised feature learning and classification based on a graph-based learning approach with attention mechanism for cancer subtyping. More specifically, we train graph convolutional network models on each omics dataset to extract latent representations, which are then concatenated to form a comprehensive multi-omics feature embedding. We further develop sample fusion network based on the omics-specific graphs, incorporating the derived features and feeding them into a graph attention model for subtype classification. This two-stage multi-omics framework is applied to eight cancer types, with performance evaluated in terms of test accuracy, training time, macro-averaged precision, recall, and F-score. Experimental results show that the proposed method outperforms state-of-the-art approaches across various cancer types. Additionally, we provide empirical evidence supporting the hypothesis that retaining a limited number of high-confidence edges and utilizing enriched embeddings from intermediate graph neural network layers can improve predictive performance. AVAILABILITY AND IMPLEMENTATION: Data and the code are available at https://github.com/YD-00/MO-GCAN-Updated.git.

Neoplasms

Cross-tissue multi-omics integration highlights BPHL and mitochondrial targets in Alzheimer's disease.

BACKGROUND: Mitochondrial dysfunction is a hallmark of Alzheimer's disease (AD), yet specific molecular targets remain to be fully characterized. METHODS: A summary-data-based Mendelian randomization (SMR) framework integrated AD genome-wide association study (GWAS) statistics (39,918 cases) with blood DNA methylation quantitative trait loci (mQTL), gene expression (eQTL), and protein (pQTL) data for 1136 mitochondria-related genes. Associations were assessed using Bayesian colocalization and HEIDI testing. Tissue relevance was evaluated in four brain regions (hippocampus, amygdala, cortex, frontal cortex) using GTEx and external transcriptomic datasets. RESULTS: Screening identified eight candidates supported across blood mQTL and eQTL layers. Stepwise central nervous system (CNS) evaluation singled out biphenyl hydrolase-like (BPHL) as the consistent candidate. Higher genetically predicted BPHL expression was associated with reduced AD risk across the hippocampus (OR=0.920, 95% CI 0.873-0.970), amygdala (OR=0.925, 95%CI 0.880-0.973), cortex (OR=0.943, 95% CI 0.908-0.978), and frontal cortex (OR=0.938, 95%CI 0.901-0.976). These findings aligned with protein-protein interactions connecting BPHL to respiratory complexes and lower BPHL expression in independent AD brains. Functional enrichment converged on oxidative phosphorylation pathways. CONCLUSIONS: By integrating multi-omics data with tissue-specific validation, this study nominates BPHL as a consistent protective candidate in the brain. These findings provide genetic support for mitochondrial molecular perturbations in AD, offering insights for future validation.

Alzheimer Disease

In vivo porcine multi-omics integration identifies microbiome-driven histamine elevation and lasting gut perturbations following Ascaris suum infection and fenbendazole treatment.

Ascaris roundworms impair human and swine health. While treatments using anthelmintic drugs are generally effective in eliminating worms, their effects on the gut microenvironment remain poorly understood. Here we applied integrated multi-omics to characterize infection- and treatment-associated alterations in the pig-Ascaris system. In vitro anaerobic cultures were conducted as supportive validation of selected observations. Ascaris suum infection altered microbial composition and dysregulated 182 serum and fecal metabolites, including histamine and p-cresol sulfate. Compared with time-matched uninfected controls, infected pigs treated with fenbendazole showed marked differences in gut microbial composition 13&#x2009;days after confirmed worm clearance. Eleven microbial pathways were enriched in successfully treated pigs, including peptidoglycan biosynthesis and histidine metabolism, indicating that infection-associated alterations may persist after treatment. In vitro co-exposure of Lactobacillus reuteri to fenbendazole and A. suum proteins increased histamine production by approximately 79% at 48&#x2009;h (p&#x2009;<&#x2009;0.05), serving as supportive evidence of a microbiome contribution. Collectively, our in vivo findings support that host-microbiota-parasite interactions are multifaceted. Microbiota-derived metabolites were associated with regulation of host gene expression, such as TFF2 and IL8. Microbiota plasticity allows the exploitation of the niche differentiated upon infection, resulting in the proliferation of certain Lactobacillus strains in treated animals. Nevertheless, interpretations of treatment effects are made cautiously given the absence of an uninfected drug-only group and the cross-sectional design. Understanding these complex interactions will be important for the design of next-generation functional anthelmintics.

Animals

IGCN: integrative graph convolution networks for patient level insights and biomarker discovery in multi-omics integration.

MOTIVATION: Developing computational tools for integrative analysis across multiple types of omics data has been of immense importance in cancer molecular biology and precision medicine research. While recent advancements have yielded integrative prediction solutions for multi-omics data, these methods lack a comprehensive and cohesive understanding of the rationale behind their specific predictions. To shed light on personalized medicine and unravel previously unknown characteristics within integrative analysis of multi-omics data, we introduce a novel integrative neural network approach for cancer molecular subtype and biomedical classification applications, named Integrative Graph Convolutional Networks (IGCN). RESULTS: To demonstrate the superiority of IGCN, we compare its performance with other state-of-the-art approaches across different cancer subtype and biomedical classification tasks. Our experimental results show that our proposed model outperforms the state-of-the-art and baseline methods. IGCN identifies which types of omics data receive more emphasis for each patient when predicting a specific class. Additionally, IGCN has the capability to pinpoint significant biomarkers from a range of omics data types. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/bozdaglab/IGCN.

Humans

X-intNMF: a cross- and intra-omics regularized NMF framework for multi-omics integration.

MOTIVATION: The rapid accumulation of multi-omics data presents a valuable opportunity to advance our understanding of complex diseases and biological systems, driving the development of integrative computational methods. However, the complexity of biological processes, spanning multiple molecular layers and involving intricate regulatory interactions, requires models that can capture both intra- and cross-omics relationships. Most existing integration methods primarily focus on sample-level similarities or intra-omics feature interactions, often neglecting the interactions across different omics layers. This limitation can result in the loss of critical biological information and suboptimal performance. To address this gap, we propose X-intNMF, a network-regularized non-negative matrix factorization (NMF) framework that simultaneously integrates intra- and cross-omics feature interactions into a shared low-dimensional representation (see Fig.&#xa0;1). By modeling these multi-layered relationships, X-intNMF enhances the representation of biological interactions and improves integration quality and prediction accuracy. RESULTS: For evaluation, we applied X-intNMF to predict breast cancer phenotypes and classify clinical outcomes in lung and ovarian cancers using mRNA expression, microRNA expression, and DNA methylation data from TCGA. The results show that X-intNMF consistently outperforms state-of-the-art methods. Ablation studies confirm that incorporating both cross-omics and intra-omics interactions contributes significantly to the model's improved performance. Additionally, survival analysis on 25 TCGA cancer datasets demonstrates that the integrated multi-omics representation offers strong prognostic value for both overall survival and disease-free status. These findings highlight X-intNMF's ability to effectively model multi-layered molecular interactions while maintaining interpretability, robustness, and scalability within the NMF framework. AVAILABILITY AND IMPLEMENTATION: The source code and datasets supporting this study are publicly available at GitHub (https://github.com/compbiolabucf/X-intNMF) and archived on Zenodo (https://doi.org/10.5281/zenodo.18238385).

Multiomics

Machine learning-based clinical prediction model and multi-omics integration for assessing pancreatic cancer risk in new-onset diabetes.

BACKGROUND: Given that pancreatic cancer (PC) is typically diagnosed at an advanced stage but is often preceded by new-onset diabetes mellitus (NODM), providing a window for early detection, we sought to develop and validate an interpretable machine-learning model integrated with multi-omics profiling to identify early biomarkers of NODM-associated PC. METHODS: In a population-based cohort, individuals with NODM-associated PC and NODM without PC were identified and randomly divided (70:30) into training and validation sets after feature selection. Eight machine learning (ML) classifiers were compared using fivefold cross-validation, and model performance was evaluated in terms of discrimination, calibration, and decision curve&#x2013;based clinical utility. We evaluated interpretability using the Shapley additive explanations (SHAP) analyses. Mechanistically, Olink proteomic profiling and metabolomics were analyzed through clinical classifications and model-defined risk strata. RESULTS: Categorical boosting achieved the best performance in the independent validation set (AUROC&#x2009;=&#x2009;0.844). The NODM cohort was stratified into high- (n&#x2009;=&#x2009;2,362) and low-risk (n&#x2009;=&#x2009;5,030) groups, and internal validation together with SHAP analyses demonstrated consistent model performance and identified clinically interpretable predictors. Proteomic and metabolomic analyses under clinical and risk-based grouping identified 39 overlapping differentially expressed proteins and 145 overlapping metabolites with enriched across 11 shared KEGG pathways. Cross-platform validation highlighted PLTP, CRTAC1, and ITGAV as serum biomarkers with a strong potential for early NODM-PC detection. CONCLUSIONS: We developed an interpretable ML framework centered on NODM enables practical risk stratification for early PC detection by multi-omics and provides a pathway of ML-based triage followed by biomarker confirmation for earlier detection and diagnosis.

Humans

Algorithms and tools for data-driven omics integration to achieve multilayer biological insights: a narrative review.

Systems biology is a holistic approach to biological sciences that combines experimental and computational strategies, aimed at integrating information from different scales of biological processes to unravel pathophysiological mechanisms and behaviours. In this scenario, high-throughput technologies have been playing a major role in providing huge amounts of omics data, whose integration would offer unprecedented possibilities in gaining insights on diseases and identifying potential biomarkers. In the present review, we focus on strategies that have been applied in literature to integrate genomics, transcriptomics, proteomics, and metabolomics in the year range 2018-2024. Integration approaches were divided into three main categories: statistical-based approaches, multivariate methods, and machine learning/artificial intelligence techniques. Among them, statistical approaches (mainly based on correlation) were the ones with a slightly higher prevalence, followed by multivariate approaches, and machine learning techniques. Integrating multiple biological layers has shown great potential in uncovering molecular mechanisms, identifying putative biomarkers, and aid classification, most of the time resulting in better performances when compared to single omics analyses. However, significant challenges remain. The high-throughput nature of omics platforms introduces issues such as variable data quality, missing values, collinearity, and dimensionality. These challenges further increase when combining multiple omics datasets, as the complexity and heterogeneity of the data increase with integration. We report different strategies that have been found in literature to cope with these challenges, but some open issues still remain and should be addressed to disclose the full potential of omics integration.

Algorithms

Trans-omics integration underscores distinct roles of polyunsaturated phospholipids in bidirectional offspring birth weight deviations.

BACKGROUND: Abnormal birth weights are associated with adverse pregnancy outcomes and future metabolic consequences. We aimed to examine cord blood lipidomes from low, normal and high birth weight (LBW, NBW, HBW) infants to identify core lipid signatures associated with non-optimum birth weight, and to derive biological insights through trans-omics data integration with placental proteome, maternal plasma lipidome and clinical phenome. METHODS: We conducted quantitative lipidomics of cord blood samples from two independent cohorts: a retrospective discovery cohort (n = 147) and a prospective validation cohort (n = 73). Integration with placental proteomics, maternal plasma lipidomics and clinical phenomics was conducted to elucidate potential biological implications. FINDINGS: We identified substantial reductions in cord blood polyunsaturated phospholipids (PUFA-PLs) (FDR <0.05) associated with placental vesicle trafficking and formation in LBW, and altered neutrophil degranulation in HBW. Combinatorial analyses of paired maternal plasma and cord blood samples indicated that cord blood PUFA-PL reductions were not attributable to deficient maternal supply, but rather to impeded assimilation (LBW) and increased utilisation (HBW). INTERPRETATION: Our findings provide biological insights that may inform targetable, lipid-oriented nutritional and/or pharmacological strategies to modulate foetal growth and development, with the goal of optimising clinical outcomes for both mother and child. FUNDING: This work was supported by the National Natural Science Foundation of China (82170854, 81870579, 81870545, 82571043, 2357308); National High Level Hospital Clinical Research Funding (2022-PUMCH-C-019); Noncommunicable Chronic Diseases-National Science and Technology Major Project (2024ZD0530200 and 2024ZD0530204); Beijing Municipal Science & Technology Commission (Z201100005520011); Peking University Clinical Scientist Training Program (No. BMU2023PYJH022); Beijing Municipal Natural Science Foundation (7202163, 7184252).

Humans

NMFProfiler: a multi-omics integration method for samples stratified in groups.

MOTIVATION: The development of high-throughput sequencing enabled the massive production of "omics" data for various applications in biology. By analyzing simultaneously paired datasets collected on the same samples, integrative statistical approaches allow researchers to get a global picture of such systems and to highlight existing relationships between various molecular types and levels. Here, we introduce NMFProfiler, an integrative supervised NMF that accounts for the stratification of samples into groups of biological interest. RESULTS: NMFProfiler was shown to successfully extract signatures characterizing groups with performances comparable to or better than state-of-the-art approaches. In particular, NMFProfiler was used in a clinical study on atopic dermatitis (AD) and to analyze a multi-omic cancer dataset. In the first case, it successfully identified signatures combining known AD protein biomarkers and novel transcriptomic biomarkers. In addition, it was also able to extract signatures significantly associated to cancer survival. AVAILABILITY AND IMPLEMENTATION: NMFProfiler is released as a Python package, NMFProfiler (v0.3.0), available on PyPI.

Humans