PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “multi-omics data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

317 records · Page 18Linked to original sources

scBSP: a fast and accurate tool for identifying spatially variable features from high-resolution spatial omics data.

MOTIVATION: Emerging spatial omics technologies empower comprehensive exploration of biological systems from multi-omics perspectives in their native tissue location in 2D and 3D space. However, the limited sequencing depth, increasing spatial resolution, and growing spatial spots in spatial omics technologies present significant computational challenges in identifying biologically meaningful molecules with variable spatial distributions across various omics modalities. RESULTS: We introduce scBSP, an open-source, versatile, and user-friendly package for identifying spatially variable features in large-scale spatial omics data. scBSP demonstrates significantly enhanced computational efficiency, processing high-resolution spatial omics data within seconds, and exhibits robust cross-platform performance by consistently identifying spatially variable features with high reproducibility across various sequencing platforms. AVAILABILITY AND IMPLEMENTATION: scBSP is available for download from R CRAN at https://cran.r-project.org/web/packages/scBSP/index.html and PyPI at https://pypi.org/project/scbsp/.

Software↗

HoloFoodR: a statistical programming framework for holo-omics data integration workflows.

SUMMARY: Holo-omics is an emerging research area that integrates multi-omic datasets from the host organism and its microbiome to study their interactions. Recently, curated and openly accessible holo-omic databases have been developed. The HoloFood database, for instance, provides nearly 10 000 holo-omic profiles for salmon and chicken under controlled treatments. However, bridging the gap between holo-omic data resources and algorithmic frameworks remains a challenge. Combining the latest advances in statistical programming with curated holo-omic data sets can facilitate the design of open and reproducible research workflows in the emerging field of holo-omics. AVAILABILITY AND IMPLEMENTATION: HoloFoodR R/Bioconductor package and the source code are available under the open-source Artistic License 2.0 at the package homepage https://doi.org/10.18129/B9.bioc.HoloFoodR.

Software↗

Genetic evidence that advanced COVID-19 accelerates longitudinal brain atrophy: A Mendelian randomization study.

Coronavirus disease 2019 (COVID-19) was reported to persist long-term in the brain and leave several long-term neurologic sequelae. However, the causal relationship between COVID-19 and brain aging is still unknown. The genome-wide association study (GWAS) data on COVID-19 phenotypes (susceptibility, hospitalization, and severity), involving a total of 5,779,391 participants, were collected from the COVID-19 Host Genetics Initiative. In addition, GWAS data on longitudinal changes in 15 brain structures, assessed via magnetic resonance imaging across the lifespan, were sourced from the ENIGMA Consortium and involved 15,640 participants. Two-sample Mendelian randomization was conducted to infer the causal relationship between COVID-19 and longitudinal brain changes. Multi-trait GWAS meta-analysis, colocalization, and fine-mapping analyses were performed to identify shared genetic etiologies. H3K27me3 ChIP-seq was used to evaluate the regulatory effect of colocalized loci. Two-step Mendelian randomization was applied to explore potential mediating mechanisms across multi-omics layers, including proteomics, metabolomics, and immunomics. Our results showed that COVID-19 hospitalization (β = -262.405, P = .041) and severity (β = -177.676, P = .049) were genetically associated with atrophied volume of total brain during longitudinal change. This suggests that individuals with advanced COVID-19 may be more susceptible to accelerated global brain aging. Caudate was genetically affected by all COVID-19 phenotypes. Seven variants were shared between advanced COVID-19 and global brain aging. rs117169628 was colocalized between advanced COVID-19 and global brain aging, and exerted an inhibitory effect on CDH15 expression, further strengthening the causality. Six metabolites, 1 protein, and 1 immune trait were identified as potential mediators. Our study indicates that advanced COVID-19 might be genetically associated with accelerated brain aging. Brain health should be paid more attention in long COVID-19.

Humans↗

scooby: Modeling multi-modal genomic profiles from DNA sequence at single-cell resolution.

Understanding how regulatory DNA elements shape gene expression across individual cells is a fundamental challenge in genomics. Joint RNA-seq and epigenomic profiling provides opportunities to build unifying models of gene regulation capturing sequence determinants across steps of gene expression. However, current models, developed primarily for bulk omics data, fail to capture the cellular heterogeneity and dynamic processes revealed by single-cell multi-modal technologies. Here, we introduce scooby, the first framework to model scRNA-seq coverage and scATAC-seq insertion profiles along the genome from sequence at single-cell resolution. For this, we leverage the pre-trained multi-omics profile predictor Borzoi as a foundation model, equip it with a cell-specific decoder, and fine-tune its sequence embeddings. Specifically, we condition the decoder on the cell position in a precomputed single-cell embedding resulting in strong generalization capability. Applied to a hematopoiesis dataset, scooby recapitulates cell-specific expression levels of held-out genes, and identifies regulators and their putative target genes through in silico motif deletion. Moreover, accurate variant effect prediction with scooby allows for breaking down bulk eQTL effects into single-cell effects and delineating their impact on chromatin accessibility and gene expression. We anticipate scooby to aid unraveling the complexities of gene regulation at the resolution of individual cells.

Journal Article↗

Integrative metabolomic and proteomic analysis of diabetic kidney disease progression with younger-onset type 2 diabetes.

AIM: Younger-onset type 2 diabetes (YT2D) confers a disproportionately high risk of diabetic kidney disease (DKD), yet early biomarkers and underlying mechanisms remain poorly defined. We aimed to identify metabolites associated with DKD progression and integrate metabolomic and proteomic data to elucidate pathways involved in a multi-ethnic Asian cohort. MATERIALS AND METHODS: In this prospective study, 787 YT2D patients (diagnosed at ≤ age 40) were followed for a median of 5.7 years. DKD progression was defined as an annual decline in estimated glomerular filtration rate (eGFR) of ≥3 mL/min/1.73 m2 or ≥ 40% reduction in eGFR from baseline. Plasma metabolites were measured by nuclear magnetic resonance spectroscopy. Multivariable regression analysis was performed in a discovery (N = 550) and internal validation cohort (N = 237). Integrative metabolomic-proteomic analysis (N = 428) was performed using sparse partial least squares discriminant analysis (sPLS-DA). RESULTS: Ninety-eight metabolites were differentially expressed between DKD progressors and non-progressors, of which total branched-chain amino acids (BCAAs) (OR = 0.60, 95% CI 0.46-0.79), valine (OR = 0.62, 95% CI 0.48-0.81), and leucine (OR = 0.56, 95% CI 0.43-0.74) associated with DKD progression, independent of metabolic risk factors. Integrative analysis identified three components comprising 23 proteins and 30 metabolites, involved in the citrate cycle and apoptosis, which improved prediction of DKD progression beyond clinical risk factors (AUC 0.69-0.83). CONCLUSION: Lower plasma BCAA levels are independently associated with DKD progression in YT2D. Integrative multi-omics analysis highlights disruptions in metabolic and apoptotic pathways, providing insights into DKD pathophysiology and potential biomarkers for early risk stratification.

Humans↗

Algorithms and tools for data-driven omics integration to achieve multilayer biological insights: a narrative review.

Systems biology is a holistic approach to biological sciences that combines experimental and computational strategies, aimed at integrating information from different scales of biological processes to unravel pathophysiological mechanisms and behaviours. In this scenario, high-throughput technologies have been playing a major role in providing huge amounts of omics data, whose integration would offer unprecedented possibilities in gaining insights on diseases and identifying potential biomarkers. In the present review, we focus on strategies that have been applied in literature to integrate genomics, transcriptomics, proteomics, and metabolomics in the year range 2018-2024. Integration approaches were divided into three main categories: statistical-based approaches, multivariate methods, and machine learning/artificial intelligence techniques. Among them, statistical approaches (mainly based on correlation) were the ones with a slightly higher prevalence, followed by multivariate approaches, and machine learning techniques. Integrating multiple biological layers has shown great potential in uncovering molecular mechanisms, identifying putative biomarkers, and aid classification, most of the time resulting in better performances when compared to single omics analyses. However, significant challenges remain. The high-throughput nature of omics platforms introduces issues such as variable data quality, missing values, collinearity, and dimensionality. These challenges further increase when combining multiple omics datasets, as the complexity and heterogeneity of the data increase with integration. We report different strategies that have been found in literature to cope with these challenges, but some open issues still remain and should be addressed to disclose the full potential of omics integration.

Algorithms↗

Radiogenomics predicts immune microenvironment heterogeneity and response to combination immunotherapy in hepatocellular carcinoma.

BACKGROUND: The combination of immune checkpoint inhibitors (ICIs) with anti-angiogenic agents is the preferred first-line therapy option for patients with advanced hepatocellular carcinoma (HCC), yet only a subset of patients responds, urging the quest for prediction biomarkers. We aimed to integrate genomics with radiology to propose an immune-derived radiogenomics biomarker of response to such combination immunotherapy and evaluate its added value in clinical context. METHODS: We integrated bulk RNA sequencing (RNA-seq) and proteomics data of 994 HCC patients with single-cell RNA-seq data of 11 samples across multiple datasets to identify an immune-related signature (IRS) that may influence sensitivity or resistance to such combined immunotherapy strategy, followed by verification of selected marker genes using immunohistochemistry and cytological experiments. We then trained/validated a cross-modality radiogenomics biomarker using machine learning based on TCIA database that was further tested in multi-scale independent cohorts covering 754 HCC patients. RESULTS: Integrative multi-omics analysis identifed a parsimonious 2-gene prognostic signature including KPNA2 and SMG5 that was significantly associated with immune heterogeneity and response to combination immunotherapy. Machine-learning pipeline exported the optimal 4-feature radiogenomics biomarker using support vector machine that significantly discriminated prognosis (hazard ratio 1.415&#x2013;1.890; p&#x2009;<&#x2009;0.05 for all) and modestly predicted response to ICI plus anti-angiogenic therapy (area under the curve 0.720&#x2013;0.829) in independent retrospective series across major imaging modalities (computed tomography/magnetic resonance imaging). In a prospective neoadjuvant cohort, this biomarker also showed favorable performance for predicting pathological response and tumor recurrence, accompanied by biological validation through single-cell RNA-seq analysis of pre-treatment biopsies. CONCLUSIONS: Our study provides a cross-device-cross-modal radiogenomics biomarker that can improve patient selection for emerging ICI plus anti-angiogenic therapy with novel potential therapeutic targets in HCC.

Humans↗

Rhythm profiling using COFE reveals multi-omic circadian rhythms in human cancers in vivo.

The study of ubiquitous circadian rhythms in human physiology requires regular measurements across time. Repeated sampling of the different internal tissues that house circadian clocks is both practically and ethically infeasible. Here, we present a novel unsupervised machine learning approach (COFE) that can use single high-throughput omics samples (without time labels) from individuals to reconstruct circadian rhythms across cohorts. COFE can simultaneously assign time labels to samples and identify rhythmic data features used for temporal reconstruction, while also detecting invalid orderings. With COFE, we discovered widespread de novo circadian gene expression rhythms in 11 different human adenocarcinomas using data from The Cancer Genome Atlas (TCGA) database. The arrangement of peak times of core clock gene expression was conserved across cancers and resembled a healthy functional clock except for the mistiming of a few key genes. Moreover, rhythms in the transcriptome were strongly associated with the cancer-relevant proteome. The rhythmic genes and proteins common to all cancers were involved in metabolism and the cell cycle. Although these rhythms were synchronized with the cell cycle in many cancers, they were uncoupled with clocks in healthy matched tissue. The targets of most of FDA-approved and potential anti-cancer drugs were rhythmic in tumor tissue with different amplitudes and peak times. These findings emphasize the utility of considering "time" in cancer therapy, and suggest a focus on clocks in healthy tissue rather than free-running clocks in cancer tissue. Our approach thus creates new opportunities to repurpose data without time labels to study circadian rhythms.

Humans↗

Anticancer drug response prediction integrating multi-omics pathway-based difference features and multiple deep learning techniques.

Individualized prediction of cancer drug sensitivity is of vital importance in precision medicine. While numerous predictive methodologies for cancer drug response have been proposed, the precise prediction of an individual patient's response to drug and a thorough understanding of differences in drug responses among individuals continue to pose significant challenges. This study introduced a deep learning model PASO, which integrated transformer encoder, multi-scale convolutional networks and attention mechanisms to predict the sensitivity of cell lines to anticancer drugs, based on the omics data of cell lines and the SMILES representations of drug molecules. First, we use statistical methods to compute the differences in gene expression, gene mutation, and gene copy number variations between within and outside biological pathways, and utilized these pathway difference values as cell line features, combined with the drugs' SMILES chemical structure information as inputs to the model. Then the model integrates various deep learning technologies multi-scale convolutional networks and transformer encoder to extract the properties of drug molecules from different perspectives, while an attention network is devoted to learning complex interactions between the omics features of cell lines and the aforementioned properties of drug molecules. Finally, a multilayer perceptron (MLP) outputs the final predictions of drug response. Our model exhibits higher accuracy in predicting the sensitivity to anticancer drugs comparing with other methods proposed recently. It is found that PARP inhibitors, and Topoisomerase I inhibitors were particularly sensitive to SCLC when analyzing the drug response predictions for lung cancer cell lines. Additionally, the model is capable of highlighting biological pathways related to cancer and accurately capturing critical parts of the drug's chemical structure. We also validated the model's clinical utility using clinical data from The Cancer Genome Atlas. In summary, the PASO model suggests potential as a robust support in individualized cancer treatment. Our methods are implemented in Python and are freely available from GitHub (https://github.com/queryang/PASO).

Deep Learning↗

The Progress of Gout Prediction Models Based on Multi-source Data.

INTRODUCTION: Gout, a highly serious inflammatory disease that is caused by monosodium urate crystals, is becoming an increasingly significant health concern. Artificial Intelligence and multi-omics-based research have made significant gains for the early detection and prevention of gout based on diverse approaches. This review intends to summarize current advances in forecasting gout susceptibility and gout-related symptoms, evaluate the predictive efficacy of different features, and ascertain which clinical and omics characteristics are most effective in these prediction models. METHODS: We explored the PubMed database after 2010 using keywords such as "gout", "predictive model", "risk prediction", and "machine learning", and confined our search to Englishlanguage articles. The original peer-reviewed research articles that developed gout models were selected. Research that was not original or lacked internal validation was excluded. RESULTS: Clinical features, genomics, microbiomics, radiomics, and metabolomics have been utilized to construct models related to gout and have demonstrated excellent predictive performance. Multisource data prediction models usually exhibit better effectiveness. DISCUSSION: Gout-oriented models performed excellently in predictive performance but present limitations in certain clinical and omics domains. However, if they are to affect actual patient care, they must overcome some external confirmation roadblocks and the fiscal and practical implications they will face ahead of time. CONCLUSION: This review indicates that clinical and multi-omics models of gout are significant instruments for clinical decision-making. The models constructed in these studies may be crucial for the treatment of gout and its practical benefits.

Gout↗

Bioinformatics in crop research: using genomic data for crop improvement.

Sustainable crop development aims to maintain or increase yields while reducing environmental impact and managing the challenges imposed by climate change. As the global population grows and arable land becomes scarcer, the integration of molecular breeding with bioinformatics has emerged as an effective strategy for long-term crop improvement. Bioinformatics enables researchers to analyze and interpret the vast quantities of genetic data generated by high-throughput sequencing, making it possible to identify molecular markers, candidate genes, and regulatory networks linked to specific agronomic traits, which breeders then translate into focused, ecologically sustainable breeding programs. This approach has enabled major progress across several fronts: the identification of genes conferring resistance to biotic stressors (pests, pathogens) and abiotic stressors (drought, salinity, heat); the development of nutrient-efficient, low-input crop varieties; the improvement of agronomic performance and nutritional quality through identification of yield- and quality-related genes; and the conservation and deployment of genetic diversity to safeguard long-term breeding sustainability. By combining genomic data with precision breeding techniques, researchers are developing crops that are better adapted to a growing population and a changing climate, positioning the integration of molecular breeding and bioinformatics as a central pillar of future global food security.

bioinformatics↗