PubMed HealthSearch

SEARCH · PubMed Health

Results for “High-dimensional”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Whole-genome phenotype prediction with machine learning: open problems in bacterial genomics.

MOTIVATION: How can we identify causal genetic mechanisms governing bacterial traits? Initial efforts entrusting machine learning models to handle the task of predicting phenotype from genotype yield high accuracy scores. However, attempts to extract meaningful interpretations from the predictive models are found to be corrupted by falsely identified 'causal' features. Relying solely on pattern recognition and correlations is unreliable, significantly so in bacterial genomics settings where high-dimensionality and spurious associations are the norm. Though it is not yet clear whether we can overcome this hurdle, significant efforts are being made towards discovering potential high-risk bacterial genetic variants. In view of this, we set up open problems surrounding phenotype prediction from bacterial whole-genome datasets and extending those approaches to learning causal effects, and discuss challenges that impact the reliability of a machine's decision-making when faced with datasets of this nature. RESULTS: We identify major sources of non-injectivity in the formulation of the genotype-to-phenotype mapping function-linkage-disequilibrium, limited sampling, information loss in representations, unmeasured confounders and observational noise-and analyse their implications for machine learning applications. Using a collection of 4,140 Staphylococcus aureus isolates, we illustrate challenges surrounding the defined open problems. AVAILABILITY AND IMPLEMENTATION: Raw sequencing data are available from the European Nucleotide Archive (ENA) under project accessions ERP001012, PRJEB3174, PRJEB2655, PRJEB2756, and PRJEB2944. Assemblies and annotations were generated with the Sanger bacterial pipeline (https://github.com/sanger-pathogens/vr-codebase) and unitigs extracted using DBGWAS (https://gitlab.com/leoisl/dbgwas).

Machine Learning

MO-GCAN: multi-omics integration based on graph convolutional and attention networks.

MOTIVATION: Cancer subtypes play a critical role in disease progression, prognosis, and treatment, making their detection essential for tailoring precision medicine. Studies have shown that multi-omics integration outperforms single-omics approaches in cancer subtyping tasks. However, due to the high-dimensionality of multi-omics data, many existing studies either fail to capture the correlation between true labels and learned features, or lack sufficient capacity to model complex biological representations. These limitations hinder the full potential of leveraging the rich and complementary information embedded in multi-omics datasets. RESULT: We propose a framework that leverages supervised feature learning and classification based on a graph-based learning approach with attention mechanism for cancer subtyping. More specifically, we train graph convolutional network models on each omics dataset to extract latent representations, which are then concatenated to form a comprehensive multi-omics feature embedding. We further develop sample fusion network based on the omics-specific graphs, incorporating the derived features and feeding them into a graph attention model for subtype classification. This two-stage multi-omics framework is applied to eight cancer types, with performance evaluated in terms of test accuracy, training time, macro-averaged precision, recall, and F-score. Experimental results show that the proposed method outperforms state-of-the-art approaches across various cancer types. Additionally, we provide empirical evidence supporting the hypothesis that retaining a limited number of high-confidence edges and utilizing enriched embeddings from intermediate graph neural network layers can improve predictive performance. AVAILABILITY AND IMPLEMENTATION: Data and the code are available at https://github.com/YD-00/MO-GCAN-Updated.git.

Neoplasms

LimROTS: a hybrid method integrating empirical Bayes and reproducibility-optimized statistics for robust differential expression analysis.

MOTIVATION: Differential expression analysis plays a vital role in omics research enabling precise identification of features that associate with different phenotypes. This process is critical for uncovering biological differences between conditions, such as disease versus healthy states. In proteomics, several statistical methods have been used, ranging from simple t-tests to more advanced methods like DEqMS, limma and ROTS. However, a flexible method for reproducibility-optimized statistics tailored for clinical omics data has been lacking. RESULTS: In this study, we developed LimROTS, a hybrid method that integrates a linear regression model and the empirical Bayes approach with reproducibility optimized statistics, to create a novel moderated ranking statistic, for robust and flexible analysis of proteomics data. We validated its performance using twenty-one proteomics gold standard spike-in datasets with different protein mixtures, MS instruments, and techniques for benchmarking. This hybrid approach improves accuracy and reproducibility of complex proteomics data, making LimROTS a powerful tool for high-dimensional omics data analysis. AVAILABILITY AND IMPLEMENTATION: LimROTS has been implemented as an R/Bioconductor package, available at https://doi.org/doi:10.18129/B9.bioc.LimROTS. Additionally, the code used in this study is available in GitHub repository https://github.com/AliYoussef96/LimROTSmanuscript.

Bayes Theorem

Sparse polygenic risk score inference with the spike-and-slab LASSO.

MOTIVATION: Large-scale biobanks, with rich phenotypic and genomic data across hundreds of thousands of samples, provide ample opportunities to elucidate the genetics of complex traits and diseases. Consequently, there is growing demand for robust and scalable methods for disease risk prediction from genotype data. Inference in this setting is challenging due to the high-dimensionality of genomic data, especially when coupled with smaller sample sizes. Popular Polygenic Risk Score (PRS) inference methods address this challenge by adopting sparse Bayesian priors or penalized regression techniques, such as the Least Absolute Shrinkage and Selection Operator (LASSO). However, the former class of methods are not as scalable and do not produce exact sparsity, while the latter tends to over-shrink large coefficients. RESULTS: In this study, we present SSLPRS, a novel PRS method based on the Spike-and-Slab LASSO (SSL) prior, which offers a theoretical bridge between the two frameworks. We extend previous work to derive a coordinate-ascent inference algorithm that operates on GWAS summary statistics, which is orders-of-magnitude more efficient than corresponding individual-level-based implementations. To illustrate the statistical properties of the proposed model, we conducted experiments involving nine simulation configurations and nine quantitative phenotypes from the UK Biobank. Our results demonstrate that SSLPRS is competitive with state-of-the-art methods in terms of prediction accuracy and exhibits superior variable selection performance, especially in sparse genetic architectures. In simulations, this translates to upwards of 50% improvement in positive predictive value. In analysis of real phenotypes, we show that selected variants are highly enriched for meaningful genomic annotations and have better replication rates in larger meta-analyses. AVAILABILITY AND IMPLEMENTATION: SSLPRS is available in the open-source package https://github.com/li-lab-mcgill/penprs.

Multifactorial Inheritance

Malaria-GENOMAP: a web-based tool for exploring genomic variation of malaria parasites.

MOTIVATION: Malaria, caused by Plasmodium parasites, imposes a significant public health burden. While Plasmodium falciparum remains the primary target of elimination strategies due to its high mortality rate, lesser-known species such as P. malariae, P. vivax, and P. knowlesi continue to contribute to substantial human morbidity. Genomic approaches, including whole-genome sequencing, offer powerful tools for understanding the biology, transmission, and emerging drug resistance of these neglected Plasmodium species. However, there is an urgent need for informatic tools to summarize and visualize the high-dimensional and complex genomic data generated. RESULTS: We developed Malaria-GENOMAP, a user-friendly web-based tool, which integrates genomic variant data, such as allele frequencies, with geographical maps and chromosome-wide to gene views for in-depth exploration. The tool includes variation from P. knowlesi (n = 139), P. malariae (n = 158), P. ovale curtisi (n = 36), P. ovale wallikeri (n = 47), P. simium (n = 38), and P. vivax (n = 1359). It enables the investigation of population structure, geographic associations of mutations, and putative drug resistance markers, offering valuable insights for malaria control efforts. AVAILABILITY AND IMPLEMENTATION: Malaria-GENOMAP is available online at https://genomics.lshtm.ac.uk/malaria-genomaps.

Internet

PEARL: integrative multi-omics classification and omics feature discovery via deep graph learning.

MOTIVATION: Integrating multi-omics data provides valuable insights into biological processes by capturing information across multiple molecular layers, enabling a comprehensive understanding of complex diseases and driving advancements in precision medicine. However, existing computational methods for multi-omics integration face significant challenges, such as low reliability and poor generalizability, due to the high dimensionality and low sample size nature of omics data. RESULTS: To address these challenges, we present PEARL (Pearson-Enhanced spectrAl gRaph convoLutional networks), a novel deep graph learning method for biomedical classification and functional important omics features identification. PEARL leverages a simple yet effective learning architecture to achieve superior and robust performance in high-dimensional, low-sample-size multi-omics settings. Our results demonstrate that PEARL significantly outperforms existing state-of-the-art methods on both synthetic and real biomedical datasets. Furthermore, applied to Alzheimer's disease (AD) brain multi-omics data, features prioritized by PEARL lead to functionally important genes that demonstrate significant enrichment in AD-related pathways. These findings highlight PEARL's practical utility in biomedical research and its potential to enhance biological interpretability in multi-omics studies. AVAILABILITY AND IMPLEMENTATION: The source code of our computational framework is available at https://github.com/zqq121017/PEARL.

Multiomics

Odon: an ultra-fast viewer for spatial proteomics.

MOTIVATION: Multiplexed spatial proteomics and spatial transcriptomics generate large, high-dimensional imaging datasets that are challenging to visualize efficiently, particularly at whole-slide and cohort scale. Visualization is an essential step for rapid detection of staining artefacts, such as protein aggregates or non-specific staining. RESULTS: Here, we present Odon, a native Rust desktop viewer designed for rapid, interactive exploration of multiplex imaging data on a standard laptop. Odon is primarily built around the OME-Zarr imaging format, and supports annotations via GeoJSON and GeoParquet, with secondary support for SpatialData, Xenium containers, and TIFF. Data can be stored locally or streamed directly from HTTP or S3-compatible object storage using viewport-driven tile loading. Odon incorporates a highly optimized rendering engine designed for viewport-driven tile loading and GPU-based compositing. In scripted benchmarks using synthetic multiplex OME-Zarr datasets, Odon showed lower peak memory use, lower affine-derived zoom-step error, and faster warm-start image loading than napari and QuPath under the tested conditions. Its GPU-based compositing pipeline also enables smooth rendering and interaction with >1 000 000 segmented cells. Odon further supports integrated visual analytics, including live thresholding and cell selection, and a mosaic mode for simultaneous viewing of hundreds of regions of interest in cohort and tissue microarray studies. Together, these features establish Odon as a high-performance platform for scalable visualization of spatial proteomics data. AVAILABILITY AND IMPLEMENTATION: Source code and compiled installers are available at https://github.com/alexcoulton/odon.

Proteomics

Dynamic Alterations in the Blood Transcriptome Characterize Drug Use Behavior and Co-Morbidities in Cocaine Use Disorder: A Preliminary Study.

Individuals with cocaine use disorder (CUD) who attempt abstinence experience craving and relapse that can benefit from multimodal treatment monitoring. Longitudinal studies linking behavioral manifestations in CUD to the blood transcriptome are not only limited but also computationally complex. Therefore, we developed an analytical pipeline to investigate the connection between drug use behaviors during abstinence and change in the blood transcriptome. We conducted a longitudinal study with CUD (n = 12 subjects) and collected behavioral metrics and blood RNA-seq at baseline, 3, 6, and 9 months. Our analytical pipeline of the high-dimensional data encompasses hierarchical k-means clustering to classify subjects to responder groups based on behavioral scores and abstinence duration, in silico cell deconvolution, differential analysis with correlated multivariate testing over time, gene set enrichment analysis, and gene co-expression with time splines and RNA-seq data. The pipeline captured dynamic changes in behavioral scores and abstinence duration in responder groups. Genes showing differential transcript-level expression were enriched in substance use and cardiovascular disease-associated genetic risk loci in responder groups. Lastly, time-dependent gene co-expression revealed dynamic changes related to immune processes, cell cycle, RNA-protein synthesis, and second messenger signaling for days of abstinence. This is a preliminary investigation, providing an innovative and scalable pipeline for blood-based longitudinal RNA-seq studies in CUD, potentially applicable to other substance use disorders. It outlines a data-driven approach for analyzing composite longitudinal drug use behavioral phenotypes with blood-based transcriptomics. We also demonstrate changes in drug use behaviors and the blood transcriptome during drug abstinence.

Humans

Mul-PheG2P: decoupled learning and prediction-space fusion enables robust and interpretable multi-phenotype genomic prediction.

Genomic prediction of multiple phenotypes is crucial in modern plant breeding; however, existing methods struggle with negative transfer and lack interpretability, particularly across high-dimensional small-sample data and diverse species. To address this, we propose Mul-PheG2P, a novel paradigm based on decoupled learning and predictive space fusion. It employs a two-stage design: first training phenotype-specific encoders using genetic data, then decoupling phenotype-specific learning from cross-phenotype aggregation via an interpretable prediction layer. Mul-PheG2P outperforms existing methods across diverse crop datasets, including maize (Zea mays), wheat (Triticum aestivum), and tomato (Solanum lycopersicum). It provides a multi-scale interpretability chain: at the macro level, it quantifies phenotypic contributions via attention-based weighting; at the micro level, Integrated Gradients reveal the genetic basis of predictions. Notably, the model successfully identified the CCT (CONSTANS, CO-like, and TOC) motif regulating photoperiodism and the SQUAMOSA (SQUAMOSA promoter binding protein) promoter for inflorescence development, confirming its ability to capture functional biological mechanisms. These results highlight the high performance and interpretability of Mul-PheG2P, showcasing its value for low-cost, large-scale screening to advance precision breeding.

Phenotype

Multiomics approaches to cardiovascular disease: technological innovations and clinical translation.

Cardiovascular diseases (CVDs) remain the leading cause of global morbidity and mortality, reflecting a persistent gap between clinical phenotyping and the molecular mechanisms that govern disease initiation, progression, and interindividual variability. Recent advances in emerging technologies have fundamentally reshaped cardiovascular physiology by enabling high-resolution, cross-layer profiling of the heart and vasculature across genomic, epigenomic, transcriptomic, proteomic, metabolomic, lipidomic, glycomic, and fluxomic layers, increasingly at single-cell and spatial resolution. These approaches reveal CVD as a coordinated, multilayered process driven by dynamic interactions among cell types, regulatory programs, and metabolic states, rather than isolated gene-level defects. In this review, we synthesize how emerging multiomic, computational, and functional genomic technologies are redefining the study of cardiovascular disease across molecular, cellular, and tissue levels. We highlight recent innovations in single-cell and spatial atlases, long-read sequencing, proteomics and metabolomics, integrative data modeling, and functional omics approaches, including genome-scale perturbation screens and single-cell perturbation frameworks. These platforms enable mechanistic dissection of regulatory circuits, distinguish primary disease drivers from secondary adaptations, and directly assess therapeutic reversibility, advancing the field beyond associative biomarker discovery toward mechanism-guided target prioritization. We further discuss key methodological and translational challenges accompanying high-dimensional cardiovascular data, including preanalytical variability, control selection, temporal misalignment across molecular layers, population diversity, and reference bias. By integrating technological innovation with computational rigor and functional validation, this review frames emerging omics-enabled strategies as a unified, physiologically grounded framework for translating molecular insight into clinically meaningful cardiovascular phenotypes and advancing precision cardiovascular medicine.

Humans

Integrated functional genomics and safety assessment of plant-growth-promoting Caryophanales from post-maize-cultivation soils.

This study aimed to evaluate six environmental bacterial strains isolated from post-maize cultivation soils as candidates for agricultural biopreparation development, using an integrated functional genomic and safety assessment framework. Building on experimental validation of plant-growth-promoting activities, the analysis included: plant-growth-promoting traits (PGPT-Pred) using PLABase; carbohydrate-active enzymes (CAZymes) relevant for lignocellulosic crop residue degradation (dbCAN3); secondary metabolite profiles (antiSMASH); and screening for virulence factors and antibiotic resistance genes (ABRicate, BTyper3).All analyzed strains possess 1,449-1,617 predicted PGPT-encoding genes (24.1-35.9% of total genes), which are strongly shaped by taxonomic relatedness, as confirmed by congruence testing against ANI-based genomic divergence. Paenibacillus amylolyticus 5mez and Priestia megaterium 7psych showed distinct functional profiles compared to Bacillus spp., while Bacillus subtilis sensu lato strains were most similar to each other. Genomic predictions suggest involvement in nutrient acquisition (N, P, K, Fe) and stress mitigation. Secondary metabolite analysis revealed high biosynthetic potential, with non-Bacillus species harbouring a large proportion of unknown gene clusters, indicating underexplored metabolite diversity. CAZyme profiling identified P. amylolyticus 5mez as the most enzyme-rich strain, while B. cereus s.s. zielonkawy showed ligninolytic potential despite low overall CAZyme abundance. The safety assessment identified B. cereus s.s. zielonkawy as toxigenic and unsuitable for use. Of the remaining strains, P. amylolyticus 5mez and Pr. megaterium 7psych demonstrated the most favourable safety profiles, exhibiting no detectable virulence factors or antibiotic resistance genes, justifying their priority use in agricultural biopreparations, pending phenotypic validation. Given the high-dimensional, low-sample-size nature of multi-trait datasets in applied microbial genomics, tailored statistical approaches, including noise-reduction-validated PCA and distance-based congruence testing, were applied; their rationale and limitations are discussed.

Soil Microbiology

A weakly supervised deep learning-based recurrence prediction and risk stratification of lung adenocarcinoma from pathology whole-slide images.

BACKGROUND: Accurate prediction of postoperative recurrence in lung adenocarcinoma (LUAD) is essential for guiding clinical decision-making and improving patient outcomes. Although various predictive models have been developed, most rely on complex genomic analyses and high-dimensional clinical data. The complexity of these approaches substantially limits their feasibility for routine clinical use. To address this clinical challenge, this study aims to predict postoperative recurrence using routinely available hematoxylin and eosin (H&E)-stained images and characterize the associated biological features. METHODS: A total of 329 patients who underwent curative resection at the First Affiliated Hospital of Wenzhou Medical University (FHWMU) were retrospectively enrolled and randomly assigned to training and internal validation cohorts in a 7:3 ratio. An independent external validation cohort comprising 70 patients from the Clinical Proteomic Tumor Analysis Consortium (CPTAC) was included. Three patch-level feature extractors (Inception_V3, ResNet18, and DenseNet121) were evaluated within a weakly supervised multiple-instance learning (MIL) framework incorporating automated region-of-interest (ROI) detection on segmented whole-slide images (WSIs). Model performance was assessed using the area under the receiver operating characteristic curve (AUC), Kaplan-Meier (KM) survival analysis, and multivariable Cox proportional hazards regression. Transcriptomic profiling and gene set enrichment analysis (GSEA) were conducted to investigate biological differences between risk groups. RESULTS: The model achieved AUCs of 0.923 in the training cohort, 0.891 in the internal validation cohort, and 0.847 in the external validation cohort. The model effectively stratified patients into high- and low-risk groups with significantly different recurrence-free survival (RFS) across all cohorts (all P&#x2009;<&#x2009;0.001) and retained prognostic value within AJCC stages I-III. Transcriptomic analyses revealed consistent enrichment of cell cycle-related pathways and neutrophil extracellular trap (NET) formation in high-risk patients across both institutional and CPTAC cohorts, aligning with distinct biological profiles of the model-derived risk stratification. CONCLUSIONS: This weakly supervised deep learning framework enables accurate and externally validated prediction of postoperative recurrence in LUAD using routinely available histopathological images, and integration of histopathological features with molecular analyses enhances biological interpretability. This work provides a clinically accessible and cost-effective tool for postoperative risk assessment in LUAD patients.

Humans

Exploring the use of machine and deep learning in genome-wide association studies: a comprehensive review.

The advent of high-throughput sequencing technologies has generated increasingly large and complex genomic datasets, necessitating analytical approaches capable of capturing high-dimensional and potentially nonlinear genetic interactions. This situation has significantly impacted the entire field of Genome-Wide Association Study (GWAS), whose primary goal is the identification of genomic traits and variants that are statistically associated with the risk of a disease. However, traditional GWAS methods may show reduced performance when applied to highly polygenic and nonlinear genetic architectures. Computational strategies from Artificial Intelligence (AI) and, in particular, from machine- and deep-learning may provide a powerful tool to overcome such limitations, especially by capturing nonlinear interactions and complex hidden regularities in large-scale data, which traditional GWAS approaches might overlook. To date, only a few approaches have been introduced and systematically assessed. In this review, we describe the main characteristics and limitations of standard statistical approaches for GWAS, the main uses of AI methods in computational genomics, and recent attempts to leverage AI strategies in GWAS. Particular attention will be devoted to key issues, such as the interpretability of methods and results, and the curse of dimensionality. More specifically, the review presents 30 methods designed to leverage AI in GWAS, as well as presenting a comprehensive set of evaluation metrics for their performance, also providing references to the most frequently used databases, and biobanks. Overall, this work may serve as a starting point for both dry- and wet-lab researchers, aiming to extract deeper insights from genomic data by moving beyond traditional linear additive assumptions, and leveraging large-scale datasets through AI-driven approaches.

Artificial intelligence

[Applications and Challenges of Deep Learning in Human Genome Research].

In recent years, the advent of high-throughput omics technologies has fueled an explosive growth in human genomic data. Uncovering the latent functions within this vast data has become a significant challenge in functional genomics research. While traditional statistical methods have proved successful for analyzing smaller-scale datasets in the past, they exhibit clear limitations in analytical efficiency and integrating multi-dimensional data, struggling to meet the escalating demands of contemporary genomic analysis. The introduction of deep learning (DL) technologies offers a novel paradigm for this field. This review systematically examines the advances in applying deep learning to human genomics research. Studies demonstrate that when ample labeled data is available, discriminative DL computational methods-such as Convolutional Neural Networks (CNNs) and Long Short-Term Memory networks (LSTMs)-achieve high accuracy and efficiency in genomic variant discovery tasks. Furthermore, generative DL methods, particularly Large Language Models (LLMs) leveraging self-supervised pre-training strategies, effectively integrate complex genomic information and exhibit superior performance in functional genomic sequence annotation and gene regulation studies. This review also explores the application of LLMs in multi-omics data integration and prediction. Looking ahead, the continued accumulation of long-read sequencing and high-dimensional data is expected to enable DL technologies to integrate increasingly complex and heterogeneous genomic information, playing an increasingly crucial role in human genomics research.

Deep Learning

Tensor decomposition of multi-dimensional splicing events across multiple tissues to identify splicing-mediated risk genes associated with complex traits.

Identifying risk genes associated with complex traits remains challenging. Integrating gene expression data with Genome-Wide Association Study (GWAS) through Transcriptome-Wide Association Study (TWAS) methods has discovered candidate risk genes for various complex traits. Splicing, which explains a comparable heritability of complex traits as gene expression, is&#xa0;under-explored&#xa0;due to its multidimensionality. To leverage multiple splicing events in a gene and shared splicing across tissues, we develop Multi-tissue Splicing Gene (MTSG), which employs tensor decomposition and sparse Canonical Correlation Analysis (sCCA) to extract meaningful information from high-dimensional multiple splicing events across multiple tissues. We build MTSG models using GTEx data and apply them to GWAS summary statistics of Alzheimer's disease (AD) (111,326 cases and 677,663 controls) and schizophrenia (SCZ) (36,989 cases and 113,075 controls). We identify 174 and 497 significant splicing-mediated risk genes for AD and SCZ, respectively, at Bonferroni correction. For AD, our results demonstrate significant enrichment of AD related pathways and identify additional AD risk genes not detected in the single-tissue analysis, while preserving most top genes identified in the brain frontal cortex. Consistently, for SCZ, genes identified by our brain-wide MTSG model, built from a cluster of 13 brain tissues, exhibit stronger enrichment in SCZ-relevant genes and MTSG identifies unique SCZ risk genes compared to single-tissue models. These results showcase that our MTSG models capture distinctive splicing events across tissues, which might be overlooked when using single tissue alone. Our MTSG models can be applied to other complex traits to help identify splicing-mediated disease risk genes.

Humans

A novel and robust feature selection method with FDR control for omics-wide association analysis.

Omics-wide association analysis is a very important tool for medicine and human health study. However, the modern omics data sets collected often exhibit the high-dimensionality, unknown distribution response, unknown distribution features and unknown complex association relationships between the response and its explanatory features. Reliable association analysis results depend on an accurate modeling for such data sets. Most of the existing association analysis methods rely on the specific model assumptions and lack effective false discovery rate (FDR) control. To address these limitations, the paper firstly applies a single index model for omics data. The model shows robust performance in allowing the relationships between the response variable and linear combination of covariates to be connected by any unknown monotonic link function, and both the random error and the covariates can follow any unknown distribution. Then based on this model, the paper combines rank-based approach and symmetrized data aggregation approach to develop a novel and robust feature selection method for achieving fine-mapping of risk features while controlling the false positive rate of selection. The theoretical results support the proposed method and the analysis results of simulated data show the new method possesses effective and robust performance for all the scenarios. The new method is also used to analyze the two real datasets and identifies some risk features unreported by the existing finds.

Humans

A framework for block-wise missing data in multi-omics.

High-throughput technologies have generated vast amounts of omic data. It is a consensus that the integration of diverse omics sources improves predictive models and biomarker discovery. However, managing multiple omics data poses challenges such as data heterogeneity, noise, high-dimensionality and missing data, especially in block-wise patterns. This study addresses the challenges of high dimensionality and block-wise missing data through a regularization and constrained-based approach. The methodology is implemented in the R package bwm for binary and continuous response variables, and applied to breast cancer and exposome multi-omics datasets, achieving strong performance even in scenarios with missing data present in all omics. In binary classification task, our proposed model achieves accuracy in the range of 86% to 92%, and F1 in the range of 68% to 79%. And, in regression task the correlation between true and predicted responses is in the range of 72% to 76%. However, there is a slight decline in performance metrics as the percentage of missing data increases. In scenarios where block-wise missing data affects multiple omics, the model performance actually surpasses that of scenarios where missing data is present in only one omics. One possible explanation for this might be that the other scenarios introduce a greater diversity of observation profiles, leading to a more robust model. Depending on the specific omics being studied, there is greater consistency in feature selection when comparing block-wise missing data scenarios.

Humans

Development of methodology to support molecular endotype discovery from synovial fluid of individuals with knee osteoarthritis: The STEpUP OA consortium.

OBJECTIVES: To develop a protocol for largescale analysis of synovial fluid proteins, for the identification of biological networks associated with subtypes of osteoarthritis. METHODS: Synovial Fluid To detect molecular Endotypes by Unbiased Proteomics in Osteoarthritis (STEpUP OA) is an international consortium utilising clinical data (capturing pain, radiographic severity and demographic features) and knee synovial fluid from 17 participating cohorts. 1746 samples from 1650 individuals comprising OA, joint injury, healthy and inflammatory arthritis controls, divided into discovery (n = 1045) and replication (n = 701) datasets, were analysed by SomaScan Discovery Plex V4.1 (>7000 SOMAmers/proteins). An optimised approach to standardisation was developed. Technical confounders and batch-effects were identified and adjusted for. Poorly performing SOMAmers and samples were excluded. Variance in the data was determined by principal component (PC) analysis. RESULTS: A synovial fluid standardised protocol was optimised that had good reliability (<20% co-efficient of variation for >80% of SOMAmers in pooled samples) and overall good correlation with immunoassay. 1720 samples and >6290 SOMAmers met inclusion criteria. 48% of data variance (PC1) was strongly correlated with individual SOMAmer signal intensities, particularly with low abundance proteins (median correlation coefficient 0.70), and was enriched for nuclear and non-secreted proteins. We concluded that this component was predominantly intracellular proteins, and could be adjusted for using an 'intracellular protein score' (IPS). PC2 (7% variance) was attributable to processing batch and was batch-corrected by ComBat. Lesser effects were attributed to other technical confounders. Data visualisation revealed clustering of injury and OA cases in overlapping but distinguishable areas of high-dimensional proteomic space. CONCLUSIONS: We have developed a robust method for analysing synovial fluid protein, creating a molecular and clinical dataset of unprecedented scale to explore potential patient subtypes and the molecular pathogenesis of OA. Such methodology underpins the development of new approaches to tackle this disease which remains a huge societal challenge.

Humans