PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “High-dimensional”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Natural killer (NK) cells with downregulated activating receptors and IL-10 production promote Trypanosoma cruzi T-cell responses in subjects in the Chronic phase of Trypanosoma cruzi infection: An exploratory immunological analysis.

BACKGROUND: Subjects with chronic Chagas disease and no signs of heart disease exhibit decreased NKp46 expression, high CD57 expression, and IL-10 production by NK cells. This study provides a detailed characterization of the phenotype and function of NK cells according to the severity of heart disease, and evaluates how these changes following treatment with benznidazole, as well as their association with Trypanosoma cruzi-specific T-cell responses. METHODS: The phenotype and function of NK cells in a cohort of 51 subjects infected with Trypanosoma cruzi and exhibiting varying degrees of heart disease were evaluated using high-dimensional flow cytometry and Boolean gating analysis. ELISPOT assays were performed to measure IFN-γ and IL-2 production in response to T. cruzi antigens, using CD56, CD4 and CD8-depleted PBMC. RESULTS: In contrast to individuals without heart disease, those with advanced cardiomyopathy have an increased number of NK cells that express the activating receptors NKp46 and CD16, as well as the differentiation marker CD57. NK cells in subjects with advanced cardiomyopathy were also found to be enriched in CD56+CD107+granzyme B+TNF-α+ cells and depleted of CD56+IL-10+ cells. The function of NK cells shifted towards monofunctionality in subjects with declining T. cruzi-specific antibodies following treatment with benznidazole. Eliminating CD56+ cells significantly decreased the number of IFN-γ-producing cells in response to T. cruzi. CONCLUSIONS: In chronic T. cruzi infection, NK-cell function may be balanced by the downregulation of activating receptors and IL-10 production, However, parasite persistence may desensitize the regulation of NK activating receptors, enabling NK cells to exert a potent polyfunctional cytotoxicity that potentially induce tissue damage.

Humans↗

A novel and robust feature selection method with FDR control for omics-wide association analysis.

Omics-wide association analysis is a very important tool for medicine and human health study. However, the modern omics data sets collected often exhibit the high-dimensionality, unknown distribution response, unknown distribution features and unknown complex association relationships between the response and its explanatory features. Reliable association analysis results depend on an accurate modeling for such data sets. Most of the existing association analysis methods rely on the specific model assumptions and lack effective false discovery rate (FDR) control. To address these limitations, the paper firstly applies a single index model for omics data. The model shows robust performance in allowing the relationships between the response variable and linear combination of covariates to be connected by any unknown monotonic link function, and both the random error and the covariates can follow any unknown distribution. Then based on this model, the paper combines rank-based approach and symmetrized data aggregation approach to develop a novel and robust feature selection method for achieving fine-mapping of risk features while controlling the false positive rate of selection. The theoretical results support the proposed method and the analysis results of simulated data show the new method possesses effective and robust performance for all the scenarios. The new method is also used to analyze the two real datasets and identifies some risk features unreported by the existing finds.

Humans↗

A framework for block-wise missing data in multi-omics.

High-throughput technologies have generated vast amounts of omic data. It is a consensus that the integration of diverse omics sources improves predictive models and biomarker discovery. However, managing multiple omics data poses challenges such as data heterogeneity, noise, high-dimensionality and missing data, especially in block-wise patterns. This study addresses the challenges of high dimensionality and block-wise missing data through a regularization and constrained-based approach. The methodology is implemented in the R package bwm for binary and continuous response variables, and applied to breast cancer and exposome multi-omics datasets, achieving strong performance even in scenarios with missing data present in all omics. In binary classification task, our proposed model achieves accuracy in the range of 86% to 92%, and F1 in the range of 68% to 79%. And, in regression task the correlation between true and predicted responses is in the range of 72% to 76%. However, there is a slight decline in performance metrics as the percentage of missing data increases. In scenarios where block-wise missing data affects multiple omics, the model performance actually surpasses that of scenarios where missing data is present in only one omics. One possible explanation for this might be that the other scenarios introduce a greater diversity of observation profiles, leading to a more robust model. Depending on the specific omics being studied, there is greater consistency in feature selection when comparing block-wise missing data scenarios.

Humans↗

Development of methodology to support molecular endotype discovery from synovial fluid of individuals with knee osteoarthritis: The STEpUP OA consortium.

OBJECTIVES: To develop a protocol for largescale analysis of synovial fluid proteins, for the identification of biological networks associated with subtypes of osteoarthritis. METHODS: Synovial Fluid To detect molecular Endotypes by Unbiased Proteomics in Osteoarthritis (STEpUP OA) is an international consortium utilising clinical data (capturing pain, radiographic severity and demographic features) and knee synovial fluid from 17 participating cohorts. 1746 samples from 1650 individuals comprising OA, joint injury, healthy and inflammatory arthritis controls, divided into discovery (n = 1045) and replication (n = 701) datasets, were analysed by SomaScan Discovery Plex V4.1 (>7000 SOMAmers/proteins). An optimised approach to standardisation was developed. Technical confounders and batch-effects were identified and adjusted for. Poorly performing SOMAmers and samples were excluded. Variance in the data was determined by principal component (PC) analysis. RESULTS: A synovial fluid standardised protocol was optimised that had good reliability (<20% co-efficient of variation for >80% of SOMAmers in pooled samples) and overall good correlation with immunoassay. 1720 samples and >6290 SOMAmers met inclusion criteria. 48% of data variance (PC1) was strongly correlated with individual SOMAmer signal intensities, particularly with low abundance proteins (median correlation coefficient 0.70), and was enriched for nuclear and non-secreted proteins. We concluded that this component was predominantly intracellular proteins, and could be adjusted for using an 'intracellular protein score' (IPS). PC2 (7% variance) was attributable to processing batch and was batch-corrected by ComBat. Lesser effects were attributed to other technical confounders. Data visualisation revealed clustering of injury and OA cases in overlapping but distinguishable areas of high-dimensional proteomic space. CONCLUSIONS: We have developed a robust method for analysing synovial fluid protein, creating a molecular and clinical dataset of unprecedented scale to explore potential patient subtypes and the molecular pathogenesis of OA. Such methodology underpins the development of new approaches to tackle this disease which remains a huge societal challenge.

Humans↗

Integrated single-cell transcriptomics, Mendelian randomization, and machine learning identify CEBPZ as an immune-related biomarker in oral lichen planus.

BACKGROUND: Oral lichen planus (OLP) is a chronic, immune-mediated oral mucosal disease with complex pathophysiology and potential for malignant transformation. Understanding its molecular basis is critical for the development of precise diagnostic and therapeutic strategies. OBJECTIVES: We aimed to identify key immune-related biomarkers and characterize cellular dynamics in OLP, with a particular focus on the role of CEBPZ in disease pathogenesis. MATERIAL AND METHODS: We analyzed single-cell RNA sequencing (scRNA-seq) data from OLP lamina propria samples (GSE211630) to identify disease-specific T-cell subpopulations using high-dimensional weighted gene co-expression network analysis (hdWGCNA) for oxidative stress-related gene modules.-data-based Mendelian randomization (SMR) integrated FinnGen genome-wide association study (GWAS; 342,499 Europeans) data with Genotype-Tissue Expression (GTEx) expression quantitative trait loci (eQTL) data to identify causal genes. Machine learning (ML) models (least absolute shrinkage and selection operator (LASSO) and convolutional neural network (CNN)) were developed using bulk RNA-seq datasets (GSE52130 and GSE38616) for diagnostic purposes. RESULTS: We identified OLP-specific T-cell populations (clusters 0, 3, 5, 7, 13, and 15) with enhanced migration inhibition factor (MIF) pathway signaling toward B cells and monocytes. Two oxidative stress-associated modules contained hub genes, including CEBPZ. Summary-data-based Mendelian randomization analysis identified 231 OLP-associated genes, with CEBPZ uniquely intersecting LASSO-selected markers (odds ratio (OR) = 1.057, 95% confidence interval (95% CI) = 1.013-1.102, p = 0.010). Machine learning models achieved area under the curve (AUC) values ranging from 0.653 to 0.745, with the CNN model reaching a validation accuracy of 0.735. CEBPZ showed elevated expression in OLP T cells and correlated with enhanced MIF-(CD74+CXCR4) signaling. CONCLUSIONS: This integrative approach identifies CEBPZ as a pivotal biomarker linking genetic susceptibility, oxidative stress, and immune dysregulation in OLP. Our diagnostic models offer promising tools for OLP management.

CEBPZ↗

CCNA2 orchestrates the PI3K/AKT signaling axis to propel prostate cancer metastasis.

BACKGROUND: Prostate cancer (PCa) remains one of the most common malignancies in men, posing a persistent global burden in terms of both public health and socioeconomic costs. Although early detection is essential for improving patient outcomes, existing clinical tools, including prostate-specific antigen (PSA) screening, digital rectal examination, and transrectal ultrasound-guided biopsy, are hampered by suboptimal specificity and positive predictive value, resulting in frequent overdiagnosis and overtreatment of indolent lesions while missing a subset of aggressive tumors at an early stage. In this context, the rapid advancement of high-throughput omics technologies, coupled with sophisticated machine learning (ML) algorithms, provides a powerful computational framework to dissect high-dimensional genomic data, uncover latent gene expression signatures, and identify candidate biomarkers with superior discriminative performance over conventional clinicopathological parameters. Therefore, in this study, we sought to screen for crucial ML-based biomarkers associated with PCa, with a particular focus on systematically assessing the diagnostic and prognostic value of CCNA2. Leveraging large-scale transcriptomic cohorts from public repositories, we employed an ensemble of ML approaches to prioritize candidate genes and subsequently evaluated the diagnostic performance of CCNA2 through receiver operating characteristic curve analysis, as well as its prognostic utility via Kaplan-Meier survival estimation and multivariate Cox proportional hazards modeling. Our findings are anticipated to elucidate the molecular landscape of PCa and offer a promising biomarker candidate for early detection and risk stratification. METHODS: This study integrated single-cell RNA sequencing, bulk transcriptomic data from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) repositories, immunofluorescence, and multiple ML algorithms with in vitro functional assays to evaluate CCNA2 expression, clinical relevance, and biological behavior in PCa. RESULTS: CCNA2 was linked to metastasis and poor prognosis. High CCNA2 expression significantly correlated with adverse survival outcomes, and knockdown of CCNA2 suppressed proliferation, migration, and invasion in PCa cell lines. Mechanistically, CCNA2 modulated the PI3K/AKT signaling pathway. An ML-based diagnostic model incorporating CCNA2 demonstrated high predictive accuracy across multiple validation cohorts. CONCLUSIONS: CCNA2 serves as a promising prognostic biomarker and therapeutic target in prostate adenocarcinoma, driving tumor progression potentially via the PI3K/AKT axis.

CCNA2↗

Artificial Intelligence for Natural Products Discovery and Development.

Natural products (NPs) remain a cornerstone of modern drug discovery, offering stereochemical complexity and diverse bioactivities that precisely modulate therapeutic targets, refined through billions of years of evolution. However, their research has long been hindered by inefficient, empirical workflows, high resource consumption, structural complexity, and the "multicomponent, multi-target" nature of their mechanisms. The exponential growth of genomic, metabolomic, and spectral data has overwhelmed conventional analytical methods, exposing critical bottlenecks in handling high-dimensional, heterogeneous datasets that exceed human interpretive capacity. Artificial intelligence (AI) is emerging as a transformative paradigm to address these challenges, integrating multi-omics and chemical data to shift NP research from fragmented empiricism toward mechanism-driven, precision-oriented development. By leveraging deep learning architectures- including graph neural networks, Transformers, and diffusion-based generative models-AI enables systematic decoding of NP biosynthesis, automated structure elucidation, rational target identification, knowledge extraction from vast unstructured scientific literature, and de novo molecular design. This review comprehensively surveys recent advances in AI applications across the full NP discovery and development pipeline, encompassing genome mining, structure-based and ligand-based virtual screening, multimodal structural characterization, lead optimization, and biosynthetic pathway engineering. We further examine the emerging roles of protein-centric, molecule- centric, and multimodal foundation models, as well as large language models, in bridging genotype-to-chemotype gaps and unlocking unstructured scientific knowledge. Finally, we discuss critical challenges including data scarcity, representational limitations for complex stereochemistry, physical plausibility in generative models, and the urgent need for experimental validation, while outlining future directions toward autonomous experimentation, closed-loop optimization, and human-AI collaborative discovery.

Artificial intelligence↗

A Comprehensive Review of Radiomics in Pulmonary Nodule Management: Clinical Applications and Standardization Dilemmas.

Lung cancer is the most common and fatal malignant tumour. Early detection and treatment are likely to reduce mortality, but most pulmonary nodules identified during routine health checks are harmless. Consequently, a clear distinction between benign and malignant nodules is vital to improve early detection and reduce unnecessary interventions. Radiomics, a new omics technology, can be used to extract high-dimensional quantitative features from medical images, providing a profound understanding of tumour pathophysiology. Radiomics has attracted the attention of medical researchers since its formal definition by the Dutch researcher Lambin et al. in 2012. The number of research papers on radiomics has grown tremendously over the past few years. At present, it is used to predict pulmonary nodule malignancy, for noninvasive risk stratification, for integration with genomics to identify genetic mutations associated with lung cancer, and for evaluation of therapeutic responses. With this review, we summarise the literature on radiomics of pulmonary nodules, discuss how it could be used in nodule management, and address the current challenges and future directions for improving precision oncology.

Humans↗

Integrating machine learning and GWAS for variant prioritization in the INCIPE cohort highlights ABC transporter genes in chronic kidney disease.

INTRODUCTION: Chronic kidney disease (CKD) is a major public health challenge, affecting approximately 674 million people worldwide and representing one of the fastest-growing causes of mortality. Since CKD is frequently asymptomatic in its early stages, the identification of novel genetic biomarkers may improve early detection and risk stratification. Genome-Wide Association Studies (GWAS) have identified numerous genetic loci associated with CKD and related traits; however, their performance is often limited in small and imbalanced cohorts, where reduced statistical power increases both false-positive and false-negative findings. Machine learning (ML) approaches can complement conventional GWAS by prioritizing biologically relevant genetic signals from high-dimensional genomic data. METHODS: In this study, we implemented a nested ensemble (NCBC) model composed of an undersampler and a CatBoostClassifier (CBC) to prioritize candidate genetic variants associated with CKD in the INCIPE cohort. Prioritized variants were functionally annotated and evaluated through enrichment analyses, GTEx gene expression profiling, and protein-protein interaction network analyses. Genes identified by the CKDGen Consortium were analysed as an external reference set and used to validate the biological relevance of the prioritized results. RESULTS: The NCBC model outperformed conventional ML classifiers, achieving a ROC AUC score of 87.77%, compared to 50%-53% for the other evaluated models. Among the prioritized genes, 56.25% showed protein-protein interactions with genes previously reported by the CKDGen Consortium, whereas only 1.9% of randomly generated gene sets showed interactions. DISCUSSION: Our study demonstrates that the NCBC model improves the prioritization of biologically plausible candidate variants in a small and imbalanced CKD cohort. Functional analyses suggested ABC transporter-related genes, including ABCA13, ABCA4, and ABCC4 genes, as promising candidate for future validation, with ABCA4 showing substantial expression in kidney tissues. Overall, these findings support the integration of ML with GWAS to prioritize candidate genes and investigate the genetic architecture of complex diseases.

SNP prioritization↗

AI-integrated digital breeding for crop improvement.

Crop breeding increasingly depends on the effective integration and interpretation of large, heterogeneous datasets spanning genomic, phenotypic, multi-omics, and environmental layers. Conventional breeding approaches are often insufficient to capture the complex relationships among these data or to support timely selection decisions. Digital breeding can help address this limitation by complementing field experimentation, mixed models, and genomic prediction with the integration of biological data and computational prediction throughout the breeding process. In particular, the rapid advancement of artificial intelligence (AI) has improved the analysis of high-dimensional datasets and broadened its application to trait prediction, selection, and breeding design. Here, we review recent developments in AI-enabled digital breeding, encompassing genomic, phenomic, and multi-omics data generation and analysis, predictive modeling, explainable and generative AI, and data-driven breeding decision support. We further discuss emerging AI applications, their current contributions to crop research and breeding, and the major considerations affecting their reliable and practical implementation. Collectively, this review provides a structured understanding of the roles of AI across the digital breeding process and offers guidance for future methodological development and practical application in crop improvement.

artificial intelligence↗

Metaphor comprehension: a computational theory.

Metaphor comprehension involves an interaction between the meaning of the topic and the vehicle terms of the metaphor. Meaning is represented by vectors in a high-dimensional semantic space. Predication modifies the topic vector by merging it with selected features of the vehicle vector. The resulting metaphor vector can be evaluated by comparing it with known landmarks in the semantic space. Thus, metaphorical prediction is treated in the present model in exactly the same way as literal predication. Some experimental results concerning metaphor comprehension are simulated within this framework, such as the nonreversibility of metaphors, priming of metaphors with literal statements, and priming of literal statements with metaphors.

Algorithms↗

Cancer Immunotherapy: Therapeutic Limitations and Next-Generation Precision Strategies.

Cancer immunotherapy has reshaped oncology, largely through immune checkpoint inhibitors that release the brakes on tumor-reactive T cells. Yet the benefit remains uneven, and that unevenness traces back to a few basic biological limits. Checkpoint blockade amplifies immunity that is already present; it does not create tumor specificity de novo. Poor Ag quality, defective Ag presentation, a suppressive microenvironment, and epigenetically fixed T-cell exhaustion together set a ceiling on what checkpoint release can achieve. Next-generation strategies try to move past these limits by reorganizing immunotherapy around the functional layers of the immune response. Cancer vaccines define tumor-specific neoantigens and expand the responses against them. Ab-based approaches tune inhibitory signaling, draw immune cells toward the tumor, and trigger immunogenic cell death. Cellular therapies-chimeric Ag receptor T cell, TCR-engineered T cells, and tumor-infiltrating lymphocytes (TILs)-boost effector potency, with TIL therapy notable for preserving tumor-reactive repertoires shaped in vivo. Rather than rivals, these modalities are best seen as complementary layers-Ag definition, immune priming, effector optimization, and microenvironmental conditioning-to be combined in a programmable way. As genomic profiling, immunopeptidomics, and high-dimensional immune monitoring mature, the field is shifting from checkpoint-centered release toward precision immunoengineering, in which tumor-specific immunity is deliberately designed, aligned, and sustained.

Cancer vaccines↗

A rapid method for exploring the protein structure universe.

We have developed an automatic protein fingerprinting method for the evaluation of protein structural similarities based on secondary structure element compositions, spatial arrangements, lengths, and topologies. This method can rapidly identify proteins sharing structural homologies as we demonstrate with five test cases: the globins, the mammalian trypsinlike serine proteases, the immunoglobulins, the cupredoxins, and the actinlike ATPase domain-containing proteins. Principal components analysis of the similarity distance matrix calculated from an all-by-all comparison of 1,031 unique chains in the Protein Data Bank has produced a distribution of structures within a high-dimensional structural space. Fifty percent of the variance observed for this distribution is bounded by six axes, two of which encode structural variability within two large families, the immunoglobulins and the trypsinlike serine proteases. Many aspects of the spatial distribution remain stable upon reduction of the database to 140 proteins with minimal family overlap. The axes correlated with specific structural families are no longer observed. A clear hierarchy of organization is seen in the arrangement of protein structures in the universe. At the highest level, protein structures populate regions corresponding to the all-alpha, all-beta, and alpha/beta superfamilies. Large protein families are arranged along family-specific axes, forming local densely populated regions within the space. The lowest level of organization is intrafamilial; homologous structures are ordered by variations in peripheral secondary structure elements or by conformational shifts in the tertiary structure.

Algorithms↗

EM mixed model analysis of data from informatively censored normal distributions.

Maximum likelihood techniques using the EM algorithm are applied to correlated normally distributed survival data from a placebo-controlled, double-blind, dose-ranging crossover study to assess the short-term efficacy of an antianginal drug in patients with chronic stable angina. Censoring was informative and nonterminal and was not due to death or withdrawal from the study. Unlike previous approaches these techniques are mathematically and computationally tractable, do not require computations of high-dimensional integrals, and do not require the inversion of large matrices.

Algorithms↗

Kromoscopic analysis: a possible alternative to spectroscopic analysis for noninvasive measurement of analytes in vivo.

Light that penetrates scattering media shows nonlinearities that mask the broad and shallow perturbations made by trace analytes on the background illuminant spectrum. Narrow-band spectroscopic decomposition and deconvolution of such weak bands is a formidable analytical task that pushes the fundamentally linear spectroscopic method beyond practical limits. Kromoscopy is a high-dimensional analog of human color perception; it has broad-band spectrally overlapping detectors similar to those of the visual system, but in the infrared. Analyte bands are integrated fully in two or more detectors with different relative weightings. As in color vision, the analyte information is coded in the direct correlations between detector signals, which individually have higher signal-to-noise ratios than their spectroscopic counterparts. Our Kromoscopic instrument responds directly to glucose in aqueous solution, is not affected by temperature disturbances, and is fast enough to measure physiologically induced Kromoscopic changes in the arterial pulse waveform with high precision.

Blood Glucose↗

Three-dimensional linear and nonlinear transformations: an integration of light microscopical and MRI data.

The registration of image volumes derived from different imaging modalities such as MRI, PET, SPECT, and CT has been described in numerous studies in which functional and morphological data are combined on the basis of macrostructural information. However, the exact topography of architectural details is defined by microstructural information derived from histological sections. Therefore, a technique is developed for integrating micro- and macrostructural information based on 1) a three-dimensional reconstruction of the histological volume which accounts for linear and nonlinear histological deformations, and 2) a two-step procedure which transforms these volumes to a reference coordinate system. The two-step procedure uses an extended principal axes transformation (PAT) generalized to affine transformations and a fast, automated full-multigrid method (FMG) for determining high-dimensional three-dimensional nonlinear deformations in order to account for differences in the morphology of individuals. With this technique, it is possible to define upwards of 1,000 times the resolution of approximately 1 mm in MRI, making possible the identification of geometric and texture features of microscopically defined brain structures.

Brain Mapping↗

Spatial and temporal pattern analysis via spiking neurons.

Spiking neurons, receiving temporally encoded inputs, can compute radial basis functions (RBFs) by storing the relevant information in their delays. In this paper we show how these delays can be learned using exclusively locally available information (basically the time difference between the pre- and postsynaptic spikes). Our approach gives rise to a biologically plausible algorithm for finding clusters in a high-dimensional input space with networks of spiking neurons, even if the environment is changing dynamically. Furthermore, we show that our learning mechanism makes it possible that such RBF neurons can perform some kind of feature extraction where they recognize that only certain input coordinates carry relevant information. Finally we demonstrate that this model allows the recognition of temporal sequences even if they are distorted in various ways.

Action Potentials↗

Systematic evaluation of one-dimensional-to-two-dimensional near-infrared spectroscopy transformations with deep learning for quantifying coconut sap adulteration.

Near-infrared (NIR) spectroscopy have limitations when combined with deep learning (DL) algorithms because they rely on low-dimensional datasets. Therefore, we investigated the potential of transforming one-dimensional (1D) NIR spectra into two-dimensional (2D) spectrograms using synchronous and asynchronous techniques and the continuous wavelet transform (CWT) and their effectiveness by integrating with DL for detecting adulteration in coconut sap. NIR spectra (12,500-4000&#xa0;cm-1) were collected from binary mixtures (0%-100%;w/w). The performance of all DL (convolutional neural networks-CNN, AlexNet and ResNet) models was compared with that of partial least squares (PLS). The models were ranked in the mentioned order based on their performances: 2D-CWT&#xa0;>&#xa0;2D-asynchronous > 2D-synchronous > 1D/2D-PLS. The important features of the best model can be explained and visualized using gradient-weighted-class-activation-mapping. The findings highlight that the 1D-to-2D NIR data transformation combined with DL is a highly robust approach because it addresses the feature representation gap in NIR data and effectively captures the spatial-spectral correlations.

Spectroscopy, Near-Infrared↗