PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Dimensionality Reduction”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Single-Cell Proteomics Reveals Proteome Remodeling and Cellular Heterogeneity During NGF-Induced PC12 Neuronal Differentiation.

Single-cell proteomics enables direct measurement of cellular heterogeneity during dynamic biological processes, but its application to fragile and highly adherent neuronal models remains challenging. Here, we developed and applied an optimized single-cell proteomics workflow to characterize proteome remodeling during nerve growth factor (NGF)-induced differentiation of PC12 cells. To enable reliable single-cell analysis, we implemented gentle dissociation, antiaggregation strategies, and thermal inkjet-based cell dispensing, achieving high accuracy in single-cell isolation. Inclusion of n-dodecyl-β-d-maltoside (DDM) improved recovery of membrane-associated and low-solubility proteins. Coupled with LC-ion mobility-mass spectrometry, this workflow enabled quantification of 2,000-3,000 proteins per cell across the differentiation time course. Single-cell proteomic analysis revealed progressive and heterogeneous proteome remodeling during differentiation. While undifferentiated cells formed a relatively homogeneous population, later stages (Days 4-6) exhibited increased variability, including multimodal protein abundance distributions and separation into distinct subpopulations. Dimensionality reduction, clustering, and non-negative matrix factorization identified multiple coexisting proteomic states within the same time points, reflecting asynchronous differentiation trajectories. These subpopulations were characterized by coordinated differences in pathways related to intracellular trafficking, protein translation, cytoskeletal organization, and neuronal maturation. Comparison with bulk proteomics demonstrated that proteins associated with differentiated neuronal states, including those involved in neurite formation and structural remodeling, are underrepresented in population-averaged measurements but are enriched within specific single-cell subpopulations. Temporal and cluster-resolved analyses further revealed distinct protein expression trajectories, including early decreases in cell cycle and metabolic pathways and later increases in neuronal structural and regulatory proteins. Together, this study establishes an optimized workflow for single-cell proteomics of neuronal systems and demonstrates that NGF-induced PC12 differentiation proceeds through heterogeneous and divergent proteomic states that are not resolved by bulk analysis.

Animals↗

Nonlinear mapping networks.

Among the many dimensionality reduction techniques that have appeared in the statistical literature, multidimensional scaling and nonlinear mapping are unique for their conceptual simplicity and ability to reproduce the topology and structure of the data space in a faithful and unbiased manner. However, a major shortcoming of these methods is their quadratic dependence on the number of objects scaled, which imposes severe limitations on the size of data sets that can be effectively manipulated. Here we describe a novel approach that combines conventional nonlinear mapping techniques with feed-forward neural networks, and allows the processing of data sets orders of magnitude larger than those accessible with conventional methodologies. Rooted on the principle of probability sampling, the method employs a classical algorithm to project a small random sample, and then "learns" the underlying nonlinear transform using a multilayer neural network trained with the back-propagation algorithm. Once trained, the neural network can be used in a feed-forward manner to project the remaining members of the population as well as new, unseen samples with minimal distortion. Using examples from the fields of image processing and combinatorial chemistry, we demonstrate that this method can generate projections that are virtually indistinguishable from those derived by conventional approaches. The ability to encode the nonlinear transform in the form of a neural network makes nonlinear mapping applicable to a wide variety of data mining applications involving very large data sets that are otherwise computationally intractable.

Journal Article↗

Advances in diversity profiling and combinatorial series design.

Rapid advances in synthetic and screening technology have recently enabled the simultaneous synthesis and biological evaluation of large chemical libraries containing hundreds to tens of thousands of compounds, using molecular diversity as a means to design and prioritize experiments. This paper reviews some of the most important computational work in the field of diversity profiling and combinatorial library design, with particular emphasis on methodology and applications. It is divided into four sections that address issues related to molecular representation, dimensionality reduction, compound selection, and visualization.

Chemistry, Pharmaceutical↗

Association of AGER genetic variants with chronic obstructive pulmonary disease susceptibility in Southern Chinese Han populations.

OBJECTIVE: Chronic obstructive pulmonary disease (COPD) remains a leading cause of disability and mortality among elderly populations. Studies indicate that AGER plays a critical regulatory role in the pathogenesis of respiratory disorders. However, the genetic variations in AGER to COPD susceptibility remain incompletely understood. This study employs a case-control design to investigate associations between AGER genetic variants and COPD risk in the Southern Chinese Han population. METHODS: This study enrolled 270 COPD patients and 271 healthy controls. AGER single-nucleotide polymorphisms (SNPs) were analysed using the MassARRAY iPLEX platform. Logistic regression models evaluated associations between AGER polymorphisms and COPD susceptibility, with false discovery rate (FDR) correction applied to mitigate multiple testing errors. SNP-SNP interactions were investigated through multifactor dimensionality reduction (MDR) analysis. Expression quantitative trait locus (eQTL) data from the GTEx database were further analysed to assess regulatory relationships between SNPs and AGER gene expression levels. RESULTS: This study showed that rs3134941 (G allele, OR = 0.21, 95% CI = 0.10-0.41, p (FDR) = 0.001) and rs3131300 (G allele, OR = 0.32, 95% CI = 0.20-0.49, p (FDR) = 0.0001) were significantly associated with a reduced susceptibility to COPD. MDR indicated that rs3131300 was the optimal predictive model for COPD risk. Additionally, initial mechanistic investigations utilizing the GTEx database identify rs3134941 (C > G) and rs3131300 (A > G) as significant expression quantitative trait loci for AGER mRNA in cell-cultured fibroblasts and whole blood. CONCLUSION: Our study demonstrated that AGER genetic variants might play a protective role in the progression of COPD.

Aged↗

Linear transformations of data space in MEG.

Magnetoencephalography (MEG) is a method which allows the non-invasive measurement of the minute magnetic field which is generated by ion currents in the brain. Due to the complex sensitivity profile of the sensors, the measured data are a non-trivial representation of the currents where information specific to local generators is distributed across many channels and each channel contains a mixture of contributions from many such generators. We propose a framework which generates a new representation of the data through a linear transformation which is designed so that some desired property is optimized in one or more new virtual channel(s). First figures of merit are suggested to describe the relation between the measured data and the underlying currents. Within this context the new framework is established by first showing how the transformation matrix itself is designed and then by its application to real and simulated data. The results demonstrate that the proposed linear transformations of data space provide a computationally efficient tool for analysis and a very much needed dimensional reduction of the data.

Brain↗

Negative dataset selection impacts machine learning-based predictors for multiple bacterial species promoters.

MOTIVATION: Advances in bacterial promoter predictors based on machine learning have greatly improved identification metrics. However, existing models overlooked the impact of negative datasets, previously identified in GC-content discrepancies between positive and negative datasets in single-species models. This study aims to investigate whether multiple-species models for promoter classification are inherently biased due to the selection criteria of negative datasets. We further explore whether the generation of synthetic random sequences (SRS) that mimic GC-content distribution of promoters can partly reduce this bias. RESULTS: Multiple-species predictors exhibited GC-content bias when using CDS as a negative dataset, suggested by specificity and sensibility metrics in a species-specific manner, and investigated by dimensionality reduction. We demonstrated a reduction in this bias by using the SRS dataset, with less detection of background noise in real genomic data. In both scenarios DNABERT showed the best metrics. These findings suggest that GC-balanced datasets can enhance the generalizability of promoter predictors across Bacteria. AVAILABILITY AND IMPLEMENTATION: The source code of the experiments is freely available at https://github.com/maigonzalezh/MultispeciesPromoterClassifier.

Machine Learning↗

Multimodal deep learning for immunotherapy response prediction and biomarker discovery in non-small cell lung cancer.

OBJECTIVE: Immunotherapy has emerged as a promising treatment for advanced non-small cell lung cancer (NSCLC), but accurately predicting which patients will benefit from it remains a major clinical challenge. To address this, we aim to develop a novel multimodal method, DeepAFM, that integrates histopathology, genomic features, and clinical information to predict patient responses to anti-PD-(L)1 immunotherapy. MATERIALS AND METHODS: A total of 93 patients with advanced NSCLC were included in this study. Histopathological whole-slide images were processed using a self-supervised VQVAE2 for representation learning. PCA and K-means clustering were then applied for dimensionality reduction and feature grouping. Key regions of interest were visualized through permutation importance evaluation and color-coding techniques. The extracted histopathological features, along with genomic alterations and clinical variables, were integrated into the DeepAFM multimodal prediction model. RESULTS: The DeepAFM achieved a high predictive performance with an area under the curve (AUC) of 0.77 (95% confidence interval: 0.69-1.00). Attention-based heatmaps revealed that the model could identify critical pathological patterns, genomic mutations, and clinical indicators associated with patient responses to immunotherapy. DISCUSSION: The integration of multimodal data enabled the model to capture complex interactions among pathology, genomics, and clinical characteristics, enhancing the interpretability and predictive power of immunotherapy response prediction. The visualization techniques facilitated the identification of biologically meaningful features and potential biomarkers. CONCLUSION: This study demonstrates the effectiveness of the DeepAFM in predicting responses to immunotherapy in advanced NSCLC. The approach not only improves prediction accuracy but also provides valuable insights for personalized treatment strategies and biomarker discovery.

Humans↗

Genomic hallmarks of depot medroxyprogesterone acetate-associated meningiomas.

BACKGROUND: Population-based studies have linked progestin exposure to increased meningioma risk. However, the molecular basis of meningiomas associated with depot medroxyprogesterone acetate (DMPA)-a common injectable contraceptive-remains undefined. METHODS: We performed an integrated clinicopathologic and genomic analysis of meningiomas from 10 women with long-term DMPA exposure. Tumors underwent histopathological analysis, targeted sequencing, and DNA methylation profiling. Data were integrated with reference cohorts (Baylor and Heidelberg) and analyzed through classifier assignment, consensus clustering, copy number analysis, differential methylation testing, and dimensionality reduction. RESULTS: Depot medroxyprogesterone acetate-associated meningiomas were all newly diagnosed, World Health Organization grade 1 tumors with a predilection for the anterior and central skull base (n = 6). Nine patients harbored multiple meningiomas. Four experienced regression of untreated meningiomas following DMPA cessation, while 5 demonstrated stabilization. Histopathology demonstrated relative overrepresentation of metaplastic morphology, an uncommon meningioma subtype. All DMPA-associated meningiomas mapped to benign molecular groups, and most exhibited low copy number alteration burden. Targeted sequencing revealed enrichment for TRAF7 mutations (n = 5), with no NF2 mutations detected. Eight tumors shared consensus cluster identity, with cohesive grouping on principal component analysis and t-distributed stochastic neighbor embedding. No differential methylation was identified at the progesterone receptor locus. CONCLUSIONS: Depot medroxyprogesterone acetate-associated meningiomas represent a recognizable phenotype within the broader NF2-wildtype/TRAF7-enriched spectrum of benign meningiomas, characterized by chromosomal stability, a shared methylation profile, tumor multiplicity, and regression or stabilization following DMPA cessation. While derived from a small single-institution cohort, these findings provide a molecular framework for understanding progestin-associated meningioma biology, reinterpreting epidemiologic literature, and informing population-level risk stratification.

Humans↗

Probability-based current dipole localization from biomagnetic fields.

Focal biomagnetic sources are described as pointlike current dipoles. The dipole parameters, position, and moment coordinates are commonly determined from biomagnetic data using iterative nonlinear optimization algorithms such as the Levenberg-Marquardt algorithm. However, even for single-dipole sources, mislocalizations can occur due to side minima of the cost function or due to a wrong choice of the start vector. This can be shown by introducing a cost function where the independent variables are only the position coordinates instead of position and moment coordinates. This dimensional reduction--which is also possible for multiple dipole sources--is achieved by calculating the cost function at each position with the position and data-dependent, optimum dipole moments. We call these dipoles with--in a least squares sense--optimum moments, locally optimal dipoles. The visualization of such a single-dipole cost function and of the iteration steps of the Levenberg-Marquardt algorithm show why mislocalizations cannot be avoided. Therefore, we propose an alternative noniterative localization algorithm for single-dipole sources without this drawback. It uses localization probabilities calculated by means of the locally optimal dipoles. Besides the determination of the dipole parameters, the proposed algorithm furnishes a reliable error for each localization. Its effectiveness is shown with simulated and real patient data.

Algorithms↗

Use of IR Biotyper as a feasible methodology to type Klebsiella pneumoniae.

UNLABELLED: Klebsiella pneumoniae is one of the most frequently reported healthcare-associated pathogens. The current gold standard approach to perform the epidemiological typing of these bacteria is Whole Genome Sequencing (WGS), which is an expensive and challenging procedure. IR Biotyper (Bruker Daltonics, GmbH) is a new equipment based on Fourier transform infrared spectroscopy, which allows a rapid, low-cost, and user-friendly method to type bacterial isolates. However, there is a need for studies that evaluate the efficacy of the IR Biotyper. The aim of this study was to evaluate the capability of IR Biotyper to type K. pneumoniae according to sequence type (ST) and capsular type-using K locus (KL)-as well as to develop a classifier using machine learning. Seventy-three isolates of K. pneumoniae previously characterized by WGS were selected for IR Biotyper analysis using principal component analysis for dimensionality reduction, Euclidean, and unweighted pair group method with arithmetic mean (UPGMA) for clustering method, and spectra were analyzed in the 1,300-800 cm⁻¹ wavenumber range. Among these, 54 isolates were used to create a classifier, and 19 were used to validate the classifier. When considering the ST, ST307 was grouped in the same cluster as ST11. When KL was considered for the analysis, the clusters were 100% correctly grouped according to their KL type. Furthermore, the classifier developed was able to classify the isolates according to KL with a high concordance. This study showed that KL correlates well with KL for typing K. pneumoniae isolates using the IR Biotyper. Additionally, IR Biotyper demonstrated to be a cost-effective method and a promising tool to classify isolates within minutes. IMPORTANCE: Klebsiella pneumoniae is a major cause of severe hospital infections, and controlling its spread requires quick identification and comparison of bacterial strains. WGS is accurate but expensive, slow, and technically demanding. In this study, we evaluated the IR Biotyper, a device that uses infrared light to analyze bacteria and group them by capsule type-a key feature linked to their spread. The IR Biotyper matched WGS results with high accuracy, delivering results in minutes instead of days. This fast, affordable method can help hospitals detect outbreaks earlier and respond more effectively. Our findings suggest that the IR Biotyper is a valuable tool for routine use in microbiology laboratories, supporting epidemiological surveillance and outbreak control.

Klebsiella pneumoniae↗

Mixtures of probabilistic principal component analyzers.

Principal component analysis (PCA) is one of the most popular techniques for processing, compressing, and visualizing data, although its effectiveness is limited by its global linearity. While nonlinear variants of PCA have been proposed, an alternative paradigm is to capture data complexity by a combination of local linear PCA projections. However, conventional PCA does not correspond to a probability density, and so there is no unique way to combine PCA models. Therefore, previous attempts to formulate mixture models for PCA have been ad hoc to some extent. In this article, PCA is formulated within a maximum likelihood framework, based on a specific form of gaussian latent variable model. This leads to a well-defined mixture model for probabilistic principal component analyzers, whose parameters can be determined using an expectation-maximization algorithm. We discuss the advantages of this model in the context of clustering, density modeling, and local dimensionality reduction, and we demonstrate its application to image compression and handwritten digit recognition.

Algorithms↗

Self-organization as an iterative kernel smoothing process.

Kohonen's self-organizing map, when described in a batch processing mode, can be interpreted as a statistical kernel smoothing problem. The batch SOM algorithm consists of two steps. First, the training data are partitioned according to the Voronoi regions of the map unit locations. Second, the units are updated by taking weighted centroids of the data falling into the Voronoi regions, with the weighing function given by the neighborhood. Then, the neighborhood width is decreased and steps 1, 2 are repeated. The second step can be interpreted as a statistical kernel smoothing problem where the neighborhood function corresponds to the kernel and neighborhood width corresponds to kernel span. To determine the new unit locations, kernel smoothing is applied to the centroids of the Voronoi regions in the topological space. This interpretation leads to some new insights concerning the role of the neighborhood and dimensionality reduction. It also strengthens the algorithm's connection with the Principal Curve algorithm. A generalized self-organizing algorithm is proposed, where the kernel smoothing step is replaced with an arbitrary nonparametric regression method.

Algorithms↗

scGPA: an LLM-assisted workflow for directional virtual gene perturbation analysis from single-cell transcriptomes.

BACKGROUND: Existing virtual perturbation methods can often infer directional changes by comparing predicted post-perturbation expression profiles with control cells. However, workflows that directly return direction-specific downstream candidate genes together with confidence scores, evidence support and interpretable summaries remain limited. We developed scGPA, an LLM-assisted workflow system for directional single-cell virtual gene perturbation analysis. METHODS: scGPA starts from raw single-cell RNA sequencing data and performs quality control, normalization, dimensionality reduction, clustering and cell-group selection. It then constructs cell-group-specific wild-type regulatory networks using repeated subsampling, principal component regression (PCR)/Ridge-based network inference and CP tensor denoising. Based on these networks, scGPA simulates dose-aware virtual knockdown of the target gene and applies signed perturbation propagation to estimate the magnitude and direction of downstream transcriptional responses. LLM assistance is used for marker-based cell-type annotation, evidence-guided candidate prioritization and user-facing biological summarization. RESULTS: We benchmarked scGPA across five public Perturb-seq datasets and compared its performance with GEARS, scGPT and a random baseline. The overall correct prediction rate of scGPA was 23.0%, exceeding those of GEARS (20.7%), scGPT (15.1%) and the random baseline (13.6%). These results indicate that scGPA achieved a higher correct prediction rate than the two comparator models and the random baseline. We subsequently evaluated scGPA using a public osteosarcoma single-cell dataset and performed qRT-PCR validation in 143B osteosarcoma cells. Among genes with significant experimental changes, scGPA achieved a directional concordance of 76.9%. When all tested downstream genes were counted, 37.0% were directionally correct, 51.9% showed no significant change and 11.1% changed in the opposite direction. CONCLUSIONS: scGPA provides a practical workflow system for predicting and prioritizing direction-specific downstream transcriptional responses after target-gene perturbation. By integrating single-cell regulatory network inference, signed virtual perturbation and LLM-assisted interpretation, scGPA supports target-gene function inference and downstream mechanistic investigation from single-cell transcriptomic data.

Single-Cell Gene Expression Analysis↗

Interaction between toll-like receptor 4 polymorphism and abdominal obesity on ovarian cancer risk in Chinese women.

OBJECTIVES: the aim of this study was to evaluate the impact of TLR4 gene single nucleotide polymorphisms (SNPs) and additional TLR4 gene SNP- SNP and SNP- abdominal obesity (AO) interaction on ovarian cancer (OC) risk. METHODS: Generalized multifactor dimensionality reduction method were utilized to identify the most informative interactions between four SNPs in the TLR4 gene and abdominal obesity. Logistic regression was employed to investigate the association between 4 SNPs within TLR4 gene and OC risk, and additional SNP- SNP and gene- AO interaction on OC risk, ORs (95%CI) were calculated. RESULTS: The analysis of logistic regression indicated a markedly elevated risk of OC in individuals carrying either the rs4986790-G or rs11536889-C alleles in the TLR4 gene compared to those with the standard genetic variations, adjusted ORs (95%CI) were 1.61 (1.28-1.96) and 1.48 (1.09-1.91). GMDR analysis indicated a significant two-locus model (p = 0.018) involving rs4986790 and rs11536889, and a significant two-locus model (p = 0.001) involving rs4986790 and AO. Participants with rs4986790- AG/GG and rs11536889GC/ CC genotype has the highest OC risk, compared to participants with rs4986790-AA and rs11536889-GG genotype, OR (95%CI) = 2.58 (1.46-3.71), and abdominal obese participants with rs4986790- AG/GG genotype have the highest OC risk, compared to non- abdominal obese participants with rs4986790-AA genotype, OR (95%CI) = 3.17 (1.78-4.58). CONCLUSIONS: The findings suggested that TLR4 gene rs4986790 and rs11536889 polymorphisms were associated with increased OC risk. Significant interaction also existed between rs4986790 and AO, which means that the WC levels may influence the impact of rs4986790 on OC risk.

Adult↗

ABO exon polymorphisms are related to ischemic stroke in a Chinese Han population.

BACKGROUND: Recent research have underscored the relation of ABO blood group system to cerebrovascular disorders predisposition. The present investigation endeavors to delve into the relationship between ABO polymorphisms and ischemic stroke (IS) risk. METHODS: A cohort of 646 IS patients and 649 matched healthy controls was recruited. Genotyping of five SNPs within ABO were conducted by Agena MassARRAY platform. Logistic regression models were employed to estimate odds ratios (ORs) and 95% confidence intervals (CIs). Additionally, SNP-SNP interaction was assessed by multifactor dimensionality reduction (MDR) method. Furthermore, Analysis of Variance (ANOVA) was utilized to explore the association between genotypes and blood lipid profiles. RESULTS: The study identified an elevated IS risk associated with rs8176740 and rs8176720 in the overall population. Notably, ABO rs8176720 emerged as the most informative single-locus model for IS susceptibility. These variants were related to an elevated IS risk, specifically in female subjects, the subgroup aged > 64 years, non-smokers, drinkers or non-drinkers. Moreover, rs8176749 and rs8176745 were associated with red blood cell count levels and total bilirubin levels. CONCLUSION: This study firstly demonstrated the association of ABO rs8176740 and rs8176720 with IS incidence, which increased the understanding regarding the effect of ABO on IS pathogenesis.

Aged↗

PLNMFG: Pseudo-label guided non-negative matrix factorization model with graph constraint for single-cell multi-omics data clustering.

The development of single-cell multi-omics sequencing technologies has enabled the simultaneous analysis of multi-omics data within the same cell. Accurate clustering of these cells is crucial for downstream analyses of complex biological functions. Despite significant advances in multi-omics integration approaches, current methodologies exhibit two major limitations. First, they inadequately incorporate prior biological knowledge from various omic layers. Second, these methods often conduct independent dimensionality reduction on individual omic datasets, thereby failing to capture the intrinsic complementary information and potentially overlooking crucial cross-platform interactions. Motivated by these, this study investigates a non-negative matrix factorization model called PLNMFG, which integrates the unified latent representation learning that retains the features between and within omics and the cluster structure learning that retains the intrinsic structure of the data into one joint framework. Specially, PLNMFG performs adaptive imputation to handle dropout events and uses prior pseudo-labels as constraints during the process of collective non-negative matrix factorization, as a result, a more robust latent representation that preserves the double similarity information is obtained. Graph Laplacian constraint is applied during clustering which further preserves structure characteristic of multi-omics data. In addition, the weight of each omic is adaptively learned based on the omic contribution. A series of experiments on 8 benchmark datasets show that our model performs well in terms of clustering accuracy and computational efficiency.

Single-Cell Analysis↗

Myeloid landscape of BRAF-mutant papillary thyroid cancer and thyroiditis.

Papillary thyroid cancer (PTC) is less aggressive when associated with lymphocytic thyroiditis (LT), even in the presence of oncogenic BRAF, including smaller tumours, less lymph node involvement and reduced extrathyroidal extension. To investigate possible immune mechanisms underlying this association, we compared the tumour microenvironment of PTC-BRAF with LT and that without LT using single-cell RNA sequencing (scRNA-seq). Single-cell libraries were generated from fresh and fixed tumour samples with post-dissociation viability >70% using the 10x Genomics Chromium Platform and sequenced on an Illumina NovaSeq 6000. We analysed scRNA-seq data from 11 PTC-BRAF tumours: four with LT (one publicly available sample) and seven without LT. Downstream analyses included quality control, batch correction, dimensionality reduction, and differential gene expression analysis. We found that neutrophils were the predominant myeloid cell type in PTCs without LT. Thyrocytes without LT showed significant expression of the neutrophil recruitment chemokine ECRG4. In the absence of LT, neutrophils expressed oncogenic genes with poor clinical outcomes. In contrast, thyrocytes from tumours with LT showed increased expression of MHC-II antigen presentation, consistent with effective immune surveillance. Thyrocytes and macrophages in the presence of LT showed enrichment of interferon gamma response pathways. Our data suggest that LT in thyroid cancer is associated with enhanced antigen presentation and fewer features of pro-tumourigenic innate immune activity. These results identify previously under-recognised innate immune cell population and associated transcriptomic features, which suggest new mechanisms to target immune treatments in PTC refractory to other therapies.

Humans↗