PubMed HealthSearch

SEARCH · PubMed Health

Results for “Dimensionality Reduction”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Negative dataset selection impacts machine learning-based predictors for multiple bacterial species promoters.

MOTIVATION: Advances in bacterial promoter predictors based on machine learning have greatly improved identification metrics. However, existing models overlooked the impact of negative datasets, previously identified in GC-content discrepancies between positive and negative datasets in single-species models. This study aims to investigate whether multiple-species models for promoter classification are inherently biased due to the selection criteria of negative datasets. We further explore whether the generation of synthetic random sequences (SRS) that mimic GC-content distribution of promoters can partly reduce this bias. RESULTS: Multiple-species predictors exhibited GC-content bias when using CDS as a negative dataset, suggested by specificity and sensibility metrics in a species-specific manner, and investigated by dimensionality reduction. We demonstrated a reduction in this bias by using the SRS dataset, with less detection of background noise in real genomic data. In both scenarios DNABERT showed the best metrics. These findings suggest that GC-balanced datasets can enhance the generalizability of promoter predictors across Bacteria. AVAILABILITY AND IMPLEMENTATION: The source code of the experiments is freely available at https://github.com/maigonzalezh/MultispeciesPromoterClassifier.

Machine Learning

Multimodal deep learning for immunotherapy response prediction and biomarker discovery in non-small cell lung cancer.

OBJECTIVE: Immunotherapy has emerged as a promising treatment for advanced non-small cell lung cancer (NSCLC), but accurately predicting which patients will benefit from it remains a major clinical challenge. To address this, we aim to develop a novel multimodal method, DeepAFM, that integrates histopathology, genomic features, and clinical information to predict patient responses to anti-PD-(L)1 immunotherapy. MATERIALS AND METHODS: A total of 93 patients with advanced NSCLC were included in this study. Histopathological whole-slide images were processed using a self-supervised VQVAE2 for representation learning. PCA and K-means clustering were then applied for dimensionality reduction and feature grouping. Key regions of interest were visualized through permutation importance evaluation and color-coding techniques. The extracted histopathological features, along with genomic alterations and clinical variables, were integrated into the DeepAFM multimodal prediction model. RESULTS: The DeepAFM achieved a high predictive performance with an area under the curve (AUC) of 0.77 (95% confidence interval: 0.69-1.00). Attention-based heatmaps revealed that the model could identify critical pathological patterns, genomic mutations, and clinical indicators associated with patient responses to immunotherapy. DISCUSSION: The integration of multimodal data enabled the model to capture complex interactions among pathology, genomics, and clinical characteristics, enhancing the interpretability and predictive power of immunotherapy response prediction. The visualization techniques facilitated the identification of biologically meaningful features and potential biomarkers. CONCLUSION: This study demonstrates the effectiveness of the DeepAFM in predicting responses to immunotherapy in advanced NSCLC. The approach not only improves prediction accuracy but also provides valuable insights for personalized treatment strategies and biomarker discovery.

Humans

Genomic hallmarks of depot medroxyprogesterone acetate-associated meningiomas.

BACKGROUND: Population-based studies have linked progestin exposure to increased meningioma risk. However, the molecular basis of meningiomas associated with depot medroxyprogesterone acetate (DMPA)-a common injectable contraceptive-remains undefined. METHODS: We performed an integrated clinicopathologic and genomic analysis of meningiomas from 10 women with long-term DMPA exposure. Tumors underwent histopathological analysis, targeted sequencing, and DNA methylation profiling. Data were integrated with reference cohorts (Baylor and Heidelberg) and analyzed through classifier assignment, consensus clustering, copy number analysis, differential methylation testing, and dimensionality reduction. RESULTS: Depot medroxyprogesterone acetate-associated meningiomas were all newly diagnosed, World Health Organization grade 1 tumors with a predilection for the anterior and central skull base (n = 6). Nine patients harbored multiple meningiomas. Four experienced regression of untreated meningiomas following DMPA cessation, while 5 demonstrated stabilization. Histopathology demonstrated relative overrepresentation of metaplastic morphology, an uncommon meningioma subtype. All DMPA-associated meningiomas mapped to benign molecular groups, and most exhibited low copy number alteration burden. Targeted sequencing revealed enrichment for TRAF7 mutations (n = 5), with no NF2 mutations detected. Eight tumors shared consensus cluster identity, with cohesive grouping on principal component analysis and t-distributed stochastic neighbor embedding. No differential methylation was identified at the progesterone receptor locus. CONCLUSIONS: Depot medroxyprogesterone acetate-associated meningiomas represent a recognizable phenotype within the broader NF2-wildtype/TRAF7-enriched spectrum of benign meningiomas, characterized by chromosomal stability, a shared methylation profile, tumor multiplicity, and regression or stabilization following DMPA cessation. While derived from a small single-institution cohort, these findings provide a molecular framework for understanding progestin-associated meningioma biology, reinterpreting epidemiologic literature, and informing population-level risk stratification.

Humans

Use of IR Biotyper as a feasible methodology to type Klebsiella pneumoniae.

UNLABELLED: Klebsiella pneumoniae is one of the most frequently reported healthcare-associated pathogens. The current gold standard approach to perform the epidemiological typing of these bacteria is Whole Genome Sequencing (WGS), which is an expensive and challenging procedure. IR Biotyper (Bruker Daltonics, GmbH) is a new equipment based on Fourier transform infrared spectroscopy, which allows a rapid, low-cost, and user-friendly method to type bacterial isolates. However, there is a need for studies that evaluate the efficacy of the IR Biotyper. The aim of this study was to evaluate the capability of IR Biotyper to type K. pneumoniae according to sequence type (ST) and capsular type-using K locus (KL)-as well as to develop a classifier using machine learning. Seventy-three isolates of K. pneumoniae previously characterized by WGS were selected for IR Biotyper analysis using principal component analysis for dimensionality reduction, Euclidean, and unweighted pair group method with arithmetic mean (UPGMA) for clustering method, and spectra were analyzed in the 1,300-800 cm⁻¹ wavenumber range. Among these, 54 isolates were used to create a classifier, and 19 were used to validate the classifier. When considering the ST, ST307 was grouped in the same cluster as ST11. When KL was considered for the analysis, the clusters were 100% correctly grouped according to their KL type. Furthermore, the classifier developed was able to classify the isolates according to KL with a high concordance. This study showed that KL correlates well with KL for typing K. pneumoniae isolates using the IR Biotyper. Additionally, IR Biotyper demonstrated to be a cost-effective method and a promising tool to classify isolates within minutes. IMPORTANCE: Klebsiella pneumoniae is a major cause of severe hospital infections, and controlling its spread requires quick identification and comparison of bacterial strains. WGS is accurate but expensive, slow, and technically demanding. In this study, we evaluated the IR Biotyper, a device that uses infrared light to analyze bacteria and group them by capsule type-a key feature linked to their spread. The IR Biotyper matched WGS results with high accuracy, delivering results in minutes instead of days. This fast, affordable method can help hospitals detect outbreaks earlier and respond more effectively. Our findings suggest that the IR Biotyper is a valuable tool for routine use in microbiology laboratories, supporting epidemiological surveillance and outbreak control.

Klebsiella pneumoniae

scGPA: an LLM-assisted workflow for directional virtual gene perturbation analysis from single-cell transcriptomes.

BACKGROUND: Existing virtual perturbation methods can often infer directional changes by comparing predicted post-perturbation expression profiles with control cells. However, workflows that directly return direction-specific downstream candidate genes together with confidence scores, evidence support and interpretable summaries remain limited. We developed scGPA, an LLM-assisted workflow system for directional single-cell virtual gene perturbation analysis. METHODS: scGPA starts from raw single-cell RNA sequencing data and performs quality control, normalization, dimensionality reduction, clustering and cell-group selection. It then constructs cell-group-specific wild-type regulatory networks using repeated subsampling, principal component regression (PCR)/Ridge-based network inference and CP tensor denoising. Based on these networks, scGPA simulates dose-aware virtual knockdown of the target gene and applies signed perturbation propagation to estimate the magnitude and direction of downstream transcriptional responses. LLM assistance is used for marker-based cell-type annotation, evidence-guided candidate prioritization and user-facing biological summarization. RESULTS: We benchmarked scGPA across five public Perturb-seq datasets and compared its performance with GEARS, scGPT and a random baseline. The overall correct prediction rate of scGPA was 23.0%, exceeding those of GEARS (20.7%), scGPT (15.1%) and the random baseline (13.6%). These results indicate that scGPA achieved a higher correct prediction rate than the two comparator models and the random baseline. We subsequently evaluated scGPA using a public osteosarcoma single-cell dataset and performed qRT-PCR validation in 143B osteosarcoma cells. Among genes with significant experimental changes, scGPA achieved a directional concordance of 76.9%. When all tested downstream genes were counted, 37.0% were directionally correct, 51.9% showed no significant change and 11.1% changed in the opposite direction. CONCLUSIONS: scGPA provides a practical workflow system for predicting and prioritizing direction-specific downstream transcriptional responses after target-gene perturbation. By integrating single-cell regulatory network inference, signed virtual perturbation and LLM-assisted interpretation, scGPA supports target-gene function inference and downstream mechanistic investigation from single-cell transcriptomic data.

Single-Cell Gene Expression Analysis

Interaction between toll-like receptor 4 polymorphism and abdominal obesity on ovarian cancer risk in Chinese women.

OBJECTIVES: the aim of this study was to evaluate the impact of TLR4 gene single nucleotide polymorphisms (SNPs) and additional TLR4 gene SNP- SNP and SNP- abdominal obesity (AO) interaction on ovarian cancer (OC) risk. METHODS: Generalized multifactor dimensionality reduction method were utilized to identify the most informative interactions between four SNPs in the TLR4 gene and abdominal obesity. Logistic regression was employed to investigate the association between 4 SNPs within TLR4 gene and OC risk, and additional SNP- SNP and gene- AO interaction on OC risk, ORs (95%CI) were calculated. RESULTS: The analysis of logistic regression indicated a markedly elevated risk of OC in individuals carrying either the rs4986790-G or rs11536889-C alleles in the TLR4 gene compared to those with the standard genetic variations, adjusted ORs (95%CI) were 1.61 (1.28-1.96) and 1.48 (1.09-1.91). GMDR analysis indicated a significant two-locus model (p = 0.018) involving rs4986790 and rs11536889, and a significant two-locus model (p = 0.001) involving rs4986790 and AO. Participants with rs4986790- AG/GG and rs11536889GC/ CC genotype has the highest OC risk, compared to participants with rs4986790-AA and rs11536889-GG genotype, OR (95%CI) = 2.58 (1.46-3.71), and abdominal obese participants with rs4986790- AG/GG genotype have the highest OC risk, compared to non- abdominal obese participants with rs4986790-AA genotype, OR (95%CI) = 3.17 (1.78-4.58). CONCLUSIONS: The findings suggested that TLR4 gene rs4986790 and rs11536889 polymorphisms were associated with increased OC risk. Significant interaction also existed between rs4986790 and AO, which means that the WC levels may influence the impact of rs4986790 on OC risk.

Adult

ABO exon polymorphisms are related to ischemic stroke in a Chinese Han population.

BACKGROUND: Recent research have underscored the relation of ABO blood group system to cerebrovascular disorders predisposition. The present investigation endeavors to delve into the relationship between ABO polymorphisms and ischemic stroke (IS) risk. METHODS: A cohort of 646 IS patients and 649 matched healthy controls was recruited. Genotyping of five SNPs within ABO were conducted by Agena MassARRAY platform. Logistic regression models were employed to estimate odds ratios (ORs) and 95% confidence intervals (CIs). Additionally, SNP-SNP interaction was assessed by multifactor dimensionality reduction (MDR) method. Furthermore, Analysis of Variance (ANOVA) was utilized to explore the association between genotypes and blood lipid profiles. RESULTS: The study identified an elevated IS risk associated with rs8176740 and rs8176720 in the overall population. Notably, ABO rs8176720 emerged as the most informative single-locus model for IS susceptibility. These variants were related to an elevated IS risk, specifically in female subjects, the subgroup aged > 64 years, non-smokers, drinkers or non-drinkers. Moreover, rs8176749 and rs8176745 were associated with red blood cell count levels and total bilirubin levels. CONCLUSION: This study firstly demonstrated the association of ABO rs8176740 and rs8176720 with IS incidence, which increased the understanding regarding the effect of ABO on IS pathogenesis.

Aged

PLNMFG: Pseudo-label guided non-negative matrix factorization model with graph constraint for single-cell multi-omics data clustering.

The development of single-cell multi-omics sequencing technologies has enabled the simultaneous analysis of multi-omics data within the same cell. Accurate clustering of these cells is crucial for downstream analyses of complex biological functions. Despite significant advances in multi-omics integration approaches, current methodologies exhibit two major limitations. First, they inadequately incorporate prior biological knowledge from various omic layers. Second, these methods often conduct independent dimensionality reduction on individual omic datasets, thereby failing to capture the intrinsic complementary information and potentially overlooking crucial cross-platform interactions. Motivated by these, this study investigates a non-negative matrix factorization model called PLNMFG, which integrates the unified latent representation learning that retains the features between and within omics and the cluster structure learning that retains the intrinsic structure of the data into one joint framework. Specially, PLNMFG performs adaptive imputation to handle dropout events and uses prior pseudo-labels as constraints during the process of collective non-negative matrix factorization, as a result, a more robust latent representation that preserves the double similarity information is obtained. Graph Laplacian constraint is applied during clustering which further preserves structure characteristic of multi-omics data. In addition, the weight of each omic is adaptively learned based on the omic contribution. A series of experiments on 8 benchmark datasets show that our model performs well in terms of clustering accuracy and computational efficiency.

Single-Cell Analysis

Myeloid landscape of BRAF-mutant papillary thyroid cancer and thyroiditis.

Papillary thyroid cancer (PTC) is less aggressive when associated with lymphocytic thyroiditis (LT), even in the presence of oncogenic BRAF, including smaller tumours, less lymph node involvement and reduced extrathyroidal extension. To investigate possible immune mechanisms underlying this association, we compared the tumour microenvironment of PTC-BRAF with LT and that without LT using single-cell RNA sequencing (scRNA-seq). Single-cell libraries were generated from fresh and fixed tumour samples with post-dissociation viability >70% using the 10x Genomics Chromium Platform and sequenced on an Illumina NovaSeq 6000. We analysed scRNA-seq data from 11 PTC-BRAF tumours: four with LT (one publicly available sample) and seven without LT. Downstream analyses included quality control, batch correction, dimensionality reduction, and differential gene expression analysis. We found that neutrophils were the predominant myeloid cell type in PTCs without LT. Thyrocytes without LT showed significant expression of the neutrophil recruitment chemokine ECRG4. In the absence of LT, neutrophils expressed oncogenic genes with poor clinical outcomes. In contrast, thyrocytes from tumours with LT showed increased expression of MHC-II antigen presentation, consistent with effective immune surveillance. Thyrocytes and macrophages in the presence of LT showed enrichment of interferon gamma response pathways. Our data suggest that LT in thyroid cancer is associated with enhanced antigen presentation and fewer features of pro-tumourigenic innate immune activity. These results identify previously under-recognised innate immune cell population and associated transcriptomic features, which suggest new mechanisms to target immune treatments in PTC refractory to other therapies.

Humans

Recent Advances in Multi-Omics of Systemic Lupus Erythematosus.

This comprehensive narrative review examines recent advances in multi-omics research for Systemic Lupus Erythematosus (SLE), emphasizing integrated approaches over single-omics studies. The review critically evaluates technological advancements, methodological innovations, and clinical applications while identifying current limitations and future research directions. We conducted a comprehensive narrative review following SANRA guidelines, searching PubMed, Web of Science, Scopus, and Embase, covering publications from January 2018 to June 2025. The review focuses on studies integrating two or more omics layers in SLE research, with emphasis on computational methods, biomarker validation, and clinical applications. Multi-omics integration has revealed critical insights into SLE pathogenesis, including immune cell heterogeneity, gene-environment interactions, and metabolic dysregulation. However, significant challenges remain in data integration methodologies, small sample sizes, and biomarker reproducibility. Current computational approaches include early integration (concatenation), intermediate integration (joint dimensionality reduction), and late integration (ensemble methods). While multi-omics approaches offer unprecedented insights into SLE complexity, standardized integration protocols and robust validation frameworks are urgently needed. Small sample sizes and heterogeneity issues limit reproducibility, particularly affecting biomarker discovery and clinical translation. Multi-omics integration represents a paradigm shift toward precision medicine in SLE, but realizing this potential requires addressing current methodological limitations, standardizing validation processes, and developing robust computational frameworks for reliable clinical applications.

Humans

MKMC enables reference-free transcriptomic analysis using k-mer representations.

Traditional RNA-seq analysis depends heavily on genome alignment and gene annotation, limiting its utility in non-model organisms and introducing biases that can obscure regulatory complexity. We present MKMC (Multi-sample Kmer Counter), a scalable, reference-free toolkit for RNA-seq analysis that leverages k-mer-based statistics to detect biological variation without requiring alignment. MKMC integrates fast k-mer counting, abundance matrix generation, normalization, dimensionality reduction, and differential analysis into a unified workflow. Across diverse datasets, MKMC recapitulates key biological signals-including sex differences in killifish liver-and matches alignment-based pipelines in differential expression analysis and transcriptomic age prediction. Notably, MKMC detects isoform-specific events missed by traditional methods, one of which we validated using in situ hybridization. These results reveal previously hidden isoform-level regulatory events that contribute to sex- and age-associated transcriptional programs. MKMC offers a robust, extensible alternative to alignment-based approaches, enabling transcriptomic discovery across both model and non-model systems. While we focus here on RNA-seq as a primary application, MKMC is broadly applicable to any k-mer-based analysis of next-generation sequencing data.

MKMC

[Dorsal stabilization of thoracic and lumbar vertebral injuries].

The concept of angle-stable transpedicular screw-rod instrumentation, realized in the different models of internal spine fixators with intrinsic stability, allows secure stabilization of the most unstable fracture patterns, limited-segment fixation and three-dimensional reduction of the fragments. The canal diameter is improved by ligamentotaxis and, if necessary, by hemilaminectomy and fragment impaction. Late collapse of the upper disk space must be anticipated and may lead to some increase in kyphotic deformity. The situations are identified where an additional formal interbody fusion is recommended.

Fracture Fixation, Internal

[Reverberation of the hindlimb rudimentation on its innervation in squamate reptiles].

When the dimensional reduction of the hind limb begins, a first caudal displacement of the lombar part of the lombo-sacral plexus - which involves the loss of the first root of the sacral part -- appears with a threshold in the increase in the number of presacral vertebrae. This a first indication of the serpentiform tendancy. Others thresholds can conduct to produce the disappearance of the sacral vertebrae and sacral root. The qualitative reduction only concerns the terminal branches of the plexus and does not seem to be associated with the vertebral elongation. If a caudo-proximal reduction of the brachial plexus occurs early in the lepidosaurian line and exists in all the Squamata, even in the Iguana which have well developed limbs, it is not the same for the reduction of the lombo-sacral plexus which does not appear in these Iguana. At last, if the reduction modalities of the both plexus are often differents, their supposed displacements facilitate the extension of the intermediate vertebral region.

Animals

Three-dimensional analysis of infarct size reduction after administration of gallopamil in dogs.

Several interventions have been shown to protect the ischemic myocardium from evolving to necrosis. However, three-dimensional evaluation of infarct size reduction using anatomical measurements is lacking. Therefore, Tc-99m labeled albumin microspheres were injected into the left atrium of 16 dogs, 1 min after coronary artery occlusion, to assess the hypoperfused zone. After 15 min, 8 dogs received, intravenously, gallopamil (Procorum) (Ga: 0.08 mg/k as a bolus + 0.2 mg/k/h for 6 h) and the remaining 8 dogs served as controls. At sacrifice (6 h), the left ventricle was cut into 3 mm slices, stained with triphenyltetrazolium chloride to delineate the size of infarction and autoradiographed to delineate the hypoperfused zone. Lateral reduction was calculated by the difference distance between hypoperfused zone and infarct size on the endocardium, mid-myocardium and epicardium; radial reduction was calculated by the difference of the transmural distance taken at mid-infarction and mid-hypoperfusion and the longitudinal reduction was calculated by the distance of the hypoperfused slices that did not infarct. With this technique it could be demonstrated that gallopamil reduced the size of infarction in all three dimensions.

Animals

Lipid fluidity and membrane protein dynamics.

Membrane fluidity plays an important role in cellular functions. Membrane proteins are mobile in the lipid fluid environment; lateral diffusion of membrane proteins is slower than expected by theory, due to both the effect of protein crowding in the membrane and to constraints from the aqueous matrix. A major aspect of diffusion is in macromolecular associations: reduction of dimensionality for membrane diffusion facilitates collisional encounters, as those concerned with receptor-mediated signal transduction and with electron transfer chains. In mitochondrial electron transfer, diffusional control is prevented by the excess of collisional encounters between fast-diffusing ubiquinone and the respiratory complexes. Another aspect of dynamics of membrane proteins is their conformational flexibility. Lipids may induce the optimal conformation for catalytic activity. Breaks in Arrhenius plots of membrane-bound enzymes may be related to lipid fluidity: the break could occur when a limiting viscosity is reached for catalytic activity. Viscosity can affect protein conformational changes by inhibiting thermal fluctuations to the inner core of the protein molecule.

Diffusion

Elemental classification in multi-detector stem images using image analysis clustering techniques.

The multi-detector STEM images form a multi-dimensional measurement space where picture elements of morphologically distinct regions cluster and can be separated for classification using image processing clustering techniques. These images are highly correlated and improve the classification only at the high cost of adding dimensionality. Data reduction techniques which take the class separation into account can be used to compress the useful information carried by these images into a few components to which clustering techniques can be successfully applied. At high magnification, a slight displacement is sometimes observed, which can slightly impair the classification result. The redundancy of the information in the quadrant images suggests their use for phase retrieval - which, added to an energy loss data channel, may greatly improve the classification.

Animals

Stimulus equalization: temporary reduction of stimulus complexity to facilitate discrimination learning.

Learners with limited behavioral repertoires often have difficulty discriminating complex, multidimensional stimuli. Procedures that use gradual stimulus change have been developed to facilitate such discrimination, but these procedures are often difficult to implement and costly in terms of teacher time and expertise. This study investigated the effectiveness of stimulus equalization, an error reduction procedure involving an abrupt but temporary reduction of dimensional complexity. A microcomputer was used for stimulus presentation, data collection, and response analyses. Preschool children's responding to groups of four stimuli differing along several dimensions was analyzed on four discrimination tasks under several conditions. On each task, one element from a different dimension was programmed as correct. When trial-and-error training failed to establish the discrimination, equalization training began in which differences in the irrelevant dimensions were eliminated. When correct responding developed, the differences were reinstated, and correct performance was maintained in all but one instance. In a repeated acquisition design, stimulus equalization was found to be generally effective and superior to the trial-and-error method. Implications for computer-assisted instruction are discussed.

Attention

Binding of basic peptides to acidic lipids in membranes: effects of inserting alanine(s) between the basic residues.

We studied the binding of peptides containing five basic residues to membranes containing acidic lipids. The peptides have five arginine or lysine residues and zero, one, or two alanines between the basic groups. The vesicles were formed from mixtures of a zwitterionic lipid, phosphatidylcholine, and an acidic lipid, either phosphatidylserine or phosphatidylglycerol. Measuring the binding using equilibrium dialysis, ultrafiltration, and electrophoretic mobility techniques, we found that all peptides bind to the membranes with a sigmoidal dependence on the mole fraction of acidic lipid. The sigmoidal dependence (Hill coefficient greater than 1 or apparent cooperativity) is due to both electrostatics and reduction of dimensionality and can be described by a simple model that combines Gouy-Chapman-Stern theory with mass action formalism. The adjustable parameter in this model is the microscopic association constant k between a basic residue and an acidic lipid (1 less than k less than 10 M-1). The addition of alanine residues decreases the affinity of the peptides for the membranes; two alanines inserted between the basic residues reduces k 2-fold. Equivalently, the affinity of the peptide for the membrane decreases 10-fold, probably due to a combination of local electrostatic effects and the increased loss of entropy that may occur when the more massive alanine-containing peptides bind to the membrane. The arginine peptides bind more strongly than the lysine peptides: k for an arginine residue is 2-fold higher than for a lysine residue. Our results imply that a cluster of arginine and lysine residues with interspersed electrically neutral amino acids can bind a significant fraction of a cytoplasmic protein to the plasma membrane if the cluster contains more than five basic residues.

Alanine