PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Dimensionality Reduction”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Mapping peroxidase in plant tissues by scanning electrochemical microscopy.

Scanning electrochemical microscopy has been firstly used to map the enzymatic activity in natural plant tissues. The peroxidase (POD) was maintained in its original state in the celery (Apium graveolens L.) tissues and electrochemically visualized under its native environment. Ferrocenemethanol (FMA) was selected as a mediator to probe the POD in celery tissues based on the fact that POD catalyzed the oxidation of FMA by H(2)O(2) to increase FMA(+) concentration. Two-dimensional reduction current profiles for FMA(+) produced images indicating the distribution and activity of the POD at the surface of the celery tissues. These images showed that the POD was widely distributed in the celery tissues, and larger amounts were found in some special regions such as the center of celery stem and around some vascular bundles.

Apium↗

Genetic heterogeneity affects the risk of incident depression, comorbidity, and response to environment: A prospective trajectory study.

BACKGROUND: Depression exhibits significant heterogeneity in its genetic underpinnings. The role of genetic components in the development of depression and its comorbidities remains insufficiently explored. METHODS: First, depression risk loci from a large-scale genome-wide meta-analysis were annotated to Gene Ontology (GO) terms by functional enrichment. GO-based polygenic risk scores (GO-PRS) were then calculated for individuals in the UK Biobank. Principal component analysis (PCA) was applied for dimensionality reduction, followed by cluster analysis to identify genetic subtypes of depression. Multistate models were applied to assess the impact of genetic patterns on the trajectory from healthy status to incident depression, and depression to 26 subsequent diseases, as well as the associations between environmental factors and disease trajectories across genetic subtypes. RESULTS: Participants were categorized into three genetic subtypes: immune-dominant, neuro-dominant, and comprehensive-risk. Significant differences in risk of depression and subsequent diseases, and susceptibility to environmental factors were observed across subtypes. Comprehensive-risk subtype showed higher risks of depression compared to immune-dominant (HR: 1.10, 95% CI: 1.05-1.15) and neuro-dominant subtype (HR: 1.12, 95% CI: 1.08-1.16). Comprehensive-risk subtype exhibited higher risks of transition from depression to subsequent diseases, such as anemia compared to immune-dominant subtype, and diseases of the digestive system compared to neuro-dominant subtype. Environmental factors were more strongly associated with the transition from depression to subsequent diseases in immune-dominant and comprehensive-risk subtypes, including cardiovascular, respiratory, and metabolic diseases. CONCLUSIONS: Our findings highlight the genetic heterogeneity of depression and comorbidities, and shed light on how genetic components modulate responses to environmental factors.

Humans↗

The development of features in object concepts.

According to one productive and influential approach to cognition, categorization, object recognition, and higher level cognitive processes operate on a set of fixed features, which are the output of lower level perceptual processes. In many situations, however, it is the higher level cognitive process being executed that influences the lower level features that are created. Rather than viewing the repertoire of features as being fixed by low-level processes, we present a theory in which people create features to subserve the representation and categorization of objects. Two types of category learning should be distinguished. Fixed space category learning occurs when new categorizations are representable with the available feature set. Flexible space category learning occurs when new categorizations cannot be represented with the features available. Whether fixed or flexible, learning depends on the featural contrasts and similarities between the new category to be represented and the individuals existing concepts. Fixed feature approaches face one of two problems with tasks that call for new features: If the fixed features are fairly high level and directly useful for categorization, then they will not be flexible enough to represent all objects that might be relevant for a new task. If the fixed features are small, subsymbolic fragments (such as pixels), then regularities at the level of the functional features required to accomplish categorizations will not be captured by these primitives. We present evidence of flexible perceptual changes arising from category learning and theoretical arguments for the importance of this flexibility. We describe conditions that promote feature creation and argue against interpreting them in terms of fixed features. Finally, we discuss the implications of functional features for object categorization, conceptual development, chunking, constructive induction, and formal models of dimensionality reduction.

Child↗

Single-Cell Proteomics Reveals Proteome Remodeling and Cellular Heterogeneity During NGF-Induced PC12 Neuronal Differentiation.

Single-cell proteomics enables direct measurement of cellular heterogeneity during dynamic biological processes, but its application to fragile and highly adherent neuronal models remains challenging. Here, we developed and applied an optimized single-cell proteomics workflow to characterize proteome remodeling during nerve growth factor (NGF)-induced differentiation of PC12 cells. To enable reliable single-cell analysis, we implemented gentle dissociation, antiaggregation strategies, and thermal inkjet-based cell dispensing, achieving high accuracy in single-cell isolation. Inclusion of n-dodecyl-β-d-maltoside (DDM) improved recovery of membrane-associated and low-solubility proteins. Coupled with LC-ion mobility-mass spectrometry, this workflow enabled quantification of 2,000-3,000 proteins per cell across the differentiation time course. Single-cell proteomic analysis revealed progressive and heterogeneous proteome remodeling during differentiation. While undifferentiated cells formed a relatively homogeneous population, later stages (Days 4-6) exhibited increased variability, including multimodal protein abundance distributions and separation into distinct subpopulations. Dimensionality reduction, clustering, and non-negative matrix factorization identified multiple coexisting proteomic states within the same time points, reflecting asynchronous differentiation trajectories. These subpopulations were characterized by coordinated differences in pathways related to intracellular trafficking, protein translation, cytoskeletal organization, and neuronal maturation. Comparison with bulk proteomics demonstrated that proteins associated with differentiated neuronal states, including those involved in neurite formation and structural remodeling, are underrepresented in population-averaged measurements but are enriched within specific single-cell subpopulations. Temporal and cluster-resolved analyses further revealed distinct protein expression trajectories, including early decreases in cell cycle and metabolic pathways and later increases in neuronal structural and regulatory proteins. Together, this study establishes an optimized workflow for single-cell proteomics of neuronal systems and demonstrates that NGF-induced PC12 differentiation proceeds through heterogeneous and divergent proteomic states that are not resolved by bulk analysis.

Animals↗

An adaptive strategy for single- and multi-cluster gene assignment.

Strict assignment of genes to one class, dimensionality reduction, a priori specification of the number of classes, the need for a training set, nonunique solution, and complex learning mechanisms are some of the inadequacies of current clustering algorithms. Existing algorithms cluster genes on the basis of high positive correlations between their expression patterns. However, genes with strong negative correlations can also have similar functions and are most likely to have a role in the same pathways. To address some of these issues, we propose the adaptive centroid algorithm (ACA), which employs an analysis of variance (ANOVA)-based performance criterion. The ACA also uses Euclidian distances, the center-of-mass principle for heterogeneously distributed mass elements, and the given data set to give unique solutions. The proposed approach involves three stages. In the first stage a two-way ANOVA of the gene expression matrix is performed. The two factors in the ANOVA are gene expression and experimental condition. The residual mean squared error (MSE) from the ANOVA is used as a performance criterion in the ACA. Finally, correlated clusters are found based on the Pearson correlation coefficients. To validate the proposed approach, a two-way ANOVA is again performed on the discovered clusters. The results from this last step indicate that MSEs of the clusters are significantly lower compared to that of the fibroblast-serum gene expression matrix. The ACA is employed in this study for single- as well as multi-cluster gene assignments.

Algorithms↗

Nonlinear mapping networks.

Among the many dimensionality reduction techniques that have appeared in the statistical literature, multidimensional scaling and nonlinear mapping are unique for their conceptual simplicity and ability to reproduce the topology and structure of the data space in a faithful and unbiased manner. However, a major shortcoming of these methods is their quadratic dependence on the number of objects scaled, which imposes severe limitations on the size of data sets that can be effectively manipulated. Here we describe a novel approach that combines conventional nonlinear mapping techniques with feed-forward neural networks, and allows the processing of data sets orders of magnitude larger than those accessible with conventional methodologies. Rooted on the principle of probability sampling, the method employs a classical algorithm to project a small random sample, and then "learns" the underlying nonlinear transform using a multilayer neural network trained with the back-propagation algorithm. Once trained, the neural network can be used in a feed-forward manner to project the remaining members of the population as well as new, unseen samples with minimal distortion. Using examples from the fields of image processing and combinatorial chemistry, we demonstrate that this method can generate projections that are virtually indistinguishable from those derived by conventional approaches. The ability to encode the nonlinear transform in the form of a neural network makes nonlinear mapping applicable to a wide variety of data mining applications involving very large data sets that are otherwise computationally intractable.

Journal Article↗

Advances in diversity profiling and combinatorial series design.

Rapid advances in synthetic and screening technology have recently enabled the simultaneous synthesis and biological evaluation of large chemical libraries containing hundreds to tens of thousands of compounds, using molecular diversity as a means to design and prioritize experiments. This paper reviews some of the most important computational work in the field of diversity profiling and combinatorial library design, with particular emphasis on methodology and applications. It is divided into four sections that address issues related to molecular representation, dimensionality reduction, compound selection, and visualization.

Chemistry, Pharmaceutical↗

Association of AGER genetic variants with chronic obstructive pulmonary disease susceptibility in Southern Chinese Han populations.

OBJECTIVE: Chronic obstructive pulmonary disease (COPD) remains a leading cause of disability and mortality among elderly populations. Studies indicate that AGER plays a critical regulatory role in the pathogenesis of respiratory disorders. However, the genetic variations in AGER to COPD susceptibility remain incompletely understood. This study employs a case-control design to investigate associations between AGER genetic variants and COPD risk in the Southern Chinese Han population. METHODS: This study enrolled 270 COPD patients and 271 healthy controls. AGER single-nucleotide polymorphisms (SNPs) were analysed using the MassARRAY iPLEX platform. Logistic regression models evaluated associations between AGER polymorphisms and COPD susceptibility, with false discovery rate (FDR) correction applied to mitigate multiple testing errors. SNP-SNP interactions were investigated through multifactor dimensionality reduction (MDR) analysis. Expression quantitative trait locus (eQTL) data from the GTEx database were further analysed to assess regulatory relationships between SNPs and AGER gene expression levels. RESULTS: This study showed that rs3134941 (G allele, OR = 0.21, 95% CI = 0.10-0.41, p (FDR) = 0.001) and rs3131300 (G allele, OR = 0.32, 95% CI = 0.20-0.49, p (FDR) = 0.0001) were significantly associated with a reduced susceptibility to COPD. MDR indicated that rs3131300 was the optimal predictive model for COPD risk. Additionally, initial mechanistic investigations utilizing the GTEx database identify rs3134941 (C > G) and rs3131300 (A > G) as significant expression quantitative trait loci for AGER mRNA in cell-cultured fibroblasts and whole blood. CONCLUSION: Our study demonstrated that AGER genetic variants might play a protective role in the progression of COPD.

Aged↗

Linear transformations of data space in MEG.

Magnetoencephalography (MEG) is a method which allows the non-invasive measurement of the minute magnetic field which is generated by ion currents in the brain. Due to the complex sensitivity profile of the sensors, the measured data are a non-trivial representation of the currents where information specific to local generators is distributed across many channels and each channel contains a mixture of contributions from many such generators. We propose a framework which generates a new representation of the data through a linear transformation which is designed so that some desired property is optimized in one or more new virtual channel(s). First figures of merit are suggested to describe the relation between the measured data and the underlying currents. Within this context the new framework is established by first showing how the transformation matrix itself is designed and then by its application to real and simulated data. The results demonstrate that the proposed linear transformations of data space provide a computationally efficient tool for analysis and a very much needed dimensional reduction of the data.

Brain↗

Negative dataset selection impacts machine learning-based predictors for multiple bacterial species promoters.

MOTIVATION: Advances in bacterial promoter predictors based on machine learning have greatly improved identification metrics. However, existing models overlooked the impact of negative datasets, previously identified in GC-content discrepancies between positive and negative datasets in single-species models. This study aims to investigate whether multiple-species models for promoter classification are inherently biased due to the selection criteria of negative datasets. We further explore whether the generation of synthetic random sequences (SRS) that mimic GC-content distribution of promoters can partly reduce this bias. RESULTS: Multiple-species predictors exhibited GC-content bias when using CDS as a negative dataset, suggested by specificity and sensibility metrics in a species-specific manner, and investigated by dimensionality reduction. We demonstrated a reduction in this bias by using the SRS dataset, with less detection of background noise in real genomic data. In both scenarios DNABERT showed the best metrics. These findings suggest that GC-balanced datasets can enhance the generalizability of promoter predictors across Bacteria. AVAILABILITY AND IMPLEMENTATION: The source code of the experiments is freely available at https://github.com/maigonzalezh/MultispeciesPromoterClassifier.

Machine Learning↗

Approximate geodesic distances reveal biologically relevant structures in microarray data.

MOTIVATION: Genome-wide gene expression measurements, as currently determined by the microarray technology, can be represented mathematically as points in a high-dimensional gene expression space. Genes interact with each other in regulatory networks, restricting the cellular gene expression profiles to a certain manifold, or surface, in gene expression space. To obtain knowledge about this manifold, various dimensionality reduction methods and distance metrics are used. For data points distributed on curved manifolds, a sensible distance measure would be the geodesic distance along the manifold. In this work, we examine whether an approximate geodesic distance measure captures biological similarities better than the traditionally used Euclidean distance. RESULTS: We computed approximate geodesic distances, determined by the Isomap algorithm, for one set of lymphoma and one set of lung cancer microarray samples. Compared with the ordinary Euclidean distance metric, this distance measure produced more instructive, biologically relevant, visualizations when applying multidimensional scaling. This suggests the Isomap algorithm as a promising tool for the interpretation of microarray data. Furthermore, the results demonstrate the benefit and importance of taking nonlinearities in gene expression data into account.

Algorithms↗

Multimodal deep learning for immunotherapy response prediction and biomarker discovery in non-small cell lung cancer.

OBJECTIVE: Immunotherapy has emerged as a promising treatment for advanced non-small cell lung cancer (NSCLC), but accurately predicting which patients will benefit from it remains a major clinical challenge. To address this, we aim to develop a novel multimodal method, DeepAFM, that integrates histopathology, genomic features, and clinical information to predict patient responses to anti-PD-(L)1 immunotherapy. MATERIALS AND METHODS: A total of 93 patients with advanced NSCLC were included in this study. Histopathological whole-slide images were processed using a self-supervised VQVAE2 for representation learning. PCA and K-means clustering were then applied for dimensionality reduction and feature grouping. Key regions of interest were visualized through permutation importance evaluation and color-coding techniques. The extracted histopathological features, along with genomic alterations and clinical variables, were integrated into the DeepAFM multimodal prediction model. RESULTS: The DeepAFM achieved a high predictive performance with an area under the curve (AUC) of 0.77 (95% confidence interval: 0.69-1.00). Attention-based heatmaps revealed that the model could identify critical pathological patterns, genomic mutations, and clinical indicators associated with patient responses to immunotherapy. DISCUSSION: The integration of multimodal data enabled the model to capture complex interactions among pathology, genomics, and clinical characteristics, enhancing the interpretability and predictive power of immunotherapy response prediction. The visualization techniques facilitated the identification of biologically meaningful features and potential biomarkers. CONCLUSION: This study demonstrates the effectiveness of the DeepAFM in predicting responses to immunotherapy in advanced NSCLC. The approach not only improves prediction accuracy but also provides valuable insights for personalized treatment strategies and biomarker discovery.

Humans↗

Genomic hallmarks of depot medroxyprogesterone acetate-associated meningiomas.

BACKGROUND: Population-based studies have linked progestin exposure to increased meningioma risk. However, the molecular basis of meningiomas associated with depot medroxyprogesterone acetate (DMPA)-a common injectable contraceptive-remains undefined. METHODS: We performed an integrated clinicopathologic and genomic analysis of meningiomas from 10 women with long-term DMPA exposure. Tumors underwent histopathological analysis, targeted sequencing, and DNA methylation profiling. Data were integrated with reference cohorts (Baylor and Heidelberg) and analyzed through classifier assignment, consensus clustering, copy number analysis, differential methylation testing, and dimensionality reduction. RESULTS: Depot medroxyprogesterone acetate-associated meningiomas were all newly diagnosed, World Health Organization grade 1 tumors with a predilection for the anterior and central skull base (n = 6). Nine patients harbored multiple meningiomas. Four experienced regression of untreated meningiomas following DMPA cessation, while 5 demonstrated stabilization. Histopathology demonstrated relative overrepresentation of metaplastic morphology, an uncommon meningioma subtype. All DMPA-associated meningiomas mapped to benign molecular groups, and most exhibited low copy number alteration burden. Targeted sequencing revealed enrichment for TRAF7 mutations (n = 5), with no NF2 mutations detected. Eight tumors shared consensus cluster identity, with cohesive grouping on principal component analysis and t-distributed stochastic neighbor embedding. No differential methylation was identified at the progesterone receptor locus. CONCLUSIONS: Depot medroxyprogesterone acetate-associated meningiomas represent a recognizable phenotype within the broader NF2-wildtype/TRAF7-enriched spectrum of benign meningiomas, characterized by chromosomal stability, a shared methylation profile, tumor multiplicity, and regression or stabilization following DMPA cessation. While derived from a small single-institution cohort, these findings provide a molecular framework for understanding progestin-associated meningioma biology, reinterpreting epidemiologic literature, and informing population-level risk stratification.

Humans↗

Functional renormalization group and the field theory of disordered elastic systems.

We study elastic systems, such as interfaces or lattices, pinned by quenched disorder. To escape triviality as a result of "dimensional reduction," we use the functional renormalization group. Difficulties arise in the calculation of the renormalization group functions beyond one-loop order. Even worse, observables such as the two-point correlation function exhibit the same problem already at one-loop order. These difficulties are due to the nonanalyticity of the renormalized disorder correlator at zero temperature, which is inherent to the physics beyond the Larkin length, characterized by many metastable states. As a result, two-loop diagrams, which involve derivatives of the disorder correlator at the nonanalytic point, are naively "ambiguous." We examine several routes out of this dilemma, which lead to a unique renormalizable field theory at two-loop order. It is also the only theory consistent with the potentiality of the problem. The beta function differs from previous work and the one at depinning by novel "anomalous terms." For interfaces and random-bond disorder we find a roughness exponent zeta=0.208 298 04epsilon+0.006 858epsilon(2), epsilon=4-d. For random-field disorder we find zeta=epsilon/3 and compute universal amplitudes to order O(epsilon(2)). For periodic systems we evaluate the universal amplitude of the two-point function. We also clarify the dependence of universal amplitudes on the boundary conditions at large scale. All predictions are in good agreement with numerical and exact results and are an improvement over one loop. Finally we calculate higher correlation functions, which turn out to be equivalent to those at depinning to leading order in epsilon.

Journal Article↗

Black holes in Gödel universes and pp waves.

We find exact solutions for rotating and nonrotating neutral black holes in the Gödel universe of five-dimensional minimal supergravity theory. We also describe the embedding of this solution in M-theory. After dimensional reduction and T-duality, we obtain a supergravity solution corresponding to placing a black string in a pp-wave background.

Journal Article↗

Probability-based current dipole localization from biomagnetic fields.

Focal biomagnetic sources are described as pointlike current dipoles. The dipole parameters, position, and moment coordinates are commonly determined from biomagnetic data using iterative nonlinear optimization algorithms such as the Levenberg-Marquardt algorithm. However, even for single-dipole sources, mislocalizations can occur due to side minima of the cost function or due to a wrong choice of the start vector. This can be shown by introducing a cost function where the independent variables are only the position coordinates instead of position and moment coordinates. This dimensional reduction--which is also possible for multiple dipole sources--is achieved by calculating the cost function at each position with the position and data-dependent, optimum dipole moments. We call these dipoles with--in a least squares sense--optimum moments, locally optimal dipoles. The visualization of such a single-dipole cost function and of the iteration steps of the Levenberg-Marquardt algorithm show why mislocalizations cannot be avoided. Therefore, we propose an alternative noniterative localization algorithm for single-dipole sources without this drawback. It uses localization probabilities calculated by means of the locally optimal dipoles. Besides the determination of the dipole parameters, the proposed algorithm furnishes a reliable error for each localization. Its effectiveness is shown with simulated and real patient data.

Algorithms↗

A computer-aided diagnostic system to characterize CT focal liver lesions: design and optimization of a neural network classifier.

In this paper, a computer-aided diagnostic (CAD) system for the classification of hepatic lesions from computed tomography (CT) images is presented. Regions of interest (ROIs) taken from nonenhanced CT images of normal liver, hepatic cysts, hemangiomas, and hepatocellular carcinomas have been used as input to the system. The proposed system consists of two modules: the feature extraction and the classification modules. The feature extraction module calculates the average gray level and 48 texture characteristics, which are derived from the spatial gray-level co-occurrence matrices, obtained from the ROIs. The classifier module consists of three sequentially placed feed-forward neural networks (NNs). The first NN classifies into normal or pathological liver regions. The pathological liver regions are characterized by the second NN as cyst or "other disease." The third NN classifies "other disease" into hemangioma or hepatocellular carcinoma. Three feature selection techniques have been applied to each individual NN: the sequential forward selection, the sequential floating forward selection, and a genetic algorithm for feature selection. The comparative study of the above dimensionality reduction methods shows that genetic algorithms result in lower dimension feature vectors and improved classification performance.

Algorithms↗

Probabilistic independent component analysis for functional magnetic resonance imaging.

We present an integrated approach to probabilistic independent component analysis (ICA) for functional MRI (FMRI) data that allows for nonsquare mixing in the presence of Gaussian noise. In order to avoid overfitting, we employ objective estimation of the amount of Gaussian noise through Bayesian analysis of the true dimensionality of the data, i.e., the number of activation and non-Gaussian noise sources. This enables us to carry out probabilistic modeling and achieves an asymptotically unique decomposition of the data. It reduces problems of interpretation, as each final independent component is now much more likely to be due to only one physical or physiological process. We also describe other improvements to standard ICA, such as temporal prewhitening and variance normalization of timeseries, the latter being particularly useful in the context of dimensionality reduction when weak activation is present. We discuss the use of prior information about the spatiotemporal nature of the source processes, and an alternative-hypothesis testing approach for inference, using Gaussian mixture models. The performance of our approach is illustrated and evaluated on real and artificial FMRI data, and compared to the spatio-temporal accuracy of results obtained from classical ICA and GLM analyses.

Algorithms↗