PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Dimensionality Reduction”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Advances in diversity profiling and combinatorial series design.

Rapid advances in synthetic and screening technology have recently enabled the simultaneous synthesis and biological evaluation of large chemical libraries containing hundreds to tens of thousands of compounds, using molecular diversity as a means to design and prioritize experiments. This paper reviews some of the most important computational work in the field of diversity profiling and combinatorial library design, with particular emphasis on methodology and applications. It is divided into four sections that address issues related to molecular representation, dimensionality reduction, compound selection, and visualization.

Chemistry, Pharmaceutical↗

Association of AGER genetic variants with chronic obstructive pulmonary disease susceptibility in Southern Chinese Han populations.

OBJECTIVE: Chronic obstructive pulmonary disease (COPD) remains a leading cause of disability and mortality among elderly populations. Studies indicate that AGER plays a critical regulatory role in the pathogenesis of respiratory disorders. However, the genetic variations in AGER to COPD susceptibility remain incompletely understood. This study employs a case-control design to investigate associations between AGER genetic variants and COPD risk in the Southern Chinese Han population. METHODS: This study enrolled 270 COPD patients and 271 healthy controls. AGER single-nucleotide polymorphisms (SNPs) were analysed using the MassARRAY iPLEX platform. Logistic regression models evaluated associations between AGER polymorphisms and COPD susceptibility, with false discovery rate (FDR) correction applied to mitigate multiple testing errors. SNP-SNP interactions were investigated through multifactor dimensionality reduction (MDR) analysis. Expression quantitative trait locus (eQTL) data from the GTEx database were further analysed to assess regulatory relationships between SNPs and AGER gene expression levels. RESULTS: This study showed that rs3134941 (G allele, OR = 0.21, 95% CI = 0.10-0.41, p (FDR) = 0.001) and rs3131300 (G allele, OR = 0.32, 95% CI = 0.20-0.49, p (FDR) = 0.0001) were significantly associated with a reduced susceptibility to COPD. MDR indicated that rs3131300 was the optimal predictive model for COPD risk. Additionally, initial mechanistic investigations utilizing the GTEx database identify rs3134941 (C > G) and rs3131300 (A > G) as significant expression quantitative trait loci for AGER mRNA in cell-cultured fibroblasts and whole blood. CONCLUSION: Our study demonstrated that AGER genetic variants might play a protective role in the progression of COPD.

Aged↗

Linear transformations of data space in MEG.

Magnetoencephalography (MEG) is a method which allows the non-invasive measurement of the minute magnetic field which is generated by ion currents in the brain. Due to the complex sensitivity profile of the sensors, the measured data are a non-trivial representation of the currents where information specific to local generators is distributed across many channels and each channel contains a mixture of contributions from many such generators. We propose a framework which generates a new representation of the data through a linear transformation which is designed so that some desired property is optimized in one or more new virtual channel(s). First figures of merit are suggested to describe the relation between the measured data and the underlying currents. Within this context the new framework is established by first showing how the transformation matrix itself is designed and then by its application to real and simulated data. The results demonstrate that the proposed linear transformations of data space provide a computationally efficient tool for analysis and a very much needed dimensional reduction of the data.

Brain↗

Negative dataset selection impacts machine learning-based predictors for multiple bacterial species promoters.

MOTIVATION: Advances in bacterial promoter predictors based on machine learning have greatly improved identification metrics. However, existing models overlooked the impact of negative datasets, previously identified in GC-content discrepancies between positive and negative datasets in single-species models. This study aims to investigate whether multiple-species models for promoter classification are inherently biased due to the selection criteria of negative datasets. We further explore whether the generation of synthetic random sequences (SRS) that mimic GC-content distribution of promoters can partly reduce this bias. RESULTS: Multiple-species predictors exhibited GC-content bias when using CDS as a negative dataset, suggested by specificity and sensibility metrics in a species-specific manner, and investigated by dimensionality reduction. We demonstrated a reduction in this bias by using the SRS dataset, with less detection of background noise in real genomic data. In both scenarios DNABERT showed the best metrics. These findings suggest that GC-balanced datasets can enhance the generalizability of promoter predictors across Bacteria. AVAILABILITY AND IMPLEMENTATION: The source code of the experiments is freely available at https://github.com/maigonzalezh/MultispeciesPromoterClassifier.

Machine Learning↗

Approximate geodesic distances reveal biologically relevant structures in microarray data.

MOTIVATION: Genome-wide gene expression measurements, as currently determined by the microarray technology, can be represented mathematically as points in a high-dimensional gene expression space. Genes interact with each other in regulatory networks, restricting the cellular gene expression profiles to a certain manifold, or surface, in gene expression space. To obtain knowledge about this manifold, various dimensionality reduction methods and distance metrics are used. For data points distributed on curved manifolds, a sensible distance measure would be the geodesic distance along the manifold. In this work, we examine whether an approximate geodesic distance measure captures biological similarities better than the traditionally used Euclidean distance. RESULTS: We computed approximate geodesic distances, determined by the Isomap algorithm, for one set of lymphoma and one set of lung cancer microarray samples. Compared with the ordinary Euclidean distance metric, this distance measure produced more instructive, biologically relevant, visualizations when applying multidimensional scaling. This suggests the Isomap algorithm as a promising tool for the interpretation of microarray data. Furthermore, the results demonstrate the benefit and importance of taking nonlinearities in gene expression data into account.

Algorithms↗

Multimodal deep learning for immunotherapy response prediction and biomarker discovery in non-small cell lung cancer.

OBJECTIVE: Immunotherapy has emerged as a promising treatment for advanced non-small cell lung cancer (NSCLC), but accurately predicting which patients will benefit from it remains a major clinical challenge. To address this, we aim to develop a novel multimodal method, DeepAFM, that integrates histopathology, genomic features, and clinical information to predict patient responses to anti-PD-(L)1 immunotherapy. MATERIALS AND METHODS: A total of 93 patients with advanced NSCLC were included in this study. Histopathological whole-slide images were processed using a self-supervised VQVAE2 for representation learning. PCA and K-means clustering were then applied for dimensionality reduction and feature grouping. Key regions of interest were visualized through permutation importance evaluation and color-coding techniques. The extracted histopathological features, along with genomic alterations and clinical variables, were integrated into the DeepAFM multimodal prediction model. RESULTS: The DeepAFM achieved a high predictive performance with an area under the curve (AUC) of 0.77 (95% confidence interval: 0.69-1.00). Attention-based heatmaps revealed that the model could identify critical pathological patterns, genomic mutations, and clinical indicators associated with patient responses to immunotherapy. DISCUSSION: The integration of multimodal data enabled the model to capture complex interactions among pathology, genomics, and clinical characteristics, enhancing the interpretability and predictive power of immunotherapy response prediction. The visualization techniques facilitated the identification of biologically meaningful features and potential biomarkers. CONCLUSION: This study demonstrates the effectiveness of the DeepAFM in predicting responses to immunotherapy in advanced NSCLC. The approach not only improves prediction accuracy but also provides valuable insights for personalized treatment strategies and biomarker discovery.

Humans↗

Genomic hallmarks of depot medroxyprogesterone acetate-associated meningiomas.

BACKGROUND: Population-based studies have linked progestin exposure to increased meningioma risk. However, the molecular basis of meningiomas associated with depot medroxyprogesterone acetate (DMPA)-a common injectable contraceptive-remains undefined. METHODS: We performed an integrated clinicopathologic and genomic analysis of meningiomas from 10 women with long-term DMPA exposure. Tumors underwent histopathological analysis, targeted sequencing, and DNA methylation profiling. Data were integrated with reference cohorts (Baylor and Heidelberg) and analyzed through classifier assignment, consensus clustering, copy number analysis, differential methylation testing, and dimensionality reduction. RESULTS: Depot medroxyprogesterone acetate-associated meningiomas were all newly diagnosed, World Health Organization grade 1 tumors with a predilection for the anterior and central skull base (n = 6). Nine patients harbored multiple meningiomas. Four experienced regression of untreated meningiomas following DMPA cessation, while 5 demonstrated stabilization. Histopathology demonstrated relative overrepresentation of metaplastic morphology, an uncommon meningioma subtype. All DMPA-associated meningiomas mapped to benign molecular groups, and most exhibited low copy number alteration burden. Targeted sequencing revealed enrichment for TRAF7 mutations (n = 5), with no NF2 mutations detected. Eight tumors shared consensus cluster identity, with cohesive grouping on principal component analysis and t-distributed stochastic neighbor embedding. No differential methylation was identified at the progesterone receptor locus. CONCLUSIONS: Depot medroxyprogesterone acetate-associated meningiomas represent a recognizable phenotype within the broader NF2-wildtype/TRAF7-enriched spectrum of benign meningiomas, characterized by chromosomal stability, a shared methylation profile, tumor multiplicity, and regression or stabilization following DMPA cessation. While derived from a small single-institution cohort, these findings provide a molecular framework for understanding progestin-associated meningioma biology, reinterpreting epidemiologic literature, and informing population-level risk stratification.

Humans↗

Functional renormalization group and the field theory of disordered elastic systems.

We study elastic systems, such as interfaces or lattices, pinned by quenched disorder. To escape triviality as a result of "dimensional reduction," we use the functional renormalization group. Difficulties arise in the calculation of the renormalization group functions beyond one-loop order. Even worse, observables such as the two-point correlation function exhibit the same problem already at one-loop order. These difficulties are due to the nonanalyticity of the renormalized disorder correlator at zero temperature, which is inherent to the physics beyond the Larkin length, characterized by many metastable states. As a result, two-loop diagrams, which involve derivatives of the disorder correlator at the nonanalytic point, are naively "ambiguous." We examine several routes out of this dilemma, which lead to a unique renormalizable field theory at two-loop order. It is also the only theory consistent with the potentiality of the problem. The beta function differs from previous work and the one at depinning by novel "anomalous terms." For interfaces and random-bond disorder we find a roughness exponent zeta=0.208 298 04epsilon+0.006 858epsilon(2), epsilon=4-d. For random-field disorder we find zeta=epsilon/3 and compute universal amplitudes to order O(epsilon(2)). For periodic systems we evaluate the universal amplitude of the two-point function. We also clarify the dependence of universal amplitudes on the boundary conditions at large scale. All predictions are in good agreement with numerical and exact results and are an improvement over one loop. Finally we calculate higher correlation functions, which turn out to be equivalent to those at depinning to leading order in epsilon.

Journal Article↗

Black holes in Gödel universes and pp waves.

We find exact solutions for rotating and nonrotating neutral black holes in the Gödel universe of five-dimensional minimal supergravity theory. We also describe the embedding of this solution in M-theory. After dimensional reduction and T-duality, we obtain a supergravity solution corresponding to placing a black string in a pp-wave background.

Journal Article↗

Probability-based current dipole localization from biomagnetic fields.

Focal biomagnetic sources are described as pointlike current dipoles. The dipole parameters, position, and moment coordinates are commonly determined from biomagnetic data using iterative nonlinear optimization algorithms such as the Levenberg-Marquardt algorithm. However, even for single-dipole sources, mislocalizations can occur due to side minima of the cost function or due to a wrong choice of the start vector. This can be shown by introducing a cost function where the independent variables are only the position coordinates instead of position and moment coordinates. This dimensional reduction--which is also possible for multiple dipole sources--is achieved by calculating the cost function at each position with the position and data-dependent, optimum dipole moments. We call these dipoles with--in a least squares sense--optimum moments, locally optimal dipoles. The visualization of such a single-dipole cost function and of the iteration steps of the Levenberg-Marquardt algorithm show why mislocalizations cannot be avoided. Therefore, we propose an alternative noniterative localization algorithm for single-dipole sources without this drawback. It uses localization probabilities calculated by means of the locally optimal dipoles. Besides the determination of the dipole parameters, the proposed algorithm furnishes a reliable error for each localization. Its effectiveness is shown with simulated and real patient data.

Algorithms↗

A computer-aided diagnostic system to characterize CT focal liver lesions: design and optimization of a neural network classifier.

In this paper, a computer-aided diagnostic (CAD) system for the classification of hepatic lesions from computed tomography (CT) images is presented. Regions of interest (ROIs) taken from nonenhanced CT images of normal liver, hepatic cysts, hemangiomas, and hepatocellular carcinomas have been used as input to the system. The proposed system consists of two modules: the feature extraction and the classification modules. The feature extraction module calculates the average gray level and 48 texture characteristics, which are derived from the spatial gray-level co-occurrence matrices, obtained from the ROIs. The classifier module consists of three sequentially placed feed-forward neural networks (NNs). The first NN classifies into normal or pathological liver regions. The pathological liver regions are characterized by the second NN as cyst or "other disease." The third NN classifies "other disease" into hemangioma or hepatocellular carcinoma. Three feature selection techniques have been applied to each individual NN: the sequential forward selection, the sequential floating forward selection, and a genetic algorithm for feature selection. The comparative study of the above dimensionality reduction methods shows that genetic algorithms result in lower dimension feature vectors and improved classification performance.

Algorithms↗

Probabilistic independent component analysis for functional magnetic resonance imaging.

We present an integrated approach to probabilistic independent component analysis (ICA) for functional MRI (FMRI) data that allows for nonsquare mixing in the presence of Gaussian noise. In order to avoid overfitting, we employ objective estimation of the amount of Gaussian noise through Bayesian analysis of the true dimensionality of the data, i.e., the number of activation and non-Gaussian noise sources. This enables us to carry out probabilistic modeling and achieves an asymptotically unique decomposition of the data. It reduces problems of interpretation, as each final independent component is now much more likely to be due to only one physical or physiological process. We also describe other improvements to standard ICA, such as temporal prewhitening and variance normalization of timeseries, the latter being particularly useful in the context of dimensionality reduction when weak activation is present. We discuss the use of prior information about the spatiotemporal nature of the source processes, and an alternative-hypothesis testing approach for inference, using Gaussian mixture models. The performance of our approach is illustrated and evaluated on real and artificial FMRI data, and compared to the spatio-temporal accuracy of results obtained from classical ICA and GLM analyses.

Algorithms↗

Use of IR Biotyper as a feasible methodology to type Klebsiella pneumoniae.

UNLABELLED: Klebsiella pneumoniae is one of the most frequently reported healthcare-associated pathogens. The current gold standard approach to perform the epidemiological typing of these bacteria is Whole Genome Sequencing (WGS), which is an expensive and challenging procedure. IR Biotyper (Bruker Daltonics, GmbH) is a new equipment based on Fourier transform infrared spectroscopy, which allows a rapid, low-cost, and user-friendly method to type bacterial isolates. However, there is a need for studies that evaluate the efficacy of the IR Biotyper. The aim of this study was to evaluate the capability of IR Biotyper to type K. pneumoniae according to sequence type (ST) and capsular type-using K locus (KL)-as well as to develop a classifier using machine learning. Seventy-three isolates of K. pneumoniae previously characterized by WGS were selected for IR Biotyper analysis using principal component analysis for dimensionality reduction, Euclidean, and unweighted pair group method with arithmetic mean (UPGMA) for clustering method, and spectra were analyzed in the 1,300-800 cm⁻¹ wavenumber range. Among these, 54 isolates were used to create a classifier, and 19 were used to validate the classifier. When considering the ST, ST307 was grouped in the same cluster as ST11. When KL was considered for the analysis, the clusters were 100% correctly grouped according to their KL type. Furthermore, the classifier developed was able to classify the isolates according to KL with a high concordance. This study showed that KL correlates well with KL for typing K. pneumoniae isolates using the IR Biotyper. Additionally, IR Biotyper demonstrated to be a cost-effective method and a promising tool to classify isolates within minutes. IMPORTANCE: Klebsiella pneumoniae is a major cause of severe hospital infections, and controlling its spread requires quick identification and comparison of bacterial strains. WGS is accurate but expensive, slow, and technically demanding. In this study, we evaluated the IR Biotyper, a device that uses infrared light to analyze bacteria and group them by capsule type-a key feature linked to their spread. The IR Biotyper matched WGS results with high accuracy, delivering results in minutes instead of days. This fast, affordable method can help hospitals detect outbreaks earlier and respond more effectively. Our findings suggest that the IR Biotyper is a valuable tool for routine use in microbiology laboratories, supporting epidemiological surveillance and outbreak control.

Klebsiella pneumoniae↗

Application of multilevel models to morphometric data. Part 2. Correlations.

Multilevel organization of morphometric data (cells are "nested" within patients) requires special methods for studying correlations between karyometric features. The most distinct feature of these methods is that separate correlation (covariance) matrices are produced for every level in the hierarchy. In karyometric research, the cell-level (i.e., within-tumor) correlations seem to be of major interest. Beside their biological importance, these correlation coefficients (CC) are compulsory when dimensionality reduction is required. Using MLwiN, a dedicated program for multilevel modeling, we show how to use multivariate multilevel models (MMM) to obtain and interpret CC in each of the levels. A comparison with two usual, "single-level" statistics shows that MMM represent the only way to obtain correct cell-level correlation coefficients. The summary statistics method (take average values across each patient) produces patient-level CC only, and the "pooling" method (merge all cells together and ignore patients as units of analysis) yields incorrect CC at all. We conclude that multilevel modeling is an indispensable tool for studying correlations between morphometric variables.

Algorithms↗

The ubiquitous nature of epistasis in determining susceptibility to common human diseases.

There is increasing awareness that epistasis or gene-gene interaction plays a role in susceptibility to common human diseases. In this paper, we formulate a working hypothesis that epistasis is a ubiquitous component of the genetic architecture of common human diseases and that complex interactions are more important than the independent main effects of any one susceptibility gene. This working hypothesis is based on several bodies of evidence. First, the idea that epistasis is important is not new. In fact, the recognition that deviations from Mendelian ratios are due to interactions between genes has been around for nearly 100 years. Second, the ubiquity of biomolecular interactions in gene regulation and biochemical and metabolic systems suggest that relationship between DNA sequence variations and clinical endpoints is likely to involve gene-gene interactions. Third, positive results from studies of single polymorphisms typically do not replicate across independent samples. This is true for both linkage and association studies. Fourth, gene-gene interactions are commonly found when properly investigated. We review each of these points and then review an analytical strategy called multifactor dimensionality reduction for detecting epistasis. We end with ideas of how hypotheses about biological epistasis can be generated from statistical evidence using biochemical systems models. If this working hypothesis is true, it suggests that we need a research strategy for identifying common disease susceptibility genes that embraces, rather than ignores, the complexity of the genotype to phenotype relationship.

Disease↗

Multilocus analysis of hypertension: a hierarchical approach.

While hypertension is a complex disease with a well-documented genetic component, genetic studies often fail to replicate findings. One possibility for such inconsistency is that the underlying genetics of hypertension is not based on single genes of major effect, but on interactions among genes. To test this hypothesis, we studied both single locus and multilocus effects, using a case-control design of subjects from Ghana. Thirteen polymorphisms in eight candidate genes were studied. Each candidate gene has been shown to play a physiological role in blood pressure regulation and affects one of four pathways that modulate blood pressure: vasoconstriction (angiotensinogen, angiotensin converting enzyme - ACE, angiotensin II receptor), nitric oxide (NO) dependent and NO independent vasodilation pathways and sodium balance (G protein-coupled receptor kinase, GRK4). We evaluated single site allelic and genotypic associations, multilocus genotype equilibrium and multilocus genotype associations, using multifactor dimensionality reduction (MDR). For MDR, we performed systematic reanalysis of the data to address the role of various physiological pathways. We found no significant single site associations, but the hypertensive class deviated significantly from genotype equilibrium in more than 25% of all multilocus comparisons (2,162 of 8,178), whereas the normotensive class rarely did (11 of 8,178). The MDR analysis identified a two-locus model including ACE and GRK4 that successfully predicted blood pressure phenotype 70.5% of the time. Thus, our data indicate epistatic interactions play a major role in hypertension susceptibility. Our data also support a model where multiple pathways need to be affected in order to predispose to hypertension.

Alleles↗

Renin-angiotensin system gene polymorphisms and atrial fibrillation.

BACKGROUND: The activated local atrial renin-angiotensin system (RAS) has been reported to play an important role in the pathogenesis of atrial fibrillation (AF). We hypothesized that RAS genes might be among the susceptibility genes of nonfamilial structural AF and conducted a genetic case-control study to demonstrate this. METHODS AND RESULTS: A total of 250 patients with documented nonfamilial structural AF and 250 controls were selected. The controls were matched to cases on a 1-to-1 basis with regard to age, gender, presence of left ventricular dysfunction, and presence of significant valvular heart disease. The ACE gene insertion/deletion polymorphism, the T174M, M235T, G-6A, A-20C, G-152A, and G-217A polymorphisms of the angiotensinogen gene, and the A1166C polymorphism of the angiotensin II type I receptor gene were genotyped. In multilocus haplotype analysis, the angiotensinogen gene haplotype profile was significantly different between cases and controls (chi2=62.5, P=0.0002). In single-locus analysis, M235T, G-6A, and G-217A were significantly associated with AF. Frequencies of the M235, G-6, and G-217 alleles were significantly higher in cases than in controls (P=0.000, 0.005, and 0.002, respectively). The odds ratios for AF were 2.5 (95% CI 1.7 to 3.3) with M235/M235 plus M235/T235 genotype, 3.3 (95% CI 1.3 to 10.0) with G-6/G-6 genotype, and 2.0 (95% CI 1.3 to 2.5) with G-217/G-217 genotype. Furthermore, significant gene-gene interactions were detected by the multifactor-dimensionality reduction method and multilocus linkage disequilibrium tests. CONCLUSIONS: This study demonstrates the association of RAS gene polymorphisms with nonfamilial structural AF and may provide the rationale for clinical trials to investigate the use of ACE inhibitor or angiotensin II antagonist in the treatment of structural AF.

Aged↗

What causes a neuron to spike?

The computation performed by a neuron can be formulated as a combination of dimensional reduction in stimulus space and the nonlinearity inherent in a spiking output. White noise stimulus and reverse correlation (the spike-triggered average and spike-triggered covariance) are often used in experimental neuroscience to "ask" neurons which dimensions in stimulus space they are sensitive to and to characterize the nonlinearity of the response. In this article, we apply reverse correlation to the simplest model neuron with temporal dynamics-the leaky integrate-and-fire model-and find that for even this simple case, standard techniques do not recover the known neural computation. To overcome this, we develop novel reverse-correlation techniques by selectively analyzing only "isolated" spikes and taking explicit account of the extended silences that precede these isolated spikes. We discuss the implications of our methods to the characterization of neural adaptation. Although these methods are developed in the context of the leaky integrate-and-fire model, our findings are relevant for the analysis of spike trains from real neurons.

Action Potentials↗