PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Principal Component Analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

The classification of solvents by combining classical QSPR methodology with principal component analysis.

The results of a quantitative structure-property relationship (QSPR) analysis of 127 different solvent scales and 774 solvents using the CODESSA PRO program are presented. QSPR models for each scale were constructed using only theoretical descriptors. The high quality of the models is reflected by the squared multiple correlation coefficients that range from 0.726 to 0.999; only 18 models have R2< 0.800. This enables direct theoretical calculation of predicted values for any scale and/or for any organic solvent, including those previously unmeasured. The molecular descriptors involved in the models are classified and discussed according to (i) the origin of their calculation (i.e., constitutional, geometric, charge-related, etc.) and (ii) the commonly accepted classification of physical interactions between the solute and solvent molecules in liquid (condensed) media. A reduced matrix 774 (solvents) x 100 (solvent scales) was selected for the principal component analysis (PCA) by taking into account only the solvent scales with more than 20 experimental data points. The first 5 principal components account for 75% of the total variance. The robustness of the PCA model obtained was validated by the comparison models development for restricted submatrices of data and with the results obtained for the full data set. The total variance accounted for by the first three PCs, for the submatrices with the same number of solvent scales but different numbers of solvents, varies from 68.2% to 59.0%. This demonstrates that the total variance described by the first 3 components is essentially stable as the number of solvents involved varies from 100 to 774. Subsequently, a matrix with 703 diverse solvents and 100 solvent scales was selected for the general classification of the solvents and scales according to the scores and loadings obtained from the PCA treatment. Classification of the theoretical molecular descriptors, derived from the chemical structure alone, according to their relevance to specific types of intermolecular interaction (cavity formation, electrostatic polarization, dispersion, and hydrogen bonding) in liquid media enables a more easily comprehensible physical interpretation of the QSPR of molecular properties in liquids and solutions. The reported QSPR models for solvent scales with theoretical molecular descriptors and the results of the PCA analysis are potentially of great practical importance, as they extend the applicability of correlations with empirical solvent scales to many previously unmeasured systems.

Journal Article↗

Characterisation of acute myocardial ischaemia in a canine model based on principal component analysis of unipolar endocardial electrograms.

The study presents a method for identifying endocardial electrical features relevant to local ischaemia detection at rest. The method consists of, first, normalisation of electrograms to a uniform representation; secondly, the use of principal component analysis to reduce the dimensionality of the electrogram vector space; and, thirdly, a search for a classification axis that matches the degree of ischaemia present in the tissue. Left ventricular myocardial states were assessed by echocardiography and NOGA mapping in eight dogs at baseline and then immediately after, 5h after and 3 days after occlusion of the left anterior descending coronary artery. Five principal components were required to approximate electrograms with an average error of less than 10% of the peak-to-peak amplitude. Correlations of 0.77, 0.80 and 0.84 were obtained between the principal component-based parameters and the echocardiography scores at the three ischaemic stages, respectively. Expression of these parameters in the time domain showed that the major changes occurred in the depolarisation segment of the endocardial electrogram as well as in the ST-segment. In conclusion, the proposed method provides a suitable alternative co-ordinate system for the classification of ischaemic regions and highlights signal segments that change as a result of pathology.

Acute Disease↗

Finding haplotype tagging SNPs by use of principal components analysis.

The immense volume and rapid growth of human genomic data, especially single nucleotide polymorphisms (SNPs), present special challenges for both biomedical researchers and automatic algorithms. One such challenge is to select an optimal subset of SNPs, commonly referred as "haplotype tagging SNPs" (htSNPs), to capture most of the haplotype diversity of each haplotype block or gene-specific region. This information-reduction process facilitates cost-effective genotyping and, subsequently, genotype-phenotype association studies. It also has implications for assessing the risk of identifying research subjects on the basis of SNP information deposited in public domain databases. We have investigated methods for selecting htSNPs by use of principal components analysis (PCA). These methods first identify eigenSNPs and then map them to actual SNPs. We evaluated two mapping strategies, greedy discard and varimax rotation, by assessing the ability of the selected htSNPs to reconstruct genotypes of non-htSNPs. We also compared these methods with two other htSNP finders, one of which is PCA based. We applied these methods to three experimental data sets and found that the PCA-based methods tend to select the smallest set of htSNPs to achieve a 90% reconstruction precision.

Chromosome Mapping↗

Principal component analysis for selection of optimal SNP-sets that capture intragenic genetic variation.

Candidate gene association studies often utilize one single nucleotide polymorphism (SNP) for analysis, with an initial report typically not being replicated by subsequent studies. The failure to replicate may result from incomplete or poor identification of disease-related variants or haplotypes, possibly due to naive SNP selection. A method for identification of linkage disequilibrium (LD) groups and selection of SNPs that capture sufficient intra-genic genetic diversity is described. We assume all SNPs with minor allele frequency above a pre-determined frequency have been identified. Principal component analysis (PCA) is applied to evaluate multivariate SNP correlations to infer groups of SNPs in LD (LD-groups) and to establish an optimal set of group-tagging SNPs (gtSNPs) that provide the most comprehensive coverage of intra-genic diversity while minimizing the resources necessary to perform an informative association analysis. This PCA method differs from haplotype block (HB) and haplotype-tagging SNP (htSNP) methods, in that an LD-group of SNPs need not be a contiguous DNA fragment. Results of the PCA method compared well with existing htSNP methods while also providing advantages over those methods, including an indication of the optimal number of SNPs needed. Further, evaluation of the method over multiple replicates of simulated data indicated PCA to be a robust method for SNP selection. Our findings suggest that PCA may be a powerful tool for establishing an optimal SNP set that maximizes the amount of genetic variation captured for a candidate gene using a minimal number of SNPs.

Adult↗

Localization of the event-related potential novelty response as defined by principal components analysis.

Recent research indicates that novel stimuli elicit at least two distinct components, the Novelty P3 and the P300. The P300 is thought to be elicited when a context updating mechanism is activated by a wide class of deviant events. The functional significance of the Novelty P3 is uncertain. Identification of the generator sources of the two components could provide additional information about their functional significance. Previous localization efforts have yielded conflicting results. The present report demonstrates that the use of principal components analysis (PCA) results in better convergence with knowledge about functional neuroanatomy than did previous localization efforts. The results are also more convincing than that obtained by two alternative methods, MUSIC-RAP and the Minimum Norm. Source modeling on 129-channel data with BESA and BrainVoyager suggests the P300 has sources in the temporal-parietal junction whereas the Novelty P3 has sources in the anterior cingulate.

Acoustic Stimulation↗

Improving image contrast using principal component analysis for subsequent image segmentation.

This article presents a technique for improving MR image contrast by linearly combining multiple MR images with different tissue contrast. The weighting coefficients of the linear combination are derived using principal component analysis. The contrast-enhanced composite image is segmented subsequently using gray level-based 1D segmentation methods. The technique reduces a multispectral image set to composite eigenimages and allows application of appropriate 1D segmentation methods that do not have equivalent counterparts in multispectral methods.

Brain↗

Relationships between induction of anesthesia and mitotic spindle disturbances studied by means of principal component analysis.

A dataset comprising the activity of 30 compounds in 4 biological tests--anesthesia of tadpoles, anesthesia of frog heart, abnormal growth and spindle disturbances in Allium root tips--was re-evaluated by means of principal component analysis. A two-component model is required to explain the variation in biological activity of the compounds. It is found that abnormal growth is different from the other biological responses. When this test is excluded, as much as 90% of the variation is explained by a one-component model, the determining factor most probably being the lipophilic character of the compounds. Mammalian mitotic cells respond in a similar way to mitotic cells of Allium root tips. It is suggested that possible regularities in the dose-response relationships for anesthesia, teratogenic effects and generation of abnormal chromosome numbers require further exploration.

Anesthetics↗

Functionalization of hydrocolloids: principal component analysis applied to the study of correlations between parameters describing the consistency of hydrogels.

This work was part of a pure research project on the functionalization of three families of hydrocolloids: cellulose derivatives, carrageenates, and alginates. Principal component analysis (PCA), a powerful statistical method, was used to demonstrate the relations existing among these different parameters that describe the consistency of hydrogels and their spreadability. This approach therefore provides a basis for modeling hydrogel consistency. PCA also afforded a classification of hydrogels that demonstrated the remarkable adhesiveness of very stiff gels based on cellulose derivatives and sodium or potassium alginates. The corresponding semi-fluid gels and all the gels based on carrageenates and mixed sodium-calcium alginates, whatever their spreadability, were found to be very poorly adhesive. Generalized to all the many colloids currently marketed, this approach can be used to set up a databank for the formulation of mucoadhesive excipients.

Adhesiveness↗

Principal component analysis of proton nuclear magnetic resonance spectra of lipoprotein fractions from patients with coronary heart disease and healthy subjects.

Blood plasma was drawn from 12 healthy subjects and 12 patients with coronary heart disease (CHD). The lipoproteins were fractionated by serial ultracentrifugation. The methyl and methylene regions of proton nuclear magnetic resonance (1 H NMR) spectra of the lipoproteins very low density lipoprotein (VLDL), low density lipoprotein (LDL) and high density lipoprotein (HDL) were analysed by principal component analysis (PCA). Grouping patterns in the score plots and the profiles of the principal components revealed several characteristics of the spectra. LDL subparticle size among the CHD group was consistently skewed against smaller, denser subparticles. This feature was independent of the concentration of LDL cholesterol. Analysis of the LDL spectra by soft independent modelling of class analogy (SIMCA) showed that none of the samples from the CHD group could be assigned as healthy subjects (p < 0.05). We also found that the samples from the healthy subjects were associated with a higher concentration of HDL cholesterol and larger VLDL subparticles. The approach presented, in which PCA is used in combination with NMR spectroscopy, might be implemented in clinical studies to give information about lipoprotein subparticle distribution and lipid content.

Adult↗

Principal components analysis for the visualisation of multidimensional chemical data acquired by scanning Raman microspectroscopy.

Raman microspectroscopy is ideally suited to surface analysis as it allows detailed chemical information to be acquired from surfaces at a relatively high spatial resolution (typically 1 microm). Using a motorised sample table or probe, it is possible to raster scan a surface to obtain spatially resolved chemical information. Visualisation of the acquired data is a problem, however, as the spectrum acquired at each point can contain several hundred individual intensity measurements. Existing visualisation methods are limited to plotting each scanned point with an intensity determined from the measured intensity at a single wavenumber, or the similarly between the point's spectrum and a reference spectrum. Such methods are wasteful as a lot of acquired information is discarded, and results are prone to misinterpretation due to background variance and instrumental noise. In this paper we introduce a new method that uses principal components analysis (PCA) to reduce the spectrum at each point to three factors that are then used to define the red, green and blue components of the corresponding point on a false colour map. To increase the effective resolution, interpolation is used to approximate the colours corresponding to points between those actually scanned. To demonstrate the technique, the internal surface of a beverage can, contaminated with a 40 microm diameter carbonised oven impurity, consisting mainly of sp2- and sp3-hybridised saturated carbon bonds, has been used as a case study.

Beverages↗

Principal components analysis as an aid to classification of renal dynamic studies.

Four hundred renal dynamic studies obtained using 99mTc-(Sn) DTPA were classified clinically into four classes (normal, pre-renal lesion, intrarenal lesion and urinary tract obstruction). Principal components analysis was then applied to the kidney activity/time curves and yielded good class separation using only three components. The separation of normal and obstructed or damaged kidneys using only the first component was better than that obtained using the calculated mean transit time.

Diagnosis, Differential↗

Principal component analysis of ERP differences related to the meaning of an ambiguous word.

Event-related potentials (ERPs) to the noun and verb meanings of/'led/in the single ambiguous phrase 'it was/'led/' were re-analyzed using principal component analysis (PCA). These data had previously been analyzed by SWDA and reported in this journal. PCA defined 3 meaning-related components, comprising 40.3% of the entire data variance. The N150 component was shown to be larger for the noun meaning than for the verb meaning; the P230 epoch differed in its anterior-posterior distribution according to meaning; and N370 for noun responses was relatively more negative at the right posterior lead and positive at the left anterior. All components taken together, the left anterior lead showed the greatest meaning-related difference. Previous analysis by SWDA had resulted in significant discriminant functions for left hemisphere ERPs, but this analysis did not yield a clear definition of the effects of meaning on specific ERP components or of the scalp distributions of meaning-related components. Thus, while the results of both analyses support the interpretation that the perceived meaning of words has a substantial effect on ERP wave forms, PCA appears to provide the clearest definition of the ERP component effects.

Adult↗

Monte Carlo sampling and principal component analysis of flux distributions yield topological and modular information on metabolic networks.

The work presented here uses Monte Carlo random sampling combined with flux balance analysis and linear programming to analyse the steady-state flux distributions on the surface of the glucose-ammonia phenotypic phase plane of an Escherichia coli system grown on glucose-minimal medium. The distribution of allowable glucose and ammonia uptake rates showed a triangular shape, the apex corresponding to maximum growth rate. The exact shape, e.g. the diagonal boundary is determined by the relative amounts of nutrients required for growth. The logarithm of flux values has a normal distribution, e.g. there is a log normal distribution, and most of the reactions have an order of magnitude between 10(-1) and 1. The increase in the number of blocked reactions as growth switched from aerobic to micro-aerobic phase and the presence of alternate networks for a single optimal solution were both reflections of the variability of pathway utilization for survival and growth. Principal component analysis (PCA) provided us with significant clues on the correlations between individual reactions and correlations between sets of reactions. Furthermore, PCA identified the most influential reactions of the system. The PCA score plots clearly distinguish two different growth phases, micro-aerobic and aerobic. The loading plots for each growth phase showed both the impact of the reactions on the model and the clustering of reactions that are highly correlated. These results have proved that PCA is a promising way to analyse correlations in high-dimensional solution spaces and to detect modular patterns among reactions in a network.

Ammonia↗

Principal component analysis of polarity and interaction parameters in inverse gas chromatography.

Inverse gas chromatography is used in the characterization of aliphatic-aromatic and aromatic ketones, their oximes, and ketone-oxime or oxime-oxime mixtures. All these organic materials are used as liquid stationary phases in gas chromatographic columns. A series of polarity and Flory-Huggins interaction parameters are determined and used to describe the physicochemical properties of examined materials, metal extractants, and products of their degradation. Principal component analysis (PCA) is performed on a data matrix consisting of polarity and interaction parameters for ketones, their oximes, and mixtures. The calculations are carried out on the correlation matrix. It is found that seven principal components account for more than 95% of the total variance in the data, indicating that the polarity (interaction) parameters are not correlating well. Physical meanings are attributed to the principal components, the most influential ones being that the first and the second principal components account for several Flory-Huggins interaction parameters, whereas the fifth is correlated with criterion "A". The plots of component loadings show characteristic groupings of polarity indicators, whereas that of component scores show several groupings of stationary phases. Cluster analysis provides mainly the same groupings. PCA allows for the grouping of polarity and solubility parameters based on the information carried within those parameters. There is no need to use more than one parameter from each cluster. McReynolds polarity and the partial molar excess Gibbs free energy of solution per methylene group carry the same information. The groups of ketones, oximes, and their mixtures can be distinguished with the use of PCA on the basis of the measured polarity, solubility parameters, or both.

Chemical Phenomena↗

Principal component analysis of some oxidative stress parameters and their relationships in hemodialytic and transplanted patients.

BACKGROUND: Oxidative stress profoundly influences the biochemistry of proteins and many other molecules in tissues of uremic patients. In three different groups of uremic patients, the concentrations of the free and bound pentosidine and low-molecular-weight-advanced glycoxydation end products (LMW-AGEs), carbonyls (LMW-C), advanced oxidation protein products (AOPP) and the total antioxidant power of serum were studied in order to determine the relationships between these factors in hemodialytic and transplanted patients. PATIENTS AND METHODS: The above-mentioned parameters were determined in 10 subjects who were currently in hemodialysis (HD) treatment, 10 kidney transplanted patients with chronic renal failure (Tx-CRF), 10 kidney transplanted patients with normal renal function (Tx-N) and 10 healthy subjects (Ctr). The data matrix (40x7) was analyzed using the principal component analysis (PCA). RESULTS: AGEs, carbonyls and AOPP were strongly correlated, while the total antioxidative serum capacity was not related to the other oxidative stress parameters. All the oxidative stress-related parameter values (AGEs, AOPP and LMW-C) in the Tx patients were similar to those of the control group, but were higher in the patients with chronic renal failure. CONCLUSIONS: The correlation between early and advanced oxidative stress markers indicates that reactive oxygen species are involved in a common step in the mechanism of protein modification in all the patient examined. The relationships between carbonyls and AGEs (free, bound pentosidine and LMW-AGEs) support the hypothesis of "carbonyl stress". The common mechanism of the formation of oxidation products in healthy and diseased subject suggests their role of detoxification within kidney function. The total antioxidant power of the serum is not related to the other parameters, which indicates a possible role of molecule interfering.

Antioxidants↗

Inter- and intra-individual variability of ground reaction forces during sit-to-stand with principal component analysis.

Variable reduction is an important issue in biomechanics, because the definition of a non-redundant set of variables necessary for a complete description of a given motor act provides information about the motor strategy. A systematic tool for dealing with variable reduction problems is Principal Component Analysis. In this paper, as an example of an application of this technique, the set of Ground Reaction Forces (GRFs) provided by a six-component force plate, gained during standing up in a heterogeneous population of 82 normal individuals, was reduced to a set of fewer variables. Each subject was required to stand up from a chair five times at different, randomly self selected, speeds, obtaining a data set of 410 trials. Principal Components (PCs) of GRFs were computed for each trial. On average, over the ensemble of trials, first and second PCs (PC1 and PC2) explained together 90% of PCs. Inter- and intra-individual repeatability of the first two PCs was investigated by examining the correlation coefficient between PC waveforms obtained from the whole set of trials and within the set of trials performed by the same subject, respectively. While the PC1 exhibited repeatable patterns, the second one, although repeatable within the group of trials performed by the same subject, displayed marked inter-individual variability. Therefore, PC1 was related to intrinsic aspects of the motor task and PC2 to inter-subject features.

Adult↗

Molecular descriptors for effective classification of biologically active compounds based on principal component analysis identified by a genetic algorithm.

We have evaluated combinations of 111 descriptors that were calculated from two-dimensional representations of molecules to classify 455 compounds belonging to seven biological activity classes using a method based on principal component analysis. The analysis was facilitated by application of a genetic algorithm. Using scoring functions that related the number of compounds in pure classes (i.e., compounds with the same biological activity), singletons, and mixed classes, effective descriptor sets were identified. A combination of only four molecular descriptors accounting for aromatic character, hydrogen bond acceptors, estimated polar van der Waals surface area, and a single structural key gave overall best results. At this performance level, approximately 91% of the compounds occurred in pure classes and mixed classes were absent. The results indicate that combinations of only a few critical descriptors are preferred to partition compounds according to their biological activity, at least in the test cases studied here.

Algorithms↗

Ambient air particulate concentrations and metallic elements principal component analysis at Taichung Harbor (TH) and WuChi Traffic (WT) near Taiwan Strait during 2004-2005.

The purpose of this study is to characterize metallic elements associated with atmospheric particulate matter of total suspended particulate (TSP), fine particle (particle matter with aerodynamical diameter <2.5 microm, PM(2.5)), coarse particle (particle matter with aerodynamical diameter 2.5-10 microm, PM (2.5-10)) at the Taichung Harbor (TH) and WuChi Traffic (WT) sampling site of central Taiwan during March 2004 to February 2005. The result indicated the average total suspended particulate concentration in 1 year was 157.31 and 112.58 microg m(-3) at TH and WT sampling site, respectively. Fine particle (PM(2.5)) size was the dominant species at TH and WT sampling site. In TH sampling site, higher correlation coefficient was observed on total suspended particulates of metallic elements Fe and Zn. And in WT sampling site, higher correlation coefficients displayed on total suspended particulates of metallic elements Fe and Zn, Fe and Mn. Ambient airborne particle principal component analysis of metallic metals was used to identify the possible pollutant sources in this study. At the TH sampling site, 50.81% of the total variance of the data was observed in factor 1. Higher loading of Fe (0.86), Zn (0.79), Pb (0.76), and Mn (0.68) were contributed by traffic emission and the soil source. At the WT sampling site, factor 1 explained 53.74% of the total variance of the data and had high loading for Zn (0.86) and Cu (0.85), which were identified as industrial/traffic emission sources.

Air↗