PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Principal Component Analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Generalized principal component analysis (GPCA).

This paper presents an algebro-geometric solution to the problem of segmenting an unknown number of subspaces of unknown and varying dimensions from sample data points. We represent the subspaces with a set of homogeneous polynomials whose degree is the number of subspaces and whose derivatives at a data point give normal vectors to the subspace passing through the point. When the number of subspaces is known, we show that these polynomials can be estimated linearly from data; hence, subspace segmentation is reduced to classifying one point per subspace. We select these points optimally from the data set by minimizing certain distance function, thus dealing automatically with moderate noise in the data. A basis for the complement of each subspace is then recovered by applying standard PCA to the collection of derivatives (normal vectors). Extensions of GPCA that deal with data in a high-dimensional space and with an unknown number of subspaces are also presented. Our experiments on low-dimensional data show that GPCA outperforms existing algebraic algorithms based on polynomial factorization and provides a good initialization to iterative techniques such as K-subspaces and Expectation Maximization. We also present applications of GPCA to computer vision problems such as face clustering, temporal video segmentation, and 3D motion segmentation from point correspondences in multiple affine views.

Algorithms↗

Differential diagnosis in dementia. Principal components analysis of clinical data from a population survey.

OBJECTIVE: To reduce all the clinical data, collected from an unselected group of subjects, to a small set of factors and to see how these factors correspond to standard clinical diagnosis of dementing disorders. DESIGN: Population survey. SETTING: General community: elderly older than 74 years, from an area in Stockholm, Sweden. SUBJECTS: Population-based sample including (1) all the screened positive subjects using the Mini-Mental State examination; and (2) a random sample of the screened negative subjects, matched by age and sex. A clinical examination and an informant interview were carried out. Cases were identified using Diagnostic and Statistical Manual of Mental Disorders, Revised Third Edition diagnostic criteria for Alzheimer's disease (AD) and other dementias. MAIN OUTCOME MEASURE: Independently from the clinical diagnosis, a principal components factor analysis was carried out to investigate groupings among the clinical data (factors). Factor scores, calculated as a weighted sum of the symptom variables and converted to a standard score form. RESULTS: Four major factors were found: cognitive impairment, cerebrovascular disease, disturbed behavior, and depressive symptoms. The comparison of these factors with the clinical diagnoses showed that (1) the cognitive impairment factor discriminated demented cases from nondemented; (2) the cerebrovascular disease factor discriminated vascular dementia from AD cases and nondemented; (3) the disturbed behavior factor discriminated AD cases from vascular dementia cases and nondemented, indicating behavioral changes characteristic of AD. CONCLUSIONS: This finding, if replicated, would have implications for the construction of diagnostic criteria for AD.

Aged↗

Experimental comparison of data transformation procedures for analysis of principal components.

Results of principal component analysis depend on data scaling. Recently, based on theoretical considerations, several data transformation procedures have been suggested in order to improve the performance of principal component analysis of image data with respect to the optimum separation of signal and noise. The aim of this study was to test some of those suggestions, and to compare several procedures for data transformation in analysis of principal components experimentally. The experiment was performed with simulated data and the performance of individual procedures was compared using the non-parametric Friedman's test. The optimum scaling found was that which unifies the variance of noise in the observed images. In data with a Poisson distribution, the optimum scaling was the norm used in correspondence analysis. Scaling mainly affected the definition of the signal space. Once the dimension of the signal space was known, the differences in error of data and signal reproduction were small. The choice of data transformation depends on the amount of available prior knowledge (level of noise in individual images, number of components, etc), on the type of noise distribution (Gaussian, uniform, Poisson, other), and on the purpose of analysis (data compression, filtration, feature extraction).

Computer Simulation↗

Application of principal component analysis for the estimation of source of heavy metal contamination in surface sediments from the Rybnik Reservoir.

The concentrations of metals, loss of ignition and nutrient (N, P) were determined in the bottom sediments of the Rybnik Reservoir (southern Poland). The mean concentrations of the metals in the bottom sediments were: Cd 25.8 microgram/g, Cu 451.7 microgram/g, Zn 1583.4 microgram/g, Ni 71.1 microgram/g, Pb 118.6 microgram/g, Cr 129.8 microgram/g, Fe 38782 microgram/g and Mn 2018.7 microgram/g. The bottom sediments are very heavily loaded with zinc, manganese, copper, nickel, phosphorus and lead (percentage enrichment factor), and cadmium, phosphorus and zinc (index of geoaccumulation). The increase of cadmium, lead, nickel and zinc concentrations was connected with the inflow of the contaminated water of the river Ruda and long-range transport. The contamination of the reservoir with copper and manganese resulted mainly from atmospheric precipitation. The variability of the bottom sediment loading with metals during the investigations was affected in the first place by changes in the concentration of iron, but also those elements whose concentrations in the bottom sediment were elevated compared to the concentrations in shale--cadmium, nickel and lead.

Environmental Monitoring↗

Oriented principal component analysis for large margin classifiers.

Large margin classifiers (such as MLPs) are designed to assign training samples with high confidence (or margin) to one of the classes. Recent theoretical results of these systems show why the use of regularisation terms and feature extractor techniques can enhance their generalisation properties. Since the optimal subset of features selected depends on the classification problem, but also on the particular classifier with which they are used, global learning algorithms for large margin classifiers that use feature extractor techniques are desired. A direct approach is to optimise a cost function based on the margin error, which also incorporates regularisation terms for controlling capacity. These terms must penalise a classifier with the largest margin for the problem at hand. Our work shows that the inclusion of a PCA term can be employed for this purpose. Since PCA only achieves an optimal discriminatory projection for some particular distribution of data, the margin of the classifier can then be effectively controlled. We also propose a simple constrained search for the global algorithm in which the feature extractor and the classifier are trained separately. This allows a degree of flexibility for including heuristics that can enhance the search and the performance of the computed solution. Experimental results demonstrate the potential of the proposed method.

Algorithms↗

Multivariate analysis of microarray data by principal component discriminant analysis: prioritizing relevant transcripts linked to the degradation of different carbohydrates in Pseudomonas putida S12.

The value of the multivariate data analysis tools principal component analysis (PCA) and principal component discriminant analysis (PCDA) for prioritizing leads generated by microarrays was evaluated. To this end, Pseudomonas putida S12 was grown in independent triplicate fermentations on four different carbon sources, i.e. fructose, glucose, gluconate and succinate. RNA isolated from these samples was analysed in duplicate on an anonymous clone-based array to avoid bias during data analysis. The relevant transcripts were identified by analysing the loadings of the principal components (PC) and discriminants (D) in PCA and PCDA, respectively. Even more specifically, the relevant transcripts for a specific phenotype could also be ranked from the loadings under an angle (biplot) obtained after PCDA analysis. The leads identified in this way were compared with those identified using the commonly applied fold-difference and hierarchical clustering approaches. The different data analysis methods gave different results. The methods used were complementary and together resulted in a comprehensive picture of the processes important for the different carbon sources studied. For the more subtle, regulatory processes in a cell, the PCDA approach seemed to be the most effective. Except for glucose and gluconate dehydrogenase, all genes involved in the degradation of glucose, gluconate and fructose were identified. Moreover, the transcriptomics approach resulted in potential new insights into the physiology of the degradation of these carbon sources. Indications of iron limitation were observed with cells grown on glucose, gluconate or succinate but not with fructose-grown cells. Moreover, several cytochrome- or quinone-associated genes seemed to be specifically up- or downregulated, indicating that the composition of the electron-transport chain in P. putida S12 might change significantly in fructose-grown cells compared to glucose-, gluconate- or succinate-grown cells.

Carbohydrate Metabolism↗

Principal component analysis of various respiratory function tests: the relationship between factor score and severity of pulmonary circulatory disorder in chronic obstructive pulmonary disease.

The relationship between pulmonary haemodynamics and values of various respiratory function tests was studied in patients with mild chronic obstructive pulmonary disease (COPD) and the following results were obtained. (1) The value of mean pulmonary artery pressure (mPAP) at rest in COPD patients was slightly elevated to 19.4 mmHg on average compared with our control value of less than 18 mmHg. (2) Analysis of the data of 11 routine respiratory function tests in 88 COPD patients extracted two principal components: an index of the expiratory function and an index for overinflation of the lung. (3) In individual patients, mPAP expressed the severity of pulmonary circulatory disorder roughly inverse to the factor score of the first principal component (index of expiratory function) but not to that of the second principal component (overinflation of the lung). (4) Discriminant analysis was performed in all 88 COPD patients according to data from the 11 respiratory function tests. The probability of mPAP being above or below 18 mmHg was 18.2%. (5) The relationship between the predicted EPOI value and the factor score was similar to that between mPAP and the factor score. EPOI (exercise pulmonary artery pressure-oxygen consumption index) was calculated with the following equation: EPOI = (mPAPex#-mPAPrest)/[VO2ex-VO2rest)/BSA##). On the other hand, EPOIpred was calculated with the prediction equation obtained from multiple linear regression (dependent variable; EPOI, independent variable; respiratory function).

Adult↗

Geometrical principal component analysis of planar-segments of the three-channel Lissajous' trajectory of human auditory brain stem evoked potentials.

Three-Channel Lissajous' Trajectories (3CLTs) of Auditory Brain Stem Evoked Potentials (ABEP) were obtained from 15 normal humans. Planar-segments of 3CLT were identified and the orientations of the first two geometrical principal components, which interact to produce the planar-segments, were calculated. Each principal component's orientation in voltage space was quantified by its coefficients (A, B and C). Intersubject variability of these orientations was comparable to the variability of plane orientations. The principal components of planar-segments can indicate the type of generator activity that is involved in the formation of planar-segments. The results of this analysis indicate that planarity of each 3CLT component is produced by the interaction of simultaneous multiple generators, or by a single synchronous generator which changes its orientation. The coefficients of these principal components may complement plane coefficients as quantitative indices of 3CLT of ABEP.

Adult↗

Comparison of male adolescent-report of attention-deficit/hyperactivity disorder (ADHD) symptoms across two cultures using latent class and principal components analysis.

BACKGROUND: The goal of this study is to gauge the consistency of Attention Deficit/Hyperactivity Disorder (ADHD) latent class models that are generated by different informants such as adolescents and parents. The consistency of adolescent-derived latent classes from two different samples was assessed and these results were then compared to the class structure generated by parent-report ADHD information. METHODS: Self-reported DSM-IV Criterion A ADHD symptoms of 497 adolescent males from a population-based twin study in the state of Missouri (USA) were subjected to principal components and latent class analysis, and findings were compared to previous results obtained from identical analyses using an adolescent sample from Porto Alegre, Brazil (N = 483). RESULTS: The bi-dimensional structure of self-reported ADHD symptoms was similar for both male adolescent groups, but explained less than 40% of the symptom variance in either sample. Two factors, one with loadings on inattention symptoms only and the other with loadings on hyperactive-impulsive symptoms only, were identified in the Missouri sample. Specific ADHD latent classes did not replicate well across the Missouri and Brazilian samples, and both groups were characterized by the presence of several combined symptom classes but few inattentive or hyperactive-impulsive classes. CONCLUSIONS: While adolescent-report information across two different cultures can at least in part reproduce the two-factor structure of ADHD, results from latent class analysis suggest that adolescent reporting on their own symptoms is markedly different from the type of information parents provide about ADHD symptoms in their offspring. The current findings indicate that if male adolescents endorse any ADHD symptoms there is a tendency for them to report combined type problems.

Adolescent↗

Principal component analysis of trace elements in Serbian wheat.

Trace elements (Cu, Fe, Pb, Hg, Cd, As, Mn, Zn) were analyzed quantitatively in 14 wheat samples collected from fields in all Serbian growing regions, harvested in 2002. Microelements were determined according to an atomic absorption spectrophotometric method. Principal component analyses (PCA) were performed on data matrices consisting of contents of trace elements in wheats (columns) and all Serbian wheat-growing regions (rows). It was found that four principal components account for 87.2% of the total variance in the data. The plot of component loadings showed significant groupings for concentration of some microelements. The component scores indicated the similarities among the Serbian wheat-growing regions. The loading plot reveals that there is no need to measure all of the variables to achieve the same classification. It is enough to measure one variable per group. Naturally, this conclusion is valid only within the limits of the present study of wheat grain samples from different parts of Serbia.

Seeds↗

EEG bands during wakefulness, slow-wave, and paradoxical sleep as a result of principal component analysis in the rat.

Rat EEG has been empirically divided in bands that frequently do not correspond with EEG generators nor with the functional meaning of EEG rhythms. Power spectra from wakefulness (W), slow-wave sleep (SWS), and paradoxical sleep (PS) of Wistar rats were submitted to Principal Component Analyses (PCA) to investigate which frequencies are covariant. Three independent eigenvectors were identified for SWS: a band between 1-6, an intermediate band between 7-15, and a fast band between 16-32 Hz (90.74% of the variance); two independent eigenvectors were extracted for PS: slow frequencies between 1-6 covarying together with frequencies between 11-16 Hz, and activity between 6-10 covarying together with fast frequencies between 17-32 Hz (80.38% of the variance); four eigen-vectors were obtained for W: 3-7, 8-9, 10-21 and 21-32 Hz (81.47% of the variance). Vigilance states showed significant differences in AP from 1 to 22 Hz. PCA extracted broad bands different for each vigilance state, which included the most representative EEG activities characteristic of them. These results indicate that during SWS, slow oscillations include frequencies up to 6 Hz, and spindle oscillations frequencies down to 7 Hz. No alpha frequencies were identified as an independent band. Frequencies within theta and beta were gathered in the same eigenvector during PS and in different eigenvectors during W suggesting coordinated activation of hippocampal and cortical systems during PS. These bands are consistent with the underlying neurophysiological mechanisms of sleep and wakefulness and with firing frequencies of generators of rhythmic activity obtained in cellular studies in animals.

Animals↗

Principal component analysis of biogenic amines and polyphenols in Hungarian wines.

Biogenic amines, polyphenols, and resveratrol were analyzed quantitatively in 25 different Hungarian wines from the same wine-making region, harvest of 1998. Polyphenols were determined according to a spectrophotometric method, whereas other substrates were analyzed using overpressured-layer chromatography (OPLC). Principal component analyses (PCA) were performed on data matrices consisting of substrates (columns) and different sorts of wines (rows) from the region of Pécs (southern Hungary). It was found that four (unrotated) principal components account for >80% of the total variance in the data. The plots of component loadings showed significant groupings for concentrations of biogenic amines (and polyphenols). Similarly, the component scores grouped according to the different sorts of wines. The loading plots reveal that there is no need to measure all of the variables to achieve the same characterization. It is enough to measure one variable per group. Naturally, this conclusion is valid only within the limits of the present study; wines from other regions may behave differently.

Biogenic Amines↗

Positive and negative symptoms in the psychoses: principal components analysis of items from the Scale for the Assessment of Positive Symptoms and the Scale for the Assessment of Negative Symptoms.

The present study investigated the factor structure of the items contained in Andreasen's scales for the assessment of positive and negative symptoms (SAPS and SANS) by use of a series of principal components analyses (PCAs) with oblique rotations of the axes. It was found that the structure could be summarized by three major components labeled negative symptoms, thought disorder, and delusions/hallucinations. Dimensionality could meaningfully be increased to five components. Negative symptoms was found to separate into two components that we labeled negative signs and social dysfunctions. The delusions/hallucinations factor could be separated into two components, delusions and hallucinations, with "loss of boundary" delusions being related to both factors. Delusions of persecution were independent of other symptoms. The thought disorder factor did not decompose meaningfully within the investigated dimensionality. A two-factor solution did not explain the correlation between symptoms adequately. The results do not support the simple dichotomy between positive and negative symptoms in psychosis, but suggest that a wider dimensional concept may be more useful in future studies.

Delusions↗

Action potential classifiers: a functional comparison of template matching, principal components analysis and an artificial neural network.

Multiunit neural activity occurs often in electrophysiological studies when utilizing extracellular electrodes. In order to estimate the activity of the individual neurons each action potential in the recording must be classified to its neuron of origin. This paper compares the accuracy of two traditional methods of action potential classification--template matching and principal components--against the performance of an artificial neural network (ANN). Both traditional methods use averages of action potential shapes to form their corresponding classifiers while the artificial neural network 'learns' a nonlinear relationship between a set of prototype action potentials and assigned classes. The set of prototypic action potentials and the assigned classes is termed the training set. The training set contained action potentials from each class which exhibited the full range of amplitude variability. The ANN provided better classification results and was more robust in analysis of across-animal data sets than either of the traditional action potential classification methods.

Action Potentials↗

Principal components analysis of physical growth in savannah baboons.

Morphometric data collected from 118 male and 169 female savannah baboons (Papio cynocephalus anubis) aged between birth and 5.5 years were analyzed to describe the morphology and physical growth of this species. Measurements included weight, crown-rump length, triceps circumference, and skinfolds at the neck, subscapular, suprailiac, and triceps anatomical sites. Principal components analyses were applied to the data to provide multivariate assessments of morphological patterning among the variables. These analyses resulted in the extraction of two unrotated orthogonal components that accounted for 88% of the overall sample variation. The first component accounted for 77% of the variation and represents an axis of overall body size. The second component represents an axis of shape variation that contrasts body size with fat patterning, and was interpreted as a measure of body leanness. Individual component scores were computed for determining age, gender, and age-by-gender interaction effects. Both components were found to be age dependent for both genders. Males and females shared similar age patterning along the two components; however, gender differences did occur in patterning along the two components; however, gender differences did occur in respect to leanness. The multivariate measure of overall body size increased for both genders similarly with advancing age. Age patterning along the leanness component was described as a decrease from birth to 1 year, followed by an increase in leanness in older ages. Females had a delayed and significantly less intense increase in leanness relative to males.

Aging↗

Principal component analysis and large-scale correlations in non-coding sequences of human DNA.

We have calculated a full set of second-order correlation functions of nucleotides in noncoding DNA. They are found to be independently invariant in regard to permutations of A and T, and also C and G. Considering correlation functions as a 4 x 4 matrix with a symmetrical basis, we have found the principal components-objects with zero cross-correlations. These three principal components are present the base compositions: (A + T - C - G), (A - T), (C - G). The long-range behavior of these principal components yields power-law dependencies with different critical exponents.

Base Composition↗

Molecular diversity sample generation on the basis of quantum-mechanical computations and principal component analysis.

The present study introduces a new strategy of selection of a maximum diversity sample of n compounds from N available in a molecular database. This strategy can be useful in pharmacological screening, combinatorial chemistry or parallel synthesis planning. It consists of first describing the compounds by means of parameters derived from quantum mechanical computations (water solvation deltaG, benzene solvation deltaG, octanol solvation deltaG, dipolar moment), as well as standard molecular parameters such as solvent-accessible surface area and molecular weight. Solvation parameters are used because of the importance of this phenomenon in the pharmacological behaviour. Redundant information in the description of the compounds is eliminated by using principal components (PC) instead of the original descriptors. Based on the similarity between the N compounds in the PC space, they are classified into n groups by k-means cluster analysis. The compounds that are nearest to the centroid of each cluster constituted the maximum diversity sample. When practical difficulties exist for the use of one of the proposed compounds, another also close to the cluster centroid can substitute for it. This strategy has been tested in the selection of a sample of 50 amines from the 923 available in the Aldrich catalogue. The results have been contrasted with those obtained from an optimal, distance-based experimental design, resulting in an 86% of agreement between both approaches. An R(2)-like diversity coefficient has been used to assess the quality of the proposed solutions.

Amines↗

Multipoint dissolution specification and acceptance sampling rule based on profile modeling and principal component analysis.

In dissolution testing, multiple dissolution measurements at specific time points are needed in quality control when the compliance of the product requires controlled dissolution throughout the time course. The dissolution specification based on general multivariate confidence region was proposed by Chen and Tsong (8). This paper presents two alternative procedures when the dissolution profile consists of important measurements at more than 4 time points. In the first procedure, when the dissolution profile can be described by a physical curve through modeling, the dissolution specification is developed based on the confidence region of the parameters of the physical curve. In the second procedure, the principal components (PCS) as the linear combinations of the dissolution measurements are identified and dissolution specification is set based by the confidence intervals of the values of principal components. In both approaches the specification can be set at lower dimensions than the general multivariate confidence region approach. A single-stage acceptance rule can be used in both approaches by first projecting the dissolution values of each tablet in the new testing batch onto the determined parameters axes (through modeling in modeling approach and through projection on the selected PCS in principal component approach). Then check if the projections of the new tablet fall within the specifications. Finally, count the number of tablets that fall outside the specification limits and reject the batch if the proportion of out-of-specification tablet is high and accept the lot for release if the proportion is low.

Chemical Phenomena↗