PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Principal Component Analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Evaluation of the effect of data pre-treatment procedures on classical pattern recognition and principal components analysis: a case study for the geographical classification of tea.

A simple transformation that uses the half-range and central value has been used as a data pre-treatment procedure for principal component analysis (PCA) and pattern recognition techniques. The results obtained have been compared with the results from classical normalisation of data (mean normalisation, maximum normalisation and range normalisation), autoscaling and the minimum-maximum transformation. Three data sets were used in the study. The first was formed by determining 17 elements in 53 tea samples (901 pieces of data). The second and third data sets arose from two long-term drift studies performed to examine instrumental stability at standard and robust conditions. The instruments used were an inductively coupled plasma atomic emission spectrometer and an inductively coupled plasma mass spectrometer. Each drift diagnosis experiment consisted of replicate determinations of a test solution containing 15 analytes at 10 mg l-1 over 8 h without recalibration. Twenty-nine emission lines were determined 99 times, thus, each data set was formed by 2881 pieces of data. Data pre-treatment was applied to the three data sets prior to the use of principal component analysis, cluster analysis, linear discrimination analysis and soft independent modelling of class analogy. The study revealed that the half-range and central value transformation resulted in a better classification of the tea samples than that achieved using the classical normalisation. The loadings in the PCA for the long-term stability study, under both standard and robust conditions, were found to be similar to the drift trends only when the minimum-maximum transformation and the mean or maximum normalizations were used as data pre-treatments.

Humans↗

Assignment of enzyme substrate specificity by principal component analysis of aligned protein sequences: an experimental test using DNA glycosylase homologs.

We have studied the relationship between amino acid sequence and substrate specificity in a DNA glycosylase family by characterizing experimentally the specificity of four new members of the family. We show that principal component analysis (PCA) of the sequence family correctly predicts the substrate specificity of one of the novel homologs even though conventional sequence analysis methods fail to group this homolog with other sequences of the same specificity. PCA also suggested, correctly, that another homolog characterized previously differs in its specificity from those sequences with which it clusters by conventional criteria. These results suggest that principal component analysis of sequence families can be a useful tool in annotating genome sequences when there is ambiguity concerning which subfamily a new homolog belongs to. Published 2000 Wiley-Liss, Inc.

Amino Acid Sequence↗

Folding dynamics of proteins from denatured to native state: principal component analysis.

Several trajectories starting from random configurations and ending in the native state for chymotrypsin inhibitor 2, CI2, are generated using a Go-type model where the backbone torsional angles execute random jumps on which a drift towards their native values is superposed. Bond lengths and bond angles are kept fixed, and the size of the backbone atoms and side groups are recognized. The large datasets obtained are analyzed using a particular type of principal component analysis known as Karhunen-Loeve expansion (KLE). Trajectories are decomposed separately into modes in residue space and time space. General features of different folding trajectories are compared in the modal space and relationships between the structure of CI2 and its folding dynamics are obtained. Dynamic scaling and order reduction of the folding trajectories are discussed. A continuous wavelet transform is used to decompose the nonstationary folding trajectories into windows exhibiting different features of folding dynamics. Analysis of correlations confirms the known two-state nature of folding of CI2. All of the conserved residues of the protein are shown to be stationary in the small modes of the residue space. The sequential nature of folding is shown by examining the slow modes of the trajectories. The present model of protein folding dynamics is compared with the simple Rouse model of polymer dynamics. Principal component analysis is shown to be a very effective tool for the characterization of the general folding features of proteins.

Computational Biology↗

[Principal component analysis and integral methods of cerebral vascular hemodynamic parameters].

OBJECTIVE: To establish a predicting model for stroke according to cerebral vascular hemodynamic indexes and major risk factors of stroke. METHODS: Participants selected from a stroke cohort with 25,355 population in China. The first step was to carry out principal component analysis using CVHI. Logistic regression with principal component and main risk factors of stroke were then served as independent variables and stroke come on as dependent variables. The predictive model was established according to coefficient of regression and probability of each participant was also estimated. Finally, ROC curve was protracted and predictive efficacy was measured. RESULTS: The accumulative contribution rates of four principal components were 58.1%, 79.4%, 88.4% and 94.6% respectively. Seven variables were being selected into the equation with the first to fourth principal component as history of hypertension, age and sex. Area under ROC curve was 0.855 and optimal cut-off point was probability over 0.05. Sensitivity, specificity and accuracy of stroke prediction were 80.7%, 78.5% and 78.5% respectively. CONCLUSION: The model established by principal component and regression could effectively predict the incidence of stroke coming on.

Brain↗

Classification of gasoline data obtained by gas chromatography using a piecewise alignment algorithm combined with feature selection and principal component analysis.

A fast and objective chemometric classification method is developed and applied to the analysis of gas chromatography (GC) data from five commercial gasoline samples. The gasoline samples serve as model mixtures, whereas the focus is on the development and demonstration of the classification method. The method is based on objective retention time alignment (referred to as piecewise alignment) coupled with analysis of variance (ANOVA) feature selection prior to classification by principal component analysis (PCA) using optimal parameters. The degree-of-class-separation is used as a metric to objectively optimize the alignment and feature selection parameters using a suitable training set thereby reducing user subjectivity, as well as to indicate the success of the PCA clustering and classification. The degree-of-class-separation is calculated using Euclidean distances between the PCA scores of a subset of the replicate runs from two of the five fuel types, i.e., the training set. The unaligned training set that was directly submitted to PCA had a low degree-of-class-separation (0.4), and the PCA scores plot for the raw training set combined with the raw test set failed to correctly cluster the five sample types. After submitting the training set to piecewise alignment, the degree-of-class-separation increased (1.2), but when the same alignment parameters were applied to the training set combined with the test set, the scores plot clustering still did not yield five distinct groups. Applying feature selection to the unaligned training set increased the degree-of-class-separation (4.8), but chemical variations were still obscured by retention time variation and when the same feature selection conditions were used for the training set combined with the test set, only one of the five fuels was clustered correctly. However, piecewise alignment coupled with feature selection yielded a reasonably optimal degree-of-class-separation for the training set (9.2), and when the same alignment and ANOVA parameters were applied to the training set combined with the test set, the PCA scores plot correctly classified the gasoline fingerprints into five distinct clusters.

Algorithms↗

NMR spectral quantitation by principal-component analysis. II. Determination of frequency and phase shifts.

This paper extends the use of principal-component analysis in spectral quantification to the estimation of frequency and phase shifts in a single resonant peak across a series of spectra. The estimated parameters can be used to correct the spectra accordingly, resulting in more accurate peak-area estimation. Further, the removal of the variations in phase and frequency cause by instrumental and experimental fluctuations makes it possible to determine more accurately the remaining variations, which bear biological significance. The procedure is demonstrated on simulated data, a 3D chemical-shift-imaging dataset acquired from a cylinder of inorganic phosphate (Pi), and a set of 736 31P NMR in vivo spectra taken from a kinetic study of rate muscle energetics. In all cases, the procedure rapidly and automatically identifies the frequency and phase shifts present in the individual spectra. In the kinetic study, the procedure is used twice, first to adjust the phase and frequency of a reference peak (phosphocreatine) and then to determine the individual frequencies of the Pi peak in each of the spectra which further can be used for estimation of pH changes during the experiment.

Computer Simulation↗

A principal component analysis of multifocal pattern reversal VEP.

Multifocal visual evoked potentials (mfVEP) were recorded with three channels from 31 control subjects. A principal component analysis was applied to all local responses. The first principal component reversed polarity above and below the horizontal meridian in the case of the midline channel and across the vertical meridian in the case of the lateral channel. In addition, the first principal components of the responses around the vertical meridian were reversed in polarity compared to those around the horizontal meridian, consistent with the region near the vertical meridian lying outside the calcarine fissure. A model was proposed that allowed for the construction of a coronal section of V1 based on the distribution of the first principal component. This approach provides a means of deriving a V1 component from mfVEP recordings with only three recording channels.

Adult↗

Quantitative gait evaluation of hip diseases using principal component analysis.

The measurement of five gait parameters, namely, joint angular displacement of lower extremities, floor reaction forces, trajectory for a point of force application, temporal factor and distance factor has been performed with ease and high speed using mini-computer on-line real-time processing. Gait data of 211 patients with hip diseases was normalized, quantified and summarized by the principal component analysis. A 'gait evaluation plane' was formed according to the results obtained by the principal component analysis. The gait evaluation using the plane was compared with clinical conditions of patients, and it was evident that this system can evaluate the recovery of the gait by treatment.

Biomechanical Phenomena↗

PCAVR: a portable laboratory program for performing varimax-rotated principal components analysis of event-related potentials.

A portable laboratory computer program for performing varimax-rotated principal components analysis (PCA) of event-related potentials (ERPs) is described. The program is written in FORTRAN 77; its compiled version requires 429, 140 bytes of memory. The program reads a matrix of numbers from an input file. The PCA can be performed on either the variance/covariance or the correlation matrix. The program computes the first six principal components. The user is given the option of rotating as many of the principal components as desired based upon the percentages of variance that they account for. Eigenvalues, percentages of variance, cumulative percentages of variance, factor loadings, and factor scores are written to an output file. Description of the program is preceded by a conceptual overview both of PCA as a factor analytic technique and of the application of PCA to ERP data analysis.

Brain↗

The Influence Function of Principal Component Analysis by Self-Organizing Rule.

This article is concerned with a neural network approach to principal component analysis (PCA). An algorithm for PCA by the self-organizing rule has been proposed and its robustness observed through the simulation study by Xu and Yuille (1995). In this article, the robustness of the algorithm against outliers is investigated by using the theory of influence function. The influence function of the principal component vector is given in an explicit form. Through this expression, the method is shown to be robust against any directions orthogonal to the principal component vector. In addition, a statistic generated by the self-organizing rule is proposed to assess the influence of data in PCA.

Journal Article↗

Principal components analysis of sources of variability in retinal ganglion cell responses.

An approach to the functional organization of retinal ganglion cell processing in terms of correlation matrices (Levine and Shefner, 1975; 1977a, b; Shefner and Levine, 1979) is extended by applying Principal Components Analysis. This analysis reduces each correlation matrix to a few components which are implicit in the data. The component loadings describe properties of the system in terms of loadings (correlations) of time bins on the underlying components. Each component is identified by the experimental conditions associated with the highest loadings. Mixed conditions are quantitatively interpreted as weighted contributions from the various identified components. This has led to new interpretations of existing data. Several properties of ganglion cell inputs are analyzed in this manner, including ON and OFF processes, center and surround mechanisms, rod and cone inputs, and spatially distinct areas within the receptive field center. Although the details vary, generally one of the components is highly associated with ON processes and the other with OFF and/or MAINTAINED processes. several advantages may be realized through the use of Principal Components Analysis: (1) all of the data contribute to the analysis, (2) the number and relative importance of contributing processes may be assessed, (3) the relative contribution of underlying processes to mixed responses may be assessed, and (4) the most parsimonious representation of the data is obtained.

Animals↗

Independence of soleus H-reflex tests in control and spastic subjects shown by principal components analysis.

Different soleus H-reflex tests are used in the study of neurophysiological mechanisms of motor control. We studied the interdependence pattern of a number of soleus H-reflex tests, i.e., vibratory inhibition, the ratio of the reflex response to direct muscle potential (H/M ratio) and the homonymous recovery curve with a principal components analysis in 48 healthy controls and 38 patients with signs of the upper motoneuron syndrome. In controls, the analysis showed 3 independent principal components (PCs). Vibratory inhibition and H/M ratio loaded on separate components. Late facilitation and late inhibition variables of the recovery curve loaded on the third component due to the positive correlation (P < 0.001) between these variables. In spastic patients the analysis identified 4 independent PCs corresponding with vibratory inhibition, H/M ratio, late facilitation and late inhibition variables, respectively. The findings suggest that the mutual independence of the different soleus H-reflex tests in patients with the upper motoneuron syndrome has retained the control situation to a large extent.

Adolescent↗

Identification of Tibicen cicada species by a Principal Components Analysis of their songs.

Specific identification of three Tibicen cicadas, T. japonicus, T. flammatus and T. bihamatus, by their chirping sounds was carried out using Principal Components Analysis (PCA). High quality recordings of each species were used as the standards. The peak and mean frequencies and the pulse rate were used as the variables. Out of 12 samples recorded in the fields one fell in the vicinity of T. japonicus and all other were positioned near T. bihamatus. Then the cluster analysis of the PCA scores clearly separated each species and allocated the samples in the same way.

Acoustic Stimulation↗

Genome screen for a combined bone phenotype using principal component analysis: the Framingham study.

Genetic factors substantially contribute to variation in bone mass. There is a controversy as to whether shared genetic factors exist for bone mass at different sites. We hypothesize that using a composite phenotypic score of several correlated bone mass measures may provide complementary results for linkage studies. In the members of 323 pedigrees from the Framingham Osteoporosis Study, bone mineral density (BMD) was measured at the lumbar spine and three femoral sites (Lunar DPX-L), and quantitative ultrasound (QUS) measured at the calcaneus (Hologic Sahara). Data on age, sex, anthropometry, alcohol and caffeine intake, smoking status, physical activity, menopause, and estrogen use (in females) were also obtained. Principal component analyses of BMD and QUS phenotypes were performed in each sex and generation (parents and offspring). The principal component analyses yielded two components, whose loadings were extracted as principal component scores (PC1 and PC2) for each individual, with PC1 explaining up to 66% of the total variation of all bone mass measurements, and PC2 an additional 24%. Principal component analysis of the three femoral BMD measures resulted in one component (PC_hip) that explained 89-91% of the common variation of hip BMD measures. Quantitative genetic analysis (using the variance components method) revealed that both principal component scores were under significant genetic influences (covariate-adjusted heritabilities of PC1, PC2, and PC_hip were 0.66 +/- 0.07, 0.44 +/- 0.07, and 0.61 +/- 0.06, respectively). For PC1, loci of suggestive linkage were identified on chromosomes 1q21.3 and 8q24.3 with the maximum multipoint LOD scores 2.5 and 2.4, respectively. For PC2, multipoint LOD score was 2.1 on 1p36. Suggestive linkage of PC_hip was found on 8q24.3 and 16p13.2 (LODs>1.9). In conclusion, an approach to linkage analysis using the linear combination of several correlated bone phenotypes suggests that there are chromosomal loci regulating bone mass, with seemingly pleiotropic effects at different skeletal sites.

Adult↗

[Principal component analysis and cluster analysis of inorganic elements in Panax quinque folium. L].

The contents of elements such as Mg, Al, P, Ca, V, Cr, Mn, Fe, Co, Ni, Cu, Zn, As, Se, Sr, Mo, Cd and Pb in twelve Panax quinque folium. L samples were determined by means of ICP/MS. The results were used for the development of element fingerprint chromatogram. The principal component analysis of SPSS was applied for the study of characteristic elements in Panax quinque folium. L. Five principal components which accounted for over 90% of the total variance were extracted from the original data. The analysis results show that Fe, Al, V, Mn, Mg, Sr, Mo, Ca and Cu may be the characteristic elements in Panax quinque folium. L. The results of Q-type cluster analysis show that the samples could be clustered reasonably into five groups, and the elemental distribution characteristics are related to the breeds of Panax quinque folium. L.

Cluster Analysis↗

Headspace-solid-phase microextraction fast GC in combination with principal component analysis as a tool to classify different chemotypes of chamomile flower-heads (Matricaria recutita l.).

Headspace-solid-phase microextraction gas chromatography-principal component analysis (HS-SPME GC-PCA) is proposed as a complementary or alternative method to essential oil (EO) GC-PCA in order to discriminate between flower-heads of chamomile of different chemotypes. Ninety-two EOs and the headspaces sampled by HS-SPME of the corresponding chamomile flower-heads were examined by conventional GC and fast GC (F-GC) and the results submitted to statistical analysis by PCA. HS-SPME F-GC-PCA showed itself to be a rapid technique by which to distinguish chamomile flower-head chemotypes a produced results in agreement with the accepted EO classification. Using this method, the analysis time was reduced from at least 4.5 h with EO conventional GC to less than 1 h with HS-SPME F-GC. This approach can thus successfully be used as an analytical decision maker in order to reduce the number of time-consuming EO conventional GC analyses by limiting them to those samples that cannot unequivocally be classified. The EO conventional GC and HS-SPME F-GC results of PCA were very uniform, but they did not provide quantitative correlations between the components as determined by the two methods. A different statistical approach and a larger number of samples will be needed in order to correlate components in the headspace sampled by SPME and those in the corresponding EO quantitatively through a function.

Chromatography, Gas↗

Dynamic monitoring and control of patient anaesthetic and dose levels: time-delay, moving-average neural networks, and principal components analysis.

The goal of this study was to examine the capabilities of neural network models for dynamic monitoring and control of patient anaesthetic and dose levels. The network models that we considered are split into two basic groups: static networks and dynamic networks. Static networks are characterised by equations that are memoryless. On the other hand, dynamic networks are systems with memory. Additionally, principal components analysis was used to introduce a further improvement to network design by reducing the dimensionality of the encoded temporal information. Principal components analysis was applied as both pre-processing and post-processing techniques. In the first instance it was used to reduce the dimensionality of the data to more manageable intrinsic information. In the second instance it was employed to understand how the hidden layers separate the data, in order to optimise the network architecture.

Anesthesia↗

Spectral coherence in normal adults: unrestricted principal components analysis; relation of factors to age, gender, and neuropsychologic data.

This paper demonstrates, by means of Principal Components Analysis (PCA), an objective approach to the reduction of large data sets produced by multichannel spectral coherence analyses. Coherence data, gathered from 371 normal healthy adults using Hjorth/Laplacian referencing during waking eyes-open and eyes-closed states, were analyzed by "unrestricted" PCA where neither spatial nor temporal variance was folded into among subject variance. There was substantial data reduction with our 4416 initial coherence variables for each state reduced to just 150 factors containing approximately 80% of the variance reflecting a 30 fold concentration of information content. Varimax rotation of the first 40 factors, encompassing 50% of the total variance for both states, revealed loading patterns primarily bilateral with no hemispheric bias, relationships primarily between distant single electrode pairs, (although a single electrode to multiple electrode pattern was also observed), and involvement of all spectral bands. Elemental left to right and anterior to posterior coherence patterns, often used on an a priori basis for coherence studies, were not evident among the rotated factor loading patterns. On the basis of high loadings upon extra bipolar artifact channels, 32 factors accounting for approximately 40% of the variance were identified as reflecting artifactual coherence relationships. By multiple regression the 48 non-artifactual factor scores successfully predicted subject age. In general, coherence diminished with age, which may partly explain age-related EEG desynchronization in healthy adults. Coherence factors also predicted 6 of 10 neuropsychologic variables. Gender was successfully predicted by discriminant analysis. No global interpretations about coherence and gender or neuropsychologic function were possible, i.e., almost equal numbers of factors increased as decreased in males as females. PCA derived coherence factor scores are useful for subsequent statistical analyses, but their factor loading plots of cortical coupling may require more experience to fully interpret.

Adult↗