PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Protein Structure, Secondary”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Skewed distribution of protein secondary structure contents over the conformational triangle.

A conformational triangle method is presented to analyze the secondary structure contents of 1028 structurally known proteins in the non-redundant data set of the recent 25% PDB_SELECT. The secondary structure contents of each protein are mapped on to a point in the triangle. It was found that the distribution of the 1028 points is strongly skewed in the triangle and about 42% of the whole area is empty, which is called the forbidden area. The detailed border between the allowable and forbidden areas was calculated. The possible explanation of the skewed distribution is discussed. The distributions of the mapping points for enzymes and non-enzymes in this non-redundant data set are compared. It was found that a necessary rather than a sufficient condition for an enzyme molecule is that its coil content must be >/=0.223. It is hoped that the skewed distribution observed here could be used to test the secondary structure and threading predictions.

Computational Biology↗

Specific correlations between relative synonymous codon usage and protein secondary structure.

We found significant species-specific correlations between the use of two synonymous codons and protein secondary structure units by comparing the three-dimensional structures of human and Escherichia coli proteins with their mRNA sequences. The correlations are not explained by codon-context, expression level, GC/AU content, or positional effects. The E. coli correlation is between Asn AAC and the C-terminal regions of beta-sheet segments; it may result from selection for translational accuracy, suggesting the hypothesis that downstream Asn residues are important for beta-sheet formation. The correlation in human proteins is between Asp GAU and the N termini of alpha-helices; it may be important for eukaryote-specific sequential, cotranslational folding. The kingdom-specific correlations may reflect kingdom-specific differences in translational mechanisms. The correlations may help identify residues that are important for secondary structure formation, be useful in secondary structure prediction algorithms, and have implications for recombinant gene expression.

Bacterial Proteins↗

Protein secondary structure prediction with dihedral angles.

We present DESTRUCT, a new method of protein secondary structure prediction, which achieves a three-state accuracy (Q3) of 79.4% in a cross-validated trial on a nonredundant set of 513 proteins. An iterative set of cascade-correlation neural networks is used to predict both secondary structure and psi dihedral angles, with predicted values enhancing the subsequent iteration. Predictive accuracies of 80.7% and 81.7% are achieved on the CASP4 and CASP5 targets, respectively. Our approach is significantly more accurate than other contemporary methods, due to feedback and a novel combination of structural representations.

Amino Acid Sequence↗

Prediction of protein secondary structure using the 3D-1D compatibility algorithm.

A new method for the prediction of protein secondary structure is proposed, which relies totally on the global aspect of a protein. The prediction scheme is as follows. A structural library is first scanned with a query sequence by the 3D-1D compatibility method developed before. All the structures examined are sorted with the compatibility score and the top 50 in the list are picked out. Then, all the known secondary structures of the 50 proteins are globally aligned against the query sequence, according to the 3D-1D alignments. Prediction of either alpha helix, beta strand or coil is made by taking the majority among the observations at each residue site. Besides 325 proteins in the structural library, 77 proteins were selected from the latest release of the Brookhaven Protein Data Bank, and they were divided into three data sets. Data set 1 was used as a training set for which several adjustable parameters in the method were optimized. Then, the final form of the method was applied to a testing set (data set 2) which contained proteins of chain length < or = 400 residues. The average prediction accuracy was as high as 69% in the three-state assessment of alpha, beta and coil. On the other hand, data set 3 contains only those proteins of length > 400 residues, for which the present method would not work properly because of the size effect inherent in the 3D-1D compatibility method. The proteins in data set 3 were, therefore, subdivided into constituent domains (data set 4) before being fed into the prediction program. The prediction accuracy for data set 4 was 66% on average, a few percent lower than that for data set 2. Possible causes for this discrepancy are discussed.

Algorithms↗

Evaluation of the information content in infrared spectra for protein secondary structure determination.

Fourier-transform infrared spectroscopy is a method of choice for the experimental determination of protein secondary structure. Numerous approaches have been developed during the past 15 years. A critical parameter that has not been taken into account systematically is the selection of the wavenumbers used for building the mathematical models used for structure prediction. The high quality of the current Fourier-transform infrared spectrometers makes the absorbance at every single wavenumber a valid and almost noiseless type of information. We address here the question of the amount of independent information present in the infrared spectra of proteins for the prediction of the different secondary structure contents. It appears that, at most, the absorbance at three distinct frequencies of the spectra contain all the nonredundant information that can be related to one secondary structure content. The ascending stepwise method proposed here identifies the relevance of each wavenumber of the infrared spectrum for the prediction of a given secondary structure and yields a particularly simple method for computing the secondary structure content. Using the 50-protein database built beforehand to contain as little fold redundancy as possible, the standard error of prediction in cross-validation is 5.5% for the alpha-helix, 6.6% for the beta-sheet, and 3.4% for the beta-turn.

Algorithms↗

Predicting protein secondary structure by cascade-correlation neural networks.

The back-propagation neural network algorithm is a commonly used method for predicting the secondary structure of proteins. Whilst popular, this method can be slow to learn and here we compare it with an alternative: the cascade-correlation architecture. Using a constructive algorithm, cascade-correlation achieves predictive accuracies comparable to those obtained by back-propagation, in shorter time.

Algorithms↗

SOMCD: method for evaluating protein secondary structure from UV circular dichroism spectra.

This article presents SOMCD, an improved method for the evaluation of protein secondary structure from circular dichroism spectra, based on Kohonen's self-organizing maps (SOM). Protein circular dichroism (CD) spectra are used to train a SOM, which arranges the spectra on a two-dimensional map. Location in the map reflects the secondary structure composition of a protein. With SOMCD, the prediction of beta-turn has been included. The number of spectra in the training set has been increased, and it now includes 39 protein spectra and 6 reference spectra. Finally, SOM parameters have been chosen to minimize distortion and make the network produce clusters with known properties. Estimation results show improvements compared with the previous version, K2D, which, in addition, estimated only three secondary structure components; the accuracy of the method is more uniform over the different secondary structures.

Algorithms↗

Protein secondary structure: category assignment and predictability.

In the last decade, the prediction of protein secondary structure has been optimized using essentially one and the same assignment scheme known as DSSP. We present here a different scheme, which is more predictable. This scheme predicts directly the hydrogen bonds, which stabilize the secondary structures. Single sequence prediction of the new three category assignment gives an overall prediction improvement of 3.1% and 5.1% compared to the DSSP assignment and schemes where the helix category consists of alpha-helix and 3(10)-helix, respectively. These results were achieved using a standard feed-forward neural network with one hidden layer on a data set identical to the one used in earlier work.

Algorithms↗

GOR V server for protein secondary structure prediction.

SUMMARY: We have created the GOR V web server for protein secondary structure prediction. The GOR V algorithm combines information theory, Bayesian statistics and evolutionary information. In its fifth version, the GOR method reached (with the full jack-knife procedure) an accuracy of prediction Q3 of 73.5%. Although GOR V has been among the most successful methods, its online unavailability has been a deterrent to its popularity. Here, we remedy this situation by creating the GOR V server.

Algorithms↗

The SSEA server for protein secondary structure alignment.

SUMMARY: We present a web server that computes alignments of protein secondary structures. The server supports both performing pairwise alignments and searching a secondary structure against a library of domain folds. It can calculate global and local secondary structure element alignments. A combination of local and global alignment steps can be used to search for domains inside the query sequence or help in the discrimination of novel folds. Both the SCOP and PDB fold libraries, clustered at 95 and 40% sequence identity, are available for alignment. AVAILABILITY: The web server interface is freely accessible to academic users at http://protein.cribi.unipd.it/ssea/. The executable version and benchmarking data are available from the same web page.

Algorithms↗

Chemometric tools for classification and elucidation of protein secondary structure from infrared and circular dichroism spectroscopic measurements.

Protein classification and characterization often rely on the information contained in the protein secondary structure. Protein class assignment is usually based on X-ray diffraction measurements, which need the protein in a crystallized form, or on NMR spectra, to obtain the structure of a protein in solution. Simple spectroscopic techniques, such as circular dichroism (CD) and infrared (IR) spectroscopies, are also known to be related to protein secondary structure, but they have seldom been used for protein classification. To see the potential of CD, IR, and combined CD/IR measurements for protein classification, unsupervised pattern recognition methods, Principal Component Analysis (PCA) and cluster analysis, are proposed first to check for natural grouping tendencies of proteins according to their measured spectra. Partial Least Squares Discriminant Analysis (PLS-DA), a supervised pattern recognition method, is used afterwards to test the possibility to model explicitly each protein class and to test these models in class assignment of unknown proteins. Determination of the protein secondary structure, understood as the prediction of the abundance of the different secondary structure motifs in the biomolecule, was carried out with the local regression method interval Partial Least Squares (iPLS). CD, IR, and CD/IR measurements were correlated to the fraction of the motif to be predicted, determined from X-ray measurements. iPLS builds models extracting the spectral information most correlated to a specific secondary motif and avoids the use of irrelevant spectral regions. Spectral intervals chosen by iPLS models provide structural information which can be used to confirm previous biochemical assignments or identify new motif-related spectral features. The predictive ability of the models built with the selected spectral regions has a quality similar to previous classical approaches.

Circular Dichroism↗

Direct observation of protein secondary structure in gas vesicles by atomic force microscopy.

The protein that forms the gas vesicle in the cyanobacterium Anabaena flos-aquae has been imaged by atomic force microscopy (AFM) under liquid at room temperature. The protein constitutes "ribs" which, stacked together, form the hollow cylindrical tube and conical end caps of the gas vesicle. By operating the microscope in deflection mode, it has been possible to achieve sub-nanometer resolution of the rib structure. The lateral spacing of the ribs was found to be 4.6 +/- 0.1 nm. At higher resolution the ribs are observed to consist of pairs of lines at an angle of approximately 55 degrees to the rib axis, with a repeat distance between each line of 0.57 +/- 0.05 nm along the rib axis. These observed dimensions and periodicities are consistent with those determined from previous x-ray diffraction studies, indicating that the protein is arranged in beta-chains crossing the rib at an angle of 55 degrees to the rib axis. The AFM results confirm the x-ray data and represent the first direct images of a beta-sheet protein secondary structure using this technique. The orientation of the GvpA protein component of the structure and the extent of this protein across the ribs have been established for the first time.

Anabaena↗

Protein secondary structure prediction in different structural classes.

Information about the secondary structure of a protein can be helpful in understanding its native folded state. In previous work, it was shown that the medium-range interactions predominate in all-alpha class and the long-range interactions predominate in all-beta class proteins. Based on this, in this work the performance of several structure prediction methods in different structural classes of globular proteins was analyzed. It was found that all the methods predict the secondary structures of all-alpha proteins more accurately than other classes.

Protein Structure, Secondary↗

Estimation of protein secondary structure from circular dichroism spectra: comparison of CONTIN, SELCON, and CDSSTR methods with an expanded reference set.

We have expanded the reference set of proteins used in SELCON3 by including 11 additional proteins (selected from the reference sets of Yang and co-workers and Keiderling and co-workers). Depending on the wavelength range and whether or not denatured proteins are included in the reference set, five reference sets were constructed with the number of reference proteins varying from 29 to 48. The performance of three popular methods for estimating protein secondary structure fractions from CD spectra (implemented in software packages CONTIN, SELCON3, and CDSSTR) and a variant of CONTIN, CONTIN/LL, that incorporates the variable selection method in the locally linearized model in CONTIN, were examined using the five reference sets described here, and a 22-protein reference set. Secondary structure assignments from DSSP were used in the analysis. The performances of all three methods were comparable, in spite of the differences in the algorithms used in the three software packages. While CDSSTR performed the best with a smaller reference set and larger wavelength range, and CONTIN/LL performed the best with a larger reference set and smaller wavelength range, the performances for individual secondary structures were mixed. Analyzing protein CD spectra using all three methods should improve the reliability of predicted secondary structural fractions. The three programs are provided in CDPro software package and have been modified for easier use with the different reference sets described in this paper. CDPro software is available at the website: http://lamar.colostate.edu/ approximately sreeram/CDPro.

Circular Dichroism↗

Single-pass attenuated total reflection Fourier transform infrared spectroscopy for the prediction of protein secondary structure.

Principal component regression (PCR) was applied to a spectral library of proteins in H2O solution acquired by single-pass attenuated total reflectance (ATR) Fourier transform infrared (FT-IR) spectroscopy. PCR was used to predict the secondary structure content, principally alpha-helical and the beta-sheet content, of proteins within a spectral library. Quantitation of protein secondary structure content was performed as a proof of principle that use of single-pass ATR-FT-IR is an appropriate method for protein secondary structure analysis. The ATR-FT-IR method permits acquisition of the entire spectral range from 700 to 3900 cm(-1) without significant interference from water bands. An "inside model space" bootstrap and a genetic algorithm (GA) were used to improve prediction results. Specifically, the bootstrap was utilized to increase the number of replicates for adequate training and validation of the PCR model. The GA was used to optimize PCR parameters, particularly wavenumber selection. The use of the bootstrap allowed for adequate representation of variability in the amide A, amide B, and C-H stretching regions due to differing levels of sample hydration. Implementation of the bootstrap improved the robustness of the PCR models significantly; however, the use of a GA only slightly improved prediction results. Two spectral libraries are presented where one was better suited for beta-sheet content prediction and the other for alpha-helix content prediction. The GA-optimized PCR method for alpha-helix content prediction utilized 120 wavenumbers within the amide I, II, A, B, and IV and the C-H stretching regions and 18 factors. For beta-sheet content predictions, 580 wavenumbers within the amide I, II, A, and B and the C-H stretching regions and 18 factors were used. The validation results using these two methods yielded an average absolute error of 1.7% for alpha-helix content prediction and an average absolute error of 2.3% for beta-sheet content prediction. After the PCR models were developed and validated, they were used to predict the alpha-helix and beta-sheet content of two unknowns, casein and immunoglobulin G.

Multivariate Analysis↗

Effect of ethanol on the protein secondary structure of the human gastric mucosa, in vitro.

The effect of ethanol on the secondary conformational structure of proteins of the human gastric mucosa was investigated by attenuated total reflection/Fourier transform infrared (ATR/FT-IR) spectroscopy. The IR peak intensity and position of each structural component of gastric mucosa was found to change significantly with the ethanol concentration and length of exposure. The peak intensity due to the beta-sheet and/or beta-turn conformational structure in amide I and II bands of gastric mucosa clearly increased after treatment with ethanol. Moreover, the peak at 1635 cm-1 shifted to 1630 cm-1 after treatment with 40% ethanol for 3 h, or 80% ethanol for 1 h, and a distinct shoulder also appeared at 1643 cm-1. This shift occurred more rapidly and was more pronounced after exposure of mucosa to 80% ethanol, compared with the effect of 40% ethanol, but the alpha-helical structure at the amide I and II bands was not influenced by either concentration of ethanol. Ethanol treatment might also transform the secondary structure of amide III in gastric mucosa from an alpha-helix to a mainly random coil with extensive unfolding. The absorption between 1180 and 980 cm-1, which is assigned to glycoprotein structure, was also reduced after treatment with ethanol. This strongly indicates that ethanol influences the conformation of the lipids and proteins of human gastric mucosa, leading to their deformation.

Ethanol↗

Clustering of amino acids for protein secondary structure prediction.

Simple hidden Markov models are proposed for predicting secondary structure of a protein from its amino acid sequence. Since the length of protein conformation segments varies in a narrow range, we ignore the duration effect of length distribution, and focus on inclusion of short range correlations of residues and of conformation states in the models. Conformation-independent and -dependent amino acid coarse-graining schemes are designed for the models by means of proper mutual information. We compare models of different level of complexity, and establish a practical model with a high prediction accuracy.

Algorithms↗

Prediction of protein secondary structure by neural networks: encoding short and long range patterns of amino acid packing.

A complex, cascaded neural network designed to predict the secondary structure of globular proteins has been developed. Information about the local buried-unburied pattern and the average tendency of the particular types of amino acids to be buried inside the globule were used. Nonspecific information about long distance contact maps was also employed. These modifications result in a noticeable improvement (3-9%) of prediction accuracy. The best result for the average success ratio for the testing set of nonhomologous proteins was 68.3% (with corresponding Matthews' coefficients, C alpha,beta,coil equal to 0.60, 0.47, 0.43, respectively).

Amino Acid Sequence↗