PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Protein Structure, Secondary”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Statistical correlation between protein secondary structure and messenger RNA stem-loop structure.

A new integrated sequence-structure database, called IADE (Integrated ASTRAL-DSSP-EMBL), incorporating matching mRNA sequence, amino acid sequence, and protein secondary structural data, is constructed. It includes 648 protein domains. Based on the IADE database, we studied the relation between RNA stem-loop frequencies and protein secondary structure. It was found that the alpha-helices and beta-strands on proteins tend to be preferably "coded" by mRNA stem region, while the coils on proteins tend to be preferably "coded" by mRNA loop region. These tendencies are more obvious if we observe the structural words (SWs). An SW is defined by a four-amino-acid-fragment that shows the pronounced secondary structural (alpha-helix or beta-strand) propensity. It is demonstrated that the deduced correlation between protein and mRNA structure can hardly be explained as the stochastic fluctuation effect.

Databases as Topic↗

Studies on the relationships between the synonymous codon usage and protein secondary structural units.

The relationship between the synonymous codon usage and protein secondary structural elements (alpha helices and beta sheets) were reinvestigated by taking structural information of proteins from Protein Data Bank (PDB) and their corresponding mRNA sequences from GenBank for four different organisms E. coli, B. subtilis, S. cerevisiae, and Homo sapiens. It was observed that synonymous codon families have non-random codon usage, but there does not exist any species invariant universal correlation between the synonymous codon usage and protein secondary structural elements. The secondary structural units of proteins can be distinguished from the occurrences of bases at the second codon position.

Amino Acids↗

Sequence representation and prediction of protein secondary structure for structural motifs in twilight zone proteins.

Characterizing and classifying regularities in protein structure is an important element in uncovering the mechanisms that regulate protein structure, function and evolution. Recent research concentrates on analysis of structural motifs that can be used to describe larger, fold-sized structures based on homologous primary sequences. At the same time, accuracy of secondary protein structure prediction based on multiple sequence alignment drops significantly when low homology (twilight zone) sequences are considered. To this end, this paper addresses a problem of providing an alternative sequences representation that would improve ability to distinguish secondary structures for the twilight zone sequences without using alignment. We consider a novel classification problem, in which, structural motifs, referred to as structural fragments (SFs) are defined as uniform strand, helix and coil fragments. Classification of SFs allows to design novel sequence representations, and to investigate which other factors and prediction algorithms may result in the improved discrimination. Comprehensive experimental results show that statistically significant improvement in classification accuracy can be achieved by: (1) improving sequence representations, and (2) removing possible noise on the terminal residues in the SFs. Combining these two approaches reduces the error rate on average by 15% when compared to classification using standard representation and noisy information on the terminal residues, bringing the classification accuracy to over 70%. Finally, we show that certain prediction algorithms, such as neural networks and boosted decision trees, are superior to other algorithms.

Algorithms↗

Estimation of protein secondary structure from circular dichroism spectra: inclusion of denatured proteins with native proteins in the analysis.

We have expanded our reference set of proteins used in the estimation of protein secondary structure by CD spectroscopy from 29 to 37 proteins by including 3 additional globular proteins with known X-ray structure and 5 denatured proteins. We have also modified the self-consistent method for analyzing protein CD spectra, SELCON3, by including a new selection criterion developed by W. C. Johnson, Jr. (Proteins Struct. Funct. Genet. 35, 307-312, 1999). The secondary structure corresponding to the denatured proteins was approximated to be 90% unordered, owing to the spectral similarity of the denatured proteins and unordered structures. We examined the thermal denaturation of ribonuclease T1 by CD using both the original and expanded sets of reference proteins and obtained more consistent results with the expanded set. The expanded set of reference proteins will be helpful for the determination of protein secondary structure from protein CD spectra with higher reliability, especially of proteins with significant unordered structure content and/or in the course of denaturation.

Circular Dichroism↗

Integrating protein secondary structure prediction and multiple sequence alignment.

Modern protein secondary structure prediction methods are based on exploiting evolutionary information contained in multiple sequence alignments. Critical steps in the secondary structure prediction process are (i) the selection of a set of sequences that are homologous to a given query sequence, (ii) the choice of the multiple sequence alignment method, and (iii) the choice of the secondary structure prediction method. Because of the close relationship between these three steps and their critical influence on the prediction results, secondary structure prediction has received increased attention from the bioinformatics community over the last few years. In this treatise, we discuss recent developments in computational methods for protein secondary structure prediction and multiple sequence alignment, focus on the integration of these methods, and provide some recommendations for state-of-the-art secondary structure prediction in practice.

Amino Acid Sequence↗

Automatic amide I frequency selection for rapid quantification of protein secondary structure from Fourier transform infrared spectra of proteins.

Here we report the development of a new neural network based approach for rapid quantification of protein secondary structure from Fourier transform infrared (FTIR) spectra of proteins. A technique for efficiently reducing the amount of spectral data by almost 90% is suggested to facilitate faster neural network analysis. Additionally, an automatic procedure is introduced for selecting only those regions within the amide I band of protein FTIR spectra, which can be best related to secondary structure contents by subsequent neural network analysis. Based on a given reference set of FTIR spectra from proteins with known secondary structure, a subset of merely 29 out of 101 amide I absorbance values could be identified, which lead to an improved prediction accuracy. The average prediction accuracy achieved for helix, sheet, turn, bend, and other is 4.96% which is better than that achieved by alternative methods that have been previously reported indicating the significant potential of this approach. Our suggested automatic amide I frequency selection procedure may be easily extended to identify promising regions from spectral data recorded by other spectroscopic techniques, like for example circular dichroism spectroscopy.

Algorithms↗

Designed molecules that fold to mimic protein secondary structures.

Molecules that fold to mimic protein secondary structures have emerged as important targets of bioorganic chemistry. Recently, a variety of compounds that mimic helices, turns, and sheets have been developed, with notable advances in the design of beta-peptides that mimic each of these structures. These compounds hold promise as a step toward synthetic molecules with protein-like properties and as drugs that block protein-protein interactions.

Carbohydrate Sequence↗

Computational methods for protein secondary structure prediction using multiple sequence alignments.

Efforts to use computers in predicting the secondary structure of proteins based only on primary structure information started over a quarter century ago [1-3]. Although the results were encouraging initially, the accuracy of the pioneering methods generally did not attain the level required for using predictions of secondary structures reliably in modelling the three-dimensional topology of proteins. During the last decade, however, the introduction of new computational techniques as well as the use of multiple sequence information has lead to a dramatic increase in the success rate of prediction methods, such that successful 3D modelling based on predicted secondary structure has become feasible [e.g., Ref 4]. This review is aimed at presenting an overview of the scale of the secondary structure prediction problem and associated pitfalls, as well as the history of the development of computational prediction methods. As recent successful strategies for secondary structure prediction all rely on multiple sequence information, some methods for accurate protein multiple sequence alignments will also be described. While the main focus is on prediction methods for globular proteins, also the prediction of trans-membrane segments within membrane proteins will be briefly summarised. Finally, an integrated iterative approach tying secondary structure prediction and multiple alignment will be introduced [5].

Algorithms↗

MUPRED: a tool for bridging the gap between template based methods and sequence profile based methods for protein secondary structure prediction.

Predicting secondary structures from a protein sequence is an important step for characterizing the structural properties of a protein. Existing methods for protein secondary structure prediction can be broadly classified into template based or sequence profile based methods. We propose a novel framework that bridges the gap between the two fundamentally different approaches. Our framework integrates the information from the fuzzy k-nearest neighbor algorithm and position-specific scoring matrices using a neural network. It combines the strengths of the two methods and has a better potential to use the information in both the sequence and structure databases than existing methods. We implemented the framework into a software system MUPRED. MUPRED has achieved three-state prediction accuracy (Q3) ranging from 79.2 to 80.14%, depending on which benchmark dataset is used. A higher Q3 can be achieved if a query protein has a significant sequence identity (>25%) to a template in PDB. MUPRED also estimates the prediction accuracy at the individual residue level more quantitatively than existing methods. The MUPRED web server and executables are freely available at http://digbio.missouri.edu/mupred.

Algorithms↗

Protein secondary structure prediction using dynamic programming.

In the present paper, we describe how a directed graph was constructed and then searched for the optimum path using a dynamic programming approach, based on the secondary structure propensity of the protein short sequence derived from a training data set. The protein secondary structure was thus predicted in this way. The average three-state accuracy of the algorithm used was 76.70%.

Algorithms↗

New method for protein secondary structure assignment based on a simple topological descriptor.

A simple, five-element descriptor, derived from the Delaunay tessellation of a protein structure in a single point per residue representation, can be assigned to each residue in the protein. The descriptor characterizes main-chain topology and connectivity in the neighborhood of the residue and does not explicitly depend on putative hydrogen bonds or any geometric parameter, including bond length, angles, and areas. Rules based on this descriptor can be used for accurate, robust, and computationally efficient secondary structure assignment that correlates well with the existing methods.

Algorithms↗

Bayesian network multi-classifiers for protein secondary structure prediction.

Successful secondary structure predictions provide a starting point for direct tertiary structure modelling, and also can significantly improve sequence analysis and sequence-structure threading for aiding in structure and function determination. Hence the improvement of predictive accuracy of the secondary structure prediction becomes essential for future development of the whole field of protein research. In this work we present several multi-classifiers that combine the predictions of the best current classifiers available on Internet. Our results prove that combining the predictions of a set of classifiers by creating composite classifiers is a fruitful one. We have created multi-classifiers that are more accurate than any of the component classifiers. The multi-classifiers are based on Bayesian networks. They are validated with 9 different datasets. Their predictive accuracy results outperform the best secondary structure predictors by 1.21% on average. Our main contributions are: (i) we improved the best know predictive accuracy by 1.21%, (ii) our best results have been obtained with a new semi naïve Bayes approach named Pazzani-EDA and (iii) our multi-classifiers combine results of previously build classifiers predictions obtained through Internet, thanks to our development of a Java application.

Bayes Theorem↗

[The possible role of the elements of protein secondary structure in adaptation to the action of ionizing radiation].

Changes in the secondary structure of enzymes induced by gamma-rays 60Co at doses not exceeding one ionization per macromolecule were studied to elucidate a possible role of radiation-chemical processes in the evolution of proteins. The data on the comparative radioresistance of various types of secondary protein structures, alpha-helix, parallel and anti-parallel beta-structures, and beta-turn, were obtained by the method of circular dichroism. It was shown that beta-turns were resistant against radiation, alpha-helix was relatively stable, and beta-layer underwent significant changes. The importance of these structural types in the evolution of proteins is discussed. A special role of beta-turn as structural elements fixing the confirmation of macromolecules and therefore responsible for adaptation of the protein structure against a constant radiation background is proposed.

Adaptation, Physiological↗

Near-infrared analysis of protein secondary structure in aqueous solutions and freeze-dried solids.

Near-infrared spectroscopy (NIR) of various proteins (bovine serum albumin, lysozyme, ovalbumin, gamma-globulin, beta-lactoglobulin, myoglobin, cytochrome-c) was investigated as a possible analytical method of the protein secondary structure in various physical states. The spectra of proteins in aqueous solutions (transmission mode, solvent-compensated) and those in freeze-dried solids (nondestructive diffuse reflection mode) showed several bands at similar frequencies in the combination (4000-5000 cm(-1)) and first overtone (5600-6600 cm(-1)) spectral regions. The normalized second-derivative near-infrared spectra of proteins in aqueous solutions suggested that some bands indicated alpha-helix (4090, 4365-4370, 4615, and 5755 cm(-1)) and beta-sheet (4060, 4405, 4525-4540, 4865, and 5915-5925 cm(-1)) structures. The proteins mostly maintained spectra characteristic of their native structure after freeze-drying, although some reductions in alpha-helical structure and increase in unordered or beta-sheet structures were observed. The near-infrared analysis also showed beta-sheet formation of heat-treated BSA in aqueous solutions and in subsequently freeze-dried solids. The present results thus indicated that the nondestructive near-infrared analysis can be used for the investigation of dehydration-induced changes in protein secondary structures.

Freeze Drying↗

SOPMA: significant improvements in protein secondary structure prediction by consensus prediction from multiple alignments.

Recently a new method called the self-optimized prediction method (SOPM) has been described to improve the success rate in the prediction of the secondary structure of proteins. In this paper we report improvements brought about by predicting all the sequences of a set of aligned proteins belonging to the same family. This improved SOPM method (SOPMA) correctly predicts 69.5% of amino acids for a three-state description of the secondary structure (alpha-helix, beta-sheet and coil) in a whole database containing 126 chains of non-homologous (less than 25% identity) proteins. Joint prediction with SOPMA and a neural networks method (PHD) correctly predicts 82.2% of residues for 74% of co-predicted amino acids. Predictions are available by Email to deleage@ibcp.fr or on a Web page (http:@www.ibcp.fr/predict.html).

Databases, Factual↗

Skewed distribution of protein secondary structure contents over the conformational triangle.

A conformational triangle method is presented to analyze the secondary structure contents of 1028 structurally known proteins in the non-redundant data set of the recent 25% PDB_SELECT. The secondary structure contents of each protein are mapped on to a point in the triangle. It was found that the distribution of the 1028 points is strongly skewed in the triangle and about 42% of the whole area is empty, which is called the forbidden area. The detailed border between the allowable and forbidden areas was calculated. The possible explanation of the skewed distribution is discussed. The distributions of the mapping points for enzymes and non-enzymes in this non-redundant data set are compared. It was found that a necessary rather than a sufficient condition for an enzyme molecule is that its coil content must be >/=0.223. It is hoped that the skewed distribution observed here could be used to test the secondary structure and threading predictions.

Computational Biology↗

Predicting protein secondary structure by cascade-correlation neural networks.

The back-propagation neural network algorithm is a commonly used method for predicting the secondary structure of proteins. Whilst popular, this method can be slow to learn and here we compare it with an alternative: the cascade-correlation architecture. Using a constructive algorithm, cascade-correlation achieves predictive accuracies comparable to those obtained by back-propagation, in shorter time.

Algorithms↗

Chemometric tools for classification and elucidation of protein secondary structure from infrared and circular dichroism spectroscopic measurements.

Protein classification and characterization often rely on the information contained in the protein secondary structure. Protein class assignment is usually based on X-ray diffraction measurements, which need the protein in a crystallized form, or on NMR spectra, to obtain the structure of a protein in solution. Simple spectroscopic techniques, such as circular dichroism (CD) and infrared (IR) spectroscopies, are also known to be related to protein secondary structure, but they have seldom been used for protein classification. To see the potential of CD, IR, and combined CD/IR measurements for protein classification, unsupervised pattern recognition methods, Principal Component Analysis (PCA) and cluster analysis, are proposed first to check for natural grouping tendencies of proteins according to their measured spectra. Partial Least Squares Discriminant Analysis (PLS-DA), a supervised pattern recognition method, is used afterwards to test the possibility to model explicitly each protein class and to test these models in class assignment of unknown proteins. Determination of the protein secondary structure, understood as the prediction of the abundance of the different secondary structure motifs in the biomolecule, was carried out with the local regression method interval Partial Least Squares (iPLS). CD, IR, and CD/IR measurements were correlated to the fraction of the motif to be predicted, determined from X-ray measurements. iPLS builds models extracting the spectral information most correlated to a specific secondary motif and avoids the use of irrelevant spectral regions. Spectral intervals chosen by iPLS models provide structural information which can be used to confirm previous biochemical assignments or identify new motif-related spectral features. The predictive ability of the models built with the selected spectral regions has a quality similar to previous classical approaches.

Circular Dichroism↗