PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Protein Structure, Secondary”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Prediction of protein secondary structure from circular dichroism spectra: an attempt to solve the problem of the best-fitting reference protein subsets.

In least-squares fitting of protein circular dichroism (CD) spectra using basis CD spectra for the respective secondary structure components, as given by reference proteins of known structural composition, good fits of the CD spectrum do not necessarily correspond to appropriate fits of the underlying structural composition of a test protein. In an attempt to overcome this problem, CD similarity measures were used to construct a subset of five reference CD spectra which permitted three-component fitting of the CD and prediction of the relative magnitude of total beta-sheet (antiparallel + parallel) and total other structure (beta-turn + remainder), relative to helix, in the test protein. A backpropagation neural network (BPN) was also trained to make this prediction. In subsequent five-component fitting of the CD spectrum, using a Monte Carlo method to generate subsets of reference proteins from the working data set, only those secondary structure fits which conformed to the consensus prediction of the CD similarity measures and the BPN were accepted. The method enhanced the fitting of antiparallel beta-sheet and beta-turn for 16 proteins, compared to the variable selection method of P. Manavalan and W. C. Johnson (1987, Anal. Biochem. 167, 76-85). Some other proteins were less well fitted. Improvement in results is expected with a larger representation of feasible basis CD spectra in the working reference proteins.

Circular Dichroism↗

A study of protein secondary structure by Fourier transform infrared/photoacoustic spectroscopy and its application for recombinant proteins.

FT-IR/PAS (Fourier transform infrared/photoacoustic spectroscopy) was used to evaluate the secondary structure of proteins. Four well-studied proteins, concanavalin A, hemoglobin, lysozyme, and trypsin, which have different distributions of secondary structures, were used for assignments of the infrared bands and evaluating the accuracy of FT-IR/PAS methods. Secondary structure contents estimated from FT-IR/PAS and other physical methods (e.g., X-ray diffraction, CD, and traditional FT-IR) show good agreement. In addition, the secondary structure can be evaluated with as little as 0.5 micrograms of protein (concanavalin A), suggesting that FT-IR/PAS is a sensitive and useful technique that could be applied to studies of the folding of recombinant and mutant proteins where only small amounts of material are available. Recombinant phosphorylase kinase gamma 1-300 subunit expressed in Escherichia coli was found in the inclusion bodies. We found that renatured phosphorylase kinase gamma 1-300 subunit has two kinase forms: one has a 10-fold higher activity than the other one. Both fractions, however, are the same as judged from sodium dodecylsulfate-polyacrylamide gel electrophoresis. Differences in conformation were demonstrated by using the FT-IR/PAS method, which showed that the low-activity form has more beta-sheet structure than the form with high activity. Analysis of these kinase forms by CD confirms the interpretation made by the FT-IR/PAS method.

Protein Structure, Secondary↗

A new approach to the evaluation of protein secondary structure predictions at the level of the elements of secondary structure.

For many purposes, such as the prediction of the class of protein folds, the existence of an element of secondary structure rather than its precise position and length must be defined correctly. However, most methods for the evaluation of secondary structure prediction consider success in terms of the percentage of individual amino acids predicted correctly. In this paper the success in predicting elements of secondary structure is discussed. The number of overlapping residues in the predicted and observed secondary structures were considered as a function of the total number of amino acids in the observed and predicted secondary structures. A matrix search procedure was used to remove the ambiguity which similar studies may have had in defining the equivalent secondary structures between predicted and observed structures. In this study a loop was treated in the same way as an alpha-helix and a beta-strand. To describe the accuracy at the level of elements of secondary structure, a set of parameters was defined, similar to those used commonly at the level of individual amino acids. This approach was used to assess the methods of Chou and Fasman (1974b, Biochemistry, 13, 222-245), Lim (1974b, J. Mol. Biol., 88, 873-894) and Garnier et al. (1978, J. Mol. Biol., 120, 97-120). It was found that these methods were much poorer at the secondary structure level than at the amino acid level. This approach can be used generally for secondary structure prediction methods.

Amino Acids↗

The future of protein secondary structure prediction accuracy.

BACKGROUND: The accuracy of secondary structure prediction for a protein from knowledge of its sequence has been significantly improved by about 7% to the 70-75% range by inclusion of information residing in sequences similar to the query sequence. The scientific literature has been inconsistent, if not negative, regarding chances for further improvement from the vast knowledge to be provided by genome sequencing efforts. RESULTS: By applying a prediction technique that is particularly sensitive to added sequence information to a standard set of query sequences with related primary structures taken from chronologically successive releases of the SWISS-PROT database, it is shown that prediction accuracy can be expected to reach 80-85% with a large 10-fold increase in present sequence knowledge. CONCLUSIONS: Even with present prediction approaches, improvement in prediction accuracy can still be expected, albeit limited to no more than 10%.

Algorithms↗

Accurate and automated classification of protein secondary structure with PsiCSI.

PsiCSI is a highly accurate and automated method of assigning secondary structure from NMR data, which is a useful intermediate step in the determination of tertiary structures. The method combines information from chemical shifts and protein sequence using three layers of neural networks. Training and testing was performed on a suite of 92 proteins (9437 residues) with known secondary and tertiary structure. Using a stringent cross-validation procedure in which the target and homologous proteins were removed from the databases used for training the neural networks, an average 89% Q3 accuracy (per residue) was observed. This is an increase of 6.2% and 5.5% (representing 36% and 33% fewer errors) over methods that use chemical shifts (CSI) or sequence information (Psipred) alone. In addition, PsiCSI improves upon the translation of chemical shift information to secondary structure (Q3 = 87.4%) and is able to use sequence information as an effective substitute for sparse NMR data (Q3 = 86.9% without (13)C shifts and Q3 = 86.8% with only H(alpha) shifts available). Finally, errors made by PsiCSI almost exclusively involve the interchange of helix or strand with coil and not helix with strand (<2.5 occurrences per 10000 residues). The automation, increased accuracy, absence of gross errors, and robustness with regards to sparse data make PsiCSI ideal for high-throughput applications, and should improve the effectiveness of hybrid NMR/de novo structure determination methods. A Web server is available for users to submit data and have the assignment returned.

Automation↗

Evaluation and improvement of multiple sequence methods for protein secondary structure prediction.

A new dataset of 396 protein domains is developed and used to evaluate the performance of the protein secondary structure prediction algorithms DSC, PHD, NNSSP, and PREDATOR. The maximum theoretical Q3 accuracy for combination of these methods is shown to be 78%. A simple consensus prediction on the 396 domains, with automatically generated multiple sequence alignments gives an average Q3 prediction accuracy of 72.9%. This is a 1% improvement over PHD, which was the best single method evaluated. Segment Overlap Accuracy (SOV) is 75.4% for the consensus method on the 396-protein set. The secondary structure definition method DSSP defines 8 states, but these are reduced by most authors to 3 for prediction. Application of the different published 8- to 3-state reduction methods shows variation of over 3% on apparent prediction accuracy. This suggests that care should be taken to compare methods by the same reduction method. Two new sequence datasets (CB513 and CB251) are derived which are suitable for cross-validation of secondary structure prediction methods without artifacts due to internal homology. A fully automatic World Wide Web service that predicts protein secondary structure by a combination of methods is available via http://barton.ebi.ac.uk/.

Algorithms↗

Continuum secondary structure captures protein flexibility.

The DSSP program assigns protein secondary structure to one of eight states. This discrete assignment cannot describe the continuum of thermal fluctuations. Hence, a continuous assignment is proposed. Technically, the continuum results from averaging over ten discrete DSSP assignments with different hydrogen bond thresholds. The final continuous assignment for a single NMR model successfully reflected the structural variations observed between all NMR models in the ensemble. The structural variations between NMR models were verified to correlate with thermal motion; these variations were captured by the continuous assignments. Because the continuous assignment reproduces the structural variation between many NMR models from one single model, functionally important variation can be extracted from a single X-ray structure. Thus, continuous assignments of secondary structure may affect future protein structure analysis, comparison, and prediction.

Computer Systems↗

SOPM: a self-optimized method for protein secondary structure prediction.

A new method called the self-optimized prediction method (SOPM) has been developed to improve the success rate in the prediction of the secondary structure of proteins. This new method has been checked against an updated release of the Kabsch and Sander database, 'DATABASE.DSSP', comprising 239 protein chains. The first step of the SOPM is to build sub-databases of protein sequences and their known secondary structures drawn from 'DATABASE.DSSP' by (i) making binary comparisons of all protein sequences and (ii) taking into account the prediction of structural classes of proteins. The second step is to submit each protein of the sub-database to a secondary structure prediction using a predictive algorithm based on sequence similarity. The third step is to iteratively determine the predictive parameters that optimize the prediction quality on the whole sub-database. The last step is to apply the final parameters to the query sequence. This new method correctly predicts 69% of amino acids for a three-state description of the secondary structure (alpha helix, beta sheet and coil) in the whole database (46,011 amino acids). The correlation coefficients are C alpha = 0.54, C beta = 0.50 and Cc = 0.48. Root mean square deviations of 10% in the secondary structure content are obtained. Implications for the users are drawn so as to derive an accuracy at the amino acid level and provide the user with a guide for secondary structure prediction. The SOPM method is available by anonymous ftp to ibcp.fr.

Algorithms↗

Improved prediction of protein secondary structure by use of sequence profiles and neural networks.

The explosive accumulation of protein sequences in the wake of large-scale sequencing projects is in stark contrast to the much slower experimental determination of protein structures. Improved methods of structure prediction from the gene sequence alone are therefore needed. Here, we report a substantial increase in both the accuracy and quality of secondary-structure predictions, using a neural-network algorithm. The main improvements come from the use of multiple sequence alignments (better overall accuracy), from "balanced training" (better prediction of beta-strands), and from "structure context training" (better prediction of helix and strand lengths). This method, cross-validated on seven different test sets purged of sequence similarity to learning sets, achieves a three-state prediction accuracy of 69.7%, significantly better than previous methods. In addition, the predicted structures have a more realistic distribution of helix and strand segments. The predictions may be suitable for use in practice as a first estimate of the structural type of newly sequenced proteins.

Amino Acid Sequence↗

Accuracy of protein secondary structure determination from circular dichroism spectra based on immunoglobulin examples.

Strong contribution of the aromatic amino acid side chain chromophores to the far-UV circular dichroism (CD) spectra substantially distorts a relatively weak CD signal originating from beta sheet, the main type of immunoglobulin secondary structure. In this study we compared the secondary structure calculated from the far-UV CD spectra with the X-ray data for three antibody Fab fragments. Calculations were performed with three different algorithms, using two sets of reference proteins. Low standard deviations between all six estimates indicate stable mathematical solutions. Despite pronounced differences in the shape and amplitude of the CD spectra, we found a strong correlation between CD and X-ray data in the secondary structure for every protein studied. The number and average length of the secondary structure elements estimated from the CD spectra closely resemble those of the X-ray data. Agreement between spectroscopic and crystallographic results demonstrates that modern methods of secondary structure calculation are resilient to distortions of the far-UV CD spectra of immunoglobulins caused by aromatic side chain chromophores.

Circular Dichroism↗

[The relation between translation speed and protein secondary structure].

Based on the statistical analysis of 119 human and 92 E. coli proteins it was found that for both human and E. coli, the mRNA sequences consisting of tri-codon and tetra-codon with high translation speed preferably code for alpha helices more than for coils. For beta strand, the preference/avoidance oscillates with the translation speed. Moreover, the non-homogeneous usages of tri-codon and tetra-codon with different translation speeds in a given secondary structure have also been found. These results cannot be simply explained by the effect of stochastic fluctuation.

Codon↗

Global characterization of protein secondary structures. Analysis of computer-modeled protein unfolding.

Analyses of structural and molecular shape changes undergone by a protein during an unfolding process are presented. The procedure, based on a spherical shape map method, provides a topological description of a three-dimensional macromolecular structure. Local properties of the backbone are used to derive a global characterization of its fold. A spherical shape map of backbone crossings is associated with a given macromolecular conformation. The map is built by classifying each point on the sphere according to the crossing pattern obtained when the backbone is observed along a direction defined by the center of the sphere and the chosen point. The surface of the sphere can be divided in equivalence classes. All the points within a given class correspond to directions from which the backbone has the same overcrossing pattern. Automatic computation and display of these equivalence classes is discussed, as is the implementation of the technique on a computer graphics workstation. The graphical manipulation simplifies the analysis of these maps when following a change in the conformation of the backbone. The procedure is illustrated with the results of a molecular dynamics computer simulation of the unfolding of the bacteriophage T4 glutaredoxin protein (in the form of its polyglycine model). The method gives a novel description of the differential structural stability for the characteristic secondary structural elements (alpha-helices and beta-sheets) present in the protein. Recognition of the persistence of structural elements over the simulation time is performed in an unbiased manner.

Algorithms↗

Protection of protein secondary structure by saccharides of different molecular weights during freeze-drying.

The protective effects of saccharides with various molecular weights (glucose, maltose, maltotriose, maltotetraose, maltopentaose, maltoheptaose, dextran 1060, dextran 4900, and dextran 10200) against lyophilization-induced structural perturbation of model proteins (BSA, ovalbumin) were studied. Fourier transform infrared (FT-IR) analysis of the proteins in initial solutions and freeze-dried solids indicated that maltose conferred the greatest protection against secondary structure change. The structure-stabilizing effect of maltooligosaccharides decreased in increasing the number of saccharide units. Larger molecules of dextran also showed a smaller structure-stabilizing effect. Increasing the effective saccharide molecular size by a borate-saccharide complexation reduced the protein structure-stabilizing effect of all of the saccharides except glucose. The results indicate that the larger saccharide molecules, and/or the complex formation with borate ion, reduce the free and accessible hydroxyl groups to interact with and stabilize the protein structure by a water-substitution mechanism.

Calorimetry, Differential Scanning↗

Elucidating protein secondary structures using alpha-carbon recurrence quantifications.

Secondary structures of proteins were studied by recurrence quantification analysis (RQA). High-resolution, 3-dimensional coordinates of alpha-carbon atoms comprising a set of 68 proteins were downloaded from the Protein Data Bank. By fine-tuning four recurrence parameters (radius, line, residue, separation), it was possible to establish excellent agreement between percent contribution of alpha-helix and beta-sheet structures determined independently by RQA and that of the DSSP algorithm (Define Secondary Structure of Proteins). These results indicate that there is an equivalency between these two techniques, which are based upon totally different pattern recognition strategies. RQA enhances qualitative contact maps by quantifying the arrangements of recurrent points of alpha carbons close in 3-dimensional space. For example, the radius was systematically increased, moving the analysis beyond local alpha-carbon neighborhoods in order to capture super-secondary and tertiary structures. However, differences between proteins could only be detected within distances up to about 6-11 A, but not higher. This result underscores the complexity of alpha-carbon spacing when super-secondary structures appear at larger distances. Finally, RQA-defined secondary structures were found to be robust against random displacement of alpha carbons upwards of 1 A. This finding has potential import for the dynamic functions of proteins in motion.

Bacterial Proteins↗

Efficient characterization of protein secondary structure in terms of screw motions.

A simple and efficient method is presented to describe the secondary structure of proteins in terms of orientational distances between consecutive peptide planes and local helix parameters. The method uses quaternion-based superposition fits of the protein peptide planes in conjunction with Chasles' theorem, which states that any rigid-body displacement can be described by a screw motion. The helix parameters are derived from the best superposition of consecutive peptide planes and the ;worst' fit is used to define the orientational distance. Applications are shown for standard secondary-structure motifs of peptide chains for several proteins belonging to different fold classes and for a description of structural changes in lysozyme under hydrostatic pressure. In the latter case, published reference data obtained by X-ray crystallography and by structural NMR measurements are used.

Algorithms↗

Bayesian segmentation of protein secondary structure.

We present a novel method for predicting the secondary structure of a protein from its amino acid sequence. Most existing methods predict each position in turn based on a local window of residues, sliding this window along the length of the sequence. In contrast, we develop a probabilistic model of protein sequence/structure relationships in terms of structural segments, and formulate secondary structure prediction as a general Bayesian inference problem. A distinctive feature of our approach is the ability to develop explicit probabilistic models for alpha-helices, beta-strands, and other classes of secondary structure, incorporating experimentally and empirically observed aspects of protein structure such as helical capping signals, side chain correlations, and segment length distributions. Our model is Markovian in the segments, permitting efficient exact calculation of the posterior probability distribution over all possible segmentations of the sequence using dynamic programming. The optimal segmentation is computed and compared to a predictor based on marginal posterior modes, and the latter is shown to provide significant improvement in predictive accuracy. The marginalization procedure provides exact secondary structure probabilities at each sequence position, which are shown to be reliable estimates of prediction uncertainty. We apply this model to a database of 452 nonhomologous structures, achieving accuracies as high as the best currently available methods. We conclude by discussing an extension of this framework to model nonlocal interactions in protein structures, providing a possible direction for future improvements in secondary structure prediction accuracy.

Algorithms↗

S curve, a graphic representation of protein secondary structure sequence and its applications.

A secondary structure sequence is a symbolic string composed of three kinds of letters, indicating the helix, strand, and coil (including turns), respectively. A graphic representation for this abstract symbolic sequence is proposed here, called the S curve. The S curve is the unique representation for a given secondary structure sequence in the sense that the sequence and the S curve can be uniquely determined from the other. Therefore, the S curve contains all the information that the secondary structure sequence contains. Different geometrical properties of the S curve are studied in details, which reflect the basic characteristics of the secondary structure sequences. The S curves are used to display, analyze, and compare the secondary structure sequences. Detailed application examples are presented. One advantage of the S curve methodology is that the main patterns of a given secondary structure sequence can be grasped quickly in a perceivable form. This is particularly useful in the cases in which longer sequences are involved and structures of proteins are unknown.

Amino Acid Sequence↗