PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Protein Structure, Secondary”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Single-pass attenuated total reflection Fourier transform infrared spectroscopy for the prediction of protein secondary structure.

Principal component regression (PCR) was applied to a spectral library of proteins in H2O solution acquired by single-pass attenuated total reflectance (ATR) Fourier transform infrared (FT-IR) spectroscopy. PCR was used to predict the secondary structure content, principally alpha-helical and the beta-sheet content, of proteins within a spectral library. Quantitation of protein secondary structure content was performed as a proof of principle that use of single-pass ATR-FT-IR is an appropriate method for protein secondary structure analysis. The ATR-FT-IR method permits acquisition of the entire spectral range from 700 to 3900 cm(-1) without significant interference from water bands. An "inside model space" bootstrap and a genetic algorithm (GA) were used to improve prediction results. Specifically, the bootstrap was utilized to increase the number of replicates for adequate training and validation of the PCR model. The GA was used to optimize PCR parameters, particularly wavenumber selection. The use of the bootstrap allowed for adequate representation of variability in the amide A, amide B, and C-H stretching regions due to differing levels of sample hydration. Implementation of the bootstrap improved the robustness of the PCR models significantly; however, the use of a GA only slightly improved prediction results. Two spectral libraries are presented where one was better suited for beta-sheet content prediction and the other for alpha-helix content prediction. The GA-optimized PCR method for alpha-helix content prediction utilized 120 wavenumbers within the amide I, II, A, B, and IV and the C-H stretching regions and 18 factors. For beta-sheet content predictions, 580 wavenumbers within the amide I, II, A, and B and the C-H stretching regions and 18 factors were used. The validation results using these two methods yielded an average absolute error of 1.7% for alpha-helix content prediction and an average absolute error of 2.3% for beta-sheet content prediction. After the PCR models were developed and validated, they were used to predict the alpha-helix and beta-sheet content of two unknowns, casein and immunoglobulin G.

Multivariate Analysis↗

Multi-class support vector machines for protein secondary structure prediction.

The solution of binary classification problems using the Support Vector Machine (SVM) method has been well developed. Though multi-class classification is typically solved by combining several binary classifiers, recently, several multi-class methods that consider all classes at once have been proposed. However, these methods require resolving a much larger optimization problem and are applicable to small datasets. Three methods based on binary classifications: one-against-all (OAA), one-against-one (OAO), and directed acyclic graph (DAG), and two approaches for multi-class problem by solving one single optimization problem, are implemented to predict protein secondary structure. Our experiments indicate that multi-class SVM methods are more suitable for protein secondary structure (PSS) prediction than the other methods, including binary SVMs, because their capacity to solve an optimization problem in one step. Furthermore, in this paper, we argue that it is feasible to extend the prediction accuracy by adding a second-stage multi-class SVM to capture the contextual information among secondary structural elements and thereby further improving the accuracies. We demonstrate that two-stage SVMs perform better than single-stage SVM techniques for PSS prediction using two datasets and report a maximum accuracy of 79.5%.

Computational Biology↗

A novel method for protein secondary structure prediction using dual-layer SVM and profiles.

A high-performance method was developed for protein secondary structure prediction based on the dual-layer support vector machine (SVM) and position-specific scoring matrices (PSSMs). SVM is a new machine learning technology that has been successfully applied in solving problems in the field of bioinformatics. The SVM's performance is usually better than that of traditional machine learning approaches. The performance was further improved by combining PSSM profiles with the SVM analysis. The PSSMs were generated from PSI-BLAST profiles, which contain important evolution information. The final prediction results were generated from the second SVM layer output. On the CB513 data set, the three-state overall per-residue accuracy, Q3, reached 75.2%, while segment overlap (SOV) accuracy increased to 80.0%. On the CB396 data set, the Q3 of our method reached 74.0% and the SOV reached 78.1%. A web server utilizing the method has been constructed and is available at http://www.bioinfo.tsinghua.edu.cn/pmsvm.

Computational Biology↗

Prediction of protein secondary structure by combining nearest-neighbor algorithms and multiple sequence alignments.

Recently Yi & Lander used a neural network and nearest-neighbor method with a scoring system that combined a sequence-similarity matrix with the local structural environment scoring scheme described by Bowie and co-workers for predicting protein secondary structure. We have improved their scoring system by taking into consideration N and C-terminal positions of alpha-helices and beta-strands and also beta-turns as distinctive types of secondary structure. Another improvement, which also decreases the time of computation, is performed by restricting a data base with a smaller subset of proteins that are similar with a query sequence. Using multiple sequence alignments rather than single sequences and a simple jury decision procedure our method reaches a sustained overall three-state accuracy of 72.2%, which is better than that observed for the most accurate multilayered neural-network approach, tested on the same data set of 126 non-homologous protein chains.

Algorithms↗

Predicting protein secondary structure and solvent accessibility with an improved multiple linear regression method.

We have improved the multiple linear regression (MLR) algorithm for protein secondary structure prediction by combining it with the evolutionary information provided by multiple sequence alignment of PSI-BLAST. On the CB513 dataset, the three states average overall per-residue accuracy, Q(3), reached 76.4%, while segment overlap accuracy, SOV99, reached 73.2%, using a rigorous jackknife procedure and the strictest reduction of eight states DSSP definition to three states. This represents an improvement of approximately 5% on overall per-residue accuracy compared with previous work. The relative solvent accessibility prediction also benefited from this combination of methods. The system achieved 77.7% average jackknifed accuracy for two states prediction based on a 25% relative solvent accessibility mode, with a Mathews' correlation coefficient of 0.548. The improved MLR secondary structure and relative solvent accessibility prediction server is available at http://spg.biosci.tsinghua.edu.cn/.

Algorithms↗

Effect of ethanol on the protein secondary structure of the human gastric mucosa, in vitro.

The effect of ethanol on the secondary conformational structure of proteins of the human gastric mucosa was investigated by attenuated total reflection/Fourier transform infrared (ATR/FT-IR) spectroscopy. The IR peak intensity and position of each structural component of gastric mucosa was found to change significantly with the ethanol concentration and length of exposure. The peak intensity due to the beta-sheet and/or beta-turn conformational structure in amide I and II bands of gastric mucosa clearly increased after treatment with ethanol. Moreover, the peak at 1635 cm-1 shifted to 1630 cm-1 after treatment with 40% ethanol for 3 h, or 80% ethanol for 1 h, and a distinct shoulder also appeared at 1643 cm-1. This shift occurred more rapidly and was more pronounced after exposure of mucosa to 80% ethanol, compared with the effect of 40% ethanol, but the alpha-helical structure at the amide I and II bands was not influenced by either concentration of ethanol. Ethanol treatment might also transform the secondary structure of amide III in gastric mucosa from an alpha-helix to a mainly random coil with extensive unfolding. The absorption between 1180 and 980 cm-1, which is assigned to glycoprotein structure, was also reduced after treatment with ethanol. This strongly indicates that ethanol influences the conformation of the lipids and proteins of human gastric mucosa, leading to their deformation.

Ethanol↗

Accurate prediction of protein secondary structural content.

An improved multiple linear regression (MLR) method is proposed to predict a protein's secondary structural content based on its primary sequence. The amino acid composition, the autocorrelation function, and the interaction function of side-chain mass derived from the primary sequence are taken into account. The average absolute errors of prediction over 704 unrelated proteins with the jackknife test are 0.088, 0.081, and 0.059 with standard deviations 0.073, 0.066, and 0.055 for alpha-helix, beta-sheet, and coil, respectively. That the sum of predicted secondary structure content should be close to 1.0 was introduced as a criterion to evaluate whether the prediction is acceptable. While only the predictions with the sum of predicted secondary structure content between 0.99 and 1.01 are accepted (about 11% of all proteins), the absolute errors are 0.058 for alpha-helix, 0.054 for beta-sheet, and 0.045 for coil.

Algorithms↗

Electronic and vibrational second-order nonlinear optical properties of protein secondary structural motifs.

A perturbation theory approach was developed for predicting the vibrational and electronic second-order nonlinear optical (NLO) polarizabilities of materials and macromolecules comprised of many coupled chromophores, with an emphasis on common protein secondary structural motifs. The polarization-dependent NLO properties of electronic and vibrational transitions in assemblies of amide chromophores comprising the polypeptide backbones of proteins were found to be accurately recovered in quantum chemical calculations by treating the coupling between adjacent oscillators perturbatively. A novel diagrammatic approach was developed to provide an intuitive visual means of interpreting the results of the perturbation theory calculations. Using this approach, the chiral and achiral polarization-dependent electronic SHG, isotropic SFG, and vibrational SFG nonlinear optical activities of protein structures were predicted and interpreted within the context of simple orientational models.

Algorithms↗

Clustering of amino acids for protein secondary structure prediction.

Simple hidden Markov models are proposed for predicting secondary structure of a protein from its amino acid sequence. Since the length of protein conformation segments varies in a narrow range, we ignore the duration effect of length distribution, and focus on inclusion of short range correlations of residues and of conformation states in the models. Conformation-independent and -dependent amino acid coarse-graining schemes are designed for the models by means of proper mutual information. We compare models of different level of complexity, and establish a practical model with a high prediction accuracy.

Algorithms↗

Prediction of protein secondary structure by neural networks: encoding short and long range patterns of amino acid packing.

A complex, cascaded neural network designed to predict the secondary structure of globular proteins has been developed. Information about the local buried-unburied pattern and the average tendency of the particular types of amino acids to be buried inside the globule were used. Nonspecific information about long distance contact maps was also employed. These modifications result in a noticeable improvement (3-9%) of prediction accuracy. The best result for the average success ratio for the testing set of nonhomologous proteins was 68.3% (with corresponding Matthews' coefficients, C alpha,beta,coil equal to 0.60, 0.47, 0.43, respectively).

Amino Acid Sequence↗

Redefining the goals of protein secondary structure prediction.

Secondary structure prediction recently has surpassed the 70% level of average accuracy, evaluated on the single residue states helix, strand and loop (Q3). But the ultimate goal is reliable prediction of tertiary (three-dimensional, 3D) structure, not 100% single residue accuracy for secondary structure. A comparison of pairs of structurally homologous proteins with divergent sequences reveals that considerable variation in the position and length of secondary structure segments can be accommodated within the same 3D fold. It is therefore sufficient to predict the approximate location of helix, strand, turn and loop segments, provided they are compatible with the formation of 3D structure. Accordingly, we define here a measure of segment overlap (Sov) that is somewhat insensitive to small variations in secondary structure assignments. The new segment overlap measure ranges from an ignorance level of 37% (random protein pairs) via a current level of 72% for a prediction method based on sequence profile input to neural networks (PHD) to an average 90% level for homologous protein pairs. We conclude that the highest scores one can reasonably expect for secondary structure prediction are a single residue accuracy of Q3 > 85% and a fractional segment overlap of Sov > 90%.

Amino Acid Sequence↗

Generation of a substructure library for the description and classification of protein secondary structure. II. Application to spectra-structure correlations in Fourier transform infrared spectroscopy.

Fourier transform infrared spectroscopy has become well known as a sensitive and informative tool for studying secondary structure in proteins. Present analysis of the conformation-sensitive amide I region in protein infrared spectra, when combined with band narrowing techniques, provides more information concerning protein secondary structure than can be meaningfully interpreted. This is due in part to limited models for secondary structure. Using the algorithm described in the previous paper of this series, we have generated a library of substructures for several trypsin-like serine proteases. This library was used as a basis for spectra-structure correlations with infrared spectra in the amide I' region, for five homologous proteins for which spectra were collected. Use of the substructure library has allowed correlations not previously possible with template-based methods of protein conformational analysis.

Algorithms↗

G-factor analysis of protein secondary structure in solutions and thin films.

The biological activity of proteins is structure dependent. In this discussion, we describe development of g-factor analysis for characterizing the secondary structure of proteins in solutions and films. In g-factor analysis, experimental circular dichroism (CD) and UV absorbance spectra are converted to dimensionless g-factor spectra by dividing the differential absorbance of circularly polarized light (AL - AR) by the UV absorbance (A) at each wavelength. The spectra can be deconvolved and the secondary structure estimated without information on the protein molecular weight, sample concentration, or sample path length. We have refined g-factor spectral acquisition and deconvolution parameters to improve the precision and accuracy of experimental g-factor spectra and the subsequent secondary structure estimations. In general, slower scan rates and longer response times are required to obtain high quality g-factor spectra than to obtain ordinary CD spectra, particularly for data at wavelengths longer than 230 nm. The spectral acquisition and deconvolution procedures have been validated for aqueous bovine serum albumin (BSA) and poly(L-proline). The secondary structures of fibronectin and laminin in buffer solutions are predicted. When BSA and poly(L-proline) form films on quartz, their secondary structures change significantly: 13% for BSA and 32% for poly(L-proline). By contrast, the secondary structure of fibronectin is the same in solution and films. The g-factor method is an easy, rapid, accurate and precise method for determining secondary structure and structural changes in protein solutions and films. Potential applications range from proteomics and structure-based drug discovery, to the design and fabrication of biosensors, biomaterials and biofluidic devices.

Algorithms↗

Prediction of protein secondary structure from circular dichroism spectra: an attempt to solve the problem of the best-fitting reference protein subsets.

In least-squares fitting of protein circular dichroism (CD) spectra using basis CD spectra for the respective secondary structure components, as given by reference proteins of known structural composition, good fits of the CD spectrum do not necessarily correspond to appropriate fits of the underlying structural composition of a test protein. In an attempt to overcome this problem, CD similarity measures were used to construct a subset of five reference CD spectra which permitted three-component fitting of the CD and prediction of the relative magnitude of total beta-sheet (antiparallel + parallel) and total other structure (beta-turn + remainder), relative to helix, in the test protein. A backpropagation neural network (BPN) was also trained to make this prediction. In subsequent five-component fitting of the CD spectrum, using a Monte Carlo method to generate subsets of reference proteins from the working data set, only those secondary structure fits which conformed to the consensus prediction of the CD similarity measures and the BPN were accepted. The method enhanced the fitting of antiparallel beta-sheet and beta-turn for 16 proteins, compared to the variable selection method of P. Manavalan and W. C. Johnson (1987, Anal. Biochem. 167, 76-85). Some other proteins were less well fitted. Improvement in results is expected with a larger representation of feasible basis CD spectra in the working reference proteins.

Circular Dichroism↗

A study of protein secondary structure by Fourier transform infrared/photoacoustic spectroscopy and its application for recombinant proteins.

FT-IR/PAS (Fourier transform infrared/photoacoustic spectroscopy) was used to evaluate the secondary structure of proteins. Four well-studied proteins, concanavalin A, hemoglobin, lysozyme, and trypsin, which have different distributions of secondary structures, were used for assignments of the infrared bands and evaluating the accuracy of FT-IR/PAS methods. Secondary structure contents estimated from FT-IR/PAS and other physical methods (e.g., X-ray diffraction, CD, and traditional FT-IR) show good agreement. In addition, the secondary structure can be evaluated with as little as 0.5 micrograms of protein (concanavalin A), suggesting that FT-IR/PAS is a sensitive and useful technique that could be applied to studies of the folding of recombinant and mutant proteins where only small amounts of material are available. Recombinant phosphorylase kinase gamma 1-300 subunit expressed in Escherichia coli was found in the inclusion bodies. We found that renatured phosphorylase kinase gamma 1-300 subunit has two kinase forms: one has a 10-fold higher activity than the other one. Both fractions, however, are the same as judged from sodium dodecylsulfate-polyacrylamide gel electrophoresis. Differences in conformation were demonstrated by using the FT-IR/PAS method, which showed that the low-activity form has more beta-sheet structure than the form with high activity. Analysis of these kinase forms by CD confirms the interpretation made by the FT-IR/PAS method.

Protein Structure, Secondary↗

A new approach to the evaluation of protein secondary structure predictions at the level of the elements of secondary structure.

For many purposes, such as the prediction of the class of protein folds, the existence of an element of secondary structure rather than its precise position and length must be defined correctly. However, most methods for the evaluation of secondary structure prediction consider success in terms of the percentage of individual amino acids predicted correctly. In this paper the success in predicting elements of secondary structure is discussed. The number of overlapping residues in the predicted and observed secondary structures were considered as a function of the total number of amino acids in the observed and predicted secondary structures. A matrix search procedure was used to remove the ambiguity which similar studies may have had in defining the equivalent secondary structures between predicted and observed structures. In this study a loop was treated in the same way as an alpha-helix and a beta-strand. To describe the accuracy at the level of elements of secondary structure, a set of parameters was defined, similar to those used commonly at the level of individual amino acids. This approach was used to assess the methods of Chou and Fasman (1974b, Biochemistry, 13, 222-245), Lim (1974b, J. Mol. Biol., 88, 873-894) and Garnier et al. (1978, J. Mol. Biol., 120, 97-120). It was found that these methods were much poorer at the secondary structure level than at the amino acid level. This approach can be used generally for secondary structure prediction methods.

Amino Acids↗

The future of protein secondary structure prediction accuracy.

BACKGROUND: The accuracy of secondary structure prediction for a protein from knowledge of its sequence has been significantly improved by about 7% to the 70-75% range by inclusion of information residing in sequences similar to the query sequence. The scientific literature has been inconsistent, if not negative, regarding chances for further improvement from the vast knowledge to be provided by genome sequencing efforts. RESULTS: By applying a prediction technique that is particularly sensitive to added sequence information to a standard set of query sequences with related primary structures taken from chronologically successive releases of the SWISS-PROT database, it is shown that prediction accuracy can be expected to reach 80-85% with a large 10-fold increase in present sequence knowledge. CONCLUSIONS: Even with present prediction approaches, improvement in prediction accuracy can still be expected, albeit limited to no more than 10%.

Algorithms↗

Accurate and automated classification of protein secondary structure with PsiCSI.

PsiCSI is a highly accurate and automated method of assigning secondary structure from NMR data, which is a useful intermediate step in the determination of tertiary structures. The method combines information from chemical shifts and protein sequence using three layers of neural networks. Training and testing was performed on a suite of 92 proteins (9437 residues) with known secondary and tertiary structure. Using a stringent cross-validation procedure in which the target and homologous proteins were removed from the databases used for training the neural networks, an average 89% Q3 accuracy (per residue) was observed. This is an increase of 6.2% and 5.5% (representing 36% and 33% fewer errors) over methods that use chemical shifts (CSI) or sequence information (Psipred) alone. In addition, PsiCSI improves upon the translation of chemical shift information to secondary structure (Q3 = 87.4%) and is able to use sequence information as an effective substitute for sparse NMR data (Q3 = 86.9% without (13)C shifts and Q3 = 86.8% with only H(alpha) shifts available). Finally, errors made by PsiCSI almost exclusively involve the interchange of helix or strand with coil and not helix with strand (<2.5 occurrences per 10000 residues). The automation, increased accuracy, absence of gross errors, and robustness with regards to sparse data make PsiCSI ideal for high-throughput applications, and should improve the effectiveness of hybrid NMR/de novo structure determination methods. A Web server is available for users to submit data and have the assignment returned.

Automation↗