[Protein secondary structure prediction].
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to Akira R Kinjo.
Explore the source record for details and available documents.
BACKGROUND: One-dimensional protein structures such as secondary structures or contact numbers are useful for three-dimensional structure prediction and helpful for intuitive understanding of the sequence-structure relationship. Accurate prediction methods will serve as a basis for these and other purposes. RESULTS: We implemented a program CRNPRED which predicts secondary structures, contact numbers and residue-wise contact orders. This program is based on a novel machine learning scheme called critical random networks. Unlike most conventional one-dimensional structure prediction methods which are based on local windows of an amino acid sequence, CRNPRED takes into account the whole sequence. CRNPRED achieves, on average per chain, Q3 = 81% for secondary structure prediction, and correlation coefficients of 0.75 and 0.61 for contact number and residue-wise contact order predictions, respectively. CONCLUSION: CRNPRED will be a useful tool for computational as well as experimental biologists who need accurate one-dimensional protein structure predictions.
Human transcriptional regulation factors, such as activators, repressors, and enhancer-binding factors are quite different from their prokaryotic counterparts in two respects: the average sequence in human is more than twice as long as that in prokaryotes, while the fraction of sequence aligned to domains of known structure is 31% in human transcription factors (TFs), less than half of that in bacterial TFs (72%). Intrinsically disordered (ID) regions were identified by a disorder-prediction program, and were found to be in good agreement with available experimental data. Analysis of 401 human TFs with experimental evidence from the Swiss-Prot database showed that as high as 49% of the entire sequence of human TFs is occupied by ID regions. More than half of the human TFs consist of a small DNA binding domain (DBD) and long ID regions frequently sandwiching unassigned regions. The remaining TFs have structural domains in addition to DBDs and ID regions. Experimental studies, particularly those with NMR, revealed that the transactivation domains in unbound TFs are usually unstructured, but become structured upon binding to their partners. The sequences of human and mouse TF orthologues are 90.5% identical despite a high incidence of ID regions, probably reflecting important functional roles played by ID regions. In general ID regions occupy a high fraction in TFs of eukaryotes, but not in prokaryotes. Implications of this dichotomy are discussed in connection with their functional roles in transcriptional regulation and evolution.
One-dimensional (1D) structures of proteins such as secondary structure and contact number provide intuitive pictures to understand how the native three-dimensional (3D) structure of a protein is encoded in the amino acid sequence. However, it is still not clear whether a given set of 1D structures contains sufficient information for recovering the underlying 3D structure. Here we show that the 3D structure of a protein can be recovered from a set of three types of 1D structures, namely, secondary structure, contact number and residue-wise contact order which is introduced here for the first time. Using simulated annealing molecular dynamics simulations, the structures satisfying the given native 1D structural restraints were sought for 16 proteins of various structural classes and of sizes ranging from 56 to 146 residues. By selecting the structures best satisfying the restraints, all the proteins showed a coordinate RMS deviation of <4 A from the native structure, and, for most of them, the deviation was even <2 A. The present result opens a new possibility to protein structure prediction and our understanding of the sequence-structure relationship.
The contact number of an amino acid residue in a protein structure is defined by the number of C(beta) atoms around the C(beta) atom of the given residue, a quantity similar to, but different from, solvent accessible surface area. We present a method to predict the contact numbers of a protein from its amino acid sequence. The method is based on a simple linear regression scheme and predicts the absolute values of contact numbers. When single sequences are used for both parameter estimation and cross-validation, the present method predicts the contact numbers with a correlation coefficient of 0.555 on average. When multiple sequence alignments are used, the correlation increases to 0.627, which is a significant improvement over previous methods. In terms of discrete states prediction, the accuracies for 2-, 3-, and 10-state predictions are, respectively, 71.4%, 54.1%, and 18.9% with residue type-dependent unbiased thresholds, and 76.3%, 59.2%, and 21.8% with residue type-independent unbiased thresholds. The difference between accessible surface area and contact number from a prediction viewpoint and the application of contact number prediction to three-dimensional structure prediction are discussed.
The pattern of amino acid substitutions and sequence conservation over many structure-based alignments of protein sequences was analyzed as a function of percentage sequence identity. The statistics of the amino acid substitutions were converted into the form of log-odds amino acid substitution matrices to which eigenvalue decomposition was applied. It was found that the most important component of the substitution matrices exhibited a sharp transition at the sequence identity of 30-35%, which coincides with the twilight zone. Above the transition point, the most dominant component is related to the mutability of amino acids and it acts to disfavor any substitutions, whereas below the transition point, the most dominant component is related to the hydrophobicity of amino acids and substitutions between residues of similar hydrophobic character are positively favored. Implications for protein evolution and sequence analysis are discussed.
The living cell is inherently crowded with proteins and macromolecules. To avoid aggregation of denatured proteins in the living cell, molecular chaperones play important roles. Here we introduce a simple model to describe crowded protein solutions with chaperone-like species based on a dynamic density functional theory. As predicted by others, our simulations show that macromolecular crowding enhances the association of proteins and chaperones. However, when the intrinsic folding rate of the protein is slow, it is possible that crowding also enhances aggregation of proteins. The results of simulation suggest that, when the concentration of the crowding agent is as high as that in the cell, the association of the protein and unbound chaperone becomes correlated with the aggregation process, and that the protein-bound chaperones efficiently destroy the potential nuclei of aggregates and thus prevent the aggregation.
Inside the living cell is inherently crowded with proteins and other macromolecules. Thus, it is indispensable to take into account various interactions between the protein and other macromolecules for thorough understanding of protein functions in cellular contexts. Here we focus on the excluded volume interaction imposed on the protein by surrounding macromolecules or "crowding agents." We have presented a theoretical framework for describing equilibrium properties of proteins in crowded solutions [A. R. Kinjo and S. Takada, Phys. Rev. E (to be published)]. In the present paper, we extend the theory to describe nonequilibrium properties of proteins in crowded solutions. Dynamics simulations exhibit qualitatively different morphologies depending on the aggregating conditions, and it was found that macromolecular crowding accelerates the onset of aggregation while stabilizing the native protein in the quasiuniform phase before the onset of aggregation. It is also observed, however, that the aggregation may be kinetically inhibited in highly crowded conditions. The effects of crowding on folding and unfolding of proteins are also examined, and the results suggest that fast folding is an important factor in preventing aggregation of denatured proteins.
Proteins are neither purified nor diluted inside the living cell. Thus it is indispensable to take into account various interactions between the protein of interest and other macromolecules for understanding the properties of proteins in physiological conditions. Here we focus on excluded volume interactions which are omnipresent in dense or crowded solutions of proteins and macromolecules or "crowding agents." A protein solution with macromolecular crowding agents is modeled by means of a density functional theory. Effects of macromolecular crowding on protein aggregation and stability are investigated in particular. Phase diagrams are obtained in various parameter spaces by solving the equation of state. Two generic features are found: the addition of the crowding agent (1) enhances the aggregation of the denatured proteins, and (2) stabilizes the native protein unless the aggregation occurs. The present theory is qualitatively in good agreement with experimental observations and unifies previous theories regarding the crowding effects on protein stability and aggregation.