PubMed Health⌕ Search

Biomedical subjects

Akinori Sarai

Publications and source records attributed to Akinori Sarai.

At least 19 recordsLinked to original sources

Atom-wise statistics and prediction of solvent accessibility in proteins.

In this work, we explore a novel method to broaden the scope of sequence-based predictions of solvent accessibility or accessible surface area (ASA) to the atomic level. All 167 heavy atoms from the 20 types of amino acid residues in proteins have been studied. An analysis of ASA distribution of these atomic groups in different proteins has been performed and rotamer-style libraries have been developed. We observe that the ASA of some atomic groups (e.g., backbone C and N atoms) can be estimated from the sequence environment within a mean absolute error of 2-3 angstroms(2). However, some side chain atoms such as CG in Pro, NH1 in Arg and NE2 in Gln show a strong variability making it more difficult to estimate their ASA from sequence environment. In general, the prediction of ASA becomes more difficult for atomic positions at the side chain extremities of long amino acid residues (aromatic side chain terminals being the exception). Several atomic groups are frequently exposed to solvent. Some of them have a bimodal distribution, suggesting two stable conformations in terms of their solvent exposure. More detailed understanding and prediction of solvent accessibility, i.e., at an atomic level is expected to help in bioinformatics approaches to structure prediction, functional relevance of atomic solvent accessibilities and other interaction analyses.

Amino Acids↗

ReadOut: structure-based calculation of direct and indirect readout energies and specificities for protein-DNA recognition.

Protein-DNA interactions play a central role in regulatory processes at the genetic level. DNA-binding proteins recognize their targets by direct base-amino acid interactions and indirect conformational energy contribution from DNA deformations and elasticity. Knowledge-based approach based on the statistical analysis of protein-DNA complex structures has been successfully used to calculate interaction energies and specificities of direct and indirect readouts in protein-DNA recognition. Here, we have implemented the method as a webserver, which calculates direct and indirect readout energies and Z-scores, as a measure of specificity, using atomic coordinates of protein-DNA complexes. This server is freely available at http://gibk26.bse.kyutech.ac.jp/jouhou/readout/. The only input to this webserver is the Protein Data Bank (PDB) style coordinate data of atoms or the PDB code itself. The server returns total energy Z-scores, which estimate the degree of sequence specificity of the protein-DNA complex. This webserver is expected to be useful for estimating interaction energy and DNA conformation energy, and relative contributions to the specificity from direct and indirect readout. It may also be useful for checking the quality of protein-DNA complex structures, and for engineering proteins and target DNAs.

DNA↗

Dimensionality of amino acid space and solvent accessibility prediction with neural networks.

Solvent accessibility prediction from amino acid sequences has been pursued by several researchers. Such a prediction typically starts by transforming the amino acid category (or type) information into numerical representations. All twenty amino acids can be completely and uniquely represented by 20-dimensional vectors. Here, we investigate if the amino acid space defined in this way really requires twenty dimensions. We tried to develop corresponding representations in fewer dimensions. A method for searching optimal codification schema in an arbitrary space using neural networks was developed. The method is used to obtain optimal encoding of amino acids at various levels of dimensionality, and applied to optimize the amino acid codifications for the prediction of the solvent accessibility values of the proteins using feed-forward neural networks. The traditional 20-dimensional codification seems to be redundant in solving the solvent accessibility prediction problem, since a 1-dimensional codification is able to achieve almost the same degree of accuracy as the 20-dimensional codification. Optimal coding in much fewer dimensions could be used to make the predictions of accessible surface area with almost the same degree of accuracy as that obtained by a fully unique 20-dimensional coding. The 1-dimensional amino acid codification for solvent accessibility prediction obtained by a purely mathematical way based on neural networks is highly correlated with a physical property of the amino acids, namely their average solvent accessibility. The method developed to find the optimal codification is general, although the codification thus produced is dependent on the type of estimated property.

Amino Acid Sequence↗

ProTherm and ProNIT: thermodynamic databases for proteins and protein-nucleic acid interactions.

ProTherm and ProNIT are two thermodynamic databases that contain experimentally determined thermodynamic parameters of protein stability and protein-nucleic acid interactions, respectively. The current versions of both the databases have considerably increased the total number of entries and enhanced search interface with added new fields, improved search, display and sorting options. As on September 2005, ProTherm release 5.0 contains 17,113 entries from 771 proteins, retrieved from 1497 scientific articles (approximately 20% increase in data from the previous version). ProNIT release 2.0 contains 4900 entries from 273 research articles, representing 158 proteins. Both databases can be queried using WWW interfaces. Both quick search and advanced search are provided on this web page to facilitate easy retrieval and display of the data from these databases. ProTherm is freely available online at http://gibk26.bse.kyutech.ac.jp/jouhou/Protherm/protherm.html and ProNIT at http://gibk26.bse.kyutech.ac.jp/jouhou/pronit/pronit.html.

DNA↗

Classification of protein-DNA complexes based on structural descriptors.

We attempt to classify protein-DNA complexes by using a set of 11 descriptors, mainly characterizing protein-DNA interactions, including the number of atomic contacts at major and minor grooves, conformational deviations from standard B- and A-DNA forms, widths of DNA grooves, GC content, specificity measures of direct and indirect readouts, and buried surface area at the complex interface. The cluster analyses were carried out for a unique set of 62 complexes including a variety of protein motifs, and 7 distinct clusters were revealed from the analyses. We found that some proteins with the same motif are classified into different clusters, whereas different proteins with distinct motifs are classified into the same cluster. These results suggest that the conventional motif-based classification of DNA binding proteins may not necessarily correspond to structural and functional properties of protein-DNA complexes, and that the present classification will help to identify common properties and rules that govern protein-DNA recognition.

Cluster Analysis↗

Sequence-dependent conformational energy of DNA derived from molecular dynamics simulations: toward understanding the indirect readout mechanism in protein-DNA recognition.

Sequence dependence of DNA conformation plays a crucial role in its recognition by proteins and ligands. To clarify the relationship between sequence and conformation, it is necessary to quantify the conformational energy and specificity of DNA. Here, we make a systematic analysis of dodecamer DNA structures including all the 136 unique tetranucleotide sequences at the center by molecular dynamics simulations. Using a simplified conformational model with six parameters to describe the geometry of adjacent base pairs and harmonic potentials along these coordinates, we estimated the equilibrium conformational parameters and the harmonic potentials of mean force for the central base-pair steps from many trajectories of the simulations. This enabled us to estimate the conformational energy and the specificity for any given DNA sequence and structure. We tested our method by using sequence-structure threading to estimate the conformational energy and the Z-score as a measure of specificity for many B-DNA and A-DNA crystal structures. The average Z-scores were negative for both kinds of structures, indicating that the potential of mean force from the simulation is capable of predicting sequence specificity for the crystal structures and that it may be used to study the sequence specificity of both types of DNA. We also estimated the positional distribution of conformational energy and Z-score within DNA and showed that they are strongly position dependent. This analysis enabled us to identify particular conformations responsible for the specificity. The presented results will provide an insight into the mechanisms of DNA sequence recognition by proteins and ligands.

Base Sequence↗

Role of inter and intramolecular interactions in protein-DNA recognition.

Protein-DNA recognition plays an essential role in the regulation of gene expression. Regulatory proteins are known to recognize specific DNA sequences directly through atomic contacts between protein and DNA, and/or indirectly through the conformational properties of the DNA. In this work, we have analyzed the specificity of intermolecular interactions by statistical analysis of base-amino acid interactions within protein-DNA complexes as well as the computer simulations of base-amino acid interactions. The specificity of the intramolecular interactions was studied by statistical analysis of the sequence-dependent DNA conformational parameters and the elastic properties of DNA. Systematic comparison of these specificities in a large number of protein-DNA complexes revealed that both intermolecular and intramolecular interactions contribute to the specificity of protein-DNA recognition, and their relative contributions vary depending upon the protein-DNA complex. We demonstrated that combination of the intermolecular and intramolecular energies leads to enhanced specificity and the combined energy could explain experimental data on binding affinity changes caused by base mutations. These results provided new insight into the relationship between specificity and structure in the process of protein-DNA recognition, which would lead to prediction of specific protein-DNA binding sites.

Base Sequence↗

The Diamond STING server.

Diamond STING is a new version of the STING suite of programs for a comprehensive analysis of a relationship between protein sequence, structure, function and stability. We have added a number of new functionalities by both providing more structure parameters to the STING Database and by improving/expanding the interface for enhanced data handling. The integration among the STING components has also been improved. A new key feature is the ability of the STING server to handle local files containing protein structures (either modeled or not yet deposited to the Protein Data Bank) so that they can be used by the principal STING components: (Java)Protein Dossier ((J)PD) and STING Report. The current capabilities of the new STING version and a couple of biologically relevant applications are described here. We have provided an example where Diamond STING identifies the active site amino acids and folding essential amino acids (both previously determined by experiments) by filtering out all but those residues by selecting the numerical values/ranges for a set of corresponding parameters. This is the fundamental step toward a more interesting endeavor-the prediction of such residues. Diamond STING is freely accessible at http://sms.cbi.cnptia.embrapa.br and http://trantor.bioc.columbia.edu/SMS.

Acid Anhydride Hydrolases↗

The role of alternative translation start sites in the generation of human protein diversity.

According to the scanning model, 40S ribosomal subunits initiate translation at the first (5' proximal) AUG codon they encounter. However, if the first AUG is in a suboptimal context, it may not be recognized, and translation can then initiate at downstream AUG(s). In this way, a single RNA can produce several variant products. Earlier experiments suggested that some of these additional protein variants might be functionally important. We have analysed human mRNAs that have AUG triplets in 5' untranslated regions and mRNAs in which the annotated translational start codon is located in a suboptimal context. It was found that 3% of human mRNAs have the potential to encode N-terminally extended variants of the annotated proteins and 12% could code for N-truncated variants. The predicted subcellular localizations of these protein variants were compared: 31% of the N-extended proteins and 30% of the N-truncated proteins were predicted to localize to subcellular compartments that differed from those targeted by the annotated protein forms. These results suggest that additional AUGs may frequently be exploited for the synthesis of proteins that possess novel functional properties.

5' Untranslated Regions↗

PSSM-based prediction of DNA binding sites in proteins.

BACKGROUND: Detection of DNA-binding sites in proteins is of enormous interest for technologies targeting gene regulation and manipulation. We have previously shown that a residue and its sequence neighbor information can be used to predict DNA-binding candidates in a protein sequence. This sequence-based prediction method is applicable even if no sequence homology with a previously known DNA-binding protein is observed. Here we implement a neural network based algorithm to utilize evolutionary information of amino acid sequences in terms of their position specific scoring matrices (PSSMs) for a better prediction of DNA-binding sites. RESULTS: An average of sensitivity and specificity using PSSMs is up to 8.7% better than the prediction with sequence information only. Much smaller data sets could be used to generate PSSM with minimal loss of prediction accuracy. CONCLUSION: One problem in using PSSM-derived prediction is obtaining lengthy and time-consuming alignments against large sequence databases. In order to speed up the process of generating PSSMs, we tried to use different reference data sets (sequence space) against which a target protein is scanned for PSI-BLAST iterations. We find that a very small set of proteins can actually be used as such a reference data without losing much of the prediction value. This makes the process of generating PSSMs very rapid and even amenable to be used at a genome level. A web server has been developed to provide these predictions of DNA-binding sites for any new protein from its amino acid sequence. AVAILABILITY: Online predictions based on this method are available at http://www.netasa.org/dbs-pssm/

Algorithms↗

Integration of bioinformatics and computational biology to understand protein-DNA recognition mechanism.

Transcription factors play essential role in the gene regulation in higher organisms, binding to multiple target sequences and regulating multiple genes in a complex manner. In order to decipher the mechanism of gene regulation, it is important to understand the molecular mechanism of protein-DNA recognition. Here we describe a strategy to approach this problem, using various methods in bioinformatics and computational biology. We have used a knowledge-based approach, utilizing rapidly increasing structural data of protein-DNA complexes, to derive empirical potential functions for the specific interactions between bases and amino acids as well as for DNA conformation, from the statistical analyses on the structural data. Then these statistical potentials are used to quantify the specificity of protein-DNA recognition. The quantification of specificity has enabled us to establish the structure-function analysis of transcription factors, such as the effects of binding cooperativity on target recognition. The method is also applied to real genome sequences, predicting potential target sites. We are also using computer simulations of protein-DNA interactions and DNA conformation in order to complement the empirical method. The integration of these approaches together will provide deeper insight into the mechanism of protein-DNA recognition and improve the target prediction of transcription factors.

Binding Sites↗

Protein-DNA recognition patterns and predictions.

Structural data on protein-DNA complexes provide clues for understanding the mechanism of protein-DNA recognition. Although the structures of a large number of protein-DNA complexes are known, the mechanisms underlying their specific binding are still only poorly understood. Analysis of these structures has shown that there is no simple one-to-one correspondence between bases and amino acids within protein-DNA complexes; nevertheless, the observed patterns of interaction carry important information on the mechanisms of protein-DNA recognition. In this review, we show how the patterns of interaction, either observed in known structures or derived from computer simulations, confer recognition specificity, and how they can be used to examine the relationship between structure and specificity and to predict target DNA sequences used by regulatory proteins.

Adenine↗

Knowledge-based prediction of DNA atomic structure from nucleic sequence.

A simple knowledge-based method for DNA atomic structure prediction from nucleic sequence is presented. We used free B-DNA crystal structures to estimate the distribution of trinucleotide base pairs and tetranucleotide base-pair steps conformational coordinates. We used these distributions as a basis to predict the 3D position of the non-hydrogen atoms of the nucleic bases of any arbitrary DNA sequence of any length. The only constraint imposed was that the structure is a B-DNA one with Watson-Crick complementary base pairs. The method was tested on not seen DNA structures with sequence lengths varying from 6bp to 12bp. The obtained predictions have RMSE around 0.5 A for the translational conformational coordinates, and around 5 degrees for the rotational. For the estimation of the nucleic base non-hydrogen atom coordinates the RMSE is around 1.1 A. The knowledge-based method outperformed a technique based on genetic algorithms in the prediction of B-DNA structures.

Base Sequence↗

Look-up tables for protein solvent accessibility prediction and nearest neighbor effect analysis.

We developed dictionaries of two-, three-, and five-residue patterns in proteins and computed the average solvent accessibility of the central residues in their native proteins. These dictionaries serve as a look-up table for making subsequent predictions of solvent accessibility of amino acid residues. We find that predictions made in this way are very close to those made using more sophisticated methods of solvent accessibility prediction. We also analyzed the effect of immediate neighbors on the solvent accessibility of residues. This helps us in understanding how the same residue type may have different accessible surface areas in different proteins and in different positions of the same protein. We observe that certain residues have a tendency to increase or decrease the solvent accessibility of their neighboring residues in C- or N-terminal positions. Interestingly, the C-terminal and N-terminal neighbor residues are found to have asymmetric roles in modifying solvent accessibility of residues. As expected, similar neighbors enhance the hydrophobic or hydrophilic character of residues. Detailed look-up tables are provided on the web at www.netasa.org/look-up/.

Amino Acid Sequence↗

Bound peptide-dependent thermal stability of major histocompatibility complex class II molecule I-Ek.

We used differential scanning calorimetry to study the thermal denaturation of murine major histocompatibility complex class II, I-E(k), accommodating hemoglobin (Hb) peptide mutants possessing a single amino acid substitution of the chemically conserved amino acids buried in the I-Ek pocket (positions 71 and 73) and exposed to the solvent (position 72). All of the I-Ek-Hb(mut) molecules exhibited greater thermal stability at pH 5.5 than at pH 7.4, as for the I-Ek-Hb(wt) molecule, which can explain the peptide exchange function of MHC II. The thermal stability was strongly dependent on the bound peptide sequences; the I-Ek-Hb(mut) molecules were less stable than the I-Ek-Hb(wt) molecules, in good correlation with the relative affinity of each peptide for I-Ek. This supports the notion that the bound peptide is part of the completely folded MHC II molecule. The thermodynamic parameters for I-Ek-Hb(mut) folding can explain the thermodynamic origin of the stability difference, in correlation with the crystal structural analysis, and the limited contributions of the residues to the overall conformation of the I-Ek-peptide complex. We found a linear relationship between the denaturation temperature and the calorimetric enthalpy change. Thus, although the MHC II-peptide complex could have a diverse thermal stability spectrum, depending on the amino acid sequences of the bound peptides, the conformational perturbations are limited. The variations in the MHC II-peptide complex stability would function in antigen recognition by the T cell receptor by affecting the stability of the MHC II-peptide-T cell receptor ternary complex.

Animals↗

Moment-based prediction of DNA-binding proteins.

Net charge, electric dipole moment and quadrupole moment tensors were calculated for 78 amino acid sequences from 62 representative DNA-binding proteins with known structures. It was found that the magnitudes of the moments of electric charge distribution in these chains differ significantly from those of a non-binding control data set. Net charge, net dipole moment and quadrupole moment could each distinguish binding and non-binding proteins with 82.6%, 77.4% and 73.7% accuracy by single-variable predictors without cross-validation. Using hybrid predictors with information of charge and both moments, the best predictions were 85.6% without cross-validation and 83.9% for the cross-validated data sets. This level of prediction accuracy obtained with these simple descriptors competes with the results obtained using more complex models including many descriptors. The coarse graining of atomic charges onto C(alpha) atoms did not reduce the prediction accuracy significantly. This result suggests that we can use C(alpha) coordinates derived from homology modeling to predict DNA-binding proteins. The speed and accuracy of this method, in combination with homology-based methods of structure prediction, should enhance genome-wide recognition of DNA-binding proteins.

Animals↗

Qgrid: clustering tool for detecting charged and hydrophobic regions in proteins.

We have developed a simple but powerful method and web server to quickly locate charged and hydrophobic clusters in proteins (http://www.netasa.org/qgrid/index.html). For the charged clusters, each atom in the protein is first assigned a charge according to a standard force field. Then a box is created with dimensions corresponding to the range of atomic coordinates. This box is then divided into cubic grids of selected size, which now have one or more charged atoms in them. This leaves each grid with a certain amount of charge. Cubic grids with more than a cutoff charge are then clustered using a hierarchical clustering method based on Euclidean distance. A tree diagram made from the resulting clusters indicates the distribution of charged and hydrophobic regions of the protein. Hydrophobic clusters are developed by grouping the positions of C(alpha) atoms of such residues. We propose that such a tree representation will be helpful in detecting protein-protein interfaces, structure similarity and motif detection.

Hydrophobic and Hydrophilic Interactions↗

ASAView: database and tool for solvent accessibility representation in proteins.

BACKGROUND: Accessible surface area (ASA) or solvent accessibility of amino acids in a protein has important implications. Knowledge of surface residues helps in locating potential candidates of active sites. Therefore, a method to quickly see the surface residues in a two dimensional model would help to immediately understand the population of amino acid residues on the surface and in the inner core of the proteins. RESULTS: ASAView is an algorithm, an application and a database of schematic representations of solvent accessibility of amino acid residues within proteins. A characteristic two-dimensional spiral plot of solvent accessibility provides a convenient graphical view of residues in terms of their exposed surface areas. In addition, sequential plots in the form of bar charts are also provided. Online plots of the proteins included in the entire Protein Data Bank (PDB), are provided for the entire protein as well as their chains separately. CONCLUSIONS: These graphical plots of solvent accessibility are likely to provide a quick view of the overall topological distribution of residues in proteins. Chain-wise computation of solvent accessibility is also provided.

Computer Graphics↗