PubMed Health⌕ Search

Biomedical subjects

Shandar Ahmad

Publications and source records attributed to Shandar Ahmad.

At least 19 recordsLinked to original sources

Sequence and structural features of carbohydrate binding in proteins and assessment of predictability using a neural network.

BACKGROUND: Protein-Carbohydrate interactions are crucial in many biological processes with implications to drug targeting and gene expression. Nature of protein-carbohydrate interactions may be studied at individual residue level by analyzing local sequence and structure environments in binding regions in comparison to non-binding regions, which provide an inherent control for such analyses. With an ultimate aim of predicting binding sites from sequence and structure, overall statistics of binding regions needs to be compiled. Sequence-based predictions of binding sites have been successfully applied to DNA-binding proteins in our earlier works. We aim to apply similar analysis to carbohydrate binding proteins. However, due to a relatively much smaller region of proteins taking part in such interactions, the methodology and results are significantly different. A comparison of protein-carbohydrate complexes has also been made with other protein-ligand complexes. RESULTS: We have compiled statistics of amino acid compositions in binding versus non-binding regions- general as well as in each different secondary structure conformation. Binding propensities of each of the 20 residue types and their structure features such as solvent accessibility, packing density and secondary structure have been calculated to assess their predisposition to carbohydrate interactions. Finally, evolutionary profiles of amino acid sequences have been used to predict binding sites using a neural network. Another set of neural networks was trained using information from single sequences and the prediction performance from the evolutionary profiles and single sequences were compared. Best of the neural network based prediction could achieve an 87% sensitivity of prediction at 23% specificity for all carbohydrate-binding sites, using evolutionary information. Single sequences gave 68% sensitivity and 55% specificity for the same data set. Sensitivity and specificity for a limited galactose binding data set were obtained as 63% and 79% respectively for evolutionary information and 62% and 68% sensitivity and specificity for single sequences. Propensity and other sequence and structural features of carbohydrate binding sites have also been compared with our similar extensive studies on DNA-binding proteins and also with protein-ligand complexes. CONCLUSION: Carbohydrates typically show a preference to bind aromatic residues and most prominently tryptophan. Higher exposed surface area of binding sites indicates a role of hydrophobic interactions. Neural networks give a moderate success of prediction, which is expected to improve when structures of more protein-carbohydrate complexes become available in future.

Amino Acid Sequence↗

Atom-wise statistics and prediction of solvent accessibility in proteins.

In this work, we explore a novel method to broaden the scope of sequence-based predictions of solvent accessibility or accessible surface area (ASA) to the atomic level. All 167 heavy atoms from the 20 types of amino acid residues in proteins have been studied. An analysis of ASA distribution of these atomic groups in different proteins has been performed and rotamer-style libraries have been developed. We observe that the ASA of some atomic groups (e.g., backbone C and N atoms) can be estimated from the sequence environment within a mean absolute error of 2-3 angstroms(2). However, some side chain atoms such as CG in Pro, NH1 in Arg and NE2 in Gln show a strong variability making it more difficult to estimate their ASA from sequence environment. In general, the prediction of ASA becomes more difficult for atomic positions at the side chain extremities of long amino acid residues (aromatic side chain terminals being the exception). Several atomic groups are frequently exposed to solvent. Some of them have a bimodal distribution, suggesting two stable conformations in terms of their solvent exposure. More detailed understanding and prediction of solvent accessibility, i.e., at an atomic level is expected to help in bioinformatics approaches to structure prediction, functional relevance of atomic solvent accessibilities and other interaction analyses.

Amino Acids↗

ReadOut: structure-based calculation of direct and indirect readout energies and specificities for protein-DNA recognition.

Protein-DNA interactions play a central role in regulatory processes at the genetic level. DNA-binding proteins recognize their targets by direct base-amino acid interactions and indirect conformational energy contribution from DNA deformations and elasticity. Knowledge-based approach based on the statistical analysis of protein-DNA complex structures has been successfully used to calculate interaction energies and specificities of direct and indirect readouts in protein-DNA recognition. Here, we have implemented the method as a webserver, which calculates direct and indirect readout energies and Z-scores, as a measure of specificity, using atomic coordinates of protein-DNA complexes. This server is freely available at http://gibk26.bse.kyutech.ac.jp/jouhou/readout/. The only input to this webserver is the Protein Data Bank (PDB) style coordinate data of atoms or the PDB code itself. The server returns total energy Z-scores, which estimate the degree of sequence specificity of the protein-DNA complex. This webserver is expected to be useful for estimating interaction energy and DNA conformation energy, and relative contributions to the specificity from direct and indirect readout. It may also be useful for checking the quality of protein-DNA complex structures, and for engineering proteins and target DNAs.

DNA↗

Dimensionality of amino acid space and solvent accessibility prediction with neural networks.

Solvent accessibility prediction from amino acid sequences has been pursued by several researchers. Such a prediction typically starts by transforming the amino acid category (or type) information into numerical representations. All twenty amino acids can be completely and uniquely represented by 20-dimensional vectors. Here, we investigate if the amino acid space defined in this way really requires twenty dimensions. We tried to develop corresponding representations in fewer dimensions. A method for searching optimal codification schema in an arbitrary space using neural networks was developed. The method is used to obtain optimal encoding of amino acids at various levels of dimensionality, and applied to optimize the amino acid codifications for the prediction of the solvent accessibility values of the proteins using feed-forward neural networks. The traditional 20-dimensional codification seems to be redundant in solving the solvent accessibility prediction problem, since a 1-dimensional codification is able to achieve almost the same degree of accuracy as the 20-dimensional codification. Optimal coding in much fewer dimensions could be used to make the predictions of accessible surface area with almost the same degree of accuracy as that obtained by a fully unique 20-dimensional coding. The 1-dimensional amino acid codification for solvent accessibility prediction obtained by a purely mathematical way based on neural networks is highly correlated with a physical property of the amino acids, namely their average solvent accessibility. The method developed to find the optimal codification is general, although the codification thus produced is dependent on the type of estimated property.

Amino Acid Sequence↗

Classification of protein-DNA complexes based on structural descriptors.

We attempt to classify protein-DNA complexes by using a set of 11 descriptors, mainly characterizing protein-DNA interactions, including the number of atomic contacts at major and minor grooves, conformational deviations from standard B- and A-DNA forms, widths of DNA grooves, GC content, specificity measures of direct and indirect readouts, and buried surface area at the complex interface. The cluster analyses were carried out for a unique set of 62 complexes including a variety of protein motifs, and 7 distinct clusters were revealed from the analyses. We found that some proteins with the same motif are classified into different clusters, whereas different proteins with distinct motifs are classified into the same cluster. These results suggest that the conventional motif-based classification of DNA binding proteins may not necessarily correspond to structural and functional properties of protein-DNA complexes, and that the present classification will help to identify common properties and rules that govern protein-DNA recognition.

Cluster Analysis↗

Sequence-dependent conformational energy of DNA derived from molecular dynamics simulations: toward understanding the indirect readout mechanism in protein-DNA recognition.

Sequence dependence of DNA conformation plays a crucial role in its recognition by proteins and ligands. To clarify the relationship between sequence and conformation, it is necessary to quantify the conformational energy and specificity of DNA. Here, we make a systematic analysis of dodecamer DNA structures including all the 136 unique tetranucleotide sequences at the center by molecular dynamics simulations. Using a simplified conformational model with six parameters to describe the geometry of adjacent base pairs and harmonic potentials along these coordinates, we estimated the equilibrium conformational parameters and the harmonic potentials of mean force for the central base-pair steps from many trajectories of the simulations. This enabled us to estimate the conformational energy and the specificity for any given DNA sequence and structure. We tested our method by using sequence-structure threading to estimate the conformational energy and the Z-score as a measure of specificity for many B-DNA and A-DNA crystal structures. The average Z-scores were negative for both kinds of structures, indicating that the potential of mean force from the simulation is capable of predicting sequence specificity for the crystal structures and that it may be used to study the sequence specificity of both types of DNA. We also estimated the positional distribution of conformational energy and Z-score within DNA and showed that they are strongly position dependent. This analysis enabled us to identify particular conformations responsible for the specificity. The presented results will provide an insight into the mechanisms of DNA sequence recognition by proteins and ligands.

Base Sequence↗

Prediction and evolutionary information analysis of protein solvent accessibility using multiple linear regression.

A multiple linear regression method was applied to predict real values of solvent accessibility from the sequence and evolutionary information. This method allowed us to obtain coefficients of regression and correlation between the occurrence of an amino-acid residue at a specific target and its sequence neighbor positions on the one hand, and the solvent accessibility of that residue on the other. Our linear regression model based on sequence information and evolutionary models was found to predict residue accessibility with 18.9% and 16.2% mean absolute error respectively, which is better than or comparable to the best available methods. A correlation matrix for several neighbor positions to examine the role of evolutionary information at these positions has been developed and analyzed. As expected, the effective frequency of hydrophobic residues at target positions shows a strong negative correlation with solvent accessibility, whereas the reverse is true for charged and polar residues. The correlation of solvent accessibility with effective frequencies at neighboring positions falls abruptly with distance from target residues. Longer protein chains have been found to be more accurately predicted than their smaller counterparts.

Amino Acids↗

TMBETA-NET: discrimination and prediction of membrane spanning beta-strands in outer membrane proteins.

We have developed a web-server, TMBETA-NET for discriminating outer membrane proteins and predicting their membrane spanning beta-strand segments. The amino acid compositions of globular and outer membrane proteins have been systematically analyzed and a statistical method has been proposed for discriminating outer membrane proteins. The prediction of membrane spanning segments is mainly based on feed forward neural network and refined with beta-strand length. Our program takes the amino acid sequence as input and displays the type of the protein along with membrane-spanning beta-strand segments as a stretch of highlighted amino acid residues. Further, the probability of residues to be in transmembrane beta-strand has been provided with a coloring scheme. We observed that outer membrane proteins were discriminated with an accuracy of 89% and their membrane spanning beta-strand segments at an accuracy of 73% just from amino acid sequence information. The prediction server is available at http://psfs.cbrc.jp/tmbeta-net/.

Amino Acids↗

PSSM-based prediction of DNA binding sites in proteins.

BACKGROUND: Detection of DNA-binding sites in proteins is of enormous interest for technologies targeting gene regulation and manipulation. We have previously shown that a residue and its sequence neighbor information can be used to predict DNA-binding candidates in a protein sequence. This sequence-based prediction method is applicable even if no sequence homology with a previously known DNA-binding protein is observed. Here we implement a neural network based algorithm to utilize evolutionary information of amino acid sequences in terms of their position specific scoring matrices (PSSMs) for a better prediction of DNA-binding sites. RESULTS: An average of sensitivity and specificity using PSSMs is up to 8.7% better than the prediction with sequence information only. Much smaller data sets could be used to generate PSSM with minimal loss of prediction accuracy. CONCLUSION: One problem in using PSSM-derived prediction is obtaining lengthy and time-consuming alignments against large sequence databases. In order to speed up the process of generating PSSMs, we tried to use different reference data sets (sequence space) against which a target protein is scanned for PSI-BLAST iterations. We find that a very small set of proteins can actually be used as such a reference data without losing much of the prediction value. This makes the process of generating PSSMs very rapid and even amenable to be used at a genome level. A web server has been developed to provide these predictions of DNA-binding sites for any new protein from its amino acid sequence. AVAILABILITY: Online predictions based on this method are available at http://www.netasa.org/dbs-pssm/

Algorithms↗

Application of residue distribution along the sequence for discriminating outer membrane proteins.

Discriminating outer membrane proteins from other folding types of globular and membrane proteins is an important problem both for detecting outer membrane proteins from genomic sequences and for the successful prediction of their secondary and tertiary structures. In this work, we have systematically analyzed the distribution of amino acid residues in the sequences of globular and outer membrane proteins. We observed that the occurrence of two neighboring aliphatic and polar residues is significantly higher in outer membrane proteins than in globular proteins. From the information about the dipeptide composition we have devised a statistical method for discriminating outer membrane proteins from other globular and membrane proteins. Our approach correctly picked up the outer membrane proteins with an accuracy of 95% for the training set of 337 proteins. On the other hand, our method has correctly excluded the globular proteins at an accuracy of 79% in a non-redundant dataset of 674 proteins. Furthermore, the present method is able to correctly exclude alpha-helical membrane proteins up to an accuracy of 87%. These accuracy levels are comparable to other methods in the literature. The influence of protein size and structural class for discrimination is discussed.

Amino Acid Sequence↗

Solvent accessibility in native and isolated domain environments: general features and implications to interface predictability.

A non-redundant database of 4536 structural domains, comprising more than 790,000 residues, has been used for the calculation of their solvent accessibility in the native protein environment and then in the isolated domain environment. Nearly 140,000 (18%) residues showed a change in accessible surface area in the above two conditions. General features of this change under these two circumstances have been pointed out. Propensities of these interfacing amino acid residues have been calculated and their variation for different secondary structure types has been analyzed. Actual amount of surface area lost by different secondary structures is higher in the case of helix and strands compared to coil and other conformations. Overall change in surface area in hydrophobic and uncharged residues is higher than that in charged residues. An attempt has been made to know the predictability of interface residues from sequence environments. This analysis and prediction results have significant implications towards determining interacting residues in proteins and for the prediction of protein-protein, protein-ligand, protein-DNA and similar interactions.

Amino Acids↗

Look-up tables for protein solvent accessibility prediction and nearest neighbor effect analysis.

We developed dictionaries of two-, three-, and five-residue patterns in proteins and computed the average solvent accessibility of the central residues in their native proteins. These dictionaries serve as a look-up table for making subsequent predictions of solvent accessibility of amino acid residues. We find that predictions made in this way are very close to those made using more sophisticated methods of solvent accessibility prediction. We also analyzed the effect of immediate neighbors on the solvent accessibility of residues. This helps us in understanding how the same residue type may have different accessible surface areas in different proteins and in different positions of the same protein. We observe that certain residues have a tendency to increase or decrease the solvent accessibility of their neighboring residues in C- or N-terminal positions. Interestingly, the C-terminal and N-terminal neighbor residues are found to have asymmetric roles in modifying solvent accessibility of residues. As expected, similar neighbors enhance the hydrophobic or hydrophilic character of residues. Detailed look-up tables are provided on the web at www.netasa.org/look-up/.

Amino Acid Sequence↗

Moment-based prediction of DNA-binding proteins.

Net charge, electric dipole moment and quadrupole moment tensors were calculated for 78 amino acid sequences from 62 representative DNA-binding proteins with known structures. It was found that the magnitudes of the moments of electric charge distribution in these chains differ significantly from those of a non-binding control data set. Net charge, net dipole moment and quadrupole moment could each distinguish binding and non-binding proteins with 82.6%, 77.4% and 73.7% accuracy by single-variable predictors without cross-validation. Using hybrid predictors with information of charge and both moments, the best predictions were 85.6% without cross-validation and 83.9% for the cross-validated data sets. This level of prediction accuracy obtained with these simple descriptors competes with the results obtained using more complex models including many descriptors. The coarse graining of atomic charges onto C(alpha) atoms did not reduce the prediction accuracy significantly. This result suggests that we can use C(alpha) coordinates derived from homology modeling to predict DNA-binding proteins. The speed and accuracy of this method, in combination with homology-based methods of structure prediction, should enhance genome-wide recognition of DNA-binding proteins.

Animals↗

Qgrid: clustering tool for detecting charged and hydrophobic regions in proteins.

We have developed a simple but powerful method and web server to quickly locate charged and hydrophobic clusters in proteins (http://www.netasa.org/qgrid/index.html). For the charged clusters, each atom in the protein is first assigned a charge according to a standard force field. Then a box is created with dimensions corresponding to the range of atomic coordinates. This box is then divided into cubic grids of selected size, which now have one or more charged atoms in them. This leaves each grid with a certain amount of charge. Cubic grids with more than a cutoff charge are then clustered using a hierarchical clustering method based on Euclidean distance. A tree diagram made from the resulting clusters indicates the distribution of charged and hydrophobic regions of the protein. Hydrophobic clusters are developed by grouping the positions of C(alpha) atoms of such residues. We propose that such a tree representation will be helpful in detecting protein-protein interfaces, structure similarity and motif detection.

Hydrophobic and Hydrophilic Interactions↗

ASAView: database and tool for solvent accessibility representation in proteins.

BACKGROUND: Accessible surface area (ASA) or solvent accessibility of amino acids in a protein has important implications. Knowledge of surface residues helps in locating potential candidates of active sites. Therefore, a method to quickly see the surface residues in a two dimensional model would help to immediately understand the population of amino acid residues on the surface and in the inner core of the proteins. RESULTS: ASAView is an algorithm, an application and a database of schematic representations of solvent accessibility of amino acid residues within proteins. A characteristic two-dimensional spiral plot of solvent accessibility provides a convenient graphical view of residues in terms of their exposed surface areas. In addition, sequential plots in the form of bar charts are also provided. Online plots of the proteins included in the entire Protein Data Bank (PDB), are provided for the entire protein as well as their chains separately. CONCLUSIONS: These graphical plots of solvent accessibility are likely to provide a quick view of the overall topological distribution of residues in proteins. Chain-wise computation of solvent accessibility is also provided.

Computer Graphics↗

Neural network-based prediction of transmembrane beta-strand segments in outer membrane proteins.

Prediction of transmembrane beta-strands in outer membrane proteins (OMP) is one of the important problems in computational chemistry and biology. In this work, we propose a method based on neural networks for identifying the membrane-spanning beta-strands. We introduce the concept of "residue probability" for assigning residues in transmembrane beta-strand segments. The performance of our method is evaluated with single-residue accuracy, correlation, specificity, and sensitivity. Our predicted segments show a good agreement with experimental observations with an accuracy level of 73% solely from amino acid sequence information. Further, the predictive power of N- and C-terminal residues in each segments, number of segments in each protein, and the influence of cutoff probability for identifying membrane-spanning beta-strands will be discussed. We have developed a Web server for predicting the transmembrane beta-strands from the amino acid sequence, and the prediction results are available at http://psfs.cbrc.jp/tmbeta-net/.

Algorithms↗

Role of non-covalent interactions for determining the folding rate of two-state proteins.

Understanding the factors influencing the folding rate of proteins is a challenging problem. In this work, we have analyzed the role of non-covalent interactions for the folding rate of two-state proteins by free-energy approach. We have computed the free-energy terms, hydrophobic, electrostatic, hydrogen-bonding and van der Waals free energies. The hydrophobic free energy has been divided into the contributions from different atoms, carbon, neutral nitrogen and oxygen, charged nitrogen and oxygen, and sulfur. All the free-energy terms have been related with the folding rates of 28 two-state proteins with single and multiple correlation coefficients. We found that the hydrophobic free energy due to carbon atoms and hydrogen-bonding free energy play important roles to determine the folding rate in combination with other free energies. The normalized energies with total number of residues showed better results than the total energy of the protein. The comparison of amino acid properties with free-energy terms indicates that the energetic terms explain better the folding rate than amino acid properties. Further, the combination of free energies with topological parameters yielded the correlation of 0.91. The present study demonstrates the importance of topology for determining the folding rate of two-state proteins.

Binding Sites↗

Analysis and prediction of DNA-binding proteins and their binding residues based on composition, sequence and structural information.

MOTIVATION: Though vitally important to cell function, the mechanism of protein-DNA binding has not yet been completely understood. We therefore analysed the relationship between DNA binding and protein sequence composition, solvent accessibility and secondary structure. Using non-redundant databases of transcription factors and protein-DNA complexes, neural network models were developed to utilize the information present in this relationship to predict DNA-binding proteins and their binding residues. RESULTS: Sequence composition was found to provide sufficient information to predict the probability of its binding to DNA with nearly 69% sensitivity at 64% accuracy for the considered proteins; sequence neighbourhood and solvent accessibility information were sufficient to make binding site predictions with 40% sensitivity at 79% accuracy. Detailed analysis of binding residues shows that some three- and five-residue segments frequently bind to DNA and that solvent accessibility plays a major role in binding. Although, binding behaviour was not associated with any particular secondary structure, there were interesting exceptions at the residue level. Over-representation of some residues in the binding sites was largely lost at the total sequence level, but a different kind of compositional preference was observed in DNA-binding proteins.

Algorithms↗