PubMed Health⌕ Search

Biomedical subjects

Jerry Tsai

Publications and source records attributed to Jerry Tsai.

13 recordsLinked to original sources

Detecting protein dissimilarities in multiple alignments using Bayesian variable selection.

MOTIVATION: We present an application of Bayesian variable selection to the novel detection of sequence elements that confer negative design to protein structure and function. As an illustration, we analyze the different dimer interfaces between the CXCL8 chemokine family with the CCL4 and CCL2 chemokine families to discover the changes that disfavor CXCL8 of quaternary structure. RESULTS: In comparison with known experimental results, our method identifies evolutionarily conserved sequence changes in the CC families that inhibit CXCL8 quaternary structure. Therefore, we find positive selection of negative design elements. Furthermore, our approach predicts that a two-residue deletion conserved in the CCL4 chemokine family disfavors CXCL8 dimerization. AVAILABILITY: The Matlab code for the Bayesian variable selection is freely available at http://stat.tamu.edu/~mvannucci/webpages/codes.html

Algorithms↗

Terminal ion pairs stabilize the second beta-hairpin of the B1 domain of protein G.

The effects of terminal ion pairs on the stability of a beta-hairpin peptide corresponding to the C-terminal residues of the B1 domain of protein G were determined using thermal unfolding as monitored by nuclear magnetic resonance and circular dichroism spectroscopy. Molecular dynamics (MD) simulations were also performed to examine the effect of ion pairs on the structures. Eight peptides were studied including the wild type (G41) and the N-terminal modified sequences that had the first residue deleted (E42), replaced with a Lys (K41), or extended by an additional Gly (G40). Acetylated variants were made to examine the effect of removing the positive N-terminal charge on beta-hairpin stability. The rank in stability determined experimentally is K41 > E42 approximately G41 approximately G40 > Ac-K41 > Ac-E42 approximately Ac-G41 > Ac-G40. The Tm of the K41 peptide is 12 degrees C higher than G41, while the Tm values for the acetylated peptides are less than their unacetylated forms by more than 15 degrees C. NOE cross-peaks between side-chain methylene groups at the N- and C-termini and larger CalphaH shifts compared to random values are seen for K41. The addition of 20% methanol increases the stability in K41 and G41. The MD studies complement these results by showing that the charged N-terminus is important to stability. The type of ion pair observed varies with peptide, and when formed the simulations show that the ion pair can prevent fraying of the beta-strands through electrostatic and hydrophobic contacts. Therefore, introducing favorable electrostatic interactions at the N- and C-termini can substantially enhance beta-hairpin stability and help define the structure.

Amino Acid Sequence↗

Hydrogen bonding increases packing density in the protein interior.

The contribution of hydrogen bonds and the burial of polar groups to protein stability is a controversial subject. Theoretical studies suggest that burying polar groups in the protein interior makes an unfavorable contribution to the stability, but experimental studies show that burying polar groups, especially those that are hydrogen bonded, contributes favorably to protein stability. Understanding the factors that are not properly accounted for by the theoretical models would improve the models so that they more accurately describe experimental results. It has been suggested that hydrogen bonds may contribute to protein stability, in part, by increasing packing density in the protein interior, and thereby increasing the contribution of van der Waals interactions to protein stability. To investigate the influence of hydrogen bonds on packing density, we analyzed 687 crystal structures and determined the volume of buried polar groups as a function of their extent of hydrogen bonding. Our findings show that peptide groups and polar side chains that form hydrogen bonds occupy a smaller volume than the same groups when they do not form hydrogen bonds. For example, peptide groups in which both polar groups are hydrogen bonded occupy a volume, on average, 5.2 A3 less than a peptide group that is not hydrogen bonded.

Crystallography, X-Ray↗

Cataloging the relationships between proteins: a review of interaction databases.

By organizing and making widely accessible the increasing amounts of data from high-throughput analyses, protein interaction databases have become an integral resource for the biological community in relating sequence data with higher-order function. To provide a sense of the use and applicability of these databases, we describe each of the major comprehensive interaction databases as well as some of the more specialized ones. Content description, search/browse functionalities, and data presentation are discussed. A succinct explanation of database contents helps the user quickly identify whether the database contains applicable information to their research interest. Broad levels of search/browse functions as well as descriptions/examples allow users to quickly find and access pertinent data. At this point, clear presentation of search results as well as the primary content is necessary. Many databases display information graphically or divided into smaller digestible parts over a number of tabbed/linked pages. In addition, cross-linking between the databases promotes interconnectivity of the data and is an added layer of relational data for the user. Overall, although these protein interaction databases are under continual improvement, their current state shows that much time and effort has gone into organizing and presenting these large sets of data-describing protein interactions.

Database Management Systems↗

Assessing methods for identifying pair-wise atomic contacts across binding interfaces.

An essential step in understanding the molecular basis of protein-protein interactions is the accurate identification of inter-protein contacts. We evaluate a number of common methods used in analyzing protein-protein interfaces: a Voronoi polyhedra-based approach, changes in solvent accessible surface area (DeltaSASA) and various radial cutoffs (closest atom, Cbeta, and centroid). First, we compared the Voronoi polyhedra-based analysis to the DeltaSASA and show that using Voronoi polyhedra finds knob-in-hole contacts. To assess the accuracy between the Voronoi polyhedra-based approach and the various radial cutoff methods, two sets of data were used: a small set of 75 experimental mutants and a larger one of 592 structures of protein-protein interfaces. In an assessment using the small set, the Voronoi polyhedra-based methods, a solvent accessible surface area method, and the closest atom radial method identified 100% of the direct contacts defined by mutagenesis data, but only the Voronoi polyhedra-based method found no false positives. The other radial methods were not able to find all of the direct contacts even using a cutoff of 9A. With the larger set of structures, we compared the overall number contacts using the Voronoi polyhedra-based method as a standard. All the radial methods using a 6-A cutoff identified more interactions, but these putative contacts included many false positives as well as missed many false negatives. While radial cutoffs are quicker to calculate as well as to implement, this result highlights why radial cutoff methods do not have the proper resolution to detail the non-homogeneous packing within protein interfaces, and suggests an inappropriate bias in pair-wise contact potentials. Of the radial cutoff methods, using the closest atom approach exhibits the best approximation to the more intensive Voronoi calculation. Our version of the Voronoi polyhedra-based method QContacts is available at .

Databases, Protein↗

Characterizing conserved structural contacts by pair-wise relative contacts and relative packing groups.

To adequately deal with the inherent complexity of interactions between protein side-chains, we develop and describe here a novel method for characterizing protein packing within a fold family. Instead of approaching side-chain interactions absolutely from one residue to another, we instead consider the relative interactions of contacting residue pairs. The basic element, the pair-wise relative contact, is constructed from a sequence alignment and contact analysis of a set of structures and consists of a cluster of similarly oriented, interacting, side-chain pairs. To demonstrate this construct's usefulness in analyzing protein structure, we used the pair-wise relative contacts to analyze two sets of protein structures as defined by SCOP: the diverse globin-like superfamily (126 structures) and the more uniform heme binding globin family (a 94 structure subset of the globin-like superfamily). The superfamily structure set produced 1266 unique pair-wise relative contacts, whereas the family structure subset gave 1001 unique pair-wise relative contacts. For both sets, we show that these constructs can be used to accurately and automatically differentiate between fold classes. Furthermore, these pair-wise relative contacts correlate well with sequence identity and thus provide a direct relationship between changes in sequence and changes in structure. To capture the complexity of protein packing, these pair-wise relative contacts can be superimposed around a single residue to create a multi-body construct called a relative packing group. Construction of convex hulls around the individual packing groups provides a measure of the variation in packing around a residue and defines an approximate volume of space occupied by the groups interacting with a residue. We find that these relative packing groups are useful in understanding the structural quality of sequence or structure alignments. Moreover, they provide context to calculate a value for structural randomness, which is important in properly assessing the quality of a structural alignment. The results of this study provide the framework for future analysis for correlating sequence changes to specific structure changes.

Algorithms↗

Practical conversion from torsion space to Cartesian space for in silico protein synthesis.

Many applications require a method for translating a large list of bond angles and bond lengths to precise atomic Cartesian coordinates. This simple but computationally consuming task occurs ubiquitously in modeling proteins, DNA, and other polymers as well as in many other fields such as robotics. To find an optimal method, algorithms can be compared by a number of operations, speed, intrinsic numerical stability, and parallelization. We discuss five established methods for growing a protein backbone by serial chain extension from bond angles and bond lengths. We introduce the Natural Extension Reference Frame (NeRF) method developed for Rosetta's chain extension subroutine, as well as an improved implementation. In comparison to traditional two-step rotations, vector algebra, or Quaternion product algorithms, the NeRF algorithm is superior for this application: it requires 47% fewer floating point operations, demonstrates the best intrinsic numerical stability, and offers prospects for parallel processor acceleration. The NeRF formalism factors the mathematical operations of chain extension into two independent terms with orthogonal subsets of the dependent variables; the apparent irreducibility of these factors hint that the minimal operation set may have been identified. Benchmarks are made on Intel Pentium and Motorola PowerPC CPUs.

Algorithms↗

Some fundamental aspects of building protein structures from fragment libraries.

We have investigated some of the basic principles that influence generation of protein structures using a fragment-based, random insertion method. We tested buildup methods and fragment library quality for accuracy in constructing a set of known structures. The parameters most influential in the construction procedure are bond and torsion angles with minor inaccuracies in bond angles alone causing >6 A CalphaRMSD for a 150-residue protein. Idealization to a standard set of values corrects this problem, but changes the torsion angles and does not work for every structure. Alternatively, we found using Cartesian coordinates instead of torsion angles did not reduce performance and can potentially increase speed and accuracy. Under conditions simulating ab initio structure prediction, fragment library quality can be suboptimal and still produce near-native structures. Using various clustering criteria, we created a number of libraries and used them to predict a set of native structures based on nonnative fragments. Local CalphaRMSD fit of fragments, library size, and takeoff/landing angle criteria weakly influence the accuracy of the models. Based on a fragment's minimal perturbation upon insertion into a known structure, a seminative fragment library was created that produced more accurate structures with fragments that were less similar to native fragments than the other sets. These results suggest that fragments need only contain native-like subsections, which when correctly overlapped, can recreate a native-like model. For fragment-based, random insertion methods used in protein structure prediction and design, our findings help to define the parameters this method needs to generate near-native structures.

Computer Simulation↗

An improved protein decoy set for testing energy functions for protein structure prediction.

We have improved the original Rosetta centroid/backbone decoy set by increasing the number of proteins and frequency of near native models and by building on sidechains and minimizing clashes. The new set consists of 1,400 model structures for 78 different and diverse protein targets and provides a challenging set for the testing and evaluation of scoring functions. We evaluated the extent to which a variety of all-atom energy functions could identify the native and close-to-native structures in the new decoy sets. Of various implicit solvent models, we found that a solvent-accessible surface area-based solvation provided the best enrichment and discrimination of close-to-native decoys. The combination of this solvation treatment with Lennard Jones terms and the original Rosetta energy provided better enrichment and discrimination than any of the individual terms. The results also highlight the differences in accuracy of NMR and X-ray crystal structures: a large energy gap was observed between native and non-native conformations for X-ray structures but not for NMR structures.

Algorithms↗

Evidence of turn and salt bridge contributions to beta-hairpin stability: MD simulations of C-terminal fragment from the B1 domain of protein G.

We ran and analyzed a total of eighteen, 10 ns molecular dynamics simulations of two C-terminal beta-hairpins from the B1 domain of Protein G: twelve runs for the last 16 residues and six runs for the last 15 residues, G41-E56 and E42-E56, respectively. Based on their CalphaRMS deviation from the starting structure and the pattern of stabilizing interactions (hydrogen bonds, hydrophobic contacts, and salt bridges), we were able to classify the twelve runs on G41-E56 into one of three general states of the beta-hairpin ensemble: 'Stable', 'Unstable', and 'Unfolded'. Comparing the specific interactions between these states, we find that on average the stable beta-hairpin buries 287 A(2) of hydrophobic surface area, makes 13 hydrogen bonds, and forms 3 salt-bridges. We find that the hydrophobic core prefers to make some specific contacts; however, this core does not require optimal packing. Side-chain hydrogen bonds stabilize the beta-hairpin turn with strong stabilizing interactions primarily due to the carboxyl of D46 with contributions from T49 hydroxyl. Buoyed by the strength of the hydrophobic core, other hydrogen bonds, primarily main-chain, guide the beta-hairpin into registration by forming a loose network of interactions, making an approximately constant number of hydrogen bonds from a pool of possible candidates. In simulations on E42-E56, where the salt bridge closing the termini is not favored, we observe that all the simulations show no 'Stable' behavior, but are 'Unstable' or 'Unfolded'. We can estimate that the salt-bridge between the termini provides approximately 1.3 kcal/mol. Altogether, the results suggest that the beta-hairpin folds beginning at the turn, followed by hydrophobic collapse, and then hydrogen bond formation. Salt bridges help to stabilize the folded conformations by inhibiting unfolded states.

Hydrogen Bonding↗

Calculations of protein volumes: sensitivity analysis and parameter database.

MOTIVATION: The precise sizes of protein atoms in terms of occupied packing volume are of great importance. We have previously presented standard volumes for protein residues based on calculations with Voronoi-like polyhedra. To understand the applicability and limitations of our set, we investigated, in detail, the sensitivity of the volume calculations to a number of factors: (i) the van der Waals radii set, (ii) the criteria for including buried atoms in the calculations or atom selection, (iii) the method of positioning the dividing plane in polyhedra construction, and (iv) the set of structures used in the averaging. RESULTS: We find that different radii sets have only moderate affects to the distribution and mean of volumes. Atom selection and dividing plane methods cause larger changes in protein atoms volumes. More significantly, we show how the variation in volumes appears to be clearly related to the quality of the structures analyzed, with higher quality structures giving consistently smaller average volumes with less variance.

Algorithms↗

Contact order and ab initio protein structure prediction.

Although much of the motivation for experimental studies of protein folding is to obtain insights for improving protein structure prediction, there has been relatively little connection between experimental protein folding studies and computational structural prediction work in recent years. In the present study, we show that the relationship between protein folding rates and the contact order (CO) of the native structure has implications for ab initio protein structure prediction. Rosetta ab initio folding simulations produce a dearth of high CO structures and an excess of low CO structures, as expected if the computer simulations mimic to some extent the actual folding process. Consistent with this, the majority of failures in ab initio prediction in the CASP4 (critical assessment of structure prediction) experiment involved high CO structures likely to fold much more slowly than the lower CO structures for which reasonable predictions were made. This bias against high CO structures can be partially alleviated by performing large numbers of additional simulations, selecting out the higher CO structures, and eliminating the very low CO structures; this leads to a modest improvement in prediction quality. More significant improvements in predictions for proteins with complex topologies may be possible following significant increases in high-performance computing power, which will be required for thoroughly sampling high CO conformations (high CO proteins can take six orders of magnitude longer to fold than low CO proteins). Importantly for such a strategy, simulations performed for high CO structures converge much less strongly than those for low CO structures, and hence, lack of simulation convergence can indicate the need for improved sampling of high CO conformations. The parallels between Rosetta simulations and folding in vivo may extend to misfolding: The very low CO structures that accumulate in Rosetta simulations consist primarily of local up-down beta-sheets that may resemble precursors to amyloid formation.

Algorithms↗