PubMed Health⌕ Search

Biomedical subjects

Jun-tao Guo

Publications and source records attributed to Jun-tao Guo.

9 recordsLinked to original sources

Photochemical surface mapping of C14S-Sml1p for constrained computational modeling of protein structure.

Photochemically generated hydroxyl radicals were used to map solvent-exposed regions in the C14S mutant of the protein Sml1p, a regulator of the ribonuclease reductase enzyme Rnr1p in Saccharomyces cerevisiae. By using high-performance mass spectrometry to characterize the oxidized peptides created by the hydroxyl radical reactions, amino acid solvent-accessibility data for native and denatured C14S Sml1p that revealed a solvent-excluding tertiary structure in the native state were obtained. The data on solvent accessibilities of various amino acids within the protein were then utilized to evaluate the de novo computational models generated by the HMMSTR/Rosetta server. The top five models initially generated by the server all disagreed with both published nuclear magnetic resonance (NMR) data and the solvent-accessibility data obtained in this study. A structural model adjusted to fit the previously reported NMR data satisfied most of the solvent-accessibility constraints. Through minor adjustment of the rotamers of two amino acid side chains for this latter structure, a model that not only provided a lower energy conformation but also completely satisfied previously reported data from NMR and tryptophan fluorescence measurements, in addition to the solvent-accessibility data presented here, was generated.

Amino Acid Sequence↗

Quantitative evaluation of protein-DNA interactions using an optimized knowledge-based potential.

Computational evaluation of protein-DNA interaction is important for the identification of DNA-binding sites and genome annotation. It could validate the predicted binding motifs by sequence-based approaches through the calculation of the binding affinity between a protein and DNA. Such an evaluation should take into account structural information to deal with the complicated effects from DNA structural deformation, distance-dependent multi-body interactions and solvation contributions. In this paper, we present a knowledge-based potential built on interactions between protein residues and DNA tri-nucleotides. The potential, which explicitly considers the distance-dependent two-body, three-body and four-body interactions between protein residues and DNA nucleotides, has been optimized in terms of a Z-score. We have applied this knowledge-based potential to evaluate the binding affinities of zinc-finger protein-DNA complexes. The predicted binding affinities are in good agreement with the experimental data (with a correlation coefficient of 0.950). On a larger test set containing 48 protein-DNA complexes with known experimental binding free energies, our potential has achieved a high correlation coefficient of 0.800, when compared with the experimental data. We have also used this potential to identify binding motifs in DNA sequences of transcription factors (TF). The TFs in 79.4% of the known TF-DNA complexes have accurately found their native binding sequences from a large pool of DNA sequences. When tested in a genome-scale search for TF-binding motifs of the cyclic AMP regulatory protein (CRP) of Escherichia coli, this potential ranks all known binding motifs of CRP in the top 15% of all candidate sequences.

Base Sequence↗

Molecular modeling of the core of Abeta amyloid fibrils.

Amyloid fibrils, a key pathological feature of Alzheimer's disease (AD) and other amyloidosis implicated in neurodegeneration, have a characteristic cross-beta structure. Here we present a structural model for the core of amyloid fibrils formed by the Abeta peptide using computational approaches and experimental data. Abeta(15-36) was threaded against the parallel beta-helical proteins. Our multi-layer model was constructed using the top scoring template 1lxa, a left-handed parallel beta-helical protein. This six-rung helical model has in-register repeats of the Abeta(15-36) sequence. Each rung has three beta-strands separated by two turns. The model was tested using molecular dynamics simulations in explicit water, and is in good agreement with a number of experimental observations. In addition, a model based on right-handed helical proteins is also described. The core structural model described here might serve as the building block of the Abeta(1-40) amyloid fibril as well as some other amyloid fibrils.

Amyloid beta-Peptides↗

Sml1p is a dimer in solution: characterization of denaturation and renaturation of recombinant Sml1p.

Sml1p is a small 104-amino acid protein from Saccharomyces cerevisiae that binds to the large subunit (Rnr1p) of the ribonucleotide reductase complex (RNR) and inhibits its activity. During DNA damage, S phase, or both, RNR activity must be tightly regulated, since failure to control the cellular level of dNTP pools may lead to genetic abnormalities, such as genome rearrangements, or even cell death. Structural characterization of Sml1p is an important step in understanding the regulation of RNR. Until now the oligomeric state of Sml1p was unknown. Mass spectrometric analysis of wild-type Sml1p revealed an intermolecular disulfide bond involving the cysteine residue at position 14 of the primary sequence. To determine whether disulfide bonding is essential for Sml1p oligomerization, we mutated the Cys14 to serine. Sedimentation equilibrium measurements in the analytical ultracentrifuge show that both wild-type and C14S Sml1p exist as dimers in solution, indicating that the dimerization is not a result of a disulfide bond. Further studies of several truncated Sml1p mutants revealed that the N-terminal 8-20 residues are responsible for dimerization. Unfolding/refolding studies of wild-type and C14S Sml1p reveal that both proteins refold reversibly and have almost identical unfolding/refolding profiles. It appears that Sml1p is a two-domain protein where the N-terminus is responsible for dimerization and the C-terminus for binding and inhibiting Rnr1p activity.

Amino Acid Sequence↗

PROSPECT-PSPP: an automatic computational pipeline for protein structure prediction.

Knowledge of the detailed structure of a protein is crucial to our understanding of the biological functions of that protein. The gap between the number of solved protein structures and the number of protein sequences continues to widen rapidly in the post-genomics era due to long and expensive processes for solving structures experimentally. Computational prediction of structures from amino acid sequence has come to play a key role in narrowing the gap and has been successful in providing useful information for the biological research community. We have developed a prediction pipeline, PROSPECT-PSPP, an integration of multiple computational tools, for fully automated protein structure prediction. The pipeline consists of tools for (i) preprocessing of protein sequences, which includes signal peptide prediction, protein type prediction (membrane or soluble) and protein domain partition, (ii) secondary structure prediction, (iii) fold recognition and (iv) atomic structural model generation. The centerpiece of the pipeline is our threading-based program PROSPECT. The pipeline is implemented using SOAP (Simple Object Access Protocol), which makes it easier to share our tools and resources. The pipeline has an easy-to-use user interface and is implemented on a 64-node dual processor Linux cluster. It can be used for genome-scale protein structure prediction. The pipeline is accessible at http://csbl.bmb.uga.edu/protein_pipeline.

Computational Biology↗

Protein structure prediction using sparse dipolar coupling data.

Residual dipolar coupling (RDC) represents one of the most exciting emerging NMR techniques for protein structure studies. However, solving a protein structure using RDC data alone is still a highly challenging problem. We report here a computer program, RDC-PROSPECT, for protein structure prediction based on a structural homolog or analog of the target protein in the Protein Data Bank (PDB), which best aligns with the (15)N-(1)H RDC data of the protein recorded in a single ordering medium. Since RDC-PROSPECT uses only RDC data and predicted secondary structure information, its performance is virtually independent of sequence similarity between a target protein and its structural homolog/analog, making it applicable to protein targets beyond the scope of current protein threading techniques. We have tested RDC-PROSPECT on all (15)N-(1)H RDC data (representing 43 proteins) deposited in the BioMagResBank (BMRB) database. The program correctly identified structural folds for 83.7% of the target proteins, and achieved an average alignment accuracy of 98.1% residues within a four-residue shift.

Computational Biology↗

Mapping abeta amyloid fibril secondary structure using scanning proline mutagenesis.

Although the amyloid fibrils formed from the Alzheimer's disease amyloid peptide Abeta are rich in cross-beta sheet, the peptide likely also exhibits turn and unstructured regions when it becomes incorporated into amyloid. We generated a series of single-proline replacement mutants of Abeta(1-40) and determined the thermodynamic stabilities of amyloid fibrils formed from these mutants to characterize the susceptibility of different residue positions of the Abeta sequence to proline substitution. The results suggest that the Abeta peptide, when engaged in the amyloid fibril, folds into a conformation containing three highly structured segments, consisting of contiguous sequence elements 15-21, 24-28, and 31-36, that are sensitive to proline replacement and likely to include the beta-sheet portions of the fibrils. Residues relatively insensitive to proline replacement fall into two groups: (a) residues 1-14 and 37-40 are likely to exist in relatively unstructured, flexible elements extruded from the beta-sheet-rich amyloid core; (b) residues 22, 23, 29 and 30 are likely to occupy turn positions between these three structured elements. Although destabilized, fibrils formed from Abeta(1-40) proline mutants are very similar in structure to wild-type fibrils, as indicated by hydrogen-deuterium exchange and other analysis. Interestingly, however, some proline mutations destabilize fibrils while at the same time increasing the number of amide protons protected from hydrogen exchange. This suggests that the stability of amyloid fibrils, rather than being driven exclusively by the formation of H-bonded beta-sheet, is achieved, as in globular proteins, through a balance of stabilizing and destabilizing forces. The proline scanning data are most compatible with a model for amyloid protofilament structure loosely resembling the parallel beta-helix folding motif, such that each Abeta(15-36) core region occupies a single layer of a prismatic, H-bonded stack of peptides.

Amino Acid Substitution↗

Improving the performance of DomainParser for structural domain partition using neural network.

Structural domains are considered as the basic units of protein folding, evolution, function and design. Automatic decomposition of protein structures into structural domains, though after many years of investigation, remains a challenging and unsolved problem. Manual inspection still plays a key role in domain decomposition of a protein structure. We have previously developed a computer program, DomainParser, using network flow algorithms. The algorithm partitions a protein structure into domains accurately when the number of domains to be partitioned is known. However the performance drops when this number is unclear (the overall performance is 74.5% over a set of 1317 protein chains). Through utilization of various types of structural information including hydrophobic moment profile, we have developed an effective method for assessing the most probable number of domains a structure may have. The core of this method is a neural network, which is trained to discriminate correctly partitioned domains from incorrectly partitioned domains. When compared with the manual decomposition results given in the SCOP database, our new algorithm achieves higher decomposition accuracy (81.9%) on the same data set.

Algorithms↗

PROSPECT II: protein structure prediction program for genome-scale applications.

A new method for fold recognition is developed and added to the general protein structure prediction package PROSPECT (http://compbio.ornl.gov/PROSPECT/). The new method (PROSPECT II) has four key features. (i) We have developed an efficient way to utilize the evolutionary information for evaluating the threading potentials including singleton and pairwise energies. (ii) We have developed a two-stage threading strategy: (a) threading using dynamic programming without considering the pairwise energy and (b) fold recognition considering all the energy terms, including the pairwise energy calculated from the dynamic programming threading alignments. (iii) We have developed a combined z-score scheme for fold recognition, which takes into consideration the z-scores of each energy term. (iv) Based on the z-scores, we have developed a confidence index, which measures the reliability of a prediction and a possible structure-function relationship based on a statistical analysis of a large data set consisting of threadings of 600 query proteins against the entire FSSP templates. Tests on several benchmark sets indicate that the evolutionary information and other new features of PROSPECT II greatly improve the alignment accuracy. We also demonstrate that the performance of PROSPECT II on fold recognition is significantly better than any other method available at all levels of similarity. Improvement in the sensitivity of the fold recognition, especially at the superfamily and fold levels, makes PROSPECT II a reliable and fully automated protein structure and function prediction program for genome-scale applications.

Algorithms↗