PubMed Health⌕ Search

Biomedical subjects

Naigong Zhang

Publications and source records attributed to Naigong Zhang.

3 recordsLinked to original sources

MMDB: annotating protein sequences with Entrez's 3D-structure database.

Three-dimensional (3D) structure is now known for a large fraction of all protein families. Thus, it has become rather likely that one will find a homolog with known 3D structure when searching a sequence database with an arbitrary query sequence. Depending on the extent of similarity, such neighbor relationships may allow one to infer biological function and to identify functional sites such as binding motifs or catalytic centers. Entrez's 3D-structure database, the Molecular Modeling Database (MMDB), provides easy access to the richness of 3D structure data and its large potential for functional annotation. Entrez's search engine offers several tools to assist biologist users: (i) links between databases, such as between protein sequences and structures, (ii) pre-computed sequence and structure neighbors, (iii) visualization of structure and sequence/structure alignment. Here, we describe an annotation service that combines some of these tools automatically, Entrez's 'Related Structure' links. For all proteins in Entrez, similar sequences with known 3D structure are detected by BLAST and alignments are recorded. The 'Related Structure' service summarizes this information and presents 3D views mapping sequence residues onto all 3D structures available in MMDB (http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=structure).

Databases, Protein↗

Inhibition of HIV-1 virus replication using small soluble Tat peptides.

Although the introduction of highly active antiretroviral therapy (HAART) has led to a significant reduction in AIDS-related morbidity and mortality, unfortunately, many patients discontinue their initial HAART regimen, resulting in development of viral resistance. During HIV infection, the viral activator Tat is needed for viral progeny formation, and the basic and core domains of Tat are the most conserved parts of the protein. Here, we show that a Tat 41/44 peptide from the core domain can inhibit HIV-1 gene expression and replication. The peptides are not toxic to cells and target the Cdk2/Cyclin E complex, inhibiting the phosphorylation of serine 5 of RNAPII. Using the Cdk2 X-ray crystallography structure, we found that the low-energy wild-type peptides could bind to the ATP binding pocket, whereas the mutant peptide bound to the Cdk2 interface. Finally, we show that these peptides do not allow loading of the catalytic domain of the cdk/cyclin complex onto the HIV-1 promoter in vivo.

Amino Acid Sequence↗

Fast accurate evaluation of protein solvent exposure.

Protein solvation energies are often taken to be proportional to solvent-accessible surface areas. Computation of these areas is numerically demanding and may become a bottleneck for folding and design applications. Fast graph-based methods, such as dead-end elimination (DEE), become possible if all energies, including solvation energies, are expressed as single-residue and pair-residue terms. To this end, Street and Mayo originated a pair-residue approximation for solvent-accessible surface areas (Street AG, Mayo SL. Pairwise calculation of protein solvent accessible surface areas. Fold Des 1998;3:253-258). The dominant source of error in this method is the overlapping burial of side-chain surfaces in the protein core. Here we report a new pair-residue approximation, which greatly reduces this overlap error by the use of optimized generic side-chains. We have tested the generic-side-chain method for the ten proteins studied by Street and Mayo and for 377 single-domain proteins from the CATH database (Orengo CA, Michie AD, Jones S, Jones DT, Swindells MB, Thornton JM. CATH-A hierarchic classification of protein domain structures. Structure 1997;5:1093-1108). With little additional cost in computation, the new method consistently reduces error for total areas and residue-by-residue areas by more than a factor of two. For example, the residue-by-residue error (for buried area) is reduced from 7.42 A(2) to 3.70 A(2). This difference translates into a solvation energy difference of approximately 0.2 kcal/mol per residue, amounting to a reduction in root-mean-square energy error of 2 kcal/mol for a 100 residue chain, a potentially critical difference for both protein folding and design applications.

Computational Biology↗