PubMed Health⌕ Search

Biomedical subjects

Zhirong Sun

Publications and source records attributed to Zhirong Sun.

9 recordsLinked to original sources

DBSubLoc: database of protein subcellular localization.

We have built a protein subcellular localization annotation database, the DBSubLoc database, which is available at http://www.bioinfo.tsinghua. edu.cn/dbsubloc.html. Annotations were taken from primary protein databases, model organism genome projects and literature texts, and then were analyzed to dig out the subcellular localization features of the proteins. The proteins are also classified into different categories. Based on sequence alignment, non-redundant subsets of the database have been built, which may provide useful information for subcellular localization prediction. The database now contains >60,000 protein sequences including approximately 30,000 protein sequences in the non-redundant data sets. Online download, search and Blast tools are also available.

Animals↗

Structural and functional characterization of the human CCR5 receptor in complex with HIV gp120 envelope glycoprotein and CD4 receptor by molecular modeling studies.

The entry of human immunodeficiency virus (HIV) into cells depends on a sequential interaction of the gp120 envelope glycoprotein with the cellular receptors CD4 and members of the chemokine receptor family. The CC chemokine receptor CCR5 is such a receptor for several chemokines and a major coreceptor for the entry of R5 HIV type-1 (HIV-1) into cells. Although many studies focus on the interaction of CCR5 with HIV-1, the corresponding interaction sites in CCR5 and gp120 have not been matched. Here we used an approach combining protein structure modeling, docking and molecular dynamics simulation to build a series of structural models of the CCR5 in complexes with gp120 and CD4. Interactions such as hydrogen bonds, salt bridges and van der Waals contacts between CCR5 and gp120 were investigated. Three snapshots of CCR5-gp120-CD4 models revealed that the initial interactions of CCR5 with gp120 are involved in the negatively charged N-terminus (Nt) region of CCR5 and positively charged bridging sheet region of gp120. Further interactions occurred between extracellular loop2 (ECL2) of CCR5 and the base of V3 loop regions of gp120. These interactions may induce the conformational changes in gp120 and lead to the final entry of HIV into the cell. These results not only strongly support the two-step gp120-CCR5 binding mechanism, but also rationalize extensive biological data about the role of CCR5 in HIV-1 gp120 binding and entry, and may guide efforts to design novel inhibitors.

CD4 Antigens↗

Mining gene expression data using a novel approach based on hidden Markov models.

In this work we have developed a new framework for microarray gene expression data analysis. This framework is based on hidden Markov models. We have benchmarked the performance of this probability model-based clustering algorithm on several gene expression datasets for which external evaluation criteria were available. The results showed that this approach could produce clusters of quality comparable to two prevalent clustering algorithms, but with the major advantage of determining the number of clusters. We have also applied this algorithm to analyze published data of yeast cell cycle gene expression and found it able to successfully dig out biologically meaningful gene groups. In addition, this algorithm can also find correlation between different functional groups and distinguish between function genes and regulation genes, which is helpful to construct a network describing particular biological associations. Currently, this method is limited to time series data. Supplementary materials are available at http://www.bioinfo.tsinghua.edu.cn/~rich/hmmgep_supp/.

Algorithms↗

An approach to identify over-represented cis-elements in related sequences.

Computational identification of transcription factor binding sites is an important research area of computational biology. Positional weight matrix (PWM) is a model to describe the sequence pattern of binding sites. Usually, transcription factor binding sites prediction methods based on PWMs require user-defined thresholds. The arbitrary threshold and also the relatively low specificity of the algorithm prevent the result of such an analysis from being properly interpreted. In this study, a method was developed to identify over-represented cis-elements with PWM-based similarity scores. Three sets of closely related promoters were analyzed, and only over- represented motifs with high PWM similarity scores were reported. The thresholds to evaluate the similarity scores to the PWMs of putative transcription factors binding sites can also be automatically determined during the analysis, which can also be used in further research with the same PWMs. The online program is available on the website: http://www.bioinfo.tsinghua.edu.cn/- zhengjsh/OTFBS/.

Actins↗

Proteins with class alpha/beta fold have high-level participation in fusion events.

Now that complete genome sequences are available for a variety of organisms, the elucidation of potential gene products function is a central goal in the post-genome era. Domain fusion analysis has been proposed recently to infer the functional association of the component proteins. Here, we took a new approach to the analysis of the structural features of the proteins involved in fusion events. An exhaustive survey of fusion events within 30 completely sequenced genomes and subsequent structure annotations to the component proteins at a SCOP superfamily level with hidden Markov models was carried out. A domain fusion map was then constructed. The results revealed that proteins with the class alpha/beta fold are frequently involved in fusion events, around 86% of the total 676 assigned single-domain fusion pairs including at least one component protein belonging to the alpha/beta fold class. Moreover, the domain fusion map in our work may offer an attractive framework for designing chimeric enzymes following Nature's lead, and may give useful hints for exploring the evolutionary history of proteins. (c) 2002 Elsevier Science Ltd.

Archaeal Proteins↗

Identifying genes related to drug anticancer mechanisms using support vector machine.

In an effort to identify genes related to the cell line chemosensitivity and to evaluate the functional relationships between genes and anticancer drugs acting by the same mechanism, a supervised machine learning approach called support vector machine was used to label genes into any of the five predefined anticancer drug mechanistic categories. Among dozens of unequivocally categorized genes, many were known to be causally related to the drug mechanisms. For example, a few genes were found to be involved in the biological process triggered by the drugs (e.g. DNA polymerase epsilon was the direct target for the drugs from DNA antimetabolites category). DNA repair-related genes were found to be enriched for about eight-fold in the resulting gene set relative to the entire gene set. Some uncharacterized transcripts might be of interest in future studies. This method of correlating the drugs and genes provides a strategy for finding novel biologically significant relationships for molecular pharmacology.

Antineoplastic Agents↗

Mining functional relationships in feature subspaces from gene expression profiles and drug activity profiles.

In an effort to determine putative functional relationships between gene expression patterns and drug activity patterns of 60 human cancer cell lines, a novel method was developed to discover local associations within cell line subsets. The association of drug-gene pairs is an explorative way of discovering gene markers that predict clinical tumor sensitivity to therapy. Nine drug-gene networks were discovered, as well as dozens of gene-gene and drug-drug networks. Three drug-gene networks with well studied members were discussed and the literature shows that hypothetical functional relationships exist. Therefore, this method enables the gathering of new information beyond global associations.

Antineoplastic Agents↗

Wastewater minimization in indirect electrochemical synthesis of phenylacetaldehyde.

Wastewater minimization in phenylacetaldehyde production by using indirect electrochemical oxidation of phenylethane instead of the seriously polluting traditional chemical process is described in this paper. Results show that high current efficiency of Mn(III) and high yield of phenylacetaldehyde can be obtained at the same sulfuric acid concentration (60%). The electrolytic mediator can be recycled and there will be no waste discharged.

Acetaldehyde↗

Sequence-dependent flexibility in promoter sequences.

The non-neighbor interactions between base-pairs were taken into account to calculate the angular parameters (Omega, rho and tau) describing the orientation of successive base-pair planes and the translation parameters (D(y)) along the long axis of base-pair steps for 36 independent tetramers. A statistical mechanical model was proposed to predict the DNA flexibility that is mainly related to the thermal fluctuations at individual base-pair steps. The DNA flexibility can be described by the root-mean-square deviation of the end-to-end distance of DNA helical structure. The present model was then used to investigate the extreme flexible pattern in prokaryotic and eukaryotic promoter sequences. The results demonstrated several extreme flexible regions related to functionally important elements exist both in prokaryotic promoters and in eukaryotic promoters, DNA flexibility and AT content are highly correlated. The probabilities finding flexibility pattern in promoter sequences were also estimated statistically. The biological implications were discussed briefly.

Animals↗