PubMed Health⌕ Search

Biomedical subjects

Zhengchang Su

Publications and source records attributed to Zhengchang Su.

15 recordsLinked to original sources

Operon prediction using both genome-specific and general genomic information.

We have carried out a systematic analysis of the contribution of a set of selected features that include three new features to the accuracy of operon prediction. Our analyses have led to a number of new insights about operon prediction, including that (i) different features have different levels of discerning power when used on adjacent gene pairs with different ranges of intergenic distance, (ii) certain features are universally useful for operon prediction while others are more genome-specific and (iii) the prediction reliability of operons is dependent on intergenic distances. Based on these new insights, our newly developed operon-prediction program achieves more accurate operon prediction than the previous ones, and it uses features that are most readily available from genomic sequences. Our prediction results indicate that our (non-linear) decision tree-based classifier can predict operons in a prokaryotic genome very accurately when a substantial number of operons in the genome are already known. For example, the prediction accuracy of our program can reach 90.2 and 93.7% on Bacillus subtilis and Escherichia coli genomes, respectively. When no such information is available, our (linear) logistic function-based classifier can reach the prediction accuracy at 84.6 and 83.3% for E.coli and B.subtilis, respectively.

Bacillus subtilis↗

Operon prediction in Pyrococcus furiosus.

Identification of operons in the hyperthermophilic archaeon Pyrococcus furiosus represents an important step to understanding the regulatory mechanisms that enable the organism to adapt and thrive in extreme environments. We have predicted operons in P.furiosus by combining the results from three existing algorithms using a neural network (NN). These algorithms use intergenic distances, phylogenetic profiles, functional categories and gene-order conservation in their operon prediction. Our method takes as inputs the confidence scores of the three programs, and outputs a prediction of whether adjacent genes on the same strand belong to the same operon. In addition, we have applied Gene Ontology (GO) and KEGG pathway information to improve the accuracy of our algorithm. The parameters of this NN predictor are trained on a subset of all experimentally verified operon gene pairs of Bacillus subtilis. It subsequently achieved 86.5% prediction accuracy when applied to a subset of gene pairs for Escherichia coli, which is substantially better than any of the three prediction programs. Using this new algorithm, we predicted 470 operons in the P.furiosus genome. Of these, 349 were validated using DNA microarray data.

Algorithms↗

Novel changes in discoidal high density lipoprotein morphology: a molecular dynamics study.

ApoA-I is a uniquely flexible lipid-scavenging protein capable of incorporating phospholipids into stable particles. Here we report molecular dynamics simulations on a series of progressively smaller discoidal high density lipoprotein particles produced by incremental removal of palmitoyloleoylphosphatidylcholine via four different pathways. The starting model contained 160 palmitoyloleoylphosphatidylcholines and a belt of two antiparallel amphipathic helical lipid-associating domains of apolipoprotein (apo) A-I. The results are particularly compelling. After a few nanoseconds of molecular dynamics simulation, independent of the starting particle and method of size reduction, all simulated double belts of the four lipidated apoA-I particles have helical domains that impressively approximate the x-ray crystal structure of lipid-free apoA-I, particularly between residues 88 and 186. These results provide atomic resolution models for two of the particles produced by in vitro reconstitution of nascent high density lipoprotein particles. These particles, measuring 95 angstroms and 78 angstroms by nondenaturing gradient gel electrophoresis, correspond in composition and in size/shape (by negative stain electron microscopy) to the simulated particles with molar ratios of 100:2 and 50:2, respectively. The lipids of the 100:2 particle family form minimal surfaces at their monolayer-monolayer interface, whereas the 50:2 particle family displays a lipid pocket capable of binding a dynamic range of phospholipid molecules.

Apolipoprotein A-I↗

Computational inference and experimental validation of the nitrogen assimilation regulatory network in cyanobacterium Synechococcus sp. WH 8102.

Deciphering the regulatory networks encoded in the genome of an organism represents one of the most interesting and challenging tasks in the post-genome sequencing era. As an example of this problem, we have predicted a detailed model for the nitrogen assimilation network in cyanobacterium Synechococcus sp. WH 8102 (WH8102) using a computational protocol based on comparative genomics analysis and mining experimental data from related organisms that are relatively well studied. This computational model is in excellent agreement with the microarray gene expression data collected under ammonium-rich versus nitrate-rich growth conditions, suggesting that our computational protocol is capable of predicting biological pathways/networks with high accuracy. We then refined the computational model using the microarray data, and proposed a new model for the nitrogen assimilation network in WH8102. An intriguing discovery from this study is that nitrogen assimilation affects the expression of many genes involved in photosynthesis, suggesting a tight coordination between nitrogen assimilation and photosynthesis processes. Moreover, for some of these genes, this coordination is probably mediated by NtcA through the canonical NtcA promoters in their regulatory regions.

Bacterial Proteins↗

Mapping of orthologous genes in the context of biological pathways: An application of integer programming.

Mapping biological pathways across microbial genomes is a highly important technique in functional studies of biological systems. Existing methods mainly rely on sequence-based orthologous gene mapping, which often leads to suboptimal mapping results because sequence-similarity information alone does not contain sufficient information for accurate identification of orthology relationship. Here we present an algorithm for pathway mapping across microbial genomes. The algorithm takes into account both sequence similarity and genomic structure information such as operons and regulons. One basic premise of our approach is that a microbial pathway could generally be decomposed into a few operons or regulons. We formulated the pathway-mapping problem to map genes across genomes to maximize their sequence similarity under the constraint that the mapped genes be grouped into a few operons, preferably coregulated in the target genome. We have developed an integer-programming algorithm for solving this constrained optimization problem and implemented the algorithm as a computer software program, p-map. We have tested p-map on a number of known homologous pathways. We conclude that using genomic structure information as constraints could greatly improve the pathway-mapping accuracy over methods that use sequence-similarity information alone.

Algorithms↗

The evolution of microbial phosphonate degradative pathways.

Phosphonate utilization by microbes provides a potential source of phosphorus for their growth. Homologous genes for both C-P lyase and phosphonatase degradative pathways are distributed in distantly related bacterial species. The phn gene clusters for the C-P lyase pathway show great structural and compositional variation among organisms, but all contain phnG-phnM genes that are essential for C-P bond cleavage. In the gamma-proteobacterium Erwinia carotovora, genes common to phosphonate biosyntheses were found in neighboring positions of those for the C-P lyase degradative pathway and in the same transcriptional direction. A gene encoding a hypothetical protein DUF1045 was found predominantly associated with the phn gene cluster and was predicted functionally related to C-P bond cleavage. Genes for phosphonate degradation are frequently located in close proximity of genes encoding transposases or other mobile elements. Phylogenetic analyses suggest that both degradative pathways have been subject to extensive lateral gene transfers during their evolution. The implications of plasmids and transposition in the evolution of phosphonate degradation are also discussed.

Bacteria↗

Comparative genomics analysis of NtcA regulons in cyanobacteria: regulation of nitrogen assimilation and its coupling to photosynthesis.

We have developed a new method for prediction of cis-regulatory binding sites and applied it to predicting NtcA regulated genes in cyanobacteria. The algorithm rigorously utilizes concurrence information of multiple binding sites in the upstream region of a gene and that in the upstream regions of its orthologues in related genomes. A probabilistic model was developed for the evaluation of prediction reliability so that the prediction false positive rate could be well controlled. Using this method, we have predicted multiple new members of the NtcA regulons in nine sequenced cyanobacterial genomes, and showed that the false positive rates of the predictions have been reduced on an average of 40-fold compared to the conventional methods. A detailed analysis of the predictions in each genome showed that a significant portion of our predictions are consistent with previously published results about individual genes. Intriguingly, NtcA promoters are found for many genes involved in various stages of photosynthesis. Although photosynthesis is known to be tightly coordinated with nitrogen assimilation, very little is known about the underlying mechanism. We postulate for the fist time that these genes serve as the regulatory points to orchestrate these two important processes in a cyanobacterial cell.

Algorithms↗

Prediction of functional modules based on comparative genome analysis and Gene Ontology application.

We present a computational method for the prediction of functional modules encoded in microbial genomes. In this work, we have also developed a formal measure to quantify the degree of consistency between the predicted and the known modules, and have carried out statistical significance analysis of consistency measures. We first evaluate the functional relationship between two genes from three different perspectives--phylogenetic profile analysis, gene neighborhood analysis and Gene Ontology assignments. We then combine the three different sources of information in the framework of Bayesian inference, and we use the combined information to measure the strength of gene functional relationship. Finally, we apply a threshold-based method to predict functional modules. By applying this method to Escherichia coli K12, we have predicted 185 functional modules. Our predictions are highly consistent with the previously known functional modules in E.coli. The application results have demonstrated that our approach is highly promising for the prediction of functional modules encoded in a microbial genome.

Bayes Theorem↗

A store-operated nonselective cation channel in human lymphocytes.

1. Agonist interaction with phospholipase C-linked receptors at the plasma membrane can elicit both Ca2+ and Na+ influxes in lymphocytes. While Ca2+ influx is mediated by Ca2+ release-activated Ca2+ (CRAC) channels, the pathway responsible for Na+ influx is largely unknown. 2. We show that thapsigargin, ionomycin, ADP-ribose and IP3 activated a nonselective cation channel in lymphocytes that had a slightly outwardly rectifying I-V relationship, and a single channel conductance of 23.1 pS. We termed this channel a Ca2+ release-activated nonselective cation (CRANC) channel. 3. On activation in cell-attached configuration, switching to an inside-out configuration abolished CRANC channel activity. 4. Transfection of Jurkat T cells with antisense oligonucleotides for LTRPC2 reduced capacitative Ca2+ entry. 5. These results suggest that CRANC channels are responsible for the Na+ influx as well as a portion of the Ca2+ influx in lymphocytes induced by store depletion, that sustained activation of CRANC channels requires some property of the environment of a cell depleted of its Ca2+ stores; and that LTRPC2 protein is a likely component of the CRANC channel.

Animals↗

Prediction of functional modules based on gene distributions in microbial genomes.

We present a computational method for prediction of functional modules that can be directly applied to the newly sequenced microbial genomes for predicting gene functions and the component genes of biological pathways. We first quantify the functional relatedness among genes based on their distribution (i.e., their existences and orders) across multiple microbial genomes, and obtain a gene network in which every pair of genes is associated with a score representing their functional relatedness. We then apply a threshold-based clustering algorithm to this gene network, and obtain modules for each of which the number of genes is bounded from above by a pre-specified value and the component genes are more strongly functionally related to each other than genes across the predicted modules. Particularly, when the module size is bounded by 130, we obtain 167 functional modules covering 813 genes for Escherichia coli K12, and 138 functional modules covering 731 genes for Bacillus subtilis subsp. subtilis str. 168. We have used the gene ontology (GO) information to assess the prediction results. The GO similarities among the genes of the same functional module are compared with the GO similarities among the genes that are randomly clustered together. This comparison reveals that our predicted functional modules are statistically and biologically significant, and the genes of the same functional module share more commonality in terms of biological process than in terms of molecular function or cellular component. We have also examined the predicted functional modules that are common to both Escherichia coli K12 and Bacillus subtilis subsp. subtilis str. 168, and provide explanations for some functional modules.

Cluster Analysis↗

Ca2+ modulation of Ca2+ release-activated Ca2+ channels is responsible for the inactivation of its monovalent cation current.

The Ca(2+) release-activated Ca(2+) (CRAC) channel is the most well documented of the store-operated ion channels that are widely expressed and are involved in many important biological processes. However, the regulation of the CRAC channel by intracellular or extracellular messengers as well as its molecular identity is largely unknown. Specifically, in the absence of extracellular divalent cations it becomes permeable to monovalent cations with a larger conductance, however this monovalent cation current inactivates rapidly by an unknown mechanism. Here we found that Ca(2+) dissociation from a site on the extracellular side of the CRAC channel is responsible for the inactivation of its Na(+) current, and Ca(2+) occupancy of this site otherwise potentiates its Ca(2+) as well as Na(+) currents. This Ca(2+)-dependent potentiation is required for the normal functioning of CRAC channels.

Animals↗

Mapping of microbial pathways through constrained mapping of orthologous genes.

We present a novel computer algorithm for mapping biological pathways from one prokaryotic genome to another. The algorithm maps genes in a known pathway to their homologous genes (if any) in a target genome that is most consistent with (a) predicted orthologous gene relationship, (b) predicted operon structures, and (c) predicted co-regulation relationship of operons. Mathematically, we have formulated this problem as a constrained minimum spanning tree problem (called a Steiner network problem), and demonstrated that this formulation has the desired property through applications. We have solved this mapping problem using a combinatorial optimization algorithm, with guaranteed global optimality. We have implemented this algorithm as a computer program, called PMAP. Our test results on pathway mapping are highly encouraging -- we have mapped a number of pathways of H. influenzae, B. subtilis, H. pylori, and M. tuberculosis to E. coli using P-MAP, whose homologous pathways in E coli. are known and hence the mapping accuracy could be checked. We have then mapped known E. coli pathways in the EcoCyc database to the newly sequenced organism Synechococcus sp WH8102, and predicted 158 Synechococcus pathways. Detailed analyses on the predicted pathways indicate that P-MAP's mapping results are consistent with our general knowledge about (local) pathways. We believe that P-MAP will be a useful tool for microbial genome annotation projects and inference of individual microbial pathways.

Algorithms↗

Computational prediction of operons in Synechococcus sp. WH8102.

We computationally predict operons in the Synechococcus sp. WH8102 genome based on three types of genomic data: intergenic distances, COG gene functions and phylogenetic profiles. In the proposed method, we first estimate a log-likelihood distribution for each type of genomic data, and then fuse these distribution information by a perceptron to discriminate pairs of genes within operons (WO pairs) from those across transcription unit borders (TUB pairs). Computational experiments demonstrated that WO pairs tend to have shorter intergenic distances, a higher probability being in the same COG functional categories and more similar phylogenetic profiles than TUB pairs, indicating their powerful capabilities for operon prediction. By testing the method on 236 known operons of Escherichia coli K12, an overall accuracy of 83.8% is obtained by joint learning from multiple types of genomic data, whereas individual information source yields accuracies of 80.4%, 74.4%, and 70.6% respectively. We have applied this new approach, in conjunction with our previous comparative genome analysis-based approach, to predict 556 (putative) operons in WH8102. All predicted data are available at (http://www.cs.ucr.edu/~xin/operons.htm) for public use.

Computational Biology↗

Computational inference of regulatory pathways in microbes: an application to phosphorus assimilation pathways in Synechococcus sp. WH8102.

We present a computational protocol for inference of regulatory and signaling pathways in a microbial cell, through literature search, mining "high-throughput'' biological data of various types, and computer-assisted human inference. This protocol consists of four key components: (a) construction of template pathways for microbial organisms related to the target genome, which either have been extensively studied and/or have a significant amount of (relevant) experimental data, (b) inference of initial pathway models for the target genome, through combining the template pathway models and target genome-specific information, (c) refinement and expansion of the initial pathway models through applications of various data mining tools, including phylogenetic profile analysis, inference of protein-protein interactions, and prediction of transcription factor binding sites, and (d) validation and refinement of the pathway models using pathway-specific experimental data or other information. To demonstrate the effectiveness of this procedure, we have applied it to the construction of the phosphorus assimilation pathways in cyanobacterium sp. WH8102. We present, in this paper, a model of the core components of this pathway.

Bacterial Proteins↗

Regulation of Ca2+ release-activated Ca2+ channels by INAD and Ca2+ influx factor.

The coupling mechanism between depletion of Ca(2+) stores in the endoplasmic reticulum and plasma membrane store-operated ion channels is fundamental to Ca(2+) signaling in many cell types and has yet to be completely elucidated. Using Ca(2+) release-activated Ca(2+) (CRAC) channels in RBL-2H3 cells as a model system, we have shown that CRAC channels are maintained in the closed state by an inhibitory factor rather than being opened by the inositol 1,4,5-trisphosphate receptor. This inhibitory role can be fulfilled by the Drosophila protein INAD (inactivation-no after potential D). The action of INAD requires Ca(2+) and can be reversed by a diffusible Ca(2+) influx factor. Thus the coupling between the depletion of Ca(2+) stores and the activation of CRAC channels may involve a mammalian homologue of INAD and a low-molecular-weight, diffusible store-depletion signal.

Animals↗