PubMed Health⌕ Search

Biomedical subjects

Xiaopeng Zhu

Publications and source records attributed to Xiaopeng Zhu.

17 recordsLinked to original sources

Diversity of the O-superfamily conotoxins from Conus miles.

Conopeptides display prominent features of hypervariability and high selectivity of large gene families that mediate interactions between organisms. Remarkable sequence diversity of O-superfamily conotoxins was found in a worm-hunting cone snail Conus miles. Five novel cDNA sequences encoding O-superfamily precursor peptides were identified in C. miles native to Hainan by RT-PCR and 3'-RACE. They share the common cysteine pattern of the O-superfamily conotoxin (C-C-CC-C-C, with three disulfide bridges). The predicted peptides consist of 27-33 amino acids. We then performed a phylogenetic analysis of the new and published homologue sequences from C. miles and the other Conus species. Sequence divergence (%) and residue substitutions to view evolutionary relationships of the precursors' signal, propeptide, and mature toxin regions were analyzed. Percentage divergence of the amino acid sequences of the prepro region exhibited high conservation, whereas the sequences of the mature peptides ranged from almost identical with to highly divergent from inter- and intra-species. Despite the O-superfamily being a large and diverse group of peptides, widely distributed in the venom ducts of all major feeding types of Conus and discovered in several Conus species, it was for the first time that the newly found five O-superfamily peptides in this research came from the vermivorous C. miles. So far, conotoxins of the O-superfamily whose properties have been characterized are from piscivorous and molluscivorous Conus species, and their amino acid sequences and mode of action have been discussed in detail. The elucidated cDNAs of the five toxins are new and of importance and should attract the interest of researchers in the field, which would pave the way for a better understanding of the relationship of their structure and function.

Amino Acid Sequence↗

Sequence diversity of O-superfamily conopetides from Conus marmoreus native to Hainan.

The full-length cDNAs of six new O-superfamily conotoxins (CTX) were cloned and sequenced from Conus marmoreus native to Hainan in China South Sea using RT-PCR and 3'-RACE. Six novel conotoxin precursors encoded by these cDNAs consist of three typical regions of signal, pro-peptide and mature peptide. All the six toxin regions share a common O-superfamily cysteine pattern (C-C-CC-C-C, with three disulfide bridges). The predicted precursors are composed of 73-88 amino acids, and the predicted mature peptides consist of 26-34 amino acids. Phylogenetic analysis of new conotoxins from C. marmoreus from the present study and published homologue T-superfamily sequences from other Conus species was performed systematically. Patterns of sequence divergence for three regions of signal, pro-region and mature peptides, as well as Cys codon usage define the major O-superfamily branches and suggest how these separate branches arose. Percent identities of the amino acid sequences of the signal region exhibited high conservation, whereas the sequences of the mature peptides ranged from almost identical to highly divergent between inter- and intra-species. Notably, the diversity of the pro-region was also high with intermediate divergence between that observed in signal and toxin regions. Amino acid sequences and their mode of action (target) of previously identified conotoxins from molluscivorous C. marmoreus for the known conotoxins classes are discussed in detail. The data presented are new and should pave the way for chemical synthesis of these unique conotoxins for to allow determination of the molecular targets of these peptides, and also to provide clues for a better understanding of the phylogeny of these peptides.

Amino Acid Sequence↗

Prediction of structured non-coding RNAs in the genomes of the nematodes Caenorhabditis elegans and Caenorhabditis briggsae.

We present a survey for non-coding RNAs and other structured RNA motifs in the genomes of Caenorhabditis elegans and Caenorhabditis briggsae using the RNAz program. This approach explicitly evaluates comparative sequence information to detect stabilizing selection acting on RNA secondary structure. We detect 3,672 structured RNA motifs, of which only 678 are known non-translated RNAs (ncRNAs) or clear homologs of known C. elegans ncRNAs. Most of these signals are located in introns or at a distance from known protein-coding genes. With an estimated false positive rate of about 50% and a sensitivity on the order of 50%, we estimate that the nematode genomes contain between 3,000 and 4,000 RNAs with evolutionary conserved secondary structures. Only a small fraction of these belongs to the known RNA classes, including tRNAs, snoRNAs, snRNAs, or microRNAs. A relatively small class of ncRNA candidates is associated with previously observed RNA-specific upstream elements.

Animals↗

Profiling Caenorhabditis elegans non-coding RNA expression with a combined microarray.

Small non-coding RNAs (ncRNAs) are encoded by genes that function at the RNA level, and several hundred ncRNAs have been identified in various organisms. Here we describe an analysis of the small non-coding transcriptome of Caenorhabditis elegans, microRNAs excepted. As a substantial fraction of the ncRNAs is located in introns of protein-coding genes in C.elegans, we also analysed the relationship between ncRNA and host gene expression. To this end, we designed a combined microarray, which included probes against ncRNA as well as host gene mRNA transcripts. The microarray revealed pronounced differences in expression profiles, even among ncRNAs with housekeeping functions (e.g. snRNAs and snoRNAs), indicating distinct developmental regulation and stage-specific functions of a number of novel transcripts. Analysis of ncRNA-host mRNA relations showed that the expression of intronic ncRNA loci with conserved upstream motifs was not correlated to (and much higher than) expression levels of their host genes. Even promoter-less intronic ncRNA loci, though showing a clear correlation to host gene expression, appeared to have a surprising amount of 'expressional freedom', depending on host gene function. Taken together, our microarray analysis presents a more complete and detailed picture of a non-coding transcriptome than hitherto has been presented for any other multicellular organism.

Animals↗

Dynamic changes in subgraph preference profiles of crucial transcription factors.

Transcription factors with a large number of target genes--transcription hub(s), or THub(s)--are usually crucial components of the regulatory system of a cell, and the different patterns through which they transfer the transcriptional signal to downstream cascades are of great interest. By profiling normalized abundances (A(N)) of basic regulatory patterns of individual THubs in the yeast Saccharomyces cerevisiae transcriptional regulation network under five different cellular states and environmental conditions, we have investigated their preferences for different basic regulatory patterns. Subgraph-normalized abundances downstream of individual THubs often differ significantly from that of the network as a whole, and conversely, certain over-represented subgraphs are not preferred by any THub. The THub preferences changed substantially when the cellular or environmental conditions changed. This switching of regulatory pattern preferences suggests that a change in conditions does not only elicit a change in response by the regulatory network, but also a change in the mechanisms by which the response is mediated. The THub subgraph preference profile thus provides a novel tool for description of the structure and organization between the large-scale exponents and local regulatory patterns.

Cell Cycle↗

Phylophenetic properties of metabolic pathway topologies as revealed by global analysis.

BACKGROUND: As phenotypic features derived from heritable characters, the topologies of metabolic pathways contain both phylogenetic and phenetic components. In the post-genomic era, it is possible to measure the "phylophenetic" contents of different pathways topologies from a global perspective. RESULTS: We reconstructed phylophenetic trees for all available metabolic pathways based on topological similarities, and compared them to the corresponding 16S rRNA-based trees. Similarity values for each pair of trees ranged from 0.044 to 0.297. Using the quartet method, single pathways trees were merged into a comprehensive tree containing information from a large part of the entire metabolic networks. This tree showed considerably higher similarity (0.386) to the corresponding 16S rRNA-based tree than any tree based on a single pathway, but was, on the other hand, sufficiently distinct to preserve unique phylogenetic information not reflected by the 16S rRNA tree. CONCLUSION: We observed that the topology of different metabolic pathways provided different phylogenetic and phenetic information, depicting the compromise between phylogenetic information and varying evolutionary pressures forming metabolic pathway topologies in different organisms. The phylogenetic information content of the comprehensive tree is substantially higher than that of any tree based on a single pathway, which also gave clues to constraints working on the topology of the global metabolic networks, information that is only partly reflected by the topologies of individual metabolic pathways.

Chromosome Mapping↗

Integrated analysis of multiple data sources reveals modular structure of biological networks.

It has been a challenging task to integrate high-throughput data into investigations of the systematic and dynamic organization of biological networks. Here, we presented a simple hierarchical clustering algorithm that goes a long way to achieve this aim. Our method effectively reveals the modular structure of the yeast protein-protein interaction network and distinguishes protein complexes from functional modules by integrating high-throughput protein-protein interaction data with the added subcellular localization and expression profile data. Furthermore, we take advantage of the detected modules to provide a reliably functional context for the uncharacterized components within modules. On the other hand, the integration of various protein-protein association information makes our method robust to false-positives, especially for derived protein complexes. More importantly, this simple method can be extended naturally to other types of data fusion and provides a framework for the study of more comprehensive properties of the biological network and other forms of complex networks.

Algorithms↗

A novel scoring schema for peptide identification by searching protein sequence databases using tandem mass spectrometry data.

BACKGROUND: Tandem mass spectrometry (MS/MS) is a powerful tool for protein identification. Although great efforts have been made in scoring the correlation between tandem mass spectra and an amino acid sequence database, improvements could be made in three aspects, including characterization ofpeaks in spectra, adoption of effective scoring functions and access to thereliability of matching between peptides and spectra. RESULTS: A novel scoring function is presented, along with criteria to estimate the performance confidence of the function. Through learning the typesof product ions and the probability of generating them, a hypothetic spectrum was generated for each candidate peptide. Then relative entropy was introduced to measure the similarity between the hypothetic and the observed spectra. Based on the extreme value distribution (EVD) theory, a threshold was chosen to distinguish a true peptide assignment from a random one. Tests on a public MS/MS dataset demonstrated that this method performs better than the well-known SEQUEST. CONCLUSION: A reliable identification of proteins from the spectra promises a more efficient application of tandem mass spectrometry to proteomes with high complexity.

Algorithms↗

Identifying Hfq-binding small RNA targets in Escherichia coli.

The Hfq-binding small RNAs (sRNAs) have recently drawn much attention as regulators of translation in Escherichia coli. We attempt to identify the targets of this class of sRNAs in genome scale and gain further insight into the complexity of translational regulation induced by Hfq-binding sRNAs. Using a new alignment algorithm, most known negatively regulated targets of Hfq-binding sRNAs were identified. The results also show several interesting aspects of the regulatory function of Hfq-binding sRNAs.

Algorithms↗

NPInter: the noncoding RNAs and protein related biomacromolecules interaction database.

The noncoding RNAs and protein related biomacromolecules interaction database (NPInter; http://bioinfo.ibp.ac.cn/NPInter or http://www.bioinfo.org.cn/NPInter) is a database that documents experimentally determined functional interactions between noncoding RNAs (ncRNAs) and protein related biomacromolecules (PRMs) (proteins, mRNAs or genomic DNAs). NPInter intends to provide the scientific community with a comprehensive and integrated tool for efficient browsing and extraction of information on interactions between ncRNAs and PRMs. Beyond cataloguing details of these interactions, the NPInter will be useful for understanding ncRNA function, as it adds a very important functional element, ncRNAs, to the biomolecule interaction network and sets up a bridge between the coding and the noncoding kingdoms.

Animals↗

Identification and molecular diversity of T-superfamily conotoxins from Conus lividus and Conus litteratus.

The T-superfamily conotoxins comprise a large and diverse group of biologically active peptides and are widely distributed in venom ducts of all major feeding types of Conus. Six novel T-superfamily peptides from the two worm-hunting cone snail species of Conus lividus andConus. litteratus native to Hainan were identified and determined to share a common signal sequence as well as a conserved arrangement of cysteine residues (CC-CC). The predicted mature peptides consist of 11-15 amino acids only. Phylogenetic analyses of new conotoxins from C. lividus andC. litteratus in present study and published homologue T-superfamily sequences from the other Conus species was systematically performed. Phylogenetic trees, residue substitutions to view evolutionary relationships of the precursors' signal, propeptide, and mature toxin regions were explored, as well as residue frequency component and cystine codon usage. Percent divergence of the amino acid sequences of the signal-region exhibited high conservation, whereas the sequences of the mature peptides ranged from high similarity to high divergence between inter- and intro-species. Notably, diversity of pro-peptide region was also high with intermediate percent divergence between that observed in signal and toxin-regions. Consensus hydrophobic residues Leu, Val, Ala, Ile and Pro of signal regions were abundant, whereas among propeptides, basic residues Arg and Lys and acidic residue Asp, addition of hydrophilic residues Thr and Ser were abundant. Residue frequency components were hypervariable in mature toxin region except for highly conservative cystine frame residues. The T-superfamily conotoxins have been previously found mainly in piscivorous and molluscivorous cone snails. The newly identified six T-superfamily peptides described in this investigation exemplify the first to be found from vermivorousC. lividus andC. litteratus. The elucidated cDNAs of the six toxins will facilitate a better understanding of the relationship between structure and function as well as provide a framework for their further research and development.

Amino Acid Sequence↗

Novel O-superfamily conotoxins identified by cDNA cloning from three vermivorous Conus species.

The O-superfamily of conotoxins includes several subfamilies with different pharmacological targets, all of which are voltage-gated ion channels and distributed widely in varied Conus species. The venom components from any Conus species are quite distinct from those of other species. Seven novel O-superfamily peptides were identified by cDNA cloning from the three vermivorous Conus species of C. betulinus, C. lividus and C. caracteristicus native to Hainan. They share three common signal sequences, and a conserved arrangement of cysteine residues (C-C-CC-C-C). Phylogenetic analysis of newly found conotoxins in this study and known homologue O-superfamily sequences from the other Conus species was performed systematically. Divergence and percentage identity of the amino acid sequences of the signal regions suggest that the novel conotoxins described in this investigation belong to the three broad clades: MSGL, ME-QK and MKLT, each of which has its own characteristic signature signal sequence and cysteine codon conservation. Relative to this work, it is noted that O-superfamily conotoxins are not well represented from vermivorous species. The elucidated cDNAs of these newly found vermivorous toxins would facilitate a better understanding for basic research and drug discovery.

Animals↗

Organization of the Caenorhabditis elegans small non-coding transcriptome: genomic features, biogenesis, and expression.

Recent evidence points to considerable transcription occurring in non-protein-coding regions of eukaryote genomes. However, their lack of conservation and demonstrated function have created controversy over whether these transcripts are functional. Applying a novel cloning strategy, we have cloned 100 novel and 61 known or predicted Caenorhabditis elegans full-length ncRNAs. Studying the genomic environment and transcriptional characteristics have shown that two-thirds of all ncRNAs, including many intronic snoRNAs, are independently transcribed under the control of ncRNA-specific upstream promoter elements. Furthermore, the transcription levels of at least 60% of the ncRNAs vary with developmental stages. We identified two new classes of ncRNAs, stem-bulge RNAs (sbRNAs) and snRNA-like RNAs (snlRNAs), both featuring distinct internal motifs, secondary structures, upstream elements, and high and developmentally variable expression. Most of the novel ncRNAs are conserved in Caenorhabditis briggsae, but only one homolog was found outside the nematodes. Preliminary estimates indicate that the C. elegans transcriptome contains approximately 2700 small non-coding RNAs, potentially acting as regulatory elements in nematode development.

Animals↗

Performance analysis of very-small-aperture lasers.

The fabrication of very-small-aperture lasers is demonstrated, and their performance is analyzed. Because of strong optical feedback caused by a gold film on the front facet of the laser, its behavior changes: The threshold current decreases, the density of light inside the laser diode and the redshift effect of the spectra are enhanced, and the laser diode's lifetime is shorter than that of common laser diodes with large driving current.

Journal Article↗

The interactome as a tree--an attempt to visualize the protein-protein interaction network in yeast.

The refinement and high-throughput of protein interaction detection methods offer us a protein-protein interaction network in yeast. The challenge coming along with the network is to find better ways to make it accessible for biological investigation. Visualization would be helpful for extraction of meaningful biological information from the network. However, traditional ways of visualizing the network are unsuitable because of the large number of proteins. Here, we provide a simple but information-rich approach for visualization which integrates topological and biological information. In our method, the topological information such as quasi-cliques or spoke-like modules of the network is extracted into a clustering tree, where biological information spanning from protein functional annotation to expression profile correlations can be annotated onto the representation of it. We have developed a software named PINC based on our approach. Compared with previous clustering methods, our clustering method ADJW performs well both in retaining a meaningful image of the protein interaction network as well as in enriching the image with biological information, therefore is more suitable in visualization of the network.

Algorithms↗

Date of origin of the SARS coronavirus strains.

BACKGROUND: A new respiratory infectious epidemic, severe acute respiratory syndrome (SARS), broke out and spread throughout the world. By now the putative pathogen of SARS has been identified as a new coronavirus, a single positive-strand RNA virus. RNA viruses commonly have a high rate of genetic mutation. It is therefore important to know the mutation rate of the SARS coronavirus as it spreads through the population. Moreover, finding a date for the last common ancestor of SARS coronavirus strains would be useful for understanding the circumstances surrounding the emergence of the SARS pandemic and the rate at which SARS coronavirus diverge. METHODS: We propose a mathematical model to estimate the evolution rate of the SARS coronavirus genome and the time of the last common ancestor of the sequenced SARS strains. Under some common assumptions and justifiable simplifications, a few simple equations incorporating the evolution rate (K) and time of the last common ancestor of the strains (T0) can be deduced. We then implemented the least square method to estimate K and T0 from the dataset of sequences and corresponding times. Monte Carlo stimulation was employed to discuss the results. RESULTS: Based on 6 strains with accurate dates of host death, we estimated the time of the last common ancestor to be about August or September 2002, and the evolution rate to be about 0.16 base/day, that is, the SARS coronavirus would on average change a base every seven days. We validated our method by dividing the strains into two groups, which coincided with the results from comparative genomics. CONCLUSION: The applied method is simple to implement and avoid the difficulty and subjectivity of choosing the root of phylogenetic tree. Based on 6 strains with accurate date of host death, we estimated a time of the last common ancestor, which is coincident with epidemic investigations, and an evolution rate in the same range as that reported for the HIV-1 virus.

China↗

Topological structure analysis of the protein-protein interaction network in budding yeast.

Interaction detection methods have led to the discovery of thousands of interactions between proteins, and discerning relevance within large-scale data sets is important to present-day biology. Here, a spectral method derived from graph theory was introduced to uncover hidden topological structures (i.e. quasi-cliques and quasi-bipartites) of complicated protein-protein interaction networks. Our analyses suggest that these hidden topological structures consist of biologically relevant functional groups. This result motivates a new method to predict the function of uncharacterized proteins based on the classification of known proteins within topological structures. Using this spectral analysis method, 48 quasi-cliques and six quasi-bipartites were isolated from a network involving 11,855 interactions among 2617 proteins in budding yeast, and 76 uncharacterized proteins were assigned functions.

Algorithms↗