PubMed Health⌕ Search

Biomedical subjects

Lun Cai

Publications and source records attributed to Lun Cai.

9 recordsLinked to original sources

Profiling Caenorhabditis elegans non-coding RNA expression with a combined microarray.

Small non-coding RNAs (ncRNAs) are encoded by genes that function at the RNA level, and several hundred ncRNAs have been identified in various organisms. Here we describe an analysis of the small non-coding transcriptome of Caenorhabditis elegans, microRNAs excepted. As a substantial fraction of the ncRNAs is located in introns of protein-coding genes in C.elegans, we also analysed the relationship between ncRNA and host gene expression. To this end, we designed a combined microarray, which included probes against ncRNA as well as host gene mRNA transcripts. The microarray revealed pronounced differences in expression profiles, even among ncRNAs with housekeeping functions (e.g. snRNAs and snoRNAs), indicating distinct developmental regulation and stage-specific functions of a number of novel transcripts. Analysis of ncRNA-host mRNA relations showed that the expression of intronic ncRNA loci with conserved upstream motifs was not correlated to (and much higher than) expression levels of their host genes. Even promoter-less intronic ncRNA loci, though showing a clear correlation to host gene expression, appeared to have a surprising amount of 'expressional freedom', depending on host gene function. Taken together, our microarray analysis presents a more complete and detailed picture of a non-coding transcriptome than hitherto has been presented for any other multicellular organism.

Animals↗

Faster and more accurate global protein function assignment from protein interaction networks using the MFGO algorithm.

MOTIVATION: Predicting protein function accurately is an important issue in the post-genomic era. To achieve this goal, several approaches have been proposed deduce the function of unclassified proteins through sequence similarity, co-expression profiles, and other information. Among these methods, the global optimization method (GOM) is an interesting and powerful tool that assigns functions to unclassified proteins based on their positions in a physical interactions network [Vazquez, A., Flammini, A., Maritan, A. and Vespignani, A. (2003) Global protein function prediction from protein-protein interaction networks, Nat. Biotechnol., 21, 697-700]. To boost both the accuracy and speed of GOM, a new prediction method, MFGO (modified and faster global optimization) is presented in this paper, which employs local optimal repetition method to reduce calculation time, and takes account of topological structure information to achieve a more accurate prediction. CONCLUSION: On four proteins interaction datasets, including Vazquez dataset, YP dataset, DIP-core dataset, and SPK dataset, MFGO was tested and compared with the popular MR (majority rule) and GOM methods. Experimental results confirm MFGO's improvement on both speed and accuracy. Especially, MFGO method has a distinctive advantage in accurately predicting functions for proteins with few neighbors. Moreover, the robustness of the approach was validated both in a dataset containing a high percentage of unknown proteins and a disturbed dataset through random insertion and deletion. The analysis shows that a moderate amount of misplaced interactions do not preclude a reliable function assignment.

Algorithms↗

Organization of the Caenorhabditis elegans small non-coding transcriptome: genomic features, biogenesis, and expression.

Recent evidence points to considerable transcription occurring in non-protein-coding regions of eukaryote genomes. However, their lack of conservation and demonstrated function have created controversy over whether these transcripts are functional. Applying a novel cloning strategy, we have cloned 100 novel and 61 known or predicted Caenorhabditis elegans full-length ncRNAs. Studying the genomic environment and transcriptional characteristics have shown that two-thirds of all ncRNAs, including many intronic snoRNAs, are independently transcribed under the control of ncRNA-specific upstream promoter elements. Furthermore, the transcription levels of at least 60% of the ncRNAs vary with developmental stages. We identified two new classes of ncRNAs, stem-bulge RNAs (sbRNAs) and snRNA-like RNAs (snlRNAs), both featuring distinct internal motifs, secondary structures, upstream elements, and high and developmentally variable expression. Most of the novel ncRNAs are conserved in Caenorhabditis briggsae, but only one homolog was found outside the nematodes. Preliminary estimates indicate that the C. elegans transcriptome contains approximately 2700 small non-coding RNAs, potentially acting as regulatory elements in nematode development.

Animals↗

NONCODE: an integrated knowledge database of non-coding RNAs.

NONCODE is an integrated knowledge database dedicated to non-coding RNAs (ncRNAs), that is to say, RNAs that function without being translated into proteins. All ncRNAs in NONCODE were filtered automatically from literature and GenBank, and were later manually curated. The distinctive features of NONCODE are as follows: (i) the ncRNAs in NONCODE include almost all the types of ncRNAs, except transfer RNAs and ribosomal RNAs. (ii) All ncRNA sequences and their related information (e.g. function, cellular role, cellular location, chromosomal information, etc.) in NONCODE have been confirmed manually by consulting relevant literature: more than 80% of the entries are based on experimental data. (iii) Based on the cellular process and function, which a given ncRNA is involved in, we introduced a novel classification system, labeled process function class, to integrate existing classification systems. (iv) In addition, some 1100 ncRNAs have been grouped into nine other classes according to whether they are specific to gender or tissue or associated with tumors and diseases, etc. (v) NONCODE provides a user-friendly interface, a visualization platform and a convenient search option, allowing efficient recovery of sequence, regulatory elements in the flanking sequences, secondary structure, related publications and other information. The first release of NONCODE (v1.0) contains 5339 non-redundant sequences from 861 organisms, including eukaryotes, eubacteria, archaebacteria, virus and viroids. Access is free for all users through a web interface at http://noncode.bioinfo.org.cn.

Base Sequence↗

The interactome as a tree--an attempt to visualize the protein-protein interaction network in yeast.

The refinement and high-throughput of protein interaction detection methods offer us a protein-protein interaction network in yeast. The challenge coming along with the network is to find better ways to make it accessible for biological investigation. Visualization would be helpful for extraction of meaningful biological information from the network. However, traditional ways of visualizing the network are unsuitable because of the large number of proteins. Here, we provide a simple but information-rich approach for visualization which integrates topological and biological information. In our method, the topological information such as quasi-cliques or spoke-like modules of the network is extracted into a clustering tree, where biological information spanning from protein functional annotation to expression profile correlations can be annotated onto the representation of it. We have developed a software named PINC based on our approach. Compared with previous clustering methods, our clustering method ADJW performs well both in retaining a meaningful image of the protein interaction network as well as in enriching the image with biological information, therefore is more suitable in visualization of the network.

Algorithms↗

Date of origin of the SARS coronavirus strains.

BACKGROUND: A new respiratory infectious epidemic, severe acute respiratory syndrome (SARS), broke out and spread throughout the world. By now the putative pathogen of SARS has been identified as a new coronavirus, a single positive-strand RNA virus. RNA viruses commonly have a high rate of genetic mutation. It is therefore important to know the mutation rate of the SARS coronavirus as it spreads through the population. Moreover, finding a date for the last common ancestor of SARS coronavirus strains would be useful for understanding the circumstances surrounding the emergence of the SARS pandemic and the rate at which SARS coronavirus diverge. METHODS: We propose a mathematical model to estimate the evolution rate of the SARS coronavirus genome and the time of the last common ancestor of the sequenced SARS strains. Under some common assumptions and justifiable simplifications, a few simple equations incorporating the evolution rate (K) and time of the last common ancestor of the strains (T0) can be deduced. We then implemented the least square method to estimate K and T0 from the dataset of sequences and corresponding times. Monte Carlo stimulation was employed to discuss the results. RESULTS: Based on 6 strains with accurate dates of host death, we estimated the time of the last common ancestor to be about August or September 2002, and the evolution rate to be about 0.16 base/day, that is, the SARS coronavirus would on average change a base every seven days. We validated our method by dividing the strains into two groups, which coincided with the results from comparative genomics. CONCLUSION: The applied method is simple to implement and avoid the difficulty and subjectivity of choosing the root of phylogenetic tree. Based on 6 strains with accurate date of host death, we estimated a time of the last common ancestor, which is coincident with epidemic investigations, and an evolution rate in the same range as that reported for the HIV-1 virus.

China↗

Topological structure analysis of the protein-protein interaction network in budding yeast.

Interaction detection methods have led to the discovery of thousands of interactions between proteins, and discerning relevance within large-scale data sets is important to present-day biology. Here, a spectral method derived from graph theory was introduced to uncover hidden topological structures (i.e. quasi-cliques and quasi-bipartites) of complicated protein-protein interaction networks. Our analyses suggest that these hidden topological structures consist of biologically relevant functional groups. This result motivates a new method to predict the function of uncharacterized proteins based on the classification of known proteins within topological structures. Using this spectral analysis method, 48 quasi-cliques and six quasi-bipartites were isolated from a network involving 11,855 interactions among 2617 proteins in budding yeast, and 76 uncharacterized proteins were assigned functions.

Algorithms↗

[Sequence analysis of bacterial transposon in NHX gene of Populus euphratica].

The United Nations Environment Program estimates that approximately 20% of agricultural land and 50% of cropland in the world is salt-stressed. The gene NHX (Na+/H+ exchanger) encodes functional protein that catalyzes the countertransport of Na+ and H+ across membranes and may play an important role in plant salt tolerance. To clone the NHX from the wild plant Populus euphratica collected in Tarim basin and Xinjiang Wujiaqu district into a T-vector, designed primer was used to amplify 1kb NHX cDNA fragment with RT-PCR. Total RNA was extracted from Populus euphratica tissue (plant tissue was collected from Tarim basin and Xinjiang Wujiaqu district and stored in liquid nitrogen) according to the Plant RNA Mini Kits of Omega. First cDNAs were synthesized from 1 microg total RNA of Populus euphratica seedling. A pair of primers were used to perform RT-PCR. The amplified DNA fragment was purified and cloned into pMD18-T vector. However, 1kb and 2.3kb fragment were obtained from Tarim basin and Xinjiang Wujiaqu district and named as PtNHX and PwNHX, respectively. Sequence analysis reveals that the cloned PtNHX fragment of Populus euphratica contains partial NHX coding region with 98%, 86%, 84% and 80% identity comparing with Atriplex gemelini, Suaeda maritima, Arabidopsis thaliana and Oryza sativa, respectively. This analysis suggests that NHX gene would be highly conserved in terms of evolution in plant; and it also suggests that the NHX gene of Populus euphratica also would have the similarity with that of Arabidopsis. It may be of great importance in improvement of the plant salt tolerance and breed of crop. At the same time, sequence analysis shows that PwNHX gene includes a coding region about 1350bp with 99% identity comparing with transposon Tn10 IS10-left transposase of Shigella flexneri. On the one hand, the NHX gene may lose its function because it was inserted a fragment in coding region. On the other hand, its product may play a important role in salt tolerance. Populus grow in saline soil. It speculates that it may have other salt tolerance mechanism in Populus. The transposon can be used as transposon tagging to clone other genes and it will help us to understand farther the salt tolerance mechanism.

Amino Acid Sequence↗

[Studies on the properties of Cecropin-XJ expressed in yeast from Xinjiang silkworm].

The purpose of this study is to investigate the properties of recombinant Cecropin-XJ isolated from Xinjing silkworm and expressed in Pichia yeast. According to Agarose Diffusion Assay, this recombinant Cecropin-XJ has exhibited an extreme heat-stable characteristic and the ability to kill ampicillin-resistant S. aureus and S. enterica. Moreover, we have observed that the Cecropin-XJ well tolerant to extreme acidic, basic, and high salt environments as well as resistant to 24 hours digestion by artificial gastric juice. The inhibition capability to S. aureus with 1mg Cecropin-XJ is equal to 1200U ampicillin. With a broad spectrum of antibacterial activities, the Cecropin-XJ is able to inhibit the Gram-positive bacteria and Gram-negative bacteria. These findings could lead it to a broad applications for agriculture, medical, domestic animal and food industry. The further investigation of the antibacterial mechanism of cecropin-XJ is needed.

Animals↗