PubMed Health⌕ Search

Biomedical subjects

Jingchu Luo

Publications and source records attributed to Jingchu Luo.

12 recordsLinked to original sources

Systematic high-yield production of human secreted proteins in Escherichia coli.

Human secreted proteins play a very important role in signal transduction. In order to study all potential secreted proteins identified from the human genome sequence, systematic production of large amounts of biologically active secreted proteins is a prerequisite. We selected 25 novel genes as a trial case for establishing a reliable expression system to produce active human secreted proteins in Escherichia coli. Expression of proteins with or without signal peptides was examined and compared in E. coli strains. The results indicated that deletion of signal peptides, to a certain extent, can improve the expression of these proteins and their solubilities. More importantly, under expression conditions such as induction temperature, N-terminus fusion peptides need to be optimized in order to express adequate amounts of soluble proteins. These recombinant proteins were characterized as well-folded proteins. This system enables us to rapidly obtain soluble and highly purified human secreted proteins for further functional studies.

Chromosome Mapping↗

GBA server: EST-based digital gene expression profiling.

Expressed Sequence Tag-based gene expression profiling can be used to discover functionally associated genes on a large scale. Currently available web servers and tools focus on finding differentially expressed genes in different samples or tissues rather than finding co-expressed genes. To fill this gap, we have developed a web server that implements the GBA (Guilt-by-Association) co-expression algorithm, which has been successfully used in finding disease-related genes. We have also annotated UniGene clusters with links to several important databases such as GO, KEGG, OMIM, Gene, IPI and HomoloGene. The GBA server can be accessed and downloaded at http://gba.cbi.pku.edu.cn.

Algorithms↗

DATF: a database of Arabidopsis transcription factors.

UNLABELLED: We have probably developed the most comprehensive database of Arabidopsis transcription factors (DATF). The DATF contains known and predicted Arabidopsis transcription factors (1827 genes in 56 families) with the unique information of 1177 cloned sequences and many other features including 3D structure templates, EST expression information, transcription factor binding sites and nuclear location signals. AVAILABILITY: DATF is freely available at http://datf.cbi.pku.edu.cn

Amino Acid Sequence↗

SPD--a web-based secreted protein database.

With the improved secreted protein prediction approach and comprehensive data sources, including Swiss-Prot, TrEMBL, RefSeq, Ensembl and CBI-Gene, we have constructed secretomes of human, mouse and rat, with a total of 18 152 secreted proteins. All the entries are ranked according to the prediction confidence. They were further annotated via a proteome annotation pipeline that we developed. We also set up a secreted protein classification pipeline and classified our predicted secreted proteins into different functional categories. To make the dataset more convincing and comprehensive, nine reference datasets are also integrated, such as the secreted proteins from the Gene Ontology Annotation (GOA) system at the European Bioinformatics Institute, and the vertebrate secreted proteins from Swiss-Prot. All these entries were grouped via a TribeMCL based clustering pipeline. We have constructed a web-based secreted protein database, which has been publicly available at http://spd.cbi.pku.edu.cn. Users can browse the database via a GO assignment or chromosomal-location-based interface. Moreover, text query and sequence similarity search are also provided, and the sequence and annotation data can be downloaded freely from the SPD website.

Animals↗

Identification of a naturally occurring recombinant isolate of Sugarcane mosaic virus causing maize dwarf mosaic disease.

The complete nucleotide sequence of a potyvirus causing severe maize dwarf mosaic disease in Shaanxi province, northwestern China was determined (GenBank accession No. AY569692). The full genome is 9596 nucleotides in length excluding the 3 '-terminal poly (A) sequence. It contains a large open reading frame (ORF) flanked by a 149 nt 5'-untranslated region (UTR) and a 255 nt 3'-UTR. The putative polyprotein encoded by this large ORF comprises of 3063 amino acid residues. Sequence comparisons and phylogenetic analyses showed that this potyvirus is an isolate of Sugarcane mosaic virus (SCMV). The entire sequences shared identities of 89.6-97.6 % and 79.3-93.3% with 9 sequenced SCMV isolates at the nucleotide and deduced amino acid levels, respectively. But it showed much lower identities with Maize dwarf mosaic virus (MDMV), Sorghum mosaic virus (SrMV) and Johnsongrass mosaic virus (JGMV) isolates. The putative coat protein sequence is identical to that of a Chinese maize isolate SCMV-HZ. However, partition comparisons and phylogenetic profile analyses of the viral nucleotide sequences indicated that it is a recombinant isolate of SCMV. The recombination sites are located within the 6K1 and CI coding regions.

3' Untranslated Regions↗

Duplication and DNA segmental loss in the rice genome: implications for diploidization.

* Large-scale duplication events have been recently uncovered in the rice genome, but different interpretations were proposed regarding the extent of the duplications. * Through analysing the 370 Mb genome sequences assembled into 12 chromosomes of Oryza sativa subspecies indica, we detected 10 duplicated blocks on all 12 chromosomes that contained 47% of the total predicted genes. Based on the phylogenetic analysis, we inferred that this was a result of a genome duplication that occurred c. 70 million years ago, supporting the polyploidy origin of the rice genome. In addition, a segmental duplication was also identified involving chromosomes 11 and 12, which occurred c. 5 million years ago. * Following the duplications, there have been large-scale chromosomal rearrangements and deletions. About 30-65% of duplicated genes were lost shortly after the duplications, leading to a rapid diploidization. * Together with other lines of evidence, we propose that polyploidization is still an ongoing process in grasses of polyploidy origins.

Biological Evolution↗

RDfolder: a web server for prediction of RNA secondary structure.

Prediction of RNA secondary structure is important in the functional analysis of RNA molecules. The RDfolder web server described in this paper provides two methods for prediction of RNA secondary structure: random stacking of helical regions and helical regions distribution. The random stacking method predicts secondary structure by Monte Carlo simulations. The method of helical regions distribution predicts secondary structure based on the helices that appear most frequently in the set of structures, which are generated by the random stacking method. The RDfolder web server can be accessed at http://rna.cbi.pku.edu.cn.

Internet↗

PCAS--a precomputed proteome annotation database resource.

BACKGROUND: Many model proteomes or "complete" sets of proteins of given organisms are now publicly available. Much effort has been invested in computational annotation of those "draft" proteomes. Motif or domain based algorithms play a pivotal role in functional classification of proteins. Employing most available computational algorithms, mainly motif or domain recognition algorithms, we set up to develop an online proteome annotation system with integrated proteome annotation data to complement existing resources. RESULTS: We report here the development of PCAS (ProteinCentric Annotation System) as an online resource of pre-computed proteome annotation data. We applied most available motif or domain databases and their analysis methods, including hmmpfam search of HMMs in Pfam, SMART and TIGRFAM, RPS-PSIBLAST search of PSSMs in CDD, pfscan of PROSITE patterns and profiles, as well as PSI-BLAST search of SUPERFAMILY PSSMs. In addition, signal peptide and TM are predicted using SignalP and TMHMM respectively. We mapped SUPERFAMILY and COGs to InterPro, so the motif or domain databases are integrated through InterPro. PCAS displays table summaries of pre-computed data and a graphical presentation of motifs or domains relative to the protein. As of now, PCAS contains human IPI, mouse IPI, and rat IPI, A. thaliana, C. elegans, D. melanogaster, S. cerevisiae, and S. pombe proteome.PCAS is available at http://pak.cbi.pku.edu.cn/proteome/gca.php CONCLUSION: PCAS gives better annotation coverage for model proteomes by employing a wider collection of available algorithms. Besides presenting the most confident annotation data, PCAS also allows customized query so users can inspect statistically less significant boundary information as well. Therefore, besides providing general annotation information, PCAS could be used as a discovery platform. We plan to update PCAS twice a year. We will upgrade PCAS when new proteome annotation algorithms identified.

Algorithms↗

PepPat, a pattern-based oligopeptide homology search method and the identification of a novel tachykinin-like peptide.

UNLABELLED: PepPat, a hybrid method that combines pattern matching with similarity scoring, is described. We also report PepPat's application in the identification of a novel tachykinin-like peptide. PepPat takes as input a query peptide and a user-specified regular expression pattern within the peptide. It first performs a database pattern match and then ranks candidates on the basis of their similarity to the query peptide. PepPat calculates similarity over the pattern spanning region, enhancing PepPat's sensitivity for short query peptides. PepPat can also search for a user-specified number of occurrences of a repeated pattern within the target sequence. We illustrate PepPat's application in short peptide ligand mining. As a validation example, we report the identification of a novel tachykinin-like peptide, C14TKL-1, and show it is an NK1 (neuokinin receptor 1) agonist whose message is widely expressed in human periphery. AVAILABILITY: PepPat is offered online at: http://peppat.cbi.pku.edu.cn.

Algorithms↗

Secreted protein prediction system combining CJ-SPHMM, TMHMM, and PSORT.

To increase the coverage of secreted protein prediction, we describe a combination strategy. Instead of using a single method, we combine Hidden Markov Model (HMM)-based methods CJ-SPHMM and TMHMM with PSORT in secreted protein prediction. CJ-SPHMM is an HMM-based signal peptide prediction method, while TMHMM is an HMM-based transmembrane (TM) protein prediction algorithm. With CJ-SPHMM and TMHMM, proteins with predicted signal peptide and without predicted TM regions are taken as putative secreted proteins. This HMM-based approach predicts secreted protein with Ac (Accuracy) at 0.82 and Cc (Correlation coefficient) at 0.75, which are similar to PSORT with Ac at 0.82 and Cc at 0.76. When we further complement the HMM-based method, i.e., CJ-SPHMM + TMHMM with PSORT in secreted protein prediction, the Ac value is increased to 0.86 and the Cc value is increased to 0.81. Taking this combination strategy to search putative secreted proteins from the International Protein Index (IPI) maintained at the European Bioinformatics Institute (EBI), we constructed a putative human secretome with 5235 proteins. The prediction system described here can also be applied to predicting secreted proteins from other vertebrate proteomes.

Computational Biology↗

PGAAS: a prokaryotic genome assembly assistant system.

MOTIVATION: In order to accelerate the finishing phase of genome assembly, especially for the whole genome shotgun approach of prokaryotic species, we have developed a software package designated prokaryotic genome assembly assistant system (PGAAS). The approach upon which PGAAS is based is to confirm the order of contigs and fill gaps between contigs through peptide links obtained by searching each contig end with BLASTX against protein databases. RESULTS: We used the contig dataset of the cyanobacterium Synechococcus sp. strain PCC7002 (PCC7002), which was sequenced with six-fold coverage and assembled using the Phrap package. The subject database is the protein database of the cyanobacterium, Synechocystis sp. strain PCC6803 (PCC6803). We found more than 100 non-redundant peptide segments which can link at least 2 contigs. We tested one pair of linked contigs by sequencing and obtained satisfactory result. PGAAS provides a graphic user interface to show the bridge peptides and pier contigs. We integrated Primer3 into our package to design PCR primers at the adjacent ends of the pier contigs. AVAILABILITY: We tested PGAAS on a Linux (Redhat 6.2) PC machine. It is developed with free software (MySQL, PHP and Apache). The whole package is distributed freely and can be downloaded as UNIX compress file: ftp://ftp.cbi.pku.edu.cn/pub/software/unix/pgaas1.0.tar.gz. The package is being continually updated.

Algorithms↗