PubMed Health⌕ Search

Biomedical subjects

Yuri Kawai

Publications and source records attributed to Yuri Kawai.

3 recordsLinked to original sources

Signal sequence and keyword trap in silico for selection of full-length human cDNAs encoding secretion or membrane proteins from oligo-capped cDNA libraries.

We have developed an in silico method of selection of human full-length cDNAs encoding secretion or membrane proteins from oligo-capped cDNA libraries. Fullness rates were increased to about 80% by combination of the oligo-capping method and ATGpr, software for prediction of translation start point and the coding potential. Then, using 5'-end single-pass sequences, cDNAs having the signal sequence were selected by PSORT ('signal sequence trap'). We also applied 'secretion or membrane protein-related keyword trap' based on the result of BLAST search against the SWISS-PROT database for the cDNAs which could not be selected by PSORT. Using the above procedures, 789 cDNAs were primarily selected and subjected to full-length sequencing, and 334 of these cDNAs were finally selected as novel. Most of the cDNAs (295 cDNAs: 88.3%) were predicted to encode secretion or membrane proteins. In particular, 165(80.5%) of the 205 cDNAs selected by PSORT were predicted to have signal sequences, while 70 (54.2%) of the 129 cDNAs selected by 'keyword trap' preserved the secretion or membrane protein-related keywords. Many important cDNAs were obtained, including transporters, receptors, and ligands, involved in significant cellular functions. Thus, an efficient method of selecting secretion or membrane protein-encoding cDNAs was developed by combining the above four procedures.

5' Flanking Region↗

Complete sequencing and characterization of 21,243 full-length human cDNAs.

As a base for human transcriptome and functional genomics, we created the "full-length long Japan" (FLJ) collection of sequenced human cDNAs. We determined the entire sequence of 21,243 selected clones and found that 14,490 cDNAs (10,897 clusters) were unique to the FLJ collection. About half of them (5,416) seemed to be protein-coding. Of those, 1,999 clusters had not been predicted by computational methods. The distribution of GC content of nonpredicted cDNAs had a peak at approximately 58% compared with a peak at approximately 42%for predicted cDNAs. Thus, there seems to be a slight bias against GC-rich transcripts in current gene prediction procedures. The rest of the cDNAs unique to the FLJ collection (5,481) contained no obvious open reading frames (ORFs) and thus are candidate noncoding RNAs. About one-fourth of them (1,378) showed a clear pattern of splicing. The distribution of GC content of noncoding cDNAs was narrow and had a peak at approximately 42%, relatively low compared with that of protein-coding cDNAs.

Chromosomes, Human, 21-22 and Y↗

Database and analysis system for cDNA clones obtained from full-length enriched cDNA libraries.

We have developed an efficient sequence-analysis system and a database system for clones obtained from full-length enriched cDNA libraries made by using the oligo-capping method. We developed a semi-automatic analysis system for 5'- and 3'-end sequences. It pre-processes raw sequences (vector cut and accurate-sequence region extraction), clusters the sequences, searches for similarities through public databases, annotates completeness of clones and analyzes the ORFs in the sequences. Newly developed or improved programs are used in each step. A new program, ESTiMateFull is used to evaluate and to predict the sequence-fullness based on comparisons with mRNA and EST sequences, respectively. The ATGpr program is used to predict sequence-fullness based on statistical information. The combination of full-length enriched cDNA clones and ATGpr fullness prediction resulted in 70% accuracy in the specificity and the sensitivity of the fullness predictions. For the ORFs predicted by the ATGpr, the signal peptides are predicted and a motif search is performed by our new system. We also developed a program that assembles our sequences with dbEST sequences and developed a system to retrieve clones by the characteristics of the ORFs. As keywords, combination of various results of the analyses can be used for retrieval. And various results such as ORF features and database search results can be shown on the same screen by multiple displays. Full-length clones having interesting functions can thus be retrieved efficiently by using this system.

Amino Acid Sequence↗