PubMed Health⌕ Search

Biomedical subjects

Rintaro Saito

Publications and source records attributed to Rintaro Saito.

12 recordsLinked to original sources

Computational analysis of stop codon readthrough in D.melanogaster.

MOTIVATION: Readthrough is an unusual process in which a stop codon is misread or skipped. Recently it has been shown that some translation is regulated by the readthrough reactions although the complete mechanism is not clear. Therefore, the discovery of 'readthrough genes' is important for further investigation of their cellular roles, which may provide additional insights into the mechanism of translational regulation. RESULTS: We constructed a system that lists candidates of readthrough genes based on the existence of a 'protein motif' at the 3' untranslated region (UTR). Using this system, we extracted 85 candidates from 4082 nucleic acid sequences of Drosophila melanogaster in GenBank database. The sequences of these candidates had a slightly more stable secondary structure and different base preferences compared to the non-candidates. As these features are known to have an effect on readthrough events, we would like to suggest that these candidates contain actual readthrough genes. AVAILABILITY: Source code of the system is available upon request.

Algorithms↗

Collection, mapping, and annotation of over 28,000 cDNA clones from japonica rice.

We collected and completely sequenced 28,469 full-length complementary DNA clones from Oryza sativa L. ssp. japonica cv. Nipponbare. Through homology searches of publicly available sequence data, we assigned tentative protein functions to 21,596 clones (75.86%). Mapping of the cDNA clones to genomic DNA revealed that there are 19,000 to 20,500 transcription units in the rice genome. Protein informatics analysis against the InterPro database revealed the existence of proteins presented in rice but not in Arabidopsis. Sixty-four percent of our cDNAs are homologous to Arabidopsis proteins.

Alternative Splicing↗

Construction of reliable protein-protein interaction networks with a new interaction generality measure.

MOTIVATION: Recent screening techniques have made large amounts of protein-protein interaction data available, from which biologically important information such as the function of uncharacterized proteins, the existence of novel protein complexes, and novel signal-transduction pathways can be discovered. However, experimental data on protein interactions contain many false positives, making these discoveries difficult. Therefore computational methods of assessing the reliability of each candidate protein-protein interaction are urgently needed. RESULTS: We developed a new 'interaction generality' measure (IG2) to assess the reliability of protein-protein interactions using only the topological properties of their interaction-network structure. Using yeast protein-protein interaction data, we showed that reliable protein-protein interactions had significantly lower IG2 values than less-reliable interactions, suggesting that IG2 values can be used to evaluate and filter interaction data to enable the construction of reliable protein-protein interaction networks.

Binding Sites↗

Global insights into protein complexes through integrated analysis of the reliable interactome and knockout lethality.

We performed an integrated computational analysis of data derived from a comprehensive set of protein-protein interactions (interactome) and a phenotype dataset on lethality in Saccharomyces cerevisiae. For the analysis, we selected reliable interactome data using our previous 'interaction generality,' a computational approach to assess reliability of interactions. Those efforts gave clear evidence that proteins with lethal phenotypes in knockout studies (lethal proteins) may interact with each other to form functional protein complexes to perform their cellular roles. However, our analysis indicates that interactions between lethal proteins are rather restricted to the same cellular pathway or function, and it is quite unlikely that they interact with other lethal proteins functioning in different cellular roles. Furthermore, our results allowed us predictions on the functions of thus far uncharacterized lethal proteins with an estimated 93% accuracy. Thus, the analysis described in here can provide global insights into the biological features of the protein complexes.

Computational Biology↗

Identification of putative noncoding RNAs among the RIKEN mouse full-length cDNA collection.

With the sequencing and annotation of genomes and transcriptomes of several eukaryotes, the importance of noncoding RNA (ncRNA)-RNA molecules that are not translated to protein products-has become more evident. A subclass of ncRNA transcripts are encoded by highly regulated, multi-exon, transcriptional units, are processed like typical protein-coding mRNAs and are increasingly implicated in regulation of many cellular functions in eukaryotes. This study describes the identification of candidate functional ncRNAs from among the RIKEN mouse full-length cDNA collection, which contains 60,770 sequences, by using a systematic computational filtering approach. We initially searched for previously reported ncRNAs and found nine murine ncRNAs and homologs of several previously described nonmouse ncRNAs. Through our computational approach to filter artifact-free clones that lack protein coding potential, we extracted 4280 transcripts as the largest-candidate set. Many clones in the set had EST hits, potential CpG islands surrounding the transcription start sites, and homologies with the human genome. This implies that many candidates are indeed transcribed in a regulated manner. Our results demonstrate that ncRNAs are a major functional subclass of processed transcripts in mammals.

Animals↗

Inferring higher functional information for RIKEN mouse full-length cDNA clones with FACTS.

FACTS (Functional Association/Annotation of cDNA Clones from Text/Sequence Sources) is a semiautomated knowledge discovery and annotation system that integrates molecular function information derived from sequence analysis results (sequence inferred) with functional information extracted from text. Text-inferred information was extracted from keyword-based retrievals of MEDLINE abstracts and by matching of gene or protein names to OMIM, BIND, and DIP database entries. Using FACTS, we found that 47.5% of the 60,770 RIKEN mouse cDNA FANTOM2 clone annotations were informative for text searches. MEDLINE queries yielded molecular interaction-containing sentences for 23.1% of the clones. When disease MeSH and GO terms were matched with retrieved abstracts, 22.7% of clones were associated with potential diseases, and 32.5% with GO identifiers. A significant number (23.5%) of disease MeSH-associated clones were also found to have a hereditary disease association (OMIM Morbidmap). Inferred neoplastic and nervous system disease represented 49.6% and 36.0% of disease MeSH-associated clones, respectively. A comparison of sequence-based GO assignments with informative text-based GO assignments revealed that for 78.2% of clones, identical GO assignments were provided for that clone by either method, whereas for 21.8% of clones, the assignments differed. In contrast, for OMIM assignments, only 28.5% of clones had identical sequence-based and text-based OMIM assignments. Sequence, sentence, and term-based functional associations are included in the FACTS database (http://facts.gsc.riken.go.jp/), which permits results to be annotated and explored through web-accessible keyword and sequence search interfaces. The FACTS database will be a critical tool for investigating the functional complexity of the mouse transcriptome, cDNA-inferred interactome (molecular interactions), and pathome (pathologies).

Animals↗

CDS annotation in full-length cDNA sequence.

The identification of coding sequences (CDS) is an important step in the functional annotation of genes. CDS prediction for mammalian genes from genomic sequence is complicated by the vast abundance of intergenic sequence in the genome, and provides little information about how different parts of potential CDS regions are expressed. In contrast, mammalian gene CDS prediction from cDNA sequence offers obvious advantages, yet encounters a different set of complexities when performed on high-throughput cDNA (HTC) sequences, such as the set of 60,770 cDNAs isolated from full-length enriched libraries of the FANTOM2 project. We developed a CDS annotation strategy that uses a variety of different CDS prediction programs to annotate the CDS regions of FANTOM2 cDNAs. These include rsCDS, which uses sequence similarity to known proteins; ProCrest; Longest-ORF and Truncated-ORF, which are ab initio based predictors; and finally, DECODER and NCBI CDS predictor, which use a combination of both principles. Aided by graphical displays of these CDS prediction results in the context of other sequence similarity results for each cDNA, FANTOM2 CDS inspection by curators and follow-up quality control procedures resulted in high quality CDS predictions for a total of 14,345 FANTOM2 clones.

Animals↗

Targeting a complex transcriptome: the construction of the mouse full-length cDNA encyclopedia.

We report the construction of the mouse full-length cDNA encyclopedia,the most extensive view of a complex transcriptome,on the basis of preparing and sequencing 246 libraries. Before cloning,cDNAs were enriched in full-length by Cap-Trapper,and in most cases,aggressively subtracted/normalized. We have produced 1,442,236 successful 3'-end sequences clustered into 171,144 groups, from which 60,770 clones were fully sequenced cDNAs annotated in the FANTOM-2 annotation. We have also produced 547,149 5' end reads,which clustered into 124,258 groups. Altogether, these cDNAs were further grouped in 70,000 transcriptional units (TU),which represent the best coverage of a transcriptome so far. By monitoring the extent of normalization/subtraction, we define the tentative equivalent coverage (TEC),which was estimated to be equivalent to >12,000,000 ESTs derived from standard libraries. High coverage explains discrepancies between the very large numbers of clusters (and TUs) of this project,which also include non-protein-coding RNAs,and the lower gene number estimation of genome annotations. Altogether,5'-end clusters identify regions that are potential promoters for 8637 known genes and 5'-end clusters suggest the presence of almost 63,000 transcriptional starting points. An estimate of the frequency of polyadenylation signals suggests that at least half of the singletons in the EST set represent real mRNAs. Clones accounting for about half of the predicted TUs await further sequencing. The continued high-discovery rate suggests that the task of transcriptome discovery is not yet complete.

Animals↗

The mammalian protein-protein interaction database and its viewing system that is linked to the main FANTOM2 viewer.

Here, we describe the development of a mammalian protein-protein interaction (PPI) database and of a PPI Viewer application to display protein interaction networks (http://fantom21.gsc.riken.go.jp/PPI/). In the database, we stored the mammalian PPIs identified through our PPI assays (internal PPIs), as well as those we extracted and processed (external PPIs) from publicly available data sources, the DIP and BIND databases and MEDLINE abstracts by using FACTS, a new functional inference and curation system. We integrated the internal and external PPIs into the PPI database, which is linked to the main FANTOM2 viewer. In addition, we incorporated into the PPI Viewer information regarding the luciferase reporter activity of internal PPIs and the data confidence of external PPIs; these data enable visualization and evaluation of the reliability of each interaction. Using the described system, we successfully identified several interactions of biological significance. Therefore, the PPI Viewer is a useful tool for exploring FANTOM2 clone-related protein interactions and their potential effects on signaling and cellular communication.

Animals↗

Interaction generality, a measurement to assess the reliability of a protein-protein interaction.

Here we introduce the 'interaction generality' measure, a new method for computationally assessing the reliability of protein-protein interactions obtained in biological experiments. This measure is basically the number of proteins involved in a given interaction and also adopts the idea that interactions observed in a complicated interaction network are likely to be true positives. Using a group of yeast protein-protein interactions identified in various biological experiments, we show that interactions with low generalities are more likely to be reproducible in other independent assays. We constructed more reliable networks by eliminating interactions whose generalities were above a particular threshold. The rate of interactions with common cellular roles increased from 63% in the unadjusted estimates to 79% in the refined networks. As a result, the rate of cross-talk between proteins with different cellular roles decreased, enabling very clear predictions of the functions of some unknown proteins. The results suggest that the interaction generality measure will make interaction data more useful in all organisms and may yield insights into the biological roles of the proteins studied.

Computational Biology↗

T2BP, a novel TRAF2 binding protein, can activate NF-kappaB and AP-1 without TNF stimulation.

TRAF2 is a key molecule involved in TNF signaling, which is crucial for the regulation of inflammatory processes. We have identified a novel TRAF2 binding protein, designated as T2BP (TRAF2 binding protein), by a mammalian two-hybrid screening approach. T2BP is a relatively small protein of 184 amino acids, which includes a forkhead-associated domain, the phosphopeptide binding motif. The interaction domain search showed that the TRAF domain in TRAF2 is required for the binding to T2BP whereas almost the entire protein in T2BP binds to TRAF2. The interaction was further confirmed by co-immunoprecipitation. Expression profiling for T2BP and TRAF2 revealed an ubiquitous expression in adult mouse tissues. Overexpression of T2BP in HEK293 cells activated NF-kappaB and AP-1 in a dose dependent manner as well as seen in the TNF-treated control cells. Our results suggest that T2BP is involved in the TNF-mediated signaling by its interaction with TRAF2.

Adaptor Proteins, Signal Transducing↗

Inferring alternative splicing patterns in mouse from a full-length cDNA library and microarray data.

Although many studies on alternative splicing of specific genes have been reported in the literature, the general mechanism that regulates alternative splicing has not been clearly understood. In this study, we systematically aligned each pair of the 21,076 cDNA sequences of Mus musculus, searched for putative alternative splicing patterns, and constructed a list of potential alternative splicing sites. Two cDNAs are suspected to be alternatively spliced and originating from a common gene if they share most of their region with a high degree of sequence homology, but parts of the sequences are very distinctive or deleted in either cDNA. The list contains the following information: (1) tissue, (2) developmental stage, (3) sequences around splice sites, (4) the length of each gapped region, and (5) other comments. The list is available at http://www.bioinfo.sfc.keio.ac.jp/intron. Our results have predicted a number of unreported alternatively spliced genes, some of which are expressed only in a specific tissue or at a specific developmental stage.

Alternative Splicing↗