PubMed Health⌕ Search

Biomedical subjects

Naoki Osato

Publications and source records attributed to Naoki Osato.

8 recordsLinked to original sources

A genome-wide and nonredundant mouse transcription factor database.

Here we describe the development of a genome-wide and nonredundant mouse transcription factor database and its viewer (http://genome.gsc.riken.gp/TFdb/). We systematically selected transcription factors with DNA-binding properties and their regulators on the basis of their LocusLink and Gene Ontology annotations. We also incorporated into our database information regarding the corresponding available cDNA clones and their structural properties. Because of these features, our database is unique and may provide useful information for systematic genome-wide studies of transcriptional regulation.

Animals↗

Identification of region-specific transcription factor genes in the adult mouse brain by medium-scale real-time RT-PCR.

We established a medium-scale real-time RT-PCR system focusing on transcription factors and applied it to their expression profiles in the adult mouse 11 brain regions (http://genome.gsc.riken.jp/qRT-PCR/). Almost 90% of the examined genes showed significant expression in at least one region. We successfully extracted 179 region-specific genes by clustering analysis. Interestingly, the transcription factors involved in the development of the pituitary were still expressed in the adult pituitary, suggesting that they also play important roles in maintenance of the pituitary. These results provide unique molecular markers that may account for the molecular basis of the unique functions of specific brain regions.

Aging↗

Comprehensive analysis of NAC family genes in Oryza sativa and Arabidopsis thaliana.

The NAC domain was originally characterized from consensus sequences from petunia NAM and from Arabidopsis ATAF1, ATAF2, and CUC2. Genes containing the NAC domain (NAC family genes) are plant-specific transcriptional regulators and are expressed in various developmental stages and tissues. We performed a comprehensive analysis of NAC family genes in Oryza sativa (a monocot) and Arabidopsis thaliana (a dicot). We found 75 predicted NAC proteins in full-length cDNA data sets of O. sativa (28,469 clones) and 105 in putative genes (28,581 sequences) from the A. thaliana genome. NAC domains from both predicted and known NAC family proteins were classified into two groups and 18 subgroups by sequence similarity. There were a few differences in amino acid sequences in the NAC domains between O. sativa and A. thaliana. In addition, we found 13 common sequence motifs from transcriptional activation regions in the C-terminal regions of predicted NAC proteins. These motifs probably diverged having correlations with NAC domain structures. We discuss the relationship between the structure and function of the NAC family proteins in light of our results and the published data. Our results will aid further functional analysis of NAC family genes.

Arabidopsis↗

Antisense transcripts with rice full-length cDNAs.

BACKGROUND: Natural antisense transcripts control gene expression through post-transcriptional gene silencing by annealing to the complementary sequence of the sense transcript. Because many genome and mRNA sequences have become available recently, genome-wide searches for sense-antisense transcripts have been reported, but few plant sense-antisense transcript pairs have been studied. The Rice Full-Length cDNA Sequencing Project has enabled computational searching of a large number of plant sense-antisense transcript pairs. RESULTS: We identified sense-antisense transcript pairs from 32,127 full-length rice cDNA sequences produced by this project and public rice mRNA sequences by aligning the cDNA sequences with rice genome sequences. We discovered 687 bidirectional transcript pairs in rice, including sense-antisense transcript pairs. Both sense and antisense strands of 342 pairs (50%) showed homology to at least one expressed sequence tag other than that of the pair. Microarray analysis showed 82 pairs (32%) out of 258 pairs on the microarray were more highly expressed than the median expression intensity of 21,938 rice transcriptional units. Both sense and antisense strands of 594 pairs (86%) had coding potential. CONCLUSIONS: The large number of plant sense-antisense transcript pairs suggests that gene regulation by antisense transcripts occurs in plants and not only in animals. On the basis of our results, experiments should be carried out to analyze the function of plant antisense transcripts.

DNA, Antisense↗

Collection, mapping, and annotation of over 28,000 cDNA clones from japonica rice.

We collected and completely sequenced 28,469 full-length complementary DNA clones from Oryza sativa L. ssp. japonica cv. Nipponbare. Through homology searches of publicly available sequence data, we assigned tentative protein functions to 21,596 clones (75.86%). Mapping of the cDNA clones to genomic DNA revealed that there are 19,000 to 20,500 transcription units in the rice genome. Protein informatics analysis against the InterPro database revealed the existence of proteins presented in rice but not in Arabidopsis. Sixty-four percent of our cDNAs are homologous to Arabidopsis proteins.

Alternative Splicing↗

Targeting a complex transcriptome: the construction of the mouse full-length cDNA encyclopedia.

We report the construction of the mouse full-length cDNA encyclopedia,the most extensive view of a complex transcriptome,on the basis of preparing and sequencing 246 libraries. Before cloning,cDNAs were enriched in full-length by Cap-Trapper,and in most cases,aggressively subtracted/normalized. We have produced 1,442,236 successful 3'-end sequences clustered into 171,144 groups, from which 60,770 clones were fully sequenced cDNAs annotated in the FANTOM-2 annotation. We have also produced 547,149 5' end reads,which clustered into 124,258 groups. Altogether, these cDNAs were further grouped in 70,000 transcriptional units (TU),which represent the best coverage of a transcriptome so far. By monitoring the extent of normalization/subtraction, we define the tentative equivalent coverage (TEC),which was estimated to be equivalent to >12,000,000 ESTs derived from standard libraries. High coverage explains discrepancies between the very large numbers of clusters (and TUs) of this project,which also include non-protein-coding RNAs,and the lower gene number estimation of genome annotations. Altogether,5'-end clusters identify regions that are potential promoters for 8637 known genes and 5'-end clusters suggest the presence of almost 63,000 transcriptional starting points. An estimate of the frequency of polyadenylation signals suggests that at least half of the singletons in the EST set represent real mRNAs. Clones accounting for about half of the predicted TUs await further sequencing. The continued high-discovery rate suggests that the task of transcriptome discovery is not yet complete.

Animals↗

Antisense transcripts with FANTOM2 clone set and their implications for gene regulation.

We have used the FANTOM2 mouse cDNA set (60,770 clones), public mRNA data, and mouse genome sequence data to identify 2481 pairs of sense-antisense transcripts and 899 further pairs of nonantisense bidirectional transcription based upon genomic mapping. The analysis greatly expands the number of known examples of sense-antisense transcript and nonantisense bidirectional transcription pairs in mammals. The FANTOM2 cDNA set appears to contain substantially large numbers of noncoding transcripts suitable for antisense transcript analysis. The average proportion of loci encoding sense-antisense transcript and nonantisense bidirectional transcription pairs on autosomes was 15.1 and 5.4%, respectively. Those on the X chromosome were 6.3 and 4.2%, respectively. Sense-antisense transcript pairs, rather than nonantisense bidirectional transcription pairs, may be less prevalent on the X chromosome, possibly due to X chromosome inactivation. Sense and antisense transcripts tended to be isolated from the same libraries, where nonantisense bidirectional transcription pairs were not apparently coregulated. The existence of large numbers of natural antisense transcripts implies that the regulation of gene expression by antisense transcripts is more common that previously recognized. The viewer showing mapping patterns of sense-antisense transcript pairs and nonantisense bidirectional transcription pairs on the genome and other related statistical data is available on our Web site.

Animals↗

A computer-based method of selecting clones for a full-length cDNA project: simultaneous collection of negligibly redundant and variant cDNAs.

We describe a computer-based method that selects representative clones for full-length sequencing in a full-length cDNA project. Our method classifies end sequences using two kinds of criteria, grouping, and clustering. Grouping places together variant cDNAs, family genes, and cDNAs with sequencing errors. Clustering separates those cDNA clones into distinct clusters. The full-length sequences of the clones selected by grouping are determined preferentially, and then the sequences selected by clustering are determined. Grouping reduced the number of rice cDNA clones for full-length sequencing to 21% and mouse cDNA clones to 25%. Rice full-length sequences selected by grouping showed a 1.07-fold redundancy. Mouse full-length sequences showed a 1.04-fold redundancy, which can be reduced by approximately 30% from the selection using our previous method. To estimate the coverage of unique genes, we used FANTOM (Functional Annotation of RIKEN Mouse cDNA Clones) clusters (). Grouping covered almost all unique genes (93% of FANTOM clusters), and clustering covered all genes. Therefore, our method is useful for the selection of appropriate representative clones for full-length sequencing, thereby greatly reducing the cost, labor, and time necessary for this process.

Animals↗