PubMed Health⌕ Search

Biomedical subjects

Shuichi Kawashima

Publications and source records attributed to Shuichi Kawashima.

11 recordsLinked to original sources

EGassembler: online bioinformatics service for large-scale processing, clustering and assembling ESTs and genomic DNA fragments.

Expressed sequence tag (EST) sequencing has proven to be an economically feasible alternative for gene discovery in species lacking a draft genome sequence. Ongoing large-scale EST sequencing projects feel the need for bioinformatics tools to facilitate uniform EST handling. This brings about a renewed importance for a universal tool for processing and functional annotation of large sets of ESTs. EGassembler (http://egassembler.hgc.jp/) is a web server, which provides an automated as well as a user-customized analysis tool for cleaning, repeat masking, vector trimming, organelle masking, clustering and assembling of ESTs and genomic fragments. The web server is publicly available and provides the community a unique all-in-one online application web service for large-scale ESTs and genomic DNA clustering and assembling. Running on a Sun Fire 15K supercomputer, a significantly large volume of data can be processed in a short period of time. The results can be used to functionally annotate genes, to facilitate splice alignment analysis, to link the transcripts to genetic and physical maps, design microarray chips, to perform transcriptome analysis and to map to KEGG metabolic pathways. The service provides an excellent bioinformatics tool to research groups in wet-lab as well as an all-in-one-tool for sequence handling to bioinformatics researchers.

Computational Biology↗

ODB: a database of operons accumulating known operons across multiple genomes.

Operon structures play an important role in co-regulation in prokaryotes. Although over 200 complete genome sequences are now available, databases providing genome-wide operon information have been limited to certain specific genomes. Thus, we have developed an ODB (Operon DataBase), which provides a data retrieval system of known operons among the many complete genomes. Additionally, putative operons that are conserved in terms of known operons are also provided. The current version of our database contains about 2000 known operon information in more than 50 genomes and about 13 000 putative operons in more than 200 genomes. This system integrates four types of associations: genome context, gene co-expression obtained from microarray data, functional links in biological pathways and the conservation of gene order across the genomes. These associations are indicators of the genes that organize an operon, and the combination of these indicators allows us to predict more reliable operons. Furthermore, our system validates these predictions using known operon information obtained from the literature. This database integrates known literature-based information and genomic data. In addition, it provides an operon prediction tool, which make the system useful for both bioinformatics researchers and experimental biologists. Our database is accessible at http://odb.kuicr.kyoto-u.ac.jp/.

Databases, Nucleic Acid↗

From genomics to chemical genomics: new developments in KEGG.

The increasing amount of genomic and molecular information is the basis for understanding higher-order biological systems, such as the cell and the organism, and their interactions with the environment, as well as for medical, industrial and other practical applications. The KEGG resource (http://www.genome.jp/kegg/) provides a reference knowledge base for linking genomes to biological systems, categorized as building blocks in the genomic space (KEGG GENES) and the chemical space (KEGG LIGAND), and wiring diagrams of interaction networks and reaction networks (KEGG PATHWAY). A fourth component, KEGG BRITE, has been formally added to the KEGG suite of databases. This reflects our attempt to computerize functional interpretations as part of the pathway reconstruction process based on the hierarchically structured knowledge about the genomic, chemical and network spaces. In accordance with the new chemical genomics initiatives, the scope of KEGG LIGAND has been significantly expanded to cover both endogenous and exogenous molecules. Specifically, RPAIR contains curated chemical structure transformation patterns extracted from known enzymatic reactions, which would enable analysis of genome-environment interactions, such as the prediction of new reactions and new enzyme genes that would degrade new environmental compounds. Additionally, drug information is now stored separately and linked to new KEGG DRUG structure maps.

Biotransformation↗

Extracting sequence motifs and the phylogenetic features of SNARE-dependent membrane traffic.

The SNARE proteins are required for membrane fusion during intracellular vesicular transport and for its specificity. Only the unique combination of SNARE proteins (cognates) can be bound and can lead to membrane fusion, although the characteristics of the possible specificity of the binding combinations encoded in the SNARE sequences have not yet been determined. We discovered by whole genome sequence analysis that sequence motifs (conserved sequences) in the SNARE motif domains for each protein group correspond to localization sites or transport pathways. We claim that these motifs reflect the specificity of the binding combinations of SNARE motif domains. Using these motifs, we could classify SNARE proteins from 48 organisms into their localization sites or transport pathways. The classification result shows that more than 10 SNARE subgroups are kingdom specific and that the SNARE paralogs involved in the plasma membrane-related transport pathways have developed greater variations in higher animals and higher plants than those involved in the endoplasmic reticulum-related transport pathways throughout eukaryotic evolution.

Amino Acid Motifs↗

Alteration of gene expression in human hepatocellular carcinoma with integrated hepatitis B virus DNA.

PURPOSE: Integration of hepatitis B virus (HBV) DNA into the human genome is one of the most important steps in HBV-related carcinogenesis. This study attempted to find the link between HBV DNA, the adjoining cellular sequence, and altered gene expression in hepatocellular carcinoma (HCC) with integrated HBV DNA. EXPERIMENTAL DESIGN: We examined 15 cases of HCC infected with HBV by cassette ligation-mediated PCR. The human DNA adjacent to the integrated HBV DNA was sequenced. Protein coding sequences were searched for in the human sequence. In five cases with HBV DNA integration, from which good quality RNA was extracted, gene expression was examined by cDNA microarray analysis. RESULTS: The human DNA sequence successive to integrated HBV DNA was determined in the 15 HCCs. Eight protein-coding regions were involved: ras-responsive element binding protein 1, calmodulin 1, mixed lineage leukemia 2 (MLL2), FLJ333655, LOC220272, LOC255345, LOC220220, and LOC168991. The MLL2 gene was expressed in three cases with HBV DNA integrated into exon 3 of MLL2 and in one case with HBV DNA integrated into intron 3 of MLL2. Gene expression analysis suggested that two HCCs with HBV integrated into MLL2 had similar patterns of gene expression compared with three HCCs with HBV integrated into other loci of human chromosomes. CONCLUSIONS: HBV DNA was integrated at random sites of human DNA, and the MLL2 gene was one of the targets for integration. Our results suggest that HBV DNA might modulate human genes near integration sites, followed by integration site-specific expression of such genes during hepatocarcinogenesis.

Adult↗

Conservation of gene co-regulation between two prokaryotes: Bacillus subtilis and Escherichia coli.

We measured conservation of gene co-regulation between two distantly related prokaryotes, B. subtilis and E. coli. The co-regulation between genes was extracted from knowledge of regulation of genes stored in databases. For B. subtilis operons, we obtained the data set from ODB which we have developed and, for the regulons, we used DBTBS. For E. coli data set, we used known regulons derived from RegulonDB. We obtained a reliable data set of co-regulated genes in B. subtilis and E. coli. About 60-80 % of gene pairs conserved co-regulation relationships, so co-regulation between genes are highly conserved even between distantly related species. To measure the functional relationship between these conserved genes, we used KEGG PATHWAY and COG. When two co-regulated genes are in the same biological pathway in KEGG or share the same functional category in COG, we assume that they have the same function. As a result, we also found that many conserved co-regulated gene pairs share the same functions. These observations would help to predict gene co-regulation and protein functions.

Bacillus subtilis↗

Comprehensive analysis and prediction of synthetic lethality using subcellular locations.

The lethality of a gene is a fundamental and representative measure for understanding the function of a gene and its associated bio-systems. Recently, many research groups have started focusing on the concept of synthetic lethality. The synthetic lethality between genes is defined by the combination of mutations in two genes causing cell death. Here, we confirm that synthetic lethality and cellular location have close relationships among the Saccharomyces cerevisiae genes. Furthermore, we attempt the prediction of candidate gene pairs with synthetic lethality. The prediction is based on the hierarchical aspect model (HAM) which learns from a data set of cellular location to estimate a likelihood value indicating the synthetic lethality between genes.

Cell Death↗

Autoimmune diseases and peptide variations.

The immune system plays an essential role in the defense of the host against invaders. Enormous numbers of lymphocytes are recruited and immense numbers of antibodies or cytokines are secreted in various kinds of immune response. But the system also has the possibility of being the cause of tissue injury or some kind of diseases. For example, when their functions target their host, an autoimmune disease occurs. Although the pathogenesis of various autoimmune diseases has been scrutinized intensively, there is little evidence as of yet. But it has been reported that in most of the disease subjects, a broad spectrum of antibodies recognizing components of self tissue or circulating self antigens that normally should be ignored are observed. In this study, we come to the conclusion that proteins targeted by these autoreactive antibodies share the same peptides with some kind of proteins of viruses known to infect human. This result supports the fact that viral infection is a speculative cause of the disease in some subjects.

Amino Acid Motifs↗

The KEGG resource for deciphering the genome.

A grand challenge in the post-genomic era is a complete computer representation of the cell and the organism, which will enable computational prediction of higher-level complexity of cellular processes and organism behavior from genomic information. Toward this end we have been developing a knowledge-based approach for network prediction, which is to predict, given a complete set of genes in the genome, the protein interaction networks that are responsible for various cellular processes. KEGG at http://www.genome.ad.jp/kegg/ is the reference knowledge base that integrates current knowledge on molecular interaction networks such as pathways and complexes (PATHWAY database), information about genes and proteins generated by genome projects (GENES/SSDB/KO databases) and information about biochemical compounds and reactions (COMPOUND/GLYCAN/REACTION databases). These three types of database actually represent three graph objects, called the protein network, the gene universe and the chemical universe. New efforts are being made to abstract knowledge, both computationally and manually, about ortholog clusters in the KO (KEGG Orthology) database, and to collect and analyze carbohydrate structures in the GLYCAN database.

Animals↗

Update of MAGEST: Maboya Gene Expression patterns and Sequence Tags.

MAGEST is a database for maternal gene expression information for an ascidian, Halocynthia roretzi. The ascidian has become an animal model in developmental biological research because it shows a simple developmental process, and belongs to one of the chordate groups. Various data are deposited into the MAGEST database, e.g. the 3'- and 5'-tag sequences from the fertilized egg cDNA library, the results of similarity searches against GenBank and the expression data from whole mount in situ hybridization. Over the last 2 years, the data retrieval systems have been improved in several aspects, and the tag sequence entries have increased to over 20 000 clones. Additionally, we constructed a database, translated MAGEST, for the amino acid fragment sequences predicted from the EST data sets. Using this information comprehensively, we should obtain new information on gene functions. The MAGEST database is accessible via the Internet at http://www.genome.ad.jp/magest/.

Amino Acid Sequence↗

The KEGG databases at GenomeNet.

The Kyoto Encyclopedia of Genes and Genomes (KEGG) is the primary database resource of the Japanese GenomeNet service (http://www.genome.ad.jp/) for understanding higher order functional meanings and utilities of the cell or the organism from its genome information. KEGG consists of the PATHWAY database for the computerized knowledge on molecular interaction networks such as pathways and complexes, the GENES database for the information about genes and proteins generated by genome sequencing projects, and the LIGAND database for the information about chemical compounds and chemical reactions that are relevant to cellular processes. In addition to these three main databases, limited amounts of experimental data for microarray gene expression profiles and yeast two-hybrid systems are stored in the EXPRESSION and BRITE databases, respectively. Furthermore, a new database, named SSDB, is available for exploring the universe of all protein coding genes in the complete genomes and for identifying functional links and ortholog groups. The data objects in the KEGG databases are all represented as graphs and various computational methods are developed to detect graph features that can be related to biological functions. For example, the correlated clusters are graph similarities which can be used to predict a set of genes coding for a pathway or a complex, as summarized in the ortholog group tables, and the cliques in the SSDB graph are used to annotate genes. The KEGG databases are updated daily and made freely available (http://www.genome.ad.jp/kegg/).

Animals↗