PubMed Health⌕ Search

Biomedical subjects

Shigehiko Kanaya

Publications and source records attributed to Shigehiko Kanaya.

At least 19 recordsLinked to original sources

Comparative analysis of argK-tox clusters and their flanking regions in phaseolotoxin-producing Pseudomonas syringae pathovars.

DNA fragments containing argK-tox clusters and their flanking regions were cloned from the chromosomes of Pseudomonas syringae pathovar (pv.) actinidiae strain KW-11 (ACT) and P. syringae pv. phaseolicola strain MAFF 302282 (PHA), and then their sequences were determined. Comparative analysis of these sequences and the sequences of P. syringae pv. tomato DC3000 (TOM) (Buell et al., Proc Natl Acad Sci USA 100:10181-10186, 2003) and pv. syringae B728a (SYR) (Feil et al., Proc Natl Acad Sci USA 102:11064-11069, 2005) revealed that the chromosomal backbone regions of ACT and TOM shared a high similarity to each other but presented a low similarity to those of PHA and SYR. Nevertheless, almost-identical DNA regions of about 38 kb were confirmed to be present on the chromosomes of both ACT and PHA, which we named "tox islands." The facts that the GC content of such tox islands was 6% lower than that of the chromosomal backbone regions of P. syringae, and that argK-tox clusters, which are considered to be of exogenous origin based on our previous studies (Sawada et al., J Mol Evol 54:437-457, 2002), were confirmed to be contained within the tox islands, suggested that the tox islands were an exogenous, mobile genetic element inserted into the chromosomes of P. syringae strains. It was also predicted that the tox islands integrated site-specifically into the homologous sites of the chromosomes of ACT and PHA in the same direction, respectively, wherein 34 common gene coding sequences (CDSs) existed. Furthermore, at the left end of the tox islands were three CDSs, which encoded polypeptides and had similarities to the members of the tyrosine recombinase family, suggesting that these putative site-specific recombinases were involved in the recent horizontal transfer of tox islands.

3' Flanking Region↗

Clarification of pathway-specific inhibition by Fourier transform ion cyclotron resonance/mass spectrometry-based metabolic phenotyping studies.

We have developed a metabolic profiling scheme based on direct-infusion Fourier transform ion cyclotron resonance mass spectrometry (FT-ICR/MS). The scheme consists of: (1) reproducible data collection under optimized FT-ICR/MS analytical conditions; (2) automatic mass-error correction and multivariate analyses for metabolome characterization using a newly developed metabolomics tool (DMASS software); (3) identification of marker metabolite candidates by searching a species-metabolite relationship database, KNApSAcK; and (4) structural analyses by an MS/MS method. The scheme was applied to metabolic phenotyping of Arabidopsis (Arabidopsis thaliana) seedlings treated with different herbicidal chemical classes for pathway-specific inhibitions. Arabidopsis extracts were directly infused into an electrospray ionization source on an FT-ICR/MS system. Acquired metabolomics data were comprised of mass-to-charge ratio values with ion intensity information subjected to principal component analysis, and metabolic phenotypes from the herbicide treatments were clearly differentiated from those of the herbicide-free treatment. From each herbicide treatment, candidate metabolites representing such metabolic phenotypes were found through the KNApSAcK database search. The database search and MS/MS analyses suggested dose-dependent accumulation patterns of specific metabolites including several flavonoid glycosides. The metabolic phenotyping scheme on the basis of FT-ICR/MS coupled with the DMASS program is discussed as a general tool for high throughput metabolic phenotyping studies.

Arabidopsis↗

Development and implementation of an algorithm for detection of protein complexes in large interaction networks.

BACKGROUND: After complete sequencing of a number of genomes the focus has now turned to proteomics. Advanced proteomics technologies such as two-hybrid assay, mass spectrometry etc. are producing huge data sets of protein-protein interactions which can be portrayed as networks, and one of the burning issues is to find protein complexes in such networks. The enormous size of protein-protein interaction (PPI) networks warrants development of efficient computational methods for extraction of significant complexes. RESULTS: This paper presents an algorithm for detection of protein complexes in large interaction networks. In a PPI network, a node represents a protein and an edge represents an interaction. The input to the algorithm is the associated matrix of an interaction network and the outputs are protein complexes. The complexes are determined by way of finding clusters, i. e. the densely connected regions in the network. We also show and analyze some protein complexes generated by the proposed algorithm from typical PPI networks of Escherichia coli and Saccharomyces cerevisiae. A comparison between a PPI and a random network is also performed in the context of the proposed algorithm. CONCLUSION: The proposed algorithm makes it possible to detect clusters of proteins in PPI networks which mostly represent molecular biological functional units. Therefore, protein complexes determined solely based on interaction data can help us to predict the functions of proteins, and they are also useful to understand and explain certain biological processes.

Algorithms↗

Large-scale identification of protein-protein interaction of Escherichia coli K-12.

Protein-protein interactions play key roles in protein function and the structural organization of a cell. A thorough description of these interactions should facilitate elucidation of cellular activities, targeted-drug design, and whole cell engineering. A large-scale comprehensive pull-down assay was performed using a His-tagged Escherichia coli ORF clone library. Of 4339 bait proteins tested, partners were found for 2667, including 779 of unknown function. Proteins copurifying with hexahistidine-tagged baits on a Ni2+-NTA column were identified by MALDI-TOF MS (matrix-assisted laser desorption ionization time of flight mass spectrometry). An extended analysis of these interacting networks by bioinformatics and experimentation should provide new insights and novel strategies for E. coli systems biology.

Escherichia coli K12↗

Global landscape of protein complexes in the yeast Saccharomyces cerevisiae.

Identification of protein-protein interactions often provides insight into protein function, and many cellular processes are performed by stable protein complexes. We used tandem affinity purification to process 4,562 different tagged proteins of the yeast Saccharomyces cerevisiae. Each preparation was analysed by both matrix-assisted laser desorption/ionization-time of flight mass spectrometry and liquid chromatography tandem mass spectrometry to increase coverage and accuracy. Machine learning was used to integrate the mass spectrometry scores and assign probabilities to the protein-protein interactions. Among 4,087 different proteins identified with high confidence by mass spectrometry from 2,357 successful purifications, our core data set (median precision of 0.69) comprises 7,123 protein-protein interactions involving 2,708 proteins. A Markov clustering algorithm organized these interactions into 547 protein complexes averaging 4.9 subunits per complex, about half of them absent from the MIPS database, as well as 429 additional interactions between pairs of complexes. The data (all of which are available online) will help future studies on individual proteins as well as functional genomics and systems biology.

Biological Evolution↗

A flexible representation of omic knowledge for thorough analysis of microarray data.

BACKGROUND: In order to understand microarray data reasonably in the context of other existing biological knowledge, it is necessary to conduct a thorough examination of the data utilizing every aspect of available omic knowledge libraries. So far, a number of bioinformatics tools have been developed. However, each of them is restricted to deal with one type of omic knowledge, e.g., pathways, interactions or gene ontology. Now that the varieties of omic knowledge are expanding, analysis tools need a way to deal with any type of omic knowledge. Hence, we have designed the Omic Space Markup Language (OSML) that can represent a wide range of omic knowledge, and also, we have developed a tool named GSCope3, which can statistically analyze microarray data in comparison with the OSML-formatted omic knowledge data. RESULTS: In order to test the applicability of OSML to represent a variety of omic knowledge specifically useful for analysis of Arabidopsis thaliana microarray data, we have constructed a Biological Knowledge Library (BiKLi) by converting eight different types of omic knowledge into OSML-formatted datasets. We applied GSCope3 and BiKLi to previously reported A. thaliana microarray data, so as to extract any additional insights from the data. As a result, we have discovered a new insight that lignin formation resists drought stress and activates transcription of many water channel genes to oppose drought stress; and most of the 20S proteasome subunit genes show similar expression profiles under drought stress. In addition to this novel discovery, similar findings previously reported were also quickly confirmed using GSCope3 and BiKLi. CONCLUSION: GSCope3 can statistically analyze microarray data in the context of any OSML-represented omic knowledge. OSML is not restricted to a specific data type structure, but it can represent a wide range of omic knowledge. It allows us to convert new types of omic knowledge into datasets that can be used for microarray data analysis with GSCope3. In addition to BiKLi, by collecting various types of omic knowledge as OSML libraries, it becomes possible for us to conduct detailed thorough analysis from various biological viewpoints. GSCope3 and BiKLi are available for academic users at our web site http://omicspace.riken.jp.

Journal Article↗

Novel phylogenetic studies of genomic sequence fragments derived from uncultured microbe mixtures in environmental and clinical samples.

A self-organizing map (SOM) was developed as a novel bioinformatics strategy for phylogenetic classification of sequence fragments obtained from pooled genome samples of uncultured microbes in environmental and clinical samples. This phylogenetic classification was possible without either orthologous sequence sets or sequence alignments. We first constructed SOMs for tetranucleotide frequencies in 210,000 5 kb sequence fragments obtained from 1502 prokaryotes for which at least 10 kb of genomic sequence has been deposited in public DNA databases. The sequences could be classified primarily according to phylogenetic groups without information regarding the species. We used the SOM method to classify sequence fragments derived from environmental samples of the Sargasso Sea and of an acidophilic biofilm growing in acid mine drainage. Phylogenetic diversity of the environmental sequences was effectively visualized on a single map. Sequences that were derived from a single genome but cloned independently could be reassociated in silico. G + C% has been used for a long period as a fundamental parameter for phylogenetic classification of microbes, but the G + C% is apparently too simple a parameter to differentiate a wide variety of known species. Oligonucleotide frequency can be used to distinguish the species because oligonucleotide frequencies vary significantly among their genomes.

Algorithms↗

Self-Organizing Map (SOM) unveils and visualizes hidden sequence characteristics of a wide range of eukaryote genomes.

Novel tools are needed for comprehensive comparisons of interspecies characteristics of massive amounts of genomic sequences currently available. An unsupervised neural network algorithm, Self-Organizing Map (SOM), is an effective tool for clustering and visualizing high-dimensional complex data on a single map. We modified the conventional SOM, on the basis of batch-learning SOM, for genome informatics making the learning process and resulting map independent of the order of data input. We generated the SOMs for tri- and tetranucleotide frequencies in 10- and 100-kb sequence fragments from 38 eukaryotes for which almost complete genome sequences are available. SOM recognized species-specific characteristics (key combinations of oligonucleotide frequencies) in the genomic sequences, permitting species-specific classification of the sequences without any information regarding the species. We also generated the SOM for tetranucleotide frequencies in 1-kb sequence fragments from the human genome and found sequences for four functional categories (5' and 3' UTRs, CDSs and introns) were classified primarily according to the categories. Because the classification and visualization power is very high, SOM is an efficient and powerful tool for extracting a wide range of genome information.

3' Untranslated Regions↗

Elucidation of gene-to-gene and metabolite-to-gene networks in arabidopsis by integration of metabolomics and transcriptomics.

Since the completion of genome sequences of model organisms, functional identification of unknown genes has become a principal challenge in biology. Post-genomics sciences such as transcriptomics, proteomics, and metabolomics are expected to discover gene functions. This report outlines the elucidation of gene-to-gene and metabolite-to-gene networks via integration of metabolomics with transcriptomics and presents a strategy for the identification of novel gene functions. Metabolomics and transcriptomics data of Arabidopsis grown under sulfur deficiency were combined and analyzed by batch-learning self-organizing mapping. A group of metabolites/genes regulated by the same mechanism clustered together. The metabolism of glucosinolates was shown to be coordinately regulated. Three uncharacterized putative sulfotransferase genes clustering together with known glucosinolate biosynthesis genes were candidates for involvement in biosynthesis. In vitro enzymatic assays of the recombinant gene products confirmed their functions as desulfoglucosinolate sulfotransferases. Several genes involved in sulfur assimilation clustered with O-acetylserine, which is considered a positive regulator of these genes. The genes involved in anthocyanin biosynthesis clustered with the gene encoding a transcriptional factor that up-regulates specifically anthocyanin biosynthesis genes. These results suggested that regulatory metabolites and transcriptional factor genes can be identified by this approach, based on the assumption that they cluster with the downstream genes they regulate. This strategy is applicable not only to plant but also to other organisms for functional elucidation of unknown genes.

Arabidopsis↗

Accurate extraction of functional associations between proteins based on common interaction partners and common domains.

MOTIVATION: Genomic and proteomic approaches have accumulated a huge amount of data which provide clues to protein function. However, interpreting single omic data for predicting uncharacterized protein functions has been a challenging task, because the data contain a lot of false positives. To overcome this problem, methods for integrating data from various omic approaches are needed for more accurate function prediction. RESULT: In this paper, we have developed a method which extracts functionally similar proteins with high confidence by integrating protein-protein interaction data and domain information. We used this method to analyze publicly available data from Saccharomyces cerevisiae. We identified 1042 functional associations, involving 765 proteins of which 98 (12.8%) had no previously ascribed function. Our method extracts functionally similar protein pairs more accurately than conventional methods, and predicting function for previously uncharacterized proteins can be achieved. Our method can of course be applied to protein-protein interaction data for any species.

Algorithms↗

GeneLook: a novel ab initio gene identification system suitable for automated annotation of prokaryotic sequences.

With the rapid increases in the amounts of sequence data for prokaryotic genomes, it has become important to develop systems for automated and accurate genome annotation. We present herein a novel ab initio gene identification system, GeneLook, that predicts protein-coding open reading frames (ORFs) with high sensitivity and specificity with no prior knowledge of the sequence composition. The system predicts protein-coding ORFs in two stages, seed ORF selection and main prediction. In the selection of reliable seed ORFs containing at least 200 codons, GeneLook predicts translation start sites and operon structures through searches for ribosome-binding sites and a novel operon prediction algorithm. The codon and nucleotide frequencies of seed ORFs are then used to determine values for two new coding-potential parameters for identification of protein-coding ORFs of at least 34 codons and for another parameter that improves the prediction accuracy for GC-rich genomes. In the main prediction, GeneLook uses these parameters to identify the most likely genes of a given minimal length. We assessed the performance of GeneLook with two indices, sensitivity and specificity that are defined as true positives (TP)/(TP+false negatives) and TP/(TP+false positives), respectively. This system predicted protein-coding ORFs for Escherichia coli and Bacillus subtilis with sensitivities of 96.5% and 96.2%, respectively, and specificities of 96.9% and 96.1%, respectively. The system also identified 94.1% of annotated genes of the Pseudomonas aeruginosa genome, which is GC-rich, with high specificity (97.2%). Furthermore, GeneLook identified protein-coding ORFs with high accuracy from a wide variety of prokaryotic genomes.

Bacillus subtilis↗

Sequential binding of SeqA protein to nascent DNA segments at replication forks in synchronized cultures of Escherichia coli.

To demonstrate that sequestration A (SeqA) protein binds preferentially to hemimethylated GATC sequences at replication forks and forms clusters in Escherichia coli growing cells, we analysed, by the chromatin immunoprecipitation (ChIP) assay using anti-SeqA antibody, a synchronized culture of a temperature-sensitive dnaC mutant strain in which only one round of chromosomal DNA replication was synchronously initiated. After synchronized initiation of chromosome replication, the replication origin oriC was first detected by the ChIP assay, and other six chromosomal regions having multiple GATC sequences were sequentially detected according to bidirectional replication of the chromosome. In contrast, DNA regions lacking the GATC sequence were not detected by the ChIP assay. These results indicate that SeqA binds hemimethylated nascent DNA segments according to the proceeding of replication forks in the chromosome, and SeqA releases from the DNA segments when fully methylated. Immunofluorescence microscopy reveals that a single SeqA focus containing paired replication apparatuses appears at the middle of the cell immediately after initiation of chromosome replication and the focus is subsequently separated into two foci that migrate to 1/4 and 3/4 cellular positions, when replication forks proceed bidirectionally an approximately one-fourth distance from the replication origin towards the terminus. This supports the translocating replication apparatuses model.

Bacterial Outer Membrane Proteins↗

Integration of transcriptomics and metabolomics for understanding of global responses to nutritional stresses in Arabidopsis thaliana.

Plant metabolism is a complex set of processes that produce a wide diversity of foods, woods, and medicines. With the genome sequences of Arabidopsis and rice in hands, postgenomics studies integrating all "omics" sciences can depict precise pictures of a whole-cellular process. Here, we present, to our knowledge, the first report of investigation for gene-to-metabolite networks regulating sulfur and nitrogen nutrition and secondary metabolism in Arabidopsis, with integration of metabolomics and transcriptomics. Transcriptome and metabolome analyses were carried out, respectively, with DNA macroarray and several chemical analytical methods, including ultra high-resolution Fourier transform-ion cyclotron MS. Mathematical analyses, including principal component analysis and batch-learning self-organizing map analysis of transcriptome and metabolome data suggested the presence of general responses to sulfur and nitrogen deficiencies. In addition, specific responses to either sulfur or nitrogen deficiency were observed in several metabolic pathways: in particular, the genes and metabolites involved in glucosinolate metabolism were shown to be coordinately modulated. Understanding such gene-to-metabolite networks in primary and secondary metabolism through integration of transcriptomics and metabolomics can lead to identification of gene function and subsequent improvement of production of useful compounds in plants.

Arabidopsis↗

Informatics for unveiling hidden genome signatures.

With the increasing amount of available genome sequences, novel tools are needed for comprehensive analysis of species-specific sequence characteristics for a wide variety of genomes. We used an unsupervised neural network algorithm, a self-organizing map (SOM), to analyze di-, tri-, and tetranucleotide frequencies in a wide variety of prokaryotic and eukaryotic genomes. The SOM, which can cluster complex data efficiently, was shown to be an excellent tool for analyzing global characteristics of genome sequences and for revealing key combinations of oligonucleotides representing individual genomes. From analysis of 1- and 10-kb genomic sequences derived from 65 bacteria (a total of 170 Mb) and from 6 eukaryotes (460 Mb), clear species-specific separations of major portions of the sequences were obtained with the di-, tri-, and tetranucleotide SOMs. The unsupervised algorithm could recognize, in most 10-kb sequences, the species-specific characteristics (key combinations of oligonucleotide frequencies) that are signature features of each genome. We were able to classify DNA sequences within one and between many species into subgroups that corresponded generally to biological categories. Because the classification power is very high, the SOM is an efficient and fundamental bioinformatic strategy for extracting a wide range of genomic information from a vast amount of sequences.

Animals↗

Periodicity in prokaryotic and eukaryotic genomes identified by power spectrum analysis.

We used a power spectrum method to identify periodic patterns in nucleotide sequence, and characterized nucleotide sequences that confer periodicities to prokaryotic and eukaryotic genomes and genomes. A 10-bp periodicity was prevalent in hyperthermophilic bacteria and archaebacteria, and an 11-bp periodicity was prevalent in eubacteria. The 10-bp periodicity was also prevalent in the eukaryotes such as the worm Caenorhabditis elegans. Additionally, in the worm genome, a 68-bp periodicity in chromosome I, a 59-bp periodicity in chromosome II, and a 94-bp periodicity in chromosome III were found. In human chromosomes 21 and 22, approximately 167- or 84-bp periodicity was detected along the entire length of these chromosomes. Because the 167-bp is identical to the length of DNA that forms two complete helical turns in nucleosome organization, we speculated that the respective sequences may correspond to arrays of a special compact form of nucleosomes clustered in specific regions of the human chromosomes. This periodic element contained a high frequency of TGG. TGG-rich sequences are known to form a specific subset of folded DNA structures, and therefore, the sequences might have potential to form specific higher order structures related to the clustered occurrence of a specific form of the speculated nucleosomes.

Animals↗

Distribution of repetitive sequences on the leading and lagging strands of the Escherichia coli genome: comparative study of Long Direct Repeat (LDR) sequences.

In the present study, we developed a method for detecting sequences whose similarity to a target sequence is statistically significant and we examined the distribution of these sequences in the E. coli K-12 genome. Target sequences examined are as follows: (i) short repeat: Crossover hot-spot instigator (Chi) sequence, replication termination (Ter) sequence, and DnaA binding sequence (DnaA box); (ii) potential stem-loop structure repeats: palindromic unit (PU), boxC sequences, and intergenic repeat unit (IRU); (iii) potential RNA coding repeats: rRNAs, PAIR, TRIP, and QUAD; and (iv) potential protein coding repeats: insertion elements (ISs) and Long Direct Repeats (LDRs). We also examined the distribution of these sequences on leading and lagging strands. We obtained another four statistically significant LDR sequences with more than 187 bp matched to LDR-A near the LDR loci, suggesting that these regions might be used as high recombination hot spots for LDR. Adaptation of individual LDRs to E. coli genome is also discussed on the basis of codon usage.

Bacterial Proteins↗

Conservation of translation initiation sites based on dinucleotide frequency and codon usage in Escherichia coli K-12 (W3110): non-random distribution of A/T-rich sequences immediately upstream of the translation initiation codon.

Dinucleotide frequencies are useful for characterizing consensus elements as a minimum unit of nucleotide sequence because the neighborhood relations of nucleotide sequences are reflected in dinucleotides. Using a consensus score based on dinucleotide frequencies and intra-species codon usage heterogeneity, denoted by the Z1 parameter, we report the relationship between nucleotide conservation at the translation initiation sites of genes in the Escherichia coli K-12 genome (W3110) and codon usage in its downstream genes. Significant positive correlations were obtained in three regions centered at -13, -4, and +7, which correspond to the Shine-Dalgarno element, the A + T element immediately upstream of the translation initiation site, and the downstream box, respectively.

Base Sequence↗

A phylogenomic study of the OCTase genes in Pseudomonas syringae pathovars: the horizontal transfer of the argK-tox cluster and the evolutionary history of OCTase genes on their genomes.

Phytopathogenic Pseudomonas syringae is subdivided into about 50 pathovars due to their conspicuous differentiation with regard to pathogenicity. Based on the results of a phylogenetic analysis of four genes (gyrB, rpoD, hrpL, and hrpS), Sawada et al. (1999) showed that the ancestor of P. syringae had diverged into at least three monophyletic groups during its evolution. Physical maps of the genomes of representative strains of these three groups were constructed, which revealed that each strain had five rrn operons which existed on one circular genome. The fact that the structure and size of genomes vary greatly depending on the pathovar shows that P. syringae genomes are quite rich in plasticity and that they have undergone large-scale genomic rearrangements. Analyses of the codon usage and the GC content at the codon third position, in conjunction with phylogenomic analyses, showed that the gene cluster involved in phaseolotoxin synthesis (argK-tox cluster) expanded its distribution by conducting horizontal transfer onto the genomes of two P. syringae pathovars (pv. actinidiae and pv. phaseolicola) from bacterial species distantly related to P. syringae and that its acquisition was quite recent (i.e., after the ancestor of P. syringae diverged into the respective pathovars). Furthermore, the results of a detailed analysis of argK [an anabolic ornithine carbamoyltransferase (anabolic OCTase) gene], which is present within the argK-tox cluster, revealed the plausible process of generation of an unusual composition of the OCTase genes on the genomes of these two phaseolotoxin-producing pathovars: a catabolic OCTase gene (equivalent to the orthologue of arcB of P. aeruginosa) and an anabolic OCTase gene (argF), which must have been formed by gene duplication, have first been present on the genome of the ancestor of P. syringae; the catabolic OCTase gene has been deleted; the ancestor has diverged into the respective pathovars; the foreign-originated argK-tox cluster has horizontally transferred onto the genomes of pv. actinidiae and pv. phaseolicola; and hence two copies of only the anabolic OCTase genes (argK and argF) came to exist on the genomes of these two pathovars. Thus, the horizontal gene transfer and the genomic rearrangement were proven to have played an important role in the pathogenic differentiation and diversification of P. syringae.

Evolution, Molecular↗