PubMed Health⌕ Search

Biomedical subjects

Sachiyo Aburatani

Publications and source records attributed to Sachiyo Aburatani.

12 recordsLinked to original sources

ASIAN: a web server for inferring a regulatory network framework from gene expression profiles.

The standard workflow in gene expression profile analysis to identify gene function is the clustering by various metrics and techniques, and the following analyses, such as sequence analyses of upstream regions. A further challenging analysis is the inference of a gene regulatory network, and some computational methods have been intensively developed to deduce the gene regulatory network. Here, we describe our web server for inferring a framework of regulatory networks from a large number of gene expression profiles, based on graphical Gaussian modeling (GGM) in combination with hierarchical clustering (http://eureka.ims.u-tokyo.ac.jp/asian). GGM is based on a simple mathematical structure, which is the calculation of the inverse of the correlation coefficient matrix between variables, and therefore, our server can analyze a wide variety of data within a reasonable computational time. The server allows users to input the expression profiles, and it outputs the dendrogram of genes by several hierarchical clustering techniques, the cluster number estimated by a stopping rule for hierarchical clustering and the network between the clusters by GGM, with the respective graphical presentations. Thus, the ASIAN (Automatic System for Inferring A Network) web server provides an initial basis for inferring regulatory relationships, in that the clustering serves as the first step toward identifying the gene function.

Cluster Analysis↗

Relationship between segmental duplications and repeat sequences in human chromosome 7.

Various types of repeat sequences are abundant in genomic sequences, and they are associated with the biological phenomena at distinct levels. In particular, comparative analyses of whole-genome-sized sequence data have revealed that repeat sequences cause segmental duplications, which are a type of chromosomal structural arrangement. In this study, we analyzed the relationships between segmental duplications and repeat sequences in human chromosome 7. For this purpose, three methods for detecting repeat sequences were applied to the genomic sequences of human chromosome 7: RepeatMasker for the dispersed repeats, TRF for the tandem repeats, and STEPSTONE for the inter-spread repeats. By plotting the detected repeat sequences against the locations on the chromosome, all three types of repeats were found to be concentrated around the regions of segmental duplications, as a macroscopic feature of their distributions. Furthermore, the latter two repeat sequences were classified in terms of their periods, and the distribution bias of the detected repeat sequences was statistically tested between the segmental duplication regions and the other regions. As a result, the periods of two repeats were biased, with less than a 5% level of significance probability by the chi(2) test, and the repeats with long periods, about 130bp and more than 400bp, were attributed to a bias with a 5% level of significance probability by the normalized residual test. The mechanism of segmental duplications is discussed based on the present results.

Base Composition↗

Elucidation of the relationships between LexA-regulated genes in the SOS response.

Monitoring the expression of many genes under different conditions is a common approach for investigating gene relationships. In particular, the monitoring sheds light on the biological phenomena in which many genes are coordinately expressed. In this study, we analyzed the expression profiles of LexA-regulated genes after UV irradiation, to elucidate the genes related to the SOS response, which involves coordinately regulated gene expression. By the two-gene relationship analysis, the LexA-regulated genes were highly correlated with the genes involved in the DNA repair functions. The LexA-regulated genes with highly significant probability were divided into two groups: the LexA-regulated genes that were mutually related within them were related to the genes with DNA repair functions, while the LexA-regulated genes that were less related within them showed lower relation to the genes with DNA repair functions. By a multiple gene relationship analysis, the two types of LexA-regulated genes were clearly clustered, and the inferred network between the clusters indicated their sequential relationship of clusters in the two groups of LexA-regulated genes in the SOS response; the former type of genes emerged in the early stage of the SOS response upon the signal transduction by membrane proteins, cessation of cell division and recognition of DNA damage, and the latter type emerged in a later stage, and functioned in the repair mechanism and the resumption of DNA replication.

Bacterial Proteins↗

ASIAN: a website for network inference.

UNLABELLED: We constructed a website for inferring a network by applying the graphical Gaussian model, from a large amount of data, including redundant information. The available tools on the website are based on a system, named ASIAN (Automatic System for Inferring A Network), in combination with the two methods in our previous papers, which were designed to analyze gene expression profiles on a genomic scale. One of the remarkable features of the website is its ability to infer a network, concomitant with hierarchical clustering and the following estimation of cluster boundaries. AVAILABILITY: http://eureka.ims.u-tokyo.ac.jp/asian

Computer Simulation↗

An integrated comprehensive workbench for inferring genetic networks: voyagene.

We propose an integrated, comprehensive network-inferring system for genetic interactions, named VoyaGene, which can analyze experimentally observed expression profiles by using and combining the following five independent inferring models: Clustering, Threshold-Test, Bayesian, multi-level digraph and S-system models. Since VoyaGene also has effective tools for visualizing the inferred results, researchers may evaluate the combination of appropriate inferring models, and can construct a genetic network to an accuracy that is beyond the reach of a single inferring model. Through the use of VoyaGene, the present study demonstrates the effectiveness of combining different inferring models.

Algorithms↗

Detection of inter-spread repeat sequence in genomic DNA sequence.

Various types of periodic patterns in nucleotide sequences are known to be very abundant in a genomic DNA sequence, and to play important biological roles such as gene expression, genome structural stabilization, and recombination. We present a new method, named "STEPSTONE", to find a specific periodic pattern of repeat sequence, inter-spread repeat, in which the tandem repeats of the conserved and the not-conserved regions appear periodically. In our method, at first, the data on periods of short repeat sequences found in a target sequence are stored as a hash data, and then are selected by application of an auto-correlation test in time series analysis. Among the statistically selected sequences, the inter-spread repeats are obtained by usual alignment procedures through two steps. To test the performance of our method, we examined the inter-spread repeats in Mycobacterium tuberculosis and Zamia paucijuga genomic sequences. As a result, our method exactly detected the repeats in the two sequences, being useful for identifying systematically the inter-spread repeats in DNA sequence.

Algorithms↗

Causes for the large genome size in a cyanobacterium Anabaena sp. PCC7120.

Three possible causes responsible for the large genome size of a cyanobacterium Anabaena sp. PCC7120 are investigated: 1) sequential tandem duplications of gene segments, genes or genomic segments, 2) horizontal gene transfers from other organisms, and 3) whole-genome duplication. We evaluated the frequency distribution of angles between paralog locations for the possibility 1), the fraction of genes deviated in GC content, GC skew, AT skew and codon adaptation index for the 2) and the gene-configuration comparison of paralogs for the 3). As a result, the possibility 3), the whole-genome duplication, was more reasonable as a molecular cause than the other causes for the large genome size in Anabaena sp. PCC7120. In addition, the whole-genome duplication was supported by the analysis of distribution pattern of protein genes with respect to functional categories.

Anabaena↗

Discovery of novel transcription control relationships with gene regulatory networks generated from multiple-disruption full genome expression libraries.

Gene regulatory networks elucidated from strategic, genome-wide experimental data can aid in the discovery of novel gene function information and expression regulation events from observation of transcriptional regulation among genes of known and unknown biological function. To create a reliable and comprehensive data set for the elucidation of transcription regulation networks, we conducted systematic genome-wide disruption expression experiments of yeast on 118 genes with known involvement in transcription regulation. We report several novel regulatory relationships between known transcription factors and other genes with previously unknown biological function discovered with this expression library. Here we report the downstream regulatory subnetworks for UME6 and MET28. The elucidated network topology among these genes demonstrates MET28's role as a nodal point between genes involved in cell division and those involved in DNA repair mechanisms.

Algorithms↗

Use of gene networks from full genome microarray libraries to identify functionally relevant drug-affected genes and gene regulation cascades.

We developed an extensive yeast gene expression library consisting of full-genome cDNA array data for over 500 yeast strains, each with a single-gene disruption. Using this data, combined with dose and time course expression experiments with the oral antifungal agent griseofulvin, whose exact molecular targets were previously unknown, we used Boolean and Bayesian network discovery techniques to determine the gene expression regulatory cascades affected directly by this drug. Using this method we identified CIK1 as an important affected target gene related to the functional phenotype induced by griseofulvin. Cellular functional analysis of griseofulvin showed similar tubulin-specific morphological effects on mitotic spindle formation to those of the drug, in agreement with the known function of CIK1p. Further, using the nonparametric, nonlinear Bayesian gene networks we were able to identify alternative ligand-dependant transcription factors and G protein homologues upstream of CIK1 that regulate CIK1 expression and might therefore serve as alternative molecular targets to induce the same molecular response as griseofulvin.

Bayes Theorem↗

Bayesian network and nonparametric heteroscedastic regression for nonlinear modeling of genetic network.

We propose a new statistical method for constructing a genetic network from microarray gene expression data by using a Bayesian network. An essential point of Bayesian network construction is the estimation of the conditional distribution of each random variable. We consider fitting nonparametric regression models with heterogeneous error variances to the microarray gene expression data to capture the nonlinear structures between genes. Selecting the optimal graph, which gives the best representation of the system among genes, is still a problem to be solved. We theoretically derive a new graph selection criterion from Bayes approach in general situations. The proposed method includes previous methods based on Bayesian networks. We demonstrate the effectiveness of the proposed method through the analysis of Saccharomyces cerevisiae gene expression data newly obtained by disrupting 100 genes.

Bayes Theorem↗

Use of gene networks for identifying and validating drug targets.

We propose a new method for identifying and validating drug targets by using gene networks, which are estimated from cDNA microarray gene expression profile data. We created novel gene disruption and drug response microarray gene expression profile data libraries for the purpose of drug target elucidation. We use two types of microarray gene expression profile data for estimating gene networks and then identifying drug targets. The estimated gene networks play an essential role in understanding drug response data and this information is unattainable from clustering methods, which are the standard for gene expression analysis. In the construction of gene networks, we use the Bayesian network model. We use an actual example from analysis of the Saccharomyces cerevisiae gene expression profile data to express a concrete strategy for the application of gene network information to drug discovery.

Algorithms↗

Bayesian network and nonparametric heteroscedastic regression for nonlinear modeling of genetic network.

We propose a new statistical method for constructing genetic network from microarray gene expression data by using a Bayesian network. An essential point of Bayesian network construction is in the estimation of the conditional distribution of each random variable. We consider fitting nonparametric regression models with heterogeneous error variances to the microarray gene expression data to capture the nonlinear structures between genes. A problem still remains to be solved in selecting an optimal graph, which gives the best representation of the system among genes. We theoretically derive a new graph selection criterion from Bayes approach in general situations. The proposed method includes previous methods based on Bayesian networks. We demonstrate the effectiveness of the proposed method through the analysis of Saccharomyces cerevisiae gene expression data newly obtained by disrupting 100 genes.

Artificial Intelligence↗