PubMed Health⌕ Search

Biomedical subjects

Katsuhisa Horimoto

Publications and source records attributed to Katsuhisa Horimoto.

11 recordsLinked to original sources

Partial correlation coefficient between distance matrices as a new indicator of protein-protein interactions.

MOTIVATION: The computational prediction of protein-protein interactions is currently a major issue in bioinformatics. Recently, a variety of co-evolution-based methods have been investigated toward this goal. In this study, we introduced a partial correlation coefficient as a new measure for the degree of co-evolution between proteins, and proposed its use to predict protein-protein interactions. RESULTS: The accuracy of the prediction by the proposed method was compared with those of the original mirror tree method and the projection method previously developed by our group. We found that the partial correlation coefficient effectively reduces the number of false positives, as compared with other methods, although the number of false negatives increased in the prediction by the partial correlation coefficient. AVAILABILITY: The R script for the prediction of protein-protein interactions reported in this manuscript is available at http://timpani.genome.ad.jp/~parco/

Amino Acid Sequence↗

Protein threading with profiles and distance constraints using clique based algorithms.

With the advent of experimental technologies like chemical cross-linking, it has become possible to obtain distances between specific residues of a newly sequenced protein. These types of experiments usually are less time consuming than X-ray crystallography or NMR. Consequently, it is highly desired to develop a method that incorporates this distance information to improve the performance of protein threading methods. However, protein threading with profiles in which constraints on distances between residues are given is known to be NP-hard. By using the notion of a maximum edge-weight clique finding algorithm, we introduce a more efficient method called FTHREAD for profile threading with distance constraints that is 18 times faster than its predecessor CLIQUETHREAD. Moreover, we also present a novel practical algorithm NTHREAD for profile threading with Non-strict constraints. The overall performance of FTHREAD on a data set shows that although our algorithm uses a simple threading function, our algorithm performs equally well as some of the existing methods. Particularly, when there are some unsatisfied constraints, NTHREAD (Non-strict constraints threading algorithm) performs better than threading with FTHREAD (Strict constraints threading algorithm). We have also analyzed the effects of using a number of distance constraints. This algorithm helps the enhancement of alignment quality between the query sequence and template structure, once the corresponding template structure is determined for the target sequence.

Algorithms↗

Symbolic-numeric estimation of parameters in biochemical models by quantifier elimination.

The sequencing of complete genomes allows analyses of the interactions between various biological molecules on a genomic scale, which prompted us to simulate the global behaviors of biological phenomena on the molecular level. One of the basic mathematical problems in the simulation is the parameter optimization in the kinetic model for complex dynamics, and many estimation methods have been designed. We introduce a new approach to estimate the parameters in biological kinetic models by quantifier elimination (QE), in combination with numerical simulation methods. The estimation method was applied to a model for the inhibition kinetics of HIV proteinase with ten parameters and nine variables, and attained the goodness of fit to 300 points of observed data with the same magnitude as that obtained by the previous estimation methods, remarkably by using only one or two points of data. Furthermore, the utilization of QE demonstrated the feasibility of the present method for elucidating the behavior of the parameters and the variables in the analyzed model. Therefore, the present symbolic-numeric method is a powerful approach to reveal the fundamental mechanisms of kinetic models, in addition to being a computational engine.

Algorithms↗

ASIAN: a web server for inferring a regulatory network framework from gene expression profiles.

The standard workflow in gene expression profile analysis to identify gene function is the clustering by various metrics and techniques, and the following analyses, such as sequence analyses of upstream regions. A further challenging analysis is the inference of a gene regulatory network, and some computational methods have been intensively developed to deduce the gene regulatory network. Here, we describe our web server for inferring a framework of regulatory networks from a large number of gene expression profiles, based on graphical Gaussian modeling (GGM) in combination with hierarchical clustering (http://eureka.ims.u-tokyo.ac.jp/asian). GGM is based on a simple mathematical structure, which is the calculation of the inverse of the correlation coefficient matrix between variables, and therefore, our server can analyze a wide variety of data within a reasonable computational time. The server allows users to input the expression profiles, and it outputs the dendrogram of genes by several hierarchical clustering techniques, the cluster number estimated by a stopping rule for hierarchical clustering and the network between the clusters by GGM, with the respective graphical presentations. Thus, the ASIAN (Automatic System for Inferring A Network) web server provides an initial basis for inferring regulatory relationships, in that the clustering serves as the first step toward identifying the gene function.

Cluster Analysis↗

Predicting absolute contact numbers of native protein structure from amino acid sequence.

The contact number of an amino acid residue in a protein structure is defined by the number of C(beta) atoms around the C(beta) atom of the given residue, a quantity similar to, but different from, solvent accessible surface area. We present a method to predict the contact numbers of a protein from its amino acid sequence. The method is based on a simple linear regression scheme and predicts the absolute values of contact numbers. When single sequences are used for both parameter estimation and cross-validation, the present method predicts the contact numbers with a correlation coefficient of 0.555 on average. When multiple sequence alignments are used, the correlation increases to 0.627, which is a significant improvement over previous methods. In terms of discrete states prediction, the accuracies for 2-, 3-, and 10-state predictions are, respectively, 71.4%, 54.1%, and 18.9% with residue type-dependent unbiased thresholds, and 76.3%, 59.2%, and 21.8% with residue type-independent unbiased thresholds. The difference between accessible surface area and contact number from a prediction viewpoint and the application of contact number prediction to three-dimensional structure prediction are discussed.

Amino Acid Sequence↗

Relationship between segmental duplications and repeat sequences in human chromosome 7.

Various types of repeat sequences are abundant in genomic sequences, and they are associated with the biological phenomena at distinct levels. In particular, comparative analyses of whole-genome-sized sequence data have revealed that repeat sequences cause segmental duplications, which are a type of chromosomal structural arrangement. In this study, we analyzed the relationships between segmental duplications and repeat sequences in human chromosome 7. For this purpose, three methods for detecting repeat sequences were applied to the genomic sequences of human chromosome 7: RepeatMasker for the dispersed repeats, TRF for the tandem repeats, and STEPSTONE for the inter-spread repeats. By plotting the detected repeat sequences against the locations on the chromosome, all three types of repeats were found to be concentrated around the regions of segmental duplications, as a macroscopic feature of their distributions. Furthermore, the latter two repeat sequences were classified in terms of their periods, and the distribution bias of the detected repeat sequences was statistically tested between the segmental duplication regions and the other regions. As a result, the periods of two repeats were biased, with less than a 5% level of significance probability by the chi(2) test, and the repeats with long periods, about 130bp and more than 400bp, were attributed to a bias with a 5% level of significance probability by the normalized residual test. The mechanism of segmental duplications is discussed based on the present results.

Base Composition↗

Elucidation of the relationships between LexA-regulated genes in the SOS response.

Monitoring the expression of many genes under different conditions is a common approach for investigating gene relationships. In particular, the monitoring sheds light on the biological phenomena in which many genes are coordinately expressed. In this study, we analyzed the expression profiles of LexA-regulated genes after UV irradiation, to elucidate the genes related to the SOS response, which involves coordinately regulated gene expression. By the two-gene relationship analysis, the LexA-regulated genes were highly correlated with the genes involved in the DNA repair functions. The LexA-regulated genes with highly significant probability were divided into two groups: the LexA-regulated genes that were mutually related within them were related to the genes with DNA repair functions, while the LexA-regulated genes that were less related within them showed lower relation to the genes with DNA repair functions. By a multiple gene relationship analysis, the two types of LexA-regulated genes were clearly clustered, and the inferred network between the clusters indicated their sequential relationship of clusters in the two groups of LexA-regulated genes in the SOS response; the former type of genes emerged in the early stage of the SOS response upon the signal transduction by membrane proteins, cessation of cell division and recognition of DNA damage, and the latter type emerged in a later stage, and functioned in the repair mechanism and the resumption of DNA replication.

Bacterial Proteins↗

ASIAN: a website for network inference.

UNLABELLED: We constructed a website for inferring a network by applying the graphical Gaussian model, from a large amount of data, including redundant information. The available tools on the website are based on a system, named ASIAN (Automatic System for Inferring A Network), in combination with the two methods in our previous papers, which were designed to analyze gene expression profiles on a genomic scale. One of the remarkable features of the website is its ability to infer a network, concomitant with hierarchical clustering and the following estimation of cluster boundaries. AVAILABILITY: http://eureka.ims.u-tokyo.ac.jp/asian

Computer Simulation↗

Detection of inter-spread repeat sequence in genomic DNA sequence.

Various types of periodic patterns in nucleotide sequences are known to be very abundant in a genomic DNA sequence, and to play important biological roles such as gene expression, genome structural stabilization, and recombination. We present a new method, named "STEPSTONE", to find a specific periodic pattern of repeat sequence, inter-spread repeat, in which the tandem repeats of the conserved and the not-conserved regions appear periodically. In our method, at first, the data on periods of short repeat sequences found in a target sequence are stored as a hash data, and then are selected by application of an auto-correlation test in time series analysis. Among the statistically selected sequences, the inter-spread repeats are obtained by usual alignment procedures through two steps. To test the performance of our method, we examined the inter-spread repeats in Mycobacterium tuberculosis and Zamia paucijuga genomic sequences. As a result, our method exactly detected the repeats in the two sequences, being useful for identifying systematically the inter-spread repeats in DNA sequence.

Algorithms↗

Causes for the large genome size in a cyanobacterium Anabaena sp. PCC7120.

Three possible causes responsible for the large genome size of a cyanobacterium Anabaena sp. PCC7120 are investigated: 1) sequential tandem duplications of gene segments, genes or genomic segments, 2) horizontal gene transfers from other organisms, and 3) whole-genome duplication. We evaluated the frequency distribution of angles between paralog locations for the possibility 1), the fraction of genes deviated in GC content, GC skew, AT skew and codon adaptation index for the 2) and the gene-configuration comparison of paralogs for the 3). As a result, the possibility 3), the whole-genome duplication, was more reasonable as a molecular cause than the other causes for the large genome size in Anabaena sp. PCC7120. In addition, the whole-genome duplication was supported by the analysis of distribution pattern of protein genes with respect to functional categories.

Anabaena↗

Inference of a genetic network by a combined approach of cluster analysis and graphical Gaussian modeling.

MOTIVATION: Recent advances in DNA microarray technologies have made it possible to measure the expression levels of thousands of genes simultaneously under different conditions. The data obtained by microarray analyses are called expression profile data. One type of important information underlying the expression profile data is the 'genetic network,' that is, the regulatory network among genes. Graphical Gaussian Modeling (GGM) is a widely utilized method to infer or test relationships among a plural of variables. RESULTS: In this study, we developed a method combining the cluster analysis with GGM for the inference of the genetic network from the expression profile data. The expression profile data of 2467 Saccharomyces cerevisiae genes measured under 79 different conditions (Eisen et al., PROC: Natl Acad. Sci. USA, 95, 14683-14868, 1998) were used for this study. At first, the 2467 genes were classified into 34 clusters by a cluster analysis, as a preprocessing for GGM. Then, the expression levels of the genes in each cluster were averaged for each condition. The averaged expression profile data of 34 clusters were subjected to GGM, and a partial correlation coefficient matrix was obtained as a model of the genetic network of S. cerevisiae. The accuracy of the inferred network was examined by the agreement of our results with the cumulative results of experimental studies.

Bayes Theorem↗