PubMed Health⌕ Search

Biomedical subjects

Kenji Satou

Publications and source records attributed to Kenji Satou.

9 recordsLinked to original sources

Computational discovery of transcriptional regulatory rules.

MOTIVATION: Even in a simple organism like yeast Saccharomyces cerevisiae, transcription is an extremely complex process. The expression of sets of genes can be turned on or off by the binding of specific transcription factors to the promoter regions of genes. Experimental and computational approaches have been proposed to establish mappings of DNA-binding locations of transcription factors. However, although location data obtained from experimental methods are noisy owing to imperfections in the measuring methods, computational approaches suffer from over-prediction problems owing to the short length of the sequence motifs bound by the transcription factors. Also, these interactions are usually environment-dependent: many regulators only bind to the promoter region of genes under specific environmental conditions. Even more, the presence of regulators at a promoter region indicates binding but not necessarily function: the regulator may act positively, negatively or not act at all. Therefore, identifying true and functional interactions between transcription factors and genes in specific environment conditions and describing the relationship between them are still open problems. RESULTS: We developed a method that combines expression data with genomic location information to discover (1) relevant transcription factors from the set of potential transcription factors of a target gene; and (2) the relationship between the expression behavior of a target gene and that of its relevant transcription factors. Our method is based on rule induction, a machine learning technique that can efficiently deal with noisy domains. When applied to genomic location data with a confidence criterion relaxed to P-value = 0.005, and three different expression datasets of yeast S.cerevisiae, we obtained a set of regulatory rules describing the relationship between the expression behavior of a specific target gene and that of its relevant transcription factors. The resulting rules provide strong evidence of true positive gene-regulator interactions, as well as of protein-protein interactions that could serve to identify transcription complexes. AVAILABILITY: Supplementary files are available from http://www.jaist.ac.jp/~h-pham/regulatory-rules

Base Sequence↗

Detection and normalization of biases present in spotted cDNA microarray data: a composite method addressing dye, intensity-dependent, spatially-dependent, and print-order biases.

Microarrays are often used to identify target genes that trigger specific diseases, to elucidate the mechanisms of drug effects, and to check SNPs. However, data from microarray experiments are well known to contain biases resulting from the experimental protocols. Therefore, in order to elucidate biological knowledge from the data, systematic biases arising from their protocols must be removed prior to any data analysis. To remove these biases, many normalization methods are used by researchers. However, not all biases are eliminated from the microarray data because not all types of errors from experimental protocols are known. In this paper, we report an effective way of removing various types of biases by treating each microarray dataset independently to detect biases present in the dataset. After the biases contained in each dataset were identified, a combination of normalization methods specifically made for each dataset was applied to remove biases one at a time.

Algorithms↗

Support vector machines for prediction and analysis of beta and gamma-turns in proteins.

Tight turns have long been recognized as one of the three important features of proteins, together with alpha-helix and beta-sheet. Tight turns play an important role in globular proteins from both the structural and functional points of view. More than 90% tight turns are beta-turns and most of the rest are gamma-turns. Analysis and prediction of beta-turns and gamma-turns is very useful for design of new molecules such as drugs, pesticides, and antigens. In this paper we investigated two aspects of applying support vector machine (SVM), a promising machine learning method for bioinformatics, to prediction and analysis of beta-turns and gamma-turns. First, we developed two SVM-based methods, called BTSVM and GTSVM, which predict beta-turns and gamma-turns in a protein from its sequence. When compared with other methods, BTSVM has a superior performance and GTSVM is competitive. Second, we used SVMs with a linear kernel to estimate the support of amino acids for the formation of beta-turns and gamma-turns depending on their position in a protein. Our analysis results are more comprehensive and easier to use than the previous results in designing turns in proteins.

Algorithms↗

Utilizing weakly controlled vocabulary for sentence segmentation in biomedical literature.

Since biomedical texts contain a wide variety of domain specific terms, building a large dictionary to perform term matching is of great relevance. However, due to the existence of null boundary between adjacent terms, this matching is not a trivial problem. Moreover, it is known that generative words cannot be comprehensively included in a dictionary because their possible variations are infinite. In this study, we report our approach to dictionary building and term matching in biomedical texts. Large amount of terms with/without part-of-speech (POS) and/or category information were gathered, and a completion program generated approximately 1.36 million term variants to avoid stemming problems when matching terms. The dictionary was stored in a relational database management system (RDBMS) for quick lookup, and used by a matching program. Since the matching operation is not restricted to a substring surrounded by space characters, we can avoid the problem of null boundaries. This feature is also useful for generative words. Experimental results on GENIA corpus are promising: nearly half of the possible terms were correctly recognized as a meaningful segment, and most of the remaining half could be correctly recognized by some post-processing process, like chunking and further decomposition. It should be remarked that although we have not used term cost, connectivity cost, or syntactic information, reasonable segmentation and dictionary lookup were performed in most cases.

Abstracting and Indexing↗

Qualitatively predicting acetylation and methylation areas in DNA sequences.

Eukaryotic genomes are packaged by the wrapping of DNA around histone octamers to form nucleosomes. Nucleosome occupancy, acetylation, and methylation, which have a major impact on all nuclear processes involving DNA, have been recently mapped across the yeast genome using chromatin immunoprecipitation and DNA microarrays. However, this experimental protocol is laborious and expensive. Moreover, experimental methods often produce noisy results. In this paper, we introduce a computational approach to the qualitative prediction of nucleosome occupancy, acetylation, and methylation areas in DNA sequences. Our method uses support vector machines to discriminate between DNA areas with high and low relative occupancy, acetylation, or methylation, and rank k-gram features based on their support for these DNA modifications. Experimental results on the yeast genome reveal genetic area preferences of nucleosome occupancy, acetylation, and methylation that are consistent with previous studies. Supplementary files are available from http://www.jaist.ac.jp/~tran/nucleosome/.

Acetylation↗

Reconstruction of phylogenetic relationships from metabolic pathways based on the enzyme hierarchy and the gene ontology.

There has been much interest in the structural comparison and alignment of metabolic pathways. Several techniques have been conceived to assess the similarity of metabolic pathways of different organisms. In this paper, we show that the combination of a new heuristic algorithm for the comparison of metabolic pathways together with any of three enzyme similarity measures (hierarchical, information content, and gene ontology) can be used to derive a metabolic pathway similarity measure that is suitable for reconstructing phylogenetic relationships from metabolic pathways. Experimental results on the Glycolysis pathway of 73 organisms representing the three domains of life show that our method outperforms previous techniques.

Enzymes↗

Drug interaction ontology (DIO) for inferences of possible drug-drug interactions.

Drug Interaction Ontology (DIO) was developed for formal representation of pharmacological knowledge. It provides a fundamental framework for accumulation of reusable knowledge components in molecular pharmacology. Ontology was employed and implemented as a relational model. Some features include: 1) Drug-biomolecule interaction was assumed as a primitive knowledge element. 2) Symbolic representation was developed for drug-biomolecule interaction. Consequences of two conjugated units of interaction were defined by using symbols. These are applied for query development for identification of possible drug-drug interaction. 3) The triadic relationship model was developed as a ground model for bio-logical interactions and/or function, including semantic ones. One application of DIO is to support hypothesis generation of drug interaction by providing new hypotheses from a structured database storing literature information on known drug-biomolecule interactions. A knowledge base using DIO that contains information beginning with anti-cancer drugs is now under development. Detection of possible drug interaction was tested and its capacity to lead clinically known ones was confirmed. The system generated theoretically possible drug-drug interactions, which implies potential usefulness of new drugs to be tested before actual clinical application. In this paper, sorivudine and 5-fluorouracil mediated by dihydropyrimidine dehydrogenase are presented.

Arabinofuranosyluracil↗

Mining yeast transcriptional regulatory modules from factor DNA-binding sites and gene expression data.

UNLABELLED: In eukaryotes, gene expression is controlled by various transcription factors that bind to the promoter regions. Transcription factors may act positively, negatively or not at all. Different combinations of them may also activate or repress gene expression, and form regulatory networks of transcription. Uncovering such regulatory networks is a central challenge in genomic biology. In this study, we first defined a new kind of motifs in regulatory networks, transcriptional regulatory modules (TRMs), with the form factorset --> geneset, which emphasizes the combinatorial gene control of the group of factors factorset on the group of genes geneset. Second, we developed an efficient method based on a closed itemset mining technique for finding the two most informative kinds of TRMs, closed inf-TRMs and closed sup-TRMs, from factor DNA-binding sites and gene expression profiles data. The set of all closed inf-TRMs and closed sup-TRMs is often orders of magnitude smaller than the set of all TRMs but does not loss any information. When being applied to yeast data, our method produced results that are more compact, concise and comprehensive than those from previous studies to identify and interpret the transcriptional role of regulator combinations on sets of genes. AVAILABILITY: Supplementary files: http://www.jaist.ac.jp/~h-pham/regulation/.

Algorithms↗

Prediction and analysis of beta-turns in proteins by support vector machine.

Tight turn has long been recognized as one of the three important features of proteins after the alpha-helix and beta-sheet. Tight turns play an important role in globular proteins from both the structural and functional points of view. More than 90% tight turns are beta-turns. Analysis and prediction of beta-turns in particular and tight turns in general are very useful for the design of new molecules such as drugs, pesticides, and antigens. In this paper, we introduce a support vector machine (SVM) approach to prediction and analysis of beta-turns. We have investigated two aspects of applying SVM to the prediction and analysis of beta-turns. First, we developed a new SVM method, called BTSVM, which predicts beta-turns of a protein from its sequence. The prediction results on the dataset of 426 non-homologous protein chains by sevenfold cross-validation technique showed that our method is superior to the other previous methods. Second, we analyzed how amino acid positions support (or prevent) the formation of beta-turns based on the "multivariable" classification model of a linear SVM. This model is more general than the other ones of previous statistical methods. Our analysis results are more comprehensive and easier to use than previously published analysis results.

Amino Acid Sequence↗