PubMed Health⌕ Search

Biomedical subjects

S S Hannenhalli

Publications and source records attributed to S S Hannenhalli.

3 recordsLinked to original sources

Identification and cross-species comparison of canine osteoarthritic gene regulatory cis-elements.

OBJECTIVE: To better understand transcription regulation of osteoarthritis (OA) by examining common promoter motifs in canine osteoarthritic genes, to identify other genes containing these motifs and to assess the conservation of these motifs between canine, human, mouse and rat. DESIGN: Differentially expressed transcripts in canine OA were mapped to the human genome. We thus identified 20 orthologous human transcripts representing 19 up-regulated genes and 62 orthologous transcripts representing 60 down-regulated genes. The 5 kbp upstream regions of these transcripts were used to identify binding sites and build promoter models based on those sites. The human genome was subsequently searched for other transcripts likely to be regulated by the same promoter models. Orthologous transcripts were then identified in canine, rat and mouse for determination of potential cross-species conservation of binding sites comprising the promoter model. RESULTS: Four promoter models containing 5-6 transcripts and 5-8 common transcription factor binding sites were developed. They include binding sites for AP-4, AP-2alpha and gamma, and E2F. Several hundred other human genes were found to contain these promoter motifs. Furthermore these motifs were significantly over represented in the orthologous genes in canine, rat and mouse genomes. CONCLUSIONS: We have developed and applied a computational methodology to identify common promoter elements implicated in OA and shared amongst four higher vertebrates. The transcription factors associated with these binding sites and other genes driven by these promoter motifs have been implicated in OA, chondrocyte development and with other biological factors involved in the disease.

Animals↗

Analysis and prediction of functional sub-types from protein sequence alignments.

The increasing number and diversity of protein sequence families requires new methods to define and predict details regarding function. Here, we present a method for analysis and prediction of functional sub-types from multiple protein sequence alignments. Given an alignment and set of proteins grouped into sub-types according to some definition of function, such as enzymatic specificity, the method identifies positions that are indicative of functional differences by comparison of sub-type specific sequence profiles, and analysis of positional entropy in the alignment. Alignment positions with significantly high positional relative entropy correlate with those known to be involved in defining sub-types for nucleotidyl cyclases, protein kinases, lactate/malate dehydrogenases and trypsin-like serine proteases. We highlight new positions for these proteins that suggest additional experiments to elucidate the basis of specificity. The method is also able to predict sub-type for unclassified sequences. We assess several variations on a prediction method, and compare them to simple sequence comparisons. For assessment, we remove close homologues to the sequence for which a prediction is to be made (by a sequence identity above a threshold). This simulates situations where a protein is known to belong to a protein family, but is not a close relative of another protein of known sub-type. Considering the four families above, and a sequence identity threshold of 30 %, our best method gives an accuracy of 96 % compared to 80 % obtained for sequence similarity and 74 % for BLAST. We describe the derivation of a set of sub-type groupings derived from an automated parsing of alignments from PFAM and the SWISSPROT database, and use this to perform a large-scale assessment. The best method gives an average accuracy of 94 % compared to 68 % for sequence similarity and 79 % for BLAST. We discuss implications for experimental design, genome annotation and the prediction of protein function and protein intra-residue distances.

Adenylyl Cyclases↗

Bacterial start site prediction.

With the growing number of completely sequenced bacterial genes, accurate gene prediction in bacterial genomes remains an important problem. Although the existing tools predict genes in bacterial genomes with high overall accuracy, their ability to pinpoint the translation start site remains unsatisfactory. In this paper, we present a novel approach to bacterial start site prediction that takes into account multiple features of a potential start site, viz., ribosome binding site (RBS) binding energy, distance of the RBS from the start codon, distance from the beginning of the maximal ORF to the start codon, the start codon itself and the coding/non-coding potential around the start site. Mixed integer programing was used to optimize the discriminatory system. The accuracy of this approach is up to 90%, compared to 70%, using the most common tools in fully automated mode (that is, without expert human post-processing of results). The approach is evaluated using Bacillus subtilis, Escherichia coli and Pyrococcus furiosus. These three genomes cover a broad spectrum of bacterial genomes, since B.subtilis is a Gram-positive bacterium, E.coli is a Gram-negative bacterium and P. furiosus is an archaebacterium. A significant problem is generating a set of 'true' start sites for algorithm training, in the absence of experimental work. We found that sequence conservation between P. furiosus and the related Pyrococcus horikoshii clearly delimited the gene start in many cases, providing a sufficient training set.

Algorithms↗