PubMed Health⌕ Search

Biomedical subjects

Alexander E Kel

Publications and source records attributed to Alexander E Kel.

7 recordsLinked to original sources

FeatureScan: revealing property-dependent similarity of nucleotide sequences.

FeatureScan is a software package aiming to reveal novel types of DNA sequence similarity by comparing physico-chemical properties. Thirty-eight different parameters of DNA double strands such as charge, melting enthalpy, conformational parameters and the like are provided. As input FeatureScan requires two sequences, a pattern sequence and a target sequence, search conditions are set by selecting a specific DNA parameter and a threshold value. Search results are displayed in FASTA format and directly linked to external genome databases/browsers (ENSEMBL, NCBI, UCSC). An Internet version of FeatureScan is accessible at http://genome.gbf.de/featurescan/. As part of the HOBIT initiative (http://hobit.sourceforge.net/) FeatureScan is also accessible as a web service at its above home page. Currently, several preloaded genomes are provided at this Internet website (Homo sapiens, Mus musculus, Rattus norvegicus and four strains of Escherichia coli) as target sequences. Standalone executables of FeatureScan are available on request.

Animals↗

Signal-theoretical DNA similarity measure revealing unexpected similarities of E. coli promoters.

We present an implementation of the signal theory based approach for detection of novel types of DNA similarity which are based on physical properties of DNA. Systematic study of the sensitivity of the new similarity measure revealed qualitative differences to letter-based similarity. A variety of physical parameters of DNA double strands, which in a straightforward way reflect different kinds of information hidden behind the primary structure of DNA, showed a wide range of recognition power of the signal similarity measure. We applied the novel DNA similarity measure for the analysis of promoters of E.coli genes. We found that promoter similarities revealed by our approach correlate with their transcription regulatory responsivenesses to different antibiotic and osmotic treatments. Accelerated by special hardware for fast Fourier transformations, the method is easily applicable for the analysis of entire eukaryotic genomes in minutes.

DNA, Bacterial↗

Large-scale collection and characterization of promoters of human and mouse genes.

We report the generation and initial characterization of a large-scale collection of sequences of putative promoter regions (PPRs) of human and mouse genes. Based on our unique collection of 400,225 and 580,209 human and mouse full-length cDNAs, we determined exact transcriptional start sites (TSSs). Using positional information of the TSSs, we could retrieve adjacent sequences as PPRs for 8,793 and 6,875 human and mouse genes, respectively. The positions of the PPRs were 4 kb upstream to previously reported 5'-ends of cDNAs on average, demonstrating that full-length cDNA information is indispensable for this purpose. Among those PPRs supported by experimentally validated TSSs, 3,324 could be paired as mutually homologous genes between human and mouse and were used for the comprehensive comparative studies. The sequence identities in the proximal regions of the TSSs were 45% on average, and 22,794 putative transcription factor binding sites that are conserved between human and mouse were identified. The data resource created in the present work and the results of the sequences' initial characterization should lay the firm foundation for deciphering the transcriptional modulations of human genes. All the data were deposited and made available through a database for comparative studies, DBTSS.

Animals↗

Systematic DNA-binding domain classification of transcription factors.

Based on the manual annotation of transcription factors stored in the TRANSFAC database, we developed a library of hidden Markov models (HMM) to represent their DNA-binding domains and used it for a comprehensive classification. The models constructed were applied on the UniProt/Swiss-Prot database, leading to a systematic classification of further DNA-binding protein entries. The HMM library obtained can be used to classify any newly discovered transcription factor according to its DNA-binding domain and, thus, to generate hypotheses about its DNA-binding specificity.

Binding Sites↗

Composition-sensitive analysis of the human genome for regulatory signals.

Known transcription regulatory signals which generally act as transcription factor binding sites (TFs) differ significantly in their base composition. Therefore, their occurrence in a genome largely depends on the local base composition. In an attempt to initiate an all human genome analysis for the occurrence of potential TFs, we systematically analyzed the GC-content of distinct functional regions (e. g., upstream and downstream gene regions, exons, long and short introns, repetitive elements) and correlated the frequencies of potential binding sites of a representative set of TFs in these regions. For these analyses, we used the pattern collection of the TRANSFAC database on transcriptional regulation, the information about functionally relevant combinations of them from the database TRANSCompel, and our new resource, TRANSGenomeTM, which provides an overall annotation of the human genome with emphasis on its regulatory characteristics. We show that the occurrence of sequence patterns with regulatory potential may be supported by, but cannot be fully explained by either the GC content of a whole chromosome or its putative promoter regions, nor by the information content of the patterns. Several patterns, HNF-3, NFAT, and GC box, show a clear overrepresentation in all promoter groups as well as in all chromosomes. Other patterns, like E2F and CRE-BP1, are underrepresented in all promoter groups as well as in all chromosomes in comparison with random sequences. Simultaneously, both patterns are over-represented in promoters in comparison with repetitive elements. We define several structural characteristics of the proximal promoters that differentiate them from other functional genomic regions. Two well-known promoter elements, GC- and TATA-boxes, are statistically enriched in promoters in comparison with random sequences, repetitive elements and exons. Altogether, our findings provide insights into the macroheterogeneity amongst the individual chromosomes, into the microheterogeneity among different functional regions of individual chromosomes, contribute to further understanding of structural organization of gene regulatory regions, and give first hints on the development of regulatory features during evolution.

Animals↗

Prediction of potential C/EBP/NF-kappaB composite elements using matrix-based search methods.

Bacterial infections trigger a wide range of host cell responses. For the interaction of Pseudomonas aeruginosa and epithelial cells it is known that transcription factor NF-kappaB plays a central role, but its effects have to be specified by cooperation with additional factors. NF-B containing composite elements, e. g. with C/EBP, may be appropriate indicators for new antibacterial response genes. We refined matrix-based search methods for C/EBP, which was necessary because of weak consensi of the previously existing C/EBP matrices, established a model for C/EBP/ NF-kappaB composite element, used it for scanning all known human 5'-flanking sequences and identified 135 new candidate genes. The newly constructed C/EBP binding patterns will be available with one of the next releases of the TRANSFAC database (http://www.gene-regulation.de).

Base Sequence↗

TRANSCompel: a database on composite regulatory elements in eukaryotic genes.

Originating from COMPEL, the TRANSCompel database emphasizes the key role of specific interactions between transcription factors binding to their target sites providing specific features of gene regulation in a particular cellular content. Composite regulatory elements contain two closely situated binding sites for distinct transcription factors and represent minimal functional units providing combinatorial transcriptional regulation. Both specific factor--DNA and factor--factor interactions contribute to the function of composite elements (CEs). Information about the structure of known CEs and specific gene regulation achieved through such CEs appears to be extremely useful for promoter prediction, for gene function prediction and for applied gene engineering as well. Each database entry corresponds to an individual CE within a particular gene and contains information about two binding sites, two corresponding transcription factors and experiments confirming cooperative action between transcription factors. The COMPEL database, equipped with the search and browse tools, is available at http://www.gene-regulation.com/pub/databases.html#transcompel. Moreover, we have developed the program CATCH for searching potential CEs in DNA sequences. It is freely available as CompelPatternSearch at http://compel.bionet.nsc.ru/FunSite/CompelPatternSearch.html.

Animals↗