PubMed Health⌕ Search

Biomedical subjects

Naum I Gershenzon

Publications and source records attributed to Naum I Gershenzon.

4 recordsLinked to original sources

The features of Drosophila core promoters revealed by statistical analysis.

BACKGROUND: Experimental investigation of transcription is still a very labor- and time-consuming process. Only a few transcription initiation scenarios have been studied in detail. The mechanism of interaction between basal machinery and promoter, in particular core promoter elements, is not known for the majority of identified promoters. In this study, we reveal various transcription initiation mechanisms by statistical analysis of 3393 nonredundant Drosophila promoters. RESULTS: Using Drosophila-specific position-weight matrices, we identified promoters containing TATA box, Initiator, Downstream Promoter Element (DPE), and Motif Ten Element (MTE), as well as core elements discovered in Human (TFIIB Recognition Element (BRE) and Downstream Core Element (DCE)). Promoters utilizing known synergetic combinations of two core elements (TATA_Inr, Inr_MTE, Inr_DPE, and DPE_MTE) were identified. We also establish the existence of promoters with potentially novel synergetic combinations: TATA_DPE and TATA_MTE. Our analysis revealed several motifs with the features of promoter elements, including possible novel core promoter element(s). Comparison of Human and Drosophila showed consistent percentages of promoters with TATA, Inr, DPE, and synergetic combinations thereof, as well as most of the same functional and mutual positions of the core elements. No statistical evidence of MTE utilization in Human was found. Distinct nucleosome positioning in particular promoter classes was revealed. CONCLUSION: We present lists of promoters that potentially utilize the aforementioned elements/combinations. The number of these promoters is two orders of magnitude larger than the number of promoters in which transcription initiation was experimentally studied. The sequences are ready to be experimentally tested or used for further statistical analysis. The developed approach may be utilized for other species.

Animals↗

Computational technique for improvement of the position-weight matrices for the DNA/protein binding sites.

Position-weight matrices (PWMs) are broadly used to locate transcription factor binding sites in DNA sequences. The majority of existing PWMs provide a low level of both sensitivity and specificity. We present a new computational algorithm, a modification of the Staden-Bucher approach, that improves the PWM. We applied the proposed technique on the PWM of the GC-box, binding site for Sp1. The comparison of old and new PWMs shows that the latter increase both sensitivity and specificity. The statistical parameters of GC-box distribution in promoter regions and in the human genome, as well as in each chromosome, are presented. The majority of commonly used PWMs are the 4-row mononucleotide matrices, although 16-row dinucleotide matrices are known to be more informative. The algorithm efficiently determines the 16-row matrices and preliminary results show that such matrices provide better results than 4-row matrices.

Algorithms↗

Promoter classifier: software package for promoter database analysis.

Promoter Classifier is a package of seven stand-alone Windows-based C++ programs allowing the following basic manipulations with a set of promoter sequences: (i) calculation of positional distributions of nucleotides averaged over all promoters of the dataset; (ii) calculation of the averaged occurrence frequencies of the transcription factor binding sites and their combinations; (iii) division of the dataset into subsets of sequences containing or lacking certain promoter elements or combinations; (iv) extraction of the promoter subsets containing or lacking CpG islands around the transcription start site; and (v) calculation of spatial distributions of the promoter DNA stacking energy and bending stiffness. All programs have a user-friendly interface and provide the results in a convenient graphical form. The Promoter Classifier package is an effective tool for various basic manipulations with eukaryotic promoter sequences that usually are necessary for analysis of large promoter datasets. The program Promoter Divider is described in more detail as a representative component of the package.

Algorithms↗

Synergy of human Pol II core promoter elements revealed by statistical sequence analysis.

MOTIVATION: The subject of our paper is bioinformatics analysis of the distinguishing features of human promoter DNA sequences, in particular of synergetic combinations of core promoter elements therein. We suppose that specific scenarios of transcription initiation are essentially related to various particular implementations of the interaction of basal transcription machinery with promoter DNA, depending on the presence and mutual positioning of core promoter elements. RESULTS: In addition to the combinations of core promoter elements previously experimentally confirmed [TATA box and Initiator (Inr), Downstream Promoter Element (DPE) and Inr, and TFIIB recognition element (BRE) and TATA box] we propose other alternate synergetic combinations: BRE and Inr, BRE and DPE, and TATA and DPE with respective models. The suggestion is based on a high statistical significance of the alternate combinations in promoters, comparable with the significance of the known combinations. We also present arguments that the BRE element is statistically more important than previously thought, and suggest possible mechanisms of action of the core elements in the promoters with multiple transcription start sites. CONTACT: ioschikhes-1@medctr.osu.edu SUPPLEMENTARY INFORMATION: Supplementary information is available at http://bmi.osu.edu/~ilya/synergy/Gershenzon_SuppMat-R.pdf.

DNA Polymerase II↗