Evolution of eukaryotic gene repertoire and gene structure: discovering the unexpected dynamics of genome evolution.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to V N Babenko.
Explore the source record for details and available documents.
Completion of the human genome sequence provides evidence for a gene count with lower bound 30,000-40,000. Significant protein complexity may derive in part from multiple transcript isoforms. Recent EST based studies have revealed that alternate transcription, including alternative splicing, polyadenylation and transcription start sites, occurs within at least 30-40% of human genes. Transcript form surveys have yet to integrate the genomic context, expression, frequency, and contribution to protein diversity of isoform variation. We determine here the degree to which protein coding diversity may be influenced by alternate expression of transcripts by exhaustive manual confirmation of genome sequence annotation, and comparison to available transcript data to accurately associate skipped exon isoforms with genomic sequence. Relative expression levels of transcripts are estimated from EST database representation. The rigorous in silico method accurately identifies exon skipping using verified genome sequence. 545 genes have been studied in this first hand-curated assessment of exon skipping on chromosome 22. Combining manual assessment with software screening of exon boundaries provides a highly accurate and internally consistent indication of skipping frequency. 57 of 62 exon skipping events occur in the protein coding regions of 52 genes. A single gene, (FBXO7) expresses an exon repetition. 59% of highly represented multi-exon genes are likely to express exon-skipped isoforms in ratios that vary from 1:1 to 1:>100. The proportion of all transcripts corresponding to multi-exon genes that exhibit an exon skip is estimated to be 5%.
Nucleotide sequences of the mitochondrial DNA (mtDNA) control region were studied in Germans living in the Altai, Russia. Although this ethnic group has been living in Russia for a long time, the obtained data indicate that its mitochondrial gene pool retains the main characteristics of the Western and Central European gene pools. Regarding the mitochondrial gene pool, Russian Germans were more similar to Germans living in Germany than to Russians with regard to the frequency of the Cambridge nucleotide sequence, frequencies and composition of five European haplotypic groups (classification of Richards et al.), and average intra- and interpopulation pairwise nucleotide differences. However, the mitochondrial gene pool of Altaian Germans also differed from that of Western European populations. The gene pool of Altaian Germans contained the ancestral variants of the main haplotypic groups. To date, these variants have not been found in modern Western and Central European populations, which is apparently due to their lower frequencies. In addition, some previously unknown mtDNA variants with specific nucleotide substitutions were found in Altaian Germans. The obtained results suggest that the modern mitochondrial gene pool of Europeans, including Germans from Germany, was largely affected by the demographic processes that occurred in the past two centuries. The Germans that lived in Russia were relatively isolated and, hence, retained more characteristics of the ancestral gene pool.
It is well known that non-coding mRNA sequences are dissimilar in many structural features. For individual mRNAs correlations were found for some of these features and their translational efficiency. However, no systematic statistical analysis was undertaken to relate protein abundance and structural characteristics of mRNA encoding the given protein. We have demonstrated that structural and contextual features of eukaryotic mRNAs encoding high- and low-abundant proteins differ in the 5' untranslated regions (UTR). Statistically, 5' UTRs of low-expression mRNAs are longer, their guanine plus cytosine content is higher, they have a less optimal context of the translation initiation codons of the main open reading frames and contain more frequently upstream AUG than 5' UTRs of high-expression mRNAs. Apart from the differences in 5' UTRs, high-expression mRNAs contain stronger termination signals. Structural features of low- and high-expression mRNAs are likely to contribute to the yield of their protein products.
Restriction fragment length polymorphism (RFLP) was studied in restriction sites AvaII, BamHI, EcoRV, KpnI, HaeIII, and RsaI of the mitochondrial DNA (mtDNA) D-loop in populations of Old Believers (Starovery) and in Slavic migrants in northern Siberia. Frequencies of rare variants of all polymorphic sites studied were estimated. The results were compared with the published data on mtDNA polymorphism sites studied were estimated. The results were compared with the published data on mtDNA polymorphism in Russian populations of central and southern Russia and in other Eastern Slavic, Caucasoid, and Mongolian populations. Significance of interpopulation differences with respect to distributions of variants of polymorphism sites was estimated with the use of the chi2 test. The comparison did not reveal significant differences between any groups of Eastern Slavs, including Old Believers. However, they significantly differed from both Mongols and Europeans (P < 0.05). To date, stable estimates of the polymorphism level in most of the restriction sites have been obtained for the Russian population. Regarding the distribution of the restriction-site variants of mtDNA, Russians significantly differ from the total European population, as well as from Mongols. In the populations of Old Believers, the effect of isolation on the diversity of the mitochondrial genome was demonstrated.
GeneExpress system has been designed to integrate description, analysis, and recognition of eukaryotic regulatory sequences. The system includes 5 basic units: (1) GeneNet contains an object-oriented database for accumulation of data on gene networks and signal transduction pathways and a Java-based viewer that allows an exploration and visualization of the GeneNet information; (2) Transcription Regulation combines the database on transcription regulatory regions of eukaryotic genes (TRRD) and TRRD Viewer; (3) Transcription Factor Binding Site Recognition contains a compilation of transcription factor binding sites (TFBSC) and programs for their analysis and recognition; (4) mRNA Translation is designed for analysis of structural and contextual features of mRNA 5'UTRs and prediction of their translation efficiency; and (5) ACTIVITY is the module for analysis and site activity prediction of a given nucleotide sequence. Integration of the databases in the GeneExpress is based on the Sequence Retrieval System (SRS) created in the European Bioinformatics Institute.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
A computer tool has been developed for revealing sets of oligonucleotides invariant for isofunctional families of DNA (RNA) and for using these in functional identification of nucleotide sequences. The tool allows one to: build up vocabularies of invariant oligonucleotides for the families of isofunctional nucleotide sequences; assess significance of the vocabularies; identify nucleotide sequences with the vocabularies of invariant oligonucleotides; determine the most effective identification parameters to minimize first and second type errors; assess the efficiency of identification of individual isofunctional families with the oligonucleotide vocabularies; determine the evolutionary characteristics of the families of isofunctional sequences on which vocabulary volume depends. Based on the system mentioned, we have analyzed a total of 322 protein-encoding gene families and have built up sets of invariant oligonucleotides, or again, oligonucleotide vocabularies that are characteristic of gene families and subfamilies. Identification of nucleotide sequences belonging to these families with the sets of invariant oligonucleotides revealed has been shown. Under the most effective identification parameters, the first type error (false negative) on control (independent) data was 10-15%, the second type error (false positive) was just 1-2 redundant sequences per sequence being examined. As has been shown, the volume of a vocabulary of invariant oligonucleotides depends on the percentage of variable positions in the multiple alignment within a family.
The results of genetic analysis of the differences in dopamine (DA) content between two lines of Drosophila virilis in response to heat stress are presented. The gene(s) controlling the differences in DA content under normal conditions is located on the X-chromosome, and the gene(s) controlling the differences in responsiveness of DA system to stress was found to be autosomal. By means of polyacrylamide gel electrophoresis was also examined phenoloxidase activity in these lines under normal conditions and short-term heat stress (60 min, 38 degrees C). Of the five phenol oxidase fractions identified, three were major, A1, A2 and A3 components, and two appeared occasionally. It seems likely that A1 is monophenol oxidase and A2 is diphenol oxidase. It was shown that neither mono, nor diphenol oxidase activities change under heat stress. It was found that D. virilis lines 147 and 101 differ in diphenol oxidase activity under normal conditions. Genetic analysis of these differences revealed that they are controlled by a single gene (or a group of closely linked genes). The gene controlling diphenol oxidase activity in D. virilis is located on the X-chromosome.
Nine asymptotic tests of the Hardy-Weinberg distribution have been analysed. It is shown that the usual chi 2 test is the most appropriate one, considering type I and II errors.
Explore the source record for details and available documents.
MOTIVATION: Despite the growing volume of data on primary nucleotide sequences, the regulatory regions remain a major puzzle with regard to their function. Numerous recognising programs considering a diversity of properties of regulatory regions have been developed. The system proposed here allows the specific contextual, conformational and physico-chemical properties to be revealed based on analysis of extended DNA regions. RESULTS: The Internet-accessible computer system RegScan, designed to analyse the extended regulatory regions of eukaryotic genes, has been developed. The computer system comprises the following software: (i) programs for classification dividing a set of promoters into TATA-containing and TATA-less promoters and promoters with and without CpG islands; (ii) programs for constructing (a) nucleotide frequency profiles, (b) sequence complexity profiles and (c) profiles of conformational and physico-chemical properties; (iii) the program for constructing the sets of degenerate oligonucleotide motifs of a specified length; and (iv) the program searching for and visualising repeats in nucleotide sequences. The system has allowed us to demonstrate the following characteristic patterns of vertebrate promoter regions: the TATA box region is flanked by regions with an increased G+C content and increased bending stiffness, the TATA box content is asymmetric and promoter regions are saturated with both direct and inverted repeats. AVAILABILITY: The computer system RegScan is available via the Internet at http://www.mgs.bionet.nsc. ru/Systems/RegScan, http://www.cbil.upenn.edu/mgs/systems/r egscan/.
It is suggest to use the Kendall's rank correlation coefficient tau for the analysis of correlation between mutation spectra. The results of computer simulations and analysis of real spectra showed that the approach can be recommended for the analysis of mutation spectra.
Promoter regions of eukaryotic genes were analyzed for the presence of repeated fragments. It was found that the promotor sequences are abundant in direct, symmetrical, and inverted repeats. A computer system for searching and visualizing the repeats was developed.