PubMed Health⌕ Search

Biomedical subjects

I B Rogozin

Publications and source records attributed to I B Rogozin.

At least 19 recordsLinked to original sources

Mutagenic specificity of the base analog 6-N-hydroxylaminopurine in the LYS2 gene of yeast Saccharomyces cerevisiae.

We used the LYS2 gene mutational system to study mutation specificity of the base analog 6-N-hydroxylaminopurine (HAP) in yeast. We characterized phenotypes of mutations using codon-specific nonsense suppressors and the test employing inactivation of the release factor Sup35 due to overexpression and formation of prion-like derivative [PSI]. We have shown that HAP induces predominantly nonsense mutations. While the tests using codon-specific nonsense-suppressors allowed to identify only about 50% of nonsense-mutations, all the nonsense-mutations were identified in the test with defective Sup35. We determined and analyzed the spectrum of HAP-induced nucleotide changes in two regions of the gene. HAP induces predominantly GC-->AT transitions in a hotspots of a central position of trinucleotide GGA or AGG. Directionality of these transitions is consistent with the idea that initial dHAPMP incorporation in the leading strand is more genetically dangerous than in lagging DNA strand. We revealed a specific context inhibitory for HAP mutagenesis, a "T" in -1 position to mutation site.

Adenine↗

Comparative study and prediction of DNA fragments associated with various elements of the nuclear matrix.

Scaffold/matrix-associated region (S/MAR) sequences are DNA regions that are attached to the nuclear matrix, and participate in many cellular processes. The nuclear matrix is a complex structure consisting of various elements. In this paper we compared frequencies of simple nucleotide motifs in S/MAR sequences and in sequences extracted directly from various nuclear matrix elements, such as nuclear lamina, cores of rosette-like structures, synaptonemal complex. Multivariate linear discriminant analysis revealed significant differences between these sequences. Based on this result we have developed a program, ChrClass (Win/NT version, ftp.bionet.nsc.ru/pub/biology/chrclass/chrclass.zip), for the prediction of the regions associated with various elements of the nuclear matrix in a query sequence. Subsequently, several test samples were analyzed by using two S/MAR prediction programs (a ChrClass and MAR-Finder) and a simple MRS criterion (S/MAR recognition signature) indicating the presence of S/MARs. Some overlap between the predictions of all MAR prediction tools has been found. Simultaneous use of the ChrClass, MRS criterion and MAR-Finder programs may help to obtain a more clearcut picture of S/MAR distribution in a query sequence. In general, our results suggest that the proportion of missed S/MARs is lower for ChrClass, whereas the proportion of wrong S/MARs is lower for MAR-Finder and MRS.

Animals↗

Somatic mutation hotspots correlate with DNA polymerase eta error spectrum.

Mutational spectra analysis of 15 immunoglobulin genes suggested that consensus motifs RGYW and WA were universal descriptors of somatic hypermutation. Highly mutable sites, "hotspots", that matched WA were preferentially found in one DNA strand and RGYW hotspots were found in both strands. Analysis of base-substitution hotspots in DNA polymerase error spectra showed that 33 of 36 hotspots in the human polymerase eta spectrum conformed to the WA consensus. This and four other characteristics of polymerase eta substitution specificity suggest that errors introduced by this enzyme during synthesis of the nontranscribed DNA strand in variable regions may contribute to strand-specific somatic hypermutagenesis of immunoglobulin genes at A-T base pairs.

Amino Acid Motifs↗

Cloning and functional analysis of SEL1L promoter region, a pancreas-specific gene.

We examined the promoter activity of SEL1L, the human ortholog of the C. elegans gene sel-1, a negative regulator of LIN-12/NOTCH receptor proteins. To understand the relation in SEL1L transcription pattern observed in different epithelial cells, we determined the transcription start site and sequenced the 5' flanking region. Sequence analysis revealed the presence of consensus promoter elements--GC boxes and a CAAT box--but the absence of a TATA motif. Potential binding sites for transcription factors that are involved in tissue-specific gene expression were identified, including: activator protein-2 (AP-2), hepatocyte nuclear factor-3 (HNF3 beta), homeobox Nkx2-5 and GATA-1. Transcription activity of the TATA-less SEL1L promoter was analyzed by transient transfection using luciferase reporter gene constructs. A core basal promoter of 302 bp was sufficient for constitutive promoter activity in all the cell types studied. This genomic fragment contains a CAAT and several GC boxes. The activity of the SEL1L promoter was considerably higher in mouse pancreatic beta cells (beta TC3) than in several human pancreatic neoplastic cell lines; an even greater reduction of its activity was observed in cells of nonpancreatic origin. These results suggest that SEL1L promoter may be a useful tool in gene therapy applications for pancreatic pathologies.

Animals↗

Characterization of the genomic Xist locus in rodents reveals conservation of overall gene structure and tandem repeats but rapid evolution of unique sequence.

The Xist locus plays a central role in the regulation of X chromosome inactivation in mammals, although its exact mode of action remains to be elucidated. Evolutionary studies are important in identifying conserved genomic regions and defining their possible function. Here we report cloning, sequence analysis, and detailed characterization of the Xist gene from four closely related species of common vole (field mouse), Microtus arvalis. Our analysis reveals that there is overall conservation of Xist gene structure both between different vole species and relative to mouse and human Xist/XIST. Within transcribed sequence, there is significant conservation over five short regions of unique sequence and also over Xist-specific tandem repeats. The majority of unique sequences, however, are evolving at an unexpectedly high rate. This is also evident from analysis of flanking sequences, which reveals a very high rate of rearrangement and invasion of dispersed repeats. We discuss these results in the context of Xist gene function and evolution.

3' Untranslated Regions↗

Genome alignment, evolution of prokaryotic genome organization, and prediction of gene function using genomic context.

Gene order in prokaryotes is conserved to a much lesser extent than protein sequences. Only several operons, primarily those that code for physically interacting proteins, are conserved in all or most of the bacterial and archaeal genomes. Nevertheless, even the limited conservation of operon organization that is observed can provide valuable evolutionary and functional clues through multiple genome comparisons. A program for constructing gapped local alignments of conserved gene strings in two genomes was developed. The statistical significance of the local alignments was assessed using Monte Carlo simulations. Sets of local alignments were generated for all pairs of completely sequenced bacterial and archaeal genomes, and for each genome a template-anchored multiple alignment was constructed. In most pairwise genome comparisons, <10% of the genes in each genome belonged to conserved gene strings. When closely related pairs of species (i.e., two mycoplasmas) are excluded, the total coverage of genomes by conserved gene strings ranged from <5% for the cyanobacterium Synechocystis sp to 24% for the minimal genome of Mycoplasma genitalium, and 23% in Thermotoga maritima. The coverage of the archaeal genomes was only slightly lower than that of bacterial genomes. The majority of the conserved gene strings are known operons, with the ribosomal superoperon being the top-scoring string in most genome comparisons. However, in some of the bacterial-archaeal pairs, the superoperon is rearranged to the extent that other operons, primarily those subject to horizontal transfer, show the greatest level of conservation, such as the archaeal-type H+-ATPase operon or ABC-type transport cassettes. The level of gene order conservation among prokaryotic genomes was compared to the cooccurrence of genomes in clusters of orthologous genes (COGs) and to the conservation of protein sequences themselves. Only limited correlation was observed between these evolutionary variables. Gene order conservation shows a much lower variance than the cooccurrence of genomes in COGs, which indicates that intragenome homogenization via recombination occurs in evolution much faster than intergenome homogenization via horizontal gene transfer and lineage-specific gene loss. The potential of using template-anchored multiple-genome alignments for predicting functions of uncharacterized genes was quantitatively assessed. Functions were predicted or significantly clarified for approximately 90 COGs (approximately 4% of the total of 2414 analyzed COGs). The most significant predictions were obtained for the poorly characterized archaeal genomes; these include a previously uncharacterized restriction-modification system, a nuclease-helicase combination implicated in DNA repair, and the probable archaeal counterpart of the eukaryotic exosome. Multiple genome alignments are a resource for studies on operon rearrangement and disruption, which is central to our understanding of the evolution of prokaryotic genomes. Because of the rapid evolution of the gene order, the potential of genome alignment for prediction of gene functions is limited, but nevertheless, such predictions information significantly complements the results obtained through protein sequence and structure analysis.

Computational Biology↗

Purifying selection and birth-and-death evolution in the ubiquitin gene family.

Ubiquitin is a highly conserved protein that is encoded by a multigene family. It is generally believed that this gene family is subject to concerted evolution, which homogenizes the member genes of the family. However, protein homogeneity can be attained also by strong purifying selection. We therefore studied the proportion (p(S)) of synonymous nucleotide differences between members of the ubiquitin gene family from 28 species of fungi, plants, and animals. The results have shown that p(S) is generally very high and is often close to the saturation level, although the protein sequence is virtually identical for all ubiquitins from fungi, plants, and animals. A small proportion of species showed a low level of p(S) values, but these values appeared to be caused by recent gene duplication. It was also found that the number of repeat copies of the gene family varies considerably with species, and some species harbor pseudogenes. These observations suggest that the members of this gene family evolve almost independently by silent nucleotide substitution and are subjected to birth-and-death evolution at the DNA level.

Animals↗

Similarity pattern analysis in mutational distributions.

The validity and applicability of the statistical procedure - similarity pattern analysis (SPAN) - to the study of mutational distributions (MDs) was demonstrated with two sets of data. The first was mutational spectra (MS) for 697 GC to AT transitions produced with eight alkylating agents (AAs) in the lacI gene of Escherichia coli. The second was a recently summarized data on the distributions of 11562 spontaneous, radiation- and chemical-induced forward mutations in the ad-3 region of heterokaryon 12 of Neurospora crassa. They were analyzed as large two-way contingency tables (CTs) where two kinds of profiles were compared: site (or genotypic class) profiles and origin (or mutagen) profiles. To measure similarity (homogeneity) between any pair of profiles, the relevant sufficient statistics, Kastenbaum-Hirotsu squared distance (KHi(2)), was used. Collapsing the similar profiles into distinct internally homogeneous clusters named 'collapsets' revealed their similarity pattern. To facilitate the procedure, the computer program, COLLAPSE, was elaborated. The results of SPAN for the lacI spectra were found comparable with the results of their previous analysis with two multivariate statistical methods, the factor and cluster analyses. In the ad-3 data set, five collapsets were revealed among origin profiles (OPs): (I) ENU = 4NQO = 4HAQO = FANFT = SQ18506; (II) AF-2 = EI = MMS = DEP; (III) ETO = UV; (IV) AHA = PROCARB; and (V) He ions = protons. Moreover, the previous observation that MDs are dose-dependent was confirmed for X-ray-induced MDs. Profiles induced with the low doses of X-rays are similar to that induced with 85Sr, and profiles induced with the medium X-ray doses to those induced with protons and He ions. Evaluated similarities appear to be rather reasonable: mutagens with similar mode of action induce similar MDs. Similarity pattern revealed among genotypic class profiles (GCPs) seems to be also interpretable. When supplemented with descriptive cluster analysis, SPAN appears to be a fruitful methodology in MS analysis.

Chromosome Mapping↗

Protein-coding regions prediction combining similarity searches and conservative evolutionary properties of protein-coding sequences.

The gene identification procedure in a completely new gene with no good homology with protein sequences can be a very complex task. In order to identify the protein-coding region, a new method, 'SYNCOD', based on the analysis of conservative evolutionary properties of coding regions, has been realized. This program is able to identify and use the coding region homologies of the non-annotated (unknown) protein-coding sequences already present in the nucleotide sequence databases by using the alignment produced by BLASTN. The ratio of number mismatches resulting in synonymous codons to the number of mismatches resulting in non-synonymous codons is estimated for each open reading frame. Monte Carlo simulations are then used to estimate the significance of the ratio deviation from random behavior. The SYNCOD program has been tested on generated random sequences and on different control sets. The high accuracy of predicting protein-coding regions (the correlation coefficient, CC, varies from 0.67 to 0.79) and the high specificity (the portion of wrong exons, WE, varies from 0.06 to 0.07) have proved to be important features of the suggested approach. The SYNCOD program is resident on the ITBA-CNR Web Server and can be used via the Internet (URL: www.itba.mi.cnr.it/webgene).

Algorithms↗

Characterization of several LINE-1 elements in Microtus kirgisorum.

Several LINE-1s have been isolated and characterized from genomic DNA of the vole, Microtus kirgisorum. Blot hybridization revealed specific restriction patterns of L1 elements in vole genomes. Rehybridization of the genomic blot with a cloned 5'-end fragment revealed two major bands indicating the presence of two different L1 subfamilies. The copy numbers are estimated for different parts of M. kirgisorum L1 elements. Data also demonstrate that most vole L1 elements are truncated at the 5'-end; however, in contrast to mouse, the ORF1 copy number is higher in vole. A difference between the substitution rates of the ORF1 5'-region (approximately 330 nucleotides) and the rest of the L1 coding regions is revealed.

Amino Acid Sequence↗

The subclass approach for mutational spectrum analysis: application of the SEM algorithm.

Analysis and comparison of mutational spectra represents an important problem in molecular biology. To analyse a mutational spectra we apply an algorithm based on the SEM subclass approach (Simulation, Expectation, Maximization). The algorithm tries to classify the mutational sites according to different mutation probabilities, and each site should belong to one class. Each class is approximated by binomial distribution and thus any real mutational spectrum is regarded as a mixture of binomial distributions. The separation process runs iteratively. Each iteration includes the simulation, maximization and estimation procedures. To evaluate the quality of the classification results, the X2 test is used. The algorithm has been checked on random spectra with preset parameters and on real mutational spectra. As has been shown, 17 out of 19 analysed real mutational spectra can be divided into two or more classes of sites, of which one contains hotspots of mutation. For the G:C-->A:T mutational spectra induced by Sn1 alkylating mutagenes (11 spectra) the classification accuracy was 0.95. To test different site volumes, each Sn1-induced spectrum was divided into the G-->A and C-->T spectra. The classification accuracy for these spectra was 0.96. From the analysis of classification errors it is possible to suggest that at least part of them cannot be ascribed to the faults of the algorithm but are caused by some special features of the mutagenesis itself. The results of the real data are in good relation with existing knowledge. The approach we present is an attempt to formalize the concept of a "mutational hotspot". The program implementing the SEM algorithm is available on the Web server (http:/(/)www.itba.mi.cnr.it/webmutation).

Algorithms↗

Multiple antimutagenesis mechanisms affect mutagenic activity and specificity of the base analog 6-N-hydroxylaminopurine in bacteria and yeast.

Base analog 6-N-hydroxylaminopurine is a potent mutagen in variety of prokaryotic and eukaryotic organisms. In the review, we discuss recent results of the studies of HAP mutagenic activity, genetic control and specificity in bacteria and yeast with the emphasis to the mechanisms protecting living cells from mutagenic and toxic effects of this base analog.

Adenine↗

A novel subfamily of LINE-derived elements in mice.

Hybrid sequences have been described previously that consist of a 5' region homologous to ORF2 of LINEs and a 3' end that shares homology with a sequence located in the first intron of Cepsilon immunoglobulin. The present investigation has revealed 14 new sequences from seven murine species, that show high homology to those observed earlier. Database search has found several new homologous hybrid sequences including one located in the mouse T-cell receptor (Tcra) locus. Several interesting features of this sequence include identical 15-bp flanking short direct repeats as well as poly-A signal and A-rich sequence at the 3' end. We have classified this set of sequences as LINE-derived elements (LDEs), which constitute a newly observed subfamily. Comparative analysis of these sequences suggests that a single recombination event was responsible for the production of an LDE progenitor. The phylogenetic tree shows a number of elements that pre-existed in the common ancestor of murine species and displays different evolutionary rates. The time of LDE origin is estimated at approximately 10-15 MYA.

Animals↗

Repetitive DNA sequences in the common vole: cloning, characterization and chromosome localization of two novel complex repeats MS3 and MS4 from the genome of the East European vole Microtus rossiaemeridionalis.

We have characterized two novel, complex, heterochromatic repeat sequences, MS3 and MS4, isolated from Microtus rossiaemeridionalis genomic DNA. Sequence analysis indicates that both repeats consist of unique sequences interrupted by repeat elements of different origin and can be classified as long complex repeat units (LCRUs). A unique feature of both repeat units is the presence of short interspersed repeat elements (SINEs), which are usually characteristic of the euchromatic part of the genome. Comparative analysis revealed no significant stretches of homology in the nucleotide sequences between the two repeats, suggesting that the repeats originated independently during the course of vole genome evolution. Fluorescence in situ hybridization analysis demonstrates that MS3 and MS4 occupy distinct domains in the heterochromatic regions of the sex chromosomes in M. transcaspicus and M. arvalis but collocalize in M. rossiaemeridionalis and M. kirgisorum heterochromatic blocks. The localization pattern of the repeats on the vole chromosomes confirms the independent origin of the two repeats and suggests that expansion of the heterochromatic blocks has occurred subsequent to speciation.

Animals↗

Identification of a U7snRNA homologue mapping to the human Xq27.1 region, between the DXS1232 and DXS119 loci.

To contribute to the identification and analysis of novel genes, we undertook the study of a cosmid clone in the Xq27 region of human DNA. The cloned fragment was previously observed to have a high number of evolutionarily conserved sequences. In this genomic stretch of DNA we have identified sequence homologous to the U7 RNA gene including its potential regulatory elements. This paper describes the genomic organisation of this gene and its mapping to the Xq27.1 genomic sub-interval between the DXS1232 and DXS119 loci.

Animals↗

Analysis of donor splice sites in different eukaryotic organisms.

We present here a new algorithm for functional site analysis. It is based on four main assumptions: each variation of nucleotide composition makes a different contribution to the overall binding free energy of interaction between a functional site and another molecule; nonfunctioning site-like regions (pseudosites) are absent or rare in genomes; there may be errors in the sample of sites; and nucleotides of different site positions are considered to be mutually dependent. In this algorithm, the site set is divided into subsets, each described by a certain consensus. Donor splice sites of the human protein-coding genes were analyzed. Comparing the results with other methods of donor splice site prediction has demonstrated a more accurate prediction of consensus sequences AG/GU(A,G), G/GUnAG, /GU(A,G)AG, /GU(A,G)nGU, and G/GUA than is achieved by weight matrix and consensus (A,C)AG/GU(A,G)AGU with mismatches. The probability of the first type error, E1, for the obtained consensus set was about 0.05, and the probability of the second type error, E2, was 0.15. The analysis demonstrated that accuracy of the functional site prediction could be improved if one takes into account correlations between the site positions. The accuracy of prediction by using human consensus sequences was tested on sequences from different organisms. Some differences in consensus sequences for the plant Arabidopsis sp., the invertebrate Caenorhabditis sp., and the fungus Aspergillus sp. were revealed. For the yeast Saccharomyces sp. only one conservative consensus, /GUA(U,A,C)G(U,A,C), was revealed (E1 = 0.03, E2 = 0.03). Yeast is a very interesting model to use for analysis of molecular mechanisms of splicing.

Algorithms↗