PubMed Health⌕ Search

Biomedical subjects

Alexander Bolshoy

Publications and source records attributed to Alexander Bolshoy.

12 recordsLinked to original sources

Involvement of DNA curvature in intergenic regions of prokaryotes.

It is known that DNA curvature plays a certain role in gene regulation. The distribution of curved DNA in promoter regions is evolutionarily preserved, and it is mainly determined by temperature of habitat. However, very little is known on the distribution of DNA curvature in termination sites. Our main objective was to comprehensively analyze distribution of curved sequences upstream and downstream to the coding genes in prokaryotic genomes. We applied CURVATURE software to 170 complete prokaryotic genomes in a search for possible typical distribution of DNA curvature around starts and ends of genes. Performing cluster analyses and other statistical tests, we obtained novel results regarding various factors influencing curvature distribution in intergenic regions, such as growth temperature, A+T composition and genome size. We also analyzed intergenic regions between converging genes in 15 selected genomes. The results show that six genomes presented peaks of curvature excess larger than 3 SDs. Insufficient statistics did not allow us to draw further conclusion. Our hypothesis is that DNA curvature could affect transcription termination in many prokaryotes either directly, through contacts with RNA polymerase, or indirectly, via contacts with some regulatory proteins.

Cluster Analysis↗

Large-scale genome clustering across life based on a linguistic approach.

With the availability of genome sequences, the possibility of new phylogenetic reconstructions arises in order to reveal genomic relationships among organisms. According to the compositional-spectra (CS) approach proposed in our previous studies, any genomic sequence can be characterized by a distribution of frequencies of imperfect matching of words (oligonucleotides). In the current application of CS-analysis, we attempted to analyze the cluster structure of genomes across life. It appeared that compositional spectra show a clear three-group clustering of the compared prokaryotic and eukaryotic genomes. Unexpectedly, this grouping seriously differs from the classical Universal Tree of Life structure represented by common kingdoms known as Eubacteria, Archaebacteria, and Eukarya. The revealed CS-clustering displays high stability, putatively reflecting its objective nature, and still enigmatic biological significance that may result from convergent evolution driven by ecological selection. We believe that our approach provides a new and wider (compared to traditional methods) perspective of extracting genomic information of high evolutionary relevance.

Base Composition↗

Centromere parC of plasmid R1 is curved.

The centromere sequence parC of Escherichia coli low-copy-number plasmid R1 consists of two sets of 11 bp iterated sequences. Here we analysed the intrinsic sequence-directed curvature of parC by its migration anomaly in polyacrylamide gels. The 159 bp long parC is strongly curved with anomaly values (k-factors) close to 2. The properties of the parC curvature agree with those of other curved DNA sequences. parC contains two regions of 5-fold repeated iterons separated by 39 bp. We modified 4 bp within this intermediate sequence so that we could analyse the two 5-fold repeated regions independently. The analysis shows that the two repeat regions are not independently curved parts of parC but that the overall curvature is a property of the whole fragment. Since the centromere sequence of an E.coli plasmid as well as eukaryotic centromere sequences show DNA curvature, we speculate that curvature might be a general property of centromeres.

Base Sequence↗

Sequence periodicity of Escherichia coli is concentrated in intergenic regions.

BACKGROUND: Sequence periodicity with a period close to the DNA helical repeat is a very basic genomic property. This genomic feature was demonstrated for many prokaryotic genomes. The Escherichia coli sequences display the period close to 11 base pairs. RESULTS: Here we demonstrate that practically only ApA/TpT dinucleotides contribute to overall dinucleotide periodicity in Escherichia coli. The noncoding sequences reveal this periodicity much more prominently compared to protein-coding sequences. The sequence periodicity of ApC/GpT, ApT and GpC dinucleotides along the Escherichia coli K-12 is found to be located as well mainly within the intergenic regions. CONCLUSIONS: The observed concentration of the dinucleotide sequence periodicity in the intergenic regions of E. coli suggests that the periodicity is a typical property of prokaryotic intergenic regions. We suppose that this preferential distribution of dinucleotide periodicity serves many biological functions; first of all, the regulation of transcription.

Base Composition↗

Overlapping messages and survivability.

The phenomenon of overlapping of various sequence messages in genomes is a puzzle for evolutionary theoreticians, geneticists, and sequence researchers. The overlapping is possible due to degeneracy of the messages, in particular, degeneracy of codons. It is often observed in organisms with a limited size of genome, possessing polymerases of low fidelity. The most accepted view considers the overlapping as a mechanism to increase the amount of information per unit length. Here we present a model that suggests direct evolutionary advantage of the message overlapping. Two opposing drives are considered: (a) reduction in the amount of vulnerable points when the overlapping of two messages involves common critical points and (b) cumulative compromising cost of coexistence of messages at the same site. Over a broad range of conditions the reduction of the target size prevails, thus making the overlapping of messages advantageous.

Animals↗

Large retrotransposon derivatives: abundant, conserved but nonautonomous retroelements of barley and related genomes.

Retroviruses and LTR retrotransposons comprise two long-terminal repeats (LTRs) bounding a central domain that encodes the products needed for reverse transcription, packaging, and integration into the genome. We describe a group of retrotransposons in 13 species and four genera of the grass tribe Triticeae, including barley, with long, approximately 4.4-kb LTRs formerly called Sukkula elements. The approximately 3.5-kb central domains include reverse transcriptase priming sites and are conserved in sequence but contain no open reading frames encoding typical retrotransposon proteins. However, they specify well-conserved RNA secondary structures. These features describe a novel group of elements, called LARDs or large retrotransposon derivatives (LARDs). These appear to be members of the gypsy class of LTR retrotransposons. Although apparently nonautonomous, LARDs appear to be transcribed and can be recombinationally mapped due to the polymorphism of their insertion sites. They are dispersed throughout the genome in an estimated 1.3 x 10(3) full-length copies and 1.16 x 10(4) solo LTRs, indicating frequent recombinational loss of internal domains as demonstrated also for the BARE-1 barley retrotransposon.

3' Untranslated Regions↗

Curvature distribution in prokaryotic genomes.

DNA curvature is known to play a biological role in gene regulation, in particular, initiation of transcription. We applied the software CURVATURE based on the wedge model to predict whether promoter regions of certain prokaryotes may be characterized by higher intrinsic DNA curvature located within or upstream to these regions. The main purpose was to verify our earlier hypothesis that the DNA curvature plays a biological role in gene regulation in mesophilic as compared to hyperthermophilic prokaryotes, i.e., DNA curvature presumably has a functional adaptive significance determined by temperature selection. Therefore, we analyzed all available complete prokaryotic genomes. The analysis showed that there is a group of genomes with a relatively high average DNA curvature upstream of start of genes. Remarkably, all organisms of this group appeared to be mesophilic, which is a full confirmation of the former hypothesis. The conservative patterns of genomic curvature distribution across different mesophilic bacterial and archaeal genomes presented in this study provide a new, convincing indication that curved DNA is evolutionarily preserved and determined by temperature selection. Moreover, we found a rather peculiar property of hyperthermophilic prokaryotes: the coding regions are predicted to be significantly more curved than it would be expected from their dinucleotide composition.

DNA, Archaeal↗

Hidden messages in the nef gene of human immunodeficiency virus type 1 suggest a novel RNA secondary structure.

The coexistence of multiple codes in the genome of human immunodeficiency virus type 1 (HIV-1) was analyzed. We explored factors constraining the variability of the virus genome primarily in relation to conserved RNA secondary structures overlapping coding sequences, and used a simple combination of algorithms for RNA secondary structure prediction based on the nearest-neighbor thermodynamic rules and a statistical approach. In our previous study, we applied this combination to a non- redundant data set of env nucleotide sequences, confirmed the conservative secondary structure of the rev-responsive element (RRE) and found a new RNA structure in the first conserved (C1) region of the env gene. In this study, we analyzed the variability of putative RNA secondary structures inside the nef gene of HIV-1 by applying these algorithms to a non-redundant data set of 104 nef sequences retrieved from the Los Alamos HIV database, and predicted the existence of a novel functional RNA secondary structure in the beta3/beta4 regions of nef. The predicted RNA fold in the beta3/beta4 region of nef appears in two forms with different loop sizes. The loop of the first fold consists of seven nucleotides (positions 494-500), with consensus UCAAGCU appearing in 79% of sequences. The other has a five-base loop (positions 495-499) with consensus CAAGC. The difference in size between these two loops may reflect the difference between respective counterparts in the hairpin recognition. This may also have an adaptive biological significance.

Algorithms↗

A large-scale comparison of genomic sequences: one promising approach.

We introduce a novel, linguistic-like method of genome analysis. We propose a natural approach to characterizing genomic sequences based on occurrences of fixed length words from a predefined, sufficiently large set of words (strings over the alphabet [A, C, G, T]). A measure based on this approach is called compositional spectrum and is actually a histogram of imperfect word occurrences. Our results assert that the compositional spectrum is an overall characteristic of a long sequence i.e., a complete genome or an uninterrupted part of a chromosome. This attribute is manifested in the similarity of spectra obtained on different stretches of the same genome, and simultaneously in a broad range of dissimilarities between spectral representations of different genomes. High flexibility characterizes this approach due to imperfect matching and as a result sets of relatively long words can be considered. The proposed approach may have various applications in intra- and intergenomic sequence comparisons.

Algorithms↗

DNA sequence analysis linguistic tools: contrast vocabularies, compositional spectra and linguistic complexity.

This is a review of the methods based on counting oligomers in nucleotide and amino acid sequences. Such methods are analogous to the formal linguistic analysis of human texts. This review includes methods based on the calculation of observed occurrences (frequencies) of oligomers and their distribution, as well as those based on deviations between the observed and the expected occurrences (contrast words, genome signatures) in biological sequences. Both types of methods have a wide range of sensitivity and can identify homologous as well as functionally and taxonomically related sequences.

Algorithms↗

RNA secondary structure and squence conservation in C1 region of human immunodeficiency virus type 1 env gene.

We have analyzed amino acid, nucleotide sequence, and RNA secondary structure variability in the env gene of human immunodeficiency virus type (HIV-1). In applying algorithms for computing optimal RNA-folding patterns to a nonredundant data set of 178 env nucleotide sequences, we found a conserved RNA stem-loop structure in the first conserved (C1) region of the env gene. This detailed examination also revealed the known secondary structure conservation of the Rev-responsive element (RRE). This finding is also supported by a higher third position conservation of the translatable reading frame along these subregions. The typical folding of the C1 region consists of two isolated stem-loop structures. These highly conserved structures are likely to have a biological function. This assumption is supported by the conservation of the third position along the coding region of these structures. The third position retains a conservation level above what would be statistically expected.

Algorithms↗

Sequence complexity profiles of prokaryotic genomic sequences: a fast algorithm for calculating linguistic complexity.

MOTIVATION: One of the major features of genomic DNA sequences, distinguishing them from texts in most spoken or artificial languages, is their high repetitiveness. Variation in the repetitiveness of genomic texts reflects the presence and density of different biologically important messages. Thus, deviation from an expected number of repeats in both directions indicates a possible presence of a biological signal. Linguistic complexity corresponds to repetitiveness of a genomic text, and potential regulatory sites may be discovered through construction of typical patterns of complexity distribution. RESULTS: We developed software for fast calculation of linguistic sequence complexity of DNA sequences. Our program utilizes suffix trees to compute the number of subwords present in genomic sequences, thereby allowing calculation of linguistic complexity in time linear in genome size. The measure of linguistic complexity was applied to the complete genome of Haemophilus influenzae. Maps of complexity along the entire genome were obtained using sliding windows of 40, 100, and 2000 nucleotides. This approach provided an efficient way to detect simple sequence repeats in this genome. In addition, local profiles of complexity distribution around the starts of translation were constructed for 21 complete prokaryotic genomes. We hypothesize that complexity profiles correspond to evolutionary relationships between organisms. We found principal differences in profiles of the GC-rich and other (non-GC-rich) genomes. We also found characteristic differences in profiles of AT genomes, which probably reflect individual species variations in translational regulation. AVAILABILITY: The program is available upon request from Alexander Bolshoy or at http://csweb.haifa.ac.il/library/#complex.

Algorithms↗