PubMed HealthSearch

Biomedical subjects

E N Trifonov

Publications and source records attributed to E N Trifonov.

At least 19 recordsLinked to original sources

Periodic recurrence of methionines: fossil of gene fusion?

As we have recently shown, approximately 20% of proteins are made of uniform size units of approximately 123 aa for eukaryotes and approximately 152 aa for prokaryotes. Such regularity may reflect certain past events in protein evolution by fusion (molecular recombination) of a spectrum of standard-size protein-coding DNA segments--the early genes. Consequently, methionines, as start residues, would mark those locations in proteins that correspond to the DNA recombination sites--the borders between the fused genes. This positional preference of the methionines may still survive as a fossil of the early protein sequence organization. In this study we address the question how methionines are distributed in modern protein sequences. This analysis of eukaryotic sequences shows that methionine residues do preferentially appear at the positions corresponding to the multiples of the unit size, as predicted.

Biological Evolution

Segmented structure of protein sequences and early evolution of genome by combinatorial fusion of DNA elements.

A theory of an early stage of genome evolution by combinatorial fusion of circular DNA units is suggested, based on protein sequence "fossil" evidence. The evidence includes preference of protein sequence lengths for certain sizes--multiples of 123 aa for eukaryotes and multiples of 152 aa for prokaryotes. At the DNA level these sizes correspond to 350-450 base pairs--the known optimal range for DNA ring closure. The methionine residues repeatedly appear along the sequences with the same period of about 120 aa (in eukaryotes), presumably marking the sites of insertion of the early genes--rings of protein-coding DNA. No torsional constraint in this DNA results in very sharp estimate of the helical periodicity of the early DNA, indistinguishable from the experimental mean value for extant DNA. According to the combinatorial fusion theory, based on the above evidence, in the pregenomic, prerecombinational stage the genes and the noncoding sequences existed in form of autonomously replicating DNA rings of close to standard size, randomly segregating between dividing cells, like modern plasmids do. In the recombinational early genomic stage the rings started to fuse, forming larger DNA molecules consisting of several unit genes connected in various combinations and forming long protein-coding sequences (combinatorial fusion). This process, which involved, perhaps, noncoding sequences as well, eventually resulted in the formation of large genomes. The dispersed circular DNA--or, rather, evolutionarily advanced derivatives thereof--may still exist in the form of various mobile DNA elements.

Biological Evolution

Underlying order in protein sequence organization.

The idea of a possible standard modular structure of proteins has been known since 1929 when it was introduced by Svedberg. It still remains an idea with no quantitative confirmation of universality of such hypothetical organization. From a large collection of nonredundant protein sequences representing > 100 eukaryotic and prokaryotic species, we have obtained the protein sequence length distributions. Mere inspection of these distributions, as well as spectral analysis, shows that 15-30% of proteins, depending on species and sequence types, indeed appear to be made of sequence units with characteristic lengths of approximately 125 aa for eukaryotes and approximately 150 aa for prokaryotes. This underlying order in protein sequence organization is shown to be universal--that is, the weak regularity observed is not caused by a particular dominant species or protein group. Possible mechanisms are discussed that may be responsible for the observed regularity, including a hypothesis about the recombinational nature of such protein sequence organization.

Amino Acid Sequence

On the recombinational origin of protein-sequence-subunit structure.

Since 1929 the concept that proteins are built from subunits of certain standard size (Svedberg 1929) has been revisited several times, each time with a new demonstration that, indeed, there are certain preferred protein sizes. According to recent estimates the overrepresented sizes are close to multiples of 125 amino acid (aa) residues for eukaryotes and 150 residues for prokaryotes. To explain these preferences, a hypothesis is suggested, and quantitatively developed, on the recombinational nature of this regularity. The protein-coding sequences are assumed to evolve at some early stage via recombinational events--insertions of DNA circles of a certain optimal size. The contour lengths of the protein-coding DNA circles had to be simultaneously divisible by three and, to minimize torsional constraint, by the DNA helical repeat. With these two conditions satisfied, the calculated contour lengths of the DNA circles, 250-500 base pairs (bp), turn out to correspond well to known optimal DNA circularization sizes and to the predicted range of the protein sequence subunit sizes: 80-170 aa residues, which covers experimentally observed values. The subunit size is found to be strongly influenced by the helical repeat of DNA. The sizes 125 and 150 aa are derived when the corresponding helical repeats of DNA are set within fractions of promilles from the 10.54 bp/turn value. This fits to the experimentally estimated mean for natural mixed DNA sequences, 10.53-10.57 bp/turn.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence

CURVATURE: software for the analysis of curved DNA.

Software is presented to plot the sequence-dependent spatial trajectory of the DNA double helix and/or distribution of curvature along the DNA molecule. The nearest-neighbor wedge model is implemented to calculate overall DNA path using local helix parameters: helix twist angle, wedge (deflection) angle and direction (of deflection) angle. The procedures described proved to be very convenient as tools for investigation of a relationship between overall DNA curvature and its gel electrophoretic mobility. All parameters of the model had been estimated from experimental data. Using these wedge parameters the program takes, as input, any DNA sequence and calculates the likely degree of curvature at each point along the molecule. This information is displayed both graphically and in the form of simplified representations of curved double helices. The Software, CURVATURE, can thus be used to investigate possible roles of curvature in modulation of gene expression and for location of curved portions of DNA, which may play an important role in sequence-specific protein--DNA interactions.

Algorithms

Imported sequences in the mitochondrial yeast genome identified by nucleotide linguistics.

In addition to universally appearing mitochondrial (mt) genes, origins of replication and transcription start regions typical of all mt genome variants of the yeast Saccharomyces cerevisiae, the mt genomes of some of the strains contain variable sequences. These sequences are apparently largely dispensable. They are mainly composed of group-I and -II introns and intergenic open reading frames (ORFs). Many of the introns contain ORFs, some of which were shown by genetic and biochemical means to be involved in splicing and transposition of the mt introns. Some of the optional sequences are hypothesized to be mobile genetic elements. Nucleotide (nt) sequences of the mt genome of S. cerevisiae were examined by analyzing occurrences of oligodeoxyribonucleotide (oligo) 'words'. This linguistic technique had been found to be sensitive to both function and origin of the sequence [Pietrokovski et al., J. Biomol. Struct. Dyn. 7 (1990) 1251-1268]. A clear difference is found between the oligo vocabularies of the optional and basic yeast mt sequences. The difference is mainly located in protein coding segments of the optional sequences which contain conserved amino acid motifs, characteristic of intronic and intergenic ORFs. The use of nt linguistics to detect the sequence dissimilarity and its causes in yeast mitochondria provides fast and straightforward results, identifying the intronic and intergenic ORFs as DNA sequences of foreign, non-mt origin.

Amino Acid Sequence

Recognition of correct reading frame by the ribosome.

The translation frame-monitoring mechanism has been suggested earlier, based on transient complementary contacts, between mRNA and rRNA. Recent studies related to the frame-monitoring mechanism are reviewed. The mechanism is well supported by both new experimental and sequence analysis data. Experiments are suggested for further elucidation of the structural details of the mRNA-rRNA interaction in the ribosome.

Base Sequence

Preferred positions of AA and TT dinucleotides in aligned nucleosomal DNA sequences.

Multiple alignment of 118 nucleosomal DNA sequences by maximizing simultaneously match of AA dinucleotides and match of TT dinucleotides results in a pattern of the dinucleotide distributions which is characteristic of the nucleosomal DNA sequences. The AA dinucleotides are found to be distributed symmetrically relative to the TT dinucleotide distribution, around the middle point of the nucleosomal DNA sequence. The distances between major peaks of the distributions are multiples of about 10.4 bases. The peaks of the TT distribution are shifted by 6 bases downstream from the peaks of the AA distribution.

Adenine

mRNA periodical infrastructure complementary to the proof-reading site in the ribosome.

Virtually all mRNA sequences carry a 3-base periodical pattern, presumably involved in the translation frame monitoring mechanism (Trifonov, E.N., J. Mol. Biol. 194, 643-652, 87). The hidden pattern, 5'-(GHN)n-3' (H representing nonG, N any base), is further refined by extensive computational analysis of mRNA sequences. According to mononucleotide preferences in the three positions of coding triplets, it appears now as 5'-(GHU)n-3'. Dinucleotide frequencies independent of mononucleotides (contrast dinucleotides, 2) generate the motif 5'-(GCU)n-3'. The same motif is found by regarding the expected avoidance of destabilizing base oppositions in hypothetical transient complementary complexes between mRNA and rRNA. This hidden pattern, in its refined consensus form, 5'-(GCU)n-3', is an almost perfect complementary match to a unique site in small subunit rRNA, the universally conserved (3) proofreading loop at position 525 (of E.coli small subunit rRNA): [formula: see text] This strongly suggests that the 525 site is a major structural component of the previously proposed frame-keeping mechanism which is based on the in-frame contacts between mRNA and three segments of rRNA. Consistent with the original proposition, this site is one of three believed to interact with mRNA.

Base Sequence

Curved DNA without A-A: experimental estimation of all 16 DNA wedge angles.

The principal sequence feature responsible for intrinsic DNA curvature is generally assumed to be runs of adenines. However, according to the wedge model of DNA curvature, each dinucleotide step is associated with a characteristic deflection of the local helix axis. Thus, an important test of a more general view of sequence-dependent DNA curvature is whether sequence elements other than A-A cause the DNA axis to deflect. To address this question, we have applied the wedge model to a large body of experimental data. The axial path of DNA can be described at each step by three Eulerian angles: the helical twist, the deflection angle (wedge angle), and the direction of the deflection. Circularization and gel electrophoretic mobility data on 54 synthetic DNA fragments, both from other laboratories and from our own, were used to compare the theoretical predictions of the wedge model with experiment. By minimizing misfit between calculated and observed DNA curvature, we have found that the stacks AG/CT, CG/CG, GA/TC, and GC/GC, in addition to AA/TT, have large wedge values. We have also synthesized seven sequences without AA/TT elements but with these other wedges correctly phased to cause appreciable predicted curvature. All appear curved as demonstrated by anomalous gel mobilities. The full set of 16 roll and tilt wedge angles is estimated and, together with the known 10 helical twists, these allow prediction of the general sequence-dependent trajectory of the DNA axis.

Adenine

Splice junctions follow a 205-base ladder.

Of all aspects of mRNA maturation the accuracy of intervening sequence excision and exon ligation is, perhaps, the most enigmatic. Attempts to identify the essential elements involved in this process have thus far not yielded any satisfactory answer as to what structural (sequence) features are prerequisite for the vital precision of this process. In our search for underlying structural orders we asked whether exons and introns had any positional preferences within a gene. This analysis led to the unexpected discovery that the DNA length is synchronized between successive 3' splicing sites as well as between successive 5' splicing sites, with a frame of approximately 205 base pairs. This observation reveals additional organization of genes in eukaryotes and, perhaps, links gene splicing with chromatin structure.

Animals

DNA in profile.

The double helix structure of DNA is not necessarily straight but rather can be curved at almost every base pair. Thus, each piece of DNA possesses a unique silhouette, as individual as its base sequence.

Base Sequence

Sequence-dependent kinks induced in curved DNA.

In certain curved DNA fragments without AA dinucleotides, the gel retardation anomaly associated with curvature passes through a maximum with fragment length, indicating length (and electric field) dependent structural transitions in the DNA. We suggest that thermally induced stereochemical kinks in DNA are stabilized in the gel, thus relieving the effects of curvature. These kinks are shown to occur specifically at CA/TG and TA/TA stacks. Other physical and biological evidence points to frequent structural dislocations at CA and TA steps. These reversible sequence dependent kinks may therefore represent a novel class of structural protein-DNA recognition elements.

Base Sequence

Linguistic measure of taxonomic and functional relatedness of nucleotide sequences.

The frequencies of "words", oligonucleotides within nucleotide sequences, reflect the genetic information contained in the sequence "texts". Nucleotide sequences are characteristically represented by their contrast word vocabularies. Comparison of the sequences by correlating their contrast vocabularies is shown to reflect well the relatedness (unrelatedness) between the sequences. A single value, the linguistic similarity between the sequences, is suggested as a measure of sequence relatedness. Sequences as short as 1000 bases can be characterized and quantitatively related to other sequences by this technique. The linguistic sequence similarity value is used for analysis of taxonomically and functionally diverse nucleotide sequences. The similarity value is shown to be very sensitive to the relatedness of the source species, thus providing a convenient tool for taxonomic classification of species by their sequence vocabularies. Functionally diverse sequences appear distinct by their linguistic similarity values. This can be a basis for a quick screening technique for functional characterization of the sequences and for mapping functionally distinct regions in long sequences.

Animals

The multiple codes of nucleotide sequences.

Nucleotide sequences carry genetic information of many different kinds, not just instructions for protein synthesis (triplet code). Several codes of nucleotide sequences are discussed including: (1) the translation framing code, responsible for correct triplet counting by the ribosome during protein synthesis; (2) the chromatin code, which provides instructions on appropriate placement of nucleosomes along the DNA molecules and their spatial arrangement; (3) a putative loop code for single-stranded RNA-protein interactions. The codes are degenerate and corresponding messages are not only interspersed but actually overlap, so that some nucleotides belong to several messages simultaneously. Tandemly repeated sequences frequently considered as functionless "junk" are found to be grouped into certain classes of repeat unit lengths. This indicates some functional involvement of these sequences. A hypothesis is formulated according to which the tandem repeats are given the role of weak enhancer-silencers that modulate, in a copy number-dependent way, the expression of proximal genes. Fast amplification and elimination of the repeats provides an attractive mechanism of species adaptation to a rapidly changing environment.

Amino Acid Sequence

CCAAT box revisited: bidirectionality, location and context.

The so-called CCAAT box is believed to be a major promoter element of higher eukaryotes though it is ill-defined being deduced from very limited sequence data. The comprehensive computer analysis of an unbiased set of 168 promoters presented here removes several of the persisting uncertainties. In particular, it delineates the region of preferential occurrence of the CCAAT element to -110 to -50 relative to the initiation site, suggests that integrity of this pentamer is essential, and confirms bidirectionality as a general property of this element. Within the above region the signal is found to occur in a specific sequence context which is an important supplement to its description.

Animal Population Groups