PubMed Health⌕ Search

Biomedical subjects

Edward N Trifonov

Publications and source records attributed to Edward N Trifonov.

12 recordsLinked to original sources

Dinucleosome DNA of human K562 cells: experimental and computational characterizations.

Dinucleosome formation is the first step in the organization of the higher order chromatin structure. With the ultimate aim of elucidating the dinucleosome structure, we constructed a library of human dinucleosome DNA. The library consists of PCR-amplifiable DNA fragments obtained by treatment of nuclei of erythroid K562 cells with micrococcal nuclease followed by extraction of DNA and adaptor ligation to the blunt-ended DNA fragments. The library was then cloned using a plasmid vector and the sequences of the clones were determined. The dominating clones containing the Alu elements were removed. A total of 1002 clones, which comprised a dinucleosome database, contained 84 and 918 clones from the clones before and after removing Alu elements, respectively. Approximately 70% of the clones were between 300 and 400 bp in size and they were distributed to various locations of all chromosomes except the Y chromosome. The clones containing A(2)N(8)A(2)N(8)A(2) or T(2)N(8)T(2)N(8)T(2) sequences were classified into three types, Type I (N shape), Type II (V shape) and Type III (M shape) according to DNA curvature plots. The locations of experimentally determined curved DNA segments matched well with the calculated ones though the clones of Types I and III showed additional curved DNA segments as revealed by the curvature plots. The distributions of complementary dinucleotides in the nucleosome DNA, at the ends of the dinucleosome DNA clones, allowed us to predict the positions of the nucleosome dyad axis, and estimate the size of the nucleosome core DNA, 125nt. The distributions of AA and TT dinucleotides, as well as other RR and YY dinucleotides, showed a periodicity with an average period of 10.4 bases, close to the values observed before. Mapping of nucleosome positions in the dinucleosome database based on the observed periodicity revealed that the nucleosomes were separated by a linker of 7.5+ approximately 10 x n nt. This indicates that the nucleosome-nucleosome orientations are, typically, halfway between parallel and antiparallel. Also an important finding is that the distributions of AA/TT and other RR/YY dinucleotides, apparently, reflect both DNA curvature and DNA bendability, cooperatively contributing to the nucleosome formation.

Base Sequence↗

Hidden messages in the nef gene of human immunodeficiency virus type 1 suggest a novel RNA secondary structure.

The coexistence of multiple codes in the genome of human immunodeficiency virus type 1 (HIV-1) was analyzed. We explored factors constraining the variability of the virus genome primarily in relation to conserved RNA secondary structures overlapping coding sequences, and used a simple combination of algorithms for RNA secondary structure prediction based on the nearest-neighbor thermodynamic rules and a statistical approach. In our previous study, we applied this combination to a non- redundant data set of env nucleotide sequences, confirmed the conservative secondary structure of the rev-responsive element (RRE) and found a new RNA structure in the first conserved (C1) region of the env gene. In this study, we analyzed the variability of putative RNA secondary structures inside the nef gene of HIV-1 by applying these algorithms to a non-redundant data set of 104 nef sequences retrieved from the Los Alamos HIV database, and predicted the existence of a novel functional RNA secondary structure in the beta3/beta4 regions of nef. The predicted RNA fold in the beta3/beta4 region of nef appears in two forms with different loop sizes. The loop of the first fold consists of seven nucleotides (positions 494-500), with consensus UCAAGCU appearing in 79% of sequences. The other has a five-base loop (positions 495-499) with consensus CAAGC. The difference in size between these two loops may reflect the difference between respective counterparts in the hairpin recognition. This may also have an adaptive biological significance.

Algorithms↗

Evolutionary aspects of protein structure and folding.

The traditional reconstruction of molecular events of the past based on sequence conservation becomes very vague beyond one to two billion years ago. There are certain molecular features, however, such as polymer flexibility and loop closure, that are conserved merely because of their physical nature. This allows one to penetrate the earliest stages of protein evolution.

Amino Acid Sequence↗

Protein sequences yield a proteomic code.

Analysis of crystallized protein structures suggests that globular proteins are organized as consecutively connected units of 25-35 residues. These units are closed loops, that is returns of the polypeptide chain trajectory to a close contact with itself. This universal feature of apparently polymer-statistical nature is a basis for a principally novel view on the globular proteins as loop fold structures. The same unit size has been detected in protein sequences translated from complete prokaryotic genomes by positional autocorrelation analysis, which strongly indicates the evolutionary connection of the units. The units are further characterized by prototype sequences matching to their numerous derivatives in the translated genomes. The matches to five strongest prokaryotic prototypes and three prototypes of C. elegans are identified in the sequences of crystallized proteins, and their structures analyzed. Corresponding segments of the polypeptide chains in majority of cases form closed loops, though evolutionary fate of every prototype element is shown to be rather diverse. Then loop ends can be separated by a sequence-wise distant segments and stabilized by the spatial interactions in the context of the overall globular structure. The units belong to a presumably limited spectrum of the sequence prototypes, full repertoire of which would constitute a proteomic code.

Amino Acid Motifs↗

Spelling protein structure.

Recent sequence analysis of complete prokaryotic proteomes suggests that in early evolutionary stages proteins were rather small, of the size 25-35 amino acids. Corroborating evidence comes from protein crystal data, which indicate this size for closed loops--universal structural units of globular proteins. In the latest development we were able to derive and structurally characterize several sequence/structure prototypes apparently representing early protein units. Structurally the prototypes appear as closed loops stabilized by end-to-end van der Waals interactions. While nearly standard in size the loops are highly diverse in terms of their secondary structure. A presentation of the protein as an assembly of descendants of the prototypes, the first of its kind, is described in detail here. The sequence and structure of the ATP-binding subunit of histidine permease of S. typhimurium is shown to contain several modified copies of different prototype elements, closed loops, and, thus, can be spelled as: x-PI-x-PIV-PVI-PII-PVII-x, where PI-PVII are the prototype elements. This study sets up the basic principles for the sequence/structure prototype spelling of globular proteins.

ATP-Binding Cassette Transporters↗

RNA secondary structure and squence conservation in C1 region of human immunodeficiency virus type 1 env gene.

We have analyzed amino acid, nucleotide sequence, and RNA secondary structure variability in the env gene of human immunodeficiency virus type (HIV-1). In applying algorithms for computing optimal RNA-folding patterns to a nonredundant data set of 178 env nucleotide sequences, we found a conserved RNA stem-loop structure in the first conserved (C1) region of the env gene. This detailed examination also revealed the known secondary structure conservation of the Rev-responsive element (RRE). This finding is also supported by a higher third position conservation of the translatable reading frame along these subregions. The typical folding of the C1 region consists of two isolated stem-loop structures. These highly conserved structures are likely to have a biological function. This assumption is supported by the conservation of the third position along the coding region of these structures. The third position retains a conservation level above what would be statistically expected.

Algorithms↗

What positions nucleosomes?--A model.

Here we propose a new determinant for localization of nucleosomes along genomic DNA, in addition to sequence-dependent features. The new specific class of chromatin scaling signals involves curved DNA. According to the observed positional distribution of DNA curvature, the new synchronizing signal occurs once per four nucleosomes on average. This new factor in nucleosome positioning should substantially influence the efficiency of biological reactions through regulatory factors microscopically and the entire chromatin structure through the 30 nm fiber structure macroscopically. Allocation of the new type of signals is found to be fixed evolutionarily although they could be shifted in accordance with the hierarchy of functional genomic structures.

Chromatin↗

Loop fold structure of proteins: resolution of Levinthas paradox.

According to Levinthal a protein chain of ordinary size would require enormous time to sort its conformational states before the final fold is reached. Experimentally observed time of folding suggests an estimate of the chain length for which the time would be sufficient. This estimate by order of magnitude fits to experimentally observed universal closed loop elements of globular proteins - 25-30 residues.

Models, Chemical↗

Back to units of protein folding.

In response to the criticism by A. Finkelstein (J Biomol Struct Dyn 20, 311-314, 2002) of our Communication (J Biomol Struct Dyn 20, 5-6, 2002) several issues are dealt with. Importance of the notion of elementary folding unit, its size and structure, and the necessity of further characterization of the units for the elucidation of the protein folding in vivo are discussed. The criticism (J Biomol Struct Dyn 20, 311-314, 2002) on the hierarchical protein folding is also briefly addressed.

Kinetics↗

Distribution of rare triplets along mRNA and their relation to protein folding.

It is believed that pausing during mRNA translation plays some role in ensuring proper folding of newly synthesized sections of a protein chain. Such pausing occurs when rare triplets are encountered in the mRNA, as it takes additional time for the corresponding rare species of tRNA to be delivered. To determine whether pause sites are non-randomly distributed along prokaryotic mRNA (cDNA), we have located clusters of rare triplets in cDNA sequences from 21 different bacteria. From the individual profiles of local codon frequencies calculated with various windows, the positions of the clusters of the rarest codons were taken for generation of the combined histograms of positional preferences of the pause sites. The histograms show that in the prokaryotic sequences, the pause sites are located preferentially at the start positions and at about 155 triplets from the starts. To verify the generality of these observations, the data are grouped in six independent sets about 500 sequences each, all revealing the same features. A less prominent maximum is also seen at the triplet position 75. Judging by the amplitude of the peak at 155 triplets, an optimal cluster size is estimated to equal 18 triplets. The distance 155 closely corresponds to the sizes of typical protein folds and to earlier estimated prokaryotic protein sequence segments. This supports the suggestion of a role for translation pausing in the cotranslational folding of protein domains. The profiles of rare codons in mRNA can serve in the detection or prediction of boundaries between protein domains.

Amino Acid Sequence↗

Spectral analysis of distributions: finding periodic components in eukaryotic enzyme length data.

We introduce the spectral analysis of distributions (SAD), a method for detecting and evaluating possible periodicity in experimental data distributions (histograms) of arbitrary shape. SAD determines whether a given empirical distribution contains a periodic component. We also propose a system of probabilistic mixture distributions to model a histogram consisting of a smooth background together with peaks at periodic intervals, with each peak corresponding to a fixed number of subunits added together. This mixture distribution model allows us to estimate the parameters of the data and to test the statistical significance of the estimated peaks. The analysis is applied to the length distribution of eukaryotic enzymes.

Algorithms↗

Closed loops: persistence of the protein chain returns.

It has recently been discovered that globular proteins are universally built from standard loop-n-lock units of about 30 amino acid residues. The hypothesis has been put forward on the loop stage in the protein evolution when the units were autonomous. Later they joined together making longer chains. One would expect that the early individual loop-n-lock elements might still be detected in modern protein sequences as remnants of the hypothetical 30-residue sequence prototypes. Among several strong sequence motifs, extracted from protein sequences of 23 complete bacterial proteomes, one 32-residue prototype was studied here in detail. Numerous sequence segments related to the prototype are identified in the crystal structures of proteins of a PDB_SELECT database. Analysis of the respective chain trajectories for the cases with different degrees of sequence conservation confirms that the majority of the segments correspond to the closed loops. In the evolutionary diversification of the prototypes the secondary structure yields first, while the sequence is still moderately conserved. The last feature to go is the chain return property. Apparently, the opening of the loops would severely destabilize the protein fold, which explains their conservation.

Amino Acid Motifs↗