PubMed Health⌕ Search

Biomedical subjects

M D Ermolaeva

Publications and source records attributed to M D Ermolaeva.

8 recordsLinked to original sources

Prediction of operons in microbial genomes.

Operon structure is an important organization feature of bacterial genomes. Many sets of genes occur in the same order on multiple genomes; these conserved gene groupings represent candidate operons. This study describes a computational method to estimate the likelihood that such conserved gene sets form operons. The method was used to analyze 34 bacterial and archaeal genomes, and yielded more than 7600 pairs of genes that are highly likely (P: >/= 0.98) to belong to the same operon. The sensitivity of our method is 30-50% for the Escherichia coli genome. The predicted gene pairs are available from our World Wide Web site http://www.tigr.org/tigr-scripts/operons/operons.cgi.

Algorithms↗

A probabilistic method for identifying start codons in bacterial genomes.

As the pace of genome sequencing has accelerated, the need for highly accurate gene prediction systems has grown. Computational systems for identifying genes in prokaryotic genomes have sensitivities of 98-99% or higher (Delcher et al., Nucleic Acids Res., 27, 4636-4641, 1999). These accuracy figures are calculated by comparing the locations of verified stop codons to the predictions. Determining the accuracy of start codon prediction is more problematic, however, due to the relatively small number of start sites that have been confirmed by independent, non-computational methods. Nonetheless, the accuracy of gene finders at predicting the exact gene boundaries at both the 5' and 3' ends of genes is of critical importance for microbial genome annotation, especially in light of the important signaling information that is sometimes found on the 5' end of a protein coding region. In this paper we propose a probabilistic method to improve the accuracy of gene identification systems at finding precise translation start sites. The new system, RBSfinder, is tested on a validated set of genes from Escherichia coli, for which it improves the accuracy of start site locations predicted by computational gene finding systems from the range 67-77% to 90% correct.

Algorithms↗

Synonymous codon usage in bacteria.

In most bacteria, synonymous codons are not used with equal frequencies. Different factors have been proposed to contribute to codon usage preference, including translational selection, GC composition, strand-specific mutational bias, amino acid conservation, protein hydropathy, transcriptional selection and even RNA stability. The review discusses these factors and their contribution to bias in synonymous codon usage in bacterial genomes.

Bacteria↗

Prediction of transcription terminators in bacterial genomes.

This study describes an algorithm that finds rho-independent transcription terminators in bacterial genomes and evaluates the accuracy of its predictions. The algorithm identifies terminators by searching for a common mRNA motif: a hairpin structure followed by a short uracil-rich region. For each terminator, an energy-scoring function that reflects hairpin stability, and a tail-scoring function based on the number of U nucleotides and their proximity to the stem, are computed. A confidence value can be assigned to each terminator by analyzing candidate terminators found both within and between genes, and taking into account the energy and tail scores. The confidence is an empirical estimate of the probability that the sequence is a true terminator. The algorithm was used to conduct a comprehensive analysis of 12 bacterial genomes to identify likely candidates for rho-independent transcription terminators. Four of these genomes (Deinococcus radiodurans, Escherichia coli, Haemophilus influenzae and Vibrio cholerae) were found to have large numbers of rho-independent terminators. Among the other genomes, most appear to have no transcription terminators of this type, with the exception of Thermotoga maritima. A set of 131 experimentally determined E. coli terminators was used to evaluate the sensitivity of the method, which ranges from 89 % to 98 %, with corresponding false positive rates of 2 % and 18 %.

Algorithms↗

DNA sequence of both chromosomes of the cholera pathogen Vibrio cholerae.

Here we determine the complete genomic sequence of the gram negative, gamma-Proteobacterium Vibrio cholerae El Tor N16961 to be 4,033,460 base pairs (bp). The genome consists of two circular chromosomes of 2,961,146 bp and 1,072,314 bp that together encode 3,885 open reading frames. The vast majority of recognizable genes for essential cell functions (such as DNA replication, transcription, translation and cell-wall biosynthesis) and pathogenicity (for example, toxins, surface antigens and adhesins) are located on the large chromosome. In contrast, the small chromosome contains a larger fraction (59%) of hypothetical genes compared with the large chromosome (42%), and also contains many more genes that appear to have origins other than the gamma-Proteobacteria. The small chromosome also carries a gene capture system (the integron island) and host 'addiction' genes that are typically found on plasmids; thus, the small chromosome may have originally been a megaplasmid that was captured by an ancestral Vibrio species. The V. cholerae genomic sequence provides a starting point for understanding how a free-living, environmental organism emerged to become a significant human bacterial pathogen.

Base Sequence↗

[Molecular dynamics of oligopeptides. 3. Maps of levels of free energy of modified dipeptides and dynamic correlation in amino acid residues].

A method of free energy maps for studying the dynamic correlations of fluctuations in molecules with conformational mobility is proposed. An agreement between the structure of the free energy level map and the type of the corresponding cross-correlation function (in the presence and absence of the correlation of fluctuations in conformational freedom degree in modified dipeptides) was established.

Amino Acids↗

[Molecular dynamics of oligopeptides. 1. The use of long trajectory and high temperature data for determination of statistical weight of conformational substates].

The method of molecular dynamics investigations of the particularities of macromolecule physical-chemical properties is discussed. The results, obtained from the calculations of modified dipeptides in different regimes (different temperatures, length of trajectories and ways of thermostat simulation) are compared. The optimal conditions for this peptides calculations are determined: collisional dynamics regime, trajectories not less than 5000 ps, temperature about 1000 K. In this case the figurative point is able to scan the molecule configuration space and the statistically reliable results could be obtained.

Dipeptides↗

[Molecular dynamics of oligopeptides. 2. Dynamic correlation functions of the conformational modes of the modified dipeptides].

The method of investigation of dynamic correlation functions in molecules with conformational mobility on an example of modified dypeptides by molecular dynamics is proposed. Comparison of results for different types torsion angles and internuclear distances are carried out. Connection between the chemical structure of peptides and the peculiarities of internal dynamic correlations is established. A question about dynamic isomorphism of amino acids is discussed.

Dipeptides↗