PubMed Health⌕ Search

Biomedical subjects

Stanislaw Cebrat

Publications and source records attributed to Stanislaw Cebrat.

4 recordsLinked to original sources

Higher mutation rate helps to rescue genes from the elimination by selection.

Directional mutation pressure associated with replication processes is the main cause of the asymmetry between the leading and lagging DNA strands in bacterial genomes. On the other hand, the asymmetry between sense and antisense strands of protein coding sequences is a result of both mutation and selection pressures. Thus, there are two different ways of superposition of the sense strand, on the leading or lagging strand. Besides many other implications of these two possible situations, one seems to be very important - because of the asymmetric replication-associated mutation pressure, the mutation rate of genes depends on their location. Using Monte Carlo methods, we have simulated, under experimentally determined directional mutation pressure, the divergence rate and the elimination rate of genes depending on their location in respect to the leading/lagging DNA strands in the asymmetric prokaryotic genome. We have found that the best survival strategy for the majority of genes is to sometimes switch between DNA strands. Paradoxically, this strategy results in higher substitution rates but remains in agreement with observations in bacterial genomes that such inversions are very frequent and divergence rate between homologs lying on different DNA strands is very high.

Amino Acid Substitution↗

Where does bacterial replication start? Rules for predicting the oriC region.

Three methods, based on DNA asymmetry, the distribution of DnaA boxes and dnaA gene location, were applied to identify the putative replication origins in 120 chromosomes. The chromosomes were classified according to the agreement of these methods and the applicability of these methods was evaluated. DNA asymmetry is the most universal method of putative oriC identification in bacterial chromosomes, but it should be applied together with other methods to achieve better prediction. The three methods identify the same region as a putative origin in all Bacilli and Clostridia, many Actinobacteria and gamma Proteobacteria. The organization of clusters of DnaA boxes was analysed in detail. For 76 chromosomes, a DNA fragment containing multiple DnaA boxes was identified as a putative origin region. Most bacterial chromosomes exhibit an overrepresentation of DnaA boxes; many of them contain at least two clusters of DnaA boxes in the vicinity of the oriC region. The additional clusters of DnaA boxes are probably involved in controlling replication initiation. Surprisingly, the characteristic features of the initiation of replication, i.e. a cluster of DnaA boxes, a dnaA gene and a switch in asymmetry, were not found in some of the analysed chromosomes, particularly those of obligatory intracellular parasites or endosymbionts. This is presumably connected with many mechanisms disturbing DNA asymmetry, translocation or disappearance of the dnaA gene and decay of the Escherichia coli perfect DnaA box pattern.

Bacteria↗

Representation of mutation pressure and selection pressure by PAM matrices.

This paper analyses the relationship between the mutation data matrix 1PAM/PET91, representing the effect of both mutation and selection pressures exerted on 16130 homologous proteins of different organisms, and a mutation probability matrix (1PAM/MPM) representing the effect of pure mutation pressure on protein coding of the Borrelia burgdorferi genome. The 1PAM/PMP matrix was derived with the help of computer simulations, which used empirical nucleotide substitution rates found for the B. burgdorferi genome. Here, it is shown that the frequency of amino acid occurrence is strongly related to their effective survival time. We found that the shorter the turnover time of an amino acid under pure mutation pressure, the lower its fraction in the proteins coded by the genome and the more protected by selection pressure is its position in proteins. Results of analyses suggest that during evolution the mutational pressure has been optimised to some extent to the selection requirements.

Algorithms↗

How many protein-coding genes are there in the Saccharomyces cerevisiae genome?

We have compared the results of estimations of the total number of protein-coding genes in the Saccharomyces cerevisiae genome, which have been obtained by many laboratories since the yeast genome sequence was published in 1996. We propose that there are 5300-5400 genes in the genome. This makes the first estimation of the number of intronless ORFs longer than 100 codons, based on the features of the set of genes with phenotypes known in 1997 to be correct. This estimation assumed that the set of the first 2300 genes with known phenotypes was representative for the whole set of protein-coding genes in the genome. The same method used in this paper for the approximation of the total number of protein-coding sequences among more than 40 000 ORFs longer than 20 codons gives a result that is only slightly higher. This suggests that there are still some non-coding ORFs in the databases and a few dozen small ORFs, not yet annotated, which probably code for proteins.

Databases, Factual↗