PubMed HealthSearch

Biomedical subjects

G Pesole

Publications and source records attributed to G Pesole.

12 recordsLinked to original sources

WORDUP: an efficient algorithm for discovering statistically significant patterns in DNA sequences.

We present here a fast and sensitive method designed to isolate short nucleotide sequences which have non-random statistical properties and may thus be biologically active. It is based on a first order Markov analysis and allows us to detect statistically significant sequence motifs from six to ten nucleotides long which are significantly shared (or avoided) in the sequences under investigation. This method has been tested on a set of 521 sequences extracted from the Eukaryotic Promoter Database (2). Our results demonstrate the accuracy and the efficiency of the method in that the sequence motifs which are known to act as eukaryotic promoters, such as the TATA-box and the CAAT-box, were clearly identified. In addition we have found other statistically significant motifs, the biological roles of which are yet to be clarified.

Algorithms

A statistical method for detecting regions with different evolutionary dynamics in multialigned sequences.

We describe a stochastic method for tracing the evolutionary pattern of multialigned sequences. This method allows us to detect gene regions with distinct evolutionary dynamics, e.g., regions that significantly deviate from the expected behavior. Accurate detection of hypervariable or hyperconstrained regions may provide useful information on the structure/function relationship of biosequences. This information can help localize functional constraints. In addition, the selection of distinct evolutionary dynamics may assist in the correct use of biosequences as reliable molecular clocks.

Animals

The evolution of the mitochondrial D-loop region and the origin of modern man.

The origin of modern man is a highly debated issue that has recently been tackled by using mitochondrial DNA sequences. The limited genetic variability of human mtDNA has been explained in terms of a recent common genetic ancestry, thus implying that all modern-population mtDNAs originated from a single woman who lived in Africa less than 0.2 Mya. This divergence time is based on both the estimation of the rate of mtDNA change and its calibration date. Because different estimates of the rate of mtDNA evolution can completely change the scenario of the origin of modern man, we have reanalyzed the available mitochondrial sequence data by using an improved version of the statistical model, the "Markov clock," devised in our laboratory. Our analysis supports the African origin of modern man, but we found that the ancestral female from which all extant human mtDNAs originated lived in a time span of 0.3-0.8 Mya. Pushing back the date of the deepest root of the human implies that the earliest divergence would have been in the Homo erectus population.

Animals

Glutamine synthetase gene evolution: a good molecular clock.

Glutamine synthetase (EC 6.3.1.2) gene evolution in various animals, plants, and bacteria was evaluated by a general stationary Markov model. The evolutionary process proved to be unexpectedly regular even for a time span as long as that between the divergence of prokaryotes from eukaryotes. This enabled us to draw phylogenetic trees for species whose phylogeny cannot be easily reconstructed from the fossil record. Our calculation of the times of divergence of the various organelle-specific enzymes led us to hypothesize that the pea and bean chloroplast genes for these enzymes originated from the duplication of nuclear genes as a result of the different metabolic needs of the various species. Our data indicate that the duplication of plastid glutamine synthetase genes occurred long after the endosymbiotic events that produced the organelles themselves.

Amino Acid Sequence

Evolutionary analysis of the nucleus-encoded subunits of mammalian cytochrome c oxidase.

The cytochrome c oxidase enzyme complex of eukaryotes is made up of three mitochondrial-coded subunits and a variable number of nuclear-coded subunits. Some nuclear-coded subunits are present in multiple forms and probably perform a tissue- or development-specific function. A detailed evolutionary analysis of the cytochrome c oxidase subunits that have been sequenced to date is reported here. We have found that gene duplication events from which the liver and heart isoforms of rat subunits VIa and subunit VIII originated can both be dated at about 240 +/- 90 million years ago, long before the radiation of mammalian lineages. Sequence divergence between the processed-type pseudogenes for the subunits IV, VIc and VIII have been estimated. Our results indicate that they arose fairly recently, thus suggesting that retroposition is a continuing process. We show that the rate of silent substitution in mitochondrial-coded subunits is 5-10 times higher than in nuclear-coded subunits; on the other hand replacement rates, although differing from gene to gene, are roughly of the same order of magnitude in both nuclear and mitochondrial genes. In the case of most of the nuclear-coded proteins we observed a slightly greater similarity between rats and cow, which agrees with the data obtained for mitochondrial-coded subunits.

Amino Acid Sequence

The main regulatory region of mammalian mitochondrial DNA: structure-function model and evolutionary pattern.

The evolution of the main regulatory region (D-loop) of the mammalian mitochondrial genome was analyzed by comparing the sequences of eight mammalian species: human, common chimpanzee, pygmy chimpanzee, dolphin, cow, rat, mouse, and rabbit. The best alignment of the sequences was obtained by optimization of the sequence similarities common to all these species. The two peripheral left and right D-loop domains, which contain the main regulatory elements so far discovered, evolved rapidly in a species-specific manner generating heterogeneity in both length and base composition. They are prone to the insertion and deletion of elements and to the generation of short repeats by replication slippage. However, the preservation of some sequence blocks and similar cloverleaf-like structures in these regions, indicates a basic similarity in the regulatory mechanisms of the mitochondrial genome in all mammalian species. We found, particularly in the right domain, significant similarities to the telomeric sequences of the mitochondrial (mt) and nuclear DNA of Tetrahymena thermophila. These sequences may be interpreted as relics of telomeres present in ancestral linear forms of mtDNA or may simply represent efficient templates of RNA primase-like enzymes. Due to their peculiar evolution, the two peripheral domains cannot be used to estimate in a quantitative way the genetic distances between mammalian species. On the other hand the central domain, highly conserved during evolution, behaves as a good molecular clock. Reliable estimates of the times of divergence between closely and distantly related species were obtained from the central domain using a Markov model and assuming nonhomogeneous evolution of nucleotide sites.

Animals

The branching order of mammals: phylogenetic trees inferred from nuclear and mitochondrial molecular data.

In order to clarify some controversial phylogenies such as those regarding the triplet of human, rodent, and cow and the evolutionary position of Lagomorpha with respect to other mammals, we have analyzed both nuclear and mitochondrial genes using the stationary Markov model developed in our laboratory. We found that the two sets of genes give different results. In particular the mitochondrial tree showed rabbit linked first to rodents and the rabbit-rodents branch linked to artiodactyls with human as the outgroup. The most favorite nuclear tree showed human linked first to artiodactyls and the human-artiodactyls branch linked to rabbit with rodents as the outgroup. The obvious questions, (1) which tree is the correct one, or (2) both trees can be incorrect, and (3) how can we explain such an evolutionary pattern, are discussed on the basis of our limited knowledge of factors that influence the clocklike behavior of biological macromolecules.

Animals

Direct evidence that restriction endonucleases may under estimate the degree of divergence between molecules.

We studied two polymorphic forms of mtDNA extracted from A. lixula eggs. In order to compare and to quantitate the variability, we sequenced specific regions of the two molecules. In this way, we obtained a precise measurement of the variability within two haplotypes. We also obtained a direct demonstration that some differences in nucleotide sequence can escape detection when restriction endonuclease analysis is used. Our results underline the unreliability of the use of restriction mapping to estimate divergence between relatively short and closely related DNA sequences.

Animals

DNA microenvironments and the molecular clock.

A few years ago we presented a stationary Markov model of gene evolution according to which only homologous genes from not too divergent species obeying the condition of being stationary may behave as reliable molecular clocks. A compartmentalized model of the nuclear genome in which the genes are distributed in compartments, the isochores, defined by their G + C content has been proposed recently. We have found that only homologous gene pairs that are stationary, and belong to the same isochore, can be used consistently for the determination of phylogeny and base substitution rate. In particular, for the rodent-human couple, only about half of the homologous gene pairs are stationary. Stationary genes evolve at the third silent codon position with the same velocity independent of the genes and base composition. By contrast, nonstationary genes display apparent rate values (pseudovelocities) that are significantly higher. Our results cast doubt upon recent claims of a large acceleration in the rate of molecular evolution in rodents.

Biological Evolution

A backtranslation method based on codon usage strategy.

This study describes a method for the backtranslation of an aminoacidic sequence, an extremely useful tool for various experimental approaches. It involves two computer programs CLUSTER and BACKTR written in Fortran 77 running on a VAX/VMS computer. CLUSTER generates a reliable codon usage table through a cluster analysis, based on a chi 2-like distance between the sequences. BACKTR produces backtranslated sequences according to different options when use is made of the codon usage table obtained in addition to selecting the least ambiguous potential oligonucleotide probes within an aminoacidic sequence. The method was tested by applying it to 158 yeast genes.

Amino Acid Sequence