Model of evolution of molecular sequences.
Explore the source record for details and available documents.
SEARCH · PubMed Health
Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
A search for genes located on human chromosome 21 resulted in the isolation of a HeLa cDNA clone, pUNC724, which hybridized to 3.7 and 2.5 kilobase (kb) EcoRI fragments on each of the human acrocentric chromosomes. In situ hybridization further localized pUNC724 to the pericentromeric region of the human acrocentrics. Two other EcoRI fragments that hybridized to pUNC724 were assigned to the long arms of chromosomes 1 and 18. The pUNC724 sequence does not appear to be related to ribosomal or satellite DNA sequences. The juxtaposition of DNA sequences homologous to pUNC724 and ribosomal DNA sequences presumably occurred within the past thirty-five million years, following the divergence of the lines leading to man and the New World owl monkey, Aotus trivirgatus--pUNC724 is not syntenic with the single chromosome containing ribosomal DNA sequences in the owl monkey.
Evolutionary models for estimating the number of nucleotide substitutions between DNA sequences are evaluated with data from histone gene sequences of 7 remote species. It is found that the nucleotide compositions at the third codon position of H2A genes vary greatly among species and are highly correlated with the compositions at the first position of H2A genes, with those at the first and third position of H4 genes and with those in the up- and downstream sequences of H2A genes. This implies the existence of regional constraints over DNA sequences during the evolutionary process, which is different over species. Possible causes for the variations are increment of G + C content in higher eukarotypes, and chromosomal recombination, which brought the histone genes onto different isochores and thus under different selective or mutational pressures. Substitutions at different positions in a codon have been found not to be independent, probably due to multiple substitutions, i.e., single substitution events involving multiple sites. The implication of these results to phylogeny inferring is discussed.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Sequences of 47 members of the Zn-containing alcohol dehydrogenase (ADH) family were aligned progressively, and an evolutionary tree with detailed branch order and branch lengths was produced. The alignment shows that only 9 amino acid residues (of 374 in the horse liver ADH sequence) are conserved in this family; these include eight Gly and one Val with structural roles. Three residues that bind the catalytic Zn and modulate its electrostatic environment are conserved in 45 members. Asp 223, which determines specificity for NAD, is found in all but the two NADP-dependent enzymes, which have Gly or Ala. Ser or Thr 48, which makes a hydrogen bond to the substrate, is present in 46 members. The four Cys ligands for the structural zinc are conserved except in zeta-crystallin, the sorbitol dehydrogenases, and two bacterial enzymes. Analysis of the evolutionary tree gives estimates of the times of divergence for different animal ADHs. The human class II (pi) and class III (chi) ADHs probably diverged about 630 million years ago, and the newly identified human ADH6 appeared about 520 million years ago, implying that these classes of enzymes may exist or have existed in all vertebrates. The human class I ADH isoenzymes (alpha, beta, and gamma) diverged about 80 million years ago, suggesting that these isoenzymes may exist or have existed in all primates. Analysis of branch lengths shows that these plant ADHs are more conserved than the animal ones and that class III ADHs are more conserved than class I ADHs. The rate of acceptance of point mutations (PAM units) shows that selection pressure has existed for ADHs, implying that these enzymes play definite metabolic roles.
A full-length cDNA encoding mouse annexin V (ANX5) was cloned, sequenced, and utilized for chromosomal mapping. The gene lies on mouse chromosome 3 in close linkage with the fibroblast growth factor 2 (basic) gene and is syntenic with other genes known to have orthologous counterparts on human chromosome 4q. The open reading frame encoded a protein of 319 amino acids (aa), with 92-96% identity to ANX5 in other species. Internal repeat 3 of mouse ANX5 exhibited the highest level of nonconservative aa replacements with respect to other annexin subfamilies, but the greatest sequence conservation among ANX5 species members. This region may thus contain features that distinguish ANX5 from other annexins in properties or function. Phylogenetic analysis and homology testing of ANX5 members indicated that the 34-kDa annexin from Torpedo marmorata may also belong to this subfamily. Comparison of nine species of ANX5 led to an estimation of the unit evolutionary mutation rate at 1% aa replacements every 8 million years, comparable to other annexins.
The sequences of immunoglobulin (Ig) heavy chain constant (C) region genes do not appear to be well conserved in evolution. The nature of C region change was examined by comparing VH and CH segment variability in the shark, which has numerous C mu as well as VH genes. The sequence diversity among VH was found to be similar to that among C mu 1, suggesting that one Ig segment does not change faster than the other. Although the frequencies of nucleotide substitution were similar, changes in the form of insertions or deletions in loop segments occurred more often in C; the resulting loss of sequence continuity in C exons would make it more difficult to recover C region sequences by crosshybridization. Thus the C regions would seem to be less well conserved than they are. The C region sequences isolated from different animals are described, and sequence and structure are discussed with respect to functions found in mammalian Ig. It is suggested that, although fish, amphibians and reptiles have fewer Ig classes than mammals, heterogeneous gene products are in fact produced and may mediate different effector functions. A theory based on C region selection is presented to explain antibody maturation in animals with a multicluster Ig gene organization.
A polymorphism of the variable number of tandem repeat (VNTR) type is located 97 bp downstream of exon VI of the parathyroid hormone-related peptide (PTHrP) gene in humans. The repeat unit has the general sequence G(TA)nC, where n equals 4-11. In order to characterize the evolutionary history of this VNTR, we initially tested for its presence in 13 different species representing four main groups of living primates. The sequence is present in the human, great apes, and Old World monkeys, but not in New World monkeys; and this region failed to PCR amplify in the Loris group. Thus, the evolution of the sequence as part of the PTHrP gene started at least 25-35 millions years ago, after divergence of the Old World and New World monkeys, but before divergence of Old World monkeys and great apes and humans. The structural changes occurring during evolution are characterized by a relatively high degree of sequence divergence. In general, the tandem repeat region tends to be longer and more complex in higher primates with the repeat unit motifs all being based on a TA-dinucleotide repeat sequence. Intra-species variability of the locus was demonstrated only in humans and gorilla. The divergence of the TA-dinucleotide repeat sequence and the variable mutation rates observed in different primate species are in contrast to the relative conservation of the flanking sequences during primate evolution. This suggests that the nature of the TA-dinucleotide repeat sequence, rather than its flanking sequences, is responsible for generating variability.(ABSTRACT TRUNCATED AT 250 WORDS)
BACKGROUND: A large number of bioinformatics applications in the fields of bio-sequence analysis, molecular evolution and population genetics typically share input/output methods, data storage requirements and data analysis algorithms. Such common features may be conveniently bundled into re-usable libraries, which enable the rapid development of new methods and robust applications. RESULTS: We present Bio++, a set of Object Oriented libraries written in C++. Available components include classes for data storage and handling (nucleotide/amino-acid/codon sequences, trees, distance matrices, population genetics datasets), various input/output formats, basic sequence manipulation (concatenation, transcription, translation, etc.), phylogenetic analysis (maximum parsimony, markov models, distance methods, likelihood computation and maximization), population genetics/genomics (diversity statistics, neutrality tests, various multi-locus analyses) and various algorithms for numerical calculus. CONCLUSION: Implementation of methods aims at being both efficient and user-friendly. A special concern was given to the library design to enable easy extension and new methods development. We defined a general hierarchy of classes that allow the developer to implement its own algorithms while remaining compatible with the rest of the libraries. Bio++ source code is distributed free of charge under the CeCILL general public licence from its website http://kimura.univ-montp2.fr/BioPP.
Nucleotide sequences of all genomes are subject to compositional constraints that affect, to about the same extent, both coding and noncoding sequences; influence not only the structure and function of the genome, but also those of transcripts and proteins; are the result of environmental pressures; and largely control the fixation of mutations. These findings indicate that noncoding sequences are associated with biological functions; that the organismal phenotype comprises two components, the classical phenotype, corresponding to the "gene products," and a "genome phenotype," which is defined by the compositional constraints; and that natural selection plays a more important role in genome evolution than do random events.
BACKGROUND: Structure conservation constrains evolutionary sequence divergence, resulting in observable sequence patterns. Most current models of protein evolution do not take structure into account explicitly, being unsuitable for investigating the effects of structure conservation on sequence divergence. To this end, we recently developed the Structurally Constrained Protein Evolution (SCPE) model. The model starts with the coding sequence of a protein with known three-dimensional structure. At each evolutionary time-step of an SCPE simulation, a trial sequence is generated by introducing a random point mutation in the current coding DNA sequence. Then, a "score" for the trial sequence is calculated and the mutation is accepted only if its score is under a given cutoff, lambda. The SCPE score measures the distance between the trial sequence and a given reference sequence, given the structure. In our first brief report we used a "global score", in which the same reference sequence, the ancestral one, was used at each evolutionary step. Here, we introduce a new scoring function, the "local score", in which the sequence accepted at the previous evolutionary time-step is used as the reference. We assess the model on the UDP-N-acetylglucosamine acyltransferase (LPXA) family, as in our previous report, and we extend this study to all other members of the left-handed parallel beta helix fold (LbetaH) superfamily whose structure has been determined. RESULTS: We studied site-dependent entropies, amino acid probability distributions, and substitution matrices predicted by SCPE and compared with experimental data for several members of the LbetaH superfamily. We also evaluated structure conservation during simulations. Overall, SCPE outperforms JTT in the description of sequence patterns observed in structurally constrained sites. Maximum Likelihood calculations show that the local-score and global-score SCPE substitution matrices obtained for LPXA outperform the JTT model for the LPXA family and for the structurally constrained sites of class i of other members within the LbetaH superfamily. CONCLUSION: We extended the SCPE model by introducing a new scoring function, the local score. We performed a thorough assessment of the SCPE model on the LPXA family and extended it to all other members of known structure of the LbetaH superfamily.
5S rRNA sequences were determined for the myxobacteria Cystobacter fuscus, Myxococcus coralloides, Sorangium cellulosum, and Nannocystis exedens and for the radioresistant bacteria Deinococcus radiodurans and Deinococcus radiophilus. A dendrogram was constructed by using weighted pairwise grouping based on these and all other previously known eubacterial 5S rRNA sequences, and this dendrogram showed differences as well as similarities compared with results derived from 16S rRNA analyses. In the dendrogram, Deinococcus 5S rRNA sequences clustered with 5S rRNA sequences of the genus Thermus, as suggested by the results of 16S rRNA analyses. However, in contrast to the 16S rRNA results, the Deinococcus-Thermus cluster divided the 5S rRNA sequences of the alpha subdivision of the class Proteobacteria from the 5S rRNA sequences of the beta and gamma subgroups of the Proteobacteria. The myxobacterial 5S rRNA sequence data failed to confirm the existence of a delta subgroup of the class Proteobacteria, which was suggested by the results of 16S rRNA analyses.
The genome sequences of multiple species has enabled functional inferences from comparative genomics. A primary objective is to infer biological functions from the conservation of homologous DNA sequences between species. A second, more difficult, objective is to understand what functional DNA sequences have changed over time and are responsible for species' phenotypic differences. The neutral theory of molecular evolution provides a theoretical framework in which both objectives can be explicitly tested. Development of statistical tests within this framework has provided insight into the evolutionary forces that constrain and in some cases change DNA sequences and the resulting patterns that emerge. In this article, we review recent work on how functional constraint and changes in protein function are inferred from protein polymorphism and divergence data. We relate these studies to our understanding of the neutral theory and adaptive evolution.
Retroviral replication is a very error-prone process. Replication of retroviruses gives rise to populations of closely related but different genomes referred to as 'quasispecies'. This huge swarm of different sequences constitutes a reservoir of potentially useful genomes in case of an environmental change, endowing retroviruses with extreme adaptability. Retrotransposons are mobile genetic elements closely related to retroviruses, and retrotransposition is as error prone as retroviral replication. The Tnt1 retrotransposon is present in hundreds of copies in the genome of tobacco that show a high level of sequence heterogeneity. When Tnt1 is expressed, its RNA is not a single sequence but a population of sequences displaying a quasispecies-like structure. This population structure gives to Tnt1, as in the case of retroviruses, a high sequence plasticity and an adaptive capacity. We propose this adaptivity as the major reason for Tnt1 maintenance in Nicotiana genomes and we discuss in this paper the importance of sequence variability for Tnt1 evolution.
The nucleotide sequences of the 5S ribosomal RNAs of the bacteria Agrobacterium tumefaciens, Alcaligenes faecalis, Pseudomonas cepacia, Aquaspirillum serpens and Acinetobacter calcoaceticus have been determined. The sequences fit in a generally accepted model for 5S RNA secondary structure. However, a closer comparative examination of these and other bacterial 5S RNA primary structures reveals the potential of additional base pairing and of multiple equilibria between a set of slightly different alternative secondary structures in one area of the molecule. The phylogenetic position of the examined bacteria is derived from a 5S RNA sequence alignment by a clustering method and compared with the position derived on the basis of 16S ribosomal RNA oligonucleotide catalogs.
A large fraction, sometimes the largest fraction, of a eukaryotic genome consists of repeated DNA sequences. Copy numbers range from several thousand to millions per diploid genome. All classes of repetitive DNA sequences examined to date exhibit apparently general, but little studied, patterns of "concerted evolution." Historically, concerted evolution has been defined as the nonindependent evolution of repetitive DNA sequences, resulting in a sequence similarity of repeating units that is greater within than among species. This intraspecific homogenization of repetitive sequence arrays is said to take place via the poorly understood mechanisms of "molecular drive." The evolutionary population dynamics of molecular drive remains largely unstudied in natural populations, and thus the potential significance of these evolutionary dynamics for population differentiation is unknown. This review attempts to demonstrate the potential importance of the mechanisms responsible for concerted evolution in the differentiation of populations. It contends that any natural grouping that is characterized by reproductive isolation and limited gene flow is capable of exhibiting concerted evolution of repetitive DNA arrays. Such effects are known to occur in protein and RNA-coding repetitive sequences, as well as in so-called "junk DNA," and thus have important implications for the differentiation and discrimination of natural populations.
Sequence folding is known to determine the spatial structure and catalytic function of proteins and nucleic acids. We show here that folding also plays a key role in enhancing the evolutionary stability of the intermolecular recognition necessary for the prevalent mode of catalytic action in replication, namely, in trans, one molecule catalyzing the replication of another copy, rather than itself. This points to a novel aspect of why molecular life is structured as it is, in the context of life as it could be: folding allows limited, structurally localized recognition to be strongly sensitive to global sequence changes, facilitating the evolution of cooperative interactions. RNA secondary structure folding, for example is shown to be able to stabilize the evolution of prolonged functional sequences, using only a part of this length extension for intermolecular recognition, beyond the limits of the (cooperative) error threshold. Such folding could facilitate the evolution of polymerases in spatially heterogeneous systems. This facilitation is, in fact, vital because physical limitations prevent complete sequence-dependent discrimination for any significant-size biopolymer substrate. The influence of partial sequence recognition between biopolymer catalysts and complex substrates is investigated within a stochastic, spatially resolved evolutionary model of trans catalysis. We use an analytically tractable nonlinear master equation formulation called PRESS (McCaskill et al., Biol. Chem. 382: 1343-1363), which makes use of an extrapolation of the spatial dynamics down from infinite dimensional space, and compare the results with Monte Carlo simulations.