PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

The genetic architecture of resistance.

Plant resistance genes (R genes), especially the nucleotide binding site leucine-rich repeat (NBS-LRR) family of sequences, have been extensively studied in terms of structural organization, sequence evolution and genome distribution. These studies indicate that NBS-LRR sequences can be split into two related groups that have distinct amino-acid motif organizations, evolutionary histories and signal transduction pathways. One NBS-LRR group, characterized by the presence of a Toll/interleukin receptor domain at the amino-terminal end, seems to be absent from the Poaceae. Phylogenetic analysis suggests that a small number of NBS-LRR sequences existed among ancient Angiosperms and that these ancestral sequences diversified after the separation into distinct taxonomic families. There are probably hundreds, perhaps thousands, of NBS-LRR sequences and other types of R gene-like sequences within a typical plant genome. These sequences frequently reside in 'mega-clusters' consisting of smaller clusters with several members each, all localized within a few million base pairs of one another. The organization of R-gene clusters highlights a tension between diversifying and conservative selection that may be relevant to gene families that are unrelated to disease resistance.

Amino Acid Motifs↗

Factors affecting the errors in the estimation of evolutionary distances between sequences.

Phylogenetic methods that use matrices of pairwise distances between sequences (e.g., neighbor joining) will only give accurate results when the initial estimates of the pairwise distances are accurate. For many different models of sequence evolution, analytical formulae are known that give estimates of the distance between two sequences as a function of the observed numbers of substitutions of various classes. These are often of a form that we call "log transform formulae". Errors in these distance estimates become larger as the time t since divergence of the two sequences increases. For long times, the log transform formulae can sometimes give divergent distance estimates when applied to finite sequences. We show that these errors become significant when t approximately 1/2 |lambda(max)|(-1) logN, where lambda(max) is the eigenvalue of the substitution rate matrix with the largest absolute value and N is the sequence length. Various likelihood-based methods have been proposed to estimate the values of parameters in rate matrices. If rate matrix parameters are known with reasonable accuracy, it is possible to use the maximum likelihood method to estimate evolutionary distances while keeping the rate parameters fixed. We show that errors in distances estimated in this way only become significant when t approximately 1/2 |lambda(1)|(-1) logN, where lambda(1) is the eigenvalue of the substitution rate matrix with the smallest nonzero absolute value. The accuracy of likelihood-based distance estimates is therefore much higher than those based on log transform formulae, particularly in cases where there is a large range of timescales involved in the rate matrix (e.g., when the ratio of transition to transversion rates is large). We discuss several practical ways of estimating the rate matrix parameters before distance calculation and hence of increasing the accuracy of distance estimates.

Evolution, Molecular↗

Divergent evolution of an "orphon" histone gene cluster in Chironomus.

The histone genes of the midge Chironomus thummi thummi are organized in tandemly repeated gene groups, each containing the four core histone genes plus an H1 gene. These repetitive gene groups are found at five different loci, linked on one chromosomal arm. In addition to the clustered gene groups an isolated histone gene group exists which is found spatially separated on a different chromosome ("orphon" gene group). These orphon genes have been cloned and analysed in detail. Nucleotide sequence and in situ hybridization data suggest that the orphon gene group was established early during chironomid speciation, possibly by a transposition-like mechanism. This allowed the genes to be moved as an integer group. The comparison of orphon and "clustered" histone genes in C. thummi thummi indicates that the early spatial separation of the orphon genes from their tandemly organized relatives may have stimulated divergent sequence evolution. This is particularly true for the orphon H1 gene, which has diverged considerably by unusual mutation mechanisms. The translocation of normally clustered genes to new genomic sites may favour the generation of sequence variants, which could fulfill specialized functional tasks.

Animals↗

Evolution of structure and function of V-ATPases.

Proton pumping ATPases/ATPsynthases are found in all groups of present-day organisms. The structure of V- and F-type ATPases/ATP synthases is very conserved throughout evolution. Sequence analysis shows that the V- and F-type ATPases evolved from the same enzyme already present in the last common ancestor of all known extant life forms. The catalytic and noncatalytic subunits found in the dissociable head groups of the V/F-type ATPases are paralogous subunits, i.e., these two types of subunits evolved from a common ancestral gene. The gene duplication giving rise to these two genes (i.e., encoding the catalytic and noncatalytic subunits) predates the time of the last common ancestor. Mapping of gene duplication events that occurred in the evolution of the proteolipid, the noncatalytic and the catalytic subunits, onto the tree of life leads to a prediction for the likely subunit structure of the encoded ATPases. A correlation between structure and function of V/F-ATPases has been established for present-day organisms. Implications resulting from this correlation for the bioenergetics operative in proto-eukaryotes and in the last common ancestor are presented. The similarities of the V/F-ATPase subunits to an ATPase-like protein that was implicated to play a role in flagellar assembly are evaluated. Different V-ATPase isoforms have been detected in some higher eukaryotes. These data are analyzed with respect to the possible function of the different isoforms (tissue specific, organelle specific) and with respect to the point in their evolution when these gene duplications giving rise to the isoforms had occurred, i.e., how far these isoforms are distributed.

Amino Acid Sequence↗

What can and what cannot be inferred from pairwise sequence comparisons?

We address questions of identifiability in molecular phylogeny, the art of reconstructing the history of a sample of sequences given just the sequences at the leaves of the phylogenetic tree. Here, the 'history' consists of the tree topology, plus the transition probabilities which define the Markov process of sequence evolution along the branches of the tree. It is assumed that sequences have infinite length, and the pairwise joint distributions of letters at the leaves is taken to be known. We focus on two cases: (1) If the sites of a sequence evolve identically and independently, the topology can be reconstructed, but the one-way edge transition matrices cannot. However, the return-trip transition matrices are reconstructible for every edge, up to conjugation in the case of internal edges. (2) If a rate factor varies from site to site, different topologies may produce identical pairwise joint distributions, even under the same distribution of rate factors. Consequently, identifiability of the topology is lost on the basis of pairwise sequence comparisons, even if the distribution of rate factors is known. The results are discussed in the context of additive measures of phylogenetic distance.

Base Sequence↗

Molecular phylogenetics: state-of-the-art methods for looking into the past.

As the amount of molecular sequence data in the public domain grows, so does the range of biological topics that it influences through evolutionary considerations. In recent years, a number of developments have enabled molecular phylogenetic methodology to keep pace. Likelihood-based inferential techniques, although controversial in the past, lie at the heart of these new methods and are producing the promised advances in the understanding of sequence evolution. They allow both a wide variety of phylogenetic inferences from sequence data and robust statistical assessment of all results. It cannot remain acceptable to use outdated data analysis techniques when superior alternatives exist. Here, we discuss the most important and exciting methods currently available to the molecular phylogeneticist.

Animals↗

Molecular phylogeny: pitfalls and progress.

Molecular phylogeny based on nucleotide or amino acid sequence comparison has become a widespread tool for general taxonomy and evolutionary analyses. It seems the only means to establish a natural classification of microorganisms, since their phenotypic traits are not always consistent with genealogy. After an optimistic period during which comprehensive microbial evolutionary pictures appeared, the discovery of several pitfalls affecting molecular phylogenetic reconstruction challenged the general validity of this approach. In addition to biological factors, such as horizontal gene transfer, some methodological problems may produce misleading phylogenies. They are essentially (i) loss of phylogenetic signal by the accumulation of overlapping mutations, (ii) incongruity between the real evolutionary process and the assumed models of sequence evolution, and (iii) differences of evolutionary rates among species or among positions within a sequence. Here, we discuss these problems and some strategies proposed to overcome their effects.

Artifacts↗

Structure and evolution of plant disease resistance genes.

This article reviews recent advances that shed light on plant disease resistance genes, beginning with a brief overview of their structure, followed by their genomic organization and evolution. Plant disease resistance genes have been exhaustively investigated in terms of their structural organization, sequence evolution and genome distribution. There are probably hundreds of NBS-LRR sequences and other types of R-gene-like sequences within a typical plant genome. Recent studies revealed positive selection and selective maintenance of variation in plant resistance and defence-related genes. Plant resistance genes are highly polymorphic and have diverse recognition specificities. R-genes occur as members of clustered gene families that have evolved through duplication and diversification. These genes appear to evolve more rapidly than other regions of the genome, and domains such as the leucine-rich repeat, are subject to adaptive selection

Evolution, Molecular↗

Contrasting patterns of nonneutral evolution in proteins encoded in nuclear and mitochondrial genomes.

We report that patterns of nonneutral DNA sequence evolution among published nuclear and mitochondrially encoded protein-coding loci differ significantly in animals. Whereas an apparent excess of amino acid polymorphism is seen in most (25/31) mitochondrial genes, this pattern is seen in fewer than half (15/36) of the nuclear data sets. This differentiation is even greater among data sets with significant departures from neutrality (14/15 vs. 1/6). Using forward simulations, we examined patterns of nonneutral evolution using parameters chosen to mimic the differences between mitochondrial and nuclear genetics (we varied recombination rate, population size, mutation rate, selective dominance, and intensity of germ line bottleneck). Patterns of evolution were correlated only with effective population size and strength of selection, and no single genetic factor explains the empirical contrast in patterns. We further report that in Arabidopsis thaliana, a highly self-fertilizing plant with effectively low recombination, five of six published nuclear data sets also exhibit an excess of amino acid polymorphism. We suggest that the contrast between nuclear and mitochondrial nonneutrality in animals stems from differences in rates of recombination in conjunction with a distribution of selective effects. If the majority of mutations segregating in populations are deleterious, high linkage may hinder the spread of the occasional beneficial mutation.

Animals↗

Evolutionarily different alphoid repeat DNA on homologous chromosomes in human and chimpanzee.

Centromeric alphoid DNA in primates represents a class of evolving repeat DNA. In humans, chromosomes 13 and 21 share one subfamily of alphoid DNA while chromosomes 14 and 22 share another subfamily. We show that similar pairwise homogenizations occur in the chimpanzee (Pan troglodytes), where chromosomes 14 and 22, homologous to human chromosomes 13 and 21, share one partially homogenized alphoid DNA subfamily and chromosomes 15 and 23, homologous to human chromosomes 14 and 22, share another extensively homogenized subfamily. Such a pattern of homogenization presumably predates speciation 3-10 million years ago. However, the alphoid DNA on these human and chimpanzee chromosomes is not orthologous but originates from two evolutionarily different repeat families. It follows that dramatic sequence evolution has occurred in a concerted fashion among the chromosomes in one or both species during or after separation.

Animals↗

Evolutionary implications of multiple SINE insertions in an intronic region from diverse mammals.

An analysis of the nuclear beta-fibrinogen intron 7 locus from 30 taxa representing 12 placental orders of mammals reveals the enriched occurrences of short interspersed element (SINE) insertion events. Mammalian-wide interspersed repeats (MIRs) are present at orthologous sites of all examined species except those in the order Rodentia. The higher substitution rate in mouse and a rare MIR deletion from rat account for the absence of MIR in the rodents. A minimum of five lineage-specific SINE sequences are also found to have independently inserted into this intron in Carnivora, Artiodactyla and Lagomorpha. In the case of Carnivora, the unique amplification pattern of order-specific CAN SINE provides important evidence for the "pan-carnivore" hypothesis of this repeat element and reveals that the CAN SINE family may still be active today. Particularly interesting is the finding that all identified lineage-specific SINE elements show a strong tendency to insert within or in very close proximity to the preexisting MIRs for their efficient integrations, suggesting that the MIR element is a hot spot for successive insertions of other SINEs. The unexpected MIR excision as a result of a random deletion in the rat intron locus and the non-random site targeting detected by this study indicate that SINEs actually have a greater insertional flexibility and regional specificity than had previously been recognized. Implications for SINE sequence evolution upon and following integration, as well as the fascinating interactions between retroposons and the host genomes are discussed.

Animals↗

A priori estimation of phylogenetic information conserved in aligned sequences.

A new phenomenological approach to explorative data analysis, the estimation of spectra of supporting positions, allows the search for conserved tracks left by phylogeny in DNA sequences. Spectra of supporting positions can be generated without reference to a tree topology or a model of sequence evolution and are therefore an ideal tool for a priori estimation of information content of data sets. Analysis of published 18S rDNA alignments shows that signal to noise relationship varies greatly in a way not detected by conventional tree-construction methods.

Animals↗

Does recombination shape the distribution and evolution of tandemly arrayed genes (TAGs) in the Arabidopsis thaliana genome?

Tandemly arrayed genes (TAGs) are an important genomic component. However, most previous studies have focused on individual TAG families, and a broader characterization of their genomic distribution is not yet available. In this study, we examined the distribution of TAGs in the Arabidopsis thaliana genome and examined TAG density with relation to recombination rates. Recombination rates along A. thaliana chromosomes were estimated by comparing a genetic map with the genome sequence. Average recombination rates in A. thaliana are high, and rates vary more than threefold among chromosomal regions. Comparisons between TAG density and recombination indicate a positive correlation on chromosomes 1, 2, and 3. Moreover, there is a consistent centromeric effect. Relative to single-copy genes, TAGs are proportionally less frequent in centromeres than on chromosomal arms. We also examined several factors that have been proposed to affect the sequence evolution of TAG members. Sequence divergence is related to the number of members in the TAG, but genomic location has no obvious effect on TAG sequence divergence, nor does the presence of unrelated genes within a TAG. Overall, the distribution of TAGs in the genome is not consistent with theoretical models predicting the accumulation of repeats in regions of low recombination but may be consistent with stabilizing selection models of TAG evolution.

Arabidopsis↗

Phylogenetic relations of humans and African apes from DNA sequences in the psi eta-globin region.

Sequences from the upstream and downstream flanking DNA regions of the psi eta-globin locus in Pan troglodytes (common chimpanzee), Gorilla gorilla (gorilla), and Pongo pygmaeus (orangutan, the closest living relative to Homo, Pan, and Gorilla) provided further data for evaluating the phylogenetic relations of humans and African apes. These newly sequenced orthologs [an additional 4.9 kilobase pairs (kbp) for each species] were combined with published psi eta-gene sequences and then compared to the same orthologous stretch (a continuous 7.1-kbp region) available for humans. Phylogenetic analysis of these nucleotide sequences by the parsimony method indicated (i) that human and chimpanzee are more closely related to each other than either is to gorilla and (ii) that the slowdown in the rate of sequence evolution evident in higher primates is especially pronounced in humans. These results indicate that features (for example, knuckle-walking) unique to African apes (but not to humans) are primitive and that even local molecular clocks should be applied with caution.

Animals↗

Molecular epidemiology of rabies virus in France: comparison with vaccine strains.

A molecular epidemiological study of the rabies virus currently prevalent in France was carried out by directly sequencing polymerase chain reaction-amplified genes. The rabies virus pseudogene psi was chosen as the most divergent genomic area, and as such the best 'clock' for measuring virus evolution. Sequence comparisons between 12 wild rabies virus isolates indicated strong conservation whatever the host and wherever the virus had been isolated. This holds true for a unique wild reservoir, the fox. On the other hand, a good correlation between genetic and geographical criteria indicates a slow evolution of the wild virus in parallel with the spatio-temporal progression of the epizootic. In contrast to their intrinsic homogeneity (about 2% divergence), the wild isolate sequences showed a marked divergence from those of vaccine seed strains (about 14.7%). This finding invites world-wide molecular epidemiological studies, particularly in countries in which vaccination failures have been reported.

Animals↗

The roles of positive and negative selection in the molecular evolution of insect endosymbionts.

The evolutionary rate acceleration observed in most endosymbiotic bacteria may be explained by higher mutation rates, changes in selective pressure, and increased fixation of deleterious mutations by genetic drift. Here, we explore the forces influencing molecular evolution in Blochmannia, an obligate endosymbiont of Camponotus and related ant genera. Our goals were to compare rates of sequence evolution in Blochmannia with related bacteria, to explore variation in the strength and efficacy of negative (purifying) selection, and to evaluate the effect of positive selection. For six Blochmannia pairs, plus Buchnera and related enterobacteria, estimates of sequence divergence at four genes confirm faster rates of synonymous evolution in the ant mutualist. This conclusion is based on higher dS between Blochmannia lineages despite their more recent divergence. Likewise, generally higher dN in Blochmannia indicates faster rates of nonsynonymous substitution in this group. One exception is the groEL gene, for which lower dN and dN/dS compared to Buchnera indicate exceptionally strong negative selection in Blochmannia. In addition, we explored evidence for positive selection in Blochmannia using both site-and lineage-based maximum likelihood models. These approaches confirmed heterogeneity of dN/dS among codon sites and revealed significant variation in dN/dS across Blochmannia lineages for three genes. Lineage variation affected genes independently, with no evidence of parallel changes in dN/dS across genes along a given branch. Our data also reveal instances of dN/dS greater than one; however, we do not interpret these large dN/dS ratios as evidence for positive selection. In sum, while drift may contribute to an overall rate acceleration at nonsynonymous sites in Blochmannia, variable selective pressures best explain the apparent gene-specific changes in dN/dS across lineages of this ant mutualist. In the course of this study, we reanalyzed variation at Buchnera groEL and found no evidence of positive selection that was previously reported.

Animals↗