PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

Functional and evolutionary implications of natural variation in clock genes.

Nearly all studies of natural variation within clock genes involve the period (per) locus, which was originally isolated in the fruit-fly. Intra- and interspecific work on per has focused mostly on a region of Thr-Gly or Ser-Gly repeats, which show rapid length and sequence evolution. The functional implications of nucleotide variation in this repetitive array have been characterised using behavioural, molecular, ecological, structural and statistical analyses. A population genetics approach to variation in per has also been useful in defining species histories within Drosophilids and, in some cases, in implicating selective processes in the evolution of the per gene. Interspecific analysis of per expression patterns reveals evolutionary alterations in this clock gene's regulation.

Animals↗

A phylogeny of the megapodes (Aves: Megapodiidae) based on nuclear and mitochondrial DNA sequences.

DNA sequences from the first intron of the nuclear gene rhodopsin (RDP1) and from the mitochondrial gene ND2 were used to construct a phylogeny of the avian family Megapodiidae. RDP1 sequences evolved about six times more slowly than ND2 and showed less homoplasy, substitution bias, and rate heterogeneity across sites. Analysis of RDP1 produced a phylogeny that was well resolved at the genus level, but RDP1 did not evolve rapidly enough for intrageneric comparisons. The ND2 phylogeny resolved intrageneric relationships and was congruent with the RDP1 phylogeny except for a single node: this node was the only aspect of tree topology sensitive to weighting in parsimony analyses. Despite differences in sequence evolution, RDP1 and ND2 contained congruent phylogenetic signal and were combined to produce a phylogeny that reflects the resolving power of both genes. This phylogeny shows an early split within the megapodes, leading to two major clades: (1) Macrocephalon and the mound-building genera Talegalla, Leipoa, Aepypodius, and Alectura, and (2) Eulipoa and Megapodius. It differs significantly from previous hypotheses based on morphology but is consistent with affiliations suggested by a recent study of parasitic chewing lice.

Animals↗

The evolutionary origin of Indian Ocean tortoises (Dipsochelys).

Today, the only surviving wild population of giant tortoises in the Indian Ocean occurs on the island of Aldabra. However, giant tortoises once inhabited islands throughout the western Indian Ocean. Madagascar, Africa, and India have all been suggested as possible sources of colonization for these islands. To address the origin of Indian Ocean tortoises (Dipsochelys, formerly Geochelone gigantea), we sequenced the 12S, 16S, and cyt b genes of the mitochondrial DNA. Our phylogenetic analysis shows Dipsochelys to be embedded within the Malagasy lineage, providing evidence that Indian Ocean giant tortoises are derived from a common Malagasy ancestor. This result points to Madagascar as the source of colonization for western Indian Ocean islands by giant tortoises. Tortoises are known to survive long oceanic voyages by floating with ocean currents, and thus, currents flowing northward towards the Aldabra archipelago from the east coast of Madagascar would have provided means for the colonization of western Indian Ocean islands. Additionally, we found an accelerated rate of sequence evolution in the two Malagasy Pyxis species examined. This finding supports previous theories that shorter generation time and smaller body size are related to an increase in mitochondrial DNA substitution rate in vertebrates.

Animals↗

Bacterial endosymbionts in animals.

Molecular phylogenetic studies reveal that many endosymbioses between bacteria and invertebrate hosts result from ancient infections followed by strict vertical transmission within host lineages. Endosymbionts display a distinctive constellation of genetic properties including AT-biased base composition, accelerated sequence evolution, and, at least sometimes, small genome size; these features suggest increased genetic drift. Molecular genetic characterization also has revealed adaptive, host-beneficial traits such as amplification of genes underlying nutrient provision.

Animals↗

Age and rate of diversification of the Hawaiian silversword alliance (Compositae).

Comparisons between insular and continental radiations have been hindered by a lack of reliable estimates of absolute diversification rates in island lineages. We took advantage of rate-constant rDNA sequence evolution and an "external" calibration using paleoclimatic and fossil data to determine the maximum age and minimum diversification rate of the Hawaiian silversword alliance (Compositae), a textbook example of insular adaptive radiation in plants. Our maximum-age estimate of 5.2 +/- 0.8 million years ago for the most recent common ancestor of the silversword alliance is much younger than ages calculated by other means for the Hawaiian drosophilids, lobelioids, and honeycreepers and falls approximately within the history of the modern high islands (</=5.1 +/- 0.2 million years ago). By using a statistically efficient estimator that reduces error variance by incorporating clock-based estimates of divergence times, a minimum diversification rate for the silversword alliance was estimated to be 0.56 +/- 0.17 species per million years. This exceeds average rates of more ancient continental radiations and is comparable to peak rates in taxa with sufficiently rich fossil records that changes in diversification rate can be reconstructed.

Journal Article↗

An empirical assessment of long-branch attraction artefacts in deep eukaryotic phylogenomics.

In the context of exponential growing molecular databases, it becomes increasingly easy to assemble large multigene data sets for phylogenomic studies. The expected increase of resolution due to the reduction of the sampling (stochastic) error is becoming a reality. However, the impact of systematic biases will also become more apparent or even dominant. We have chosen to study the case of the long-branch attraction artefact (LBA) using real instead of simulated sequences. Two fast-evolving eukaryotic lineages, whose evolutionary positions are well established, microsporidia and the nucleomorph of cryptophytes, were chosen as model species. A large data set was assembled (44 species, 133 genes, and 24,294 amino acid positions) and the resulting rooted eukaryotic phylogeny (using a distant archaeal outgroup) is positively misled by an LBA artefact despite the use of a maximum likelihood-based tree reconstruction method with a complex model of sequence evolution. When the fastest evolving proteins from the fast lineages are progressively removed (up to 90%), the bootstrap support for the apparently artefactual basal placement decreases to virtually 0%, and conversely only the expected placement, among all the possible locations of the fast-evolving species, receives increasing support that eventually converges to 100%. The percentage of removal of the fastest evolving proteins constitutes a reliable estimate of the sensitivity of phylogenetic inference to LBA. This protocol confirms that both a rich species sampling (especially the presence of a species that is closely related to the fast-evolving lineage) and a probabilistic method with a complex model are important to overcome the LBA artefact. Finally, we observed that phylogenetic inference methods perform strikingly better with simulated as opposed to real data, and suggest that testing the reliability of phylogenetic inference methods with simulated data leads to overconfidence in their performance. Although phylogenomic studies can be affected by systematic biases, the possibility of discarding a large amount of data containing most of the nonphylogenetic signal allows recovering a phylogeny that is less affected by systematic biases, while maintaining a high statistical support.

Animals↗

The systematic component of phylogenetic error as a function of taxonomic sampling under parsimony.

The effect of taxonomic sampling on phylogenetic accuracy under parsimony is examined by simulating nucleotide sequence evolution. Random error is minimized by using very large numbers of simulated characters. This allows estimation of the consistency behavior of parsimony, even for trees with up to 100 taxa. Data were simulated on 8 distinct 100-taxon model trees and analyzed as stratified subsets containing either 25 or 50 taxa, in addition to the full 100-taxon data set. Overall accuracy decreased in a majority of cases when taxa were added. However, the magnitude of change in the cases in which accuracy increased was larger than the magnitude of change in the cases in which accuracy decreased, so, on average, overall accuracy increased as more taxa were included. A stratified sampling scheme was used to assess accuracy for an initial subsample of 25 taxa. The 25-taxon analyses were compared to 50- and 100-taxon analyses that were pruned to include only the original 25 taxa. On average, accuracy for the 25 taxa was improved by taxon addition, but there was considerable variation in the degree of improvement among the model trees and across different rates of substitution.

Classification↗

Linking dynamical and population genetic models of persistent viral infection.

This article develops a theoretical framework to link dynamical and population genetic models of persistent viral infection. This linkage is useful because, while the dynamical and population genetic theories have developed independently, the biological processes they describe are completely interrelated. Parameters of the dynamical models are important determinants of evolutionary processes such as natural selection and genetic drift. We develop analytical methods, based on coupled differential equations and Markov chain theory, to predict the accumulation of genetic diversity within the viral population as a function of dynamical parameters. These methods are first applied to the standard model of viral dynamics and then generalized to consider the infection of multiple host cell types by the viral population. Each cell type is characterized by specific parameter values. Inclusion of multiple cell types increases the likelihood of persistent infection and can increase the amount of genetic diversity within the viral population. However, the overall rate of gene sequence evolution may actually be reduced.

Biological Evolution↗

Disparity index: a simple statistic to measure and test the homogeneity of substitution patterns between molecular sequences.

A common assumption in comparative sequence analysis is that the sequences have evolved with the same pattern of nucleotide substitution (homogeneity of the evolutionary process). Violation of this assumption is known to adversely impact the accuracy of phylogenetic inference and tests of evolutionary hypotheses. Here we propose a disparity index, ID, which measures the observed difference in evolutionary patterns for a pair of sequences. On the basis of this index, we have developed a Monte Carlo procedure to test the homogeneity of the observed patterns. This test does not require a priori knowledge of the pattern of substitutions, extent of rate heterogeneity among sites, or the evolutionary relationship among sequences. Computer simulations show that the ID-test is more powerful than the commonly used chi2-test under a variety of biologically realistic models of sequence evolution. An application of this test in an analysis of 3789 pairs of orthologous human and mouse protein-coding genes reveals that the observed evolutionary patterns in neutral sites are not homogeneous in 41% of the genes, apparently due to shifts in G + C content. Thus, the proposed test can be used as a diagnostic tool to identify genes and lineages that have evolved with substantially different evolutionary processes as reflected in the observed patterns of change. Identification of such genes and lineages is an important early step in comparative genomics and molecular phylogenetic studies to discover evolutionary processes that have shaped organismal genomes.

Animals↗

Unequally Abundant Chromosomes and Unusual Collections of Transferred Sequences Characterize Mitochondrial Genomes of Gastrodia (Orchidaceae), One of the Largest Mycoheterotrophic Plant Genera.

The mystery of genomic alternations in heterotrophic plants is among the most intriguing in evolutionary biology. Compared to plastid genomes (plastomes) with parallel size reduction and gene loss, mitochondrial genome (mitogenome) variation in heterotrophic plants remains underexplored in many aspects. To further unravel the evolutionary outcomes of heterotrophy, we present a comparative mitogenomic study with 13 de novo assemblies of Gastrodia (Orchidaceae), one of the largest fully mycoheterotrophic plant genera, and its relatives. Analyzed Gastrodia mitogenomes range from 0.56 to 2.1 Mb, each consisting of numerous, unequally abundant chromosomes or contigs. Size variation might have evolved through chromosome rearrangements followed by stochastic loss of "dispensable" chromosomes, with deletion-biased mutations. The discovery of a hyper-abundant (&#x223c;15 times intragenomic average) chromosome in two assemblies represents the hitherto most extreme copy number variation in any mitogenomes, with similar architectures discovered in two metazoan lineages. Transferred sequence contents highlight asymmetric evolutionary consequences of heterotrophy: despite drastically reduced intracellular plastome transfers convergent across heterotrophic plants, their rarity of horizontally acquired sequences sharply contrasts parasitic plants, where massive transfers from their hosts prevail. Rates of sequence evolution are markedly elevated but not explained by copy number variation, extending prior findings of accelerated molecular evolution from parasitic to heterotrophic plants. Putative evolutionary scenarios for these mitogenomic convergence and divergence fit well with the common (e.g. plastome contraction) and specific (e.g. host identity) aspects of the two heterotrophic types. These idiosyncratic mycoheterotrophs expand known architectural variability of plant mitogenomes and provide mechanistic insights into their content and size variation.

Genome, Mitochondrial↗

Significantly different patterns of amino acid replacement after gene duplication as compared to after speciation.

We have performed a large-scale analysis of amino acid sequence evolution after gene duplication by comparing evolution after gene duplication with evolution after speciation in over 1,800 phylogenetic trees constructed from manually curated alignments of protein domains downloaded from the PFAM database. The site-specific rate of evolution is significantly altered by gene duplication. A significant increase in the proportion of amino acid substitutions at constrained (slowly evolving) sites after duplication was observed. An increase in the proportion of replacements at normally constrained amino acid sites could result from relaxation of purifying selective pressure. However, the proportion of amino acid replacements involving radical changes in amino acid properties after duplication does not appear to be significantly increased by relaxed selective pressure. The increased proportion of replacements at constrained sites was observed over a relatively large range of protein change (up to 25% amino acid replacements per site). These findings have implications for our understanding of the nature of evolution after duplication and may help to shed light on the evolution of novel protein functions through gene duplication.

Algorithms↗

Mammalian housekeeping genes evolve more slowly than tissue-specific genes.

Do housekeeping genes, which are turned on most of the time in almost every tissue, evolve more slowly than genes that are turned on only at specific developmental times or tissues? Recent large-scale gene expression studies enable us to have a better definition of housekeeping genes and to address the above question in detail. In this study, we examined 1581 human-mouse orthologous gene pairs for their patterns of sequence evolution, contrasting housekeeping genes with tissue-specific genes. Our results show that, in comparison to tissue-specific genes, housekeeping genes on average evolve more slowly and are under stronger selective constraints as reflected by significantly smaller values of Ka/Ks. Besides stronger purifying selection, we explored several other factors that can possibly slow down nonsynonymous rates in housekeeping genes. Although mutational bias might slightly slow the nonsynonymous rates in housekeeping genes, it is unlikely to be the major cause of the rate difference between the two types of genes. The codon usage pattern of housekeeping genes does not seem to differ from that of tissue-specific genes. Moreover, contrary to the old textbook concept, we found that approximately 74% of the housekeeping genes in our study belong to multigene families, not significantly different from that of the tissue-specific genes ( approximately 70%). Therefore, the stronger selective constraints on housekeeping genes are not due to a lower degree of genetic redundancy.

Animals↗

Bayesian model adequacy and choice in phylogenetics.

Bayesian inference is becoming a common statistical approach to phylogenetic estimation because, among other reasons, it allows for rapid analysis of large data sets with complex evolutionary models. Conveniently, Bayesian phylogenetic methods use currently available stochastic models of sequence evolution. However, as with other model-based approaches, the results of Bayesian inference are conditional on the assumed model of evolution: inadequate models (models that poorly fit the data) may result in erroneous inferences. In this article, I present a Bayesian phylogenetic method that evaluates the adequacy of evolutionary models using posterior predictive distributions. By evaluating a model's posterior predictive performance, an adequate model can be selected for a Bayesian phylogenetic study. Although I present a single test statistic that assesses the overall (global) performance of a phylogenetic model, a variety of test statistics can be tailored to evaluate specific features (local performance) of evolutionary models to identify sources failure. The method presented here, unlike the likelihood-ratio test and parametric bootstrap, accounts for uncertainty in the phylogeny and model parameters.

Animals↗

Evolutionary distances for protein-coding sequences: modeling site-specific residue frequencies.

Estimation of evolutionary distances from coding sequences must take into account protein-level selection to avoid relative underestimation of longer evolutionary distances. Current modeling of selection via site-to-site rate heterogeneity generally neglects another aspect of selection, namely position-specific amino acid frequencies. These frequencies determine the maximum dissimilarity expected for highly diverged but functionally and structurally conserved sequences, and hence are crucial for estimating long distances. We introduce a codon-level model of coding sequence evolution in which position-specific amino acid frequencies are free parameters. In our implementation, these are estimated from an alignment using methods described previously. We use simulations to demonstrate the importance and feasibility of modeling such behavior; our model produces linear distance estimates over a wide range of distances, while several alternative models underestimate long distances relative to short distances. Site-to-site differences in rates, as well as synonymous/nonsynonymous and first/second/third-codon-position differences, arise as a natural consequence of the site-to-site differences in amino acid frequencies.

Amino Acids↗

Weighted neighbor joining: a likelihood-based approach to distance-based phylogeny reconstruction.

We introduce a distance-based phylogeny reconstruction method called "weighted neighbor joining," or "Weighbor" for short. As in neighbor joining, two taxa are joined in each iteration; however, the Weighbor criterion for choosing a pair of taxa to join takes into account that errors in distance estimates are exponentially larger for longer distances. The criterion embodies a likelihood function on the distances, which are modeled as correlated Gaussian random variables with different means and variances, computed under a probabilistic model for sequence evolution. The Weighbor criterion consists of two terms, an additivity term and a positivity term, that quantify the implications of joining the pair. The first term evaluates deviations from additivity of the implied external branches, while the second term evaluates confidence that the implied internal branch has a positive branch length. Compared with maximum-likelihood phylogeny reconstruction, Weighbor is much faster, while building trees that are qualitatively and quantitatively similar. Weighbor appears to be relatively immune to the "long branches attract" and "long branch distracts" drawbacks observed with neighbor joining, BIONJ, and parsimony.

Animals↗

Parsimony, likelihood, and the role of models in molecular phylogenetics.

Methods such as maximum parsimony (MP) are frequently criticized as being statistically unsound and not being based on any "model." On the other hand, advocates of MP claim that maximum likelihood (ML) has some fundamental problems. Here, we explore the connection between the different versions of MP and ML methods, particularly in light of recent theoretical results. We describe links between the two methods--for example, we describe how MP can be regarded as an ML method when there is no common mechanism between sites (such as might occur with morphological data and certain forms of molecular data). In the process, we clarify certain historical points of disagreement between proponents of the two methodologies, including a discussion of several forms of the ML optimality criterion. We also describe some additional results that shed light on how much needs to be assumed about underlying models of sequence evolution in order to successfully reconstruct evolutionary trees.

Evolution, Molecular↗

The origins of acquired immune deficiency syndrome viruses: where and when?

In the absence of direct epidemiological evidence, molecular evolutionary studies of primate lentiviruses provide the most definitive information about the origins of human immunodeficiency virus (HIV)-1 and HIV-2. Related lentiviruses have been found infecting numerous species of primates in sub-Saharan Africa. The only species naturally infected with viruses closely related to HIV-2 is the sooty mangabey (Cercocebus atys) from western Africa, the region where HIV-2 is known to be endemic. Similarly, the only viruses very closely related to HIV-1 have been isolated from chimpanzees (Pan troglodytes), and in particular those from western equatorial Africa, again coinciding with the region that appears to be the hearth of the HIV-1 pandemic. HIV-1 and HIV-2 have each arisen several times: in the case of HIV-1, the three groups (M, N and O) are the result of independent cross-species transmission events. Consistent with the phylogenetic position of a 'fossil' virus from 1959, molecular clock analyses using realistic models of HIV-1 sequence evolution place the last common ancestor of the M group prior to 1940, and several lines of evidence indicate that the jump from chimpanzees to humans occurred before then. Both the inferred geographical origin of HIV-1 and the timing of the cross-species transmission are inconsistent with the suggestion that oral polio vaccines, putatively contaminated with viruses from chimpanzees in eastern equatorial Africa in the late 1950s, could be responsible for the origin of acquired immune deficiency syndrome.

Africa↗