PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Pyrococcus genome comparison evidences chromosome shuffling-driven evolution.

The genomes of three Pyrococcus species, P.abyssi, P.furiosus and P.horikoshii, were compared at the DNA level, taking advantage of our identification of their replication origins. Three types of rearrangements have been identified: (i) inversion and translation across the replication axis (origin/terminus), (ii) inversion and translocation restricted to a replichore (the half chromosome divided by the replication axis) and (iii) apparent mobility of long clusters of repeated sequences. Rearrangements restricted within a replichore were more common between P.furiosus and the two other Pyrococcus species than between P.horikoshii and P.abyssi. A strong correlation was found between 23 homologous insertion sequence elements, present only in P.furiosus, and recombined segment boundaries, suggesting that transposition events have been a major cause of genomic disruption in this species. Moreover, gene orientation bias was much more disrupted than strand composition biases in fragments that switched their orientation within a replichore upon recombination. This allowed us to conclude that one reversion and one translation occurred in P.abyssi after its divergence from P.horikoshii, and that a smaller segment has specifically recombined in P.furiosus. Whereas a majority of genes are transcribed in the same direction as DNA replication in P.horikoshii and P.abyssi, the colinearity of transcription and replication is only maintained for highly transcribed genes in P.furiosus. We discuss the implications of genomic rearrangements on gene orientation and composition biases, and their consequences on sequence evolution.

Chromosome Inversion↗

Sex-determination gene and pathway evolution in nematodes.

The pathway that controls sexual fate in the nematode Caenorhabditis elegans has been well characterized at the molecular level. By identifying differences between the sex-determination mechanisms in C. elegans and other nematode species, it should be possible to understand how complex sex-determining pathways evolve. Towards this goal, orthologues of many of the C. elegans sex regulators have been isolated from other members of the genus Caenorhabditis. Rapid sequence evolution is observed in every case, but several of the orthologues appear to have conserved sex-determining roles. Thus extensive sequence divergence does not necessarily coincide with changes in pathway structure, although the same forces may contribute to both. This review summarizes recent findings and, with reference to results from other animals, offers explanations for why sex-determining genes and pathways appear to be evolving rapidly. Experimental strategies that hold promise for illuminating pathway differences between nematodes are also discussed.

Animals↗

The standard genetic code enhances adaptive evolution of proteins.

The standard genetic code, by which most organisms translate genetic material into protein metabolism, is non-randomly organized. The Error Minimization hypothesis interprets this non-randomness as an adaptation, proposing that natural selection produced a pattern of codon assignments that buffers genomes against the impact of mutations. Indeed, on the average any given point mutation has a lesser effect on the chemical properties of the utilized amino acid than expected by chance. Might it also, however, be the case that the non-random nature of the code effects the rate of adaptive evolution? To investigate this, here we develop population genetic simulations to test the rate of adaptive gene evolution under different genetic codes. We identify two independent properties of a genetic code that profoundly influence the speed of adaptive evolution. Noting that the standard genetic code exhibits both, we offer a new insight into the effects of the "error minimizing" code: such a code enhances the efficacy of adaptive sequence evolution.

Evolution, Molecular↗

Ross River virus genetic variants in Australia and the Pacific Islands.

HaeIII and TaqI restriction digest profiles of cDNA to infected cell RNA or virion RNA were used as a guide to genetic relationships between fourteen isolates of Ross River virus (RRV) obtained from mosquitoes collected in various localities in eastern Australia where the virus is endemic. RRV isolates from Fiji, American Samoa, the Cook Islands and the Wallis Islands where major outbreaks of epidemic polyarthritis took place in 1979-1980 were also examined. Among these RRV isolates we have identified three genetic types (I-III) on the basis of differences between their restriction digest profiles. We estimate that 1.5-5% nucleotide sequence diversity exists between genetic types. Within each genetic type strain differentiation gave rise to small but significant differences in restriction digest profiles. No clear pattern of geographic distribution of RRV genetic types could be established from the limited number of RRV isolates examined. Genetic types I, II and III, respectively, were isolated from three, three and one different mosquito species, indicating there is no strong association between genetic type and the species of mosquito vector. HaeIII restriction digest analysis did not detect any genetic difference between the four Pacific Island isolates, suggesting that a single RRV variant was involved in the epidemics. Genetically, this variant was closely related to isolates of genetic type II. Virtually identical HaeIII restriction digest profiles were observed for isolates obtained at various stages of the Pacific Island epidemics, suggesting that extensive sequence evolution did not accompany Ross River virus spread.

Alphavirus↗

A low rate of simultaneous double-nucleotide mutations in primates.

The occurrence of double-nucleotide (doublet) mutations is contrary to the normal assumption that point mutations affect single nucleotides. Here we develop a new method for estimating the doublet mutation rate and apply it to more than a megabase of human-chimpanzee-baboon genomic DNA alignments and more than a million human single-nucleotide polymorphisms. The new method accounts for the effect of regional variation in evolutionary rates, which may be a confounding factor in previous estimates of the doublet mutation rate. Furthermore we determine sequence context effects by using sequence comparisons over a variety of lineage lengths. This approach yields a new estimate of the doublet mutation rate of 0.3% of the singleton rate, indicating that doublet mutations are far rarer than previously thought. Our results suggest that doublet mutations are unlikely to have caused the correlation between synonymous and nonsynonymous substitution rates in mammals, and also show that regional variation and sequence context effects play an important role in primate DNA sequence evolution.

Animals↗

Di-, tri-, and tetranucleotide frequencies covary with lifespan and genome size across protostome invertebrates.

Animal lifespans span orders of magnitude, yet how genome sequence covaries with lifespan remains poorly characterized outside vertebrates. Although promoter CpG density has been linked to vertebrate longevity due to its gene-regulatory function through DNA methylation, it is unclear whether such patterns are promoter- and CpG-specific, or if they reflect broader sequence evolution. We curated maximum lifespan estimates for 466 protostome species spanning eight phyla with available genome assemblies and quantified mono-, di-, tri-, and tetranucleotide composition across whole genomes, intergenic regions, and six gene-associated regions (two upstream regions, exons, introns, and two downstream regions) defined using Benchmarking Universal Single-Copy Orthologs. Dinucleotide observed/expected ratios showed significant associations with lifespan and genome size in different ways. Lifespan-associated motifs were most pronounced in gene-associated non-coding regions, especially in introns and downstream regions, whereas genome-size effects were strongest in whole-genome and intergenic sequence. Tri- and tetranucleotide observed/expected ratios broadly recapitulated this regional organization. In contrast, GC content was not associated with lifespan across regions, indicating that the observed signals are not explained by mononucleotide composition but instead by how those nucleotides are arranged into short sequence motifs. These results suggest that lifespan and genome size show distinct but overlapping associations with regional sequence composition across invertebrate species and that lifespan-associated motif evolution extends beyond vertebrate promoter methylation architectures.

CpG density↗

Vestige: maximum likelihood phylogenetic footprinting.

BACKGROUND: Phylogenetic footprinting is the identification of functional regions of DNA by their evolutionary conservation. This is achieved by comparing orthologous regions from multiple species and identifying the DNA regions that have diverged less than neutral DNA. Vestige is a phylogenetic footprinting package built on the PyEvolve toolkit that uses probabilistic molecular evolutionary modelling to represent aspects of sequence evolution, including the conventional divergence measure employed by other footprinting approaches. In addition to measuring the divergence, Vestige allows the expansion of the definition of a phylogenetic footprint to include variation in the distribution of any molecular evolutionary processes. This is achieved by displaying the distribution of model parameters that represent partitions of molecular evolutionary substitutions. Examination of the spatial incidence of these effects across regions of the genome can identify DNA segments that differ in the nature of the evolutionary process. RESULTS: Vestige was applied to a reference dataset of the SCL locus from four species and provided clear identification of the known conserved regions in this dataset. To demonstrate the flexibility to use diverse models of molecular evolution and dissect the nature of the evolutionary process Vestige was used to footprint the Ka/Ks ratio in primate BRCA1 with a codon model of evolution. Two regions of putative adaptive evolution were identified illustrating the ability of Vestige to represent the spatial distribution of distinct molecular evolutionary processes. CONCLUSION: Vestige provides a flexible, open platform for phylogenetic footprinting. Underpinned by the PyEvolve toolkit, Vestige provides a framework for visualising the signatures of evolutionary processes across the genome of numerous organisms simultaneously. By exploiting the maximum-likelihood statistical framework, the complex interplay between mutational processes, DNA repair and selection can be evaluated both spatially (along a sequence alignment) and temporally (for each branch of the tree) providing visual indicators to the attributes and functions of DNA sequences.

Algorithms↗

Manuscript evolution.

Frequently, letters, words and sentences are used in undergraduate textbooks and the popular press as an analogy for the coding, transfer and corruption of information in DNA. We discuss here how the converse can be exploited, by using programs designed for biological analysis of sequence evolution to uncover the relationships between different manuscript versions of a text. We point out similarities between the evolution of DNA and the evolution of texts.

Evolution, Molecular↗

Manuscript evolution.

Frequently, letters, words and sentences are used in undergraduate textbooks and the popular press as an analogy for the coding, transfer and corruption of information in DNA. We discuss here how the converse can be exploited, by using programs designed for biological analysis of sequence evolution to uncover the relationships between different manuscript versions of a text. We point out similarities between the evolution of DNA and the evolution of texts.

DNA↗

Rates of nucleotide substitution and mammalian nuclear gene evolution. Approximate and maximum-likelihood methods lead to different conclusions.

Rates and patterns of synonymous and nonsynonymous substitutions have important implications for the origin and maintenance of mammalian isochores and the effectiveness of selection at synonymous sites. Previous studies of mammalian nuclear genes largely employed approximate methods to estimate rates of nonsynonymous and synonymous substitutions. Because these methods did not account for major features of DNA sequence evolution such as transition/transversion rate bias and unequal codon usage, they might not have produced reliable results. To evaluate the impact of the estimation method, we analyzed a sample of 82 nuclear genes from the mammalian orders Artiodactyla, Primates, and Rodentia using both approximate and maximum-likelihood methods. Maximum-likelihood analysis indicated that synonymous substitution rates were positively correlated with GC content at the third codon positions, but independent of nonsynonymous substitution rates. Approximate methods, however, indicated that synonymous substitution rates were independent of GC content at the third codon positions, but were positively correlated with nonsynonymous rates. Failure to properly account for transition/transversion rate bias and unequal codon usage appears to have caused substantial biases in approximate estimates of substitution rates.

Animals↗

Testing the chromosomal speciation hypothesis for humans and chimpanzees.

Fixed differences of chromosomal rearrangements between isolated populations may promote speciation by preventing between-population gene flow upon secondary contact, either because hybrids suffer from lowered fitness or, more likely, because recombination is reduced in rearranged chromosomal regions. This chromosomal speciation hypothesis thus predicts more rapid genetic divergence on rearranged than on colinear chromosomes because the former are less porous to gene flow. A number of studies of fungi, plants, and animals, including limited genetic data of humans and chimpanzees, support the hypothesis. Here we reexamine the hypothesis for humans and chimpanzees with substantially more genomic data than were used previously. No difference is observed between rearranged and colinear chromosomes in the level of genomic DNA sequence divergence between species. The same is also true for protein sequences. When the gorilla is used as an outgroup, no acceleration in protein sequence evolution associated with chromosomal rearrangements is found. Furthermore, divergence in expression pattern between orthologous genes is not significantly different for rearranged and colinear chromosomes. These results, showing that chromosomal rearrangements did not affect the rate of genetic divergence between humans and chimpanzees, are expected if incipient species on the evolutionary lineages separating humans and chimpanzees did not hybridize.

Animals↗

Restriction fragment polymorphism in the sex-determining region of the Y chromosomal DNA of European wild mice.

Using 32P-labeled probe consisting mainly of (GATA)n we have shown that a male specific Alu1 DNA blot pattern which defines the Y chromosome sex-determining locus in inbred mice is highly polymorphic in wild mice, indicating substantial sequence evolution in this region under field conditions. In all cases examined by in situ hybridization, the region concerned is paracentromeric. In contrast, the blot pattern of another probe (M 34) which detects repeated sequences specific to the mouse Y chromosome but outside the sex-determining locus, remains constant between different isolates.

Animals↗

Mutational analysis of viroid pathogenicity: tomato apical stunt viroid.

A series of nucleotide substitutions within the pathogenicity domain of tomato apical stunt viroid have been evaluated for their effects upon infectivity and symptom expression. None of the 12 A----G substitutions and one C----U substitution that were examined abolished infectivity in a whole plant bioassay, and the resulting progeny were characterized by nucleotide sequence analysis of cDNAs amplified by the polymerase chain reaction. Four of the 13 substitutions gave rise to altered progeny, but the patterns of sequence changes observed were unexpectedly complex. Mutations that did not rapidly revert to the wild-type sequence are located near the right border of the pathogenicity domain, a region which shows considerable natural sequence variability. None had a detectable effect upon symptom expression. The ability to observe viroid sequence evolution in vivo may provide insight into the molecular interactions responsible for viroid host range and symptom formation.

Base Sequence↗

The evolution of insertion sequences within enteric bacteria.

To identify mechanisms that influence the evolution of bacterial transposons, DNA sequence variation was evaluated among homologs of insertion sequences IS1, IS3 and IS30 from natural strains of Escherichia coli and related enteric bacteria. The nucleotide sequences within each class of IS were highly conserved among E. coli strains, over 99.7% similar to a consensus sequence. When compared to the range of nucleotide divergence among chromosomal genes, these data indicate high turnover and rapid movement of the transposons among clonal lineages of E. coli. In addition, length polymorphism among IS appears to be far less frequent than in eukaryotic transposons, indicating that nonfunctional elements comprise a smaller fraction of bacterial transposon populations than found in eukaryotes. IS present in other species of enteric bacteria are substantially divergent from E. coli elements, indicating that IS are mobilized among bacterial species at a reduced rate. However, homologs of IS1 and IS3 from diverse species provide evidence that recombination events and horizontal transfer of IS among species have both played major roles in the evolution of these elements. IS3 elements from E. coli and Shigella show multiple, nested, intragenic recombinations with a distantly related transposon, and IS1 homologs from diverse taxa reveal a mosaic structure indicative of multiple recombination and horizontal transfer events.

Base Sequence↗

Models of amino acid substitution and applications to mitochondrial protein evolution.

Models of amino acid substitution were developed and compared using maximum likelihood. Two kinds of models are considered. "Empirical" models do not explicitly consider factors that shape protein evolution, but attempt to summarize the substitution pattern from large quantities of real data. "Mechanistic" models are formulated at the codon level and separate mutational biases at the nucleotide level from selective constraints at the amino acid level. They account for features of sequence evolution, such as transition-transversion bias and base or codon frequency biases, and make use of physicochemical distances between amino acids to specify nonsynonymous substitution rates. A general approach is presented that transforms a Markov model of codon substitution into a model of amino acid replacement. Protein sequences from the entire mitochondrial genomes of 20 mammalian species were analyzed using different models. The mechanistic models were found to fit the data better than empirical models derived from large databases. Both the mutational distance between amino acids (determined by the genetic code and mutational biases such as the transition-transversion bias) and the physicochemical distance are found to have strong effects on amino acid substitution rates. A significant proportion of amino acid substitutions appeared to have involved more than one codon position, indicating that nucleotide substitutions at neighboring sites may be correlated. Rates of amino acid substitution were found to be highly variable among sites.

Amino Acid Substitution↗

High intron sequence conservation across three mammalian orders suggests functional constraints.

Several studies have demonstrated high levels of sequence conservation in noncoding DNA compared between two species (e.g., human and mouse), and interpreted this conservation as evidence for functional constraints. If this interpretation is correct, it suggests the existence of a hidden class of abundant regulatory elements. However, much of the noncoding sequence conserved between two species may result from chance or from small-scale heterogeneity in mutation rates. Stronger inferences are expected from sequence comparisons using more than two taxa, and by testing for spatial patterns of conservation in addition to primary sequence similarity. We used a Bayesian local alignment method to compare approximately 10 kb of intron sequence from nine genes in a pairwise manner between human, whale, and seal to test whether the degree and pattern of conservation is consistent with neutral divergence. Comparison of the three sets of conserved gapless pairwise blocks revealed the following patterns: The proportion of identical intron nucleotides averaged 47% in pairwise comparisons and 28% across the three taxa. Proportions of conserved sequence were similar in unique sequence and general mammalian repetitive elements. We simulated sequence evolution under a neutral model using published estimates of substitution rate heterogeneity for noncoding DNA and found pairwise identity at 33% and three-taxon identity at 16% of nucleotide sites. Spatial patterns of primary sequence conservation were also nonrandomly distributed within introns. Overall, segments of intron sequence closer to flanking exons were significantly more conserved than interior intron sequence. This level of intron sequence conservation is above that expected by chance and strongly suggests that intron sequences are playing a larger functional role in gene regulation than previously realized.

Animals↗

Functional divergence of duplicated genes formed by polyploidy during Arabidopsis evolution.

To study the evolutionary effects of polyploidy on plant gene functions, we analyzed functional genomics data for a large number of duplicated gene pairs formed by ancient polyploidy events in Arabidopsis thaliana. Genes retained in duplicate are not distributed evenly among Gene Ontology or Munich Information Center for Protein Sequences functional categories, which indicates a nonrandom process of gene loss. Genes involved in signal transduction and transcription have been preferentially retained, and those involved in DNA repair have been preferentially lost. Although the two members of each gene pair must originally have had identical transcription profiles, less than half of the pairs formed by the most recent polyploidy event still retain significantly correlated profiles. We identified several cases where groups of duplicated gene pairs have diverged in concert, forming two parallel networks, each containing one member of each gene pair. In these cases, the expression of each gene is strongly correlated with the other nonhomologous genes in its network but poorly correlated with its paralog in the other network. We also find that the rate of protein sequence evolution has been significantly asymmetric in >20% of duplicate pairs. Together, these results suggest that functional diversification of the surviving duplicated genes is a major feature of the long-term evolution of polyploids.

Arabidopsis↗

Evidence for human immunodeficiency virus type 1 replication in vivo in CD14(+) monocytes and its potential role as a source of virus in patients on highly active antiretroviral therapy.

In vitro studies show that human immunodeficiency virus type 1 (HIV-1) does not replicate in freshly isolated monocytes unless monocytes differentiate to monocyte-derived macrophages. Similarly, HIV-1 may replicate in macrophages in vivo, whereas it is unclear whether blood monocytes are permissive to productive infection with HIV-1. We investigated HIV-1 replication in CD14(+) monocytes and resting and activated CD4(+) T cells by measuring the levels of cell-associated viral DNA and mRNA and the genetic evolution of HIV-1 in seven acutely infected patients whose plasma viremia had been <100 copies/ml for 803 to 1,544 days during highly active antiretroviral therapy (HAART). HIV-1 DNA was detected in CD14(+) monocytes as well as in activated and resting CD4(+) T cells throughout the course of study. While significant variation in the decay slopes of HIV-1 DNA was seen among individual patients, viral decay in CD14(+) monocytes was on average slower than that in activated and resting CD4(+) T cells. Measurements of HIV-1 sequence evolution and the concentrations of unspliced and multiply spliced mRNA provided evidence of ongoing HIV-1 replication, more pronounced in CD14(+) monocytes than in resting CD4(+) T cells. Phylogenetic analyses of HIV-1 sequences indicated that after prolonged HAART, viral populations related or identical to those found only in CD14(+) monocytes were seen in plasma from three of the seven patients. In the other four patients, HIV-1 sequences in plasma and the three cell populations were identical. CD14(+) monocytes appear to be one of the potential in vivo sources of HIV-1 in patients receiving HAART.

Amino Acid Sequence↗