PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Contrasting patterns of nonneutral evolution in proteins encoded in nuclear and mitochondrial genomes.

We report that patterns of nonneutral DNA sequence evolution among published nuclear and mitochondrially encoded protein-coding loci differ significantly in animals. Whereas an apparent excess of amino acid polymorphism is seen in most (25/31) mitochondrial genes, this pattern is seen in fewer than half (15/36) of the nuclear data sets. This differentiation is even greater among data sets with significant departures from neutrality (14/15 vs. 1/6). Using forward simulations, we examined patterns of nonneutral evolution using parameters chosen to mimic the differences between mitochondrial and nuclear genetics (we varied recombination rate, population size, mutation rate, selective dominance, and intensity of germ line bottleneck). Patterns of evolution were correlated only with effective population size and strength of selection, and no single genetic factor explains the empirical contrast in patterns. We further report that in Arabidopsis thaliana, a highly self-fertilizing plant with effectively low recombination, five of six published nuclear data sets also exhibit an excess of amino acid polymorphism. We suggest that the contrast between nuclear and mitochondrial nonneutrality in animals stems from differences in rates of recombination in conjunction with a distribution of selective effects. If the majority of mutations segregating in populations are deleterious, high linkage may hinder the spread of the occasional beneficial mutation.

Animals↗

Evolutionarily different alphoid repeat DNA on homologous chromosomes in human and chimpanzee.

Centromeric alphoid DNA in primates represents a class of evolving repeat DNA. In humans, chromosomes 13 and 21 share one subfamily of alphoid DNA while chromosomes 14 and 22 share another subfamily. We show that similar pairwise homogenizations occur in the chimpanzee (Pan troglodytes), where chromosomes 14 and 22, homologous to human chromosomes 13 and 21, share one partially homogenized alphoid DNA subfamily and chromosomes 15 and 23, homologous to human chromosomes 14 and 22, share another extensively homogenized subfamily. Such a pattern of homogenization presumably predates speciation 3-10 million years ago. However, the alphoid DNA on these human and chimpanzee chromosomes is not orthologous but originates from two evolutionarily different repeat families. It follows that dramatic sequence evolution has occurred in a concerted fashion among the chromosomes in one or both species during or after separation.

Animals↗

A priori estimation of phylogenetic information conserved in aligned sequences.

A new phenomenological approach to explorative data analysis, the estimation of spectra of supporting positions, allows the search for conserved tracks left by phylogeny in DNA sequences. Spectra of supporting positions can be generated without reference to a tree topology or a model of sequence evolution and are therefore an ideal tool for a priori estimation of information content of data sets. Analysis of published 18S rDNA alignments shows that signal to noise relationship varies greatly in a way not detected by conventional tree-construction methods.

Animals↗

Phylogenetic relations of humans and African apes from DNA sequences in the psi eta-globin region.

Sequences from the upstream and downstream flanking DNA regions of the psi eta-globin locus in Pan troglodytes (common chimpanzee), Gorilla gorilla (gorilla), and Pongo pygmaeus (orangutan, the closest living relative to Homo, Pan, and Gorilla) provided further data for evaluating the phylogenetic relations of humans and African apes. These newly sequenced orthologs [an additional 4.9 kilobase pairs (kbp) for each species] were combined with published psi eta-gene sequences and then compared to the same orthologous stretch (a continuous 7.1-kbp region) available for humans. Phylogenetic analysis of these nucleotide sequences by the parsimony method indicated (i) that human and chimpanzee are more closely related to each other than either is to gorilla and (ii) that the slowdown in the rate of sequence evolution evident in higher primates is especially pronounced in humans. These results indicate that features (for example, knuckle-walking) unique to African apes (but not to humans) are primitive and that even local molecular clocks should be applied with caution.

Animals↗

Molecular epidemiology of rabies virus in France: comparison with vaccine strains.

A molecular epidemiological study of the rabies virus currently prevalent in France was carried out by directly sequencing polymerase chain reaction-amplified genes. The rabies virus pseudogene psi was chosen as the most divergent genomic area, and as such the best 'clock' for measuring virus evolution. Sequence comparisons between 12 wild rabies virus isolates indicated strong conservation whatever the host and wherever the virus had been isolated. This holds true for a unique wild reservoir, the fox. On the other hand, a good correlation between genetic and geographical criteria indicates a slow evolution of the wild virus in parallel with the spatio-temporal progression of the epizootic. In contrast to their intrinsic homogeneity (about 2% divergence), the wild isolate sequences showed a marked divergence from those of vaccine seed strains (about 14.7%). This finding invites world-wide molecular epidemiological studies, particularly in countries in which vaccination failures have been reported.

Animals↗

Molecular footprints of human immunoglobulin gene evolution: a new sequence family.

Analysis of the human VK (ref. 2) gene locus led to the detection of a new sequence family (L sequences). Its copy number is in the range of 10(2). The L sequences, which are about 500 bp long, are found as part of the 3' flanking regions of a clustered set of human VKI genes but they occur also separate from the genes. Models are discussed in which L sequences are viewed as molecular footprints of amplification and transposition processes of VK genes.

Amino Acid Sequence↗

The evolution of DNA sequences in Escherichia coli.

It is proposed that certain families of transposable elements originally evolved in plasmids and functioned in forming replicon fusions to aid in the horizontal transmission of non-conjugational plasmids. This hypothesis is supported by the finding that the transposable elements Tn3 and gamma delta are found almost exclusively in plasmids, and also by the distribution of the unrelated insertion sequences IS4 and IS5 among a reference collection of 67 natural isolates of Escherichia coli. Each insertion sequence was found to be present in only about one-third of the strains. Among the ten strains found to contain both insertion sequences, the number of copies of the elements was negatively correlated. With respect to IS5, approximately half of the strains containing a chromosomal copy of the insertion element also contained copies within the plasmid complement of the strain.

Base Sequence↗

Sequence variation and evolution of nuclear DNA in man and the primates.

Recent advances in nucleic acid technology have facilitated the detection and detailed structural analysis of a wide variety of genes in higher organisms, including those in man. This in turn has opened the way to an examination of the evolution of structural genes and their surrounding and intervening sequences. In a study of the evolution of haemoglobin genes and neighbouring sequences in man and the primates, we have investigated gene arrangement and DNA sequence divergence both within and between species ranging from Old World monkeys to man. This analysis is beginning to reveal the evolutionary constraints that have acted on this region of the genome during primate evolution. Furthermore, DNA sequence variation, both within and between species, provides, in principle, a novel and powerful method for determining interspecific phylogenetic distances and also for analysing the structure of present-day human populations. Application of this new branch of molecular biology to other areas of the human genome should prove important in unravelling the history of genetic changes that have occurred during the evolution of man.

Animals↗

Substitution bias, rapid saturation, and the use of mtDNA for nematode systematics.

Only relatively recently have researchers turned to molecular methods for nematode phylogeny reconstruction. Thus, we lack the extensive literature on evolutionary patterns and phylogenetic usefulness of different DNA regions for nematodes that exists for other taxa. Here, we examine the usefulness of mtDNA for nematode phylogeny reconstruction and provide data that can be used for a priori character weighting or for parameter specification in models of sequence evolution. We estimated the substitution pattern for the mitochondrial ND4 gene from intraspecific comparisons in four species of parasitic nematodes from the family Trichostrongylidae (38-50 sequences per species). The resulting pattern suggests a strong mutational bias toward A and T, and a lower transition/transversion ratio than is typically observed in other taxa. We also present information on the relative rates of substitution at first, second, and third codon positions and on relative rates of saturation of different types of substitutions in comparisons ranging from intraspecific to interordinal. Silent sites saturate extremely quickly, presumably owing to the substitution bias and, perhaps, to an accelerated mutation rate. Results emphasize the importance of using only the most closely related sequences in order to infer patterns of substitution accurately for nematodes or for other taxa having strongly composition-biased DNA. ND4 also shows high amino acid polymorphism at both the intra- and interspecific levels, and in higher level comparisons, there is evidence of saturation at variable amino acid sites. In general, we recommend using mtDNA coding genes only for phylogenetics of relatively closely related nematode species and, even then, using only nonsynonymous substitutions and the more conserved mitochondrial genes (e.g., cytochrome oxidases). On the other hand, the high substitution rate in genes such as ND4 should make them excellent for population genetics studies, identifying cryptic species, and resolving relationships among closely related congeners when other markers show insufficient variation.

Amino Acid Sequence↗

Evolution of the secondary structures and compensatory mutations of the ribosomal RNAs of Drosophila melanogaster.

This paper examines the effects of DNA sequence evolution on RNA secondary structures and compensatory mutations. Models of the secondary structures of Drosophila melanogaster 18S ribosomal RNA (rRNA) and of the complex between 2S, 5.8S, and 28S rRNAs have been drawn on the basis of comparative and energetic criteria. The overall AU richness of the D. melanogaster rRNAs allows the resolution of some ambiguities in the structures of both large rRNAs. Comparison of the sequence of expansion segment V2 in D. melanogaster 18S rRNA with the same region in three other Drosophila species and the tsetse fly (Glossina morsitans morsitans) allows us to distinguish between two models for the secondary structure of this region. The secondary structures of the expansion segments of D. melanogaster 28S rRNA conform to a general pattern for all eukaryotes, despite having highly divergent sequences between D. melanogaster and vertebrates. The 70 novel compensatory mutations identified in the 28S rRNA show a strong (70%) bias toward A-U base pairs, suggesting that a process of biased mutation and/or biased fixation of A and T point mutations or AT-rich slippage-generated motifs has occurred during the evolution of D. melanogaster rDNA. This process has not occurred throughout the D. melanogaster genome. The processes by which compensatory pairs of mutations are generated and spread are discussed, and a model is suggested by which a second mutation is more likely to occur in a unit with a first mutation as such a unit begins to spread through the family and concomitantly through the population. Alternatively, mechanisms of proofreading in stem-loop structures at the DNA level, or between RNA and DNA, might be involved. The apparent tolerance of noncompensatory mutations in some stems which are otherwise strongly supported by comparative criteria within D. melanogaster 28S rRNA must be borne in mind when compensatory mutations are used as a criterion in secondary-structure modeling. Noncompensatory mutation may extend to the production of unstable structures where a stem is stabilized by RNA-protein or additional RNA-RNA interactions in the mature ribosome. Of motifs suggested to be involved in rRNA processing, one (CGAAAG) is strongly overrepresented in the 28S rRNA sequence. The data are discussed both in the context of the forces involved with the evolution of multigene families and in the context of molecular coevolution in the rDNA family in particular.

Animals↗

Characterization of the boundaries between adjacent rapidly and slowly evolving genomic regions in Drosophila.

The site of a dramatic change in the rate of DNA sequence evolution exists near the 68C glue gene clusters of several Drosophila species. We have previously determined the approximate location of this transition site by comparison of restriction maps of the regions flanking the 68C-like glue gene cluster of five members of the melanogaster species subgroup. In the present work we report the sequence of the transition region in three of these Drosophila species: D. melanogaster, D. yakuba, and D. erecta. Using a best-fit alignment of these sequences, we find that the site of transition from slowly to rapidly evolving sequences occurs abruptly within a region less than 50 nucleotides in length. Although frequency of nucleotide substitutions changes as much as 10-fold across this boundary, frequency of small insertion/deletion events stays nearly constant.

Animals↗

Phylogenetically enhanced statistical tools for RNA structure prediction.

MOTIVATION: Methods that predict the structure of molecules by looking for statistical correlation have been quite effective. Unfortunately, these methods often disregard phylogenetic information in the sequences they analyze. Here, we present a number of statistics for RNA molecular-structure prediction. Besides common pair-wise comparisons, we consider a few reasonable statistics for base-triple predictions, and present an elaborate analysis of these methods. All these statistics incorporate phylogenetic relationships of the sequences in the analysis to varying degrees, and the different nature of these tests gives a wide choice of statistical tools for RNA structure prediction. RESULTS: Starting from statistics that incorporate phylogenetic information only as independent sequence evolution models for each position of a multiple alignment, and extending this idea to a joint evolution model of two positions, we enhance the usual purely statistical methods (e.g. methods based on the Mutual Information statistic) with the use of phylogenetic information available in the sequences. In particular, we present a joint model based on the HKY evolution model, and consequently a X(2) test of independence for two positions. A significant part of this work is devoted to some mathematical analysis of these methods. We tested these statistics on regions of 16S and 23S rRNA, and tRNA.

Base Sequence↗

Deep Sequencing Reveals Dual Evolution of SARS-CoV-2: Insights Into Defective Genomes From Wuhan-Hu-1 Variants to Omicron Subvariants.

SARS-CoV-2 has evolved from early variants dominating the first (B.1.5, B.1.1) and second (B.1.177) pandemic waves, which exhibited a higher frequency of minority mutants with deletions leading to Defective Viral Genomes (DVGs) in the spike region near the S1/S2 cleavage site than the Alpha, Beta, and Delta variants. The emergence of Omicron has significantly altered the dominant variant profile, with Omicron subvariants now representing 100% of circulating viruses. To monitor the evolution and adaptation of Omicron in the human population, a deep-sequencing study was performed in RNA samples of BA.1, BA.1.1, BA.2, BA.5, BQ.1.1, XBB.1.5 and BA.2.86 Omicron subvariants. The findings reveal two occurrences of similar evolutionary patterns within SARS-CoV-2 characterized by a shift from a significant to a very low production of DVGs. This event suggests that DVGs might play a role in the virus's spread and adaptation for persistence in infected humans.

SARS-CoV-2↗

Estimation of evolutionary distances between nucleotide sequences.

A formal mathematical analysis of the substitution process in nucleotide sequence evolution was done in terms of the Markov process. By using matrix algebra theory, the theoretical foundation of Barry and Hartigan's (Stat. Sci. 2:191-210, 1987) and Lanave et al.'s (J. Mol. Evol. 20:86-93, 1984) methods was provided. Extensive computer simulation was used to compare the accuracy and effectiveness of various methods for estimating the evolutionary distance between two nucleotide sequences. It was shown that the multiparameter methods of Lanave et al.'s (J. Mol. Evol. 20:86-93, 1984), Gojobori et al.'s (J. Mol. Evol. 18:414-422, 1982), and Barry and Hartigan's (Stat. Sci. 2:191-210, 1987) are preferable to others for the purpose of phylogenetic analysis when the sequences are long. However, when sequences are short and the evolutionary distance is large, Tajima and Nei's (Mol. Biol. Evol. 1:269-285, 1984) method is superior to others.

Base Sequence↗

Detection of convergent and parallel evolution at the amino acid sequence level.

Adaptive evolution at the molecular level can be studied by detecting convergent and parallel evolution at the amino acid sequence level. For a set of homologous protein sequences, the ancestral amino acids at all interior nodes of the phylogenetic tree of the proteins can be statistically inferred. The amino acid sites that have experienced convergent or parallel changes on independent evolutionary lineages can then be identified by comparing the amino acids at the beginning and end of each lineage. At present, the efficiency of the methods of ancestral sequence inference in identifying convergent and parallel changes is unknown. More seriously, when we identify convergent or parallel changes, it is unclear whether these changes are attributable to random chance. For these reasons, claims of convergent and parallel evolution at the amino acid sequence level have been disputed. We have conducted computer simulations to assess the efficiencies, of the parsimony and Bayesian methods of ancestral sequence inference in identifying convergent and parallel-change sites. Our results showed that the Bayesian method performs better than the parsimony method in identifying parallel changes, and both methods are inefficient in identifying convergent changes. However, the Bayesian method is recommended for estimating the number of convergent-change sites because it gives a conservative estimate. We have developed statistical tests for examining whether the observed numbers of convergent and parallel changes are due to random chance. As an example, we reanalyzed the stomach lysozyme sequences of foregut fermenters and found that parallel evolution is statistically significant, whereas convergent evolution is not well supported.

Amino Acid Sequence↗