PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Sequence variation and evolution of nuclear DNA in man and the primates.

Recent advances in nucleic acid technology have facilitated the detection and detailed structural analysis of a wide variety of genes in higher organisms, including those in man. This in turn has opened the way to an examination of the evolution of structural genes and their surrounding and intervening sequences. In a study of the evolution of haemoglobin genes and neighbouring sequences in man and the primates, we have investigated gene arrangement and DNA sequence divergence both within and between species ranging from Old World monkeys to man. This analysis is beginning to reveal the evolutionary constraints that have acted on this region of the genome during primate evolution. Furthermore, DNA sequence variation, both within and between species, provides, in principle, a novel and powerful method for determining interspecific phylogenetic distances and also for analysing the structure of present-day human populations. Application of this new branch of molecular biology to other areas of the human genome should prove important in unravelling the history of genetic changes that have occurred during the evolution of man.

Animals↗

Substitution bias, rapid saturation, and the use of mtDNA for nematode systematics.

Only relatively recently have researchers turned to molecular methods for nematode phylogeny reconstruction. Thus, we lack the extensive literature on evolutionary patterns and phylogenetic usefulness of different DNA regions for nematodes that exists for other taxa. Here, we examine the usefulness of mtDNA for nematode phylogeny reconstruction and provide data that can be used for a priori character weighting or for parameter specification in models of sequence evolution. We estimated the substitution pattern for the mitochondrial ND4 gene from intraspecific comparisons in four species of parasitic nematodes from the family Trichostrongylidae (38-50 sequences per species). The resulting pattern suggests a strong mutational bias toward A and T, and a lower transition/transversion ratio than is typically observed in other taxa. We also present information on the relative rates of substitution at first, second, and third codon positions and on relative rates of saturation of different types of substitutions in comparisons ranging from intraspecific to interordinal. Silent sites saturate extremely quickly, presumably owing to the substitution bias and, perhaps, to an accelerated mutation rate. Results emphasize the importance of using only the most closely related sequences in order to infer patterns of substitution accurately for nematodes or for other taxa having strongly composition-biased DNA. ND4 also shows high amino acid polymorphism at both the intra- and interspecific levels, and in higher level comparisons, there is evidence of saturation at variable amino acid sites. In general, we recommend using mtDNA coding genes only for phylogenetics of relatively closely related nematode species and, even then, using only nonsynonymous substitutions and the more conserved mitochondrial genes (e.g., cytochrome oxidases). On the other hand, the high substitution rate in genes such as ND4 should make them excellent for population genetics studies, identifying cryptic species, and resolving relationships among closely related congeners when other markers show insufficient variation.

Amino Acid Sequence↗

Evolution of the secondary structures and compensatory mutations of the ribosomal RNAs of Drosophila melanogaster.

This paper examines the effects of DNA sequence evolution on RNA secondary structures and compensatory mutations. Models of the secondary structures of Drosophila melanogaster 18S ribosomal RNA (rRNA) and of the complex between 2S, 5.8S, and 28S rRNAs have been drawn on the basis of comparative and energetic criteria. The overall AU richness of the D. melanogaster rRNAs allows the resolution of some ambiguities in the structures of both large rRNAs. Comparison of the sequence of expansion segment V2 in D. melanogaster 18S rRNA with the same region in three other Drosophila species and the tsetse fly (Glossina morsitans morsitans) allows us to distinguish between two models for the secondary structure of this region. The secondary structures of the expansion segments of D. melanogaster 28S rRNA conform to a general pattern for all eukaryotes, despite having highly divergent sequences between D. melanogaster and vertebrates. The 70 novel compensatory mutations identified in the 28S rRNA show a strong (70%) bias toward A-U base pairs, suggesting that a process of biased mutation and/or biased fixation of A and T point mutations or AT-rich slippage-generated motifs has occurred during the evolution of D. melanogaster rDNA. This process has not occurred throughout the D. melanogaster genome. The processes by which compensatory pairs of mutations are generated and spread are discussed, and a model is suggested by which a second mutation is more likely to occur in a unit with a first mutation as such a unit begins to spread through the family and concomitantly through the population. Alternatively, mechanisms of proofreading in stem-loop structures at the DNA level, or between RNA and DNA, might be involved. The apparent tolerance of noncompensatory mutations in some stems which are otherwise strongly supported by comparative criteria within D. melanogaster 28S rRNA must be borne in mind when compensatory mutations are used as a criterion in secondary-structure modeling. Noncompensatory mutation may extend to the production of unstable structures where a stem is stabilized by RNA-protein or additional RNA-RNA interactions in the mature ribosome. Of motifs suggested to be involved in rRNA processing, one (CGAAAG) is strongly overrepresented in the 28S rRNA sequence. The data are discussed both in the context of the forces involved with the evolution of multigene families and in the context of molecular coevolution in the rDNA family in particular.

Animals↗

Characterization of the boundaries between adjacent rapidly and slowly evolving genomic regions in Drosophila.

The site of a dramatic change in the rate of DNA sequence evolution exists near the 68C glue gene clusters of several Drosophila species. We have previously determined the approximate location of this transition site by comparison of restriction maps of the regions flanking the 68C-like glue gene cluster of five members of the melanogaster species subgroup. In the present work we report the sequence of the transition region in three of these Drosophila species: D. melanogaster, D. yakuba, and D. erecta. Using a best-fit alignment of these sequences, we find that the site of transition from slowly to rapidly evolving sequences occurs abruptly within a region less than 50 nucleotides in length. Although frequency of nucleotide substitutions changes as much as 10-fold across this boundary, frequency of small insertion/deletion events stays nearly constant.

Animals↗

Phylogenetically enhanced statistical tools for RNA structure prediction.

MOTIVATION: Methods that predict the structure of molecules by looking for statistical correlation have been quite effective. Unfortunately, these methods often disregard phylogenetic information in the sequences they analyze. Here, we present a number of statistics for RNA molecular-structure prediction. Besides common pair-wise comparisons, we consider a few reasonable statistics for base-triple predictions, and present an elaborate analysis of these methods. All these statistics incorporate phylogenetic relationships of the sequences in the analysis to varying degrees, and the different nature of these tests gives a wide choice of statistical tools for RNA structure prediction. RESULTS: Starting from statistics that incorporate phylogenetic information only as independent sequence evolution models for each position of a multiple alignment, and extending this idea to a joint evolution model of two positions, we enhance the usual purely statistical methods (e.g. methods based on the Mutual Information statistic) with the use of phylogenetic information available in the sequences. In particular, we present a joint model based on the HKY evolution model, and consequently a X(2) test of independence for two positions. A significant part of this work is devoted to some mathematical analysis of these methods. We tested these statistics on regions of 16S and 23S rRNA, and tRNA.

Base Sequence↗

Deep Sequencing Reveals Dual Evolution of SARS-CoV-2: Insights Into Defective Genomes From Wuhan-Hu-1 Variants to Omicron Subvariants.

SARS-CoV-2 has evolved from early variants dominating the first (B.1.5, B.1.1) and second (B.1.177) pandemic waves, which exhibited a higher frequency of minority mutants with deletions leading to Defective Viral Genomes (DVGs) in the spike region near the S1/S2 cleavage site than the Alpha, Beta, and Delta variants. The emergence of Omicron has significantly altered the dominant variant profile, with Omicron subvariants now representing 100% of circulating viruses. To monitor the evolution and adaptation of Omicron in the human population, a deep-sequencing study was performed in RNA samples of BA.1, BA.1.1, BA.2, BA.5, BQ.1.1, XBB.1.5 and BA.2.86 Omicron subvariants. The findings reveal two occurrences of similar evolutionary patterns within SARS-CoV-2 characterized by a shift from a significant to a very low production of DVGs. This event suggests that DVGs might play a role in the virus's spread and adaptation for persistence in infected humans.

SARS-CoV-2↗

Estimation of evolutionary distances between nucleotide sequences.

A formal mathematical analysis of the substitution process in nucleotide sequence evolution was done in terms of the Markov process. By using matrix algebra theory, the theoretical foundation of Barry and Hartigan's (Stat. Sci. 2:191-210, 1987) and Lanave et al.'s (J. Mol. Evol. 20:86-93, 1984) methods was provided. Extensive computer simulation was used to compare the accuracy and effectiveness of various methods for estimating the evolutionary distance between two nucleotide sequences. It was shown that the multiparameter methods of Lanave et al.'s (J. Mol. Evol. 20:86-93, 1984), Gojobori et al.'s (J. Mol. Evol. 18:414-422, 1982), and Barry and Hartigan's (Stat. Sci. 2:191-210, 1987) are preferable to others for the purpose of phylogenetic analysis when the sequences are long. However, when sequences are short and the evolutionary distance is large, Tajima and Nei's (Mol. Biol. Evol. 1:269-285, 1984) method is superior to others.

Base Sequence↗

Detection of convergent and parallel evolution at the amino acid sequence level.

Adaptive evolution at the molecular level can be studied by detecting convergent and parallel evolution at the amino acid sequence level. For a set of homologous protein sequences, the ancestral amino acids at all interior nodes of the phylogenetic tree of the proteins can be statistically inferred. The amino acid sites that have experienced convergent or parallel changes on independent evolutionary lineages can then be identified by comparing the amino acids at the beginning and end of each lineage. At present, the efficiency of the methods of ancestral sequence inference in identifying convergent and parallel changes is unknown. More seriously, when we identify convergent or parallel changes, it is unclear whether these changes are attributable to random chance. For these reasons, claims of convergent and parallel evolution at the amino acid sequence level have been disputed. We have conducted computer simulations to assess the efficiencies, of the parsimony and Bayesian methods of ancestral sequence inference in identifying convergent and parallel-change sites. Our results showed that the Bayesian method performs better than the parsimony method in identifying parallel changes, and both methods are inefficient in identifying convergent changes. However, the Bayesian method is recommended for estimating the number of convergent-change sites because it gives a conservative estimate. We have developed statistical tests for examining whether the observed numbers of convergent and parallel changes are due to random chance. As an example, we reanalyzed the stomach lysozyme sequences of foregut fermenters and found that parallel evolution is statistically significant, whereas convergent evolution is not well supported.

Amino Acid Sequence↗

Phylogenetic relationships of the liverworts (Hepaticae), a basal embryophyte lineage, inferred from nucleotide sequence data of the chloroplast gene rbcL.

Sequence data from the chloroplast-encoded gene rbcL were obtained for 24 liverworts, a basal group of embryophytes. Maximum likelihood and parsimony analyses of these data, along with data from other major green plant lineages, confirm hypotheses based on morphological data, such as the paraphyly of bryophytes, and the basal position of liverworts. Molecular data corroborate the deep separation between the complex thalloid and leafy/simple thalloid liverworts implied by morphological data, but the monophyly of liverworts could not be rejected. The effects of accounting for site-to-site rate heterogeneity in these data were examined using maximum likelihood methods. Comparison of trees obtained with and without rate heterogeneity showed that simply allowing for heterogeneity had a greater improvement on likelihood score than optimization of transition/transversion bias. Incorporation of site-to-site rate heterogeneity in the larger analysis, however, did not necessarily change which topology was favored. Properties of rbcL sequences from the two liverwort groups were compared. Significantly different substitution rates were found between leafy/simple thalloid and complex thalloid liverwort taxa, with rates of rbcL sequence evolution in leafy/simple thalloid taxa being higher and more indicative of those of vascular plants, and with those of complex thalloid taxa (such as Marchantia) being slower. Codon usage in rbcL in complex thalloid liverworts was biased toward NNU and NNA, compared to the leafy/simple thalloid liverworts. Although base composition and relative substitution rates differed between the two groups, no significant differences were detected within each of the two groups of liverworts. The signal present in first and second codon sites versus third codon sites was compared. While the third codon positions in rbcL across this taxon sampling are highly variable (with only 15 constant sites of 439), the trees obtained were in general agreement with trees from the entire data set and with trees obtained from independent sources of data. The presence of signal in third codon positions across greater than 400 MY of plant evolution means that definitions of saturation based on pair-wise comparisons of sequences inadequately assess phylogenetic signal.

Chloroplasts↗

Mouse IgA heavy chain gene sequence: implications for evolution of immunoglobulin hinge axons.

The complete nucleotide sequence of the gene and mRNA coding for the constant (C) region of the secreted form of the BALB/c mouse IgA immunoglobulin alpha heavy (H) chain has been determined. As in other immunoglobulins, the three C region domains of the alpha protein, C alpha 1, C alpha 2, and C alpha 3 are coded in separate exons. However, the hinge region of C alpha is not coded on a separate exon as it is in other hinge-containing immunoglobulins. Instead, the alpha hinge is coded as a 5' extension of the C alpha 2 exon, and we suggest that it may have evolved by duplication leading to incorporation of an acceptor RNA splice site into the coding portion of the C alpha 2 exon. Extensions of this concept could provide an explanation for duplications in the human alpha 1 chain.

Amino Acid Sequence↗

Precise sequence assignment of replication origin in the control region of chick mitochondrial DNA relative to 5' and 3' D-loop ends, secondary structure, DNA synthesis, and protein binding.

The data reported identify for the first time the sequence of an avian mitochondrial heavy-strand replication origin, OH, located only about 12 nucleotides (nt) downstream from the conserved sequence block CSB-1, as well as the sequence of premature synthesis arrest of the 781 (+/-1) nt D-loop strand, only 6-7 nt downstream from a TAS-like (termination-associated) element. Both sites are associated with putative cruciform secondary structures. A major sequence-specific DNA-binding/cleavage site of a potential regulatory protein, the approximately 36-kDa aMDP1 (shown previously to stimulate mtDNA synthesis), is located about 90 nt upstream of OH. Correlated in vivo analysis of avian genome-length mtDNA replication provides missing evidence on the functional equivalence of D-loop origin with nascent initiation, and on the direction, asymmetry and temporal aspects of a full round of replication. The importance of the results to understanding the regulation of linked replication/transcription and the unusual sequence evolution of avian mtDNA is

Animals↗

Kinetoplast DNA minicircles: regions of extensive sequence divergence.

Previous work has shown that the kinetoplast minicircle DNA of Leishmania species exhibits species-specific sequence divergence and this observation has led to the development of a DNA probe-based diagnostic test for leishmaniasis. In the work reported here, we demonstrate that the minicircle is composed of three types of DNA sequences with differing specificities reflecting different rates of DNA sequence change. A library of cloned fragments of kinetoplast DNA (kDNA) from Leishmania mexicana amazonensis was prepared and the cloned subfragments were found to contain DNA sequences with different taxonomic specificities based on hybridization analysis with various species of Leishmania. Four groups of subfragments were found, those that hybridized with a large number of Leishmania sp. as well as sequences unique to the species, subspecies, or isolate. Analysis of nested deletions of a single, full-length minicircle demonstrates that these different taxonomic specificities are contained within a single minicircle. This implies that different regions of a single minicircle have DNA sequences that diverge at different rates. These sequences represent potentially valuable tools in diagnostic, epidemiologic, and ecological studies of leishmaniasis and provide the basis for a model of kDNA sequence evolution.

Animals↗

Analytical expression of the purine/pyrimidine autocorrelation function after and before random mutations.

The mutation process is a classical evolutionary genetic process. The type of mutations studied here is the random substitutions of a purine base R (adenine or guanine) by a pyrimidine base Y (cytosine or thymine) and reciprocally (transversions). The analytical expressions derived allow us to analyze in genes the occurrence probabilities of motifs and d-motifs (two motifs separated by any d bases) on the R/Y alphabet under transversions. These motif probabilities can be obtained after transversions (in the evolutionary sense; from the past to the present) and, unexpectedly, also before transversions (after back transversions, in the inverse evolutionary sense, from the present to the past). This theoretical part in Section 2 is a first generalization of a particular formula recently derived. The application in Section 3 is based on the analytical expression giving the autocorrelation function (the d-motif probabilities) before transversions. It allows us to study primitive genes from actual genes. This approach solves a biological problem. The protein coding genes of chloroplasts and mitochondria have a preferential occurrence of the 6-motif YRY(N)6YRY (maximum of the autocorrelation function for d = 6, N = R or Y) with a periodicity modulo 3. The YRY(N)6YRY preferential occurrence without the periodicity modulo 3 is also observed in the RNA coding genes (ribosomal, transfer, and small nuclear RNA genes) and in the noncoding genes (introns and 5' regions of eukaryotic nuclei). However, there are two exceptions to this YRY(N)6YRY rule: the protein coding genes of eukaryotic nuclei, and prokaryotes, where YRY(N)6YRY has the second highest value after YRY(N)0YRY (YRYYRY) with a periodicity modulo 3. When we go backward in time with the analytical expression, the protein coding genes of both eukaryotic nuclei and prokaryotes retrieve the YRY(N)6YRY preferential occurrence with a periodicity modulo 3 after 0.2 back transversions per base. In other words, the actual protein coding genes of chloroplasts and mitochondria are similar to the primitive protein coding genes of eukaryotic nuclei and prokaryotes. On the other hand, this application represents the first result concerning the mutation process in the model of DNA sequence evolution we recently proposed. According to this model, the actual genes on the R/Y alphabet derive from two successive evolutionary genetic processes: an independent mixing of a few nonrandom types of oligonucleotides leading to genes called primitive followed by a mutation process in these primitive genes.(ABSTRACT TRUNCATED AT 400 WORDS)

Base Sequence↗

Inference of population history using a likelihood approach.

We introduce an approach to revealing the likelihood of different population histories that utilizes an explicit model of sequence evolution for the DNA segment under study. Based on a phylogenetic tree reconstruction method we show that a Tamura-Nei model with heterogeneous mutation rates is a fair description of the evolutionary process of the hypervariable region I of the mitochondrial DNA from humans. Assuming this complex model still allows the estimation of population history parameters, we suggest a likelihood approach to conducting statistical inference within a class of expansion models. More precisely, the likelihood of the data is based on the mean pairwise differences between DNA sequences and the number of variable sites in a sample. The use of likelihood ratios enables comparison of different hypotheses about population history, such as constant population size during the past or an increase or decrease of population size starting at some point back in time. This method was applied to show that the population of the Basques has expanded, whereas that of the Biaka pygmies is most likely decreasing. The Nuu-Chah-Nulth data are consistent with a model of constant population.

Base Composition↗

Molecular evolution of the wingless gene and its implications for the phylogenetic placement of the butterfly family Riodinidae (Lepidoptera: papilionoidea).

The sequence evolution of the nuclear gene wingless was investigated among 34 representatives of three lepidopteran families (Riodinidae, Lycaenidae, and Nymphalidae) and four outgroups, and its utility for inferring phylogenetic relationships among these taxa was assessed. Parsimony analysis yielded a well-resolved topology supporting the monophyly of the Riodinidae and Lycaenidae, respectively, and indicating that these two groups are sister lineages, with strong nodal support based on bootstrap and decay indices. Although wingless provides robust support for relationships within and between the riodinids and the lycaenids, it is less informative about nymphalid relationships. Wingless does not consistently recover nymphalid monophyly or traditional subfamilial relationships within the nymphalids, and nodal support for all but the most recent branches in this family is low. Much of the phylogenetic information in this data set is derived from first- and second-position substitutions. However, third positions, despite showing uncorrected pairwise divergences up to 78%, also contain consistent signal at deep nodes within the family Riodinidae and at the node defining the sister relationship between the riodinids and lycaenids. Several hypotheses about how third-position signal has been retained in deep nodes are discussed. These include among-site rate variation, identified as a significant factor by maximum likelihood analyses, and nucleotide bias, a prominent feature of third positions in this data set. Understanding the mechanisms which underlie third-position signal is a first step in applying appropriate models to accommodate the specific evolutionary processes involved in each lineage.

Animals↗

Statistical tests of models of DNA substitution.

Penny et al. have written that "The most fundamental criterion for a scientific method is that the data must, in principle, be able to reject the model. Hardly any [phylogenetic] tree-reconstruction methods meet this simple requirement." The ability to reject models is of such great importance because the results of all phylogenetic analyses depend on their underlying models--to have confidence in the inferences, it is necessary to have confidence in the models. In this paper, a test statistic suggested by Cox is employed to test the adequacy of some statistical models of DNA sequence evolution used in the phylogenetic inference method introduced by Felsenstein. Monte Carlo simulations are used to assess significance levels. The resulting statistical tests provide an objective and very general assessment of all the components of a DNA substitution model; more specific versions of the test are devised to test individual components of a model. In all cases, the new analyses have the additional advantage that values of phylogenetic parameters do not have to be assumed in order to perform the tests.

Animals↗

Evolutionary variants of the human immunodeficiency virus type 1 V3 region characterized by using a heteroduplex tracking assay.

Syncytium-inducing (SI) variants of human immunodeficiency virus type 1 (HIV-1) are evolutionary variants that are associated with rapid CD4+ cell loss and rapid disease progression. The heteroduplex tracking assay (HTA) was used to detect evolutionary V3 variants by amplifying the V3 sequences from viral RNA derived from 50 samples of patient plasma. For this V3-specific HTA (V3-HTA), heteroduplexes were formed between the patient V3 sequences and a probe with the subtype B consensus V3 sequence. Evolution was then measured by divergence from the consensus. The presence of evolutionary variants was correlated with SI detection data on the same samples from the MT-2 cell culture assay. Evolutionary variants were correlated with the SI phenotype in 88% of the samples, and 96% of the SI samples contained evolutionary variants. In most cases the evolutionary V3 variants represented discrete clonal outgrowths of virus. Sequence analysis of the six discordant samples that did not show this correlation indicated that three non-syncytium-inducing (NSI) samples had V3 sequences that had evolved away from the consensus sequence but not toward an SI genotype. A fourth sample showed little evolution away from the consensus but was SI, which indicates that not all SI variants require basic substitutions in V3. The other two samples had SI-like genotypes and NSI phenotypes, suggesting that V3-HTA was able to detect SI emergence in these samples in the absence of their detection in vitro. V3-HTA was also used to confirm SI variant selection in MT-2 cells and to examine the possibility of variant selection during virus culture in peripheral blood cells.

Acquired Immunodeficiency Syndrome↗

PASSML: combining evolutionary inference and protein secondary structure prediction.

MOTIVATION: Evolutionary models of amino acid sequences can be adapted to incorporate structure information; protein structure biologists can use phylogenetic relationships among species to improve prediction accuracy. Results : A computer program called PASSML ('Phylogeny and Secondary Structure using Maximum Likelihood') has been developed to implement an evolutionary model that combines protein secondary structure and amino acid replacement. The model is related to that of Dayhoff and co-workers, but we distinguish eight categories of structural environment: alpha helix, beta sheet, turn and coil, each further classified according to solvent accessibility, i.e. buried or exposed. The model of sequence evolution for each of the eight categories is a Markov process with discrete states in continuous time, and the organization of structure along protein sequences is described by a hidden Markov model. This paper describes the PASSML software and illustrates how it allows both the reconstruction of phylogenies and prediction of secondary structure from aligned amino acid sequences. AVAILABILITY: PASSML 'ANSI C' source code and the example data sets described here are available at http://ng-dec1.gen.cam.ac.uk/hmm/Passml.html and 'downstream' Web pages. CONTACT: P.Lio@gen.cam.ac.uk

Adenylate Kinase↗