PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Mouse IgA heavy chain gene sequence: implications for evolution of immunoglobulin hinge axons.

The complete nucleotide sequence of the gene and mRNA coding for the constant (C) region of the secreted form of the BALB/c mouse IgA immunoglobulin alpha heavy (H) chain has been determined. As in other immunoglobulins, the three C region domains of the alpha protein, C alpha 1, C alpha 2, and C alpha 3 are coded in separate exons. However, the hinge region of C alpha is not coded on a separate exon as it is in other hinge-containing immunoglobulins. Instead, the alpha hinge is coded as a 5' extension of the C alpha 2 exon, and we suggest that it may have evolved by duplication leading to incorporation of an acceptor RNA splice site into the coding portion of the C alpha 2 exon. Extensions of this concept could provide an explanation for duplications in the human alpha 1 chain.

Amino Acid Sequence↗

Benchmarking tools for the alignment of functional noncoding DNA.

BACKGROUND: Numerous tools have been developed to align genomic sequences. However, their relative performance in specific applications remains poorly characterized. Alignments of protein-coding sequences typically have been benchmarked against "correct" alignments inferred from structural data. For noncoding sequences, where such independent validation is lacking, simulation provides an effective means to generate "correct" alignments with which to benchmark alignment tools. RESULTS: Using rates of noncoding sequence evolution estimated from the genus Drosophila, we simulated alignments over a range of divergence times under varying models incorporating point substitution, insertion/deletion events, and short blocks of constrained sequences such as those found in cis-regulatory regions. We then compared "correct" alignments generated by a modified version of the ROSE simulation platform to alignments of the simulated derived sequences produced by eight pairwise alignment tools (Avid, BlastZ, Chaos, ClustalW, DiAlign, Lagan, Needle, and WABA) to determine the off-the-shelf performance of each tool. As expected, the ability to align noncoding sequences accurately decreases with increasing divergence for all tools, and declines faster in the presence of insertion/deletion evolution. Global alignment tools (Avid, ClustalW, Lagan, and Needle) typically have higher sensitivity over entire noncoding sequences as well as in constrained sequences. Local tools (BlastZ, Chaos, and WABA) have lower overall sensitivity as a consequence of incomplete coverage, but have high specificity to detect constrained sequences as well as high sensitivity within the subset of sequences they align. Tools such as DiAlign, which generate both local and global outputs, produce alignments of constrained sequences with both high sensitivity and specificity for divergence distances in the range of 1.25-3.0 substitutions per site. CONCLUSION: For species with genomic properties similar to Drosophila, we conclude that a single pair of optimally diverged species analyzed with a high performance alignment tool can yield accurate and specific alignments of functionally constrained noncoding sequences. Further algorithm development, optimization of alignment parameters, and benchmarking studies will be necessary to extract the maximal biological information from alignments of functional noncoding DNA.

Animals↗

Precise sequence assignment of replication origin in the control region of chick mitochondrial DNA relative to 5' and 3' D-loop ends, secondary structure, DNA synthesis, and protein binding.

The data reported identify for the first time the sequence of an avian mitochondrial heavy-strand replication origin, OH, located only about 12 nucleotides (nt) downstream from the conserved sequence block CSB-1, as well as the sequence of premature synthesis arrest of the 781 (+/-1) nt D-loop strand, only 6-7 nt downstream from a TAS-like (termination-associated) element. Both sites are associated with putative cruciform secondary structures. A major sequence-specific DNA-binding/cleavage site of a potential regulatory protein, the approximately 36-kDa aMDP1 (shown previously to stimulate mtDNA synthesis), is located about 90 nt upstream of OH. Correlated in vivo analysis of avian genome-length mtDNA replication provides missing evidence on the functional equivalence of D-loop origin with nascent initiation, and on the direction, asymmetry and temporal aspects of a full round of replication. The importance of the results to understanding the regulation of linked replication/transcription and the unusual sequence evolution of avian mtDNA is

Animals↗

Using the tangle: a consistent construction of phylogenetic distance matrices for quartets.

Distance based algorithms are a common technique in the construction of phylogenetic trees from taxonomic sequence data. The first step in the implementation of these algorithms is the calculation of a pairwise distance matrix to give a measure of the evolutionary change between any pair of the extant taxa. A standard technique is to use the log det formula to construct pairwise distances from aligned sequence data. We review a distance measure valid for the most general models, and show how the log det formula can be used as an estimator thereof. We then show that the foundation upon which the log det formula is constructed can be generalized to produce a previously unknown estimator which improves the consistency of the distance matrices constructed from the log det formula. This distance estimator provides a consistent technique for constructing quartets from phylogenetic sequence data under the assumption of the most general Markov model of sequence evolution.

Algorithms↗

Kinetoplast DNA minicircles: regions of extensive sequence divergence.

Previous work has shown that the kinetoplast minicircle DNA of Leishmania species exhibits species-specific sequence divergence and this observation has led to the development of a DNA probe-based diagnostic test for leishmaniasis. In the work reported here, we demonstrate that the minicircle is composed of three types of DNA sequences with differing specificities reflecting different rates of DNA sequence change. A library of cloned fragments of kinetoplast DNA (kDNA) from Leishmania mexicana amazonensis was prepared and the cloned subfragments were found to contain DNA sequences with different taxonomic specificities based on hybridization analysis with various species of Leishmania. Four groups of subfragments were found, those that hybridized with a large number of Leishmania sp. as well as sequences unique to the species, subspecies, or isolate. Analysis of nested deletions of a single, full-length minicircle demonstrates that these different taxonomic specificities are contained within a single minicircle. This implies that different regions of a single minicircle have DNA sequences that diverge at different rates. These sequences represent potentially valuable tools in diagnostic, epidemiologic, and ecological studies of leishmaniasis and provide the basis for a model of kDNA sequence evolution.

Animals↗

Analytical expression of the purine/pyrimidine autocorrelation function after and before random mutations.

The mutation process is a classical evolutionary genetic process. The type of mutations studied here is the random substitutions of a purine base R (adenine or guanine) by a pyrimidine base Y (cytosine or thymine) and reciprocally (transversions). The analytical expressions derived allow us to analyze in genes the occurrence probabilities of motifs and d-motifs (two motifs separated by any d bases) on the R/Y alphabet under transversions. These motif probabilities can be obtained after transversions (in the evolutionary sense; from the past to the present) and, unexpectedly, also before transversions (after back transversions, in the inverse evolutionary sense, from the present to the past). This theoretical part in Section 2 is a first generalization of a particular formula recently derived. The application in Section 3 is based on the analytical expression giving the autocorrelation function (the d-motif probabilities) before transversions. It allows us to study primitive genes from actual genes. This approach solves a biological problem. The protein coding genes of chloroplasts and mitochondria have a preferential occurrence of the 6-motif YRY(N)6YRY (maximum of the autocorrelation function for d = 6, N = R or Y) with a periodicity modulo 3. The YRY(N)6YRY preferential occurrence without the periodicity modulo 3 is also observed in the RNA coding genes (ribosomal, transfer, and small nuclear RNA genes) and in the noncoding genes (introns and 5' regions of eukaryotic nuclei). However, there are two exceptions to this YRY(N)6YRY rule: the protein coding genes of eukaryotic nuclei, and prokaryotes, where YRY(N)6YRY has the second highest value after YRY(N)0YRY (YRYYRY) with a periodicity modulo 3. When we go backward in time with the analytical expression, the protein coding genes of both eukaryotic nuclei and prokaryotes retrieve the YRY(N)6YRY preferential occurrence with a periodicity modulo 3 after 0.2 back transversions per base. In other words, the actual protein coding genes of chloroplasts and mitochondria are similar to the primitive protein coding genes of eukaryotic nuclei and prokaryotes. On the other hand, this application represents the first result concerning the mutation process in the model of DNA sequence evolution we recently proposed. According to this model, the actual genes on the R/Y alphabet derive from two successive evolutionary genetic processes: an independent mixing of a few nonrandom types of oligonucleotides leading to genes called primitive followed by a mutation process in these primitive genes.(ABSTRACT TRUNCATED AT 400 WORDS)

Base Sequence↗

Inference of population history using a likelihood approach.

We introduce an approach to revealing the likelihood of different population histories that utilizes an explicit model of sequence evolution for the DNA segment under study. Based on a phylogenetic tree reconstruction method we show that a Tamura-Nei model with heterogeneous mutation rates is a fair description of the evolutionary process of the hypervariable region I of the mitochondrial DNA from humans. Assuming this complex model still allows the estimation of population history parameters, we suggest a likelihood approach to conducting statistical inference within a class of expansion models. More precisely, the likelihood of the data is based on the mean pairwise differences between DNA sequences and the number of variable sites in a sample. The use of likelihood ratios enables comparison of different hypotheses about population history, such as constant population size during the past or an increase or decrease of population size starting at some point back in time. This method was applied to show that the population of the Basques has expanded, whereas that of the Biaka pygmies is most likely decreasing. The Nuu-Chah-Nulth data are consistent with a model of constant population.

Base Composition↗

Maximum-likelihood methods for phylogeny estimation.

Maximum-likelihood (ML) estimation of phylogenies has reached a rather high level of sophistication because of algorithmic advances, improvements in models of sequence evolution, and improvements in statistical approaches and application of cluster computing. Here, I provide a brief basic background in application of the general principle of ML estimation to phylogenetics and provide an example of selecting among a nested set of ML models using a dynamic approach to hierarchical likelihood-ratio tests. I focus attention on PAUP* because it provides unique ease of switching among alternative optimality criteria (e.g., minimum evolution, parsimony, and ML). Further, examples of parametric bootstrap tests are provided that demonstrate statistical tests of phylogenetic hypotheses and model adequacy, in an absolute rather than relative sense. The increasing availability of clustered, parallelized computation makes use of such parametric approaches feasible.

Algorithms↗

A conservative test of genetic drift in the endosymbiotic bacterium Buchnera: slightly deleterious mutations in the chaperonin groEL.

The obligate endosymbiotic bacterium Buchnera aphidicola shows elevated rates of sequence evolution compared to free-living relatives, particularly at nonsynonymous sites. Because Buchnera experiences population bottlenecks during transmission to the offspring of its aphid host, it is hypothesized that genetic drift and the accumulation of slightly deleterious mutations can explain this rate increase. Recent studies of intraspecific variation in Buchnera reveal patterns consistent with this hypothesis. In this study, we examine inter- and intraspecific nucleotide variation in groEL, a highly conserved chaperonin gene that is constitutively overexpressed in Buchnera. Maximum-likelihood estimates of nonsynonymous substitution rates across Buchnera species are strikingly low at groEL compared to other loci. Despite this evidence for strong purifying selection on groEL, our intraspecific analysis of this gene documents reduced synonymous polymorphism, elevated nonsynonymous polymorphism, and an excess of rare alleles relative to the neutral expectation, as found in recent studies of other Buchnera loci. Comparisons with Escherichia coli generally show patterns predicted by their differences in N(e). The sum of these observations is not expected under relaxed or balancing selection, selective sweeps, or increased mutation rate. Rather, they further support the hypothesis that drift is an important force driving accelerated protein evolution in this obligate mutualist.

Buchnera↗

Molecular evolution of the wingless gene and its implications for the phylogenetic placement of the butterfly family Riodinidae (Lepidoptera: papilionoidea).

The sequence evolution of the nuclear gene wingless was investigated among 34 representatives of three lepidopteran families (Riodinidae, Lycaenidae, and Nymphalidae) and four outgroups, and its utility for inferring phylogenetic relationships among these taxa was assessed. Parsimony analysis yielded a well-resolved topology supporting the monophyly of the Riodinidae and Lycaenidae, respectively, and indicating that these two groups are sister lineages, with strong nodal support based on bootstrap and decay indices. Although wingless provides robust support for relationships within and between the riodinids and the lycaenids, it is less informative about nymphalid relationships. Wingless does not consistently recover nymphalid monophyly or traditional subfamilial relationships within the nymphalids, and nodal support for all but the most recent branches in this family is low. Much of the phylogenetic information in this data set is derived from first- and second-position substitutions. However, third positions, despite showing uncorrected pairwise divergences up to 78%, also contain consistent signal at deep nodes within the family Riodinidae and at the node defining the sister relationship between the riodinids and lycaenids. Several hypotheses about how third-position signal has been retained in deep nodes are discussed. These include among-site rate variation, identified as a significant factor by maximum likelihood analyses, and nucleotide bias, a prominent feature of third positions in this data set. Understanding the mechanisms which underlie third-position signal is a first step in applying appropriate models to accommodate the specific evolutionary processes involved in each lineage.

Animals↗

A statistical characterization of consistent patterns of human immunodeficiency virus evolution within infected patients.

Within-patient HIV populations evolve rapidly because of a high mutation rate, short generation time, and strong positive selection pressures. Previous studies have identified "consistent patterns" of viral sequence evolution. Just before HIV infection progresses to AIDS, evolution seems to slow markedly, and the genetic diversity of the viral population drops. This evolutionary slowdown could be caused either by a reduction in the average viral replication rate or because selection pressures weaken with the collapse of the immune system. The former hypothesis (which we denote "cellular exhaustion") predicts a simultaneous reduction in both synonymous and nonsynonymous evolution, whereas the latter hypothesis (denoted "immune relaxation") predicts that only nonsynonymous evolution will slow. In this paper, we present a set of statistical procedures for distinguishing between these alternative hypotheses using DNA sequences sampled over the course of infection. The first component is a new method for estimating evolutionary rates that takes advantage of the temporal information in longitudinal DNA sequence samples. Second, we develop a set of probability models for the analysis of evolutionary rates in HIV populations in vivo. Application of these models to both synonymous and nonsynonymous evolution affords a comparison of the cellular-exhaustion and immune-relaxation hypotheses. We apply the procedures to longitudinal data sets in which sequences of the env gene were sampled over the entire course of infection. Our analyses (1) statistically confirm that an evolutionary slowdown occurs late in infection, (2) strongly support the immune-relaxation hypothesis, and (3) indicate that the cessation of nonsynonymous evolution is associated with disease progression.

Acquired Immunodeficiency Syndrome↗

Statistical tests of models of DNA substitution.

Penny et al. have written that "The most fundamental criterion for a scientific method is that the data must, in principle, be able to reject the model. Hardly any [phylogenetic] tree-reconstruction methods meet this simple requirement." The ability to reject models is of such great importance because the results of all phylogenetic analyses depend on their underlying models--to have confidence in the inferences, it is necessary to have confidence in the models. In this paper, a test statistic suggested by Cox is employed to test the adequacy of some statistical models of DNA sequence evolution used in the phylogenetic inference method introduced by Felsenstein. Monte Carlo simulations are used to assess significance levels. The resulting statistical tests provide an objective and very general assessment of all the components of a DNA substitution model; more specific versions of the test are devised to test individual components of a model. In all cases, the new analyses have the additional advantage that values of phylogenetic parameters do not have to be assumed in order to perform the tests.

Animals↗

Evolutionary variants of the human immunodeficiency virus type 1 V3 region characterized by using a heteroduplex tracking assay.

Syncytium-inducing (SI) variants of human immunodeficiency virus type 1 (HIV-1) are evolutionary variants that are associated with rapid CD4+ cell loss and rapid disease progression. The heteroduplex tracking assay (HTA) was used to detect evolutionary V3 variants by amplifying the V3 sequences from viral RNA derived from 50 samples of patient plasma. For this V3-specific HTA (V3-HTA), heteroduplexes were formed between the patient V3 sequences and a probe with the subtype B consensus V3 sequence. Evolution was then measured by divergence from the consensus. The presence of evolutionary variants was correlated with SI detection data on the same samples from the MT-2 cell culture assay. Evolutionary variants were correlated with the SI phenotype in 88% of the samples, and 96% of the SI samples contained evolutionary variants. In most cases the evolutionary V3 variants represented discrete clonal outgrowths of virus. Sequence analysis of the six discordant samples that did not show this correlation indicated that three non-syncytium-inducing (NSI) samples had V3 sequences that had evolved away from the consensus sequence but not toward an SI genotype. A fourth sample showed little evolution away from the consensus but was SI, which indicates that not all SI variants require basic substitutions in V3. The other two samples had SI-like genotypes and NSI phenotypes, suggesting that V3-HTA was able to detect SI emergence in these samples in the absence of their detection in vitro. V3-HTA was also used to confirm SI variant selection in MT-2 cells and to examine the possibility of variant selection during virus culture in peripheral blood cells.

Acquired Immunodeficiency Syndrome↗

PASSML: combining evolutionary inference and protein secondary structure prediction.

MOTIVATION: Evolutionary models of amino acid sequences can be adapted to incorporate structure information; protein structure biologists can use phylogenetic relationships among species to improve prediction accuracy. Results : A computer program called PASSML ('Phylogeny and Secondary Structure using Maximum Likelihood') has been developed to implement an evolutionary model that combines protein secondary structure and amino acid replacement. The model is related to that of Dayhoff and co-workers, but we distinguish eight categories of structural environment: alpha helix, beta sheet, turn and coil, each further classified according to solvent accessibility, i.e. buried or exposed. The model of sequence evolution for each of the eight categories is a Markov process with discrete states in continuous time, and the organization of structure along protein sequences is described by a hidden Markov model. This paper describes the PASSML software and illustrates how it allows both the reconstruction of phylogenies and prediction of secondary structure from aligned amino acid sequences. AVAILABILITY: PASSML 'ANSI C' source code and the example data sets described here are available at http://ng-dec1.gen.cam.ac.uk/hmm/Passml.html and 'downstream' Web pages. CONTACT: P.Lio@gen.cam.ac.uk

Adenylate Kinase↗

Annotation, nomenclature and evolution of four novel homeobox genes expressed in the human germ line.

The homeobox genes comprise a large gene superfamily characterised by a conserved DNA motif encoding the homeodomain. Most homeodomain proteins function as transcription factors, and many have important roles in embryonic development and cell differentiation. Here we describe, annotate and name four novel homeobox genes in the human genome: ARGFX, DPRX, TPRX1 and DUXA. Each has generated multiple retrotransposed (processed) pseudogenes; these are reliable indicators of germ-line expression because only in germ-line cells can retrotransposition result in inheritance to the next generation. The retrotransposed sequences were exploited here as a novel means to deduce exon-intron boundaries. All four novel genes show accelerated rates of protein sequence evolution. This fast rate of sequence change may be connected with roles in human reproductive biology. Deducing the evolutionary origins of these genes is not straightforward, but we propose that TPRX1, DPRX and DUXA are highly divergent derivatives of the CRX gene, itself a member of the Otx homeobox gene family.

Evolution, Molecular↗

Subgenome-specific markers in allopolyploid cotton Gossypium hirsutum: implications for evolutionary analysis of polyploids.

We developed a set of genetic markers specific to the A and D genome types of cotton using representational difference analysis (RDA). These markers produce amplification products with genomic DNA from allotetraploid cotton Gossypium hirsutum. One of the markers is a polymorphic amplified restriction fragment (PARF) - a sequence found in both A and D genomes but differently flanked by restriction sites. Results of phylogenetic analysis of the PARF sequences from diploid cottons and from allotetraploid G. hirsutum agree with a previous observation of the interlocus concerted evolution (sequences corresponding to A and D genomes are homogenized to a D genome-type sequence). Our study shows how RDA can be used to develop genome-specific markers that can be used to study molecular evolution of allopolyploids.

DNA, Plant↗

Simple and accurate estimation of ancestral protein sequences.

There are a variety of reasons to reconstruct the sequences of ancient proteins, but whatever the reason, the value of the reconstructed protein depends on the accuracy with which the ancient sequence is inferred. This study uses sequences simulated by a sequence-evolution simulation program that compares parsimony, maximum likelihood, and the Bayesian methods of inferring ancestral sequences and concludes that the Bayesian method, as implemented by MRBAYES 3.11, is preferred. Estimated ancestral sequences are of necessity the same length as the alignment on which the underlying phylogeny is based. A highly accurate method for correcting the estimated sequences is introduced, and it is shown that the correction permits inferring the sequences of ancient protein sequences with a very high degree of accuracy.

Base Sequence↗

Molecular evolution of the GapC gene family in Amsinckia spectabilis populations that differ in outcrossing rate.

Molecular evolutionary analysis of the glyceraldehyde 3-phosphate dehydrogenase (GapC) gene family was conducted in the plant genus Amsinckia (Boraginaceae), a group that exhibits marked variation in the mating system. GapC genes in this group differ from those of Arabidopsis thaliana in terms of both intron size and number. Phylogenetic and Southern hybridization analyses suggest the presence of multiple GapC loci, each defined by a set of base substitutions that are in strong linkage disequilibrium. One species of Amsinckia, A. spectabilis, was studied in some detail. This species consists of selfing (A. s. spectabilis) and outcrossing (A. s. microcarpa) varieties. Two selfing populations and one outcrossing population sample were analyzed in detail for variation at one of the members of this gene family. GapC3. A reduction in number of GapC3 haplotypes and level of genetic diversity was observed in the selfing populations of A. spectabilis. GapC3 in the outcrossing population (but not the two selfing populations) exhibited a significant departure from neutrality in the direction of an excess of singletons. These results are discussed in the context of forces acting on sequence evolution in populations with different mating systems.

Amsinckia↗