PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

A phylogenetic analysis of woodpeckers and their allies using 12S, Cyt b, and COI nucleotide sequences (class Aves; order Piciformes).

Although the woodpeckers have long been recognized as a natural, monophyletic taxon, morphological analyses of their intra- and intergeneric relationships have produced conflicting results. To clarify this issue, and as part of a larger study of piciform relationships, nucleotide sequences for the 12S ribosomal RNA (12S; 1123 bp), cytochrome b (Cyt b; 1022 bp), and cytochrome oxidase c subunit 1 (COI; 1512 bp) mitochondrial genes were obtained from 34 piciform species that included 16 of the 23 currently recognized woodpecker genera (subfamily Picinae), three piculets (subfamily Picumninae), a wryneck (subfamily Jynginae), a honeyguide (family Indicatoridae), and three barbets (infraorder Ramphastides). Analyses were conducted on the individual and combined 12S, Cyt b, and COI sequences with maximum parsimony, neighbor-joining, maximum likelihood, and Bayesian algorithms. Based on the strong, congruent support among the different data partitions and models of sequence evolution, a highly resolved consensus of the relationships among woodpeckers and their allies could be formed. The monophyly of Indicatoridae + Picidae (infraorder Picides), Picidae, Picinae + Picumninae, and Picinae was strongly supported in all analyses. However, the tribes Colaptini, Picini, Campephilini, and Campetherini were shown to be paraphyletic as were the genera of Colaptes and Piculus. A revision of the tribal-level classification of woodpeckers is proposed and the importance of plumage convergence among woodpeckers is discussed.

Animals↗

Follow-up study of intrahost HIV type 2 variability reveals discontinuous evolution of C2V3 sequences.

Sequence change dynamics in the highly polymorphic C2V3 region of the envelope glycoprotein gp105 coding sequence of HIV-2 was studied for two HIV-2 seropositive Guinean individuals over a 5-year period. Proviral DNA was amplified by PCR from uncultured peripheral blood mononuclear cells (PBMC), cloned, and sequenced. Amino acid sequence alignment and phylogenetic analysis showed distinct viral populations for the three collected samples (1992, 1995, and 1997) for both individuals. A discontinuous replacement of sequences suggests activation of latent viruses present in reservoirs, a process that has been documented for the HIV-1 infection but unreported so far for HIV-2. The nucleotide sequences presented in this work have been assigned accession numbers AJ400276 to AJ400319.

Amino Acid Sequence↗

Simulation studies on the evolution of amino acid sequences.

A model of molecular evolution in which the parameter (intrinsic rate of amino acid substitution) fluctuates from time to time was investigated by simulating the process. It was found that the usual method of estimation such as Poisson fitting underestimates this variation of the parameter when remote comparisons are made. At the same time, four distance measures (minimun base difference, Poisson fitting, random nucleotide substitutions and negative binomial fitting) were tested for their accuracy. When the substitution rate is not uniform among the amino acid sites, the negative bionomial fitting gives most satisfactory results, however, one needs to know the parameter beforehand in order to use this method. It was pointed out that the fluctuation of the evolutionary rate is expected if the nearly neutral but very slightly deleterious mutations play an important role on molecular evolution.

Amino Acid Sequence↗

Codon bias and noncoding GC content correlate negatively with recombination rate on the Drosophila X chromosome.

The patterns and processes of molecular evolution may differ between the X chromosome and the autosomes in Drosophila melanogaster. This may in part be due to differences in the effective population size between the two chromosome sets and in part to the hemizygosity of the X chromosome in Drosophila males. These and other factors may lead to differences both in the gene complements of the X and the autosomes and in the properties of the genes residing on those chromosomes. Here we show that codon bias and recombination rate are correlated strongly and negatively on the X chromosome, and that this correlation cannot be explained by indirect relationships with other known determinants of codon bias. This is in dramatic contrast to the weak positive correlation found on the autosomes. We explored possible explanations for these patterns, which required a comprehensive analysis of the relationships among multiple genetic properties such as protein length and expression level. This analysis highlights conserved features of coding sequence evolution on the X and the autosomes and illuminates interesting differences between these two chromosome sets.

Animals↗

Tandem repeats in plant mitochondrial genomes: application to the analysis of population differentiation in the conifer Norway spruce.

Mitochondrial DNA, widely applied in studies of population differentiation in animals, is rarely used in plants because of its slow rate of sequence evolution and its complex genomic organization. We demonstrate the utility of two polymorphic mitochondrial tandem repeats located in the second intron of the nad1 gene of Norway spruce. Most of the size variants showed pronounced population differentiation and a distinct geographical distribution. A GenBank search revealed that mitochondrial tandem repeats occur in a broad range of plant species and may serve as a novel molecular marker for unravelling population processes in plants.

Cycadopsida↗

Mutation, selection, and ancestry in branching models: a variational approach.

We consider the evolution of populations under the joint action of mutation and differential reproduction, or selection. The population is modelled as a finite-type Markov branching process in continuous time, and the associated genealogical tree is viewed both in the forward and the backward direction of time. The stationary type distribution of the reversed process, the so-called ancestral distribution, turns out as a key for the study of mutation-selection balance. This balance can be expressed in the form of a variational principle that quantifies the respective roles of reproduction and mutation for any possible type distribution. It shows that the mean growth rate of the population results from a competition for a maximal long-term growth rate, as given by the difference between the current mean reproduction rate, and an asymptotic decay rate related to the mutation process; this tradeoff is won by the ancestral distribution. We then focus on the case when the type is determined by a sequence of letters (like nucleotides or matches/mismatches relative to a reference sequence), and we ask how much of the above competition can still be seen by observing only the letter composition (as given by the frequencies of the various letters within the sequence). If mutation and reproduction rates can be approximated in a smooth way, the fitness of letter compositions resulting from the interplay of reproduction and mutation is determined in the limit as the number of sequence sites tends to infinity. Our main application is the quasispecies model of sequence evolution with mutation coupled to reproduction but independent across sites, and a fitness function that is invariant under permutation of sites. In this model, the fitness of letter compositions is worked out explicitly. In certain cases, their competition leads to a phase transition.

Animals↗

Crypthecodinium and Tetrahymena: an exercise in comparative evolution.

Nucleotide sequences have been determined for the highly variable D2 region of the large rRNA molecule for over 60 strains of dinoflagellates. These strains were selected from a worldwide collection that represents all the known sibling species (compatibility groups, Mendelian species) in the sibling swarm referred to as Crypthecodinium cohnii. A phylogenetic tree has been constructed from an analysis of the variations in a length of about 180 bases, using PHYLOGEN string analysis programs. The Crypthecodinium tree is compared with the previously published but here augmented tree constructed upon the same rRNA region for the sibling species of a worldwide collection of ciliated protozoa related to the genus Tetrahymena. The first reported sequence of Lambornella clarki, the parasite of tree-hole mosquitoes, is included. The dinoflagellate species complex is much more homogeneous with respect to ribosomal variation. The mean number of differences among sequences from different Crypthecodinium species is about 7, in comparison with 22 differences among the ciliate species examined. Moreover, all the diversity in the dinoflagellates can be explained by base substitutions, whereas insertions and deletions are common in the ciliates. The dinoflagellates are also much more uniform with respect to nutritional and genetic economies. The two complexes differ also in the relationship between molecular variations and breeding compatibility. All tetrahymenine sibling species thus far examined are monomorphic in the D2 region, but several dinoflagellate species are polymorphic. Several different dinoflagellate species, moreover, have identical D2 regions. This kind of ribosomal identity of incompatible strains is found in these ciliates only in one tight cluster of species--Group C. The tetrahymenine swarm is apparently much older than the Crypthecodinium swarm, and the dinoflagellate species produce incompatible progeny species much more readily than do the ciliates, perhaps by the acquisition of mutations that potentiate incompatibility in sympatric populations.

Animals↗

Extreme variations in the ratios of non-synonymous to synonymous nucleotide substitution rates in signal peptide evolution.

Nucleotide sequences encoding signal peptides from the precursors of alpha-amylase/trypsin inhibitors from cereals are homologous to those corresponding to the precursors of thaumatin II and of plastocyanins. Non-synonymous (KA) and synonymous (KS) rates of nucleotide substitutions have been calculated for all possible binary combinations. Extreme variation in KA/KS ratios has been observed; from the 0.167 average found within the plastocyanin family to an average of 1.90 calculated for the inhibitors/thaumatin II transition. A similar calculation has been carried out for the signal peptide sequences of thionins, which are unrelated to those of the alpha-amylase/trypsin inhibitor family, and an average KA/KS of 0.12 has been obtained. This variation can be largely explained in terms of an empirical index of stability related to amino acid composition and seems to be independent of functional constraints.

Amino Acid Sequence↗

Defective accessory genes in a human immunodeficiency virus type 1-infected long-term survivor lacking recoverable virus.

We have been studying a patient who acquired human immunodeficiency virus (HIV) infection via a blood transfusion 13 years ago. She has remained asymptomatic since that time. The blood donor and two other recipients have all died of AIDS. Although this patient has shown persistently strong seroreactivity to HIV type 1 (HIV-1) antigens by Western blot (immunoblot), she has been continually HIV culture negative in results from multiple laboratories over the last 6 years and has a very low viral burden. Her CD4+ T-cell count has fluctuated around a mean of 399 cells per microliters, with little change in lymphocyte subset percentages. Strong cellular immune responses to HIV-1 epitopes by this patient have been demonstrated. We now report the results of an intensive molecular genetic analysis of the HIV-1 proviral quasispecies from this patient sampled over 5 years. Long terminal repeat region sequences supported the argument for normal basal and Tat-mediated promoter activities. Sequential sequencing of the nef gene revealed a low frequency (8.3%) of defective genes and a striking lack of sequence evolution. Functional analysis of predominant nef genes by both a cell surface CD4 downregulation and a viral infectivity complementation assay showed wild-type function. In contrast, sequential analysis of an amplicon containing the vif, vpr, vpu, tat1, and rev1 genes revealed the presence of inactivating mutations in 64% of the clones. These data suggest that this patient, initially infected with a virulent swarm of HIV-1, is presently infected with a more-attenuated viral quasispecies as a result of effective host immunity.

Acquired Immunodeficiency Syndrome↗

Theoretical foundation of the balanced minimum evolution method of phylogenetic inference and its relationship to weighted least-squares tree fitting.

Due to its speed, the distance approach remains the best hope for building phylogenies on very large sets of taxa. Recently (R. Desper and O. Gascuel, J. Comp. Biol. 9:687-705, 2002), we introduced a new "balanced" minimum evolution (BME) principle, based on a branch length estimation scheme of Y. Pauplin (J. Mol. Evol. 51:41-47, 2000). Initial simulations suggested that FASTME, our program implementing the BME principle, was more accurate than or equivalent to all other distance methods we tested, with running time significantly faster than Neighbor-Joining (NJ). This article further explores the properties of the BME principle, and it explains and illustrates its impressive topological accuracy. We prove that the BME principle is a special case of the weighted least-squares approach, with biologically meaningful variances of the distance estimates. We show that the BME principle is statistically consistent. We demonstrate that FASTME only produces trees with positive branch lengths, a feature that separates this approach from NJ (and related methods) that may produce trees with branches with biologically meaningless negative lengths. Finally, we consider a large simulated data set, with 5,000 100-taxon trees generated by the Aldous beta-splitting distribution encompassing a range of distributions from Yule-Harding to uniform, and using a covarion-like model of sequence evolution. FASTME produces trees faster than NJ, and much faster than WEIGHBOR and the weighted least-squares implementation of PAUP*. Moreover, FASTME trees are consistently more accurate at all settings, ranging from Yule-Harding to uniform distributions, and all ranges of maximum pairwise divergence and departure from molecular clock. Interestingly, the covarion parameter has little effect on the tree quality for any of the algorithms. FASTME is freely available on the web.

Algorithms↗

Quantifying residual HIV-1 replication in patients receiving combination antiretroviral therapy.

BACKGROUND: In patients infected with human immunodeficiency virus type 1 (HIV-1), combination antiretroviral therapy can result in sustained suppression of plasma levels of the virus. However, replication-competent virus can still be recovered from latently infected resting memory CD4 lymphocytes; this finding raises serious doubts about whether antiviral treatment can eradicate HIV-1. METHODS: We looked for evidence of residual HIV-1 replication in eight patients who began treatment soon after infection and in whom plasma levels of HIV-1 RNA were undetectable after two to three years of antiretroviral therapy. We examined whether there had been changes over time in HIV-1 proviral sequences in peripheral-blood mononuclear cells, which would indicate residual viral replication. We also performed in situ hybridization studies on tissues from one patient to identify cells actively expressing HIV-1 RNA. We estimated the rate of decrease of latent, replication-competent HIV-1 in resting CD4 lymphocytes on the basis of the decrease in the numbers of proviral sequences identified during primary infection and direct sequential measurements of the size of the latent reservoir. RESULTS: Six of the eight patients had no significant variations in proviral sequences during treatment. However, in two patients there was sequence evolution but no evidence of drug-resistant viral genotypes. In one patient, extensive in situ studies provided additional evidence of persistent viral replication in lymphoid tissues. Using two independent approaches, we estimated that the half-life of the latent, replication-competent virus in resting CD4 lymphocytes was approximately six months. CONCLUSIONS: These findings suggest that combination antiretroviral regimens suppress HIV-1 replication in some but not all patients. Given the half-life of latently infected CD4 lymphocytes of about six months, it may require many years of effective antiretroviral treatment to eliminate this reservoir of HIV-1.

Adult↗

The complete sequence of the mitochondrial genome of Daphnia pulex (Cladocera: Crustacea).

The sequence of the mitochondrial DNA (mtDNA) of the branchiopod crustacean Daphnia pulex has been completed. It is 15333bp with an A+T content of 62.3%, and contains the typical complement of 13 protein-coding, 22 transfer RNA (tRNA) and two ribosomal RNA (rRNA) genes. Comparison of this sequence with the sequences of the other eight completely sequenced arthropod mtDNAs showed that gene order and orientation are identical to that of Drosophila but different from Artemia due to the rearrangement of two tRNA genes. Nucleotide composition, codon usage, and amino acid composition are very similar in the crustaceans, but divergent from insects and chelicerates which show a much higher bias towards A+T. However, with few exceptions, the mitochondrial proteins of Daphnia are more similar to those of the dipteran insects (Drosophila and Anopheles) than to those of Artemia, at both the nucleotide and amino acid levels, suggesting that Artemia mtDNA is evolving at an accelerated rate. These results also show that sequence evolution and the evolution of nucleotide composition can be decoupled. Analysis of nucleotide substitution patterns in COII showed that there has been an unbiased acceleration of the overall substitution rate in Artemia. In contrast, the accelerated substitution rate in Apis is due partly to extreme A+T mutation pressure. Secondary structures are proposed for the Daphnia tRNAs and rRNAs. The tRNAs are similar to those of other arthropods but tend to have TPsiC arms that are only 4bp long. The rRNA secondary structures are similar to those proposed for insects except for the absence of a small number of helices in Daphnia. Phylogenetic analysis of second codon positions grouped Daphnia with Artemia, as expected, despite the latter's accelerated divergence rate. In contrast, the unusual pattern of mtDNA divergence in Apis led to a topology in which the holometabolous insects (Anopheles, Drosophila, Apis) appeared to be paraphyletic with respect to the hemimetabolous insect, Locusta, due to the early branching of Apis.

Animals↗

Nucleotide sequence of yellow fever virus: implications for flavivirus gene expression and evolution.

The sequence of the entire RNA genome of the type flavivirus, yellow fever virus, has been obtained. Inspection of this sequence reveals a single long open reading frame of 10,233 nucleotides, which could encode a polypeptide of 3411 amino acids. The structural proteins are found within the amino-terminal 780 residues of this polyprotein; the remainder of the open reading frame consists of nonstructural viral polypeptides. This genome organization implies that mature viral proteins are produced by posttranslational cleavage of a polyprotein precursor and has implications for flavivirus RNA replication and for the evolutionary relation of this virus family to other RNA viruses.

Base Sequence↗

Plant conserved non-coding sequences and paralogue evolution.

Genome duplication is a powerful evolutionary force and is arguably most prominent in plants, where several ancient whole-genome duplication events have been documented. Models of gene evolution predict that functional divergence between duplicates (subfunctionalization) is caused by the loss of regulatory elements. Studies of conserved non-coding sequences (CNSs), which are putative regulatory elements, indicate that plants have far fewer CNSs per gene than mammals, suggesting that plants have less complex regulatory mechanisms. Furthermore, a recent study of a duplicated gene pair in maize suggests that CNSs are lost in a complementary fashion, perhaps driving subfunctionalization. If subfunctionalization is common, one expects duplicate genes to diverge in expression; recent microarray analyses in Arabidopsis thalinia suggest that this is the case. Plant genomes are relatively complex on a genomic level because of the prevalence of whole-genome duplication and, paradoxically, subfunctionalization after duplication can lead to relatively simple regulatory regions on a per gene basis.

Base Sequence↗

Monophyly and relationships of wrens (Aves: Troglodytidae): a congruence analysis of heterogeneous mitochondrial and nuclear DNA sequence data.

The wrens (Aves: Troglodytidae) are a group of primarily New World insectivorous birds, the monophyly of which has long been recognized, but whose intergeneric relationships are essentially unknown. In order to test the monophyly of the group, and to attempt to resolve relationships among genera within it, sequences from the mitochondrial cytochrome b gene and the fourth intron of the nuclear beta-fibrinogen gene were obtained from nearly all genera of wrens, from their relatives as suggested by traditional taxonomy and DNA-DNA hybridization analyses, and from additional passerines. Maximum likelihood analysis of the two data sets yielded maximal congruence between independently derived estimates of relationship, outperforming a variety of weighted parsimony methods. Hierarchical likelihood ratio tests indicated that the two gene regions differed significantly in every estimated parameter of sequence evolution, and combined analysis of the two data sets was accomplished using a heterogeneous-model Bayesian approach. Independent and simultaneous analyses of both data sets supported monophyly of the wrens (excluding one recently added member, the monotypic genus Donacobius) and a sister-group relationship between wrens and the gnatcatchers (Polioptila). Additionally, strong support was found for paraphyly of the genus Thryothorus, and for a sister-group relationship between the genera Cistothorus and Troglodytes. Analyses of these data failed to resolve basal relationships within wrens, possibly due to ambiguity in rooting with a distant, species-poor outgroup. Analysis of the combined data for wrens alone yielded results which were largely congruent with relationships inferred using the complete data set, with the benefit of stronger support for relationships within the group. However, alternative rootings of this ingroup tree were weakly supported by nucleotide substitution data. Insertion-deletion events suggest that the genus Salpinctes may be sister to all other wrens.

Animals↗

Evolution of homologous sequences on the human X and Y chromosomes, outside of the meiotic pairing segment.

A sequence isolated from the long arm of the human Y chromosome detects a highly homologous locus on the X. This homology extends over at least 50 kb of DNA and is postulated to be the result of a transposition event between the X and Y chromosomes during recent human evolution, since homologous sequences are shown to be present on the X chromosome alone in the chimpanzee and gorilla.

Animals↗

Analytical expression of the purine/pyrimidine codon probability after and before random mutations.

Recently, we proposed a new model of DNA sequence evolution (Arquès and Michel. 1990b. Bull. math. Biol. 52, 741-772) according to which actual genes on the purine/pyrimidine (R/Y) alphabet (R = purine = adenine or guanine, Y = pyrimidine = cytosine or thymine) are the result of two successive evolutionary genetic processes: (i) a mixing (independent) process of non-random oligonucleotides (words of base length less than 10: YRY(N)6, YRYRYR and YRYYRY are so far identified; N = R or Y) leading to primitive genes (words of several hundreds of base length) and followed by (ii) a random mutation process, i.e., transformations of a base R (respectively Y) into the base Y (respectively R) at random sites in these primitive genes. Following this model the problem investigated here is the study of the variation of the 8 R/Y codon probabilities RRR, ..., YYY under random mutations. Two analytical expressions solved here allow analysis of this variation in the classical evolutionary sense (from the past to the present, i.e., after random mutations), but also in the inverted evolutionary sense (from the present to the past, i.e., before random mutations). Different properties are also derived from these formulae. Finally, a few applications of these formulae are presented. They prove the proposition in Arquès and Michel (1990b. Bull. math. Biol. 52, 741-772), Section 3.3.2, with the existence of a maximal mean number of random mutations per base of the order 0.3 in the protein coding genes. They also confirm the mixing process of oligonucleotides by excluding the purine/pyrimidine contiguous and alternating tracts from the formation process of primitive genes.

Base Sequence↗

Data decisiveness, data quality, and incongruence in phylogenetic analysis: an example from the monocotyledons using mitochondrial atp A sequences.

We examined three parallel data sets with respect to qualities relevant to phylogenetic analysis of 20 exemplar monocotyledons and related dicotyledons. The three data sets represent restriction-site variation in the inverted repeat region of the chloroplast genome, and nucleotide sequence variation in the chloroplast-encoded gene rbcL and in the mitochondrion-encoded gene atpA, the latter of which encodes the alpha-subunit of mitochondrial ATP synthase. The plant mitochondrial genome has been little used in plant systematics, in part because nucleotide sequence evolution in enzyme-encoding genes of this genome is relatively slow. The three data sets were examined in separate and combined analyses, with a focus on patterns of congruence, homoplasy, and data decisiveness. Data decisiveness (described by P. Goloboff) is a measure of robustness of support for most parsimonious trees by a data set in terms of the degree to which those trees are shorter than the average length of all possible trees. Because indecisive data sets require relatively fewer additional steps than decisive ones to be optimized on nonparsimonious trees, they will have a lesser tendency to be incongruent with other data sets. One consequence of this relationship between decisiveness and character incongruence is that if incongruence is used as a criterion of noncombinability, decisive data sets, which provide robust support for relationships, are more likely to be assessed as noncombinable with other data sets than are indecisive data sets, which provide weak support for relationships. For the sampling of taxa in this study, the atpA data set has about half as many cladistically informative nucleotides as the rbcL data set per site examined, and is less homoplastic and more decisive. The rbcL data set, which is the least decisive of the three, exhibits the lowest levels of character incongruence. Whatever the molecular evolutionary cause of this phenomenon, it seems likely that the poorer performance of rbcL than atpA, in terms of data decisiveness, is due to both its higher overall level of homoplasy and the fact that it is performing especially poorly at nonsynonymous sites.

Adenosine Triphosphatases↗