PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

K-casein gene phylogeny of higher ruminants (Pecora, Artiodactyla).

To assess phylogenetic relationships among the higher ruminants (infraorder Pecora, order Artiodactyla), we analyzed K-casein DNA sequences, including 434 nucleotides of the fourth exon. The higher ruminant families Bovidae, Cervidae, Giraffidae, and Antilocapridae each have monophyletic K-casein sequences. Maximum parsimony and distance analyses identify Giraffidae as a sister group to either Cervidae or a Bovidae-Cervidae clade and Antilocapridae as a sister group to a Bovidae-Cervidae-Giraffidae clade. At a higher level these four families occur as a monophyletic clade relative to Tragulidae and Suidae. Within Cervidae, the subfamily Odocoileinae is monophyletic and Cervinae and Muntiacinae occur as independent lineages within a separate clade. Within Bovidae, the subfamilies Bovinae and Caprinae are monophyletic. Genera within Cervinae (Cervus, Elaphurus) and Bovinae (Bison, Bos) are paraphyletic. There is intraspecific allelic variation in Cervus elaphus, Odocoileus hemionus, and Bison bison. The rate of K-casein fourth exon DNA sequence evolution is estimated to be about 0.004 nucleotide substitutions per million years. The K-casein phylogeny is discussed relative to other molecular and morphological data.

Animals↗

Complete nucleotide sequence of pSK41: evolution of staphylococcal conjugative multiresistance plasmids.

The 46.4-kb nucleotide sequence of pSK41, a prototypical multiresistance plasmid from Staphylococcus aureus, has been determined, representing the first completely sequenced conjugative plasmid from a gram-positive organism. Analysis of the sequence has enabled the identification of the probable replication, maintenance, and transfer functions of the plasmid and has provided insights into the evolution of a clinically significant group of plasmids. The basis of deletions commonly associated with pSK41 family plasmids has been investigated, as has the observed insertion site specificity of Tn552-like beta-lactamase transposons within them. Several of the resistance determinants carried by pSK41-like plasmids were found to be located on up to four smaller cointegrated plasmids. pSK41 and related plasmids appear to represent a consolidation of antimicrobial resistance functions, collected by a preexisting conjugative plasmid via transposon insertion and IS257-mediated cointegrative capture of other plasmids.

Amino Acid Sequence↗

Viral evolution and interferon resistance of hepatitis C virus RNA replication in a cell culture model.

Hepatitis C virus (HCV) replicates through an error-prone process that may support the evolution of genetic variants resistant to the host cell antiviral response and interferon (IFN)-based therapy. We evaluated HCV-IFN interactions within a long-term culture system of Huh7 cell lines harboring different variants of an HCV type 1b subgenomic RNA replicon that differed at only two sites within the NS5A-encoding region. A replicon with a K insertion at HCV codon 2040 replicated efficiently and exhibited sequence stability in the absence of host antiviral pressure. In contrast, a replicon with an L2198S point mutation replicated poorly and triggered a cellular response characterized by IFN-beta production and low-level IFN-stimulated gene (ISG) expression. When maintained in long term-culture, the L2198S RNA evolved into a stable high-passage (HP) variant with six additional point mutations throughout the HCV protein-encoding region that enhanced viral replication. The HP RNA transduced Huh7 cells with more than 1,000-fold greater efficiency than its L2198S progenitor or the K2040 sequence. Replication of the HP RNA resisted suppression by IFN-alpha treatment and was associated with virus-directed reduction in host cell expression of ISG56, an antagonist of HCV RNA translation. Accordingly, the HP RNA was retained within polyribosome complexes in vivo that were refractory to IFN-induced disassembly. These results identify ISG56 as a translational control effector of the host response to HCV and provide direct evidence to link this response to viral sequence evolution, ISG regulation, and selection of the IFN-resistant viral phenotype.

Biological Evolution↗

Human immunodeficiency virus type 1 coreceptor switching: V1/V2 gain-of-fitness mutations compensate for V3 loss-of-fitness mutations.

Human immunodeficiency virus type 1 (HIV-1) entry into target cells is mediated by the virus envelope binding to CD4 and the conformationally altered envelope subsequently binding to one of two chemokine receptors. HIV-1 envelope glycoprotein (gp120) has five variable loops, of which three (V1/V2 and V3) influence the binding of either CCR5 or CXCR4, the two primary coreceptors for virus entry. Minimal sequence changes in V3 are sufficient for changing coreceptor use from CCR5 to CXCR4 in some HIV-1 isolates, but more commonly additional mutations in V1/V2 are observed during coreceptor switching. We have modeled coreceptor switching by introducing most possible combinations of mutations in the variable loops that distinguish a previously identified group of CCR5- and CXCR4-using viruses. We found that V3 mutations entail high risk, ranging from major loss of entry fitness to lethality. Mutations in or near V1/V2 were able to compensate for the deleterious V3 mutations and may need to precede V3 mutations to permit virus survival. V1/V2 mutations in the absence of V3 mutations often increased the capacity of virus to utilize CCR5 but were unable to confer CXCR4 use. V3 mutations were thus necessary but not sufficient for coreceptor switching, and V1/V2 mutations were necessary for virus survival. HIV-1 envelope sequence evolution from CCR5 to CXCR4 use is constrained by relatively frequent lethal mutations, deep fitness valleys, and requirements to make the right amino acid substitution in the right place at the right time.

Amino Acid Sequence↗

Potential sources of the 1995 Venezuelan equine encephalitis subtype IC epidemic.

Venezuelan equine encephalitis viruses (VEEV) belonging to subtype IC have caused three (1962-1964, 1992-1993 and 1995) major equine epizootics and epidemics. Previous sequence analyses of a portion of the envelope glycoprotein gene demonstrated a high degree of conservation among isolates from the 1962-1964 and the 1995 outbreaks, as well as a 1983 interepizootic mosquito isolate from Panaquire, Venezuela. However, unlike subtype IAB VEEV that were used to prepare inactivated vaccines that probably initiated several outbreaks, subtype IC viruses have not been used for vaccine production and their conservation cannot be explained in this way. To characterize further subtype IC VEEV conservation and to evaluate potential sources of the 1995 outbreak, we sequenced the complete genomes of three isolates from the 1962-1964 outbreak, the 1983 Panaquire interepizootic isolate, and two isolates from 1995. The sequence of the Panaquire isolate, and that of virus isolated from a mouse brain antigen prepared from subtype IC strain P676 and used in the same laboratory, suggested that the Panaquire isolate represents a laboratory contaminant. Some authentic epizootic IC strains isolated 32 years apart showed a greater degree of sequence identity than did isolates from the same (1962-1964 or 1995) outbreak. If these viruses were circulating and replicating between 1964 and 1995, their rate of sequence evolution was at least 10-fold lower than that estimated during outbreaks or that of closely related enzootic VEEV strains that circulate continuously. Current understanding of alphavirus evolution is inconsistent with this conservation. This subtype IC VEEV conservation, combined with phylogenetic relationships, suggests the possibility that the 1995 outbreak was initiated by a laboratory strain.

Amino Acid Sequence↗

Exploration of phylogenetic data using a global sequence analysis method.

BACKGROUND: Molecular phylogenetic methods are based on alignments of nucleic or peptidic sequences. The tremendous increase in molecular data permits phylogenetic analyses of very long sequences and of many species, but also requires methods to help manage large datasets. RESULTS: Here we explore the phylogenetic signal present in molecular data by genomic signatures, defined as the set of frequencies of short oligonucleotides present in DNA sequences. Although violating many of the standard assumptions of traditional phylogenetic analyses--in particular explicit statements of homology inherent in character matrices--the use of the signature does permit the analysis of very long sequences, even those that are unalignable, and is therefore most useful in cases where alignment is questionable. We compare the results obtained by traditional phylogenetic methods to those inferred by the signature method for two genes: RAG1, which is easily alignable, and 18S RNA, where alignments are often ambiguous for some regions. We also apply this method to a multigene data set of 33 genes for 9 bacteria and one archea species as well as to the whole genome of a set of 16 gamma-proteobacteria. In addition to delivering phylogenetic results comparable to traditional methods, the comparison of signatures for the sequences involved in the bacterial example identified putative candidates for horizontal gene transfers. CONCLUSION: The signature method is therefore a fast tool for exploring phylogenetic data, providing not only a pretreatment for discovering new sequence relationships, but also for identifying cases of sequence evolution that could confound traditional phylogenetic analysis.

Algorithms↗

Short-wavelength sensitive opsin (SWS1) as a new marker for vertebrate phylogenetics.

BACKGROUND: Vertebrate SWS1 visual pigments mediate visual transduction in response to light at short wavelengths. Due to their importance in vision, SWS1 genes have been isolated from a surprisingly wide range of vertebrates, including lampreys, teleosts, amphibians, reptiles, birds, and mammals. The SWS1 genes exhibit many of the characteristics of genes typically targeted for phylogenetic analyses. This study investigates both the utility of SWS1 as a marker for inferring vertebrate phylogenetic relationships, and the characteristics of the gene that contribute to its phylogenetic utility. RESULTS: Phylogenetic analyses of vertebrate SWS1 genes produced topologies that were remarkably congruent with generally accepted hypotheses of vertebrate evolution at both higher and lower taxonomic levels. The few exceptions were generally associated with areas of poor taxonomic sampling, or relationships that have been difficult to resolve using other molecular markers. The SWS1 data set was characterized by a substantial amount of among-site rate variation, and a relatively unskewed substitution rate matrix, even when the data were partitioned into different codon sites and individual taxonomic groups. Although there were nucleotide biases in some groups at third positions, these biases were not convergent across different taxonomic groups. CONCLUSION: Our results suggest that SWS1 may be a good marker for vertebrate phylogenetics due to the variable yet consistent patterns of sequence evolution exhibited across fairly wide taxonomic groups. This may result from constraints imposed by the functional role of SWS1 pigments in visual transduction.

Animals↗

Developmental differences in methylation of human Alu repeats.

Alu repeats are especially rich in CpG dinucleotides, the principal target sites for DNA methylation in eukaryotes. The methylation state of Alus in different human tissues is investigated by simple, direct genomic blot analysis exploiting recent theoretical and practical advances concerning Alu sequence evolution. Whereas Alus are almost completely methylated in somatic tissues such as spleen, they are hypomethylated in the male germ line and tissues which depend on the differential expression of the paternal genome complement for development. In particular, we have identified a subset enriched in young Alus whose CpGs appear to be almost completely unmethylated in sperm DNA. The existence of this subset potentially explains the conservation of CpG dinucleotides in active Alu source genes. These profound, sequence-specific developmental changes in the methylation state of Alu repeats suggest a function for Alu sequences at the DNA level, such as a role in genomic imprinting.

5-Methylcytosine↗

A likelihood approach for comparing synonymous and nonsynonymous nucleotide substitution rates, with application to the chloroplast genome.

A model of DNA sequence evolution applicable to coding regions is presented. This represents the first evolutionary model that accounts for dependencies among nucleotides within a codon. The model uses the codon, as opposed to the nucleotide, as the unit of evolution, and is parameterized in terms of synonymous and nonsynonymous nucleotide substitution rates. One of the model's advantages over those used in methods for estimating synonymous and nonsynonymous substitution rates is that it completely corrects for multiple hits at a codon, rather than taking a parsimony approach and considering only pathways of minimum change between homologous codons. Likelihood-ratio versions of the relative-rate test are constructed and applied to data from the complete chloroplast DNA sequences of Oryza sativa, Nicotiana tabacum, and Marchantia polymorpha. Results of these tests confirm previous findings that substitution rates in the chloroplast genome are subject to both lineage-specific and locus-specific effects. Additionally, the new tests suggest tha the rate heterogeneity is due primarily to differences in nonsynonymous substitution rates. Simulations help confirm previous suggestions that silent sites are saturated, leaving no evidence of heterogeneity in synonymous substitution rates.

Chloroplasts↗

[Tree reconciliation: reconstruction of species evolution by phylogenetic gene trees].

It is well known that phylogenetic trees derived from different protein families are often incongruent. This is explained by mapping errors and by the essential processes of gene duplication, loss, and horizontal transfer. Therefore, the problem is to derive a "consensus" tree best fitting the given set of gene trees. This work presents a new method of deriving this tree. The method is different from the existing ones, since it considers not only the topology of the initial gene trees, but also the reliability of their branches. Thereby one can explicitly take into account the possible errors in the gene trees caused by the absence of reliable models of sequence evolution, by uneven evolution of different gene families and taxonomic groups, etc.

Algorithms↗

Evolution of structural shape in bacterial globin-related proteins.

The globin family of proteins has a characteristic structural pattern of helix interactions that nonetheless exhibits some variation. A simplified model for globin structural evolution was developed in which protein shape evolved by random change of contacts between helices. A conserved globin domain of 15 bacterial proteins representing four structural families was studied. Using a parsimony approach ancestral structural states could be reconstructed. The distribution of number of contact changes per site for a fixed topology tree fit a gamma distribution. Homoplasy was high, with multiple changes per site and no support for an invariant class of residue-residue contacts. Contacts changed more slowly than sequence. A phylogenetic reconstruction using a distance measure based on the proportion of shared contacts was generally consistent with a sequence-based phylogeny but not highly resolved. Contact pattern convergence between members of different globin family proteins could not be detected. Simulation studies indicated the convergence test was sensitive enough to have detected convergence involving only 10% of the contacts, suggesting a limit on the extent of selection for a specific contact pattern. Contact site methods may provide additional approaches to study the relationship between protein structure and sequence evolution.

Amino Acids↗

Likelihood analysis of phylogenetic networks using directed graphical models.

A method for computing the likelihood of a set of sequences assuming a phylogenetic network as an evolutionary hypothesis is presented. The approach applies directed graphical models to sequence evolution on networks and is a natural generalization of earlier work by Felsenstein on evolutionary trees, including it as a special case. The likelihood computation involves several steps. First, the phylogenetic network is rooted to form a directed acyclic graph (DAG). Then, applying standard models for nucleotide/amino acid substitution, the DAG is converted into a Bayesian network from which the joint probability distribution involving all nodes of the network can be directly read. The joint probability is explicitly dependent on branch lengths and on recombination parameters (prior probability of a parent sequence). The likelihood of the data assuming no knowledge of hidden nodes is obtained by marginalization, i.e., by summing over all combinations of unknown states. As the number of terms increases exponentially with the number of hidden nodes, a Markov chain Monte Carlo procedure (Gibbs sampling) is used to accurately approximate the likelihood by summing over the most important states only. Investigating a human T-cell lymphotropic virus (HTLV) data set and optimizing both branch lengths and recombination parameters, we find that the likelihood of a corresponding phylogenetic network outperforms a set of competing evolutionary trees. In general, except for the case of a tree, the likelihood of a network will be dependent on the choice of the root, even if a reversible model of substitution is applied. Thus, the method also provides a way in which to root a phylogenetic network by choosing a node that produces a most likely network.

Computer Graphics↗

Optimal gene trees from sequences and species trees using a soft interpretation of parsimony.

Gene duplication and gene loss as well as other biological events can result in multiple copies of genes in a given species. Because of these gene duplication and loss dynamics, in addition to variation in sequence evolution and other sources of uncertainty, different gene trees ultimately present different evolutionary histories. All of this together results in gene trees that give different topologies from each other, making consensus species trees ambiguous in places. Other sources of data to generate species trees are also unable to provide completely resolved binary species trees. However, in addition to gene duplication events, speciation events have provided some underlying phylogenetic signal, enabling development of algorithms to characterize these processes. Therefore, a soft parsimony algorithm has been developed that enables the mapping of gene trees onto species trees and modification of uncertain or weakly supported branches based on minimizing the number of gene duplication and loss events implied by the tree. The algorithm also allows for rooting of unrooted trees and for removal of in-paralogues (lineage-specific duplicates and redundant sequences masquerading as such). The algorithm has also been made available for download as a software package, Softparsmap.

Algorithms↗

Phylogeny of the bee genus Halictus (Hymenoptera: halictidae) based on parsimony and likelihood analyses of nuclear EF-1alpha sequence data.

We investigated higher-level phylogenetic relationships within the genus Halictus based on parsimony and maximum likelihood (ML) analysis of elongation factor-1alpha DNA sequence data. Our data set includes 41 OTUs representing 35 species of halictine bees from a diverse sample of outgroup genera and from the three widely recognized subgenera of Halictus (Halictus s.s., Seladonia, and Vestitohalictus). We analyzed 1513 total aligned nucleotide sites spanning three exons and two introns. Equal-weights parsimony analysis of the overall data set yielded 144 equally parsimonious trees. Major conclusions supported in this analysis (and in all subsequent analyses) included the following: (1) Thrincohalictus is the sister group to Halictus s.l., (2) Halictus s.l. is monophyletic, (3) Vestitohalictus renders Seladonia paraphyletic but together Seladonia + Vestitohalictus is monophyletic, (4) Michener's Groups 1 and 3 are monophyletic, and (5) Michener's Group 1 renders Group 2 paraphyletic. In order to resolve basal relationships within Halictus we applied various weighting schemes under parsimony (successive approximations character weighting and implied weights) and employed ML under 17 models of sequence evolution. Weighted parsimony yielded conflicting results but, in general, supported the hypothesis that Seladonia + Vestitohalictus is sister to Michener's Group 3 and renders Halictus s.s. paraphyletic. ML analyses using the GTR model with site-specific rates supported an alternative hypothesis: Seladonia + Vestitohalictus is sister to Halictus s.s. We mapped social behavior onto trees obtained under ML and parsimony in order to reconstruct the likely historical pattern of social evolution. Our results are unambiguous: the ancestral state for the genus Halictus is eusociality. Reversal to solitary behavior has occurred at least four times among the species included in our analysis.

Animals↗

Similarity between putative ATP-binding sites in land plant plastid ORF2280 proteins and the FtsH/CDC48 family of ATPases.

Plastid ORF2280 proteins from five species of land plant are shown to have limited amino-acid sequence similarity to a family of proteins that includes the yeast CDC48, SEC18, PAS1 and SUG1 proteins, three subunits of the mammalian 26S protease, and the Escherichia coli FtsH protein. These proteins all contain one or two ATPase domains and many are involved in cell division, transport of proteins across membranes, or proteolysis. Similarity with the ORF2280 proteins is restricted to a single region of about 130 amino acids that contains: (1) sequences resembling a nucleotide binding site but lacking two normally conserved residues, and (2) a downstream conserved motif with the consensus sequence VIX2TX2PX3DPALX2P. Most of the rest of ORF2280 is very poorly conserved among land plants, even though other family members such as CDC48 have slow rates of protein sequence evolution. In contrast, a protein encoded by plastid DNA of the rhodophyte alga Porphyra purpurea is very similar to E. coli FtsH. Phylogenetic analysis suggests that the red and green plastid genes are not true homologues (orthologues) but distinct members of an ancient gene family.

ATP-Dependent Proteases↗

Switch to unusual amino acids at codon 215 of the human immunodeficiency virus type 1 reverse transcriptase gene in seroconvertors infected with zidovudine-resistant variants.

Sequences of the human immunodeficiency virus type 1 (HIV-1) reverse transcriptase (RT) domain were determined by direct sequencing of HIV-1 RNA in successive plasma samples from eight seroconverting patients infected with virus bearing the T215Y/F amino acid substitution associated with zidovudine (ZDV) resistance. At baseline, additional mutations associated with ZDV resistance were detected. Three patients had the M41L amino acid change, which persisted. Two patients had both the D67N and the K70R amino acid substitutions; reversion to the wild type was seen at both positions in one of these patients and at codon 70 in the other one. Reversion to the wild type at codon 215 was observed in only one of eight patients. Unusual amino acids, such as aspartic acid (D) and cysteine (C), appeared at position 215 in four patients during follow-up. These variants isolated by coculturing were sensitive to ZDV. Overgrowth of these variants suggests that they have better fitness than the original T215Y variant. Intraindividual nucleoside substitutions over time were 10 times more frequent in codons associated with ZDV resistance (41, 67, 70, 215, and 219) than in other codons of the RT domain. The predominance of nonsynonymous substitutions observed over time suggests that most changes reflect adaptation of the RT function. The variance in sequence evolution observed among patients, in particular at codon 215, supports a role for chance in the evolution of the RT domain.

Amino Acids↗

Evolutionary shift in the site of cleavage of prelysozyme.

Sequences are presented for the signal peptides of prelysozymes from 6 species of birds and compared to the known sequence for chicken prelysozyme c. The sequencing was done with synthetic oligonucleotides as primers and oviduct mRNA as the template, obviating the need to clone DNA from these species. Ring-necked pheasant prelysozyme c differs from all other prelysozymes c and pre-alpha-lactalbumins examined by being cleaved in vivo between amino acid residues 17 and 18 instead of between residues 18 and 19. The feature unique to the signal peptide of pheasant prelysozyme c is proline at position 17. Besides showing that proline is acceptable as the carboxyl-terminal amino acid of the signal peptide, our finding implies that it cannot occur as the penultimate amino acid in the signal peptide. This supports the view that unless a polypeptide has the proper secondary structure, signal peptidase will not cleave it, and that this secondary structure is a beta-turn. Another outcome of this comparative study is an estimate that the mean rate of sequence evolution in the prelysozyme signal peptide is 1%/two million years of divergence, similar to that calculated for the insulin signal peptide. Because this rate is a third of the silent substitution rate, it is likely that one out of every three amino acid substitutions is compatible with signal peptide function.

Amino Acid Sequence↗

The mitochondrial DNA of land plants: peculiarities in phylogenetic perspective.

Land plants exhibit a significant evolutionary plasticity in their mitochondrial DNA (mtDNA), which contrasts with the more conservative evolution of their chloroplast genomes. Frequent genomic rearrangements, the incorporation of foreign DNA from the nuclear and chloroplast genomes, an ongoing transfer of genes to the nucleus in recent evolutionary times and the disruption of gene continuity in introns or exons are the hallmarks of plant mtDNA, at least in flowering plants. Peculiarities of gene expression, most notably RNA editing and trans-splicing, are significantly more pronounced in land plant mitochondria than in chloroplasts. At the same time, mtDNA is generally the most slowly evolving of the three plant cell genomes on the sequence level, with unique exceptions in only some plant lineages. The slow sequence evolution and a variable occurrence of introns in plant mtDNA provide an attractive reservoir of phylogenetic information to trace the phylogeny of older land plant clades, which is as yet not fully resolved. This review attempts to summarize the unique aspects of land plant mitochondrial evolution from a phylogenetic perspective.

DNA, Mitochondrial↗