PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,351 records · Page 75Linked to original sources

Evolutionary genetics of the capsular locus of serogroup 6 pneumococci.

The evolution of the capsular biosynthetic (cps) locus of serogroup 6 Streptococcus pneumoniae was investigated by analyzing sequence variation within three serotype-specific cps genes from 102 serotype 6A and 6B isolates. Sequence variation within these cps genes was related to the genetic relatedness of the isolates, determined by multilocus sequence typing, and to the inferred patterns of recent evolutionary descent, explored using the eBURST algorithm. The serotype-specific cps genes had a low percent G+C, and there was a low level of sequence diversity in this region among serotype 6A and 6B isolates. There was also little sequence divergence between these serotypes, suggesting a single introduction of an ancestral cps sequence, followed by slight divergence to create serotypes 6A and 6B. A minority of serotype 6B isolates had cps sequences (class 2 sequences) that were approximately 5% divergent from those of other serotype 6B isolates (class 1 sequences) and which may have arisen by a second, more recent introduction from a related but distinct source. Expression of a serotype 6A or 6B capsule correlated perfectly with a single nonsynonymous polymorphism within wciP, the rhamnosyl transferase gene. In addition to ample evidence of the horizontal transfer of the serotype 6A and 6B cps locus into unrelated lineages, there was evidence for relatively frequent changes from serotype 6A to 6B, and vice versa, among very closely related isolates and examples of recent recombinational events between class 1 and 2 cps serogroup 6 sequences.

Amino Acid Sequence↗

Evolution of a triplet repeat in a conifer.

The opportunity to trace the evolution of a triplet repeat is rare, especially for seed-plant lineages with a well-defined fossil record. Microsatellite PtTX2133 sequences from 18 species in 2 conifer genera were used to calibrate the birth of a CAGn repeat, from its protomicrosatellite origins to its repeat expansion. Birth occurred in the hard-pine genome ~ 136 million years ago, or 14 million generations ago, then expanded as a polymorphic triplet repeat 136-100 million years before a major North American vicariance event. Calibration of the triplet-repeat birth and expansion is supported by the shared allelic lineages among Old and New World hard pines and the shared alleles solely among North American diploxylon or hard pines. Five CAGn repeat units appeared to be the expansion threshold for Old and New World diploxylon pines. Haploxylon pine species worldwide did not undergo birth and repeat expansion, remaining monomorphic, with a single imperfect 198-bp allele. A sister genus, Picea, had only a region of cryptic simplicity, preceding a proto-microsatellite region. The polymorphic triplet repeat in hard pines is older than some long-lived microsatellites reported for reptiles, yet younger than those reported for insects. Some cautionary points are raised about phylogenetic applications for this long-lived microsatellite.

Alleles↗

Positive selection dictates the choice between kinetic and thermodynamic protein folding and stability in subtilases.

Subtilisin E (SbtE) is a member of the ubiquitous superfamily of serine proteases called subtilases and serves as a model for understanding propeptide-mediated protein folding mechanisms. Unlike most proteins that adopt thermodynamically stable conformations, the native state of SbtE is trapped into a kinetically stable conformation. While kinetic stability offers distinct functional advantages to the native state, the constraints that dictate the selection between kinetic and thermodynamic folding and stability remain unknown. Using highly conserved subtilases, we demonstrate that adaptive evolution of sequence dictates selection of folding pathways. Intracellular and extracellular serine proteases (ISPs and ESPs, respectively) constitute two subfamilies within the family of subtilases that have highly conserved sequences, structures, and catalytic activities. Our studies on the folding pathways of subtilisin E (SbtE), an ESP, and its homologue intracellular serine protease 1 (ISP1), an ISP, show that although topology, contact order, and hydrophobicity that drive protein folding reactions are conserved, ISP1 and SbtE fold through significantly different pathways and kinetics. While SbtE absolutely requires the propeptide to fold into a kinetically trapped conformer, ISP1 folds to a thermodynamically stable state more than 1 million times faster and independent of a propeptide. Furthermore, kinetics establish that ISP1 and SbtE fold through different intermediate states. An evolutionary analysis of folding constraints in subtilases suggests that observed differences in folding pathways may be mediated through positive selection of specific residues that map mostly onto the protein surface. Together, our results demonstrate that closely related subtilases can fold through distinct pathways and mechanisms, and suggest that fine sequence details can dictate the choice between kinetic and thermodynamic folding and stability.

Amino Acid Sequence↗

Hemagglutinin sequence clusters and the antigenic evolution of influenza A virus.

Continual mutations to the hemagglutinin (HA) gene of influenza A virus generate novel antigenic strains that cause annual epidemics. Using a database of 560 viral RNA sequences, we study the structure and tempo of HA evolution over the past two decades. We detect a critical length scale, in amino acid space, at which HA sequences aggregate into clusters, or swarms. We investigate the spatio-temporal distribution of viral swarms and compare it to the time series of the influenza vaccines recommended by the World Health Organization. We introduce a method for predicting future dominant HA amino acid sequences and discuss its potential relevance to vaccine choice. We also investigate the relationship between cluster structure and the primary antibody-combining regions of the HA protein.

Antibodies↗

Novel phytochrome sequences in Arabidopsis thaliana: structure, evolution, and differential expression of a plant regulatory photoreceptor family.

Phytochrome is a plant regulatory photoreceptor that mediates red light effects on a wide variety of physiological and molecular responses. DNA blot analysis indicates that the Arabidopsis thaliana genome contains four to five phytochrome-related gene sequences. We have isolated and sequenced cDNA clones corresponding to three of these genes and have deduced the amino acid sequence of the full-length polypeptide encoded in each case. One of these proteins (phyA) shows 65-80% amino acid sequence identity with the major, etiolated-tissue phytochrome apoproteins described previously in other plant species. The other two polypeptides (phyB and phyC) are unique in that they have low sequence identity (approximately 50%) with each other, with phyA, and with all previously described phytochromes. The phyA, phyB, and phyC proteins are of similar molecular mass, have related hydropathic profiles, and contain a conserved chromophore attachment region. However, the sequence comparison data indicate that the three phy genes diverged early in plant evolution, well before the divergence of the two major groups of angiosperms, the monocots and dicots. The steady-state level of the phyA transcript is high in dark-grown A. thaliana seedlings and is down-regulated by light. In contrast, the phyB and phyC transcripts are present at lower levels and are not strongly light-regulated. These findings indicate that the red/far light-responsive phytochrome photoreceptor system in A. thaliana, and perhaps in all higher plants, consists of a family of chromoproteins that are heterogeneous in structure and regulation.

Amino Acid Sequence↗

Point mutations with positive selection were a major force during the evolution of a receptor-kinase resistance gene family of rice.

The rice (Oryza sativa) Xa26 gene, which confers resistance to bacterial blight disease and encodes a leucine-rich repeat (LRR) receptor kinase, resides at a locus clustered with tandem homologous genes. To investigate the evolution of this family, four haplotypes from the two subspecies of rice, indica and japonica, were analyzed. Comparative sequence analysis of 34 genes of 10 types of paralogs of the family revealed haplotype polymorphisms and pronounced paralog diversity. The orthologs in different haplotypes were more similar than the paralogs in the same haplotype. At least five types of paralogs were formed before the separation of indica and japonica subspecies. Only 7% of amino acid sites were detected to be under positive selection, which occurred in the extracytoplasmic domain. Approximately 74% of the positively selected sites were solvent-exposed amino acid residues of the LRR domain that have been proposed to be involved in pathogen recognition, and 73% of the hypervariable sites detected in the LRR domain were subject to positive selection. The family is formed by tandem duplication followed by diversification through recombination, deletion, and point mutation. Most variation among genes in the family is caused by point mutations and positive selection.

Amino Acid Sequence↗

Homology-dependent gene silencing and host defense in plants.

Analyses of transgene silencing phenomena in plants and other organisms have revealed the existence of epigenetic silencing mechanisms that are based on recognition of nucleic acid sequence homology at either the DNA or RNA level. Common triggers of homology-dependent gene silencing include inverted DNA repeats and double-stranded RNA, a versatile silencing molecule that can induce both degradation of homologous RNA in the cytoplasm and methylation of homologous DNA sequences in the nucleus. Inverted repeats might be frequently associated with silencing because they can potentially interact in cis and in trans to trigger DNA methylation via homologous DNA pairing, or they can be transcribed to produce double-stranded RNA. Homology-dependent gene silencing mechanisms are ideally suited for countering natural parasitic sequences such as transposable elements and viruses, which are usually present in multiple copies and/or produce double-stranded RNA during replication. These silencing mechanisms can thus be regarded as host defense strategies to foreign or invasive nucleic acids. The high content of transposable elements and, in some cases, endogenous viruses in many plant genomes suggests that host defenses do not always prevail over invasive sequences. During evolution, slightly faulty genome defense responses probably allowed transposable elements and viral sequences to accumulate gradually in host chromosomes and to invade host genes. Possible beneficial consequences of this "foreign" DNA buildup include the establishment of genome defense-derived epigenetic control mechanisms for regulating host gene expression and acquired hereditary immunity to some viruses.

Animals↗

Molecular evidence on the origin and evolution of glutinous rice.

Glutinous rice is a major type of cultivated rice with long-standing cultural importance in Asia. A mutation in an intron 1 splice donor site of the Waxy gene is responsible for the change in endosperm starch leading to the glutinous phenotype. Here we examine an allele genealogy of the Waxy locus to trace the evolutionary and geographical origins of this phenotype. On the basis of 105 glutinous and nonglutinous landraces from across Asia, we find evidence that the splice donor mutation has a single evolutionary origin and that it probably arose in Southeast Asia. Nucleotide diversity measures indicate that the origin of glutinous rice is associated with reduced genetic variation characteristic of selection at the Waxy locus; comparison with an unlinked locus, RGRC2, confirms that this pattern is specific to Waxy. In addition, we find that many nonglutinous varieties in Northeast Asia also carry the splice donor site mutation, suggesting that partial suppression of this mutation may have played an important role in the development of Northeast Asian nonglutinous rice. This study demonstrates the utility of phylogeographic approaches for understanding trait diversification in crops, and it contributes to growing evidence on the importance of modifier loci in the evolution of domestication traits.

Base Sequence↗

Biosynthesis of isoprenoids via mevalonate in Archaea: the lost pathway.

Isoprenoid compounds are ubiquitous in living species and diverse in biological function. Isoprenoid side chains of the membrane lipids are biochemical markers distinguishing archaea from the rest of living forms. The mevalonate pathway of isoprenoid biosynthesis has been defined completely in yeast, while the alternative, deoxy-D-xylulose phosphate synthase pathway is found in many bacteria. In archaea, some enzymes of the mevalonate pathway are found, but the orthologs of three yeast proteins, accounting for the route from phosphomevalonate to geranyl pyrophosphate, are missing, as are the enzymes from the alternative pathway. To understand the evolution of isoprenoid biosynthesis, as well as the mechanism of lipid biosynthesis in archaea, sequence motifs in the known enzymes of the two pathways of isoprenoid biosynthesis were analyzed. New sequence relationships were detected, including similarities between diphosphomevalonate decarboxylase and kinases of the galactokinase superfamily, between the metazoan phosphomevalonate kinase and the nucleoside monophosphate kinase superfamily, and between isopentenyl pyrophosphate isomerases and MutT pyrophosphohydrolases. Based on these findings, orphan members of the galactokinase, nucleoside monophosphate kinase, and pyrophosphohydrolase families in archaeal genomes were evaluated as candidate enzymes for the three missing steps. Alternative methods of finding these missing links were explored, including physical linkage of open reading frames and patterns of ortholog distribution in different species. Combining these approaches resulted in the generation of a short list of 13 candidate genes for the three missing functions in archaea, whose participation in isoprenoid biosynthesis is amenable to biochemical and genetic investigation.

Amino Acid Motifs↗

The evolution of defective and autonomous parvoviruses.

Because of the small size and genetic simplicity of small DNA viruses, parvoviruses would appear to be excellent models for studying viral evolution and adaptation. In an earlier publication we hypothesized the evolution of sequences of cellular "junk" DNA into protective interfering transposons. These transposons would interfere with invading pathogenic viruses by competing with the pathogen DNA for replicative enzymes. We speculated that a small, defective parvovirus, the adeno-associated virus (AAV), which usually requires the presence of a pathogenic helper virus to replicate, may have evolved from such a piece of cellular "junk" DNA. Our theory predicted that AAVs, as a consequence of their defective nature, developed under pressures favoring maintenance of their transposon like qualities. In contrast, disease-causing, autonomous, non-defective parvoviruses such as the B19 agent of humans and the canine parvovirus, even though their origins may have been in cellular DNA, would appear to have developed under totally different evolutionary pressures. In this paper we will present evidence for a common ancestry for the defective and autonomous parvoviruses and discuss the divergent paths this evolution may have taken in establishing the two genera.

Base Sequence↗

Evolution of a multigene family that encodes the Kunitz chymotrypsin inhibitor in winged bean: a possible intermediate in the generation of a new gene with a distinct pattern of expression.

Winged bean Kunitz chymotrypsin inhibitor (WCI) accumulates in an organ-specific and temporally regulated manner. The protein is encoded by a multigene family that includes at least four putative inhibitor-coding genes and three pseudogenes. The structure of the WCI genes indicates that an insertion at a 5' proximal site occurred after duplication of the ancestral WCI gene and that several gene conversion events subsequently contributed to the evolution of this gene family. Analysis of the promoter activity of the 5' regions of the WCI genes in transgenic tobacco showed that only the 5' regions of the WCI-3a and WCI-3b genes, which encode the major WCI protein in winged bean, promoted the organ-specific and temporally regulated expression of a reporter gene. The 5' region of a pseudogene, the WCI-P1 gene which contains frameshift mutations, exhibited constitutive promoter activity in tobacco, an indication that the 5' region of the WCI-P1 gene might spontaneously have acquired new regulatory sequences during evolution. Since gene conversion is a relatively frequent event and since the homology between the WCI-P1 and WCI-3a/b genes is disrupted at a 5' proximal site by remnants of an inserted sequence, the WCI-P1 gene appears to be a possible intermediate that could be converted into a new functional gene with a distinct pattern of expression by a single gene-conversion event.

Base Sequence↗

Variation in the ribosomal internal transcribed spacers and 5.8S rDNA among five species of Acropora (Cnidaria; Scleractinia): patterns of variation consistent with reticulate evolution.

The ITS sequences of Acropora spp. are the shortest so far identified in any metazoan and are among the shortest seen in eukaryotes; ITS1 was 70-80 bases, and ITS2 was 100-112 bases. The ITS sequences were also highly variable, but base composition and secondary structure prediction indicate that divergent sequence variants are unlikely to be pseudogenes. The pattern of variation was unusual in several other respects: (1) two distinct ITS2 types were detected in both A. hyacinthus and A. cytherea, species known to hybridize in vitro with high success rates, and a putative intermediate ITS2 form was also detected in A. cytherea; (2) A. valida was found to contain highly (29%) diverged ITS1 variants; and (3) A. longicyathus contained two distinct 5.8S rDNA types. These data are consistent with a reticulate evolutionary history for the genus Acropora.

Animals↗

Arthropod and mollusk defensins--evolution by exon-shuffling.

Arthropod and mollusk defensins are secreted antibacterial proteins that exhibit similarity in sequence, mode of action and structure and are expressed ubiquitously. Comparison of the gene organization of a newly cloned scorpion defensin gene, with that of other arthropods and the mussel, revealed that all exons and introns, aside from the exon encoding the mature protein, differ widely in number, size and sequence. This variability suggests that the exon encoding the mature defensin has undergone exon-shuffling and integrated downstream of unrelated leader sequences during evolution. Unlike other exon-shuffling events, in which modules are added into existing proteins, arthropod and mollusk defensins represent the first instance of exon-shuffling of autonomous modules.

Animals↗

The plasticity of immunoglobulin gene systems in evolution.

The mechanism of recombination-activating gene (RAG)-mediated rearrangement exists in all jawed vertebrates, but the organization and structure of immunoglobulin (Ig) genes, as they differ in fish and among fish species, reveal their capability for rapid evolution. In systems where there can exist 100 Ig loci, exon restructuring and sequence changes of the constant regions led to divergence of effector functions. Recombination among these loci created hybrid genes, the strangest of which encode variable (V) regions that function as part of secreted molecules and, as the result of an ancient translocation, are also grafted onto the T-cell receptor. Genomic changes in V-gene structure, created by RAG recombinase acting on germline recombination signal sequences, led variously to the generation of fixed receptor specificities, pseudogene templates for gene conversion, and ultimately to Ig sequences that evolved away from Ig function. The presence of so many Ig loci in fishes raises interesting questions not only as to how their regulation is achieved but also how successive whole-locus duplications are accommodated by a system whose function in other vertebrates is based on clonal antigen receptor expression.

Animals↗

The random character of protein evolution and its effects on the reliability of phylogenetic information deduced from amino acid sequences and compositions.

Because evolution occurs by random events, the actual number of substitutions that occur in any period is not exactly equal to the number expected from the mean rate of substitution, but is statistically distributed about it. In consequence, even if rates of evolution are constant in different lineages, 'trees' deduced from descendant protein sequences contain random errors. When there are fewer than about eight differences between the sequences of the most distantly related pair from a set of proteins, this random effect is very large. It can then render trivial the statistical disadvantage inherent in using a crude measure of protein difference, such as amino acid composition or immunological cross-reactivity, in preference to a measure based the sequences of the most distantly related pair from a set of proteins, this random effect is very large. It can then render trivial the statistical disadvantage inherent in using a crude measure of protein difference, such as amino acid composition or immunological cross-reactivity, in preference to a measure based the sequences of the most distantly related pair from a set of proteins, this random effect is very large. It can then render trivial the statistical disadvantage inherent in using a crude measure of protein difference, such as amino acid composition or immunological cross-reactivity, in preference to a measure based on amino acid sequence. In some cases, such as classification of mammals on the basis of cytochrome c structure, it appears to make little difference to the reliability of the results whether the sequences of the protein concerned are known or not. It may also be possible to obtain more reliable phylogenetic information from composition measurements on several kinds of protein than one could obtain from sequence measurements on a single kind of protein.

Amino Acid Sequence↗

Identification of an IL-8 homolog in lamprey (Lampetra fluviatilis): early evolutionary divergence of chemokines.

Subtractive hybridization was used to study river lamprey (Lampetra fluviatilis) leukocyte-specific cDNA. A clone representing the most abundant component (12%) of the leukocyte library subtracted with liver cDNA was isolated and characterized. The cDNA encodes a presumably secreted polypeptide of 101 residues. The 3' untranslated region of the cDNA contains motifs characteristic of the transiently expressing genes. Comparison of the deduced amino acid sequence with known protein sequences revealed its homology to the members of the chemokine superfamily. Designated as LFCA-1, the lamprey protein contains four conserved cysteines, of which the first two are separated by a residue, and a number of other CXC family characteristic residues. LFCA-1 has the highest similarity to the chicken EMF-1 (40%) and to the mammalian IL-8 (32-33%). However, it lacks the ELR motif essential for the function of the mammalian IL-8-related chemokines. Based on the phylogenetic analysis of the LFCA-1 relationship to the higher vertebrate chemokines, it is concluded that the evolutionary origin of the chemokine superfamily is ancient, and that the divergence of the CXC and CC families most likely occurred at the time or before the first vertebrates emerged.

Amino Acid Sequence↗

An improved algorithm for statistical alignment of sequences related by a star tree.

The insertion-deletion model developed by Thorne, Kishino and Felsenstein (1991, J. Mol. Evol., 33, 114-124; the TKF91 model) provides a statistical framework of two sequences. The statistical alignment of a set of sequences related by a star tree is a generalization of this model. The known algorithm computes the probability of a set of such sequences in O(l2k) time, where l is the geometric mean of the sequence lengths and k is the number of sequences. An improved algorithm is presented whose running time is only O(2(2k)lk).

Algorithms↗

Fold prediction and evolutionary analysis of the POZ domain: structural and evolutionary relationship with the potassium channel tetramerization domain.

Using iterative database searches, a statistically significant sequence similarity was detected between the POZ (poxvirus and zinc finger) domains found in a variety of proteins involved in animal transcription regulation, cytoskeleton organization, and development, and the tetramerization domain of animal potassium channels. Using the crystal structure of the Aplysia Shaker channel tetramerization domain as a template, the common structure of the POZ domain class was predicted. Examination of the structure resulted in the identification of several structural features and specific amino acid residues that may be involved in conserved protein-protein interactions mediated by the POZ domains as well as those that may contribute to the specificity of these interactions. Phylogenetic analysis of the POZ domains suggests that the common ancestor of the crown group eukaryotes already possessed this domain; POZ domains have undergone independent expansion in plants and in different animal lineages.

Amino Acid Sequence↗