PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53Linked to original sources

Optimal alignments in linear space.

Space, not time, is often the limiting factor when computing optimal sequence alignments, and a number of recent papers in the biology literature have proposed space-saving strategies. However, a 1975 computer science paper by Hirschberg presented a method that is superior to the new proposals, both in theory and in practice. The goal of this paper is to give Hirschberg's idea the visibility it deserves by developing a linear-space version of Gotoh's algorithm, which accommodates affine gap penalties. A portable C-software package implementing this algorithm is available on the BIONET free of charge.

Algorithms↗

Conservation versus parallel gains in intron evolution.

Orthologous genes from distant eukaryotic species, e.g. animals and plants, share up to 25-30% intron positions. However, the relative contributions of evolutionary conservation and parallel gain of new introns into this pattern remain unknown. Here, the extent of independent insertion of introns in the same sites (parallel gain) in orthologous genes from phylogenetically distant eukaryotes is assessed within the framework of the protosplice site model. It is shown that protosplice sites are no more conserved during evolution of eukaryotic gene sequences than random sites. Simulation of intron insertion into protosplice sites with the observed protosplice site frequencies and intron densities shows that parallel gain can account but for a small fraction (5-10%) of shared intron positions in distantly related species. Thus, the presence of numerous introns in the same positions in orthologous genes from distant eukaryotes, such as animals, fungi and plants, appears to reflect mostly bona fide evolutionary conservation.

Animals↗

Insights into vertebrate evolution from the chicken genome sequence.

The chicken has recently joined the ever-growing list of fully sequenced animal genomes. Its unique features include expanded gene families involved in egg and feather production as well as more surprising large families, such as those for olfactory receptors. Comparisons with other vertebrate genomes move us closer to defining a set of essential vertebrate genes.

Animals↗

Evolution of satellite DNAs in a radiation of endemic Hawaiian spiders: does concerted evolution of highly repetitive sequences reflect evolutionary history?

Satellite DNAs are known for an unusual and nonuniform evolution characterized by rapid evolutionary change between species and concerted evolution leading to molecular homogeneity within species. In this paper we use satellite DNAs for phylogenetic analysis of a rapidly evolving lineage of spiders and compare the phylogeny with a hypothesis previously generated based on mitochondrial DNA and allozymes. The spiders examined include almost all species within a monophyletic clade of endemic Hawaiian Tetragnatha species, the spiny-leg clade. The phylogeny based on satellite sequences is largely congruent to those produced by mtDNA and allozymes, except that the satellite DNA yields much longer branches, with higher levels of support for any given node. Closely related species that have differentiated ecologically within an island are well resolved with satellite DNA but much less so with mtDNA. These results suggest that Tetragnatha stDNA repeats seem to be evolving gradually and cohesively during the diversification of these endemic Hawaiian spiders. The study also reveals gain-loss of satellite DNA copies during species diversification. We conclude that satellite DNA sequences may potentially be very useful for resolving relationships between rapidly evolving taxa within an adaptive radiation. In addition, satellite DNA as a nuclear marker suggests that hybridization or peripatry could play a possible role in species formation that cannot be revealed by mitochondrial markers due to its maternal inheritance.

Animals↗

Sequence of cowpea chlorotic mottle virus RNAs 2 and 3 and evidence of a recombination event during bromovirus evolution.

The genomic sequence of cowpea chlorotic mottle virus (CCMV) was completed by sequencing biologically active cDNA clones of CCMV RNA2 (2774 bases) and RNA3 (2173 bases). While only the central core of the encoded 94-kDa CCMV 2a protein contains features conserved among known and putative RNA replication proteins from many viruses, both flanking regions of CCMV 2a show substantial similarity to the corresponding protein of the related brome mosaic virus (BMV). The 3a proteins of CCMV and BMV, implicated as contributors to the distinct host specificities of the two viruses, show lower levels of conservation but are still discernibly related throughout. Major differences occur in the organization of noncoding sequences in CCMV and BMV RNA3. With respect to an otherwise similar region preceding the BMV 3a gene, the CCMV RNA3 5' noncoding sequence contains a clearly bounded 111-base insertion that must reflect a sequence rearrangement in evolution of at least one of the two viruses. The presence of a subgenomic promoter-like sequence near the end of the novel CCMV sequence makes the organization of genes in CCMV RNA3 reminiscent of the 3' end of tobacco mosaic virus RNA, suggesting that CCMV or its 3a gene might have been derived from an ancestor with fewer genomic RNAs. Sequence similarities between the CCMV and BMV RNA3 intercistronic regions include the subgenomic mRNA promoter and an oligo(A), but not an intercistronic segment required for BMV RNA3 amplification, implying that replication signals on the two RNA3s may be organized quite differently.

Amino Acid Sequence↗

Chicken apolipoprotein A-I: cDNA sequence, tissue expression and evolution.

Using an antibody against chicken apolipoprotein (apo) A-I, we identified multiple cDNA clones for the protein in two intestinal cDNA libraries in lambda gt11. The complete nucleotide sequence of chicken apoA-I cDNA was determined. The sequence predicts a mature protein of 240 amino acids, a 6-amino acid propeptide and an 18-amino acid signal peptide. Using a 32P-cDNA probe, we detected the presence of apoA-I mRNA in 21 day old chicken intestine, liver, kidney, spleen, breast muscle and brain. The primary sequence of apoA-I contains numerous tandem repeats of 11 and 22 residues in a manner similar to the mammalian proteins. Our analysis of apoA-I sequences from human, rabbit, dog, rat, and chicken indicates that the rate of amino acid substitution is considerably faster in the rat lineage than in other mammalian lineages.

Amino Acid Sequence↗

N-Terminal polyhedrin sequences and occluded Baculovirus evolution.

A phylogenetic tree for occluded baculoviruses was constructed based on the N-terminal amino acid sequence of occlusion body proteins from six baculoviruses including three lepidopteran nuclear polyhedrosis viruses (NPVs), [two unicapsid (Bombyx mori and Orgyia pseudotsugata) and one multicapsid (Orgyia pseudotsugata)]; one granulosis virus (Pieris brassicae); and NPVs from a hymenopteran (Neodiprion sertifer) and a dipteran (Tipula paludosa). Amino acid sequence data for the B. mori NPV were from a report by Serebryani et al. (1977) and that for the O. pseudotsugata NPVs were reported previously by us (Rohrmann et al. 1979). The other N-terminal amino acid sequences are presented in this paper. The phylogenetic relationships determined based on the molecular evolution of polyhedrin were also investigated by antigenic comparisons of the proteins using a solid phase radioimmune assay. The results indicate that the lepidopteran NPVs are the most closely related of the above group of viruses and are related to these viruses in the following order: N. sertifer NPV, P. brassicae granulosis virus, and T. paludosa NPV. These data, in conjunction with Baculovirus distribution and evidence concerning insect phylogeny, suggest that the Baculovirus have an ancient association with insects and may havae evolved along with them.

Amino Acid Sequence↗

Neutral theory of molecular evolution.

DNA sequence data are generally interpreted as favouring Kimura's neutral theory but not without dissent and often with a great deal of controversy with respect to molecular clocks, DNA polymorphism, adaptive evolution, and gene genealogy. Although the theory serves as a guiding principle, many issues concerning mutation, recombination, and selection remain unsettled. Of particular importance is the need for more knowledge about the function and structure of molecules.

Adaptation, Biological↗

DNA sequence arrangement and preliminary evidence on its evolution.

Some recent measurements of the sequence arrangement and evolution of the eukaryotic genome are reviewed. The range of genome sizes and extent of sequence transcribed into nuclear and messenger RNA indicate that the majority of the single copy DNA is not made up of structural genes. The rate of base substitution in the single copy DNA among the primates is similar to that of the codons for certain rapidly changing amino acid residues. This leads to the hypothesis that there is a "basal" rate of change in the genome not strongly affected by selection. The DNA of most higher animals shows a large amount of short period interspersion of repetitive and single copy DNA sequences and a smaller amount of long repetitive regions. The sequence divergence among the short interspersed repetitive sequences is greater than that of the sequences in long repetitive regions. The long repetitive regions are most probably recent additions to the genome and the short interspersed repetitive sequences result from a history of base substitution and translocation. The process of sequence rearrangement appears to be a significant part of the evolution of the genome and may have a much greater effect on the evolution of the phenotype than sequence alteration by base substitution.

Alleles↗

Codon reiteration and the evolution of proteins.

Sequence data banks have been searched for proteins possessing uninterrupted reiterations of any amino acid. Hydrophilic amino acids, and particularly glutamine, account for a large proportion of the longer reiterants. In the genes for these proteins, the most common reiterants are those that contain poly(CAG), even out-of-frame or, to a lesser degree, those that contain repeated doublets of CA, AG, or GC. The preferential generation of such reiterants requires that DNA strand-specific signals predispose to reiteration and thus to the extension of coding regions.

Amino Acid Sequence↗

Rooting the archaebacterial tree: the pivotal role of Thermococcus celer in archaebacterial evolution.

The sequence of the 16S ribosomal RNA gene from the archaebacterium Thermococcus celer shows the organism to be related to the methanogenic archaebacteria rather than to its phenotypic counterparts, the extremely thermophilic archaebacteria. This conclusion turns on the position of the root of the archaebacterial phylogenetic tree, however. The problems encountered in rooting this tree are analyzed in detail. Under conditions that suppress evolutionary noise both the parsimony and evolutionary distance methods yield a root location (using a number of eubacterial or eukaryotic outgroup sequences) that is consistent with that determined by an "internal rooting" method, based upon an (approximate) determination of relative evolutionary rates.

Archaea↗

Functional tuning of a salvaged green fluorescent protein variant with a new sequence space by directed evolution.

We previously reported a method, designated functional salvage screen (FSS), to generate protein lineages with new sequence spaces through the functional or structural salvage of a defective protein by employing green fluorescent protein (GFP) as a model protein. Here, in an attempt to mimic a step in the natural evolution process of proteins, the functionally salvaged mutant GFP-I5 with new sequence space, but showing low fluorescence intensity and stability, was selected and fine-tuned by directed evolution. During a course of functional tuning, GFP-I5 was found to evolve rapidly, recovering the spectral traits to those of the parent GFPuv. The mutant 3E4 from the third round of directed evolution possessed four substitutions; three (F64L, E111V and K166Q) were at the original GFP gene and the other (K8N) at the inserted segment. The fluorescence intensity of 3E4 was approximately 28-fold stronger than GFP-I5, and other spectral properties were retained. Biochemical and biophysical investigations suggested that the fine-tuning by directed evolution led the salvaged variant GFP-I5 to a functionally favorable structure, resulting in recovery of stability and fluorescence. Site-directed mutagenesis of the mutated amino acid residues in both GFPuv and GFP-I5 revealed that each amino acid residue has a different effect on the fluorescence intensity, which implies that 3E4 adopted a new evolutionary path with respect to fluorescence characteristics compared with the parent GFPuv. Directed evolution in conjunction with FSS is expected to be used for generating protein lineages with new fitness landscapes.

Directed Molecular Evolution↗

Origin of noncoding DNA sequences: molecular fossils of genome evolution.

The total amount of noncoding sequences on chromosomes of contemporary organisms varies significantly from species to species. We propose a hypothesis for the origin of these noncoding sequences that assumes that (i) an approximately equal to 0.55-kilobase (kb)-long reading frame composed the primordial gene and (ii) a 20-kb-long single-stranded polynucleotide is the longest molecule (as a genome) that was polymerized at random and without a specific template in the primordial soup/cell. The statistical distribution of stop codons allows examination of the probability of generating reading frames of approximately equal to 0.55 kb in this primordial polynucleotide. This analysis reveals that with three stop codons, a run of at least 0.55-kb equivalent length of nonstop codons would occur in 4.6% of 20-kb-long polynucleotide molecules. We attempt to estimate the total amount of noncoding sequences that would be present on the chromosomes of contemporary species assuming that present-day chromosomes retain the prototype primordial genome structure. Theoretical estimates thus obtained for most eukaryotes do not differ significantly from those reported for these specific organisms, with only a few exceptions. Furthermore, analysis of possible stop-codon distributions suggests that life on earth would not exist, at least in its present form, had two or four stop codons been selected early in evolution.

Animals↗

Nucleotide sequence, functional characterization and evolution of pFKN, a virulence plasmid in Pseudomonas syringae pathovar maculicola.

Pseudomonas syringae pv. maculicola strain M6 (Psm M6) carries the avrRpm1 gene, encoding a type III effector, on a 40 kb plasmid, pFKN. We hypothesized that this plasmid might carry additional genes required for pathogenesis on plants. We report the sequence and features of pFKN. In addition to avrRpm1, pFKN carries an allele of another type III effector, termed avrPphE, and a gene of unknown function (ORF8), expression of which is induced in planta, suggesting a role in the plant-pathogen interaction. The region of pFKN carrying avrRpm1, avrPphE and ORF8 exhibits several features of pathogenicity islands (PAIs). Curing of pFKN (creating Psm M6C) caused a significant reduction in virulence on Arabidopsis leaves. However, complementation studies using Psm M6C demonstrated an obvious virulence function only for avrRpm1. pFKN can integrate and excise from the chromosome of Psm M6 at low frequency via homologous recombination between identical sequence segments located on the chromosome and on pFKN. These segments are part of two nearly identical transposons carrying avrPphE. The avrPphE transposon was also detected in other strains of P. s. pv. maculicola and in P. s. tomato strain DC3000. The avrPphE transposon was found inserted at different loci in different strains. The analysis of sequences surrounding the avrPphE transposon insertion site in the chromosome of Psm M6 indicates that pFKN integrates into a PAI that encodes type III effectors. The integration of pFKN into this chromosomal region may therefore be seen as an evolutionary process determining the formation of a new PAI in the chromosome of Psm M6.

Arabidopsis↗

Long and short repeats of sea urchin DNA and their evolution.

Repeated sequences cloned from the DNA of the sea urchin S. purpuratus were used as probes to measure the lengths of individual families of repeats. Some probes reassociated much more rapidly with preparations of long repeats than with short repeats while others reassociated more rapidly with short repeats than with long repeats. In this way two of five cloned repeats were shown to represent families with a great majority of sequences in the long class. One represented a family with similar number of long and short class members. Two were members of predominantly short class families - The cloned repeats representing long class families, formed more precise duplexes than those representing short class families. Thermal stability measurements using S. purpuratus or S. franciscanus driver DNA showed that precise repetitive sequences have as great an interspecies sequence difference as the less precise repeats. Thus the precision of many families may result from recent multiplication rather than from selective pressure on the DNA sequences. Measurements of evolutionary frequency change show a clear correlation between the frequency change and the size of families of repeats in S. purpuratus. Comparison with S. franciscanus indicates that many of the large size families in S. purpuratus are those that have grown in size since these two species diverged.

Animals↗

Evolution of DNA or amino acid sequences with dependent sites.

A framework is outlined to study the evolution of DNA or amino acid sequences, if sequence sites do not evolve independently. The units of evolution are nonoverlapping subsequences of length l. Each subsequence evolves independently of the others, but within a subsequence the sequences show a Markov order one dependency. We describe an algorithm to mimic the evolution of such sequences. The influence of dependencies between sites on distance estimates and the reliability of tree reconstruction methods is investigated. We show that an inappropriate model of sequence evolution in the tree reconstruction process will lead to a nonempty Felsenstein zone. Finally, we describe a method to infer l from sequence data. Examples from the evolution of DNA sequences as well as from amino acids are given.

Algorithms↗