PubMed HealthSearch

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Variance to mean ratio, R(t), for poisson processes on phylogenetic trees.

The ratio of expected variance to mean, R(t), of numbers of DNA base substitutions for contemporary sequences related by a "star" phylogeny is widely seen as a measure of the adherence of the sequences' evolution to a Poisson process with a molecular clock, as predicted by the "neutral theory" of molecular evolution under certain conditions. A number of estimators of R(t) have been proposed, all predicted to have mean 1 and distributions based on the chi 2. Various genes have previously been analyzed and found to have values of R(t) far in excess of 1, calling into question important aspects of the neutral theory. In this paper, I use Monte Carlo simulation to show that the previously suggested means and distributions of estimators of R(t) are highly inaccurate. The analysis is applied to star phylogenies and to general phylogenetic trees, and well-known gene sequences are reanalyzed. For star phylogenies the results show that Kimura's estimators ("The Neutral Theory of Molecular Evolution," Cambridge Univ. Press, Cambridge, 1983) are unsatisfactory for statistical testing of R(t), but confirm the accuracy of Bulmer's correction factor (Genetics 123: 615-619, 1989). For all three nonstar phylogenies studied, attained values of all three estimators of R(t), although larger than 1, are within their true confidence limits under simple Poisson process models. This shows that lineage effects can be responsible for high estimates of R(t), restoring some limited confidence in the molecular clock and showing that the distinction between lineage and molecular clock effects is vital.(ABSTRACT TRUNCATED AT 250 WORDS)

Analysis of Variance

Intra-Host Evolution Provides for the Continuous Emergence of SARS-CoV-2 Variants.

Variants of concern (VOC) in SARS-CoV-2 refer to viruses whose viral genomes differ from the ancestor virus by ≥3 single-nucleotide variants (SNVs) and that show the potential for higher transmissibility and/or worse clinical progression. VOC have the potential to disrupt ongoing public health measures and vaccine efforts. Still, too little is known regarding how frequently new viral variants emerge and under what circumstances. We report a study to determine the degree of SARS-CoV-2 sequence evolution in 94 patients and to estimate the frequency at which highly diverse variants emerge. Two cases accumulated ≥9 SNVs over a 2-week period and one case accumulated 23 SNVs over 3 weeks, including three nonsynonymous mutations in the spike protein (D138H, E554D, D614G). The remainder of the infected patients did not show signs of intra-host evolution. We estimate that in as much as 2% of hospitalized COVID-19 cases, variants with multiple mutations in the spike glycoprotein emerge in as little as 1 month of persistent intra-host virus replication. This suggests the continued local emergence of variants with multiple nonsynonymous SNVs, even in patients without overt immune deficiency. Surveillance by sequencing for (i) viremic COVID-19 patients, (ii) patients suspected of reinfection, and (iii) patients with diminished immune function may offer broad public health benefits. IMPORTANCE New SARS-CoV-2 variants can potentially disrupt ongoing public health measures and vaccine efforts. Still, little is known regarding how frequently new viral variants emerge and under what circumstances. Based on this study, we estimate that in hospitalized COVID-19 cases, variants with multiple mutations may emerge locally in as little as 1 month, even in patients without overt immune deficiency. Surveillance by sequencing for continuously shedding patients, patients suspected of reinfection, and patients with diminished immune function may offer broad public health benefits.

Humans

Nucleotide sequence and molecular evolution of mouse retrovirus-like IAP elements.

We determined the nucleotide (nt) sequences of cDNA and genomic clones for murine intracisternal type A particle (IAP) elements, which are retrovirus-like repetitive sequences in rodent genomes. The nucleotide sequence of the cDNA resembled that of retrovirus RNA genomes in its lack of the U5 sequence within the 3' long terminal repeat. By sequence comparison of our clones with reported rodent IAP elements, we located the probable gag, pol and env gene regions. The sequences for the pol, env and the 3' two-thirds of the gag region were conserved among the IAP elements. In the regions, synonymous substitutions occurred more frequently than non-synonymous ones, which suggested that the regions in question were functionally constrained until fairly recently. The rate of nucleotide substitutions in the regions was estimated to be 6-10 X 10(-9) nt per site per year, and significantly higher than that of the cellular genes. These rates may exemplify a characteristic of the nucleotide substitutions for an endogenous retrovirus. The sequence homology between the IAP element and IgE-binding factor gene is discussed.

Animals

Two distinct ferredoxins from Rhodobacter capsulatus: complete amino acid sequences and molecular evolution.

Two distinct ferredoxins were purified from Rhodobacter capsulatus SB1003. Their complete amino acid sequences were determined by a combination of protease digestion, BrCN cleavage and Edman degradation. Ferredoxins I and II were composed of 64 and 111 amino acids, respectively, with molecular weights of 6,728 and 12,549 excluding iron and sulfur atoms. Both contained two Cys clusters in their amino acid sequences. The first cluster of ferredoxin I and the second cluster of ferredoxin II had a sequence, CxxCxxCxxxCP, in common with the ferredoxins found in Clostridia. The second cluster of ferredoxin I had a sequence, CxxCxxxxxxxxCxxxCM, with extra amino acids between the second and third Cys, which has been reported for other photosynthetic bacterial ferredoxins and putative ferredoxins (nif-gene products) from nitrogen-fixing bacteria, and with a unique occurrence of Met. The first cluster of ferredoxin II had a CxxCxxxxCxxxCP sequence, with two additional amino acids between the second and third Cys, a characteristics feature of Azotobacter-[3Fe-4S] [4Fe-4S]-ferredoxin. Ferredoxin II was also similar to Azotobacter-type ferredoxins with an extended carboxyl (C-) terminal sequence compared to the common Clostridium-type. The evolutionary relationship of the two together with a putative one recently found to be encoded in nifENXQ region in this bacterium [Moreno-Vivian et al. (1989) J. Bacteriol. 171, 2591-2598] is discussed.

Amino Acid Sequence

A likelihood approach for comparing synonymous and nonsynonymous nucleotide substitution rates, with application to the chloroplast genome.

A model of DNA sequence evolution applicable to coding regions is presented. This represents the first evolutionary model that accounts for dependencies among nucleotides within a codon. The model uses the codon, as opposed to the nucleotide, as the unit of evolution, and is parameterized in terms of synonymous and nonsynonymous nucleotide substitution rates. One of the model's advantages over those used in methods for estimating synonymous and nonsynonymous substitution rates is that it completely corrects for multiple hits at a codon, rather than taking a parsimony approach and considering only pathways of minimum change between homologous codons. Likelihood-ratio versions of the relative-rate test are constructed and applied to data from the complete chloroplast DNA sequences of Oryza sativa, Nicotiana tabacum, and Marchantia polymorpha. Results of these tests confirm previous findings that substitution rates in the chloroplast genome are subject to both lineage-specific and locus-specific effects. Additionally, the new tests suggest tha the rate heterogeneity is due primarily to differences in nonsynonymous substitution rates. Simulations help confirm previous suggestions that silent sites are saturated, leaving no evidence of heterogeneity in synonymous substitution rates.

Chloroplasts

Similarity between putative ATP-binding sites in land plant plastid ORF2280 proteins and the FtsH/CDC48 family of ATPases.

Plastid ORF2280 proteins from five species of land plant are shown to have limited amino-acid sequence similarity to a family of proteins that includes the yeast CDC48, SEC18, PAS1 and SUG1 proteins, three subunits of the mammalian 26S protease, and the Escherichia coli FtsH protein. These proteins all contain one or two ATPase domains and many are involved in cell division, transport of proteins across membranes, or proteolysis. Similarity with the ORF2280 proteins is restricted to a single region of about 130 amino acids that contains: (1) sequences resembling a nucleotide binding site but lacking two normally conserved residues, and (2) a downstream conserved motif with the consensus sequence VIX2TX2PX3DPALX2P. Most of the rest of ORF2280 is very poorly conserved among land plants, even though other family members such as CDC48 have slow rates of protein sequence evolution. In contrast, a protein encoded by plastid DNA of the rhodophyte alga Porphyra purpurea is very similar to E. coli FtsH. Phylogenetic analysis suggests that the red and green plastid genes are not true homologues (orthologues) but distinct members of an ancient gene family.

ATP-Dependent Proteases

Evolutionary shift in the site of cleavage of prelysozyme.

Sequences are presented for the signal peptides of prelysozymes from 6 species of birds and compared to the known sequence for chicken prelysozyme c. The sequencing was done with synthetic oligonucleotides as primers and oviduct mRNA as the template, obviating the need to clone DNA from these species. Ring-necked pheasant prelysozyme c differs from all other prelysozymes c and pre-alpha-lactalbumins examined by being cleaved in vivo between amino acid residues 17 and 18 instead of between residues 18 and 19. The feature unique to the signal peptide of pheasant prelysozyme c is proline at position 17. Besides showing that proline is acceptable as the carboxyl-terminal amino acid of the signal peptide, our finding implies that it cannot occur as the penultimate amino acid in the signal peptide. This supports the view that unless a polypeptide has the proper secondary structure, signal peptidase will not cleave it, and that this secondary structure is a beta-turn. Another outcome of this comparative study is an estimate that the mean rate of sequence evolution in the prelysozyme signal peptide is 1%/two million years of divergence, similar to that calculated for the insulin signal peptide. Because this rate is a third of the silent substitution rate, it is likely that one out of every three amino acid substitutions is compatible with signal peptide function.

Amino Acid Sequence

Evolutionary aspects of urea cycle enzyme genes.

The functions and expression pattern of urea cycle enzymes have undergone considerable changes during the course of evolution. Sequence analyses shows that urea cycle enzymes from mammals are homologous to microbial enzymes of the arginine-metabolic pathway. Recently, an unexpected relationship was found between argininosuccinate lyase (EC 4.3.2.1), the fourth enzyme of the cycle, and delta-crystallin, a lens structural protein of birds and reptiles.

Animals

De novo Genes in Plants: Origins, Mechanisms, and Functional Implications.

De novo genes originate from previously non-coding genomic regions. They provide an important source of lineage-specific innovation. In plants, these genes may contribute to adaptation, trait diversity and crop evolution. This review summarizes recent progress in plant de novo gene research. It first discusses major routes of gene birth, including transcription-first, open reading frame (ORF)-first and concurrent models. It also examines how nascent loci acquire regulatory control and enter existing biological networks. The review then summarizes their evolutionary features, including weak early constraint, rapid molecular change, restricted expression and structural refinement. It further discusses plant de novo genes involved in stress responses, seed germination, kernel dehydration, subspecies divergence, reproductive isolation and floral scent diversification. Current methods for identifying de novo genes remain limited by rapid sequence evolution, genome annotation quality, polyploidy and transposable elements. Whole-genome synteny alignment, multi-omics evidence and machine-learning approaches can improve candidate discovery. However, each method has important limitations. Finally, this review highlights key future questions in functional validation, latent coding potential in long non-coding RNAs, epigenetic activation, regulatory-network integration and crop improvement. These perspectives clarify how de novo genes shape plant adaptation and how they may be used in precision breeding and synthetic biology.

adaptive evolution

New approaches to dating suggest a recent age for the human mtDNA ancestor.

The most critical and controversial feature of the African origin hypothesis of human mitochondrial DNA (mtDNA) evolution is the relatively recent age of about 200 ka inferred for the human mtDNA ancestor. If this age is wrong, and the actual age instead approaches 1 million years ago, then the controversy abates. Reliable estimates of the age of the human mtDNA ancestor and the associated standard error are therefore crucial. However, more recent estimates of the age of the human ancestor rely on comparisons between human and chimpanzee mtDNAs that may not be reliable and for which standard errors are difficult to calculate. We present here two approaches for deriving an intraspecific calibration of the rate of human mtDNA sequence evolution that allow standard errors to be readily calculated. The estimates resulting from these two approaches for the age of the human mtDNA ancestor (and approximate 95% confidence intervals) are 133 (63-356) and 137 (63-416) ka ago. These results provide the strongest evidence yet for a relatively recent origin of the human mtDNA ancestor.

Animals

Differential pattern of sequence heterogeneity in the hepatitis C virus E1 and E2/NS1 proteins.

The E1 and E2/NS1 genes, encoding the putative hepatitis C virus envelope proteins, show a high rate of sequence variations. We analyzed the degree and distribution of sequence heterogeneity in serum samples from hepatitis C virus-infected subjects. The mutations in the E1 region were mainly type-specific and the rate of variability was apparently not linked to the clinical phase of the infection. The sequence evolution of the E1 region during interferon treatment was low, regardless of the response to therapy. In contrast, an increased degree of variation, apparently related to the stage of viral replication, was present in E2 region derived from patients undergoing interferon treatment. These results are consistent with the hypothesis that the E2 protein represents a major target of the immune response.

Adolescent

Evolution and isoforms of V-ATPase subunits.

The structure of V- and F-ATPases/ATP synthases is remarkably conserved throughout evolution. Sequence analyses show that the V- and F-ATPases evolved from the same enzyme that was already present in the last common ancestor of all known extant life forms. The catalytic and non-catalytic subunits found in the dissociable head groups of both V-ATPases and F-ATPases are paralogous subunits, i.e. these two types of subunits evolved from a common ancestral gene. The gene duplication giving rise to these two genes (i.e. those encoding the catalytic and non-catalytic subunits) pre-dates the time of the last common ancestor. Similarities between the V- and F-ATPase subunits and an ATPase-like protein that is implicated in flagellar assembly are evaluated with regard to the early evolution of ATPases. Mapping of gene duplication events that occurred in the evolution of the proteolipid, the non-catalytic and the catalytic subunits onto the tree of life leads to a prediction of the likely quaternary structure of the encoded ATPases. The phylogenetic implications of V-ATPases found in eubacteria are discussed. Different V-ATPase isoforms have been detected in some higher eukaryotes, whereas others were shown to have only a single gene encoding the catalytic V-ATPase subunit. These data are analyzed with respect to the possible function of the different isoforms (tissue-specific, organelle-specific). The point in evolution at which the different isoforms arose is mapped by phylogenetic analysis.

Adenosine Triphosphatases

Evolution of the mouse H-2K region: a hot spot of mutation associated with genes transcribed in embryos and/or germ cells.

Active gene transcription is known to promote genetic change in neighboring DNA. We reasoned that the change would be readily heritable if transcription was occurring in germ cells or early embryonic cells before the germ cells are set aside. The H-2K region of the major histocompatibility complex (MHC) provides a good vehicle for testing this hypothesis because it is replete with such genes. We have compared the amount of polymorphism in 240 kb of DNA contiguous with H-2K and 150 kb of DNA flanking a homologous duplicated region in t-haplotypes and inbred strains. Using 90 probes and three restriction enzymes, we find a staggering difference in the amount of polymorphism in the H-2K region vs. the duplicated region (26% vs. 0%) of t-haplotypes. The disparity in the rate of divergence between the two regions indicates that the spatial distribution of genes and their expression pattern might be important factors in sequence evolution. Since t-haplotypes normally show extremely limited variability among themselves due to their recent divergence from a single ancestor, these results imply that the mutation rate in the H-2K region is unusually high. This is in apparent contradiction to the current view that the MHC loci have evolved at the same rate as other loci. The implications for the evolution of the H-2K gene are discussed.

Animals

On the possibility of the production of many rare proteins by higher eukaryotes.

Arguments are presented which support the idea that the large amount of single copy DNA in the haploid genome of higher eukaryotes is coding for many (approximately 10(6) rare proteins, that is, proteins present at the level of no more than a few molecules per cell. It is asserted that a large number of such rare protein species is consistent with (1) apparent slow rates of DNA sequence evolution as shown by DNA annealing experiments, (2) physiological complexity and the DNA content of organisms, and (3) evidence that there are many proteins already known whose primary purpose is the regulation of the function of other proteins. It is shown that the best available detection techniques would not uncover these proteins unless a concerted effort were made to do so. An outline of such an effort, based on immune precipitation and two-dimensional gel electrophoresis, is presented and discussed.

Animals

The size distribution of insertions and deletions in human and rodent pseudogenes suggests the logarithmic gap penalty for sequence alignment.

The size distributions of deletions, insertions, and indels (i.e., insertions or deletions) were studied, using 78 human processed pseudogenes and other published data sets. The following results were obtained: (1) Deletions occur more frequently than do insertions in sequence evolution; none of the pseudogenes studied shows significantly more insertions than deletions. (2) Empirically, the size distributions of deletions, insertions, and indels can be described well by a power law, i.e., fk = Ck-b, where fk is the frequency of deletion, insertion, or indel with gap length k, b is the power parameter, and C is the normalization factor. (3) The estimates of b for deletions and insertions from the same data set are approximately equal to each other, indicating that the size distributions for deletions and insertions are approximately identical. (4) The variation in the estimates of b among various data sets is small, indicating that the effect of local structure exists but only plays a secondary role in the size distribution of deletions and insertions. (5) The linear gap penalty, which is most commonly used in sequence alignment, is not supported by our analysis; rather, the power law for the size distribution of indels suggests that an appropriate gap penalty is wk = a + b ln k, where a is the gap creation cost and blnk is the gap extension cost. (6) The higher frequency of deletion over insertion suggests that the gap creation cost of insertion (ai) should be larger than that of deletion (ad); that is, ai - ad = ln R, where R is the frequency ratio of deletions to insertions.

Animals

Simulation studies on the evolution of amino acid sequences.

A model of molecular evolution in which the parameter (intrinsic rate of amino acid substitution) fluctuates from time to time was investigated by simulating the process. It was found that the usual method of estimation such as Poisson fitting underestimates this variation of the parameter when remote comparisons are made. At the same time, four distance measures (minimun base difference, Poisson fitting, random nucleotide substitutions and negative binomial fitting) were tested for their accuracy. When the substitution rate is not uniform among the amino acid sites, the negative bionomial fitting gives most satisfactory results, however, one needs to know the parameter beforehand in order to use this method. It was pointed out that the fluctuation of the evolutionary rate is expected if the nearly neutral but very slightly deleterious mutations play an important role on molecular evolution.

Amino Acid Sequence

Crypthecodinium and Tetrahymena: an exercise in comparative evolution.

Nucleotide sequences have been determined for the highly variable D2 region of the large rRNA molecule for over 60 strains of dinoflagellates. These strains were selected from a worldwide collection that represents all the known sibling species (compatibility groups, Mendelian species) in the sibling swarm referred to as Crypthecodinium cohnii. A phylogenetic tree has been constructed from an analysis of the variations in a length of about 180 bases, using PHYLOGEN string analysis programs. The Crypthecodinium tree is compared with the previously published but here augmented tree constructed upon the same rRNA region for the sibling species of a worldwide collection of ciliated protozoa related to the genus Tetrahymena. The first reported sequence of Lambornella clarki, the parasite of tree-hole mosquitoes, is included. The dinoflagellate species complex is much more homogeneous with respect to ribosomal variation. The mean number of differences among sequences from different Crypthecodinium species is about 7, in comparison with 22 differences among the ciliate species examined. Moreover, all the diversity in the dinoflagellates can be explained by base substitutions, whereas insertions and deletions are common in the ciliates. The dinoflagellates are also much more uniform with respect to nutritional and genetic economies. The two complexes differ also in the relationship between molecular variations and breeding compatibility. All tetrahymenine sibling species thus far examined are monomorphic in the D2 region, but several dinoflagellate species are polymorphic. Several different dinoflagellate species, moreover, have identical D2 regions. This kind of ribosomal identity of incompatible strains is found in these ciliates only in one tight cluster of species--Group C. The tetrahymenine swarm is apparently much older than the Crypthecodinium swarm, and the dinoflagellate species produce incompatible progeny species much more readily than do the ciliates, perhaps by the acquisition of mutations that potentiate incompatibility in sympatric populations.

Animals

Extreme variations in the ratios of non-synonymous to synonymous nucleotide substitution rates in signal peptide evolution.

Nucleotide sequences encoding signal peptides from the precursors of alpha-amylase/trypsin inhibitors from cereals are homologous to those corresponding to the precursors of thaumatin II and of plastocyanins. Non-synonymous (KA) and synonymous (KS) rates of nucleotide substitutions have been calculated for all possible binary combinations. Extreme variation in KA/KS ratios has been observed; from the 0.167 average found within the plastocyanin family to an average of 1.90 calculated for the inhibitors/thaumatin II transition. A similar calculation has been carried out for the signal peptide sequences of thionins, which are unrelated to those of the alpha-amylase/trypsin inhibitor family, and an average KA/KS of 0.12 has been obtained. This variation can be largely explained in terms of an empirical index of stability related to amino acid composition and seems to be independent of functional constraints.

Amino Acid Sequence