PubMed HealthSearch

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Comparative mitogenomics of Ocnus glacialis reveals lineage-specific evolutionary rates and complex gene rearrangements in Dendrochirotida.

The order Dendrochirotida (Class Holothuroidea) is a species-rich echinoderm group, yet its internal evolutionary history remains poorly resolved due to limited mitogenomic resources. In this study, we characterized the first complete mitochondrial genome of Ocnus glacialis and conducted comparative analyses to elucidate its phylogenetic position and molecular evolutionary patterns. The circular mitogenome of O. glacialis is 16,776 bp in length, containing the canonical set of 37 genes. Among the analyzed dendrochirotids, O. glacialis exhibited the highest A + T content (70.88%) and a near-zero AT-skew, a compositional profile often linked to lineage-specific evolution in specialized environments. Selection pressure analyses, including branch-model tests, revealed that these compositional features are associated with relaxed purifying selection and an accelerated rate of sequence evolution. Branch-site analyses further identified specific codon sites in cytb, nad2, nad4l, nad5, and nad6 under positive or relaxed constraints. Structurally, O. glacialis displayed the most complex gene rearrangement pattern among the studied species, characterized by multiple tandem duplication-random loss (TDRL) events and extensive intergenic sequences. Furthermore, divergence time estimation suggests that these structural and compositional shifts occurred in tandem with the lineage's diversification. We propose that these mitogenomic signatures reflect a synergistic outcome of habitat transition toward Arctic cold-water and deep-sea environments, coupled with demographic factors such as reduced effective population sizes inherent to its benthic life history. By resolving taxonomic uncertainties, this study provides a robust temporal and molecular framework for understanding the evolutionary history and ecological diversification of the Ocnus lineage.

Animals

Modeling residue usage in aligned protein sequences via maximum likelihood.

A computational method is presented for characterizing residue usage, i.e., site-specific residue frequencies, in aligned protein sequences. The method obtains frequency estimates that maximize the likelihood of the sequences in a simple model for sequence evolution, given a tree or a set of candidate trees computed by other methods. These maximum-likelihood frequencies constitute a profile of the sequences, and thus the method offers a rigorous alternative to sequence weighting for constructing such a profile. The ability of this method to discard misleading phylogenetic effects allows the biochemical propensities of different positions in a sequence to be more clearly observed and interpreted.

Amino Acids

Length variation and secondary structure of introns in the Mlc1 gene in six species of Drosophila.

A nearly universal feature of intron sequences is that even closely related species exhibit a large number of insertion/deletion differences. The goal of the analysis described here is to test whether the observed pattern of insertion/deletion events in the genealogy of the myosin alkali light chain (Mlc1) gene is consistent with neutrality, and if not, to determine the underlying forces of evolutionary change. Mlc1 pre-mRNA is alternatively spliced, and one constraint is that signals necessary for tissue-specificity of directed splicing must be conserved. If the total length of an intron is functionally constrained, then the distribution of indels on branches of the gene genealogy should reflect a departure from randomness. Here we perform a phylogenetic analysis, inferring ancestral states wherever possible on a phylogeny of 29 alleles of Mlc1 from six species of Drosophila. Observed patterns of indels on the genealogy were compared to those from simulated data, with the result that we cannot reject the null hypothesis of neutrality. A clear departure from a neutral prediction was seen in the excess folding free energy predicted for the introns flanking the alternatively spliced exon. Relative rate tests also suggest a retardation in the rate of Mlc1 sequence evolution in the simulans clade.

Animals

DIPLOMO: the tool for a new type of evolutionary analysis.

A package of computer programs called DIPLOMO (DIstance PLOt MOnitor) has been developed for making pairwise comparisons of different estimates of the distances between a set of taxa by plotting them against each other in a simple scatter plot. Taxa with similar relative distance characteristics are thereby grouped graphically. Groupings of different taxa may be directly identified, and the distance characteristics of chosen groups visualised and compared using devices to give them different colours or symbols. The program is particularly useful for detecting and analysing subtle trends in gene sequence evolution. This is done by comparing different components of change, for example synonymous versus non-synonymous nucleotide changes, transversions versus transitions and changes in different genes of the same set of taxa, etc. The program has a wide range of other uses, for example comparing different methods of sequence analysis, assessing which components of genetic change correlate best with phenotypic change or with geographical separation. This paper describes the DIPLOMO package, and illustrates typical DIPLOMO analyses using lentivirus gene sequence data.

Biological Evolution

Molecular evolution of ruminant lysozymes.

The evolution of a new digestive enzyme, stomach lysozyme, from an antibacterial host defense enzyme provides a link between molecular evolution and organismal evolution. Lysozymes have been recruited at least three times (twice from a conventional lysozyme c and once from a calcium-binding lysozyme c) in vertebrates for functioning in the stomach. The recruitment of lysozyme for its new biological function involved many molecular changes, beyond those required to adapt the protein to function in the stomach. The evolution of the stomach lysozyme gene has been extensively studied in ruminant artiodactyls. In ruminants, the lysozyme c gene has duplicated to yield a family of about ten genes. These duplications allowed: (1) specialization of gene function and (2) increased levels of expression. The ruminant stomach lysozyme genes have evolved in an episodic fashion - there was a period of rapid adaptive sequence evolution, driven by positive selection in the early ruminant, that was followed by an increase in purifying selection upon the well-adapted stomach lysozyme sequence among modern species. Recombination of small portions (exons) of the genes between members of the lysozyme gene family may have aided in adaptive evolution. Evolution to a stomach lysozyme is not irreversible; at least one member of the ruminant stomach lysozyme gene family appears to have reverted to a more ancestral function, yet retains hallmarks of its history as a stomach lysozyme.

Animals

Ross River virus genetic variants in Australia and the Pacific Islands.

HaeIII and TaqI restriction digest profiles of cDNA to infected cell RNA or virion RNA were used as a guide to genetic relationships between fourteen isolates of Ross River virus (RRV) obtained from mosquitoes collected in various localities in eastern Australia where the virus is endemic. RRV isolates from Fiji, American Samoa, the Cook Islands and the Wallis Islands where major outbreaks of epidemic polyarthritis took place in 1979-1980 were also examined. Among these RRV isolates we have identified three genetic types (I-III) on the basis of differences between their restriction digest profiles. We estimate that 1.5-5% nucleotide sequence diversity exists between genetic types. Within each genetic type strain differentiation gave rise to small but significant differences in restriction digest profiles. No clear pattern of geographic distribution of RRV genetic types could be established from the limited number of RRV isolates examined. Genetic types I, II and III, respectively, were isolated from three, three and one different mosquito species, indicating there is no strong association between genetic type and the species of mosquito vector. HaeIII restriction digest analysis did not detect any genetic difference between the four Pacific Island isolates, suggesting that a single RRV variant was involved in the epidemics. Genetically, this variant was closely related to isolates of genetic type II. Virtually identical HaeIII restriction digest profiles were observed for isolates obtained at various stages of the Pacific Island epidemics, suggesting that extensive sequence evolution did not accompany Ross River virus spread.

Alphavirus

Di-, tri-, and tetranucleotide frequencies covary with lifespan and genome size across protostome invertebrates.

Animal lifespans span orders of magnitude, yet how genome sequence covaries with lifespan remains poorly characterized outside vertebrates. Although promoter CpG density has been linked to vertebrate longevity due to its gene-regulatory function through DNA methylation, it is unclear whether such patterns are promoter- and CpG-specific, or if they reflect broader sequence evolution. We curated maximum lifespan estimates for 466 protostome species spanning eight phyla with available genome assemblies and quantified mono-, di-, tri-, and tetranucleotide composition across whole genomes, intergenic regions, and six gene-associated regions (two upstream regions, exons, introns, and two downstream regions) defined using Benchmarking Universal Single-Copy Orthologs. Dinucleotide observed/expected ratios showed significant associations with lifespan and genome size in different ways. Lifespan-associated motifs were most pronounced in gene-associated non-coding regions, especially in introns and downstream regions, whereas genome-size effects were strongest in whole-genome and intergenic sequence. Tri- and tetranucleotide observed/expected ratios broadly recapitulated this regional organization. In contrast, GC content was not associated with lifespan across regions, indicating that the observed signals are not explained by mononucleotide composition but instead by how those nucleotides are arranged into short sequence motifs. These results suggest that lifespan and genome size show distinct but overlapping associations with regional sequence composition across invertebrate species and that lifespan-associated motif evolution extends beyond vertebrate promoter methylation architectures.

CpG density

Restriction fragment polymorphism in the sex-determining region of the Y chromosomal DNA of European wild mice.

Using 32P-labeled probe consisting mainly of (GATA)n we have shown that a male specific Alu1 DNA blot pattern which defines the Y chromosome sex-determining locus in inbred mice is highly polymorphic in wild mice, indicating substantial sequence evolution in this region under field conditions. In all cases examined by in situ hybridization, the region concerned is paracentromeric. In contrast, the blot pattern of another probe (M 34) which detects repeated sequences specific to the mouse Y chromosome but outside the sex-determining locus, remains constant between different isolates.

Animals

Mutational analysis of viroid pathogenicity: tomato apical stunt viroid.

A series of nucleotide substitutions within the pathogenicity domain of tomato apical stunt viroid have been evaluated for their effects upon infectivity and symptom expression. None of the 12 A----G substitutions and one C----U substitution that were examined abolished infectivity in a whole plant bioassay, and the resulting progeny were characterized by nucleotide sequence analysis of cDNAs amplified by the polymerase chain reaction. Four of the 13 substitutions gave rise to altered progeny, but the patterns of sequence changes observed were unexpectedly complex. Mutations that did not rapidly revert to the wild-type sequence are located near the right border of the pathogenicity domain, a region which shows considerable natural sequence variability. None had a detectable effect upon symptom expression. The ability to observe viroid sequence evolution in vivo may provide insight into the molecular interactions responsible for viroid host range and symptom formation.

Base Sequence

The evolution of insertion sequences within enteric bacteria.

To identify mechanisms that influence the evolution of bacterial transposons, DNA sequence variation was evaluated among homologs of insertion sequences IS1, IS3 and IS30 from natural strains of Escherichia coli and related enteric bacteria. The nucleotide sequences within each class of IS were highly conserved among E. coli strains, over 99.7% similar to a consensus sequence. When compared to the range of nucleotide divergence among chromosomal genes, these data indicate high turnover and rapid movement of the transposons among clonal lineages of E. coli. In addition, length polymorphism among IS appears to be far less frequent than in eukaryotic transposons, indicating that nonfunctional elements comprise a smaller fraction of bacterial transposon populations than found in eukaryotes. IS present in other species of enteric bacteria are substantially divergent from E. coli elements, indicating that IS are mobilized among bacterial species at a reduced rate. However, homologs of IS1 and IS3 from diverse species provide evidence that recombination events and horizontal transfer of IS among species have both played major roles in the evolution of these elements. IS3 elements from E. coli and Shigella show multiple, nested, intragenic recombinations with a distantly related transposon, and IS1 homologs from diverse taxa reveal a mosaic structure indicative of multiple recombination and horizontal transfer events.

Base Sequence

Primary structures of dehydrogenases. Evolutionary characteristics related to functional aspects; models for isozyme developments and ancestral connections.

This chapter describes known characteristics of evolutionary changes in individual dehydrogenases, as well as possible relationships among this group of enzymes. Data from primary structures are correlated with those from other observations. Variations in the amino acid sequences demonstrate functional properties, and can be interpreted in relation to conformational aspects, subunit arrangements and enzyme stabilities. Different types of isozyme developments have occurred and show functional fixations at various levels. They define isozyme patterns of general significance in protein evolution. Sequence similarities may be found between different segments. They are analyzed in relation to known conformations, subunit sizes, species divergence and genetic mechanisms. A wide-ranging evolutionary model is discussed relating dehydrogenases and some other oligomeric enzymes to a distant, frequently remodelled ancestral building unit of repetitive occurrence.

Alcohol Oxidoreductases

Comparative Genomics of Sex-Determination-Related Genes Reveals Shared Evolutionary Patterns Between Bivalves and Mammals, but Not Fruit Flies.

The molecular basis of sex determination (SD), while being extensively studied in model organisms, remains poorly understood in many animal groups. Bivalves, a diverse class of molluscs with a variety of reproductive modes, represent an ideal yet challenging clade for investigating SD and the evolution of sexual systems. However, the absence of a comprehensive framework has limited progress in this field, particularly regarding the study of sex-determination-related genes (SRGs). In this study, we performed a genome-wide sequence evolutionary analysis of the Dmrt, Sox and Fox gene families in more than 40 bivalve species. For the first time, we provide an extensive and phylogenetically aware dataset of these SRGs, and we find support for the hypothesis that Dmrt-1L and Sox-H may act as primary sex-determining genes by showing their high levels of sequence diversity within the bivalve genomic context. To validate our findings, we studied the same gene families in two well-characterised systems, mammals and fruit flies (genus Drosophila). In the former, we found that the male sex-determining gene Sry exhibits a pattern of amino acid sequence diversity similar to that of Dmrt-1L and Sox-H in bivalves, consistent with its role as master SD regulator. In contrast, no such pattern was observed among genes of the fruit fly SD cascade, which is controlled by a chromosomic mechanism. Overall, our findings highlight similarities in the sequence evolution of some mammal and bivalve SRGs, possibly driven by a comparable architecture of SD cascades. This work underscores once again the importance of employing a comparative approach when investigating understudied and non-model systems.

Animals

Molecular evolution of the hepatitis B virus genome.

The hepatitis B virus (HBV) has a circular DNA genome of about 3,200 base pairs. Economical use of the genome with overlapping reading frames may have led to severe constraints on nucleotide substitutions along the genome and to highly variable rates of substitution among nucleotide sites. Nucleotide sequences from 13 complete HBV genomes were compared to examine such variability of substitution rates among sites and to examine the phylogenetic relationships among the HBV variants. The maximum likelihood method was employed to fit models of DNA sequence evolution that can account for the complexity of the pattern of nucleotide substitution. Comparison of the models suggests that the rates of substitution are different in different genes and codon positions; for example, the third codon position changes at a rate over ten times higher than the second position. Furthermore, substantial variation of substitution rates was detected even after the effects of genes and codon positions were corrected; that is, rates are different at different sites of the same gene or at the same codon position. Such rates after the correction were also found to be positively correlated at adjacent sites, which indicated the existence of conserved and variable domains in the proteins encoded by the viral genome. A multiparameter model validates the earlier finding that the variation in nucleotide conservation is not random around the HBV genome. The test for the existence of a molecular clock suggests that substitution rates are more or less constant among lineages. The phylogenetic relationships among the viral variants were examined. Although the data do not seem to contain sufficient information to resolve the details of the phylogeny, it appears quite certain that the serotypes of the viral variants do not reflect their genetic relatedness.

Codon

Variance to mean ratio, R(t), for poisson processes on phylogenetic trees.

The ratio of expected variance to mean, R(t), of numbers of DNA base substitutions for contemporary sequences related by a "star" phylogeny is widely seen as a measure of the adherence of the sequences' evolution to a Poisson process with a molecular clock, as predicted by the "neutral theory" of molecular evolution under certain conditions. A number of estimators of R(t) have been proposed, all predicted to have mean 1 and distributions based on the chi 2. Various genes have previously been analyzed and found to have values of R(t) far in excess of 1, calling into question important aspects of the neutral theory. In this paper, I use Monte Carlo simulation to show that the previously suggested means and distributions of estimators of R(t) are highly inaccurate. The analysis is applied to star phylogenies and to general phylogenetic trees, and well-known gene sequences are reanalyzed. For star phylogenies the results show that Kimura's estimators ("The Neutral Theory of Molecular Evolution," Cambridge Univ. Press, Cambridge, 1983) are unsatisfactory for statistical testing of R(t), but confirm the accuracy of Bulmer's correction factor (Genetics 123: 615-619, 1989). For all three nonstar phylogenies studied, attained values of all three estimators of R(t), although larger than 1, are within their true confidence limits under simple Poisson process models. This shows that lineage effects can be responsible for high estimates of R(t), restoring some limited confidence in the molecular clock and showing that the distinction between lineage and molecular clock effects is vital.(ABSTRACT TRUNCATED AT 250 WORDS)

Analysis of Variance

Accelerated evolution and Muller's rachet in endosymbiotic bacteria.

Many bacteria live only within animal cells and infect hosts through cytoplasmic inheritance. These endosymbiotic lineages show distinctive population structure, with small population size and effectively no recombination. As a result, endosymbionts are expected to accumulate mildly deleterious mutations. If these constitute a substantial proportion of new mutations, endosymbionts will show (i) faster sequence evolution and (ii) a possible shift in base composition reflecting mutational bias. Analyses of 16S rDNA of five independently derived endosymbiont clades show, in every case, faster evolution in endosymbionts than in free-living relatives. For aphid endosymbionts (genus Buchnera), coding genes exhibit accelerated evolution and unusually low ratios of synonymous to nonsynonymous substitutions compared to ratios for the same genes for enterics. This concentration of the rate increase in nonsynonymous substitutions is expected under the hypothesis of increased fixation of deleterious mutations. Polypeptides for all Buchnera genes analyzed have accumulated amino acids with codon families rich in A+T, supporting the hypothesis that substitutions are deleterious in terms of polypeptide function. These observations are best explained as the result of Muller's ratchet within small asexual populations, combined with mutational bias. In light of this explanation, two observations reported earlier for Buchnera, the apparent loss of a repair gene and the overproduction of a chaperonin, may reflect compensatory evolution. An alternative hypothesis, involving selection on genomic base composition, is contradicted by the observation that the speedup is concentrated at nonsynonymous sites.

Animals

Compositional heterogeneity and patterns of molecular evolution in the Drosophila genome.

The rates and patterns of molecular evolution in many eukaryotic organisms have been shown to be influenced by the compartmentalization of their genomes into fractions of distinct base composition and mutational properties. We have examined the Drosophila genome to explore relationships between the nucleotide content of large chromosomal segments and the base composition and rate of evolution of genes within those segments. Direct determination of the G + C contents of yeast artificial chromosome clones containing inserts of Drosophila melanogaster DNA ranging from 140-340 kb revealed significant heterogeneity in base composition. The G + C content of the large segments studied ranged from 36.9% G + C for a clone containing the hunchback locus in polytene region 85, to 50.9% G + C for a clone that includes the rosy region in polytene region 87. Unlike other organisms, however, there was no significant correlation between the base composition of large chromosomal regions and the base composition at fourfold degenerate nucleotide sites of genes encompassed within those regions. Despite the situation seen in mammals, there was also no significant association between base composition and rate of nucleotide substitution. These results suggest that nucleotide sequence evolution in Drosophila differs from that of many vertebrates and does not reflect distinct mutational biases, as a function of base composition, in different genomic regions. Significant negative correlations between codon-usage bias and rates of synonymous site divergence, however, provide strong support for an argument that selection among alternative codons may be a major contributor to variability in evolutionary rates within Drosophila genomes.

Animals

Intra-Host Evolution Provides for the Continuous Emergence of SARS-CoV-2 Variants.

Variants of concern (VOC) in SARS-CoV-2 refer to viruses whose viral genomes differ from the ancestor virus by ≥3 single-nucleotide variants (SNVs) and that show the potential for higher transmissibility and/or worse clinical progression. VOC have the potential to disrupt ongoing public health measures and vaccine efforts. Still, too little is known regarding how frequently new viral variants emerge and under what circumstances. We report a study to determine the degree of SARS-CoV-2 sequence evolution in 94 patients and to estimate the frequency at which highly diverse variants emerge. Two cases accumulated ≥9 SNVs over a 2-week period and one case accumulated 23 SNVs over 3 weeks, including three nonsynonymous mutations in the spike protein (D138H, E554D, D614G). The remainder of the infected patients did not show signs of intra-host evolution. We estimate that in as much as 2% of hospitalized COVID-19 cases, variants with multiple mutations in the spike glycoprotein emerge in as little as 1 month of persistent intra-host virus replication. This suggests the continued local emergence of variants with multiple nonsynonymous SNVs, even in patients without overt immune deficiency. Surveillance by sequencing for (i) viremic COVID-19 patients, (ii) patients suspected of reinfection, and (iii) patients with diminished immune function may offer broad public health benefits. IMPORTANCE New SARS-CoV-2 variants can potentially disrupt ongoing public health measures and vaccine efforts. Still, little is known regarding how frequently new viral variants emerge and under what circumstances. Based on this study, we estimate that in hospitalized COVID-19 cases, variants with multiple mutations may emerge locally in as little as 1 month, even in patients without overt immune deficiency. Surveillance by sequencing for continuously shedding patients, patients suspected of reinfection, and patients with diminished immune function may offer broad public health benefits.

Humans

Nucleotide sequence and molecular evolution of mouse retrovirus-like IAP elements.

We determined the nucleotide (nt) sequences of cDNA and genomic clones for murine intracisternal type A particle (IAP) elements, which are retrovirus-like repetitive sequences in rodent genomes. The nucleotide sequence of the cDNA resembled that of retrovirus RNA genomes in its lack of the U5 sequence within the 3' long terminal repeat. By sequence comparison of our clones with reported rodent IAP elements, we located the probable gag, pol and env gene regions. The sequences for the pol, env and the 3' two-thirds of the gag region were conserved among the IAP elements. In the regions, synonymous substitutions occurred more frequently than non-synonymous ones, which suggested that the regions in question were functionally constrained until fairly recently. The rate of nucleotide substitutions in the regions was estimated to be 6-10 X 10(-9) nt per site per year, and significantly higher than that of the cellular genes. These rates may exemplify a characteristic of the nucleotide substitutions for an endogenous retrovirus. The sequence homology between the IAP element and IgE-binding factor gene is discussed.

Animals