PubMed HealthSearch

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

MCALIGN: stochastic alignment of noncoding DNA sequences based on an evolutionary model of sequence evolution.

A method is described for performing global alignment of noncoding DNA sequences based on an evolutionary model parameterized by the frequency distribution of lengths of insertion/deletion events (indels) and their rate relative to nucleotide substitutions. A stochastic hill-climbing algorithm is used to search for the most probable alignment between a pair of sequences or three sequences of known phylogenetic relationship. The performance of the procedure, parameterized according to the empirical distribution of indel lengths in noncoding DNA of Drosophila species, is investigated by simulation. We show that there is excellent agreement between true and estimated alignments over a wide range of sequence divergences, and that the method outperforms other available alignment methods.

Algorithms

Direct link between convergent evolution at sequence level and phenotypic level of septal pore cap in Agaricomycotina.

Several homologous morphological characters, despite sharing apparently similar features, are known to have independently evolved in different lineages multiple times. However, the genetic backgrounds of such morphological convergences remain poorly understood. To detect any correlated amino acid substitutions potentially responsible for morphological convergence at the phenotypic level, we focused on the morphology of the septal pore cap (SPC), a structure involved in mycelia's complex multicellularity in fungi. SPCs are classified into 3 morphological types: perforate, imperforate, and vesiculate. To understand the evolutionary events that occurred at the sequence level during the morphological convergence of perforate SPCs in Agaricomycotina, we examined sequence differences among species with different SPC types by comparative genomic analysis using a single-copy gene dataset from 12 Agaricomycotina genomes with morphological literature of SPC. Our analysis revealed that sequences of 8 genes, including an SPC-related gene spc33, were clustered based on SPC morphology rather than species relationship. Additionally, same amino acid substitutions independently occurred in both lineages in which species with perforate SPCs emerged. These findings suggest that specific amino acid substitutions in spc33 were critical for the emergence of perforate SPCs in multiple lineages. Further, our gene search for spc33 across organisms suggests that spc33 evolved shortly before the emergence of imperforate SPC. This study represents the first step toward elucidating the genetic basis of the morphological evolution of SPC. It contributes to both clarifying the genetic basis underlying morphological convergence and advances the study of fungal evolutionary morphology.

Evolution, Molecular

Long-range mRNA folding shapes expression and sequence of bacterial genes.

Bacterial gene expression is strongly influenced by local mRNA secondary structure, yet the impact of long-range folding remains poorly understood. Here, we show that sequences hundreds of nucleotides from the mRNA 5' end can act as potent repressors of gene expression through long-range base pairing to the ribosome binding site (RBS), subjecting anti-RBS sequences to negative selection. Using massively parallel reporter assays in Bacillus subtilis, we identify anti-RBS sequences as among the strongest determinants of reduced mRNA abundance across the transcript body. We demonstrate that distal anti-RBS elements engage in long-range folding with the Shine-Dalgarno sequence, blocking ribosome entry and promoting mRNA decay. Consistent with these repressive effects, anti-RBS-like sequences are depleted throughout diverse bacterial coding sequences but not from leaderless transcripts, and introducing distal anti-RBS to native genes reduces expression. Our findings establish that long-range mRNA folding is a conserved force shaping gene expression and constrains coding sequence evolution.

Bacillus subtilis

Deep Sequencing Reveals Dual Evolution of SARS-CoV-2: Insights Into Defective Genomes From Wuhan-Hu-1 Variants to Omicron Subvariants.

SARS-CoV-2 has evolved from early variants dominating the first (B.1.5, B.1.1) and second (B.1.177) pandemic waves, which exhibited a higher frequency of minority mutants with deletions leading to Defective Viral Genomes (DVGs) in the spike region near the S1/S2 cleavage site than the Alpha, Beta, and Delta variants. The emergence of Omicron has significantly altered the dominant variant profile, with Omicron subvariants now representing 100% of circulating viruses. To monitor the evolution and adaptation of Omicron in the human population, a deep-sequencing study was performed in RNA samples of BA.1, BA.1.1, BA.2, BA.5, BQ.1.1, XBB.1.5 and BA.2.86 Omicron subvariants. The findings reveal two occurrences of similar evolutionary patterns within SARS-CoV-2 characterized by a shift from a significant to a very low production of DVGs. This event suggests that DVGs might play a role in the virus's spread and adaptation for persistence in infected humans.

SARS-CoV-2

Exploring the Mitochondrial Genomes of Phoebe Species (Lauraceae): Structural Dynamics and Functional Conservation.

Plant mitochondrial genomes (mitogenomes) vary markedly in size and architecture despite generally slow rates of sequence evolution. Phoebe is an ecologically and economically valuable genus of Lauraceae, yet its mitogenome diversity remains poorly characterized. In this study, we newly sequenced, assembled, and annotated the mitogenomes of three nationally protected Class II wild plants (P. bournei, P. chekiangensis, P. zhennan) from China and compared their mitogenomic characteristics. The three assemblies were resolved into representative circular configurations ranging from 808 to 864 kb, with similar GC contents and conserved protein-coding capacity. Each mitogenome contained distinct 41 protein-coding genes, 27-28 transfer RNAs, and three ribosomal RNAs. Synteny analysis revealed extensive changes in homologous-block order and orientation despite substantial sequence homology among the three species. Abundant repeats occurred predominantly in noncoding regions, while plastid-derived fragments documented historical intracellular DNA transfer. The three species exhibited similar codon usage and predicted RNA-editing patterns, whereas low synonymous divergence limited inference from pairwise ratios. Phylogenetic analysis based on mitochondrial protein-coding genes recovered Phoebe as a well-supported monophyletic lineage. These results reveal substantial structural divergence accompanied by conserved nucleotide composition and coding capacity, providing valuable data for further understanding the evolutionary variation of plant mitogenomes of Phoebe and the Lauraceae.

Phoebe

Comparative mitogenomics of Ocnus glacialis reveals lineage-specific evolutionary rates and complex gene rearrangements in Dendrochirotida.

The order Dendrochirotida (Class Holothuroidea) is a species-rich echinoderm group, yet its internal evolutionary history remains poorly resolved due to limited mitogenomic resources. In this study, we characterized the first complete mitochondrial genome of Ocnus glacialis and conducted comparative analyses to elucidate its phylogenetic position and molecular evolutionary patterns. The circular mitogenome of O. glacialis is 16,776 bp in length, containing the canonical set of 37 genes. Among the analyzed dendrochirotids, O. glacialis exhibited the highest A + T content (70.88%) and a near-zero AT-skew, a compositional profile often linked to lineage-specific evolution in specialized environments. Selection pressure analyses, including branch-model tests, revealed that these compositional features are associated with relaxed purifying selection and an accelerated rate of sequence evolution. Branch-site analyses further identified specific codon sites in cytb, nad2, nad4l, nad5, and nad6 under positive or relaxed constraints. Structurally, O. glacialis displayed the most complex gene rearrangement pattern among the studied species, characterized by multiple tandem duplication-random loss (TDRL) events and extensive intergenic sequences. Furthermore, divergence time estimation suggests that these structural and compositional shifts occurred in tandem with the lineage's diversification. We propose that these mitogenomic signatures reflect a synergistic outcome of habitat transition toward Arctic cold-water and deep-sea environments, coupled with demographic factors such as reduced effective population sizes inherent to its benthic life history. By resolving taxonomic uncertainties, this study provides a robust temporal and molecular framework for understanding the evolutionary history and ecological diversification of the Ocnus lineage.

Animals

Di-, tri-, and tetranucleotide frequencies covary with lifespan and genome size across protostome invertebrates.

Animal lifespans span orders of magnitude, yet how genome sequence covaries with lifespan remains poorly characterized outside vertebrates. Although promoter CpG density has been linked to vertebrate longevity due to its gene-regulatory function through DNA methylation, it is unclear whether such patterns are promoter- and CpG-specific, or if they reflect broader sequence evolution. We curated maximum lifespan estimates for 466 protostome species spanning eight phyla with available genome assemblies and quantified mono-, di-, tri-, and tetranucleotide composition across whole genomes, intergenic regions, and six gene-associated regions (two upstream regions, exons, introns, and two downstream regions) defined using Benchmarking Universal Single-Copy Orthologs. Dinucleotide observed/expected ratios showed significant associations with lifespan and genome size in different ways. Lifespan-associated motifs were most pronounced in gene-associated non-coding regions, especially in introns and downstream regions, whereas genome-size effects were strongest in whole-genome and intergenic sequence. Tri- and tetranucleotide observed/expected ratios broadly recapitulated this regional organization. In contrast, GC content was not associated with lifespan across regions, indicating that the observed signals are not explained by mononucleotide composition but instead by how those nucleotides are arranged into short sequence motifs. These results suggest that lifespan and genome size show distinct but overlapping associations with regional sequence composition across invertebrate species and that lifespan-associated motif evolution extends beyond vertebrate promoter methylation architectures.

CpG density

Comparative Genomics of Sex-Determination-Related Genes Reveals Shared Evolutionary Patterns Between Bivalves and Mammals, but Not Fruit Flies.

The molecular basis of sex determination (SD), while being extensively studied in model organisms, remains poorly understood in many animal groups. Bivalves, a diverse class of molluscs with a variety of reproductive modes, represent an ideal yet challenging clade for investigating SD and the evolution of sexual systems. However, the absence of a comprehensive framework has limited progress in this field, particularly regarding the study of sex-determination-related genes (SRGs). In this study, we performed a genome-wide sequence evolutionary analysis of the Dmrt, Sox and Fox gene families in more than 40 bivalve species. For the first time, we provide an extensive and phylogenetically aware dataset of these SRGs, and we find support for the hypothesis that Dmrt-1L and Sox-H may act as primary sex-determining genes by showing their high levels of sequence diversity within the bivalve genomic context. To validate our findings, we studied the same gene families in two well-characterised systems, mammals and fruit flies (genus Drosophila). In the former, we found that the male sex-determining gene Sry exhibits a pattern of amino acid sequence diversity similar to that of Dmrt-1L and Sox-H in bivalves, consistent with its role as master SD regulator. In contrast, no such pattern was observed among genes of the fruit fly SD cascade, which is controlled by a chromosomic mechanism. Overall, our findings highlight similarities in the sequence evolution of some mammal and bivalve SRGs, possibly driven by a comparable architecture of SD cascades. This work underscores once again the importance of employing a comparative approach when investigating understudied and non-model systems.

Animals

Intra-Host Evolution Provides for the Continuous Emergence of SARS-CoV-2 Variants.

Variants of concern (VOC) in SARS-CoV-2 refer to viruses whose viral genomes differ from the ancestor virus by ≥3 single-nucleotide variants (SNVs) and that show the potential for higher transmissibility and/or worse clinical progression. VOC have the potential to disrupt ongoing public health measures and vaccine efforts. Still, too little is known regarding how frequently new viral variants emerge and under what circumstances. We report a study to determine the degree of SARS-CoV-2 sequence evolution in 94 patients and to estimate the frequency at which highly diverse variants emerge. Two cases accumulated ≥9 SNVs over a 2-week period and one case accumulated 23 SNVs over 3 weeks, including three nonsynonymous mutations in the spike protein (D138H, E554D, D614G). The remainder of the infected patients did not show signs of intra-host evolution. We estimate that in as much as 2% of hospitalized COVID-19 cases, variants with multiple mutations in the spike glycoprotein emerge in as little as 1 month of persistent intra-host virus replication. This suggests the continued local emergence of variants with multiple nonsynonymous SNVs, even in patients without overt immune deficiency. Surveillance by sequencing for (i) viremic COVID-19 patients, (ii) patients suspected of reinfection, and (iii) patients with diminished immune function may offer broad public health benefits. IMPORTANCE New SARS-CoV-2 variants can potentially disrupt ongoing public health measures and vaccine efforts. Still, little is known regarding how frequently new viral variants emerge and under what circumstances. Based on this study, we estimate that in hospitalized COVID-19 cases, variants with multiple mutations may emerge locally in as little as 1 month, even in patients without overt immune deficiency. Surveillance by sequencing for continuously shedding patients, patients suspected of reinfection, and patients with diminished immune function may offer broad public health benefits.

Humans

De novo Genes in Plants: Origins, Mechanisms, and Functional Implications.

De novo genes originate from previously non-coding genomic regions. They provide an important source of lineage-specific innovation. In plants, these genes may contribute to adaptation, trait diversity and crop evolution. This review summarizes recent progress in plant de novo gene research. It first discusses major routes of gene birth, including transcription-first, open reading frame (ORF)-first and concurrent models. It also examines how nascent loci acquire regulatory control and enter existing biological networks. The review then summarizes their evolutionary features, including weak early constraint, rapid molecular change, restricted expression and structural refinement. It further discusses plant de novo genes involved in stress responses, seed germination, kernel dehydration, subspecies divergence, reproductive isolation and floral scent diversification. Current methods for identifying de novo genes remain limited by rapid sequence evolution, genome annotation quality, polyploidy and transposable elements. Whole-genome synteny alignment, multi-omics evidence and machine-learning approaches can improve candidate discovery. However, each method has important limitations. Finally, this review highlights key future questions in functional validation, latent coding potential in long non-coding RNAs, epigenetic activation, regulatory-network integration and crop improvement. These perspectives clarify how de novo genes shape plant adaptation and how they may be used in precision breeding and synthetic biology.

adaptive evolution

Comprehensive analysis of synonymous codon usage bias and evolutionary dynamics in the chloroplast genomes of eight Coptis species.

Coptis is a medically important genus renowned for producing valuable isoquinoline alkaloids. Although its chloroplast genomes encode key components for photosynthesis and plastid gene expression, the evolutionary constraints acting on their coding sequences and synonymous codon usage remain poorly resolved. Here, we combined a transparent taxon-level sampling strategy with comparative analyses of chloroplast CDSs from eight Coptis taxa. We quantified nucleotide composition, relative synonymous codon usage, effective number of codons, neutrality and PR2 patterns, and correspondence analysis, and then integrated these results with a core-CDS distance analysis and gene-wise pairwise dN/dS estimates. The chloroplast genomes showed a conserved AT-rich composition, especially at the third codon position (GC3 approximately 30.3-30.8%), with a consistent GC1 > GC2 > GC3 trend. Thirty preferred codons were detected, 28 ending in A/T, and eleven optimal codons were shared across the genus. The core-CDS distance analysis recovered a close relationship between C. chinensis and C. chinensis var. brevisepala, whereas most coding genes showed dN/dS values below one, consistent with pervasive purifying constraint. Across 48 consistently filtered CDSs, GC3s was negatively associated with mean dN (Spearman rho = -0.404, P = 0.00439) and CAI was positively associated with mean dN (rho = 0.303, P = 0.0361), whereas the remaining associations were not significant (all P > = 0.0972). These results extend codon-usage analysis by linking synonymous-site composition to coding-sequence evolution within Coptis, while providing a hypothesis-generating resource for future plastid engineering studies.

Genome, Chloroplast

Unequally Abundant Chromosomes and Unusual Collections of Transferred Sequences Characterize Mitochondrial Genomes of Gastrodia (Orchidaceae), One of the Largest Mycoheterotrophic Plant Genera.

The mystery of genomic alternations in heterotrophic plants is among the most intriguing in evolutionary biology. Compared to plastid genomes (plastomes) with parallel size reduction and gene loss, mitochondrial genome (mitogenome) variation in heterotrophic plants remains underexplored in many aspects. To further unravel the evolutionary outcomes of heterotrophy, we present a comparative mitogenomic study with 13 de novo assemblies of Gastrodia (Orchidaceae), one of the largest fully mycoheterotrophic plant genera, and its relatives. Analyzed Gastrodia mitogenomes range from 0.56 to 2.1 Mb, each consisting of numerous, unequally abundant chromosomes or contigs. Size variation might have evolved through chromosome rearrangements followed by stochastic loss of "dispensable" chromosomes, with deletion-biased mutations. The discovery of a hyper-abundant (∼15 times intragenomic average) chromosome in two assemblies represents the hitherto most extreme copy number variation in any mitogenomes, with similar architectures discovered in two metazoan lineages. Transferred sequence contents highlight asymmetric evolutionary consequences of heterotrophy: despite drastically reduced intracellular plastome transfers convergent across heterotrophic plants, their rarity of horizontally acquired sequences sharply contrasts parasitic plants, where massive transfers from their hosts prevail. Rates of sequence evolution are markedly elevated but not explained by copy number variation, extending prior findings of accelerated molecular evolution from parasitic to heterotrophic plants. Putative evolutionary scenarios for these mitogenomic convergence and divergence fit well with the common (e.g. plastome contraction) and specific (e.g. host identity) aspects of the two heterotrophic types. These idiosyncratic mycoheterotrophs expand known architectural variability of plant mitogenomes and provide mechanistic insights into their content and size variation.

Genome, Mitochondrial

Algorithms to reconstruct past indels: The deletion-only parsimony problem.

Ancestral sequence reconstruction is an important task in bioinformatics, with applications ranging from protein engineering to the study of genome evolution. When sequences can only undergo substitutions, optimal reconstructions can be efficiently computed using well-known algorithms. However, accounting for indels in ancestral reconstructions is much harder. First, for biologically-relevant problem formulations, no polynomial-time exact algorithms are available. Second, multiple reconstructions are often equally parsimonious or likely, making it crucial to correctly display uncertainty in the results. Here, we consider a parsimony approach where only deletions are allowed, while addressing the aforementioned limitations. First, we describe an exact algorithm to obtain all the optimal solutions. The algorithm runs in polynomial time if only one solution is sought. Second, we show that all possible optimal reconstructions for a fixed node can be represented using a graph computable in polynomial time. While previous studies have proposed graph-based representations of ancestral reconstructions, this result is the first to offer a solid mathematical justification for this approach. Finally we provide arguments for the relevance of the deletion-only case for the general case.

Algorithms

Uce-based phylogeny and classification of Megachilini.

The generic-level classification of the bee tribe Megachilini (Megachilidae) has remained controversial due to poor phylogenetic resolution at the base of the group, particularly among the brood parasitic genera and the numerous dauber ("Chalicodoma s. l.") lineages. We present a phylogenomic analysis of Megachilini based on ultraconserved elements (UCEs), sampling 52 ingroup taxa with emphasis on the dauber lineages. We also present a combined UCE + six-gene analysis to improve taxon coverage, resulting in a dataset with 127 ingroup taxa. Maximum likelihood, coalescent, and Bayesian analyses of multiple UCE matrices recover largely congruent topologies with substantially improved support relative to previous studies. Our results strongly support the monophyly of Megachilini, the early divergence of Noteriades and Gronoceras, and a single origin of brood parasitism. All remaining non-parasitic Megachilini form a moderately supported clade sister to the brood parasitic lineage. The leafcutter bees are monophyletic and nested within dauber lineages. Several major dauber clades are consistently recovered, including an exclusively Australian clade corresponding to the Hackeriapis group of subgenera, while several recognized subgenera are paraphyletic. The lineage known as Morphella, previously placed in synonymy with the subgenus Callomegachile, was not closely related to that subgenus and is here treated as a valid subgenus. Divergence-time analyses place the crown age of Megachilini in the late Eocene to early Oligocene, with major extant lineages diversifying during the Miocene. Limited morphological diagnosability of several clades indicates that splitting non-parasitic lineages into numerous genera would result in an impractical classification that would widen the gap between taxonomists and non-specialists and exacerbate the taxonomic impediment in bees. We therefore advocate retaining a single genus Megachile for non-parasitic Megachilini (excluding Noteriades and Gronoceras), as the classification best supported by phylogenomic evidence and most robust to future taxon sampling.

Animals

Evolutionary conservation and adaptability of cholecystokinin neuropeptide signaling in the sea cucumber Apostichopus japonicus.

BACKGROUND: Food ingestion is fundamental for animal survival and growth, with the cessation of feeding upon nutrient fulfillment being tightly regulated by a variety of satiety factors. Notably, sulfakinin/cholecystokinin (SK/CCK)-type neuropeptide signaling has been identified as an inhibitory regulator of food intake across the animal kingdom. However, its regulatory mechanism in feeding in deuterostome invertebrates remains unclear. Here, we characterized SK/CCK-type signaling in a deuterostome invertebrate, the sea cucumber Apostichopus japonicus (phylum Echinodermata). RESULTS: A single SK/CCK-type precursor in A. japonicus generates two mature peptides (AjSK/CCK1, AjSK/CCK2) that activate a shared receptor (AjSK/CCKR), triggering Ca2+ mobilization via the Gαq-dependent pathway and extracellular signal regulated kinase 1/2 (ERK1/2) phosphorylation. Both peptides induce dose-dependent contraction of longitudinal muscles, while AjSK/CCK2 additionally elicits sustained contraction of the posterior intestine, an effect absent in other gut regions. Long-term injection of both peptides reduces food intake and significantly downregulates orexin-type neuropeptide genes (AjOrexin1P, AjOrexin2P) in the circumoral nerve ring (CNR) and intestine. CONCLUSIONS: Unlike mammals, where CCK inhibits feeding by contracting the pyloric sphincter to delay gastric emptying, SK/CCK-type peptides in sea cucumbers exert their anorexic effect in part by selectively contracting the posterior intestine, thereby inhibiting intestinal emptying. This divergence in action sites highlights the evolutionary adaptability of SK/CCK-type signaling as a conserved inhibitory regulator of feeding across bilaterian animals. Elucidating these mechanisms in the economically important A. japonicus may inform development of appetite-promoting agents for sustainable aquaculture.

Animals

Phylogenetic and Genetic Evolution Analysis of Complete SFTSV Genome Sequences in Shandong Province, China.

Severe fever with thrombocytopenia syndrome (SFTS) is an emerging infectious disease caused by SFTS virus (SFTSV). Shandong province is one of the epidemic regions with high incidence rate of SFTS. To investigate phylogenetical and genetic evolution characteristics of SFTSV in Shandong province, we isolated SFTSV from suspected patients between April 2023 and October 2024, and then whole SFTSV genomes were amplified and sequenced in this study. A total of 25 new strains were analyzed together 56 strains submitted in Genbank from Shandong province. Phylogenetical and genetic analyses of the data set revealed that four genotypes were co-circulating in Shandong province. C3 genotype was the most common genotype in each year with lower genetic divergence. 298 amino acid substitutions were detected in the four proteins of SFTSV, but only two substitutions (Arg624Lys and Arg962Ser) had been proven to have potential impacts on biological functions. In addition, one reassortment strain (C3/C4/C4 for L, M and S segments) and three recombinant strains were identified. Analysis of selection pressure at the level of amino acid substitutions indicated genes within the four ORFs of SFTSV were all subjected to negative selection. In conclusion, the genetic characteristics and evolutionary mechanism of SFTSV was complex in Shandong province. It is necessary to conduct continuous surveillance to grasp the genetic evolution patterns, and to discover novel prevalent variants in a timely manner.

China

The complete sequence of the silkworm W chromosome uncovers its rapid evolution by large-scale duplications/deletions and translocation of W-linked genes.

The complete sequence of the W chromosome, which carries feminization activity in the silkworm, is crucial for understanding the sex-determination system in Lepidoptera. However, extensive accumulation of transposons due to lack of recombination, the very rare protein-coding genes and almost no information about molecular markers has hindered full W sequencing. We report the first complete silkworm W sequence (T2T_W, 11683305 bp) obtained by combining sequencing-assembly technologies and newly developed error detection methods, evaluated with genetically mapped W-RAPD markers, W-mutants, and W-derived BAC clones. The T2T_W sequence showed that the W is composed of a massive 92% accumulation of transposons and repeat sequences, among which the main constituents are intact LTR/LINE retrotransposons indicating recent expansions. In addition to Fem clusters producing Fem piRNA (Feminizer-derived PIWI-interacting RNA), we found 26 protein-coding genes in the W sequence. These include four gene pairs encoding zinc-finger motifs designated z1:z20 and a gene encoding serine/arginine repetitive matrix protein 1-like (SRRM1-like). To identify candidate genes for female sex-determination and differentiation we also sequenced the shortest W (3.8 Mb) from a translocation mutant with feminizing activity, which harbored four conventional genes: a Fem cluster, a pair of z1:z20 isoforms, z20-S, and a SRRM1-like gene. Phylogenetic analysis revealed that z1:z20 originated from a copy of an autosomal zinc-finger gene pair, z2:z21, translocated onto the W around 2.43 Mya and subsequently amplified to yield 4 W-linked zinc-finger gene pairs. The complete W sequence revealed that large-scale deletions and amplifications played a significant role in W chromosome evolution.

Animals