PubMed HealthSearch

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Algorithms to reconstruct past indels: The deletion-only parsimony problem.

Ancestral sequence reconstruction is an important task in bioinformatics, with applications ranging from protein engineering to the study of genome evolution. When sequences can only undergo substitutions, optimal reconstructions can be efficiently computed using well-known algorithms. However, accounting for indels in ancestral reconstructions is much harder. First, for biologically-relevant problem formulations, no polynomial-time exact algorithms are available. Second, multiple reconstructions are often equally parsimonious or likely, making it crucial to correctly display uncertainty in the results. Here, we consider a parsimony approach where only deletions are allowed, while addressing the aforementioned limitations. First, we describe an exact algorithm to obtain all the optimal solutions. The algorithm runs in polynomial time if only one solution is sought. Second, we show that all possible optimal reconstructions for a fixed node can be represented using a graph computable in polynomial time. While previous studies have proposed graph-based representations of ancestral reconstructions, this result is the first to offer a solid mathematical justification for this approach. Finally we provide arguments for the relevance of the deletion-only case for the general case.

Algorithms

Uce-based phylogeny and classification of Megachilini.

The generic-level classification of the bee tribe Megachilini (Megachilidae) has remained controversial due to poor phylogenetic resolution at the base of the group, particularly among the brood parasitic genera and the numerous dauber ("Chalicodoma s. l.") lineages. We present a phylogenomic analysis of Megachilini based on ultraconserved elements (UCEs), sampling 52 ingroup taxa with emphasis on the dauber lineages. We also present a combined UCE + six-gene analysis to improve taxon coverage, resulting in a dataset with 127 ingroup taxa. Maximum likelihood, coalescent, and Bayesian analyses of multiple UCE matrices recover largely congruent topologies with substantially improved support relative to previous studies. Our results strongly support the monophyly of Megachilini, the early divergence of Noteriades and Gronoceras, and a single origin of brood parasitism. All remaining non-parasitic Megachilini form a moderately supported clade sister to the brood parasitic lineage. The leafcutter bees are monophyletic and nested within dauber lineages. Several major dauber clades are consistently recovered, including an exclusively Australian clade corresponding to the Hackeriapis group of subgenera, while several recognized subgenera are paraphyletic. The lineage known as Morphella, previously placed in synonymy with the subgenus Callomegachile, was not closely related to that subgenus and is here treated as a valid subgenus. Divergence-time analyses place the crown age of Megachilini in the late Eocene to early Oligocene, with major extant lineages diversifying during the Miocene. Limited morphological diagnosability of several clades indicates that splitting non-parasitic lineages into numerous genera would result in an impractical classification that would widen the gap between taxonomists and non-specialists and exacerbate the taxonomic impediment in bees. We therefore advocate retaining a single genus Megachile for non-parasitic Megachilini (excluding Noteriades and Gronoceras), as the classification best supported by phylogenomic evidence and most robust to future taxon sampling.

Animals

Evolutionary conservation and adaptability of cholecystokinin neuropeptide signaling in the sea cucumber Apostichopus japonicus.

BACKGROUND: Food ingestion is fundamental for animal survival and growth, with the cessation of feeding upon nutrient fulfillment being tightly regulated by a variety of satiety factors. Notably, sulfakinin/cholecystokinin (SK/CCK)-type neuropeptide signaling has been identified as an inhibitory regulator of food intake across the animal kingdom. However, its regulatory mechanism in feeding in deuterostome invertebrates remains unclear. Here, we characterized SK/CCK-type signaling in a deuterostome invertebrate, the sea cucumber Apostichopus japonicus (phylum Echinodermata). RESULTS: A single SK/CCK-type precursor in A. japonicus generates two mature peptides (AjSK/CCK1, AjSK/CCK2) that activate a shared receptor (AjSK/CCKR), triggering Ca2+ mobilization via the Gαq-dependent pathway and extracellular signal regulated kinase 1/2 (ERK1/2) phosphorylation. Both peptides induce dose-dependent contraction of longitudinal muscles, while AjSK/CCK2 additionally elicits sustained contraction of the posterior intestine, an effect absent in other gut regions. Long-term injection of both peptides reduces food intake and significantly downregulates orexin-type neuropeptide genes (AjOrexin1P, AjOrexin2P) in the circumoral nerve ring (CNR) and intestine. CONCLUSIONS: Unlike mammals, where CCK inhibits feeding by contracting the pyloric sphincter to delay gastric emptying, SK/CCK-type peptides in sea cucumbers exert their anorexic effect in part by selectively contracting the posterior intestine, thereby inhibiting intestinal emptying. This divergence in action sites highlights the evolutionary adaptability of SK/CCK-type signaling as a conserved inhibitory regulator of feeding across bilaterian animals. Elucidating these mechanisms in the economically important A. japonicus may inform development of appetite-promoting agents for sustainable aquaculture.

Animals

Phylogenetic and Genetic Evolution Analysis of Complete SFTSV Genome Sequences in Shandong Province, China.

Severe fever with thrombocytopenia syndrome (SFTS) is an emerging infectious disease caused by SFTS virus (SFTSV). Shandong province is one of the epidemic regions with high incidence rate of SFTS. To investigate phylogenetical and genetic evolution characteristics of SFTSV in Shandong province, we isolated SFTSV from suspected patients between April 2023 and October 2024, and then whole SFTSV genomes were amplified and sequenced in this study. A total of 25 new strains were analyzed together 56 strains submitted in Genbank from Shandong province. Phylogenetical and genetic analyses of the data set revealed that four genotypes were co-circulating in Shandong province. C3 genotype was the most common genotype in each year with lower genetic divergence. 298 amino acid substitutions were detected in the four proteins of SFTSV, but only two substitutions (Arg624Lys and Arg962Ser) had been proven to have potential impacts on biological functions. In addition, one reassortment strain (C3/C4/C4 for L, M and S segments) and three recombinant strains were identified. Analysis of selection pressure at the level of amino acid substitutions indicated genes within the four ORFs of SFTSV were all subjected to negative selection. In conclusion, the genetic characteristics and evolutionary mechanism of SFTSV was complex in Shandong province. It is necessary to conduct continuous surveillance to grasp the genetic evolution patterns, and to discover novel prevalent variants in a timely manner.

China

Evolution of lipoproteins deduced from protein sequence data.

1. Human serum apolipoprotein A-I contains a prominent 11-residue sequence periodicity. 2. Similar 11-residue segments occur in the other sequenced human apolipoproteins, C-I, C-III, and A-II. 3. Computer analyses of the sequences support the hypothesis that they evolved from a common ancestor. 4. An evolutionary history of these proteins is proposed. 5. The estimated rate of change of these proteins indicates that all four types will be found throughout the vertebrates and that related proteins will also be found in invertebrates.

Amino Acid Sequence

Evolution of the transfer RNA molecule.

Base sequences of many transfer RNA (tRNA) species obtained from different sources contain homologous regions. These homologies, which are 6 to 20 nucleotides long, occur both within the same tRNA molecule and between many different tRNA molecules repeatedly. Since it is very unlikely an 80 or so nucleotide long tRNA molecule could have been formed at once, under primordial conditions, we propose that the homologous oligonucleotides found within the tRNA molecules to-day represent the earliest adapter from which tRNA molecules have evolved.

Base Sequence

The complete sequence of the silkworm W chromosome uncovers its rapid evolution by large-scale duplications/deletions and translocation of W-linked genes.

The complete sequence of the W chromosome, which carries feminization activity in the silkworm, is crucial for understanding the sex-determination system in Lepidoptera. However, extensive accumulation of transposons due to lack of recombination, the very rare protein-coding genes and almost no information about molecular markers has hindered full W sequencing. We report the first complete silkworm W sequence (T2T_W, 11683305 bp) obtained by combining sequencing-assembly technologies and newly developed error detection methods, evaluated with genetically mapped W-RAPD markers, W-mutants, and W-derived BAC clones. The T2T_W sequence showed that the W is composed of a massive 92% accumulation of transposons and repeat sequences, among which the main constituents are intact LTR/LINE retrotransposons indicating recent expansions. In addition to Fem clusters producing Fem piRNA (Feminizer-derived PIWI-interacting RNA), we found 26 protein-coding genes in the W sequence. These include four gene pairs encoding zinc-finger motifs designated z1:z20 and a gene encoding serine/arginine repetitive matrix protein 1-like (SRRM1-like). To identify candidate genes for female sex-determination and differentiation we also sequenced the shortest W (3.8 Mb) from a translocation mutant with feminizing activity, which harbored four conventional genes: a Fem cluster, a pair of z1:z20 isoforms, z20-S, and a SRRM1-like gene. Phylogenetic analysis revealed that z1:z20 originated from a copy of an autosomal zinc-finger gene pair, z2:z21, translocated onto the W around 2.43 Mya and subsequently amplified to yield 4 W-linked zinc-finger gene pairs. The complete W sequence revealed that large-scale deletions and amplifications played a significant role in W chromosome evolution.

Animals

Satellite DNA evolution in Tytonidae (Aves: Strigiformes): dynamic repeat landscapes despite conserved karyotypes.

The elevated chromosome numbers observed in Tytonidae relative to the putative ancestral avian karyotype suggest that lineage-specific chromosomal fissions may have played an important role in the evolutionary history of this family. Here, we provide the first cytogenetic characterization of the American barn owl (Tyto furcata) and performs a comparative repeatome analysis across members of the Tytonidae, including other two species, the Western barn owl (Tyto alba), and the Oriental bay owl (Phodilus badius). The karyotype of T. furcata showed a 2n = 92, closely resembling that previously described for T. alba, indicating a high degree of chromosomal conservation within Tytonidae. Although T. furcata and T. alba exhibit similar karyotypic organization, comparative repeatomic analyses revealed differences in their composition, including variation in satellite DNA (satDNA) repertoires and abundance. Eight satDNA families were identified in T. furcata, nine in T. alba, and 28 in P. badius, highlighting the dynamic evolution of repetitive sequences. Several satDNA families were shared between T. furcata and T. alba, whereas some appeared species-specific, supporting the library hypothesis of satDNA evolution. In P. badius, multiple satDNAs exhibited similarity to transposable elements, suggesting that mobile elements contributed to their diversification. Cytogenetic analyses demonstrated centromeric heterochromatin distribution in T. furcata, as well as a large heterochromatic W chromosome enriched in DNA repeats. The localization of satDNAs in centromeric regions and the apparent accumulation of repeats on the W chromosome reinforce the role of repetitive sequences in chromosome organization and sex chromosome differentiation. Together, these findings reveal repeatome diversification despite conserved macrochromosomal structure and provide new insights into genome evolution and chromosomal dynamics in birds.

Animals

Conserved protein folds underpin the diversification of secreted proteins in a fungal pathogen.

BACKGROUND: During host colonization, fungal plant pathogens secrete effector-like proteins that alter host cell physiology and target plant-associated microbes. However, rapid evolution and low sequence conservation hinder the study and characterization of these proteins. The fungus Zymoseptoria passerinii infects Hordeum spp. and includes lineages adapted to wild and domesticated barley. To date, the evolution of effector-like proteins in this species has not been addressed. RESULTS: We combined multiple structure-based and network analyses to unravel the secretome of Z. passerinii. We first compared AlphaFold2 and ESMFold predictions to establish the baseline for structural analyses. We identified 72 structural clusters in the secretome, revealing fold-level relationships across divergent sequences. We showed that effector-like proteins with predicted host immune-interfering functions evolved from a limited group of protein folds, whereas proteins with predicted antimicrobial properties were distributed across fold groups. Physicochemical comparisons indicate that putative antimicrobial effectors predominantly emerged through amino acid replacements on common effector-enriched scaffolds in Z. passerinii, reconfiguring surface charge and electrostatics. We analyzed intra- and interspecific variation in selected effector-enriched families by comparing Z. passerinii proteins and homologs across the genus Zymoseptoria. We describe constrained core folds, with local variation in loop and surface-exposed regions, consistent with fold stability while still enabling protein diversification. We further report that putative antimicrobial effector homologs are broadly distributed across the genus despite sequence divergence. CONCLUSIONS: The secretome of Z. passerinii is organized around common structural folds that support diverse biological roles, including host manipulation and host-associated microbial interactions. Conserved scaffolds combined with surface and physicochemical variation likely contribute to rapid adaptive evolution of effector-like proteins in Z. passerinii.

Fungal Proteins

FUSE-PhyloTree: linking functions and sequence conservation modules of a protein family through phylogenomic analysis.

SUMMARY: FUSE-PhyloTree is a phylogenomic analysis software for identifying local sequence conservation associated with the different functions of a multi-functional (e.g. paralogous or multi-domain) protein family. FUSE-PhyloTree introduces an original approach that combines advanced sequence analysis with phylogenetic methods. First, local sequence conservation modules within the family are identified using partial local multiple sequence alignment. Next, the evolution of the detected modules and known protein functions is inferred within the family's phylogenetic tree using three-level phylogenetic reconciliation and ancestral state reconstruction. As a result, FUSE-PhyloTree provides a gene tree annotated with both predicted sequence modules and ancestral gene functions, enabling the association of functions with specific sequence regions based on their co-emergence. AVAILABILITY AND IMPLEMENTATION: FUSE-PhyloTree is provided as Docker and Singularity images including all the required software tools. Images, source code, test data, and documentation are available at https://github.com/OcMalde/fuse-phylotree and https://zenodo.org/records/15855068.

Phylogeny

Structure and evolution of transplantation antigens: partial amino-acid sequences of H-2K and H-2D alloantigens.

Techniques for the amino acid sequence analysis of subnanomole quantities of polypeptides have been applied to characterize beta2-microglobulin and transplantation antigens of the mouse isolated from spleen cells by indirect immunoprecipitation. Eleven residues were identified throughout the NH2-terminal 27 residues of the beta2-microglobulin; all were identical to residues seen at the corresponding positions of beta2-microglobulins from other species. Two K and two D transplantation antigens were examined and the following generalizations emerged from the limited partial amino-acid sequence data: (1) the K and D molecules are homologous to one another; (2) they do not show amino acid sequence homology with immunoglobulins; (3) the two K and two D molecules differ from one another by multiple amino acid substitutions; and (4) the K molecules as a class cannot be distinguished from the D molecules as a class. The genetic and evolutionary implications of these observations are discussed.

Amino Acid Sequence

The molecular evolution of cytochrome c in eukaryotes.

Using many more cytochrome sequences than previously available, we have confirmed: 1, the eukaryotic cytochrome c diverged from a common ancestor; 2, the ancestral eukaryotic cytochrome c was not greatly different in character from those present today; 3, fixations are non-randomly distributed among the codons, there being evidence for at least four classes of variability; 4, there are similar classes of variability when the data are considered according to the nucleotide position within the codon; 5, the number of covarions (concomitantly variable codons) in mammalian cytochrome c genes is about 12 and the same value has been obtained for dicotyledenous plants as well; 6, all of the hyper- and most highly variable codons are for external residues, nearly 60 per cent of the invariable codons are for internal residues and nearly half of the codons for internal residues are invariable; 7, the first nucleotide position of a codon is more likely and the second position less likely to fix mutations than would be expected on the basis of the number of ways that alternative amino acids can be reached; 8, the character of nucleotide replacements is enormously non-random, with G-A interchanges representing 42% of those observed in the first nucleotide position, but the observation does not stem from a bias in the DNA strand receiving the mutation, nor from the presence of a compositional equilibrium, nor from a bias in the frequency with which different nucleotides mutate, but rather from a bias in the acceptability of an alternative nucleotide as circumscribed by the functional acceptability of the new amino acid encoded; and 9, the unit evolutionary period is approximately 150 million years/observable (amino acid changing) nucleotide replacement/cytochrome c covarion in two diverging lines. Wherever non-randomness has been observed, it has always been consistent with the consideration that an alternative amino acid at any location is more likely to be acceptable the more closely it resembles the present amino acid in its physico-chemical properties. Finally, in no case did the a priori assumption of a biologically realistic phylogeny lead to any observations or conclusions that were in any way significantly different from those obtained when the phylogeny was based solely upon the sequences, proving that the earlier results were not a consequence of some internal circularity.

Amino Acid Sequence

The biological origin of antibody diversity.

Antibody diversity has a compelling fascination for many scientists and over the years speculations have sometimes seemed more numerous than facts. Now the structural basis of antibody specificity is well defined. Amino acid sequences and recently three-dimensional structures of various immunoglobulins provide the most solid basis for discussing the origin of diversity. The novel pattern of variable (V) and Constant (C) regions of amino acid sequence has been resolved further to show the functional pattern of variability. Inheritance of separate V and C genes is accepted, but attempts to define more than one gene coding for each V region are considered here to be unnecessary. The pattern of variability is still best understood in terms of mutation and the presence or absence of various selective pressures. The major area of debate still hinges around the extent to which mutation and selection operate during evolution or somatically. Sequence data have now been generally interpreted to require multiple V genes carried in the germ line. A few individual VH genes have been mapped in close linkage to CH genes in the mouse. The apparent existence of three VH alleles in rabbits was a strong argument against multiple V genes. Now the three phenotypes have been shown to be due to alleles controlling the expression of three sets of VH genes all present on the same chromosome. That V-gene expression requires rejoining of V and C genes at the DNA level is now almost certain. Models for the joining process can draw on the precedents of transposable genetic elements, which are widespread in Nature. The total extent of antibody diversity remains a philosophical point. Estimates of the number of antibody molecules required for observed diversity are reduced by two recently documented proposals. Each antibody combining site apparently has many (estimated at 100) different specificities and most combinations of VH and VL regions probably form a viable site. A given combining site can be defined by its pattern of shared specificities. Several specific antibody repertoires have been measured and the size in each case is consistent with the stringency with which the specificity is selected. Repertoire size appears to be under genetic control, but there are problems in viewing the genotype through the veil of clonal selection. Molecular hybridization has been used recently in an attempt to count V and C genes directly. C genes are seen in DNA having nonreiterated sequences, as formal genetics predicts. Each V-region probe hybridizes at a similar rate to C-region probes. Interpretation of this result depends on the extent to which one V-region probe will reveal nonhomologous V genes. Previous estimates that many cross-hybridizing genes should have been seen if present are possibly exaggerated. It is argued here that the data are compatible with a germ-line gene for each probe studied. Maximum estimates for the number of germ-line genes are sufficient to account for antibody diversity...

Amino Acid Sequence

Clonal evolution of marker chromosomes in a case of myelofibrosis with myeloid metaplasia and myeloblastic transformation.

The diverse spectrum of acquired chromosome abnormalities in a female patient with myelofibrosis and myeloid metaplasia is described. A sequence of karyotypic evolution involving a ring chromosome is postulated. The terminal clinical picture was unusual in that there was obstructive renal failure from extramedullary myeloblastic transformation and infiltration of the bladder, and this was also present in other sites. Initially neutrophils showed low alkaline phosphatases activity but latterly two distinct populations in which cells had either high activity or none.

Alkaline Phosphatase

Genome-scale evolution and phylodynamics of swine influenza A viruses in China: a genomic epidemiology study.

BACKGROUND: Pigs are recognised as crucial intermediate hosts for the emergence of influenza viruses of pandemic potential. As the largest pork-producing nation, China hosts a complex ecosystem of swine influenza viruses (SIVs). We aimed to investigate the evolutionary processes, spatiotemporal dynamics, and biological characteristics of SIVs in China. METHODS: From Jan 15, 2016, to Dec 22, 2020, we collected nasal swabs from pigs at eight abattoirs and 16 swine farms in the Guangdong, Henan, and Shandong provinces of China, as part of SIV surveillance. SIVs were detected with RT-PCR. Positive samples underwent viral isolation and genome sequencing. We analysed evolution and spatiotemporal dynamics using the whole genomes of isolated SIVs, as well as genome sequences of SIV isolates from human infections worldwide retrieved from the Global Initiative on Sharing All Influenza Data and GenBank Flu databases up to April 28, 2024. Viral sequences without a sample collection area or date were excluded from the analysis. Viral receptor-binding properties and in-vitro replication of strains isolated in this study were evaluated with a solid-phase binding assay and various cell lines, including Madin-Darby canine kidney cells, porcine alveolar macrophages, primary porcine trachea epithelial cells, human bronchial epithelioid, and human lung adenocarcinoma epithelial (A549) cells. Viral replication and transmission studies were conducted in 33 guinea pigs and 13 pigs. Additionally, we collected serum samples from pig farm workers and members of the general public recruited by the Third Affiliated Hospital of Sun Yat-sen University between Feb 28 and May 11, 2023, to detect specific antibodies against Eurasian avian-like A(H1) and human-like A(H3N2) SIVs using the haemagglutination inhibition assay. FINDINGS: 23 (1·3%) of 1818 nasal swabs collected in abattoirs had SIVs; 22 (0·9%) of 2375 swabs from swine farms had SIVs. Further viral isolation yielded 39 strains of SIV. We identified 534 A(H1N1), 69 A(H1N2), and 92 A(H3N2) SIVs, representing 20 genotypes within the Eurasian avian-like lineage, 14 within the classical swine A(H1) lineage, and 16 within the human-like A(H3N2) lineage. The introduction of the A(H1N1)pdm/09 virus significantly influenced the internal gene pool of SIVs, enhancing genotypic diversity in China. Notably, the Eurasian avian-like A(H1), classical swine A(H1), and human-like A(H3N2) lineages showed human-mediated spread over long distances between provinces, with the Eurasian avian-like A(H1) lineage showing the most prevalent spread pathways. Eurasian avian-like A(H1) SIVs showed a preference for binding to sialic acid α-2,6 glycan receptors, predominantly found in humans, resulting in an increased production of progeny viruses in human airway epithelial cells, as well as effective transmission and infectivity among guinea pigs and pigs. Among 54 eligible serum samples collected from pig farm workers (24 from slaughterhouses and 30 from swine farms), 23 (43%) were seropositive for Eurasian avian-like A(H1) SIVs and 46 (85%) for human-like A(H3N2) SIVs. Among 100 eligible samples from members of the general public, 14 (14%) were seropositive for Eurasian avian-like A(H1) SIVs and 85 (85%) for human-like A(H3N2) SIVs. INTERPRETATION: This study elucidates the evolutionary processes and spatiotemporal patterns of SIVs, highlighting potential risks to public health. These findings are crucial for informing public health interventions that aim to prevent future SIV epidemics in China and other countries worldwide. FUNDING: Scientific Innovation Strategy-Construction of High-Level Academy of Agriculture Science-Distinguished Scholar (R2020PY-JC001).

Animals

Balancing under constraint: Structural insights into norovirus evolution and antigenic innovation.

Norovirus is the leading cause of acute viral gastroenteritis worldwide. While genomic studies have revealed its diversity and evolutionary patterns, the structural mechanisms driving viral adaptation remain poorly understood. Here, we establish a comprehensive structural database of norovirus VP1 P-domains across nine genogroups (GI-GIX) through large-scale AlphaFold2 predictions. By integrating phylogenetic analysis of VP1 sequences and structures, we demonstrate that sequence and structural evolution show overall concordance under purifying selection, yet significant local discrepancies reveal distinct patterns of convergent evolution shaped by structural constraints and functional divergence. Focusing on the predominant GII.4 genotype, we found that compared to near-full-genome and nucleotide trees, only the VP1 amino acid tree reliably clustered GII.4 variants in chronological order as monophyletic groups. We further identify a hierarchical evolutionary strategy: positive selection may drive structural hypervariability in major antigenic epitopes D and C for immune escape, with epitope D exhibiting pronounced structural flexibility that complicates its structural characterization, whereas coevolutionary analysis uncovers a broad network of compensatory interactions spanning multiple epitopes, with striking enrichment in epitope A. These epitopes exhibited a pattern of "sequence plasticity with structural conservation", maintained by coevolutionary constraints that preserve conformational integrity. Together, these findings suggest that norovirus vaccine strategies targeting the structurally conserved conformations of epitopes A and G could overcome the limitations of traditional strain-specific approaches, offering a pathway toward broad protection against evolving viral diversity.

Norovirus

Transition of Staphylococcus aureus tetracycline resistance plasmid pT181 from independent multicopy replicon to predominantly integrated chromosomal element over 65 years.

Mobile genetic elements (MGEs), including plasmids, phages and genome islands, are major sources of bacterial genetic diversity. The small plasmid pT181 confers tetracycline resistance in bacterial pathogen Staphylococcus aureus via an efflux pump, TetK. pT181 was one of the earliest sequenced S. aureus plasmids, and has been isolated in both clinical and livestock-associated strains for decades, both as an independent replicon and integrated in the chromosome as part of staphylococcal cassette chromosome mec (SCCmec). Bacterial genome analysis tools and high-quality sequences with metadata are publicly available, but these resources remain underleveraged for examining historical data, especially when studying the spread of MGEs across a species and over time. Using publicly available reads and metadata, we explored the evolution of pT181 over almost seven decades of samples to identify temporal trends in sequence evolution, copy number changes, and spread across S. aureus and beyond. pT181 was prevalent across S. aureus (found in 9.5% of 83,366 genomes tested), with a conserved sequence outside of three hypervariable regions. The history of pT181 since 1954 is characterized by spread across strains, significant variation in plasmid copy number of the independent replicon, and increasing frequency of integration of the plasmid into the S. aureus chromosome. We have identified multiple chromosomal integration locations of the plasmid, including outside of the previously characterized SCCmec. We find that pT181 has been transferred across staphylococcaceae and into a Gram-negative species. The repeated integration of pT181 into the chromosome may indicate co-evolution of the plasmid and the host, potentially to facilitate increased antibiotic resistance.

Journal Article

The physical and evolutionary energy landscapes of devolved protein sequences corresponding to pseudogenes.

Protein evolution is guided by structural, functional, and dynamical constraints ensuring organismal viability. Pseudogenes are genomic sequences identified in many eukaryotes that lack translational activity due to sequence degradation and thus over time have undergone "devolution." Previously pseudogenized genes sometimes regain their protein-coding function, suggesting they may still encode robust folding energy landscapes despite multiple mutations. We study both the physical folding landscapes of protein sequences corresponding to human pseudogenes using the Associative Memory, Water Mediated, Structure and Energy Model, and the evolutionary energy landscapes obtained using direct coupling analysis (DCA) on their parent protein families. We found that generally mutations that have occurred in pseudogene sequences have disrupted their native global network of stabilizing residue interactions, making it harder for them to fold if they were translated. In some cases, however, energetic frustration has apparently decreased when the functional constraints were removed. We analyzed this unexpected situation for Cyclophilin A, Profilin-1, and Small Ubiquitin-like Modifier 2 Protein. Our analysis reveals that when such mutations in the pseudogene ultimately stabilize folding, at the same time, they likely alter the pseudogenes' former biological activity, as estimated by DCA. We localize most of these stabilizing mutations generally to normally frustrated regions required for binding to other partners.

Cyclophilin A