PubMed HealthSearch

SEARCH · PubMed Health

Results for “subgenome evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

18 recordsLinked to original sources

The genome of Lespedeza potaninii reveals biased subgenome evolution and drought adaptation.

Lespedeza potaninii, a xerophytic subshrub belonging to the legume family, is native to the Tengger Desert and is highly adapted to drought. It has important ecological value due to its drought adaptability, but the underlying molecular mechanisms remain largely unknown. Here, we report a 1.24 Gb chromosome-scale assembly of the L. potaninii genome (contig N50 = 15.75 Mb). Our results indicate that L. potaninii underwent an allopolyploid event with 2 subgenomes, A and B, presenting asymmetric evolution and B subgenome dominance. We estimate that the 2 diploid progenitors of L. potaninii diverged around 3.6 million years ago (MYA) and merged around 1.0 MYA. We revealed that the expansion of hub genes associated with drought responses, such as the binding partner 1 of accelerated cell death 11 (ACD11) (BPA1), facilitated environmental adaptations of L. potaninii to desert habitats. We found a novel function of the BPA1 family in abiotic stress tolerance in addition to the known role in regulating the plant immune response, which could improve drought tolerance by positively regulating reactive oxygen species homeostasis in plants. We revealed that bZIP transcription factors could bind to the BPA1 promoter and activate its transcription. Our work fills the genomic data gap in the Lespedeza genus and the tribe Desmodieae, which should provide theoretical support both in the study of drought tolerance and in the molecular breeding of legume crops.

Genome, Plant

Chromosome-level assembly and annotation of the yellow-shelled fish (Barbodes Wynaadensis).

Barbodes wynaadensis, a unique cyprinid species native to Yunnan Province in China, stands out as an allotetraploid (AABB) fish with a complex evolutionary history. Leveraging a multi-platform sequencing strategy combining MGI short-read, PacBio long-read, and Hi-C scaffolding technologies, we assembled the first chromosome-level genome for B. wynaadensis. The final assembled genome spans 1.76 Gb in length with a contig N50 of 33.53 Mb, demonstrating high assembly continuity. Hi-C scaffolding enabled the reconstruction of 50 pseudochromosomes, representing 99.94% of the total genome assembly. Genome annotation identified 46,121 protein-coding genes, with a functional annotation rate of 99.76%. Repetitive elements constituted 48.26% of the genomic sequences, including lineage-specific expansions of DNA transposons (29.26%) and LTRs (6.36%). This high-quality assembly resolves challenges in polyploid genome reconstruction and provides a critical resource for investigating Cyprinidae evolution, particularly subgenome divergence and adaptation. The dataset also enables practical applications, such as molecular marker development for population monitoring, supporting conservation efforts for this threatened endemic species amid habitat degradation in the Nujiang River basin.

Animals

Doubled Genomes, Divergent Fates: Genomic Insights Into Diversification in an Allotetraploid Cavefish.

Cave environments impose unique challenges that drive remarkable genetic and phenotypic changes in cave-dwelling organisms. In this study, we investigated the genomic basis of adaptation in the small eye golden-line fish (Sinocyclocheilus microphthalmus), an allotetraploid cavefish endemic to Guangxi, China. Using whole-genome resequencing data from 47 individuals across six cave locations, we examined how neutral and selective forces influence diversification. Our analyses uncovered significant population structure indicative of allopatric divergence, along with evidence of locus-specific selection contributing to genomic differentiation. We identified seven single outlier clusters (SOCs), each tied to the divergence of specific populations, underscoring the role of local processes in driving diversity. Genes associated with vision showed relaxed selection, likely reflecting adaptation to darkness, while positive selection on other loci revealed additional functional shifts. Notably, allopolyploidy was found to fuel divergence through subgenome-specific patterns and asymmetric evolution within SOCs and among homoeologs. Taken together, these findings provide valuable insights into mechanisms of cave evolution and illustrate how allotetraploid genomes can facilitate diversification, potentially contributing to speciation in extreme environments.

Animals

Contrasting regulation of protein-coding genes and lncRNA homeologs in allotetraploid Coffea arabica.

A chromosome-level Bourbon assembly revealed that protein-coding homeologs are predominantly co-regulated between subgenomes. In contrast, intergenic lncRNAs display a modest, but statistically consistent bias toward subgenome E across diverse developmental and stress contexts. Coffea arabica is an allotetraploid species derived from natural hybridization between C. canephora and C. eugenioides, which contributed the C and E subgenomes, respectively. This genomic origin poses major challenges for genome assembly, annotation, and the interpretation of gene regulation. In this study, a high-quality genome assembly of C. arabica was generated and annotated, with particular emphasis on identifying protein-coding genes and intergenic long non-coding RNAs (lincRNAs). Homeologous relationships between genes from the C and E subgenomes were established, providing a robust framework to investigate subgenomic conservation and regulatory divergence. Using an extensive collection of publicly available RNA-seq libraries spanning multiple developmental stages, tissues, and environmental conditions, the relative transcriptional contribution of each subgenome was evaluated. On a global scale, gene expression was largely balanced between subgenomes, with no consistent evidence of subgenome dominance. While protein-coding genes showed comparable regulatory behavior across subgenomes, lincRNAs exhibited a more asymmetric expression pattern, suggesting higher subgenome-specific expression that is interpreted here as a consistent directional tendency rather than as evidence of subgenome dominance. Together, these results provide new insights into the regulatory architecture of the C. arabica genome and establish a foundational genomic and transcriptomic resource for future functional studies and crop improvement efforts.

Coffea

Subgenomic divergence and functional innovation following whole-genome duplication in Maleae species of Rosaceae.

Whole-genome duplication (WGD) drives plant evolution by inducing karyotype rearrangements and gene loss through subgenome fractionation. In this study, we investigate post-WGD evolutionary dynamics in Rosaceae, focusing on Maleae species, which uniquely experienced an additional WGD. Using phylogenetic and synteny analyses, we reveal that chromosomal breakpoints act as hotspots for localized fractionation, contributing to blurred homoeologous origins and influencing gene retention patterns. Here, we reconstruct karyotype evolution across Rosaceae subfamilies, highlighting chromosome reductions and lineage-specific rearrangements in Dryadoideae, Rosoideae, and Amygdaloideae. We also identify a bias for retaining transcription factors and hormone-related genes from older WGDs in subsequent polyploidy events. Transcriptome analysis classifies WGD-derived genes in Maleae species, such as apple and loquat, into three expression groups, with hormone-enriched genes playing roles in lignification and fruit-related innovations. These findings demonstrate the interplay between chromosomal breakpoints, biased retention, and functional divergence, revealing their contributions to genomic and phenotypic evolution in Maleae and their adaptive success within Rosaceae.

Genome, Plant

Distinct evolutionary trajectories of subgenomic centromeres in polyploid wheat.

BACKGROUND: Centromeres are crucial for precise chromosome segregation and maintaining genome stability during cell division. However, their evolutionary dynamics, particularly in polyploid organisms with complex genomic architectures, remain largely enigmatic. Allopolyploid wheat, with its well-defined hierarchical ploidy series and recent polyploidization history, serves as an excellent model to explore centromere evolution. RESULTS: In this study, we perform a systematic comparative analysis of centromeres in common wheat and its corresponding ancestral species, utilizing the latest comprehensive reference genome assembly available. Our findings reveal that wheat centromeres predominantly consist of five types of centromeric-specific retrotransposon elements (CRWs), with CRW1 and CRW2 being the most prevalent. We identify distinct evolutionary trajectories in the functional centromeres of each subgenome, characterized by variations in copy number, insertion age, and CRW composition. By utilizing CENH3-ChIP data across various ploidy levels, we uncover a series of CRW invasion events that have shaped the evolution of AA subgenome centromeres. Conversely, the evolutionary process of the DD subgenome centromeres involves their expansion from diploid to hexaploid wheat, facilitating adaptation to a larger genomic context. Integration of complete einkorn centromere assemblies and Aegilops tauschii pan-genomes further revealed subgenome-specific centromere evolutionary trajectories. By inclusion of synthetic hexaploid from S2-S3 generations, alongside 2x/6 × natural accessions, we demonstrate that DD subgenome centromere expansion represents a gradual evolutionary process rather than an immediate response to polyploidization. CONCLUSIONS: Our study provides a comprehensive landscape of centromere adaptation, evolution, and maturation, along with insights into how retrotransposon invasions drive centromere evolution in polyploid wheat.

Centromere

The centromere landscapes of four karyotypically diverse Papaver species provide insights into chromosome evolution and speciation.

Understanding the roles played by centromeres in chromosome evolution and speciation is complicated by the fact that centromeres comprise large arrays of tandemly repeated satellite DNA, which hinders high-quality assembly. Here, we used long-read sequencing to generate nearly complete genome assemblies for four karyotypically diverse Papaver species, P. setigerum (2n = 44), P. somniferum (2n = 22), P. rhoeas (2n = 14), and P. bracteatum (2n = 14), collectively representing 45 gapless centromeres. We identified four centromere satellite (cenSat) families and experimentally validated two representatives. For the two allopolyploid genomes (P. somniferum and P. setigerum), we characterized the subgenomic distribution of each satellite and identified a "homogenizing" phase of centromere evolution in the aftermath of hybridization. An interspecies comparison of the peri-centromeric regions further revealed extensive centromere-mediated chromosome rearrangements. Taking these results together, we propose a model for studying cenSat competition after hybridization and shed further light on the complex role of the centromere in speciation.

Centromere

Transposable elements as modulators of homoeologous gene expression in bread wheat: lessons from the pan-transcriptome era.

Bread wheat (Triticum aestivum L.) is an allohexaploid (AABBDD) whose three ancestral subgenomes generate complex patterns of gene regulation. Most genes exist as homoeologous triads, and the relative expression balance among copies, homoeolog expression bias, is central to polyploid evolution and adaptation. Recent high-quality assemblies, long-read transcriptomics, and pan-transcriptome resources have uncovered extensive cultivar-specific transcriptional diversity. Because transposable elements (TEs) compose over 80% of the wheat genome, they are prime candidates for shaping subgenome asymmetry. We synthesize recent pan-genomic and transcriptomic evidence, including genome-wide associations between TE insertions and genome-specific expression, and propose a unifying framework in which TEs modulate homoeolog expression by donating cis-regulatory sequences, altering chromatin states, producing small RNAs, and driving structural variation. We discuss experimental and computational challenges for establishing causality, and outline future functional and translational strategies to leverage TE-associated regulatory diversity in wheat breeding.

Triticum

Genome evolution of the ancient hexaploid Platanus × acerifolia (London planetree).

Whole-genome duplication (WGD; i.e., polyploidy) and chromosomal rearrangement (i.e., genome shuffling) significantly influence genome structure and organization. Many polyploids show extensive genome shuffling relative to their pre-WGD ancestors. No reference genome is currently available for Platanaceae (Proteales), one of the sister groups to the core eudicots. Moreover, Platanus × acerifolia (London planetree; Platanaceae) is a widely used street tree. Given the pivotal phylogenetic position of Platanus and its 2-y flowering transition, understanding its flowering-time regulatory mechanism has significant evolutionary implications; however, the impact of Platanus genome evolution on flowering-time genes remains unknown. Here, we assembled a high-quality, chromosome-level reference genome for P. × acerifolia using a phylogeny-based subgenome phasing method. Comparative genomic analyses revealed that P. × acerifolia (2n = 42) is an ancient hexaploid with three subgenomes resulting from two sequential WGD events; Platanus does not seem to share any WGD with other Proteales or with core eudicots. Each P. × acerifolia subgenome is highly similar in structure and content to the reconstructed pre-WGD ancestral eudicot genome without chromosomal rearrangements. The P. × acerifolia genome exhibits karyotypic stasis and gene sub-/neo-functionalization and lacks subgenome dominance. The copy number of flowering-time genes in P. × acerifolia has undergone an expansion compared to other noncore eudicots, mainly via the WGD events. Sub-/neo-functionalization of duplicated genes provided the genetic basis underlying the unique flowering-time regulation in P. × acerifolia. The P. × acerifolia reference genome will greatly expand understanding of the evolution of genome organization, genetic diversity, and flowering-time regulation in angiosperms.

Polyploidy

Discovery of the order 'Quisvirales' redefines the evolution of RNA replication and transcription in the phylum Pisuviricota.

Genome replication in positive-stranded RNA (ssRNA+) viruses is mediated by cognate enzymes, including ubiquitous RNA-dependent RNA polymerase (RdRp). In ssRNA+ viruses with multiple open reading frames (ORFs) in their genomes, replication often is accompanied by synthesis of subgenomic RNAs (transcription) for expression of 3'-proximal ORFs. In addition, all ssRNA+ viruses with genomes larger than ~7 kb encode helicases, linking helicases to RNA genome expansion. Helicases are essential ATPases that unwind nucleic acids and are classified into six recognized superfamilies (SF1-SF6). In the phylum Pisuviricota that includes important pathogens, helicases of SF1-SF3 are integrated into multi-enzyme replicase polyprotein(s) including 3C(-like) protease (3CLpro) and RdRp. Here, large-scale mining of invertebrate metatranscriptomes and targeted genome sequence assembly uncovered six spider-associated ssRNA+ viruses that, based on their conserved 3CLpro-RdRp module in replicase polyproteins, genome size (20-22 kb), and phylogeny, form a family-like cluster in a putative order, named 'Quisvirales'. Quisviruses have similar genome and replicase architectures to enveloped coronaviruses and other nidoviruses. Notably, quisviruses encode ORFs 1a and 1b with predicted -1 programmed ribosomal frameshifting elements in the ORF1a/b overlap region. Using an original mapping approach for detecting chimeric sequencing reads, we obtained evidence that 3'-proximal ORFs are expressed via 5'-coterminal, leader-containing subgenomic RNAs. This suggests that the quisvirus subgenomic RNAs are generated through discontinuous transcription-a mechanism otherwise exclusively found in nidoviruses among the many ssRNA+ virus orders that synthesize subgenomic RNAs. Striking differences between nido- and quisviruses are, however, the RdRp being the only common core ORF1b-encoded enzyme and the replacement of the nidovirus SF1 helicase by a novel superfamily helicase. This quisvirus SF7 helicase, like the Picornavirales SF3 helicase, comprises an AAA+ (ATPase-like) domain typical for ring-forming helicases and thus must play an essential role in replication. The discovery of the order 'Quisvirales' demonstrates that viruses employing large replicase polyproteins of nidovirus-like complexity and discontinuous transcription may have evolved repeatedly from an 3CLpro-RdRp-encoding ancestor.

AAA+/RecA-like ATPase

A single hybrid origin of cultivated peanut.

This study, the first in a three-part series, lays the foundation for understanding the origin of the peanut crop (Arachis hypogaea). Its subsequent evolution is explored in the two papers that follow. The evidence that A. hypogaea originated from a single hybridization event between Arachis duranensis and Arachis ipaënsis less than 10 000 years ago was already very strong. Here, we extend this evidence using more than 1600 single-nucleotide polymorphisms to make an almost exhaustive comparison of wild Arachis section germplasm conserved ex situ with the A and B subgenomes of divergent, sequenced cultivated peanuts. The wild relatives of peanut are highly selfing and their geocarpy means they plant their own seeds, allowing them to persist as discrete populations for millennia. This unusual biology creates a rare opportunity for genetic archaeology: ancestral lineages can be identified with exceptional precision. Our results reaffirm a single origin for the cultigen, identifying A. duranensis from Río Seco and A. ipaënsis K 30076 as the closest known relatives of the A and B subgenomes of peanut. As a genomic resource, we generated a chromosome-scale assembly of the Río Seco A. duranensis K 30065 and confirmed that it is more closely related to the A subgenome of peanut than the current reference genome (V14167). Even if somewhat closer wild accessions were found through new field collections, they would still belong to the same ancestral lineage. With this level of evidence, the origin of peanut is now known in greater detail than that of any other ancient polyploid crop.

Arachis

A high-quality draft genome assembly of Johnsongrass illuminates relationships between polyploidization, crop-wild hybridization, and reproductive biology.

Johnsongrass [Sorghum halepense (L.) Pers.] is an allopolyploid, rhizomatous, perennial grass species and one of the most troublesome weeds in global agriculture. We assembled the first Johnsongrass genome to clarify poorly understood genetic factors influencing variable rates of crop-wild hybridization with cultivated sorghum [S. bicolor (L.) Moench]. The draft genome assembly has a total size of 3.26 Gb and BUSCO completeness of 95.3%. We also report the first evolutionary analysis of INHIBITION OF ALIEN POLLEN (IAP), the only known cross-(in)compatibility locus in the genus. Our results reveal an evolutionary history of genome instability, including the loss of distinct parental subgenomes, and suggest that Nebraska accession 'J-37,' the genome donor, is a segmental allotetraploid that may function as a diploid or aneuploid during meiosis. Genome instability could explain observations of variable ploidies in Johnsongrass and facilitate ongoing hybridization with sorghum where gamete ploidies and IAP alleles match. Given this information, we provide a suggested research framework for studying evolution and gene expression in the Sorghum genus where crop-wild hybridization occurs and for predicting the potential for hybridization between specific crossing partners. Collectively, this work will bolster efforts to study and manage reproductive biology in other crop-wild polyploid complexes.

Sorghum

Fishing for a reelGene: evaluating gene models with evolution and machine learning.

Assembled genomes and their associated annotations have transformed our study of gene function. However, each new annotated assembly generates new gene models. Inconsistencies between annotations likely arise from biological and technical causes, including pseudogene misclassification, transposon activity, and intron retention from sequencing of unspliced transcripts. To evaluate gene model predictions, we developed reelGene, a pipeline of machine learning models focused on (1) transcription boundaries, (2) mRNA integrity, and (3) protein structure. The first two models leverage sequence characteristics and evolutionary conservation across related taxa to learn the grammar of conserved transcription boundaries and mRNA sequences, while the third uses the conserved evolutionary grammar of protein sequences to predict whether a gene can produce a protein. Evaluating 1.8 million transcript models in Zea mays ssp. mays (maize), reelGene classified 28% as incorrectly annotated or non-functional. We find that reelGene classifies 92.2% of genes in the maize proteome and 99.2% of genes within the maize classical gene list as functional. reelGene also provides a way to further investigate genome biology- for instance, reelGene indicates that 10.3% of dispensable genes in B73 are functional, and within retained duplicate genes, reelGene identifies a 30% bias toward the retention of the M1 subgenome when one copy is functional and the other is non-functional. As an annotation-evaluating tool, reelGene is directly applicable to species of the Andropogoneae tribe, including other important crops like sorghum and miscanthus. As a community resource, reelGene has been integrated onto MaizeGDB both as a browser track and as an individual Shiny App, allowing researchers to evaluate gene model accuracy and further investigate genome biology.

Machine Learning

3D chromatin remodeling during domestication defines novel targets for crop improvement.

Three-dimensional (3D) genome folding shapes gene regulation, yet the genetic underpinnings linking 3D genome evolution to phenotypic innovation during domestication remain elusive. Using population-scale Hi-C profiling of 34 semi-wild and 267 cultivated allotetraploid cottons, we generated a pan-3D genome atlas capturing extensive diversity in topologically associating domains (TADs) and chromatin loops. Chromatin interactome-wide association studies identified 105 TAD reconfigurations and 58 loop rewirings that were established as the 3D chromatin basis of fiber quality, boosting heritability estimates for fiber strength by 16% and fiber length by 20%. We reveal that domestication selection within sequence-defined sweeps fixed 57% of 3D conformation signatures, thereby decoupling sequence-level from chromatin-level selection and shifting the subgenome expression balance of 39 homoeologs in cultivated cotton. Sequence-based modeling and mutational analyses identified the C2H2 zinc-finger protein YY1 as a conserved mediator of 3D genome organization. This study provides a resource for redefining precision-breeding paradigms by harnessing cryptic 3D chromatin targets.

3D genome

Mutation accumulation in a hybrid parthenogenetic vertebrate.

Asexual lineages are thought to experience elevated extinction rates compared with sexual species, yet direct evidence for the underlying genetic causes remains scarce. Muller's ratchet predicts that the absence of recombination in asexual organisms facilitates the accumulation of deleterious mutations, thereby reducing long-term fitness. Here, we test this hypothesis in the hybrid-origin, parthenogenetic whiptail lizard Aspidoscelis tesselatus by integrating short-read RNAseq and long-read IsoSeq data from both the asexual lineage and its parental sexual species. We reconstructed phased transcripts for A. tesselatus to quantify mutation accumulation relative to the parental sexual species. Comparative analyses revealed elevated ω ratios in both parental genomic complements (subgenomes) of the parthenogenetic lineage, consistent with accelerated accumulation of nonsynonymous mutations. Structural variant analyses identified multiple indels in expressed transcripts predicted to disrupt protein domains. Functional annotation indicated that genes affected by both single-nucleotide variants and indels were enriched for roles in chromatin organization, apoptosis regulation, and transcriptional control. While both parental subgenomes showed similar evolutionary patterns, the maternal complement exhibited more structural and missense mutations than the paternal complement. Together, these results provide evidence that mutations accumulate in asexual A. tesselatus in genes involved in core cellular functions, supporting theoretical predictions that Muller's ratchet contributes to mutation accumulation in asexual lineages.

Animals

Genomic surveillance of enterovirus D68 circulating in 2025 reveals the emergence of a novel A2/B3 recombinant lineage.

Enterovirus D68 (EV-D68) has re-emerged over the past decade as a significant respiratory pathogen associated with severe respiratory disease and acute flaccid myelitis. Its circulation has typically followed a biennial pattern, with predominance in late summer and early fall, a pattern that was temporarily disrupted during the COVID-19 pandemic. Surveillance in 2025 revealed off-season circulation of EV-D68. This study describes the genomic characteristics of the 2025 EV-D68 viruses and the clinical features of affected patients. Between May and December 2025, remnant respiratory specimens positive for rhinovirus/enterovirus were screened for EV-D68 and subjected to whole-genome sequencing. Phylogenetic analyses were performed using maximum-likelihood methods. Recombination was assessed using subgenomic phylogenies, SimPlot similarity and BootScan analyses, and read-level inspection. Among 1,321 patients tested, 147 (11.1%) were EV-D68-positive, and 119 (81.0%) yielded complete genomes. EV-D68 positivity increased in July 2025, peaked in August (~21%), and remained elevated through September and October, exceeding levels observed in 2024. Patients had a median age of 36 years, with infections disproportionately affecting older adults. Phylogenetic analysis demonstrated exclusive circulation of subclade A2. Five genomes formed a distinct recombinant lineage (A2-Re). Subgenomic phylogenies showed clustering with A2 viruses in the P1 region and with B3 viruses in the P2-P3 regions. SimPlot and BootScan analyses identified a recombination breakpoint near the 2A/2B junction (~nt 3,700). The recombinant lineage was associated with temporally clustered cases in September-October. These findings demonstrate recombination between distinct EV-D68 subclades and underscore the importance of whole-genome surveillance for accurate viral characterization. Continued genomic monitoring is essential for detecting emerging variants with potential implications for transmissibility, pathogenicity, and public health preparedness.IMPORTANCEThis study highlights an increased off-season circulation of Enterovirus D68 (EV-D68) and a higher burden of disease in adults in 2025. The identification of a novel A2-B3 recombinant lineage provides evidence of ongoing viral evolution through recombination, a mechanism that may alter transmissibility, virulence, or immune responses. Detection of this lineage in temporally clustered cases suggests local transmission and underscores the potential for rapid spread of newly emerged variants. These findings emphasize the limitations of partial genomic approaches and the critical role of whole-genome sequencing in accurately characterizing circulating strains and identifying recombination events. Enhanced genomic surveillance is essential to detect emerging variants in real time, inform diagnostic assay performance, and support public health responses. Continued monitoring of EV-D68 evolution will be important for anticipating changes in disease burden, guiding clinical awareness, and strengthening preparedness for future outbreaks.

Humans

A Young ahsg/fetuin-a Inactive Retrocopy Reflects Recent Retrotransposon Activity in the Xenopus laevis Lineage.

The vertebrate ahsg (alpha 2-HS glycoprotein, also coined fetuin-a) homologs are highly expressed in the liver, and their secreted protein products exert complex systemic effects, including the regulation of biomineralization of soft and skeletal tissues. Here, we report a previously uncharacterized ahsg retrocopy in the allotetraploid frog species Xenopus laevis. We show that this young retrocopy was born from the ahsg.L homeologue less than 10 Mya, and landed in the S subgenome in a locus located between asic2.S and smarcd2.S. The ahsg.L-retrocopy ends with a poly(A) tail, is intronless, and is flanked by target site duplications. While the ahsg.L-retrocopy's ORF is devoid of frameshifts and nonsense mutations, it suffers from a short 5' deletion, eliminating the original start codon and the signal peptide. Remarkably, this truncated ORF lies in frame with an ATG codon contributed by the neighboring genomic sequence, suggesting that the ahsg.L-retrocopy might potentially be expressed and translated into a protein product. Nevertheless, examination of RNA-Seq and proteomic experiments respectively performed on liver and bone tissues did not provide expression evidence for the ahsg.L-retrocopy. We propose that, in spite of its rescued ORF, the ahsg.L-retrocopy is non-functional and can be considered a young pseudogene born from recent retrotransposon activity in the Xenopus laevis lineage.

Animals

Direct RNA nanopore sequencing of full-length coronavirus genomes provides novel insights into structural variants and enables modification analysis.

Sequence analyses of RNA virus genomes remain challenging owing to the exceptional genetic plasticity of these viruses. Because of high mutation and recombination rates, genome replication by viral RNA-dependent RNA polymerases leads to populations of closely related viruses, so-called "quasispecies." Standard (short-read) sequencing technologies are ill-suited to reconstruct large numbers of full-length haplotypes of (1) RNA virus genomes and (2) subgenome-length (sg) RNAs composed of noncontiguous genome regions. Here, we used a full-length, direct RNA sequencing (DRS) approach based on nanopores to characterize viral RNAs produced in cells infected with a human coronavirus. By using DRS, we were able to map the longest (∼26-kb) contiguous read to the viral reference genome. By combining Illumina and Oxford Nanopore sequencing, we reconstructed a highly accurate consensus sequence of the human coronavirus (HCoV)-229E genome (27.3 kb). Furthermore, by using long reads that did not require an assembly step, we were able to identify, in infected cells, diverse and novel HCoV-229E sg RNAs that remain to be characterized. Also, the DRS approach, which circumvents reverse transcription and amplification of RNA, allowed us to detect methylation sites in viral RNAs. Our work paves the way for haplotype-based analyses of viral quasispecies by showing the feasibility of intra-sample haplotype separation. Even though several technical challenges remain to be addressed to exploit the potential of the nanopore technology fully, our work illustrates that DRS may significantly advance genomic studies of complex virus populations, including predictions on long-range interactions in individual full-length viral RNA haplotypes.

Cell Line