PubMed HealthSearch

SEARCH · PubMed Health

Results for “Transposable element identification”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

15 recordsLinked to original sources

Hide and seek: de novo identification in sugar beet reveals impact of non-autonomous LTR retrotransposons.

Plant genomes are filled with retrotransposons and their derivatives, constantly undergoing sequence diversification and structural rearrangement. Among them, short, non-autonomous retrotransposons lack full coding capacity and often form subfamilies. As a result, non-autonomous retrotransposons are incompletely identified in most to all genome assemblies.Here, we capitalize on our comprehensive understanding of the transposable element (TE) landscape in sugar beet (Beta vulgaris) to assess the extent of the blind spot for non-autonomous long terminal repeat (LTR) retrotransposons. This use case serves to answer if all of these sequences are derivatives of easier-to-identify full-length elements or if there is more variability that is currently overlooked.For this we applied a semi-automated structural discovery workflow followed by in-depth manual verification to characterize non-autonomous LTR retrotransposons in sugar beet. We retrieve more than 100 non-autonomous LTR retrotransposon families that lack complete autonomous coding capacity, including canonical terminal-repeat retrotransposons in miniature (TRIMs), elongated non-coding derivatives and families retaining fragmented coding remnants. The identified families span a broad range, including elements exceeding 15,000 bp in length and display evidence for reshuffling and modular evolution. Only a subset of families could be confidently linked to autonomous retrotransposons, showing sequence diversification within the non-autonomous LTR retrotransposon fraction beyond the autonomous genomic templates.We highlight that a large fraction of non-autonomous LTR retrotransposons is incompletely recovered with the current TE identification workflows, even if the output is well-curated and condensed into TE libraries and suggest procedures to remedy this gap. This study gives a genome-wide view into the non-autonomous LTR retrotransposon landscape of a single plant genome and highlights the importance of structure-based approaches for their identification and classification.

LTR retrotransposons

Hydrolytic endonucleolytic ribozyme (HYER): Systematic identification, characterization and potential application in nucleic acid manipulation.

Group II introns are transposable elements that can propagate in host genomes through the "copy and paste" mechanism. They usually comprise RNA and protein components for effective propagation. Recently, we found that some bacterial GII-C introns without protein components had multiple copies in their resident genomes, implicating their potential transposition activity. We demonstrated that some of these systems are active for hydrolytic DNA cleavage and proved their DNA manipulation capability in bacterial or mammalian cells. These introns are therefore named HYdrolytic Endonucleolytic Ribozymes (HYERs). Here, we provide a detailed protocol for the systematic identification and characterization of HYERs and present our perspectives on its potential application in nucleic acid manipulation.

RNA, Catalytic

Long-read sequencing reveals a hidden Alu-mediated splice defect in CPLANE1, causing orofaciodigital syndrome type VI.

Orofaciodigital syndrome type VI (OFD VI) is a recessive ciliopathy characterized by excessive polydactyly, molar tooth sign, cleft lip, and developmental delay, caused by pathogenic variants in CPLANE1. Here, we present a patient with OFD VI that remained genetically unexplained after routine genetic testing, including short-read whole genome sequencing (WGS). Using long-read sequencing, we found two biallelic splice-site variants in CPLANE1, c.8633-4_8633-3del, and an Alu element insertion close to an exon-intron boundary. Transcript analysis showed that each variant independently resulted in exon skipping, and quantitative expression studies revealed reduced total CPLANE1 mRNA levels in patient-derived fibroblasts. Based on these findings, we were able to re-classify the c.8633-4_8633-3del variant from a variant of uncertain significance (VUS) to likely pathogenic. The identification of an Alu element insertion missed by short-read WGS highlights the added diagnostic value of long-read sequencing in uncovering cryptic, transposable element-associated pathogenic variants.

Journal Article

NextLongIso: a comprehensive Nextflow pipeline for multi-dimensional long-read RNA-seq analysis.

SUMMARY: Long-read RNA sequencing technologies, including Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (ONT), enable direct characterization of full-length transcripts and transcriptome complexity. However, analysis of long-read RNA-seq data remains fragmented across multiple tools, limiting the ability to obtain a unified view of transcript structure, expression, and regulatory variation in long-read transcriptomes. We present NextLongIso, a scalable and reproducible Nextflow pipeline that enables coordinated analysis of multiple layers of transcript regulation. Rather than focusing solely on transcript reconstruction, NextLongIso integrates transcript discovery with downstream regulatory analyses to jointly characterize alternative splicing, isoform switching, transcript boundary dynamics (including alternative promoters and polyadenylation), and transposable element-associated transcription from both PacBio and ONT datasets. By eliminating complex cross-tool data harmonization, this unified framework facilitates the transition from transcript identification to functional interpretation of transcriptomic variation. AVAILABILITY AND IMPLEMENTATION: NextLongIso is implemented in Nextflow and is freely available at github: https://github.com/YidanSunResearchLab/nf-LongIso.git and Zenodo: https://doi.org/10.5281/zenodo.21049837.

Software

Beyond Canonical Neoantigens: Emerging Technologies for Identification of Noncanonical Antigens and Implications for Personalized Cancer Vaccines.

Over the past decade, advances in sequencing technologies and computational pipelines enabled the development of personalized cancer vaccines (PCVs). Current PCV strategies primarily target cancer neoantigens generated by non-synonymous DNA mutations, which can result in altered amino acid sequences capable of eliciting tumor-specific immune responses. More recently, a distinct class of tumor-specific antigens (TSA), termed noncanonical or cryptic antigens, has emerged as an additional source of immunogenic targets. Unlike canonical neoantigens, noncanonical antigens typically cannot be identified by tumor/normal whole-exome sequencing, as they do not arise from classical DNA mutations. Instead, they are often associated with less well recognized and/or aberrant processes in the pathways from DNA to human leukocyte antigen (HLA)-presented peptides. Examples include transposable elements, circular RNA, translation of alternative open reading frames and/or long non-coding RNA, among others. Emerging evidence suggests that noncanonical antigens represent a substantial portion of the tumor-specific immunopeptidome and, similar to canonical neoantigens, are absent during thymic selection and can evade central tolerance and elicit T cell responses. Technological advances have increasingly facilitated the identification of noncanonical antigens. Long-read RNA sequencing reveals noncanonical transcripts by improving transcriptome assembly, while ribosome profiling provides genome-wide maps of actively translated regions, facilitating the discovery of peptides from aberrant translation events. Specialized molecular approaches enable enrichment and sequencing of circular RNAs, and immunopeptidomics using mass spectrometry allows for direct characterization of HLA-presented peptides. Together, these technological advances have led to an increasing interest in prioritizing and targeting noncanonical antigens in the next generation of PCVs. This review provides an overview of the diverse origins of TSAs beyond classical neoantigens and discusses emerging approaches that may enable the integration of these antigens in future clinical trials.

circular RNA

Enzymatic depletion of transposable elements in sequencing libraries and its application for genotyping multiplexed CRISPR-edited plants.

Whole-genome sequencing has become a common strategy to genotype individual plants of interest. Although a limited number of genomic regions usually need to be surveyed with this strategy, excess sequencing information is almost always generated at an appreciable financial cost. Repetitive sequences (e.g., transposons), which can account for more than 80% of the genome of some plants, are often not required in these genotyping projects. Therefore, strategies that enrich DNA coding for the protein-coding genes prior to sequencing can lower the cost to obtain sufficient sequence information. Here, we present the development and application of methylation-sensitive reduced representation sequencing (MsRR-Seq), which relies on the cytosine methylation-sensitive restriction enzyme MspJI to deplete constitutive heterochromatic DNA before library construction. By applying MsRR-Seq to citrus and maize, we show that protein-coding genes can be enriched in sequencing datasets. We then describe the application of MsRR-Seq to facilitate the identification of complex mutants from populations of citrus plants resulting from multiplex CRISPR/Cas9 editing of four genes. Overall, this work demonstrates an easy and low-cost method to enrich non-repetitive DNA in high-throughput sequencing libraries, an approach that is especially useful for large plant genomes with an excessively high proportion of methylated repetitive sequences.

DNA Transposable Elements

Genome-wide analysis of Enterococcus faecalis genes that facilitate interspecies competition with Lactobacillus crispatus.

Enterococci are opportunistic pathogens notorious for causing a variety of infections. While both Enterococcus faecalis and Lactobacillus crispatus are commensal residents of the vaginal tract, the molecular mechanisms that enable E. faecalis to take advantage of a vaginal biome with lower counts of lactobacilli to colonize the vaginal tract and induce aerobic vaginitis remain unknown. Here, we show that L. crispatus eradicates E. faecalis in a contact-independent manner. Using transposon sequencing to identify E. faecalis OG1RF transposon (Tn) mutants that are either under-represented or over-represented when co-cultured with L. crispatus, we found that Tn mutants with disruption in the dltABCD operon, that encodes the proteins responsible for the D-alanylation of teichoic acids, and OG1RF_11697 encoding for an uncharacterized hypothetical protein are more susceptible to killing by L. crispatus. Inversely, Tn mutants with disruption in ldh1, which encodes for L-lactate dehydrogenase, are more resistant to L. crispatus killing. Using the Galleria mellonella infection model, we show that co-injection of L. crispatus with E. faecalis OG1RF enhances larvae survival while this L. crispatus-mediated protection was lost in larvae co-infected with either L. crispatus and E. faecalisΔldh1 or Δldh1Δldh2 strains. Last, using RNA sequencing to identify E. faecalis genes that are differently expressed in the presence of L. crispatus, we found major changes in the expression of genes associated with glycerophospholipid metabolism, central metabolism, and general stress responses. The findings in this study provide insights into how E. faecalis mitigate assaults by L. crispatus.IMPORTANCEEnterococcus faecalis is an opportunistic pathogen notorious for causing a multitude of infections. As vaginal commensals, E. faecalis must interact with Lactobacillus crispatus, but how E. faecalis overcomes or mitigate assaults by L. crispatus killing remains unknown. We show that L. crispatus eradicates E. faecalis temporally in a contact-independent manner. Using high-throughput molecular approaches, we identified genetic determinants that enable E. faecalis to compete with L. crispatus. This study represents an important first step for the identification of adaptive genetic traits required for enterococci to tolerate assaults by lactobacilli.

Enterococcus faecalis

PAT: An Image Analysis Tool for Automated Scoring of Pollen in Alexander-Stained Anthers.

Quantitative pollen viability analysis is a critical but labor-intensive step in plant reproductive biology. Existing deep-learning Segment Anything Models (SAM) fail to reliably segment viable pollen in Alexander-stained anthers. To address this, we fine-tuned an existing Cellpose-SAM model for pollen segmentation. We integrated it into PAT (Pollen Analysis Tool), a cross-platform desktop application. PAT features instance segmentation with interactive quality control, an in-app model retraining module, and publication-ready statistical outputs. We deployed PAT in an EMS suppressor screen of semi-sterile Arabidopsis smg7-6 mutants, enabling efficient candidate prioritization for whole-genome sequencing and mapping of the candidate mutation. This screen led to the identification of a point mutation in CAP-D2 (capd2-2), a Condensin I subunit, that rescues the smg7-6 meiotic phenotype. Notably, mutation in a Condensin II subunits (CAP-D3 and CAP-H2) does not confer rescue. Further characterization suggests the capd2-2 allele is hypomorphic, showing no defects in vegetative growth, chromocenter compaction, or transposable element silencing. Collectively, we demonstrate that accessible AI tools have the potential to bridge gaps in plant phenotyping and accelerate the pace of biological discovery.

Alexander staining

Whole-genome characterization and phylogenetic placement of Fusarium oxysporum f. sp. vasinfectum isolates.

Fusarium wilt of cotton, caused by Fusarium oxysporum f. sp. vasinfectum (Fov), remains a persistent threat to cotton production worldwide. Among the known races, Fov race 4 and its extra-virulent variants cause particularly severe losses in Upland cotton. Although several Fov genome assemblies have been assigned to races, the genomic diversity and evolutionary relationships among pathogenic and non-pathogenic isolates associated with cotton outbreaks remain poorly understood at the whole-genome level. This study addressed these gaps by generating and comparing high-quality genome assemblies of four Fusarium isolates collected from Texas cotton fields: two pathogenic (TX17-24 and TX18-9) and two non-pathogenic (TX17-6 and TX18-6). Draft assemblies were generated using Oxford Nanopore long reads and polished with Illumina reads. Comparative genomic analyses showed that pathogenic isolates possessed larger genomes and more conserved orthologous families, whereas non-pathogenic isolates contained more unique genes. Analyses of predicted secreted effectors, transposable elements, and carbohydrate-active enzymes further distinguished pathogenic and non-pathogenic lineages, suggesting roles in virulence adaptation and genome plasticity. Phylogenomic analyses using k-mer-based, assembly- and alignment-free methods incorporated all available long-read Fov genomes and revealed substantial genetic diversity within races 1 and 4, clustering isolates into multiple sublineages. These findings show that Fov race diversification is underestimated when based on traditional classification schemes and may be shaped by host specialization, geographic separation, or horizontal gene transfer. This work advances our understanding of the genomic diversity and evolutionary dynamics of Fov and establishes a foundation for improved race identification and characterization of Fusarium wilt pathogenesis in cotton.

Fusarium oxysporum

LncRAnalyzer: a robust workflow for long non-coding RNA discovery using RNA-Seq.

Long non-coding RNA (lncRNA) is a major transcript category that lacks protein-coding capabilities, with relatively low abundance and complex expression patterns. Distinguishing lncRNAs from protein-coding genes is a complex process involving multiple filtering steps. We developed an automated pipeline named LncRAnalyzer featuring retrained models for 60 species. This workflow aims to reduce the likelihood of obtaining protein-coding or partial protein-coding transcripts during lncRNA identification by utilizing eight distinct approaches. We conducted a 10-fold cross-validation of the sorghum models and training sets with their standard ones and other approaches using real-life RNA-Seq datasets and known lncRNA and CDS sequences of sorghum. The results showed that the sorghum models and training sets were outperformed. The pipeline output comprises upset plots illustrating the number of lncRNA/NPCTs identified by the approaches, commonly identified lncRNA and their classes, NPCTs, and expression count tables. A feature-level comparison and benchmarking analysis of LncRAnalyzer with four existing pipelines, namely, LncPipe, LncEvo, lncRNA-Annotation, and Plant-LncPipe, demonstrated that LncRAnalyzer is more comprehensive, easier to implement, and accurate in lncRNA predictions. This workflow also ascertains lncRNA origins from various Transposable Elements (TEs) in plants using TE annotations from APTEdb [http://apte.cp.utfpr.edu.br/]. LncRAnalyzer is publicly available on GitLab [https://gitlab.com/nikhilshinde0909/LncRAnalyzer.git] for academic users.

RNA, Long Noncoding

Applications of transposon-insertion sequencing for understanding bacterial physiology.

Transposon-insertion sequencing (Tn-seq) couples transposon mutagenesis with next-generation sequencing to identify the transposon insertion site for thousands of mutants in parallel. It is a powerful technology with a myriad of uses beyond the identification of essential genes required for a cell to grow and divide. Tn-seq is particularly useful as a high-throughput method to assign function to function-unknown genes, which have increased steadily with the abundance of newly sequenced bacterial genomes. Tn-seq has now been adapted for use in over 100 bacterial species. Here, we summarize the applications of Tn-seq for querying bacterial physiology and discuss some of the possible applications for the future.

DNA Transposable Elements

Evolutionary relationships among strains of Mycobacterium tuberculosis with few copies of IS6110.

Molecular typing of Mycobacterium tuberculosis by using IS6110 shows low discrimination when there are fewer than five copies of the insertion sequence. Using a collection of such isolates from a study of the epidemiology of tuberculosis in London, we have shown a substantial degree of congruence between IS6110 patterns and both spoligotype and PGRS type. This indicates that the IS6110 types mainly represent distinct families of strains rather than arising through the convergent insertion of IS6110 into favored positions. This is supported by identification of the genomic sites of the insertion of IS6110 in these strains. The combined data enable identification of the putative evolutionary relationships of these strains, comprising three lineages broadly associated with patients born in South Asia (India and Pakistan), Africa, and Europe, respectively. These lineages appear to be quite distinct from M. tuberculosis isolates with multiple copies of IS6110.

Africa

Mobilization of blaVIM genes via the Tn6292 transposon among carbapenem-resistant Enterobacter cloacae complex isolates from colonized patients in a Spanish hospital.

UNLABELLED: The aim of this study was to perform molecular characterization of the carbapenem-resistant Enterobacter cloacae complex (ECC) isolates from colonized patients in a hospital using whole-genome sequencing (WGS) technology. As part of routine surveillance for multidrug-resistant bacterial colonization, 21 ECC isolates were recovered from patients at San Carlos Hospital in Madrid (Spain) between December 2020 and November 2024. WGS was used to determine their genetic relatedness. Furthermore, species identification, sequence type (ST), resistome, plasmid content, and flanking mobile genetic elements (MGEs) of the carbapenemase genes were derived from the WGS data. The most prevalent carbapenemase gene identified was blaVIM-1 (n = 18, 85.7%), with other notable genes including blaKPC-2 (n = 1, 4.8%), blaKPC-3 (n = 1, 4.8%), and blaOXA-48 (n = 1, 4.8%). Several blaACT and blaESBL variants were also found among the carbapenem-resistant ECC isolates. All of them carried at least one blaACT gene, with blaACT-7 (11/21) and blaTEM-type (14/21) genes being the most common AmpC and ESBL-encoding genes, respectively. Additionally, two isolates exhibited the presence of the mcr-9 gene. Overall, E. hormaechei subsp. steigerwaltii (ST93), followed by E. hormaechei subsp. hoffmanii (ST78 and ST50), were the predominant species and STs circulating among the carbapenem-resistant ECC strains. The blaVIM-1 gene was part of class 1 integrons located within a Tn3-family transposon, Tn6292. blaKPC and blaOXA-48 were linked to Tn4401 and Tn1999 transposons, respectively. In conclusion, the presence of the blaVIM within a transposon Tn6292 enhances its mobility across bacterial genomes, underscoring the value of high-throughput sequencing in monitoring the spread of carbapenem-resistant ECC isolates. IMPORTANCE: This study highlights why monitoring the spread of antibiotic-resistant bacteria in hospitals is critical. By analyzing the complete DNA of carbapenem-resistant bacteria, antibiotics were considered a last line of treatment. We found that the resistance genes are not isolated. Instead, they are embedded within mobile elements called transposons. This means that they can "jump" between different bacteria, accelerating the spread of resistance. These findings emphasize the importance of high-resolution genomic technologies to track and control the spread of these dangerous bacteria in clinical settings, helping preserve the effectiveness of life-saving treatments.

Humans

Identification of sporulation genes in Bacillus anthracis highlights similarities and significant differences with Bacillus subtilis.

The molecular basis of endospore formation in the model gram-positive bacterium Bacillus subtilis has been investigated for over half a century. Here, using high throughput and classical genetic approaches, we performed a comparative analysis of sporulation in the human pathogen Bacillus anthracis. A transposon-sequencing screen identified >150 genes required for B. anthracis sporulation. As anticipated, many of the genes that are critical for sporulation in B. subtilis were also required for B. anthracis sporulation. However, we identified >50 genes that are important for sporulation in B. anthracis but not in B. subtilis, and 22 B. anthracis sporulation genes that are absent from the B. subtilis genome. To validate the hits from our screen, we generated an ordered transposon-mutant library using Knockout Sudoku. Cytological analysis of a subset of the canonical sporulation-defective mutants revealed similar but not identical phenotypes in the pathogen compared to the model. We investigated several of the newly identified sporulation genes, with an in-depth analysis of one, ORF 04167, renamed ipdA. Sporulating cells lacking ipdA are blocked in the morphological process of engulfment, generating septal bulges. An AlphaFold-Multimer screen and a classical genetic enrichment revealed that IpdA is a secreted inhibitor of the polysaccharide deacetylase PdaN. Our data support a model in which induction of IpdA at the onset of sporulation inhibits deacetylation of the cell wall peptidoglycan (PG), enabling the sporulation-specific PG hydrolases to catalyze engulfment. Altogether, our studies reveal that B. subtilis is an excellent model for endospore formation in B. anthracis, while underscoring the importance of direct analysis in B. anthracis. The suite of tools that we have generated will catalyze the molecular dissection of sporulation and other cell biological processes in this important human pathogen.

Bacillus anthracis

Rapid screening and identification of genes involved in bacterial extracellular membrane vesicle production using a curvature-sensing peptide.

Bacteria secrete extracellular membrane vesicles (EMVs). Physiological functions and biotechnological applications of these lipid nanoparticles have been attracting significant attention. However, the details of the molecular basis of EMV biogenesis have not yet been fully elucidated. In our previous work, an N-terminus-substituted FAAV peptide labeled with nitrobenzoxadiazole (NBD; nFAAV5-NBD) was developed. This peptide can sense the curvature of a lipid bilayer and selectively bind to EMVs even in the presence of cells. Here, we applied nFAAV5-NBD to a genome-wide screening of hyper- and hypo-vesiculation transposon mutants of a Gram-negative bacterium, Shewanella vesiculosa HM13, to identify the genes involved in EMV production. We analyzed the transposon insertion sites in hyper- and hypo-vesiculation mutants and identified 16 and six genes, respectively, with a transposon inserted within or near them. Targeted gene-disrupted mutants of the identified genes showed that the lack of putative dipeptidyl carboxypeptidase, glutamate synthase β-subunit, LapG protease, metallohydrolase, RNA polymerase sigma-54 factor, inactive transglutaminase, PepSY domain-containing protein, and Rhs-family protein caused EMV overproduction. On the other hand, disruption of the genes encoding putative phosphoenolpyruvate synthase, d-hexose-6-phosphate epimerase, NAD-specific glutamate dehydrogenase, and sensory box histidine kinase/response regulator decreased EMV production. This study demonstrates the utility of a novel screening method using a curvature-sensing peptide for mutants with altered EMV productivity and provides information on the genes related to EMV production.IMPORTANCEConventional methods for isolation and quantification of extracellular membrane vesicles (EMVs) are generally time-consuming. nFAAV5-NBD can detect EMVs in the culture without separating EMVs from cells. In situ detection of EMVs using this peptide facilitated screening of the genes related to EMV production. We succeeded in identifying various genes associated with EMV production of Shewanella vesiculosa HM13, which would contribute to the elucidation of bacterial EMV formation mechanisms. Additionally, the hyper-vesiculating mutants obtained in this study would be valuable for EMV applications, such as secreting useful substances as EMV cargoes and producing artificially functionalized EMVs.

Shewanella