PubMed HealthSearch

SEARCH · PubMed Health

Results for “Genome, Plant”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Hide and seek: de novo identification in sugar beet reveals impact of non-autonomous LTR retrotransposons.

Plant genomes are filled with retrotransposons and their derivatives, constantly undergoing sequence diversification and structural rearrangement. Among them, short, non-autonomous retrotransposons lack full coding capacity and often form subfamilies. As a result, non-autonomous retrotransposons are incompletely identified in most to all genome assemblies.Here, we capitalize on our comprehensive understanding of the transposable element (TE) landscape in sugar beet (Beta vulgaris) to assess the extent of the blind spot for non-autonomous long terminal repeat (LTR) retrotransposons. This use case serves to answer if all of these sequences are derivatives of easier-to-identify full-length elements or if there is more variability that is currently overlooked.For this we applied a semi-automated structural discovery workflow followed by in-depth manual verification to characterize non-autonomous LTR retrotransposons in sugar beet. We retrieve more than 100 non-autonomous LTR retrotransposon families that lack complete autonomous coding capacity, including canonical terminal-repeat retrotransposons in miniature (TRIMs), elongated non-coding derivatives and families retaining fragmented coding remnants. The identified families span a broad range, including elements exceeding 15,000 bp in length and display evidence for reshuffling and modular evolution. Only a subset of families could be confidently linked to autonomous retrotransposons, showing sequence diversification within the non-autonomous LTR retrotransposon fraction beyond the autonomous genomic templates.We highlight that a large fraction of non-autonomous LTR retrotransposons is incompletely recovered with the current TE identification workflows, even if the output is well-curated and condensed into TE libraries and suggest procedures to remedy this gap. This study gives a genome-wide view into the non-autonomous LTR retrotransposon landscape of a single plant genome and highlights the importance of structure-based approaches for their identification and classification.

LTR retrotransposons

Enhanced exonuclease-Cas9 systems promote multiple nucleotide deletions with higher efficiency and broader targeting scope in plants.

CRISPR-Cas9 is a widely used platform for plant genome editing, but its outcomes are typically dominated by small insertions and deletions (indels). Such limited mutation profiles restrict its utility in functional studies of non-coding RNAs and regulatory elements, such as microRNAs (miRNAs), untranslated regions (UTRs), and promoter sequences, where larger sequence disruptions are often required. Here, we developed enhanced exonuclease-Cas9 platforms, termed multiple nucleotide deletion Cas9 (MND-Cas9) systems, for efficient generation of large deletions in rice. By screening four exonucleases (RecJ, T5, TREX2, and SbcB), we established MND-Cas9v1 systems based on TREX2 or SbcB that produced substantially larger deletions without reducing editing efficiency. Further optimization with an inserted DNA-binding domain (DBD) between Cas9 and exonuclease yielded MND-Cas9v2, which simultaneously enhanced efficiency and deletion size. To expand PAM compatibility, we introduced PAM-relaxed Cas9-NG and SpG variants, generating MND-Cas9-NG/SpGv2 systems with broader targeting scope and superior performance compared to their parental nucleases. Finally, we demonstrated the utility of these systems in two applications: MND-Cas9v2 efficiently knocked out the miRNA gene OsMIR530, producing larger seeds, and generated extended deletions in the 3'UTR of OsGhd2, which upregulated its expression and increased grain size. These results demonstrate that MND-Cas9 systems enable high-efficiency generation of extended deletions and facilitate functional analyses of non-coding RNAs and regulatory sequences. Overall, this work establishes a versatile and expandable exonuclease-Cas9 platform that substantially broadens the mutational spectrum and application potential of CRISPR-Cas9 for plant genome engineering.

CRISPR-Cas Systems

The Rise of Plant Pan-Genomes: From Genome Variation to Predictive Breeding.

Plant pan-genomics is entering a new phase beyond genome variation discovery, requiring a shift from cataloguing genomic diversity toward understanding how variation generates biological function and breeding value. Here, we propose that the future of plant pan-genomics will be shaped by three conceptual transitions. First, structural variation (SV), presence-absence variation (PAV), and haplotype diversity should be interpreted not merely as genomic differences, but as regulatory components that influence gene networks, chromatin organization, and complex traits. Second, the expansion from species-level pan-genomes to genus-level super pan-genomes provides an evolutionary framework for uncovering adaptive genetic modules preserved in wild relatives and overlooked during domestication. Third, integrating pan-genomes with pan-omics, three-dimensional genome analyses, and artificial intelligence will enable the transformation of genomic variation into predictive models for crop improvement. We further propose that the ultimate value of pan-genomes lies not in generating increasingly complete genome collections, but in establishing a mechanistic bridge between genome diversity, biological function, and breeding decisions. This transition will move crop improvement from empirical selection toward rational genome design, where evolutionary diversity can be systematically interpreted, predicted, and engineered.

Journal Article

Plant species identification by genome skimming across the vascular plant tree of life.

Accurate species identification is essential for biodiversity conservation and sustainable use, yet standard plant DNA barcoding often fails to achieve species-level resolution. We present a large-scale empirical evaluation of genome skimming as a tool to improve plant species discrimination. Using standardised data from 1969 individuals representing 475 species from 32 genera across major lineages of the vascular plant tree of life, we compare conventional plastid + internal transcribed spacer (ITS) barcodes with genome skimming approaches. Standard barcoding using rbcL, matK, trnH-psbA and ITS resolved about half of species (49.3%), with six genera showing <&#x2009;25% species discrimination. By contrast, genome skimming enabled the recovery of complete plastid genomes, yielding 57.6% species discrimination. It also generated sufficient nuclear genomic data for additional resolution from k-mer analysis, achieving 66.8% species discrimination - an average gain of 17.5% over standard barcodes - while eliminating cases of extreme failure (<&#x2009;25% resolution). The recovery of complete plastomes and ribosomal DNAs from genome skims also ensures backward compatibility with existing barcode datasets. Our results demonstrate that genome skimming provides data that substantially improves species-level resolution across diverse plant lineages and offers a scalable, high-throughput approach for building comprehensive reference resources to support global biodiversity initiatives.

DNA Barcoding, Taxonomic

Haplotype-resolved telomere-to-telomere genome assembly of Populus lasiocarpa unveils retrotransposon-driven centromere evolution.

Centromeres, essential for chromosome segregation, exhibit remarkable evolutionary dynamism in sequence composition and structural organization. Here, we report the first haplotype-resolved, telomere-to-telomere genome assembly of Populus lasiocarpa (PLAS) and precisely map all 38 functional centromeres through CENH3 ChIP-Seq. Unlike classical satellite-rich centromeres in model plants, PLAS centromeres lack abundant satellite arrays but are dominated by retrotransposons, particularly RLG and RIL elements, which form intricate nested TE arrays within the functional centromeric regions, disrupting their structural integrity and driving their evolution. Comparative analysis with P. trichocarpa reveals a conserved retrotransposon-dominated architecture, despite minimal sequence conservation. We propose a cyclic model of centromere evolution in which autonomous retrotransposons destabilize functional centromeres through epigenetic erosion, triggering neocentromere formation at pericentromeric sites enriched in transposable elements (TEs) and tandem repeats (TRs). These neocentromeres either succumb to recurrent retrotransposon invasions or stabilize through KARMA-mediated TR expansion, ultimately giving rise to satellite-rich centromeres. Our work redefines centromeres as dynamic, epigenetically plastic domains shaped by retrotransposon-TR antagonism, challenging the satellite-centric paradigm and offering novel insights into plant genome evolution.

Retroelements

Efficient CRISPR/Cas-SF01 genome editing tools with high editing efficiency in allotetraploid oilseed rape.

CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats)-Cas9 has been widely utilized for plant genome editing, but the protospacer adjacent motif (PAM) requirement limits its editing scope. CRISPR/Cas12i3 belongs to the type-VI Cas system that has gained extensive attention due to its smaller size and less restricted canonical TTN PAM sequence. In this study, we explored the newly developed Cas-SF01 system (Cas12i3 variant) for genome editing in oilseed rape. We established an efficient protoplast transformation system in oilseed rape to compare editing efficiency between Cas-SF01 and Cas9. Cas-SF01 shows cleavage activities at the tested 5'-TTN-3' PAM sites with editing outcomes sharing considerable similarities with the CRISPR-Cas9 system in protoplast. Cas-SF01 also induces high efficiency mutagenesis for multiple target sites in stable transformed oilseed rape lines, generating mutants with multilocular silique and male sterile phenotypes. Furthermore, Cas-SF01-derived cytosine base editors (CBEs) were developed to produce targeted C-to-T base edits. Compared to SpCas9, Cas-SF01 has an expanded PAM range and effectively recognizes TTN PAMs, which has substantially broadened the scope of editable sites within the rapeseed genome. No mutations were identified at the putative off-target sites among the edited plants. This study developed a robust, first-of-its-kind Cas12 system in the allotetraploid Brassica napus, expanding the scope of editing and enriching genome-editing toolkits for biological research and genetic improvement.

Brassica napus

AEGIS: an annotation extraction and genomic integration resource.

MOTIVATION: Genome annotation files (GFF3/GTF) are the standard for storing genomic feature data, yet their flexibility often results in formatting inconsistencies that create bottlenecks for downstream bioinformatics analyses. A robust, unified framework is required to parse, standardise, and validate these files to ensure interoperability and facilitate complex comparative genomic tasks. RESULTS: We present AEGIS (Annotation Extraction and Genomic Integration Suite), a comprehensive toolkit designed to parse, correct, and standardise genome annotations. Beyond quality control, AEGIS provides advanced modules for flexible feature extraction (e.g., coding sequences, promoters) and comparative genomic analysis. Uniquely, it integrates multiple lines of evidence, including sequence homology, synteny, and coordinate-based lift-overs, to assess gene model correspondence and infer orthology. We demonstrate the utility of AEGIS by quantifying complex structural changes between Arabidopsis annotation versions and identifying high-confidence orthologues across diverse plant genomes. AVAILABILITY: AEGIS is implemented in Python. Source code and documentation are freely available under the GPL-3 license at https://github.com/Tomsbiolab/aegis and as a Docker container at https://hub.docker.com/r/tomsbiolab/aegis. The package is also available on PyPI (pip install aegis-bio).

Software

Recent gene duplication and structural remodeling drive rapid lineage-specific gene family evolution in plants.

Gene duplication promotes the generation of novel gene functions and trait diversity across species. Here, we present DupHIST, a computational pipeline that reconstructs the hierarchical timing of gene duplications by integrating maximum likelihood (ML)-based phylogeny with substitution-derived timing via statistical smoothing. Applied to over 4.5 million genes from 114 plant genomes, we successfully inferred duplication histories across nearly 130,000 orthogroups. This large-scale analysis showed that 53.0% of genes arose from recent, lineage-specific duplications, with high concentrations in particular multi-copy families. Among these, NLR, C48, and P450 families exemplified how recently duplicated genes undergo rapid stepwise structural remodeling. This process was primarily driven by small-scale mutations, including insertions, deletions, and frameshifts, that rapidly accumulated shortly after duplication. By resolving the precise duplication order, we reconstructed these architectural changes, thereby enabling both the inference of putative ancestral structures and the exploration of functional diversification arising from structural remodeling. Structure-based clustering further uncovered that recently duplicated, uncharacterized genes retain core domain structures resembling known functional proteins even across phylogenetically distant species lacking sequence homology. Our findings reveal that recent gene duplications and subsequent structural remodeling represent a widespread and lineage-specific force driving rapid diversification of gene families in plants.

Gene duplication history

Features affecting Cas9-induced editing efficiency and patterns in tomato: evidence from a large CRISPR dataset.

CRISPR/Cas9 is a cornerstone of plant genome editing, yet the determinants of editing efficiency for a given single-guide RNAs (sgRNAs) and DNA double-strand break (DSB) repair outcomes remain poorly understood, particularly in plants. Here, we generated a large experimental dataset comprising 420 sgRNAs targeting promoters, exons, and introns of 137 genes in tomato protoplasts, and quantified editing efficiency and repair footprints together with chromatin accessibility and transcriptional state in the same cellular context. Editing efficiency was consistently higher at targets in accessible chromatin and modestly higher in promoters and introns than in exons, whereas transcriptional activity had no detectable effect. Editing efficiencies were more similar among sgRNAs targeting the same gene than among different genes, revealing a local genomic influence on Cas9 activity. A distinct subset of sgRNAs achieved near-complete editing and produced characteristic repair footprints dominated by long deletions with extended microhomology tracts, indicative of microhomology-mediated end joining (MMEJ), resembling patterns associated with high-efficiency guides in human cells, and suggesting conserved sequence-driven repair biases across species. In contrast, widely used human-trained prediction models failed to accurately rank sgRNA performance in plants, highlighting the limits of cross-species predictability. Together, this dataset provides a resource for improving guide design and mechanistic understanding of plant DNA repair.

Solanum lycopersicum

PlantPan: A comprehensive multi-species plant pan-genome database.

The pan-genome represents the complete genomic diversity of specific species, serving as a valuable resource for studying species evolution, crop domestication, and guiding crop breeding and improvement. While there are several single-species-specific plant pan-genome databases, the availability of multi-species pan-genome databases is limited. Additionally, variations in methods and data types used for plant pan-genome analysis across different databases hinder the comparison and integration of pan-genome information from various projects at multi-species or single-species levels. To tackle this challenge, we introduce PlantPan, a comprehensive database housing the results of pan-genome analysis for 195 genomes from 11 plant species. PlantPan aims to provide extensive information, including gene-centric and sequence-centric pan-genome information, graph-based pan-genome, pan-genome openness profiles, gene functions and its variation characteristics, homologous genes, and gene clusters across different species. Statistically, PlantPan incorporates 9&#x2009;163&#x2009;011 genes, 694&#x2009;191 gene clusters, 526&#x2009;973&#x2009;370 genome variations, and 1&#x2009;616&#x2009;089 non-redundant genome variation groups at the species level, 33&#x2009;455,098 genome synteny, and 177&#x2009;827 non-redundant genome synteny groups at the species level. Regarding functional genes, PlantPan contains 5&#x2009;222&#x2009;720 genes related to transcription factors, 395&#x2009;247 literature-reported resistance genes, 455&#x2009;748 predicted microbial/disease resistance genes, and 1&#x2009;612&#x2009;112 genes related to molecular pathways. In summary, PlantPan is a vital platform for advancing the application of pan-genomes in molecular breeding for crops and evolutionary research for plants.

Genome, Plant

Nuclear DNA of plastid origin (NUPTs), neglected driver of genome variation and evolutionary innovation.

Plant nuclear genomes contain a variable, though typically minor, fraction of DNA sequences of plastid origin known as NUPTs. Unlike the massive transfer of DNA and genes from the proto-organelle genome to the nucleus that occurred during the endosymbiotic event that gave rise to plastids, the formation of NUPTs is an ongoing process that does not imply concomitant DNA loss. Although NUPTs are generally considered to be potentially deleterious insertions that are continuously generated and rapidly eliminated at near-constant turnover rates, accumulating evidence reveals alternative evolutionary trajectories. In this review, we discuss recent findings that highlight the episodic formation of NUPTs, their subsequent proliferation, and their eventual long-term fixation within the nuclear genome. We also explore their non-random spatial association with specific genomic elements. NUPTs show preferential overlap with specific superfamilies of transposable elements, which may facilitate their proliferation and dispersal throughout the nuclear genome. Regarding protein-coding genes, the contribution of NUPTs varies among species. In contrast, NUPTs are found to be consistently enriched among certain classes of non-coding RNA genes, notably rRNA, tRNA, and specific regulatory RNA families, suggesting that they are involved in the evolution of gene regulation and translational machinery. Overall, these findings underscore the unexpected complexity of the mechanisms underlying NUPT formation and support the idea that they are a significant source of genome variation and evolutionary innovation. Further research is necessary to fully elucidate the mechanisms underlying NUPT formation, as well as to determine their potential adaptive significance in plant genome evolution.

Plastids

Simultaneous Visualization of Protein and Genomic Regions in Plant Nuclei.

Immunohistostaining (IHS) is a widely used technique in diagnostic and research laboratories in which specific antibodies are used to detect and visualize a protein of interest in cells or tissues. Similarly, with specific oligonucleotide probes, the fluorescence in situ hybridization (FISH) method allows one to visualize genomic regions and to analyze its localization in the nuclear space. Here, we describe a combined FISH-IHS technique that enables researchers to determine the localization of protein and genomic loci in plant nuclei simultaneously. This method can be applied to extracted nuclei and sections of paraffin-embedded tissues. It provides a valuable tool to improve our understanding of nuclear dynamics by revealing the spatial relationship between specific genomic loci and target proteins.

In Situ Hybridization, Fluorescence

Plant cis-regulatory grammar: Decoding the multidimensional code of transcriptional regulation for programmable crop engineering.

Cis-regulatory elements (CREs) orchestrate the spatiotemporal precision of gene expression that underlies plant development, adaptation, and domestication. Decoding the cis-regulatory grammar of plant genomes remains a central challenge in modern biology, with profound implications for programmable crop engineering. Here, recent conceptual and technological advances are synthesized to reshape our understanding of plant CREs. This review first argues that CRE function is not only an intrinsic property of DNA sequence alone but also emerges from a multidimensional context, including chromatin accessibility, histone modifications, three-dimensional genome topology, and cell type-specific regulatory landscapes. Furthermore, the convergence of single-cell epigenomics, high-throughput functional assays, and CRISPR-based dissection has begun to unravel this contextual grammar, revealing the computational principles governing transcriptional regulation. Critically, we propose that artificial intelligence (AI) platforms are catalyzing an ongoing transition from descriptive discovery to predictive engineering, wherein these platforms outperform natural evolution in designing synthetic CREs. Finally, a roadmap is outlined toward a plant regulatory grammar foundation model, which will enable truly predictive engineering of gene expression when fine-tuned for specific tasks. Collectively, the integration of single-cell resolution maps, precise genome editing, AI-driven design, and regulatory-compliant delivery systems promises to transform our ability to reprogram plant gene regulation for next-generation agriculture, bridging the gap between foundational regulatory biology and tangible crop improvement.

artificial intelligence

Rice transcription factor bHLH25 confers resistance to multiple diseases by sensing H2O2.

Hydrogen peroxide (H2O2) is a ubiquitous signal regulating many biological processes, including innate immunity, in all eukaryotes. However, it remains largely unknown that how transcription factors directly sense H2O2 in eukaryotes. Here, we report that rice basic/helix-loop-helix transcription factor bHLH25 directly senses H2O2 to confer resistance to multiple diseases caused by fungi or bacteria. Upon pathogen attack, rice plants increase the production of H2O2, which directly oxidizes bHLH25 at methionine 256 in the nucleus. Oxidized bHLH25 represses miR397b expression to activate lignin biosynthesis for plant cell wall reinforcement, preventing pathogens from penetrating plant cells. Lignin biosynthesis consumes H2O2 causing accumulation of non-oxidized bHLH25. Non-oxidized bHLH25 switches to promote the expression of Copalyl Diphosphate Synthase 2 (CPS2), which increases phytoalexin biosynthesis to inhibit expansion of pathogens that escape into plants. This oxidization/non-oxidation status change of bHLH25 allows plants to maintain H2O2, lignin and phytoalexin at optimized levels to effectively fight against pathogens and prevents these three molecules from over-accumulation that harms plants. Thus, our discovery reveals a novel mechanism by which a single protein promotes two independent defense pathways against pathogens. Importantly, the bHLH25 orthologues from available plant genomes all contain a conserved M256-like methionine suggesting the broad existence of this mechanism in the plant kingdom. Moreover, this Met-oxidation mechanism may also be employed by other eukaryotic transcription factors to sense H2O2 to change functions.

Hydrogen Peroxide

Haplotype-aware long-read error correction.

Error correction of long reads is an important initial step in genome assembly workflows. For organisms with ploidy greater than one, it is important to preserve haplotype-specific variation during read correction. This challenge has driven the development of several haplotype-aware correction methods. However, existing methods are based on either ad-hoc heuristics or deep learning approaches. In this paper, we introduce a rigorous formulation for this problem. Our approach builds on the minimum error correction framework used in reference-based haplotype phasing. We prove that the proposed formulation for error correction of reads in de novo context, i.e., without using a reference genome, is NP-hard. To make our exact algorithm scale to large datasets, we introduce practical heuristics. Experiments using PacBio HiFi sequencing datasets from human and plant genomes show that our approach achieves accuracy comparable to state-of-the-art methods. Implementation: https://github.com/at-cg/HALE .

Clustering

Unraveling evolutionary relationships in the Sida generic alliance (Malvaceae, Malvoideae): a phylogenetic and cytotaxonomic overview.

Sida (Malvaceae), the largest Malveae-Abutilinae member, has poorly defined morphological limits which overlaps with 11 phylogenetically closely related genera that comprises the "Sida generic alliance". The 12 genera are distributed in the tropics especially in Brazil where one third of its species diversity is found. Evolutionary relationships within Sida generic alliance remain unresolved due to morphological convergence, limited taxon sampling, and lack of integrative approaches including cytogenetic data. We reconstructed the phylogeny of Sida and allied genera using a multilocus dataset (nuclear ITS and seven plastid loci) including 193 species classified in 19 genera and analyzed chromosome evolution using cytogenetic data (chromosome number) for 79 species of the 19 genera. The phylogeny recovered seven clades-Abutilon, Bakeridesia, Callianthe, Gaya, and three Sida clades (I-III)-and confirmed the polyphyly of Sida, the largest genera. We detected reticulate evolution, with incongruence between nuclear and plastid topologies. Chromosome number ranged from 2n&#xa0;=&#xa0;12 to 60 and represented synapomorphies for most clades. Ancestral character reconstruction indicated that ascending dysploidy and polyploidy predominated in karyotype evolution of Sida and allied genera. Our results reveal taxonomic incongruence in current classifications probably related to reticulate evolution. A generic-level taxonomic revision is necessary and should rely on integrated phylogenetic and karyotypic evidence. This study provides a framework for phylogenetic systematics and emphasizes the role of Brazil as a hotspot for plant genomic research.

Phylogeny

Published Database Resources for Traditional, Complementary, and Integrative Medicine: Update of a Systematic Review.

BACKGROUND: Traditional, Complementary, and Integrative Medicine (TCIM) has been established in the academic context of universities. In recent years, strategies have been developed worldwide to strengthen the role of TCIM in supporting the health of the population. Online databases are a common way for obtaining evidence-based information. This article is an update of a former systematic review from 2010 on published databases resources for TCIM. METHODS: The databases CINAHL, CAMbase, Web of Science, MEDLINE/PubMed, and Google Scholar search engine were searched for databases related to TCIM published in peer-reviewed journals between 2010 and November 2024. All included databases were visited online, and information on the origin, content, and scope of the database was extracted. RESULTS: A total of 6579 articles were identified through the literature search. After exclusion of irrelevant articles, full-text screening of 127 articles yielded 37 new databases. Together with 16 still available old databases, these mainly contained information on herbal therapies (n = 15) and Traditional Chinese Medicine (n = 11) from 18 different countries. Newly identified medicinal plant databases offer various scientific resources such as crude drugs, indigenous plants, and structures for natural and phytochemical components with molecular biological content. CONCLUSIONS: This literature review illustrates the dynamic development in the database landscape over the last 15 years. While the number of bibliographic databases is shrinking, databases in the field of medical plants/herbal therapy content are on the rise, which might be due to advances in plant genomics and molecular biology.

Humans