PubMed HealthSearch

SEARCH · PubMed Health

Results for “Structured Barcoding”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

15 recordsLinked to original sources

ClarID: A Human-Readable and Compact Identifier Specification for Biomedical Metadata Integration.

BACKGROUND: In biomedical research, subjects and biospecimens are commonly tracked using simple IDs or UUIDs, which guarantee uniqueness but convey no embedded semantic information. Contextual metadata (such as tissue type, diagnosis, or assay) is often stored separately, making integration, cohort selection, and downstream analysis cumbersome. While structured barcoding systems exist in large consortia (e.g., TCGA, GTEx) or domain-specific contexts (e.g., SPREC, GOLD), no unified, extensible framework currently spans both subjects and biosamples in a human- and machine-readable way. METHODS: We developed ClarID, a domain-agnostic specification that supports two identifier formats: (i) a human-readable form (e.g., 'CNAG_Test-HomSap-00001-LIV-TUM-RNA-C22.0-TRT-P1W' that encodes key metadata such as project, species, subject_id, tissue, assay, disease, timepoint and duration (from that event); and (ii) a compact version named 'stub' (e.g., 'CT01001LTR0N401T1W') optimized for filenames, pipelines, and labeling.ClarID is implemented through an open-source command-line tool, ClarID-Tools, which processes tabular metadata files (CSV/TSV) and uses a YAML-based codebook to generate, decode, and validate identifiers, as well as to create and read QR codes. The tool supports bulk and single-sample processing and allows easy integration with institutional workflows. RESULTS: To demonstrate ClarID's utility, we applied it to datasets from the Genomic Data Commons (GDC), generating interpretable identifiers for more than 113,000 clinical records (subjects) and 4,255 biospecimen records. All materials, including pre-processing scripts, input and encoded data, are publicly available and fully reproducible via the accompanying GitHub repository and Google Colab. CONCLUSIONS: ClarID fills a critical gap between opaque accession numbers and rich metadata schemas by embedding key context directly into structured identifiers. It enhances traceability, facilitates downstream analysis, and remains adaptable to project-specific needs through a configurable codebook. The accompanying ClarID-Tools software is freely available, together with full documentation and reproducible pipelines, at https://github.com/CNAG-Biomedical-Informatics/clarid-tools.

Biosample identifiers

Directed evolution of engineered virus-like particles with improved production and transduction efficiencies.

Engineered virus-like particles (eVLPs) are promising vehicles for transient delivery of proteins and RNAs, including gene editing agents. We report a system for the laboratory evolution of eVLPs that enables the discovery of eVLP variants with improved properties. The system uses barcoded guide RNAs loaded within DNA-free eVLP-packaged cargos to uniquely label each eVLP variant in a library, enabling the identification of desired variants following selections for desired properties. We applied this system to mutate and select eVLP capsids with improved eVLP production properties or transduction efficiencies in human cells. By combining beneficial capsid mutations, we developed fifth-generation (v5) eVLPs, which exhibit a 2-4-fold increase in cultured mammalian cell delivery potency compared to previous-best v4 eVLPs. Analyses of v5 eVLPs suggest that these capsid mutations optimize packaging and delivery of desired ribonucleoprotein cargos rather than native viral genomes and substantially alter eVLP capsid structure. These findings suggest the potential of barcoded eVLP evolution to support the development of improved eVLPs.

Humans

A systematic capsid evolution approach performed in vivo for the design of AAV vectors with tailored properties and tropism.

Adeno-associated virus (AAV) capsid modification enables the generation of recombinant vectors with tailored properties and tropism. Most approaches to date depend on random screening, enrichment, and serendipity. The approach explored here, called BRAVE (barcoded rational AAV vector evolution), enables efficient selection of engineered capsid structures on a large scale using only a single screening round in vivo. The approach stands in contrast to previous methods that require multiple generations of enrichment. With the BRAVE approach, each virus particle displays a peptide, derived from a protein, of known function on the AAV capsid surface, and a unique molecular barcode in the packaged genome. The sequencing of RNA-expressed barcodes from a single-generation in vivo screen allows the mapping of putative binding sequences from hundreds of proteins simultaneously. Using the BRAVE approach and hidden Markov model-based clustering, we present 25 synthetic capsid variants with refined properties, such as retrograde axonal transport in specific subtypes of neurons, as shown for both rodent and human dopaminergic neurons.

barcoding

Is There a Fly in My Soup? To What Extent Do Metabarcoding and Individual Barcoding Tell the Same Story?

Metabarcoding has become the method of choice for characterizing complex arthropod communities. The extent to which metabarcoded bulk samples will recover the same community composition as individual sequencing of all individuals in the sample remains poorly quantified. Biases such as unequal extraction of DNA from different taxa, primer mismatches and non-random PCR may cause the selective drop-out of species from metabarcoding data. At the same time, DNA metabarcoding may reveal arthropod taxa present not as individuals, but as DNA residues on the surface or in the gut of insects. To quantify the consistency in sample contents established by different means, we metabarcoded 45 bulk insect samples, then extracted all arthropods and sequenced them individually. Metabarcoding targeted 418 bp at the 3' end of the Folmer barcoding region, while individual barcodes captured the entire 658 bp Folmer region. The metabarcoding workflow, including PCR amplification, sequencing and bioinformatics, was performed in three replicates from three separate lysate aliquots per sample. For the main analyses, sequences were assigned to Barcode Index Numbers (BINs) as identical taxonomic categories across data types, thereby allowing the detection of even rare but biologically true taxa. Since such reference-based validation will be unavailable to any researcher dealing with metabarcoding data alone, we validated our key findings through an alternative workflow, i.e., de novo clustering of sequences. We found that metabarcoding is replicable, as different replicates of the same sample recover similar species richness and composition. Individual barcoding and metabarcoding provide similar impressions of relative differences in community structure: species-rich vs. species-poor samples rank similarly among data types (Spearman's ⍴ = 0.88-0.99) as do differences in relative dissimilarity between sample pairs (Spearman's ⍴ = 0.55-0.90). Dissimilarity between data types varies with BIN richness in the sample, but this relationship reflects nestedness rather than turnover: metabarcoding recovers the same set of core species as individual barcoding but adds hundreds of species on top. Any BIN recovered as an individual occurred with high probability in the metabarcoding data, and any BIN found in high read abundances by metabarcoding was likely found as an individual (p > 0.8). In terms of abundances, the number of individual insects per BIN was well predicted by the number of metabarcoding reads (R2 > 0.68 for a model including taxonomy as a random effect). Our analysis suggests that metabarcoding data will be informative of the sample contents in terms of arthropod species richness, composition and taxon-specific abundances. Taxa recovered in low copy numbers in metabarcoding sequence data will likely represent DNA left as residues from past biotic interactions. Barring sequencing errors, both types of data yield biologically relevant insights into the taxa present in the source community.

Animals

Assessment of Genetic Diversity and Population Structure on Azadirachta indica A. Juss. in an Urban Metropolitan: Ahmedabad, India.

Azadirachta indica (A. indica) A. Juss., commonly known as Neem, is a valuable multipurpose tree with profound medicinal properties and socioeconomic importance, widely recognized since ancient Ayurvedic times. Despite its prominence, knowledge about its genetic diversity within the metropolitan area of Ahmedabad is limited. This study marks the first in-depth exploration of the genetic diversity and population structure of A. indica in Ahmedabad. The authenticity of the species was validated through DNA barcoding, and a Geographical Information System (GIS) was used to collect the samples. A total of 35 A. indica accessions were analyzed using five Inter Simple Sequence Repeat (ISSR) primers. Genetic diversity and population structure were evaluated using Inter Simple Sequence Repeat (ISSR) markers through polymorphism assessment, clustering, ordination, and Bayesian population structure analyses. ISSRs revealed a high level of polymorphism (75.66%), indicating substantial genetic variability among accessions. An analysis of genetic diversity indices revealed low to moderate diversity (Hs = 0.14, Ht = 0.217, I = 0.217). Analysis of Molecular Variance (AMOVA) analysis depicted 81% variation within the population and 19% among the population. Low to moderate genetic differentiation (Gst = 0.319) and moderate gene flow (Nm = 1.06) indicated that urban development has not hindered gene flow among populations. Mantel's test revealed a weak but significant correlation between genetic and geographic distances, suggesting limited isolation by distance. The estimated ΔK using STRUCTURE exhibited two subpopulations, representing two gene pools for A. indica accessions (K = 2). Collectively, these patterns indicate that urbanization has not severely disrupted genetic connectivity in A. indica, reflecting its resilience and adaptive potential in a metropolitan environment. These findings provide pivotal knowledge for further understanding the genetic diversity and population structure of A. indica in one of the fastest-growing cities in India, which can be utilized for new breeding programmes, sustainable development and future conservation strategies around the globe.

India

A genetic atlas for the butterflies of continental Canada and United States.

Multi-locus genetic data for phylogeographic studies is generally limited in geographic and taxonomic scope as most studies only examine a few related species. The strong adoption of DNA barcoding has generated large datasets of mtDNA COI sequences. This work examines the butterfly fauna of Canada and United States based on 13,236 COI barcode records derived from 619 species. It compiles i) geographic maps depicting the spatial distribution of haplotypes, ii) haplotype networks (minimum spanning trees), and iii) standard indices of genetic diversity such as nucleotide diversity (π), haplotype richness (H), and a measure of spatial genetic structure (GST). High intraspecific genetic diversity and marked spatial structure were observed in the northwestern and southern North America, as well as in proximity to mountain chains. While species generally displayed concordance between genetic diversity and spatial structure, some revealed incongruence between these two metrics. Interestingly, most species falling in this category shared their barcode sequences with one at least other species. Aside from revealing large-scale phylogeographic patterns and shedding light on the processes underlying these patterns, this work also exposed cases of potential synonymy and hybridization.

Animals

Comparative chloroplast genomics of six Bupleurum (Apiaceae) accessions: candidate barcodes, phylogeny based on available plastomes, and candidate RNA-editing sites.

INTRODUCTION: Bupleurum L. (Apiaceae), a taxonomically intricate genus of about 190 species and a source of Radix Bupleuri (Chai Hu), is difficult to discriminate because of convergent morphology, infraspecific variation, and limited genomic sampling. This study aimed to characterize plastome variation, identify and validate candidate molecular markers, reconstruct plastid phylogenetic relationships, and assess candidate plastid RNA-editing sites in Bupleurum. METHODS: We assembled six plastomes from subgenus Bupleurum, screened 51 Bupleurum plastomes for diagnostic loci, reconstructed whole-plastome and partitioned protein-coding-sequence phylogenies, and predicted plastid C-to-U RNA-editing candidates across the six newly assembled plastomes using a PREP-Cp-compatible workflow. Candidate barcode performance was evaluated against the reference plastome phylogenies, and codon-based models were used to test for positive selection. RESULTS: The plastomes were 154,496-155,778 bp with the canonical quadripartite structure and GC contents of 37.67-37.73%. Gene content was stable (131-132 genes; 86-87 protein-coding genes); B. falcatum subsp. cernuum lacked ycf15 but contained an additional inverted-repeat-associated ycf1 annotation. A/U-ending synonymous codons were favoured. Finite pairwise Ka/Ks estimates were below 1 for most genes, and site-specific codon models detected no positive selection. Each plastome contained 55-61 pure microsatellites, dominated by A/T mononucleotide motifs. MarkerSeek ranked 265 features and identified atpF-atpH, petA-psbJ, rpl32-trnL-UAG, and ycf1 as leading candidate barcodes. ycf1 recovered 38 of 41 nodes strongly supported by both reference trees, whereas a partitioned four-locus analysis recovered 40 of 41 and distinguished all 51 accession sequences. However, only one of seven multi-accession operational binomial groups was monophyletic, and only one showed a positive local barcode gap. The whole-plastome phylogeny recovered Bupleurum as monophyletic relative to Chamaesium. The two sampled Penninervia accessions occupied early-diverging positions without forming an exclusive clade. B. falcatum subsp. cernuum was sister to B. ranunculoides, with B. ranunculoides subsp. telonense sister to that pair. A partitioned 74-CDS analysis recovered the same key relationships and 45 of 50 internal bipartitions. Across the six newly assembled plastomes, 57-63 nonsynonymous C-to-U candidates were predicted per accession (367 total) in 21-22 genes; 269 affected the second codon position and 98 the first. DISCUSSION: Bupleurum plastomes are structurally conservative but retain localised divergence useful for marker development. Concordant whole-plastome and CDS genealogies support genus monophyly, whereas sparse Penninervia sampling and maternal plastid inheritance preclude rejecting traditional subgeneric classification. The predicted RNA-editing sites represent candidates for future experimental validation rather than an established Bupleurum editome. These genomic resources support authentication, conservation, and evolutionary research in Bupleurum.

Apiaceae

Novel antibodies for identification, selection, and manipulation of T cells expressing Whitlow linker-containing CARs.

BACKGROUND: The translational study of chimeric antigen receptor (CAR) T-cell function, persistence, immunophenotype, and spatial localization after infusion is crucial for understanding factors that influence clinical outcomes. However, research has been limited by a lack of optimized tools to reliably detect CAR-engineered cells. To address this, we developed a novel platform to generate monoclonal antibodies (mAbs) targeting a linker peptide incorporated in single-chain variable fragments (scFvs) of most CAR constructs. METHODS: Using recombinant proteins and scFv linker peptides as immunogens, we generated murine mAbs against the Whitlow linker peptide, capable of binding cells expressing Whitlow linker-containing CARs in both fresh and formalin-fixed paraffin-embedded (FFPE) tissues. We evaluated these antibodies in multiple in vitro translational applications relevant to CAR T-cell research and manufacturing. RESULTS: We identified five unique mAbs reactive against the Whitlow linker and characterized their binding properties and three-dimensional structural conformation. One clone was evaluated in depth, demonstrating comparable capacity to identify CAR T cells in peripheral blood relative to other methods using anti-idiotype antibodies or recombinant CAR-target proteins. In contrast to these reagents, the anti-Whitlow mAb detects cells expressing Whitlow linker-containing CARs with different antigen specificities, including those harboring the widely employed anti-CD19 FMC63-derived scFv as well as other scFvs, such as those targeting B-cell maturation antigen (BCMA) or CD33. Importantly, the anti-Whitlow mAb identified CAR T cells in situ in archival FFPE tissues, and a DNA-barcoded format enabled their spatial characterization and immunophenotyping in highly multiplexed immunohistochemistry. We also assessed the functional consequences of antibody binding on CAR T cells in vitro and demonstrated the feasibility of anti-Whitlow mAb-mediated selective enrichment of CAR-expressing T cells for potential utility in manufacturing workflows. CONCLUSIONS: Anti-Whitlow mAb clones exhibited distinct structural and functional properties that can be leveraged for multiple applications, providing versatile tools for detection, selection and manipulation of a broad range of clinical and preclinical CAR T-cell products.

Humans

Comprehensive plastome variation and RNA editing in Mentha: insights into phylogenetic relationships and candidate DNA barcodes.

INTRODUCTION: Mentha is an economically and medicinally important genus in Lamiaceae, but its taxonomy and species delimitation remain challenging because of frequent hybridization, polyploidy, and marked morphological plasticity. METHODS: In this study, we comparatively analyzed 12 plastomes representing major Mentha species, hybrid taxa, and unresolved accessions, including four newly assembled genomes, to characterize plastome structure, repeat composition, sequence divergence, phylogenetic relationships, and plastid RNA editing. The M. arvensis plastome and RNA-seq datasets originated from independent Swiss and Indian accessions, respectively. RESULTS: The plastomes were highly conserved in overall organization, ranging from 151,824 to 152,154 bp and displaying the typical quadripartite structure. Gene content and order were largely stable across taxa, with only minor variation likely associated with annotation differences at IR/SC boundary regions. Codon usage analysis revealed a clear bias toward A/U-ending synonymous codons, and most shared protein-coding genes showed low Ka/Ks ratios, indicating predominant purifying selection. Repeat analyses showed that simple sequence repeats were mainly composed of A/T-rich mononucleotide motifs, whereas long repeats were concentrated in the 30-40 bp size class. Comparative analyses identified six hypervariable regions, namely ccsA-ndhD, ycf1, ndhD, rpl32-trnL-UAG, rbcL-accD, and petA-psbJ, which represent promising candidate plastid markers for species discrimination. Phylogenetic analysis based on complete plastomes provided strong support for relationships among the sampled taxa and recovered a close affinity among M. aquatica, M. arvensis, and M. canadensis. In addition, RNA-seq analysis of M. arvensis identified 17 candidate plastid RNA editing sites, most of which were C-to-U conversions and nonsynonymous events. DISCUSSION: Together, these results expand plastid genomic resources for Mentha and provide a useful framework for phylogenetic inference, species identification, and future germplasm utilization.

RNA editing

Bayesian inference of lineage trees by joint analysis of single-cell multimodal lineage-tracing data with BiLinT.

The advent of single-cell lineage-tracing technologies has enabled the simultaneous profiling of gene expression and lineage barcodes. However, accurate, high-resolution reconstruction of cell lineage trees remains challenging because most existing approaches treat these modalities separately and therefore fail to fully exploit their complementary information. Here we present BiLinT, a Bayesian framework that jointly models multimodal single-cell lineage-tracing data for lineage tree reconstruction. BiLinT integrates barcode evolution (a continuous-time Markov chain) with gene expression dynamics (an Ornstein-Uhlenbeck process) within a unified probabilistic model. Across synthetic and real data sets, BiLinT provides accurate lineage-tree reconstruction and reveals differentiation-associated clonal structure and developmental fate biases.

Journal Article

ARCADIA reveals spatially dependent transcriptional programs through integration of scRNA-seq and spatial proteomics.

MOTIVATION: Cellular states are strongly influenced by spatial context, but single-cell RNA sequencing (scRNA-seq) loses information about local tissue organization, while spatial proteomic assays capture limited marker panels that constrain transcriptomic inference. Integrating these modalities can elucidate how spatial niches shape transcriptional programs, yet existing approaches depend on either feature-level correspondence such as gene-protein linkage or cell-level barcode pairing, which is often unavailable. RESULTS: We present ARCADIA (ARchetype-based Clustering and Alignment with Dual Integrative Autoencoders), a generative framework for cross-modal integration that operates without cell barcode pairing and does not assume direct feature-to-feature correspondence. ARCADIA identifies modality-specific archetypes, that is, convex combinations of cells representing extreme phenotypic states, and aligns these anchors across modalities by minimizing the discrepancy between their cell-type composition profiles. The aligned archetypes define a shared coordinate system that anchors dual variational autoencoders (VAEs) trained with cross-modal geometric regularization, preserving archetype structure and spatial neighborhood information while enabling bidirectional translation between modalities. On semi-synthetic CITE-seq data, ARCADIA outperforms existing weak-linkage methods. Applied to independent human tonsil scRNA-seq and CODEX data, ARCADIA reconstructs known tissue architecture and reveals spatially dependent transcriptional programs linking B-cell maturation and T-cell activation or exhaustion to microenvironmental niches. AVAILABILITY AND IMPLEMENTATION: Source code is accessible at https://github.com/azizilab/ARCADIA_public. Reproducibility scripts and data are available at https://github.com/azizilab/arcadia_reproducibility.

Proteomics

Performance Profiles of Short DNA Barcode Segments for Family Level Detection of Asteraceae Within Asterales.

Short DNA barcodes may facilitate sequence recovery from degraded material, but their ability to retain target-family identity while excluding related taxa varies among genomic regions. We computationally evaluated 16 nuclear, plastid, and mitochondrial marker regions from 11 Asterales families using 279,956 NCBI locus-record matches and an accession-disjoint discovery/test design. Thirty-one candidate segments of 50-200 bp (mean, 98.55 bp) were screened in discovery data and evaluated for within-Asteraceae sequence recall, differentiation from non-Asteraceae Asterales, in silico primer behavior, phylogenetic placement, and exploratory matching across 808 metadata-defined metagenomic samples. Conserved regions such as matR and rbcL showed high within-Asteraceae identity, whereas ITS1, ITS, and trnH-psbA showed larger differences from related-family backgrounds; ITS2 and ycf1 showed intermediate profiles. Candidate segments were placed within or immediately adjacent to Asteraceae reference branches in segment-specific maximum-likelihood analyses, although support and topology varied among regions. Metadata-defined target-containing groups had higher mean query coverage and identity than background groups; because target presence was not independently verified and no classifier was fitted, these comparisons were descriptive and did not estimate diagnostic accuracy. Definitionally linked sequence statistics were interpreted as structural associations rather than evidence of causal evolutionary mechanisms. These results provide a family-level computational comparison of candidate short segments for Asteraceae detection within Asterales. Species identification, operational marker combinations, threshold robustness, and laboratory performance require validation using taxonomically dense, voucher-linked, and experimentally characterized datasets.

Asteraceae

Comparative Analysis of Chloroplast Genomes Reveals Molecular Evolution and Phylogenetic Relationships in Fraxinus (Fraxinus mandshurica).

Fraxinus mandshurica (Manchurian ash) is an ecologically and economically valuable hardwood tree native to Northeast Asia, yet its genomic resources remain limited. We assembled its complete chloroplast (cp) genome (155,559 bp) using hybrid PacBio and Illumina sequencing and performed comparative, phylogenetic, and evolutionary analyses. The cp genome exhibits a typical quadripartite structure encoding 132 gene copies, comprising 114 unique genes (80 protein-coding, 30 tRNA, and 4 rRNA genes), with 18 genes duplicated in the inverted repeat (IR) regions. Simple sequence repeat analysis revealed dominance of mononucleotide A/T repeats. Phylogenetic analysis of 53 complete cp genomes strongly supported the monophyly of Oleaceae and resolved F. mandshurica as sister to the North American F. nigra, consistent with previously proposed Miocene intercontinental dispersal scenarios between East Asia and North America. Most protein-coding genes were under strong purifying selection (Ka/Ks << 1), whereas petB, rpl2, and several ndh genes showed elevated Ka/Ks values that are suggestive of altered selective constraint but are based on very few substitutions and are therefore not, on their own, evidence of positive selection. Nucleotide diversity (Pi) analysis identified 15 hypervariable intergenic spacers (mean Pi = 0.067), among which trnM-CAU-rps14, ndhJ-ndhK, and petL-petG represent promising candidate barcode regions requiring further validation. This study provides a high-quality, fully annotated cp genome of F. mandshurica and a valuable genomic resource for future phylogenetic, population genetic, and conservation studies of this important genus.

Fraxinus

A systematic strategy for identifying causal single nucleotide polymorphisms and their target genes on Juvenile arthritis risk haplotypes.

BACKGROUND: Although genome-wide association studies (GWAS) have identified multiple regions conferring genetic risk for juvenile idiopathic arthritis (JIA), we are still faced with the task of identifying the single nucleotide polymorphisms (SNPs) on the disease haplotypes that exert the biological effects that confer risk. Until we identify the risk-driving variants, identifying the genes influenced by these variants, and therefore translating genetic information to improved clinical care, will remain an insurmountable task. We used a function-based approach for identifying causal variant candidates and the target genes on JIA risk haplotypes. METHODS: We used a massively parallel reporter assay (MPRA) in myeloid K562 cells to query the effects of 5,226 SNPs in non-coding regions on JIA risk haplotypes for their ability to alter gene expression when compared to the common allele. The assay relies on 180&#xa0;bp oligonucleotide reporters ("oligos") in which the allele of interest is flanked by its cognate genomic sequence. Barcodes were added randomly by PCR to each oligo to achieve&#x2009;>&#x2009;20 barcodes per oligo to provide a quantitative read-out of gene expression for each allele. Assays were performed in both unstimulated K562 cells and cells stimulated overnight with interferon gamma (IFNg). As proof of concept, we then used CRISPRi to demonstrate the feasibility of identifying the genes regulated by enhancers harboring expression-altering SNPs. RESULTS: We identified 553 expression-altering SNPs in unstimulated K562 cells and an additional 490 in cells stimulated with IFNg. We further filtered the SNPs to identify those plausibly situated within functional chromatin, using open chromatin and H3K27ac ChIPseq peaks in unstimulated cells and open chromatin plus H3K4me1 in stimulated cells. These procedures yielded 42 unique SNPs (total&#x2009;=&#x2009;84) for each set. Using CRISPRi, we demonstrated that enhancers harboring MPRA-screened variants in the TRAF1 and LNPEP/ERAP2 loci regulated multiple genes, suggesting complex influences of disease-driving variants. CONCLUSION: Using MPRA and CRISPRi, JIA risk haplotypes can be queried to identify plausible candidates for disease-driving variants. Once these candidate variants are identified, target genes can be identified using CRISPRi informed by the 3D chromatin structures that encompass the risk haplotypes.

Humans

Recurrent and niche-specific functional bacteriome of maize hybrid revealed by integrated metabarcoding and culturomics.

The plant microbiome plays a pivotal role in plant survival in natural habitats by facilitating nutrient acquisition, stress adaptation, and disease suppression, while also offering opportunities to enhance crop productivity and climate resilience. However, the distribution of persistent and culturable bacteriome across maize-associated niches and their functional potential remain poorly resolved. This study integrated metagenomic next-generation sequencing (mNGS-based metabarcoding) and culturomics to characterise the maize-associated bacteriome of bulk soil, rhizoplane, phylloplane, and cob of the maize hybrid PHM-1 under contrasting cropping and tillage systems, and to identify recurrent and agriculturally promising bacteriome components. The bacteriome exhibited pronounced niche-specific structuring, whereas overall bacterial community composition did not differ significantly across cropping and tillage treatments (ANOSIM, R&#x2009;=&#x2009;0.038, p&#x2009;=&#x2009;0.306). Proteobacteria predominated in the culturable bacteriome (69-84%; mean, 76.2%) but accounted for only 1% of the total bacteriome, whereas Patescibacteria and Firmicutes were relatively enriched. Niche-specific dominance was evident, with Pantoea accounting for 40.79% of the total and 56.27% of the culturable phylloplane bacteriome under cereal monocropping, while Serratia represented 31.59% and 59.40% of the total and culturable cob bacteriomes, respectively. Across niches, mNGS captured substantially greater bacteriome diversity, particularly uncultured and unidentified taxa in soil-associated compartments, whereas culturomics recovered a narrower but functionally accessible fraction. Culturomics yielded 99 isolates representing 32 species across 12 genera, including six genera shared with the mNGS-derived recurrent bacteriome: Bacillus, Enterobacter, Pantoea, Pseudomonas, Serratia, and Stenotrophomonas. Functional screening identified strong biocontrol and plant-beneficial traits among core-associated isolates. Pseudomonas oryzihabitans ZM-DL-PA10 inhibited Rhizoctonia solani, Macrophomina phaseolina, and Bipolaris maydis by up to 40.6%, 43.9%, and 45.2%, respectively, through secreted and volatile metabolites; exhibited P, K, and Zn solubilisation; and produced IAA and siderophores. It also recorded the lowest B. maydis disease index (ADI) of 1.00. Pantoea ananatis ZM-BH-EA4 showed 52.4% and 68.5% inhibition of R. solani and B. maydis, respectively, through volatile metabolites. Collectively, the integration of mNGS and culturomics revealed a strongly compartmentalised maize bacteriome and identified recurrent, culturable, and functionally promising bacterial taxa, providing a targeted resource for microbiome-based crop protection and climate-resilient maize production.

Zea mays