PubMed HealthSearch

SEARCH · PubMed Health

Results for “Scale Reliant Inference”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

3 recordsLinked to original sources

Scale reliant mixed effects models enhance microbiome data analysis.

Linear models, including those used for differential abundance analyses, are frequently used in microbiome research to assess how experimental conditions (e.g., disease state or age) affect microbial abundance. Linear mixed-effects models (MEMs) extend linear models to accommodate complex designs, such as longitudinal sampling or hierarchical study structures. However, when applied to microbiome data, existing MEM approaches suffer from high false positive and false negative rates because sequence counts are compositional - they reflect relative rather than absolute abundances. Current methods attempt to overcome this limitation through normalization, but these approaches rely on strong, often unrealistic assumptions about the unmeasured biological scale (e.g., total microbial load). Here we introduce scale-reliant mixed-effects models (SR-MEM), which extend our earlier scale-reliant inference framework by explicitly modeling uncertainty in the unmeasured scale via user-defined probability distributions. By treating scale as a latent variable rather than fixing it through normalization, SR-MEM enables robust inference for complex experimental designs. SR-MEM can incorporate external scale measurements (e.g., flow cytometry, qPCR) or leverage scale information from independent studies to further improve inference. Across simulations and multiple real-world case studies, SR-MEM consistently controls the false discovery rate while maintaining comparable or higher power than standard approaches relying on normalization or bias correction. In reanalyses of published datasets, SR-MEM yields results that are more reproducible across studies and more consistent with known biological and pharmacological effects. SR-MEM provides a principled and practical framework for mixed-effects modeling of microbiome sequence count data in the presence of unmeasured biological scale. By avoiding normalization-based assumptions and instead propagating scale uncertainty through inference, SR-MEM improves error control and reproducibility in longitudinal and hierarchical studies. An accessible implementation is provided in the ALDEx3 R package.

Microbiota

Uncertainty Modeling Outperforms Machine Learning for Microbiome Data Analysis.

Microbiome sequencing measures relative rather than absolute abundances, providing no direct information about total microbial load. Normalization methods attempt to compensate, but rely on strong, often untestable assumptions that can bias inference. Experimental measurements of load (e.g., qPCR, flow cytometry) offer a solution, but remain costly and uncommon. A recent high-profile study proposed that machine learning could bypass this limitation by predicting microbial load from sequencing data alone. To evaluate this claim, we assembled mutt, the largest public database of paired sequencing and load measurements, spanning 35 studies and over 15,000 samples. Using mutt, we show that published machine learning models fail to generalize: on average they perform worse than a naive baseline that always predicted the training set mean. These failures stem from covariate shift-limited shared taxa between studies, differences in community composition, and differences in preprocessing pipelines-that silently derail model inputs. In contrast, Bayesian partially identified models do not attempt to impute microbial load, but instead propagate scale uncertainty through downstream analyses. Across 30 benchmark datasets, Bayesian partially identified models consistently outperformed normalization and machine learning approaches, providing a principled and reproducible foundation for microbiome inference.

16S rRNA-seq

GenBank mining reveals novel insights into Rhizobium phylogeny: Identical 16S rRNA sequences are mainly uncoupled from species designation, host plant, and geographic origin: How this search suggested the definition of a direct 'microbial h-index'.

16S rDNA is the historical gold standard for bacterial identification, particularly in metabarcoding approaches reliant on sequence similarity thresholds. We analyzed 6,660 Rhizobium 16S rRNA gene sequences from GenBank to examine the relationship between sequence identity and three metadata: species name, host plant, and geographic origin. Using an iterative BLAST-based pipeline, we detected 116,069 pairwise matches and assessed concordance among sequences (average length 1,328 bp) sharing 100% identity. For those in which the organism name, host plant and country of isolation were present in the record, surprisingly, 66.59% of identical sequence pairs showed full discordance across all three metadata, while only 1.40% shared the same name, host, and country. The most widespread sequence, detected 371 times, was associated with over 56 different host plants across 25 countries and bore multiple species name designations. These results highlight a striking mismatch between the 16S barcode and the taxonomic, ecological, and phenotypic variability it is assumed to reflect, likely arising from the slow evolution of rRNA genes contrasted with the mobility of ecologically relevant genes via horizontal transfer on plasmids, transposons, and phages. Our findings further challenge the limitations of relying on 16S rRNA alone for fine-scale taxonomic and metadata-based inference in capturing the true functional and ecological diversity of bacteria, endorsing the critical importance of polyphasic taxonomic approaches that integrate genomic, phenotypic, and ecological data. An interesting byproduct of the analysis was to realize the possibility of treating these data as if they were 'citations.' The more one finds the same query sequence, the more that sequence can be considered biologically 'cited', i.e., re-proposed elsewhere in the world. Thus, one can also analyze the h-index of such a ranking. In our Rhizobium dataset, we calculated an h-index = 201, meaning the sequence ranked 201st had 202 identical homologues in GenBank. Although the research effort on given species is directly connected with it, this number provides a quantitative indicator of a taxon's sequence recurrence and distribution within public databases, independent of nomenclatural inconsistencies, offering a novel framework for assessing bacterial representation across global datasets.

RNA, Ribosomal, 16S