PubMed HealthSearch

SEARCH · PubMed Health

Results for “Models, Statistical”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Inferring the sensitivity of wastewater metagenomic sequencing for early detection of viruses: a statistical modelling study.

BACKGROUND: Metagenomic sequencing of wastewater (W-MGS) can in principle detect any known or novel pathogen in a population. We aimed to quantify the sensitivity and cost of W-MGS for viral pathogen detection by jointly analysing W-MGS and epidemiological data for a range of human-infecting viruses. METHODS: In this statistical modelling study, we analysed sequencing data from four studies of untargeted W-MGS to estimate the relative abundance of 11 human-infecting viruses. Corresponding prevalence and incidence estimates were obtained or calculated from academic and public health reports. We combined these estimates using a hierarchical Bayesian model to predict relative abundance at set prevalence or incidence values, allowing comparison across studies and viruses. These predictions were then used to estimate the sequencing depth and concomitant cost required for pathogen detection using W-MGS with or without use of a hybridisation capture enrichment panel. FINDINGS: After controlling for variation in local infection rates, relative abundance varied by orders of magnitude across studies for a given virus. For instance, a local SARS-CoV-2 weekly incidence of 1% corresponded to a predicted SARS-CoV-2 relative abundance ranging from 3·8 × 10-10 to 2·4 × 10-7 across studies, translating to orders-of-magnitude variation in the cost of operating a system able to detect a SARS-CoV-2-like pathogen at a given sensitivity. Use of a respiratory virus enrichment panel in two studies greatly increased predicted relative abundance of SARS-CoV-2, lowering yearly costs by 27-fold (from US$7·87 million to $287 000) and 29-fold (from $1·98 million to $69 100) for a system able to detect a SARS-CoV-2-like pathogen before reaching 0·01% cumulative incidence. INTERPRETATION: The large variation in viral relative abundance after controlling for epidemiological factors indicates that other sources of inter-study variation, such as differences in sewershed hydrology and laboratory protocols, have a substantial impact on the sensitivity and cost of W-MGS. Well chosen hybridisation capture panels can greatly increase sensitivity and reduce cost for viruses in the panel, but might reduce sensitivity to unknown or unexpected pathogens. FUNDING: The Wellcome Trust, Open Philanthropy, and Musk Foundation.

Humans

A statistical simulation model to guide the choices of analytical methods in arrayed CRISPR screen experiments.

An arrayed CRISPR screen is a high-throughput functional genomic screening method, which typically uses 384 well plates and has different gene knockouts in different wells. Despite various computational workflows, there is currently no systematic way to find what is a good workflow for arrayed CRISPR screening data analysis. To guide this choice, we developed a statistical simulation model that mimics the data generating process of arrayed CRISPR screening experiments. Our model is flexible and can simulate effects on phenotypic readouts of various experimental factors, such as the effect size of gene editing, as well as biological and technical variations. With two examples, we showed that the simulation model can assist making principled choice of normalization and hit calling method for the arrayed CRISPR data analysis. This simulation model is implemented in an R package and can be downloaded from Github.

CRISPR-Cas Systems

transfactor: transcription factor activity estimation via probabilistic gene expression deconvolution.

Gene expression is a primary modality being studied to differentiate between biological cells. Contemporary single-cell studies simultaneously measure genome-wide transcription levels for thousands of individual cells in a single experiment. While the characterization of cell population differences has often occurred through differential gene expression analysis, tiny effect sizes become statistically significant when thousands of cells are available for each population, compromising biological interpretation. Moreover, these large studies have spurred the development of methods to infer gene regulatory networks (GRNs) directly from the data, and GRN databases are becoming more comprehensive. In this work, we propose a statistical model for gene expression measures and an inference method that leverage GRNs to deconvolve transcription factor (TF) activity from gene expression, by probabilistically assigning mRNA molecules to TFs. This shifts the paradigm from investigating gene expression differences to regulatory differences at the level of TF activity, aiding interpretation and allowing prioritization of a limited number of TFs responsible for significant contributions to the observed gene expression differences. The inferred TF activities result in intuitive prioritization of TFs in terms of the (difference in) estimated number of molecules they produce, in contrast to other widely used methods relying on arbitrary enrichment scores. Our model allows the incorporation of prior information on the regulatory potential between each TF and target gene and is able to deal with both repressing and activating interactions. We compare our approach to other TF activity estimation methods using two simulation experiments and two case studies. Single-cell RNA-sequencing; TF activity; bioinformatics; GRN.

Transcription Factors

Enhancing detection of polygenic adaptation: a comparative study of machine learning and statistical approaches using simulated evolve-and-resequence data.

BACKGROUND: Detecting signals of polygenic adaptation remains a significant challenge in population genomics, as traditional methods often struggle to identify the associated subtle, multi-locus allele-frequency shifts. Here, we introduced and tested several novel approaches combining machine learning techniques with traditional statistical tests to detect polygenic adaptation patterns in time-series of allele frequency changes from whole genome data. We implemented a Naive Bayesian Classifier (NBC) and One-Class Support Vector Machines (OCSVM), and compared their performance against the classical Fisher's Exact Test (FET). Furthermore, we combined machine learning and statistical models (OCSVM-FET and NBC-FET), resulting in 5 competing approaches. The framework is mainly designed and validated for evolve-and-resequence (EaR) experimental designs, where defined selection pressures and temporal sampling are feasible, but might be applicable for certain natural experiments as well. RESULTS: Using a simulated dataset based on empirical C. riparius Pool-Seq data, we evaluated methods across evolutionary scenarios varying in generation, selection strength, and number of loci under selection. Our results demonstrate that the combined OCSVM-FET approach consistently outperformed competing methods, achieving the lowest false positive rate, highest area under the curve, and high accuracy. The performance peak aligned with what we term the 'late dynamic phase' of adaptation - the period after initial selection has occurred but before fixation - highlighting the method's sensitivity to ongoing selective processes. CONCLUSIONS: Furthermore, we emphasize the critical role of parameter tuning, balancing biological assumptions with methodological rigor. While broader applicability remains an important direction for future work, the present benchmarking is intentionally scoped to EaR experimental contexts.

Machine Learning

Use of a dense single nucleotide polymorphism map for in silico mapping in the mouse.

Rapid expansion of available data, both phenotypic and genotypic, for multiple strains of mice has enabled the development of new methods to interrogate the mouse genome for functional genetic perturbations. In silico mapping provides an expedient way to associate the natural diversity of phenotypic traits with ancestrally inherited polymorphisms for the purpose of dissecting genetic traits. In mouse, the current single nucleotide polymorphism (SNP) data have lacked the density across the genome and coverage of enough strains to properly achieve this goal. To remedy this, 470,407 allele calls were produced for 10,990 evenly spaced SNP loci across 48 inbred mouse strains. Use of the SNP set with statistical models that considered unique patterns within blocks of three SNPs as an inferred haplotype could successfully map known single gene traits and a cloned quantitative trait gene. Application of this method to high-density lipoprotein and gallstone phenotypes reproduced previously characterized quantitative trait loci (QTL). The inferred haplotype data also facilitates the refinement of QTL regions such that candidate genes can be more easily identified and characterized as shown for adenylate cyclase 7.

Adenylyl Cyclases

The Mediating Role of Immune Cells in the Genetically Predicted Relationship between Gut Microbiota and Puerperal Sepsis: A Mendelian Randomization Study.

INTRODUCTION: The causal relationship between gut microbiota and puerperal sepsis (PS) remains unclear, and there is a lack of in-depth research regarding the potential mediating role of immune cells in this context. OBJECTIVE: This study aims to investigate the causal relationship between gut microbiota and PS using Mendelian randomization (MR) analysis and to assess the mediating effects of immune cells on the risk of PS onset through mediation analysis. MATERIALS AND METHODS: We selected data from large-scale genome-wide association studies (GWAS) involving 473 gut microbiota species, 731 immune cell phenotypes, and PS datasets. Univariate MR (UVMR) analysis was employed to explore the causal relationship between gut microbiota and PS, with the primary statistical method being inverse variance weighting (IVW). Multiple statistical models were applied for sensitivity analysis to minimize the confounding effects of horizontal pleiotropy and heterogeneity. Subsequently, a two-step mediation MR analysis was conducted to evaluate whether immune cells mediate the relationship between gut microbiota and PS. RESULTS AND DISCUSSION: Analysis using various statistical models indicated that 11 gut microbiota species (e.g., Azorhizobium, Bacillus velezensis, CAG-245 sp000435175, Lentimicrobiaceae, and Providencia) exhibited a causal relationship with PS. Further reverse causal analysis between PS and gut microbiota ruled out the possibility of reverse causality. The two-step mediation MR analysis demonstrated that the percentage of IgD-CD27- B cells (10.26%) and CD62L- monocytes (17.29%) partially mediated the effect of CAG-245 sp000435175 on PS risk. CONCLUSION: This study provides evidence of a causal relationship between the abundance of certain gut microbiota species and PS, while also revealing a potential mediating role of immune cells. These findings offer valuable theoretical insights into personalized treatment strategies and the development of novel diagnostic biomarkers for PS.

Humans

Beyond predictive performance: A systematic review and critical methodological appraisal of AI/ML and conventional modelling strategies in breast, colorectal, and pancreatic Cancer.

BACKGROUND: Predictive modelling for cancer risk, treatment-related complications, and survival is central to precision oncology. Conventional logistic regression (LR) and Cox proportional hazards (CoxPH) regression remain widely used but are limited when modelling nonlinear interactions, high-dimensional imaging features, and multimodal clinical-metabolic predictors. Artificial intelligence (AI) and machine learning (ML) methods offer expanded capability through automated feature extraction, ensemble learning, and flexible survival modelling, but the evidence on when AI/ML adds value over conventional models across cancer sites and predictive tasks remains fragmented. OBJECTIVE: To systematically evaluate the methodological performance, validation strategies, and translational limitations of AI/ML models compared with conventional statistical models in published predictive-modelling studies for breast, colorectal, or pancreatic cancer. METHODS: PubMed, Scopus, and Web of Science were searched for studies published between January 2019 and March 2025. Two reviewers independently conducted title-and-abstract screening, full-text eligibility assessment, and PROBAST risk-of-bias assessment. Sixty-five studies (n = 907,567 participants) were narratively synthesised by cancer site, predictive task, model family, comparator, validation strategy, predictor modality, and calibration or explainability reporting. RESULTS: The 65 studies comprised breast cancer (n = 35), colorectal cancer (n = 21), and pancreatic cancer (n = 9). AI/ML superiority over LR and CoxPH was task- and data-dependent. CNN- and U-Net-based models predominated in imaging and body-composition tasks, tree-based ensembles consistently outperformed LR for tabular perioperative complication prediction, and CoxPH remained competitive, and in the largest pancreatic risk study, superior to XGBoost (C-index 0.802 vs 0.723) in well-structured datasets. PROBAST analysis-domain risk was moderate in 54 of 65 studies (83%), driven by limited external validation, sparse calibration reporting (11/65), and few decision-curve analyses (7/65). CONCLUSION: AI/ML adds the most methodological value in imaging-derived feature extraction and nonlinear perioperative prediction, while conventional regression remains preferable in large, structured datasets with linear predictors. Clinical translation requires standardised body-composition definitions, external validation, calibration assessment, decision-curve analysis, and explainability, in line with TRIPOD+AI and CLAIM standards.

Humans

TreeFlow: Probabilistic Modelling and Automatic Differentiation for Phylogenetics.

Probabilistic modelling frameworks are powerful tools for statistical modelling and inference. They are not immediately generalizable to phylogenetic problems due to the particular computational properties of the phylogenetic tree object. TreeFlow is a software library for probabilistic modelling and automatic differentiation with phylogenetic trees. It embeds phylogenetic trees in the TensorFlow Probability framework, and implements inference algorithms for phylogenetic models given a fixed tree topology. We demonstrate how TreeFlow can be used to quickly implement and assess new models. We also show that it provides reasonable performance for gradient-based inference algorithms compared to specialized computational libraries for phylogenetics.

Bayesian inference

Trends in the Prevalence of Foods High in Saturated Fats, Sodium, and Added Sugars among U.S. adults, NHANES 2007-2018.

BACKGROUND: Foods and beverages high in saturated fats, sodium, and added sugars (HFSS) are often ultra-processed and linked to poor health outcomes, but few studies have investigated their intake. OBJECTIVE: To describe the trends in the intake of HFSS foods and beverages between 2007 and 2018 in a nationally representative sample of U.S. adults, by sociodemographic characteristics and What We Eat in America food groups. DESIGN: This is a secondary, cross-sectional analysis of the National Health and Nutrition Examination Survey (NHANES) between 2007 and 2018. PARTICIPANTS/SETTING: The final sample included 27,984 adults 19 years of age or older from NHANES with at least one complete dietary recall. MAIN OUTCOME MEASURES: The primary outcomes are the percentage of total energy intake from foods classified as HFSS according to the Pan American Health Organization (PAHO) Nutrient Profile Model. STATISTICAL ANALYSES PERFORMED: To estimate the percentage of energy intake from foods and beverages HFSS, linear regression models with interaction terms between cycles and covariates were used. RESULTS: The overall intake of foods and beverages HFSS did not change, representing over 60% of the total energy intake between 2007-2010 and 2015-2018. The intake of foods and beverages high in sodium increased by 2.0 percentage points (95% CI: 0.5, 3.5) and 3.5 percentage points (95% CI: 1.1, 5.8), respectively. The intake of foods and beverages high in saturated fats increased by 6.1 percentage points (95% CI: 4.5, 7.6) and 6.1 percentage points (95% CI: 3.9, 8.2), respectively. The intake of foods and beverages high in added sugars did not change. CONCLUSION: In the U.S., intake of HFSS foods and beverages is high. Future research should focus on whether public health interventions and policies might reduce the intake of foods high in nutrients of concern.

Added sugars

Polygenic risk score and its role in cancer susceptibility.

BACKGROUND: Polygenic risk score (PRS) has the ability to stratify inherited susceptibility to cancer and, as a complement to monogenic testing, can identify individuals at increased genetic risk even when no pathogenic variant is detected in high- or moderate-penetrance genes. It reflects the combined additive effects of a large number of low-penetrance variants across the genome, and represents a continuum of genetic susceptibility with an approximately normal distribution. Clinically relevant differences are typically observed in individuals in the highest and lowest percentiles of the PRS distribution, while relative risk gradients depend on the cancer type, the specific PRS model, and the reference population used. PRS is not a single test but rather a family of statistical models that differ in their design, predictive performance, and transferability across populations, underscoring the need for external validation and population-specific calibration of absolute risk. Broader implementation is thus still held back by differences between individual PRS models, limited transferability, and the lack of harmonized guidance on indication, reporting, and clinical decision-making. Consequently, clinical use in the European Union remains largely confined to pilot studies and local projects. Within these initiatives, PRS is most commonly applied in two main ways - either as a triage tool for intensified diagnostics or screening in higher-risk groups, or as a component of multifactorial absolute-risk models (e. g. BOADICEA/CanRisk) that integrate PRS with other risk factors such as pathogenic variants in moderate-penetrance genes (e. g. ATM or CHEK2), family history, or lifestyle factors. By refining absolute-risk estimates, PRS may shift individuals across clinical decision thresholds for more intensive surveillance and preventive strategies. AIM: This review summarizes the principles of PRS, the main sources of variability between models, and its potential applications in risk stratification and personalized cancer screening. It also addresses limitations in transferability, the need for calibration, and the currently limited evidence for improvements in hard clinical outcomes.

Humans

Comparing artificial and convolutional neural networks with traditional models for Genomic prediction in wheat.

With the rapid development of sequencing technology, the application of genomic prediction has become more and more common in breeding schemes of livestocks and crops. Selecting an appropriate statistical model is of central importance to achieve high prediction accuracy. Recently, machine learning models have been expected to upgrade genomic prediction into a new era. However, the perspective still suffers from lack of evidence that machine learning models can generally outperform the traditional ones on empirical data sets. In this study, we compared two machine learning models based on artificial neural network (ANN) and convolutional neural network (CNN) with four traditional models, including genomic best linear unbiased prediction (GBLUP), Bayesian ridge regression (BRR), BayesA and BayesB, using three published data sets for grain yield in wheat. For each model, we considered two variants: modeling and ignoring the genotype-by-environment ([Formula: see text]) interaction. In the comparison, we considered two strategies of cross-validation: predicting genotypes that have not been evaluated in any environment (CV1) and predicting genotypes that have been tested in other environments (CV2). Our results showed that traditional Bayesian models (BayesA, BayesB, and BRR) outperformed GBLUP, ANN and CNN when considering [Formula: see text] interaction. The accuracies of ANN and CNN were higher than traditional models only in CV1 and when [Formula: see text] interaction was ignored. It was also found that the performance of the two machine learning models was significantly affected by the interaction between the CV strategy and the way of treating the [Formula: see text] interaction, while that of the four traditional models was only influenced by whether the [Formula: see text] interaction was considered or not. Thus, machine learning models can be a powerful complementary to the traditional ones and their superiority may depend on the prediction scenario. Among the two machine learning models, we observed that the accuracy of ANN was higher than CNN in most cases, indicating that it is still challenging to adapt complex machine learning models such as CNN to genomic prediction.

ANN

Context-specific genetic effects inform endotypes and treatment in asthma.

BACKGROUND: Asthma has heterogeneous risk factors, subtypes, and treatments. It is often unclear how to stratify this heterogeneity in scientific studies and clinical care. Genetics could explain root causes of this clinical heterogeneity, called endotypes, but prior studies have used models that are not designed for complex diseases like asthma. OBJECTIVE: We aimed to find genetic effects that partly explain different asthma endotypes. METHODS: We used recent powerful and robust statistical models of context-specific genetic effects in complex traits. We identified genetic subtypes by clustering clinical asthma features in a case-control cohort, GALA II. We replicated the genetic endotypes in the UK Biobank with gene-context interaction tests. RESULTS: Asthma-associated single nucleotide polymorphisms, polygenic scores, and genome-wide heritability revealed subtype-specific genetic endotypes correlated with type 2 inflammation, allergy, and neuroticism. We validated the type 2 associations with molecular data including nasal RNA sequencing. In the UK Biobank, we replicated these endotypes and found they interact with several polygenic scores and drug-relevant genes. CONCLUSION: Our results show how context-specific genetic effects can unravel biomedically meaningful endotypes of complex disease and suggest novel precision treatment strategies.

Humans

Drug repurposing using transcriptomics: principles and unmet needs in cardiovascular disease.

Although cardiovascular disease is the leading cause of death globally, therapeutic development in this field is slow. Given the high cost of developing new drugs and running clinical trials for cardiovascular disease, repurposing of drugs with approved safety profiles is an attractive strategy for therapeutic development that can significantly reduce the time and cost investment before phase II clinical trials. In the era of "Omics," various new methods and several large databases have been developed to enable the use of transcriptomics data for drug repurposing. This review summarizes the principles and workflow of signature mapping, which forms the foundation of statistical models used for transcriptome-based drug repurposing. We highlight the features of different analysis pipelines and databases that have been developed for signature mapping. These analysis pipelines prioritize genes that are statistically important, an approach that fundamentally differs from the pharmacological approach of identifying disease-driving and therapeutically targetable pathways. Outcomes of signature mapping pipelines are sensitive to the quality of input data, and results are not always reproducible. Moreover, all widely used RNA-seq databases are derived from cancer research and lack high-quality molecular data for cardiovascular disease. These unmet needs call for interdisciplinary collaboration and large networks of cardiovascular research-oriented biobanks to create the databases needed for transcriptomic-based signature mapping for drug repurposing efforts.

Drug Repositioning

Modeling Alternative Conformational States in CASP16.

The CASP16 Ensemble Prediction experiment assessed advances in methods for modeling proteins, nucleic acids, and their complexes in multiple conformational states. Targets included systems with experimental structures determined in two or three states, evaluated by direct comparison to experimental coordinates, as well as domain-linker-domain (D-L-D) targets assessed against statistical models from NMR and SAXS data. This paper focuses on the former class of multi-state targets. Ten ensembles were released as community challenges, including ligand-induced conformational changes, protein-DNA complexes, a trimeric protein, a stem-loop RNA, and multiple oligomeric states of a single RNA. For five targets, some groups produced reasonably accurate models of both reference states (best TM-score >0.75). However, with the exception of one protein-ligand complex (T1214), where an apo structure was available as a template, predictors generally failed to capture key structural details distinguishing the states. Overall, accuracy was significantly lower than for single-state targets in other CASP experiments. The most successful approaches generated multiple AlphaFold2 models using enhanced multiple sequence alignments and sampling protocols, followed by model quality based selection. While the AlphaFold3 server performed well on several targets, individual groups outperformed it in specific cases. By contrast, predictions for one protein-DNA complex, three RNA targets, and multiple oligomeric RNA states consistently fell short (TM-score <0.75). These results highlight both progress and persistent challenges in multi-state prediction. Despite recent advances, accurate modeling of conformational ensembles, particularly RNA and large multimeric assemblies, remains a critical frontier for structural biology.

AlphaFold2

Pilea: profiling bacterial growth dynamics from metagenomes with sketching.

BACKGROUND: Quantifying bacteria's growth rates is essential for understanding their ecological roles and for building predictive models in environmental and clinical settings. Peak-to-trough ratios (PTRs) derived from shotgun metagenomes offer a culture-independent proxy for in situ growth rates of bacterial species, yet their reliable computation remains challenging. RESULTS: We introduce Pilea&#xa0;( https://github.com/xinehc/pilea ), an alignment-free, sketching-based method that incorporates statistical models for robust PTR estimation. Pilea achieves speed improvements over existing methods while also enhancing accuracy, as demonstrated on both simulated and real datasets. CONCLUSIONS: By scaling efficiently to comprehensive reference collections such as the Genome Taxonomy Database (GTDB), Pilea enables large-scale analyses of bacterial growth dynamics across biomes, unlocking new insights for ecological research. Video Abstract.

Bacteria

Dissecting adult plant resistance to stem rust through multi-model GWAS in a diverse barley germplasm panel.

INTRODUCTION: Stem rust (SR), caused by Puccinia graminis f. sp. tritici (Pgt), remains a major threat to global barley production, particularly in regions with conducive environments and evolving pathogen populations. Despite progress in understanding seedling resistance, adult plant resistance (APR) to SR remains underexplored in diverse barley germplasm. This study aimed to dissect the genetic architecture of APR to SR in a panel of diverse origins of two-row spring barley using a genome-wide association study (GWAS). METHODS: A total of 273 barley accessions were evaluated for APR to SR in two distinct environments in Kazakhstan. Phenotypic data were combined with high-density SNP genotyping to perform GWAS using five statistical models (GLM, MLM, MLMM, FarmCPU, and BLINK). Population structure and kinship were accounted for to identify robust marker-trait associations (MTAs), followed by haplotype-based QTL delineation. Transcriptomic data from 16 barley tissues were used to identify candidate genes within major QTL regions. Substantial phenotypic variation in SR severity was observed across environments. RESULTS: A total of 204 MTAs were identified, among which 96 were stable across models, resulting in 19 model-stable QTLs spanning all seven barley chromosomes. Six QTLs co-localized with known SR-resistance QTLs and genes, including Rpg1 and Rpg6. Q_rpg_7H.1 (coinciding with Rpg1) was one of the strongest and most consistent QTL, harboring 42 highly expressed candidate genes. A novel major-effect QTL on chromosome 5H, Q_rpg_5H.1 (3.5 - 9.9 Mb), not previously associated with known resistance loci, contained 10 highly expressed genes grouped into three co-expression clusters, including WRKY transcription factors and PR-5 proteins. CONCLUSION: This study provides new insights into the complex, multilayered genetic control of SR resistance in barley. The discovery of both known and novel QTLs offers valuable targets for marker-assisted selection and lays the foundation for breeding durable SR-resistant barley adapted to diverse agroecological conditions.

Hordeum vulgare L.

Nationwide Survey Using Real-Time PCR in 2024 and 2025 Supports the Absence of Xylella fastidiosa in Korea.

Xylella fastidiosa is a plant-pathogenic bacterium that causes severe diseases in economically important crops, such as citrus and grapevine, thereby posing a significant threat to global agriculture. Although X. fastidiosa has not yet been reported in Korea, the increase in international trade and its presence in neighboring countries highlight the necessity of continued surveillance. The objective of this study was to verify the absence of X. fastidiosa in Korea and to establish a reliable diagnostic framework through a nationwide survey conducted in 2024 and 2025. The sampling design was generated using the RiBESS+ statistical model to ensure the reliability of the survey results. Host plants, including grapevines (Vitis vinifera), mandarin oranges (Citrus unshiu), and cherry blossoms (Prunus yedoensis), were selected and sampled from urban and agricultural areas throughout the country for a nationwide survey. Genomic DNA was extracted from plant petioles and analyzed using real-time PCR with an optimized primer set (XF16S-F/R). Over a period of two years, a total of 2,314 samples were collected, exceeding the required sample size of 843 per year. X. fastidiosa was not detected in any of the collected and tested samples. These results confirm the absence of X. fastidiosa in Korea throughout the study period with high statistical confidence. This study provides evidence confirming the absence of X. fastidiosa in Korea and proposes a standardized methodology for future surveillance and early detection of other invasive prohibited quarantine pests.

X. fastidiosa

Human Systems Immunology in the Omics Era: Challenges, Methods, and Emerging Directions.

The human immune system is a highly complex, dynamic, and heterogeneous network shaped by genetic, environmental, and temporal influences. Advances in high-throughput omics technologies have transformed our ability to study this complexity directly and comprehensively in human cohorts. These developments have positioned systems immunology as a powerful framework for investigating coordinated immune responses, identifying regulatory mechanisms, and linking molecular patterns to clinical phenotypes. However, the analytical challenges inherent to large-scale, multimodal datasets-including batch effects, small sample sizes, high dimensionality, and substantial interindividual heterogeneity-require rigorous study design, robust statistical modeling, and thoughtful data analysis strategies. In this review, we summarize key technological foundations enabling modern human systems immunology, outline common analytical pitfalls and effective mitigation approaches, discuss data integration concepts, and highlight emerging opportunities in the field. Together, these technological and analytical advances are redefining how immune function is measured and interpreted in real-world human biology and hold significant promise for enhancing mechanistic insight, biomarker discovery, and precision medicine across immunological diseases and interventions.

Humans