PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “cross-validation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Striping artifact removal in VisiumHD data through nuclear counts modeling.

MOTIVATION: 10x Genomics VisiumHD enables spatial transcriptomics at 2 µm × 2 µm resolution but exhibits slide-specific, non-periodic striping artifacts due to lane-width variability. These multiplicative row/column effects distort bin total counts and can bias downstream analyses. The state-of-the-art destriping approach is the normalization procedure used as a preprocessing step in bin2cell; it applies sequential high-quantile row- then column-wise normalization, which is asymmetric and can introduce edge effects/macro-stripes and distortions of large-scale total-count structure. RESULTS: We propose a statistical destriping approach that leverages nuclei segmentation from the co-registered H&E image. Assuming transcript abundance is constant within each nucleus, we model bin counts with a negative binomial distribution whose mean is a product of a nucleus-specific concentration and row- and column-specific stripe-factors reflecting lane-width variation. We fit all parameters in a generalized linear modeling framework with cross-validated regularization on stripe-factors and iterative dispersion estimation, and use the fitted parameters to correct the observed counts into a destriped image. On synthetic data with known ground truth, our method improves stripe-factor estimation accuracy and reduces error in corrected counts relative to bin2cell and bin2cell-derived baselines. Across four public VisiumHD slides, it consistently lowers striping intensity while substantially better preserving biological signal present in the large-scale global count structure and avoiding the artifacts introduced by other methods. AVAILABILITY AND IMPLEMENTATION: All source code and links to publicly available data used for this study are available at https://github.com/paolamalsot/destriping-GLM.

Artifacts↗

ExoShorkie: predicting RNA-seq coverage of exogenous genomes in yeast by transfer learning.

MOTIVATION: Predicting the RNA-seq coverage of native and exogenous sequences is central to many molecular- and synthetic-biology applications. Substantial progress has been made in developing methods to predict the RNA-seq coverage of native genomic sequences, with the recently developed Shorkie achieving state-of-the-art performance in yeast. However, prediction performance of these methods over exogenous DNA is still unknown. Recent studies measured RNA-seq coverage of large exogenous genomes in yeast, providing a unique opportunity to train machine-learning models on a large exogenous sequence space and to improve both prediction performance and our understanding of regulatory mechanisms. RESULTS: We introduce ExoShorkie, a method we developed by extending Shorkie through transfer learning across multiple exogenous RNA-seq datasets. We demonstrate that ExoShorkie significantly improves prediction performance on held-out exogenous genomes and outperforms both a native-genome-trained Shorkie baseline and Yorzoi, the only competing method for predicting exogenous RNA-seq coverage in yeast, in cross-validation and in leave-one-genome-out evaluations. Furthermore, through interpretability analyses we reveal biologically meaningful regulatory motifs and distinct regulatory rules in exogenous genomes in yeast, providing new insights into transcriptional regulation. AVAILABILITY AND IMPLEMENTATION: ExoShorkie is available at https://github.com/OrensteinLab/ExoShorkie.

Genome, Fungal↗

GiantHost: a domain-adaptive and uncertainty-aware framework for giant virus host prediction.

MOTIVATION: Nucleocytoplasmic large DNA viruses (NCLDVs) play crucial roles in global ecosystems. Although metagenomics has vastly accelerated the discovery of novel NCLDVs, predicting their hosts from fragmented contigs remains a critical bottleneck, with no dedicated end-to-end computational tools currently available. Addressing this gap requires overcoming three fundamental challenges: the extreme scarcity of labeled reference genomes, the severe domain shift between laboratory isolates and diverse environmental metagenomes, and the inability of traditional deterministic models to quantify prediction uncertainty-a crucial requirement for reliable ecological profiling where novel, divergent viruses are prevalent. RESULTS: We present GiantHost, the first NCLDV host prediction tool with domain adaptation and uncertainlty awareness. GiantHost employs a dual-tower neural network to integrate dense genome traits and sparse GVOG profiles, allowing better integration of heterogeneous features. To overcome label scarcity and domain shift, we leverage 1400 environmental viral genomes (GVMAGs) via semi-supervised multi-task learning and Domain Adversarial Neural Networks (DANN), effectively bridging the distributional gap between RefSeq and environmental data. Additionally, GiantHost incorporates Conformal Prediction (CP) to output statistically guaranteed prediction sets rather than overconfident single labels. Evaluated under rigorous genome-level cross-validation, GiantHost demonstrates robust predictive power. Applied to the Tara Ocean dataset, GiantHost successfully captured the vertical stratification of NCLDV hosts-revealing a depth-dependent decline of phytoplankton-infecting viruses and a relative enrichment of Amoebozoa-infecting viruses in the mesopelagic zone. AVAILABILITY: The source code of GiantHost is available via: https://github.com/FuchuanQu/GiantHost.

Giant Viruses↗

abCRISPR: deep learning-based design of abasic gRNA sequences for specific CRISPR-Cas9 genome editing.

SUMMARY: CRISPR-Cas9 has become a widely used tool for genome editing. However, its off-target cleavage caused by partial sequence matches with guide RNAs (gRNAs) remains a critical limitation. Recently, abasic gRNAs (ØXØ) have been developed to enhance target specificity, but their effects vary depending on the positional sequence context. Here, we present abCRISPR, a deep neural network (DNN) framework for the rational design of ØXØ sequences with minimized off-target activity. abCRISPR leverages informative few-shot training with paired datasets of abasic and unmodified gRNAs, using high-quality random mismatch target libraries, exhaustively sequenced for mismatched off-target substrates (n = 97583) in in vitro CRISPR-Cas9 cleavage experiments. Predicted off-target activities for both abasic and unmodified gRNAs showed strong correlation with experimental data (r ≥ 0.95, 10-fold cross-validation). Notably, these comprehensive training sets provide robust ground-truth negatives, enabling accurate and sensitive prediction of off-targets. For unmodified gRNAs, abCRISPR (AUC = 0.98) was validated to outperform existing deep learning-based methods (AUC = 0.45-0.68). When applied to the human genome, abCRISPR generated ØXØ sequences, covering 58 875 004 potent CRISPR-targetable sites with improved target specificity. Together, this work provides a comprehensive bioinformatics resource for safe and precise CRISPR-Cas9 genome editing. AVAILABILITY AND IMPLEMENTATION: The source code for abCRISPR and training data are available at https://doi.org/10.5281/zenodo.20398246. abCRISPR results for the human genome are available at http://clip.korea.ac.kr/abCRISPR/.

Deep Learning↗

Metabolomics Reveals Metabolic Characteristics of Functional Cure in Chronic Hepatitis B Treated With Entecavir Combined With Pegylated Interferon Alpha.

BACKGROUND: Entecavir (ETV) combined with pegylated interferon alpha (PEG-IFNα) improves chronic hepatitis B (CHB) functional cure rates, but therapeutic heterogeneity and underlying metabolic mechanisms remain unclear. This study used untargeted metabolomics to identify metabolic signatures, mechanisms, and predictive biomarkers of functional cure with ETV-PEG-IFNα. METHODS: Thirty-eight CHB patients were grouped into ETV monotherapy (Group E, n = 12) and ETV-PEG-IFNα combination therapy (Group Z, n = 26); Group Z was subdivided into cured (Group A, n = 13) and noncured (Group B, n = 13). Serum metabolomic profiling, multivariate statistics, and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analysis identified differential metabolites. A random forest model was built using key metabolites. RESULTS: Three hundred eighty-eight metabolites were identified. Four differential metabolites distinguished Group A and B (upregulated guanidinoacetic acid, uracil 5-carboxylate; downregulated L-methionine S-oxide, oleamide), enriching amino acid metabolism pathways. Nine differential metabolites between Group E and Z implicated amino acid, immune, and fatty acid pathways. The random forest model based on the four Group A/B metabolites showed 88.5% cross-validation accuracy (AUC = 0.920), with L-methionine S-oxide and oleamide as key predictors. CONCLUSIONS: This study reveals metabolic rewiring in CHB functional cure via ETV-PEG-IFNα therapy, involving energy metabolism, oxidative stress, and immunomodulation, based on which we propose a tentative metabolism-immunity synergy model to guide future research. Key metabolites, especially L-methionine S-oxide and oleamide, show exploratory predictive potential for functional cure that warrants further validation in independent cohorts.

Humans↗

Early Transcriptional Changes in Neutrophil-Mediated Processes Following Recanalization After Ischemic Stroke.

BACKGROUND: Ischemic stroke is a leading cause of death and long-term disability worldwide. Recanalization therapies, including thrombolysis and mechanical thrombectomy, restore blood flow, yet many patients experience poor outcomes, a phenomenon known as futile recanalization. Given the short therapeutic window for ischemic stroke, identifying early biomarkers to guide targeted interventions and improve outcomes is critical. METHODS: Using a murine middle cerebral occlusion model that mimics a large vessel occlusion with recanalization, a comprehensive microarray analysis from blood samples collected immediately and 3 hours after recanalization (N=44) was performed. Differentially expressed genes, enrichment pathways, immune cell proportions, enriched cell markers, predicted micro-RNAs, and transcription factors were identified using RStudio. Findings in mice were validated with rat middle cerebral artery occlusion (GSE21136) and patients with stroke (GSE16561) data sets to confirm transcriptional changes in peripheral blood postrecanalization. RESULTS: Il1r2, Cd55, Mmp8, Cd14, and Cd69 were early biomarkers poststroke and postrecanalization. Cross-validation revealed Vcan as a differentially expressed gene conserved across species, making it a novel ischemic marker detected as early as 3 hours postrecanalization (4 hours after middle cerebral artery occlusion) in mice, 24 hours after recanalization in rats (middle cerebral artery occlusion-thrombectomy), and within 24 hours from onset in humans receiving recombinant tissue plasminogen activator-thrombolysis. CIBERSORTx and ImmuCellAI-mouse deconvolution showed neutrophil elevation postrecanalization. Leukocyte and neutrophil activation pathways were enriched early after stroke in mice and humans, with stronger upregulation in the female sex. Several regulatory micro-RNAs were identified, and Nuclear Factor Erythroid 4 (NFE4) and Metal Regulatory Transcription Factor 1 (MTF1) emerged as key transcription factors. A coregulatory network underlying neutrophil activity was constructed, highlighting its central role in early responses to ischemia and recanalization, which was enriched in the female sex. CONCLUSIONS: We identified novel early genomic markers for ischemia and recanalization, including the conserved marker Vcan, and highlighted age- and sex-specific immune responses. Mapping a neutrophil-centered coregulatory network provides mechanistic insight into futile recanalization and supports the development of targeted therapies to improve clinical outcomes.

Animals↗

CAGNet: a structure-aware clustering-alternated graph network for cell-cell interaction inference in spatial transcriptomics.

MOTIVATION: Understanding cell-cell interactions (CCIs) in spatial transcriptomics is crucial for uncovering the spatial organization and functional heterogeneity of tissues. However, existing graph-based models typically rely on static clustering or fixed adjacency structures, which limits their ability to capture dynamic cellular relationships. RESULTS: We propose CAGNet, a two-stage framework for CCI inference from spatial transcriptomics data. In Stage 1, a Graph Attention Network encoder with joint feature and graph reconstruction learns structure-aware node embeddings from spatial gene expression profiles. In Stage 2, an alternating optimization mechanism iteratively updates cluster centers via KL-guided soft assignment and refines node embeddings through spatial graph reconstruction, establishing a closed-loop between representation learning and clustering. Experiments on three 10x Genomics Visium datasets demonstrate that CAGNet consistently outperforms six CCI inference baselines across ACC, AUC, AP, Precision, Recall, and F1. CAGNet also achieves the highest Adjusted Rand Index on all three datasets against six spatial domain identification methods, confirming that the learned embeddings capture biologically relevant spatial organization. Information-theoretic analysis further shows that CAGNet retains the highest mutual information between input features and learned embeddings among all compared methods. Ablation studies and 5-fold cross-validation confirm the contribution of each component and the reproducibility of the results. AVAILABILITY: The proposed method is implemented in the CAGNet package available at http://github.com/mahan1233333-maker/CAGNet .

Spatial Transcriptomics↗

Comprehensive in silico genomics analysis of global trends and host-specific emergence of aminoglycoside resistance in Staphylococcus aureus: a One-Health perspective.

BACKGROUND: Aminoglycosides remain clinically valuable against Staphylococcus aureus. Aminoglycoside resistance in S. aureus represents a critical One Health concern and is primarily driven by aminoglycoside-modifying enzymes (AMEs), which are frequently plasmid-encoded. Although regional studies have provided valuable insights, the global epidemiology of aminoglycoside resistance determinants remains poorly characterized because comprehensive data integrating human, animal, and environmental reservoirs are still lacking. This study addresses this gap by analyzing over 110,000 S. aureus genomes (2000-2025) to map the global resistome, quantify temporal and host-specific trends, and assess the association between genetic determinants and phenotypic resistance. METHODS: We performed a retrospective One Health meta-analysis of 110,309 S. aureus genomes collected between 2000 and 2025 from 128 countries. Genomes were quality-filtered and aminoglycoside resistance determinants were identified using NCBI AMRFinderPlus (v4.0.23). Multilocus sequence typing and host-source harmonization (Human, Animal, Environment, Unknown) enabled clonal and reservoir stratification. Temporal trends in gene prevalence and resistance burden were modeled with robust regression. Geographic and host-associated structuring of key genes was assessed via &#x3c7;2 and enrichment tests. Machine-learning models (elastic-net, random forests, XGBoost) were benchmarked for minimum inhibitory concentration (MIC) prediction via nested cross-validation, with performance evaluated by mean absolute error, RMSE, and SHAP-based feature importance. All analyses were conducted in R and Python using publicly available, de-identified genomic data. RESULTS: Aminoglycoside resistance-associated genes were dominated by modifying enzyme determinants, with ant(6)-Ia, ant(9)-Ia, aph(3')-IIIa, sat4, aadD1, and aac(6')-Ie/aph(2'')-Ia occurring in 14-22% of isolates worldwide. Temporal analysis revealed significant declines in several major determinants, most notably ant(9)-Ia (-2.22 percentage points per year, p&#x2009;<&#x2009;0.001), whereas apmA exhibited a non-significant decreasing trend in animal isolates. Host structuring was marked: human clinical isolates concentrated common determinants, while animal and environmental isolates harbored rare alleles (apmA, spw, str, spd). Geographic mapping confirmed near-universal distribution of common genes but focal restriction of rare ones. Publicly available phenotypic data indicated strong activity of amikacin, whereas gentamicin showed a distinct resistant subpopulation that closely corresponded with AME gene carriage. Genotype-phenotype analyses demonstrated strong concordance, with gene-rich complements predicting resistant MIC strata and absence of determinants predicting susceptibility. Analysis across different gene classes revealed frequent co-occurrence of aminoglycoside resistance genes with determinants from other classes, such as mecA, blaZ, and MLS_B, embedding them within multidrug-resistant (MDR) genomic contexts. CONCLUSION: Over 25&#xa0;years, the prevalence of aminoglycoside resistance-associated genes in S. aureus has declined for several common determinants, while rare veterinary-linked alleles are emerging in animal isolates. Strong genotype-phenotype concordance supports genomic prediction for gentamicin and amikacin, where MIC data are available, although phenotypic confirmation remains essential. The frequent co-occurrence of aminoglycoside resistance genes with other antimicrobial resistance determinants indicates their integration within co-occurrence patterns of MDR genes, defined here as clusters of co-occurring resistance genes often carried on shared mobile genetic elements. These patterns highlight the need for integrated One Health surveillance combining clinical, veterinary, and environmental monitoring with plasmid-context resolution to anticipate emerging threats.

Aminoglycosides↗

N6-methyladenine identification using deep learning and discriminative feature integration.

N6-methyladenine (6&#xa0;mA) is a pivotal DNA modification that plays a crucial role in epigenetic regulation, gene expression, and various biological processes. With advancements in sequencing technologies and computational biology, there is an increasing focus on developing accurate methods for 6&#xa0;mA site identification to enhance early detection and understand its biological significance. Despite the rapid progress of machine learning in bioinformatics, accurately detecting 6&#xa0;mA sites remains a challenge due to the limited generalizability and efficiency of existing approaches. In this study, we present Deep-N6mA, a novel Deep Neural Network (DNN) model incorporating optimal hybrid features for precise 6&#xa0;mA site identification. The proposed framework captures complex patterns from DNA sequences through a comprehensive feature extraction process, leveraging k-mer, Dinucleotide-based Cross Covariance (DCC), Trinucleotide-based Auto Covariance (TAC), Pseudo Single Nucleotide Composition (PseSNC), Pseudo Dinucleotide Composition (PseDNC), and Pseudo Trinucleotide Composition (PseTNC). To optimize computational efficiency and eliminate irrelevant or noisy features, an unsupervised Principal Component Analysis (PCA) algorithm is employed, ensuring the selection of the most informative features. A multilayer DNN serves as the classification algorithm to identify N6-methyladenine sites accurately. The robustness and generalizability of Deep-N6mA were rigorously validated using fivefold cross-validation on two benchmark datasets. Experimental results reveal that Deep-N6mA achieves an average accuracy of 97.70% on the F. vesca dataset and 95.75% on the R. chinensis dataset, outperforming existing methods by 4.12% and 4.55%, respectively. These findings underscore the effectiveness of Deep-N6mA as a reliable tool for early 6&#xa0;mA site detection, contributing to epigenetic research and advancing the field of computational biology.

Deep Learning↗

Proteomic discovery analysis of quantitatively assessed emphysema in the general population. The MESA Lung Study.

BACKGROUND: Pulmonary emphysema occurs frequently in older adults, often without airflow limitation. Its presence predicts symptoms, respiratory hospitalizations and deaths, and all-cause mortality. Proteomics may provide further insights into emphysema pathogenesis and inform therapeutic targets. OBJECTIVE: We performed a proteomic discovery analysis of percent emphysema on computed tomography (CT) in a population-based, multiethnic sample from the Multi-Ethnic Study of Atherosclerosis (MESA) Lung Study. Replication was performed in two chronic obstructive pulmonary disease (COPD)-based studies, the SubPopulations and InteRmediate Outcome Measures in COPD Study (SPIROMICS) and the Genetic Epidemiology of COPD (COPDGene) Study. METHODS: MESA recruited participants from the general population in 2000-02. The MESA Lung Study performed full-lung CT scans in 2010-12. Percent emphysema was defined as the percentage of lung voxels&#x2009;<&#x2009;-950 Hounsfield units. Over 7,200 plasma aptamers were measured via SomaScan. Cross-sectional linear and least absolute shrinkage and selection operator (LASSO) regression models were adjusted for demographics, anthropometrics, smoking, renal function, and scanner parameters. Statistical significance was defined as a false discovery rate p-value&#x2009;<&#x2009;0.05. Gene Ontology (GO)/Reactome enrichment analyses were performed. LASSO-selected proteins' predictive performance was evaluated. RESULTS: Among 2,504 participants in the MESA Lung Study, mean age was 69.4&#xa0;years, 1,291 had ever smoked, and median percent emphysema-like lung was 1.4%. In total, 1,234 aptamers were significantly associated with percent emphysema in the MESA Lung Study, and 35 replicated in the SPIROMICS and COPDGene Studies. Novel associations included protein family with sequence similarity (FAM) 177A1, syntenin-2, ubiquitin carboxyl-terminal hydrolase 25, and uncharacterized protein C20orf173. Previously identified emphysema-associated proteins included soluble advanced glycosylation end product-specific receptor (sRAGE), protein S100-A12, high mobility group protein B1, and roundabout homolog 2. Enrichment analyses identified 40 GO biological processes, including chemokine production and regulation and cell-cell adhesion and regulation, and two Reactome pathways, including RAGE signaling. In tenfold cross-validation, novel proteins were largely retained by LASSO (R2&#x2009;=&#x2009;5.4%), improved overall model performance (R2&#x2009;=&#x2009;24.8%), and uniquely explained greater variance in percent emphysema. CONCLUSIONS: This analysis in a general population sample identified novel and previously characterized proteins whose functional roles were validated by GO/Reactome enriched pathways, offering new insights into emphysema pathophysiology and therapeutics.

Humans↗

Plasma inflammatory proteome profiles identify MASLD among children with overweight or obesity.

BACKGROUND & AIMS: Pediatric metabolic dysfunction-associated steatotic liver disease (MASLD) is increasingly prevalent among children with overweight or obesity, yet its early diagnosis remains a major clinical challenge. This study aimed to identify circulating inflammatory proteins associated with MASLD and to develop a proteomic risk score (ProScore) to improve diagnostic accuracy. METHODS: In this cross-sectional study of 161 children (median age 8.5&#xa0;years) with overweight or obesity, MASLD was assessed by vibration-controlled transient elastography, with 42 cases identified. Plasma concentrations of 92 inflammation-related proteins were quantified using a high-throughput proximity extension assay. The ProScore was compared with eleven conventional anthropometric/metabolic indices (WHtR, METS-IR, SPISE, PNFI, VAI, LAP, TyG, TyG-ALT, TyG-WC, TyG-WHtR, and TyG-BMI) and a genetic risk score (GRS). Six machine learning algorithms were employed and diagnostic performance was assessed using area under the curve (AUC) with fivefold cross-validation. RESULTS: Fifteen proteins were significantly associated with MASLD. A six-protein panel (FGF-21, CDCP1, CD244, OPG, Flt3L, MCP-1) achieved the highest diagnostic accuracy (AUC&#x2009;=&#x2009;0.84), exceeding that of all conventional indices (AUC&#x2009;=&#x2009;0.65-0.78; all P&#x2009;<&#x2009;0.05). ProScore performance remained robust in school-based validation (AUC&#x2009;=&#x2009;0.83), with no substantial improvement when combined with conventional indices. Diagnostic accuracy was higher in children with lower GRS (AUC&#x2009;=&#x2009;0.92) than in those with higher GRS (AUC&#x2009;=&#x2009;0.80; P&#x2009;=&#x2009;0.003). CONCLUSIONS: A proteomic signature of systemic inflammation provides accurate, non-invasive identification of MASLD in at-risk children, outperforming conventional metabolic and genetic tools, and may have utility in clinical and public health settings.

Humans↗

Exploration of predictive and prognostic alternative splicing signatures in lung adenocarcinoma using machine learning methods.

BACKGROUND: Alternative splicing (AS) plays critical roles in generating protein diversity and complexity. Dysregulation of AS underlies the initiation and progression of tumors. Machine learning approaches have emerged as efficient tools to identify promising biomarkers. It is meaningful to explore pivotal AS events (ASEs) to deepen understanding and improve prognostic assessments of lung adenocarcinoma (LUAD) via machine learning algorithms. METHOD: RNA sequencing data and AS data were extracted from The Cancer Genome Atlas (TCGA) database and TCGA SpliceSeq database. Using several machine learning methods, we identified 24 pairs of LUAD-related ASEs implicated in splicing switches and a random forest-based classifiers for identifying lymph node metastasis (LNM) consisting of 12 ASEs. Furthermore, we identified key prognosis-related ASEs and established a 16-ASE-based prognostic model to predict overall survival for LUAD patients using Cox regression model, random survival forest analysis, and forward selection model. Bioinformatics analyses were also applied to identify underlying mechanisms and associated upstream splicing factors (SFs). RESULTS: Each pair of ASEs was spliced from the same parent gene, and exhibited perfect inverse intrapair correlation (correlation coefficient&#x2009;=&#x2009;-&#x2009;1). The 12-ASE-based classifier showed robust ability to evaluate LNM status of LUAD patients with the area under the receiver operating characteristic (ROC) curve (AUC) more than 0.7 in fivefold cross-validation. The prognostic model performed well at 1, 3, 5, and 10&#xa0;years in both the training cohort and internal test cohort. Univariate and multivariate Cox regression indicated the prognostic model could be used as an independent prognostic factor for patients with LUAD. Further analysis revealed correlations between the prognostic model and American Joint Committee on Cancer stage, T stage, N stage, and living status. The splicing network constructed of survival-related SFs and ASEs depicts regulatory relationships between them. CONCLUSION: In summary, our study provides insight into LUAD researches and managements based on these AS biomarkers.

Adenocarcinoma of Lung↗

Machine learning-based clinical prediction model and multi-omics integration for assessing pancreatic cancer risk in new-onset diabetes.

BACKGROUND: Given that pancreatic cancer (PC) is typically diagnosed at an advanced stage but is often preceded by new-onset diabetes mellitus (NODM), providing a window for early detection, we sought to develop and validate an interpretable machine-learning model integrated with multi-omics profiling to identify early biomarkers of NODM-associated PC. METHODS: In a population-based cohort, individuals with NODM-associated PC and NODM without PC were identified and randomly divided (70:30) into training and validation sets after feature selection. Eight machine learning (ML) classifiers were compared using fivefold cross-validation, and model performance was evaluated in terms of discrimination, calibration, and decision curve&#x2013;based clinical utility. We evaluated interpretability using the Shapley additive explanations (SHAP) analyses. Mechanistically, Olink proteomic profiling and metabolomics were analyzed through clinical classifications and model-defined risk strata. RESULTS: Categorical boosting achieved the best performance in the independent validation set (AUROC&#x2009;=&#x2009;0.844). The NODM cohort was stratified into high- (n&#x2009;=&#x2009;2,362) and low-risk (n&#x2009;=&#x2009;5,030) groups, and internal validation together with SHAP analyses demonstrated consistent model performance and identified clinically interpretable predictors. Proteomic and metabolomic analyses under clinical and risk-based grouping identified 39 overlapping differentially expressed proteins and 145 overlapping metabolites with enriched across 11 shared KEGG pathways. Cross-platform validation highlighted PLTP, CRTAC1, and ITGAV as serum biomarkers with a strong potential for early NODM-PC detection. CONCLUSIONS: We developed an interpretable ML framework centered on NODM enables practical risk stratification for early PC detection by multi-omics and provides a pathway of ML-based triage followed by biomarker confirmation for earlier detection and diagnosis.

Humans↗

Genetics-Informed Mapping Identifies a CRIM1-Associated Endocardial Inflammatory Remodeling State in Acute Myocardial Infarction.

BACKGROUND Acute myocardial infarction (AMI) reflects inherited susceptibility and inflammatory remodeling, but the cellular contexts linking genetic risk to disease remain unclear. MATERIAL AND METHODS We integrated a meta-transcriptome-wide association study (TWAS) with a human cardiac single-nucleus RNA-sequencing atlas contained 11 individuals (5 AMI and 6 donor) to identify genetics-informed cellular programs. Composite program states were defined by global score quartiles. A fixed 5-gene panel was evaluated for nucleus-level endocardial low-transcriptional-state (Endo_LTS) vs endocardial high-transcriptional-state (Endo_HTS) discrimination within the AMI endocardium using 5-fold leave-1-patient-out cross-validation. Functional follow-up used CRIM1 silencing in hypoxia-treated human induced pluripotent stem cell (hiPSC)-derived endocardial endothelial-like cells and complementary peripheral blood analyses. RESULTS The endocardium exhibited the most prominent infarction-associated increase in TWAS-anchored program activity, with expansion of program-high states and higher CytoTRACE scores. A consensus 5-gene panel (RPS8, PLEC, CFDP1, CRIM1, TNS2) was identified. Among 2163 AMI endocardial nuclei from 5 patients, the state classifier included 364 Endo_LTS and 751 Endo_HTS nuclei; 1048 Endo_MTS nuclei were excluded. Pooled out-of-fold ROC-AUCs ranged from 0.665 to 0.831. The panel also showed discriminatory value in an independent peripheral-blood AMI-vs-control cohort. CRIM1 was prioritized as a candidate linked to the remodeling program. CRIM1 silencing attenuated ACTA2/alpha-SMA, vimentin, LDHA, CCL2, and VEGFA and partially restored CD31, whereas TGF-&#xdf; remained elevated. CONCLUSIONS These findings identify a genetics-informed endocardial inflammatory remodeling state in AMI and define a 5-gene surrogate of its activated state. CRIM1 is prioritized as a candidate linked to selected inflammatory, metabolic, and structural outputs. Persistent TGF-b elevation after CRIM1 silencing argues against a simple linear regulatory model and indicates that further mechanistic validation is required.

Humans↗

VarPPUD: Pinpointing diagnostic variants from sets of prioritized, strong candidate variants.

Rare and ultra-rare genetic conditions are estimated to impact nearly 1 in 17 people worldwide, yet accurately pinpointing the diagnostic variants underlying each of these conditions remains a formidable challenge. Because comprehensive, in vivo functional assessment of all possible genetic variants is infeasible, clinicians instead consider in silico variant pathogenicity predictions to distinguish plausibly disease-causing from benign variants across the genome. However, in the most difficult undiagnosed cases, such as those accepted to the Undiagnosed Diseases Network (UDN), existing pathogenicity predictions cannot reliably discern true etiological variant(s) from other deleterious candidate variants that were prioritized through case- or family-level analyses. Pinpointing the disease-causing variant from a small pool of plausible candidates remains a largely manual effort requiring extensive clinical workups, functional and experimental assays, and eventual identification of genotype- and phenotype-matched individuals. Here, we introduce VarPPUD, a tool trained on prioritized variants from UDN cases, that leverages gene-, amino acid-, and nucleotide-level features to discern pathogenic (disease causative) variants from other damaging or deleterious variants that are unlikely to be confirmed as relevant to the disease. VarPPUD achieves a cross-validated accuracy of 79.3% and precision of 77.5% on a held-out subset of uniquely challenging UDN cases, respectively representing an average 18.6% and 23.4% improvement over nine existing state-of-the-art pathogenicity prediction tools on this task. We validate VarPPUD's ability to discriminate likely from unlikely pathogenic variants using both synthetic data generated via a GAN-based framework and a temporally held-out set of UDN patients evaluated between 2022 and 2024. The model was trained exclusively on data available through 2021 and applied without retraining to the post-2021 cohort, demonstrating strong generalizability to newly accrued cases. Finally, we show how VarPPUD can be probed to evaluate each input feature's importance and contribution toward prediction-an essential step toward understanding the distinct characteristics of newly-uncovered disease-causing variants.

Humans↗

Opportunities for machine learning to predict cross-neutralization in FMDV serotype O.

Accurately estimating cross-neutralization between serotype O foot-and-mouth disease viruses (FMDVs) is critical for guiding vaccine selection and disease management. In this study, we developed a machine learning approach to estimate r1 values-an established measure of antigenic similarity-using VP1 sequence data and published virus neutralization titer (VNT) results. Our dataset comprised 108 serum-virus pairs representing 73 distinct FMDV strains. We applied Boruta feature selection and random forest classifiers, optimizing model performance through tenfold cross-validation and sub-sampling to address class imbalance. Predictors included pairwise amino acid distances, site-specific polymorphisms, and differences in potential N-glycosylation sites. Using a 0.3 r1 threshold to define cross-neutralization, the final model achieved high accuracy (0.96), sensitivity (0.93), and specificity (0.96) in training, and performed robustly on independent test sets - accuracy was 0.75 (95% CI 0.60 and 0.90), F1 score 0.86% and PPV 0.77. Importantly, key VP1 residues-positions 48, 100, 135, 150, and 151-emerged as strong predictors of antigenic relationships. Our results demonstrate the utility of integrating routinely generated genomic data with machine learning to inform vaccine candidate selection and anticipate immune interactions among circulating FMDV strains. This approach offers a practical tool for accelerating vaccine decision-making and can be adapted to other FMDV serotypes. The latest version of the r1 predictive model is available for access via a Shiny dashboard (https://dmakau.shinyapps.io/PredImmune-FMD/).

Foot-and-Mouth Disease Virus↗

Combination of computational techniques and RNAi reveal targets in Anopheles gambiae for malaria vector control.

Increasing reports of insecticide resistance continue to hamper the gains of vector control strategies in curbing malaria transmission. This makes identifying new insecticide targets or alternative vector control strategies necessary. CLassifier of Essentiality AcRoss EukaRyote (CLEARER), a leave-one-organism-out cross-validation machine learning classifier for essential genes, was used to predict essential genes in Anopheles gambiae and selected predicted genes experimentally validated. The CLEARER algorithm was trained on six model organisms: Caenorhabditis elegans, Drosophila melanogaster, Homo sapiens, Mus musculus, Saccharomyces cerevisiae and Schizosaccharomyces pombe, and employed to identify essential genes in An. gambiae. Of the 10,426 genes in An. gambiae, 1,946 genes (18.7%) were predicted to be Cellular Essential Genes (CEGs), 1716 (16.5%) to be Organism Essential Genes (OEGs), and 852 genes (8.2%) to be essential as both OEGs and CEGs. RNA interference (RNAi) was used to validate the top three highly expressed non-ribosomal predictions as probable vector control targets, by determining the effect of these genes on the survival of An. gambiae G3 mosquitoes. In addition, the effect of knockdown of arginase (AGAP008783) on Plasmodium berghei infection in mosquitoes was evaluated, an enzyme we computationally inferred earlier to be essential based on chokepoint analysis. Arginase and the top three genes, AGAP007406 (Elongation factor 1-alpha, Elf1), AGAP002076 (Heat shock 70kDa protein 1/8, HSP), AGAP009441 (Elongation factor 2, Elf2), had knockdown efficiencies of 91%, 75%, 63%, and 61%, respectively. While knockdown of HSP or Elf2 significantly reduced longevity of the mosquitoes (p<0.0001) compared to control groups, Elf1 or arginase knockdown had no effect on survival. However, arginase knockdown significantly reduced P. berghei oocytes counts in the midgut of mosquitoes when compared to LacZ-injected controls. The study reveals HSP and Elf2 as important contributors to mosquito survival and arginase as important for parasite development, hence placing them as possible targets for vector control.

Animals↗

Shared genetic basis and spatial cellular atlas of psoriasis and metabolic syndrome.

BACKGROUND: Psoriasis (PS) and metabolic syndrome (MetS) frequently co-occur. Characterizing their shared genetic architecture and spatially enriched cellular populations may clarify the context of their co-occurrence and generate hypotheses for functional validation. METHODS: We integrated genome-wide association study (GWAS) summary statistics for PS, MetS, and five related components with spatially resolved single-cell transcriptomic data. Global and local genetic correlations were assessed using linkage disequilibrium score regression, genetic covariance analysis, high-definition likelihood, and local analysis of variant association. A bivariate causal mixture model quantified polygenic overlap. Conditional/conjunctional false discovery rate and composite-null pleiotropy analyses identified shared susceptibility loci. Finally, gsMap evaluated trait-associated enrichment across annotated embryonic tissues at single-cell resolution. RESULTS: Genetic approaches identified significant genome-wide correlations and polygenic sharing between PS, MetS, and its components. Local and cross-trait analyses identified region-specific signals and cross-validated shared loci. gsMap revealed trait-specific tissue enrichment. PS showed the strongest enrichment in the epidermis (pCauchy&#x2009;=&#x2009;1.0573&#x2009;&#xd7;&#x2009;10&#x2009; -&#x2009;&#x2074;), adipose tissue (pCauchy&#x2009;=&#x2009;1.5366&#x2009;&#xd7;&#x2009;10&#x2009;-&#x2009;&#x2074;), and liver (pCauchy&#x2009;=&#x2009;1.0167&#x2009;&#xd7;&#x2009;10&#x2009;-&#x2009;&#xb3;). Across MetS, FBG, HDL-C, hypertension, and TG, enriched regions mainly involved the liver, adipose tissue, and epidermis. WC enrichment was predominantly observed in adipose tissue (pCauchy&#x2009;=&#x2009;1.7823&#x2009;&#xd7;&#x2009;10&#x2009;-&#x2009;&#x2074;), with no significant liver or epidermal enrichment. CONCLUSION: Integrating GWAS with single-cell transcriptomic and spatial information characterized shared genetic architecture between PS and MetS-related phenotypes and their spatial enrichment patterns. These findings provide a framework for generating testable hypotheses about comorbidity biology and guiding future functional and clinical validation.

Psoriasis↗