PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “cross-validation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Assay-dependent variability in peptide biomarker quantification: experimental evidence from renalase in chronic kidney disease.

BACKGROUND: Renalase is a promising biomarker for kidney disease, but published levels vary widely between studies. We hypothesised that variability in commercial enzyme-linked immunosorbent assays (ELISAs) kits and matrix effects (serum vs plasma) drive these inconsistencies. METHODS: Paired serum and plasma samples from 56 participants (28 chronic kidney disease (CKD) stages 2-5, 28 healthy controls) were tested using three commercial renalase ELISAs (BTLAB, Cloud-Clone, EIAab). We assessed intra-assay precision, inter-assay agreement (Spearman's rank correlation and Bland-Altman analysis on log10-transformed values), matrix effects, and associations with estimated glomerular filtration rate (eGFR). Diagnostic performance was evaluated by Receiver operating characteristic (ROC) analysis. RESULTS: Inter-assay renalase concentrations differed markedly (up to orders of magnitude), with weak inter-assay correlations (r&#x2009;&#x2264;&#x2009;0.25). Bland-Altman analyses revealed large, systematic biases between kits. Only the BTLAB assay showed consistent serum/plasma agreement, a significant correlation with eGFR (&#x3c1;&#x2009;&#x2248;&#x2009;0.32-0.42, p&#x2009;<&#x2009;0.05), and moderate discriminatory performance for CKD in serum (AUC = 0.70) and plasma (AUC = 0.68). Cloud-Clone and EIAab produced divergent results and strong matrix-dependent biases. CONCLUSIONS: Observed variability among commercial ELISA platforms may compromise comparability between studies. Harmonisation, standardised reference materials, and cross-validation are necessary before renalase assays can be used reliably in clinical practice.

Humans↗

Metagenome-scale modeling to assess microbiome metabolic complementarity for precision microbiota transplantation therapies.

Fecal microbiota transplantation (FMT) holds therapeutic promise beyond recurrent Clostridioides difficile infection, but clinical outcomes remain unpredictable and donor-selection strategies remain limited, in part because the role of donor&#x2012;recipient metabolic interactions in shaping the post-FMT community remains poorly understood. Here, we leverage metagenome-scale metabolic modeling to quantify metabolic niche complementarity between donor and recipient microbiomes and predict post-FMT community composition. Using MICOM-derived metabolic models, we show that donor genomes whose metabolic flux profiles are more dissimilar from the recipient community colonize at significantly higher rates in a murine FMT model. In a human IBS trial, the same metric predicted post-FMT community composition via leave-one-out cross-validation and captured known disease-associated alterations in short-chain fatty acid, sulfur, and gas metabolism. We then performed 2,548 in silico FMT simulations between IBS-D/M patients and donors from the OpenBiome biobank to evaluate personalized donor screening, identifying super-donors characterized by high taxonomic diversity, broad metabolic niche coverage, and community interaction networks dominated by cross-feeding rather than competition. Together, these results support metabolic niche complementarity as a potential determinant of post-FMT community composition and provide a mechanistic basis for evaluating donor-recipient metabolic compatibility. This framework offers a scalable approach for generating testable hypotheses for personalized donor selection.

Fecal Microbiota Transplantation↗

RR-interval-based atrial fibrillation detection and burden estimation: cross-dataset validation and calibration-aware probability analysis.

Objective.Atrial fibrillation (AF) burden has become an increasingly important endpoint in long-duration rhythm monitoring, but reliable burden estimation requires more than accurate AF detection alone. In particular, when burden is derived by aggregating predicted AF probabilities over time, probability calibration may directly affect burden validity under external dataset shift.Approach.This study developed an interpretable-interval feature model for AF detection and evaluated it using record-wise cross-validation on a development cohort and independent cross-dataset external validation on public Holter electrocardiographic databases. Window-level performance was assessed using the area under the receiver operating characteristic curve (ROC-AUC), area under the precision-recall curve (PR-AUC), Brier score, expected calibration error (ECE), and calibration intercept and calibration slope. Recording-level AF burden was estimated using both probability-based and hard-label aggregation and evaluated using mean absolute error (MAE) and agreement analyses.Main results.The model showed high discrimination in both development and external evaluation, with external ROC-AUC ofand PR-AUC of. However, external calibration deteriorated despite preserved ranking performance, with Brier score of, ECE(15) of, calibration intercept of, and calibration slope of. In the external cohort, probability-based burden estimation preserved strong association with reference burden but showed weaker raw agreement than hard-label aggregation, with MAE ofversus, consistent with systematic probability underprediction. Repeated external recalibration across record-level splits substantially improved probability quality and probability-based burden estimation. Median probability-burden MAE decreased fromwithout recalibration toafter Platt recalibration andafter isotonic recalibration, while median ECE(15) decreased fromtoand, respectively.Significance.These findings indicate that-interval-based AF detection maintained strong ranking performance in the tested external cohort, but probability calibration should be evaluated explicitly when predicted probabilities are aggregated into AF-burden estimates.

Atrial Fibrillation↗

PLAID: ultrafast single-sample gene set enrichment scoring.

SUMMARY: In recent years, computational methods have emerged that calculate enrichment of gene signatures within individual samples. These signatures offer critical insights into the coordinated activity of functionally related genes, proteins or metabolites, enabling the identification of unique molecular profiles in individual cells and patients. This strategy is pivotal for patient stratification and advancement of personalized medicine. However, the rise of large-scale datasets, including single-cell profiles and population biobanks, has exposed significant computational inefficiencies in existing methods. Current methods often demand excessive runtime and memory resources, becoming impractical for large datasets. Overcoming these limitations is a focus of current efforts by bioinformatics teams in academia and the pharmaceutical industry, as essential to support basic and clinical biomedical research. To address this critical need, we developed PLAID (Pathway Level Average Intensity Detection), an ultrafast and memory optimized single sample gene set enrichment algorithm that utilizes sparse matrix computation. PLAID delivers highly accurate gene set scoring and surpasses the performance of current methods in single-cell and bulk transcriptomics, and proteomics data. PLAID uniquely integrates the most widely used gene set scoring algorithms, enabling researchers to apply multiple methods for cross-validation with outstanding runtime efficiency and minimal memory requirement. AVAILABILITY AND IMPLEMENTATION: PLAID is implemented in the R language for statistical computing. PLAID source code and installation instructions are available with no restrictions at https://github.com/bigomics/plaid.

Algorithms↗

MegaPlantTF: a machine learning framework for comprehensive identification and classification of plant transcription factors.

MOTIVATION: Understanding the role of transcription factors (TFs) in plants is essential for the study of gene regulation and various biological processes. However, both TF detection and classification remain challenging due to the great diversity and complexity of these proteins. Conventional approaches, such as BLAST, often suffer from high computational complexity and limited performance on less common TF families. RESULTS: We introduce MegaPlantTF, the first comprehensive machine learning and deep learning framework for the prediction (TF versus non-TF) and classification (family-level) of plant TFs. Our method employs k-mer-based protein representations and a two-stage architecture combining a deep feed-forward neural network with a stacking ensemble classifier. To ensure robust performance assessment, we report micro-, macro-, and weighted-average performance metrics, providing a holistic evaluation of both frequent and underrepresented TF families. Additionally, we employ threshold-based evaluation to calibrate confidence in TF detection. The results show that MegaPlantTF achieves strong accuracy and precision, particularly with a k-mer size of 3 and a classification threshold of 0.5, and maintains stable performance even under stringent thresholds. In addition to the standard cross-validation tests, a use case study on Sorghum bicolor confirms that our method performs strongly in the genome-wide analysis, making it highly suitable for large-scale TF identification and classification tasks. MegaPlantTF represents a novel contribution by integrating k-mer encoding, binary family-specific classifiers, and a two-stage stacking ensemble into a unified, reproducible framework for large-scale plant TF identification and classification. AVAILABILITY AND IMPLEMENTATION: MegaPlantTF is freely accessible through a public web server available at https://bioinformatics.um6p.ma/MegaPlantTF. The complete source code, including pretrained models and example datasets, is available at https://github.com/Bioinformatics-UM6P/MegaPlantTF.

Transcription Factors↗

Striping artifact removal in VisiumHD data through nuclear counts modeling.

MOTIVATION: 10x Genomics VisiumHD enables spatial transcriptomics at 2&#x2009;&#xb5;m &#xd7; 2&#x2009;&#xb5;m resolution but exhibits slide-specific, non-periodic striping artifacts due to lane-width variability. These multiplicative row/column effects distort bin total counts and can bias downstream analyses. The state-of-the-art destriping approach is the normalization procedure used as a preprocessing step in bin2cell; it applies sequential high-quantile row- then column-wise normalization, which is asymmetric and can introduce edge effects/macro-stripes and distortions of large-scale total-count structure. RESULTS: We propose a statistical destriping approach that leverages nuclei segmentation from the co-registered H&E image. Assuming transcript abundance is constant within each nucleus, we model bin counts with a negative binomial distribution whose mean is a product of a nucleus-specific concentration and row- and column-specific stripe-factors reflecting lane-width variation. We fit all parameters in a generalized linear modeling framework with cross-validated regularization on stripe-factors and iterative dispersion estimation, and use the fitted parameters to correct the observed counts into a destriped image. On synthetic data with known ground truth, our method improves stripe-factor estimation accuracy and reduces error in corrected counts relative to bin2cell and bin2cell-derived baselines. Across four public VisiumHD slides, it consistently lowers striping intensity while substantially better preserving biological signal present in the large-scale global count structure and avoiding the artifacts introduced by other methods. AVAILABILITY AND IMPLEMENTATION: All source code and links to publicly available data used for this study are available at https://github.com/paolamalsot/destriping-GLM.

Artifacts↗

ExoShorkie: predicting RNA-seq coverage of exogenous genomes in yeast by transfer learning.

MOTIVATION: Predicting the RNA-seq coverage of native and exogenous sequences is central to many molecular- and synthetic-biology applications. Substantial progress has been made in developing methods to predict the RNA-seq coverage of native genomic sequences, with the recently developed Shorkie achieving state-of-the-art performance in yeast. However, prediction performance of these methods over exogenous DNA is still unknown. Recent studies measured RNA-seq coverage of large exogenous genomes in yeast, providing a unique opportunity to train machine-learning models on a large exogenous sequence space and to improve both prediction performance and our understanding of regulatory mechanisms. RESULTS: We introduce ExoShorkie, a method we developed by extending Shorkie through transfer learning across multiple exogenous RNA-seq datasets. We demonstrate that ExoShorkie significantly improves prediction performance on held-out exogenous genomes and outperforms both a native-genome-trained Shorkie baseline and Yorzoi, the only competing method for predicting exogenous RNA-seq coverage in yeast, in cross-validation and in leave-one-genome-out evaluations. Furthermore, through interpretability analyses we reveal biologically meaningful regulatory motifs and distinct regulatory rules in exogenous genomes in yeast, providing new insights into transcriptional regulation. AVAILABILITY AND IMPLEMENTATION: ExoShorkie is available at https://github.com/OrensteinLab/ExoShorkie.

Genome, Fungal↗

GiantHost: a domain-adaptive and uncertainty-aware framework for giant virus host prediction.

MOTIVATION: Nucleocytoplasmic large DNA viruses (NCLDVs) play crucial roles in global ecosystems. Although metagenomics has vastly accelerated the discovery of novel NCLDVs, predicting their hosts from fragmented contigs remains a critical bottleneck, with no dedicated end-to-end computational tools currently available. Addressing this gap requires overcoming three fundamental challenges: the extreme scarcity of labeled reference genomes, the severe domain shift between laboratory isolates and diverse environmental metagenomes, and the inability of traditional deterministic models to quantify prediction uncertainty-a crucial requirement for reliable ecological profiling where novel, divergent viruses are prevalent. RESULTS: We present GiantHost, the first NCLDV host prediction tool with domain adaptation and uncertainlty awareness. GiantHost employs a dual-tower neural network to integrate dense genome traits and sparse GVOG profiles, allowing better integration of heterogeneous features. To overcome label scarcity and domain shift, we leverage 1400 environmental viral genomes (GVMAGs) via semi-supervised multi-task learning and Domain Adversarial Neural Networks (DANN), effectively bridging the distributional gap between RefSeq and environmental data. Additionally, GiantHost incorporates Conformal Prediction (CP) to output statistically guaranteed prediction sets rather than overconfident single labels. Evaluated under rigorous genome-level cross-validation, GiantHost demonstrates robust predictive power. Applied to the Tara Ocean dataset, GiantHost successfully captured the vertical stratification of NCLDV hosts-revealing a depth-dependent decline of phytoplankton-infecting viruses and a relative enrichment of Amoebozoa-infecting viruses in the mesopelagic zone. AVAILABILITY: The source code of GiantHost is available via: https://github.com/FuchuanQu/GiantHost.

Giant Viruses↗

abCRISPR: deep learning-based design of abasic gRNA sequences for specific CRISPR-Cas9 genome editing.

SUMMARY: CRISPR-Cas9 has become a widely used tool for genome editing. However, its off-target cleavage caused by partial sequence matches with guide RNAs (gRNAs) remains a critical limitation. Recently, abasic gRNAs (&#xd8;X&#xd8;) have been developed to enhance target specificity, but their effects vary depending on the positional sequence context. Here, we present abCRISPR, a deep neural network (DNN) framework for the rational design of &#xd8;X&#xd8; sequences with minimized off-target activity. abCRISPR leverages informative few-shot training with paired datasets of abasic and unmodified gRNAs, using high-quality random mismatch target libraries, exhaustively sequenced for mismatched off-target substrates (n&#x2009;=&#x2009;97583) in in vitro CRISPR-Cas9 cleavage experiments. Predicted off-target activities for both abasic and unmodified gRNAs showed strong correlation with experimental data (r&#x2009;&#x2265;&#x2009;0.95, 10-fold cross-validation). Notably, these comprehensive training sets provide robust ground-truth negatives, enabling accurate and sensitive prediction of off-targets. For unmodified gRNAs, abCRISPR (AUC = 0.98) was validated to outperform existing deep learning-based methods (AUC = 0.45-0.68). When applied to the human genome, abCRISPR generated &#xd8;X&#xd8; sequences, covering 58&#xa0;875&#xa0;004 potent CRISPR-targetable sites with improved target specificity. Together, this work provides a comprehensive bioinformatics resource for safe and precise CRISPR-Cas9 genome editing. AVAILABILITY AND IMPLEMENTATION: The source code for abCRISPR and training data are available at https://doi.org/10.5281/zenodo.20398246. abCRISPR results for the human genome are available at http://clip.korea.ac.kr/abCRISPR/.

Deep Learning↗

Metabolomics Reveals Metabolic Characteristics of Functional Cure in Chronic Hepatitis B Treated With Entecavir Combined With Pegylated Interferon Alpha.

BACKGROUND: Entecavir (ETV) combined with pegylated interferon alpha (PEG-IFN&#x3b1;) improves chronic hepatitis B (CHB) functional cure rates, but therapeutic heterogeneity and underlying metabolic mechanisms remain unclear. This study used untargeted metabolomics to identify metabolic signatures, mechanisms, and predictive biomarkers of functional cure with ETV-PEG-IFN&#x3b1;. METHODS: Thirty-eight CHB patients were grouped into ETV monotherapy (Group E, n = 12) and ETV-PEG-IFN&#x3b1; combination therapy (Group Z, n = 26); Group Z was subdivided into cured (Group A, n = 13) and noncured (Group B, n = 13). Serum metabolomic profiling, multivariate statistics, and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analysis identified differential metabolites. A random forest model was built using key metabolites. RESULTS: Three hundred eighty-eight metabolites were identified. Four differential metabolites distinguished Group A and B (upregulated guanidinoacetic acid, uracil 5-carboxylate; downregulated L-methionine S-oxide, oleamide), enriching amino acid metabolism pathways. Nine differential metabolites between Group E and Z implicated amino acid, immune, and fatty acid pathways. The random forest model based on the four Group A/B metabolites showed 88.5% cross-validation accuracy (AUC = 0.920), with L-methionine S-oxide and oleamide as key predictors. CONCLUSIONS: This study reveals metabolic rewiring in CHB functional cure via ETV-PEG-IFN&#x3b1; therapy, involving energy metabolism, oxidative stress, and immunomodulation, based on which we propose a tentative metabolism-immunity synergy model to guide future research. Key metabolites, especially L-methionine S-oxide and oleamide, show exploratory predictive potential for functional cure that warrants further validation in independent cohorts.

Humans↗

Early Transcriptional Changes in Neutrophil-Mediated Processes Following Recanalization After Ischemic Stroke.

BACKGROUND: Ischemic stroke is a leading cause of death and long-term disability worldwide. Recanalization therapies, including thrombolysis and mechanical thrombectomy, restore blood flow, yet many patients experience poor outcomes, a phenomenon known as futile recanalization. Given the short therapeutic window for ischemic stroke, identifying early biomarkers to guide targeted interventions and improve outcomes is critical. METHODS: Using a murine middle cerebral occlusion model that mimics a large vessel occlusion with recanalization, a comprehensive microarray analysis from blood samples collected immediately and 3&#x2009;hours after recanalization (N=44) was performed. Differentially expressed genes, enrichment pathways, immune cell proportions, enriched cell markers, predicted micro-RNAs, and transcription factors were identified using RStudio. Findings in mice were validated with rat middle cerebral artery occlusion (GSE21136) and patients with stroke (GSE16561) data sets to confirm transcriptional changes in peripheral blood postrecanalization. RESULTS: Il1r2, Cd55, Mmp8, Cd14, and Cd69 were early biomarkers poststroke and postrecanalization. Cross-validation revealed Vcan as a differentially expressed gene conserved across species, making it a novel ischemic marker detected as early as 3&#x2009;hours postrecanalization (4&#x2009;hours after middle cerebral artery occlusion) in mice, 24&#x2009;hours after recanalization in rats (middle cerebral artery occlusion-thrombectomy), and within 24&#x2009;hours from onset in humans receiving recombinant tissue plasminogen activator-thrombolysis. CIBERSORTx and ImmuCellAI-mouse deconvolution showed neutrophil elevation postrecanalization. Leukocyte and neutrophil activation pathways were enriched early after stroke in mice and humans, with stronger upregulation in the female sex. Several regulatory micro-RNAs were identified, and Nuclear Factor Erythroid 4 (NFE4)&#xa0;and Metal Regulatory Transcription Factor 1 (MTF1) emerged as key transcription factors. A coregulatory network underlying neutrophil activity was constructed, highlighting its central role in early responses to ischemia and recanalization, which was enriched in the female sex. CONCLUSIONS: We identified novel early genomic markers for ischemia and recanalization, including the conserved marker Vcan, and highlighted age- and sex-specific immune responses. Mapping a neutrophil-centered coregulatory network provides mechanistic insight into futile recanalization and supports the development of targeted therapies to improve clinical outcomes.

Animals↗

CAGNet: a structure-aware clustering-alternated graph network for cell-cell interaction inference in spatial transcriptomics.

MOTIVATION: Understanding cell-cell interactions (CCIs) in spatial transcriptomics is crucial for uncovering the spatial organization and functional heterogeneity of tissues. However, existing graph-based models typically rely on static clustering or fixed adjacency structures, which limits their ability to capture dynamic cellular relationships. RESULTS: We propose CAGNet, a two-stage framework for CCI inference from spatial transcriptomics data. In Stage 1, a Graph Attention Network encoder with joint feature and graph reconstruction learns structure-aware node embeddings from spatial gene expression profiles. In Stage 2, an alternating optimization mechanism iteratively updates cluster centers via KL-guided soft assignment and refines node embeddings through spatial graph reconstruction, establishing a closed-loop between representation learning and clustering. Experiments on three 10x Genomics Visium datasets demonstrate that CAGNet consistently outperforms six CCI inference baselines across ACC, AUC, AP, Precision, Recall, and F1. CAGNet also achieves the highest Adjusted Rand Index on all three datasets against six spatial domain identification methods, confirming that the learned embeddings capture biologically relevant spatial organization. Information-theoretic analysis further shows that CAGNet retains the highest mutual information between input features and learned embeddings among all compared methods. Ablation studies and 5-fold cross-validation confirm the contribution of each component and the reproducibility of the results. AVAILABILITY: The proposed method is implemented in the CAGNet package available at http://github.com/mahan1233333-maker/CAGNet .

Spatial Transcriptomics↗

Comprehensive in silico genomics analysis of global trends and host-specific emergence of aminoglycoside resistance in Staphylococcus aureus: a One-Health perspective.

BACKGROUND: Aminoglycosides remain clinically valuable against Staphylococcus aureus. Aminoglycoside resistance in S. aureus represents a critical One Health concern and is primarily driven by aminoglycoside-modifying enzymes (AMEs), which are frequently plasmid-encoded. Although regional studies have provided valuable insights, the global epidemiology of aminoglycoside resistance determinants remains poorly characterized because comprehensive data integrating human, animal, and environmental reservoirs are still lacking. This study addresses this gap by analyzing over 110,000 S. aureus genomes (2000-2025) to map the global resistome, quantify temporal and host-specific trends, and assess the association between genetic determinants and phenotypic resistance. METHODS: We performed a retrospective One Health meta-analysis of 110,309 S. aureus genomes collected between 2000 and 2025 from 128 countries. Genomes were quality-filtered and aminoglycoside resistance determinants were identified using NCBI AMRFinderPlus (v4.0.23). Multilocus sequence typing and host-source harmonization (Human, Animal, Environment, Unknown) enabled clonal and reservoir stratification. Temporal trends in gene prevalence and resistance burden were modeled with robust regression. Geographic and host-associated structuring of key genes was assessed via &#x3c7;2 and enrichment tests. Machine-learning models (elastic-net, random forests, XGBoost) were benchmarked for minimum inhibitory concentration (MIC) prediction via nested cross-validation, with performance evaluated by mean absolute error, RMSE, and SHAP-based feature importance. All analyses were conducted in R and Python using publicly available, de-identified genomic data. RESULTS: Aminoglycoside resistance-associated genes were dominated by modifying enzyme determinants, with ant(6)-Ia, ant(9)-Ia, aph(3')-IIIa, sat4, aadD1, and aac(6')-Ie/aph(2'')-Ia occurring in 14-22% of isolates worldwide. Temporal analysis revealed significant declines in several major determinants, most notably ant(9)-Ia (-2.22 percentage points per year, p&#x2009;<&#x2009;0.001), whereas apmA exhibited a non-significant decreasing trend in animal isolates. Host structuring was marked: human clinical isolates concentrated common determinants, while animal and environmental isolates harbored rare alleles (apmA, spw, str, spd). Geographic mapping confirmed near-universal distribution of common genes but focal restriction of rare ones. Publicly available phenotypic data indicated strong activity of amikacin, whereas gentamicin showed a distinct resistant subpopulation that closely corresponded with AME gene carriage. Genotype-phenotype analyses demonstrated strong concordance, with gene-rich complements predicting resistant MIC strata and absence of determinants predicting susceptibility. Analysis across different gene classes revealed frequent co-occurrence of aminoglycoside resistance genes with determinants from other classes, such as mecA, blaZ, and MLS_B, embedding them within multidrug-resistant (MDR) genomic contexts. CONCLUSION: Over 25&#xa0;years, the prevalence of aminoglycoside resistance-associated genes in S. aureus has declined for several common determinants, while rare veterinary-linked alleles are emerging in animal isolates. Strong genotype-phenotype concordance supports genomic prediction for gentamicin and amikacin, where MIC data are available, although phenotypic confirmation remains essential. The frequent co-occurrence of aminoglycoside resistance genes with other antimicrobial resistance determinants indicates their integration within co-occurrence patterns of MDR genes, defined here as clusters of co-occurring resistance genes often carried on shared mobile genetic elements. These patterns highlight the need for integrated One Health surveillance combining clinical, veterinary, and environmental monitoring with plasmid-context resolution to anticipate emerging threats.

Aminoglycosides↗

Plasma inflammatory proteome profiles identify MASLD among children with overweight or obesity.

BACKGROUND & AIMS: Pediatric metabolic dysfunction-associated steatotic liver disease (MASLD) is increasingly prevalent among children with overweight or obesity, yet its early diagnosis remains a major clinical challenge. This study aimed to identify circulating inflammatory proteins associated with MASLD and to develop a proteomic risk score (ProScore) to improve diagnostic accuracy. METHODS: In this cross-sectional study of 161 children (median age 8.5&#xa0;years) with overweight or obesity, MASLD was assessed by vibration-controlled transient elastography, with 42 cases identified. Plasma concentrations of 92 inflammation-related proteins were quantified using a high-throughput proximity extension assay. The ProScore was compared with eleven conventional anthropometric/metabolic indices (WHtR, METS-IR, SPISE, PNFI, VAI, LAP, TyG, TyG-ALT, TyG-WC, TyG-WHtR, and TyG-BMI) and a genetic risk score (GRS). Six machine learning algorithms were employed and diagnostic performance was assessed using area under the curve (AUC) with fivefold cross-validation. RESULTS: Fifteen proteins were significantly associated with MASLD. A six-protein panel (FGF-21, CDCP1, CD244, OPG, Flt3L, MCP-1) achieved the highest diagnostic accuracy (AUC&#x2009;=&#x2009;0.84), exceeding that of all conventional indices (AUC&#x2009;=&#x2009;0.65-0.78; all P&#x2009;<&#x2009;0.05). ProScore performance remained robust in school-based validation (AUC&#x2009;=&#x2009;0.83), with no substantial improvement when combined with conventional indices. Diagnostic accuracy was higher in children with lower GRS (AUC&#x2009;=&#x2009;0.92) than in those with higher GRS (AUC&#x2009;=&#x2009;0.80; P&#x2009;=&#x2009;0.003). CONCLUSIONS: A proteomic signature of systemic inflammation provides accurate, non-invasive identification of MASLD in at-risk children, outperforming conventional metabolic and genetic tools, and may have utility in clinical and public health settings.

Humans↗

Machine learning-based clinical prediction model and multi-omics integration for assessing pancreatic cancer risk in new-onset diabetes.

BACKGROUND: Given that pancreatic cancer (PC) is typically diagnosed at an advanced stage but is often preceded by new-onset diabetes mellitus (NODM), providing a window for early detection, we sought to develop and validate an interpretable machine-learning model integrated with multi-omics profiling to identify early biomarkers of NODM-associated PC. METHODS: In a population-based cohort, individuals with NODM-associated PC and NODM without PC were identified and randomly divided (70:30) into training and validation sets after feature selection. Eight machine learning (ML) classifiers were compared using fivefold cross-validation, and model performance was evaluated in terms of discrimination, calibration, and decision curve&#x2013;based clinical utility. We evaluated interpretability using the Shapley additive explanations (SHAP) analyses. Mechanistically, Olink proteomic profiling and metabolomics were analyzed through clinical classifications and model-defined risk strata. RESULTS: Categorical boosting achieved the best performance in the independent validation set (AUROC&#x2009;=&#x2009;0.844). The NODM cohort was stratified into high- (n&#x2009;=&#x2009;2,362) and low-risk (n&#x2009;=&#x2009;5,030) groups, and internal validation together with SHAP analyses demonstrated consistent model performance and identified clinically interpretable predictors. Proteomic and metabolomic analyses under clinical and risk-based grouping identified 39 overlapping differentially expressed proteins and 145 overlapping metabolites with enriched across 11 shared KEGG pathways. Cross-platform validation highlighted PLTP, CRTAC1, and ITGAV as serum biomarkers with a strong potential for early NODM-PC detection. CONCLUSIONS: We developed an interpretable ML framework centered on NODM enables practical risk stratification for early PC detection by multi-omics and provides a pathway of ML-based triage followed by biomarker confirmation for earlier detection and diagnosis.

Humans↗

Genetics-Informed Mapping Identifies a CRIM1-Associated Endocardial Inflammatory Remodeling State in Acute Myocardial Infarction.

BACKGROUND Acute myocardial infarction (AMI) reflects inherited susceptibility and inflammatory remodeling, but the cellular contexts linking genetic risk to disease remain unclear. MATERIAL AND METHODS We integrated a meta-transcriptome-wide association study (TWAS) with a human cardiac single-nucleus RNA-sequencing atlas contained 11 individuals (5 AMI and 6 donor) to identify genetics-informed cellular programs. Composite program states were defined by global score quartiles. A fixed 5-gene panel was evaluated for nucleus-level endocardial low-transcriptional-state (Endo_LTS) vs endocardial high-transcriptional-state (Endo_HTS) discrimination within the AMI endocardium using 5-fold leave-1-patient-out cross-validation. Functional follow-up used CRIM1 silencing in hypoxia-treated human induced pluripotent stem cell (hiPSC)-derived endocardial endothelial-like cells and complementary peripheral blood analyses. RESULTS The endocardium exhibited the most prominent infarction-associated increase in TWAS-anchored program activity, with expansion of program-high states and higher CytoTRACE scores. A consensus 5-gene panel (RPS8, PLEC, CFDP1, CRIM1, TNS2) was identified. Among 2163 AMI endocardial nuclei from 5 patients, the state classifier included 364 Endo_LTS and 751 Endo_HTS nuclei; 1048 Endo_MTS nuclei were excluded. Pooled out-of-fold ROC-AUCs ranged from 0.665 to 0.831. The panel also showed discriminatory value in an independent peripheral-blood AMI-vs-control cohort. CRIM1 was prioritized as a candidate linked to the remodeling program. CRIM1 silencing attenuated ACTA2/alpha-SMA, vimentin, LDHA, CCL2, and VEGFA and partially restored CD31, whereas TGF-&#xdf; remained elevated. CONCLUSIONS These findings identify a genetics-informed endocardial inflammatory remodeling state in AMI and define a 5-gene surrogate of its activated state. CRIM1 is prioritized as a candidate linked to selected inflammatory, metabolic, and structural outputs. Persistent TGF-b elevation after CRIM1 silencing argues against a simple linear regulatory model and indicates that further mechanistic validation is required.

Humans↗

Automated Classification of Lymphoma Subtypes From Histopathological Images Using a U-Net Deep Learning Model: Comparative Evaluation Study.

BACKGROUND: Accurate classification and grading of lymphoma subtypes are essential for treatment planning. Traditional diagnostic methods face challenges of subjectivity and inefficiency, highlighting the need for automated solutions based on deep learning techniques. OBJECTIVE: This study aimed to investigate the application of deep learning technology, specifically the U-Net model, in classifying and grading lymphoma subtypes to enhance diagnostic precision and efficiency. METHODS: In this study, the U-Net model was used as the primary tool for image segmentation integrated with attention mechanisms and residual networks for feature extraction and classification. A total of 620 high-quality histopathological images representing 3 major lymphoma subtypes were collected from The Cancer Genome Atlas and the Cancer Imaging Archive. All images underwent standardized preprocessing, including Gaussian filtering for noise reduction, histogram equalization, and normalization. Data augmentation techniques such as rotation, flipping, and scaling were applied to improve the model's generalization capability. The dataset was divided into training (70%), validation (15%), and test (15%) subsets. Five-fold cross-validation was used to assess model robustness. Performance was benchmarked against mainstream convolutional neural network architectures, including fully convolutional network, SegNet, and DeepLabv3+. RESULTS: The U-Net model achieved high segmentation accuracy, effectively delineating lesion regions and improving the quality of input for classification and grading. The incorporation of attention mechanisms further improved the model's ability to extract key features, whereas the residual structure of the residual network enhanced classification accuracy for complex images. In the test set (N=1250), the proposed fusion model achieved an accuracy of 92% (1150/1250), a sensitivity of 91.04% (1138/1250), a specificity of 89.04% (1113/1250), and an F1-score of 90% (1125/1250) for the classification of the 3 lymphoma subtypes, with an area under the receiver operating characteristic curve of 0.95 (95% CI 0.93-0.97). The high sensitivity and specificity of the model indicate strong clinical applicability, particularly as an assistive diagnostic tool. CONCLUSIONS: Deep learning techniques based on the U-Net architecture offer considerable advantages in the automated classification and grading of lymphoma subtypes. The proposed model significantly improved diagnostic accuracy and accelerated pathological evaluation, providing efficient and precise support for clinical decision-making. Future work may focus on enhancing model robustness through integration with advanced algorithms and validating performance across multicenter clinical datasets. The model also holds promise for deployment in digital pathology platforms and artificial intelligence-assisted diagnostic workflows, improving screening efficiency and promoting consistency in pathological classification.

Humans↗

Development and Validation of Machine Learning Models for Predicting Early Cognitive Decline Using Home Sensor-Derived Behavioral Data: Sensors in-Home for Elder Wellbeing (SINEW) Cohort Study.

BACKGROUND: As the global population continues to age, the prevalence of geriatric conditions, including dementia and frailty, is also increasing. Early identification of individuals at an elevated risk of these conditions, such as those presenting with mild cognitive impairment (MCI) or prefrailty, can provide a critical window for prompt intervention aimed at preventing or reversing disease progression. To promote such early identification, there is a burgeoning interest in the use of digital sensor technology and predictive modeling. OBJECTIVE: This study aimed to use a continuous, home-based monitoring sensor system for older adults to distinguish those exhibiting normal aging from those with MCI, early dementia, prefrailty, or frailty, and to predict their transition from normal aging to one of these conditions. METHODS: This longitudinal cohort study will recruit 200 community-dwelling adults aged &#x2265;65 years with normal cognition or MCI at baseline. A multi-sensor system will be installed in participants' homes, including passive infrared motion sensors, door contact sensors, bed sensors, medication box sensors, wearable activity bands, and Bluetooth proximity beacons. These devices will continuously capture spatiotemporal activity patterns, mobility indicators, sleep behaviors, and medication-taking routines. Annual assessments will include standardized cognitive tests (eg, Montreal Cognitive Assessment, Mini-Mental State Examination, Rey Auditory-Verbal Learning Test, digit span, Color Trails Test, semantic fluency, Stroop), frailty measures (modified Fried phenotype, gait speed, grip strength), mental health scales, sleep quality, and psychosocial indicators. Sensor-derived features-such as gait variability, activity regularity, sleep fragmentation, and medication adherence patterns-will be integrated with clinical data to develop supervised machine learning models. Planned approaches include logistic regression, random forests, gradient boosting, and deep learning. Model performance will be evaluated using cross-validation and independent test sets. Primary metrics will include area under the receiver operating characteristic curve, sensitivity, specificity, precision, recall, and F1-score. Models will be benchmarked against gold-standard clinical diagnoses and validated using temporal subsets of the dataset. RESULTS: Enrollment for this study started in November 2019 and will continue until March 2030. As of June 2025, we have enrolled 138 participants. Full data analysis has yet to begin. CONCLUSIONS: We aim to develop a reliable and effective sensor system for in-home use that will facilitate the early detection of cognitive and physical decline. In so doing, it will add to our current understanding of digital biomarkers. It is common for older adults to seek clinical intervention only when their cognitive impairment has already reached an advanced stage. The implementation of readily deployable sensor systems within community settings presents us with opportunities for prompt intervention, which holds the potential for delaying or reversing disease progression and allowing for a greater number of functional and meaningful years.

Humans↗