PubMed HealthSearch

SEARCH · PubMed Health

Results for “Repeated nested cross-validation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

4 recordsLinked to original sources

Machine learning-based analysis of oral rinse samples to identify candidate proteomic signatures for severe periodontitis: a pilot study.

This pilot study investigated whether candidate protein signatures from oral rinse samples can distinguish patients with severe periodontitis (stage III/IV) and its subtypes, generalized and localized periodontitis, from non-periodontitis controls. Participants rinsed with phosphate-buffered saline, and samples were analyzed using a Proximity Extension Assay targeting 92 inflammatory and 92 immuno-oncology proteins. A machine learning approach using repeated nested cross-validation and SHAP was implemented to identify protein signatures. The study included 38 patients (18 with localized periodontitis and 20 with generalized periodontitis) and 16 controls. After data preprocessing, 54 samples and 141 proteins were retained. Proteins Gal-1, HGF, TNFSF14, CD27, and ARG1 distinguished periodontitis from controls (ROC-AUC = 0.85, 95% CI 0.82, 0.87). For generalized periodontitis, we found a protein signature including TNFSF14, Gal-1, STAMBP, MUC-16, S100A12, HGF, CASP-8, CD27, LAP TGF-β1, TNFRSF9, and uPA (ROC-AUC = 0.92, 95% CI 0.90, 0.94). For localized periodontitis, we identified ARG1 (ROC-AUC = 0.72, 95% CI 0.68, 0.76). No proteomic signature distinguishing generalized periodontitis from localized periodontitis was identified. This pilot study indicated that oral rinses are suitable for proteomic profiling, and there was a putative protein signature that could differentiate periodontitis, generalized periodontitis, and localized periodontitis from controls. These findings warrant validation in larger independent cohorts, including a clearly defined gingivitis group, before real-world non-invasive screening applications can be considered.

Humans

Rapid glycomic analysis of serum EVs reveals altered N-glycosylation patterns in ASD.

Objective laboratory diagnostics for autism spectrum disorder (ASD) are lacking, necessitating rapid clinical screening tools. Because serum extracellular vesicle (EV) N-glycosylation captures critical neurodevelopmental signatures, we developed a fast, biologically interpretable diagnostic strategy. EVs from ASD patients with language impairment and neurotypical controls were isolated using a rapid extra-polyethylene glycol precipitation/filtration (EPF) workflow, benchmarked against ultracentrifugation. Following MALDI-TOF/MS profiling, machine learning was re-evaluated using repeated nested cross-validation to reduce optimistic bias and potential information leakage. Among five classifiers, Random Forest (RF) showed the best overall balance across discrimination, calibration, and classification metrics. RF-based SHAP analysis provided transparent interpretation, highlighting key discriminative glycans, including H4N3S1F1, H5N5S1F1, and H3N5F1. To elucidate molecular mechanisms, we integrated public EV transcriptomic data. This revealed significant dysregulation of N-glycosylation machinery genes (e.g., MAN1A1, NEU1, OSTC, RPN2), whose expression directionally aligned with observed glycan shifts in synaptic pathways. Collectively, this rapid serum EV N-glycomic workflow, combined with leakage-controlled RF-based interpretation, provides a promising foundation for non-invasive ASD biomarker discovery and future multicenter validation.

Humans

Cross-Platform Proteomics and Machine Learning Algorithms Nominate Plasma Biomarkers of Stroke Diagnosis.

BACKGROUND: Blood-based biomarkers for stroke subtyping could improve triage in emergency settings. We used cross-platform proteomics to identify plasma biomarkers differentiating major stroke diagnostic groups. METHODS: We conducted a case-control study using 2 biorepositories. Plasma was collected in the emergency department from adults with suspected stroke before therapeutic intervention. Differentially enriched proteins were identified across acute ischemic stroke, intracerebral hemorrhage, transient ischemic attack, and stroke mimics using SomaScan discovery proteomics (Grady). Differentially enriched proteins were nominated using pairwise and multigroup comparisons and adjusted for clinical covariates. Protein panels were created using least absolute shrinkage and selection operator logistic regression. Internal validation used repeated nested cross-validation (rCV) and targeted mass spectrometry (MS), while external validation used data-independent acquisition  mass spectrometry in an independent cohort (Yale). RESULTS: We included 100 subjects (40 with acute ischemic stroke, 20 with intracerebral hemorrhage, 20 with transient ischemic attack, 20 with stroke mimics) in discovery and 80 subjects (20 per group) in external validation cohorts. SomaScan quantified 7307 proteins, of which 61 differentiated stroke subtypes. We identified 7 protein classifiers for acute ischemic stroke (rCV-area under the curve, 0.82 [95% CI, 0.78-0.86]), 6 for intracerebral hemorrhage (rCV-area under the curve, 0.70 [95% CI, 0.64-0.76]), 8 for transient ischemic attack (rCV-area under the curve, 0.78 [95% CI, 0.73-0.84]), and 7 for stroke mimics (rCV-area under the curve, 0.81 [95% CI, 0.77-0.86]). Targeted proteomics internally validated 11 proteins, and data-independent acquisition-mass spectrometry externally validated 32 proteins, including VTN (vitronectin), PLG (plasminogen), and S100A9 as top stroke mimics, transient ischemic attack, and intracerebral hemorrhage classifiers. CONCLUSIONS: This study highlights plasma proteomics as a valuable tool for discovering protein biomarkers of stroke diagnosis. These findings support further validation in larger, multicenter cohorts to facilitate biomarker-guided stroke diagnosis in acute care.

Humans

A three-metabolite microbiota-associated signature for early risk stratification of gestational diabetes mellitus.

BACKGROUND: Gestational diabetes mellitus (GDM) is associated with adverse pregnancy outcomes and long-term metabolic and cardiovascular risk. However, oral glucose tolerance testing at 24-28 gestational weeks limits early risk stratification. Gut microbiota-associated metabolites may reflect early metabolic abnormalities, including those relevant to cardiometabolic health, but robust early-pregnancy biomarkers remain limited. METHODS: We conducted a multicenter nested case-control and prospective study involving 2,693 pregnant women. Untargeted metabolomics and metagenomics were integrated to identify GDM-associated metabolites and gut microbial alterations. Three consistently dysregulated metabolites, 3-hydroxydecanoic acid, γ-Glu-Leu, and propionic acid, were quantified by targeted LC-MS/MS. Candidate algorithms were compared using repeated 10-fold cross-validation, and a final generalized linear model was externally and prospectively validated. RESULTS: Women who later developed GDM showed an adverse early-pregnancy metabolic profile, including higher BMI, triglycerides, and platelet count. Untargeted metabolomics identified 14 persistently altered metabolites enriched in energy, oxidative stress, and amino acid metabolism pathways. Metagenomics revealed taxonomic restructuring and coordinated microbiota-metabolite associations. The three-metabolite model achieved AUCs of 0.838 (95% CI, 0.791-0.885) in training, 0.840 (95% CI, 0.769-0.911) in internal validation, 0.955 (95% CI, 0.925-0.985) and 0.917 (95% CI, 0.875-0.958) in two external cohorts, and 0.969 (95% CI, 0.937-1.000) in the prospective cohort. CONCLUSION: Early microbiota-associated metabolic dysregulation is detectable before routine GDM diagnosis. This compact three-metabolite panel may support early GDM risk stratification and provides metabolic evidence relevant to broader cardiometabolic risk assessment in pregnancy.

Humans