PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “cross-validation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Multivariate assessment of conflict in distressed and nondistressed mother-adolescent dyads.

A battery of measures was used to assess conflict between mothers and young adolescents (females and males, 11 to 15 years of age). Two groups of families, one composed of a distressed clinical sample (N = 38), the other a nondistressed normative sample (N = 40), participated. The assessment battery included retrospective judgments, frequency estimates, self-monitored home recording, and tape-recorded discussion of a home problem. Content of assessment measures tapped aspects of parental control, decision-making, self-reported interaction behavior, arguments, interaction behavior rated by independent "blind" observers, frequency and anger-intensity of specific problematic issues, and perceptions of positive and negative behaviors of the other family member. Based on univariate analyses, 21 of the 26 defined variables discriminated significantly in the predicted direction. Maternal and adolescent reports of behavior and independent ratings of tape-recorded interaction emerged as strong and consistent discriminators. Stepwise multivariate discriminant analysis provided successful classification of 100% of the families based on the inclusion of nine variables. In a cross-validation sample, 84% of the families were correctly classified. Implications for systematic outcome research as well as clinical application are discussed.

Adolescent↗

Identification of candidate variants in plasma associated with early versus late disease progression under anti-PD-1 therapy in metastatic NSCLC.

BACKGROUND: Immune checkpoint inhibitors (ICIs), including anti-programmed cell death protein 1 (anti-PD-1) antibodies, have significantly improved outcomes in patients with metastatic non-small cell lung cancer (mNSCLC). However, substantial heterogeneity exists in clinical benefit, with some patients exhibiting early progression (EP) and others late progression (LP). To date, no biomarkers of EP versus LP disease have been implemented in clinical practice. Circulating tumor DNA (ctDNA) analysis represents a minimally invasive strategy for identifying such biomarkers. In this proof-of-concept study, we evaluated the performance of the TruSight Oncology 500 ctDNA (TSO500 ctDNA) panel and explored its feasibility to identify candidate variants associated with early and late disease progression under anti-PD-1 therapy. METHODS: Baseline ctDNA from eight mNSCLC patients treated with pembrolizumab was extracted and sequenced using the TSO500 ctDNA assay, a 523-gene targeted next-generation sequencing panel. Patients were classified according to their response as LP or EP. Variant calling was performed using the DRAGEN Bio-IT platform, and variants were annotated and clinically interpreted using the Clinical Genomics Workspace (CGW; PierianDx) according to Association for Molecular Pathology (AMP)/American Society of Clinical Oncology (ASCO)/College of American Pathologists (CAP) guidelines. Survival outcomes were assessed using Kaplan-Meier and log-rank tests. Performance of ctDNA variants was evaluated using receiver operating characteristic (ROC) curve analysis, and multi-gene models were assessed using leave-one-out cross-validation with penalized logistic regression. RESULTS: All patients harbored detectable variants, including SNVs (100%), MNVs (87.5%), deletions (75%), and insertions (62.5%). Tier I variants were identified in 37.5% of patients, while all cases showed tier II and multiple tier III alterations. TP53 variants were associated with poorer outcomes under anti-PD-1 therapy. Individual gene alterations in TP53, ERBB3, SMC1A or LATS1 showed moderate discriminatory performance between LP and EP patients; however, combination of mutated genes improved apparent discrimination. Notably, specific two-gene combinations (SMC1A + LATS1 or ERBB3 + LATS1) showed the highest discriminatory performance between LP and EP patients in this exploratory cohort. CONCLUSIONS: This study demonstrates the feasibility and analytical performance of the TSO500 ctDNA panel and provides hypothesis-generating evidence that plasma gene variants may be useful to evaluate early versus late disease progression in patients with mNSCLC receiving immunotherapy.

TruSight Oncology 500↗

Automated Classification of Lymphoma Subtypes From Histopathological Images Using a U-Net Deep Learning Model: Comparative Evaluation Study.

BACKGROUND: Accurate classification and grading of lymphoma subtypes are essential for treatment planning. Traditional diagnostic methods face challenges of subjectivity and inefficiency, highlighting the need for automated solutions based on deep learning techniques. OBJECTIVE: This study aimed to investigate the application of deep learning technology, specifically the U-Net model, in classifying and grading lymphoma subtypes to enhance diagnostic precision and efficiency. METHODS: In this study, the U-Net model was used as the primary tool for image segmentation integrated with attention mechanisms and residual networks for feature extraction and classification. A total of 620 high-quality histopathological images representing 3 major lymphoma subtypes were collected from The Cancer Genome Atlas and the Cancer Imaging Archive. All images underwent standardized preprocessing, including Gaussian filtering for noise reduction, histogram equalization, and normalization. Data augmentation techniques such as rotation, flipping, and scaling were applied to improve the model's generalization capability. The dataset was divided into training (70%), validation (15%), and test (15%) subsets. Five-fold cross-validation was used to assess model robustness. Performance was benchmarked against mainstream convolutional neural network architectures, including fully convolutional network, SegNet, and DeepLabv3+. RESULTS: The U-Net model achieved high segmentation accuracy, effectively delineating lesion regions and improving the quality of input for classification and grading. The incorporation of attention mechanisms further improved the model's ability to extract key features, whereas the residual structure of the residual network enhanced classification accuracy for complex images. In the test set (N=1250), the proposed fusion model achieved an accuracy of 92% (1150/1250), a sensitivity of 91.04% (1138/1250), a specificity of 89.04% (1113/1250), and an F1-score of 90% (1125/1250) for the classification of the 3 lymphoma subtypes, with an area under the receiver operating characteristic curve of 0.95 (95% CI 0.93-0.97). The high sensitivity and specificity of the model indicate strong clinical applicability, particularly as an assistive diagnostic tool. CONCLUSIONS: Deep learning techniques based on the U-Net architecture offer considerable advantages in the automated classification and grading of lymphoma subtypes. The proposed model significantly improved diagnostic accuracy and accelerated pathological evaluation, providing efficient and precise support for clinical decision-making. Future work may focus on enhancing model robustness through integration with advanced algorithms and validating performance across multicenter clinical datasets. The model also holds promise for deployment in digital pathology platforms and artificial intelligence-assisted diagnostic workflows, improving screening efficiency and promoting consistency in pathological classification.

Humans↗

Development and Validation of Machine Learning Models for Predicting Early Cognitive Decline Using Home Sensor-Derived Behavioral Data: Sensors in-Home for Elder Wellbeing (SINEW) Cohort Study.

BACKGROUND: As the global population continues to age, the prevalence of geriatric conditions, including dementia and frailty, is also increasing. Early identification of individuals at an elevated risk of these conditions, such as those presenting with mild cognitive impairment (MCI) or prefrailty, can provide a critical window for prompt intervention aimed at preventing or reversing disease progression. To promote such early identification, there is a burgeoning interest in the use of digital sensor technology and predictive modeling. OBJECTIVE: This study aimed to use a continuous, home-based monitoring sensor system for older adults to distinguish those exhibiting normal aging from those with MCI, early dementia, prefrailty, or frailty, and to predict their transition from normal aging to one of these conditions. METHODS: This longitudinal cohort study will recruit 200 community-dwelling adults aged ≥65 years with normal cognition or MCI at baseline. A multi-sensor system will be installed in participants' homes, including passive infrared motion sensors, door contact sensors, bed sensors, medication box sensors, wearable activity bands, and Bluetooth proximity beacons. These devices will continuously capture spatiotemporal activity patterns, mobility indicators, sleep behaviors, and medication-taking routines. Annual assessments will include standardized cognitive tests (eg, Montreal Cognitive Assessment, Mini-Mental State Examination, Rey Auditory-Verbal Learning Test, digit span, Color Trails Test, semantic fluency, Stroop), frailty measures (modified Fried phenotype, gait speed, grip strength), mental health scales, sleep quality, and psychosocial indicators. Sensor-derived features-such as gait variability, activity regularity, sleep fragmentation, and medication adherence patterns-will be integrated with clinical data to develop supervised machine learning models. Planned approaches include logistic regression, random forests, gradient boosting, and deep learning. Model performance will be evaluated using cross-validation and independent test sets. Primary metrics will include area under the receiver operating characteristic curve, sensitivity, specificity, precision, recall, and F1-score. Models will be benchmarked against gold-standard clinical diagnoses and validated using temporal subsets of the dataset. RESULTS: Enrollment for this study started in November 2019 and will continue until March 2030. As of June 2025, we have enrolled 138 participants. Full data analysis has yet to begin. CONCLUSIONS: We aim to develop a reliable and effective sensor system for in-home use that will facilitate the early detection of cognitive and physical decline. In so doing, it will add to our current understanding of digital biomarkers. It is common for older adults to seek clinical intervention only when their cognitive impairment has already reached an advanced stage. The implementation of readily deployable sensor systems within community settings presents us with opportunities for prompt intervention, which holds the potential for delaying or reversing disease progression and allowing for a greater number of functional and meaningful years.

Humans↗

Enterocutaneous Fistula-Associated Sepsis and Mortality: Development and Validation of a Multimodal Artificial Intelligence Prediction Model.

BACKGROUND: Predicting enterocutaneous fistula (ECF)-associated sepsis and mortality poses significant challenges in digital health care due to the disease's complexity and heterogeneous clinical manifestations. Current approaches that rely on single-modal data or traditional scoring systems often fail to capture the intricate immune-inflammatory dynamics and multisystem involvement in patients with ECF. OBJECTIVE: This study aims to develop an artificial intelligence (AI)-driven multimodal fusion model integrating clinical, imaging, and transcriptomic data for early prediction of ECF-associated sepsis and 28-day mortality, addressing the limitations of conventional single-dimensional models. METHODS: This study leveraged publicly available datasets (Medical Information Mart for Intensive Care III [MIMIC-III], electronic Intensive Care Unit [eICU], and The Cancer Genome Atlas) to construct a multimodal framework. Clinical parameters were processed using Extreme Gradient Boosting, abdominal imaging features were extracted via convolutional neural networks, and transcriptomic profiles were analyzed with variational autoencoders. A Transformer-based fusion network was employed for joint prediction and validated through cross-validation and external testing. Key features were identified using Shapley Additive Explanations and Local Interpretable Model-Agnostic Explanations interpretability algorithms, while immune regulatory mechanisms were explored via weighted gene co-expression network analysis. RESULTS: The multimodal model achieved an area under the curve (AUC) of 0.89 for predicting sepsis and 28-day mortality, outperforming unimodal models (clinical-only model, AUC 0.72, and imaging-only model, AUC 0.78). Critical predictors included Sequential Organ Failure Assessment score, lactate levels, intra-abdominal free fluid on imaging, and immunoregulatory genes (programmed death-ligand 1 [PD-L1] and indoleamine 2,3-dioxygenase 1 [IDO1]). Mechanistic analysis revealed distinct immune reprogramming in patients with sepsis, characterized by increased regulatory T cells and M2 macrophages, along with downregulated cluster of differentiation 8+ (CD8+) T cells. CONCLUSIONS: This multimodal AI model offers an innovative digital solution in medical informatics, enabling precise early risk stratification for ECF-associated sepsis. By integrating multisource data and providing interpretable insights into immune-inflammatory pathways, the model enhances health care quality for patients with ECF and paves the way for personalized intervention strategies.

Humans↗

Performance of brain-damaged, schizophrenic, and normal subjects on a visual searching task.

Goldstein, Rennick, Welch, and Shelly (1973) reported on a visual searching task that generated 94.1% correct classifications when comparing brain-damaged and normal subjects, and 79.4% correct classifications when comparing brain-damaged and psychiatric patients. In the present study, representing a partial cross-validation with some modification of the test procedure, comparisons were made between brain-damaged and schizophrenic, and brain-damaged and normal subjects. There were 92.5% correct classifications for the brain-damaged vs normal comparison, and 82.5% correct classifications for the brain-damaged vs schizophrenic comparison.

Adult↗

Psychometric correlates of pain perception.

There is disagreement in the literature as to whether responsivity to painful stimuli possesses psychometric correlates. A series of methodological and statistical factors are specified in this paper which could account for the equivocality of the literature. A series of experiments were performed in which (a) various methodological and statistical issues were first resolved and (b) psychometric correlates of pain perception were then identified by means of a stepwise multiple regression procedure. The criterion variable consisted of the psychophysical judgment of pain during a 2-min. exposure to a 3,000 gm. force on the periosteum of the left fore-finger's second digit. The predictor variables consisted of selected psychological states and traits measured by the State-Trait Anxiety Inventory, Somatic Perception Questionnaire, Depression Adjective Checklist, Profile of Mood States, Eysenck Personality Inventory, and the Embedded Figures Test. The test-retest reliability of the pain test ranged from .64 to .84 across trials separated by a 3-wk. period. In the first experiment significant multiple regressions ranging between .57 and .72 were observed and psychological traits (field dependence, extraversion and trait anxiety) accounted for the variance in these analyses. In the next experiment significant multiple Rs ranging from .62 to .68 were observed. This served as cross-validation for the first experiment. The major difference was that psychological states (depression and vigor) as well as traits entered the multiple regression equations for certain of the analyses. It was concluded that selected psychological states and traits are significantly correlated with the perception of pain.

Adult↗

Lateral preference and style of cognition.

The relationship between cognitive ability and laterality was examined in terms of the relation of intelligence test scores to lateral preference. The factor analysis was performed on the variables of 12 tasks of the intelligence scale and total lateral preference. A slight relation was found between lateral preference and figure combination task. To clarify the relationship, the mean scores of tasks were tested for subjects who preferred the right and left on each preference item. Some significant differences were found. On some items, the mean scores of subjects with left preferences were inferior to those of subjects with right preferences on the figure-combination task. The result confirmed Levy's finding (1969). But cross-validation on a large sample is required.

Cognition↗

Predicting success in a smoking control program.

Using 10 independent variables, several of which had previously been related to smoking behavior, a regression equation was derived to predict success in a smoking control program. Thirty-one subjects were divided into the main and cross-validation treatment groups. Both groups participated in four 2-hour sessions. Taken individually, none of the predictor variables discriminated between successful and unsuccessful subjects. The regression equation, based on the five best predictors, accurately predicted direction of change of smoking rate 87% of the time.

Accidents, Traffic↗

Spinal meningiomas: histopathological grading using a benchmark radiomics model with notes on disease control.

OBJECTIVE: Spinal meningiomas (SMs) are common primary spinal tumors for which surgery is considered the first-line treatment when safe and feasible. The ability to extrapolate the tumor grade from preoperative imaging may significantly inform early patient expectation-setting regarding recurrence. Building on radiomics studies in cranial meningiomas, the authors aimed to construct a benchmark radiomics model to preoperatively identify the histological grade of SMs. METHODS: Institutional surgical records from May 2012 to November 2025 were queried for pathology-confirmed meningiomas below the foramen magnum, with preoperative contrast-enhanced imaging available for segmentation. SMs were classified as low-grade (WHO grade 1) and high-grade (WHO grade 2 tumors and grade 1 tumors with atypia). Tumors were manually segmented, and features were extracted using the PyRadiomics software package. An ensemble model of k-nearest neighbors, random forest, and support vector machine classifiers was trained using nested cross-validation on a subset of 10 features to differentiate tumor grades. Clinical data for the cohort were also extracted, and disease control in an adjunctive clinical series was assessed. RESULTS: Seventy-four patients were included in radiomics analysis, with an area under the receiver operating characteristic curve of 0.879 and a mean F1 score of 0.748. The model's top 5 features were all texture features that differed significantly (p < 0.05) across low- and high-grade SMs. These included measures of tumor textural and contrast-enhancement heterogeneity, with overlap with features reported in radiomics models for histological grading of intracranial meningiomas. Fifty-five patients with a median radiographic follow-up of 22.2 (range 1.9-86.4) months remained for clinical analysis after exclusion of patients with less than 1 month of follow-up and syndromic meningiomas. Four recurrences occurred at a median of 20.8 (range 1.8-41.8) months. High-grade tumor pathology did not significantly impact progression-free survival (p = 0.682, log-rank test; Cox regression high vs low grade hazard ratio [HR] 0.62, 95% CI 0.06-6.11, p = 0.685). Subtotal resection was associated with poorer progression-free survival than gross-total resection (p = 0.004, log-rank test; Cox regression subtotal vs gross-total resection HR 10.62, 95% CI 1.46-77.05, p = 0.019). These findings remain contextualized within a relatively limited follow-up window and small recurrence event count, suggesting a need to characterize the interplay between tumor grade and extent of resection as drivers of local disease control in SMs. CONCLUSIONS: A preoperative radiomics model can stratify high-grade SMs using open-source tools applied to single-institution data.

Humans↗

Identification of Critical Genes for Recurrent Aphthous Ulcer by Transcriptome Data Analysis and Mendelian Randomization.

PURPOSE: Recurrent aphthous ulcer (RAU) is a common oral mucosal disorder with a poorly understood etiology, significantly affecting patients' quality of life. This study aims to investigate critical genes linked to RAU and explore their biological mechanisms using transcriptomic data and Mendelian randomization (MR) analysis. MATERIALS AND METHODS: RAU-related gene expression data from the GEO database (GSE37265) were analyzed to identify differentially expressed genes (DEGs). A two-sample MR approach was used to assess the causal impact of expression quantitative trait loci (eQTL) on RAU. Critical genes were identified by intersecting DEGs with significant MR findings. GO and KEGG pathway enrichment analyses were performed, along with GSEA and immune cell infiltration analysis, to investigate the functions and mechanisms of these genes in RAU. RESULTS: A total of 184 differentially expressed genes (DEGs) were identified, while 339 RAU-associated genes were screened through MR analysis. Cross-validation further identified 7 critical genes. Among these, CCR1, ERP27, HCK, MICB, and SLC2A3 showed protective associations with RAU risk, whereas CD177 and IFITM1 were positively associated with increased risk. Enrichment analysis revealed that these genes are involved in specific biological processes, including cell migration, immune response, and metabolic regulation, which are closely linked to RAU pathogenesis. CONCLUSION: This systematic study comprehensively investigates the critical causative genes underlying RAU, emphasizing the intricate relationships between immune regulation and metabolic disturbances in its pathology. These findings lay a solid foundation for the development of novel biomarkers and may inform future research on targeted therapeutic strategies for RAU.

Stomatitis, Aphthous↗

Discovery and validation of a multi-protein panel for predicting non-fatal major adverse cardiovascular events in diabetic kidney disease.

OBJECTIVE: To identify plasma protein biomarkers associated with incident non-fatal major adverse cardiovascular events (MACE) in diabetic kidney disease (DKD) patients. RESEARCH DESIGN AND METHODS: We analyzed 317 DKD patients from the UK Biobank. Plasma proteomics and clinical data (demographics, metabolism, renal function) were integrated. In an exploratory discovery phase, three sequential Cox regression models (crude, socio-demographic-adjusted, socio-demographic-metabolic adjusted) screened non-fatal MACE-associated proteins. To prevent information leakage, the cohort was then randomly split into training (70%) and testing (30%) sets; machine-learning feature selection, hyperparameter optimization, and final model development were performed exclusively within the training set. The associated proteins were input into the four-step machine-learning pipeline (LASSO-Cox, random survival forest, Boruta, XGBoost-Cox). Predictive performance was validated using Kaplan-Meier survival analyses, longitudinal trajectory modeling, and ROC benchmarking. An interactive web application was deployed for clinical implementation. RESULTS: Of 1,463 plasma proteins, 561 were associated with non-fatal MACE across Cox models, with 14 overlapping proteins. Nine core proteins (ANG, IL1R1, CXCL14, ESAM, PTGDS, HAVCR1, FGFR2, IGSF8, CCL3) were validated: ANG showed the strongest non-fatal MACE association (HR&#xa0;=&#xa0;3.88, 95%CI 2.33-6.48, p<0.001), and all high-expression groups had elevated non-fatal MACE risk. GO/KEGG enrichment highlighted inflammatory-immune pathways like positive regulation of MAPK cascade, Cytokine-cytokine receptor interaction and PI3K-Akt signaling pathway as key mechanisms. The model integrating proteins, demographic factors, and clinical variables achieved the highest predictive performance across non-fatal MACE (AUC&#xa0;=&#xa0;0.768), myocardial infarction (MI) (0.808), and stroke (0.816) outcomes, with superior stability in cross-validation. CoxBoost + Elastic Net framework was selected as the optimal framework via benchmarking of 101 algorithms. The model demonstrated favorable calibration in high-risk patients and yielded positive net clinical benefit across decision thresholds of 5% to 45%. The web tool (https://jiangli2941.github.io/MACE-prediction-v2/) enables input of 28 variables, outputs non-fatal MACE risk status, risk probability, and highlights abnormal indicators. CONCLUSION: Plasma proteomics combined with machine learning identifies robust non-fatal MACE predictors in DKD.

Humans↗

Machine learning-integrated multi-omics risk prediction for pulmonary fungal infection in COPD and lung cancer: a transcriptomic and immune profiling study.

BACKGROUND: Chronic obstructive pulmonary disease (COPD) and lung cancer are major risk factors for invasive pulmonary fungal infection (IPFI), carrying an attributable mortality of 30%-80%. Their coexistence further amplifies immunosuppression, while current diagnostic criteria remain inadequate for early risk identification. METHODS: Transcriptomic data from the GEO dataset GSE296912 (scRNA-seq; 12,078 cells from normal and COPD lung tissue) and The Cancer Genome Atlas (TCGA)-lung adenocarcinoma (LUAD) bulk RNA-seq cohort (539 tumor and 59 normal samples) underwent differential expression and cross-omics integration analysis. Five machine learning models were constructed: logistic regression, SVM, random forest, XGBoost, and LASSO. Candidate genes were validated by qRT-PCR in A549 cells and THP-1-derived macrophages stimulated with heat-inactivated Aspergillus fumigatus conidia, a protocol selected to ensure BSL-2 biosafety compliance and isolate PAMP-mediated innate immune signaling. Model performance was evaluated using 5-fold stratified cross-validation with AUC, calibration curves, and decision curve analysis. RESULTS: Single-cell transcriptomic analysis of 12,078 cells identified 14 distinct cell populations, with marked myeloid expansion and immune dysregulation in COPD lung tissue. Cross-omics integration with TCGA-LUAD data identified 1,145 shared genes (79 immune-related), converging on NF-&#x3ba;B, TLR4, and cytokine receptor signaling. The random forest model achieved excellent discriminative performance (5-fold CV AUC = 0.988), with Treg infiltration, TLR4, and MMP9 as the top predictors. qRT-PCR confirmed significant upregulation of all five candidate genes (DEFB4A, S100A8, IL-8, MMP9, and TLR4) in both A549 and THP-1 cells following fungal stimulation. CONCLUSION: This multi-omics machine learning model integrating scRNA-seq and TCGA transcriptomic data demonstrates excellent discriminative performance (AUC = 0.988), with mechanistic convergence of NF-&#x3ba;B, TLR4, and oncogenic signaling pathways identified across shared immune gene signatures. In vitro qRT-PCR validation confirms the biological relevance of five key antifungal immune genes, providing a transcriptomic foundation for future prospective IPFI risk stratification in patients with COPD and lung cancer.

TLR4↗

Comparative genomics reveals hidden biosynthetic diversity in Streptomyces spp. and metal-dependent regulatory features associated with untapped specialized metabolites.

The genus Streptomyces is one of the richest sources of bioactive natural products; however, a substantial proportion of its biosynthetic gene clusters (BGCs) remain cryptic and their metabolic products are unresolved. Advances in genome mining and computational prediction now enable comprehensive exploration of this hidden biosynthetic repertoire. In this study, whole-genome sequencing and comparative genomic analyses were performed on three three newly isolated Streptomyces strains to evaluate their specialized metabolic potential. Genome assemblies were annotated and systematically analyzed using antiSMASH, DeepBGC, GECCO, and PRISM to identify, cross-validate, and functionally characterize BGCs while predicting their associated secondary metabolite scaffolds. Taxonomic analyses based on Average Nucleotide Identity (ANI), phylogenomics, and BLAST identified the isolates as Streptomyces thinghirensis, Streptomyces novocaesareae, and Streptomyces griseorubens. Applying the consensus framework across the three Streptomyces genomes yielded 43 cryptic BGCs, lacking close similarity to reference BGCs in the MIBiG database, of which 26 were classified as HIGH, 10 as MEDIUM, and 7 as LOW confidence. Notably, numerous BGCs exhibited low abundance to characterized reference clusters, indicating a high potential for previously undescribed biosynthetic pathways and novel metabolite scaffolds. Comparative analyses further revealed strain-specific biosynthetic architectures together with putative metal-responsive regulatory systems; Fur, Zur, and Nur, which were frequently associated with specialized metabolite biosynthetic loci. Collectively, these findings demonstrate the effectiveness of integrated genome-mining strategies for prioritizing cryptic biosynthetic gene clusters and highlight the remarkable biosynthetic potential of newly identified Streptomyces isolates as a source of novel natural products.

comparative genomics↗

Discovery and validation of GNA12circle as a first-trimester plasma eccDNA marker for early-onset preeclampsia.

BACKGROUND: Early-onset preeclampsia (EOPE) is a major cause of maternal and perinatal morbidity and is characterized by placental dysfunction, systemic endothelial injury, and hypertensive vascular stress. Because hypertensive disorders of pregnancy may also signal later maternal cardiovascular and cerebrovascular vulnerability, effective biomarkers for first-trimester risk assessment remain clinically important. Extrachromosomal circular DNA (eccDNA), a stable form of circulating cell-free DNA, has emerged as a potential source of disease-associated biomarkers. This study aimed to characterize first-trimester plasma eccDNA alterations associated with subsequent EOPE and to identify and validate a candidate circulating eccDNA marker for early risk assessment. METHODS: A two-stage nested case-control study was conducted within a prospective birth cohort. In the discovery stage, plasma samples collected at 11-13&#x202f;weeks of gestation from 5 women who subsequently developed EOPE and 5 matched normotensive controls were profiled by Circle-Seq to characterize genome-wide eccDNA alterations. Candidate eccDNAs were prioritized through differential abundance analysis and were further confirmed by outward PCR and Sanger sequencing. In the validation stage, the candidate selected marker was quantified by junction-specific qPCR in an independent cohort of 109 EOPE cases and 109 controls. Its potential predictive value was further evaluated alone and in combination with routine first-trimester clinical variables. RESULTS: In the exploratory discovery analysis, 410 nominally differentially abundant candidate eccDNAs were identified as a hypothesis-generating pool. Among these, GNA12circle (chr7:2876332-2,876,692) was prioritized and experimentally validated at the circular junction. In the independent validation cohort, plasma GNA12circle levels were significantly higher in women who later developed EOPE than in controls. When combined with routine first-trimester variables, GNA12circle improved predictive performance. The RF model showed the best overall cross-validated performance among the evaluated classifiers, with a mean held-out test-fold AUC of 0.843. CONCLUSION: First-trimester plasma eccDNA profiling revealed distinct alterations associated with subsequent EOPE, from which GNA12circle was identified and validated as a candidate circulating marker. These findings support further investigation of circulating eccDNA for early EOPE risk assessment in larger multicenter populations.

Humans↗

Genetic Susceptibility to Incisional Hernia Evaluation of Hernia Polygenic Risk Scores.

OBJECTIVES: Incisional hernia (IH) affects 13-30% of people after abdominal surgery, resulting in substantial morbidity and costs. While clinical risk factors have been studied extensively, genomic risk for IH is incompletely understood. We aimed to evaluate the impact of polygenic risk scores (PRS) on IH risk prediction. METHODS: We created and evaluated three PRS for abdominal hernia, ventral hernia and latent hernia susceptibility for prediction of IH in an institutional biobank. The primary outcome was defined as the diagnosis or repair of an IH based on ICD-9/10-CM/PCS and CPT codes. Clinical covariates included age, sex, body mass index (BMI), smoking status, index procedure type, and perioperative surgical site infection. A phenome-wide association study (PheWAS) was performed to assess clinical associations with increased PRS. We then tested the ability of the PRS to improve prediction for IH by modeling clinical covariates with and without PRS in patients who underwent abdominal surgery. Model performance was assessed using 10 iterations of 5-fold cross-validation to estimate Brier scores and area under the receiver operating characteristic curve (AUROC), which were compared using cross-model Bayesian analysis of variance. RESULTS: In 55,809 subjects, assessed PRS was significantly associated with incisional, umbilical, and ventral hernia on PheWAS, with 1.19 greater odds of developing IH per 1-SD increase in PRS (95% CI: 1.13-1.25, P < 0.001). Of 9,909 subjects who underwent qualifying abdominal surgery, 706 developed IH. In this cohort, the latent hernia susceptibility PRS was associated with a 16% increased hazard of developing IH per 1-SD increase (HR 1.16; 95% CI: 1.07-1.26; P < 0.001). Compared to a predictive model using clinical covariates (Brier score = 0.047, 95% CI: 0.046-0.048; AUROC = 0.660, 95% CI: 0.653-0.666), addition of the PRS showed similar Brier score and AUROC estimates (Brier score = 0.047, 95% CI: 0.046-0.048; AUROC: 0.667, 95% CI: 0.661-0.673) at five years. Cross-model Bayesian analysis demonstrated >99% probability of practical equivalence when trying to detect a difference of &#x2265; 0.02. CONCLUSION: All three PRS for hernia were independently associated with IH, suggesting that genomic factors contribute significantly to IH development. However, none of the three PRS meaningfully improved clinical IH risk prediction in patients who underwent abdominal surgery. This suggests that clinical comorbidities and surgical techniques may be equally as important as genomic architecture.

Bayesian analysis↗

Genomic signatures associated with epidemiologically defined high-risk pathogenic Escherichia coli isolates identified by interpretable machine learning.

Pathogenic Escherichia coli is a major cause of foodborne illness worldwide and includes strains capable of causing severe disease. To establish a genome-informed framework for foodborne outbreak surveillance, we analyzed 1,029 E. coli isolates from clinical, food, livestock, and environmental sources using whole-genome sequencing. Pathogenic isolates obtained from human clinical cases or linked to documented outbreaks were classified as epidemiologically defined high-risk (EpiHR), whereas the remaining pathogenic isolates were classified as non-EpiHR. Virulence-associated genomic features were extracted using a bioinformatics pipeline, and four machine learning (ML) algorithms, including gradient boosting machine, random forest (RF), and support vector machines with linear and radial basis function kernels, were evaluated. Among them, the RF model showed the best performance, achieving an area under the curve (AUC) of 0.98 and accuracy of 0.93 in 10-fold cross-validation. Additional leave-one-group-out validation showed retained discrimination across held-out sequence types and serotypes, although performance was reduced when isolates were grouped by isolation source. Evaluation using an independent test dataset of 1,908 publicly available pathogenic E. coli genomes showed an AUC of 0.97 and a sensitivity of 0.98. Feature importance analysis using Shapley additive explanations identified influential predictive features, including traT, etpB, and enterotoxin-associated genes. A reduced 10-feature model achieved an AUC of 0.79 in the independent test dataset, supporting its exploratory use for future simplified screening approaches. These results indicate that genome-based ML provides a sensitive framework for surveillance-oriented prioritization of EpiHR pathogenic E. coli isolates, with model predictions interpreted together with epidemiological information.

Escherichia coli↗

Individual temperament as a predictor of health or premature disease.

Two studies of temperament as a possible predictor of continuing good health or premature disease are reported. In earlier work to determine youthful precursors of premature disease, a number of separate characteristics distinguishing medical students who remain healthy from those with premature disorders have been identified. Characterization by temperament, an expression of innate biological endowment, provides a more global portrayal of an organism that can an aggregate of separate characteristics alone. Criteria for designating three temperament types, termed Alpha, Beta and Gamma, are presented. In 1948, 45 subjects were assigned to one of these types on the basis of youthful characteristics. The subjects in the three temperament groups had different health outcomes 30 years later, significant at the p less than 0.01 level. Gamma type had the most disorders and deaths, Beta type the fewest. A cross-validation study on 127 subjects had similar results at the p less than 0.05 level. Temperament appears to be a variable of predictive potential of individual stamina, or of vulnerability to premature disease and death.

Adult↗