PubMed HealthSearch

SEARCH · PubMed Health

Results for “Prediction”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Improving the reliability of polygenic risk score-based prediction for cardiovascular and renal complications across ancestries in type 2 diabetes using Mondrian Cross-Conformal Prediction.

Polygenic risk scores (PRS) developed in European populations often show reduced predictive performance in non-European populations, limiting their clinical utility. This lack of transferability across ancestries remains a major challenge in genomic medicine and raises concerns about health equity. We aimed to evaluate whether uncertainty-aware prediction, implemented through Mondrian Cross-Conformal Prediction, improves the performance and reliability of polygenic risk score-based predictions across ancestries for nephropathy, stroke, and myocardial infarction in individuals with type 2 diabetes in a multi-ethnic cohort. We leveraged Mondrian Cross-Conformal Prediction (MCCP), an uncertainty quantification framework, combined with logistic regression applied to a multi-polygenic risk score (multiPRS) to predict the risk of nephropathy, stroke, and myocardial infarction in individuals with type 2 diabetes. Two training frameworks were evaluated: one using 4,098 individuals with type 2 diabetes of European ancestry from the ADVANCE trial for training and 17,574 White British, 1,145 South Asian, and 749 African UK Biobank participants for testing; and another using the 17,574 White British UK Biobank participants for training and the South Asian and African participants for testing. Logistic regression provided robust baseline performance across populations. On top of this baseline, MCCP did not improve performance but added capabilities absent from probability-based stratification: for each individual, it issued a prediction together with an explicit confidence and credibility level; it allowed a tolerated error level to be set in advance and delivered prediction sets respecting it in the majority of settings; and it flagged individuals for whom no reliable prediction could be made. Applying MCCP to PRS-based prediction thus enables uncertainty-aware risk stratification and improves the reliability of risk prediction across ancestries, providing a more equitable framework for clinical use.

Female

Mul-PheG2P: decoupled learning and prediction-space fusion enables robust and interpretable multi-phenotype genomic prediction.

Genomic prediction of multiple phenotypes is crucial in modern plant breeding; however, existing methods struggle with negative transfer and lack interpretability, particularly across high-dimensional small-sample data and diverse species. To address this, we propose Mul-PheG2P, a novel paradigm based on decoupled learning and predictive space fusion. It employs a two-stage design: first training phenotype-specific encoders using genetic data, then decoupling phenotype-specific learning from cross-phenotype aggregation via an interpretable prediction layer. Mul-PheG2P outperforms existing methods across diverse crop datasets, including maize (Zea mays), wheat (Triticum aestivum), and tomato (Solanum lycopersicum). It provides a multi-scale interpretability chain: at the macro level, it quantifies phenotypic contributions via attention-based weighting; at the micro level, Integrated Gradients reveal the genetic basis of predictions. Notably, the model successfully identified the CCT (CONSTANS, CO-like, and TOC) motif regulating photoperiodism and the SQUAMOSA (SQUAMOSA promoter binding protein) promoter for inflorescence development, confirming its ability to capture functional biological mechanisms. These results highlight the high performance and interpretability of Mul-PheG2P, showcasing its value for low-cost, large-scale screening to advance precision breeding.

Phenotype

Harsh Parenting Predicts Novel HPA Receptor Gene Methylation and NR3C1 Methylation Predicts Cortisol Daily Slope in Middle Childhood.

Adverse experiences in childhood are associated with altered hypothalamic-pituitary-adrenal (HPA) axis function and negative health outcomes throughout life. It is now commonly accepted that abuse and neglect can alter epigenetic regulation of HPA genes. Accumulated evidence suggests harsh parenting practices such as spanking are also strong predictors of negative health outcomes. We predicted harsh parenting at 2.5&#xa0;years old would predict HPA gene DNA methylation similarly to abuse and neglect, and cortisol output at 8.5&#xa0;years old. Saliva samples were collected three times a day across 3 days to estimate cortisol diurnal slopes. Methylation was quantified using the Illumina Infinium MethylationEPIC array BeadChip (850&#xa0;K) with DNA collected from buccal cells. We used principal components analysis to compute a summary statistic for CpG sites across candidate genes. The first and second components were used as outcome variables in mixed linear regression analyses with harsh parenting as a predictor variable. We found harsh parenting significantly predicted methylation of several HPA axis genes, including novel gene associations with AVPRB1, CRHR1, CRHR2, and MC2R (FDR corrected p&#x2009;<&#x2009;0.05). Further, we found NR3C1 methylation predicted a steeper diurnal cortisol slope. Our results extend the current literature by demonstrating harsh parenting may influence DNA methylation similarly to more extreme early life experiences such as abuse and neglect. Further, we show NR3C1 methylation is associated with diurnal HPA function. Elucidating the molecular consequences of harsh parenting on health can inform best parenting practices and provide potential treatment targets for common complex disorders.

Child

Protocol to predict gene expression from transcriptomic data using PREDICT.

Linking DNA sequence variation to context-specific transcriptional programs is a critical challenge in regulatory genomics, especially for non-model organisms. Here, we present PREDICT, a modular Python package for discovering cis-regulatory elements and transcription factor binding motifs. We describe steps to identify enriched k-mers from differentially expressed genes, map them to known motifs, quantify their impact on gene expression, and visualize motif co-occurrences. PREDICT provides a robust, k-mer-based approach to uncover regulatory logic in diverse genomic systems. For complete details on the use and execution of this protocol, please refer to Yen et al. and Liu et al.1,2.

Gene Expression Profiling

Impact of personalised risk predictions on breast cancer risk perceptions: insights from the BREATHE study.

OBJECTIVE: Biennial mammography screening is well-established for women aged 50 and above, but guidelines for younger women are less clear. Risk-based screening may provide women with key information to make informed decisions about their breast cancer risk and screening. This study examines how predicted breast cancer (BC) risk shapes women's perception and confidence in risk prediction. METHODS: Women aged 35 to 59&#xa0;years were recruited for a prospective multi-centre cohort and stratified into above-average, average, or below-average BC risk categories based on genetic and non-genetic risk factors. Perceived risk was assessed at enrolment and after participants were informed of their predicted risk. We used ordinal models to identify predictors of perceived risk and logistic regression to examine the relationship between changes in perceived risk and confidence in the risk prediction. RESULTS: At enrolment, 43% and 47% of 4112 participants perceived their BC risk pre-result as low or average, respectively. Thirty-five percent adjusted their perceived risk to align more closely with their predicted risk. Predictors of perceived risk post-result: perceived risk pre-result, predicted risk, ethnicity and having regular menstruation. Participants who underestimated their BC risk were nearly eight times more likely to have low confidence in the accuracy of their predicted risk (OR for underestimation vs. accurate perception: 7.94 [95% CI 5.60-11.28]). Predictors of perceived risk post-result: perceived risk pre-result, predicted risk, ethnicity and having regular menstruation. Confidence in risk prediction was lowest when women's perceived risk pre-result was lower than their predicted risk (OR-2 vs 0 [95%CI] 5.06 [3.67 to 6.97]). CONCLUSION: Many women underestimated their BC risk, and their initial perceptions were influenced by the knowledge of their predicted risk. Women who underestimated their risk had less confidence in their predicted risk scores.

Humans

Predictive Models for Hypoglycemia Risk in Haemodialysis Patients With Diabetic Kidney Disease: Systematic Review and Meta-Analysis.

AIM: To provide evidence for selecting and developing reliable clinical assessment tools for hypoglycemia in diabetic kidney disease patients during haemodialysis. DESIGN: Review. METHODS: Systematic searches were performed in 9 Chinese and English databases to collect literature regarding the development of hypoglycemia risk prediction models in haemodialysis patients with diabetic kidney disease. Two reviewers independently performed literature screening, data extraction, risk-of-bias assessment, and applicability evaluation. The Prediction Model Risk of Bias Assessment Tool was used to assess the risk of bias and applicability of the included studies. Meta-analysis was conducted using R software. DATA SOURCES: CNKI, Wanfang, VIP, CBM, PubMed, Cochrane Library, EMbase, Web of Science, and CINAHL. The search period covered from the establishment date of each database to December 2025. RESULTS: Six studies, comprising six prediction models, were included. Two studies performed internal validation, and three conducted external validation. All models reported the area under the curve, ranging from 0.813 to 0.866, and calibration measures. Four studies were rated as having a high risk of bias, while all six demonstrated good overall applicability. The meta-analysis showed that the pooled AUC value of the six studies was 0.846 (95% CI: 0.823-0.867). CONCLUSION: Research on hypoglycemia risk prediction models in haemodialysis patients with diabetic kidney disease remains in the developmental stage. Although the included prediction models exhibited satisfactory apparent discriminatory ability and clinical applicability, most of the original studies suffered from a high risk of bias and lacked adequate validation. The true predictive performance and clinical application value of these models remain to be further verified. Accordingly, routine and unconditional clinical application is not recommended at this stage. Future studies should include more high-quality, multicenter external validation and develop models with high generalizability, favourable clinical applicability, and robust predictive performance to facilitate early identification of hypoglycemia risk in this population. IMPACT: This study systematically evaluated the hypoglycemia risk prediction models for diabetic kidney disease patients during haemodialysis, and the research on hypoglycemia risk prediction models for maintenance haemodialysis patients during dialysis is still in the development stage. This study provides a reference for clinical medical staff to select or develop hypoglycemia risk prediction and assessment tools for diabetic kidney disease patients during haemodialysis. REPORTING METHOD: This study was conducted in accordance with the relevant guidelines of the EQUATOR Network and followed the TRIPOD-SRMA Checklist. PATIENT OR PUBLIC CONTRIBUTION: No patient or public contribution. TRIAL REGISTRATION: PROSPERO: CRD420251243352.

Humans

CLASPP: A unified model for predicting post-translational modifications.

Post-Translational Modifications (PTMs) are a fundamental mechanism for regulating cellular pathways and increasing the functional diversity of the proteome. Accurately predicting the PTM types that are likely to occur at a given site in the primary sequence is a key challenge in functional proteomics. Existing PTM prediction models predominantly focus on either single PTM types or employ ensemble methods that combine multiple models to predict different PTM types. This fragmentation is largely driven by the vast imbalance in data availability across PTM types, making it difficult to predict multiple PTM types with a single model. To address this limitation, we present the Contrastively Learned Attention-based Stratified PTM Predictor (CLASPP), a unified PTM prediction model. CLASPP addresses imbalance challenges by leveraging unsupervised clustering-based undersampling and a novel contrastive learning framework tailored to PTM data. Additionally, our hierarchical data organization and curation are shown to improve CLASPP's performance by balancing the representation of individual PTM types and provides a standardized dataset to train and validate future model designs. Drawing inspiration from advancements in image and natural language processing, the CLASPP model employs a multi-stage training strategy and a high-quality, curated training dataset to improve PTM prediction performance. To uncover what is learned during the contrastive learning stage, the CLASPP model is shown to distinguish known protein kinase substrate specificity profiles as a form of explainability. Finally, we evaluate the application of CLASPP in predicting PTMs in different model organisms and experimentally validated ubiquitination sites in the understudied DCLK3 kinase. Overall, CLASPP represents a unified model for PTM prediction that addresses key bottlenecks in data imbalance and offers new strategies for biological data curation, thereby improving PTM-type prediction performance across diverse organisms.

Protein Processing, Post-Translational

Genome-based predictions of metabolic preferences and substrate phenotypes in psychrotrophic bacteria from permafrost environments.

Genomes reveal vast functional potential, but harbor genomic noise that obscures prediction of metabolic and environmental preferences. Genomic databases are skewed towards clinically relevant and easily cultivated bacteria, limiting predictions for diverse and underrepresented environmental taxa. Psychrotrophic bacteria, which can survive and grow in cold, nutrient-limited, dry, and saline environments, are especially underrepresented despite their relevance for understanding microbial responses to changing cold environments and potential biotechnological value given growth at low temperatures. Assembling complete genomes of 48 isolates from Alaskan permafrost, seasonally frozen active layer soils, and terrestrial ice, we used Kyoto Encyclopedia of Genes and Genomes (KEGG) ortholog annotations to evaluate the predictability of metabolic resource-use traits observed using phenotypic tests. Genome-predicted values for glycolytic versus gluconeogenic catabolic preference index, or sugar-acid preference (SAP), explained over 50% of the variance in empirically observed SAP. SAP was inversely correlated to genomic GC content, which follows phylum-level trends, indicating that coarse metabolic preference covaries with phylogeny. Regularized elastic net models offered a more granular view, linking KEGG genes to specific substrate utilization and sensitivity phenotypes and yielding moderate but reproducible accuracy (AUC 0.70-0.79) for 11 substrates, demonstrating that specific substrate responses may be predictable from relatively small subsets of KO genes. These results extend recent advances, such as the SAP metric, and highlight associations among genomic GC content, phylum, and broad metabolic strategy. Linking genomic content to phenotype using isolates is a necessary step toward predictive models of microbial function in environmental communities, and this work can be used for hypothesis generation, with applications towards more expansive data sets.IMPORTANCECold region soils and ice host psychrotrophic bacteria with metabolic traits and adaptations that enable persistence in harsh, resource-limited environments. However, these taxa are underrepresented in genomic reference databases dominated by well-studied, mesophilic organisms. This gap limits inference of ecological strategies and our ability to predict how these microbes may influence the large, thaw-vulnerable carbon reservoirs in permafrost. Here, we show that genomic GC content is associated with the sugar-versus-acid catabolic preference (SAP) of isolates across major phyla, suggesting that broad genomic features may provide a coarse signal of metabolic strategy. We demonstrate that a modified SAP metric, using binary (positive/negative) substrate utilization rather than detailed growth rate measurements, is moderately predictive, thus extending its application to slow-growing or difficult-to-culture taxa. Together, these advances broaden the toolkit for linking genome content to resource-use traits (phenotype) in poorly characterized, cold-adapted bacteria and offer a tractable entry point to broad prediction and hypothesis generation.

Genome, Bacterial

Genomic and Developmental Models to Predict Cognitive and Adaptive Outcomes in Autistic Children.

IMPORTANCE: Although early signs of autism are often observed between 18 and 36 months of age, there is considerable uncertainty regarding future development. Clinicians lack predictive tools to identify those who will later be diagnosed with co-occurring intellectual disability (ID). OBJECTIVE: To predict ID in children diagnosed with autism. DESIGN, SETTING, AND PARTICIPANTS: This prognostic study involved the development and validation of models integrating genetic variants and developmental milestones to predict ID. Models were trained, cross-validated, and tested for generalizability across 3 autism cohorts: Simons Foundation Powering Autism Research (SPARK), Simons Simplex Collection, and MSSNG. Autistic participants were assessed older than 6 years of age for ID. Study data were analyzed from January 2023 to July 2024. EXPOSURES: Ages at attaining early developmental milestones, occurrence of language regression, polygenic scores for cognitive ability and autism, rare copy number variants, de novo loss-of-function and missense variants impacting constrained genes. MAIN OUTCOMES AND MEASURES: The out-of-sample performance of predictive models was assessed using the area under the receiver operating characteristic curve (AUROC), positive predictive values (PPVs), and negative predictive values (NPVs). RESULTS: A total of 5633 autistic participants (4574 male [81.2%]) were included in this analysis. On average, participants were diagnosed with autism at 4 (IQR, 3-7) years of age and assessed for ID at 11 (8-14) years of age, with 1159 participants (20.6%) being diagnosed with ID. The model integrating all predictors yielded an AUROC of 0.653 (95% CI, 0.625-0.681), and this predictive performance was cross-validated and generalized across cohorts. This modest performance reflected that only a subset of individuals carried large-effect variants, high polygenic scores, or presented delayed milestones. However, combinations of genetic variants that are typically not considered clinically relevant by diagnostic laboratories achieved PPVs of 55% and correctly identified 10% of individuals developing ID. The addition of polygenic scores to developmental milestones specifically improved NPVs rather than PPVs. Notably, the ability to stratify ID probabilities using genetic variants was up to 2-fold higher in individuals with delayed milestones compared with those with typical development. CONCLUSIONS AND RELEVANCE: Results of this prognostic study suggest that the growing number of neurodevelopmental condition-associated variants cannot, in most cases, be used alone for predicting ID. However, models combining different classes of variants with developmental milestones provide clinically relevant individual-level predictions that could be useful for targeting early interventions.

Humans

Plasma Proteomic Profiles Predict Individual Future Osteoarthritis Risk.

OBJECTIVE: Osteoarthritis (OA) is a widespread degenerative joint disease that causes a considerable socioeconomic burden. Despite progress in genetic and environmental insights, early diagnosis is still limited by the lack of evident symptoms during the initial phases and accurate biomarkers. This study aims to identify plasma proteins associated with future risk of OA and develop a predictive model. METHODS: We conducted a large-scale proteomic analysis of 45,307 participants from the UK Biobank, excluding those with baseline OA. Plasma samples were assayed using the Olink Explore Proximity Extension Assay targeting 1,463 unique proteins. Clinical variables and OA outcomes were extracted and linked to electronic health records. A predictive model was constructed using the LightGBM machine learning method, and SHapley Additive exPlanations (SHAP) were applied to evaluate the importance of variables. RESULTS: We identified a panel of proteins significantly associated with the risk of developing OA. Notably, after adjusting for multiple confounders, collagen type IX alpha 1 chain (COL9A1) and cartilage acidic protein 1 (CRTAC1) were the most significant predictors of incident OA, with hazard ratios of 1.54 (95% confidence interval [CI] 1.48-1.61) and 1.65 (95% CI 1.54-1.78), respectively. SHAP analysis allowed a profound interpretation of the contribution of each protein and clinical variable to the model, revealing the multifactorial nature of OA risk prediction. The temporal trajectories of plasma proteins indicated that the levels of COL9A1 and CRTAC1 began to deviate from normal for more than a decade before OA onset, suggesting their potential use in early detection strategies. The predictive model, developed using the LightGBM algorithm, integrated proteins with clinical covariates and demonstrated an area under the curve (AUC) of 0.729 for 5-year OA prediction, 0.721 for 10-year prediction, and 0.723 for all incident OA. The predictive accuracy of the model was further enhanced for hip and knee OA, achieving AUCs of 0.820 and 0.803 for 5-year predictions. CONCLUSION: Our study identified the role of plasma proteomics in predicting future OA risk, which could contribute to preemptive measures. The innovative model, which integrates proteomic biomarkers with clinical data, offers a potential tool for risk assessment, potentially optimizing OA management strategies and enhancing prevention efforts.

Humans

Predicting bloodstream infection by plasma cell-free metagenomic sequencing: a prospective cohort study.

BACKGROUND: Patients receiving myelosuppressive chemotherapy or haematopoietic cell transplantation are at high risk for life-threatening bloodstream infections. A novel pre-emptive treatment paradigm guided by pathogen detection before symptoms appear might reduce this risk, but no validated screening test is available. This study evaluated the sensitivity and specificity of plasma microbial cell-free DNA metagenomic sequencing (mcfDNA-Seq) for predicting bloodstream infections in children and adolescents receiving therapy for high-risk leukaemia. METHODS: In this prospective cohort study, between Aug 9, 2017, and Feb 28, 2022, leftover clinical plasma samples were prospectively collected up to once per day from patients who were younger than 25 years, receiving care for leukaemia at St Jude Children's Research Hospital (Memphis, TN, USA), and at high risk for life-threatening bloodstream infections. mcfDNA-Seq was used to identify pathogen DNA in blood samples obtained during the 7 days before to 1 day after bloodstream infection onset, and in control samples from the same population in the absence of fever or infection. The testing laboratory was masked to sample status. Primary outcomes were predictive sensitivity of mcfDNA-Seq for detecting the expected bloodstream infection pathogen during the 3 days preceding the day of bloodstream infection onset, with a prespecified favourable sensitivity of 50%, and predictive specificity of mcfDNA-Seq in control samples. Exploratory analyses comprised assessing sensitivity and specificity restricted to bacteria or common bloodstream infection pathogens, and after applying a data-derived DNA fragment concentration cutoff; estimating the predictive sensitivity on each of the 7 days before bloodstream infection onset; identifying clinical characteristics that affected predictive sensitivity or specificity; and examining the clinical relevance of additional organisms identified by mcfDNA-Seq during bloodstream infection episodes. Diagnostic sensitivity was also assessed on samples collected on the day of, or day after, diagnosis of bloodstream infection. This study is registered with ClinicalTrials.gov, NCT03226158. FINDINGS: 94 evaluable bloodstream infections occurred in 60 (38%) of 158 enrolled participants; 19 episodes were previously described in the pilot phase of this study. The predictive sensitivity of mcfDNA-Seq was 51&#xb7;9% (95% CI 40&#xb7;5-63&#xb7;1) for all bloodstream infection episodes, 53&#xb7;8% (42&#xb7;2-65&#xb7;2) for bacterial infection only, and 51&#xb7;9% (40&#xb7;5-63&#xb7;1) when applying a DNA fragment concentration cutoff of 140 molecules per &#x3bc;L. Sensitivity was lowest at day -7 and increased daily until the day of diagnosis. Diagnostic sensitivity was 81&#xb7;3% (95% CI 71&#xb7;0-89&#xb7;1) for all bloodstream infection episodes and 83&#xb7;1% (72&#xb7;9-90&#xb7;7) for bacterial infections only. Predictive specificity was 82&#xb7;7% (95% CI 76&#xb7;0-88&#xb7;2), but improved to 88&#xb7;9% (83&#xb7;0-93&#xb7;3) for common bloodstream infection pathogens, and to 93&#xb7;8% (88&#xb7;9-97&#xb7;0) when also applying the DNA fragment concentration cutoff. Predictive sensitivity was higher in participants with acute lymphoblastic leukaemia (adjusted odds ratio [aOR] 11&#xb7;1 [1&#xb7;7-74&#xb7;2] vs those with acute myeloid leukaemia), and it was lower in polymicrobial infections (aOR 0&#xb7;0 [0&#xb7;0-0&#xb7;2] vs monomicrobial Gram-positive infections). Clinical false-positive results were positively associated with gastrointestinal disturbance alone (p=0&#xb7;037) or combined with recent administration of high-dose cytarabine (p=0&#xb7;012). Additional organisms identified by mcfDNA-Seq that were not identified by blood culture were less likely than expected organisms to have an increasing DNA concentration during the days preceding bloodstream infection diagnosis. INTERPRETATION: mcfDNA-Seq can detect causative pathogens before the onset of some bloodstream infection episodes in profoundly immunocompromised patients. Predictive specificity might be improved by restricting results to a subgroup of relevant organisms, excluding patients with high risk of false-positive results, or applying a higher concentration cutoff. Clinical trials are needed to evaluate mcfDNA-Seq-guided pre-emptive therapy for preventing life-threatening bloodstream infections in patients with high risk. FUNDING: The National Cancer Institute, American Lebanese Syrian Associated Charities, St Jude Children's Research Hospital, and Karius.

Adolescent

Machine Learning-Based Preoperative Predicting TERT Promoter Mutation and EGFR Gene Amplification Phenotype in IDH Wild-Type Glioblastoma Using Advanced MR Habitat Imaging.

BACKGROUND AND PURPOSE: The telomerase reverse transcriptase (TERT) gene promoter mutation is a crucial factor for identifying an isocitrate dehydrogenase (IDH) wild-type glioblastoma with poor prognosis, and the epidermal growth factor receptor (EGFR) amplification may be a potential prognostic factor. The purpose of this study was to investigate the value of the tumor habitats imaging model on advanced MRI in predicting TERT promoter mutation and EGFR gene amplification phenotype of IDH wild-type glioblastoma. MATERIALS AND METHODS: One hundred seventy-nine patients with pretreatment conventional MRI, DWI, and DSC-PWI were included. The data were divided into the training set (n=112), test set (n=29), and time-independent validation set (n=38). Based on the ADC and CBV map, the solid tumor area was split into several habitat subregions using the k-means clustering algorithm (hypovascular hypercellular area, hypervascular area, and hypovascular hypocellular area). In the training set, TERT promoter mutation and EGFR gene amplification phenotype prediction models were constructed using the random forest method. The reliability of prediction models was validated in the test and the time-independent validation sets. Receiver operating characteristic (ROC) curve analysis, calibration curve, and decision curve analysis (DCA) were used. RESULTS: The area under the curve (AUC) of the training, test, and validation sets of the TERT promoter prediction model was 0.877, 0.783, and 0.796, respectively. The accuracy of the TERT promoter prediction model was 82.1%, 75.9%, and 76.3%, respectively. The AUCs of the 3 sets for the EGFR gene amplification status prediction model were 0.877, 0.784, and 0.878, respectively. The accuracy of the EGFR gene amplification status prediction model was 79.5%, 75.9%, and 89.5%, respectively. Moreover, the prediction probability of these models was in good agreement with the actual result. CONCLUSIONS: The tumor habitat imaging model based on advanced MRI was useful for accurately predicting TERT promoter mutation and EGFR amplification status in IDH wild-type glioblastoma.

Humans

Pedigree-assisted genotype imputation enables cost-effective genomic prediction in Penaeus vannamei.

Genomic selection in Penaeus vannamei has long been constrained by the high cost of dense genotyping. To address this limitation, we evaluated genotype imputation from a low-density 1&#xa0;K panel to a medium-density 55&#xa0;K panel of the "Yellow Sea Array No. 1" and examined its impact on genomic prediction for harvest body weight in P. vannamei. A four-generation pedigree including 30 great-grandparents, 39 grandparents, 100 parents, and 608 offspring was genotyped using the 55&#xa0;K panel. A two-step experimental design was implemented to (i) assess the performance of different imputation algorithms under reference population scenarios with varying proportions of siblings, and (ii) compare six alternative reference population structures incorporating parents, ancestors, and siblings. Genotype imputation using the pedigree-based method FImpute v3.0 consistently achieved higher accuracy than the population-based method Beagle v5.5. Using this pedigree-assisted approach, imputation accuracy increased from 0.73 when only parental genotypes were used to 0.84 with the inclusion of 10% siblings, and subsequently plateaued at 0.87-0.90 when sibling representation reached 20%. Across the six reference population structures, imputation accuracy was primarily driven by the availability of parental genotypes, ranging from 0.50 to 0.56 in the absence of parents to 0.88-0.89 when both parents and ancestral generations were included. Accuracy remained high when both parents were available (0.84-0.87 with siblings; 0.73 without siblings) but declined substantially when only one parent was genotyped (0.65-0.68). Imputation accuracy was positively associated with both minor allele frequency (MAF) and linkage disequilibrium (max r2LD), with LD exerting the stronger influence. Heritability estimates derived from imputed 55&#xa0;K genotypes were highly consistent with those obtained from the original 55&#xa0;K data (0.39&#x2009;&#xb1;&#x2009;0.14 vs. 0.41&#x2009;&#xb1;&#x2009;0.14), indicating that genotype imputation did not compromise variance component estimation. In predictive ability analyses, pedigree-based BLUP (PBLUP) achieved higher predictive ability than genomic BLUP (GBLUP) based on the 1&#xa0;K panel, with predictive abilities of 0.42-0.44 for PBLUP compared with 0.34-0.35 for GBLUP. Using imputed genotypes for genomic prediction further improved predictive ability relative to the true 1&#xa0;K panel, yielding values ranging from 0.35 to 0.47. Notably, when parental genotypes were included in the reference population, GBLUP based on imputed genotypes surpassed the predictive ability of PBLUP and approached that achieved with the original 55&#xa0;K genotypes (0.45-0.47). Collectively, these results provide the first empirical evidence that low- to medium-density genotype imputation, combined with pedigree information, can effectively support genomic prediction in P. vannamei. This study establishes a cost-efficient and scalable framework for implementing genomic selection in P. vannamei and provides a practical reference for the application of genomic selection in other aquaculture species with constrained breeding budgets.

Animals

Comparing artificial and convolutional neural networks with traditional models for Genomic prediction in wheat.

With the rapid development of sequencing technology, the application of genomic prediction has become more and more common in breeding schemes of livestocks and crops. Selecting an appropriate statistical model is of central importance to achieve high prediction accuracy. Recently, machine learning models have been expected to upgrade genomic prediction into a new era. However, the perspective still suffers from lack of evidence that machine learning models can generally outperform the traditional ones on empirical data sets. In this study, we compared two machine learning models based on artificial neural network (ANN) and convolutional neural network (CNN) with four traditional models, including genomic best linear unbiased prediction (GBLUP), Bayesian ridge regression (BRR), BayesA and BayesB, using three published data sets for grain yield in wheat. For each model, we considered two variants: modeling and ignoring the genotype-by-environment ([Formula: see text]) interaction. In the comparison, we considered two strategies of cross-validation: predicting genotypes that have not been evaluated in any environment (CV1) and predicting genotypes that have been tested in other environments (CV2). Our results showed that traditional Bayesian models (BayesA, BayesB, and BRR) outperformed GBLUP, ANN and CNN when considering [Formula: see text] interaction. The accuracies of ANN and CNN were higher than traditional models only in CV1 and when [Formula: see text] interaction was ignored. It was also found that the performance of the two machine learning models was significantly affected by the interaction between the CV strategy and the way of treating the [Formula: see text] interaction, while that of the four traditional models was only influenced by whether the [Formula: see text] interaction was considered or not. Thus, machine learning models can be a powerful complementary to the traditional ones and their superiority may depend on the prediction scenario. Among the two machine learning models, we observed that the accuracy of ANN was higher than CNN in most cases, indicating that it is still challenging to adapt complex machine learning models such as CNN to genomic prediction.

ANN

Multimodal deep learning for immunotherapy response prediction and biomarker discovery in non-small cell lung cancer.

OBJECTIVE: Immunotherapy has emerged as a promising treatment for advanced non-small cell lung cancer (NSCLC), but accurately predicting which patients will benefit from it remains a major clinical challenge. To address this, we aim to develop a novel multimodal method, DeepAFM, that integrates histopathology, genomic features, and clinical information to predict patient responses to anti-PD-(L)1 immunotherapy. MATERIALS AND METHODS: A total of 93 patients with advanced NSCLC were included in this study. Histopathological whole-slide images were processed using a self-supervised VQVAE2 for representation learning. PCA and K-means clustering were then applied for dimensionality reduction and feature grouping. Key regions of interest were visualized through permutation importance evaluation and color-coding techniques. The extracted histopathological features, along with genomic alterations and clinical variables, were integrated into the DeepAFM multimodal prediction model. RESULTS: The DeepAFM achieved a high predictive performance with an area under the curve (AUC) of 0.77 (95% confidence interval: 0.69-1.00). Attention-based heatmaps revealed that the model could identify critical pathological patterns, genomic mutations, and clinical indicators associated with patient responses to immunotherapy. DISCUSSION: The integration of multimodal data enabled the model to capture complex interactions among pathology, genomics, and clinical characteristics, enhancing the interpretability and predictive power of immunotherapy response prediction. The visualization techniques facilitated the identification of biologically meaningful features and potential biomarkers. CONCLUSION: This study demonstrates the effectiveness of the DeepAFM in predicting responses to immunotherapy in advanced NSCLC. The approach not only improves prediction accuracy but also provides valuable insights for personalized treatment strategies and biomarker discovery.

Humans

The Progress of Gout Prediction Models Based on Multi-source Data.

INTRODUCTION: Gout, a highly serious inflammatory disease that is caused by monosodium urate crystals, is becoming an increasingly significant health concern. Artificial Intelligence and multi-omics-based research have made significant gains for the early detection and prevention of gout based on diverse approaches. This review intends to summarize current advances in forecasting gout susceptibility and gout-related symptoms, evaluate the predictive efficacy of different features, and ascertain which clinical and omics characteristics are most effective in these prediction models. METHODS: We explored the PubMed database after 2010 using keywords such as "gout", "predictive model", "risk prediction", and "machine learning", and confined our search to Englishlanguage articles. The original peer-reviewed research articles that developed gout models were selected. Research that was not original or lacked internal validation was excluded. RESULTS: Clinical features, genomics, microbiomics, radiomics, and metabolomics have been utilized to construct models related to gout and have demonstrated excellent predictive performance. Multisource data prediction models usually exhibit better effectiveness. DISCUSSION: Gout-oriented models performed excellently in predictive performance but present limitations in certain clinical and omics domains. However, if they are to affect actual patient care, they must overcome some external confirmation roadblocks and the fiscal and practical implications they will face ahead of time. CONCLUSION: This review indicates that clinical and multi-omics models of gout are significant instruments for clinical decision-making. The models constructed in these studies may be crucial for the treatment of gout and its practical benefits.

Gout

Prediction of bacterial protein-compound interactions with only positive samples.

MOTIVATION: Prediction of Compound-Protein Interactions (CPI) in bacteria is crucial to advance various pharmaceutical and chemical engineering fields, including biocatalysis, drug discovery, and industrial processing. However, current CPI models cannot be applied for bacterial CPI prediction due to the lack of curated negative interaction samples. RESULTS: We propose a novel Positive-Unlabeled (PU) learning framework, named BIN-PU, to address this limitation. BIN-PU generates pseudo positive and negative labels from known positive interaction data, enabling effective training of deep learning models for CPI prediction. We also propose a weighted positive loss function that weights to truly positive samples. We have validated BIN-PU coupled with multiple CPI backbone models, comparing the performance with the existing PU models using bacterial cytochrome P450 (CYP) data. Extensive experiments demonstrate the superiority of BIN-PU over the benchmark models in predicting CPIs with only truly positive samples. Furthermore, we have validated BIN-PU on additional bacterial proteins obtained from literature review, human CYP datasets, and uncurated data for its reproducibility. We have also validated the CPI prediction for the uncurated CYP data with biological and biophysical experiments. BIN-PU represents a significant advancement in CPI prediction for bacterial proteins, opening new possibilities for improving predictive models in related biological interaction tasks. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at https://github.com/datax-lab/CYP.

Bacterial Proteins

CROP: a feature-independent context-aware method for CRISPR-Cas9 frameshift prediction.

MOTIVATION: The CRISPR-Cas9 complex has revolutionized genome-editing technologies. By designing a 20&#x2009;nt-long guide RNA, a Cas9 nuclease can be guided to cleave almost any genomic target site (followed by NGG). The cleavage induces double-stranded DNA breaks, which are then repaired by cellular pathways. Accurate CRISPR-Cas9 repair-outcome prediction is essential for designing guide RNAs with desired genomic effects, such as gene knockout. A central challenge is quantifying the rate of frameshifts, i.e. repair-outcomes that lead to a change in the local length that is not a multiple of three. Previous methods for frameshift-rate prediction were trained on only a few experimental or cellular contexts, mostly relied on manually defined microhomology features, and were limited by sparse features and class labels. RESULTS: We developed CROP, a feature-independent context-aware repair-outcome prediction method. By aggregating specific repair outcomes as &#x394;length classes, CROP overcomes class sparsity. We designed CROP to work with variable input sequence lengths and output classes to utilize multiple datasets simultaneously. We benchmarked CROP against state-of-the-art repair-outcome prediction methods over 18 datasets, which we curated and standardized from various studies. Across all datasets, CROP outperformed all competing methods in frameshift-rate prediction. We performed cross-experiment and cross-cellular frameshift-rate predictions to investigate the generalizability of repair mechanisms. Finally, we show that CROP learned microhomology principles from raw sequences without explicit feature engineering, establishing an end-to-end architecture for CRISPR-Cas9 repair-outcome prediction that learns from multiple datasets. AVAILABILITY AND IMPLEMENTATION: CROP is available at https://github.com/OrensteinLab/CROP.

CRISPR-Cas Systems