PubMed HealthSearch

SEARCH · PubMed Health

Results for “Model performance”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Some effects of a model's performance on an observer's electromyographic activity.

It is suggested that motor reactions may be elicited in an observer as a consequence of his exposure to a model and that such reactions may become conditioned to environmental events. An experiment is reported in which observers showed greater EMG activity in the arm while watching models arm wrestle than while watching a model stutter, and greater lip EMG activity while watching a model stutter than while watching arm wrestling. Some evidence for conditioning was found in the arm activity of males watching wrestling.

Adolescent

Habitat radiomics predicts occult lymph node metastasis and uncovers immune microenvironment of head and neck cancer.

BACKGROUND: Occult lymph node metastasis (LNM) is a key prognostic factor for patients with head and neck squamous cell carcinoma (HNSCC). This study was to establish radiomics models derived from intratumoral, peritumoral, and habitat regions for identifying occult LNM in HNSCC. METHODS: Patients with pathologically confirmed HNSCC from three medical Centers (from March 2014 to April 2024) and The Cancer Genome Atlas (TCGA) were enrolled. Center 1 was split into training (n = 330) and internal test sets (n = 154), while Center 2 and Center 3 served as the external test set (n = 183). Genomic set (n = 50) from TCGA and single-cell RNA sequencing set (n = 6) from Center 1 were used for biological analysis. We used the intratumoral, peritumoral, and habitat volumes of interest (VOIs) to extract radiomics features, respectively. Based on Logistic Regression (LR), Support Vector Machine (SVM), and Random Forest (RF) classifiers, nine radiomics models were built to confirm the optimal predictive performance. The best-performing model, along with clinical-radiologic data, was combined to develop a hybrid model. The log-rank test was used to evaluate the model's prognostic performance. Additionally, bulk and single-cell RNA sequencing were applied for investigating the biological mechanisms underlying the optimal model. RESULTS: The RF-habitat radiomics model showed the best performance, achieving AUCs of 0.835-0.919 across all datasets. Survival analysis further confirmed the prognostic value of the RF-habitat radiomics model. The RF-habitat radiomics model and the hybrid model notably surpassed the clinical model in predictive performance. Moreover, the RF-habitat radiomics model was associated with the abundance level of exhaustion-associated CD8 + T cells, uncovering the immune microenvironment characteristics contributing to occult LNM in HNSCC. CONCLUSIONS: The RF-habitat radiomics model demonstrated excellent performance for predicting occult LNM in HNSCC across three cohorts, providing a non-invasive solution for occult LNM. Furthermore, radiogenomic analysis further revealed the biological associations of the model, primarily related to T cell dysfunction.

Humans

Modeling of immunosensors under nonequilibrium conditions. I. Mathematic modeling of performance characteristics.

Immunosensors for the detection of small analytes that use analyte-enzyme conjugates as signal generators require special attention if operated under nonequilibrium conditions. If the size of the analyte and the analyte-enzyme conjugate differ substantially, the two antigens do not diffuse at the same rate. This can cause time-dependent shifts in the sensitivity of competitive immunoassays. Therefore, immunosensors operating at short incubation times require precise timing that meets closely the specifications for which the sensors were calibrated. As an example, we have analyzed kinetic binding curves for the quantitative determination of progesterone with an immobilized monoclonal antibody and a conjugate between horseradish peroxidase and progesterone as signal generator. Mathematical paradigms have been developed to simulate the diffusion, antigen-antibody complex formation, and competitive binding processes in this analytical system. Dose-response curves obtained under nonequilibrium conditions can vary substantially from those obtained at equilibrium of antigen-antibody interaction. The degree of this variation depends on the performance characteristics of the major components of the immunosensor. The developed mathematical solutions reflect experimental results and can be used to model optimal conditions for immunosensors operating under nonequilibrium conditions. In this paper (Part I), we report on the mathematical modeling of the interaction between analyte, analyte-enzyme conjugate, and an immobilized antibody. In Part II (W. Schramm and S.-H. Paek (1991) Anal. Biochem. 196), we present experimental results and compare them with the theoretical models.

Antibodies

Artificial intelligence in kidney cancer: a review of clinical applications across the disease spectrum.

PURPOSE OF REVIEW: This review examines recent advances (2024-2025) in the application of artificial intelligence (AI) to kidney cancer diagnosis, prognosis, and treatment planning. It categorizes studies across 13 clinical scenarios to assess where AI offers the most clinical utility. RECENT FINDINGS: AI models have demonstrated strong performance in a range of tasks including tumor grading, subtype classification, survival prediction, and risk stratification. Integration of radiomics, genomics, and histopathology has enabled personalized, noninvasive, and timely decision-making. The highest-performing models used CT-based radiomics, particularly for predicting progression-free and recurrence-free survival. However, performance varies across tasks and tumor subtypes, with lower accuracy in detecting oncocytomas or benign vs. malignant differentiation. AI applications in metastatic and nonresected cases remain underexplored, and ultrasound remains a largely under researched modality. While some models improve diagnostic accuracy and workflow efficiency, broader validation across diverse populations is still needed. SUMMARY: AI is transforming kidney cancer care across multiple clinical stages. Although promising, real-world implementation demands ongoing validation and postdeployment monitoring to prevent performance degradation due to distributional drift. AI's integration with multimodal data offers substantial potential to improve outcomes and reduce overtreatment.

Humans

Effect of graded reductions of coronary pressure and flow on myocardial metabolism and performance: a model of "hibernating" myocardium.

The term "hibernating" myocardium has been applied to chronic left ventricular dysfunction without angina or ischemic electrocardiographic changes in patients with coronary artery disease that is reversed by therapy that increases myocardial blood flow. To investigate the relation between coronary blood flow and ventricular function experimentally, graded reductions in coronary artery pressure were produced in isolated perfused rat hearts as contractile performance (peak systolic pressure and its first derivative [dP/dt]) and metabolic variables were measured using phosphorus-31 nuclear magnetic resonance (NMR) spectroscopy. As coronary pressure and flow were reduced, significant reductions in myocardial oxygen consumption and contractile performance were observed, which returned to control levels when coronary artery pressure and flow were restored to baseline values. Two phases of metabolic abnormality were observed. With modest reductions in coronary perfusion, proportionate reductions in myocardial oxygen consumption and contractile behavior were accompanied by a slight reduction in creatine phosphate but no significant lactate production. With greater reductions in coronary artery pressure and flow, creatine phosphate decreased more, adenosine triphosphate levels and myocardial pH decreased significantly and myocardial lactate production increased. The balanced reductions in myocardial contractility and oxygen consumption without metabolic abnormalities traditionally associated with "ischemia" observed in the first phase provides evidence in normal hearts for resetting of the myocardial contractile behavior and oxygen consumption in the presence of reduced coronary flow (that is, hibernating myocardium). The data suggest that reductions in adenosine diphosphate and the index of the reduced form of nicotinamide adenine dinucleotide (NADH) (lactate formation) do not explain the coupling between coronary artery pressure and flow and myocardial oxygen consumption as contractile performance decreases.

Adenosine Diphosphate

Finerenone for Heart Failure and Risk Estimated by the PREDICT-HFpEF Model: A Secondary Analysis of FINEARTS-HF.

IMPORTANCE: Patients with heart failure (HF) and mildly reduced ejection fraction (HFmrEF) or preserved ejection fraction (HFpEF) have a spectrum of risk, and the effect of therapies may vary by risk. OBJECTIVES: To validate the Prognostic Models for Mortality and Morbidity in HFpEF (PREDICT-HFpEF) in the phase 3 randomized clinical trial Finerenone Trial to Investigate Efficacy and Safety Superior to Placebo in Patients With Heart Failure (FINEARTS-HF) and to evaluate the effect of finerenone, compared with placebo, across the spectrum of risk in these patients. DESIGN, SETTING, AND PARTICIPANTS: The FINEARTS-HF trial was conducted across 653 sites in 37 countries. Participants were adults 40 years and older with symptomatic HF and left ventricular EF of 40% or greater randomized between September 2020 and January 2023. INTERVENTION: Finerenone (titrated to 20 mg or 40 mg) or placebo. MAIN OUTCOMES AND MEASURES: The 3 PREDICT-HFpEF risk scores for the composite outcome of cardiovascular death or HF hospitalization, cardiovascular death, and all-cause death, respectively, were calculated. Predicted risk was compared with observed outcomes. Model performance was assessed using the Harrell C statistic. The rates of the predicted outcomes (plus the composite of cardiovascular death and worsening HF events, which was the primary end point in the trial) were examined according to quintiles of risk score, as was the effect of finerenone according to risk quintiles. RESULTS: A total of 6001 patients (mean [SD] age, 72 [9.6] years; 3269 male [54.5%]) were randomized in the FINEARTS-HF trial. The C statistics for cardiovascular death or HF hospitalization, cardiovascular death, and all-cause death at 2 years were 0.71 (95% CI, 0.69-0.72), 0.68 (95% CI, 0.66-0.71), and 0.69 (95% CI, 0.67-0.71), respectively. The risk of the composite outcomes was approximately 8- to 10-fold higher in those in the highest compared with the lowest risk quintile. The relative risk reduction with finerenone compared with placebo was consistent across the spectrum of risk for all outcomes examined (eg, interaction P value for primary outcome = .24). CONCLUSIONS AND RELEVANCE: Results of the FINEARTS-HF randomized clinical trial demonstrate that the PREDICT-HFpEF models performed well in terms of calibration and discrimination. Baseline risk did not modify the benefit of finerenone. TRIAL REGISTRATION: ClinicalTrials.gov Identifier: NCT04435626.

Aged

Comprehensive Evaluation and Explainable Interpretation of Peptide-HLA Binding Prediction Tools.

Accurate prediction of peptide binding to human leukocyte antigen class I (HLA-I) molecules is critical for advancing immunological research, particularly in vaccine design and immunotherapy. However, limitations in model performance, interpretability, and dataset quality impede the widespread adoption of existing predictive tools. Here, we present a comprehensive evaluation of 17 HLA-I peptide binding prediction models, utilizing a meticulously curated dataset comprising over 290,000 peptides spanning 44 HLA-I alleles. We assessed model accuracy, robustness, and interpretability, employing explainability techniques such as SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model-agnostic Explanations) to elucidate underlying prediction mechanisms. Our results reveal substantial performance disparities, with self-attention-based models, including STMHCpan and BigMHC, exhibiting superior accuracy. Notably, the capsule network model CapsNet-MHC_AN demonstrated robust performance. Models trained on eluted ligand datasets outperformed those relying on binding affinity data, underscoring the critical role of high-quality training data. Ensemble and multi-algorithm approaches further improved prediction reliability. These findings highlight the need for ongoing innovation in model architecture, integration of diverse and high-quality datasets, and incorporation of structural predictors to develop more accurate, interpretable, and clinically applicable HLA-I peptide binding prediction tools.

HLA-I binding

Evaluating the impact of modeling choices on the performance of integrated genetic and clinical models.

The value of genetic information for improving the performance of clinical risk prediction models has yielded variable conclusions. Many methodological decisions have the potential to contribute to differential results across studies. Here, we performed multiple modeling experiments integrating clinical and demographic data from electronic health records (EHR) and genetic data to understand which decision points may affect performance. Clinical data in the form of structured diagnostic codes, medications, procedural codes, and demographics were extracted from two large independent health systems and polygenic risk scores (PRS) were generated across all patients with genetic data in the corresponding biobanks. Crohn's disease was used as the model phenotype based on its substantial genetic component, established EHR-based definition, and sufficient prevalence for model training and testing. We investigated the impact of PRS integration method, as well as choices regarding training sample, model complexity, and performance metrics. Overall, our results show that including PRS resulted in higher performance by some metrics but the gain in performance was only robust when combined with demographic data alone. Improvements were inconsistent or negligible after including additional clinical information. The impact of genetic information on performance also varied by PRS integration method, with a small improvement in some cases from combining PRS with the output of a clinical model (late-fusion) compared to its inclusion an additional feature (early-fusion). The effects of other modeling decisions varied between institutions though performance increased with more compute-intensive models such as random forest. This work highlights the importance of considering methodological decision points in interpreting the impact on prediction performance when including PRS information in clinical models.

Preprint

Evaluating the impact of modeling choices on the performance of integrated genetic and clinical models.

PURPOSE: The value of genetic information for improving the performance of clinical risk prediction models has yielded variable conclusions. Many methodological decisions have the potential to contribute to differential results. We performed multiple modeling experiments integrating clinical and demographic data from electronic health records with genetic data to understand which decisions may affect performance. METHODS: Clinical data in the form of structured diagnostic codes, medications, procedural codes, and demographics were extracted from 2 large independent health systems, and polygenic risk scores (PRS) were generated across all patients of European ancestry with genetic data in the corresponding biobanks. Crohn's disease was studied based on its substantial genetic component, established electronic health records-based definition, and sufficient prevalence for training and testing. We investigated the impact of choices regarding the PRS integration method, training sample, model complexity, and performance metrics. RESULTS: Overall, our results showed that including PRS resulted in higher performance, but this gain was only robust in situations with limited clinical information. We found consistent performance increases from more compute-intensive models, such as random forest, but the impact of other decisions varied by site. CONCLUSION: This work highlights the importance of considering methodological decision points in interpreting the impact of PRS on prediction performance in clinical models.

Humans

Children's expressions of spatial knowledge.

Different expressions of spatial knowledge were examined by having groups of first-, fourth-, and sixth-grade children perform model construction, verbal description, and route reversal tasks after they learned the correct path through a pedestrian maze. Age-related improvement was found in the rate of learning the maze and in the accuracy of verbal descriptions, suggesting that maze learning may be verbally mediated. All children performed well in sequencing intersections in the model but performed poorly in choosing path options in the model. Route reversal after learning was accurate and equivalent across groups. Overall, results suggest that both general and task-specific skills are involved in different products of spatial knowledge.

Child

Machine learning-based clinical prediction model and multi-omics integration for assessing pancreatic cancer risk in new-onset diabetes.

BACKGROUND: Given that pancreatic cancer (PC) is typically diagnosed at an advanced stage but is often preceded by new-onset diabetes mellitus (NODM), providing a window for early detection, we sought to develop and validate an interpretable machine-learning model integrated with multi-omics profiling to identify early biomarkers of NODM-associated PC. METHODS: In a population-based cohort, individuals with NODM-associated PC and NODM without PC were identified and randomly divided (70:30) into training and validation sets after feature selection. Eight machine learning (ML) classifiers were compared using fivefold cross-validation, and model performance was evaluated in terms of discrimination, calibration, and decision curve–based clinical utility. We evaluated interpretability using the Shapley additive explanations (SHAP) analyses. Mechanistically, Olink proteomic profiling and metabolomics were analyzed through clinical classifications and model-defined risk strata. RESULTS: Categorical boosting achieved the best performance in the independent validation set (AUROC = 0.844). The NODM cohort was stratified into high- (n = 2,362) and low-risk (n = 5,030) groups, and internal validation together with SHAP analyses demonstrated consistent model performance and identified clinically interpretable predictors. Proteomic and metabolomic analyses under clinical and risk-based grouping identified 39 overlapping differentially expressed proteins and 145 overlapping metabolites with enriched across 11 shared KEGG pathways. Cross-platform validation highlighted PLTP, CRTAC1, and ITGAV as serum biomarkers with a strong potential for early NODM-PC detection. CONCLUSIONS: We developed an interpretable ML framework centered on NODM enables practical risk stratification for early PC detection by multi-omics and provides a pathway of ML-based triage followed by biomarker confirmation for earlier detection and diagnosis.

Humans

Simulation of mechanisms of viral interference in influenza.

Biological interference among viral agents might have significant implications for disease prevention and therapy. Field data for influenza yield conflicting evidence concerning the independence of infection rates, or disease severity, for two co-circulating viruses. To examine the effects of several assumed modes of interference for influenza, simulations of a Monte Carlo micropopulation model of influenza epidemics have been performed. Model parameters were selected so that the simulated attack rates for each of two different viral strains matched actual field data. Rates of infection were compared for single agents and for two viruses with only behavioural interference. Other simulations included temporary immunity to the other virus for the duration of the infection, and/or reduced shedding of viral particles for dual infections. Simulated viral competition had little impact on epidemic severity, duration, or size distribution. Under the conditions studied, viral interference in natural populations would be difficult to infer from field observations of attack rates. Other simulations extended a partial immunity and/or reduced viral shedding during an infection with a second virus. These indicated that interference might be suggested by field data, but it could not be demonstrated conclusively. Still other simulations showed that for epidemics with much higher attack rates for both viruses, it would be relatively easy to demonstrate interference. However, in order to observe interference between influenza strains, it would be necessary to monitor on an almost daily basis, using a method of viral detection which would have to be both highly specific and also very sensitive.

Adolescent

Computer simulations of chondrocytic clone behaviour in rabbit growth plates.

The growth behaviour of chondrocytic clones in the cell columns of the proximal tibial growth plates of young rabbits was modelled in computer simulations. Simulations were performed, modelling either clones in large groups of columns or clones in one single column. The former were based on morphological data and measurements of cell columns from an earlier study while the latter utilised previous findings of cellular kinetics in rabbit growth plates. Simulation results that resembled most closely the actual observations on rabbit growth plates were those in which a distribution of values was assumed both for clone length (ranging from 1000 to 2000 microns) and for the lengths of the discontinuities between clones. When the assumption was made in the models that the disappearing (metaphyseal) end of an 'old' clone moved more rapidly than the developing (epiphyseal) end of a 'new' clone, replacing the former, the length of the discontinuity between these two clones increased with time. This assumption, which could be modelled in the simulations of clones in a single column based on cell growth behaviour, was found to provide an explanation for an earlier finding that there are more short columns at the epiphyseal side than at the metaphyseal side of a growth plate.

Animals

Quality of laboratory performance in testing for human immunodeficiency virus type 1 antibody. Variables associated in multivariate analyses.

In May 1988, the Centers for Disease Control's Model Performance Evaluation Program (Atlanta, Ga) surveyed 1092 laboratories that performed enzyme immunoassays and Western blot tests for human immunodeficiency virus type 1 antibody on mailed plasma samples of known human immunodeficiency virus type 1 antibody reactivity and that described their laboratory characteristics and testing practices. The study objective was to evaluate the quality of laboratory performance in testing for human immunodeficiency virus type 1 antibody. After identifying relevant variables in univariate analyses, multivariate analyses were performed using stepwise logistic models. Human immunodeficiency virus type 1 antibody test performance was independently associated with analytic variables such as commercial test kit used and with nonanalytic variables such as experience, training, and degree requirements of laboratory personnel. These results validate the importance of nonanalytic variables to the quality of outcomes in laboratory testing.

AIDS Serodiagnosis

Machine learning-based clinical tool for identifying factors associated with symptomatic knee osteoarthritis: the Nagahama study.

BACKGROUND: A clinical tool that evaluates factors associated with symptomatic knee osteoarthritis (OA) based on modifiable factors is lacking. This study aimed to develop a machine learning-based clinical assessment tool using modifiable factors to identify factors associated with symptomatic knee OA and to determine its accuracy. METHODS: This study included 429 participants (81.8% women; age, 69.0&#xa0;&#xb1;&#xa0;5.3 years) from the Nagahama Study who were &#x2265;60&#xa0;years old and had radiographically confirmed knee OA. A Knee Society Knee Scoring System 2011 symptom score of <23 points defined symptomatic knee OA. Participants were randomly assigned to training (70%) and test (30%) datasets. A machine learning model was developed using Extreme Gradient Boosting with 27 variables, and the SHapley Additive exPlanation (SHAP) values were used to assess feature importance. The top 8 features were translated into a 100-point clinical scoring tool weighted by their SHAP contributions. The cutoff value indicating symptomatic knee OA in the clinical assessment tool was determined using receiver operating characteristic analysis, and model performance was evaluated in both datasets. RESULTS: The clinical assessment tool consisted of low back pain, OA severity, depressive tendencies, knee flexion/extension range of motion, knee extension and hip abduction strength, and lower limb muscle quality. The model showed moderate discriminative performance (AUC 0.771 and 0.773 in the training and test datasets, respectively), with a cutoff point of 47. CONCLUSION: The proposed clinical assessment tool may provide a structured framework for assessing modifiable factors associated with symptomatic knee OA, reflecting their contribution to current symptom status.

Humans