PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Model performance”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Evaluating Language Models for Biomedical Fact-Checking: A Benchmark Dataset for Cancer Variant Interpretation Verification.

Accurate interpretation of genomic variants is critical for precision oncology but remains slow and dependent on specialized expertise. Public knowledgebases such as the Clinical Interpretation of Variants in Cancer (CIViC) help by curating literature-backed variant interpretations in a structured form, yet verification and review have become major bottlenecks. To address this, we developed CIViC-Fact, a benchmark dataset and pipeline for testing automated systems that verify the accuracy of cancer variant claims. CIViC-Fact links structured claims to sentence-level supporting or refuting evidence from full-text articles, and includes expert annotations and explanations. We evaluated multiple language models. Proprietary models performed well without training, but a smaller open-source model, fine-tuned on CIViC-Fact, achieved the highest accuracy (89%). Applying our fact-checking pipeline to real CIViC entries showed that reviewing less than 20% of content, focusing on flagged entries, would be sufficient to catch over half of all errors. This AI-assisted triage greatly accelerates the review process without replacing or reducing expert insight, ensuring that existing careful oversight remains in place while curators can work more efficiently. CIViC-Fact provides a realistic, high-consequence framework for biomedical fact-checking and a path toward more rigorous and efficient knowledgebase curation.

Journal Article↗

A comparative model: reaction time performance in sleep-disordered breathing versus alcohol-impaired controls.

OBJECTIVES/HYPOTHESIS: Patients with sleep-disordered breathing have reaction time deficits that may lead to catastrophic accidents and loss of life. Although safety guidelines do not exist for unsafe levels of sleepiness, they have been established for unsafe levels of alcohol consumption. Since reaction time performance is altered in both, we prospectively used seven measures of reaction time performance as a comparative model in alcohol-challenged normal subjects with corresponding measures in subjects with sleep-disordered breathing. STUDY DESIGN: Institutional Review Board-approved, nonrandomized prospective controlled study. METHODS: Eighty healthy volunteers (29.1+/-7.5 y of age, 56.3% female subjects) performed four reaction time trials using a psychomotor test at baseline and at three subsequent rising alcohol-influenced time points. The same test without alcohol was given to 113 subjects (47.2+/-10.8 y of age, 19.3% female subjects) with mild to moderate sleep-disordered breathing. RESULTS: Mean blood alcohol concentrations (BACs) in the alcohol-influenced subjects at baseline and three trials were 0, 0.057, 0.080, and 0.083 g/dL. The sleep-disordered subjects had mean respiratory disturbance indices of 29.2 events per hour of sleep. On all seven reaction time measures, their performance was worse than that of the alcohol subjects when BACs were 0.057 g/dL. For three of the measures, the sleep-disordered subjects performed as poorly as or worse than the alcohol subjects when alcohol levels were 0.080 g/dL. These results could not be explained by sex or age differences. CONCLUSION: The data demonstrate that sleep-disordered subjects in this study (with a mean age of 47 y) with mild to moderate sleep-disordered breathing had worse test reaction time performance parameters than healthy, nonsleepy subjects (with a mean age of 29 y) whose BAC is illegally high for driving a commercial motor vehicle in California. This comparative model points out the potential risks of daytime sleepiness in those with sleep-disordered breathing relative to a culturally accepted standard of impairment.

Adult↗

Quality over quantity: biopsy-anchored CT radiogenomics models outperform all-lesion training in a multi-tumour cohort despite a smaller sample size.

OBJECTIVE: Radiogenomics aims to non-invasively predict tumour genotypes from imaging, but most studies assume molecular homogeneity by assigning a single biopsy-derived label to all lesions within a patient. This approach risks substantial label noise given well-documented interlesional heterogeneity. We investigated whether anchoring training to biopsy-confirmed lesions improves radiogenomic model performance and generalisability. MATERIALS AND METHODS: We retrospectively analysed 1646 patients (11473 segmented lesions) with contrast-enhanced CT and EGFR mutation status from next-generation sequencing at the Netherlands Cancer Institute, alongside an external NSCLC radiogenomics cohort (n = 158). All visible lesions were segmented, and the exact biopsy site was matched to its segmentation. Radiomic features were extracted, and machine learning models were trained with three lesion selection strategies: all lesions, non-biopsied lesions only, and biopsy-confirmed lesions only. To disentangle label quality from sample size, we created size-matched variants (one lesion per patient) for all-lesion and non-biopsied strategies. RESULTS: All models achieved significant discrimination of EGFR status on internal validation (AUC = 0.62-0.68). However, performance of the all-lesion and non-biopsied models declined on external validation (AUC = 0.55-0.63), while the biopsy-anchored model maintained stable performance (AUC = 0.62), despite having only 1/10th of the training sample size. When training sets were size-matched, the biopsy-anchored approach significantly outperformed a model trained on all available lesions on external validation (p = 0.037). CONCLUSIONS: Radiogenomic models trained on biopsy-confirmed lesions outperform conventional all-lesion strategies in external validation, despite using an order of magnitude fewer samples. Prioritising lesion-level label fidelity can mitigate heterogeneity-driven noise, enhancing robustness and clinical translation of imaging-based genomic prediction. KEY POINTS: Question Does assigning biopsy-derived molecular labels to all lesions introduce heterogeneity-driven label noise that reduces the generalisability of radiogenomic models? Findings Models trained exclusively on biopsy-confirmed lesions demonstrated superior external generalisability compared with all-lesion approaches, despite being trained on substantially fewer samples. Clinical relevance Biopsy-anchored radiogenomics improves the reliability of non-invasive mutation prediction by accounting for tumour heterogeneity, potentially supporting clinical decision-making when tissue sampling is limited or molecular results are discordant across lesions.

Humans↗

A model to describe growth patterns of the mammary gland during pregnancy and lactation.

Extensive proliferation and death of cells in the mammary gland occur during pregnancy and lactation. In this study, a mechanistic model was developed that yielded a single equation to describe the pattern of mammary growth of mammals throughout pregnancy and lactation. The model contains a single pool, which is the cell population of the mammary gland; one influx, representing cell proliferation; and one efflux, representing cell death. The parameters of the equation lend themselves to direct physiological interpretation. The model fitted data on mammary gland DNA adequately and can be related to current knowledge on factors and inhibitors of mammary gland growth. A unique definition of the parameters of the model can be difficult because of the high degree of variation among animals, an improper number of observations, or timing, as indicated by analyses of simulated data. The model can also be applied to the study of the entire lactation curve. The widely applied gamma equation and the equation that was developed in this study were compared using weekly production data from dairy cows. The new model performed well, particularly when a sharp peak in milk production occurred. The model has the advantage of providing, for the first time, a simple biological description of the lactation curve that can be used to discriminate changes in lactational performance that are associated with experimental treatments.

Animals↗

A model of human performance on the traveling salesperson problem.

A computational model is proposed of how humans solve the traveling salesperson problem (TSP). Tests of the model are reported, using human performance measures from a variety of 10-, 20-, 40-, and 60-node problems, a single 48-node problem, and a single 100-node problem. The model provided a range of solutions that approximated the range of human solutions and conformed closely to quantitative and qualitative characteristics of human performance. The minimum path lengths of subjects and model deviated by average absolute values of 0.0%, 0.9%, 2.4%, 1.4%, 3.5%, and 0.02% for the 10-, 20-, 40-, 48-, 60-, and 100-node problems, respectively. Because the model produces a range of solutions, rather than a single solution, it may find better solutions than some conventional heuristic algorithms for solving TSPs, and comparative results are reported that support this suggestion.

Algorithms↗

A test of the COPE model on motor performance and affect.

The effect of a model called COPE which suggests the use of cognitive-behavioral strategies in response to acute stress in sport was tested. Both rotary pursuit performance and changes in affect were similar for 33 subjects in groups who used the COPE model or used only a segment (one strategy) of the model. Both experimental groups performed better and experienced less negative affect after treatment than the control group.

Adaptation, Psychological↗

Inter-study variability in population pharmacokinetic meta-analysis: when and how to estimate it?

Population pharmacokinetic analysis is being increasingly applied to individual data collected in different studies and pooled in a single database. However, individual pharmacokinetic parameters may change randomly from one study to another. In this article, we show by simulation that neglecting inter-study variability (ISV) does not introduce any bias for the fixed parameters or for the residual variability but may result in an overestimation of inter-individual (IIV) variability, depending on the magnitude of the ISV. Two random study-effect (RSE) estimation methods were investigated: (i) estimation, in a single step, of the three-nested random effects (inter-study, inter-individual and residual variability); (ii) estimation of residual variability and a mixture of ISV and IIV in the first step, then separation of ISV from IIV in the second. The one-stage RSE model performed well for population parameter assessment, whereas, the two-stage model yielded good estimates of IIV only with a rich sampling design. Finally, irrespective of the method used, ISV estimates were valid only when a large number of studies was pooled. The analysis of one real data set illustrated the use of an ISV model. It showed that the fixed parameter estimates were not modified, whether an RSE model was used or not, probably because of the homogeneity of the experimental designs of the studies, and suggest no study-effect in this example.

Computer Simulation↗

Noninvasive determination of the location and distribution of DNAPL using advanced seismic reflection techniques.

Recent advances in seismic reflection amplitude analysis (e.g., amplitude versus offset-AVO, bright spot mapping) technology to directly detect the presence of subsurface DNAPL (e.g., CCl4) were applied to 216-Z-9 crib, 200 West Area, DOE Hanford Site, Washington. Modeling to determine what type of anomaly might be present was performed. Model results were incorporated in the interpretation of the seismic data to determine the location of any seismic amplitude anomalies associated with the presence of high concentrations of CCl4. Seismic reflection profiles were collected and analyzed for the presence of DNAPL. Structure contour maps of the contact between the Hanford fine unit and the Plio/Pleistocene unit and between the Plio/Pleistocene unit and the caliche layer were interpreted to determine potential DNAPL flow direction. Models indicate that the contact between the Plio/Pleistocene unit and the caliche should have a positive reflection coefficient. When high concentrations of CCl4 are present, the reflection coefficient of this interface displays a noticeable positive increase in the seismic amplitude (i.e., bright spot). Amplitude data contoured on the Plio/Pleistocene-caliche boundary display high values indicating the presence of DNAPL to the north and east of the crib area. The seismic data agree well with the well control in areas of high concentrations of CCl4.

Carbon Tetrachloride↗

Machine learning-integrated multi-omics risk prediction for pulmonary fungal infection in COPD and lung cancer: a transcriptomic and immune profiling study.

BACKGROUND: Chronic obstructive pulmonary disease (COPD) and lung cancer are major risk factors for invasive pulmonary fungal infection (IPFI), carrying an attributable mortality of 30%-80%. Their coexistence further amplifies immunosuppression, while current diagnostic criteria remain inadequate for early risk identification. METHODS: Transcriptomic data from the GEO dataset GSE296912 (scRNA-seq; 12,078 cells from normal and COPD lung tissue) and The Cancer Genome Atlas (TCGA)-lung adenocarcinoma (LUAD) bulk RNA-seq cohort (539 tumor and 59 normal samples) underwent differential expression and cross-omics integration analysis. Five machine learning models were constructed: logistic regression, SVM, random forest, XGBoost, and LASSO. Candidate genes were validated by qRT-PCR in A549 cells and THP-1-derived macrophages stimulated with heat-inactivated Aspergillus fumigatus conidia, a protocol selected to ensure BSL-2 biosafety compliance and isolate PAMP-mediated innate immune signaling. Model performance was evaluated using 5-fold stratified cross-validation with AUC, calibration curves, and decision curve analysis. RESULTS: Single-cell transcriptomic analysis of 12,078 cells identified 14 distinct cell populations, with marked myeloid expansion and immune dysregulation in COPD lung tissue. Cross-omics integration with TCGA-LUAD data identified 1,145 shared genes (79 immune-related), converging on NF-κB, TLR4, and cytokine receptor signaling. The random forest model achieved excellent discriminative performance (5-fold CV AUC = 0.988), with Treg infiltration, TLR4, and MMP9 as the top predictors. qRT-PCR confirmed significant upregulation of all five candidate genes (DEFB4A, S100A8, IL-8, MMP9, and TLR4) in both A549 and THP-1 cells following fungal stimulation. CONCLUSION: This multi-omics machine learning model integrating scRNA-seq and TCGA transcriptomic data demonstrates excellent discriminative performance (AUC = 0.988), with mechanistic convergence of NF-κB, TLR4, and oncogenic signaling pathways identified across shared immune gene signatures. In vitro qRT-PCR validation confirms the biological relevance of five key antifungal immune genes, providing a transcriptomic foundation for future prospective IPFI risk stratification in patients with COPD and lung cancer.

TLR4↗

Explorations of Cohen, Dunbar, and McClelland's (1990) connectionist model of Stroop performance.

The J. D. Cohen, K. Dunbar, and J. L. McClelland (1990) model of Stroop task performance is used to model data from a study by D. H. Spielder, D. A. Balota, and M. E. Faust (1996). The results indicate that the model fails to capture overall differences between word reading and color naming latencies when set size is increased beyond 2 response alternatives. Further empirical evidence is presented that suggests that the influence of increasing response set size in Stroop task performance is to increase the difference between overall color naming and word reading, which is in direct opposition to the decrease produced by the Cohen et al. architecture. Although the Cohen et al. model provides a useful description of meaning-level interference effects, the qualitative differences between word reading and color naming preclude a model that uses identical architectures for each process, such as that of Cohen et al., to fully capture performance in the Stroop task.

Aged↗

Verbalization in EMR children's observational learning.

The effect of descriptive verbalization during observation of a model on mentally retarded boys' retention for what they had observed was examined. Forty 9- to 12-year-old boys in public-school EMR classes were grouped on the basis of relatively high or low IQ scores. One-half of each group observed a videotaped model perform a series of novel acts, while in addition to viewing the tape, the other half described the model's actions. Observational learning was immediately tested through a set of prompts for imitation, with prizes offered commensurate with level of performance. Regardless of IQ group, the boys who were required to verbalize the model's behavior were able to imitate it significantly better than boys who merely watched the model; high and low IQ groups did not significantly differ in observational learning. Further directions for research on mentally retarded children's observational learning were suggested.

Attention↗

Comparison of the performance of two comorbidity measures, with and without information from prior hospitalizations.

OBJECTIVES: This study compares the performance of two comorbidity risk adjustment methods (the Deyo et al adaptation of the Charlson index and the Elixhauser et al method) in five groups of California hospital patients with common reasons for hospitalization, and assesses the contribution to model performance made by information drawn from prior hospital admissions. METHODS: California hospital discharge abstract data for the calendar years 1994 through 1997 were used to create a longitudinal data set for patients in the five disease groups. Eleven logistic regression models were estimated to predict the risk of in-hospital death for patients in each group, with both comorbidity risk adjustment methods applied to patient information available from only the index hospitalization, and to information available from both the index and prior hospitalizations. RESULTS: For every comparison made, the level of statistical performance (area under the receiver operating characteristics curve) demonstrated by models using the Elixhauser et al method was superior to that of models using the Deyo et al adaptation method. Although most patients have information available from prior hospital admissions, this additional information yields only small improvements in the performance of models using either comorbidity risk adjustment method. CONCLUSIONS: Better discrimination is achieved with the Elixhauser et al method using only information from the index hospitalization than is achieved with the Deyo et al adaptation using information from all identified hospital admissions. Both comorbidity risk adjustment methods achieve their best performance when information from the index hospitalization and prior admissions is separated into independent indicators of comorbid illness.

Adult↗

An evaluation of the predictive performance of distributional models for flora and fauna in north-east New South Wales.

To use models of species distributions effectively in conservation planning, it is important to determine the predictive accuracy of such models. Extensive modelling of the distribution of vascular plant and vertebrate fauna species within north-east New South Wales has been undertaken by linking field survey data to environmental and geographical predictors using logistic regression. These models have been used in the development of a comprehensive and adequate reserve system within the region. We evaluate the predictive accuracy of models for 153 small reptile, arboreal marsupial, diurnal bird and vascular plant species for which independent evaluation data were available. The predictive performance of each model was evaluated using the relative operating characteristic curve to measure discrimination capacity. Good discrimination ability implies that a model's predictions provide an acceptable index of species occurrence. The discrimination capacity of 89% of the models was significantly better than random, with 70% of the models providing high levels of discrimination. Predictions generated by this type of modelling therefore provide a reasonably sound basis for regional conservation planning. The discrimination ability of models was highest for the less mobile biological groups, particularly the vascular plants and small reptiles. In the case of diurnal birds, poor performing models tended to be for species which occur mainly within specific habitats not well sampled by either the model development or evaluation data, highly mobile species, species that are locally nomadic or those that display very broad habitat requirements. Particular care needs to be exercised when employing models for these types of species in conservation planning.

Animal Population Groups↗

Prognostic models based on literature and individual patient data in logistic regression analysis.

Prognostic models can be developed with multiple regression analysis of a data set containing individual patient data. Often this data set is relatively small, while previously published studies present results for larger numbers of patients. We describe a method to combine univariable regression results from the medical literature with univariable and multivariable results from the data set containing individual patient data. This 'adaptation method' exploits the generally strong correlation between univariable and multivariable regression coefficients. The method is illustrated with several logistic regression models to predict 30-day mortality in patients with acute myocardial infarction. The regression coefficients showed considerably less variability when estimated with the adaptation method, compared to standard maximum likelihood estimates. Also, model performance, as distinguished in calibration and discrimination, improved clearly when compared to models including shrunk or penalized estimates. We conclude that prognostic models may benefit substantially from explicit incorporation of literature data.

Age Factors↗

Modelling the growth kinetics of Phanerochaete chrysosporium in submerged static culture.

The potential commercial application of Phanerochaete chrysosporium requires methods for quantitatively predicting growth and substrate utilization. The growth kinetics of P. chrysosporium INA-12 (CNCM I-398) were investigated and modelled under nonlimiting nitrogen and carbon conditions in submerged static culture. This strain, unlike other strains, does not require nutrient limitation for induction of lignin peroxidase. Maximum levels of lignin peroxidase activity were reached 7 days after culture initiation, when almost 80% of the initial glycerol and 70% of the initial nitrogen were still present. Lignin peroxidase levels then decreased, while biomass levels increased until about day 14. The ratio of cell dry weight to wet weight was constant until the maximum biomass concentration was achieved, after which there was a decrease in the water content. The change in this ratio reflects cell lysis as it correlated with increased concentrations of nitrogen in the media, arising from cell leakage. The suitability of four growth models to predict growth, and in some cases glycerol consumption, was evaluated. A simple linear model and the Emerson model performed poorly for the early stages of growth, while a modified Williams model and the Monod model predicted substrate and biomass concentrations equally well. All models will predict biomass concentrations during the active growth phase, but they should not be used to predict biomass concentrations after the stationary growth phase, when cell lysis becomes significant.

Basidiomycota↗

Predicting duration of sleep from the three process model of regulation of alertness.

OBJECTIVES: Irregular working hours severely disturb sleep and wakefulness. This paper presents a modification of the quantitative (computerised) three process model of regulation of alertness to predict duration of sleep in connection with irregular sleep patterns. METHODS: The model uses a circadian "C" (sinusoidal) and homeostatic "S" (exponential) component (the duration of previous periods awake and asleep), which are summed to yield predicted alertness (on a scale of 1-16). It assumes that waking from sleep will occur at a given alertness level (S' + C') when recuperation is complete. Variables of electroencephalographic duration of sleep from two studies of irregular sleep were used to model the S and C variables in a regression approach to maximise prediction. The model performance was cross validated against published field and laboratory data. RESULTS: The model parameters were defined with a high degree of precision R2 = 0.99 and the validation yielded similar values R2 = 0.98-0.95, depending on the acrophase. The paper also describes a simplified graphical version of the computation model seen as a two dimensional duration of sleep nomogram. CONCLUSION: The model seems to predict group means for duration of sleep with high precision and may serve as a tool for evaluating work and rest schedules to reduce risks of sleep disturbances.

Adult↗

Aggregation bias and the use of regression in evaluating models of human performance.

Regression analyses are increasingly being used to provide confirmatory evidence for models of human performance. The amount of information made available to judge these models is reduced because clearly established standards in the techniques of performing and reporting regression analyses are lacking. This paper addresses two primary problems in regression analysis: aggregation of data and the aggregation of variables into composite models. We provide examples of the misuse of regression techniques and recommend ways in which the amount of information made available to evaluate the model being tested can be maximized in analysis and reporting.

Bias↗

Diagnostic Performance of Machine Learning for Systemic Lupus Erythematosus: Systematic Review and Meta-Analysis.

BACKGROUND: Early and accurate diagnosis of systemic lupus erythematosus (SLE) and its organ involvement is essential. Previous reviews of machine learning (ML) in SLE combined heterogeneous tasks and validation strategies and may have overinterpreted model performance. OBJECTIVE: This study evaluated the diagnostic performance of ML and deep learning (DL) models for 3 clinically distinct SLE-related tasks: SLE classification or diagnosis, lupus nephritis (LN) diagnosis, and neuropsychiatric systemic lupus erythematosus (NPSLE) discrimination. We also assessed methodological quality and certainty of evidence. METHODS: PubMed, Embase, Cochrane Library, Web of Science, and IEEE Xplore were searched from January 2014 to April 2026. Eligible peer-reviewed diagnostic accuracy studies developed or validated ML or DL models for 1 of the 3 prespecified tasks, used an accepted reference standard, and provided data for a 2×2 contingency table. Bivariate random-effects meta-analyses with the Hartung-Knapp-Sidik-Jonkman adjustment were used to pool sensitivity and specificity. We reported 95% prediction intervals (PIs), assessed risk of bias using the Quality Assessment of Diagnostic Accuracy Studies for Artificial Intelligence tool (QUADAS-AI; Viknesh Sounderajah [Imperial College London]), and evaluated certainty of evidence using the Grading of Recommendations Assessment, Development, and Evaluation framework for diagnostic test accuracy. RESULTS: Twenty-nine studies were included: 17 for SLE classification, 5 for LN diagnosis, and 7 for NPSLE discrimination. In the primary task-stratified analysis, pooled sensitivity was 0.91 (95% CI 0.86-0.94; 95% PI 0.56-0.99), and pooled specificity was 0.94 (95% CI 0.91-0.96; 95% PI 0.69-0.99), with low heterogeneity (I²=23.9% and 22.9%, respectively). DL models showed a sensitivity of 0.93 and specificity of 0.95, compared with 0.88 and 0.94 for traditional ML models. Certainty of evidence was high for most analyses but low for LN diagnosis because of inconsistency and imprecision. All studies were retrospective, and only 9 of 29 (31%) performed independent external validation. Overall risk of bias was high or unclear in 22 of 29 (75.9%) studies. No study reported model calibration, decision-curve analysis, or net clinical benefit. CONCLUSIONS: ML models showed promising diagnostic accuracy across 3 distinct SLE-related tasks, but wide PIs, limited external validation, and pervasive risk of bias restrict conclusions about real-world generalizability. Prospective multicenter studies with standardized tasks and reference standards, independent external validation, and formal assessment of calibration and clinical utility are required before clinical implementation.

Humans↗