PubMed HealthSearch

SEARCH · PubMed Health

Results for “Model performance”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Gaseous homeostasis and the circle system. Validation of a model.

The performance of a model of a subject breathing from a circle system has been examined in relation to nitrogen and helium. The ability of the model to maintain a nitrogen steady-state breathing air, the attainment of a new steady-state after perturbation of an existing nitrogen equilibrium, the washout of nitrogen from the subject model on breathing oxygen, and the estimation of functional residual capacity using a rebreathing method with helium as an indicator have been assessed. The predictable and accurate performance of the model in these studies, together with its ability to reproduce the results of a number of previously published studies in man, suggest that the model can be used to predict the behaviour of circle systems when used with inhaled anaesthetic agents.

Anesthesia, Inhalation

Perceptual studies on ultrasonic B-scan textures.

A pilot study of the perceptual characteristics of ultrasonic textured images is described. Scans of four models performed on four real-time machines optimised for display of a normal liver were used. A trial with 22 observers indicated that the model that gave images closest to the liver image varied between machines. A second, paired similarity test with five observers using all the model images was performed, with a cluster analysis of a multidimensional scaling procedure. This suggested that the prominent features of the textural images are often more closely related to the machines than to the models. Considerable further work is needed to confirm these pilot results and to identify the visual cues that are most significant in textured images.

Biophysical Phenomena

Modeling static and dynamic human cardiovascular responses to exercise.

A human performance model has been developed and described [9] which portrays the human circulatory, thermo regulatory and energy-exchange systems as an intercoupled set. In this model, steady state or static relationships are used to describe oxygen consumption and blood flow. For example, heart rate (HTRT) is calculated as a function of the oxygen and the thermo-regulatory requirements of each body compartment, using the steady state work values of cardiac output (CO, sum of all compartment blood flows) and stroke volume (SV, assumed maximal after 40% maximal oxygen consumption): HTRT=CO/SV. The steady state model has proven to be an acceptable first approximation, but the inclusion of transient characteristics are essential in describing the overall systems' adjustment to exercise stress. In the present study, the dynamic transient characteristics of heart rate, stroke volume and cardiac output were obtained from experiments utilizing step and sinusoidal forcing of work. The gain and phase relationships reveal a probable first order system with a six minute time constant, and are utilized to model the transient characteristics of these parameters. This approach leads to a more complex model but a more accurate representation of the physiology involved. The instrumentation and programming essential to these experiments are described.

Analog-Digital Conversion

MRI-based radiomics model for predicting VEGFA expression and prognosis in lower-grade glioma.

BACKGROUND: Gliomas are the most common primary tumors of the central nervous system. Their treatment remains highly challenging, with high rates of associated disability and mortality. Conventional prognostic indicators no longer adequately satisfy the clinical demands of precision medicine. Therefore, it is essential to further explore novel prognostic biomarkers to enable accurate risk stratification and to provide new reference indicators for personalized precision therapy. PURPOSES: This study aimed to investigate the prognostic significance of vascular endothelial growth factor A (VEGFA) in patients diag nosed with lower-grade gliomas (LGGs) using an MRI based radiomics model. METHODS: Data regarding VEGFA expression and clinical records of LGG patients were retrieved from The Cancer Genome Atlas (TCGA). Corresponding preoperative MRI data were obtained from The Cancer Imaging Archive (TCIA) for radiomic feature extraction. Patients were stratified into high- and low- VEGFA expression groups based on survival information from the current cohort using the survminer package. The overall survival (OS) was assessed using Kaplan-Meier analysis and Cox proportional hazards regression. Predictive models were developed using logistic regression (LR), and model performance was evaluated via receiver operating characteristic (ROC) curve analysis, with area under the curve (AUC) values reported. An optimized model incorporating the Akaike information criterion (AIC) was also constructed (AIC-LR). RESULTS: VEGFA expression was significantly associated with OS (P = 0.002). Multivariate Cox regression confirmed VEGFA as an independent prognostic factor (hazard ratio [HR] = 2.545, 95% confidence interval: 1.422-4.555). Furthermore, VEGFA expression correlated with immune infiltration levels, particularly of M1 and M2 macrophages and T follicular helper cells, and was associated with enrichment in Wnt signaling and B cell receptor signaling pathways. The LR and AIC-LR models demonstrated acceptable predictive performance, with AUCs of 0.728 (95% CI: 0.612-0.843) and 0.725(95% CI: 0.612-0.839) in the training cohort, and 0.704 (95% CI: 0.562-0.847) and 0.718(95% CI: 0.576-0.861) in the validation cohort, respectively. CONCLUSIONS: The MRI based radiomics model showed potential for noninvasive assessment of VEGFA expression and may provide auxiliary information for prognostic evaluation in LGG. Further validation in larger samples and independent external cohorts is required before clinical application.

Radiomics

Observational learning of a left-right behavioral asymmetry in mice (Mus musculus).

B6D2F1 hybrid mice that were allowed to observe a trained female mouse open a pendulum door to the right (or to the left) to enter a food compartment later solved this problem faster than pupils that had been placed behind a visual barrier. Male pupils that had observed a "left-handed" teacher performed sinistrally; males that had observed a "right-handed" model performed dextrally. Female pupils did not exhibit their demonstrator's laterality. Observational learning may provide a means to maintain certain lateralized behaviors. Such social learning may lead to the emergence of local traditions and to the cultural diffusion of behavioral asymmetries.

Animals

Evaluating Language Models for Biomedical Fact-Checking: A Benchmark Dataset for Cancer Variant Interpretation Verification.

Accurate interpretation of genomic variants is critical for precision oncology but remains slow and dependent on specialized expertise. Public knowledgebases such as the Clinical Interpretation of Variants in Cancer (CIViC) help by curating literature-backed variant interpretations in a structured form, yet verification and review have become major bottlenecks. To address this, we developed CIViC-Fact, a benchmark dataset and pipeline for testing automated systems that verify the accuracy of cancer variant claims. CIViC-Fact links structured claims to sentence-level supporting or refuting evidence from full-text articles, and includes expert annotations and explanations. We evaluated multiple language models. Proprietary models performed well without training, but a smaller open-source model, fine-tuned on CIViC-Fact, achieved the highest accuracy (89%). Applying our fact-checking pipeline to real CIViC entries showed that reviewing less than 20% of content, focusing on flagged entries, would be sufficient to catch over half of all errors. This AI-assisted triage greatly accelerates the review process without replacing or reducing expert insight, ensuring that existing careful oversight remains in place while curators can work more efficiently. CIViC-Fact provides a realistic, high-consequence framework for biomedical fact-checking and a path toward more rigorous and efficient knowledgebase curation.

Journal Article

Quality over quantity: biopsy-anchored CT radiogenomics models outperform all-lesion training in a multi-tumour cohort despite a smaller sample size.

OBJECTIVE: Radiogenomics aims to non-invasively predict tumour genotypes from imaging, but most studies assume molecular homogeneity by assigning a single biopsy-derived label to all lesions within a patient. This approach risks substantial label noise given well-documented interlesional heterogeneity. We investigated whether anchoring training to biopsy-confirmed lesions improves radiogenomic model performance and generalisability. MATERIALS AND METHODS: We retrospectively analysed 1646 patients (11473 segmented lesions) with contrast-enhanced CT and EGFR mutation status from next-generation sequencing at the Netherlands Cancer Institute, alongside an external NSCLC radiogenomics cohort (n = 158). All visible lesions were segmented, and the exact biopsy site was matched to its segmentation. Radiomic features were extracted, and machine learning models were trained with three lesion selection strategies: all lesions, non-biopsied lesions only, and biopsy-confirmed lesions only. To disentangle label quality from sample size, we created size-matched variants (one lesion per patient) for all-lesion and non-biopsied strategies. RESULTS: All models achieved significant discrimination of EGFR status on internal validation (AUC = 0.62-0.68). However, performance of the all-lesion and non-biopsied models declined on external validation (AUC = 0.55-0.63), while the biopsy-anchored model maintained stable performance (AUC = 0.62), despite having only 1/10th of the training sample size. When training sets were size-matched, the biopsy-anchored approach significantly outperformed a model trained on all available lesions on external validation (p = 0.037). CONCLUSIONS: Radiogenomic models trained on biopsy-confirmed lesions outperform conventional all-lesion strategies in external validation, despite using an order of magnitude fewer samples. Prioritising lesion-level label fidelity can mitigate heterogeneity-driven noise, enhancing robustness and clinical translation of imaging-based genomic prediction. KEY POINTS: Question Does assigning biopsy-derived molecular labels to all lesions introduce heterogeneity-driven label noise that reduces the generalisability of radiogenomic models? Findings Models trained exclusively on biopsy-confirmed lesions demonstrated superior external generalisability compared with all-lesion approaches, despite being trained on substantially fewer samples. Clinical relevance Biopsy-anchored radiogenomics improves the reliability of non-invasive mutation prediction by accounting for tumour heterogeneity, potentially supporting clinical decision-making when tissue sampling is limited or molecular results are discordant across lesions.

Humans

Machine learning-integrated multi-omics risk prediction for pulmonary fungal infection in COPD and lung cancer: a transcriptomic and immune profiling study.

BACKGROUND: Chronic obstructive pulmonary disease (COPD) and lung cancer are major risk factors for invasive pulmonary fungal infection (IPFI), carrying an attributable mortality of 30%-80%. Their coexistence further amplifies immunosuppression, while current diagnostic criteria remain inadequate for early risk identification. METHODS: Transcriptomic data from the GEO dataset GSE296912 (scRNA-seq; 12,078 cells from normal and COPD lung tissue) and The Cancer Genome Atlas (TCGA)-lung adenocarcinoma (LUAD) bulk RNA-seq cohort (539 tumor and 59 normal samples) underwent differential expression and cross-omics integration analysis. Five machine learning models were constructed: logistic regression, SVM, random forest, XGBoost, and LASSO. Candidate genes were validated by qRT-PCR in A549 cells and THP-1-derived macrophages stimulated with heat-inactivated Aspergillus fumigatus conidia, a protocol selected to ensure BSL-2 biosafety compliance and isolate PAMP-mediated innate immune signaling. Model performance was evaluated using 5-fold stratified cross-validation with AUC, calibration curves, and decision curve analysis. RESULTS: Single-cell transcriptomic analysis of 12,078 cells identified 14 distinct cell populations, with marked myeloid expansion and immune dysregulation in COPD lung tissue. Cross-omics integration with TCGA-LUAD data identified 1,145 shared genes (79 immune-related), converging on NF-κB, TLR4, and cytokine receptor signaling. The random forest model achieved excellent discriminative performance (5-fold CV AUC = 0.988), with Treg infiltration, TLR4, and MMP9 as the top predictors. qRT-PCR confirmed significant upregulation of all five candidate genes (DEFB4A, S100A8, IL-8, MMP9, and TLR4) in both A549 and THP-1 cells following fungal stimulation. CONCLUSION: This multi-omics machine learning model integrating scRNA-seq and TCGA transcriptomic data demonstrates excellent discriminative performance (AUC = 0.988), with mechanistic convergence of NF-κB, TLR4, and oncogenic signaling pathways identified across shared immune gene signatures. In vitro qRT-PCR validation confirms the biological relevance of five key antifungal immune genes, providing a transcriptomic foundation for future prospective IPFI risk stratification in patients with COPD and lung cancer.

TLR4

Verbalization in EMR children's observational learning.

The effect of descriptive verbalization during observation of a model on mentally retarded boys' retention for what they had observed was examined. Forty 9- to 12-year-old boys in public-school EMR classes were grouped on the basis of relatively high or low IQ scores. One-half of each group observed a videotaped model perform a series of novel acts, while in addition to viewing the tape, the other half described the model's actions. Observational learning was immediately tested through a set of prompts for imitation, with prizes offered commensurate with level of performance. Regardless of IQ group, the boys who were required to verbalize the model's behavior were able to imitate it significantly better than boys who merely watched the model; high and low IQ groups did not significantly differ in observational learning. Further directions for research on mentally retarded children's observational learning were suggested.

Attention

Diagnostic Performance of Machine Learning for Systemic Lupus Erythematosus: Systematic Review and Meta-Analysis.

BACKGROUND: Early and accurate diagnosis of systemic lupus erythematosus (SLE) and its organ involvement is essential. Previous reviews of machine learning (ML) in SLE combined heterogeneous tasks and validation strategies and may have overinterpreted model performance. OBJECTIVE: This study evaluated the diagnostic performance of ML and deep learning (DL) models for 3 clinically distinct SLE-related tasks: SLE classification or diagnosis, lupus nephritis (LN) diagnosis, and neuropsychiatric systemic lupus erythematosus (NPSLE) discrimination. We also assessed methodological quality and certainty of evidence. METHODS: PubMed, Embase, Cochrane Library, Web of Science, and IEEE Xplore were searched from January 2014 to April 2026. Eligible peer-reviewed diagnostic accuracy studies developed or validated ML or DL models for 1 of the 3 prespecified tasks, used an accepted reference standard, and provided data for a 2×2 contingency table. Bivariate random-effects meta-analyses with the Hartung-Knapp-Sidik-Jonkman adjustment were used to pool sensitivity and specificity. We reported 95% prediction intervals (PIs), assessed risk of bias using the Quality Assessment of Diagnostic Accuracy Studies for Artificial Intelligence tool (QUADAS-AI; Viknesh Sounderajah [Imperial College London]), and evaluated certainty of evidence using the Grading of Recommendations Assessment, Development, and Evaluation framework for diagnostic test accuracy. RESULTS: Twenty-nine studies were included: 17 for SLE classification, 5 for LN diagnosis, and 7 for NPSLE discrimination. In the primary task-stratified analysis, pooled sensitivity was 0.91 (95% CI 0.86-0.94; 95% PI 0.56-0.99), and pooled specificity was 0.94 (95% CI 0.91-0.96; 95% PI 0.69-0.99), with low heterogeneity (I²=23.9% and 22.9%, respectively). DL models showed a sensitivity of 0.93 and specificity of 0.95, compared with 0.88 and 0.94 for traditional ML models. Certainty of evidence was high for most analyses but low for LN diagnosis because of inconsistency and imprecision. All studies were retrospective, and only 9 of 29 (31%) performed independent external validation. Overall risk of bias was high or unclear in 22 of 29 (75.9%) studies. No study reported model calibration, decision-curve analysis, or net clinical benefit. CONCLUSIONS: ML models showed promising diagnostic accuracy across 3 distinct SLE-related tasks, but wide PIs, limited external validation, and pervasive risk of bias restrict conclusions about real-world generalizability. Prospective multicenter studies with standardized tasks and reference standards, independent external validation, and formal assessment of calibration and clinical utility are required before clinical implementation.

Humans

SLAM: a connectionist model for attention in visual selection tasks.

SLAM, the SeLective Attention Model, performs visual selective attention tasks, an analysis of which shows that two processes, object and attribute selection, are both necessary and sufficient. It is based upon the McClelland and Rumelhart (1981) model for visual word recognition, with the addition of a response selection and evaluation mechanism. The responses may be correct or incorrect and, in particular conditions, SLAM may not make a response at all. Moreover, it allows for the generation of specific responses in time. SLAM's main characteristics are parallelism restricted by competition within modules, heterarchical processing in a hierarchical structure, and generation of responses as a result of relaxation given the conjoint constraints of stimulation, object, and attribute selection. The model is considered to represent an individual subject performing filtering tasks and demonstrates appropriate selective behavior. It is also tested quantitatively using a single tentative set of model parameters. The study reports simulations of four different filtering experiments, modeling response latencies, and error proportions. Specifications are made to take account of instructions, previous trials, and the effect of a barmarker cue and of asynchronies in stimulus and cue onsets. The model is then extended in order to provide simulations of a number of Stroop experiments, which can be regarded as filtering tasks with nonequivalent stimuli. The extension required for Stroop simulations is the addition of direct connections between compatible stimulus and response aspects. The direct connections do not affect the simulation of simpler filtering tasks. A variety of different experiments carried out by different authors is simulated. The model is discussed in terms of how modular architecture and the interaction of excitation and inhibition generate facilitation or inhibition of response latencies.

Arousal

Speech-discrimination scores modeled as a binomial variable.

Many studies have reported variability data for tests of speech discrimination, and the disparate results of these studies have not been given a simple explanation. Arguments over the relative merits of 25- vs 50-word tests have ignored the basic mathematical properties inherent in the use of percentage scores. The present study models performance on clinical tests of speech discrimination as a binomial variable. A binomial model was developed, and some of its characteristics were tested against data from 4120 scores obtained on the CID Auditory Test W-22. A table for determining significant deviations between scores was generated and compared to observed differences in half-list scores for the W-22 tests. Good agreement was found between predicted and observed values. Implications of the binomial characteristics of speech-discrimination scores are discussed.

Humans

Three-dimensional anatomy and renal concentrating mechanism. II. Sensitivity results.

A mathematical model has been developed to simulate hypertonic urine formation in the renal medulla. The model uses published values of membrane transport parameters, as have other models, but is unique in its representation of the three-dimensional anatomy of the medulla. The model successfully predicts measured fluid flows, osmolarities, and NaCl and urea concentrations. The model results are presented in the companion to this paper [A. S. Wexler, R. E. Kalaba, D. J. Marsh. Am. J. Physiol. 260 (Renal Fluid Electrolyte Physiol. 29): F368-F383, 1991.]. In this paper we provide tests of the sensitivity of model performance to variations in the description of the anatomy and in membrane transport parameters. From these studies we conclude that 1) strict counterflow arrangements are required in the outer stripe to prevent loss of NaCl to the systemic circulation, 2) the radial organization in the inner stripe materially improves performance of the inner medulla, 3) radial organization of the inner medulla is essential to hypertonic urine formation there, 4) the model is most sensitive to variation in collecting duct parameters, and 5) reabsorption of urea in the distal tubule improves system performance. The results support the claim that the three-dimensional structure, as captured in the model, provides a crucial framework for the production of hypertonic urine.

Absorption

Neural model of adaptive hand-eye coordination for single postures.

A neural network model has been developed that achieves adaptive visual-motor coordination of a multijoint arm, without a teacher. The model learns to position an arm so that it reaches a cylinder arbitrarily positioned in space. The model uses a new neural architecture and a new algorithm for modifying neural-connection strengths. Computer simulations show that the model performs with an average position error of 4% of the arm's length and with an average orientation error of 4 degrees. The model is designed to be generalized for coordinating any number of topographic sensory inputs with limbs of any number of joints.

Humans

Dataset Readiness Assessment With Large Language Model (DRAFT-LLM): A Multi-Axis Audit Guided by LLM.

This article details the Dataset Readiness Assessment for Training (DRAFT), a systematic method for determining whether a high-dimensional biological dataset is suitable for developing reliable, equitable (i.e., the extent to which model performance, error patterns, and potential benefits or harms are evaluated and found to be acceptably distributed across relevant demographic, biological, clinical, and contextual subgroups), and scientifically meaningful machine-learning models, and DRAFT Large Language Model (DRAFT-LLM), its optional human-in-the-loop extension for calibrating study-specific audits through structured, critically reviewed LLM guidance. Standard model validation often fails to detect when apparent performance is driven by spurious correlations, technical artifacts, or hidden stratification, leading to irreproducible and inequitable findings. DRAFT-LLM addresses this gap by shifting the focus from model tuning to structured dataset auditing, organized around Support Protocols 1 to 4 that capture the scientific intent, data structure, and governance constraints of a given study. These Support Protocols: (1) elicit and formalize investigator input into a study intake and dataset card; (2) compute standardized dataset statistics and structural summaries suitable for downstream analysis and LLM context; (3) configure the language model using form-based responses, safety guardrails, and governance rules; and (4) generate personalized instructions, prompts, and code templates for running DRAFT audits. Basic Protocols 1 to 3 are instantiated from this support layer for generalization, equity, and stability: they are reusable execution patterns whose concrete behavior is determined by the cards, statistics, and configurations defined in the Support Protocols. DRAFT-LLM and DRAFT are demonstrated in this article through an end-to-end case study on The Cancer Genome Atlas (TCGA). © 2026 Wiley Periodicals LLC. Support Protocol 1: Study intake and dataset card construction Support Protocol 2: Dataset structure and advanced summary statistics for LLM context Support Protocol 3: LLM configuration using structured form responses Support Protocol 4: Generation of personalized instructions for DRAFT audits Basic Protocol 1: Generalization audit Basic Protocol 2: Equity audit Basic Protocol 3: Stability audit.

Large Language Models

Studies of parallel barrier performance by acoustical modeling.

An investigation is presented into the performance of parallel barrier configurations, using acoustical scale modeling. A realistic geometry is investigated, with the source being positioned over a paved roadway and the receiver over grass-covered ground. The grass-covered ground surface was properly modeled in terms of its impedance. Results were obtained for a range of barrier types, and demonstrate that frequency dependent effects are evident in barrier insertion loss data. In most cases, the barrier on the far side of the source did not significantly affect sound levels at the receiver. The most effective barrier design was found to be that of a gradual grass-covered slope up to an upright, thin barrier.

Acoustics

Modeling the dishabituation hierarchy: the role of the primordial hippocampus.

We present a neural model for the organization and neural dynamics of the medial pallium, the toad's homolog of mammalian hippocampus. A neural mechanism, called cumulative shrinking, is proposed for mapping temporal responses from the anterior thalamus into a form of population coding referenced by spatial positions. Synaptic plasticity is modeled as an interaction of two dynamic processes which simulates acquisition and both short-term and long-term forgetting. The structure of the medial pallium model plus the plasticity model allows us to provide an account of the neural mechanisms of habituation and dishabituation. Computer simulations demonstrate a remarkable match between the model performance and the original experimental data on which the dishabituation hierarchy was based. A set of model predictions is presented, concerning mechanisms of habituation and cellular organization of the medial pallium.

Animals

Digital separation of primary and scatter components of chest radiographs.

This article describes a technique for digital separation of the primary and scatter components of a radiographic image. The method involves mathematical modeling of the process whereby an antiscatter grid reduces scatter patterns in film radiographs. Two superimposable radiographs (one taken with and the other without an intervening antiscatter grid) are applied to the model. Performance characteristics of the grid (primary and scatter transmittance factors) are also determined and used in the model. Radiographs of a humanoid chest phantom are processed. Scatter/primary separation appears to be accurate to within 15%. Film images that are quantitatively faithful to the calculated primary and scatter fields are included.

Analog-Digital Conversion