PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Model performance”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

An evaluation of the predictive performance of distributional models for flora and fauna in north-east New South Wales.

To use models of species distributions effectively in conservation planning, it is important to determine the predictive accuracy of such models. Extensive modelling of the distribution of vascular plant and vertebrate fauna species within north-east New South Wales has been undertaken by linking field survey data to environmental and geographical predictors using logistic regression. These models have been used in the development of a comprehensive and adequate reserve system within the region. We evaluate the predictive accuracy of models for 153 small reptile, arboreal marsupial, diurnal bird and vascular plant species for which independent evaluation data were available. The predictive performance of each model was evaluated using the relative operating characteristic curve to measure discrimination capacity. Good discrimination ability implies that a model's predictions provide an acceptable index of species occurrence. The discrimination capacity of 89% of the models was significantly better than random, with 70% of the models providing high levels of discrimination. Predictions generated by this type of modelling therefore provide a reasonably sound basis for regional conservation planning. The discrimination ability of models was highest for the less mobile biological groups, particularly the vascular plants and small reptiles. In the case of diurnal birds, poor performing models tended to be for species which occur mainly within specific habitats not well sampled by either the model development or evaluation data, highly mobile species, species that are locally nomadic or those that display very broad habitat requirements. Particular care needs to be exercised when employing models for these types of species in conservation planning.

Animal Population Groups↗

Prognostic models based on literature and individual patient data in logistic regression analysis.

Prognostic models can be developed with multiple regression analysis of a data set containing individual patient data. Often this data set is relatively small, while previously published studies present results for larger numbers of patients. We describe a method to combine univariable regression results from the medical literature with univariable and multivariable results from the data set containing individual patient data. This 'adaptation method' exploits the generally strong correlation between univariable and multivariable regression coefficients. The method is illustrated with several logistic regression models to predict 30-day mortality in patients with acute myocardial infarction. The regression coefficients showed considerably less variability when estimated with the adaptation method, compared to standard maximum likelihood estimates. Also, model performance, as distinguished in calibration and discrimination, improved clearly when compared to models including shrunk or penalized estimates. We conclude that prognostic models may benefit substantially from explicit incorporation of literature data.

Age Factors↗

Modelling the growth kinetics of Phanerochaete chrysosporium in submerged static culture.

The potential commercial application of Phanerochaete chrysosporium requires methods for quantitatively predicting growth and substrate utilization. The growth kinetics of P. chrysosporium INA-12 (CNCM I-398) were investigated and modelled under nonlimiting nitrogen and carbon conditions in submerged static culture. This strain, unlike other strains, does not require nutrient limitation for induction of lignin peroxidase. Maximum levels of lignin peroxidase activity were reached 7 days after culture initiation, when almost 80% of the initial glycerol and 70% of the initial nitrogen were still present. Lignin peroxidase levels then decreased, while biomass levels increased until about day 14. The ratio of cell dry weight to wet weight was constant until the maximum biomass concentration was achieved, after which there was a decrease in the water content. The change in this ratio reflects cell lysis as it correlated with increased concentrations of nitrogen in the media, arising from cell leakage. The suitability of four growth models to predict growth, and in some cases glycerol consumption, was evaluated. A simple linear model and the Emerson model performed poorly for the early stages of growth, while a modified Williams model and the Monod model predicted substrate and biomass concentrations equally well. All models will predict biomass concentrations during the active growth phase, but they should not be used to predict biomass concentrations after the stationary growth phase, when cell lysis becomes significant.

Basidiomycota↗

Predicting duration of sleep from the three process model of regulation of alertness.

OBJECTIVES: Irregular working hours severely disturb sleep and wakefulness. This paper presents a modification of the quantitative (computerised) three process model of regulation of alertness to predict duration of sleep in connection with irregular sleep patterns. METHODS: The model uses a circadian "C" (sinusoidal) and homeostatic "S" (exponential) component (the duration of previous periods awake and asleep), which are summed to yield predicted alertness (on a scale of 1-16). It assumes that waking from sleep will occur at a given alertness level (S' + C') when recuperation is complete. Variables of electroencephalographic duration of sleep from two studies of irregular sleep were used to model the S and C variables in a regression approach to maximise prediction. The model performance was cross validated against published field and laboratory data. RESULTS: The model parameters were defined with a high degree of precision R2 = 0.99 and the validation yielded similar values R2 = 0.98-0.95, depending on the acrophase. The paper also describes a simplified graphical version of the computation model seen as a two dimensional duration of sleep nomogram. CONCLUSION: The model seems to predict group means for duration of sleep with high precision and may serve as a tool for evaluating work and rest schedules to reduce risks of sleep disturbances.

Adult↗

Aggregation bias and the use of regression in evaluating models of human performance.

Regression analyses are increasingly being used to provide confirmatory evidence for models of human performance. The amount of information made available to judge these models is reduced because clearly established standards in the techniques of performing and reporting regression analyses are lacking. This paper addresses two primary problems in regression analysis: aggregation of data and the aggregation of variables into composite models. We provide examples of the misuse of regression techniques and recommend ways in which the amount of information made available to evaluate the model being tested can be maximized in analysis and reporting.

Bias↗

Diagnostic Performance of Machine Learning for Systemic Lupus Erythematosus: Systematic Review and Meta-Analysis.

BACKGROUND: Early and accurate diagnosis of systemic lupus erythematosus (SLE) and its organ involvement is essential. Previous reviews of machine learning (ML) in SLE combined heterogeneous tasks and validation strategies and may have overinterpreted model performance. OBJECTIVE: This study evaluated the diagnostic performance of ML and deep learning (DL) models for 3 clinically distinct SLE-related tasks: SLE classification or diagnosis, lupus nephritis (LN) diagnosis, and neuropsychiatric systemic lupus erythematosus (NPSLE) discrimination. We also assessed methodological quality and certainty of evidence. METHODS: PubMed, Embase, Cochrane Library, Web of Science, and IEEE Xplore were searched from January 2014 to April 2026. Eligible peer-reviewed diagnostic accuracy studies developed or validated ML or DL models for 1 of the 3 prespecified tasks, used an accepted reference standard, and provided data for a 2×2 contingency table. Bivariate random-effects meta-analyses with the Hartung-Knapp-Sidik-Jonkman adjustment were used to pool sensitivity and specificity. We reported 95% prediction intervals (PIs), assessed risk of bias using the Quality Assessment of Diagnostic Accuracy Studies for Artificial Intelligence tool (QUADAS-AI; Viknesh Sounderajah [Imperial College London]), and evaluated certainty of evidence using the Grading of Recommendations Assessment, Development, and Evaluation framework for diagnostic test accuracy. RESULTS: Twenty-nine studies were included: 17 for SLE classification, 5 for LN diagnosis, and 7 for NPSLE discrimination. In the primary task-stratified analysis, pooled sensitivity was 0.91 (95% CI 0.86-0.94; 95% PI 0.56-0.99), and pooled specificity was 0.94 (95% CI 0.91-0.96; 95% PI 0.69-0.99), with low heterogeneity (I²=23.9% and 22.9%, respectively). DL models showed a sensitivity of 0.93 and specificity of 0.95, compared with 0.88 and 0.94 for traditional ML models. Certainty of evidence was high for most analyses but low for LN diagnosis because of inconsistency and imprecision. All studies were retrospective, and only 9 of 29 (31%) performed independent external validation. Overall risk of bias was high or unclear in 22 of 29 (75.9%) studies. No study reported model calibration, decision-curve analysis, or net clinical benefit. CONCLUSIONS: ML models showed promising diagnostic accuracy across 3 distinct SLE-related tasks, but wide PIs, limited external validation, and pervasive risk of bias restrict conclusions about real-world generalizability. Prospective multicenter studies with standardized tasks and reference standards, independent external validation, and formal assessment of calibration and clinical utility are required before clinical implementation.

Humans↗

SLAM: a connectionist model for attention in visual selection tasks.

SLAM, the SeLective Attention Model, performs visual selective attention tasks, an analysis of which shows that two processes, object and attribute selection, are both necessary and sufficient. It is based upon the McClelland and Rumelhart (1981) model for visual word recognition, with the addition of a response selection and evaluation mechanism. The responses may be correct or incorrect and, in particular conditions, SLAM may not make a response at all. Moreover, it allows for the generation of specific responses in time. SLAM's main characteristics are parallelism restricted by competition within modules, heterarchical processing in a hierarchical structure, and generation of responses as a result of relaxation given the conjoint constraints of stimulation, object, and attribute selection. The model is considered to represent an individual subject performing filtering tasks and demonstrates appropriate selective behavior. It is also tested quantitatively using a single tentative set of model parameters. The study reports simulations of four different filtering experiments, modeling response latencies, and error proportions. Specifications are made to take account of instructions, previous trials, and the effect of a barmarker cue and of asynchronies in stimulus and cue onsets. The model is then extended in order to provide simulations of a number of Stroop experiments, which can be regarded as filtering tasks with nonequivalent stimuli. The extension required for Stroop simulations is the addition of direct connections between compatible stimulus and response aspects. The direct connections do not affect the simulation of simpler filtering tasks. A variety of different experiments carried out by different authors is simulated. The model is discussed in terms of how modular architecture and the interaction of excitation and inhibition generate facilitation or inhibition of response latencies.

Arousal↗

Do severity measures explain differences in length of hospital stay? The case of hip fracture.

OBJECTIVE: To examine whether judgments about hospital length of stay (LOS) vary depending on the measure used to adjust for severity differences. DATA SOURCES/STUDY SETTING: Data on admissions to 80 hospitals nationwide in the 1992 MedisGroups Comparative Database. STUDY DESIGN: For each of 14 severity measures, LOS was regressed on patient age/sex, DRG, and severity score. Regressions were performed on trimmed and untrimmed data. R-squared was used to evaluate model performance. For each severity measure for each hospital, we calculated the expected LOS and the z-score, a measure of the deviation of observed from expected LOS. We ranked hospitals by z-scores. DATA EXTRACTION: All patients admitted for initial surgical repair of a hip fracture, defined by DRG, diagnosis, and procedure codes. PRINCIPAL FINDINGS: The 5,664 patients had a mean (s.d.) LOS of 11.9 (8.9) days. Cross-validated R-squared values from the multivariable regressions (trimmed data) ranged from 0.041 (Comorbidity Index) to 0.165 (APR-DRGs). Using untrimmed data, observed average LOS for hospitals ranged from 7.6 to 23.9 days. The 14 severity measures showed excellent agreement in ranking hospitals based on z-scores. No severity measure explained the differences between hospitals with the shortest and longest LOS. CONCLUSIONS: Hospitals differed widely in their mean LOS for hip fracture patients, and severity adjustment did little to explain these differences.

Aged↗

Stability and change in longitudinal water-level task performance.

Three longitudinal samples of children (N = 481), 8 to 16 years old, were assessed 3 times at yearly intervals on 8 water-level items. The within-child change in task performance over age is viewed as a stochastic process of the child changing or remaining in 1 of 3 latent (strategy) states: (a) bottom-parallel responders, (b) random responders, or (c) accurate responders. A random-effects binomial mixture distribution is used to model performance at each age. Change over age is gauged by a stochastic transition model. Although there was improvement in task performance over age, the more general finding is that strategy stability, not change, is most typical.

Adolescent↗

Predicting protein structure using hidden Markov models.

We discuss how methods based on hidden Markov models performed in the fold-recognition section of the CASP2 experiment. Hidden Markov models were built for a representative set of just over 1,000 structures from the Protein Data Bank (PDB). Each CASP2 target sequence was scored against this library of HMMs. In addition, an HMM was built for each of the target sequences and all of the sequences in PDB were scored against that target model, with a good score on both methods indicating a high probability that the target sequence is homologous to the structure. The method worked well in comparison to other methods used at CASP2 for targets of moderate difficulty, where the closest structure in PDB could be aligned to the target with at least 15% residue identity.

Markov Chains↗

Modeling the phenotype in parametric linkage analysis of bipolar disorder.

The definition of phenotype is a major problem in genetic studies of psychiatric disorders. Most linkage studies in bipolar disorder have defined the phenotype as a dichotomous trait and have usually employed different hierarchical classifications in order to overcome uncertainty resulting from phenotypic variability. In this study we explored the advantages of maximizing the evidence for linkage over different phenotypic definitions when conducting parametric linkage analysis of a complex trait. The GAW10 Problem 1 was used, focusing on chromosome 18 data sets. Three major phenotypic models were analyzed: quasi-quantitative, liability-based and affection-status models. Overall, no single phenotypic model performed consistently better than the others (i.e., lod scores greater than 1.0). Each model yielded higher lod scores than the others in particular instances, suggesting that it might be useful in exploratory data analysis, where the phenotype is variable, to maximize evidence for linkage over different phenotypic models.

Bipolar Disorder↗

Speech-discrimination scores modeled as a binomial variable.

Many studies have reported variability data for tests of speech discrimination, and the disparate results of these studies have not been given a simple explanation. Arguments over the relative merits of 25- vs 50-word tests have ignored the basic mathematical properties inherent in the use of percentage scores. The present study models performance on clinical tests of speech discrimination as a binomial variable. A binomial model was developed, and some of its characteristics were tested against data from 4120 scores obtained on the CID Auditory Test W-22. A table for determining significant deviations between scores was generated and compared to observed differences in half-list scores for the W-22 tests. Good agreement was found between predicted and observed values. Implications of the binomial characteristics of speech-discrimination scores are discussed.

Humans↗

Estimation of youth smoking behaviours in Canada.

This study estimated the prevalence of current smoking and smoking initiation among Canadian youth. Logistic regression was used to relate socio-demographic predictors to the occurrence of the smoking indicators among youth (15-24 years) in the 1994/95 National Population Health Survey (NPHS). Models were then applied to provincial youth populations in the 1996/97 NPHS and the 1996 census of Canada. Model-generated estimates were compared with direct estimates obtained from NPHS data. The models accurately predicted provincial rates of current youth smoking for 1994/95. When applied to the 1996/97 NPHS, the current smoking models performed reasonably well, but were less predictive when applied to 1996 census data. Modelling of youth smoking initiation was not successful. This suggests that although simple estimation models of youth smoking can be derived, these models may not be portable across different populations or time periods.

Adolescent↗

Improving health-based payment for Medicaid beneficiaries: CDPS.

This article describes the Chronic Illness and Disability Payment System (CDPS), a diagnostic classification system that Medicaid programs can use to make health-based capitated payments for TANF and disabled Medicaid beneficiaries. The authors describe the diversity of diagnoses and different burdens of illness among disabled and AFDC Medicaid beneficiaries. Claims from seven States are analyzed, and payment weights are provided that States can use when adjusting HMO payments. The authors also compare the taxonomy and statistical performance of CDPS to other leading diagnostic classification systems and find that the new model performs better in a number of respects.

Adult↗

Three-dimensional anatomy and renal concentrating mechanism. II. Sensitivity results.

A mathematical model has been developed to simulate hypertonic urine formation in the renal medulla. The model uses published values of membrane transport parameters, as have other models, but is unique in its representation of the three-dimensional anatomy of the medulla. The model successfully predicts measured fluid flows, osmolarities, and NaCl and urea concentrations. The model results are presented in the companion to this paper [A. S. Wexler, R. E. Kalaba, D. J. Marsh. Am. J. Physiol. 260 (Renal Fluid Electrolyte Physiol. 29): F368-F383, 1991.]. In this paper we provide tests of the sensitivity of model performance to variations in the description of the anatomy and in membrane transport parameters. From these studies we conclude that 1) strict counterflow arrangements are required in the outer stripe to prevent loss of NaCl to the systemic circulation, 2) the radial organization in the inner stripe materially improves performance of the inner medulla, 3) radial organization of the inner medulla is essential to hypertonic urine formation there, 4) the model is most sensitive to variation in collecting duct parameters, and 5) reabsorption of urea in the distal tubule improves system performance. The results support the claim that the three-dimensional structure, as captured in the model, provides a crucial framework for the production of hypertonic urine.

Absorption↗

Neural model of adaptive hand-eye coordination for single postures.

A neural network model has been developed that achieves adaptive visual-motor coordination of a multijoint arm, without a teacher. The model learns to position an arm so that it reaches a cylinder arbitrarily positioned in space. The model uses a new neural architecture and a new algorithm for modifying neural-connection strengths. Computer simulations show that the model performs with an average position error of 4% of the arm's length and with an average orientation error of 4 degrees. The model is designed to be generalized for coordinating any number of topographic sensory inputs with limbs of any number of joints.

Humans↗

Dataset Readiness Assessment With Large Language Model (DRAFT-LLM): A Multi-Axis Audit Guided by LLM.

This article details the Dataset Readiness Assessment for Training (DRAFT), a systematic method for determining whether a high-dimensional biological dataset is suitable for developing reliable, equitable (i.e., the extent to which model performance, error patterns, and potential benefits or harms are evaluated and found to be acceptably distributed across relevant demographic, biological, clinical, and contextual subgroups), and scientifically meaningful machine-learning models, and DRAFT Large Language Model (DRAFT-LLM), its optional human-in-the-loop extension for calibrating study-specific audits through structured, critically reviewed LLM guidance. Standard model validation often fails to detect when apparent performance is driven by spurious correlations, technical artifacts, or hidden stratification, leading to irreproducible and inequitable findings. DRAFT-LLM addresses this gap by shifting the focus from model tuning to structured dataset auditing, organized around Support Protocols 1 to 4 that capture the scientific intent, data structure, and governance constraints of a given study. These Support Protocols: (1) elicit and formalize investigator input into a study intake and dataset card; (2) compute standardized dataset statistics and structural summaries suitable for downstream analysis and LLM context; (3) configure the language model using form-based responses, safety guardrails, and governance rules; and (4) generate personalized instructions, prompts, and code templates for running DRAFT audits. Basic Protocols 1 to 3 are instantiated from this support layer for generalization, equity, and stability: they are reusable execution patterns whose concrete behavior is determined by the cards, statistics, and configurations defined in the Support Protocols. DRAFT-LLM and DRAFT are demonstrated in this article through an end-to-end case study on The Cancer Genome Atlas (TCGA). © 2026 Wiley Periodicals LLC. Support Protocol 1: Study intake and dataset card construction Support Protocol 2: Dataset structure and advanced summary statistics for LLM context Support Protocol 3: LLM configuration using structured form responses Support Protocol 4: Generation of personalized instructions for DRAFT audits Basic Protocol 1: Generalization audit Basic Protocol 2: Equity audit Basic Protocol 3: Stability audit.

Large Language Models↗

[Multifocal intraocular lenses--an assessment of current status].

UNLABELLED: Besides the diffractive multifocals, which produce a second focus for near vision by means of diffraction rings, there are different refractive multifocal IOL types with 2-7 refractive zones or an aspheric/spherical construction principle. Long-term results: 2 years after implantation of diffractive multifocal IOLs, the corrected distance and near acuities were unchanged compared to the 3-month results. The uncorrected distance acuity was, however, slightly decreased due to a minus shift of refraction to -1.2 D. The contrast sensitivity was improved after 2 years. Multi- versus monofocal IOLs: After diffractive multifocal IOL implantation, the near acuity with distance correction only was markedly improved compared to monofocal IOLs. All other acuity data did not differ between multi- or monofocal lenses. The contrast sensitivity (at low contrasts and high spatial frequencies) and mesopic visual acuity (without and with glare) were reduced compared to monofocal pseudophakic eyes. Near aniseikonia and binocular functions: In unilateral multifocal pseudophakia (monofocal IOL in fellow eye), a near aniseikonia up to 8% was found. The width of fusion was significantly lower than in bilateral multifocal pseudophakia, whereas the stereopsis showed no difference. Determinants of bifocal function: In 7.1% of our cases, no bifocal function (BFF) was present after implantation of diffractive multifocal IOLs. These patients exhibited a significantly higher age as well as higher pre- and postoperative astigmatism, when compared to patients with good BFF. Optical performance of different multifocal IOLs: By means of an optical system, described by Reiner, images of intraocular lenses can be projected into the eye ("optical implantation"); thus, the optical performance of IOLs can be judged subjectively. Using this method, the refractive 2- and 3-zone models performed best within the multifocal group (contrast sensitivity not significantly worse than that of monofocal IOL), when viewing a low-contrast chart (Regan 4%). All other multifocal lenses (diffractive, aspheric/spherical, refractive 5- and 7-zone models) were significantly inferior to the monofocal IOL. CONCLUSIONS: Implantation of multifocal IOLs should presently be restricted to special indications, particularly to the distinct patient request to dispense with wearing near or bifocal glasses, if possible. Because of the reduction in contrast sensitivity and mesopic vision and the increased glare sensibility, multifocal IOLs should not be implanted especially in professional car drivers. There are, however, differences in optical performance between the various multifocal IOL types. Further improvements, in particular concerning lens technology, will presumably extend the present spectrum of indications.

Contrast Sensitivity↗