PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Model performance”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Predicting duration of sleep from the three process model of regulation of alertness.

OBJECTIVES: Irregular working hours severely disturb sleep and wakefulness. This paper presents a modification of the quantitative (computerised) three process model of regulation of alertness to predict duration of sleep in connection with irregular sleep patterns. METHODS: The model uses a circadian "C" (sinusoidal) and homeostatic "S" (exponential) component (the duration of previous periods awake and asleep), which are summed to yield predicted alertness (on a scale of 1-16). It assumes that waking from sleep will occur at a given alertness level (S' + C') when recuperation is complete. Variables of electroencephalographic duration of sleep from two studies of irregular sleep were used to model the S and C variables in a regression approach to maximise prediction. The model performance was cross validated against published field and laboratory data. RESULTS: The model parameters were defined with a high degree of precision R2 = 0.99 and the validation yielded similar values R2 = 0.98-0.95, depending on the acrophase. The paper also describes a simplified graphical version of the computation model seen as a two dimensional duration of sleep nomogram. CONCLUSION: The model seems to predict group means for duration of sleep with high precision and may serve as a tool for evaluating work and rest schedules to reduce risks of sleep disturbances.

Adult↗

The effect of combined self- and expert-modelling on the performance of the double leg circle on the pommel horse.

In this study, we investigated whether video modelling can enhance gymnasts' performance of the circle on a pommel horse. The procedure associated expert-modelling with self-modelling and quantitative performance analysis. Sixteen gymnasts were randomly assigned to one of two groups: (1) a modelling group, which received expert- and self-modelling, and performance feedback, or (2) a control group, which received no feedback. After five sessions of training, an analysis of variance with repeated measures indicated that the gains in the back, entry, front, and exit phases of the circle were greater for the modelling group than for the control group. During the training sessions, the gymnasts in the modelling group improved their body segmental alignment during the back phase more quickly than during the other phases. As predicted, although both groups performed the same number of circles (300 in 5 days, with 10 sequences of 6 circles), the modelling group improved their body segmental alignment more than the control group. It thus appears that immediate video modelling can help to correct complex sports movements such as the circle performed on the pommel horse. However, its effectiveness seemed to be dependent on the complexity of the phase.

Adolescent↗

Modelling with constructive backpropagation.

Neural network methods have proven to be powerful tools in modelling of nonlinear processes. One crucial part of modelling is the training phase where the model parameters are adjusted so that the model performs the desired operation as well as possible. Besides parameter estimation, an important problem is to select a suitable model structure. With a bad structure we potentially run into problems like underfitting, overfitting or wasting computational resources. One approach for structure learning is to use constructive methods, where training begins with minimal structure, and then more parameters are added when needed according to some predefined rule. This kind of constructive solution has also become more attractive in neural networks literature where one of the most well known constructive techniques is cascade-correlation (CC) learning. Inspired by CC we propose and study a similar technique called constructive backpropagation (CBP). We show that CBP is computationally just as efficient as the CC algorithm even though we need to backpropagate the error through no more than one hidden layer. Further, CBP has the same constructive benefits as CC, but in addition CBP benefits from simpler implementation and the ability to utilize stochastic optimization routines. Moreover, we show how CBP can be extended to allow addition of multiple new units simultaneously and how it can be used to perform continuous automatic structure adaptation. This includes both addition and deletion of units. The performance of CBP learning is studied with time series modelling experiments which demonstrate that CBP can provide significantly better modelling capabilities compared to CC learning.

Journal Article↗

Aggregation bias and the use of regression in evaluating models of human performance.

Regression analyses are increasingly being used to provide confirmatory evidence for models of human performance. The amount of information made available to judge these models is reduced because clearly established standards in the techniques of performing and reporting regression analyses are lacking. This paper addresses two primary problems in regression analysis: aggregation of data and the aggregation of variables into composite models. We provide examples of the misuse of regression techniques and recommend ways in which the amount of information made available to evaluate the model being tested can be maximized in analysis and reporting.

Bias↗

Molecular diagnosis. Classification, model selection and performance evaluation.

OBJECTIVES: We discuss supervised classification techniques applied to medical diagnosis based on gene expression profiles. Our focus lies on strategies of adaptive model selection to avoid overfitting in high-dimensional spaces. METHODS: We introduce likelihood-based methods, classification trees, support vector machines and regularized binary regression. For regularization by dimension reduction, we describe feature selection methods: feature filtering, feature shrinkage and wrapper approaches. In small sample-size situations efficient methods of data re-use are needed to assess the predictive power of a model. We discuss two issues in using cross-validation: the difference between in-loop and out-of-loop feature selection, and estimating model parameters in nested-loop cross-validation. RESULTS: Gene selection does not reduce the dimensionality of the model. Tuning parameters enable adaptive model selection. The feature selection bias is a common pitfall in performance evaluation. Model selection and performance evaluation can be combined by nested-loop cross-validation. CONCLUSIONS: Classification of microarrays is prone to overfitting. A rigorous and unbiased assessment of the predictive power of the model is a must.

Gene Expression Profiling↗

Comparison of the Elixhauser and Charlson/Deyo methods of comorbidity measurement in administrative data.

BACKGROUND: Comorbidity risk adjustment methods have been used widely with administrative data, and the Charlson/Deyo method is perhaps the most commonly used in the literature. However, a new method defined by Elixhauser et al. has been introduced recently and could be superior, although it has not been validated widely. OBJECTIVES: We compared the Charlson/Deyo and Elixhauser methods using Canadian administrative data on patients with myocardial infarction (MI). RESEARCH DESIGN: We conducted a historical cohort study. SUBJECTS: We used administrative hospital discharge data from a large Canadian city for all cases with acute MI coded as most responsible diagnosis between January 1, 1995, and March 31, 2001. MEASURES: We used each of the 2 methods to define comorbidity variables based on the International Classification of Diseases, 9th Revision, Clinical Modification codes present in each case record. We then compared 2 models predicting in-hospital mortality based on presence or absence of the variables defined by each of the methods. Frequency tables were produced and c-statistics and changes in -2 log likelihood (-2LogL) were calculated. We also visually assessed model performance by plotting observed and expected percentages of death for increasing risk categories defined by the 2 models. RESULTS: The Elixhauser model outperformed the Charlson/Deyo model in predicting mortality, with higher c-statistic values (0.793 vs. 0.704). Superior performance of the Elixhauser method is confirmed when plotting the expected and observed risks of death across groupings of increasing risk, in which the Elixhauser method yields a wider range of predicted and observed probabilities of death across groupings (2.5%-33%) than does the Charlson/Deyo method (5%-25%). CONCLUSIONS: The Elixhauser comorbidity measurement method performs better than the widely used Charlson/Deyo method in the Canadian acute MI cases studied.

Aged↗

Diagnostic Performance of Machine Learning for Systemic Lupus Erythematosus: Systematic Review and Meta-Analysis.

BACKGROUND: Early and accurate diagnosis of systemic lupus erythematosus (SLE) and its organ involvement is essential. Previous reviews of machine learning (ML) in SLE combined heterogeneous tasks and validation strategies and may have overinterpreted model performance. OBJECTIVE: This study evaluated the diagnostic performance of ML and deep learning (DL) models for 3 clinically distinct SLE-related tasks: SLE classification or diagnosis, lupus nephritis (LN) diagnosis, and neuropsychiatric systemic lupus erythematosus (NPSLE) discrimination. We also assessed methodological quality and certainty of evidence. METHODS: PubMed, Embase, Cochrane Library, Web of Science, and IEEE Xplore were searched from January 2014 to April 2026. Eligible peer-reviewed diagnostic accuracy studies developed or validated ML or DL models for 1 of the 3 prespecified tasks, used an accepted reference standard, and provided data for a 2×2 contingency table. Bivariate random-effects meta-analyses with the Hartung-Knapp-Sidik-Jonkman adjustment were used to pool sensitivity and specificity. We reported 95% prediction intervals (PIs), assessed risk of bias using the Quality Assessment of Diagnostic Accuracy Studies for Artificial Intelligence tool (QUADAS-AI; Viknesh Sounderajah [Imperial College London]), and evaluated certainty of evidence using the Grading of Recommendations Assessment, Development, and Evaluation framework for diagnostic test accuracy. RESULTS: Twenty-nine studies were included: 17 for SLE classification, 5 for LN diagnosis, and 7 for NPSLE discrimination. In the primary task-stratified analysis, pooled sensitivity was 0.91 (95% CI 0.86-0.94; 95% PI 0.56-0.99), and pooled specificity was 0.94 (95% CI 0.91-0.96; 95% PI 0.69-0.99), with low heterogeneity (I²=23.9% and 22.9%, respectively). DL models showed a sensitivity of 0.93 and specificity of 0.95, compared with 0.88 and 0.94 for traditional ML models. Certainty of evidence was high for most analyses but low for LN diagnosis because of inconsistency and imprecision. All studies were retrospective, and only 9 of 29 (31%) performed independent external validation. Overall risk of bias was high or unclear in 22 of 29 (75.9%) studies. No study reported model calibration, decision-curve analysis, or net clinical benefit. CONCLUSIONS: ML models showed promising diagnostic accuracy across 3 distinct SLE-related tasks, but wide PIs, limited external validation, and pervasive risk of bias restrict conclusions about real-world generalizability. Prospective multicenter studies with standardized tasks and reference standards, independent external validation, and formal assessment of calibration and clinical utility are required before clinical implementation.

Humans↗

Individual and group modelling of aesthetic judgment strategies.

Individual differences in aesthetic judgments were investigated by comparing quantitative group and individual performance models of the judgment processes. Aesthetic judgments of beauty over novel, formal, graphic patterns were collected from 34 non-artist college students using a two-step ranking-rating procedure. Their judgment processes were individually modelled using Judgment Analysis. The participants showed noted individual differences. Certain features of the stimulus material, which were considered to contribute to the picture's beauty by one participant, were used in an opposing fashion by another. A group model was derived based on the average ratings of the patterns' beauty. It was concluded that the group model was not an adequate representation of the present data, whereas the data revealed systematic judgment processes at the individual subject level.

Adult↗

The accuracy of seven mathematical functions in modeling dairy cattle lactation curves based on test-day records from varying sample schemes.

Daily milk yield over the course of the lactation follows a curvilinear pattern, so a suitable function is required to model this curve. In this study, 7 functions (Wood, Wilmink, Ali and Schaeffer, cubic splines, and 3 Legendre polynomials) were used to model the lactation curve at the phenotypic level, using both daily observations and data from commonly used recording schemes. The number of observations per lactation varied from 4 to 11. Several criteria based on the analysis of the real error were used to compare models. The performance of models showed few discrepancies in the comparison criteria when daily or 4-weekly (with first test at days in milk 8) data by lactation were used. The performance of the Wood, Wilmink, and Ali and Schaeffer models were highly affected by the reduction of the sample dimension. The results of this work support the idea that the performance of these models depends on the sample properties but also shows considerable variation within the sampling groups.

Animals↗

SLAM: a connectionist model for attention in visual selection tasks.

SLAM, the SeLective Attention Model, performs visual selective attention tasks, an analysis of which shows that two processes, object and attribute selection, are both necessary and sufficient. It is based upon the McClelland and Rumelhart (1981) model for visual word recognition, with the addition of a response selection and evaluation mechanism. The responses may be correct or incorrect and, in particular conditions, SLAM may not make a response at all. Moreover, it allows for the generation of specific responses in time. SLAM's main characteristics are parallelism restricted by competition within modules, heterarchical processing in a hierarchical structure, and generation of responses as a result of relaxation given the conjoint constraints of stimulation, object, and attribute selection. The model is considered to represent an individual subject performing filtering tasks and demonstrates appropriate selective behavior. It is also tested quantitatively using a single tentative set of model parameters. The study reports simulations of four different filtering experiments, modeling response latencies, and error proportions. Specifications are made to take account of instructions, previous trials, and the effect of a barmarker cue and of asynchronies in stimulus and cue onsets. The model is then extended in order to provide simulations of a number of Stroop experiments, which can be regarded as filtering tasks with nonequivalent stimuli. The extension required for Stroop simulations is the addition of direct connections between compatible stimulus and response aspects. The direct connections do not affect the simulation of simpler filtering tasks. A variety of different experiments carried out by different authors is simulated. The model is discussed in terms of how modular architecture and the interaction of excitation and inhibition generate facilitation or inhibition of response latencies.

Arousal↗

Do severity measures explain differences in length of hospital stay? The case of hip fracture.

OBJECTIVE: To examine whether judgments about hospital length of stay (LOS) vary depending on the measure used to adjust for severity differences. DATA SOURCES/STUDY SETTING: Data on admissions to 80 hospitals nationwide in the 1992 MedisGroups Comparative Database. STUDY DESIGN: For each of 14 severity measures, LOS was regressed on patient age/sex, DRG, and severity score. Regressions were performed on trimmed and untrimmed data. R-squared was used to evaluate model performance. For each severity measure for each hospital, we calculated the expected LOS and the z-score, a measure of the deviation of observed from expected LOS. We ranked hospitals by z-scores. DATA EXTRACTION: All patients admitted for initial surgical repair of a hip fracture, defined by DRG, diagnosis, and procedure codes. PRINCIPAL FINDINGS: The 5,664 patients had a mean (s.d.) LOS of 11.9 (8.9) days. Cross-validated R-squared values from the multivariable regressions (trimmed data) ranged from 0.041 (Comorbidity Index) to 0.165 (APR-DRGs). Using untrimmed data, observed average LOS for hospitals ranged from 7.6 to 23.9 days. The 14 severity measures showed excellent agreement in ranking hospitals based on z-scores. No severity measure explained the differences between hospitals with the shortest and longest LOS. CONCLUSIONS: Hospitals differed widely in their mean LOS for hip fracture patients, and severity adjustment did little to explain these differences.

Aged↗

Stability and change in longitudinal water-level task performance.

Three longitudinal samples of children (N = 481), 8 to 16 years old, were assessed 3 times at yearly intervals on 8 water-level items. The within-child change in task performance over age is viewed as a stochastic process of the child changing or remaining in 1 of 3 latent (strategy) states: (a) bottom-parallel responders, (b) random responders, or (c) accurate responders. A random-effects binomial mixture distribution is used to model performance at each age. Change over age is gauged by a stochastic transition model. Although there was improvement in task performance over age, the more general finding is that strategy stability, not change, is most typical.

Adolescent↗

Predicting protein structure using hidden Markov models.

We discuss how methods based on hidden Markov models performed in the fold-recognition section of the CASP2 experiment. Hidden Markov models were built for a representative set of just over 1,000 structures from the Protein Data Bank (PDB). Each CASP2 target sequence was scored against this library of HMMs. In addition, an HMM was built for each of the target sequences and all of the sequences in PDB were scored against that target model, with a good score on both methods indicating a high probability that the target sequence is homologous to the structure. The method worked well in comparison to other methods used at CASP2 for targets of moderate difficulty, where the closest structure in PDB could be aligned to the target with at least 15% residue identity.

Markov Chains↗

Modeling the phenotype in parametric linkage analysis of bipolar disorder.

The definition of phenotype is a major problem in genetic studies of psychiatric disorders. Most linkage studies in bipolar disorder have defined the phenotype as a dichotomous trait and have usually employed different hierarchical classifications in order to overcome uncertainty resulting from phenotypic variability. In this study we explored the advantages of maximizing the evidence for linkage over different phenotypic definitions when conducting parametric linkage analysis of a complex trait. The GAW10 Problem 1 was used, focusing on chromosome 18 data sets. Three major phenotypic models were analyzed: quasi-quantitative, liability-based and affection-status models. Overall, no single phenotypic model performed consistently better than the others (i.e., lod scores greater than 1.0). Each model yielded higher lod scores than the others in particular instances, suggesting that it might be useful in exploratory data analysis, where the phenotype is variable, to maximize evidence for linkage over different phenotypic models.

Bipolar Disorder↗

A relaxation of the gamma frailty (Burr) model.

Frailty models are used in univariate data to account for individual heterogeneity. In the popular gamma frailty model the marginal hazard has the form of a Burr model. Although the Burr model is very useful and can offer insight on the data, it is far from perfect. The estimation of the covariate effects is linked to the baseline hazard and this makes the model coefficients hard to interpret. At the same time, the frailties are assumed constant over time, while biological reasoning in some cases may indicate that frailties may be time dependent. In this paper we present a relaxation of the Burr model which is based on loosening the link between the estimation of the covariate effects and the baseline hazard. This can be achieved by replacing the cumulative baseline hazard in the Burr model by a set of time functions, and the frailty variance by a vector of coefficients directly estimated from the data using a partial likelihood. We illustrate the similarities of the model with the Burr model and a further extension of the latter, a model with an autoregressive stochastic process for the frailty. We compare the models on simulated data sets with constant and time-dependent frailties and show how the relaxed Burr models performs on two different real data sets. We show that the relaxed Burr model serves as a good approximation to the Burr model when the frailty is constant, and furthermore it gives better results when the frailty is time dependent.

Adolescent↗

Predicting the thermal inactivation of bacteria in a solid matrix: simulation studies on the relative effects of microbial thermal resistance parameters and process conditions.

A combined mathematical model for predicting heat penetration and microbial inactivation in a solid body heated by conduction was tested experimentally by inoculating agar cylinders with Salmonella typhimurium or Enterococcus faecium and heating in a water bath. Regions of growth where bacteria had survived after heating were measured by image analysis and compared with model predictions. Visualisation of the regions of growth was improved by incorporating chromogenic metabolic indicators into the agar. Preliminary tests established that the model performed satisfactorily with both test organisms and with cylinders of different diameter. The model was then used in simulation studies in which the parameters D, z, inoculum size, cylinder diameter and heating temperature were systematically varied. These simulations showed that the biological variables D, z and inoculum size had a relatively small effect on the time needed to eliminate bacteria at the cylinder axis in comparison with the physical variables heating temperature and cylinder diameter, which had a much greater relative effect.

Agar↗

Review and assessment of models for predicting the migration of radionuclides through rivers.

The present paper summarises the results of the review and assessment of state-of-the-art models developed for predicting the migration of radionuclides through rivers. The different approaches of the models to predict the behaviour of radionuclides in lotic ecosystems are presented and compared. The models were classified and evaluated according to their main methodological approaches. The results of an exercise of model application to specific contamination scenarios aimed at assessing and comparing the model performances were described. A critical evaluation and analysis of the uncertainty of the models was carried out. The main factors influencing the inherent uncertainty of the models, such as the incompleteness of the actual knowledge and the intrinsic environmental and biological variability of the processes controlling the behaviour of radionuclides in rivers, are analysed.

Decision Making↗

Speech-discrimination scores modeled as a binomial variable.

Many studies have reported variability data for tests of speech discrimination, and the disparate results of these studies have not been given a simple explanation. Arguments over the relative merits of 25- vs 50-word tests have ignored the basic mathematical properties inherent in the use of percentage scores. The present study models performance on clinical tests of speech discrimination as a binomial variable. A binomial model was developed, and some of its characteristics were tested against data from 4120 scores obtained on the CID Auditory Test W-22. A table for determining significant deviations between scores was generated and compared to observed differences in half-list scores for the W-22 tests. Good agreement was found between predicted and observed values. Implications of the binomial characteristics of speech-discrimination scores are discussed.

Humans↗