PubMed HealthSearch

SEARCH · PubMed Health

Results for “Model performance”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Age, activity, and physical performance: an evaluation of performance models.

Two performance models relating age, physical activity, and physical performance were evaluated on 5 age cohorts involving more than 6,000 persons ranging in age from 20 to 69 years. Significant age and activity main effects, without 1 Age X Activity interaction, were obtained on 5 fitness measures and 2 activity indexes. These results are inconsistent with predictions from the moderation model, but they fit predictions from a recently formulated tonic and overpractice model of physical performance.

Activities of Daily Living

Indirect immunofluorescence test performance and questionnaire results from the Centers for Disease Control Model Performance Evaluation Program for human immunodeficiency virus type 1 testing.

Results from laboratories performing indirect immunofluorescence (IIF) testing for human immunodeficiency virus type 1 antibody and participating in the Centers for Disease Control Model Performance Evaluation Program in 1988 are presented. Approximately 90% of all laboratories receiving specimen panels or questionnaires furnished results to the Centers for Disease Control. In September 1988, 111 reports were received from IIF laboratories from 34 states and nine countries; most of these laboratories did IIF testing in conjunction with other antibody tests. Hospital laboratories were the most common type of laboratory participating in the program. Laboratories that performed IIF employed fewer personnel and performed testing less frequently than did laboratories that performed enzyme immunoassays or Western blot (immunoblot) tests and were likely to use a commercial test kit. Most of the laboratories that referred specimens for IIF testing sent them to the state laboratory. The analytic specificity for the Model Performance Evaluation Program specimens was 98.5% when indeterminate results on a negative specimen were considered correct (negative) and 89.6% when indeterminate results on a negative specimen were considered incorrect; analytic sensitivity was 94.8% when indeterminate results on a positive specimen were correct (positive) and 91.4% when indeterminate results on a positive specimen were considered incorrect. When indeterminate results were considered correct, all types of laboratories (blood bank, state, hospital, independent, and other) had analytic specificities over 96%, and all manufacturers had analytic specificities above 95%. All types of laboratories had analytic sensitivities over 92%, and analytic sensitivities were above 94% for all manufacturers and reagent sources except Cellular Products. Comparison of percentages of correct responses between IIF and Western blot assays on those samples for which there was good agreement on the target interpretation revealed no significant differences. Both individual donor and diluted materials were included in the evaluations; the diluted donor material presented the greatest testing difficulty. Within-survey reproducibility was about 93% overall and by specimen type. Between-survey reproducibility was about 81% for negative and indeterminate specimens and 88.5% for positive specimens, for an overall between-survey reproducibility of 84.3%. Differences in performance were noted when results were compared by type of laboratory and test manufacturer.

Acquired Immunodeficiency Syndrome

CDC's Model Performance Evaluation Program: assessment of the quality of laboratory performance for HIV-1 antibody testing.

In 1986, the Centers for Disease Control (CDC) implemented the Model Performance Evaluation Program (MPEP) to evaluate the performance of laboratories that test for antibody directed against human immunodeficiency virus type 1 (HIV-1). The impetus for developing this program came from the recognition of a need to assess the quality of existing and changing laboratory technology and to ensure that the quality of testing was sufficient to meet medical and public health needs. To develop the program, CDC chose HIV-1 antibody testing as the first specific application for assessing the quality of laboratory performance because (a) of the importance of accurate and reproducible test results for acquired immunodeficiency syndrome (AIDS) surveillance, prevention, and treatment programs; (b) HIV-1 testing technology is new to many laboratories; and (c) HIV-1 testing practices and applications continue to evolve. Unlike proficiency testing programs, the MPEP is not limited to assessing quality in the analytical step, alone. It will also assess quality in the preanalytical and postanalytical steps of the testing process, that is, from the time a test is requested until the clinician who ordered the test takes an action based on the test result. The participating laboratories furnish the information needed for the performance evaluation program by (a) completing questionnaires designed to describe HIV-1 testing laboratories and their testing practices, (b) analyzing specially prepared sample panels for HIV-1 antibody reactivity, and (c) reporting results to CDC.

AIDS Serodiagnosis

Centers for Disease Control perspective on quality assurance for human immunodeficiency virus type 1 antibody testing. Model Performance Evaluation Program.

Ensuring high quality in human immunodeficiency virus type 1 (HIV-1) antibody testing is an essential component of the organized public health response to epidemic HIV-1 infection. In 1986, the Centers for Disease Control designed the Model Performance Evaluation Program to assess and improve the analytic quality of HIV-1 antibody testing. In addition, the program was designed to gather information about HIV-1 antibody testing practices. The utility of this information is in identifying potential barriers to quality throughout the total testing process. Currently, 1405 laboratories participate in the program. Participating laboratories are located both within and outside the United States and consist primarily of hospitals, blood banks, health departments, and independent laboratories. The responses to a questionnaire completed by 1050 program-participant laboratories in September 1988 suggest that at several stages in the HIV-1 antibody total testing process, laboratory practices (including the interpretation of Western blot patterns) are variable and that standardization of these practices would improve quality.

Centers for Disease Control and Prevention, U.S.

Locality-aware pooling enhances protein language model performance across varied applications.

MOTIVATION: Protein language models (PLMs) are amongst the most exciting recent advances for characterizing protein sequences, and have enabled a diverse set of applications, including structure determination, functional property prediction, and mutation impact assessment, all from single protein sequences alone. State-of-the-art PLMs leverage transformer architectures originally developed for natural language processing, and are pre-trained on large protein databases to generate contextualized representations of individual amino acids. To harness the power of these PLMs to predict protein-level properties, these per-residue embeddings are typically "pooled" to fixed-size vectors that are further utilized in downstream prediction networks. Common pooling strategies include Cls-Pooling and Avg-Pooling, but neither of these approaches can capture the local substructures and long-range interactions observed in proteins. RESULTS: We propose the use of attention pooling, which can naturally capture these important features of proteins. To make the expensive attention operator (quadratic in the length of the input protein) feasible in practice, we introduce bag-of-mer pooling, or BoM-Pooling, a locality-aware hierarchical pooling technique that combines windowed average pooling with attention pooling. We empirically demonstrate that both full attention pooling and BoM-Pooling outperform previous pooling strategies on three important, diverse tasks: (i) predicting the activities of two proteins as they are varied; (ii) detecting remote homologs; and (iii) predicting signaling protein interactions with peptides. Overall, our work highlights the advantages of biologically inspired pooling techniques in protein sequence modeling and is a step toward more effective adaptations of language models in biological settings. AVAILABILITY AND IMPLEMENTATION: https://github.com/Singh-Lab/bom-pooling.

Natural Language Processing

On the correlation model: performance of a movement detecting neural element in the fly visual system.

The applicability of the basic principles of the correlation model to the description of the activity of a movement detecting neuron in the third optic ganglion of the fly's visual system has been investigated. This wide field neuron is supposed to sum the outputs of a large number of correlators (i.e. multiplying units followed by time averagers) that are distributed over almost the entire eye. The model describes and predicts the experimental results in a satisfactory way if a uniformly distributed system or correlators is assumed. The sampling base of the correlators in this system equals the interommatidial angles deltaphi. The half width of the spatial sensitivity distribution of the visual inputs of the correlators, deltarho, is equal to the half width of the retinula cells of the 1--6 system.

Action Potentials

The dynamic performance model of skeletal muscle.

Applications of electrical stimulation to the nerve or muscles associated with a defunct limb joint due to stroke or spinal cord injury are a viable means of restoring a certain level of functional movement to the patient. In this article, the currently acceptable physiology of motor control is outlined and used as a criterion for electrophysiological and biomechanical performance evaluation of contemporary electrical stimulation strategies used by various systems attempting to duplicate such motor control in an effort to restore meaningful limb function. Strategies associated with surface, nerve, intramuscular, and reflex stimulation are critically reviewed with special reference to voluntary sensory motor control of a limb joint rather than an isolated muscle.

Animals

Effects of model warmth on acquisition and performance of modeled behavior in chronic psychotics.

Hypotheses derived from an earlier study were tested in experimental dyads with 32 adult chronic psychotic state hospital residents of both sexes. Patients either interacted with an adult model who was noncontingently warm and rewarding, or were not exposed initially to a model. In a second phase, the model displayed task responses and novel behaviors incidental to the task, and the patient's subsequent imitation of the model under minimal demand conditions was recorded. In a third phase, the model again displayed the same behaviors, whereupon half the patients in each group were subjected to moderate (explicit verbal request) or high (verbal request plus material rewards) incentives to imitate. Under strong incentive conditions, initial differences between groups on incidental imitation vanished, indicating that a previously positive relationship with a model facilitates imitative performance but not learning.

Chronic Disease

Criterion-referenced testing in medical technology education: Professional Performance Situation Model.

In applying the Professional Performance Situation Model to the medical technology profession, situations describing actual laboratory performance are used as a basis for defining competence. As these definitions of competence are derived, appropriate criterion-referenced (domain-referenced) assessments are designed to measure the achievement of competence. This paper describes the process by which situations representing clinical practice are derived, the extrapolation of skill and knowledge statements reflecting expected performance, the generation of domains of competence, the design of criterion-referenced assessments, and some examples of prototype instruments used to assess attainment of the competence. The techniques include multiple choice items, checklists for use in the clinical component of the educational experience, and the adaptation of the written simulation for instruction and evaluation in medical laboratory sciences education. Validity of this approach is discussed, as well as possible implications for its use in developing assessments to measure continued competence in the profession beyond the baccalaureate education.

Clinical Laboratory Techniques

Modeling human performance in running.

This paper focuses on the characteristics of a model interpreting the effect of training on athletic performance. The model theory is presented both mathematically and graphically. In the model, a systematically quantified impulse of training produces dual responses: fitness and fatigue. In the absence of training, both decay exponentially with time. With repetitive training, these responses satisfy individual recurrence equations. Fitness and fatigue are combined in a simple linear difference equation to predict performance levels appropriate to the intensity of training being undertaken. Significant observed correlation of model-predicted performance with a measure of actual performance during both training and tapering provides validation of the model for athletes and nonathletes alike. This enables specific model parameters to be estimated and can be used to optimize future training regimens for any individual.

Adult

Integrating expert knowledge into large language models improves performance for psychiatric reasoning and diagnosis.

BACKGROUND AND METHODS: The authors sought to evaluate the performance of common large language models (LLMs) in psychiatric diagnosis, and the impact of integrating expert-derived reasoning on their performance. Clinical case vignettes and associated diagnoses were retrieved from the DSM-5-TR Clinical Cases book. Diagnostic decision trees were retrieved from the DSM-5-TR Handbook of Differential Diagnosis and refined for LLM use. Three LLMs were prompted to provide diagnosis candidates for the vignettes either by directly prompting or using the decision trees. These candidates and diagnostic categories were compared against the correct diagnoses. The positive predictive value (PPV), sensitivity, and F1 statistic were used to measure performance. RESULTS: When directly prompted to predict diagnoses, the best LLM by F1 statistic (gpt-4o) had sensitivity of 76.7 % and PPV of 40.4 %. When making use of the refined decision trees, PPV was significantly increased (65.3 %) without a significant reduction in sensitivity (70.9 %). Across all experiments, the use of the decision trees statistically significantly increased the PPV, significantly increased the F1 statistic in 5/6 experiments, and significantly reduced sensitivity in 4/6 experiments. DISCUSSION: When used to predict psychiatric diagnoses from case vignettes, direct prompting of the LLMs yielded most true positive diagnoses but had significant overdiagnosis. Integrating expert-derived reasoning into the process using decision trees improved LLM performance (as measured by F1 statistic), primarily by suppressing overdiagnosis with a lower-magnitude negative impact on sensitivity. This suggests that the integration of clinical expert-derived reasoning could improve the performance of LLM-based tools in the behavioral health setting.

Humans

Natural language processing-based model to predict radiation pneumonitis in patients with locally advanced non-small cell lung cancer undergoing chemoradiotherapy: a retrospective cohort study.

BACKGROUND: Radiation pneumonitis (RP) remains a significant treatment-related toxicity in patients with unresectable, locally advanced non-small cell lung cancer (NSCLC) undergoing chemoradiotherapy (CRT). Most existing predictive models rely on static baseline demographic or dosimetry variables and lack real-time clinical applicability. We developed a novel predictive framework that integrates longitudinal symptom data extracted from clinical notes using natural language processing (NLP) with clinical and dosimetry features to improve early RP prediction. METHODS: We retrospectively identified 227 patients with locally advanced NSCLC treated with definitive CRT at a high-volume cancer center in the United States. We included all patients older than 18 years who were diagnosed between Jan 1, 2006, and Dec 31, 2022 with histologically or cytologically confirmed unresectable Stage 2 or 3 NSCLC and treated with conformal radiotherapy to a minimum dose of ≥45 Gy with or without chemotherapy. Of these, 31 RP events were identified through manual adjudication using radiologic criteria and chart review. NLP was used to extract the temporal relationship of 16 pre-specified symptoms with treatment from over 100,000 clinical notes spanning pre- and during-treatment intervals. We trained and validated machine learning models on combinations of baseline clinical data, radiation dosimetry, and NLP-derived symptom features. Model performance was evaluated using a nested cross-validation framework, with an outer cross-validation loop reserved for performance assessment and an inner cross-validation loop used for model training and integration, and summarized using area under the receiver operating characteristic curve (AUC) and partial AUC (pAUC) at high specificity thresholds. Clinical utility was evaluated using decision curve analysis (DCA). FINDINGS: The best-performing model incorporated longitudinal NLP features and achieved a median AUC of 0.759 (90% confidence interval 0.753-0.766), significantly outperforming baseline models using only dosimetry (AUC 0.613) or clinical variables (AUC 0.635). NLP-based features such as cough trajectory, shortness of breath, and wheezing were among the most important predictors. Inclusion of NLP-derived symptom data improved early identification of high-risk patients, particularly in the clinically relevant high-specificity range (pAUC 0.021 vs. 0.010 for dosimetry alone). DCA showed that the calibrated MLP model provided greater net benefit than default strategies of treating all or no patients across clinically relevant threshold possibilities. INTERPRETATION: In this early work, NLP-based extraction of longitudinal symptoms from routine clinical documentation meaningfully enhances RP prediction in patients undergoing CRT for NSCLC. This approach leverages existing electronic health record infrastructure to deliver real-time, scalable, and interpretable risk estimates, offering a pathway toward potential early intervention and personalized toxicity management. The model and DCA requires external and prospective validation before clinical deployment; as such, future work should focus on this validation and integration into clinical decision support systems. FUNDING: AstraZeneca.

Chemoradiotherapy

Modeling: optimal marathon performance on the basis of physiological factors.

This paper examines current concepts concerning "limiting" factors in human endurance performance by modeling marathon running times on the basis of various combinations of previously reported values of maximal O2 uptake (VO2max), lactate threshold, and running economy in elite distance runners. The current concept is that VO2max sets the upper limit for aerobic metabolism while the blood lactate threshold is related to the fraction of VO2max that can be sustained in competitive events greater than approximately 3,000 m. Running economy then appears to interact with VO2max and blood lactate threshold to determine the actual running speed at lactate threshold, which is generally a speed similar to (or slightly slower than) that sustained by individual runners in the marathon. A variety of combinations of these variables from elite runners results in estimated running times that are significantly faster than the current world record (2:06:50). The fastest time for the marathon predicted by this model is 1:57:58 in a hypothetical subject with a VO2max of 84 ml.kg-1.min-1, a lactate threshold of 85% of VO2max, and exceptional running economy. This analysis suggests that substantial improvements in marathon performance are "physiologically" possible or that current concepts regarding limiting factors in endurance running need additional refinement and empirical testing.

Fatigue

Development and evaluation of a machine learning model for osteoporosis risk prediction in Korean women.

BACKGROUND: The aim of this study was to develop a machine learning (ML) model for classifying osteoporosis in Korean women based on a large-scale population cohort study. This study also aimed to assess ML model performance compared with traditional osteoporosis screening tools. Furthermore, this study aimed to examine the factors influencing the risk of osteoporosis through variable importance. METHODS: Data was collected from 4199 women aged 40-69 years in the baseline survey of the Ansan and Ansung cohort of the Korean Genome and Epidemiology Study. Osteoporosis was set as the dependent variable to develop ML classification models. Independent variables included 122 factors related to osteoporosis risk, such as socio-demographic characteristics, anthropometric parameters, lifestyle factors, reproductive factors, nutrient intakes, diet quality indices, medical history, medication history, family history, biochemical parameters, and genetic factors. The six classification models were developed using ML techniques, including decision tree, random forest, multilayer perceptron, support vector machine, light gradient boosting machine, and extreme gradient boosting (XGBoost). The six ML classification models were compared with two traditional osteoporosis screening tools, including the osteoporosis risk assessment instrument (ORAI) and the osteoporosis self-assessment tool (OST). The ML model performances were evaluated and compared using the confusion matrix and area under the curve (AUC) metrics. Variable importance was assessed using the XGBoost technique to investigate osteoporosis risk factors. RESULTS: The XGBoost model showed the highest performance out of the six ML classification models, with an accuracy of 0.705, precision of 0.664, recall of 0.830, and F1 score of 0.738. Moreover, the XGBoost model showed a higher performance on AUC than ORAI and OST. Variable importance scores were identified for 69 out of the 122 variables associated with osteoporosis risk factors. Age at menopause ranked first in variable importance. Variables of arthritis, physical activities, hypertension, education level, income level; alcohol intake, potassium intake, homeostatic model assessment for insulin resistance; energy intake, vitamin C intake, gout; and dietary inflammatory index ranked in the top 20 out of the 69 variables, using the XGBoost technique. CONCLUSIONS: This study found that an XGBoost model can be utilized to classify osteoporosis in Korean women. Age at menopause is a significant factor in osteoporosis risk, followed by arthritis, physical activities, hypertension, and education level.

Humans

Toward precision prognosis: Predicting recurrence-free survival in high-grade serous ovarian cancer patients using multi-time point clinical and computed tomography radiomics data.

OBJECTIVE: To evaluate the predictive value of clinical, genomic, and radiomics features in estimating recurrence-free survival (RFS) in patients with high-grade serous ovarian carcinoma (HGSOC) treated with neoadjuvant chemotherapy (NACT). METHODS: This single-center, retrospective study included 91 patients with HGSOC who underwent treatment with NACT followed by surgery, and who had portal venous phase contrast enhanced CT imaging at baseline and after NACT. First-order texture features based on 2D segmentation were extracted from baseline and post-NACT CT images for selected disease sites using commercially available texture software. Multivariate Cox models assessed the prognostic significance of features at baseline, after NACT, and post-surgery time points, and model performance in predicting RFS was evaluated using C-statistics. RESULTS: A model including only baseline clinical data had C-statistic 0.53, while a model including both clinical and radiomics features at baseline had C-statistic 0.63. After NACT, a model including all baseline data plus the change in radiomics features between baseline and post-NACT had C-statistic 0.63. Post-surgery, a model including all baseline data plus surgical outcome had C-statistic 0.69. Incorporating changes in radiomic features between time points did not measurably enhance model performance in the post-surgery data set (C-statistic 0.7). Age, residual disease at surgery, and kurtosis were individually associated with shorter RFS. CONCLUSIONS: Radiomic features extracted from CT imaging may offer additive prognostic value for predicting RFS in HGSOC when integrated with clinical and genetic data. Our results support the potential integration of radiomic analysis with clinical data to improve outcome prediction in HGSOC.

Humans

Some effects of a model's performance on an observer's electromyographic activity.

It is suggested that motor reactions may be elicited in an observer as a consequence of his exposure to a model and that such reactions may become conditioned to environmental events. An experiment is reported in which observers showed greater EMG activity in the arm while watching models arm wrestle than while watching a model stutter, and greater lip EMG activity while watching a model stutter than while watching arm wrestling. Some evidence for conditioning was found in the arm activity of males watching wrestling.

Adolescent

Habitat radiomics predicts occult lymph node metastasis and uncovers immune microenvironment of head and neck cancer.

BACKGROUND: Occult lymph node metastasis (LNM) is a key prognostic factor for patients with head and neck squamous cell carcinoma (HNSCC). This study was to establish radiomics models derived from intratumoral, peritumoral, and habitat regions for identifying occult LNM in HNSCC. METHODS: Patients with pathologically confirmed HNSCC from three medical Centers (from March 2014 to April 2024) and The Cancer Genome Atlas (TCGA) were enrolled. Center 1 was split into training (n = 330) and internal test sets (n = 154), while Center 2 and Center 3 served as the external test set (n = 183). Genomic set (n = 50) from TCGA and single-cell RNA sequencing set (n = 6) from Center 1 were used for biological analysis. We used the intratumoral, peritumoral, and habitat volumes of interest (VOIs) to extract radiomics features, respectively. Based on Logistic Regression (LR), Support Vector Machine (SVM), and Random Forest (RF) classifiers, nine radiomics models were built to confirm the optimal predictive performance. The best-performing model, along with clinical-radiologic data, was combined to develop a hybrid model. The log-rank test was used to evaluate the model's prognostic performance. Additionally, bulk and single-cell RNA sequencing were applied for investigating the biological mechanisms underlying the optimal model. RESULTS: The RF-habitat radiomics model showed the best performance, achieving AUCs of 0.835-0.919 across all datasets. Survival analysis further confirmed the prognostic value of the RF-habitat radiomics model. The RF-habitat radiomics model and the hybrid model notably surpassed the clinical model in predictive performance. Moreover, the RF-habitat radiomics model was associated with the abundance level of exhaustion-associated CD8 + T cells, uncovering the immune microenvironment characteristics contributing to occult LNM in HNSCC. CONCLUSIONS: The RF-habitat radiomics model demonstrated excellent performance for predicting occult LNM in HNSCC across three cohorts, providing a non-invasive solution for occult LNM. Furthermore, radiogenomic analysis further revealed the biological associations of the model, primarily related to T cell dysfunction.

Humans