PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Performance benchmarking”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Impact of patient characteristics, complications, and facility volume on the costs and time of cardiac catheterization and coronary angioplasty in 70 catheterization laboratories.

Although over 1 million procedures are performed in cardiac catheterization laboratories (CCLs) annually, little comparative data exist on costs or resource use in these settings. In this study, data from 70 CCLs were used to profile CCL times and total direct costs for 2 high-volume procedures: left heart catheterization (LHC) and percutaneous transluminal coronary angioplasty (PTCA) with or without stent placement. In total, 70,677 consecutive patient examinations for a 12-month period from January 1, 1998 to December 31, 1998 were analyzed. For LHC mean total direct costs averaged $306, whereas for PTCA catheterization laboratory costs averaged $3,172. The average total times for these procedures were 63 and 108 minutes, respectively. Seventy-two percent of the PTCA patients underwent coronary stenting with an associated incremental cost of $1,244. By multivariate linear regression, baseline patient characteristics such as age, gender, and clinical factors had little impact on total time and total costs. The major determinants of CCL time and cost were procedural factors (e.g., number and type of interventions) and in-lab complications, including profound hypotension, abrupt vessel closure, and emergency bypass surgery. Using facility procedure volume as a proxy for potential economies of scale, we found no relation between CCL volume and total direct CCL costs. There did appear to be a significant inverse relation between facility volume and total procedural time with CCLs that performed the highest volumes of LHC and PTCA procedures saving an average of 5 to 9 minutes per procedure. These findings may be useful in defining specific time and cost benchmarks for these commonly performed procedures and serve to underscore the critical role of reducing complications in both quality improvement and cost-saving efforts.

Aged↗

Ranking hospitals according to acute myocardial infarction mortality: should transfers be included?

OBJECTIVE: The objective of this population-based observational cohort study was to estimate the extent to which the inclusion/exclusion of transferred patients with acute myocardial infarction (AMI) impacts on hospital performance rankings. SUBJECTS: The authors studied 91,633 adult patients admitted to 116 acute care hospitals in Quebec, Canada, with a primary diagnosis of AMI between 1992 and 1999. MAIN OUTCOME MEASURE: Hospital performance ranks, based on 30-day AMI mortality rates, were estimated with hierarchical models and compared using 3 different methods for handling transferred patients (exclude all transfers; include transfers and assign outcome to the referring hospital; include transfers and assign outcome to the receiving hospital). The explanatory variable of interest was the hospital to which the patient's outcome was attributed. RESULTS: Using the 3 methods, 4 hospitals were ranked "best performers" once, and 1 hospital ranked among the best in 2 of the 3 analyses performed. Nine hospitals were ranked "worst performers" at least once (4 of which ranked among the "worst" once only, 2 ranked among the "worst" twice, and 3 were consistently ranked "worst performers" in all analyses). There was significant variation in mortality rates among hospitals, and the difference in the rates between the highest and lowest ranking hospitals exceeded the clinically relevant benchmark of 1%. CONCLUSIONS: Performance evaluation studies that compare hospital mortality rates typically exclude transferred patients. However, methods used to deal with AMI patient transfers influenced hospital ranks when comparing 30-day mortality rates. Excluding transfers may lead to an inaccurate depiction of the quality of healthcare services in regionalized healthcare systems that call for the timely interhospital transfer of patients with AMI.

Aged↗

Computational tradeoffs in multiplex PCR assay design for SNP genotyping.

BACKGROUND: Multiplex PCR is a key technology for detecting infectious microorganisms, whole-genome sequencing, forensic analysis, and for enabling flexible yet low-cost genotyping. However, the design of a multiplex PCR assays requires the consideration of multiple competing objectives and physical constraints, and extensive computational analysis must be performed in order to identify the possible formation of primer-dimers that can negatively impact product yield. RESULTS: This paper examines the computational design limits of multiplex PCR in the context of SNP genotyping and examines tradeoffs associated with several key design factors including multiplexing level (the number of primer pairs per tube), coverage (the % of SNP whose associated primers are actually assigned to one of several available tube), and tube-size uniformity. We also examine how design performance depends on the total number of available SNPs from which to choose, and primer stringency criterial. We show that finding high-multiplexing/high-coverage designs is subject to a computational phase transition, becoming dramatically more difficult when the probability of primer pair interaction exceeds a critical threshold. The precise location of this critical transition point depends on the number of available SNPs and the level of multiplexing required. We also demonstrate how coverage performance is impacted by the number of available snps, primer selection criteria, and target multiplexing levels. CONCLUSION: The presence of a phase transition suggests limits to scaling Multiplex PCR performance for high-throughput genomics applications. Achieving broad SNP coverage rapidly transitions from being very easy to very hard as the target multiplexing level (# of primer pairs per tube) increases. The onset of a phase transition can be "delayed" by having a larger pool of SNPs, or loosening primer selection constraints so as to increase the number of candidate primer pairs per SNP, though the latter may produce other adverse effects. The resulting design performance tradeoffs define a benchmark that can serve as the basis for comparing competing multiplex PCR design optimization algorithms and can also provide general rules-of-thumb to experimentalists seeking to understand the performance limits of standard multiplex PCR.

Algorithms↗

The quality of laboratory testing today: an assessment of sigma metrics for analytic quality using performance data from proficiency testing surveys and the CLIA criteria for acceptable performance.

To assess the analytic quality of laboratory testing in the United States, we obtained proficiency testing survey results from several national programs that comply with Clinical Laboratory Improvement Amendments (CLIA) regulations. We studied regulated tests (cholesterol, glucose, calcium, fibrinogen, and prothrombin time) and nonregulated tests (international normalized ratio [INR], glycohemoglobin, and prostate-specific antigen [PSA]). Quality was assessed on the sigma scale with a benchmark for minimum process performance of 3 sigma and a goal for world-class quality of 6 sigma. Based on the CLIA criteria for acceptable performance in proficiency testing (allowable total errors [TEa]), the national quality of cholesterol testing (TEa = 10%) estimated sigma values as 2.9 to 3.0; glucose (TEa = 10%), 2.9 to 3.3; calcium (TEa = 1.0 mg/dL), 2.8 to 3.0; prothrombin time (TEa = 15%), 1.8; INR (TEa = 20%), 2.4 to 3.5; fibrinogen (TEa = 20%), 1.8 to 3.2; glycohemoglobin (TEa = 10%), 1.9 to 2.6; and PSA (TEa = 10%), 1.2 to 1.8. The analytic quality of laboratory tests requires improvement in measurement performance and more intensive quality control monitoring than the CLIA minimum of 2 levels per day.

Benchmarking↗

Discriminative validity of the Minimally Invasive Surgical Trainer in Virtual Reality (MIST-VR) using criteria levels based on expert performance.

BACKGROUND: Increasing constraints on the time and resources needed to train surgeons have led to a new emphasis on finding innovative ways to teach surgical skills outside the operating room. Virtual reality training has been proposed as a method to both instruct surgical students and evaluate the psychomotor components of minimally invasive surgery ex vivo. METHODS: The performance of 100 laparoscopic novices was compared to that of 12 experienced (>50 minimally invasive procedures) and 12 inexperienced (<10 minimally invasive procedures) laparoscopic surgeons. The values of the experienced surgeons' performance were used as benchmark comparators (or criterion measures). Each subject completed six tasks on the Minimally Invasive Surgical Trainer-Virtual Reality (MIST-VR) three times. The outcome measures were time to complete the task, number of errors, economy of instrument movement, and economy of diathermy. RESULTS: After three trials, the mean performance of the medical students approached that of the experienced surgeons. However, 7-27% of the scores of the students fell more than two SD below the mean scores of the experienced surgeons (the criterion level). CONCLUSIONS: The MIST-VR system is capable of evaluating the psychomotor skills necessary in laparoscopic surgery and discriminating between experts and novices. Furthermore, although some novices improved their skills quickly, a subset had difficulty acquiring the psychomotor skills. The MIST-VR may be useful in identifying that subset of novices.

Adult↗

Assessment of HCFA's 1992 Medicare hospital information report of mortality following admission for hip arthroplasty.

OBJECTIVE: The Health Care Financing Administration (HCFA) produced annually from 1987 through 1994 mortality data information as part of the Medicare Hospital Information Project (MHIP) report. We assessed the validity of these data for hip arthroplasty for one state Medicare population and we analyzed the accuracy of the predictions derived from the Bailey-Makeham mortality model for this procedure. DATA SOURCES AND STUDY SETTING: The study sample consisted of claims and model data from 1,421 Medicare patients who underwent hip arthroplasty at acute care Arkansas hospitals from October 1990 through September 1991. STUDY DESIGN: Patients were stratified into two groups based on reason for surgery (fracture status): reconstruction or fracture management. Patient survival experience was compared between the two groups. The effect of fracture status on the HCFA model's predictive ability was examined empirically and via a simulation study. RESULTS: Our results indicate that hip arthroplasty patients are not uniform with regard to outcome, depending on the reason for the surgery. Patients with fracture had a much higher 30-day mortality rate than those who underwent reconstruction (p < .001). The empirical data and the simulation study suggest that the Bailey-Makeham model underestimates mortality for reconstructive surgery in fracture patients, providing a false benchmark for those institutions that perform hip arthroplasty on predominantly one category of patients. CONCLUSION: Published HCFA data concerning mortality for hip arthroplasty combines two different patient populations into one statistic. Casual examination of these data could result in a false benchmark for analysis of institutional performance. An important implication from this study for policymakers who base decisions on "report cards" or performance measurement reports is that, although they are necessary,generic case-mix, comorbidity, and severity of illness adjustments may not be sufficient to achieve accurate representations of outcomes, and that more disease/procedure--specific adjustments may be needed to avoid inappropriate conclusions.

Arkansas↗

Infection rates drop at hospitals participating in reporting system.

Data Benchmarks: The rates of nosocomial and surgical-wound infections have plummeted at hospitals participating in the National Nosocomial Infections Surveillance, a voluntary program of the federal Centers for Disease Control and Prevention. Compare your hospital's performance with the top benchmarks in a new report and learn what organizers say are the essential components of the program's success.

Centers for Disease Control and Prevention, U.S.↗

Applying Support Vector Machines for Gene Ontology based gene function prediction.

BACKGROUND: The current progress in sequencing projects calls for rapid, reliable and accurate function assignments of gene products. A variety of methods has been designed to annotate sequences on a large scale. However, these methods can either only be applied for specific subsets, or their results are not formalised, or they do not provide precise confidence estimates for their predictions. RESULTS: We have developed a large-scale annotation system that tackles all of these shortcomings. In our approach, annotation was provided through Gene Ontology terms by applying multiple Support Vector Machines (SVM) for the classification of correct and false predictions. The general performance of the system was benchmarked with a large dataset. An organism-wise cross-validation was performed to define confidence estimates, resulting in an average precision of 80% for 74% of all test sequences. The validation results show that the prediction performance was organism-independent and could reproduce the annotation of other automated systems as well as high-quality manual annotations. We applied our trained classification system to Xenopus laevis sequences, yielding functional annotation for more than half of the known expressed genome. Compared to the currently available annotation, we provided more than twice the number of contigs with good quality annotation, and additionally we assigned a confidence value to each predicted GO term. CONCLUSIONS: We present a complete automated annotation system that overcomes many of the usual problems by applying a controlled vocabulary of Gene Ontology and an established classification method on large and well-described sequence data sets. In a case study, the function for Xenopus laevis contig sequences was predicted and the results are publicly available at ftp://genome.dkfz-heidelberg.de/pub/agd/gene_association.agd_Xenopus.

Animals↗

Benchmarking large language models for extracting biobank-derived insights into health and disease.

Biobank-scale datasets such as the UK Biobank have become foundational resources for advancing biomedical discovery. Yet the complexity and heterogeneity of these resources, spanning genomics, imaging, clinical records, and metadata, pose substantial barriers to access and interpretation. Large Language Models (LLMs) offer a promising avenue for making such datasets more navigable through natural language interfaces. However, the extent to which current general-purpose LLMs can retrieve and synthesize biobank-specific insights has not yet been systematically evaluated. In this study, we present a reproducible, multi-metric evaluation framework to benchmark the capabilities of leading LLMs. We evaluated six leading large language models: Gemini 3 Pro, Claude Opus 4.5, Claude Sonnet 4.5, GPT-5.2, Mistral Large 2, and DeepSeek V3, on four benchmark tasks designed to assess biobank-related knowledge retrieval. We evaluate model performance across six dimensions (semantic accuracy, factual correctness, domain knowledge, reasoning quality, response depth, and biobank specificity) and assessed output consistency using curated UK Biobank references and a robust random baseline. All models outperformed the baseline by 2&#xd7; to 3&#xd7;&#x2009;, with strong statistical separation (p&#x2009;<&#x2009;0.001), confirming meaningful biobank-specific knowledge retrieval. Gemini 3 Pro achieved the highest overall accuracy across tasks such as keyword synthesis, institution recognition, and topic inference, while Claude Sonnet 4.5 demonstrated the most uniform performance across evaluation dimensions. Our benchmark provides a rigorous framework for evaluating LLMs in biomedical settings. Using the UK Biobank as a real-world testbed, we highlight both the capabilities and limitations of current models, measuring their capacity to recall structured biomedical knowledge consistent with authoritative biobank metadata.

Large Language Models↗