PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Performance benchmarking”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

A cross-sectional study of reproductive indices and fawn mortality in farmed white-tailed deer.

Data were obtained from a questionnaire administered to a random sample of Canadian and United States white-tailed deer (WTD) farmers. Reproductive indices and survival of fawns from birth until 1 y of age were examined. Major factors in limiting herd increase were a low reproductive rate (88 fawns per 100 does exposed to bucks) and a 30% mortality of fawns from birth until 1 y of age. The latter figure differs from reported mortality rates in fallow deer and red deer/wapiti. The unacceptably high neonatal mortality on WTD farms was determined to be as important to herd productivity as failure to produce a live fawn. Industry wide, "benchmark" estimates of reproductive performance, mortality rates, and productivity are provided, allowing farmers to compare their herds against these "benchmarks" to identify areas needing improvement.

Animal Husbandry↗

Benchmarking in healthcare: evaluating data and transforming it into action.

After the benchmarking team has accumulated data for the development of comparisons, it must be validated for completeness and consistency. Various factors can skew the analysis and should be watched: Subjective interpretations of survey questions. Lack of common definitions. Composition of input data. External factors and extraordinary events. After verifying consistency of data gathered from benchmarking partners, calculate appropriate statistics for the performance metric. Typical data tabulations are: mean, median, ratio, minimum and maximum value, normal operating range, standard deviation and correlation coefficient. Statistics derived from the data produce the benchmark against which you will measure your institution's performance. Gap analysis establishes the difference between your internal operation's performance and that of the benchmark. Information developed from properly collected data will help you determine reasons for the gap between your performance and the benchmark and project future trends. The next step is to develop an action plan based on what has been learned from benchmarking and targeted to improving performance in areas that further the strategic goals of the institution. In addition to determining performance goals, you must analyze the decision-making process involved in making changes that will move you toward those goals. Since the goal of benchmarking is to improve the organization, the team must present its analysis to those members of management who can approve an action plan. Changes resulting from benchmarking can range from incremental improvement of existing practices all the way to reenginering. Action plans are designed to effect change at levels that will vary according to the goals that have been set. The more incremental the change, the easier the implementation. The more radical the change, the greater the reward.

Data Display↗

All are not equal: a benchmark of different homology modeling programs.

Modeling a protein structure based on a homologous structure is a standard method in structural biology today. In this process an alignment of a target protein sequence onto the structure of a template(s) is used as input to a program that constructs a 3D model. It has been shown that the most important factor in this process is the correctness of the alignment and the choice of the best template structure(s), while it is generally believed that there are no major differences between the best modeling programs. Therefore, a large number of studies to benchmark the alignment qualities and the selection process have been performed. However, to our knowledge no large-scale benchmark has been performed to evaluate the programs used to transform the alignment to a 3D model. In this study, a benchmark of six different homology modeling programs- Modeller, SegMod/ENCAD, SWISS-MODEL, 3D-JIGSAW, nest, and Builder-is presented. The performance of these programs is evaluated using physiochemical correctness and structural similarity to the correct structure. From our analysis it can be concluded that no single modeling program outperform the others in all tests. However, it is quite clear that three modeling programs, Modeller, nest, and SegMod/ ENCAD, perform better than the others. Interestingly, the fastest and oldest modeling program, SegMod/ ENCAD, performs very well, although it was written more than 10 years ago and has not undergone any development since. It can also be observed that none of the homology modeling programs builds side chains as well as a specialized program (SCWRL), and therefore there should be room for improvement.

Amino Acid Sequence↗

Variability of competitive performance of distance runners.

PURPOSE: The typical variation in an athlete's performance from race to race sets a benchmark for assessing the utility of performance tests and the magnitude of factors affecting medal prospects. We report here the typical variation in competitive performance of endurance runners. METHODS: Repeated-measures analysis of log-transformed official race times provided the typical within-athlete variation in performance as coefficients of variation (CV). The types of race were cross-country runs (4 races over 9 wk), summer road runs (5 races over 4 wk), winter road runs (4 races over 9 wk), half marathons (3 races over 13 wk and 2 races over 22 wk), and marathons (2 races over 22 wk). RESULTS: Typical variation of times for the fastest quartile of male runners was 1.2-1.9% in the cross-country and road runs, 2.7% and 4.2% in half marathons, and 2.6% in marathons. Times for the slower half of runners in most events were more variable than those of the faster half (ratio of slower/faster CV, 1.0-2.3). Times of younger adult runners were more variable than times of older runners (ratio of younger/older CV, 1.1-1.8). Times of male runners were generally more variable than those of female runners (ratio of male/female CV, 0.9-1.7). CONCLUSION: Tests of endurance power suitable for assessing the smallest worthwhile changes in running performance for top runners need CV < or = 2.5% and < or = 1.5% for tests simulating half or full marathons and shorter running races, respectively. Most of the differences in variability of race times between types of race, ability groups, age groups, and sexes probably arise from differences in competitive experience and attitude toward competing.

Adolescent↗

Asymmetric subsethood-product fuzzy neural inference system (ASuPFuNIS).

This paper presents an asymmetric subsethood-product fuzzy neural inference system (ASuPFuNIS) that directly extends the SuPFuNIS model by permitting signal and weight fuzzy sets to be modeled by asymmetric Gaussian membership functions. The asymmetric subsethood-product network admits both numeric as well as linguistic inputs. Input nodes, which act as tunable feature fuzzifiers, fuzzify numeric inputs with asymmetric Gaussian fuzzy sets; and linguistic inputs are presented as is. The antecedent and consequent labels of standard fuzzy if-then rules are represented as asymmetric Gaussian fuzzy connection weights of the network. The model uses mutual subsethood based activation spread and a product aggregation operator that works in conjunction with volume defuzzification in a gradient descent learning framework. Despite the increase in the number of free parameters, the proposed model performs better than SuPFuNIS, on various benchmarking problems, both in terms of the performance accuracy and architectural economy and compares excellently with other various existing models with a performance better than most of them.

Algorithms↗

Detecting and preventing the occurrence of errors in the practices of laboratory medicine and anatomic pathology: 15 years' experience with the College of American Pathologists' Q-PROBES and Q-TRACKS programs.

This review extracts those studies from the CAP Q-PROBES and Q-TRACKS programs that have benchmarked and monitored the occurrence of errors in the practices of laboratory medicine and anatomic pathology. The outcomes of these studies represent in aggregate the analysis of millions of data points collected in thousands of hospitals throughout the United States. Also presented in this review are hospital and laboratory practices associated with improved performance (ie, fewer errors). Only those associations that were shown to be statistically significant are presented. They represent only a small fraction of the practices examined in these studies. The reader is encouraged to peruse the Q-PROBES studies cited in the reference list to learn about the wide range of practices investigated. The institution of some of these practices for which the associated error reductions were not statistically significant might nonetheless improve performance in some environments. There is no way of knowing whether some better-performing institutions compensated for not employing presumably beneficial practices by applying other practices about which the studies' authors neglected to inquire. Nor is there any way of knowing whether institutions in which performance was poor employed presumably beneficial practices, but possessed operational flaws about which the studies' authors neglected to inquire. Certainly, hospitals operating in the bottom 10% of benchmarked performances would do well to investigate the possibility that some of these practices might reduce the incidence of errors in their institutions. From the results of these studies, there emerge two complementary strategies that appear to be associated with reduction of errors. Obviously, the first strategy involves doing what is necessary to prevent the occurrence of errors in the first place. Several tactics may accomplish this goal. Healthcare workers responsible for specific tasks must be properly educated and motivated to perform those tasks with as few errors as possible. There must be written policies and protocols detailing responsibilities and providing contingencies when those responsibilities are not met. The successful completion of required tasks must be documented, especially those tasks that are performed as requisite to others. In other words, it should be impossible to move on to subsequent operations in testing processes before documenting the successful completion of previous requisite operations. Finally, the opportunities for making errors must be reduced. Specifically, the number of steps in which specimens are delivered to laboratories, tests are performed, and results are disseminated to those who use them must be reduced as much as possible. The second strategy involves the assumption that despite our best efforts to prevent them, errors will occur. No matter how smart we are, no matter how careful we try to be, we will make mistakes. It is essential that systems designed to eliminate errors include elements of redundancy to catch those mistakes. Work must be checked and verified before therapeutic decisions are finalized. This is especially true when those decisions are irrevocable and the potential damage caused by errors cannot be undone. Ideally, systems that use redundancy should include provisions to shut down the testing process altogether when the successful execution of previous steps cannot be verified. Once error detection systems are established, service providers can gauge their performance by employing tools of continuous monitoring to assess the degree to which health care workers comply with required procedures, and with which services achieve their intended outcomes.

Diagnostic Errors↗

DARKIN: a zero-shot benchmark for phosphosite-dark kinase association using protein language models.

MOTIVATION: Protein language models (pLMs) have emerged as powerful tools for capturing the intricate information encoded in protein sequences, facilitating various downstream protein prediction tasks. With numerous pLMs available, there is a critical need for diverse benchmarks to systematically evaluate their performance across biologically relevant tasks. Here, we introduce DARKIN, a zero-shot classification benchmark designed to assign phosphosites to understudied kinases, termed dark kinases. Kinases, which catalyze phosphorylation, are central to cellular signaling pathways. While phosphoproteomics enables the large-scale identification of phosphosites, determining the cognate kinase responsible for the phosphorylation event remains an experimental challenge. RESULTS: In DARKIN, we prepared training, validation, and test folds that respect the zero-shot nature of this classification problem, incorporating stratification based on kinase groups and sequence similarity. We evaluated multiple pLMs using two zero-shot classifiers: a novel, training-free k-NN-based method, and a bilinear classifier. Our findings indicate that ESM, ProtT5-XL, and SaProt exhibit superior performance on this task. DARKIN provides a challenging benchmark for assessing pLM efficacy and fosters deeper exploration of under-characterized (dark) kinases by offering a biologically relevant test bed. AVAILABILITY AND IMPLEMENTATION: The DARKIN benchmark data and the scripts for generating additional splits are publicly available at: https://github.com/tastanlab/darkin.

Protein Kinases↗

Statistical benchmarks for process measures of quality of care for mental and substance use disorders.

OBJECTIVE: Benchmarks, representing the level of performance achieved by the best-performing providers, can be used to set achievable goals for improving care, but they have not heretofore been available for mental health care. This article describes the application of a method for developing statistical benchmarks for 12 process measures of quality of care for mental and substance use disorders. METHODS: Twelve quality measures--taken from a core measure set selected by a multistakeholder panel through a formal consensus process--were constructed from 1994-1995 administrative data on care received by Medicaid beneficiaries in six states. Conformance rates were calculated at the provider level and presented as means, 90th-percentile results, and statistical benchmarks. Sample sizes for each measure ranged from 356 to 4,494 providers and from 1,205 to 78,627 cases. Three measures involved antidepressant treatment, two involved antipsychotic treatment, and one involved mood stabilizers for bipolar disorder. Six other measures involved follow-up treatment visits. RESULTS: Benchmarks for provider-level performance ranged from 59.7 percent to 97.7 percent, markedly higher than the mean results, which ranged from 9.4 percent to 65.4 percent. Benchmark results varied widely-in contrast to results for these measures at the 90th percentile of providers and in contrast to performance standards that apply the same numerical goal across varied clinical processes. CONCLUSIONS: Statistical benchmarks can be applied to results from quality assessment of mental health care. Further research should examine whether incorporating benchmarks into quality improvement activities leads to better mental health care and substance-related care and improved outcomes.

Benchmarking↗

Effective nonanatomical endoscopy training produces clinical airway endoscopy proficiency.

We studied the effectiveness of two nonanatomical endoscopic dexterity training models: "Choose the Hole" and Dexter. Effectiveness was assessed in terms of time spent training, subjective rating, performance on an anatomical manikin, and clinical performance on fellow participants who acted as awake subjects. Forty-three anesthesia specialists, trainees, and technicians volunteered. Performances were videotaped, timed, and scored with a Global Rating Scale (GRS) from 1 (very poor) to 5 (clearly superior). The Dexter group spent more time training than the Choose the Hole group (median time [range], 152 min [70-510 min] versus 75 min [17-281 min]; P < 0.01). Subjective ratings were better in the Dexter group. In clinical bronchoscopy, the Dexter group was faster (30.7 s [17.1-43.5 s] versus 36.6 s [22.8-105.1 s]; P = 0.02) and had higher GRS scores (mean [sd]: 3.0 [0.4] versus 2.6 [0.6]; P = 0.04), indicating superior performance. Clinical and manikin performance (GRS scores) were significantly correlated (rho = 0.62; P = 0.0001). Benchmark levels of clinical bronchoscopic performance can be anticipated from bench model performance without a clinical learning curve. Dexter is a more effective model for learning endoscopic dexterity than the Choose the Hole model. Airway topicalization with lidocaine in a dose range consistent with published series (490-980 mg or 7.14-14.77 mg/kg) resulted in a frequent incidence of side effects. No major adverse events occurred.

Bronchoscopy↗

Structured robotic colorectal training in a non-tertiary NHS hospital: a 502-case consecutive cohort implementation study.

Robotic-assisted colorectal surgery has expanded rapidly across NHS practice in the UK. Structured unit-wide training pathways are essential for safe technology adoption, yet published outcome data from non-tertiary hospitals remain limited. This study describes the implementation and feasibility of a unit-wide robotic colorectal program at a high-volume non-tertiary hospital, reporting outcomes across 502 consecutive resections performed by eight consultant surgeons and presenting these in the context of nationally published benchmarks. A retrospective cohort study of 502 consecutive robotic colorectal resections performed at York Teaching Hospital between May 2022 and December 2025. Eight consultant surgeons (A-H) participated in a structured four-phase training pathway incorporating simulation training, proctored cases, complexity-based case progression, and formal credentialing. Primary outcomes were 30-day mortality, unplanned return to theatre (RTT), and anastomotic leak (AL). Anastomotic leak was calculated using only patients who underwent anastomosis as the denominator. Procedure-stratified and individual surgeon outcomes with 95% confidence intervals were reported. Risk-adjusted cumulative sum (RA-CUSUM) analysis was performed to evaluate learning curves. Outcomes are presented descriptively alongside nationally published reference data; no formal statistical comparison against national benchmarks was performed. 502 robotic colorectal resections were performed. Mean patient age was 70.0 &#xb1; 11.3&#xa0;years; 58.4% were male. Median ASA grade was III. The indication was malignancy in 89.2% of cases. Length of stay was non-normally distributed and is therefore reported using median and interquartile range in the revised analysis. Key outcomes: - 30-day mortality: 1.0% (5/502; 95% CI 0.4-2.3%) - Unplanned return to theatre (RTT): 5.2% (26/502; 95% CI 3.6-7.5%) - Anastomotic leak (AL): 3.3% (15/450; 95% CI 2.0-5.5%; denominator = patients with anastomosis) - 30-day unplanned readmission: 5.0% (25/502; 95% CI 3.4-7.2%) - Conversion to open surgery: 3.6% (18/502; 95% CI 2.3-5.6%) - Lymph node yield &#x2265;12: 91.3% of cancer resections - R0 resection rate: 95.1% of cancer resections All primary outcomes fell within or below the published reference ranges used for descriptive context. RA-CUSUM trajectories were heterogeneous: no surgeon crossed the predefined upper control limit, but several curves showed later upward movement. Accordingly, the analysis is interpreted as safety surveillance rather than evidence of uniform performance improvement. RA-CUSUM monitoring showed that no surgeon crossed the predefined upper control limit; however, heterogeneous trajectories precluded a claim of uniform performance improvement.

Humans↗

Benchmarking in ambulatory surgery.

The health care industry is relatively new to benchmarking. More clinical benchmarking is needed because little is known about which practices and processes lead to which outcomes. Benchmarking is a valuable quality improvement tool that can be used to improve practices and performances when instituted properly. This article describes the benchmarking process, its usefulness, and how ambulatory surgery centers can improve performance by using benchmarking.

Ambulatory Surgical Procedures↗

Estimating changes in unrecorded alcohol consumption in Norway using indicators of harm.

AIM: To assess the value of using indicators of alcohol-related harm to estimate changes in unrecorded per capita consumption of alcohol. DESIGN: Unrecorded consumption was estimated from the discrepancy between the observed changes in a number of alcohol-related harm indicators and the changes that would be expected from changes in recorded consumption. The results were compared with estimates of unrecorded consumption from survey data. MEASUREMENTS: Four indicators of alcohol-related harm were used: alcohol-related mortality, assaults, drunken driving, and suicide. Estimates of unrecorded consumption from survey data for five different years were used as benchmarks. FINDINGS: The best performing indicators were alcohol-related mortality, suicide and assaults, in that order. Combining these indicators yielded a prediction error averaging 12% in comparison with the benchmarks. CONCLUSIONS: The method seems worthy of further applications, but it should be regarded as a supplement rather than as a substitute for other approaches.

Alcohol Drinking↗

The link between benchmarking and shareholder value.

Strategic performance can be linked to shareholder value by measuring the positive spread between a company's return on capital employed and the cost of capital. If managers see a significant performance gap in their spread, compared with that of a premier company, they should redefine their strategic goals.

Costs and Cost Analysis↗

Benchmark of biomarker identification and prognostic modeling methods on diverse censored data.

The practices of identifying biomarkers and developing prognostic models using genomic data has become increasingly prevalent. Such data often features characteristics that make these practices difficult, namely high dimensionality, correlations between predictors, and sparsity. Many modern methods have been developed to address these problematic characteristics while performing feature selection and prognostic modeling, but a large-scale comparison of their performances in these tasks on diverse right-censored time to event data (aka survival time data) is much needed. We have compiled many existing methods, including some machine learning methods, several which have performed well in previous benchmarks, primarily for comparison in regards to variable selection capability, and secondarily for survival time prediction on many synthetic datasets with varying levels of sparsity, correlation between predictors, and signal strength of informative predictors. For illustration, we have also performed multiple analyses on a publicly available and widely used cancer cohort from The Cancer Genome Atlas using these methods. We evaluated the methods through extensive simulation studies in terms of the false discovery rate, F1-score, concordance index, Brier score, root mean square error, and computation time. Of the methods compared, CoxBoost and the Adaptive LASSO performed well in all metrics, and the LASSO and elastic net excelled when evaluating concordance index and F1-score. The Benjamini-Hoschberg and q-value procedures showed volatile performances in controlling the false discovery rate. Some methods' performances were greatly affected by differences in the data characteristics. With our extensive numerical study, we have identified the best performing methods for a plethora of data characteristics using informative metrics. This will help cancer researchers in choosing the best approach for their needs when working with genomic data.

Humans↗

Benchmarking of Monte Carlo based shutdown dose rate calculations for applications to JET.

The calculation of dose rates after shutdown is an important issue for operating nuclear reactors. A validated computational tool is needed for reliable dose rate calculations. In fusion reactors neutrons induce high levels of radioactivity and presumably high doses. The complex geometries of the devices require the use of sophisticated geometry modelling and computational tools for transport calculations. Simple rule of thumb laws do not always apply well. Two computational procedures have been developed recently and applied to fusion machines. Comparisons between the two methods showed some inherent discrepancies when applied to calculation for the ITER while good agreement was found for a 14 MeV point source neutron benchmark experiment. Further benchmarks were considered necessary to investigate in more detail the reasons for the different results in different cases. In this frame the application to the Joint European Torus JET machine has been considered as a useful benchmark exercise. In a first calculational benchmark with a representative D-T irradiation history of JET the two methods differed by no more than 25%. In another, more realistic benchmark exercise, which is the subject of this paper, the real irradiation history of D-T and D-D campaigns conducted at JET in 1997-98 were used to calculate the shut-down doses at different locations, irradiation and decay times. Experimental dose data recorded at JET for the same conditions offer the possibility to check the prediction capability of the calculations and thus show the applicability (and the constraints) of the procedures and data to the rather complex shutdown dose rate analysis of real fusion devices. Calculation results obtained by the two methods are reported below, comparison with experimental results give discrepancies ranging between 2 and 10. The reasons of that can be ascribed to the high uncertainty on the experimental data and the unsatisfactory JET model used in the calculation. A new dedicated JET benchmark experiment will be performed trying to solve these issues.

Algorithms↗