PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Performance benchmarking”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Long-term air medical services system performance using APACHE-II and mortality benchmarking.

OBJECTIVE: Air medical transport programs have been in existence for two decades. During this time, no outcome measures have been developed for these services. The authors examined severity scoring and mortality data from their air medical service to characterize its performance and to identify trends in acuity and mortality over a 15-year period. METHODS: APACHE-II scores derived at the time of transport and hospital mortality data have been concurrently recorded in the flight database for adult transports since 1986. The authors analyzed these data and examined the correlation between APACHE-II score at the time of transport and hospital mortality for the 15-year period 1986-2001. RESULTS: 13,808 adult transports were identified. APACHE data were available for 8,204 patients (59%) and mortality for 10,845 (79%), respectively. The number of transports increased from 935 to 1,231 per year. Mean APACHE-II for all patients was 11.6 +/- 8.4. Overall mortality was 22%. Both patient acuity and mortality were trending upward over time. The correlation between APACHE-II and mortality was close and linear (mortality = 0.018 x APACHE-II -0.0243, R2 = 0.97). CONCLUSIONS: Both severity of illness and mortality of air-transported patients appear to be increasing slowly over time in response to changes in the health care system. The strong correlation between APACHE-II performed at the time of transport and mortality validates this technique for benchmarking. The slope of this correlation is an outcome-based characteristic of system performance that may allow monitoring of a system over time and comparisons between systems.

APACHE↗

Thresholds for human detection of patient setup errors in digitally reconstructed portal images of prostate fields.

PURPOSE: Computer-assisted methods to analyze electronic portal images for the presence of treatment setup errors should be studied in controlled experiments before use in the clinical setting. Validation experiments using images that contain known errors usually report the smallest errors that can be detected by the image analysis algorithm. This paper offers human error-detection thresholds as one benchmark for evaluating the smallest errors detected by algorithms. Unfortunately, reliable data are lacking describing human performance. The most rigorous benchmarks for human performance are obtained under conditions that favor error detection. To establish such benchmarks, controlled observer studies were carried out to determine the thresholds of detectability for in-plane and out-of-plane translation and rotation setup errors introduced into digitally reconstructed portal radiographs (DRPRs) of prostate fields. METHODS AND MATERIALS: Seventeen observers comprising radiation oncologists, radiation oncology residents, physicists, and therapy students participated in a two-alternative forced choice experiment involving 378 DRPRs computed using the National Library of Medicine Visible Human data sets. An observer viewed three images at a time displayed on adjacent computer monitors. Each image triplet included a reference digitally reconstructed radiograph displayed on the central monitor and two DRPRs displayed on the flanking monitors. One DRPR was error free. The other DRPR contained a known in-plane or out-of-plane error in the placement of the treatment field over a target region in the pelvis. The range for each type of error was determined from pilot observer studies based on a Probit model for error detection. The smallest errors approached the limit of human visual capability. The observer was told what kind of error was introduced, and was asked to choose the DRPR that contained the error. Observer decisions were recorded and analyzed using repeated-measures analysis of variance. RESULTS: The thresholds of detectability averaged over all observers were approximately 2.5 mm for in-plane translations, 1.6 degrees for in-plane rotations, 1 degrees for out-of-plane rotations, and 8% change in magnification for out-of-plane translations along the central axis. When one inexperienced observer is excluded, the average threshold for change in magnification is 5%. Experienced observers tended to perform better, but differences between groups were not statistically significant. Thresholds were computed as averages over all observers. Because of the broad range of observer capabilities, some detection tasks were too difficult for some observers, leading to missing threshold values in our data analysis. The missing values were excluded from computation of the average thresholds reported above. The effect of the missing values is to bias the average values toward the best human performance. CONCLUSIONS: Under favorable conditions, humans can detect small errors in setup geometry. The thresholds for error detection reported in this study are believed to represent rigorous but reasonable benchmarks that can be incorporated into studies evaluating algorithms for computer-assisted detection of setup errors in electronic portal images.

Diagnostic Errors↗

Predictive mortality models are not like fine wine.

The authors of a recent paper have described an updated simplified acute physiology score (SAPS) II mortality model developed on patient data from 1998 to 1999. Hospital mortality models have a limited range of applicability. SAPS II, Acute Physiology, Age, and Chronic Health Evaluation (APACHE) III, and mortality probability model (MPM)-II, which were developed in the early 1990s, have shown a decline in predictive accuracy as the models age. The deterioration in accuracy is manifested by a decline in the models' calibration. In particular, mortality tends to get over predicted when older models are applied to more contemporary data, which in turn leads to 'grade inflation' when benchmarking intensive care unit (ICU) performance. Although the authors claim that their updated SAPS II can be used for benchmarking ICU performance, it seems likely that this model might already be out of calibration for patient data collected in 2005 and beyond. Thus, the updated SAPS II model may be interesting for historical purposes, but it is doubtful that it can be an accurate tool for benchmarking data from contemporary populations.

Benchmarking↗

The master clinician project.

This article describes an internal benchmarking process developed and used by the Marshfield Clinic targeting the interface between productivity and service quality. The benchmarking first identified "better performing" physicians using production and service quality measures as benchmarks. This was followed by detailed interviews of "better performers" to discover their "best practices." Based on an analysis of the "best practices" information, a physician curriculum was designed and implemented to improve service quality and provider productivity. Optimal strategies for successful programs are discussed and, finally, recommendations for future research are identified.

Benchmarking↗

The comparative molecular surface analysis (COMSA): a novel tool for molecular design.

A new method allowing for 3-D QSAR analysis and the prediction of biological activity is presented. Unlike comparative molecular field analysis (CoMFA)-like techniques, it is based not on a comparison of the properties characterizing a discrete set of points but on the mean electrostatic potential (MEP) calculated and labeling specific areas defined on the molecular surface. A Kohonen self-organizing neural network and partial least square (PLS) analysis have been used for performing such an operation. The series of steroids complexing the corticosteroid (CBG) and testosterone (TBG) globulins, which forms a benchmark measuring the performance of the methods in molecular design, and a series of benzoic acids described by the Hammett sigma constants is used for testing the method. It is demonstrated that a method can be used efficiently to evaluate the responses determined both by the combination of electrostatic and steric effects or by electrostatic effects alone, therefore, two different schemes were developed. The first one, which involves PLS analysis of the full comparative networks, covers both steric and electrostatic effects. This scheme works well for both the CBG and TBG data. The second scheme takes into account only the properties (MEP) of these regions within molecules that can be superimposed with the template molecule. This scheme provides the best predictive power for the benzoic acids series. Comparison of the results from a CoMFA analysis proves that method is at least as effective for the responses limited by electrostatic effects, although it significantly outperforms CoMFA for CBG affinity which is dominated by steric effects.

Benzoates↗

Understanding benchmarking.

In order to meet the challenges facing health care today, organizations are turning to new approaches. Benchmarking is one such approach. Benchmarking is externally driven, encouraging organizations to look outside their own walls to learn from others and achieve exemplar performance. Organizations can benchmark within their own systems, against competitors, against "best-in-class" companies in the same general industry and against "best-in-class" companies in different industries. A four-step approach to benchmarking includes planning, collecting information, analyzing results and adapting and improving. A benchmarking study team composed of the process owner and other users of the process conducts the study. Application of benchmarking to healthcare materiel management is particularly appropriate, since many materiel management processes occur in other industries and, therefore, best practices outside the healthcare industry may be adapted. The practice, through growing in other industries, is still very new in health care.

Data Collection↗

The Telemedicine benchmark--a general tool to measure and compare the performance of video conferencing equipment in the telemedicine area.

In this paper, we describe the 'Telemedicine Benchmark' (TMB), which is a set of standard procedures, protocols and measurements to test reliability and levels of performance of data exchange in a telemedicine session. We have put special emphasis on medical imaging, i.e. digital image transfer, joint viewing and editing and 3D manipulation. With the TMB, we can compare the aptitude of different video conferencing software systems for telemedicine issues and the effect of different network technologies (ISDN, xDSL, ATM, Ethernet). The evaluation criteria used are length of delays and functionality. For the application of the TMB, a data set containing radiological images and medical reports was set up. Considering the Benchmark protocol, this data set has to be exchanged between the partners of the session. The Benchmark covers file transfer, whiteboard usage, application sharing and volume data analysis and compression. The TMB has proven to be a useful tool in several evaluation issues.

Benchmarking↗

Detection of vernier and contrast-modulated stimuli with equal Fourier energy spectra by infants and adults.

Infant and adult vernier acuity differed by a factor of only 4 to 6 when the stimuli were periodic and results were expressed in units of spatial phase. This ratio was much smaller than the factor of 50 to 100 obtained when we expressed our results and those of others in terms of the threshold spatial displacement in visual angle. We compared infant and adult vernier performance to performance on a "benchmark" contrast discrimination task, where the vernier and contrast discrimination stimuli contained identical Fourier contrast spectra. When we compared vernier performance directly to contrast discrimination performance, infant and adult data were remarkably similar, suggesting that similar parts of the visual system limit vernier and contrast performance of subjects of both ages. A control experiment on adults suggested that the superior performance of the contrast discrimination task is due to recruitment of visual pattern analyzers situated at a distance from the discontinuities in phase position and contrast.

Adult↗

Hospitalized patients with atrial fibrillation and a high risk of stroke are not being provided with adequate anticoagulation.

OBJECTIVES: The purpose of this study was to determine both treatment gaps and predictors of warfarin use in atrial fibrillation (AF) patients enrolled in a national multicenter study. BACKGROUND: The National Anticoagulation Benchmark Outcomes Report (NABOR) is a performance improvement program designed to benchmark anticoagulation prophylaxis, treatment, and outcomes among participating hospitals. METHODS: A retrospective cohort study of inpatients was performed at 21 teaching, 13 community, and 4 Veterans Administration hospitals in the U.S. Patients with an ICD-9-CM code for AF (427.31) were randomly selected. RESULTS: Among the 945 patients studied, the mean age was 71.5 (+/- 13.5) years; 43% were >75 years of age, 54.5% were men, and 67% had a history of hypertension. Most (86%) had factors that stratified them as at high risk of stroke, and only 55% of those received warfarin. Neither warfarin nor aspirin were prescribed in 21% of high-risk patients, including 18% of those with a previous stroke, transient ischemic attack, or systemic embolic event. Age >80 years (p = 0.008) and perceived bleeding risk (p = 0.022) were negative predictors of warfarin use. Persistent/permanent AF (p < 0.001) and history of stroke, transient ischemic attack, or systemic embolus (p = 0.014) were positive predictors of warfarin use, whereas high-risk stratification was not. CONCLUSIONS: This study confirms the under-use of warfarin, but also adds to published reports in several regards. It showed that risk stratification, the guidepost for treatment in international guidelines, had little effect on warfarin use, and that age >80 years and AF classification (permanent/persistent) are factors that influence warfarin use.

Aged↗