PubMed Health⌕ Search

Biomedical subjects

Turner M Osler

Publications and source records attributed to Turner M Osler.

At least 19 recordsLinked to original sources

Impact of changing the statistical methodology on hospital and surgeon ranking: the case of the New York State cardiac surgery report card.

BACKGROUND: Risk adjustment is central to the generation of health outcome report cards. It is unclear, however, whether risk adjustment should be based on standard logistic regression, fixed-effects or random-effects modeling. OBJECTIVE: The objective of this study was to determine how robust the New York State (NYS) Coronary Artery Bypass Graft (CABG) Surgery Report Card is to changes in the underlying statistical methodology. METHODS: Retrospective cohort study based on data from the NYS Cardiac Surgery Reporting System on all patient undergoing isolated CABG surgery in NYS and who were discharged between 1997 and 1999 (51,750 patients). Using the same risk factors as in the NYS models, fixed-effects and random-effects models were fitted to the NYS data. Quality outliers were identified using 1) the ratio of observed-to-expected mortality rates (O/E ratio) and confidence intervals (CIs) calculated using both parametric (Poisson distribution) and nonparametric (bootstrapping) techniques; and 2) shrinkage estimators. RESULTS: At the surgeon level, the standard logistic regression model, the fixed-effects model, and the fixed-effects component of the random-effects model demonstrated near-perfect agreement on the identity of quality outliers using a quality indicator based on the O/E ratio and the Poisson distribution. Shrinkage estimators identified the fewest outliers, whereas the O/E ratios with bootstrap CI identified the greatest number of outliers. The results were similar for hospitals, except that the fixed-effects model identified more outliers than either the NYS model or the fixed-effects component of the random-effects model. CONCLUSION: Shrinkage estimators based on random-effects models are slightly more conservative in identifying quality outliers compared with the traditional approach based on fixed-effects modeling and standard regression. Explicitly modeling surgeon provider effect (fixed-effects and random-effects models) did not significantly alter the distribution of quality outliers when compared with standard logistic regression (which does not model provider effect). Compared with the standard parametric approach, the use of a bootstrap approach to construct 95% confidence interval around the O/E ratio resulted in more providers being identified as quality outliers.

Benchmarking↗

Does date stamping ICD-9-CM codes increase the value of clinical information in administrative data?

CONTEXT: Comorbidity measures are designed to exclude complications when they map International Classification of Diseases (ICD-9-CM) codes to diagnostic categories. The use of data fields that indicates whether each secondary diagnosis was present at the time of hospital admission may lead to the more accurate identification of preexisting conditions. OBJECTIVE: To examine the rate of misclassification of ICD-9-CM codes into diagnostic categories by the Dartmouth-Manitoba adaptation of the Charlson index and by the Elixhauser comorbidity algorithm. DATA SOURCE: Analysis of 178,838 patients in the California State Inpatient Database (CA SID) admitted in 2000 for one of seven major medical and surgical conditions. The CA SID includes a condition present at admission (CPAA) modifier for each ICD-9-CM code. STUDY DESIGN: The Dartmouth/Charlson index and the Elixhauser comorbidity measure were used to map the ICD-9-CM codes into diagnostic categories for patients in each study population. We calculated the misclassification rate for each mapping algorithm, using information from the CPAA as the "gold standard." PRINCIPAL FINDINGS: The Dartmouth/Charlson index underestimated the prevalence of hemiplegia/paraplegia by 70 percent, cerebrovascular disease by 70 percent, myocardial infarction by 65 percent, congestive heart failure (CHF) by 45 percent, and peptic ulcer disease by 34 percent. The Elixhauser algorithm misclassified complications as preexisting conditions for 43 percent of the coagulopathies, 25 percent of the fluid and electrolyte disorders, 18 percent of the cardiac arrhythmias, 18 percent of the cardiac arrhythmias, and 9 percent of the cases of CHF. CONCLUSION: Adding the CPAA modifier to administrative data would significantly enhance the ability of the Dartmouth/Charlson index and of the Elixhauser algorithm to map ICD-9-CM codes to diagnostic categories accurately.

Algorithms↗

Evaluating trauma center quality: does the choice of the severity-adjustment model make a difference?

CONTEXT: The Major Trauma Outcome Study (MTOS) database was created by the American College of Surgeons over 20 years ago to establish national norms for trauma care. The primary trauma outcome prediction models used for evaluating the quality of trauma care, TRISS and ASCOT (A Severity Characterization of Trauma), were developed using the MTOS database. OBJECTIVE: First, to determine whether TRISS and ASCOT agree on hospital quality. Second, to determine whether TRISS and ASCOT accurately reflect contemporary outcomes in trauma care. DESIGN, SETTING AND PATIENTS: A retrospective cohort study based on 91,112 patients admitted to 69 hospitals between 2000 and 2001 in the National Trauma Databank. Using TRISS and ASCOT, the ratio of the observed to expected mortality rate (O/E ratio) was calculated for each hospital. Hospitals whose O/E ratio was statistically different from 1 were identified as quality outliers. Kappa analysis was used to assess the degree to which TRISS and ASCOT agreed on the identity of hospital quality outliers. RESULTS: TRISS and ASCOT disagreed on the outlier status of 35 of the 69 hospitals. Kappa analysis revealed only fair agreement (kappa = 0.23; p = 0.0015) between TRISS and ASCOT in identifying quality outliers. Thirty-eight hospitals were identified by the TRISS method as high-performance hospitals. CONCLUSION: First, TRISS and ASCOT exhibit substantial disagreement on the identity of quality outliers within the NTDB. Second, an unrealistically high number of hospitals were identified as high-performance outliers using either TRISS or ASCOT. These findings have important implications for the use of TRISS and ASCOT for benchmarking performance and quality improvement.

Calibration↗

The relation between surgeon volume and outcome following off-pump vs on-pump coronary artery bypass graft surgery.

STUDY OBJECTIVE: Off-pump coronary artery bypass graft (CABG) surgery has been recently reintroduced into clinical practice. In light of the relatively low level of experience of most cardiac surgeons with off-pump CABG surgery, and the exceptional technical challenge of working on a "beating heart," off-pump CABG surgery presents a unique opportunity to explore the effect of surgeon case volume on surgical outcome after controlling for the effects of patient case mix and hospital volume. DESIGN: A retrospective cohort study analyzing the association between surgeon volume and in-hospital mortality rate for off-pump and on-pump CABG surgery using random-effects logistic regression modeling. SETTING AND PATIENTS: The analyses were based on the New York State clinical CABG surgery registry. The study sample consisted of 36,930 patients undergoing isolated CABG surgery between 1998 and 1999 that was performed by 181 surgeons at 33 hospitals. INTERVENTIONS: None. RESULTS: There is no association between the number of CABG procedures performed off-pump by an individual surgeon and in-hospital mortality rates (p = 0.93) after controlling for hospital CABG surgery volume and patient-level risk factors. There is also no association between the off-pump CABG surgery mortality rate and the total number of both off-pump and on-pump CABG surgery cases (p = 0.78). In the on-pump CABG surgery cohort, surgeons performing a high volume of CABG procedures had significantly lower risk-adjusted mortality rates among their patients compared to those performing a very low volume, a low-volume, and a medium volume of CABG procedures (p < 0.006). CONCLUSION: For off-pump CABG surgery, surgeons performing a high volume of procedures do not have better mortality outcomes than those performing a low volume of procedures. However, higher surgeon case volumes are associated with lower mortality rates for on-pump CABG surgery. The absence of a volume-outcome association for off-pump CABG surgery is especially surprising in light of the more technically demanding nature of off-pump CABG surgery compared to on-pump CABG surgery.

Cohort Studies↗

Judging trauma center quality: does it depend on the choice of outcomes?

BACKGROUND: Trauma centers routinely benchmark their survival outcomes against a national norm using the TRISS methodology. However, the use of survival as a measure of the effectiveness of trauma care may be too limited in scope because it fails to capture information regarding functional outcomes. METHODS: The objective of this study was to develop a prediction model that allows hospitals to benchmark their functional outcomes in blunt trauma patients, and to determine whether the assessment of hospital "quality" depends on the choice of outcome measure: survival or survival combined with functional outcome. This retrospective cohort study was based on patients, aged 18 years or older, in the National Trauma Database who sustained blunt trauma in 1999 without associated head or spinal cord injury. We developed a sequential logistic model to predict the probability of a good functional outcome. The TRISS methodology was customized to this data set to obtain a survival model. Using each of these prediction models, we then obtained two standardized measures of hospital performance: one based on the number of survivors and the other based on the number of survivors with good functional outcomes. These standardized outcome measures were then used to identify low-performance and high-performance hospitals. The ranking based on these two different measures were compared. RESULTS: Fifteen of the 27 hospitals in the study cohort were categorized differently when their performance was benchmarked using survival versus functional outcome. Kappa analysis revealed minimal agreement between these two quality measures on the identity of hospital quality outliers (kappa = 0.04; p = 0.35). CONCLUSION: The evaluation of hospital quality depends on whether hospital performance is judged by looking at survival or at survival combined with functional outcome. Because functional status is an important outcome of major concern to survivors, it is important to include it in hospital performance assessment. Consideration should be given to including functional outcome in the evaluation of trauma center performance.

Benchmarking↗

The relation between trauma center outcome and volume in the National Trauma Databank.

BACKGROUND: Regionalization of trauma care services aims to improve outcomes by limiting trauma care delivery to a select group of dedicated trauma centers. However, the evidence linking trauma center volume and outcome is not conclusive. The objective of this study was to examine the volume-mortality relation for patients with severe trauma in the National Trauma Databank. METHODS: This study was based on data for adult patients 18 years of age or older in the National Trauma Databank with an Injury Severity Score (ISS) of 15 or more who sustained either blunt or penetrating trauma. The main outcome measure was in-hospital survival as a function of trauma center volume. Logistic regression modeling was used to analyze the relation between survival and hospital volume for patients sustaining either severe blunt or severe penetrating trauma. RESULTS: For the blunt trauma cohort, model diagnostics showed that the single highest-volume center was an outlier. After exclusion of the patients from this center, no association could be demonstrated between trauma volume and outcome (p = 0.465) for blunt trauma. A separate multivariate analysis of patients with penetrating trauma also could not demonstrate a significant volume-mortality association (p = 0.919). Both regression models exhibited excellent discrimination and acceptable calibration. CONCLUSION: The findings of this study do not support the position that higher trauma center volumes are associated with improved survival. The implication of this study is that the hospital volume criteria established by the American College of Surgeons may need to be reexamined.

Adolescent↗

A note on the disjointed nature of the injury severity score.

OBJECTIVE: The Injury Severity Score (ISS) is widely used for anatomic severity assessments. The ISS is the sum of the squares of a patient's three worst Abbreviated Injury Scale (AIS) severities (1-6) from three specified body regions. The set of three AIS severities (including 0s) is called a "triplet." ISS values of 9, 17, 18, 25, 26, 27, 29, 33, 34, 41, and 50 can originate from two unique triplets, but it is not clear whether the mortalities of the triplets are equal. A related question regards the monotonicity of the ISS, that is, whether mortality increases with successive values of ISS. This study sought to compare the mortality of equivalent ISS values from different triplets and to evaluate whether ISS is a monotonic function of mortality. METHODS: The ISS, its corresponding three-digit triplet, and the ICISS (an International Classification of Diseases, Ninth Revision-based competing score) were calculated for 361,381 National Trauma Data Bank patients. Fisher's exact tests were used to test for mortality differences between triplets that yield the same ISS. Plots of mortality by score value were produced to visually assess the monotonicity of the ICISS and the ISS. RESULTS: Six of the 11 triplet pairs had mortalities that differed by greater than 20%, with the largest difference being 32% for an ISS of 25 (triplets 0, 0, 5 and 0, 3, 4). Two other values (9 and 17) have triplet pairs whose mortality differences are less but still statistically different. The ISS is markedly nonmonotonic and is characterized by large spikes in mortality for successive ISS values. Plots of the ICISS show it to be largely monotonic. CONCLUSION: The ISS is a nonmonotonic, triplet-dependent function of mortality. Those who persist in using the ISS to describe populations or make risk adjustments should do so cautiously, being sure to account for triplet type. These suspect ISS values appear in approximately 25% of cases.

Humans↗

Using hierarchical modeling to measure ICU quality.

OBJECTIVE: To determine whether hierarchical modeling agrees with conventional logistic regression modeling on the identity of ICU quality outliers within a large multi-institutional database. DESIGN: Retrospective database analysis. SETTING AND PATIENTS: Subset of the Project IMPACT database consisting of 40435 adult patients admitted to surgical, medical, and mixed surgical-medical ICUs ( n=55) between 1997 and 1999 who met inclusion criteria for SAPS II. MEASUREMENTS AND RESULTS: The SAPS II score was customized to this database using conventional logistic regression and using a hierarchical (random coefficients) model. Both models exhibited excellent discrimination ( Cstatistic) and calibration (Hosmer-Lemeshow statistic). The hierarchical and nonhierarchical models had C statistics of.870 and.865, and HL statistics of 3.71 ( p>.88, df=8) and 8.94 ( p>.35, df=8), respectively. Since the random effects component of the hierarchical model accounts for between-hospital variability, only the fixed-effects coefficients were used to calculate the expected mortality rate based on the hierarchical model. The ratio and 95% confidence intervals of the observed to expected mortality rate were calculated using both models for each ICU. ICUs whose observed/expected ratio was either less than 1 or greater than 1, and whose 95% confidence interval did not include 1 were labeled as either high-performance or low-performance outliers, respectively. Analysis using kappa statistic revealed almost perfect agreement between the two models (nonhierarchical vs. hierarchical) on the identity of ICU quality outliers. CONCLUSIONS: Models obtained by customizing SAPS II using a nonhierarchical and a hierarchical approach exhibit excellent agreement on the identity of ICU quality outliers.

APACHE↗

Is the hospital volume-mortality relationship in coronary artery bypass surgery the same for low-risk versus high-risk patients?

BACKGROUND: There is evidence to support the existence of an inverse relation between mortality after coronary artery bypass graft (CABG) surgery and procedure volume. It is unclear whether all patients benefit equally from having CABG surgery performed at high-volume centers. The objective of this study was to determine whether the volume-outcome association for CABG surgery is modified by patient risk. METHODS: This retrospective cohort analysis was conducted using data from the Cardiac Surgery Reporting System database on all patients (20,078) undergoing CABG surgery in New York State who were discharged in 1996. The main outcome measure was in-hospital mortality as a function of procedure volume after adjusting for severity of disease. Logistic regression modeling was used to explore the interaction between patient risk and procedure volume. RESULTS: There is a significant interaction between procedure volume and patient risk (p = 0.01). The final model exhibits excellent discrimination (C statistic = 0.818) and goodness-of-fit (Hosmer-Lemeshow statistic = 6.02; p = 0.645). Very low (<0.5%) and low-risk (0.5%-2.0%) patients exhibit a greater reduction in CABG mortality than high (5.0%-10.0%) and very high risk (>10%) patients at high-volume centers relative to low-volume centers. Among the highest risk patients (>25% risk of mortality), higher risk patients have better outcomes at higher volume centers. CONCLUSIONS: For the vast majority of patients, low-risk patients benefit significantly more than high-risk patients from undergoing CABG surgery at high-volume centers instead of at low-volume centers. Low-risk patients benefit significantly more than high-risk patients from undergoing CABG surgery at high-volume centers instead of at low-volume centers. However, before generalizing these findings to other states, this study should be repeated using other regional population-based clinical databases.

Cohort Studies↗

Improving the Glasgow Coma Scale score: motor score alone is a better predictor.

BACKGROUND: The Glasgow Coma Scale (GCS) has served as an assessment tool in head trauma and as a measure of physiologic derangement in outcome models (e.g., TRISS and Acute Physiology and Chronic Health Evaluation), but it has not been rigorously examined as a predictor of outcome. METHODS: Using a large trauma data set (National Trauma Data Bank, N = 204,181), we compared the predictive power (pseudo R2, receiver operating characteristic [ROC]) and calibration of the GCS to its components. RESULTS: The GCS is actually a collection of 120 different combinations of its 3 predictors grouped into 12 different scores by simple addition (motor [m] + verbal [v] + eye [e] = GCS score). Problematically, different combinations summing to a single GCS score may actually have very different mortalities. For example, the GCS score of 4 can represent any of three mve combinations: 2/1/1 (survival = 0.52), 1/2/1 (survival = 0.73), or 1/1/2 (survival = 0.81). In addition, the relationship between GCS score and survival is not linear, and furthermore, a logistic model based on GCS score is poorly calibrated even after fractional polynomial transformation. The m component of the GCS, by contrast, is not only linearly related to survival, but preserves almost all the predictive power of the GCS (ROC(GCS) = 0.89, ROC(m) = 0.87; pseudo R2(GCS) = 0.42, pseudo R2(m) = 0.40) and has a better calibrated logistic model. CONCLUSION: Because the motor component of the GCS contains virtually all the information of the GCS itself, can be measured in intubated patients, and is much better behaved statistically than the GCS, we believe that the motor component of the GCS should replace the GCS in outcome prediction models. Because the m component is nonlinear in the log odds of survival, however, it should be mathematically transformed before its inclusion in broader outcome prediction models.

Algorithms↗

Independently derived survival risk ratios yield better estimates of survival than traditional survival risk ratios when using the ICISS.

BACKGROUND: The International Classification of Diseases, Ninth Revision Injury Severity Score (ICISS) is criticized because it relies on survival risk ratios (SRRs) that are contaminated by incidents with multiple injuries. An SRR for an International Classification of Diseases, Ninth Revision code is the number of patients who survive the injury divided by the number who display it. The ICISS is the product of SRRs that correspond to a patient's injuries. Traditional SRRs are derived from databases that include patients with multiple injuries and are biased toward mortality, making them nonindependent. Independent SRRs are derived from incidents where patients sustained only an isolated injury. The objective of this study is to compare the mortality prediction abilities of independent and traditional SRRs via the ICISS. METHODS: A 10-fold cross-validation design was used to estimate independent and traditional SRRs and their resulting ICISSs from 192,347 National Trauma Data Bank patients. Logistic regression modeled the scores as a function of mortality. The area under the receiver operating characteristic curve measured discrimination. Model fit was measured with the Akaike information criterion, a deviance statistic (lower is better). R2 values were compared to determine which score explained the most variance. RESULTS: The independent ICISS statistically outperforms the traditional ICISS. CONCLUSION: Traditional SRRs used by the ICISS produce less accurate estimates of mortality than independent SRRs. The ICISS can be calculated in 97.9% of incidents using independent SRRs.

Databases as Topic↗

The worst injury predicts mortality outcome the best: rethinking the role of multiple injuries in trauma outcome scoring.

BACKGROUND: The prediction of outcome after injury must incorporate measures of injury severity, but there is no consensus on how many injuries should be used in calculating these measures. Initially, the single worst injury was used to predict outcome, but the introduction of the Injury Severity Score allowed up to three injuries to contribute to outcome prediction. Subsequently, other outcome prediction approaches used many (New Injury Severity Score [NISS]) or all (ICISS and Trauma Registry Abbreviated Injury Scale Score [TRAIS], which use International Classification of Diseases, Ninth Revision [ICD-9] and Abbreviated Injury Scale [AIS] survival risk ratios [SRRs], respectively) of a patient's injuries. The ability of only the most severe injury in predicting mortality has never been studied. Our objective was to determine the ability of a patient's worst injury to predict mortality. METHODS: A 10-fold cross-validation design was used to compute six scores for each of 160,208 patients from a large trauma database (the National Trauma Data Bank [NTDB]). The scores were ICISS, TRAIS, ICISS1 (only a patient's worst ICD-9 SRR), TRAIS1 (only a patient's worst AIS SRR), NISS (sum of squares of worst three AIS severity measures), and MAXAIS (worst AIS severity measure). Discrimination was assessed using the area under the receiver operating characteristic curve. Logistic regression R2 gauged the proportion of variance each score explained. The Akaike information criterion, a deviance statistic (lower is better), assessed model fit. RESULTS: The receiver operating characteristic curve, R2, and Akaike information criterion statistics (NC_ICISS and NC_ICDSRR1 represents scores derived from the original North Carolina Hospital Discharge Database SRRs) are summarized in tabular form in the Results section. CONCLUSION: Regardless of scoring type (ICD/AIS SRRs or AIS severity), a patient's worst injury discriminates survival better, fits better, and explains more variance than currently used multiple injury scores.

Abbreviated Injury Scale↗

Complications in surgical patients.

HYPOTHESIS: Complications are common in hospitalized surgical patients. Provider error contributes to a significant proportion of these complications. DESIGN: Surgical patients were concurrently observed for the development of explicit complications. All complications were reviewed by the attending surgeon and other members of the service and evaluated for the severity of sequelae (major or minor) and for whether the complication resulted from medical error (avoidable) or not. SETTING: University teaching hospital with a level I trauma designation. PATIENTS: All inpatients (operative or nonoperative) from 4 different surgical services: general surgery, combined general surgery and trauma, vascular surgery, and cardiothoracic surgery. MAIN OUTCOME MEASURES: Total complication rate (number of complications divided by the number of patients) and the number of patients with complications. Complications were separated into those with major or minor sequelae and the proportion of each type that were due to medical error (avoidable). Rates of complications in a recent Institute of Medicine report were used as a criterion standard. RESULTS: The data for the respective groups (general surgery, vascular surgery, combined general surgery and trauma, and cardiothoracic surgery) are as follows. The number of patients was 1363, 978, 914, and 1403; number of complications, 413, 409, 295, and 378; total complication rate, 30.3%, 42.4%, 32.3%, and 26.9%; minor complication rate, 13.3%, 19.9%, 13.5%, and 13.0% (percentage of minor complications that were avoidable, 37.4%, 59.0%, 51.2%, and 49.5%); major complication rate, 16.2%, 21.1%, 18.1%, and 12.9% (percentage of major complications that were avoidable, 53.4%, 60.7%, 38.8%, and 38.7%); and mortality rate, 1.83%, 3.33%, 2.28%, and 3.34% (percentage of mortality that was avoidable, 28.0%, 44.1%, 19.0%, and 25.0%). CONCLUSIONS: Despite mortality rates that compare favorably with national benchmarks, a prospective examination of surgical patients reveals complication rates that are 2 to 4 times higher than those identified in an Institute of Medicine report. Almost half of these adverse events were judged contemporaneously by peers to be due to provider error (avoidable). Errors in care contributed to 38 (30%) of 128 deaths. Recognition that provider error contributes significantly to adverse events presents significant opportunities for improving patient outcomes.

Cardiac Surgical Procedures↗

Synergistic effect of genistein and BCNU on growth inhibition and cytotoxicity of glioblastoma cells.

OBJECTIVE: Recent experiments have shown that dietary soy isoflavones such as genistein can significantly suppress invasiveness and growth of a number of human malignancies. This study examined whether genistein, at a concentration typical of plasma levels following soy diet intake, in combination with 1,3-bis(2-chloroethyl)-1-nitrosourea (BCNU, carmustine) exhibited an additive or synergistic inhibitory effect on the growth of glioma cells. METHODS: The human glioblastoma multiforme (GBM) cell line U87 and the rodent C6 glioma were treated with genistein at 4 microM, combined with BCNU (0-50 microM). Monolayer cell growth and cytotoxicity, as measured by colonigenic survival in soft agarose, were then compared in control and drug-treated cultures. Presence of apoptosis, using the DNA ladder assay and laser scanning cytometry (LSC), was investigated in all cell lines at those concentrations where an enhancement of antiproliferative effect of BCNU in presence of genistein was observed. RESULTS: A 32-41% increase in monolayer growth inhibition and a 28-42% increase in colony cytotoxicity in the U87 cell line were observed when genistein (4 microM) was added to BCNU in the 0-10 microM dose range. In the C6 cell line, a 30-36% increase in monolayer growth inhibition and a 39-54% increase in colony cytotoxicity were observed with the BCNU dose range of 0-50 microM. All experiments showed a significant increase in growth inhibition and a decrease in colonogenic survival (P < 0.05). We were unable to detect apoptosis in any of the lines when genistein was combined with BCNU. CONCLUSION: These results indicate that genistein at typical adult dietary plasma levels can significantly enhance the antiproliferative and cytotoxic action of BCNU. The implication for treatment of GBM may be a reduction in the chemotherapeutic dose recommendations of these agents and subsequently a decrease in the risk of treatment sequelae for these patients.

Animals↗

Rating the quality of intensive care units: is it a function of the intensive care unit scoring system?

OBJECTIVE: Intensive care units (ICUs) use severity-adjusted mortality measures such as the standardized mortality ratio to benchmark their performance. Prognostic scoring systems such as Acute Physiology and Chronic Health Evaluation (APACHE) II, Simplified Acute Physiology Score II, and Mortality Probability Model II0 permit performance-based comparisons of ICUs by adjusting for severity of disease and case mix. Whether different risk-adjustment methods agree on the identity of ICU quality outliers within a single database has not been previously investigated. The objective of this study was to determine whether the identity of ICU quality outliers depends on the ICU scoring system used to calculate the standardized mortality ratio. DESIGN, SETTING, PATIENTS: Retrospective cohort study of 16,604 patients from 32 hospitals based on the outcomes database (Project IMPACT) created by the Society of Critical Care Medicine. The ICUs were a mixture of medical, surgical, and mixed medical-surgical ICUs in urban and nonurban settings. Standardized mortality ratios for each ICU were calculated using APACHE II, Simplified Acute Physiology Score II, and Mortality Probability Model II. ICU quality outliers were defined as ICUs whose standardized mortality ratio was statistically different from 1. Kappa analysis was used to determine the extent of agreement between the scoring systems on the identity of hospital quality outliers. The intraclass correlation coefficient was calculated to estimate the reliability of standardized mortality ratios obtained using the three risk-adjustment methods. MEASUREMENTS AND MAIN RESULTS: Kappa analysis showed fair to moderate agreement among the three scoring systems in identifying ICU quality outliers; the intraclass correlation coefficient suggested moderate to substantial agreement between the scoring systems. The majority of ICUs were classified as high-performance ICUs by all three scoring systems. All three scoring systems exhibited good discrimination and poor calibration in this data set. CONCLUSION: APACHE II, Simplified Acute Physiology Score II, and Mortality Probability Model II0 exhibit fair to moderate agreement in identifying quality outliers. However, the finding that most ICUs in this database were judged to be high-performing units limits the usefulness of these models in their present form for benchmarking.

APACHE↗

Identifying quality outliers in a large, multiple-institution database by using customized versions of the Simplified Acute Physiology Score II and the Mortality Probability Model II0.

OBJECTIVE: To assess whether customized versions of the Simplified Acute Physiology Score (SAPS) II and the Mortality Probability Model (MPM) II0 agree on the identity of intensive care unit quality outliers within a multiple-center database. DESIGN: Retrospective database analysis. SETTING AND PATIENTS: Patient subset of the Project IMPACT database consisting of 39,617 adult patients admitted to surgical, medical, and mixed surgical-medical intensive care units at 54 hospitals between 1995 and 1999 who met inclusion criteria for SAPS II and MPM II0. INTERVENTIONS: Customized versions of SAPS II and MPM II0 were obtained by fitting new logistic regressions to the data by using the risk score as the independent variable and outcome at hospital discharge as the dependent variable. The data set was divided randomly into a training set and a validation set. Each model was customized by using the training set; model performance was then assessed in the validation set by using the area under the receiver operating characteristic curve and the Hosmer-Lemeshow statistic. The final models were based on the entire data set. The level of agreement between the customized models on the identity of quality outliers was evaluated by using kappa analysis. MEASUREMENTS AND MAIN RESULTS: Both customized models exhibited good discrimination and good calibration in this database. The area under the receiver operating characteristic curve was 0.83 for MPM II0 and 0.872 for SAPS II following model customization. The Hosmer-Lemeshow statistic was 12.3 ( >.14) for MPM II0, and 8.17 (p >.42) for SAPS II, after customization. Kappa analysis showed only fair agreement between the two customized models with regard to the identity of the quality outliers: kappa = 0.44 (95% confidence interval, 0.24, 0.65). CONCLUSIONS: Customization of SAPS II and MPM II0 to the Project IMPACT database resulted in well-calibrated models. Despite this, the models exhibited only a moderate level of agreement in which hospitals were designated as quality outliers. Seventeen of the 54 hospitals were categorized differently depending on which of the two scoring systems was used. Therefore, the rating of quality of care appears, in part, to be a function of the prediction model used.

APACHE↗