An official ATS statement: grading the quality of evidence and strength of recommendations in ATS guidelines and recommendations.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to Roman Jaeschke.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
BACKGROUND: Systems that are used by different organisations to grade the quality of evidence and the strength of recommendations vary. They have different strengths and weaknesses. The GRADE Working Group has developed an approach that addresses key shortcomings in these systems. The aim of this study was to pilot test and further develop the GRADE approach to grading evidence and recommendations. METHODS: A GRADE evidence profile consists of two tables: a quality assessment and a summary of findings. Twelve evidence profiles were used in this pilot study. Each evidence profile was made based on information available in a systematic review. Seventeen people were given instructions and independently graded the level of evidence and strength of recommendation for each of the 12 evidence profiles. For each example judgements were collected, summarised and discussed in the group with the aim of improving the proposed grading system. Kappas were calculated as a measure of chance-corrected agreement for the quality of evidence for each outcome for each of the twelve evidence profiles. The seventeen judges were also asked about the ease of understanding and the sensibility of the approach. All of the judgements were recorded and disagreements discussed. RESULTS: There was a varied amount of agreement on the quality of evidence for the outcomes relating to each of the twelve questions (kappa coefficients for agreement beyond chance ranged from 0 to 0.82). However, there was fair agreement about the relative importance of each outcome. There was poor agreement about the balance of benefits and harms and recommendations. Most of the disagreements were easily resolved through discussion. In general we found the GRADE approach to be clear, understandable and sensible. Some modifications were made in the approach and it was agreed that more information was needed in the evidence profiles. CONCLUSION: Judgements about evidence and recommendations are complex. Some subjectivity, especially regarding recommendations, is unavoidable. We believe our system for guiding these complex judgements appropriately balances the need for simplicity with the need for full and transparent consideration of all important issues.
Explore the source record for details and available documents.
The chronic respiratory questionnaire, available as an interviewer and a self-administered instrument, includes 20 items across four domains: dyspnea (5 items), fatigue (4 items), emotional function (7 items), and mastery (4 items). When completing this instrument, patients rate their experience on a 7-point scale ranging from 1 (maximum impairment) to 7 (no impairment). The Chronic Respiratory Questionnaire has demonstrated excellent measurement properties for both discriminative and evaluative purposes and served as a model in numerous methodological studies in chronic airflow limitation and patients with chronic obstructive pulmonary disease. We performed a systematic review of the literature on the chronic respiratory questionnaire to summarize the key qualities of the chronic respiratory questionnaire and to appraise the work regarding the minimal important difference of the chronic respiratory questionnaire. This paper includes a revision of our initial definition of the minimal important difference and a methodological framework for using anchor based approaches to establish the minimal important difference pioneered by Jaeschke and colleagues. Other approaches to evaluate the minimal important difference include distribution-based methods and panel-based methods. Investigators have used all of these approaches to establish the minimal important difference for the chronic respiratory questionnaire and the results are in general agreement with the minimal important difference of 0.5 for the mean domain scores of the chronic respiratory questionnaire. As a result of this literature review and discussion at the workshop, we established several research objectives. These objectives include the exploration of presentation of quality of life information and prospective anchor-based approaches.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
OBJECTIVE: To examine the incidence and predictors of clinician discomfort with life support plans for ICU patients. DESIGN AND SETTING: Prospective cohort in 13 medical-surgical ICUs in four countries. PATIENTS: 657 mechanically ventilated adults expected to stay in ICU at least 72 h. MEASUREMENTS AND RESULTS: Daily we documented the life support plan for mechanical ventilation, inotropes and dialysis, and clinician comfort with these plans. If uncomfortable, clinicians stated whether the plan was too technologically intense (the provision of too many life support modalities or the provision of any modality for too long) or not intense enough, and why. At least one clinician was uncomfortable at least once for 283 (43.1%) patients, primarily because plans were too technologically intense rather than not intense enough (93.9% vs. 6.1%). Predictors of discomfort because plans were too intense were: patient age, medical admission, APACHE II score, poor prior functional status, organ dysfunction, dialysis in ICU, plan to withhold dialysis, plan to withhold mechanical ventilation, first week in the ICU, clinician, and city. CONCLUSIONS: Clinician discomfort with life support perceived as too technologically intense is common, experienced mostly by nurses, variable across centers, and is more likely for older, severely ill medical patients, those with acute renal failure, and patients lacking plans to forgo reintubation and ventilation. Acknowledging the sources of discomfort could improve communication and decision making.
Users of clinical practice guidelines and other recommendations need to know how much confidence they can place in the recommendations. Systematic and explicit methods of making judgments can reduce errors and improve communication. We have developed a system for grading the quality of evidence and the strength of recommendations that can be applied across a wide range of interventions and contexts. In this article we present a summary of our approach from the perspective of a guideline user. Judgments about the strength of a recommendation require consideration of the balance between benefits and harms, the quality of the evidence, translation of the evidence into specific circumstances, and the certainty of the baseline risk. It is also important to consider costs (resource utilisation) before making a recommendation. Inconsistencies among systems for grading the quality of evidence and the strength of recommendations reduce their potential to facilitate critical appraisal and improve communication of these judgments. Our system for guiding these complex judgments balances the need for simplicity with the need for full and transparent consideration of all important issues.
BACKGROUND: Intensive insulin therapy has recently been shown to decrease morbidity and mortality in the critically ill population in a large randomized clinical trial. OBJECTIVE: To determine the beliefs and attitudes of ICU clinicians about glycemic control. DESIGN: Self-administered survey. PARTICIPANTS: ICU nurses and physicians in five university-affiliated multidisciplinary ICUs. RESULTS: A total of 317 questionnaires were returned from 233 ICU nurses and 84 physicians. The reported clinically important threshold for hypoglycemia was 4 mmol/l (median, IQR 3-4 mmol/l). In non-diabetic patients, the clinically important threshold for hyperglycemia was 10 mmol/l (IQR 9-12 mmol/l); however, nurses had a significantly higher threshold than physicians (difference of 0.52 mmol/l (95% CI 0.09-0.94 mmol/l, P=0.018). In diabetic patients, the clinically important threshold for hyperglycemia was also 10 mmol/l (IQR 10-12 mmol/l), and again nurses had a significantly higher threshold than physicians (0.81 mmol/l, 95% CI 0.29-1.32 mmol/l, P=0.0023). Avoidance of hyperglycemia was judged most important for diabetic patients (87.7%, 95% CI 84.1-91.3%), patients with acute brain injury (84.5%, 95% CI 80.5-88.5%), patients with a recent seizure (74.4%, 95% CI 69.6-79.3%), patients with advanced liver disease (64.0%, 95% CI 58.7-69.3%), and for patients with acute myocardial infarction (64.0%, 95% CI 58.7-69.3%). Physicians expressed more concern than nurses about avoiding hyperglycemia in patients with acute myocardial infarction ( P=0.0004). ICU clinicians raised concerns about the accuracy of glucometer measurements in critically ill patients (46.1%, 95% CI 40.5-51.6%). CONCLUSIONS: Attention to these beliefs and attitudes could enhance the success of future clinical, educational and research efforts to modify clinician behavior and achieve better glycemic control in the ICU setting.
BACKGROUND: This review summarizes the current status of randomized trials of digitalis in treating patients with congestive heart failure who are in sinus rhythm. Methods and results Randomized double-blind placebo-controlled trials of 20 or more adult patients followed for 7 weeks or more were selected. We identified 13 trials that met the inclusion criteria, comprising a total of 7896 patients. Of this number, 7755 patients contributed to information on mortality, 7262 to information on hospitalization for worsening heart failure, and 1096 to information on clinical status. Patients treated with digitalis compared with placebo had an odds ratio and confidence intervals for mortality of 0.98 (0.89, 1.09), for hospitalization of 0.68 (0.61, 0.75), and for a lesser degree of deterioration in clinical status of 0.31 (0.21, 0.43). CONCLUSIONS: The literature indicates that the drug has no effect on long-term mortality, but reduces the incidence of hospitalization, and has a positive effect on the clinical status of symptomatic patients. The drug has beneficial effects in patients who remain symptomatic despite being appropriately treated with diuretics and angiotensin-converting enzyme inhibitors. However the effects of coadministration with beta-blockers, spironolactone, and valsartan remain uncertain.
This article about the grades of recommendation for antithrombotic and thrombolytic therapy is part of the Seventh American College of Chest Physicians Conference on Antithrombotic and Thrombolytic Therapy: Evidence-Based Guidelines. Clinicians need to know whether a recommendation is strong or weak, and about the methodological quality of the evidence underlying that recommendation. We determine the strength of a recommendation by considering the trade-off between the benefits of a treatment, on the one hand, and the risks, burdens, and costs on the other. Here, as elsewhere, we assume that a recommended treatment will increase costs (we recognize this is not always the case, but for simplicity we will continue to make this assumption). If the benefits outweigh the risks, burdens, and costs, we recommend that clinicians offer a treatment to typical patients. The uncertainty associated with the trade-off between the benefits and the risks, burdens, and costs will determine the strength of the recommendations. If we are very certain that the benefits do, or do not, outweigh the risks, burdens, and costs, we make a strong recommendation (in our formulation, Grade 1). If we are less certain of the magnitude of the benefits and the risks, burdens, and costs, and thus of their relative impact, we make a weaker Grade 2 recommendation. We grade the methodological quality of a recommendation according to the following criteria. Randomized clinical trials (RCTs) with consistent results provide evidence with a low likelihood of bias, which we classify as Grade A recommendations. RCTs with inconsistent results, or with major methodological weaknesses, warrant Grade B recommendations. Grade C recommendations come from observational studies or from a generalization from one group of patients included in randomized trials to a different, but somewhat similar, group of patients who did not participate in those trials. When we find the generalization from RCTs to be secure, or the data from observational studies overwhelmingly compelling, we choose a Grade C+. When that is not the case, we designate methodological quality as Grade C.
BACKGROUND: One of the most challenging practical and daily problems in intensive care medicine is the interpretation of the results from diagnostic tests. In neonatology and pediatric intensive care the early diagnosis of potentially life-threatening infections is a particularly important issue. FOCUS: A plethora of tests have been suggested to improve diagnostic decision making in the clinical setting of infection which is a clinical example used in this article. Several criteria that are critical to evidence-based appraisal of published data are often not adhered to during the study or in reporting. To enhance the critical appraisal on articles on diagnostic tests we discuss various measures of test accuracy: sensitivity, specificity, receiver operating characteristic curves, positive and negative predictive values, likelihood ratios, pretest probability, posttest probability, and diagnostic odds ratio. CONCLUSIONS: We suggest the following minimal requirements for reporting on the diagnostic accuracy of tests: a plot of the raw data, multilevel likelihood ratios, the area under the receiver operating characteristic curve, and the cutoff yielding the highest discriminative ability. For critical appraisal it is mandatory to report confidence intervals for each of these measures. Moreover, to allow comparison to the readers' patient population authors should provide data on study population characteristics, in particular on the spectrum of diseases and illness severity.
BACKGROUND/OBJECTIVES: The chronic respiratory questionnaire (CRQ), the St. Georges Respiratory Questionnaire (SGRQ), and the feeling thermometer (FT) evaluate change in health-related quality of life (HRQL) in patients with chronic airflow limitation (CAL). Although the interpretability, and in particular the minimal important difference (MID) in score changes, is well established for the CRQ, this is not the case for the SGRQ and FT. The objective of our study is to explore the interpretation of the SGRQ and FT. METHODS: We analyzed data from 84 patients who completed the CRQ, SGRQ, and FT before beginning pulmonary rehabilitation and 3 months later. We calculated correlations between the four CRQ domains (dyspnea, fatigue, emotional function, and mastery) and the three SGRQ domains (symptoms, activities, and impact), the SGRQ total score, and the FT. When Pearson's correlations were >/=0.5, we constructed regression equations and used the slope to calculate the change in SGRQ and FT score that corresponded to a change in CRQ score of 0.5 (the MID). Having established MID for SGRQ we than used a similar approach to examine the relation between the SGRQ and FT results. RESULTS: Comparison with the CRQ dyspnea domain suggested the MID in SGRQ total score is approximately 3.05 with a 95% confidence interval (95% CI) ranging from 0.39 to 5.71 and a change of 5.67 (95% CI 3.43-7.92) represents a moderate change (1.0 on the CRQ dyspnea domain). The MID for the FT based on the CRQ fatigue domain was 6.1 (95% CI 1.87-10.28). The FT MID based on the SGRQ activities domain, impacts domain, and total score were, respectively, 7.4 (95% CI 3.44-11.35), 5.6 (95% CI 1.6-9.64), and 5.9 (95% CI 1.97-9.78). CONCLUSIONS: An MID for the SGRQ approximates the previously suggested estimate of 4 on a scale of 0 to 100. The MID for the FT in patients with CAL is approximately 5 to 8 units on the 0 to 100 scale. These MID estimates should facilitate interpretation of clinical trials in which outcome measures include the SGRQ or FT.
We compared the diagnostic efficacy of fractionated plasma metanephrine measurements to measurements of 24-h urinary total metanephrines and catecholamines in outpatients tested for pheochromocytoma at Mayo Clinic Rochester from January 1, 1999, until November 27, 2000. Catecholaminesecreting tumors were histologically proven. The sensitivity of fractionated plasma metanephrines was 97% (30 of 31 patients), compared with a sensitivity of 90% (28 of 31) for urinary total metanephrines and catecholamines (P = 0.63). The specificity of fractionated plasma metanephrines was 85% (221 of 261), compared with 98% (257 of 261; P < 0.001) for urinary measurements. The likelihood ratios for positive tests were 6.3 (95% confidence interval, 4.7 to 8.5) for fractionated plasma metanephrines and 58.9 (95% confidence interval, 22.1 to 156.9) for urinary total metanephrines and catecholamines. An adrenal pheochromocytoma was missed by urinary testing in two patients with familial syndromes and one asymptomatic patient with an incidentally discovered adrenal mass. An extra-adrenal paraganglioma was missed by plasma testing in one patient. In conclusion, measurements of 24-h urinary total metanephrines and catecholamines yield fewer false-positive results, an attribute preferred for testing low-risk patients, but fractionated plasma metanephrine measurements may be preferred in high-risk patients with familial endocrine syndromes.
BACKGROUND AND OBJECTIVES: The chronic respiratory questionnaire (CRQ), a widely used measure of health-related quality of life (HRQL) in patients with chronic airflow limitation, includes an individualized dyspnea domain (patients identify five important activities, and report the degree of dyspnea on a 7-point scale). Because the individualized domain is unwieldy in multicenter clinical trials, we developed a standardized version and tested its discriminative and evaluative properties. METHODS: We enrolled 51 patients who completed the standardized and individualized CRQ before starting a respiratory rehabilitation program, and again 3 months later. We calculated both cross-sectional and longitudinal correlations between the two versions and a number of other HRQL instruments, and tested the relative ability of the individualized and standardized versions of the CRQ to detect improvement with rehabilitation. RESULTS: The results of the individualized questions suggested greater dysfunction (lower scores) than did the standardized questions both at baseline (3.18 vs 3.92, p < 0.001) and follow-up (4.62 vs 4.84, p = 0.051). The standardized dyspnea domain showed superior discriminative validity. While both techniques detected important, statistically significant improvement with rehabilitation (individualized domain mean change, 1.44; 95% confidence interval [CI], 1.11 to 1.77 [p < 0.001]; standardized domain mean change, 0.92; 95% CI, 0.61 to 1.24 [p < 0.01]), the difference in effect was substantial and statistically significant (mean difference, 0.52; 95% CI, 0.22 to 0.82; p = 0.001). The two versions showed comparable longitudinal validity. CONCLUSIONS: A standardized version of the CRQ dyspnea domain improves the cross-sectional validity, maintains longitudinal validity, but reduces the responsiveness. By increasing sample size, investigators can use the more efficient standardized version of the CRQ without compromising validity.
Explore the source record for details and available documents.
In this article we present factors that determine whether one can trust the results of studies that describe the properties of diagnostic tests. We review the ways in which the results of such studies could be presented to the readers and consider factors that determine the value of such reports in one's own clinical practice.