PubMed Health⌕ Search

Biomedical subjects

Nancy A Obuchowski

Publications and source records attributed to Nancy A Obuchowski.

At least 19 recordsLinked to original sources

An ROC-type measure of diagnostic accuracy when the gold standard is continuous-scale.

ROC curves and summary measures of accuracy derived from them, such as the area under the ROC curve, have become the standard for describing and comparing the accuracy of diagnostic tests. Methods for estimating ROC curves rely on the existence of a gold standard which dichotomizes patients into disease present or absent. There are, however, many examples of diagnostic tests whose gold standards are not binary-scale, but rather continuous-scale. Unnatural dichotomization of these gold standards leads to bias and inconsistency in estimates of diagnostic accuracy. In this paper, we propose a non-parametric estimator of diagnostic test accuracy which does not require dichotomization of the gold standard. This estimator has an interpretation analogous to the area under the ROC curve. We propose a confidence interval for test accuracy and a statistical test for comparing accuracies of tests from paired designs. We compare the performance (i.e. CI coverage, type I error rate, power) of the proposed methods with several alternatives. An example is presented where the accuracies of two quick blood tests for measuring serum iron concentrations are estimated and compared.

Adolescent↗

Colonic abnormalities on CT in adult hospitalized patients with Clostridium difficile colitis: prevalence and significance of findings.

OBJECTIVE: The purpose of this study was to determine the prevalence of an abnormal colon on CT in adult inpatients with Clostridium difficile colitis, compare the clinical presentation of these patients, and determine whether CT findings predicted the need for surgical treatment. MATERIALS AND METHODS: Over a 21-month period, 152 of 572 inpatients with C. difficile colitis were identified and had CT scans performed within 2 weeks of the diagnosis. These were independently and retrospectively reviewed by two reviewers. Those with colonic wall thickness greater than 4 mm were considered positive (CT-positive patients) and were further reviewed for specific findings in the colon. All 152 patients with CT scans were also retrospectively reviewed using the hospital information system for certain clinical parameters, admitting diagnoses, and reasons for scanning. The following were compared using several statistical tests: clinical parameters in CT-positive and CT-negative patients and surgical and nonsurgical groups to determine if positive scans or surgical treatment could be clinically predicted; specific CT findings in CT-positive patients to see if an association was found with clinical parameters or surgical treatment; and admitting diagnoses and reasons for scanning in scanned and unscanned populations to see which patients were more likely to undergo CT. RESULTS: Seventy-six (50%) of 152 scanned hospitalized patients with C. difficile colitis were CT-positive. These patients most often had segmental involvement (50 [66%] of 76 patients), with the rectum (60 [82%] of 73 patients) and sigmoid colon (61 [82%] of 74 patients) most often affected. Positive scans were associated with increased WBC, abdominal pain, and diarrhea. Patients with signs and symptoms of infection or abdominal complaints were more likely to be scanned. No statistical correlation was found between specific CT findings and clinical parameters or clinical parameters and patients requiring surgery. There was no predictive value of specific CT findings for surgical treatment. CONCLUSION: Half of the patients scanned had an abnormal CT, with segmental colonic disease more common than diffuse. Positive scans were more likely in patients with leukocytosis, abdominal pain, and fever. Specific CT findings did not correlate with clinical parameters and could not predict surgical treatment.

Adult↗

A comparison of the Dorfman-Berbaum-Metz and Obuchowski-Rockette methods for receiver operating characteristic (ROC) data.

There are several different statistical methods for analysing multireader ROC studies, with the Dorfman-Berbaum-Metz (DBM) method being the most frequently used. Another method is the corrected F method proposed by Obuchowski and Rockette (OR). The DBM and OR procedures at first appear quite different: DBM is a three-way ANOVA analysis of pseudovalues while OR is a two-way ANOVA analysis of accuracy estimates with correlated errors. We show that the original DBM and OR F statistics for testing the null hypothesis of equal treatments have the same form and will typically have similar values; however, differences in the denominator degrees of freedom will result in differences in p-values even when the F statistics are identical. We show how the methods can be generalized to include variations in the accuracy measure, covariance method, and degrees of freedom. Identical results are obtained when the methods agree with respect to all three of these procedure parameters; hence for a particular choice of procedure parameters the choice of method appears to depend mainly on software preference and availability. The methods are compared using data from a factorial study with two modalities, five readers, and 114 patients.

Analysis of Variance↗

Estimating and comparing diagnostic tests' accuracy when the gold standard is not binary.

RATIONALE AND OBJECTIVES: Investigators often need to assess the accuracies of diagnostic tests when the gold standard is not binary-scale. The objective of this article is to describe nonparametric estimators of diagnostic test accuracy when the gold standard is continuous, ordinal, and nominal scale. MATERIALS AND METHODS: A nonparametric method of estimating and comparing the area under receiver operating characteristic (ROC) curves, proposed by DeLong et al, is extended to situations in which the gold standard is not binary. Two examples illustrate the methods. RESULTS: Measures of diagnostic test accuracy, their variance, and tests for comparing two diagnostic tests' accuracies in paired designs are presented for situations in which the gold standard is continuous, ordinal, and nominal scale. These summary measures of diagnostic test accuracy are analogous in form and interpretation to the area under the ROC curve. CONCLUSION: Dichotomizing the outcomes of a gold standard so that traditional ROC methods can be applied can lead to bias. The methods described here are useful for assessing and comparing summary test accuracy when the gold standard is not binary scale. They have limitations similar to other summary indices.

Diagnostic Imaging↗

Acute low back pain and radiculopathy: MR imaging findings and their prognostic role and effect on outcome.

PURPOSE: To prospectively determine in patients with acute low back pain (LBP) or radiculopathy, the magnetic resonance (MR) imaging findings, prognostic role of these findings, and effect of diagnostic information on outcome. MATERIALS AND METHODS: Institutional review board approval and informed consent were obtained. This study was HIPAA compliant. A total of 246 patients with acute-onset LBP or radiculopathy were randomized to either the early information arm of the study, with MR results provided within 48 hours, or the second arm of the study, where both patients and physicians were blinded to MR results, unless this information was critical to patient care. Patients underwent 6 weeks of conservative care. Roland function scoring, visual pain analog, Short Form 36 health status survey, self-efficacy scoring, and a fear avoidance questionnaire were completed at presentation; at 2-, 4-, 6-, and 8-week follow-up; and at 6-, 12-, and 24-month follow-up. A second MR imaging examination was performed at 6-week follow-up. Multivariate logistic regression analysis was used to determine which imaging and nonimaging variables can be used to predict improvement in Roland function and patient satisfaction. The chi(2) test and repeated-measures analysis of variance were used to compare outcome of blinded and unblinded patients. RESULTS: Herniation was identified in 60% (n = 147) of patients at the initial examination. The prevalence of herniations in patients with LBP (57%) (n = 85) and those with radiculopathy (65%) (n = 62) were similar (P = .217), although patients with radiculopathy were more likely to have stenosis and nerve root compression (P < .006). There was no relationship between herniation type, size, and behavior over time with outcome. An improvement of 50% or more in Roland function score at 6-week follow-up occurred 2.7 times as often among patients with a herniation at baseline (P = .003). Improvement at 6-week follow-up was similar in unblinded (60%) (n = 55) and blinded (67%) (n = 57) patients (P = .397). Self-efficacy, fear avoidance beliefs, and the Short Form 36 subscales were similar for blinded and unblinded patients. CONCLUSION: In typical patients with LBP or radiculopathy, MR imaging does not appear to have measurable value in terms of planning conservative care. Patient knowledge of imaging findings does not alter outcome and is associated with a lesser sense of well-being.

Acute Disease↗

ROC analysis.

Explore the source record for details and available documents.

Humans↗

Assessment of suspected breast cancer by MRI: a prospective clinical trial using a combined kinetic and morphologic analysis.

OBJECTIVE: The objective of our study was to assess the incremental value of contrast-enhanced MRI in the diagnosis and treatment planning using both a three-time point kinetic and morphologic analysis in addition to mammography and sonography in patients thought to have early-stage breast cancer. SUBJECTS AND METHODS: Contrast-enhanced bilateral breast MRI was performed prospectively on 65 patients with highly suspicious imaging findings (BI-RADS category 4 or 5). All enrolled patients were believed to be candidates for breast conservation on the basis of clinical examination, mammography, and sonography. The primary index lesion's characteristics, size, and extent were assessed. Also, additional lesions detected by MRI that could represent potential malignancies in both the ipsilateral and contralateral breast were evaluated. Morphologic assessment and kinetic analysis were performed on each lesion using dedicated postprocessing and display software. The patients were reevaluated as to whether they were still candidates for breast-conservation therapy after the MRI examination and subsequent biopsies. RESULTS: There were 46 patients (71%) whose primary breast lesion (detected by mammography, sonography, or both) was found to be malignant (39 invasive breast cancers, five intraductal cancers, and two lymphomas). For the primary index lesions, the sensitivity for MRI was 100% (44/44) for predicting a breast malignancy and the specificity was 73.7% (14/19) for predicting benign lesions. MRI detected an additional 37 lesions, of which 23 were cancerous, beyond those suspected on mammography or sonography. One or more additional ipsilateral breast cancers were detected in 32% (14/44) of breast cancer patients and contralateral breast cancers in 9% (4/44) of the breast cancer patients. MRI also resulted in an incremental recommendation of mastectomy in 18% (8/44) of the pathologically confirmed breast cancer patients. MRI resulted in additional biopsy of only 14 benign lesions, six of which were shown to be atypical ductal hyperplasia. CONCLUSION: When added to the standard evaluation of clinical examination, mammography, and sonography in patients thought to have early-stage breast cancer, contrast-enhanced MRI using both a kinetic and morphologic analysis will often result in changes in recommended patient management and better treatment planning and will result in no significant increase in biopsies of benign lesions. In addition, there is a significant detection rate of occult contralateral breast cancers.

Adult↗

ROC curves in clinical chemistry: uses, misuses, and possible solutions.

BACKGROUND: ROC curves have become the standard for describing and comparing the accuracy of diagnostic tests. Not surprisingly, ROC curves are used often by clinical chemists. Our aims were to observe how the accuracy of clinical laboratory diagnostic tests is assessed, compared, and reported in the literature; to identify common problems with the use of ROC curves; and to offer some possible solutions. METHODS: We reviewed every original work using ROC curves and published in Clinical Chemistry in 2001 or 2002. For each article we recorded phase of the research, prospective or retrospective design, sample size, presence/absence of confidence intervals (CIs), nature of the statistical analysis, and major analysis problems. RESULTS: Of 58 articles, 31% were phase I (exploratory), 50% were phase II (challenge), and 19% were phase III (advanced) studies. The studies increased in sample size from phase I to III and showed a progression in the use of prospective designs. Most phase I studies were powered to assess diagnostic tests with ROC areas >/=0.70. Thirty-eight percent of studies failed to include CIs for diagnostic test accuracy or the CIs were constructed inappropriately. Thirty-three percent of studies provided insufficient analysis for comparing diagnostic tests. Other problems included dichotomization of the gold standard scale and inappropriate analysis of the equivalence of two diagnostic tests. CONCLUSION: We identify available software and make some suggestions for sample size determination, testing for equivalence in diagnostic accuracy, and alternatives to a dichotomous classification of a continuous-scale gold standard. More methodologic research is needed in areas specific to clinical chemistry.

Chemistry, Clinical↗

Multireader, multicase receiver operating characteristic analysis: an empirical comparison of five methods.

RATIONALE AND OBJECTIVES: Several statistical methods have been developed for analyzing multireader, multicase (MRMC) receiver operating characteristic (ROC) studies. The objective of this article is to increase awareness of these methods and determine if their results are concordant for published datasets. MATERIALS AND METHODS: Data from three previously published studies were reanalyzed using five MRMC methods. For each method the 95% confidence intervals (CIs) for the mean of the readers' ROC areas for each diagnostic test, the P value for the comparison of the diagnostic tests' mean accuracies, and the 95% CIs for the mean difference in ROC areas of the diagnostic tests were reported. RESULTS: Important differences in P values and CIs were seen when using parametric versus nonparametric estimates of accuracy, and there were the expected differences for random-reader versus fixed-reader models. Controlling for these differences, the Dorfman-Berbaum-Metz (DBM), Obuchowski-Rockette, Beiden-Wagner-Campbell, and Song's multivariate Wilcoxon-Mann-Whitney (WMW) methods gave almost identical results for the fixed-reader model. For the random-reader model, the DBM, Obuchowski-Rockette, and Beiden-Wagner-Campbell methods yielded approximately the same inferences, but the CIs for the Beiden-Wagner-Campbell method tend to be broader. Ishwaran's hierarchical ROC sometimes yielded significance not found with other methods. Song's modification of DBM's jack-knifing algorithm sometimes led to different conclusions than the original DBM algorithm. CONCLUSION: In choosing and applying MRMC methods, it is important to recognize: (1) the distinction between random-reader and fixed-reader models, the uncertainties accounted for by each, and thus the level of generalizeability expected from each; (2) assumptions made by the various MRMC methods; and (3) limitations of a five- or six-reader study when the reader variability is great.

Analysis of Variance↗

Special Topics III: bias.

Researchers, manuscript reviewers, and journal readers should be aware of the many potential sources of bias in radiologic studies. This article is a review of the common biases that occur in selecting patient and reader samples, choosing and applying a reference standard, performing and interpreting diagnostic examinations, and analyzing diagnostic test results. Potential implications of various biases are discussed, and practical approaches to eliminating or minimizing them are presented.

Bias↗

Nonstress delayed-enhancement magnetic resonance imaging of the myocardium predicts improvement of function after revascularization for chronic ischemic heart disease with left ventricular dysfunction.

BACKGROUND: The extent of myocardial scarring of the left ventricle (LV) is important in patients with chronic ischemic heart disease (CIHD). With delayed-enhancement magnetic resonance imaging (DE-MRI), scarred myocardium (hyper-enhanced) is easily distinguishable from viable (dark) myocardium. This investigation assessed the use of DE-MRI for predicting functional improvement after coronary artery bypass grafting (CABG) in patients with CIHD and significant LV dysfunction. METHODS: The patient population (n = 29) with CIHD and LV dysfunction (ejection fraction 28% +/- 10%) underwent both DE-MRI, to delineate scarred regions before revascularization, and echocardiography (Echo), to assess segmental function before and after CABG (interval 188 +/- 57 days). Using a 16-segment model, LV myocardium was semiquantitatively analyzed for scarring based on DE-MRI and for improvements in resting function by pre- and post-CABG Echo. RESULTS: Before CABG, 82% of targeted myocardial segments had abnormal contraction; 78% showed scarring, including 38% with greater than mild amounts (25%-100%). Normal contraction was found in 18% of segments before revascularization; scarred areas were identified in 42%, 84% of which had, at most, minimal amounts (0%-24%). Of segments with pre-CABG dysfunction, 82% with no evidence of scar recovered, compared to only 18% with > or =50% scarring. Amount of hyper-enhancement was a very good indicator of improvement of function, especially at the > or =50%/segment threshold; overall accuracy was 0.74 (95% CI 0.66-0.82, P <.001). CONCLUSIONS: In patients with CIHD and significant LV dysfunction, DE-MRI can predict likelihood of functional improvement after revascularization.

Adult↗

Calcium scoring: criteria for evaluating its effectiveness.

Engineering advances in CT have produced multi-slice instruments that can scan large areas of the body in short periods of time, and such instruments now permit high resolution examination of entire anatomic regions (eg, the chest) in a single breath hold. Alternatively, these instruments can quickly scan small areas (such as the heart) with very high resolution in a very short period of time (eg, diastole). Using such CT scanners, there is no question that coronary artery calcium can be detected in small quantities and scored accurately. However, coronary calcium screening, like all screening procedures, poses a significant dilemma: early detection in a few is almost always accompanied by negative consequences for others (eg, false positives causing anxiety and unnecessary work-up, and false negatives causing delayed treatment and false reassurance). How do we balance the benefits to a few against the negative effects to others? That is the subject of this paper. A starting point for resolving the screening dilemma is to count the number of patients needed to be screened to benefit one patient (the NNS), and conversely, to determine the number of patients screened before harming one patient (NSH). Another approach is to apply published criteria suggested for the evaluation of a screening program targeted at early disease detection. In this review article, we propose 10 criteria for evaluating the effectiveness of a screening test designed to detect a risk factor for disease (ie, calcium scoring as a risk factor for coronary artery disease). We discuss how these criteria can be used to estimate NNS and NSH. Although this work focuses on coronary calcification screening, reference is made as well to other areas, such as lung and colon cancer screening.

Calcinosis↗