PubMed Health⌕ Search

Biomedical subjects

Michael R Jiroutek

Publications and source records attributed to Michael R Jiroutek.

6 recordsLinked to original sources

Comparison of calcification specificity in digital mammography using soft-copy display versus screen-film mammography.

OBJECTIVE: The purpose of this study was to compare specificity in the interpretation of calcifications in soft-copy reviewing of digital mammograms versus hard-copy reviewing of screen-film mammograms. MATERIALS AND METHODS: A total of 130 consecutive cases with calcifications (44 malignant and 86 benign) that had been evaluated with needle or surgical biopsy were collected. Both screen-film mammography and soft-copy digital mammography were obtained in the same patients under existing research protocols using Fischer Imaging's SenoScan (n = 71), Lorad's digital mammography system (n = 35), and GE Healthcare's Senographe 2000D (n = 24). Eight trained radiologists scored all lesions--cropped or masked to display just the region of interest--both on screen-film and soft-copy digital mammography with a month between reviews to reduce the effects of learning and memory. A 5-point malignancy scale was used, with 1 as definitely not, 2 as probably not, 3 as possibly, 4 as probably, and 5 as definitely. Reviewers were randomly assigned condition order, and images within each condition were randomly ordered. Repeated measures analysis of variance was used to test for differences between conditions in specificity computed via nonparametric receiver operating characteristic (ROC) study separately for each reviewer and condition. RESULTS: Across all reviewers, the mean specificity for 1 or 2 versus 3, 4, or 5 was 0.803 for screen-film mammography (range, 0.413-0.938; SD +/- 0.166) and 0.833 for soft-copy image (range, 0.375-0.951; SD +/- 0.187). Although not statistically significant (Student's t test p values from 0.19 to 0.99 across all cut points), numeric values of specificity were consistently higher for soft-copy versus screen-film mammography. No statistical significance in specificity was seen using all possible cut points in the 5-point scale, although the primary analysis used the cutpoint for differentiation between benign and malignant cases as 1 or 2 versus 3, 4, or 5. CONCLUSION: No statistically significant difference was shown in specificity achievable using soft-copy digital versus screen-film mammography in this study.

Biopsy↗

A new method for choosing sample size for confidence interval-based inferences.

Scientists often need to test hypotheses and construct corresponding confidence intervals. In designing a study to test a particular null hypothesis, traditional methods lead to a sample size large enough to provide sufficient statistical power. In contrast, traditional methods based on constructing a confidence interval lead to a sample size likely to control the width of the interval. With either approach, a sample size so large as to waste resources or introduce ethical concerns is undesirable. This work was motivated by the concern that existing sample size methods often make it difficult for scientists to achieve their actual goals. We focus on situations which involve a fixed, unknown scalar parameter representing the true state of nature. The width of the confidence interval is defined as the difference between the (random) upper and lower bounds. An event width is said to occur if the observed confidence interval width is less than a fixed constant chosen a priori. An event validity is said to occur if the parameter of interest is contained between the observed upper and lower confidence interval bounds. An event rejection is said to occur if the confidence interval excludes the null value of the parameter. In our opinion, scientists often implicitly seek to have all three occur: width, validity, and rejection. New results illustrate that neglecting rejection or width (and less so validity) often provides a sample size with a low probability of the simultaneous occurrence of all three events. We recommend considering all three events simultaneously when choosing a criterion for determining a sample size. We provide new theoretical results for any scalar (mean) parameter in a general linear model with Gaussian errors and fixed predictors. Convenient computational forms are included, as well as numerical examples to illustrate our methods.

Biometry↗

How many cases need to be reviewed to compare performance in surgical pathology?

Recent studies have shown increased interest in measuring error rates in surgical pathology. We sought to determine how many surgical pathology cases need to be reviewed to show a significant difference from published error rates for review of routine or biopsy cases. Results of 4 series with this type of diagnostic material involving a total of 11,683 cases were reviewed to determine the range of published false-negative, false-positive, typing error, threshold error, and clinically significant error rates. Error rates ranged from 0.00% to 2.36%; clinically significant error rates ranged from 0.34% to 1.19%. Assuming a power of 0.80 and a 1-sided alpha of 0.05, the number of cases needed to be reviewed to show that a laboratory with either twice or one half the published error rate was significantly different from the range of published error rates varied from 3.30 to 50, 158. For clinically significant errors, the number of cases varied from 665 to 5,886. Because the published error rates are low, a relatively large number of cases need to be reviewed and a relatively great difference in error rate needs to exist to show a significant difference in performance in surgical pathology.

Diagnostic Errors↗

Thresholds for human detection of patient setup errors in digitally reconstructed portal images of prostate fields.

PURPOSE: Computer-assisted methods to analyze electronic portal images for the presence of treatment setup errors should be studied in controlled experiments before use in the clinical setting. Validation experiments using images that contain known errors usually report the smallest errors that can be detected by the image analysis algorithm. This paper offers human error-detection thresholds as one benchmark for evaluating the smallest errors detected by algorithms. Unfortunately, reliable data are lacking describing human performance. The most rigorous benchmarks for human performance are obtained under conditions that favor error detection. To establish such benchmarks, controlled observer studies were carried out to determine the thresholds of detectability for in-plane and out-of-plane translation and rotation setup errors introduced into digitally reconstructed portal radiographs (DRPRs) of prostate fields. METHODS AND MATERIALS: Seventeen observers comprising radiation oncologists, radiation oncology residents, physicists, and therapy students participated in a two-alternative forced choice experiment involving 378 DRPRs computed using the National Library of Medicine Visible Human data sets. An observer viewed three images at a time displayed on adjacent computer monitors. Each image triplet included a reference digitally reconstructed radiograph displayed on the central monitor and two DRPRs displayed on the flanking monitors. One DRPR was error free. The other DRPR contained a known in-plane or out-of-plane error in the placement of the treatment field over a target region in the pelvis. The range for each type of error was determined from pilot observer studies based on a Probit model for error detection. The smallest errors approached the limit of human visual capability. The observer was told what kind of error was introduced, and was asked to choose the DRPR that contained the error. Observer decisions were recorded and analyzed using repeated-measures analysis of variance. RESULTS: The thresholds of detectability averaged over all observers were approximately 2.5 mm for in-plane translations, 1.6 degrees for in-plane rotations, 1 degrees for out-of-plane rotations, and 8% change in magnification for out-of-plane translations along the central axis. When one inexperienced observer is excluded, the average threshold for change in magnification is 5%. Experienced observers tended to perform better, but differences between groups were not statistically significant. Thresholds were computed as averages over all observers. Because of the broad range of observer capabilities, some detection tasks were too difficult for some observers, leading to missing threshold values in our data analysis. The missing values were excluded from computation of the average thresholds reported above. The effect of the missing values is to bias the average values toward the best human performance. CONCLUSIONS: Under favorable conditions, humans can detect small errors in setup geometry. The thresholds for error detection reported in this study are believed to represent rigorous but reasonable benchmarks that can be incorporated into studies evaluating algorithms for computer-assisted detection of setup errors in electronic portal images.

Diagnostic Errors↗

Quantifying the value of in-house consultation in surgical pathology.

In-house consultation is a well-known method to improve diagnostic accuracy and agreement, but the technique has not been well studied. We reviewed the results of in-house consultation in a large private hospital practice setting for a 1-month period and determined its effect on diagnostic accuracy using the final sign-out as the "gold standard." During this 1-month period, 352 cases were reviewed as in-house consultations. Initial complete agreement was found in 315 (89.5%) cases. Using the initial diagnosis as the test case and the final sign-out as the gold standard, of the 37 discrepant cases, 4 (1.1%) were thought to represent false-negative results, (0.3%) a false-positive result, 3 (0.9%) differences in type, and 29 (8.2%) differences in diagnostic threshold. Disagreements in 10 cases were thought to be potentially clinically significant. Internal consultation was obtained on approximately 20% of all cases seen in the laboratory and disagreements were found in 2% of all cases. Internal consultation has a significant and measurable impact on the practice of surgical pathology.

False Negative Reactions↗

Blinded review as a method for quality improvement in surgical pathology.

CONTEXT: Several studies have shown that blinded review, because it is less biased and may improve vigilance, is an excellent method for detecting errors and improving performance in gynecologic cytology. The value of blinded review in surgical pathology is not known. OBJECTIVE: To determine the value of blinded review in surgical pathology. METHODS: Five hundred ninety-two biopsy cases were reviewed without knowledge of the original diagnosis or history, and the results were compared with those of the original diagnosis. RESULTS: Complete agreement was obtained in 567 (96%) of 592 cases. The technique of blinded review of biopsy material had a sensitivity of 98%, failing to identify a lesion in 7 cases; no cases of malignancy were missed. The specificity was 100%. Differences in diagnostic threshold were the most common source of disagreement. False-negative cases were identified by the technique and were clinically significant. Power studies show that the number of cases requiring review to identify significant errors are large, but potentially achievable by blinded review. CONCLUSION: Blinded review is a sensitive and effective method for identifying areas of disagreement, including false-negative cases, and for decreasing errors in surgical pathology biopsy material.

Diagnostic Errors↗