PubMed Health⌕ Search

Biomedical subjects

Kevin S Berbaum

Publications and source records attributed to Kevin S Berbaum.

17 recordsLinked to original sources

Breast MRI lesion classification: improved performance of human readers with a backpropagation neural network computer-aided diagnosis (CAD) system.

PURPOSE: To develop and test a computer-aided diagnosis (CAD) system to improve the performance of radiologists in classifying lesions on breast MRI (BMRI). MATERIALS AND METHODS: A CAD system was developed that uses a semiautomated segmentation method. After segmentation, 42 features based on lesion shape, texture, and enhancement kinetics were computed, and the 13 best features were selected and used as inputs to a backpropagation neural network (BNN). The BNN was trained and tested using the leave-one-out method on 80 BMRI lesions (37 benign, 43 malignant). Lesion histopathology was used as the reference standard. Five human readers classified the 80 lesions first without and then with CAD assistance. The performance of the computer classifier and the human readers was assessed using receiver operating characteristic curves; the performance of the human readers was also evaluated using multireader multicase (MRMC) analysis. RESULTS: The performance of the human readers significantly improved when aided by the CAD system (P < 0.05). MRMC analysis showed that human reader performance with and without CAD system assistance can be generalized to the population of cases (P < 0.001). CONCLUSION: A CAD system based on lesion morphology and enhancement kinetics can improve the performance of human readers in classifying lesions on breast MRI.

Breast Neoplasms↗

Peroneal tendon subluxation and dislocation: detection on volume-rendered images--initial experience.

Institutional review board approval was received and informed consent was not required for this Health Insurance Portability and Accountability Act-compliant study. The purpose of this study was to retrospectively assess the time efficiency of three-dimensional volume-rendered images obtained from multi-detector row computed tomographic data for the diagnosis of peroneal tendon subluxation or dislocation by using the consensus interpretation of multiplanar reformatted (MPR) images as the reference standard. The reference standard was provided by two musculoskeletal radiologists, and two less experienced readers evaluated 37 images in 32 patients (24 men, eight women; mean age, 41 years; age range, 18-75 years) with acute calcaneal fractures. An analysis of variance was used to compare interpretation time, and the Wilcoxon signed rank test was used to analyze diagnostic difficulty. The average time required for diagnosis was significantly shorter with volume-rendered images than with MPR images (reader 1: 42 vs 78 seconds, P<.001; reader 2: 50 vs 69 seconds, P<.01).

Adolescent↗

Interobserver agreement for Letournel acetabular fracture classification with multidetector CT: are standard Judet radiographs necessary?

PURPOSE: To retrospectively evaluate interobserver agreement for Letournel acetabular fracture classification with radiography alone and multidetector computed tomography (CT) alone and to retrospectively assess whether standard Judet views lead to a change in the classification. MATERIALS AND METHODS: Institutional review board approval was obtained; informed consent was not required for this HIPAA-compliant study, which included 101 imaging studies performed in 99 patients (78 male, 21 female; mean age, 43 years; age range, 15-86 years) with acetabular fractures. Two musculoskeletal radiologists independently classified the fractures with radiography alone and multidetector CT alone. Multiplanar reformatted and three-dimensional (3D) CT images were reviewed at a computer workstation. Readers were shown radiographs at the end of multidetector CT image reading to see if this would change the multidetector CT-based classification. kappa Values were calculated to assess interobserver agreement. For surgically treated patients, the McNemar test was used to compare the accuracy of readers' classifications. The reference standard was a combination of preoperative radiographic and multidetector CT image findings and intraoperative findings. RESULTS: Interobserver agreement was moderate (kappa = 0.42) with radiography and substantial (kappa = 0.70) with multidetector CT. Multidetector CT classification was changed in two cases (one case for each reader) after standard Judet views were added. In 73 surgically treated patients, agreement with the surgeons' classification was higher with multidetector CT than with radiography (P < .01 for one reader, P = .06 for the other reader). CONCLUSION: There is substantial interobserver agreement for Letournel acetabular fracture classification with multiplanar reformatted and 3D multidetector CT images. Standard Judet pelvic radiographs add little information for changing the multidetector CT classification.

Acetabulum↗

Adjunctive self-hypnotic relaxation for outpatient medical procedures: a prospective randomized trial with women undergoing large core breast biopsy.

Medical procedures in outpatient settings have limited options of managing pain and anxiety pharmacologically. We therefore assessed whether this can be achieved by adjunct self-hypnotic relaxation in a common and particularly anxiety provoking procedure. Two hundred and thirty-six women referred for large core needle breast biopsy to an urban tertiary university-affiliated medical center were prospectively randomized to receive standard care (n=76), structured empathic attention (n=82), or self-hypnotic relaxation (n=78) during their procedures. Patients' self-ratings at 1 min-intervals of pain and anxiety on 0-10 verbal analog scales with 0=no pain/anxiety at all, 10=worst pain/anxiety possible, were compared in an ordinal logistic regression model. Women's anxiety increased significantly in the standard group (logit slope=0.18, p<0.001), did not change in the empathy group (slope=-0.04, p=0.45), and decreased significantly in the hypnosis group (slope=-0.27, p<0.001). Pain increased significantly in all three groups (logit slopes: standard care=0.53, empathy=0.37, hypnosis=0.34; all p<0.001) though less steeply with hypnosis and empathy than standard care (p=0.024 and p=0.018, respectively). Room time and cost were not significantly different in an univariate ANOVA despite hypnosis and empathy requiring an additional professional: 46 min/161 dollars for standard care, 43 min/163 dollars for empathy, and 39 min/152 dollars for hypnosis. We conclude that, while both structured empathy and hypnosis decrease procedural pain and anxiety, hypnosis provides more powerful anxiety relief without undue cost and thus appears attractive for outpatient pain management.

Adult↗

A new software tool for removing, storing, and adding abnormalities to medical images for perception research studies.

RATIONALE AND OBJECTIVES: Image perception studies have been difficult to perform using clinical images because of the problems associated with obtaining proven abnormalities and appropriate normal controls. The objective of this research was to develop and evaluate interactive software that allows the seamless removal, archiving and insertion of abnormal areas from computed tomography (CT) lung image sets for use in image perception research. MATERIALS AND METHODS: The software tools for removing, archiving, and adding lesions are described in detail. The efficacy of the software to remove abnormal areas of lung CT studies was evaluated by having radiologists select the one altered image from a display of four. The software for adding lesions was evaluated by having radiologists classify displayed CT slices with lesions as real or artificial along with their confidence level. RESULTS: Observers could not reliably detect when images had been altered by the software. In the lesion-removal experiment, the observers correctly identified the altered display in only 15.8 +/- 2.8 of 56 sets. In the lesion-add experiment, the observers correctly identified the artificially placed lesions in 38.2 +/- 3.9 of 77 sets. The frequency distribution of the correct responses did not differ from that expected from chance selection. CONCLUSIONS: The results from both of these experiments demonstrate that radiologists could not distinguish between original and altered images. We conclude that this software can be used with volumetric CT lung images for creating normal control and target data sets for medical image perception research.

Humans↗

A comparison of the Dorfman-Berbaum-Metz and Obuchowski-Rockette methods for receiver operating characteristic (ROC) data.

There are several different statistical methods for analysing multireader ROC studies, with the Dorfman-Berbaum-Metz (DBM) method being the most frequently used. Another method is the corrected F method proposed by Obuchowski and Rockette (OR). The DBM and OR procedures at first appear quite different: DBM is a three-way ANOVA analysis of pseudovalues while OR is a two-way ANOVA analysis of accuracy estimates with correlated errors. We show that the original DBM and OR F statistics for testing the null hypothesis of equal treatments have the same form and will typically have similar values; however, differences in the denominator degrees of freedom will result in differences in p-values even when the F statistics are identical. We show how the methods can be generalized to include variations in the accuracy measure, covariance method, and degrees of freedom. Identical results are obtained when the methods agree with respect to all three of these procedure parameters; hence for a particular choice of procedure parameters the choice of method appears to depend mainly on software preference and availability. The methods are compared using data from a factorial study with two modalities, five readers, and 114 patients.

Analysis of Variance↗

Can order of report prevent satisfaction of search in abdominal contrast studies?

RATIONALE AND OBJECTIVE: A previous receiver operating characteristic (ROC) study showed a systematic shift in decision thresholds for detecting plain film abnormalities on contrast examinations rather than plain radiographs. A previous eye-position study showed that this shift was based on a relative visual neglect of plain film regions on the contrast studies. We now determine whether an intervention that changes visual search can reduce this search-based satisfaction of search effect in contrast studies of the abdomen. MATERIALS AND METHODS: The authors measured detection of 23 plain film abnormalities in 44 patients who had plain film and contrast examinations. In 2 experiments, each plain-film and contrast study was examined independently in different sessions with observers providing a confidence rating of abnormality for each interpretation. There were 13 observers in the first experiment and 10 in the second experiment. The intervention required that for the contrast studies, observers first report abnormalities in the noncontrast region of the radiograph before reporting contrast findings. ROC curve areas for each observer in each treatment condition were estimated by using a proper ROC model. The analysis focused on changes in decision thresholds among the treatment conditions. RESULTS: The SOS effect on decision thresholds in abdominal contrast studies was replicated. Although reduced, the shift in decision thresholds in detecting plain film abnormalities on contrast examinations remained when observers were required to report those abnormalities before contrast findings. CONCLUSION: Reporting plain film abnormalities before reporting abnormalities demonstrated by contrast reduced somewhat the satisfaction of search effect on decision thresholds produced by a visual neglect of noncontrast regions on contrast examinations. This suggests that interventions that direct visual search do not offer protection against satisfaction of search effects that are based on faulty visual search.

Contrast Media↗

Monte Carlo validation of the Dorfman-Berbaum-Metz method using normalized pseudovalues and less data-based model simplification.

RATIONALE AND OBJECTIVES: Two problems of the Dorfman-Berbaum-Metz (DBM) method for analyzing multireader receiver operating characteristic (ROC) studies are that it tends to be conservative and that it can produce AUC estimates outside the parameter space--ie, greater than one or less than zero. Recently it has been shown that the problem of AUC (or other accuracy) estimates outside the parameter space can be eliminated by using normalized pseudovalues, and it has been suggested that less data-based model simplification be used. Our purpose is to empirically investigate if these two modifications--normalized pseudovalues and less data-based model simplification--result in improved performance. MATERIALS AND METHODS: We examine the performance of the DBM procedure using the two proposed modifications for discrete and continuous ratings in a null simulation study comparing modalities with respect to the ROC area. The simulation study includes 144 different combinations of reader and case sample sizes, normal/abnormal case sample ratios, and variance components. The ROC area is estimated using parametric and nonparametric estimation. RESULTS: The DBM procedure with both modifications performs better than either the original DBM procedure or the DBM procedure with only one of the modifications. For parametric estimation with discrete rating data, use of both modifications resulted in the mean type I error (0.043) closest to the nominal .05 level and the smallest range (0.050) and standard deviation (0.0108) across the 144 type I error rates. CONCLUSIONS: We recommend that normalized pseudovalues and less data-based model simplification be used with the DBM procedure.

Algorithms↗

Free-response receiver operating characteristic evaluation of lossy JPEG2000 and object-based set partitioning in hierarchical trees compression of digitized mammograms.

PURPOSE: To assess the effects of two irreversible wavelet-based compression algorithms--Joint Photographic Experts Group (JPEG) 2000 and object-based set partitioning in hierarchical trees (SPIHT)--on the detection of clusters of microcalcifications and masses on digitized mammograms. MATERIALS AND METHODS: The use of the images in this retrospective image-collection study was approved by the institutional review board, and patient informed consent was not required. One hundred twelve mammographic images (28 with one or two clusters of microcalcifications, 19 with one mass, 17 with both abnormal findings, and 48 with normal findings) obtained in 60 women who ranged in age from 25 to 79 years were digitized and compressed at 40:1 and 80:1 by using the JPEG2000 and object-based SPIHT methods. Five experienced radiologists were asked to locate and rate clusters of microcalcifications and masses on the original and compressed images in a free-response receiver operating characteristic (FROC) data acquisition paradigm. Observer performance was evaluated with the jackknife FROC method. RESULTS: The mean FROC figures of merit for detecting clusters of microcalcifications, masses, and both radiographic findings on uncompressed images were 0.80, 0.81, and 0.72, respectively. With object-based SPIHT 80:1 compression, the corresponding values were larger than the values for uncompressed images by 0.005, 0.009, and -0.005, respectively. The 95% confidence interval for the differences in figures of merit between compressed and uncompressed images was -0.039, 0.033 for the microcalcification finding; -0.055, 0.034 for the mass finding; and -0.039, 0.030 for both findings. Because each of these confidence intervals includes zero, no significant difference in detection accuracy between uncompressed and object-based SPIHT 80:1 compression was observed at a P value of 5%. The F test of the null hypothesis that all of the modes (uncompressed and four compressed modes) were equivalent yielded the following results: F = 0.255, P = .903 for the microcalcification finding; F = 0.340, P = .848 for the mass finding; and F = 0.122, P = .975 for both findings. CONCLUSION: To within the accuracy of these measurements, lossy compression of digital mammographic data at 80:1 with JPEG2000 or the object-based SPIHT algorithm can be performed without decreasing the rate of detection of clusters of microcalcifications and masses.

Adult↗

Orthopedic hardware complications diagnosed with multi-detector row CT.

PURPOSE: To retrospectively evaluate multi-detector row computed tomography (CT) for the depiction of orthopedic hardware complications in the spine and appendicular skeleton. MATERIALS AND METHODS: This HIPAA-compliant study had institutional review board approval; patient informed consent was not required. Results of 114 multi-detector row CT studies performed because of possible hardware complications in 109 patients (57 men, 52 women; mean age, 44 years; age range, 12-82 years) were available for analysis. The CT studies were retrospectively reviewed and compared with clinical or surgical outcomes, which were used as the reference standard. In another experiment, detection of hardware complications on radiographs and multi-detector row CT images was compared between two readers for selected cases (18 positive and 26 negative) by using receiver operating characteristic (ROC) methods. RESULTS: For 91 (80%) of 114 multi-detector row CT studies, the complication status could be determined on the basis of clinical or surgical outcomes. Twenty-three multi-detector row CT studies were confirmed to be positive (revealing 10 cases of nonunion, five cases of hardware malplacement, three cases of hardware loosening, three perihardware fractures, and two chronic infections), and 57 were confirmed to be negative. There were three false-positive and eight false-negative multi-detector row CT studies. With clinical or surgical outcomes as the reference standard, the sensitivity, specificity, and positive and negative predictive values of multi-detector row CT were 74% (23 of 31 studies), 95% (57 of 60 studies), 88% (23 of 26 studies), and 88% (57 of 65 studies), respectively. Results of ROC analysis indicated that detection of hardware complications was much lower with radiography than with multi-detector row CT (area under ROC curve, 0.84 vs 1.00; F = 4.69, df = 1, 43; P < .05). CONCLUSION: Multi-detector row CT is an effective tool for depicting orthopedic hardware complications.

Adolescent↗

Multireader, multicase receiver operating characteristic analysis: an empirical comparison of five methods.

RATIONALE AND OBJECTIVES: Several statistical methods have been developed for analyzing multireader, multicase (MRMC) receiver operating characteristic (ROC) studies. The objective of this article is to increase awareness of these methods and determine if their results are concordant for published datasets. MATERIALS AND METHODS: Data from three previously published studies were reanalyzed using five MRMC methods. For each method the 95% confidence intervals (CIs) for the mean of the readers' ROC areas for each diagnostic test, the P value for the comparison of the diagnostic tests' mean accuracies, and the 95% CIs for the mean difference in ROC areas of the diagnostic tests were reported. RESULTS: Important differences in P values and CIs were seen when using parametric versus nonparametric estimates of accuracy, and there were the expected differences for random-reader versus fixed-reader models. Controlling for these differences, the Dorfman-Berbaum-Metz (DBM), Obuchowski-Rockette, Beiden-Wagner-Campbell, and Song's multivariate Wilcoxon-Mann-Whitney (WMW) methods gave almost identical results for the fixed-reader model. For the random-reader model, the DBM, Obuchowski-Rockette, and Beiden-Wagner-Campbell methods yielded approximately the same inferences, but the CIs for the Beiden-Wagner-Campbell method tend to be broader. Ishwaran's hierarchical ROC sometimes yielded significance not found with other methods. Song's modification of DBM's jack-knifing algorithm sometimes led to different conclusions than the original DBM algorithm. CONCLUSION: In choosing and applying MRMC methods, it is important to recognize: (1) the distinction between random-reader and fixed-reader models, the uncertainties accounted for by each, and thus the level of generalizeability expected from each; (2) assumptions made by the various MRMC methods; and (3) limitations of a five- or six-reader study when the reader variability is great.

Analysis of Variance↗

Power estimation for the Dorfman-Berbaum-Metz method.

RATIONALE AND OBJECTIVES: The authors describe how to perform power and sample size computations for the Dorfman-Berbaum-Metz method for analyzing multireader receiver operating characteristic (ROC) studies when data from a pilot study or from a previous study similar to the planned study are available. MATERIALS AND METHODS: Power and sample size computations are described in a step-by-step procedure. RESULTS: To illustrate the power computation method we treat 2 studies that did not show a significant difference between modalities as pilot studies for planning future studies. For the future studies we plan to compare the modalities with respect to the mean of the treatment-reader area under the curve (AUC) estimates. We show how to estimate the reader and case sample sizes for each future study to have 80% power to detect a specified difference in modality AUCs. CONCLUSIONS: Application of the power and sample size computation procedure for the Dorfman-Berbaum-Metz method is straightforward. For a given effect size there will be several different reader and case sample size combinations that yield the desired power. It is important that the pilot study or previous study used for the computations be comparable to the planned study with respect to the modalities, reader expertise, case selection, and ratio of normal to abnormal cases.

Analysis of Variance↗

Effect of sex and gender on drug-seeking behavior during invasive medical procedures.

RATIONALE AND OBJECTIVES: To assess how sex affects patients' drug-seeking, pain, and anxiety during interventional procedures. MATERIALS AND METHODS: Data from 159 patients were derived from two control groups of a prospective randomized trial. Seventy-six patients were male, 83 female. Patients in the standard group (n = 79) received the standard care typical for the institution; patients in the attention group (n = 80) had an additional empathic provider who stayed with them throughout the procedure. All patients were asked every 15 minutes to rate their pain and anxiety on 0-10 self-rating scales. All had access to intravenous sedatives and narcotics through a patient-controlled analgesia model. Univariate analysis of variance with a between-patient factor for group and another between-patient factor for sex was used. RESULTS: There was a significant interaction between group attribution and sex with regard to drug request and pain and anxiety ratings. Patients in the attention group requested significantly fewer drugs than patients in the standard group. Men asked for more drugs than women under standard care, but for less in the attention group. Pain and anxiety ratings for women were significantly lower in the attention group compared with standard treatment, but for men, there was no significant difference. CONCLUSION: Although both men and women benefit from the presence of an empathic provider during invasive medical procedures, men benefit more in terms of medication reduction, whereas women benefit more in terms of pain and anxiety reduction. Awareness of these gender-specific differences can aid in formulation of patient-specific treatment plans.

Adolescent↗

Observer studies involving detection and localization: modeling, analysis, and validation.

Although the receiver operating characteristic (ROC) paradigm is the accepted method for evaluation of diagnostic imaging systems, it has some serious shortcomings inasmuch as it is restricted to one observer report per image. By contrast the free-response ROC (FROC) paradigm and associated analysis method allows the observer to report multiple abnormalities within each imaging study, and uses the location of reported abnormalities to improve the measurement. Because the ROC method cannot accommodate multiple responses or use location information, its statistical power will suffer. The FROC paradigm/analysis has not enjoyed widespread acceptance because of concern about whether responses made to the same diagnostic study can be treated as independent. We propose a new jackknife FROC analysis method (JAFROC) that does not make the independence assumption. The new analysis method combines elements of FROC and the Dorfman-Berbaum-Metz (DBM) methods. To compare JAFROC to an earlier free-response analysis method (specifically the alternative free-response, or AFROC method), and to the DBM method, which uses conventional ROC scoring, we developed a model for generating simulated FROC data. The simulation model is based on an eye-movement model of how experts evaluate images. It allowed us to examine null hypothesis (NH) behavior and statistical power of the different methods. We found that AFROC analysis did not pass the NH test, being unduly conservative. Both the JAFROC method and the DBM method passed the NH test, but JAFROC had more statistical power than the DBM method. The results of this comparison suggest that future studies of diagnostic performance may enjoy improved statistical power or reduced sample size requirements through the use of the JAFROC method.

Algorithms↗

An empirical comparison of discrete ratings and subjective probability ratings.

RATIONALE AND OBJECTIVES: The authors compared receiver operating characteristic (ROC) data from a five-category discrete scale with that from a 101-category subjective probability scale to determine how well the latter categories define the ROC curve. MATERIALS AND METHODS: The authors analyzed data from a pilot study performed for another purpose in which 10 radiologists provided both a five-point confidence rating and a subjective probability rating of abnormality for each interpretation. ROC operating points were plotted for a five-category scale and a 101-category scale to determine how well the observed points covered the range of false-positive probabilities. ROC curves were fitted to the subjective probability data according to the standard ROC model. RESULTS: For these data, subjective probability ratings were somewhat more effective in populating the range of false-positive probability with ROC points. For three observers, the ROC curves inappropriately crossed the chance line. For another four, prevention of such crossing seemed to depend on one or two ROC points near the upper right corner of the ROC space, points based on discriminations within the discrete category "no abnormality to report." CONCLUSION: Subjective probability rating should provide substantially better coverage of the ROC space with operating points, preventing inappropriate crossing of the chance line. Unfortunately, the protection offered by subjective probability ratings was unreliable and depended on ROC points derived from discriminations not directly related to apparent abnormality. The use of proper ROC models to fit data may offer a better solution.

Observer Variation↗