PubMed Health⌕ Search

Biomedical subjects

Charles E Metz

Publications and source records attributed to Charles E Metz.

18 recordsLinked to original sources

Comparison of independent double readings and computer-aided diagnosis (CAD) for the diagnosis of breast calcifications.

RATIONALE AND OBJECTIVES: The aim of the study is to compare independent double readings by radiologists and computer-aided diagnosis (CAD) in diagnostic interpretation of mammographic calcifications. MATERIALS AND METHODS: Ten radiologists independently interpreted 104 mammograms containing clustered microcalcifications. Forty-six of these were malignant and 58 were benign at biopsy. Radiologists read the images with and without a computer aid by using a counterbalanced study design. Sensitivity and specificity were calculated from observer biopsy recommendations, and receiver operating characteristic (ROC) curves were computed from their diagnostic confidence ratings. Unaided double-reading sensitivity and specificity values were derived post hoc by using three different objective rules and an additional rule of simulated-optimal double reading that assumed that consultations for resolving two radiologists' different independent diagnoses always produce the correct clinical recommendation. ROC curves of unaided double readings were obtained according to the literature. RESULTS: Single reading without computer aid yielded 74% sensitivity and 32% specificity, whereas CAD reading yielded 87% sensitivity and 42% specificity and appeared on a higher ROC curve (P < .0001). Three methods of formulating independent double readings generated sensitivities between 59% and 89%, specificities between 50% and 13%, and operating points that moved essentially along the average unaided single-reading ROC curve. ROC curves of unaided independent double readings showed small, statistically insignificant improvement over those of unaided single readings. Results of the simulated-optimal double reading were similar to CAD: 89% sensitivity and 50% specificity. CONCLUSION: Independent double readings of mammographic calcifications may not improve diagnostic performance. CAD reading improves diagnostic performance to an extent approaching the maximum possible performance.

Breast Diseases↗

A quadratic model for combining quantitative diagnostic assessments from radiologist and computer in computer-aided diagnosis.

RATIONALE AND OBJECTIVES: Some computer-aided diagnosis (CAD) methods produce a quantitative diagnostic assessment (eg, likelihood of malignancy) based on computer image analysis that a radiologist who uses the computer aid must combine with his or her own assessment. Observer studies show that although CAD helps radiologists improve diagnostic performance, ad hoc use of computer aid can produce performance inferior to that of computer alone, indicating that radiologists are unable to incorporate computer assessment optimally into their final assessment. We describe a mathematical model for combining two correlated diagnostic assessments that may provide a basis for merging radiologists' ratings with computer assessments in a way that yields greater diagnostic accuracy than ad hoc merging by radiologists. MATERIALS AND METHODS: We calculate a likelihood ratio from the bivariate binormal model that describes joint probability density functions of latent decision variables of two correlated diagnostic assessments. To the extent that the bivariate binormal model is valid and that the model's parameters can be estimated reliably, results obtained in this way will be optimal because the likelihood ratio is the decision variable used by the ideal observer in any two-group classification task. We evaluated this method on two observer study datasets and in Monte Carlo simulations. RESULTS: This method produced better performance than achieved by radiologists when they incorporated computer assessment in an ad hoc way. Simulations show that with a large number of cases, this method can produce results indistinguishable from the ideal observer performance. CONCLUSIONS: This method can potentially help radiologists use quantitative computed diagnostic assessments optimally, thereby surpassing the computer in accuracy.

Analysis of Variance↗

Three-class ROC analysis--a decision theoretic approach under the ideal observer framework.

Receiver operating characteristic (ROC) analysis is well established in the evaluation of systems involving binary classification tasks. However, medical tests often require distinguishing among more than two diagnostic alternatives. The goal of this work was to develop an ROC analysis method for three-class classification tasks. Based on decision theory, we developed a method for three-class ROC analysis. In this method, the objects were classified by making the decision that provided the maximal utility relative to the other two. By making assumptions about the magnitudes of the relative utilities of incorrect decisions, we found a decision model that maximized the expected utility of the decisions when using log-likelihood ratios as decision variables. This decision model consists of a two-dimensional decision plane with log likelihood ratios as the axes and a decision structure that separates the plane into three regions. Moving the decision structure over the decision plane, which corresponds to moving the decision threshold in two-class ROC analysis, and computing the true class 1, 2, and 3 fractions defined a three-class ROC surface. We have shown that the resulting three-class ROC surface shares many features with the two-class ROC curve; i.e., using the log likelihood ratios as the decision variables results in maximal expected utility of the decisions, and the optimal operating point for a given diagnostic setting (set of relative utilities and disease prevalences) lies on the surface. The volume under the three-class surface (VUS) serves as a figure-of-merit to evaluate different data acquisition systems or image processing and reconstruction methods when the assumed utility constraints are relevant.

Algorithms↗

The hypervolume under the ROC hypersurface of "near-guessing" and "near-perfect" observers in N-class classification tasks.

We express the performance of the N-class "guessing" observer in terms of the N2-N conditional probabilities which make up an N-class receiver operating characteristic (ROC) space, in a formulation in which sensitivities are eliminated in constructing the ROC space (equivalent to using false-negative fraction and false-positive fraction in a two-class task). We then show that the "guessing" observer's performance in terms of these conditional probabilities is completely described by a degenerate hypersurface with only N-1 degrees of freedom (as opposed to the N2-N-1 required, in general, to achieve a true hypersurface in such a ROC space). It readily follows that the hypervolume under such a degenerate hypersurface must be zero when N > 2. We then consider a "near-guessing" task; that is, a task in which the N underlying data probability density functions (pdfs) are nearly identical, controlled by N-1 parameters which may vary continuously to zero (at which point the pdfs become identical). With this approach, we show that the hypervolume under the ROC hypersurface of an observer in an N-class classification task tends continuously to zero as the underlying data pdfs converge continuously to identity (a "guessing" task). The hypervolume under the ROC hypersurface of a "perfect" ideal observer (in a task in which the N data pdfs never overlap) is also found to be zero in the ROC space formulation under consideration. This suggests that hypervolume may not be a useful performance metric in N-class classification tasks for N > 2, despite the utility of the area under the ROC curve for two-class tasks.

Algorithms↗

Restrictions on the three-class ideal observer's decision boundary lines.

We are attempting to develop expressions for the coordinates of points on the three-class ideal observer's receiver operating characteristic (ROC) hypersurface as functions of the set of decision criteria used by the ideal observer. This is considerably more difficult than in the two-class classification task, because the conditional probabilities in question are not simply related to the cumulative distribution functions of the decision variables, and because the slopes and intercepts of the decision boundary lines are not independent; given the locations of two of the lines, the location of the third will be constrained depending on the other two. In this paper, we attempt to characterize those constraining relationships among the three-class ideal observer's decision boundary lines. As a result, we show that the relationship between the decision criteria and the misclassification probabilities is not one-to-one, as it is for the two-class ideal observer.

Algorithms↗

Effect of correlation on combining diagnostic information from two images of the same patient.

We have shown previously, in the context of computer-aided diagnosis (CAD), that information derived from multiple images of the same patient can be used to improve diagnostic performance. In that work, we ignored the correlation among multiple images of the same patient. In the present study, we investigate theoretically, within the framework of receiver operating characteristic (ROC) analysis, the effect of correlation on three methods for combining quantitative diagnostic information from two images: taking the average, the maximum, and the minimum of a pair of normally distributed decision variables. We assume, as in our previous work, that the quantitative diagnostic information obtained from the two images of a given patient can be transformed monotonically to two latent decision variables that are normally distributed. Similar to the situation of uncorrelated images, we found that (1) the average always improves the area under the ROC curve (AUC) compared to the single-view image; (2) the maximum and the minimum can also, but not always, improve the AUC; and (3) each method can be the best method in certain situations. In addition, as the correlation strength increases, the average performs the best less often, whereas the maximum and the minimum perform the best more often. These theoretical results are illustrated with analysis of a mammography study.

Area Under Curve↗

Robustness of computerized lesion detection and classification scheme across different breast US platforms.

PURPOSE: To evaluate the performance of a computerized detection and diagnosis method with breast ultrasonographic (US) images obtained with US equipment from two different manufacturers. MATERIALS AND METHODS: Two independent clinical breast US databases were used in this performance study. Data collection and database use were HIPAA-compliant and followed institutional review board-approved protocols, with waiver of informed consent. One database consisted of 1740 images obtained in 458 women with Philips US equipment. The other database consisted of 151 images obtained in 151 women with Siemens US equipment. The testing protocols included independent testing and round-robin analysis. The computerized scheme detects potential lesions, calculates imaging features for all candidate lesions, and subsequently classifies candidate lesions into different categories. Two separate classification tasks were evaluated: distinction between all actual lesions and false-positive detections and distinction between actual cancers and all other detected lesion candidates. Statistical analysis was performed by using both receiver operating characteristic (ROC) and free-response ROC methods. RESULTS: For the distinction between all actual lesions and false-positive detections, area under the ROC curve (A(z)) values ranged between 0.87 and 0.95 for different testing protocols. In two instances, the difference in performance between databases was significant (P < .01), but it was shown that this was due to the difference in size of the databases. In the distinction of cancer from all other detections, the A(z) values ranged between 0.80 and 0.86. No statistically significant difference was found among the different testing protocols in this instance. CONCLUSION: These results indicate that the performance of this fully automated computerized lesion detection and classification method, which demonstrated robustness over the different US equipment used, is promising.

Adolescent↗

Assessment methodologies and statistical issues for computer-aided diagnosis of lung nodules in computed tomography: contemporary research topics relevant to the lung image database consortium.

Cancer of the lung and bronchus is the leading fatal malignancy in the United States. Five-year survival is low, but treatment of early stage disease considerably improves chances of survival. Advances in multidetector-row computed tomography technology provide detection of smaller lung nodules and offer a potentially effective screening tool. The large number of images per exam, however, requires considerable radiologist time for interpretation and is an impediment to clinical throughput. Thus, computer-aided diagnosis (CAD) methods are needed to assist radiologists with their decision making. To promote the development of CAD methods, the National Cancer Institute formed the Lung Image Database Consortium (LIDC). The LIDC is charged with developing the consensus and standards necessary to create an image database of multidetector-row computed tomography lung images as a resource for CAD researchers. To develop such a prospective database, its potential uses must be anticipated. The ultimate applications will influence the information that must be included along with the images, the relevant measures of algorithm performance, and the number of required images. In this article we outline assessment methodologies and statistical issues as they relate to several potential uses of the LIDC database. We review methods for performance assessment and discuss issues of defining "truth" as well as the complications that arise when truth information is not available. We also discuss issues about sizing and populating a database.

Algorithms↗

Ideal observers and optimal ROC hypersurfaces in N-class classification.

The likelihood ratio, or ideal observer, decision rule is known to be optimal for two-class classification tasks in the sense that it maximizes expected utility (or, equivalently, minimizes the Bayes risk). Furthermore, using this decision rule yields a receiver operating characteristic (ROC) curve which is never above the ROC curve produced using any other decision rule, provided the observer's misclassification rate with respect to one of the two classes is chosen as the dependent variable for the curve (i.e., an "inversion" of the more common formulation in which the observer's true-positive fraction is plotted against its false-positive fraction). It is also known that for a decision task requiring classification of observations into N classes, optimal performance in the expected utility sense is obtained using a set of N-1 likelihood ratios as decision variables. In the N-class extension of ROC analysis, the ideal observer performance is describable in terms of an (N2-N-1)-parameter hypersurface in an (N2-N)-dimensional probability space. We show that the result for two classes holds in this case as well, namely that the ROC hypersurface obtained using the ideal observer decision rule is never above the ROC hypersurface obtained using any other decision rule (where in our formulation performance is given exclusively with respect to between-class error rates rather than within-class sensitivities).

Bayes Theorem↗

Estimating three-class ideal observer decision variables for computerized detection and classification of mammographic mass lesions.

We are using Bayesian artificial neural networks (BANNs) to classify mammographic masses in schemes for computer-aided diagnosis, and we are extending this methodology to a three-class classification task. We investigated whether a BANN can estimate ideal observer decision variables to distinguish malignant, benign, and false-positive computer detections. Five features were calculated for 63 malignant and 29 benign computer-detected mass lesions, and for 1049 false-positive computer detections, in 440 mammograms randomly divided into a training and testing set. A BANN was trained on the training set features and applied to the testing set features. We then used a known relation between three-class ideal observer decision variables and that used by a two-class ideal observer when two of three classes are grouped into one class, giving one decision variable for distinguishing malignant from nonmalignant detections, and a second for distinguishing true-positive from false-positive computer detections. For comparison, we grouped the training data into two classes in the same two ways and trained two-class BANNs for these two tasks. The three-class BANN decision variables were essentially identical in performance to the specifically trained two-class BANNs, with the average difference in area under the ROC curves being less than 0.0035 and no differences in area being statistically significant. Thus, the BANN outputs obey the same theoretical relationship as do the three-class and two-class ideal observer decision variables, which is consistent with the claim that the three-class BANN output can provide good estimates of the decision variables used by a three-class ideal observer.

Breast Neoplasms↗

An ROC comparison of four methods of combining information from multiple images of the same patient.

Variance of diagnostic information contained in an image degrades diagnostic accuracy. Acquiring multiple images of the same patient (e.g., mediolateral oblique and craniocaudal view mammograms) can, in principle, help reduce this degradation. We demonstrate how this can be accomplished in the context of computer-aided diagnosis (CAD). Assuming that computer outputs obtained from multiple images of the same patient can be transformed monotonically to the same pair of truth-conditional normal distributions and, for simplicity, ignoring correlation among images, we investigate theoretically four methods of combining the computer outputs: taking the average, the median, the maximum, or the minimum. We found, as one would expect, that both the average and the median always produce an improved area under the receiver operating characteristic (ROC) curve (AUC) compared to the single-view images, while the average always produces better performance than the median. However, the maximum and minimum also can produce improved AUCs in some situations, and under certain conditions can outperform the average. Surprisingly, we found that the maximum and minimum of normally-distributed decision variables produce nearly binormal ROC curves. These results can be used as a guide in attempting to increase the efficacy of CAD when multiple images are available from the same patient.

Algorithms↗

Short-scan SPECT imaging with non-uniform attenuation and 3D distance-dependent spatial resolution.

Image quality and quantitative accuracy in single-photon emission computed tomography (SPECT) can be degraded by, e.g., the effects of photon attenuation and finite spatial resolution. It is generally considered that adequate compensation for such effects on SPECT images requires data acquired over 2 pi. Recently, using the existing consistency condition on the data function, Noo and Wagner (2001 Inverse Problems 17 1357-72) have shown analytically that data acquired over only pi can be used to correct completely for the effect of uniform attenuation in SPECT. It remains unknown, however, whether data acquired only over pi in SPECT with non-uniform attenuation and/or 3D distance-dependent spatial resolution (DDSR) contain complete information for accurate image reconstruction. In this work, we develop a heuristic perspective, which is referred to as the potato peeler perspective to show conceptually that data in SPECT with non-uniform attenuation and/or 3D DDSR acquired over 2 pi contain redundant information and that such information can be used to reduce the scanning angle in SPECT. Specifically, we show heuristically that, in SPECT with only non-uniform attenuation, the scanning angle can be reduced from 2 pi to pi and that, in SPECT with both non-uniform attenuation and DDSR with a physically realistic form, the scanning angle can be reduced from 2 pi to pi in a practical sense. We conduct computer simulation studies, and the results from these studies corroborate the observations obtained based upon the heuristic potato peeler perspective.

Algorithms↗

Assessment of medical imaging and computer-assist systems: lessons from recent experience.

In the last 2 decades major advances have been made in the field of assessment methods for medical imaging and computer-assist systems through the use of the paradigm of the receiver operating characteristic (ROC) curve. In the most recent decade this methodology was extended to embrace the complication of reader variability through advances in the multiple-reader, multiple-case (MRMC) ROC measurement and analysis paradigm. Although this approach has been widely adopted by the imaging research community, some investigators appear averse to it, possibly from concern that it could place a greater burden on the scarce resources of patient cases and readers compared to the requirements of alternative methods. The present communication argues, however, that the MRMC ROC approach to assessment in the context of reader variability may be the most resource-efficient approach available. Moreover, alternative approaches may also be statistically uninterpretable with regard to estimated summary measures of performance and their uncertainties. The authors propose that the MRMC ROC approach be considered even more widely by the larger community with responsibilities for the introduction and dissemination of medical imaging technologies to society. General principles of study design are reviewed, and important contemporary clinical trials are used as examples.

Diagnosis, Computer-Assisted↗

Maximum likelihood fitting of FROC curves under an initial-detection-and-candidate-analysis model.

We have developed a model for FROC curve fitting that relates the observer's FROC performance not to the ROC performance that would be obtained if the observer's responses were scored on a per image basis, but rather to a hypothesized ROC performance that the observer would obtain in the task of classifying a set of "candidate detections" as positive or negative. We adopt the assumptions of the Bunch FROC model, namely that the observer's detections are all mutually independent, as well as assumptions qualitatively similar to, but different in nature from, those made by Chakraborty in his AFROC scoring methodology. Under the assumptions of our model, we show that the observer's FROC performance is a linearly scaled version of the candidate analysis ROC curve, where the scaling factors are just given by the FROC operating point coordinates for detecting initial candidates. Further, we show that the likelihood function of the model parameters given observational data takes on a simple form, and we develop a maximum likelihood method for fitting a FROC curve to this data. FROC and AFROC curves are produced for computer vision observer datasets and compared with the results of the AFROC scoring method. Although developed primarily with computer vision schemes in mind, we hope that the methodology presented here will prove worthy of further study in other applications as well.

Algorithms↗

Breast cancer: effectiveness of computer-aided diagnosis observer study with independent database of mammograms.

PURPOSE: To evaluate the effectiveness of a computerized classification method as an aid to radiologists reviewing clinical mammograms for which the diagnoses were unknown to both the radiologists and the computer. MATERIALS AND METHODS: Six mammographers and six community radiologists participated in an observer study. These 12 radiologists interpreted, with and without the computer aid, 110 cases that were unknown to both the 12 radiologist observers and the trained computer classification scheme. The radiologists' performances in differentiating between benign and malignant masses without and with the computer aid were evaluated with receiver operating characteristic (ROC) analysis. Two-tailed P values were calculated for the Student t test to indicate the statistical significance of the differences in performances with and without the computer aid. RESULTS: When the computer aid was used, the average performance of the 12 radiologists improved, as indicated by an increase in the area under the ROC curve (A(z)) from 0.93 to 0.96 (P <.001), by an increase in partial area under the ROC curve ((0.90)A(')(z)) from 0.56 to 0.72 (P <.001), and by an increase in sensitivity from 94% to 98% (P =.022). No statistically significant difference in specificity was found between readings with and those without computer aid (Delta = -0.014; P =.46; 95% CI: -0.054, 0.026), where Delta is difference in specificity. When we analyzed results from the mammographers and community radiologists as separate groups, a larger improvement was demonstrated for the community radiologists. CONCLUSION: Computer-aided diagnosis can potentially help radiologists improve their diagnostic accuracy in the task of differentiating between benign and malignant masses seen on mammograms.

Biopsy↗

Computerized analysis of digitized mammograms of BRCA1 and BRCA2 gene mutation carriers.

PURPOSE: To evaluate, by using computer image analysis, the mammographic density patterns of women with germ-line mutations in BRCA1 and BRCA2 genes in comparison with those of women at low risk of developing breast cancer. MATERIALS AND METHODS: Mammograms from 30 carriers of BRCA1 and BRCA2 mutations and from 142 low-risk women were collected retrospectively and digitized. In addition, 60 of the 142 low-risk women were randomly selected and age matched at 5-year intervals with the 30 mutation carriers. Mammographic features were extracted from the central regions of the breast images to characterize the mammographic density and heterogeneity of dense portions of the breast. These features were then merged into a single value related to the risk of breast cancer by using linear discriminant analysis. The applicability of these computer-extracted features and the output from linear discriminant analysis to differentiate between the carriers of BRCA1 and BRCA2 mutations and the low-risk women in the entire database and in an age-matched group were evaluated by using receiver operating characteristic analysis. RESULTS: Quantitative analysis of mammograms demonstrated that carriers of BRCA1 and BRCA2 mutations tended to have dense breast tissue, and their mammographic patterns tended to be low in contrast, with a coarse texture. Linear discriminant analysis resulted in values of the areas under the receiver operating characteristic curve of 0.91 and 0.92 in distinguishing between the BRCA1 and BRCA2 mutation carriers and the low-risk women in the entire database and the age-matched group, respectively. CONCLUSION: The computerized analysis of mammograms suggests that mammographic patterns in carriers of BRCA1 and BRCA2 mutations differ from those of women at low risk for breast cancer. Our computer-extracted features may be useful as radiographic markers for identifying women at high risk for breast cancer.

Adult↗

Computer-aided diagnosis in chest radiography: results of large-scale observer tests at the 1996-2001 RSNA scientific assemblies.

Since 1996, computer-aided diagnosis (CAD) schemes have been presented as interactive demonstrations on computer workstations at each scientific assembly of the Radiological Society of North America. The schemes involved (a) detection of pulmonary nodules, (b) temporal subtraction, (c) detection of interstitial lung disease, (d) differential diagnosis of interstitial lung disease, and (e) distinction between benign and malignant pulmonary nodules on chest radiographs. Large-scale observer tests were carried out to examine how radiologists can benefit from CAD systems. Observer performance was evaluated by analysis of receiver operating characteristic (ROC) curves. The statistical significance of the difference between the areas under the ROC curves without and with CAD was analyzed with the Student t test. In all of the tests, the diagnostic accuracy of the radiologists in total improved significantly when CAD was used. This result provides additional evidence that CAD has the potential to improve the performance of radiologists in their decision-making process in interpreting chest radiographs.

Diagnosis, Computer-Assisted↗