PubMed Health⌕ Search

Biomedical subjects

Robert M Nishikawa

Publications and source records attributed to Robert M Nishikawa.

14 recordsLinked to original sources

Comparison of independent double readings and computer-aided diagnosis (CAD) for the diagnosis of breast calcifications.

RATIONALE AND OBJECTIVES: The aim of the study is to compare independent double readings by radiologists and computer-aided diagnosis (CAD) in diagnostic interpretation of mammographic calcifications. MATERIALS AND METHODS: Ten radiologists independently interpreted 104 mammograms containing clustered microcalcifications. Forty-six of these were malignant and 58 were benign at biopsy. Radiologists read the images with and without a computer aid by using a counterbalanced study design. Sensitivity and specificity were calculated from observer biopsy recommendations, and receiver operating characteristic (ROC) curves were computed from their diagnostic confidence ratings. Unaided double-reading sensitivity and specificity values were derived post hoc by using three different objective rules and an additional rule of simulated-optimal double reading that assumed that consultations for resolving two radiologists' different independent diagnoses always produce the correct clinical recommendation. ROC curves of unaided double readings were obtained according to the literature. RESULTS: Single reading without computer aid yielded 74% sensitivity and 32% specificity, whereas CAD reading yielded 87% sensitivity and 42% specificity and appeared on a higher ROC curve (P < .0001). Three methods of formulating independent double readings generated sensitivities between 59% and 89%, specificities between 50% and 13%, and operating points that moved essentially along the average unaided single-reading ROC curve. ROC curves of unaided independent double readings showed small, statistically insignificant improvement over those of unaided single readings. Results of the simulated-optimal double reading were similar to CAD: 89% sensitivity and 50% specificity. CONCLUSION: Independent double readings of mammographic calcifications may not improve diagnostic performance. CAD reading improves diagnostic performance to an extent approaching the maximum possible performance.

Breast Diseases↗

Identification of simulated microcalcifications in white noise and mammographic backgrounds.

This work investigates human performance in discriminating between differently shaped simulated microcalcifications embedded in white noise or mammographic backgrounds. Human performance was determined through two alternative forced-choice (2-AFC) experiments. The signals used were computer-generated simple shapes that were designed such that they had equal signal energy. This assured equal detectability. For experiments involving mammographic backgrounds, signals were blurred to account for the imaging system modulation transfer function (MTF). White noise backgrounds were computer generated; anatomic background patches were extracted from normal mammograms. We compared human performance levels as a function of signal energy in the expected difference template. In the discrimination task, the expected difference template is the difference between the two signals shown. In white noise backgrounds, human performance in the discrimination task was degraded compared to the detection task. In mammographic backgrounds, human performance in the discrimination task exceeded that of the detection task. This indicates that human observers do not follow the optimum decision strategy of correlating the expected signal template with the image. Human observer performance was qualitatively reproduced by non-prewhitening with eye filter (NPWE) model observer calculations, in which spatial uncertainty was explicitly included by shifting the locations of the expected difference templates. The results indicate that human strategy in the discrimination task may be to match individual signal templates with the image individually, rather than to perform template matching between the expected difference template and the image.

Artifacts↗

The hypervolume under the ROC hypersurface of "near-guessing" and "near-perfect" observers in N-class classification tasks.

We express the performance of the N-class "guessing" observer in terms of the N2-N conditional probabilities which make up an N-class receiver operating characteristic (ROC) space, in a formulation in which sensitivities are eliminated in constructing the ROC space (equivalent to using false-negative fraction and false-positive fraction in a two-class task). We then show that the "guessing" observer's performance in terms of these conditional probabilities is completely described by a degenerate hypersurface with only N-1 degrees of freedom (as opposed to the N2-N-1 required, in general, to achieve a true hypersurface in such a ROC space). It readily follows that the hypervolume under such a degenerate hypersurface must be zero when N > 2. We then consider a "near-guessing" task; that is, a task in which the N underlying data probability density functions (pdfs) are nearly identical, controlled by N-1 parameters which may vary continuously to zero (at which point the pdfs become identical). With this approach, we show that the hypervolume under the ROC hypersurface of an observer in an N-class classification task tends continuously to zero as the underlying data pdfs converge continuously to identity (a "guessing" task). The hypervolume under the ROC hypersurface of a "perfect" ideal observer (in a task in which the N data pdfs never overlap) is also found to be zero in the ROC space formulation under consideration. This suggests that hypervolume may not be a useful performance metric in N-class classification tasks for N > 2, despite the utility of the area under the ROC curve for two-class tasks.

Algorithms↗

A study on several machine-learning methods for classification of malignant and benign clustered microcalcifications.

In this paper, we investigate several state-of-the-art machine-learning methods for automated classification of clustered microcalcifications (MCs). The classifier is part of a computer-aided diagnosis (CADx) scheme that is aimed to assisting radiologists in making more accurate diagnoses of breast cancer on mammograms. The methods we considered were: support vector machine (SVM), kernel Fisher discriminant (KFD), relevance vector machine (RVM), and committee machines (ensemble averaging and AdaBoost), of which most have been developed recently in statistical learning theory. We formulated differentiation of malignant from benign MCs as a supervised learning problem, and applied these learning methods to develop the classification algorithm. As input, these methods used image features automatically extracted from clustered MCs. We tested these methods using a database of 697 clinical mammograms from 386 cases, which included a wide spectrum of difficult-to-classify cases. We analyzed the distribution of the cases in this database using the multidimensional scaling technique, which reveals that in the feature space the malignant cases are not trivially separable from the benign ones. We used receiver operating characteristic (ROC) analysis to evaluate and to compare classification performance by the different methods. In addition, we also investigated how to combine information from multiple-view mammograms of the same case so that the best decision can be made by a classifier. In our experiments, the kernel-based methods (i.e., SVM, KFD, and RVM) yielded the best performance (Az = 0.85, SVM), significantly outperforming a well-established, clinically-proven CADx approach that is based on neural network (Az = 0.80).

Algorithms↗

Relevance vector machine for automatic detection of clustered microcalcifications.

Clustered microcalcifications (MC) in mammograms can be an important early sign of breast cancer in women. Their accurate detection is important in computer-aided detection (CADe). In this paper, we propose the use of a recently developed machine-learning technique--relevance vector machine (RVM)--for detection of MCs in digital mammograms. RVM is based on Bayesian estimation theory, of which a distinctive feature is that it can yield a sparse decision function that is defined by only a very small number of so-called relevance vectors. By exploiting this sparse property of the RVM, we develop computerized detection algorithms that are not only accurate but also computationally efficient for MC detection in mammograms. We formulate MC detection as a supervised-learning problem, and apply RVM as a classifier to determine at each location in the mammogram if an MC object is present or not. To increase the computation speed further, we develop a two-stage classification network, in which a computationally much simpler linear RVM classifier is applied first to quickly eliminate the overwhelming majority, non-MC pixels in a mammogram from any further consideration. The proposed method is evaluated using a database of 141 clinical mammograms (all containing MCs), and compared with a well-tested support vector machine (SVM) classifier. The detection performance is evaluated using free-response receiver operating characteristic (FROC) curves. It is demonstrated in our experiments that the RVM classifier could greatly reduce the computational complexity of the SVM while maintaining its best detection accuracy. In particular, the two-stage RVM approach could reduce the detection time from 250 s for SVM to 7.26 s for a mammogram (nearly 35-fold reduction). Thus, the proposed RVM classifier is more advantageous for real-time processing of MC clusters in mammograms.

Algorithms↗

A similarity learning approach to content-based image retrieval: application to digital mammography.

In this paper, we describe an approach to content-based retrieval of medical images from a database, and provide a preliminary demonstration of our approach as applied to retrieval of digital mammograms. Content-based image retrieval (CBIR) refers to the retrieval of images from a database using information derived from the images themselves, rather than solely from accompanying text indices. In the medical-imaging context, the ultimate aim of CBIR is to provide radiologists with a diagnostic aid in the form of a display of relevant past cases, along with proven pathology and other suitable information. CBIR may also be useful as a training tool for medical students and residents. The goal of information retrieval is to recall from a database information that is relevant to the user's query. The most challenging aspect of CBIR is the definition of relevance (similarity), which is used to guide the retrieval machine. In this paper, we pursue a new approach, in which similarity is learned from training examples provided by human observers. Specifically, we explore the use of neural networks and support vector machines to predict the user's notion of similarity. Within this framework we propose using a hierarchal learning approach, which consists of a cascade of a binary classifier and a regression module to optimize retrieval effectiveness and efficiency. We also explore how to incorporate online human interaction to achieve relevance feedback in this learning framework. Our experiments are based on a database consisting of 76 mammograms, all of which contain clustered microcalcifications (MCs). Our goal is to retrieve mammogram images containing similar MC clusters to that in a query. The performance of the retrieval system is evaluated using precision-recall curves computed using a cross-validation procedure. Our experimental results demonstrate that: 1) the learning framework can accurately predict the perceptual similarity reported by human observers, thereby serving as a basis for CBIR; 2) the learning-based framework can significantly outperform a simple distance-based similarity metric; 3) the use of the hierarchical two-stage network can improve retrieval performance; and 4) relevance feedback can be effectively incorporated into this learning framework to achieve improvement in retrieval precision based on online interaction with users; and 5) the retrieved images by the network can have predicting value for the disease condition of the query.

Algorithms↗

Estimating three-class ideal observer decision variables for computerized detection and classification of mammographic mass lesions.

We are using Bayesian artificial neural networks (BANNs) to classify mammographic masses in schemes for computer-aided diagnosis, and we are extending this methodology to a three-class classification task. We investigated whether a BANN can estimate ideal observer decision variables to distinguish malignant, benign, and false-positive computer detections. Five features were calculated for 63 malignant and 29 benign computer-detected mass lesions, and for 1049 false-positive computer detections, in 440 mammograms randomly divided into a training and testing set. A BANN was trained on the training set features and applied to the testing set features. We then used a known relation between three-class ideal observer decision variables and that used by a two-class ideal observer when two of three classes are grouped into one class, giving one decision variable for distinguishing malignant from nonmalignant detections, and a second for distinguishing true-positive from false-positive computer detections. For comparison, we grouped the training data into two classes in the same two ways and trained two-class BANNs for these two tasks. The three-class BANN decision variables were essentially identical in performance to the specifically trained two-class BANNs, with the average difference in area under the ROC curves being less than 0.0035 and no differences in area being statistically significant. Thus, the BANN outputs obey the same theoretical relationship as do the three-class and two-class ideal observer decision variables, which is consistent with the claim that the three-class BANN output can provide good estimates of the decision variables used by a three-class ideal observer.

Breast Neoplasms↗

Investigation of physical image quality indices of a bone densitometry system.

Osteoporosis is a disease characterized by a loss of bone mass and a deterioration of bone structure. Bone mineral density (BMD) measures bone mass and is currently the method used to diagnose osteoporosis, while computerized radiographic texture analysis (RTA) is being investigated as a measure of bone structure. The GE/Lunar PIXI peripheral bone densitometer (PD) system, which uses dual-energy subtraction to measure BMD, also provides a digital image of the heel or forearm. The goal of our current research was to evaluate the physical imaging properties of the PIXI system (pixel size of 0.2 mm) compared to a Fuji computed radiography (CR) system (pixel size of 0.1 mm) to determine its suitability for texture analysis from image data. Contrast was measured using a series of uniform images covering the useful clinical exposure range. Spatial resolution was characterized by the presampling modulation transfer function (MTF) determined by an edge method. Noise power spectra (NPS) for different exposures were calculated using a two-dimensional Fourier analysis method. The expectation modulation transfer function was measured and combined with the NPS data to calculate the noise-equivalent number of quanta. The slope of the characteristic curve of the peripheral densitometer (PD) system was found to be position dependent across the image, although this dependence was substantially reduced by use of the system's clinical-settings corrections. An MTF value of 0.5 was found at 0.5 cycles/mm for the densitometry system compared to the same value at 1.6 cycles/mm for the CR system. Unlike the CR system, the NPS of the densitometry system was found not to be directionally dependent and did not drop off at higher spatial frequencies.

Absorptiometry, Photon↗

Radial gradient-based segmentation of mammographic microcalcifications: observer evaluation and effect on CAD performance.

Precise segmentation of microcalcifications is essential in the development of accurate mammographic computer-aided diagnosis (CAD) schemes. We have designed a radial gradient-based segmentation method for microcalcifications, and compared it to both the region-growing segmentation method currently used in our CAD scheme and to the watershed segmentation method. Two observer studies were conducted to subjectively evaluate the proposed segmentation method. The first study (A) required observers to rate the segmentation accuracy on a 100-point scale. The second observer evaluation (B) was a preference study in which observers selected their preferred method from three displayed segmentation methods. In study A, the observers gave an average accuracy rating of 88 for the radial gradient-based and 50 for the region-growing segmentation method. In study B, the two observers selected the proposed method 56% and 62% of the time. We also investigated the effect of the proposed segmentation method on the performance of computerized classification scheme in differentiating malignant from benign clustered microcalcifications. The performances of the classification scheme using a linear discriminant analysis (LDA) or a Bayesian artificial neural network classifier both showed statistically significant improvements when using the proposed segmentation method. The areas under the receiver-operating characteristic curves for case-based performance when using the LDA classifier were 0.86 with the proposed segmentation method, 0.80 with the region-growing method, and 0.83 with the watershed method.

Algorithms↗

The use of a priori information in the detection of mammographic microcalcifications to improve their classification.

In this work, we present a calcification-detection scheme that automatically localizes calcifications in a previously detected cluster in order to generate the input for a cluster-classification scheme developed in the past. The calcification-detection scheme makes use of three pieces of a priori information: the location of the center of the cluster, the size of the cluster, and the approximate number of calcifications in the cluster. This information can be obtained either automatically from a cluster-detection scheme or manually by a radiologist. It is used to analyze only the portion of the mammogram that contains a cluster and to identify the individual calcifications more accurately, after enhancing them by means of a "Difference of Gaussians" filter. Classification performances (patient-based Az=0.92; cluster-based Az=0.72) comparable to those obtained by using manually-identified calcifications (patient-based Az=0.92; cluster-based Az=0.82) can be achieved.

Algorithms↗

Independent versus sequential reading in ROC studies of computer-assist modalities: analysis of components of variance.

RATIONALE AND OBJECTIVES: The authors analyzed two methods for arranging the temporal sequencing of unaided versus computer-assisted reading in multiple-reader, multiple-case receiver operating characteristic studies of detection of solitary pulmonary nodules on chest radiographs. In the "independent" mode the readings are separated by about I month; in the "sequential" mode, the computer-assisted reading immediately follows the unassisted reading. MATERIALS AND METHODS: The authors used the method of Beiden, Wagner, and Campbell (BWC) to decompose the components of variance of receiver operating characteristic accuracy measures into those that are correlated and those that are uncorrelated across reading conditions. Only the latter contribute to uncertainty in estimates of the difference in accuracy measures across reading conditions (unaided vs aided). This method was used to analyze data from two independent studies of the detection of solitary pulmonary nodules on chest radiographs. RESULTS: In the sequential reading mode the components that were correlated across reading conditions increased compared to the independent reading mode, as might be expected. What was not anticipated was the fact that the total reader variance was approximately the same for the two reading modes. The results were remarkably similar across the two independent studies analyzed. CONCLUSION: The sequential reading mode may thus be the more sensitive probe of the difference between unassisted and computer-assisted reading, if the mean effect is unperturbed (as here). It is also the least demanding on the logistics and investment of reader time.

Analysis of Variance↗

A support vector machine approach for detection of microcalcifications.

In this paper, we investigate an approach based on support vector machines (SVMs) for detection of microcalcification (MC) clusters in digital mammograms, and propose a successive enhancement learning scheme for improved performance. SVM is a machine-learning method, based on the principle of structural risk minimization, which performs well when applied to data outside the training set. We formulate MC detection as a supervised-learning problem and apply SVM to develop the detection algorithm. We use the SVM to detect at each location in the image whether an MC is present or not. We tested the proposed method using a database of 76 clinical mammograms containing 1120 MCs. We use free-response receiver operating characteristic curves to evaluate detection performance, and compare the proposed algorithm with several existing methods. In our experiments, the proposed SVM framework outperformed all the other methods tested. In particular, a sensitivity as high as 94% was achieved by the SVM method at an error rate of one false-positive cluster per image. The ability of SVM to out perform several well-known methods developed for the widely studied problem of MC detection suggests that SVM is a promising technique for object detection in a medical imaging application.

Algorithms↗

Maximum likelihood fitting of FROC curves under an initial-detection-and-candidate-analysis model.

We have developed a model for FROC curve fitting that relates the observer's FROC performance not to the ROC performance that would be obtained if the observer's responses were scored on a per image basis, but rather to a hypothesized ROC performance that the observer would obtain in the task of classifying a set of "candidate detections" as positive or negative. We adopt the assumptions of the Bunch FROC model, namely that the observer's detections are all mutually independent, as well as assumptions qualitatively similar to, but different in nature from, those made by Chakraborty in his AFROC scoring methodology. Under the assumptions of our model, we show that the observer's FROC performance is a linearly scaled version of the candidate analysis ROC curve, where the scaling factors are just given by the FROC operating point coordinates for detecting initial candidates. Further, we show that the likelihood function of the model parameters given observational data takes on a simple form, and we develop a maximum likelihood method for fitting a FROC curve to this data. FROC and AFROC curves are produced for computer vision observer datasets and compared with the results of the AFROC scoring method. Although developed primarily with computer vision schemes in mind, we hope that the methodology presented here will prove worthy of further study in other applications as well.

Algorithms↗