PubMed Health⌕ Search

Biomedical subjects

Keith E Muller

Publications and source records attributed to Keith E Muller.

11 recordsLinked to original sources

Comparison of calcification specificity in digital mammography using soft-copy display versus screen-film mammography.

OBJECTIVE: The purpose of this study was to compare specificity in the interpretation of calcifications in soft-copy reviewing of digital mammograms versus hard-copy reviewing of screen-film mammograms. MATERIALS AND METHODS: A total of 130 consecutive cases with calcifications (44 malignant and 86 benign) that had been evaluated with needle or surgical biopsy were collected. Both screen-film mammography and soft-copy digital mammography were obtained in the same patients under existing research protocols using Fischer Imaging's SenoScan (n = 71), Lorad's digital mammography system (n = 35), and GE Healthcare's Senographe 2000D (n = 24). Eight trained radiologists scored all lesions--cropped or masked to display just the region of interest--both on screen-film and soft-copy digital mammography with a month between reviews to reduce the effects of learning and memory. A 5-point malignancy scale was used, with 1 as definitely not, 2 as probably not, 3 as possibly, 4 as probably, and 5 as definitely. Reviewers were randomly assigned condition order, and images within each condition were randomly ordered. Repeated measures analysis of variance was used to test for differences between conditions in specificity computed via nonparametric receiver operating characteristic (ROC) study separately for each reviewer and condition. RESULTS: Across all reviewers, the mean specificity for 1 or 2 versus 3, 4, or 5 was 0.803 for screen-film mammography (range, 0.413-0.938; SD +/- 0.166) and 0.833 for soft-copy image (range, 0.375-0.951; SD +/- 0.187). Although not statistically significant (Student's t test p values from 0.19 to 0.99 across all cut points), numeric values of specificity were consistently higher for soft-copy versus screen-film mammography. No statistical significance in specificity was seen using all possible cut points in the 5-point scale, although the primary analysis used the cutpoint for differentiation between benign and malignant cases as 1 or 2 versus 3, 4, or 5. CONCLUSION: No statistically significant difference was shown in specificity achievable using soft-copy digital versus screen-film mammography in this study.

Biopsy↗

Assessment of real-time 3D visualization for cardiothoracic diagnostic evaluation and surgery planning.

RATIONALE AND OBJECTIVES: Three-dimensional (3D) real-time volume rendering has demonstrated improvements in clinical care for several areas of radiological imaging. We test whether advanced real-time rendering techniques combined with an effective user interface will allow radiologists and surgeons to improve their performance for cardiothoracic surgery planning and diagnostic evaluation. MATERIAL AND METHODS: An interactive combination 3D and 2D visualization system developed at the University of North Carolina at Chapel Hill was compared against standard tiled 2D slice presentation on a viewbox. The system was evaluated for 23 complex cardiothoracic computed tomographic (CT) cases including heart-lung and lung transplantation, tumor resection, airway stent placement, repair of congenital heart defects, aortic aneurysm repair, and resection of pulmonary arteriovenous malformation. Radiologists and surgeons recorded their impressions with and without the use of the interactive visualization system. RESULTS: The cardiothoracic surgeons reported positive benefits to using the 3D visualizations. The addition of the 3D visualization changed the surgical plan (65% of cases), increased the surgeon's confidence (on average 40% per case), and correlated well with the anatomy found at surgery (95% of cases). The radiologists reported fewer and less major changes than the surgeons in their understanding of the case due to the 3D visualization. They found new findings or additional information about existing findings in 66% of the cases; however, they changed their radiology report in only 14% of the cases. CONCLUSION: With the appropriate choice of 3D real-time volume rendering and a well-designed user interface, both surgeons and radiologists benefit from viewing an interactive 3D visualization in addition to 2D images for surgery planning and diagnostic evaluation of complex cardiothoracic cases. This study finds that 3D visualization is especially helpful to the surgeon in understanding the case, and in communicating and planning the surgery. These results suggest that including real-time 3D visualization would be of clinical benefit for complex cardiothoracic CT cases.

Confidence Intervals↗

Analyzing attributes of vessel populations.

Almost all diseases affect blood vessel attributes (vessel number, radius, tortuosity, and branching pattern). Quantitative measurement of vessel attributes over relevant vessel populations could thus provide an important means of diagnosing and staging disease. Unfortunately, little is known about the statistical properties of vessel attributes. In particular, it is unclear whether vessel attributes fit a Gaussian distribution, how dependent these values are upon anatomical location, and how best to represent the attribute values of the multiple vessels comprising a population of interest in a single patient. The purpose of this report is to explore the distributions of several vessel attributes over vessel populations located in different parts of the head. In 13 healthy subjects, we extract vessels from MRA data, define vessel trees comprising the anterior cerebral, right and left middle cerebral, and posterior cerebral circulations, and, for each of these four populations, analyze the vessel number, average radius, branching frequency, and tortuosity. For the parameters analyzed, we conclude that statistical methods employing summary measures for each attribute within each region of interest for each patient are preferable to methods that deal with individual vessels, that the distributions of the summary measures are indeed Gaussian, and that attribute values may differ by anatomical location. These results should be useful in designing studies that compare patients with suspected disease to a database of healthy subjects and are relevant to groups interested in atlas formation and in the statistics of tubular objects.

Cerebral Arteries↗

A parametric model for studying organism fitness using step-stress experiments.

We propose a method based on parametric survival analysis to analyze step-stress data. Step-stress studies are failure time studies in which the experimental stressor is increased at specified time intervals. While this protocol has been frequently employed in industrial reliability studies, it is less common in the life sciences. Possible biological applications include experiments on swimming performance of fish using a step function defining increasing water velocity over time, and treadmill tests on humans. A likelihood-ratio test is developed for comparing the failure times in two groups based on a piecewise constant hazard assumption. The test can be extended to other piecewise distributions and to include covariates. An example data set is used to illustrate the method and highlight experimental design issues. A small simulation study compares this analysis procedure to currently used methods with regard to type I error rate and power.

Analysis of Variance↗

Adjusting power for a baseline covariate in linear models.

The analysis of covariance provides a common approach to adjusting for a baseline covariate in medical research. With Gaussian errors, adding random covariates does not change either the theory or the computations of general linear model data analysis. However, adding random covariates does change the theory and computation of power analysis. Many data analysts fail to fully account for this complication in planning a study. We present our results in five parts. (i) A review of published results helps document the importance of the problem and the limitations of available methods. (ii) A taxonomy for general linear multivariate models and hypotheses allows identifying a particular problem. (iii) We describe how random covariates introduce the need to consider quantiles and conditional values of power. (iv) We provide new exact and approximate methods for power analysis of a range of multivariate models with a Gaussian baseline covariate, for both small and large samples. The new results apply to the Hotelling-Lawley test and the four tests in the "univariate" approach to repeated measures (unadjusted, Huynh-Feldt, Geisser-Greenhouse, Box). The techniques allow rapid calculation and an interactive, graphical approach to sample size choice. (v) Calculating power for a clinical trial of a treatment for increasing bone density illustrates the new methods. We particularly recommend using quantile power with a new Satterthwaite-style approximation.

Bone Density↗

Properties of internal pilots with the univariate approach to repeated measures.

Uncertainty surrounding the error covariance matrix often presents the biggest barrier to achieving accurate power analysis in the 'univariate' approach to repeated measures analysis of variance (UNIREP). A poor choice gives either an overpowered study which wastes resources, or an underpowered study with little chance of success. Internal pilot designs were introduced to resolve such uncertainty about error variance for t-tests. In earlier papers, we extended the use of internal pilots to any univariate linear model with fixed predictors and independent Gaussian errors. Here we further extend our exact and approximate results to UNIREP analysis. For a fixed treatment effect, the inaccuracy in a power calculation depends only on the ratio of the true variance to the value used for planning. The greater complexity of repeated measures requires generalizing misspecification of error variance to the misspecification of the eigenvalues of the error covariance. We recommend approximating the misspecification in terms of the first and second moments of the eigenvalues, for both fixed sample and internal pilot designs. We also describe an unadjusted approach for internal pilots with repeated measures. Simulations illustrate the fact that both positive and negative properties in the univariate setting extend to repeated measures analysis. In particular, internal pilots allow maintaining power or reducing expected sample size when the covariance matrix used for planning differs from the true value. However, an unadjusted approach can inflate test size, at least with small to moderate sample sizes. Hence new, adjusted methods must be developed for small samples. At this time, we caution against using an internal pilot design with repeated measures without first conducting simulations to document the amount of test size inflation possible for the conditions of interest.

Analysis of Variance↗

A new method for choosing sample size for confidence interval-based inferences.

Scientists often need to test hypotheses and construct corresponding confidence intervals. In designing a study to test a particular null hypothesis, traditional methods lead to a sample size large enough to provide sufficient statistical power. In contrast, traditional methods based on constructing a confidence interval lead to a sample size likely to control the width of the interval. With either approach, a sample size so large as to waste resources or introduce ethical concerns is undesirable. This work was motivated by the concern that existing sample size methods often make it difficult for scientists to achieve their actual goals. We focus on situations which involve a fixed, unknown scalar parameter representing the true state of nature. The width of the confidence interval is defined as the difference between the (random) upper and lower bounds. An event width is said to occur if the observed confidence interval width is less than a fixed constant chosen a priori. An event validity is said to occur if the parameter of interest is contained between the observed upper and lower confidence interval bounds. An event rejection is said to occur if the confidence interval excludes the null value of the parameter. In our opinion, scientists often implicitly seek to have all three occur: width, validity, and rejection. New results illustrate that neglecting rejection or width (and less so validity) often provides a sample size with a low probability of the simultaneous occurrence of all three events. We recommend considering all three events simultaneously when choosing a criterion for determining a sample size. We provide new theoretical results for any scalar (mean) parameter in a general linear model with Gaussian errors and fixed predictors. Convenient computational forms are included, as well as numerical examples to illustrate our methods.

Biometry↗

Diagnostic accuracy of digital mammography in patients with dense breasts who underwent problem-solving mammography: effects of image processing and lesion type.

PURPOSE: To determine effects of lesion type (calcification vs mass) and image processing on radiologist's performance for area under the receiver operating characteristic curve (AUC), sensitivity, and specificity for detection of masses and calcifications with digital mammography in women with mammographically dense breasts. MATERIALS AND METHODS: This study included 201 women who underwent digital mammography at seven U.S. and Canadian medical centers. Three image-processing algorithms were applied to the digital images, which were acquired with Fischer, General Electric, and Lorad digital mammography units. Eighteen readers participated in the reader study (six readers per algorithm). Baseline values for reader performance with screen-film mammograms were obtained through the additional interpretation of 179 screen-film mammograms. A repeated-measures analysis of covariance allowing unequal slopes was used in each of the nine analyses (AUC, sensitivity, and specificity for each of three machines). Bonferroni correction was used. RESULTS: Although lesion type did not affect the AUC or sensitivity for Fischer digital images, it did affect specificity (P =.0004). For the General Electric digital images, AUC, sensitivity, and specificity were not affected by lesion type. For Lorad digital images, the results strongly suggested that lesion type affected AUC and sensitivity (P <.0001). None of the three image-processing methods tested affected the AUC, sensitivity, or specificity for the Fischer, General Electric, or Lorad digital images. CONCLUSION: Findings in this study indicate that radiologist's interpretation accuracy in interpreting digital mammograms depends on lesion type. Interpretation accuracy was not influenced by the image-processing method.

Area Under Curve↗

Thresholds for human detection of patient setup errors in digitally reconstructed portal images of prostate fields.

PURPOSE: Computer-assisted methods to analyze electronic portal images for the presence of treatment setup errors should be studied in controlled experiments before use in the clinical setting. Validation experiments using images that contain known errors usually report the smallest errors that can be detected by the image analysis algorithm. This paper offers human error-detection thresholds as one benchmark for evaluating the smallest errors detected by algorithms. Unfortunately, reliable data are lacking describing human performance. The most rigorous benchmarks for human performance are obtained under conditions that favor error detection. To establish such benchmarks, controlled observer studies were carried out to determine the thresholds of detectability for in-plane and out-of-plane translation and rotation setup errors introduced into digitally reconstructed portal radiographs (DRPRs) of prostate fields. METHODS AND MATERIALS: Seventeen observers comprising radiation oncologists, radiation oncology residents, physicists, and therapy students participated in a two-alternative forced choice experiment involving 378 DRPRs computed using the National Library of Medicine Visible Human data sets. An observer viewed three images at a time displayed on adjacent computer monitors. Each image triplet included a reference digitally reconstructed radiograph displayed on the central monitor and two DRPRs displayed on the flanking monitors. One DRPR was error free. The other DRPR contained a known in-plane or out-of-plane error in the placement of the treatment field over a target region in the pelvis. The range for each type of error was determined from pilot observer studies based on a Probit model for error detection. The smallest errors approached the limit of human visual capability. The observer was told what kind of error was introduced, and was asked to choose the DRPR that contained the error. Observer decisions were recorded and analyzed using repeated-measures analysis of variance. RESULTS: The thresholds of detectability averaged over all observers were approximately 2.5 mm for in-plane translations, 1.6 degrees for in-plane rotations, 1 degrees for out-of-plane rotations, and 8% change in magnification for out-of-plane translations along the central axis. When one inexperienced observer is excluded, the average threshold for change in magnification is 5%. Experienced observers tended to perform better, but differences between groups were not statistically significant. Thresholds were computed as averages over all observers. Because of the broad range of observer capabilities, some detection tasks were too difficult for some observers, leading to missing threshold values in our data analysis. The missing values were excluded from computation of the average thresholds reported above. The effect of the missing values is to bias the average values toward the best human performance. CONCLUSIONS: Under favorable conditions, humans can detect small errors in setup geometry. The thresholds for error detection reported in this study are believed to represent rigorous but reasonable benchmarks that can be incorporated into studies evaluating algorithms for computer-assisted detection of setup errors in electronic portal images.

Diagnostic Errors↗

Improved approximate confidence intervals for the mean of a log-normal random variable.

Data analysts often compute approximate 100 (1-alpha) per cent confidence intervals for the mean of a log-normal random variable due to the computational effort required for exact intervals. We evaluate two simple approximations and demonstrate that the probabilities with which the intervals fail to capture the population mean (that is, the coverage error) can range from well above the desired level, alpha, to very near zero in small to moderate sample sizes (n < or = 100). The performance of a more sophisticated approximation, implemented via numerical integration or bootstrap sampling, is noticeably improved, but also suffers from coverage errors that are too large when n < or = 25. A new procedure is developed which outperforms existing approximations. Computing these improved intervals requires the integration of standard distribution functions. The calculations are straightforward, however, and lead to satisfactory coverage errors for n as small as 5. A related method that avoids the integration step generally outperforms existing simple approximations for n < or = 100, while maintaining the coverage error at or below alpha. Programs to implement the new procedures are provided in an Appendix.

Benzene↗

Interpretation of digital mammograms: comparison of speed and accuracy of soft-copy versus printed-film display.

PURPOSE: To compare the speed and accuracy of the interpretations of digital mammograms by radiologists by using printed-film versus soft-copy display. MATERIALS AND METHODS: After being trained in interpretation of digital mammograms, eight radiologists interpreted 63 digital mammograms, all with old studies for comparison. All studies were interpreted by all readers in soft-copy and printed-film display, with interpretations of images in the same cases at least 1 month apart. Mammograms were interpreted in cases that included six biopsy-proved cancers and 20 biopsy-proved benign lesions, 20 cases of probably benign findings in patients who underwent 6-month follow-up, and 17 cases without apparent findings. Area under the receiver operating characteristic curve (A(z)), sensitivity, and specificity were calculated for soft-copy and printed-film display. RESULTS: There was no significant difference in the speed of interpretation, but interpretations with soft-copy display were slightly faster. The differences in A(z), sensitivity, and specificity were not significantly different; A(z) and sensitivity were slightly better for interpretations with printed film, and specificity was slightly better for interpretations with soft copy. CONCLUSION: Interpretation with soft-copy display is likely to be useful with digital mammography and is unlikely to significantly change accuracy or speed.

Breast Diseases↗