PubMed Health⌕ Search

Biomedical subjects

Howard E Rockette

Publications and source records attributed to Howard E Rockette.

At least 19 recordsLinked to original sources

Tympanostomy tubes and developmental outcomes at 9 to 11 years of age.

BACKGROUND: Developmental impairments in children have been attributed to persistent middle-ear effusion in their early years of life. Previously, we reported that among children younger than 3 years of age with persistent middle-ear effusion, prompt as compared with delayed insertion of tympanostomy tubes did not result in improved cognitive, language, speech, or psychosocial development at 3, 4, or 6 years of age. However, other important components of development could not be assessed until the children were older. METHODS: We enrolled 6350 infants soon after birth and evaluated them regularly for middle-ear effusion. Before 3 years of age, 429 children with persistent effusion were randomly assigned to undergo the insertion of tympanostomy tubes either promptly or up to 9 months later if effusion persisted. We assessed literacy, attention, social skills, and academic achievement in 391 of these children at 9 to 11 years of age. RESULTS: Mean (+/-SD) scores on 48 developmental measures in the group of children who were assigned to undergo early insertion of tympanostomy tubes did not differ significantly from the scores in the group that was assigned to undergo delayed insertion. These measures included the Passage Comprehension subtest of the Woodcock Reading Mastery Tests (mean score, 98+/-12 in the early-treatment group and 99+/-12 in the delayed-treatment group); the Spelling, Writing Samples, and Calculation subtests of the Woodcock-Johnson III Tests of Achievement (96+/-13 and 97+/-16; 104+/-14 and 105+/-15; and 99+/-13 and 99+/-13, respectively); and inattention ratings on visual and auditory continuous performance tests. CONCLUSIONS: In otherwise healthy young children who have persistent middle-ear effusion, as defined in our study, prompt insertion of tympanostomy tubes does not improve developmental outcomes up to 9 to 11 years of age. (ClinicalTrials.gov number, NCT00365092 [ClinicalTrials.gov].).

Child↗

The prevalence effect in a laboratory environment: Changing the confidence ratings.

RATIONALE AND OBJECTIVES: We sought to assess whether or not prevalence levels affected the confidence ratings of readers during the interpretation of cases in a laboratory receiver operating characteristic-type observer performance study. MATERIALS AND METHODS: We reanalyzed a previously conducted observer performance study that included 14 readers and 5 different levels of prevalence. The previous study yielded the observation that in the laboratory we could not detect a "prevalence effect" in terms of differences in areas under the receiver operating characteristic curves. The detection ratings (for presence or absence) of lung nodules, interstitial disease, and pneumothorax for the five prevalence levels were compared, and a test for trend in averaged ratings as a function of abnormality prevalence was performed within a mixed-model setting that accounts for different sources of variability and correlations induced by the study design. RESULTS: The ratings of the cases in terms of confidence that the specific abnormality in question is present tend, on average, to be larger when actual disease prevalence is lower. The rate of the increase of the average confidence ratings with the decreasing prevalence of a specific abnormality is very similar for actually positive and actually negative cases for every considered abnormality. The observed trend in the changes of the average confidence ratings as a function of prevalence levels was statistically significant (p < 0.01). CONCLUSION: Expectations of disease prevalence in the case mix during a laboratory observer performance study may systematically affect the behavior of observers in terms of their actual confidence ratings.

Humans↗

Prospective study of electrical impedance scanning for identifying young women at risk for breast cancer.

BACKGROUND: One way to improve the cost-benefit ratio for breast cancer screening in younger women is to identify those at high-risk of breast cancer and manage them in an optimal manner. The purpose of this study is to evaluate the sensitivity and specificity of Electrical Impedance Scanning (EIS) for identifying young women who are at risk for having breast cancer and should be followed with directed imaging technologies. METHODS: A prospective, observational, two-arm, multi-site clinical trial was performed on women aged 30-45 years. The 'Sensitivity Arm' included Clinical Breast Examinations (CBE) and EIS (T-Scan 2000ED) on 189 women prior to scheduled breast biopsy. The 'Specificity Arm' included 1361 asymptomatic women visiting clinics for routine annual well-woman examination. Sensitivity and specificity were determined. Relative probability for a woman with a positive EIS examination was computed and compared with other approaches commonly used to define 'high-risk' in this population. RESULTS: Fifty of 189 women in the Sensitivity arm had verified cancers, 19 of whom had positive EIS examination resulting in sensitivity of 38% (19/50). Of the 1361 women in the Specificity arm, 67 had positive EIS examination resulting in a specificity of 95% (1294/1361). The relative probability of a woman with a positive EIS examination was 7.68, which compares favorably with other established risk identifiers (e.g. two first-degree relatives with breast cancer or atypical ductal hyperplasia). CONCLUSION: EIS may have an important role as a screening tool for identifying young women that should be followed more closely with advanced imaging technologies for early detection of breast cancer.

Adult↗

The effect of image display size on observer performance an assessment of variance components.

RATIONALE AND OBJECTIVE: Our goal was to investigate the effect of the displayed image size on variance components during the performance of an observer performance study to detect masses on abdominal computed tomography (CT) examinations. MATERIALS AND METHODS: A previously performed receiver operating characteristic (ROC) study with eight observers to detect abdominal masses on 166 CT examinations was reanalyzed to assess variance components when comparing two similar modes with displayed image sizes varying by a factor of 2. Case, mode, and reader-related variance components were estimated for the group of eight observers and subsets of readers after excluding each of the participants. RESULTS: There was no significant difference in the average area under the ROC curves between the two modes using the two image sizes (P > .05). Reader and reader-by-case variability were substantially larger for the mode displaying enlarged images for the group and all subsets formed by excluding a single reader. Reader variability was affected by one observer who actually performed better with the enlarged images. CONCLUSION: Sequential viewing of enlarged CT images for the detection of abdominal masses did not improve performance and increased reader variability.

Abdominal Neoplasms↗

A permutation test for comparing ROC curves in multireader studies a multi-reader ROC, permutation test.

RATIONALE AND OBJECTIVES: The aim of the study is to develop a permutation test to compare receiver operating characteristic (ROC) curves of two diagnostic modalities in a multireader paired design. MATERIALS AND METHODS: A statistical test for comparing two diagnostic modalities is developed based on all possible exchanges of the set of reader-ratings between the two modalities. An exact permutation test is formed by determining the frequency of the most extreme values of the statistic estimating the average difference in the areas under the ROC curves (AUCs). An asymptotic version of the test is constructed by obtaining the exact permutation variance and appealing to the asymptotic normality of the nonparametric estimator of the average difference in areas. Computer simulations were conducted to validate the type I error for small sample sizes. RESULTS: The new test provides a permutation approach for comparing ROC curves in a multireader paired-design setting in which effects of the readers are considered to be fixed. The type I error of the asymptotic test is close to the true value, even for samples as small as 20 normal and 20 abnormal cases. The test is designed to be sensitive to alternatives in which the AUCs of the two diagnostic modalities differ. CONCLUSIONS: The proposed test provides a powerful method for comparing two diagnostic modalities in a multireader paired-study design when the primary interest is to detect difference in average AUCs.

Analysis of Variance↗

Reader variance in ROC studies--generalizability to reader population at high and low performance levels.

RATIONALE AND OBJECTIVES: To investigate the variability between discriminative performances of readers as a function of average performance levels during receiver operating characteristic (ROC) studies. MATERIALS AND METHODS: Four subsets of cases from previously ascertained ROC rating data by 12 observers when detecting interstitial disease and pneumothorax on posteroanterior chest films were selected for each abnormality and reanalyzed to assess changes in "reader" variance component. The subsets were selected based on a prestudy subjective assessment of the subtleness of depicted abnormality (positive cases) and the difficulty in determining its absence (negative cases). Reader variance component was estimated using a bootstrap approach for each subset and the results were used to assess a general relationship between variability and average performance level. RESULTS: The reader variance component decreased substantially (from 0.007704 to 0.000426), as expected, when the areas under the ROC curves (AUC) for detecting pneumothoraces increased from 84% to 97%. On the other hand, reader variance component increased substantially (from 0.000890 to 0.005181) when AUC for detecting interstitial disease increased from 59% to 87%. The large magnitude of and changes in the reader variance component resulted in a consistent nonmonotone relationship as a function of AUC when other related variance components were included in addition to the reader component. CONCLUSION: Among several factors affecting generalizability of ROC results to the population of readers, the reader variance component depended nonmonotonically on the average diagnostic performance and is lowest at both very high and very low levels of performance.

Clinical Competence↗

Tympanometric findings and the probability of middle-ear effusion in 3686 infants and young children.

OBJECTIVE: We examined relationships between tympanometric findings and the presence or absence of middle-ear effusion in a population-based sample of children under the age of 3 years. METHODS: In a study of children's development in relation to early-life otitis media, we enrolled 6350 infants soon after birth and evaluated them regularly for the presence of middle-ear effusion. In 3686 of the children, we compared tympanometric findings with otoscopic diagnoses. We categorized tympanograms according to varying combinations of tympanometric peak height, peak pressure, and width, and calculated for each resulting category the percentage of the associated ears diagnosed as having effusion. Using these findings we developed algorithms for estimating the probability of middle-ear effusion associated with tympanograms of any configuration. RESULTS: For tympanograms generally, the lower their height and the greater their width, the greater was the probability of associated middle-ear effusion; the probability also was greater when peak pressure was negative rather than positive. Among children > or = 6 months of age, effusion was diagnosed in only 2.7% of ears with tympanometric height > or = 0.6 mL, but in 80.2% of ears with flat tympanograms. Relationships among younger infants were similar but less consistent. In both age groups, the tympanographic configurations most commonly encountered were associated with either a relatively low probability (<30%) or a relatively high probability (>70%) of the presence of middle-ear effusion. The receiver operating characteristic curve we generated using the algorithm we developed for children > or = 6 months of age gave an area under the curve of 0.84. The algorithm performed equally well when applied to a separate group of children, suggesting that it is generalizable to other unselected populations. CONCLUSIONS: The present report offers two alternative methods for estimating the probability of middle-ear effusion in children aged 6 through 35 months, given any combination of tympanometric values.

Acoustic Impedance Tests↗

A permutation test sensitive to differences in areas for comparing ROC curves from a paired design.

The area under the receiver operating characteristic (ROC) curve (AUC) is a widely accepted summary index of the overall performance of diagnostic procedures and the difference between AUCs is often used when comparing two diagnostic systems. We developed an exact non-parametric statistical procedure for comparing two ROC curves in paired design settings. The test which is based on all permutations of the subject specific rank ratings is formally a test for equality of ROC curves that is sensitive to the alternatives of AUC difference. The operating characteristics of the proposed test were evaluated using extensive simulations over a wide range of parameters. The proposed procedure can be easily implemented in experimental ROC data sets. For small samples and for underlying parameters that are common in experimental studies in diagnostic imaging the test possesses good operating characteristics and is more powerful than the conventional non-parametric procedure for AUC comparisons. We also derived an asymptotic version of the test which uses an exact estimate of the variance in the permutation space and provides a good approximation even when the sample sizes are small. This asymptotic procedure is a simple and precise approximation to the exact test and is useful for large sample sizes where the exact test may be computationally burdensome.

Biometry↗

Developmental outcomes after early or delayed insertion of tympanostomy tubes.

BACKGROUND: To prevent later developmental impairments, myringotomy with the insertion of tympanostomy tubes has often been undertaken in young children who have persistent otitis media with effusion. We previously reported that prompt as compared with delayed insertion of tympanostomy tubes in children with persistent effusion who were younger than three years of age did not result in improved developmental outcomes at three or four years of age. However, the effect on the outcomes of school-age children is unknown. METHODS: We enrolled 6350 healthy infants younger than 62 days of age and evaluated them regularly for middle-ear effusion. Before three years of age, 429 children with persistent middle-ear effusion were randomly assigned to have tympanostomy tubes inserted either promptly or up to nine months later if effusion persisted. We assessed developmental outcomes in 395 of these children at six years of age. RESULTS: At six years of age, 85 percent of children in the early-treatment group and 41 percent in the delayed-treatment group had received tympanostomy tubes. There were no significant differences in mean (+/-SD) scores favoring early versus delayed treatment on any of 30 measures, including the Wechsler Full-Scale Intelligence Quotient (98+/-13 vs. 98+/-14); Number of Different Words test, a measure of word diversity (183+/-36 vs. 175+/-36); Percentage of Consonants Correct-Revised test, a measure of speech-sound production (96+/-2 vs. 96+/-3); the SCAN test, a measure of central auditory processing (95+/-15 vs. 96+/-14); and several measures of behavior and emotion. CONCLUSIONS: In otherwise healthy children younger than three years of age who have persistent middle-ear effusion within the duration of effusion that we studied, prompt insertion of tympanostomy tubes does not improve developmental outcomes at six years of age.

Child↗

A conditional nonparametric test for comparing two areas under the ROC curves from a paired design.

RATIONALE AND OBJECTIVES: To develop a conditional nonparametric procedure for comparing two correlated areas under receiver operating characteristic (ROC) curves (AUC). MATERIALS AND METHODS: A nonparametric conditional test to compare areas under two ROC curves was developed using the distribution of the elements of the nonparametric AUC estimators in a permutation space. The conditioning is made on the observed discordances between the relative orderings of ratings of the normal and abnormal cases for the two modalities taken over all possible pairs. The type I error of the procedure was verified using computer simulations. The power of the test was compared with an existing unconditional procedure on simulated datasets from binormal distributions as well as from a mixture of binormal distributions of ratings. RESULTS: The proposed test is conservative for low sample sizes, large AUC, and high correlation between modalities. It possesses a reasonable type I error for sample sizes as low as 20 actually positive and 20 actually negative cases. In plausible situations in which the sample in observer performance studies can not be monotonically transformed into a binormal distribution, this approach may have modest power advantages over the conventional nonparametric test. CONCLUSION: The conditional nonparametric test presented here is an alternative approach to existing unconditional procedures and may offer advantages in certain types of observer performance studies.

Area Under Curve↗

Incorporating utility-weights when comparing two diagnostic systems: a preliminary assessment.

RATIONALE AND OBJECTIVES: We sought to develop a new index that incorporates utility-weights when assessing the overall performance of a diagnostic system and to provide a statistical test for comparing two indices in a paired study design. MATERIALS AND METHODS: The area under the receiver operating characteristic (ROC) curve (AUC) was used as the basis for constructing a new index. The index we propose represents a weighted average of class-specific AUCs each of which relates to a class of pairs of actually negative (normal) and actually positive (abnormal) cases with a specific predetermined utility (or clinical importance). For each pair of normal-abnormal cases, the utility is defined a priori and based on external (covariate) information. In the proposed approach utility-weights represent the relative importance (utility) of discriminating between different types of normal and abnormal cases (pairs of the same type are combined in the classes termed utility-classes). We also describe a simple nonparametric procedure for comparing the proposed indices as computed from paired data. Computer simulations were conducted to evaluate the behavior of the type I error of the proposed test in the simple albeit important instance of two utility-classes. RESULTS: The new index provides an extension of the commonly used area under the ROC curve. It allows for incorporation of utility-weights into the analysis and reduces to the conventional AUC index when all assigned utility-weights are equal to unity. Computer simulations indicate that in the considered scenario of two utility-classes, the type I error of the proposed test is comparable to that of the conventional nonparametric test for equality of AUC indices. CONCLUSIONS: The proposed index and the statistical test provide a practical approach of incorporating utilities when comparing diagnostic systems.

Algorithms↗

Variability in observer performance studies experimental observations.

RATIONALE AND OBJECTIVES: The aim of the study is to assess variance components in observer performance studies and the possible impact on study results and conclusions. MATERIALS AND METHODS: Two previously performed retrospective receiver operating characteristic-type observer performance studies to evaluate the performance of seven radiologists in detecting interstitial disease on conventional posteroanterior chest films and nine radiologists in detecting interstitial disease on a high-resolution workstation were reanalyzed by using the Beiden, Wagner, and Campbell nine-component model to estimate the different variance components. We estimated case-, reader-, and mode-related components of the variance for the group as a whole and after excluding (round robin) each reader. Overall variance was evaluated, and the effect of individual readers on overall study conclusions was assessed. RESULTS: Overall results and conclusions of the reanalysis agreed with the original one in that, as a group, radiologists performed significantly better when using conventional films (P < .05) in both studies. Reader variability was large compared with all other components, and in one study, it was substantially larger for the workstation reading mode. Reader variability was affected substantially by one observer in each study, and in one study, reader-by-mode variability was affected by another reader who performed better on the workstation. CONCLUSION: Estimates of variance components can shed light on the appropriateness of study design, as well as the sensitivity of results to the inclusion (or exclusion) of individual observers.

Humans↗

Water precautions and tympanostomy tubes: a randomized, controlled trial.

OBJECTIVES/HYPOTHESIS: The objective was to determine whether there is an increased incidence of otorrhea in young children with tympanostomy tubes who swim and bathe without water precautions as compared with children who use water precautions in the form of ear plugs. STUDY DESIGN: Prospective, randomized, investigator-blinded, controlled trial. METHODS: Two hundred one children (age range, 6 mo-6 y) who had undergone bilateral myringotomy and tube insertion were randomly assigned into one of two groups: swimming and bathing with or without ear plugs. Children were seen monthly for 1 year and whenever there was intercurrent otorrhea. RESULTS: Ninety children with and 82 children without ear plugs returned for at least one follow-up visit. Mean (SD) duration of follow-up was 9.4 (4.1) months for the children with ear plugs and 9.1 (4.4) months for the children without ear plugs. Forty-two children (47%) who wore ear plugs developed at least one episode of otorrhea, as compared with 46 (56%) who did not use ear plugs (logistic regression adjusting for stratification variables, P = .21). The mean (SD) rate of otorrhea per month was 0.07 (0.31) for the children who wore ear plugs as compared with 0.10 (0.31) for the children who did not wear ear plugs (Poisson regression adjusting for stratification variables, P = .05). CONCLUSION: There is a small but statistically significant increase in the rate of otorrhea in young children who swim and bathe without the use of ear plugs as compared with children who use ear plugs. Because the clinical impact of using ear plugs is small, their routine use may be unnecessary.

Baths↗

A comparison of two data analyses from two observer performance studies using Jackknife ROC and JAFROC.

The authors compared two methodological approaches, Jackknife ROC and JAFROC, in analyzing data ascertained during FROC (free-response receiver operating characteristics) type studies. Observer rating data obtained from two observer performance studies were analyzed. During the first study, seven radiologists interpreted 120 mammography examinations depicting 57 masses under five different conditions with and without the results of computer-aided detection (CAD). In the second study, eight radiologists interpreted 110 examinations depicting 51 masses under six different display conditions with and without CAD results. Readers rated the detection task in a FROC type response. Jackknife ROC (using the software of LABMRMC with the highest rating per case) and JAFROC were used to compute differences, if any, in summary performance levels among all reading modes in each study as well as for all paired data sets. The results of the different analytical approaches are compared. The overall results for all modes were significantly different for the first study (p < 0.05) and not significant (p > 0.05) for the second study using either analytical approach. In the first study, the performance levels represented by three paired data sets were significantly different (p < 0.05) when computed using LABMRMC and four pairs were significantly different (p < 0.05) using JAFROC. In eight of ten pairs, JAFROC produced lower p values than LABMRMC. In the second study, LABMRMC showed no significant differences for any paired data sets and JAFROC showed a significant difference for one pair. In 15 of 16 pairs, p values computed by JAFROC were lower than those computed by LABMRMC.

Breast Neoplasms↗

Computer-aided detection performance in mammographic examination of masses: assessment.

PURPOSE: To compare performance of two computer-aided detection (CAD) systems and an in-house scheme applied to five groups of sequentially acquired screening mammograms. MATERIALS AND METHODS: Two hundred nineteen film-based mammographic examinations, classified into five groups, were included in this study. Group 1 included 58 examinations in which verified malignant masses were detected during screening; group 2, 39 in which all available latest examinations were performed prior to diagnosis of these malignant masses (subset of 39 women from group 1); group 3, 22 in which findings were interpreted as negative but were verified as cancer within 1 year from the negative interpretation (missed cancers); group 4, 50 in which findings were negative and patients were not recalled for additional procedures; and group 5, 50 in which patients were recalled for additional procedures and findings were negative for cancer. In all examinations, images were processed with two Food and Drug Administration-approved commercially available CAD systems and an in-house scheme. Performance levels in terms of true-positive detection rates and number of false-positive identifications per image and per examination were compared. RESULTS: Mass detection rates in positive examinations (group 1) were 67%-72%. Detection rates among three systems were not significantly different (P > .05). In 50 negative screening examinations (group 4), false-positive rates ranged from 1.08 to 1.68 per four-view examination. Performance level differences among systems were significant for false-positive rates (P = .008). Performance of all systems was at levels lower than publicly suggested in some retrospective studies. False-positive CAD cueing rates were significantly higher for negative examinations in which patients were recalled (group 5) than they were for those in which patients were not recalled (group 4) (P < or = .002). CONCLUSION: Performance of CAD systems for mass detection at mammography varies significantly, depending on examination and system used. Actual performance of all systems in clinical environment can be improved.

Adult↗

Recall and detection rates in screening mammography.

BACKGROUND: The authors investigated the correlation between recall and detection rates in a group of 10 radiologists who had read a high volume of screening mammograms in an academic institution. METHODS: Practice-related and outcome-related databases of verified cases were used to compute recall rates and tumor detection rates for a group of 10 Mammography Quality Standard Act (MQSA)-certified radiologists who interpreted a total of 98,668 screening mammograms during the years 2000, 2001, and 2002. The relation between recall and detection rates for these individuals was investigated using parametric Pearson (r) and nonparametric Spearman (rho) correlation coefficients. The effect of the volume of mammograms interpreted by individual radiologists was assessed using partial correlations controlling for total reading volumes. RESULTS: A wide variability of recall rates (range, 7.7-17.2%) and detection rates (range, 2.6-5.4 per 1000 mammograms) was observed in the current study. A statistically significant correlation (P < 0.05) between recall and detection rates was observed in this group of 10 experienced radiologists. The results remained significant (P < 0.05) after accounting for the volume of mammograms interpreted by each radiologist. CONCLUSIONS: Optimal performance in screening mammography should be evaluated quantitatively. The general pressure to reduce recall rates through "practice guidelines" to below a fixed level for all radiologists should be assessed carefully.

Breast Neoplasms↗

Changes in breast cancer detection and mammography recall rates after the introduction of a computer-aided detection system.

BACKGROUND: Computer-aided mammography is rapidly gaining clinical acceptance, but few data demonstrate its actual benefit in the clinical environment. We assessed changes in mammography recall and cancer detection rates after the introduction of a computer-aided detection system into a clinical radiology practice in an academic setting. METHODS: We used verified practice- and outcome-related databases to compute recall rates and cancer detection rates for 24 Mammography Quality Standards Act-certified academic radiologists in our practice who interpreted 115,571 screening mammograms with (n = 59,139) or without (n = 56,432) the use of a computer-aided detection system. All statistical tests were two-sided. RESULTS: For the entire group of 24 radiologists, recall rates were similar for mammograms interpreted without and with computer-aided detection (11.39% versus 11.40%; percent difference = 0.09, 95% confidence interval [CI] = -11 to 11; P =.96) as were the breast cancer detection rates for mammograms interpreted without and with computer-aided detection (3.49% versus 3.55% per 1000 screening examinations; percent difference = 1.7, 95% CI = -11 to 19; P =.68). For the seven high-volume radiologists (i.e., those who interpreted more than 8000 screening mammograms each over a 3-year period), the recall rates were similar for mammograms interpreted without and with computer-aided detection (11.62% versus 11.05%; percent difference = -4.9, 95% CI = -21 to 4; P =.16), as were the breast cancer detection rates for mammograms interpreted without and with computer-aided detection (3.61% versus 3.49% per 1000 screening examinations; percent difference = -3.2, 95% CI = -15 to 9; P =.54). CONCLUSION: The introduction of computer-aided detection into this practice was not associated with statistically significant changes in recall and breast cancer detection rates, both for the entire group of radiologists and for the subset of radiologists who interpreted high volumes of mammograms.

Breast Neoplasms↗