PubMed Health⌕ Search

Biomedical subjects

Stuart G Baker

Publications and source records attributed to Stuart G Baker.

11 recordsLinked to original sources

A perfect correlate does not a surrogate make.

BACKGROUND: There is common belief among some medical researchers that if a potential surrogate endpoint is highly correlated with a true endpoint, then a positive (or negative) difference in potential surrogate endpoints between randomization groups would imply a positive (or negative) difference in unobserved true endpoints between randomization groups. We investigate this belief when the potential surrogate and unobserved true endpoints are perfectly correlated within each randomization group. METHODS: We use a graphical approach. The vertical axis is the unobserved true endpoint and the horizontal axis is the potential surrogate endpoint. Perfect correlation within each randomization group implies that, for each randomization group, potential surrogate and true endpoints are related by a straight line. In this scenario the investigator does not know the slopes or intercepts. We consider a plausible example where the slope of the line is higher for the experimental group than for the control group. RESULTS: In our example with unknown lines, a decrease in mean potential surrogate endpoints from control to experimental groups corresponds to an increase in mean true endpoint from control to experimental groups. Thus the potential surrogate endpoints give the wrong inference. Similar results hold for binary potential surrogate and true outcomes (although the notion of correlation does not apply). The potential surrogate endpoint would give the correct inference if either (i) the unknown lines for the two group coincided, which means that the distribution of true endpoint conditional on potential surrogate endpoint does not depend on treatment group, which is called the Prentice Criterion or (ii) if one could accurately predict the lines based on data from prior studies. CONCLUSION: Perfect correlation between potential surrogate and unobserved true outcomes within randomized groups does not guarantee correct inference based on a potential surrogate endpoint. Even in early phase trials, investigators should not base conclusions on potential surrogate endpoints in which the only validation is high correlation with the true endpoint within a group.

Endpoint Determination↗

Estimating the cumulative risk of false positive cancer screenings.

BACKGROUND: When evaluating cancer screening it is important to estimate the cumulative risk of false positives from periodic screening. Because the data typically come from studies in which the number of screenings varies by subject, estimation must take into account dropouts. A previous approach to estimate the probability of at least one false positive in n screenings unrealistically assumed that the probability of dropout does not depend on prior false positives. METHOD: By redefining the random variables, we obviate the unrealistic dropout assumption. We also propose a relatively simple logistic regression and extend estimation to the expected number of false positives in n screenings. RESULTS: We illustrate our methodology using data from women ages 40 to 64 who received up to four annual breast cancer screenings in the Health Insurance Program of Greater New York study, which began in 1963. Covariates were age, time since previous screening, screening number, and whether or not a previous false positive occurred. Defining a false positive as an unnecessary biopsy, the only statistically significant covariate was whether or not a previous false positive occurred. Because the effect of screening number was not statistically significant, extrapolation beyond 4 screenings was reasonable. The estimated mean number of unnecessary biopsies in 10 years per woman screened is.11 with 95% confidence interval of (.10,.12). Defining a false positive as an unnecessary work-up, all the covariates were statistically significant and the estimated mean number of unnecessary work-ups in 4 years per woman screened is.34 with 95% confidence interval (.32,.36). CONCLUSION: Using data from multiple cancer screenings with dropouts, and allowing dropout to depend on previous history of false positives, we propose a logistic regression model to estimate both the probability of at least one false positive and the expected number of false positives associated with n cancer screenings. The methodology can be used for both informed decision making at the individual level, as well as planning of health services.

Adult↗

Randomized trials, generalizability, and meta-analysis: graphical insights for binary outcomes.

BACKGROUND: Randomized trials stochastically answer the question. "What would be the effect of treatment on outcome if one turned back the clock and switched treatments in the given population?" Generalizations to other subjects are reliable only if the particular trial is performed on a random sample of the target population. By considering an unobserved binary variable, we graphically investigate how randomized trials can also stochastically answer the question, "What would be the effect of treatment on outcome in a population with a possibly different distribution of an unobserved binary baseline variable that does not interact with treatment in its effect on outcome?" METHOD: For three different outcome measures, absolute difference (DIF), relative risk (RR), and odds ratio (OR), we constructed a modified BK-Plot under the assumption that treatment has the same effect on outcome if either all or no subjects had a given level of the unobserved binary variable. (A BK-Plot shows the effect of an unobserved binary covariate on a binary outcome in two treatment groups; it was originally developed to explain Simpsons's paradox.) RESULTS: For DIF and RR, but not OR, the BK-Plot shows that the estimated treatment effect is invariant to the fraction of subjects with an unobserved binary variable at a given level. CONCLUSION: The BK-Plot provides a simple method to understand generalizability in randomized trials. Meta-analyses of randomized trials with a binary outcome that are based on DIF or RR, but not OR, will avoid bias from an unobserved covariate that does not interact with treatment in its effect on outcome.

Data Interpretation, Statistical↗

A simple method for analyzing data from a randomized trial with a missing binary outcome.

BACKGROUND: Many randomized trials involve missing binary outcomes. Although many previous adjustments for missing binary outcomes have been proposed, none of these makes explicit use of randomization to bound the bias when the data are not missing at random. METHODS: We propose a novel approach that uses the randomization distribution to compute the anticipated maximum bias when missing at random does not hold due to an unobserved binary covariate (implying that missingness depends on outcome and treatment group). The anticipated maximum bias equals the product of two factors: (a) the anticipated maximum bias if there were complete confounding of the unobserved covariate with treatment group among subjects with an observed outcome and (b) an upper bound factor that depends only on the fraction missing in each randomization group. If less than 15% of subjects are missing in each group, the upper bound factor is less than.18. RESULTS: We illustrated the methodology using data from the Polyp Prevention Trial. We anticipated a maximum bias under complete confounding of.25. With only 7% and 9% missing in each arm, the upper bound factor, after adjusting for age and sex, was.10. The anticipated maximum bias of.25 x.10 =.025 would not have affected the conclusion of no treatment effect. CONCLUSION: This approach is easy to implement and is particularly informative when less than 15% of subjects are missing in each arm.

Adenoma↗

Using observational data to estimate an upper bound on the reduction in cancer mortality due to periodic screening.

BACKGROUND: Because randomized cancer screening trials are very expensive, observational cancer screening studies can play an important role in the early phases of screening evaluation. Periodic screening evaluation (PSE) is a methodology for estimating the reduction in population cancer mortality from data on subjects who receive regularly scheduled screens. Although PSE does not require assumptions about natural history of cancer it requires other assumptions, particularly progressive detection - the assumption that once a cancer is detected by a screening test, it will always be detected by the screening test. METHODS: We formulate a simple version of PSE and show that it leads to an upper bound on screening efficacy if the progressive detection assumption does not hold (and any effect of birth cohort is minimal) To determine if the upper bound is reasonable, for three randomized screening trials, we compared PSE estimates based only on screened subjects with PSE estimates based on all subjects. RESULTS: In the three randomized screening trials, PSE estimates based on screened subjects gave fairly close results to PSE estimates based on all subjects. CONCLUSION: PSE has promise for obtaining an upper bound on the reduction in population cancer mortality rates based on observational screening data. If the upper bound estimate is found to be small and any birth cohort effects are likely minimal, then a definitive randomized trial would not be warranted.

Age Factors↗

A sensitivity analysis for nonrandomly missing categorical data arising from a national health disability survey.

Using data from 145,007 adults in the Disability Supplement to the National Health Interview Survey, we investigated the effect of balance difficulties on frequent depression after controlling for age, gender, race, and other baseline health status information. There were two major complications: (i) 80% of subjects were missing data on depression and the missing-data mechanism was likely related to depression, and (ii) the data arose from a complex sample survey. To adjust for (i) we investigated three classes of models: missingness in depression, missingness in depression and balance, and missingness in depression with an auxiliary variable. To adjust for (ii) we developed the first linearization variance formula for nonignorable missing-data models. Our sensitivity analysis was based on fitting a range of ignorable missing-data models along with nonignorable missing-data models that added one or two parameters. All nonignorable missing-data models that we considered fit the data substantially better than their ignorable missing-data counterparts. Under an ignorable missing-data mechanism, the odds ratio for the association between balance and depression was 2.0 with a 95% CI of (1.8, 2.2). Under 29 of the 30 selected nonignorable missing-data models, the odds ratios ranged from 2.7 with 95% CI of (2.3, 3.1) to 4.2 with 95% CI of (3.9, 4.6). Under one nonignorable missing-data model, the odds ratio was 7.4 with 95% CI of (6.3, 8.6). This is the first analysis to find a strong association between balance difficulties and frequent depression.

Adult↗

The transitive fallacy for randomized trials: if A bests B and B bests C in separate trials, is A better than C?

BACKGROUND: If intervention A bests B in one randomized trial, and B bests C in another randomized trial, can one conclude that A is better than C? The problem was motivated by the planning of a randomized trial, where A is spiral-CT screening, B is x-ray screening, and C is no screening. On its surface, this would appear to be a straightforward application of the transitive principle of logic. METHODS: We extended the graphical approach for omitted binary variables that was originally developed to illustrate Simpson's paradox, applying it to hypothetical, but plausible scenarios involving lung cancer screening, treatment for gastric cancer, and antibiotic therapy for clinical pneumonia. RESULTS: Graphical illustrations of the three examples show different ways the transitive fallacy for randomized trials can arise due to changes in an unobserved or unadjusted binary variable. In the most dramatic scenario, B bests C in the first trial, A bests B in the second trial, but C bests A at the time of the second trial. CONCLUSION: Even with large sample sizes, combining results from a previous randomized trial of B versus C with results from a new randomized trial of A versus B will not guarantee correct inference about A versus C. A three-arm trial of A, B, and C would protect against this problem and should be considered when the sequential trials are performed in the context of changing secular trends in important omitted variables such as therapy in cancer screening trials.

Antineoplastic Combined Chemotherapy Protocols↗

Statistical issues in randomized trials of cancer screening.

BACKGROUND: The evaluation of randomized trials for cancer screening involves special statistical considerations not found in therapeutic trials. Although some of these issues have been discussed previously, we present important recent and new methodologies. METHODS: Our emphasis is on simple approaches. RESULTS: We make the following recommendations: (1) Use death from cancer as the primary endpoint, but review death records carefully and report all causes of death; (2) Use a simple "causal" estimate to adjust for nonattendance and contamination occurring immediately after randomization; (3) Use a simple adaptive estimate to adjust for dilution in follow-up after the last screen CONCLUSION: The proposed guidelines combine recent methodological work on screening endpoints and noncompliance/contamination with a new adaptive method to adjust for dilution in a study where follow-up continues after the last screen. These guidelines ensure good practice in the design and analysis of randomized trials of cancer screening.

Endpoint Determination↗

Evaluating serial observations of precancerous lesions for further study as a trigger for early intervention.

Many long-term studies of the early detection of cancer involve serial observations of precancerous lesions and information as to whether or not the subject was diagnosed with cancer during the study period. Often the purpose of these studies is to decide whether or not the precancerous lesion should be studied in a future trial as a trigger for early intervention. A general approach to the analysis of cancer biomarkers is to estimate false and true positive rates to determine if they fall in the target region of those false and true positives that indicate promise for further study. The challenge with analysing serial data on precancerous lesions is estimating false and true positive rates when the number of observations varies among subjects. To solve this problem, we propose a Markov chain model in reverse time. The methodology is illustrated using serial observations of precancerous lesions found on sputum cytology.

False Negative Reactions↗

Markers for early detection of cancer: statistical guidelines for nested case-control studies.

BACKGROUND: Recently many long-term prospective studies have involved serial collection and storage of blood or tissue specimens. This has spurred nested case-control studies that involve testing some specimens for various markers that might predict cancer. Until now there has been little guidance in statistical design and analysis of these studies. METHODS: To develop statistical guidelines, we considered the purpose, the types of biases, and the opportunities for extracting additional information. RESULTS: The following guidelines: (1) For the clearest interpretation, statistics should be based on false and true positive rates - not odds ratios or relative risks (2) To avoid overdiagnosis bias, cases should be diagnosed as a result of symptoms rather than on screening. (3) To minimize selection bias, the spectrum of control conditions should be the same in study and target screening populations. (4) To extract additional information, criteria for a positive test should be based on combinations of individual markers and changes in marker levels over time. (5) To avoid overfitting, the criteria for a positive marker combination developed in a training sample should be evaluated in a random test sample from the same study and, if possible, a validation sample from another study. (6) To identify biomarkers with true and false positive rates similar to mammography, the training, test, and validation samples should each include at least 110 randomly selected subjects without cancer and 70 subjects with cancer. CONCLUSION: These guidelines ensure good practice in the design and analysis of nested case-control studies of early detection biomarkers.

Bias↗