PubMed Health⌕ Search

Biomedical subjects

G M Fitzmaurice

Publications and source records attributed to G M Fitzmaurice.

14 recordsLinked to original sources

An alternative parameterization of the general linear mixture model for longitudinal data with non-ignorable drop-outs.

This paper considers the mixture model methodology for handling non-ignorable drop-outs in longitudinal studies with continuous outcomes. Recently, Hogan and Laird have developed a mixture model for non-ignorable drop-outs which is a standard linear mixed effects model except that the parameters which characterize change over time depend also upon time of drop-out. That is, the mean response is linear in time, other covariates and drop-out time, and their interactions. One of the key attractions of the mixture modelling approach to drop-outs is that it is relatively easy to explore the sensitivity of results to model specification. However, the main drawback of mixture models is that the parameters that are ordinarily of interest are not immediately available, but require marginalization of the distribution of outcome over drop-out times. Furthermore, although a linear model is assumed for the conditional mean of the outcome vector given time of drop out, after marginalization, the unconditional mean of the outcome vector is not, in general, linear in the regression parameters. As a result, it is not possible to parsimoniously describe the effects of covariates on the marginal distribution of the outcome in terms of regression coefficients. The need to explicitly average over the distribution of the drop-out times and the absence of regression coefficients that describe the effects of covariates on the outcome are two unappealing features of the mixture modelling approach. In this paper we describe a particular parameterization of the general linear mixture model that circumvents both of these problems.

Anti-Asthmatic Agents↗

Bias in estimating association parameters for longitudinal binary responses with drop-outs.

This paper considers the impact of bias in the estimation of the association parameters for longitudinal binary responses when there are drop-outs. A number of different estimating equation approaches are considered for the case where drop-out cannot be assumed to be a completely random process. In particular, standard generalized estimating equations (GEE), GEE based on conditional residuals, GEE based on multivariate normal estimating equations for the covariance matrix, and second-order estimating equations (GEE2) are examined. These different GEE estimators are compared in terms of finite sample and asymptotic bias under a variety of drop-out processes. Finally, the relationship between bias in the estimation of the association parameters and bias in the estimation of the mean parameters is explored.

Algorithms↗

GEE with Gaussian estimation of the correlations when data are incomplete.

This paper considers a modification of generalized estimating equations (GEE) for handling missing binary response data. The proposed method uses Gaussian estimation of the correlation parameters, i.e., the estimating function that yields an estimate of the correlation parameters is obtained from the multivariate normal likelihood. The proposed method yields consistent estimates of the regression parameters when data are missing completely at random (MCAR). However, when data are missing at random (MAR), consistency may not hold. In a simulation study with repeated binary outcomes that are missing at random, the magnitude of the potential bias that can arise is examined. The results of the simulation study indicate that, when the working correlation matrix is correctly specified, the bias is almost negligible for the modified GEE. In the simulation study, the proposed modification of GEE is also compared to the standard GEE, multiple imputation, and weighted estimating equations approaches. Finally, the proposed method is illustrated using data from a longitudinal clinical trial comparing two therapeutic treatments, zidovudine (AZT) and didanosine (ddI), in patients with HIV.

Anti-HIV Agents↗

Goodness-of-fit for GEE: an example with mental health service utilization.

Suppose we use generalized estimating equations to estimate a marginal regression model for repeated binary observations. There are no established summary statistics available for assessing the adequacy of the fitted model. In this paper we propose a goodness-of-fit test statistic which has an approximate chi-squared distribution when we have specified the model correctly. The proposed statistic can be viewed as an extension of the Hosmer and Lemeshow goodness-of-fit statistic for ordinary logistic regression to marginal regression models for repeated binary responses. We illustrate the methods using data from a study of mental health service utilization by children. The repeated responses are a set of binary measures of service use. We fit a marginal logistic regression model to the data using generalized estimating equations, and we apply the proposed goodness-of-fit statistic to assess the adequacy of the fitted model.

Age Factors↗

Likelihood methods for incomplete longitudinal binary responses with incomplete categorical covariates.

We consider longitudinal studies in which the outcome observed over time is binary and the covariates of interest are categorical. With no missing responses or covariates, one specifies a multinomial model for the responses given the covariates and uses maximum likelihood to estimate the parameters. Unfortunately, incomplete data in the responses and covariates are a common occurrence in longitudinal studies. Here we assume the missing data are missing at random (Rubin, 1976, Biometrika 63, 581-592). Since all of the missing data (responses and covariates) are categorical, a useful technique for obtaining maximum likelihood parameter estimates is the EM algorithm by the method of weights proposed in Ibrahim (1990, Journal of the American Statistical Association 85, 765-769). In using the EM algorithm with missing responses and covariates, one specifies the joint distribution of the responses and covariates. Here we consider the parameters of the covariate distribution as a nuisance. In data sets where the percentage of missing data is high, the estimates of the nuisance parameters can lead to highly unstable estimates of the parameters of interest. We propose a conditional model for the covariate distribution that has several modeling advantages for the EM algorithm and provides a reduction in the number of nuisance parameters, thus providing more stable estimates in finite samples.

Affect↗

Regression models for mixed discrete and continuous responses with potentially missing values.

In this paper a likelihood-based method for analyzing mixed discrete and continuous regression models is proposed. We focus on marginal regression models, that is, models in which the marginal expectation of the response vector is related to covariates by known link functions. The proposed model is based on an extension of the general location model of Olkin and Tate (1961, Annals of Mathematical Statistics 32, 448-465), and can accommodate missing responses. When there are no missing data, our particular choice of parameterization yields maximum likelihood estimates of the marginal mean parameters that are robust to misspecification of the association between the responses. This robustness property does not, in general, hold for the case of incomplete data. There are a number of potential benefits of a multivariate approach over separate analyses of the distinct responses. First, a multivariate analysis can exploit the correlation structure of the response vector to address intrinsically multivariate questions. Second, multivariate test statistics allow for control over the inflation of the type I error that results when separate analyses of the distinct responses are performed without accounting for multiple comparisons. Third, it is generally possible to obtain more precise parameter estimates by accounting for the association between the responses. Finally, separate analyses of the distinct responses may be difficult to interpret when there is nonresponse because different sets of individuals contribute to each analysis. Furthermore, separate analyses can introduce bias when the missing responses are missing at random (MAR). A multivariate analysis can circumvent both of these problems. The proposed methods are applied to two biomedical datasets.

Air Pollution↗

The score test for independence in R x C contingency tables with missing data.

In this paper, the score test statistic for testing independence in R x C contingency tables with missing data is proposed. Under the null hypothesis of independence, the statistic has an approximate chi-squared distribution with (R - 1)(C - 1) degrees of freedom. The proposed test statistic is quite similar to the Pearson chi-squared statistic with complete data and, unlike the likelihood ratio statistic for testing independence, its computation is simple and noniterative. In addition, a score test statistic is proposed for testing independence when the rows and columns of the R x C table are ordinal. Finally, extensions of the score statistics to test for conditional independence in a set of (R x C) contingency tables with missing data are described. This yields score test statistics that are natural extensions of the Mantel-Haenszel statistic. An example, using a subset of data from the Six Cities Study, is presented to illustrate the methods.

Air Pollution↗

Estimating equations for measures of association between repeated binary responses.

Moment-based methods for analyzing repeated binary responses using the marginal odds ratio as a measure of association have been proposed by a number of authors. Carey, Zeger, and Diggle (1993, Biometrika 80, 517-526) have recently described how the marginal odds ratio can be estimated using generalized estimating equations (GEE) based on conditional residuals (deviations about conditional expectations). In this paper, we show that other measures of association between pairs of binary responses, e.g., the correlation, can also be estimated using conditional residuals. We demonstrate that the estimator of the correlation based on conditional residuals is nearly efficient when compared with maximum likelihood or second order estimating equations (GEE2) except when the correlation is large. This estimator also yields more efficient estimates of the correlation than the usual GEE estimator that is based on unconditional residuals. Furthermore, the gains in efficiency can be quite considerable when some of the responses are missing or incomplete, or, alternatively, when cluster sizes are unequal (in the clustered data setting).

Air Pollution↗

Bivariate logistic regression analysis of childhood psychopathology ratings using multiple informants.

A central issue in studies of risk factors for childhood psychopathology is utilization of the information obtained about the child's mental health status from multiple informants. In this paper, the authors propose a new approach to the analysis of risk factor data when the outcomes are binary ratings (presence/absence of symptoms). This new approach has several attractive features in this setting. The strategy taken is to perform a single analysis using multivariate modeling, in which simultaneous logistic regressions are conducted for the outcomes given by each of several informants. The advantages of this approach include the following: 1) it retains the complete information about case status for each informant; 2) it permits assessment of informant-risk factor interactions as well as "overall" risk factor effects; 3) it provides measures of association between the multiple informants and adjusts for the association between responses in the analysis; and 4) missing data on a subset of respondents can be incorporated in a straightforward way, permitting all subjects with at least one informant to be used in the analysis. To illustrate the methods, the authors present findings on risk factors for measures of "Internalizing" and "Externalizing" behaviors from two surveys using parent and teacher ratings of 6- to 11-year-old children in Connecticut between 1986 and 1989.

Child↗

Estimation methods for the join distribution of repeated binary observations.

The joint distribution of repeated binary observations is multinomial, and can be specified using a representation first suggested by Bahadur (1961, in Studies in Item Analysis and Prediction,158-168. Stanford, California: Stanford University Press), and later by Cox (1972, Applied Statistics 21, 13-120). Using the Bahadur representation, the marginal probabilities of success can be related to a set of covariates using the logistic link function, or any other suitable link function. Besides the parameters of the marginal regression model, we may also have interest in the probability of success on any of the repeated measures. For example, in the Six Cities study, a longitudinal study of the health effects of air pollution, we have interest in both the marginal probability of a child wheezing at age t (t = 10, 11, 12), and the union probability of wheezing at any of the three ages. This "union" probability can be specified in terms of the joint probabilities and the second higher-order correlations. We discuss several methods of estimating the parameters of the Bahadur model.

Adult↗

A caveat concerning independence estimating equations with multivariate binary data.

Clustered binary data occur commonly in both the biomedical and health sciences. In this paper, we consider logistic regression models for multivariate binary responses, where the association between the responses is largely regarded as a nuisance characteristic of the data. In particular, we consider the estimator based on independence estimating equations (IEE), which assumes that the responses are independent. This estimator has been shown to be nearly efficient when compared with maximum likelihood (ML) and generalized estimating equations (GEE) in a variety of settings. The purpose of this paper is to highlight a circumstance where assuming independence can lead to quite substantial losses of efficiency. In particular, when the covariate design includes within-cluster covariates, assuming independence can lead to a considerable loss of efficiency in estimating the regression parameters associated with those covariates.

Biometry↗

Sample size for repeated measures studies with binary responses.

We consider the sample size required for repeated measures studies when the response variable is binary. We propose the use of weighted least squares (WLS) for calculating the minimum sample size required to detect some minimum clinically important treatment effect. We provide tabulated values of the estimated sample sizes for a simple example and we discuss some practical considerations in determination of sample size with repeated binary responses.

Bias↗

Analysing incomplete longitudinal binary responses: a likelihood-based approach.

In this paper, we describe a likelihood-based method for analysing balanced but incomplete longitudinal binary responses that are assumed to be missing at random. Following the approach outlined in Zhao and Prentice (1990, Biometrika 77, 642-648), we focus on "marginal models" in which the marginal expectation of the response variable is related to a set of covariates. The association between binary responses is modelled in terms of conditional log odds-ratios. We describe a set of scoring equations for jointly estimating both the marginal parameters and the conditional association parameters. An outline of the EM algorithm used to obtain the maximum likelihood estimates is presented. This approach yields valid and efficient estimates when the responses are missing at random, but not necessarily missing completely at random. An example, using data from the Muscatine Coronary Risk Factor Study, is presented to illustrate this methodology.

Age Factors↗

Performance of generalized estimating equations in practical situations.

Moment methods for analyzing repeated binary responses have been proposed by Liang and Zeger (1986, Biometrika 73, 13-22), and extended by Prentice (1988, Biometrics 44, 1033-1048). In their generalized estimating equations (GEE), both Liang and Zeger (1986) and Prentice (1988) estimate the parameters associated with the expected value of an individual's vector of binary responses as well as the correlations between pairs of binary responses. In this paper, we discuss one-step estimators, i.e., estimators obtained from one step of the generalized estimating equations, and compare their performance to that of the fully iterated estimators in small samples. In simulations, we find the performance of the one-step estimator to be qualitatively similar to that of the fully iterated estimator. When the sample size is small and the association between binary responses is high, we recommend using the one-step estimator to circumvent convergence problems associated with the fully iterated GEE algorithm. Furthermore, we find the GEE methods to be more efficient than ordinary logistic regression with variance correction for estimating the effect of a time-varying covariate.

Air Pollution↗