PubMed Health⌕ Search

Biomedical subjects

Joseph L Schafer

Publications and source records attributed to Joseph L Schafer.

4 recordsLinked to original sources

Robustness of a multivariate normal approximation for imputation of incomplete binary data.

Multiple imputation has become easier to perform with the advent of several software packages that provide imputations under a multivariate normal model, but imputation of missing binary data remains an important practical problem. Here, we explore three alternative methods for converting a multivariate normal imputed value into a binary imputed value: (1) simple rounding of the imputed value to the nearer of 0 or 1, (2) a Bernoulli draw based on a 'coin flip' where an imputed value between 0 and 1 is treated as the probability of drawing a 1, and (3) an adaptive rounding scheme where the cut-off value for determining whether to round to 0 or 1 is based on a normal approximation to the binomial distribution, making use of the marginal proportions of 0's and 1's on the variable. We perform simulation studies on a data set of 206,802 respondents to the California Healthy Kids Survey, where the fully observed data on 198,262 individuals defines the population, from which we repeatedly draw samples with missing data, impute, calculate statistics and confidence intervals, and compare bias and coverage against the true values. Frequently, we found satisfactory bias and coverage properties, suggesting that approaches such as these that are based on statistical approximations are preferable in applied research to either avoiding settings where missing data occur or relying on complete-case analyses. Considering both the occurrence and extent of deficits in coverage, we found that adaptive rounding provided the best performance.

Adolescent↗

Using data augmentation to obtain standard errors and conduct hypothesis tests in latent class and latent transition analysis.

Latent class analysis (LCA) provides a means of identifying a mixture of subgroups in a population measured by multiple categorical indicators. Latent transition analysis (LTA) is a type of LCA that facilitates addressing research questions concerning stage-sequential change over time in longitudinal data. Both approaches have been used with increasing frequency in the social sciences. The objective of this article is to illustrate data augmentation (DA), a Markov chain Monte Carlo procedure that can be used to obtain parameter estimates and standard errors for LCA and LTA models. By use of DA it is possible to construct hypothesis tests concerning not only standard model parameters but also combinations of parameters, affording tremendous flexibility. DA is demonstrated with an example involving tests of ethnic differences, gender differences, and an Ethnicity x Gender interaction in the development of adolescent problem behavior.

Humans↗

On the performance of random-coefficient pattern-mixture models for non-ignorable drop-out.

Random-coefficient pattern-mixture models (RCPMMs) have been proposed for longitudinal data when drop-out is thought to be non-ignorable. An RCPMM is a random-effects model with summaries of drop-out time included among the regressors. The basis of every RCPMM is extrapolation. We review RCPMMs, describe various extrapolation strategies, and show how analyses may be simplified through multiple imputation. Using simulated and real data, we show that alternative RCPMMs that fit equally well may lead to very different estimates for parameters of interest. We also show that minor model misspecification can introduce biases that are quite large relative to standard errors, even in fairly small samples. For many scientific applications, where the form of the population model and nature of the drop-out are unknown, interval estimates from any single RCPMM may suffer from undercoverage because uncertainty about model specification is not taken into account.

Antipsychotic Agents↗

Missing data: our view of the state of the art.

Statistical procedures for missing data have vastly improved, yet misconception and unsound practice still abound. The authors frame the missing-data problem, review methods, offer advice, and raise issues that remain unresolved. They clear up common misunderstandings regarding the missing at random (MAR) concept. They summarize the evidence against older procedures and, with few exceptions, discourage their use. They present, in both technical and practical language, 2 general approaches that come highly recommended: maximum likelihood (ML) and Bayesian multiple imputation (MI). Newer developments are discussed, including some for dealing with missing data that are not MAR. Although not yet in the mainstream, these procedures may eventually extend the ML and MI methods that currently represent the state of the art.

Databases as Topic↗