PubMed HealthSearch

SEARCH · PubMed Health

Results for “missing data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Ignorability and coarse data: some biomedical examples.

Heitjan and Rubin (1991, Annals of Statistics 19, 2244-2253) define data to be "coarse" when one observes not the exact value of the data but only some set (a subset of the sample space) that contains the exact value. This definition covers a number of incomplete-data problems arising in biomedicine, including rounded, heaped, censored, and missing data. In analyzing coarse data, it is common to proceed as though the degree of coarseness is fixed in advance--in a word, to ignore the randomness in the coarsening mechanism. When coarsening is actually stochastic, however, inferences that ignore this randomness may be seriously misleading. Heitjan and Rubin (1991) have proposed a general model of data coarsening and established conditions under which it is appropriate to ignore the stochastic nature of the coarsening. The conditions are that the data be coarsened at random [a generalization of missing at random (Rubin, 1976, Biometrika 63, 581-592)] and that the parameters of the data and the coarsening process be distinct. This article presents detailed applications of the general model and the ignorability conditions to a variety of coarse-data problems arising in biomedical statistics. A reanalysis of the Stanford Heart Transplant Data (Crowley and Hu, 1977, Journal of the American Statistical Association 72, 27-36) reveals significant evidence that censoring of pretransplant survival times by transplantation was nonignorable, suggesting a greater benefit from cardiac transplantation than previous analyses had found.

Age Factors

Approaches to the analysis of quality of life data: experiences gained from a medical research council lung cancer working party palliative chemotherapy trial.

Standardization in the choice of quality of life (QOL) instruments and their application in randomised clinical trials have been advocated and generally accepted. However, there is now an urgent need to address the problems relating to the analysis and presentation of the data thus generated. There are intrinsic difficulties associated with QOL data, namely its multidimensional nature, attrition and missing data, and there is no consensus as to how these problems should be dealt with. This paper therefore considers these problems using interim data from a large Medical Research Council randomised trial in patients with small cell lung cancer and a poor prognosis, in which attrition and compliance are major concerns. Three possible approaches to the analysis of these data, which use different subsets of patients, are examined in detail. The strengths and weaknesses of these three methods are discussed, and examples of their use in the literature are given and compared with other reported approaches. The need for a standard definition of compliance is also emphasised, and a method of presentation suggested. The best current advice is that QOL data should be analysed in a number of different ways, and conclusions reached only when consistency is seen.

Antineoplastic Agents

Repeated measures designs in behavioral toxicology: application to chronic marijuana smoke exposure.

This paper discusses the application of repeated measures methods in the statistical analysis of an experiment in behavioral toxicology. The chronic marijuana smoke exposure study conducted at the National Center for Toxicological Research is used for an example of the types of problems that one encounters in analyzing these types of studies. In particular, the standard univariate analysis most frequently used for repeated measures analyses has some very restrictive assumptions on the form of the covariance matrices. These assumptions are not met in the example discussed and are rarely met in many other problems. Other possible models for analyzing repeated measures when these assumptions are not met are presented and discussed. Other problems specific to the chronic marijuana smoke exposure study that may occur in similar type studies are presented. These include pooling the experimental units into groups with comparable baselines, choosing a function of the measures to be analyzed, dealing with a large data set with many observation times and missing data, unequal group sizes and different designs for different subsets of the experimental animals. The standard univariate repeated measures analysis was chosen to analyze the data even though the violations of the covariance assumptions may lead to finding differences that do not exist (Type I or false-positive errors), since the other methods presented also had covariance assumptions that were not met or had low power. Use of Bonferroni-type multiple comparisons on the single degree of freedom contrasts of interest hopefully reduced the chances of these false-positive results.

Analysis of Variance

Using large data bases in nursing and health policy research.

Concern about the quality, cost, and outcomes of health care has become a driving force in health policy research. The growing accessibility of large clinical and administrative health care data bases has led to an interest in using such data in health policy research. Clinical data bases are created by providers of care and contain data about episodes and outcomes of care, usually organized as patient records. Administrative data bases contain data about indirect care processes such as insurance claims processing, vital event recording, and quality assurance. Clinical and administrative data bases may contain millions of records, consist of data from multiple sites, and often have missing data issues that must be considered by researchers. These and other characteristics of large data bases require special data manipulation and analytic techniques. Large data bases have been used in epidemiological studies, risk assessment, and technology assessment and to study variations in caregiver practice patterns. Because the use of large data bases by nurse researchers has been constrained by the lack of nursing-relevant data in them, there is a need to reach consensus on useful and feasible nursing data elements and to include those data in ongoing data collection efforts by government agencies and private organizations.

Bias

Validity of parental report of a child's medical history in otitis media research.

The authors compared parental reports with medical records for 157 children enrolled in a prospective study of chronic otitis media with effusion between 1987 and 1991. Parents completed a questionnaire about the child's past health history, and the research nurse abstracted history information from the clinic's medical record. Previous insertion of a tympanostomy tube (kappa = 0.96) and premature birth (kappa = 0.68) were accurately reported, but there was a substantial proportion of missing data for age at first episode of otitis media, occurrence of otitis media the previous summer, and number of episodes in the previous 18 months. Data were significantly more likely to be missing for male children, children with siblings, and those with more episodes. Parents who reported six or more previous episodes for their child overestimated the number compared with the medical record (8.7 vs. 7.4, respectively; p = 0.01), while those who reported fewer episodes underestimated the number (3.1 vs. 4.6, respectively; p = 0.01). Episodes of otitis media during the 3 months between study visits were also accurately reported (kappa = 0.94). The accuracy and completeness of parental report of the child's health history was influenced by the chronicity of otitis media, the duration of recall, and the seriousness of the event being recalled.

Child, Preschool

The impact of changing methods of data collection on the reliability of self-reported drug use of adolescents.

The purpose of this study is to determine the impact of different modes of data collection on the reliability of self-reported drug use of adolescents in a panel study. Adolescents were assigned to four groups based upon the ways they chose to respond to the survey instruments: 1) mailed questionnaires in both years, 2) survey interview in one year and mailed questionnaire in the next year, 3) mailed questionnaire in one year and survey interview in the following year, and 4) survey interview in both years. The quality of the self-reported data was examined in terms of return rates, missing data, internal consistency, and consistency of reported information over time. No significant differences were found between groups, suggesting that the mode of data collection does not affect the reliability of adolescents' self-reports of substance use.

Adolescent

Applications of multiple imputation to the analysis of censored regression data.

The first part of the article reviews the Data Augmentation algorithm and presents two approximations to the Data Augmentation algorithm for the analysis of missing-data problems: the Poor Man's Data Augmentation algorithm and the Asymptotic Data Augmentation algorithm. These two algorithms are then implemented in the context of censored regression data to obtain semiparametric methodology. The performances of the censored regression algorithms are examined in a simulation study. It is found, up to the precision of the study, that the bias of both the Poor Man's and Asymptotic Data Augmentation estimators, as well as the Buckley-James estimator, does not appear to differ from zero. However, with regard to mean squared error, over a wide range of settings examined in this simulation study, the two Data Augmentation estimators have a smaller mean squared error than does the Buckley-James estimator. In addition, associated with the two Data Augmentation estimators is a natural device for estimating the standard error of the estimated regression parameters. It is shown how this device can be used to estimate the standard error of either Data Augmentation estimate of any parameter (e.g., the correlation coefficient) associated with the model. In the simulation study, the estimated standard error of the Asymptotic Data Augmentation estimate of the regression parameter is found to be congruent with the Monte Carlo standard deviation of the corresponding parameter estimate. The algorithms are illustrated using the updated Stanford heart transplant data set.

Algorithms

Ensuring data quality in a multicenter clinical trial: remote site data entry, central coordination and feedback.

In an ongoing multicenter clinical trial, "Treatment Strategies in Schizophrenia," the five participating sites have the capacity to perform a variety of tasks or study functions independently. These tasks include (a) verification of diagnostic eligibility through the use of computerized decision algorithms; (b) assignment of patients to treatment based on prognostic indicators using a computerized randomization algorithm; (c) entry of data into a microcomputer using a clinical trial data management system that performs simple range and missing data item checks; and (d) regular transfer of all data to the central coordinating team. The clinical trial data management system employed allows for both independent site functioning and assurance of consistency across sites. The integration of a variety of software outside the main data management system provides the central coordinators with the tools to monitor critical data as it is collected, as well as the capacity to assess the flow, quality, and uniformity of the ongoing trial.

Clinical Trials as Topic

Nicardipine and propranolol in the treatment of essential hypertension.

Two hundred thirty-four patients with supine diastolic blood pressure of between 95 and 114 mm Hg were enrolled into a double-blind, randomized, parallel, multicenter trial. The patients were randomized to either nicardipine 30 mg tid, propranolol 40 mg tid, or nicardipine 30 mg tid and propranolol 40 mg tid for six weeks. Two hundred six patients yielded data for analyses. Of the 28 not included, seven had missing data, whereas the remaining 21 were excluded because they either failed to meet inclusion criteria or were noncompliant at endpoint. Both nicardipine and propranolol as monotherapies and in combination achieved statistically significant, (P less than .01), supine diastolic blood pressure reduction relative to baseline. The combination of nicardipine and propranolol showed a greater reduction in supine diastolic and systolic measurements than either of the monotherapies. Nicardipine produced greater blood pressure reductions one hour after dosing, whereas the propranolol treatment tended to produce slightly greater blood pressure decreases eight hours after dose. The combination always resulted in the greatest blood pressure reduction, independent of time after dose. Adverse experiences were reported by 26% of patients in the nicardipine-treated group, most often transient vasodilatory effects, by 17% of the propranolol-treated patients, and by 18% of the combination-treated group. This study demonstrated at the doses studied that nicardipine alone produced equivalent blood pressure reductions to those obtained by propranolol alone, but that the combination of these two drugs produced greater reductions in blood pressures than either of the monotherapies.

Adult

"Keyhole" method for accelerating imaging of contrast agent uptake.

Magnetic resonance (MR) imaging methods with good spatial and contrast resolution are often too slow to follow the uptake of contrast agents with the desired temporal resolution. Imaging can be accelerated by skipping the acquisition of data normally taken with strong phase-encoding gradients, restricting acquisition to weak-gradient data only. If the usual procedure of substituting zeroes for the missing data is followed, blurring results. Substituting instead reference data taken before or well after contrast agent injection reduces this problem. Volunteer and patient images obtained by using such reference data show that imaging can be usefully accelerated severalfold. Cortical and medullary regions of interest and whole kidney regions were studied, and both gradient- and spin-echo images are shown. The method is believed to be compatible with other acceleration methods such as half-Fourier reconstruction and reading of more than one line of k space per excitation.

Contrast Media

Crossover designs for clinical trials.

I discuss three-period crossover designs for an efficient comparison of two test treatments with special application to clinical trials which often have many practical limitations. In this paper I specify a subset of three-period crossover designs so that the investigators are not left with the problematic two-period two-sequence design, should the trials be terminated after the second period. I show that there is a dramatic reduction in variability for estimating the direct and residual treatment effects in three-period designs compared to two-period designs. I also show that the universally optimal design with ABB and BAA sequences is unsuitable when a complex form of residual effects is suspected, such as the second-order residual effects or treatment by period interactions. The design with ABB, BAA, AAB, and BBA sequences is relatively robust to these uncertain model assumptions. I also discuss missing data problems and conclude that, even with a large proportion of missing values, the three-period design is far more efficient than the two-period design.

Analysis of Variance

A PC program for classification into one of several groups on the basis of longitudinal data.

A stand-alone, menu-driven PC program, ZCLASS, written in GAUSS386i, for classifying subjects into one of several distinct, existing groups on the basis of longitudinal data is described, illustrated, and made available to interested readers. The program accepts data from studies where common times of measurement are planned, but missing data are accommodated in that one or more measurement sequences may be incomplete.

Anthropometry

Microcomputer application of Bayesean probability testing for the identification of bacteria.

A computer program (BACTID) is described which facilitates the identification of bacteria based on a priori data and Bayesean probability testing. The program is not limited to a specific format, has a short execution time, can be easily applied to a variety of situations, and can be run on almost any microcomputer system operating under either 8-bit CP/M or 16-bit MS-DOS/PC-DOS. Additionally, BACTID (1) is not limited to one type of computer (hardware independent), (2) is not limited by size of the computer's random access (RAM independent), (3) can recognize various data bases matrices (format independent), (4) is able to compensate for missing data and (5) allows for various methods of data entry. The efficacy of the program was checked against a commercially available test system and a 99.34% agreement was obtained. Also, the execution time for a 46 x 21 element data matrix was as little as 3.5 s. These results show that microcomputer identification programs are not only viable alternatives to code book registers, but also offer flexibility which is not found in commercial systems.

Bacteria

A comparison of response rate, data quality, and cost in the collection of data on sexual history and personal behaviors. Mail survey approaches and in-person interview.

The authors examined differences in rate of response, data quality, and cost between mail approaches and in-person interview in the collection of data on sexual history and personal behaviors. A sample of women from a midwestern United States university (n = 342) was identified from health service medical records as having been seen for a sexually transmitted disease (cases) or a contraceptive visit (controls) during the latter half of 1985. The women were randomly assigned to one of three data collection strategies. A total of 268 subjects (78%) participated. Results indicated no differences in validity by method of data collection or by case-control status but there were significant differences in completeness, cost, and response rates. In-person interviews resulted in more complete data than mail approaches, although all instruments had low proportions of missing data (0.001-0.006). Response rate differences were not found when data collection methodologies were compared (75-82%) but were found in case-control analyses. Cases were consistently less likely to participate and significantly less likely to respond by mail (p less than 0.05). The cost of the in-person interview was approximately four times that of the mail survey for the data collection. Implications of the case-control response rate difference suggest that mail methodologies, although low in cost, may introduce sampling bias in studies of sexually transmitted diseases.

Adolescent

Studies on occupational health: a critique.

A critique of 48 recent articles dealing with occupational health was undertaken by two readers using a set of questions devised to assess adherence to selected methodologic principles concerning data quality. Articles were read independently, responses to assessment questions were discussed and differences between readers reconciled. The greatest inattention to principles was found in the areas of sample size; definition of exposure; description, standardization, and validation of data sources; use of "blind" observers; and the possible effect of missing data on the results. Rigorus attention to these methodologic principles is necessary if the results of studies are to be accepted and applied in the prevention and control of occupational hazards.

Occupational Medicine

A method for assessing patterns of familial resemblance in complex human pedigrees, with an application to the nevus-count data in Utah kindreds.

An analytic method is described for estimating phenotypic correlations between pairs of members of specific relationships in pedigrees. In estimating correlations, this new method allows simultaneous adjustment for available covariates such as age, gender, environmental factors, and variables reflecting ascertainment mode, through mean- and variance-regression models. The estimated correlations and regression coefficients corresponding to covariates are consistent and asymptotically normally distributed. Differing from a full-likelihood approach, this new method does not require the assumption of a particular joint distribution of phenotypes from a pedigree, such as the multivariate normal distribution, but instead only requires correct specification of mean- and variance-regression models. Within this framework, missing data, if they are missing completely at random, can be ignored without biasing estimates. The method is illustrated by an application using nevus-count data from 28 Utah kinships. The results from the analysis are that covariate-adjusted nevus counts are correlated between parents and children (correlation .22; P less than .001) and between siblings (correlation .32; P less than .001), while the correlation of -.04 between husband and wife is not significantly different (P = .31) from 0. This result is consistent with a genetic etiology of nevus count.

Dysplastic Nevus Syndrome

Closed-form estimates for missing counts in two-way contingency tables.

One method for analyzing contingency tables with missing observations is to model the missing-data mechanism using log-linear models. Previous methods for obtaining estimates (of missing counts and parameters) have required an iterative algorithm. In many cases, however, one can obtain estimates by use of a simple algebraic formula. We illustrate the method with data on smoking and birth weight.

Algorithms

The use of lung function tests in identifying factors that affect lung growth and aging.

Lung function tests are used both clinically, in assessing disease, and epidemiologically, in identifying those factors which influence the growth and aging process of the lungs. The user must beware of several common pitfalls in the use of these tests, however. First, the commonly used tests of lung function can only identify patterns of dysfunction, not specific pathologic processes. Second, these tests are subject to many sources of intra- and inter-subject variability, making it difficult to dissect out the signal (for example, the rate of lung aging in adults) from the noise which may greatly exceed the signal. Finally, the analysis of longitudinal pulmonary function data is complicated by the alinearity of the growth and aging process, missing data, variable follow-up times, information censoring mechanisms and covariate processes, and problems in defining abnormality.

Adolescent