PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “missing data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Design and analysis of trials with rare outcomes: examples from trials in herpes transmission and influenza prophylaxis.

For trials with rare outcomes, the number of events to be observed drives the power of the study rather than the proportions of subjects with the event and this is an important consideration when determining sample size. The stratified version of Fisher's exact test was described by Cox as long ago as 1966, but it is only recently that computing power has allowed this to be performed routinely. The issue of stratified analysis can also be addressed through permutation tests and their implementation for multicentre trials where randomisation is stratified by site has potential value. Where the time to event is available, analysis using a proportional hazards model has a potentially valuable role but estimates of differences in proportions are often required particularly for non-inferiority studies and typical software employed for these analyses uses an asymptotic approximation rather than an exact analysis. It is important to assess the impact of missing data on such trials. This can be approached through examination of the pattern of missing data and the covariates predicting discontinuation as well as through sensitivity analysis. Sensitivity analyses need to be done carefully using realistic alternative assumptions and an appealing approach is to impute events for subjects with missing data using the observed placebo rate. Issues of design and analysis for trials with rare outcomes are discussed in the context of two examples-one from a trial designed to investigate transmission of herpes and one that studied prophylaxis of influenza.

Chemoprevention↗

Estimation of competing risks with general missing pattern in failure types.

In competing risks data, missing failure types (causes) is a very common phenomenon. In this work, we consider a general missing pattern in which, if a failure type is not observed, one observes a set of possible types containing the true type, along with the failure time. We first consider maximum likelihood estimation with missing-at-random assumption via the expectation maximization (EM) algorithm. We then propose a Nelson-Aalen type estimator for situations when certain information on the conditional probability of the true type given a set of possible failure types is available from the experimentalists. This is based on a least-squares type method using the relationships between hazards for different types and hazards for different combinations of missing types. We conduct a simulation study to investigate the performance of this method, which indicates that bias may be small, even for high proportion of missing data, for sufficiently large number of observations. The estimates are somewhat sensitive to misspecification of the conditional probabilities of the true types when the missing proportion is high. We also consider an example from an animal experiment to illustrate our methodology.

Administration, Oral↗

Estimating person locations from partial credit data containing missing responses.

Certain assessment situations produce partial credit data. For instance, performance assessment items may utilize a rubric that assigns partial credit for some not completely correct responses. In some cases examinees may choose to not answer each question. This study investigated the effect of various strategies for handling these missing responses for estimating a respondent's location. These methods included ignoring the omitted response, selecting the "midpoint" category score, treating the omitted response as incorrect, hotdecking, and a likelihood-based approach. A simulation study was performed to examine the efficacy of these methods with the partial credit and generalized partial credit models. Expected a posteriori (EAP) ability estimation was used. Results showed that the Midpoint and Likelihood procedures performed the best of methods examined. In contrast, omitted responses should not be treated as incorrect nor ignored when estimating an examinee's proficiency using EAP. Implications for practitioners are discussed.

Data Collection↗

Methodological issues in trials assessing primary prophylaxis of venous thrombo-embolism.

AIMS: Many trials have been conducted to assess the efficacy of various strategies in the prevention of venous thrombo-embolism (VTE). Some of these trials have been subject to methodological criticisms. We aimed to assess the methodological issues raised by VTE trials. METHODS AND RESULTS: We searched MEDLINE and the Cochrane Central Register of Controlled Trials for articles assessing primary thromboprophylaxis published between 1994 and 2003 in 60 general medical and specialty journals. A total of 77 articles were analysed by two independent reviewers using a list of items. No primary endpoint was defined in 20% of trials. Although the primary endpoint was collected before day 15 in 75% of trials, there were >/=20% missing data in 56% of articles and >/=30% in 24.2% of articles. The rate of missing data was 23.7+/-9.7% in studies using venography-detected deep-vein thrombosis as an endpoint compared with 5.6+/-6.0% in studies using other endpoints. Among the 47 superiority trials, 27 (57.4%) reported an intention-to-treat (ITT) analysis, but only 10 (21.3%) reported an analysis that complied with this principle. These results were consistent when limiting the analysis to articles published in high-impact journal (impact factor more than 5). CONCLUSION: Recent randomized controlled trials assessing prophylactic regimens in VTE have important methodological limitations in terms of primary endpoints, missing data, and compliance with the ITT principle. These methodological shortcomings should be addressed when planning future trials.

Humans↗

Early use of inhaled corticosteroids in the emergency department treatment of acute asthma.

BACKGROUND: Systemic corticosteroids therapy is central to the management of acute asthma The use of inhaled steroids may also be beneficial in this setting. OBJECTIVES: To determine the benefit of ICS for the treatment of patients with acute asthma managed in the emergency department (ED). SEARCH STRATEGY: Randomised controlled trials (RCTs) were identified from the Cochrane Airways Review Group register. Bibliographies from included studies, known reviews, and texts also were searched. SELECTION CRITERIA: Only RCTs or quasi-randomised trials were eligible for inclusion. Studies were included if patients presented with acute asthma to the ED or its equivalent, and were treated with ICS or placebo, in addition to standard therapy. Two reviewers independently selected potentially relevant articles, and then independently selected articles for inclusion. Methodological quality was independently assessed by two reviewers. DATA COLLECTION AND ANALYSIS: Data were extracted independently by two reviewers if the authors were unable to verify the validity of extracted information. Missing data were obtained from the authors or calculated from other data presented in the paper. MAIN RESULTS: Seven trials were selected for inclusion, but data were not available for one of them. In the six usable rials, (4 adult, 2 paediatric), a total of 352 patients were studied (179 ICS, 173 non-ICS treated). Patients treated with ICS were less likely to be admitted to hospital (OR: 0.30; 95% CI: 0.16, 0.57). This benefit was confined to patients not receiving concomitant systemic steroids. Such patients showed the same, but non-significant, trend towards reduced admissions compared to placebo treatment (OR 0.46; 95% CI: 0. 19, 1.11). In children, ICS appeared to be at least as effective as systemic steroids (OR 0.5; 95% CI: 0.24, 1.06). Patients receiving ICS demonstrated small, significant improvements in peak expiratory flows (PEFR WMD: 7%; 95% CI: 3, 13) and forced expiratory volumes (FEV-FEV1 WMD: 5.0%; 95% CI: 0.4, 9.7). The treatment was well tolerated, with few reported adverse side effects. REVIEWER'S CONCLUSIONS: Inhaled steroids reduced admission rates in patients with acute asthma who were not receiving concomitant systemic steroids. In children, inhaled steroids appear to be at least as effective as systemic steroids. Further research is needed to clarify the effect of ICS when used in addition to systemic corticosteroids, and to determine the optimal dose, agent, and frequency of ICS administration.

Acute Disease↗

Conditional pairwise estimation in the Rasch model for ordered response categories using principal components.

In the Rasch model for items with more than two ordered response categories, the thresholds that define the successive categories are an integral part of the structure of each item in that the probability of the response in any category is a function of all thresholds, not just the thresholds between any two categories. This paper describes a method of estimation for the Rasch model that takes advantage of this structure. In particular, instead of estimating the thresholds directly, it estimates the principal components of the thresholds, from which threshold estimates are then recovered. The principal components are estimated using a pairwise maximum likelihood algorithm which specialises to the well known algorithm for dichotomous items. The method of estimation has three advantageous properties. First, by considering items in all possible pairs, sufficiency in the Rasch model is exploited with the person parameter conditioned out in estimating the item parameters, and by analogy to the pairwise algorithm for dichotomous items, the estimates appear to be consistent, though unlike for the dichotomous case, no formal proof has yet been provided. Second, the estimates of each item parameter is a function of frequencies in all categories of the item rather than just a function of frequencies of two adjacent categories. This stabilizes estimates in the presence of low frequency data. Third, the procedure accounts readily for missing data. All of these properties are important when the model is used for constructing variables from large scale data sets which must account for structurally missing data. A simulation study shows that the quality of the estimates is excellent.

Algorithms↗

Extensions to the modeling of initiation and progression: applications to substance use and abuse.

Twin data can provide valuable insight into the relationship between the stages of phenomena such as disease or substance abuse. Initiation of substance use may be caused by factors that are the same as, partially shared with, or completely independent of those that cause progression from use to abuse. Comparison of rates of progression among the cotwins of twins who do vs. do not initiate provides indirect information about the relationship between initiation and progression. Existing models for this relationship have been difficult to extend because they are usually expressed in terms of explicit integrals. In this paper, the problem is overcome by regarding the analysis of twin data on initiation and progression as a special case of missing data, in which individuals who do not initiate are regarded as having missing data on progression measures. Using the general framework for the analysis of ordinal data with missing values available in Mx makes extensions that include other variables much easier. The effects of continuous covariates such as age on initiation and progression becomes simple. Also facilitated are the examination of initiation and progression in two or more substances, and transition models with two or more steps. The methods are illustrated with data on the effects of cohort on liability to cannabis use and abuse, bivariate analysis of tobacco use and dependence and cannabis use and abuse, and the relationships between initiation of smoking, regular smoking and nicotine dependence. Other suitable applications include the relationship between symptoms and diagnosis, such as fears and the progression to phobia.

Age Factors↗

Quantitative image reconstruction of GaN quantum dots from oversampled diffraction intensities alone.

The missing data problem, i.e., the intensities at the center of diffraction patterns cannot be experimentally measured, is currently a major limitation for wider applications of coherent diffraction microscopy. We report here that, when the missing data are confined within the centrospeckle, the missing data problem can be reliably solved. With an improved instrument, we recorded 27 oversampled diffraction patterns at various orientations from a GaN quantum dot nanoparticle and performed quantitative image reconstruction from the diffraction intensities alone. This work in principle clears the way for single-shot imaging experiments using x-ray free electron lasers.

Journal Article↗

Genotype relative-risks and association tests for nuclear families with missing parental data.

The development of a new method for testing the association of genetic markers with disease is presented. This approach is applicable when sampling nuclear families with one or more affected siblings and when neither, one, or both parents are missing marker genotype data. All siblings, affected and not affected, are used to probabilistically infer the missing parental marker data. A likelihood ratio statistic, which treats marker allele frequencies as nuisance parameters, is presented to test whether all marker relative risks are equal to one (i.e., no marker association). This approach offers a solution to test for marker associations when parents are difficult to obtain.

Algorithms↗

[SF-36 Health Survey in Rehabilitation Research. Findings from the North German Network for Rehabilitation Research, NVRF, within the rehabilitation research funding program].

The SF-36 Health Survey and its 12-item abridged form is an instrument for the assessment of health related quality of life that can be used with healthy persons and patient populations. Its use has been recommended within a large German multicentre rehabilitation research programme. The paper examines missing data across all five study projects of the North German Network for Rehabilitation Research (NVRF) as well as psychometric properties of the instrument. In addition, data were compared to representative norm data using the SF-36 (SF-12) in the German National Health Survey. Results showed that there were few missing data in the SF-36. Examining the impact of age, gender and health status yielded effects of higher age and female gender on missing data. Psychometric analyses showed good to excellent results of the instrument in terms of scale fit and reliability. In terms of convergent validity, medium to high correlation of the SF-36 subscales with comparable instruments (e. g. SCL-90-R) could be found. Summarizing, the SF-36/SF-12 can be recommended for use in rehabilitation research. Analyses regarding sensitivity should be conducted in future studies.

Activities of Daily Living↗

The analysis of incomplete data in the three-period two-treatment cross-over design for clinical trials.

The additional time to complete a three-period two-treatment (3P2T) cross-over trial may cause a greater number of patient dropouts than with a two-period trial. This paper develops maximum likelihood (ML), single imputation and multiple imputation missing data analysis methods for the 3P2T cross-over designs. We use a simulation study to compare and contrast these methods with one another and with the benchmark method of missing data analysis for cross-over trials, the complete case (CC) method. Data patterns examined include those where the missingness differs between the drug types and depends on the unobserved data. Depending on the missing data mechanism and the rate of missingness of the data, one can realize substantial improvements in information recovery by using data from the partially completed patients. We recommend these approaches for the 3P2T cross-over designs.

Analysis of Variance↗

National Survey of Family Growth: design, estimation, and inference.

The purpose of this report is to document the procedures used in the 1988 National Survey of Family Growth (NSFG) to select the sample, weight the data to produce national estimates, impute missing data, and estimate sampling errors. Therefore, this report necessarily contains a great deal of technical detail. For readers who do not need this level of detail, this summary briefly describes the procedures used. The National Survey of Family Growth is conducted every few years by the National Center for Health Statistics (NCHS), a part of the U.S. Department of Health and Human Services. The purpose of the survey is to collect and publish data from a national sample of women on childbearing, factors affecting childbearing (such as contraception, sterilization, and infertility), and related aspects of maternal and infant health. Interviewing for Cycle IV of the survey was done in 1988 by Westat, Inc., under a contract with NCHS. Personal interviews were conducted between January and August of 1988 with a national sample of 8,450 women in the civilian noninstitutionalized population of the United States. Interviews were conducted in person by trained female interviewers and lasted an average of 70 minutes. The interview focused on the woman's pregnancies, if any; her use of contraception; her ability to bear children (fecundity and infertility); her use of medical services for family planning, infertility, and prenatal care; her marriage and cohabitation history, if any; and a wide range of demographic and economic characteristics. This report describes some of the main methodological aspects of the survey, including the sample design, weighting, sampling errors, and imputation of missing data. These topics will be described briefly and less technically in this summary. Each topic is discussed in more detail in the rest of the report.

Adolescent↗

Identification of significant host factors for HIV dynamics modelled by non-linear mixed-effects models.

Non-linear mixed-effects models are powerful tools for modelling HIV viral dynamics. In AIDS clinical trials, the viral load measurements for each subject are often sparse. In such cases, linearization procedures are usually used for inferences. Under such linearization procedures, however, standard covariate selection methods based on the approximate likelihood, such as the likelihood ratio test, may not be reliable. In order to identify significant host factors for HIV dynamics, in this paper we consider two alternative approaches for covariate selection: one is based on individual non-linear least square estimates and the other is based on individual empirical Bayes estimates. Our simulation study shows that, if the within-individual data are sparse and the between-individual variation is large, the two alternative covariate selection methods are more reliable than the likelihood ratio test, and the more powerful method based on individual empirical Bayes estimates is especially preferable. We also consider the missing data in covariates. The commonly used missing data methods may lead to misleading results. We recommend a multiple imputation method to handle missing covariates. A real data set from an AIDS clinical trial is analysed based on various covariate selection methods and missing data methods.

Acquired Immunodeficiency Syndrome↗

GENEHUNTER versus SimWalk2 in the context of an extended kindred and a qualitative trait locus.

GENEHUNTER and SimWalk2 are among the most commonly used software for parametric multipoint linkage analysis. In the context of extended kindred analysis, GENEHUNTER has a limitation in terms of the number of individuals it can handle. One solution is to manually split the kindred into smaller pedigrees. SimWalk2 can handle a much larger number of individuals. However, its major drawback is the time it takes to process the data when compared to GENEHUNTER. Aside from the limitations of each program, when studying extended kindreds researchers are typically confronted with missing data. In this work we used simulated genotype data based on the structure of a real extended pedigree in order to compare the results obtained through GENEHUNTER and SimWalk2, evaluate the effect of discarding individuals and splitting the kindred on the logarithm of odds (lod) score, and to assess how missing data affect the performance of each program. Our results show that (1) for pedigrees of a moderate size, GENEHUNTER and SimWalk2 produce nearly the same results; (2) when using GENEHUNTER, either splitting the kindred into smaller sub-pedigrees or discarding individuals has an adverse effect when compared to the results obtained when using SimWalk2 with the whole pedigree; and (3) the performance of both programs is qualitatively similar in the missing data scenario. These conclusions are based on the sample distributions of the lod score values and of the estimates of the recombination fraction.

Computer Simulation↗

An evaluation of power and type I error of single-nucleotide polymorphism transmission/disequilibrium-based statistical methods under different family structures, missing parental data, and population stratification.

Researchers conducting family-based association studies have a wide variety of transmission/disequilibrium (TD)-based methods to choose from, but few guidelines exist in the selection of a particular method to apply to available data. Using a simulation study design, we compared the power and type I error of eight popular TD-based methods under different family structures, frequencies of missing parental data, genetic models, and population stratifications. No method was uniformly most powerful under all conditions, but type I error was appropriate for nearly every test statistic under all conditions. Power varied widely across methods, with a 46.5% difference in power observed between the most powerful and the least powerful method when 50% of families consisted of an affected sib pair and one parent genotyped under an additive genetic model and a 35.2% difference when 50% of families consisted of a single affection-discordant sibling pair without parental genotypes available under an additive genetic model. Methods were generally robust to population stratification, although some slightly less so than others. The choice of a TD-based test statistic should be dependent on the predominant family structure ascertained, the frequency of missing parental genotypes, and the assumed genetic model.

Computer Simulation↗

Multicenter trial of fluoxetine as an adjunct to behavioral smoking cessation treatment.

The authors evaluated the efficacy of fluoxetine hydrochloride (Prozac; Eli Lilly and Company, Indianapolis, IN) as an adjunct to behavioral treatment for smoking cessation. Sixteen sites randomized 989 smokers to 3 dose conditions: 10 weeks of placebo, 30 mg, or 60 mg fluoxetine per day. Smokers received 9 sessions of individualized cognitive-behavioral therapy, and biologically verified 7-day self-reported abstinence follow-ups were conducted at 1, 3, and 6 months posttreatment. Analyses assuming missing data counted as smoking observed no treatment difference in outcomes. Pattern-mixture analysis that estimates treatment effects in the presence of missing data observed enhanced quit rates associated with both the 60-mg and 30-mg doses. Results support a modest, short-term effect of fluoxetine on smoking cessation and consideration of alternative models for handling missing data.

Adult↗

Sensitivity analysis for the estimation of rates of change with non-ignorable drop-out: an application to a randomized clinical trial of the vitamin D3.

The vitamin D(3) trial was a repeated measures randomized clinical trial for secondary hyperparathyroidism in haemodialysis patients where the efficacy of the vitamin D(3) infusions for suppressing the secretion of parathyroid hormone (PTH) was compared among four dose groups over 12 weeks. In this trial, patients terminated the study before the scheduled end of the study due to their elevated serum calcium (Ca) level, that is, the administration of the vitamin D(3) was expected to cause hypercalcaemia as an adverse event. In this setting of monotone missingness, there is a potential for bias in estimation of mean rates of decline in PTH for each treatment group using the standard methods such as the generalized estimating equations (GEE) which ignore the observed past Ca histories. We estimated the treatment-group-specific mean rates of decline in PTH by the inverse probability of censoring weighted (IPCW) methods which account for the observed past histories of time-dependent factors that are both a predictor of drop-out and are correlated with the outcomes. The IPCW estimator can be viewed as an extension of the GEE estimator that allows for the data to be MAR but not MCAR. With missing data, it is rarely appropriate to analyse the data solely under the assumption that the missing data process is ignorable, because the assumption of ignorable missingness cannot be guaranteed to hold and is untestable from the observed data. We proposed a sensitivity analysis that examines how inference about the IPCW estimates of the treatment-group-specific mean rates of decline in PTH changes as we vary the non-ignorable selection bias parameter over a range of plausible values.

Bias↗

Multiple imputation technique applied to appropriateness ratings in cataract surgery.

Missing data such as appropriateness ratings in clinical research are a common problem and this often yields a biased result. This paper aims to introduce the multiple imputation method to handle missing data in clinical research and to suggest that the multiple imputation technique can give more accurate estimates than those of a complete-case analysis. The idea of multiple imputation is that each missing value is replaced with more than one plausible value. The appropriateness method was developed as a pragmatic solution to problem of trying to assess "appropriate" surgical and medical procedures for patients. Cataract surgery was selected as one of four procedures that were evaluated as a part of the Clinical Appropriateness Initiative. We created mild to high missing rates of 10%, 30% and 50% and compared the performance of logistic regression in cataract surgery. We treated the coefficients in the original data as true parameters and compared them with the other results. In the mild missing rate (10%), the deviation from the true coefficients was quite small and ignorable. After removing the missing data, the complete-case analysis did not reveal any serious bias. However, as the missing rate increased, the bias was not ignorable and it distorted the result. This simulation study suggests that a multiple imputation technique can give more accurate estimates than those of a complete-case analysis, especially for moderate to high missing rates (30 - 50%). In addition, the multiple imputation technique yields better accuracy than a single imputation technique. Therefore, multiple imputation is useful and efficient for a situation in clinical research where there is large amounts of missing data.

Cataract Extraction↗