PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “missing data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Sensitivity analysis of causal inference in a clinical trial subject to crossover.

In many clinical trials it is possible for some subjects to cross over between treatment arms. One can evaluate the effect of crossover by modeling it as a missing-data problem, where for subjects who cross over, one treats the unobserved value of the outcome in the original randomization arm as the missing data. The as-treated analysis is invalid if the crossover is nonignorable, in the sense that the crossovers represent a nonrandom sample of the randomized subjects. A recent area of general interest is the development of methods for measuring the sensitivity of inferences to nonignorability in the missing-data mechanism; one such approach is that of Troxel et al. In this paper we apply their method to the problem of measuring sensitivity to nonignorable crossover in randomized trials, extending it to the case where the crossover mechanism may differ between arms. Our method allows us to identify circumstances under which the as-treated analysis may be more or less sensitive to nonignorable crossover. We illustrate it with the example of a randomized clinical trial (RCT) in multiple sclerosis and a study of the effect of military service on income.

Causality↗

A global sensitivity analysis of performance of a medical diagnostic test when verification bias is present.

Current advances in technology provide less invasive or less expensive diagnostic tests for identifying disease status. When a diagnostic test is evaluated against an invasive or expensive gold standard test, one often finds that not all patients undergo the gold standard test. The sensitivity and specificity estimates based only on the patients with verified disease are often biased. This bias is called verification bias. Many authors have examined the consequences of verification bias and have proposed bias correction methods based on the assumption of independence between disease status and election for verification conditionally on the test result, or equivalently on the assumption that the disease status is missing at random using missing data terminology. This assumption may not be valid and one may need to consider adjustment for a possible non-ignorable verification bias resulting from the non-ignorable missing data mechanism. Such an adjustment involves ultimately uncheckable assumptions and requires sensitivity analysis. The sensitivity analysis is most often accomplished by perturbing parameters in the chosen model for the missing data mechanism, and it has a local flavour because perturbations are around the fitted model. In this paper we propose a global sensitivity analysis for assessing performance of a diagnostic test in the presence of verification bias. We derive a region of all sensitivity and specificity values consistent with the observed data and call this region a test ignorance region (TIR). The term 'ignorance' refers to the lack of knowledge due to the missing disease status for the not verified patients. The methodology is illustrated with two clinical examples.

Diagnostic Tests, Routine↗

Do quality of life assessments make a difference in the evaluation of cancer treatments?

The question posed by this set of quality of life papers is whether or not quality of life assessments in cancer clinical trials help evaluate the effects of cancer treatment on patient functioning. In this discussion, missing data problems, particularly those commonly found in advanced stage disease trials, are highlighted. Researchers are encouraged to investigate the extent of bias associated with missing data and to select analysis approaches accordingly. In the worst case, it may not be possible to analyze data longitudinally; descriptive or graphical portrayals of the data may be more appropriate. The importance of instrument reliability (minimizing measurement error) is emphasized for clinical trials research, particularly with respect to enhancing a trial's ability to detect quality of life differences by treatment arm. One strategy for addressing missing data is evaluated with respect to its impact on the measurement properties of the quality of life questionnaire. Clinical trials groups have been successful in obtaining quality of life data in multi-site settings and patients, by and large, appreciate the effort to include a systematic and standardized report of the effects of treatment on their functioning.

Bias↗

Using the national registry of HIV-infected veterans in research: lessons for the development of disease registries.

Disease-specific registries have many important applications in epidemiologic, clinical and health services research. Since 1989 the Department of Veterans Affairs has maintained a national HIV registry. VA's HIV registry is national in scope, it contains longitudinal data and detailed resource utilization and clinical information. To describe the structure, function, and limitations of VA's national HIV registry, and to test its accuracy and completeness. The VA's national HIV registry contains data that are electronically extracted from VA's computerized comprehensive clinical and administrative databases, called Veterans Integrated Health Systems Technology and Architecture (VISTA). We examined the number of AIDS patients and the number of new patients identified to the registry, by year, through December 1996. We verified data elements against information obtained from the medical records at five VA sites. By December 1996, 40,000 HIV-infected patients had been identified to the registry. We encountered missing data and problems with data classification. Missing data occurred for some elements related to the computer programming that creates the registry (e.g., pharmacy files), and for other elements because manual entry is required (e.g., ethnicity). Lack of a standardized data classification system was a problem, especially for the pharmacy and laboratory files. In using VA's national HIV registry we have learned important lessons, which, if taken into account in the future, could lead to the creation of model disease-specific registries.

HIV Infections↗

The Society of Thoracic Surgeons National Congenital Heart Surgery Database Report: analysis of the first harvest (1994-1997).

This analysis summarizes the first report of the Society of Thoracic Surgeons National Congenital Heart Surgery Database Committee in association with Summit Medical Systems. Twenty-four centers joined the program at various dates of entry resulting in 18,894 enrolled patient records. This report compiled the relevant clinical features of 18 congenital heart categories over a 4-year period (1994-1997), which included 8,149 patient records. The data analyses are largely descriptive in character. Missing data points were described and not omitted in the analysis. Statistical analysis was not performed due to missing data points in some categories. Certain trends, however, could be identified and are discussed. The first Society of Thoracic Surgeons National Congenital Heart Surgery Database Report has succeeded in establishing a finite record that can be improved to establish universal national and international utility, risk stratification, and scholarly outcome analyses.

Adolescent↗

Reasons for missing interviews in the daily electronic assessment of pain, mood, and stress.

Electronic diary assessment methods offer the potential to accurately characterize pain and other daily experiences. However, the frequent assessment of experiences over time often results in missing data. It is important to identify systematic reasons for missing data because such a pattern may bias study results and interpretations. We examined the reasons for missing electronic interviews, comparing self-report and data derived from electronic diary responses. Sixty-two patients with temporomandibular disorders were asked to rate pain intensity, pain-related activity interference, jaw use limitations, mood, and perceived stress three times a day for 8 weeks on palmtop computers. Participants also were asked the number of and reason(s) for missing electronic interviews. The average electronic diary completion rate was 91%. The correspondence between self-report and electronic data was high for the overall number of missed electronic interviews (Spearman correlation=0.77, P < 0.0001). The most common self-reported reasons for missing interviews were failure to hear the computer alarm (49%) and inconvenient time (21%). Although there was some suggestion that persistent negative mood and stress were associated with missing electronic interviews in a subgroup of patients, on the whole, the patient demographic and clinical characteristics, treatment, and daily fluctuations in pain, activity interference, mood, and stress were not associated significantly with missing daily electronic interviews. The results provide further support for the use of electronic diary methodology in pain research.

Adolescent↗

Longitudinal comparative study on the influence of computers on reporting of clinical data.

The impact of the clinical database system SISCOPE on medical services was evaluated and objective data compiled on the quality of information recording and reporting using a fully structured data entry system compared to traditional free text reporting. 1565 upper endoscopy reports produced with SISCOPE over a period of 12 months were assessed for completeness and compared to 152 and 208 free text reports done 4 months before and 1 month after the study period, respectively. Data on four common gastrointestinal findings (esophageal varices, ulcers, polyps and tumors) were evaluated. Physicians' compliance with the new system was good, as reflected by a constant level of quality of reporting over time, although a very slight decline in the ratio of computer generated reports to the total number of examinations was noted. Structured reports had an 18% missing data rate and contained 60% more relevant information than free text reports, which had a 48% missing data rate. No educational effect of the system was seen as missing data rates returned to pre-computerization levels just one month after the end of the study. It is concluded that menu-driven structured data entry systems result in production of far superior reports as compared to free text systems, probably due to their reminder effect.

Endoscopy, Digestive System↗

Multiple linear regression is a useful alternative to traditional analyses of variance.

Physiologists often wish to compare the effects of several different treatments on a continuous variable of interest, which requires an analysis of variance. Analysis of variance, as presented in most statistics texts, generally requires that there be no missing data and often that each sample group be the same size. Unfortunately, this requirement is rarely satisfied, and investigators are confronted with the problem of how to analyze data that do not strictly fit the traditional analysis of variance paradigm. One can avoid these pitfalls by recasting the analysis of variance as a multiple linear regression problem. When there are no missing data, the results of a traditional analysis of variance and the corresponding multiple regression problem are identical; when the sample sizes are unequal or there are missing data, one can use a regression formulation to analyze data that cannot be easily handled in a traditional analysis of variance paradigm and thus overcome a practical computational limitation of traditional analysis of variance. In addition to overcoming practical limitations of traditional analysis of variance, the multiple linear regression approach is more efficient because in one run of a statistics routine, not only is the analysis of variance done but also one obtains estimates of the size of the treatment effects (as opposed to just an indication of whether such effects are present or not), and many of the pairwise multiple comparisons are done (they are equivalent to t tests for significance of the regression parameter estimates). Finally, interaction between the different treatment factors is easier to interpret than it is in traditional analysis of variance.

Analysis of Variance↗

Validation of the Japanese version of the Quality of Life-Assessment of Growth Hormone Deficiency in Adults (QoL-AGHDA).

OBJECTIVE: To evaluate validity and reliability of the Japanese version of the Quality of Life-Assessment of Growth Hormone Deficiency in Adults (QoL-AGHDA). DESIGN: Observational study; cross-sectional, longitudinal. METHODS: Seventy-five adults with growth hormone deficiency completed the SF-36 (a generic health-related QOL scale) and the QoL-AGHDA before growth hormone replacement therapy and approximately 3 weeks later (when the therapy began). A sample (n=1000) of controls from the general population was also studied. We computed rates of missing data, measured reproducibility and internal consistency reliability, and tested for known-groups validity, concurrent validity, unidimensionality (by principle component analysis), and content validity. RESULTS: Rates of missing data were low (0-1.4%). The mean of QoL-AGHDA scores in the patients was 8.2 (SD, 6.4). The scores were reproducible (k=0.41-0.78), and internally consistent (alpha=0.91) and the scale was unidimensional. QoL-AGHDA scores were associated with SF-36 scores as hypothesized. Scores were significantly higher in the patients than in controls (8.1+/-0.7, and 5.6+/-0.2, P<0.001). Discrimination between patients and controls was slightly better using scores on the "General Health" and "Role Physical" subscale of the SF-36 as explanatory variables than using QoL-AGHDA scores. CONCLUSIONS: The QoL-AGHDA's reliability, validity, and rates of missing data were satisfactory, and the scale was confirmed to be unidimensional. However, because some subscales of the SF-36 were better for discriminating patients from controls, the content validity of the QoL-AGHDA may need to be re-evaluated.

Adolescent↗

The Funen Neck and Chest Pain study: analysing non-response bias by using national vital statistic data.

OBJECTIVE: To describe the Funen Neck and Chest Pain (FNCP) study and carry out a comprehensive non-response analysis of the quality of the survey. METHODS: The FNCP questionnaire was sent out to 7000 randomly selected individuals aged 20-71 years living in Funen County, Denmark. A full description of the FNCP survey, analysis of selection bias (representativeness of the background population), selective bias (non-responder bias), and item non-response bias was performed by using Danish vital statistics. RESULTS: The adjusted response rate was 60%. Men, retired individuals, and individuals with lower income tended to be late responders. Women, people aged 50+, married individuals, and individuals with two or three children, with a higher educational level, living in a single house or with a high income were more likely to participate. Conversely, men, younger individuals, singles or divorced persons, individuals living in residential caring homes, and with lower educational level were less likely to participate. Adjustments based on design and logical omissions gave an overall missing data rate of 1.0%. In general, the frequency of missing data increased with age and was higher for women. The frequency of inconsistent answers was 0.34%. CONCLUSIONS: A comprehensive non-response analysis showed some sociodemographic discrepancies between the FNCP study and the background population. The pattern of missing data was strongly associated with the design of the questionnaire and with participants' willingness to answer only questions they considered relevant. These factors must be taken into consideration when results from the FNCP study are presented.

Adult↗

GEL: a novel genotype calling algorithm using empirical likelihood.

MOTIVATION: Preliminary results on the data produced using the Affymetrix large-scale genotyping platforms show that it is necessary to construct improved genotype calling algorithms. There is evidence that some of the existing algorithms lead to an increased error rate in heterozygous genotypes, and a disproportionately large rate of heterozygotes with missing genotypes. Non-random errors and missing data can lead to an increase in the number of false discoveries in genetic association studies. Therefore, the factors that need to be evaluated in assessing the performance of an algorithm are the missing data (call) and error rates, but also the heterozygous proportions in missing data and errors. RESULTS: We introduce a novel genotype calling algorithm (GEL) for the Affymetrix GeneChip arrays. The algorithm uses likelihood calculations that are based on distributions inferred from the observed data. A key ingredient in accurate genotype calling is weighting the information that comes from each probe quartet according to the quality/reliability of the data in the quartet, and prior information on the performance of the quartet. AVAILABILITY: The GEL software is implemented in R and is available by request from the corresponding author at nicolae@galton.uchicago.edu.

Algorithms↗

A comprehensive literature review of haplotyping software and methods for use with unrelated individuals.

Interest in the assignment and frequency analysis of haplotypes in samples of unrelated individuals has increased immeasurably as a result of the emphasis placed on haplotype analyses by, for example, the International HapMap Project and related initiatives. Although there are many available computer programs for haplotype analysis applicable to samples of unrelated individuals, many of these programs have limitations and/or very specific uses. In this paper, the key features of available haplotype analysis software for use with unrelated individuals, as well as pooled DNA samples from unrelated individuals, are summarised. Programs for haplotype analysis were identified through keyword searches on PUBMED and various internet search engines, a review of citations from retrieved papers and personal communications, up to June 2004. Priority was given to functioning computer programs, rather than theoretical models and methods. The available software was considered in light of a number of factors: the algorithm(s) used, algorithm accuracy, assumptions, the accommodation of genotyping error, implementation of hypothesis testing, handling of missing data, software characteristics and web-based implementations. Review papers comparing specific methods and programs are also summarised. Forty-six haplotyping programs were identified and reviewed. The programs were divided into two groups: those designed for individual genotype data (a total of 43 programs) and those designed for use with pooled DNA samples (a total of three programs). The accuracy of programs using various criteria are assessed and the programs are categorised and discussed in light of: algorithm and method, accuracy, assumptions, genotyping error, hypothesis testing, missing data, software characteristics and web implementation. Many available programs have limitations (eg some cannot accommodate missing data) and/or are designed with specific tasks in mind (eg estimating haplotype frequencies rather than assigning most likely haplotypes to individuals). It is concluded that the selection of an appropriate haplotyping program for analysis purposes should be guided by what is known about the accuracy of estimation, as well as by the limitations and assumptions built into a program.

Algorithms↗

Pooling analysis of genetic data: the association of leptin receptor (LEPR) polymorphisms with variables related to human adiposity.

Analysis of raw pooled data from distinct studies of a single question generates a single statistical conclusion with greater power and precision than conventional metaanalysis based on within-study estimates. However, conducting analyses with pooled genetic data, in particular, is a daunting task that raises important statistical issues. In the process of analyzing data pooled from nine studies on the human leptin receptor (LEPR) gene for the association of three alleles (K109R, Q223R, and K656N) of LEPR with body mass index (BMI; kilograms divided by the square of the height in meters) and waist circumference (WC), we encountered the following methodological challenges: data on relatives, missing data, multivariate analysis, multiallele analysis at multiple loci, heterogeneity, and epistasis. We propose herein statistical methods and procedures to deal with such issues. With a total of 3263 related and unrelated subjects from diverse ethnic backgrounds such as African-American, Caucasian, Danish, Finnish, French-Canadian, and Nigerian, we tested effects of individual alleles; joint effects of alleles at multiple loci; epistatic effects among alleles at different loci; effect modification by age, sex, diabetes, and ethnicity; and pleiotropic genotype effects on BMI and WC. The statistical methodologies were applied, before and after multiple imputation of missing observations, to pooled data as well as to individual data sets for estimates from each study, the latter leading to a metaanalysis. The results from the metaanalysis and the pooling analysis showed that none of the effects were significant at the 0.05 level of significance. Heterogeneity tests showed that the variations of the nonsignificant effects are within the range of sampling variation. Although certain genotypic effects could be population specific, there was no statistically compelling evidence that any of the three LEPR alleles is associated with BMI or waist circumference in the general population.

Adipose Tissue↗

Maximum likelihood analysis of generalized linear models with missing covariates.

Missing data is a common occurrence in most medical research data collection enterprises. There is an extensive literature concerning missing data, much of which has focused on missing outcomes. Covariates in regression models are often missing, particularly if information is being collected from multiple sources. The method of weights is an implementation of the EM algorithm for general maximum-likelihood analysis of regression models, including generalized linear models (GLMs) with incomplete covariates. In this paper, we will describe the method of weights in detail, illustrate its application with several examples, discuss its advantages and limitations, and review extensions and applications of the method.

Algorithms↗

The weighted generalized estimating equations approach for the evaluation of medical diagnostic test at subunit level.

Sensitivity and specificity are common measures used to evaluate the performance of a diagnostic test. A diagnostic test is often administrated at a subunit level, e.g. at the level of vessel, ear or eye of a patient so that the treatment can be targeted at the specific subunit. Therefore, it is essential to evaluate the diagnostic test at the subunit level. Often patients with more negative subunit test results are less likely to receive the gold standard tests than patients with more positive subunit test results. To account for this type of missing data and correlation between subunit test results, we proposed a weighted generalized estimating equations (WGEE) approach to evaluate subunit sensitivities and specificities. A simulation study was conducted to evaluate the performance of the WGEE estimators and the weighted least squares (WLS) estimators (Barnhart and Kosinski, 2003) under a missing at random assumption. The results suggested that WGEE estimator is consistent under various scenarios of percentage of missing data and sample size, while the WLS approach could yield biased estimators due to a misspecified missing data mechanism. We illustrate the methodology with a cardiology example.

Biometry↗

Hospital utilization of Saskatchewan people with fetal alcohol syndrome.

We describe the hospital utilization of 194 Saskatchewan persons with Fetal Alcohol Syndrome (88% Aboriginal), born between 1973-92. Complete provincial hospitalization data were obtained for 128 patients; partial data for 29 patients. Proportionately more persons missing data were adopted, not living with biological family members or were deceased. The hospital separation rates for the children with FAS, pooled from 1987-91, compared to the 1989-90 Saskatchewan rates were significantly higher (95% level of confidence) for males and females < 1 year, 1-4 years and 5-14 years of age. Relative to Saskatchewan Registered Indians, significantly higher hospitalization rate ratios were observed for males with FAS in all age groups and for females only age 5-14 years. Rate ratios for younger females may not have achieved significance because of missing data. Higher hospitalization rates in children with FAS may not be explained solely by factors associated with ethnicity.

Adolescent↗

Highly scalable and robust rule learner: performance evaluation and comparison.

Business intelligence and bioinformatics applications increasingly require the mining of datasets consisting of millions of data points, or crafting real-time enterprise-level decision support systems for large corporations and drug companies. In all cases, there needs to be an underlying data mining system, and this mining system must be highly scalable. To this end, we describe a new rule learner called DataSqueezer. The learner belongs to the family of inductive supervised rule extraction algorithms. DataSqueezer is a simple, greedy, rule builder that generates a set of production rules from labeled input data. In spite of its relative simplicity, DataSqueezer is a very effective learner. The rules generated by the algorithm are compact, comprehensible, and have accuracy comparable to rules generated by other state-of-the-art rule extraction algorithms. The main advantages of DataSqueezer are very high efficiency, and missing data resistance. DataSqueezer exhibits log-linear asymptotic complexity with the number of training examples, and it is faster than other state-of-the-art rule learners. The learner is also robust to large quantities of missing data, as verified by extensive experimental comparison with the other learners. DataSqueezer is thus well suited to modern data mining and business intelligence tasks, which commonly involve huge datasets with a large fraction of missing data.

Algorithms↗

Prophylactic antibiotics to prevent chest infections in children with neurological impairment: the PARROT RCT.

BACKGROUND: Improvements in neonatal and paediatric care in recent decades have increased the survival of children with non-progressive neurological impairment. Respiratory disease in children with neurological impairment is common, with symptoms difficult to manage and lower respiratory tract infection occurring frequently. To reduce these, prophylactic antibiotics are being increasingly used, but the type, duration and dose of antibiotics can vary considerably, and there is limited evidence about their effectiveness in children and young people. A joint United Kingdom and Australia multicentre, randomised, double-blind, placebo-controlled trial comparing 52 weeks of azithromycin to placebo in children and young people with neurological impairment at risk of lower respiratory tract infection (PARROT) was planned to address this gap. PARROT was a multicentre, parallel group, blinded, pragmatic randomised controlled trial of 52-week duration with a planned sample size of 500 (250 in each arm) participants with neurological impairment. The primary outcome was the proportion of children and young people hospitalised with lower respiratory tract infection over the 52-week period. RESULTS: In total, 90 children and young people (62 in Australia, 28 in the United Kingdom) aged 3-17 years, with a diagnosed non-progressive, non-neuromuscular neurological impairment, who had persistent respiratory symptoms were randomised (1&#x2005;:&#x2005;1) to receive azithromycin or placebo. Baseline demographic and clinical characteristics were relatively well balanced across the two treatment groups and countries. Overall, mean (standard deviation) age was 9.2 (4.4) years, with 64% of participants having cerebral palsy, 67% being non-ambulant and 54% being totally tube-fed. At baseline, mean (standard deviation) numbers of hospital admissions with lower respiratory tract infection in the preceding year were 1.8 (2.0)/year, and general practitioner attendances 3.3 (3.0)/year. The PARROT trial was closed early to recruitment due to challenges arising from the COVID-19 pandemic. Sixty-five (72%) participants (azithromycin n&#x2005;=&#x2005;30, placebo n&#x2005;=&#x2005;35) completed 52 weeks of treatment and were not withdrawn early from the trial. Regarding the primary outcome, 11 (36.7%) in the azithromycin group were hospitalised with lower respiratory tract infection and 9 (25.7%) in the placebo group [absolute risk reduction 0.11 (95% confidence interval -0.12 to 0.33), relative risk&#xa0;1.43 (95% confidence interval 0.68 to 2.97)]. Analysis of secondary outcome data was limited by the number of missing data, but parent-reported quality of life for young person and parent, sleep amount/quality for young person and parent, and respiratory symptoms were similar between groups and countries. LIMITATIONS: As PARROT was stopped early and was consequently underpowered, it is not possible to say whether azithromycin prophylaxis is any more effective than placebo in reducing the proportion of children admitted to hospital with lower respiratory tract infection after a 52-week period. CONCLUSIONS AND FUTURE WORK: Although we cannot comment on the effectiveness of prophylactic antibiotics in this context, we can draw some useful conclusions from this trial. Thus, the importance placed by families on hospitalisation and its prevalence in both treatment groups, even during the pandemic, would suggest that this is an appropriate primary outcome measure for future trials in this high-risk group of children and young people. Furthermore, the high attrition rate and large numbers of missing data, specifically for questionnaire-based outcomes at later follow-up points, should encourage researchers to be mindful of minimising trial burden to families for any future trials wherever possible. FUNDING: This synopsis presents independent research funded by the National Institute for Health and Care Research (NIHR) Health Technology Assessment programme as award number 16/17/01.

Humans↗