PubMed HealthSearch

SEARCH · PubMed Health

Results for “missing data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Pre-natal blood lead levels and learning difficulties in children: an analysis of non-randomly missing categorical data.

This paper presents an analysis of categorical variables subject to non-response. We incorporate the incomplete data into the analysis by modelling the distribution of the variables of interest and the non-response mechanism. We discuss issues of model selection and interpretation and the effect of discarding incomplete observations. In addition, we describe how to perform all of the computations with standard statistical software. We discuss the problem of incomplete categorical data within the context of a study of the effect of lead exposure on learning difficulties in children. In this study, many of the children are not observed on some of the variables of interest. It is particularly important in this study to incorporate the incomplete data, since there is evidence that non-response is related to the variables of interest. We reach different conclusions when we incorporate the incomplete data into the analysis than we reach when we discard the incomplete data. We also examine the sensitivity of our conclusions to the choice of a model for the non-response mechanism.

Algorithms

Correcting single channel data for missed events.

Interpretation of currents recorded from single ion channels in cellular membranes or lipid bilayers is complicated by the necessarily limited time resolution of the recording and detection systems. All intervals less than a certain duration, depending on the frequency response of the system, are not detected. Such missed events produce increases in the durations of observed open and shut intervals. In order to obtain the true kinetic scheme and rate constants underlying the observed activity, it is necessary to take into account missed events. We develop methods to correct for missed events for models with two or more states, including models with multiple open and shut states, compound states, and loops. Our methods can be used in a forward direction to predict observed distributions of open and shut intervals for a given kinetic scheme and time resolution. They can also be used in a backwards direction with iterative methods to determine rate constants consistent with the observed distributions. While a given kinetic scheme with rate constants predicts unique observed distributions of open and shut intervals, rate constants determined from observed distributions are not necessarily unique. Using these correction methods, we examine the effects of missed events for a five-state model consistent with some properties of large conductance Ca-activated K channels.

Ion Channels

Malaria vaccine trials: the missing qualitative data.

Recent population-based efficacy trials of the synthetic malaria vaccine SPf66 have shown restricted, if any, clinical protection against Plasmodium falciparum infection. Despite the well-established role of antibodies in effector responses against asexual blood-stage malaria parasites, the titres of anti-SPf66 IgG antibodies do not correlate with the ability of sera from vaccine recipients to inhibit parasite growth in vitro nor with partial clinical protection which could be detected in some trials. Qualitative or functional parameters of SP66-induced antibody responses, such as IgG subclass composition and affinity, may be more predictive of clinical protection against malaria than quantitative estimates of antibody concentration or titre. Since these parameters are readily estimated by laboratory techniques currently available, and may be modulated by changes in vaccination protocols and by the use of different adjuvants, a better understanding of qualitative antibody responses induced by SPf66 and other asexual blood-stage malaria vaccine candidates, and of their relationship with clinical protection in vivo, is urgently needed for the improvement of currently used immunization schedules.

Animals

Research in physical medicine and rehabilitation. VIII. Preliminary data analysis.

This paper describes important aspects of preliminary data analysis to be taken after data are checked for clerical entry errors and before the primary statistical analysis is performed. These include description and graphic display of each variable, recoding categorical data, transforming continuous data into another continuous variable and recoding continuous to categorical data. Missing values and outlying data points are identified and several techniques are recommended to minimize mistakes in variable recoding. Related variables measured with different units may be combined by using the z transformation and converted back to one of the original units for ease of interpretation. Finally, both categorical and continuous variables are checked for reliability by using kappa or the intraclass R.

Data Collection

Interdisciplinary approach to assessing the health risk of air toxic chemicals: an overview.

To assist the regulatory branch of the Environmental Protection Agency in addressing the risk assessment of air toxics, the Health Effects Research Laboratory initiated a comprehensive inhalation toxicology program to provide key health effects data missing from the current data base. A priority ranking of chemicals based on the potential for substantial human exposure and the need for health effects data was developed to identify candidate chemicals for toxicological research. The major goal of the program is to evaluate the concentration-response from acute, intermittent and subchronic inhalation exposures to developmental, genetic, hepatic, immunologic, neurologic, pulmonary and reproductive toxicity in a manner that provides data for the regulatory health assessment of air toxic chemicals. Extrapolation and dosimetry research is also conducted to improve the basis for human risk assessment. Determination of biological endpoints to be examined will be decided on a compound-by-compound basis, depending on the physical, chemical and structural characteristics of the chemical and evaluation of the existing health data base. Although the main emphasis is on inhalation as the primary route of exposure, some of the laboratories will compare inhalation to other routes, such as oral, to better understand the influence of route of exposure and hence the potential applicability of existing health data. Acute and intermittent exposures will be done for all compounds. Upon evaluation of the acute results, a decision will be made as to whether subchronic studies are needed. Endpoints that show unusual sensitivity may be investigated in greater detail. The total length of exposure will vary from 1 to 21 days. The daily length of exposure will range from 1 to 8 hr. If adverse effects are observed at ambient levels, the time to recovery after exposure will be investigated.

Administration, Inhalation

[Molecular phylogenies and nucleotide insertion-deletion].

Molecular trees based on the analysis of sequence data can be obtained through parsimony procedures. Such approaches identify evolutionary events (homologies and homoplasies and their location in the tree). Nevertheless, it is not possible generally to analyze directly the information given by the aligned sequences. Noticeably, gaps ("indel") need special coding. Standard procedures for coding gaps offer alternatives which suffer of several weaknesses. In this article a new strategy of coding the sequences is defined in view to express the potential phylogenetic information contained in complex zones with internested insertion/deletion and substitutions, contrary to what is done up to the present. This strategy applies without loss or distortion of information to any case where gaps are present in aligned sequences. According to the hierarchy of internested states of characters (sites), this strategy introduces in the data matrix question marks, "?", which are optimized in fine in the cladogram based on all data. These "?" are not missing data, they are methodological codes, neutral to a priori phylogenetic hypotheses.

Base Sequence

Hidden bias in the use of archival data.

Nonresponses in archival data may violate the missing-at-random assumption in ways difficult to detect. Standard methods of comparing sociodemographics of respondents and nonrespondents are inappropriate when the units of analysis are not also the individuals who maintain the archival record. Under these circumstances, the distribution of missing data may be correlated with the dependent variable and traits of the record keepers. This will distort relationships, especially when listwise deletion of missing values is used in multivariate analysis. Data are used from a large clinical chart study of mentally ill patients to demonstrate the process of identifying hidden bias and the implications of such bias.

Archives

Why are missing quality of life data a problem in clinical trials of cancer therapy?

Assessment of health related quality of life has become an important endpoint in many cancer clinical trials. Because the participants of these trials often experience disease and treatment related morbidity and mortality, non-random missing assessments are inevitable. Examples are presented from several such trials that illustrate the impact of missing data on the analysis of QOL in these trials. The sensitivity of different analyses depends on the proportion of assessments that are missing and the strength of the association of the underlying reasons for missing data with disease and treatment related morbidity and mortality. In the setting of clinical trials of cancer therapy, the assumption that the data are missing completely at random (MCAR) and analyses of complete cases is usually unjustified. Further, the assumption of missing at random (MAR) may also be violated in many trials and models appropriate for non-ignorable missing data should be explored. Recommendations are presented to minimize missing data, to obtain useful documentation concerning the reasons for missing data and to perform sensitivity analyses.

Breast Neoplasms

Methodological and statistical problems in sleep apnea research: the literature on uvulopalatopharyngoplasty.

A comprehensive review of the literature on the surgical treatment of sleep apnea found 37 appropriate papers (total n = 992) on uvulopalatopharyngoplasty (UPPP). Methodological and statistical problems in these papers included the following: 1) There were no randomized studies and few (n = 4) with control groups. 2) Median sample size was only 21.5; thus statistical power was low and clinically important associations were routinely classified as "not statistically significant". 3) Only one paper presented the confidence bounds that might distinguish between statistical and clinical significance. 4) Because of short follow-up time and infrequent repeat follow-ups, little is known about whether UPPP results deteriorate with time. 5) In at least 15 papers, bias caused by retrospective designs and nonrandom loss to follow-up raised questions about the generalizability of results. 6) Few papers associated polysomnographic data with patient-based quality of life measures. 7) Missing data and missing and inconsistent definitions were common. 8) Baseline measures were often biased because the same assessment was inappropriately but routinely used for both screening and baseline. We conclude that because of these and other problems, there is much that is needlessly unknown about UPPP. It is the responsibility of the research and professional communities to define training, editorial and review procedures that will raise the methodological and statistical quality of published research.

Humans

Nonlinear Time&hyphenSeries Prediction with Missing and Noisy Data

We derive solutions for the problem of missing and noisy data in nonlinear time&hyphenseries prediction from a probabilistic point of view. We discuss different approximations to the solutions &hyphen in particular, approximations that require either stochastic simulation or the substitution of a single estimate for the missing data. We show experimentally that commonly used heuristics can lead to suboptimal solutions. We show how error bars for the predictions can be derived and how our results can be applied to K&hyphenstep prediction. We verify our solutions using two chaotic time series and the sunspot data set. In particular, we show that for K&hyphenstep prediction, stochastic simulation is superior to simply iterating the predictor.

Journal Article

Multiple imputation of missing blood pressure covariates in survival analysis.

This paper studies a non-response problem in survival analysis where the occurrence of missing data in the risk factor is related to mortality. In a study to determine the influence of blood pressure on survival in the very old (85+ years), blood pressure measurements are missing in about 12.5 per cent of the sample. The available data suggest that the process that created the missing data depends jointly on survival and the unknown blood pressure, thereby distorting the relation of interest. Multiple imputation is used to impute missing blood pressure and then analyse the data under a variety of non-response models. One special modelling problem is treated in detail; the construction of a predictive model for drawing imputations if the number of variables is large. Risk estimates for these data appear robust to even large departures from the simplest non-response model, and are similar to those derived under deletion of the incomplete records.

Aged

Long-term survival data from a clinical trial on resin-bonded bridges.

OBJECTIVES: A clinical trial, involving 203 resin-bonded bridges (RBBs) was undertaken to investigate the influence of retainer-type and luting material on the survival of these restorations. METHODS: For this evaluation, 157 patients were available (14% of the original sample was lost to follow-up or excluded from the study following the stopping criteria). Fifty per cent of the patients were questioned concerning the fate of the RBBs and 59% of questioned patients were examined clinically. The patients that were seen for examination were representatives of the experimental groups. The findings from the clinical examination were compared with the data obtained from the questionnaire. Missing data were censored at the date of the last available information. Kaplan-Meier estimates were calculated to assess the survivals at the endpoints and compared using Cox's proportional hazards procedure. RESULTS: A significant difference was found between perforated (P-type) and etched (E-type) RBBs (P = 0.05) for original bonded restorations but not when rebonded RBBs were taken into account. The results of the survival analysis were: anterior P-type, 49 +/- 7% after 10.5 years: anterior E-type, 57 +/- 7% after 10.5 years; posterior P-type, 18 +/- 11% after 6.8 years; posterior E-type, 37 +/- 13% after 10.2 years. Survivals of RBBs that were rebonded once during the evaluation period were 62 +/- 9% (11.0 years) for anterior RBBs and 51 +/- 11% (10.2 years) for posterior RBBs. CONCLUSIONS: The factor location (anterior versus posterior) was as in previous analyses, highly significant. Differences in survival between cementation materials were not significant.

Acid Etching, Dental

Poor agreement of occupational data between a hospital-based cancer registry and interview.

With occupation recognized as a risk factor for various cancers, collecting occupation and industry data by a number of vital registries, including cancer registries, has developed. Registries may be data sources for cancer etiology research and occupational disease surveillance, despite concerns that their data are fragmentary and may lack validity. To improve completeness and validity of occupational information in a hospital-based cancer registry, this study compared information obtained through abstracting medical records for the registry with information obtained through lung-cancer patient interviews. Employing the kappa statistic, agreement was generally poor, largely due to data missing in the medical record. Data quality of hospital-based cancer registries can be improved by employing trained cancer registrars to elicit occupational histories from patients.

Data Collection

Time-series analysis--cosinor analysis: a special case.

Cosinor analysis provides an accessible means of evaluating and estimating the parameter of a cyclic phenomenon. Cosinor analysis does not require that the data be equal intervals without missing data. Cosinor analysis does require that the data can reasonably be considered to take the form of a deterministic cycle with a known period.

Humans

Use of a spreadsheet program for circadian analysis of biological/physiological data.

Biological/physiological data sampled over a period of 24 h can be subjected to a mathematical analysis to determine the presence of circadian rhythmicity. Several procedures have been proposed, most being complex. To render such an analysis simpler and easy to use by non-mathematicians, we developed and tested the cosinor technique using a commonly available commercial spreadsheet (Excel). It can be used to analyze equally or unequally time-spaced data over 24 h with missing data, as well as to calculate the significance and the main limit of the resultant circadian rhythm (mesor, amplitude, acrophase and their confidence limits). Examples of its application to hourly samples of plasma cortisol and minute-by-minute rectal temperatures are shown.

Body Temperature

Regression calibration in failure time regression.

In this paper we study a regression calibration method for failure time regression analysis when data on some covariates are missing or mismeasured. The method estimates the missing data based on the data structure estimated from a validation data set, a random subsample of the study cohort in which covariates are always observed. Ordinary Cox (1972; Journal of the Royal Statistical Society, Series B 34, 187-220) regression is then applied to estimate the regression coefficients, using the observed covariates in the validation data set and the estimated covariates in the nonvalidation data set. The method can be easily implemented. We present the asymptotic theory of the proposed estimator. Finite sample performance is examined and compared with an estimated partial likelihood estimator and other related methods via simulation studies, where the proposed method performs well even though it is technically inconsistent. Finally, we illustrate the method with a mouse leukemia data set.

Animals