PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “missing data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Use of audio-enhanced personal digital assistants for school-based data collection.

PURPOSE: To review the different data collection options available to school-based researchers and to present the preliminary findings on the use of audio-enhanced personal digital assistants (APDA) for use in school-based data collection. METHODS: A newly developed APDA system was used to collect baseline data from a sample of 645 seventh grade students enrolled in a school-based intervention study. Evaluative measures included student response, time to completion, and data quality (e.g., missingness, internal consistency of responses). Differences in data administration and data quality were examined among three groups of students: students newer to the United States speaking English as a second language; special education students; and students not newer to the United States receiving regular education. RESULTS: The APDA system was well received by students and was shown to offer improvements in data administration (increased portability, time to completion) and reduced missing data. Although time to completion and proportion of missing data were similar across the three groups of students, psychometric properties of the data varied considerably. CONCLUSIONS: The APDA system offers a promising new method for collecting data in the middle school environment. Students with cognitive deficits and language barriers were able to complete the survey in a similar amount of time without additional help; however, differences in data quality suggest that limitations in comprehension of the questions remained even though the questions were read to the respondents. More research on the use of APDA is necessary to fully understand the effect of data collection mode with special populations.

Adolescent↗

Assessing the impact of chemotherapy-induced nausea and vomiting on patients' daily lives: a modified version of the Functional Living Index-Emesis (FLIE) with 5-day recall.

BACKGROUND: The Functional Living Index-Emesis (FLIE), a patient-reported outcome measure, was originally developed to assess the impact of chemotherapy-induced nausea and vomiting (CINV) on patients' daily lives over the 3 days following chemotherapy. More recent studies of CINV include assessments covering the 5 days following chemotherapy in an effort to capture information during both the acute (within 24 h) and delayed (up to 5-7 days) phases of CINV. GOALS: To assess the measurement characteristics of a modified version of the FLIE with 5-day recall. Instrument reliability, validity, and missing data were assessed. PATIENTS AND METHODS: Data were collected from 183 patients receiving cisplatin >or=70 mg/m(2) as part of a phase IIb antiemetic trial of an NK-1 receptor antagonist (MK-0869). Patients recorded the number of vomiting episodes and nausea ratings in a 5-day daily diary. RESULTS: The 5-day FLIE had: (1) excellent internal consistency within FLIE Nausea and Vomiting domains (Cronbach's alpha 0.77-0.78), (2) acceptable construct validity shown by FLIE item-total correlations stronger within domains ( r=0.74-0.97) than across domains ( r=0.52-0.76), and (3) acceptable convergent validity as shown by moderate to strong correlations between FLIE domain scores and independent endpoints of emetic episodes, nausea ratings, and use of rescue medications. The extent of missing data was within acceptable limits with less than 2% of patients missing data. CONCLUSION: The 5-day FLIE had adequate measurement characteristics for studying the impact of CINV on patients' daily lives during the period covering both the acute and delayed phases.

Activities of Daily Living↗

Methods to analyse cost data of patients who withdraw in a clinical trial setting.

BACKGROUND: Missing data resulting from premature study withdrawal are a common problem in the analysis of longitudinal data in clinical trials. To date, this subject has received little attention in the context of economic evaluations and with regard to the analysis of cost data. OBJECTIVES: To (i) demonstrate the impact of patients who drop out during the study on the outcomes of an economic evaluation, and (ii) to compare the mean and variation in costs after applying five different methods to deal with incomplete data: multiple imputation, complete cases analysis, linear extrapolation, predicted mean and hot decking. STUDY DESIGN: The study was performed using cost data collected in two randomised clinical trials comparing patients with chronic obstructive pulmonary disease receiving either tiotropium bromide or ipratropium bromide. The overall dropout rate was 17%, with the daily costs of the dropouts approximately 4 times higher than the costs of the completers. METHODS: Multiple imputation is a principled method that deals with missing observations by replacing each missing observation with a set of multiple plausible values. The variance between the resulting multiple datasets is combined with the variance between the datasets to take account of the extra uncertainty that results from missing data. The outcomes after multiple imputation were compared with the results of four naive methods to deal with missing observations: complete cases analysis, linear extrapolation, predicted mean and hot decking. All costs were expressed in 2001 euros. RESULTS: In the tiotropium bromide group, mean (standard error) costs varied from Euro 955 (137) after complete cases analysis to Euro 1298 (198) after linear extrapolation. The corresponding estimates in the ipratropium bromide group were Euro 970 (125) and Euro 1561 (244), respectively. The difference in costs between treatment groups varied from -Euro 15 (95% CI: -379 to 349) after complete cases analysis to -Euro 402 (95% CI: -883 to 79) after predicted mean, in favour of the tiotropium bromide group. The difference in costs according to the other methods varied from -Euro 263 (95% CI: -878 to 353) after linear extrapolation to -Euro 265 (95% CI: -709 to 180) after multiple imputation to -Euro 359 (95% CI: -771 to 54) after hot decking. CONCLUSION: This study showed that the method of dealing with the data of the dropouts had a large impact on the outcomes of an economic evaluation. Information about the rate of patient withdrawal and the way data of dropouts are treated is of vital importance in assessing the results of economic evaluations and should always be reported. Multiple imputation is a principled method that can be used to deal with the data of these patients.

Bronchodilator Agents↗

The multi-centre assessment of quality of life: the Interdisciplinary Group for Cancer Care Evaluation (GIVIO) experience in Italy.

One of the main issues to be considered in conducting clinical trials concerns the presence of missing data. This aspect is particularly relevant in oncology longitudinal studies, characterized by a long follow-up, and especially in quality of life studies where there is still little knowledge about patients' characteristics that predict loss of data. Since the middle of the 1980s the GIVIO (Interdisciplinary Group for Cancer Care Evaluation) co-operative group has been involved in conducting quality of life assessment studies, also focusing on the development of some strategies aimed at the minimization of missing data. In this paper we report on the results of two trials, which are now completed, concerning the quality of life assessment in a sample of breast and colon cancer patients. In order to cope with the problem of missing data, in both the trials the strategy of follow-up mailing was adopted, which proved to be an effective way to increase the response rate by nearly 50 per cent at each time point.

Adenocarcinoma↗

Marker assisted selection for the improvement of two antagonistic traits under mixed inheritance.

A Monte Carlo simulation was used to investigate the potential of Marker Assisted Selection (MAS) in a multiple-trait situation. Only additive effects were considered. The base population was assumed to be in linkage equilibrium and, next, the population was managed over 15 discrete generations, 10 males and 50 females were chosen out of the 100 candidates of each sex. Performance for two traits was simulated with an overall heritability of a given trait equal to 0.25 or 0.10 and the overall genetic correlation between traits was generally equal to -0.4 except in one case where it was equal to 0. The model involved one biallelic QTL, accounting for 10 or 20% of the genetic variance of a given trait, plus polygenes. Initial allelic frequencies at the QTL, were generally equal to 0.5 but in one case were equal to 0.1 and 0.9. A marker with 120 different alleles in the 60 founder parents was simulated in the vicinity of the QTL. Two values of the recombination rate between these two loci were considered, 0.10 and 0.02. The genetic evaluation was based on a multiple-trait BLUP animal model, accounting (MAS) or not (conventional BLUP) for marker information. Two sets of simulations were run: (1) a "missing data" case, with males having no record for one of the traits, and (2) a "secondary trait" case, with one trait having a weight in the aggregate genotype 4 times less than the other trait and the QTL acting only on this secondary trait. In the first set, evaluation methods were found to mainly affect the accuracy of overall genetic values prediction for the trait with missing data. In comparison with BLUP, MAS led to an extra overall genetic response for the trait with missing data, which was strongly penalised under the conventional BLUP, and to a deficit in response for the other trait. This more balanced evolution of the two traits was obtained, however, at the expense of the long-term overall cumulated response for the aggregate genotype, which was 1 to 2.5% lower than the one obtained under the conventional BLUP. In the second set of simulation, in the case of low initial frequency (0.1) of the QTL allele favourable to the secondary trait, MAS was found to be substantially more efficient to avoid losing this allele than BLUP only when the QTL had a large effect and the marker was close. More benefits should be expected from MAS with more specific applications,such as early selection of animals, or by applying dynamic procedures i.e. letting the respective weights to QTL and polygenic values in the selection criterion vary across generation.

Animals↗

Adjusted odds ratios for case-control studies with missing confounder data in controls.

Nonexperimental studies using computerized databases often give rise to missing or partially available information on confounders. A frequent situation occurs when data on exposure are available for all subjects of a case-control study, but data on confounders are available only for the cases but not for the controls. In that situation, the fact of confounding can be verified by assessing the association between exposure and a confounder in the cases, but the data are insufficient to produce an adjusted estimate of the relative risk if confounding is found to be present. We propose simple conditions under which an adjusted estimate of the relative risk can be obtained when data on a confounder are available only for the cases, and we derive formulae for the estimator and its confidence limits. The method requires an external estimate of the confounder prevalence or, additionally, of the confounder-exposure odds ratio. We illustrate the technique with data from a nested case-control study of the risk of acute cardiac death associated with the use of bronchodilator drugs within a cohort of 12,301 asthmatics, with smoking as the confounder of interest.

Asthma↗

Corticosteroids for treating Guillain-Barré syndrome.

BACKGROUND: The cause of Guillain-Barré syndrome (GBS) is inflammation of the peripheral nerves which corticosteroids would be expected to benefit. OBJECTIVES: To examine the efficacy of corticosteroids in hastening recovery and reducing the long term morbidity from Guillain-Barré syndrome (GBS). SEARCH STRATEGY: Search of the Cochrane Neuromuscular Disease Group register for randomised trials and enquiry from authors of trials and other experts in the field. SELECTION CRITERIA: Types of studies: quasi-randomised or randomised controlled trials TYPES OF PARTICIPANTS: patients with GBS of all ages and all degrees of severity Types of interventions: any form of corticosteroid or adrenocorticotrophic hormone Types of outcome measures: Primary: improvement in disability grade on a commonly used seven point scale four weeks after randomisation Secondary: time from randomisation until recovery of unaided walking, time from randomisation until discontinuation of ventilation (for those ventilated), mortality, proportion of patients dead or disabled (unable to walk without aid) after 12 months, improvement in disability grade after six months, improvement in disability grade after 12 months, proportion of patients who relapse, and proportion of patients with adverse events related to corticosteroid treatment DATA COLLECTION AND ANALYSIS: We identified six randomised trials. One author extracted the data and the other checked them. We obtained some missing data from investigators. MAIN RESULTS: The six eligible trials included a total of 195 corticosteroid treated patients and 187 control subjects. One trial of intravenous methylprednisolone accounted for 243 of the total 382 subjects studied (63%). This trial did not show a significant difference in any disability related outcome between the corticosteroid and placebo groups. There was no significant difference between the corticosteroid and control groups for the primary outcome measure, improvement in disability grade four weeks after randomisation. The weighted mean difference of the three trials for which this outcome was available showed no difference. The actual figure was 0.01 (95% CI -0.27 to 0.29) grade in favour of the corticosteroid group. There was also no significant difference between the groups for most of the secondary outcome measures. However in the largest trial hypertension developed less often in the intravenous methylprednisolone group (2/124, 1.6%) than in the control group (12/118, 10.2%), a significant difference in favour of corticosteroid treatment (relative risk 0.20, 95% CI 0.04 to 0.66). REVIEWER'S CONCLUSIONS: Corticosteroids should not be used in the treatment of Guillain-Barré syndrome. If a patient with Guillain-Barré syndrome needs corticosteroid treatment for some other reason its use will probably not do harm. The effect of intravenous methylprednisolone combined with intravenous immunoglobulin in Guillain-Barré syndrome is being tested with a randomised trial.

Anti-Inflammatory Agents↗

Predicting hospital mortality among injured children using a national trauma database.

PURPOSE: The purpose of this study was to develop a model that accurately predicts mortality among injured children based on components of the initial patient evaluation and that is generalizable to diverse acute care settings. Important predictive variables obtained in an emergency setting are frequently missing in even large national databases, limiting their effectiveness for developing predictions. In this study, a model predicting pediatric trauma mortality was developed using a national database and methods to handle missing data that may avoid biases that can occur restricting analyses to complete cases. METHODS: Records of pediatric patients included in the National Pediatric Trauma Registry (NPTR) between 1996 and 1999 were used as a training set in a logistic regression model to predict hospital mortality using vital signs, Glasgow Coma Scale (GCS) score, and intubation status. Multiple imputation was applied to handle missing data. The model was tested using independent data from the NPTR and National Trauma Data Bank (NTDB). RESULTS: Complete case analysis identified only GCS-eye and intubation status as predictors of mortality. A model based on complete case analysis had good discrimination (c-index = 0.784) and excellent calibration (Hosmer-Lemeshow c-statistic, 6.8) (p > 0.05). Using multiple imputation, three additional predictors of mortality (systolic blood pressure, pulse, and GCS-motor) were identified and improved model performance was observed. The model developed using multiple imputation had excellent discrimination (c-index, 0.947-0.973) in both test datasets. Calibration was better in the NPTR testing set than in the NTDB (Hosmer-Lemeshow c-statistic, 9.2 for NPTR [p > 0.05] and 258.2 for NTDB [p < 0.05]). At a probability cutoff that minimized misclassification in the training set, the false-negative and false-negative rates of the model were better than those obtained with either the Revised Trauma Score (RTS) or Pediatric Trauma Score using data from the NPTR testing set. Although the false-positive rates were lower with the RTS using data from the NTDB, the false-negative rates of the proposed model and the RTS were similar in this test dataset. CONCLUSIONS: Using multiple imputation to handle missing data, a model predicting pediatric trauma mortality was developed that compared favorably with existing trauma scores. Application of these methods may produce predictive trauma models that are more statistically reliable and applicable in clinical practice.

Child↗

A comparison of two methods for the estimation of precision with incomplete longitudinal data, jointly modelled with a time-to-event outcome.

Several methods for the estimation and comparison of rates of change in longitudinal studies with staggered entry and informative drop-outs have been recently proposed. For multivariate normal linear models, REML estimation is used. There are various approaches to maximizing the corresponding log-likelihood; in this paper we use a restricted iterative generalized least squares method (RIGLS) combined with a nested EM algorithm. An important statistical problem in such approaches is the estimation of the standard errors adjusted for the missing data (observed data information matrix). Louis has provided a general technique for computing the observed data information in terms of completed data quantities within the EM framework. The multiple imputation (MI) method for obtaining variances can be regarded as an alternative to this. The aim of this paper is to develop, apply and compare the Louis and a modified MI method in the setting of longitudinal studies where the source of missing data is either death or disease progression (informative) or end of the study (assumed non-informative). Longitudinal data are simultaneously modelled with the missingness process. The methods are illustrated by modelling CD4 count data from an HIV-1 clinical trial and evaluated through simulation studies. Both methods, Louis and MI, are used with Monte Carlo simulations of the missing data using the appropriate conditional distributions, the former with 100 simulations, the latter with 5 and 10. It is seen that naive SEs based on the completed data likelihood can be seriously biased. This bias was largely corrected by Louis and modified MI methods, which gave broadly similar estimates. Given the relative simplicity of the modified MI method, it may be preferable.

Algorithms↗

[Meta-analysis of the Italian studies on short-term effects of air pollution].

BACKGROUND: In recent years, much attention has been given to review reports on the early effects of air pollution on health, measured through daily series of deaths and/or hospital admissions. A number of large planned meta-analyses (in which methods for data retrieval and processing are commonly planned a priori for all participating centers) are on going both in the US and in Europe. The National Mortality, Morbidity and Air Pollution Study included data from 90 US cities, whereas APHEA (Air Pollution and Health, a European Approach) considers data from about 30 european cities. The present paper summarizes methods and findings of MISA, a meta-analysis of data from 8 Italian cities. It belongs to an ad hoc supplement of Epidemiologia & Prevenzione (Epidemiol Prev 2001; 25 (2) Suppl: 1-72), the official Journal of the Italian Association of Epidemiology, which contains a full description of the study. MISA was launched on March 2000, within the project "Statistics, Environment and Health" (GRASPA), funded by the Italian Ministry of Education. Additional support was given by the Authorities of the 8 participating cities (from North to South: Turin, Milan, Verona, Ravenna, Bologna, Florence, Rome and Palermo). DAILY HEALTH DATA: Deaths certificate and hospital admission data have been collected respectively from the Local Health Authority and regional files. The same programme for retrieval of data on selected hospital admissions for acute conditions was used in the 8 cities. Main data are summarized in Table 1. DAILY CONCENTRATION OF POLLUTANTS: Most data were obtained from Regional Environmental Protection Agencies, which are responsible for environmental monitoring since 1993. Verona, Palermo and Milan (1990-94) data were obtained from local sources. Monitors with more than 25% of missing data were excluded. Meteorological data were collected by the same monitors and completed with data from monitors situated in the suburbs or (in Milan and Bologna) in the airport. The monitors were selected by a group of experts to ensure comparability. For SO2 and NO2 daily averages of hourly measurements were used, whereas concentrations of ozone and CO were estimated as the maximum 8 hours moving average. Total suspended particulate or PM10 were measured as 24 hours deposition. All analyses used the whole range of observed values (Table 2). Daily data were considered as missing when more than 25% of hourly data were not available. Missing data in one monitor were imputed as average of data from the remaining monitors weighted by the ratio between the specific monitor's year average and the general year average of all the selected city monitors. Missing data in one day were imputed as average of four days (preceding and following day, the same day of the previous and following weeks). In the city of Florence and Palermo PM10 concentrations were available. For the other cities we applied a conversion factor from PTS to PM10 (0.6 for Turin and 0.8 for all the others) estimated through validation studies. Ozone concentrations were used only where background monitors were available (Turin, Verona, Bologna and Florence) and limited to the warm season (May through September). METHODS: A common protocol for the city-specific analyses was defined on the basis of a structured exploratory analysis. The adopted basic model was a Generalized Additive Model for Poisson data. Effect estimates were age-adjusted (0-64, 65-74, 75+) and formal tests of interaction pollutant-age were conducted. In the first two age groups, indicator variables for seasonality were specified, and cubic splines with fixed number of degree of freedom were specified for the last age group and for all age groups for the morbidity data. Model adequacy was checked by residual analysis and inspection of the partial autocorrelation function. In a sensitivity analysis non linear pollutant effects were considered and overdispersed [table: see text] transitional models were fitted; the analysis was conducted for all lags 0-3 and some distributed lags (0-1, 1-2, 0-3); no multipollutant models were fitted. The same model was fitted to the city data. No model selection was done: Table 3 describes the steps in model building. In the meta-analysis, for each outcome, the estimates for each pollutant and for each city were combined using fixed and random effects models. Heterogeneity of effects was tested according to DerSimonian and Laird. Results were checked using a hierarchical bayesian model, which was used to investigate heterogeneity across cities in a meta-regression phase. Non informative priors were used. Posterior distributions of parameters of interest have been obtained with WinBUGS. 10,000 iterations (excluding [table: see text] the first 2000) were retained, while for the meta-regression 100,000 iterations (excluding the first 4000) were stored. To approximate the marginal posteriors only one sample out of five were used. Achieved convergence was assessed using the Gelman and Rubin approach. In the meta-regression the models specified were the following: [formula: see text] i denotes city, j calendar period (1990-1994; 1995-1999). The first model includes only period as effect modifier, while the second model other potential variables. The ui terms (which do not vary with j) represent city specific random effects. RESULTS: For each pollutant, the meta-analysis detected a statistically significant association with mortality for natural causes. But for ozone, positive associations were commonly found for death and hospital admissions for both cardiovascular and respiratory diseases. Indeed, the only estimates whose lower 95% confidence limit bore a negative sign regarded the association between PM10 and mortality from respiratory diseases. Ozone in the warm season was positively and significantly associated with daily mortality and mortality for cardiovascular diseases whereas other estimates did not reach statistical significance and some were negative (only lag 0-1 for external comparability are reported in Table 4). Risks were highest (up to 4%) for respiratory conditions (Table 4). They were more pronounced at lag 1-2 for mortality, and at lag 0-3 for hospital admissions. Age was an effect modifier for mortality, the elderly being more susceptible. In the random effect meta-analysis, at lag 1-2, excess risks for unit increase of the pollutants at age 75+ and at age 0-64 were respectively: 4.9% and -0.4% for SO2, 1.7% and 0.6% for NO2; 2.3% and 0.2% for CO. Corresponding figures for PM10 at lag 0-1 were 1.1% and 0.2%. The effect of PM10 on mortality [table: see text] was greater during the warm season (2.8% vs 0.8%). A complete analysis is reported in the Italian text. Here we provide some details on the effects of PM10, about which the residual heterogeneity across cities was highest (Table 4). In addition, the epidemiological evidence on the hazards from this fraction of particulate matter is more controversial. Table 5 reports the excess risk estimated through the meta-analysis in 1995-99 for a 10 micrograms/m3 increase of PM10 for some outcomes. Proper prior distributions (overdispersed normal and inverse gamma) were adopted in the final bayesian analyses. The sensitivity of results to the choice of the priors were investigated (we defined proper and improper uniform, student's t), obtaining comparable results. Total natural mortality was significantly heterogeneous across cities (Q = 18.96, 5 df, p < 0.001). City-specific estimates are represented graphically in Fig. 1. As expected, the confidence (credibility) intervals are widest [table: see text] for bayesian estimates, intermediate for those obtained under a random effects model, and narrowest for those found under a fixed effects model. Nevertheless, differences in point estimates are negligible. A North-South gradient in risk is obvious. Table 6 shows, for the cities for which mortality data were available, the improvement in precision and the shrinkage of effect estimates toward the overall mean introduced by the bayesian modelling. In the meta-regression, total mortality and a deprivation score were associated with greater effects. The excess risks on hospital admission were modified by the deprivation score and by the NO2/PM10 ratio. Overall, the risk estimates were greater in the calendar period 1995-99 and there was a North-South gradient, with larger effects in cities located in Central and Southern Italy (Florence, Rome, Palermo). CONCLUSIONS: The meta-analysis of the Italian studies on short-term effects of air pollution in 8 cities, MISA, exhibits the following features: With the exception of Naples, all greatest Italian cities were included; overall a population of 7 million was enrolled. The study protocol was accurate with regard to the selection of hospital admissions for acute conditions. Monitored data of concentration of pollutant were carefully evaluated before their inclusion in the meta-analysis. City specific analyses were carried out according to a common protocol controlling for seasonality, influenza epidemics, age and meterological variables; [table: see text] the protocol derived from a structured exploratory analysis. The meta-analysis was done using fixed and random effects models; a hierarchical bayesian model was fitted in a sensitivity analysis. The heterogeneity of effects across cities was investigated using a hierarchical bayesian model for meta-regression. While mortality data are of good quality, hospital admission data are more problematic. Since the filing criteria for the latter changed around 1995, comparability of results before and after such date is limited. Moreover, hospital admissions rely on availability of beds, the offer of which may be restricted during the warm season. Comparability of pollutant concentration estimates among cities may have been influenced by differences in monitor characteristics. (ABSTRACT TRUNCATED)

Adolescent↗

Outcome prediction in trauma.

BACKGROUND: In the Trauma Audit and Research Network (TARN), currently the largest trauma network in Europe, outcome prediction is performed using the TRISS methodology since 1989. Its database contains 200,000 hospital admissions from 110 hospitals over the country, but a large amount of data is lost for the modelling because of missing data. To improve some of the shortcomings of TRISS a new model was developed. METHODS: The data for modelling consisted of 100,399 hospital trauma admissions over the period 1996 to 2001. Using the Glasgow Coma Score (GCS) instead of RTS has dramatically reduced the number of missing cases. Gender and its interaction with age have also been included in the model. The model was tested on different subsets of cases traditionally excluded, such as children, those with penetrating injuries, and ventilated and transferred patients. The new model included all those subsets using age, a transformation of ISS, GCS, gender and gender by age interaction as predictors. RESULTS: The model has shown a good discriminant ability tested by the area under the receiver operating characteristic (AROC) curve. The values of the AROC for the new model were 0.947 (95% CI: 0.943-0.951) on the prediction set and 0.952 (95% CI: 0.946-0.957) on the validation set compared respectively with 0.937 (95% CI: 0.932-0.943) and 0.941 (95% CI: 0.936-0.952) for TRISS. CONCLUSION: The new model has enabled us to include most of the cases that were excluded under the TRISS's inclusion criteria, less missing data are incurred and the predictive performance was significantly better than that of the TRISS model as shown by the AROC curves.

Adolescent↗

A comparative analysis of quality of life data from a Southwest Oncology Group randomized trial of advanced colorectal cancer.

Longitudinal quality of life measurements from an advanced-stage cancer clinical trial are analysed using a variety of methods, and the results compared. The methods used require different assumptions about the mechanism that produces the missing data. They include analyses that require the data to be missing completely at random; fixed-effects models and weighted generalized estimating equations, which require missing at random data; and a fully parametric approach where the outcomes and the missingness mechanism are jointly modelled, allowing non-ignorable missing data. The data show evidence of non-random missingness, but a formal test of non-ignorable missing data is not significant.

Colorectal Neoplasms↗

POSSUM, p-POSSUM, and Cr-POSSUM: implementation issues in a United States health care system for prediction of outcome for colon cancer resection.

PURPOSE: The Physiologic and Operative Severity Score for the enUmeration of Mortality and morbidity (POSSUM), Portsmouth revision (p)-POSSUM, and colorectal (Cr)-POSSUM scoring systems were developed as audit tools for comparing outcomes in surgical and colorectal patients on the basis of operative risk assessment. The aim of this study was to evaluate the applicability of these systems to a cohort of colon cancer patients undergoing surgery in the United States. METHODS: POSSUM factors from 890 consecutive patients undergoing major surgical procedures for colon cancer in nine United States hospitals over a two-year period from January 2000 through December 2001 were prospectively collected. The observed over the expected hospital mortality was compared by means of the POSSUM, p-POSSUM, and Cr-POSSUM scoring systems. The effect of missing data on the utility of this process for outcome assessment was assessed with three methods for data imputation. RESULTS: The number of resections per institution ranged from 13 to 437. The observed mortality rate ranged from 0.8 percent to 15.4 percent among the institutions, with an overall operative mortality of 2.3 percent. The POSSUM, p-POSSUM, and Cr-POSSUM predicted mortality was 10.7 percent, 11.2 percent, and 4.9 percent, respectively. The POSSUM and p-POSSUM models overpredicted mortality in all institutions ( P < 0.01), whereas the Cr-POSSUM demonstrated an observed over expected hospital mortality ratio of >1 in three institutions. The calculations were unaffected by the various methods of inserting missing data. CONCLUSION: An apparent overprediction of mortality for colon cancer resection was evident with all three POSSUM variants. This implies that a calibration process is required for use of these variants in the United States health care system. Missing data may be treated as normal values without influencing outcome. The Cr-POSSUM appeared to be the most promising audit tool for colorectal cancer surgery; however, it will require further refinement to provide process control graphs for identification of potential outliers and improvement in the quality of care in the United States.

Adolescent↗

On classification with incomplete data.

We address the incomplete-data problem in which feature vectors to be classified are missing data (features). A (supervised) logistic regression algorithm for the classification of incomplete data is developed. Single or multiple imputation for the missing data is avoided by performing analytic integration with an estimated conditional density function (conditioned on the observed data). Conditional density functions are estimated using a Gaussian mixture model (GMM), with parameter estimation performed using both Expectation-Maximization (EM) and Variational Bayesian EM (VB-EM). The proposed supervised algorithm is then extended to the semisupervised case by incorporating graph-based regularization. The semisupervised algorithm utilizes all available data-both incomplete and complete, as well as labeled and unlabeled. Experimental results of the proposed classification algorithms are shown.

Algorithms↗

Approaches to screening for intimate partner violence in health care settings: a randomized trial.

CONTEXT: Screening for intimate partner violence (IPV) in health care settings has been recommended by some professional organizations, although there is limited information regarding the accuracy, acceptability, and completeness of different screening methods and instruments. OBJECTIVE: To determine the optimal method for IPV screening in health care settings. DESIGN AND SETTING: Cluster randomized trial conducted from May 2004 to January 2005 at 2 each of emergency departments, family practices, and women's health clinics in Ontario, Canada. PARTICIPANTS: English-speaking women aged 18 to 64 years who were well enough to participate and could be seen individually were eligible. Of 2602 eligible women, 141 (5%) refused participation. INTERVENTION: Participants were randomized by clinic day or shift to 1 of 3 screening approaches: a face-to-face interview with a health care provider (physician or nurse), written self-completed questionnaire, and computer-based self-completed questionnaire. Two screening instruments-the Partner Violence Screen (PVS) and the Woman Abuse Screening Tool (WAST)-were administered and compared with the Composite Abuse Scale (CAS) as the criterion standard. MAIN OUTCOME MEASURES: The approaches were evaluated on prevalence, extent of missing data, and participant preference. Agreement between the screening instruments and the CAS was examined. RESULTS: The 12-month prevalence of IPV ranged from 4.1% to 17.7%, depending on screening method, instrument, and health care setting. Although no statistically significant main effects on prevalence were found for method or screening instrument, a significant interaction between method and instrument was found: prevalence was lower on the written WAST vs other combinations. The face-to-face approach was least preferred by participants. The WAST and the written format yielded significantly less missing data than the PVS and other methods. The PVS and WAST had similar sensitivities (49.2% and 47.0%, respectively) and specificities (93.7% and 95.6%, respectively). CONCLUSIONS: In screening for IPV, women preferred self-completed approaches over face-to-face questioning; computer-based screening did not increase prevalence; and written screens had fewest missing data. These are important considerations for both clinical and research efforts in IPV screening. TRIAL REGISTRATION: clinicaltrials.gov Identifier: NCT00336297.

Adult↗

Genetic Analysis Workshop 13: simulated longitudinal data on families for a system of oligogenic traits.

The Genetic Analysis Workshop 13 simulated data aimed to mimic the major features of the real Framingham Heart Study data that formed Problem 1, but under a known inheritance model and with 100 replicates, so as to allow evaluation of the statistical properties of various methods. The pedigrees used were the 330 real pedigree structures (comprising 4692 individuals) with some minor changes to protect confidentiality. Fifty trait genes and 399 microsatellite markers were simulated by gene dropping on 22 autosomal chromosomes. Assuming random ascertainment of families, a system of eight longitudinal quantitative traits (designed to be similar to those in the real data) was generated with a wide range of heritabilities, including some pleiotropic and interactive effects. Genes could affect either the baseline level or the rate of change of the phenotype. Hypertension diagnosis and treatment were simulated with treatment availability, compliance, and efficacy depending on calendar year. Nongenetic traits of smoking and alcohol were generated as covariates for other traits. Death was simulated as a hazard rate depending upon age, sex, smoking, cholesterol, and systolic blood pressure. After the complete data were simulated, missing data indicators were generated based on logistic models fitted to the real data, involving the subject's history of previous missing values, together with that of their spouses, parents, siblings, and offspring, as well as marital status, only-child indicators, current value at certain simulated traits, and the data collection pattern on the cohort into which each subject was ascertained.

Adult↗

Incomplete quality of life data in lung transplant research: comparing cross sectional, repeated measures ANOVA, and multi-level analysis.

BACKGROUND: In longitudinal studies on Health Related Quality of Life (HRQL) it frequently occurs that patients have one or more missing forms, which may cause bias, and reduce the sample size. Aims of the present study were to address the problem of missing data in the field of lung transplantation (LgTX) and HRQL, to compare results obtained with different methods of analysis, and to show the value of each type of statistical method used to summarize data. METHODS: Results from cross-sectional analysis, repeated measures on complete cases (ANOVA), and a multi-level analysis were compared. The scores on the dimension 'energy' of the Nottingham Health Profile (NHP) after transplantation were used to illustrate the differences between methods. RESULTS: Compared to repeated measures ANOVA, the cross-sectional and multi-level analysis included more patients, and allowed for a longer period of follow-up. In contrast to the cross sectional analyses, in the complete case analysis, and the multi-level analysis, the correlation between different time points was taken into account. Patterns over time of the three methods were comparable. In general, results from repeated measures ANOVA showed the most favorable energy scores, and results from the multi-level analysis the least favorable. Due to the separate subgroups per time point in the cross-sectional analysis, and the relatively small number of patients in the repeated measures ANOVA, inclusion of predictors was only possible in the multi-level analysis. CONCLUSION: Results obtained with the various methods of analysis differed, indicating some reduction of bias took place. Multi-level analysis is a useful approach to study changes over time in a data set where missing data, to reduce bias, make efficient use of available data, and to include predictors, in studies concerning the effects of LgTX on HRQL.

Adult↗

The BrainIT group: concept and core dataset definition.

INTRODUCTION: An open collaborative international network has been established which aims to improve inter-centre standards for collection of high-resolution, neurointensive care data on patients with traumatic brain injury. The group is also working towards the creation of an open access, detailed and validated database that will be useful for post-hoc hypothesis testing. In Part A, the underlying concept, the group coordination structure, membership guidelines and database access and publication criteria are described. Secondly, in part B, we describe a set of meetings funded by the EEC that allowed us to define a "Core Dataset" and we present the results of a feasibility exercise for collection of this core dataset. METHODS: Four group meetings funded by the EEC have enabled definition of a "Core Dataset" to be collected from all centres regardless of specific project aim. A paper based pilot collection of data was conducted to determine the feasibility for collection of the core dataset. Specially designed forms to collect the core dataset demographic and clinical information as well as sample the time-series data elements were distributed by both email and standard mail to 22 BrainIT centres. A deadline of two months was set to receive completed forms back from centres. A pilot data collection of minute by minute physiological monitoring data was also performed. FINDINGS: A core-dataset was defined and can be downloaded from the BrainIT web-site (go to "Core dataset" link at: www.brainit.org). Eighteen centres (82%) returned completed forms by the set deadline. Overall the feasibility for collection of the core data elements was high with only 10 of the 64 questions (16%) showing missing data. Of those 10 fields with missing data, the average number of centres not responding was 12% and the median 6%. An SQL database to hold the data has been designed and is being tested. Software tools for collection of the core dataset have been developed. Ethics approval has been granted for collection of multi-centre data as part of a pilot data collection study. INTERPRETATION: The BrainIT network provides a more standardised and higher resolution data collection mechanism for research groups, organisations and the device industry to conduct multi-centre trials of new health care technology in patients with traumatic brain injury.

Brain Injuries↗