PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “missing data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Two methods for recommending bat weights.

Baseball players swung very light and very heavy bats through our instrument and the speed of the bat was recorded. These data were used to make mathematical models for each person. Then these models were coupled with equations of physics for bat-ball collisions to compute the Ideal Bat Weight for each individual. However, these calculations required the use of a sophisticated instrument that is not conveniently available to most people. So, we tried to find items in our database that correlated with Ideal Bat Weight. However, because many cells in the database were empty, we could not use traditional statistical techniques or even neural networks. Therefore, three new methods were used to estimate the missing data: (i) a neural network was trained using subjects that had no empty cells, then that neural network was used to predict the missing data, (ii) the data patching facility of a commercial software package was used, and (iii) the empty cells were filled with random numbers. Then, using these fully populated databases, several simple models were derived for recommending bat weights.

Adolescent↗

Prevalence and patterns of same-gender sexual contact among men.

The prevalence and patterns of same-gender sexual contact among men are key components of models of the spread of HIV infection and AIDS in the U.S. population. Previous estimates by Kinsey et al. from data collected between 1938 and 1948 have been widely criticized for inadequacies of sample design. New lower-bound estimates of prevalence developed from data from a national sample survey conducted in 1970 indicate that minimums of 20.3 percent of adult men in the United States in 1970 had sexual contact to orgasm with another man at some time in life; 6.7 percent had such contact after age 19; and between 1.6 and 2.0 percent had such contact within the previous year. Although these estimates incorporate adjustments for missing data, the likelihood of underreporting suggests that these estimates might be lower bounds on the prevalence of same-gender sex among men. Two sets of alternative estimates are derived to assess the sensitivity of these estimates to the assumptions made in imputing values to missing data. Detailed estimates are presented by frequency of contact, age, education, and marital status; and supporting estimates are derived from a 1988 national survey. Data from both the 1970 and 1988 surveys indicate that never-married men are more likely than other men to have had same-gender sexual contacts within the last year. The 1970 survey also indicates, however, that approximately half the men estimated to have such contacts are found among the more numerous population of currently or previously married men.

Adult↗

Trends in trauma care in England and Wales 1989-97. UK Trauma Audit and Research Network.

BACKGROUND: In 1988, the Royal College of Surgeons reported major deficiencies in trauma care in UK hospitals. We investigated whether and how that care has changed in the last decade by use of data collected by the UK Trauma Audit and Research Network. METHODS: We analysed injury-severity, process, and outcome variables from 91602 patients' records on the database at the end of 1997, collected from 97 (49% of trauma-receiving) hospitals in England, Wales, and two in Ireland. We did longitudinal analyses of odds of death, process variables, and individual hospitals' performance. We took account of potential selection bias from missing data and recruitment of new hospitals. FINDINGS: The severity-adjusted odds of death after trauma declined gradually from 1989 (odds ratio 1997/1989 0.63 [95% CI [0.49-0.82]). In 1997, the reduction in odds of death was significant even after adjustment for missing data (ratio 1997/1989 0.72 [0.55-0.92]) and recruitment of new hospitals (0.64 [0.44-0.93]). There was significant variability in the proportion of survivors (adjusted for severity of injury and age) between the highest and lowest 10% of UK hospitals. The time between the call to the emergency services and arrival at hospital increased from 32 min in 1989 to 45 min in 1997, irrespective of injury severity. The proportion of severely injured patients seen first by senior doctors increased from 32% to 60%. INTERPRETATION: Hospital care has made a valuable but variable contribution to reductions in case fatality after injury in the UK in the past 10 years, though further improvement is possible.

Aged↗

Bayesian analysis of prevalence with covariates using simulation-based techniques: applications to HIV screening.

Ignoring the limited precision of medical diagnostic tests can incur serious bias in prevalence estimation. Conversely, treating the values of sensitivity and specificity as constants, as in most studies, inevitably underestimates the variability of prevalence estimates. Bayesian inference provides a natural framework with which to integrate the variability in the estimates of sensitivity and specificity with estimation of prevalence. However, the resulting model becomes quite complicated and presents a computational challenge. Recently, Mendoza-Blanco et al. proposed a missing-data approach with simulation-based techniques to deal with the computational difficulties. Although their approach is quite effective in reducing the computational complexity into manageable tasks, their developed methodology is not general enough for modelling the effects of covariates in prevalence estimation. In this paper, we extend their work in this direction by combining their missing-data approach with a latent variable technique for modelling discrete data. The present work also generalizes the methods of Albert and Chib for Bayesian analysis of binary response data with errors in the response. We illustrate the methodology with several real data examples extracted from the literature.

AIDS Serodiagnosis↗

Patient-assessed measures of health outcome in asthma: a comparison of four approaches.

The study compares the psychometric properties of four different approaches to patient-assessed health outcomes in asthma, including the Asthma Quality of Life Questionnaire (AQLQ), Newcastle Asthma Symptoms Questionnaire (NASQ), SF-12 and EuroQol. The instruments were administered by means of a self-completed postal questionnaire to 394 patients recruited from general practices in the North East of England. Patients completed a follow-up questionnaire at 6 months. The levels of missing data were assessed and instrument scores compared using correlational analysis. Scores were related to self-reports of smoking behaviour, socioeconomic status and health transition. Responsiveness was assessed using standardized response means. Two hundred and thirty-five patients took part in the study giving a response rate of 59.6%. There was a relatively large amount of missing data for the individualized section of the AQLQ. Correlational analysis provided evidence of convergent validity between the specific instruments; the largest correlation was found between NASQ scores and the asthma symptoms scale of the AQLQ (r = 0.84). The NASQ was found to be the most powerful at discriminating between smokers and non-smokers. All four instruments were linearly related to self-reported asthma transition (P<0.05); the specific instruments having the strongest association. The specific instruments showed good levels of responsiveness with the NASQ producing a large SRM of 0.82. SRMs for the AQLQ were of a moderate to large size (0.32-0.77) and the SRMs for the SF-12 and EuroQol were of a small size. The two specific instruments are capable of greater levels of discrimination between groups of patients and are more responsive to changes in health than the generic SF-12 and EuroQol. The greater responsiveness of the NASQ is probably due to its focus being restricted to symptoms of asthma compared to the broader focus of the AQLQ domains. The NASQ has a strong relationship with the AQLQ and is a more practical instrument that is more acceptable to patients. However, the AQLQ does measure broader patient concerns. The SF-12 and EuroQol have greater potential to capture side-effects and have wider scope for application in economic evaluation.

Asthma↗

Modelling of chronic wound healing dynamics.

Following chronic wound area over time can give a general overview of wound healing dynamics. Decrease or increase in wound area over time has been modelled using either exponential or linear models, which are two-parameter mathematical models. In many cases of chronic wound healing, a delay of healing process was noticed. Such dynamics cannot be described solely with two parameters. The reported study deals with two-, three-, and four-parameter models. Assessment of the models was based on weekly measurements of 226 chronic wounds of various aetiologies. Several quantitative fitting criteria, i.e. goodness of fit, handling missing data and prediction capability, and qualitative criteria, i.e. number of parameters and their biophysical meaning were considered. The median of goodness of fit of three- and four-parameter models was between 0.937 and 0.958, and the median of two-parameter models was 0.821 to 0.883. Two-parameter models fitted wound area over time significantly (p = 0.01) worse than three- and four-parameter models. The criterion handling missing data provided similar results, with no significant difference between three- and four-parameter models. Median prediction error of two-parameter models was between 111 and 746; three-parameter models resulted in an error of 64 to 128, and finally four-parameter models resulted in the highest prediction error of 407 and 238. Based on the values of quantitative fitting criteria obtained, three parameters were chosen as the most appropriate. Based on qualitative criteria, the delayed exponential model was selected as the most general three-parameter model. It was found to have good prediction capability and in this capacity it could be used to help physicians choose the most appropriate treatment for patients with chronic wounds after an initial three-week observation period, when the median error increase of fitting is 74%.

Chronic Disease↗

Applications of multiple imputation in medical studies: from AIDS to NHANES.

Rubin's multiple imputation is a three-step method for handling complex missing data, or more generally, incomplete-data problems, which arise frequently in medical studies. At the first step, m (> 1) completed-data sets are created by imputing the unobserved data m times using m independent draws from an imputation model, which is constructed to reasonably approximate the true distributional relationship between the unobserved data and the available information, and thus reduce potentially very serious nonresponse bias due to systematic difference between the observed data and the unobserved ones. At the second step, m complete-data analyses are performed by treating each completed-data set as a real complete-data set, and thus standard complete-data procedures and software can be utilized directly. At the third step, the results from the m complete-data analyses are combined in a simple, appropriate way to obtain the so-called repeated-imputation inference, which properly takes into account the uncertainty in the imputed values. This paper reviews three applications of Rubin's method that are directly relevant for medical studies. The first is about estimating the reporting delay in acquired immune deficiency syndrome (AIDS) surveillance systems for the purpose of estimating survival time after AIDS diagnosis. The second focuses on the issue of missing data and noncompliance in randomized experiments, where a school choice experiment is used as an illustration. The third looks at handling nonresponse in United States National Health and Nutrition Examination Surveys (NHANES). The emphasis of our review is on the building of imputation models (i.e. the first step), which is the most fundamental aspect of the method.

Acquired Immunodeficiency Syndrome↗

The effect of psychological interventions on anxiety and depression in cancer patients: results of two meta-analyses.

The findings of two meta-analyses of trials of psychological interventions in patients with cancer are presented: the first using anxiety and the second depression, as a main outcome measure. The majority of the trials were preventative, selecting subjects on the basis of a cancer diagnosis rather than on psychological criteria. For anxiety, 25 trials were identified and six were excluded because of missing data. The remaining 19 trials (including five unpublished) had a combined effect size of 0.42 standard deviations in favour of treatment against no-treatment controls (95% confidence interval (CI) 0.08-0.74, total sample size 1023). A most robust estimate is 0.36 which is based on a subset of trials which were randomized, scored well on a rating of study quality, had a sample size > 40 and in which the effect of trials with very large effects were cancelled out. For depression, 30 trials were identified, but ten were excluded because of missing data. The remaining 20 trials (including six unpublished) had a combined effect size of 0.36 standard deviations in favour of treatment against no-treatment controls (95% CI 0.06-0.66, sample size 1101). This estimate was robust for publication bias, but not study quality, and was inflated by three trials with very large effects. A more robust estimate of mean effect is the clinically weak to negligible value of 0.19. Group therapy is at least as effective as individual. Only four trials targeted interventions at those identified as at risk of, or suffering significant psychological distress, these were associated with clinically powerful effects (trend) relative to unscreened subjects. The findings suggest that preventative psychological interventions in cancer patients may have a moderate clinical effect upon anxiety but not depression. There are indications that interventions targeted at those at risk of or suffering significant psychological distress have strong clinical effects. Evidence on the effectiveness of such targeted interventions and of the feasibility and effects of group therapy in a European context is required.

Anxiety↗

A randomised controlled trial of postal versus interviewer administration of a questionnaire measuring satisfaction with, and use of, services received in the year before death.

STUDY OBJECTIVES: To develop a short form of an interview schedule used successfully in previous national surveys of care for the dying, and to investigate the effect of administering it by post on response rate, response bias and on the nature of responses to questions. DESIGN: Randomised controlled trial. SETTING: An inner London health authority. PARTICIPANTS: Informants (person registering death) of random sample of cancer deaths between June 1995 and July 1996. MAIN RESULTS: The shortened questionnaire (VOICES) has 158 questions. Response rate did not differ significantly between postal and interview groups (interview; 56% (69 of 123), postal: 52% (161 of 308). Responders in the two groups did not differ in terms of their sociodemographic characteristics. Postal questionnaires had significantly more missing data, particularly on questions about service provision and satisfaction with services. Responses to questions differed between the groups on 11 of 158 questions. Interview group respondents were more likely to give top ranking responses to questions on service satisfaction and symptom control. CONCLUSIONS: Postal questionnaires are an acceptable alternative to interviews in retrospective post-bereavement surveys of care for the dying, at least in terms of response rate and response bias. However, the increased costs of interview surveys need to be balanced against the fact that postal questionnaires result in more missing data, and possibly less reliable answers to some questions. Caution is needed in combining results from the two data collection methods as interview respondents gave more positive answers to some questions.

Adult↗

Assessing inner-city patients' hospital experiences. A controlled trial of telephone interviews versus mailed surveys.

OBJECTIVES: Obtaining accurate and representative patient-centered data may be difficult among poor, inner-city patients because of changing addresses, variable access to telephones, and a higher prevalence of illiteracy than in the populations in which many survey instruments were developed and tested. Assumptions about the usefulness of mailed surveys versus telephone interviews may not hold for the urban poor. Therefore, identifying the most efficient mode of survey administration in this population becomes an important methodological question. METHODS: We conducted a randomized trial of patients discharged from the inpatient medicine service of an urban teaching hospital to compare telephone interview with mailed self-administration of a detailed instrument for measuring patients' experiences with hospital care. Our primary outcomes were response rate, missing data, and data collection costs. Patients were excluded if they were not discharged to home or were mentally or physically unable to complete mailed or telephone interviews. The research assistant contacted eligible patients while hospitalized, informed them of the postdischarge survey, and obtained current phone numbers and addresses. Patients then were randomized to receive a 116-item satisfaction survey via one of two survey methods: mail-first (mailed surveys with follow-up on nonrespondents by telephone) or telephone-first (telephone interviews with follow-up of nonrespondents by mail). RESULTS: Of the 252 patients enrolled, 130 were randomized to the mail-first and 122 to the telephone-first method. Response rates were higher with the telephone-first (73%) compared with the mail-first method (50%; P < 0.0001). Surveys obtained by the telephone-first method had fewer missing data (0.7 +/- 2.39) for those items not involved in skip patterns compared with the mail-first method (7.1 +/- 12.3; P < 0.001) and were 42% less expensive per completed survey ($26.32 versus $37.35; P < 0.0001). CONCLUSIONS: In this survey of patients served by an urban teaching hospital, a strategy of telephone interviews with mail follow-up proved less expensive and yielded a higher response rate with more complete data than using a method where mailed surveys were followed by back-up telephone interviews. In addition, we believe that the improved response rate for telephone interviews compared with those reported in the literature for similar populations is the result of informing inpatients of the survey and obtaining telephone numbers and addresses in the hospital.

Female↗

Digital Mindfulness Intervention for Pregnant Women With Affective Disorders and Acute Stress Reactions: Prespecified Secondary Analysis of a Randomized Controlled Trial.

BACKGROUND: Pregnant women with ICD-10 (International Statistical Classification of Diseases, Tenth Revision) affective or stress-related disorders face an elevated risk of perinatal depression and anxiety, yet evidence on digital nonpharmacologic interventions for this population remains limited. OBJECTIVE: This study evaluated the effectiveness of an 8-week digital mindfulness-based intervention (eMBI) compared with treatment as usual (TAU) among pregnant women with ICD-10 affective or stress-related disorders participating in a randomized controlled trial (RCT). METHODS: This prespecified secondary analysis was conducted within a multicenter RCT in Baden-W&#xfc;rttemberg, Germany. Pregnant women aged 18 years and older with elevated depressive symptoms (Edinburgh Postnatal Depression Scale [EPDS]>9) and ICD-10-diagnosed affective or stress-related disorders were randomized 1:1 to eMBI or TAU. The intervention consisted of 8 weekly app-based mindfulness sessions (45 min each) delivered during gestational weeks 29-36, with no direct therapist contact. The primary outcome was continuous depressive symptom severity measured with the EPDS at 4-6 weeks post partum. Secondary outcomes included the EPDS at 6 months post partum, generalized anxiety (State-Trait Anxiety Inventory-State [STAI-S], State-Trait Anxiety Inventory-Trait [STAI-T]), and Pregnancy-Related Anxiety Questionnaire-Revised (PRAQ-R). Analyses followed the intention-to-treat (ITT) principle, using mixed models for repeated measures and multiple imputation. RESULTS: Of the 5299 screened women, 147 met the inclusion criteria for this subgroup analysis (intervention group [IG] had n=73 women and control group had n=74 women). Groups were comparable at baseline. The IG showed significantly greater reductions in EPDS scores at gestational week 34 (&#x394;=-2.21, P=.01), week 36 (&#x394;=-3.25, P=.01), and 4-6 weeks post partum (&#x394;=-4.81, P=.007). Treatment effects remained robust under conservative missing-data assumptions. At 4-6 weeks post partum, a higher proportion of participants in the IG achieved clinically meaningful improvement (31/73, 42.5% vs 21/74, 28.4%; adjusted odds ratio 1.56, 95% CI 1.19-2.05; P=.001). Anxiety outcomes followed a similar pattern, whereas pregnancy-related anxiety did not differ between groups. CONCLUSIONS: In this prespecified subgroup of pregnant women with ICD-10 affective or stress-related disorders, the eMBI was associated with clinically meaningful reductions in depressive symptoms from late pregnancy to 4-6 weeks post partum. Effects at 6 months post partum were attenuated and less stable across missing-data assumptions. These findings support eMBIs as a scalable, nonpharmacological adjunct to perinatal mental health care for women with affective or stress-related disorders, while confirmation in adequately powered trials with strategies to reduce postpartum attrition is warranted.

Humans↗

Performance characteristics of a composite multivariate quality control system.

We present the results of an evaluation of the performance characteristics of a composite multivariate quality control (CMQC) system that incorporates quality control rules for univariate, multivariate, and correlation conditions. The CMQC system evaluated is designed to help analysts detect unacceptable trends and systematic error in one or more variables, unacceptable random error in one or more variables, and unacceptable changes in the correlation structure of any pair of variables. It is also designed to be tolerant of missing data, to allow analysts to reject as few as one or as many as all variables in a run, and to provide analysts with control statistics and graphics that logically relate to sources of analytical error. We show that the various components of the CMQC system have adequate statistical power to detect systematic errors, random errors, and correlation changes under the conditions likely to be encountered with multivariate analytical measurement systems: (1) a single variable with increased systematic or random error; (2) all variables or a subgroup of variables affected by a common problem that increases systematic or random error; and (3) missing data for one or more variables in a run. We also show that the power of the multivariate component of the CMQC system to detect systematic and random errors is higher than the power of an alternative multivariate test criterion.

Chemistry Techniques, Analytical↗

Evaluation of the quality of life in dementia with a generic quality of life questionnaire: the Duke Health Profile.

OBJECTIVE: The study was designed to determine the acceptability, feasibility and validity of measuring quality of life in a representative sample of dementia patients with a generic instrument, the Duke Health Profile. METHOD: The French version of the Duke Health Profile was administered to 148 subjects with a mental disorder according to the DSM-III-R diagnostic criteria. The feasibility and acceptability of employing the instrument were determined by the refusal rate, the type of administration, and the percentage and distribution of missing data. Reliability was determined with Cronbach's alpha coefficient. Instrument reproducibility was assessed with the intraclass correlation coefficient for test-retest values. Internal construct validity was determined by factor analysis. Discriminant capacity was determined by comparing the average scores on each measure among patients with and without an additional chronic pathology. The measurements obtained were compared by source of information (patient, family proxy and care provider proxy). RESULTS: The feasibility and acceptability of the instrument was good. Only 2% of the patients refused to complete the questionnaire. Help from the interviewer was necessary in 79% of the cases. The average completion time was 10.6 min. Missing data exist in only 3.5% of the cases on average, except among patients with severe dementia (Mini Mental State Examination <10). For reliability, internal consistency was acceptable (Cronbach's coefficient alpha = 0.5--0.7) when the self-esteem (0.23) and social health (0.26) concepts were eliminated. Reproducibility as measured by test-retest scores was moderate to good (intraclass correlation coefficient r = 0.53--0.80), except for anxiety (0.48) and perceived health (0.45). Severity of dementia mainly affected the feasibility, acceptability and reproducibility of the instrument. The family proxy seemed to agree more with the patient than did the care provider proxy. CONCLUSION: Quality of life can be measured in patients with dementia, but special tools need to be developed for severe dementia.

Aged↗

Assessing joint pain complaints and locomotor disability in the Rotterdam study: effect of population selection and assessment mode.

OBJECTIVE: To assess the prevalence of self-assessed and physician-assessed disability and joint pain, their association, and the effect of cohort reduction and mode of assessment. DESIGN: Cross-sectional population survey. SETTING: General population, age 55 years and older. SUBJECTS: Independently living participants of the Rotterdam Study, including 1,156 men and 1,739 women. OUTCOME MEASURES: Self-reported and physician-assessed joint complaints. Patients' self-assessment of locomotor disability was by response to questions from the Stanford Health Assessment Questionnaire; physicians assessed patients' disability by administering activity tests. RESULTS: Reduction of the study cohort because of nonresponse and missing data had no influence on the frequency and effect measures. The physician-assessed prevalence of pain of the hips, knees, or feet was significantly lower than the self-assessed prevalence, with the percentage agreement being 83% for men and 74% for women, with kappa-values of approximately .40. The prevalence of physician-assessed locomotor disability was also significantly lower than the self-assessed disability, with the percentage agreement being 83% for men and 78% for women, with kappa values of .41 and .47, respectively. The associations of joint complaints with disability were similar for both modes of assessment. CONCLUSION: Cohort reduction caused by nonresponse and missing data had no influence on estimates of frequency and association. Self-assessment gives higher prevalences of joint complaints and locomotor disability than physician assessment, but the associations between complaints and disability were the same.

Aged↗

A field-compatible method for interpolating biopotentials.

Mapping of bioelectric potentials over a given surface (e.g., the torso surface, the scalp) often requires interpolation of potentials into regions of missing data. Existing interpolation methods introduce significant errors when interpolating into large regions of high potential gradients, due mostly to their incompatibility with the properties of the three-dimensional (3D) potential field. In this paper, an interpolation method, inverse-forward (IF) interpolation, was developed to be consistent with Laplace's equation that governs the 3D field in the volume conductor bounded by the mapped surface. This method is evaluated in an experimental heart-torso preparation in the context of electrocardiographic body surface potential mapping. Results demonstrate that IF interpolation is able to recreate major potential features such as a potential minimum and high potential gradients within a large region of missing data. Other commonly used interpolation methods failed to reconstruct major potential features or preserve high potential gradients. An example of IF interpolation with patient data is provided to illustrate its applicability in the actual clinical setting. Application of IF interpolation in the context of noninvasive reconstruction of epicardial potentials (the "inverse problem") is also examined.

Action Potentials↗

Data mining issues for improved birth outcomes.

Issues obstructing progress in data mining for improved health outcomes include data quality problems, data redundancy, data inconsistency, repeated measures, temporal (time-contextual) measures, and data volume. Related issues involve theoretical and technical problems involving uncertainty management, missing data and missing values, and matching appropriate data mining techniques to patient data sets. Results of data mining research in progress are reported for Duke University's perinatal database that contains nearly a decade of clinical patient data, 71,753 database (patient) records and 4-5000 variables per patient.

Artificial Intelligence↗

[Study on distribution form of mesiodistal crown diameter in large sample: Part II].

The purpose of this research was to examine the distribution of the tooth size in a large sample. The objective teeth were the left upper and lower fourteen teeth except the third molar. The tooth size of 1,000 dental casts from the Japanese female orthodontic patients was measured. On each of them, a histogram and a set of statistics (mean, standard deviation, coefficient of variation, skewness, kurtosis, Geary value) are given in order to examine the distribution. The findings are as follows: 1) Each tooth may be classified into the following four types of distribution except the congenitally missing data. TYPE I: A normal distribution was observed in the upper and lower central incisors, the lower lateral incisor, the lower canine, the upper and lower first premolars, the upper second premolar, the upper and lower first molars and the lower second molar. TYPE II: A positively skewed distribution was observed in the lower second premolar. TYPE III: A negatively skewed and leptokurtic distribution was observed in the upper canine and the upper second molar. TYPE IV: An extremely negatively skewed and leptokurtic distribution was observed in upper lateral incisor. 2) With the four teeth which were classified into TYPE II, TYPE III and TYPE IV, the distribution of the lower second premolar was concluded to be of normal distribution by logarithmic transformation. The distribution of the upper canine and the upper second molar was judged to be of lognormal distribution and the upper lateral incisor also was judged to be of three parameter lognormal distribution and four parameter lognormal distribution. 3) The distribution of thirteen teeth except the upper lateral incisor was judged to be of normal distribution, by considering the congenitally missing data and the outlier in statistical data of the tooth.

Asian People↗

Pre-natal blood lead levels and learning difficulties in children: an analysis of non-randomly missing categorical data.

This paper presents an analysis of categorical variables subject to non-response. We incorporate the incomplete data into the analysis by modelling the distribution of the variables of interest and the non-response mechanism. We discuss issues of model selection and interpretation and the effect of discarding incomplete observations. In addition, we describe how to perform all of the computations with standard statistical software. We discuss the problem of incomplete categorical data within the context of a study of the effect of lead exposure on learning difficulties in children. In this study, many of the children are not observed on some of the variables of interest. It is particularly important in this study to incorporate the incomplete data, since there is evidence that non-response is related to the variables of interest. We reach different conclusions when we incorporate the incomplete data into the analysis than we reach when we discard the incomplete data. We also examine the sensitivity of our conclusions to the choice of a model for the non-response mechanism.

Algorithms↗