PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Logistic Models”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Relation of maternal cocaine use to the risks of prematurity and low birth weight.

To determine whether maternal cocaine use at the time of delivery of the infant is an independent risk factor for low birth weight or prematurity, we performed a prospective anonymous urine toxicology screening study among 425 women in a large urban university-based maternity hospital. The data were subjected to univariate analysis with the Fisher Exact Test and odds ratio determination, and to multivariate analyses by logistic regression. Of 11 variables analyzed, cocaine use near delivery, no prenatal care, marijuana and cigarette use, black race, a previous preterm infant, and staff service were significantly associated with premature birth by univariate analysis. No prenatal care (odds ratio, 9.89; 95% confidence intervals, 3.74 to 26.17) and cocaine use (odds ratio, 7.31; 95% confidence intervals, 2.87 to 18.61) demonstrated the greatest risk associated with premature birth by univariate prediction. After analysis by multivariate logistic modeling, only cocaine use detected at birth remained a significant predictor of prematurity (odds ratio, 13.4; 95% confidence intervals, 1.23 to 145.0). Staff service, black race, cocaine use near the time of delivery, marijuana and cigarette use, a previous preterm infant, and no prenatal care were significant univariate predictors of low birth weight. Cocaine use (odds ratio, 4.14; 95% confidence intervals, 1.18 to 14.56) and marijuana use (odds ratio, 4.52; 95% confidence intervals, 1.42 to 14.39) were the strongest univariate factors. After analysis by multivariate logistic modeling, cocaine use near the time of delivery demonstrated the highest odds ratio (9.90) for predicting low birth weight, but the 95% confidence intervals included 1 (0.53 to 184.0). We conclude that independent of potentially interrelated covariables, a positive result on a cocaine urine toxicology test at the time of delivery is the most dominant factor that was tested to predict prematurity and possibly low birth weight. The effect of cocaine on the duration of gestation or fetal growth may be due to its pharmacologic properties, or cocaine use during pregnancy may identify a subgroup of women whose risk is due to as-yet-unidentifiable socioeconomic or cultural characteristics.

Analysis of Variance↗

Logit models and logistic regressions for social networks: II. Multivariate relations.

The research described here builds on our previous work by generalizing the univariate models described there to models for multivariate relations. This family, labelled p*, generalizes the Markov random graphs of Frank and Strauss, which were further developed by them and others, building on Besag's ideas on estimation. These models were first used to model random variables embedded in lattices by Ising, and have been quite common in the study of spatial data. Here, they are applied to the statistical analysis of multigraphs, in general, and the analysis of multivariate social networks, in particular. In this paper, we show how to formulate models for multivariate social networks by considering a range of theoretical claims about social structure. We illustrate the models by developing structural models for several multivariate networks.

Humans↗

Modeling the synergistic effect of high pressure and heat on inactivation kinetics of Listeria innocua: a preliminary study.

The survival curves of Listeria innocua CDW47 by high hydrostatic pressure were obtained at four pressure levels (138, 207, 276, 345 MPa) and four temperatures (25, 35, 45, 50 degrees C) in peptone solution. Tailing was observed in the survival curves. Elevated temperatures and pressures substantially promoted the inactivation of L. innocua. A linear and two non-linear (Weibull and log-logistic) models were fitted to these data and the goodness of fit of these models were compared. Regression coefficients (R2), root mean square (RMSE), accuracy factor (Af) values and residual plots suggested that linear model, although it produced good fits for some pressure-temperature combinations, was not as appropriate as non-linear models to represent the data. The residual and correlation plots strongly suggested that among the non linear models studied the log-logistic model produced better fit to the data than the Weibull model. Such pressure-temperature inactivation models form the engineering basis for design, evaluation and optimization of high hydrostatic pressure processes as a new preservation technique.

Colony Count, Microbial↗

Intra-unit correlations in seroconversion to Actinobacillus pleuropneumoniae and Mycoplasma hyopneumoniae at different levels in Danish multi-site pig production facilities.

In this paper, multilevel logistic models which take into account the multilevel structure of multi-site pig production were used to estimate the variances between pigs produced in Danish multi-site pig production facilities regarding seroconversion to Actinobacillus pleuropneumoniae serotype 2 (Ap2) and Mycoplasma hyopneumoniae (Mh). Based on the estimated variances, three newly described computational methods (model linearisation, simulation and linear modelling) and the standard method (latent-variable approach) were used to estimate the correlations (intra-class correlation components, ICCs) between pigs in the same production unit regarding seroconversion. Substantially different values of ICCs were obtained from the four methods. However, ICCs obtained by the simulation and the model linearisation were quite consistent. Data used for estimation were collected from 1161 pigs from 429 litters reared in 36 batches at six Danish multi-site farms chronically infected with the agents. At the farms, weaning age was 3-4.5 weeks, after which batches of pigs were reared using all-in/all-out management by room. Blood samples were collected shortly before: weaning, transfer from weaning-site to finishing-site, and sending the first pigs in the batch for slaughter (third sampling). Few pigs seroconverted at the weaning-sites, whereas considerable variation in seroconversion was observed at the finishing-sites. Multilevel logistic models (initially including four levels: farm, batch, litter, pig) were used to decompose the variation in seroconversion at the finishing-site. However, there was essentially no clustering at the litter level-leading to the use of three-level models. In the case of Ap2, clustering within batch was so high that the data eventually were reduced to two levels (farm, batch). For seroconversion to Ap2, ICC between pigs within batches was approximately 90%, whereas the ICC between pigs within batches for Mh was approximately 40%. This indicates that the possibility for Mh to spread between pigs within batches is lower than for Ap2. The diversity in seroconversion between batches within the same farm was large for Ap2 (ICC approximately 10%), whereas there was a relative strongly ICC (approximately 50%) between batches for Mh. This indicates that the transmission of Mh is more consistent within a farm, whereas the presence of Ap2 varies between batches within a farm.

Actinobacillus Infections↗

Prediction of intravenous immunoglobulin unresponsiveness in patients with Kawasaki disease.

BACKGROUND: In the present study, we developed models to predict unresponsiveness to intravenous immunoglobulin (IVIG) in Kawasaki disease (KD). METHODS AND RESULTS: We reviewed clinical records of 546 consecutive KD patients (development dataset) and 204 subsequent KD patients (validation dataset). All received IVIG for treatment of KD. IVIG nonresponders were defined by fever persisting beyond 24 hours or recrudescent fever associated with KD symptoms after an afebrile period. A 7-variable logistic model was constructed, including day of illness at initial treatment, age in months, percentage of white blood cells representing neutrophils, platelet count, and serum aspartate aminotransferase, sodium, and C-reactive protein, which generated an area under the receiver-operating-characteristics curve of 0.84 and 0.90 for the development and validation datasets, respectively. Using both datasets, the 7 variables were used to generate a simple scoring model that gave an area under the receiver-operating-characteristics curve of 0.85. For a cutoff of 0.15 or more in the logistic regression model and 4 points or more in the simple scoring model, sensitivity and specificity were 86% and 67% in the logistic model and 86% and 68% in the simple scoring model. The kappa statistic is 0.67, indicating good agreement between the logistic and simple scoring models. CONCLUSIONS: Our predictive models showed high sensitivity and specificity in identifying IVIG nonresponders among KD patients.

Adolescent↗

Factors affecting the occurrence of urothelial tumors in dye workers exposed to aromatic amines.

BACKGROUND: Past studies have analyzed individual jobs in dyestuff factories, materials manufactured and handled, age at exposure, and the duration of exposure in factories as factors related to the occurrence of urothelial tumors. None of these studies was based on long-term observation, and the factors involved in the occurrence of urothelial tumors remain controversial. In this study, various factors that may affect the occurrence of urothelial tumors in dye workers were assessed by multivariate analysis. METHODS: Three hundred and sixty-three workers in nine member factories of the Dyestuff Industrial Cooperative Association were included the study. Factory A is a large dyestuff chemical factory in Wakayama City with 218 dye workers. The other eight smaller factories employ a total of 145 dye workers. Correlations of tumor occurrence with a variety of factors, such as dyestuff intermediates manufactured and handled, types of job in the factory, age at the beginning of occupational exposure, and the duration of exposure were examined by multivariate analysis using multiple logistic models. RESULTS: Urothelial tumors were found in 58 (16.0%) of the 363 dye workers in the nine member factories of the Cooperative Association examined in the present study. The incidence in workers in Factory A, 5.5% (12 patients), was significantly (P < 0.01) lower than the overall incidence, while that in the eight small factories, 31.7% (46 patients), was significantly (P < 0.01) higher than the overall incidence. The risk factors significantly related to tumor occurrence in the 363 dye workers were benzidine (odds ratio, 8.302) as a dyestuff intermediate, manufacturing work (odds ratio, 4.631), and a long period of exposure (odds ratio, 1.018). Correlations of the tumor occurrence with the various factors were examined by multivariate analysis using multiple logistic models. In the total of 363 workers, benzidine as an intermediate (P < 0.05), manufacturing work (P < 0.01) and the duration of exposure (P < 0.01) were found to have contributed to the urothelial tumor occurrence. In Factory A, benzidine as an intermediate (P < 0.01) and duration of exposure (P < 0.05) contributed significantly to tumor occurrence. CONCLUSIONS: 1) The manufacturing and handling of benzidine and duration of exposure contribute significantly to the occurrence of occupational urothelial tumor, the former more strongly than the latter; 2) the contribution of different job types to tumor occurrence may be dependent upon the industrial health and safety practices in each factory.

Adolescent↗

Estimation of an expected caesarean section rate taking into account the case mix of a maternity hospital. Analysis from the AUDIPOG Sentinelle Network (France). Obstetricians of AUDIPOG. Association of Users of Computerised Files in Perinatalogy, Obstetrics and Gynaecology.

OBJECTIVE: To provide maternity unit with an expected caesarean section rate, according to its case mix (i.e. women's characteristics associated with caesarean section risk). DESIGN: Cohort study. SETTING: 149 maternity units in France. SAMPLE: 40,512 single births collected by the French Sentinelle Network, in January every year, from 1994 until 1998. METHODS: Univariate analysis was used to identify caesarean section risk factors, and multivariate analysis to adjust for the role of the maternity units' characteristics, after taking into account the women's characteristics. A two-level logistic model was used to show that the caesarean section rate varied according to maternity units' characteristics and to estimate therefore expected caesarean section rates (before and during labour), for each maternity unit, according to its case mix. MAIN OUTCOME: Caesarean section rates (before and during labour). RESULTS: Within the Sentinelle Network the caesarean section rate was 15.0% (7.6% were before labour). The joint effect of the size and juridical status on caesarean section risk was studied. The reference hospital was university maternity units with more than 2000 deliveries/year. Community or private maternity units with more than 2000 deliveries/year carried out fewer prophylactic caesarean sections than the reference hospital (ORadj = 0.7 and 0.6, respectively). Conversely, private maternity units with fewer than 2000 deliveries/year performed more prophylactic caesarean sections than the reference hospital (ORadj = 1.7). The two-level logistic model showed that a maternity unit effect still existed after taking into consideration both women's characteristics and those of the maternity unit, and estimated expected caesarean section rates. CONCLUSION: Knowledge of the expected caesarean section rate constitutes a personal reference to which the maternity hospital can compare its observed caesarean section rate, and is thus likely to have a significant effect on delivery practices.

Adult↗

Reporting bias related to an environmental hazard.

During spring 1984, second and fifth grade schoolchildren living in three Haifa Bay areas on the eastern Mediterranean coast, with different levels of air pollution, were studied. The parents of these children filled out ATS-NHLI (American Thoracic Society and the National Heart and Lung Institute) health questionnaires and the children performed PFT (Pulmonary Function Tests). A trend of higher prevalence of most reported respiratory symptoms was found for schoolchildren growing up in the medium and high polluted areas as compared with the low pollution area. Logistic models fitted for the respiratory conditions that differed significantly among the three residential areas also included background variables that could be responsible for these differences. Relative risks for respiratory conditions calculated from these models were in the range of 1.38 and 1.81 for children from the polluted area as compared to 1.00 for the low polluted area. All the measured values of PFT were within the normal range, with no consistent reduction in PFT for any residential area. During spring 1989, seventh graders (second graders in 1984) were reexamined and a new cohort of fifth grade children was studied, using the same techniques as in 1984. A very significant rise in the prevalence of most reported respiratory symptoms and diseases was observed among both fifth and seventh grade schoolchildren in 1989 compared to 1984, especially in the low and medium polluted areas and less in the polluted area. Changes over time in PFT in the older cohort were similar in the three areas. PFT of fifth graders in 1984 and in 1989 were very similar. The most significant factor in logistic models fitted for the prevalence of respiratory conditions among the studied schoolchildren in 1989, was the subjective attitude of their parents towards the deleterious effects of air pollution on their children's health, and the subjective estimate of their children's exposure to pollution rather than measured exposure. A huge campaign carried out during the survey against the main polluters in the Haifa Bay area caused both public concern and apparently reporting bias.

Air Pollutants↗

[The validity of revised death certificates (ICD-10) for ischemic heart disease in Oita City, Japan].

PURPOSE: Mortality statistics have recorded an increased number of deaths from ischemic heart disease (IHD) since death certificates were revised to reflect the International Classification of Diseases, tenth revision (ICD-10) in Japan, in 1995. However, it remains unclear whether the validity of IHD diagnosis improved after this revision. METHODS: We conducted the Oita Cardiac Death Survey to validate IHD certified deaths that occurred among residents aged 25-74 in Oita City, Japan (mean population = 273,000). Of the eligible 342 fatalities, 328 cases (95.0%) were examined by a review of the medical records and/or interviews with physicians. The MONICA criteria were applied and provided a reference standard against which to assess the validity of certified fatal IHD. Sensitivity (Se), positive predictive value (PPV), specificity (Sp) and negative predictive value (NPV) for IHD as the cause of death were analyzed, assuming that all validated IHD deaths were true. Multivariate logistic models were used to determine associations of false positive and false negative cases with sex, age at time of death and place of death. RESULTS: Vital statistics revealed 273 fatalities to be due to cardiac disease, including 143 from acute myocardial infarctions (AMI), 27 from other IHD, 52 from heart failure and 51 from other heart diseases. After validation, 25 'definite fatal AMI' and 71 'possible fatal AMI or IHD death' were identified among all subjects according to the MONICA criteria. In all, Se, PPV, Sp and NPV for IHD certified as the cause of death were 86.5% (95% Cl: 77.6-92.3), 50.3% (42.5-58.1), 64.7% (58.1-70.7), and 92.0% (86.5-95.5), respectively. PPV among persons aged 25-54 years was remarkably decreased. PPV and Sp among out-of-hospital deaths were significantly lower than for in-hospital deaths. Multivariate logistic models revealed out-of-hospital deaths and being aged 25-54 years to be significant predictors of false positive cases (odds ratio (OR) = 2.03, P < 0.001 versus in-hospital deaths and OR = 2.79, P < 0.05 versus ages of 65-74 years, respectively). CONCLUSIONS: Because false positive cases increased among certified IHD deaths after the revision, PPV and Sp percentages decreased. Out-of-hospital deaths and being aged 25-54 years were associated with increased possibility of false positive. Given our findings, IHD deaths in vital statistics may increase due to the tendency of physicians to certify IHD as the cause of death in cases without clear sign suggestive of other causes.

Adult↗

Comparative case fatality analysis of the International Tissue Plasminogen Activator/Streptokinase Mortality Trial: variation by country beyond predictive profile. The Investigators of the International Tissue Plasminogen Activator/Streptokinase Mortality Trial.

OBJECTIVES: This study was designed to examine the variation in mortality rates among countries participating in the International Tissue Plasminogen Activator/Streptokinase Mortality Trial. BACKGROUND: Despite uniform inclusion and exclusion criteria and protocol in this trial, 30-day mortality rates (irrespective of treatment allocation) ranged from 4.2% to 14.8% among the participating countries. METHODS: With use of the risk factors identified by a multi-variate logistic model, the total study group was classified into deciles on the basis of each patient's risk profile and individual probability of dying within 30 days. Expected mortality rates were then calculated and compared with actual mortality for each decile of the total study group, as well as for patients from each country. RESULTS: Independent risk factors for mortality were older age (odds ratio 1.97 for each 10-year increment), systolic hypotension (blood pressure < 95 mm Hg) at entry (odds ratio 3.7), Killip class > 1 at entry (odds ratio 3.5), history of antecedent angina (odds ratio 1.23 to 1.49), history of diabetes mellitus (odds ratio 1.64), previous infarction (odds ratio 1.23) and history of never smoking (odds ratio 1.37). The overall mortality rate among the 1,612 patients in risk deciles 9 and 10 was 26%; for the 1,606 patients in deciles 1 and 2 it was 1.2%, with a sensitivity of 58.6% and a specificity of 83.7%. The logistic model closely predicted and explained the different mortality rates for most countries (the differences between expected and actual mortality were nonsignificant). However, in the total study group, the difference between the expected and actual mortality was significant (p < 0.001). This difference was mainly ascribed to the two countries with the highest and lowest mortality rates. When the patients from these two countries were excluded from the analysis, the overall difference became nonsignificant. CONCLUSIONS: These findings suggest that the recognized risk factors associated with increased case fatality in acute myocardial infarction account only in part for mortality differences across or within populations.

Aged↗

Forecasting with growth curves: the effect of error structure.

"The main theme of this paper is an investigation into the importance of error structure as a determinant of the forecasting accuracy of the logistic model. The relationship between the variance of the disturbance term and forecasting accuracy is examined empirically. A general local logistic model is developed as a vehicle to be used in this investigation. Some brief comments are made on the assumptions about error structure, implicit or explicit, in the literature." The results suggest that "the variance of the disturbance term, when using the logistic to forecast human populations, is proportional to at least the square of population size."

Forecasting↗

Selection of women at high risk of breast cancer for initial screening.

Selective breast cancer screening refers to the intentional restriction of screening to only a high-risk subgroup of the total population of women at risk. Using data from the Canadian National Breast Screening Study, we explored methods of defining such subgroups. Discriminants were based on risk factor information collected prior to screening and were constructed using a training group of 77 cases and 400 controls. They were then tested on a separate group of 38 cases and 200 controls. Both simple risk factor counts and logistic models were utilized and separate analyses were performed for pre- and post-menopausal women. Using a logistic model, we were able to define a high-risk subgroup encompassing less than 40% of the test controls and over 85% of the test cases. Such a selection strategy, if implemented, might reduce initial visit mammography rates by up to 60% with only a small reduction in case detection. Other uses as determining the optimal age for initiation of screening are also discussed.

Age Factors↗

Comparison of available benchmark dose softwares and models using trichloroethylene as a model substance.

By using trichloroethylene as a model substance the U.S. EPA benchmark dose software was compared to the software by Crump and the software by Kalliomaa. Dose-response and dose-effect data on the liver, kidneys, central nervous system (CNS), and tumours were selected for the evaluation. Based on the present study the U.S. EPA software is preferable to the other softwares for dichotomous data. A wider range in benchmark doses was often observed for dichotomous data when the numbers of dose levels were limited. The log-logistic model in most cases gave the best fit when ranking the dichotomous models. In addition, the log-logistic model often implied a more conservative benchmark dose. For continuous data it was more difficult to find a model describing the data. The softwares by Kalliomaa and by the U.S. EPA offered the best opportunities for benchmark dose modelling of continuous data. Flexible models, like the Hill- and the Mult model, are needed for S-shaped continuous data but these models demand more dose levels in order to describe the data. Since the number of dose levels are important for model selection study design is important and should be further evaluated.

Benchmarking↗

External beam radiotherapy for painful osseous metastases: pooled data dose response analysis.

PURPOSE: Although the effectiveness of external beam irradiation in palliation of pain from osseous metastases is well established, the optimal fractionation schedule has not been determined. Clinical studies to date have failed to demonstrate an advantage for higher doses. To further address this issue, we conducted a pooled dose response analysis using data from published Phase III clinical trials. METHODS AND MATERIALS: Complete response (CR) was used as an endpoint because it was felt to be least susceptible to inconsistencies in assessment.The biological effective dose (BED) was calculated for each schedule using the linear-quadratic model and an alpha/beta of 10. Using SAS version 6.12, the data were fitted using a weighted linear regression, a logistic model, and the spline technique. Finally, BED was categorized, and odds ratios for each level were calculated. RESULTS: CR was assessed early and late in 383 and 1,007 patients, respectively. Linear regression on the early-response data yielded a poor fit and a nonsignificant dose coefficient. With the late-response data, there was an excellent fit (R-square = 0.842) and a highly significant dose coefficient (p = 0.0002). Fitting early CR to a logistic model, we could not establish a significant dose response relationship. However, with the late-response data there was an excellent fit and the dose coefficient was significantly different from zero (0.017 +/- 0.00524; p = 0.0012). Application of the spline technique or removal of an outlier resulted in an improved fit (p = 0.048 and p = 0.0001, respectively). Using BED of < 14.4 Gy as a reference level, the odds ratios for late CR were 2.29-3.32 (BED of 19.5-51.4 Gy, respectively). CONCLUSION: Our results demonstrate a clear dose-response for pain relief. Further testing of high intensity regiments is warranted.

Bone Neoplasms↗

A comparison of methods for preoperative discrimination between malignant and benign adnexal masses: the development of a new logistic regression model.

OBJECTIVE: The aim of this study was to assess the complementary use of ultrasonographic end points with the level of circulating CA 125 antigen by multivariate logistic regression analysis algorithms to distinguish malignant from benign adnexal masses before operation. STUDY DESIGN: One hundred ninety-one patients aged 18 to 93 years with overt adnexal masses were examined by transvaginal ultrasonography with color Doppler imaging and 31 variables were recorded. The end points were the histologic classification of the tumor and the areas under the receiver-operator characteristic curves of alternative algorithms. RESULTS: One hundred forty patients had benign tumors and 51 (26.7%) had malignant tumors: 31 primary invasive tumors (37% International Federation of Gynecology and Obstetrics stage I), 5 tumors of borderline malignancy (100% International Federation of Gynecology and Obstetrics stage I), and 15 tumors were metastatic and invasive. The most useful variables for the logistic regression analysis were the menopausal status, the serum CA 125 level, the presence of >/=1 papillary growth (>3 mm in length), and a color score indicative of tumor vascularity and blood flow. The optimized procedure had a sensitivity of 95.9% and a specificity of 87.1%. The area under the receiver-operator characteristic curve was significantly higher (P <.01) than the corresponding values from the independent use of serum CA 125 levels or indexes of tumor form or vascularity. CONCLUSION: Regression analysis of a few complementary variables can be used to accurately discriminate between malignant and benign adnexal masses before operation.

Adnexal Diseases↗

The effect of sample size for estimating Rasch/IRT parameters with dichotomous items.

Thirteen samples were randomly drawn from the normative database for the latest edition of Knox's Cube Test-Revised (KCT-R). Parameter estimates for the Rasch model and two and three parameter logistic models were derived and compared. Sample size influenced these estimates as might be expected. Rasch parameter estimates consistently showed the smallest values by sample size using a goodness of fit index.

Data Interpretation, Statistical↗

Relationship between height, glucose intolerance, and hypertension in an urban African black adult population: a case for the "thrifty phenotype" hypothesis?

An association between the factors of low birth weight and fetal growth retardation and subsequent risk of cardiovascular disease has been proposed; this is the basis of the "thrifty phenotype" hypothesis described in relation to type 2 diabetes mellitus. The relationship between height, presumably an indicator of early life experience, and glucose intolerance and hypertension was examined in a sample survey of noncommunicable disease in an urban African adult population. Height, other anthropometric measurements, and biosocial data were obtained in the study of 998 civil servants selected by multistage sampling in Ibadan, a major Nigerian city. Ibadan is a low-prevalence region for diabetes, with a rate of 0.8% and 2.2% for an impaired glucose tolerance. The prevalence rate of hypertension was 10.3% in the population. A significant negative correlation was found between height and blood glucose level (r = -0.14, p < 0.001), whereas there was no correlation with blood pressures. Multiple regression analyses did not demonstrate height as a determinant of either blood pressure or plasma glucose. However, in a logistic model height was found to be associated with abnormal glucose tolerance (diabetes and impaired glucose tolerance) (odds ratio, 0.01; p < 0.003). In the logistic model of the blood pressure data there was no association between height and hypertension. There was some association between height and blood glucose level and also glucose intolerance in the urban African population sample, but none with elevated blood pressure. The significance of the observed inverse relationship, though uncertain, deserves further exploration.

Adult↗

Non-ignorable missing covariates in generalized linear models.

We propose a likelihood method for estimating parameters in generalized linear models with missing covariates and a non-ignorable missing data mechanism. In this paper, we focus on one missing covariate. We use a logistic model for the probability that the covariate is missing, and allow this probability to depend on the incomplete covariate. We allow the covariates, including the incomplete covariate, to be either categorical or continuous. We propose an EM algorithm in this case. For a missing categorical covariate, we derive a closed form expression for the E- and M-steps of the EM algorithm for obtaining the maximum likelihood estimates (MLEs). For a missing continuous covariate, we use a Monte Carlo version of the EM algorithm to obtain the MLEs via the Gibbs sampler. The methodology is illustrated using an example from a breast cancer clinical trial in which time to disease progression is the outcome, and the incomplete covariate is a quality of life physical well-being score taken after the start of therapy. This score may be missing because the patients are sicker, so this covariate could be non-ignorably missing.

Algorithms↗