PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “missing data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Surrogate endpoints in clinical trials: cardiovascular diseases.

A surrogate endpoint in a cardiovascular clinical trial is defined as endpoint measured in lieu of some other so-called 'true' endpoint. A surrogate is especially useful if it is easily measured and highly correlated with the true endpoint. Often the 'true' endpoint is one with clinical importance to the patient, for example, mortality or a major clinical outcome, while a surrogate is one biologically closer to the process of disease, for example, ejection fraction. Use of the surrogate can often lead to dramatic reductions in sample size and much shorter studies than use of the true endpoint. We discuss several problems common in trials with surrogate endpoints. Most important is the effect of missing data, especially in the face of informative censoring. Possible solutions are the assignment of scores or formal penalties to missing data.

Cardiovascular Diseases↗

Evaluation of most frequent errors in daily compilation and use of a radiation treatment chart.

Between 1 March and 30 April (1994) we recorded the errors detected by the physician, the radiographer or the physicist during prescription, preparation and execution phases of 227 treatment plans. The radiation treatment modalities used were the following: (i) single or opposed fields, moulded or not; and (ii) multiple fields or kinetic techniques. The total number of sessions performed is 1613 with the cobalt unit and 2131 with the linear accelerator (total, 3744). The total number of wrong data is 155, consisting of 24/227 (10.5%) in compilation, 22/3744 (0.58%) in execution and 109/3744 (2.9%) in registration phases. The number of missing data is 140, consisting of 10/227 (4.4%) in compilation, 9/3744 (0.2%) in execution and 121/3744 (3.2%) in registration phases. Wrong data of compilation, even if in high rate (10.5%), were all found during the same compilation phase or at the first treatment, so that they did not alter the exactness of the treatment plan. Wrong and missing data, found in the registration phase (2.9% and 3.2%, respectively), depend on the repetition of daily treatment and on the registration of data on the chart after having digitized them on the display.

Cobalt Radioisotopes↗

An appraisal of echocardiography as an epidemiological tool. The Strong Heart Study.

PURPOSE: Despite the prognostic importance of left ventricular (LV) mass (LVM) by M-mode echocardiography, concern exists about bias introduced by missing data. The American Society of Echocardiography has made recommendations for linear measurements of LV wall thickness and internal dimension used to calculate LVM, but it is unknown whether their substitution for suboptimal M-modes improves measurement yield and reduces bias. METHODS: LVM measurement yield and associations of missing data with risk factors were assessed in 3487 American Indian participants in Strong Heart Study (SHS) Phase II and compared to data from other large-scale studies. RESULTS: In SHS, LVM was measurable in 3188 (91%) participants compared to 4947/6148 (80%) Framingham participants studied by classic M-mode technique, with less decrease in measurement yield with age in SHS. In univariate SHS analyses, missing LVM was significantly associated with male gender, older age, greater height, body mass index, fat-free mass, waist/hip ratio, fibrinogen and, marginally, diabetes but not smoking, blood pressure, or lipids. In logistic regression analysis, missing LVM was independently associated with male gender, older age, greater body mass index and lower forced expiratory volume (FEV(1)) (with a low multiple R(2) [.04]), but not other risk factors. Doppler stroke volume, a measure of hemodynamic volume load, was measurable in 96% of SHS participants; missing values were weakly associated with older age, higher creatinine and lower FEV(1). During 48 +/- 11 months of follow-up, inability to measures LV mass or stroke volume was not associated with higher rates of cardiovascular events or death (p = 0.25 to 0.96). CONCLUSIONS: Improvements in echocardiographic methods have increased the yield of LVM in middle-aged and older adults and allow even more consistent assessment of cardiac volume load. Despite small persistent biases, due to associations of missing LVM and Doppler stroke volume data with male gender, greater obesity, lower FEV(1) and (for LVM only) older age, individuals with missing measurement are not at higher risk of cardiovascular events.

Aged↗

Long acting beta-agonists versus theophylline for maintenance treatment of asthma.

BACKGROUND: Theophylline and long acting beta2-agonists are bronchodilators used for the management of persistent asthma symptoms, especially nocturnal asthma. They represent different classes of drug with differing side-effect profiles. OBJECTIVES: To assess the comparative efficacy, safety and side-effects of long-acting beta-agonists and theophylline in the maintenance treatment of asthma. SEARCH STRATEGY: Randomised, controlled trials (RCTs) were identified using the Cochrane Airways Group register. The register was searched using the following terms: asthma and theophylline and long acting beta-agonist or formoterol or foradile or eformoterol or salmeterol or bambuterol or bitolterol. Titles and abstracts were then screened to identify potentially relevant studies. The bibliography of each RCT was searched for additional RCTs. Authors of identified RCTs were contacted for other relevant published and unpublished studies. SELECTION CRITERIA: All included studies were RCTs involving adults and children with clinical evidence of asthma. These studies must have compared oral sustained release and/or dose adjusted theophylline with an inhaled long-acting beta-agonist. DATA COLLECTION AND ANALYSIS: Potentially relevant trials, identified by screening titles and/or abstracts, were obtained. Two reviewers independently assessed full text versions of these trials to decided whether the trial should be included in the review, and assessed its methodological quality. Where there was disagreement between reviewers, this was resolved by consensus, or reference to a third party. Data were extracted by two independent reviewers. Inter-rater reliability was assessed by simple agreement. Study authors were contacted to clarify randomisation methods, provide missing data, verify the data extracted and identify unpublished studies. Relevant pharmaceutical manufacturers were also contacted. MAIN RESULTS: Six trials met the inclusion criteria. Five used salmeterol and one, biltoterol. They were of varying quality. There was a trend for salmeterol to improve FEV1 more than theophylline in three studies and salmeterol use was associated with more symptom free nights. Bitolterol, used in only one study, was reported to be less effective than theophylline. Subjects taking salmeterol experienced fewer adverse events than those using theophylline (Relative Risk 0.38; 95%Confidence Intervals 0.25, 0.57). Significant reductions were reported for central nervous system adverse events (Relative Risk 0.51; 95%Confidence Intervals 0.30, 0.88) and gastrointestinal adverse events (Relative Risk 0.32; 95%Confidence Intervals 0.17, 0.59). REVIEWER'S CONCLUSIONS: Salmeterol may be more effective than theophylline in reducing asthma symptoms including night waking and improving lung function. More adverse events occurred in subjects using theophylline when compared to salmeterol.

Adrenergic beta-Agonists↗

Estimating the distribution of times from HIV seroconversion to AIDS using multiple imputation. Multicentre AIDS Cohort Study.

Multiple imputation is a model based technique for handling missing data problems. In this application we use the technique to estimate the distribution of times from HIV seroconversion to AIDS diagnosis with data from a cohort study of 4954 homosexual men with 4 years of follow-up. In this example the missing data are the dates of diagnosis with AIDS. The imputation procedure is performed in two stages. In the first stage, we estimate the residual AIDS-free time distribution as a function of covariates measured on the study participants with data provided by the participants who were seropositive at study entry. Specifically, we assume the residual AIDS-free times follow a log-normal regression model that depends on the covariates measured at enrolment on the seropositive participants. In the second stage we impute the date of AIDS diagnosis for the participants who seroconverted during the course of the study and are AIDS-free with use of the log-normal distribution estimated in the first stage and the covariates from each seroconverter's latest visit. The estimated proportions developing AIDS within 4 and within 7 years of seroconversion are 15 and 36 per cent respectively, with associated 95 per cent confidence intervals of (10, 21) and (26, 47) per cent. We discuss the Bayesian foundations of the multiple imputation technique and the statistical and scientific assumptions.

AIDS Serodiagnosis↗

Item response models for longitudinal quality of life data in clinical trials.

Assessment of quality of life is becoming standard in clinical trials. A popular method for measuring quality of life is with instruments which utilize multiple-item subscales, in which each item is scored on a Likert scale. Most statistical methods for the analysis of quality of life data in clinical trials do not explicity consider the properties and psychometric features which were of interest in scale development. In this regard, the measurement and statistical summarization of quality of life data, along with the clinical interpretation, can be somewhat disjoint from the psychometric concerns of the development process. The aim of this paper is to address the complicated issues present in analysing multiple-item ordinal quality of life data in clinical trials while maintaining fidelity to the psychometrical foundations upon which quality of life instruments are built. Accomplishing this will require the development of item response models which recognize the longitudinal aspects of clinical trial designs as well as the potential problem of informatively missing data. A general item response modeling approach is presented for longitudinal multiple-item quality of life data measured on ordinal scales with model components for missing data mechanisms and latent trait regression on treatment indicators and other covariates.

Clinical Trials as Topic↗

Single-channel data and missed events: analysis of a two-state Markov model.

Patch-clamp recording permits investigation of the gating kinetics of single ion channels. Careful statistical analysis of kinetic data can yield clues as to the molecular events underlying channel gating. However, it is important that such analysis should take full account of the limitations that arise from the finite time resolution of patch-clamp recording techniques. Single-ion-channel data are generally interpreted in terms of Markov process models of channel gating mechanisms. Experimental channel records suffer from time interval omission, i.e. failure to detect brief channel openings and closings. This leads to an identifiability problem when analysing single-channel data, i.e. different gating mechanisms provide equally convincing descriptions of the same experimental data. We consider a two-state Markov model of receptor-channel gating in which the channel opening rate is proportional to the agonist concentration, C in equilibrium with OA. By using computer-simulated data, the approximate likelihood of the data is maximized to yield parameter estimates for the model. At a single agonist concentration there is an identifiability problem in that two pairs of parameter estimates are obtained. The 'true' parameter estimates cannot be distinguished from the 'false' ones. By considering data corresponding to a range of agonist concentrations one may identify the 'true' parameter estimates as those that do not change as the agonist concentration is increased. Alternatively, one may identify the 'true' parameter estimates directly by maximizing a global likelihood, the latter being obtained by simultaneous consideration of data obtained at several different agonist concentrations.(ABSTRACT TRUNCATED AT 250 WORDS)

Animals↗

Estimating equations with nonignorably missing response data.

Troxel, Lipsitz, and Brennan (1997, Biometrics 53, 857-869) considered parameter estimation from survey data with nonignorable nonresponse and proposed weighted estimating equations to remove the biases in the complete-case analysis that ignores missing observations. This paper suggests two alternative modifications for unbiased estimation of regression parameters when a binary outcome is potentially observed at successive time points. The weighting approach of Robins, Rotnitzky, and Zhao (1995, Journal of the American Statistical Association 90, 106-121) is also modified to obtain unbiased estimating functions. The suggested estimating functions are unbiased only when the missingness probability is correctly specified, and misspecification of the missingness model will result in biases in the estimates. Simulation studies are carried out to assess the performance of different methods when the covariate is binary or normal. For the simulation models used, the relative efficiency of the two new methods to the weighting methods is about 3.0 for the slope parameter and about 2.0 for the intercept parameter when the covariate is continuous and the missingness probability is correctly specified. All methods produce substantial biases in the estimates when the missingness model is misspecified or underspecified. Analysis of data from a medical survey illustrates the use and possible differences of these estimating functions.

Biometry↗

Multivariate outlier detection applied to multiply imputed laboratory data.

In clinical laboratory safety data, multivariate outlier detection methods may highlight a patient whose laboratory measurements do not follow the same pattern of relationships as the majority of patients, although their individual measurements are not found to be outlying when considered one at a time. Missing data problems are often dealt with by imputing a single value as an estimate of the missing value. The completed data set may then be analysed using traditional methods. A disadvantage of using single imputation is the underestimation of variability, with a corresponding distortion of power in hypothesis testing. Multiple imputation methods attempt to overcome this problem, and in this paper a study is described which considers the application of multivariate outlier detection methods to multiply imputed clinical laboratory safety data sets. Three different proportions of missing data are generated in laboratory data sets of dimensions 4, 7, 12 and 30, and a comparison of eight multiple imputation methods is carried out. Two outlier detection techniques, Mahalanobis distance and generalized principal component analysis, are applied to the multiply imputed data sets, and their performances are discussed. Measures are introduced for assessing the accuracy of the missing data results, depending on which method of analysis is used.

Algorithms↗

Some conceptual and statistical issues in analysis of longitudinal psychiatric data. Application to the NIMH treatment of Depression Collaborative Research Program dataset.

Longitudinal studies have a prominent role in psychiatric research; however, statistical methods for analyzing these data are rarely commensurate with the effort involved in their acquisition. Frequently the majority of data are discarded and a simple end-point analysis is performed. In other cases, so called repeated-measures analysis of variance procedures are used with little regard to their restrictive and often unrealistic assumptions and the effect of missing data on the statistical properties of their estimates. We explored the unique features of longitudinal psychiatric data from both statistical and conceptual perspectives. We used a family of statistical models termed random regression models that provide a more realistic approach to analysis of longitudinal psychiatric data. Random regression models provide solutions to commonly observed problems of missing data, serial correlation, time-varying covariates, and irregular measurement occasions, and they accommodate systematic person-specific deviations from the average time trend. Properties of these models were compared with traditional approaches at a conceptual level. The approach was then illustrated in a new analysis of the National Institute of Mental Health Treatment of Depression Collaborative Research Program dataset, which investigated two forms of psychotherapy, pharmacotherapy with clinical management, and a placebo with clinical management control. Results indicated that both person-specific effects and serial correlation play major roles in the longitudinal psychiatric response process. Ignoring either of these effects produces misleading estimates of uncertainty that form the basis of statistical tests of hypotheses.

Analysis of Variance↗

Analysis of change in the presence of informative censoring: application to a longitudinal clinical trial of progressive renal disease.

The rate of change in a continuous variable, measured serially over time, is often used as an outcome in longitudinal studies or clinical trials. When patients terminate the study before the scheduled end of the study, there is a potential for bias in estimation of rate of change using standard methods which ignore the missing data mechanism. These methods include the use of unweighted generalized estimating equations methods and likelihood-based methods assuming an ignorable missing data mechanism. We present a model for analysis of informatively censored data, based on an extension of the two-stage linear random effects model, where each subject's random intercept and slope are allowed to be associated with an underlying time to event. The joint distribution of the continuous responses and the time-to-event variable are then estimated via maximum likelihood using the EM algorithm, and using the bootstrap to calculate standard errors. We illustrate this methodology and compare it to simpler approaches and usual maximum likelihood using data from a multi-centre study of the effects of diet and blood pressure control on progression of renal disease, the Modification of Diet in Renal Disease (MDRD) Study. Sensitivity analyses and simulations are used to evaluate the performance of this methodology in the context of the MDRD data, under various scenarios where the drop-out mechanism is ignorable as well as non-ignorable.

Algorithms↗

Efficacy and safety of deucravacitinib, an oral, selective tyrosine kinase 2 inhibitor, in patients with active psoriatic arthritis: 52-week results from the randomised, double-blind, placebo-controlled phase 3 POETYK PsA-1 trial.

OBJECTIVES: The randomised, double-blind, placebo-controlled, phase 3 Program fOr Evaluation of TYK2 inhibitor Psoriatic Arthritis-1 (POETYK PsA-1) trial evaluated the efficacy, safety, and tolerability of deucravacitinib, an oral, selective tyrosine kinase 2 inhibitor, in patients with PsA na&#xef;ve to biologic disease-modifying antirheumatic drugs. METHODS: Adults with active PsA, high-sensitivity C-reactive protein concentration &#x2265; 3 mg/L, and &#x2265; 1 PsA-related hand and/or foot erosion detectable via radiograph were randomised 1:1 to oral deucravacitinib 6 mg once daily or placebo through week (W) 16. At W16, patients continued receiving deucravacitinib or switched from placebo to deucravacitinib through W52. The primary endpoint was American College of Rheumatology 20% improvement in response (ACR20) at W16. Nonresponder imputation was used for missing data. Efficacy and safety were evaluated through W52. Post hoc rank analysis of covariance was used to evaluate structural damage with no missing data imputation. RESULTS: In 670 patients, a significantly greater proportion of those receiving deucravacitinib vs placebo achieved ACR20 at W16 (54.2% vs 34.1%, P < .001). Responses with deucravacitinib were increased at W52. Patients who switched from placebo to deucravacitinib achieved improvements similar to those in patients who received continuous deucravacitinib. Inhibition of structural damage was observed at W16 and W52. At W16, incidences of serious adverse events (AEs) (deucravacitinib, 1.8%; placebo, 2.4%) and discontinuations due to AEs (2.4%; 1.8%) were low and remained low through W52, without imbalances in cardiovascular events, malignancies, or opportunistic infections. No new safety signals were detected; no deaths occurred. CONCLUSIONS: Deucravacitinib demonstrated superiority vs placebo for clinical responses, patient-reported outcomes, and structural damage inhibition in patients with PsA, with favourable tolerability and safety.

Humans↗

Catquest questionnaire for use in cataract surgery care: assessment of surgical outcomes.

PURPOSE: To demonstrate the outcome for patients after cataract extraction using the Catquest cataract questionnaire and discuss the models validity in assessing outcome. SETTING: Thirty-five Swedish departments of ophthalmology. METHODS: Patients having cataract extraction performed by surgeons from 35 Swedish departments of opthalmology participated in the study. The questionnaire was given to 2970 consecutive patients having surgery during March 1995 at the participating surgical units. The questionnaire was sent by mail to patients and completed on a voluntary basis. It focuses on visual disabilities in daily life, activity level, cataract symptoms, and degree of independence. The results form the questionnaire are interpreted using a benefit matrix that credits not only a decrease in visual disabilities and cataract symptoms but also an improvement in or maintenance of a preoperative activity level. RESULTS: Complete surgical outcome data and completed preoperative and postoperative questionnaires were available in 1933 cases (65.1%). Benefit from surgery according to the model was achieved by 90.9% of the patients. Patients having their second cataract extraction had the highest frequency of the greatest benefit form surgery. There was good agreement between the different levels of benefit from surgery according to the model and the patient's global rating of his or her vision or achieved visual acuity after surgery, respectively. Patients with missing data (did not return postoperative questionnaire or had missing surgical result variables) were older and had a higher frequency of other diseases and handicaps. CONCLUSION: The Catquest cataract questionnaire allowed the outcome of cataract surgery to be graded by different levels of benefit. There seemed to be good agreement between this model of assessment and the patient's global rating of his or her vision. Missing data may be a problem when a postal questionnaire is used.

Activities of Daily Living↗

A simulation study of the effects of assignment of prior identity-by-descent probabilities to unselected sib pairs, in covariance-structure modeling of a quantitative-trait locus.

Sib pair-selection strategies, designed to identify the most informative sib pairs in order to detect a quantitative-trait locus (QTL), give rise to a missing-data problem in genetic covariance-structure modeling of QTL effects. After selection, phenotypic data are available for all sibs, but marker data-and, consequently, the identity-by-descent (IBD) probabilities-are available only in selected sib pairs. One possible solution to this missing-data problem is to assign prior IBD probabilities (i.e., expected values) to the unselected sib pairs. The effect of this assignment in genetic covariance-structure modeling is investigated in the present paper. Two maximum-likelihood approaches to estimation are considered, the pi-hat approach and the IBD-mixture approach. In the simulations, sample size, selection criteria, QTL-increaser allele frequency, and gene action are manipulated. The results indicate that the assignment of prior IBD probabilities results in serious estimation bias in the pi-hat approach. Bias is also present in the IBD-mixture approach, although here the bias is generally much smaller. The null distribution of the log-likelihood ratio (i.e., in absence of any QTL effect) does not follow the expected null distribution in the pi-hat approach after selection. In the IBD-mixture approach, the null distribution does agree with expectation.

Alleles↗

An eigenvector method for estimating item parameters of the dichotomous and polytomous Rasch models.

The purpose of this paper is to describe a technique for obtaining item parameters of the Rasch model, a technique in which the item parameters are extracted from the eigenvector of a matrix derived from comparisons between pairs of items. The technique can be applied to both dichotomous and polytomous data. In application to a previously published data set, it is shown that the technique provides item parameter estimates comparable to those produced by joint maximum likelihood estimation, and for the most difficult items, the technique appears to produce superior estimates. This method has several advantages. It easily accommodates missing data, and makes transparent the basis for item parameter estimation in the presence of missing data. Furthermore, the method provides a link to other methods in the social sciences and, in particular, provides the framework for application of graph theory to the analysis of assessment networks. Finally, it exploits several characteristics that are unique to the Rasch model.

Algorithms↗

Analysis strategies for serial multivariate ultrasonographic data that are incomplete.

Ultrasonographic measurement of intima-media thickness in the carotid artery has emerged as an important non-invasive means of assessing atherosclerosis, and has served to define primary outcome measures related to progression of arterial lesions in several large clinical trials and epidemiologic studies. It is characteristic that measurements often cannot be obtained from all sites during repeated examinations. This leads to incomplete multivariate serial data, for which the set and number of visualized sites may vary across time. We have contrasted several conditional and unconditional maximum likelihood analytical approaches, and have evaluated these with a simulation experiment based on characteristics of ultrasound measurements collected during the course of the Asymptomatic Carotid Artery Plaque Study. We examined analyses based on unweighted and generalized least squares regression in which we estimated cross-sectional summary statistics using raw means, unconditional maximum likelihood estimates and full maximum likelihood estimates. Since the genesis of missing data is not fully clear, and since the approaches we examined are based, to some degree, on the assumption that data are missing at random, we also examined the relative impact of deviations from such an assumption on each of the approaches considered. We found that maximum likelihood based approaches increased the expected efficiency of the analysis of serial ultrasound data over ignoring missing data by up to 21 per cent.

Arteriosclerosis↗

An illness-death stochastic model in the analysis of longitudinal dementia data.

A significant source of missing data in longitudinal epidemiological studies on elderly individuals is death. Subjects in large scale community-based longitudinal dementia studies are usually evaluated for disease status in study waves, not under continuous surveillance as in traditional cohort studies. Therefore, for the deceased subjects, disease status prior to death cannot be ascertained. Statistical methods assuming deceased subjects to be missing at random may not be realistic in dementia studies and may lead to biased results. We propose a stochastic model approach to simultaneously estimate disease incidence and mortality rates. We set up a Markov chain model consisting of three states, non-diseased, diseased and dead, and estimate the transition hazard parameters using the maximum likelihood approach. Simulation results are presented indicating adequate performance of the proposed approach.

Aged↗

Applications of computer-intensive statistical methods to environmental research.

Conventional statistical approaches rely heavily on the properties of the central limit theorem to bridge the gap between the characteristics of a sample and some theoretical sampling distribution. Problems associated with nonrandom sampling, unknown population distributions, heterogeneous variances, small sample sizes, and missing data jeopardize the assumptions of such approaches and cast skepticism on conclusions. Conventional nonparametric alternatives offer freedom from distribution assumptions, but design limitations and loss of power can be serious drawbacks. With the data-processing capacity of today's computers, a new dimension of distribution-free statistical methods has evolved that addresses many of the limitations of conventional parametric and nonparametric methods. Computer-intensive statistical methods involve reshuffling, resampling, or simulating a data set thousands of times to empirically define a sampling distribution for a chosen test statistic. The only assumption necessary for valid results is the random assignment of experimental units to the test groups or treatments. Application to a real data set illustrates the advantages of these methods, including freedom from distribution assumptions without loss of power, complete choice over test statistics, easy adaptation to design complexities and missing data, and considerable intuitive appeal. The illustrations also reveal that computer-intensive methods can be more time consuming than conventional methods and the amount of computer code required to orchestrate reshuffling, resampling, or simulation procedures can be appreciable.

Analysis of Variance↗