PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “missing data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Multivariate outlier detection applied to multiply imputed laboratory data.

In clinical laboratory safety data, multivariate outlier detection methods may highlight a patient whose laboratory measurements do not follow the same pattern of relationships as the majority of patients, although their individual measurements are not found to be outlying when considered one at a time. Missing data problems are often dealt with by imputing a single value as an estimate of the missing value. The completed data set may then be analysed using traditional methods. A disadvantage of using single imputation is the underestimation of variability, with a corresponding distortion of power in hypothesis testing. Multiple imputation methods attempt to overcome this problem, and in this paper a study is described which considers the application of multivariate outlier detection methods to multiply imputed clinical laboratory safety data sets. Three different proportions of missing data are generated in laboratory data sets of dimensions 4, 7, 12 and 30, and a comparison of eight multiple imputation methods is carried out. Two outlier detection techniques, Mahalanobis distance and generalized principal component analysis, are applied to the multiply imputed data sets, and their performances are discussed. Measures are introduced for assessing the accuracy of the missing data results, depending on which method of analysis is used.

Algorithms↗

Some conceptual and statistical issues in analysis of longitudinal psychiatric data. Application to the NIMH treatment of Depression Collaborative Research Program dataset.

Longitudinal studies have a prominent role in psychiatric research; however, statistical methods for analyzing these data are rarely commensurate with the effort involved in their acquisition. Frequently the majority of data are discarded and a simple end-point analysis is performed. In other cases, so called repeated-measures analysis of variance procedures are used with little regard to their restrictive and often unrealistic assumptions and the effect of missing data on the statistical properties of their estimates. We explored the unique features of longitudinal psychiatric data from both statistical and conceptual perspectives. We used a family of statistical models termed random regression models that provide a more realistic approach to analysis of longitudinal psychiatric data. Random regression models provide solutions to commonly observed problems of missing data, serial correlation, time-varying covariates, and irregular measurement occasions, and they accommodate systematic person-specific deviations from the average time trend. Properties of these models were compared with traditional approaches at a conceptual level. The approach was then illustrated in a new analysis of the National Institute of Mental Health Treatment of Depression Collaborative Research Program dataset, which investigated two forms of psychotherapy, pharmacotherapy with clinical management, and a placebo with clinical management control. Results indicated that both person-specific effects and serial correlation play major roles in the longitudinal psychiatric response process. Ignoring either of these effects produces misleading estimates of uncertainty that form the basis of statistical tests of hypotheses.

Analysis of Variance↗

Efficacy and safety of deucravacitinib, an oral, selective tyrosine kinase 2 inhibitor, in patients with active psoriatic arthritis: 52-week results from the randomised, double-blind, placebo-controlled phase 3 POETYK PsA-1 trial.

OBJECTIVES: The randomised, double-blind, placebo-controlled, phase 3 Program fOr Evaluation of TYK2 inhibitor Psoriatic Arthritis-1 (POETYK PsA-1) trial evaluated the efficacy, safety, and tolerability of deucravacitinib, an oral, selective tyrosine kinase 2 inhibitor, in patients with PsA na&#xef;ve to biologic disease-modifying antirheumatic drugs. METHODS: Adults with active PsA, high-sensitivity C-reactive protein concentration &#x2265; 3 mg/L, and &#x2265; 1 PsA-related hand and/or foot erosion detectable via radiograph were randomised 1:1 to oral deucravacitinib 6 mg once daily or placebo through week (W) 16. At W16, patients continued receiving deucravacitinib or switched from placebo to deucravacitinib through W52. The primary endpoint was American College of Rheumatology 20% improvement in response (ACR20) at W16. Nonresponder imputation was used for missing data. Efficacy and safety were evaluated through W52. Post hoc rank analysis of covariance was used to evaluate structural damage with no missing data imputation. RESULTS: In 670 patients, a significantly greater proportion of those receiving deucravacitinib vs placebo achieved ACR20 at W16 (54.2% vs 34.1%, P < .001). Responses with deucravacitinib were increased at W52. Patients who switched from placebo to deucravacitinib achieved improvements similar to those in patients who received continuous deucravacitinib. Inhibition of structural damage was observed at W16 and W52. At W16, incidences of serious adverse events (AEs) (deucravacitinib, 1.8%; placebo, 2.4%) and discontinuations due to AEs (2.4%; 1.8%) were low and remained low through W52, without imbalances in cardiovascular events, malignancies, or opportunistic infections. No new safety signals were detected; no deaths occurred. CONCLUSIONS: Deucravacitinib demonstrated superiority vs placebo for clinical responses, patient-reported outcomes, and structural damage inhibition in patients with PsA, with favourable tolerability and safety.

Humans↗

Catquest questionnaire for use in cataract surgery care: assessment of surgical outcomes.

PURPOSE: To demonstrate the outcome for patients after cataract extraction using the Catquest cataract questionnaire and discuss the models validity in assessing outcome. SETTING: Thirty-five Swedish departments of ophthalmology. METHODS: Patients having cataract extraction performed by surgeons from 35 Swedish departments of opthalmology participated in the study. The questionnaire was given to 2970 consecutive patients having surgery during March 1995 at the participating surgical units. The questionnaire was sent by mail to patients and completed on a voluntary basis. It focuses on visual disabilities in daily life, activity level, cataract symptoms, and degree of independence. The results form the questionnaire are interpreted using a benefit matrix that credits not only a decrease in visual disabilities and cataract symptoms but also an improvement in or maintenance of a preoperative activity level. RESULTS: Complete surgical outcome data and completed preoperative and postoperative questionnaires were available in 1933 cases (65.1%). Benefit from surgery according to the model was achieved by 90.9% of the patients. Patients having their second cataract extraction had the highest frequency of the greatest benefit form surgery. There was good agreement between the different levels of benefit from surgery according to the model and the patient's global rating of his or her vision or achieved visual acuity after surgery, respectively. Patients with missing data (did not return postoperative questionnaire or had missing surgical result variables) were older and had a higher frequency of other diseases and handicaps. CONCLUSION: The Catquest cataract questionnaire allowed the outcome of cataract surgery to be graded by different levels of benefit. There seemed to be good agreement between this model of assessment and the patient's global rating of his or her vision. Missing data may be a problem when a postal questionnaire is used.

Activities of Daily Living↗

A simulation study of the effects of assignment of prior identity-by-descent probabilities to unselected sib pairs, in covariance-structure modeling of a quantitative-trait locus.

Sib pair-selection strategies, designed to identify the most informative sib pairs in order to detect a quantitative-trait locus (QTL), give rise to a missing-data problem in genetic covariance-structure modeling of QTL effects. After selection, phenotypic data are available for all sibs, but marker data-and, consequently, the identity-by-descent (IBD) probabilities-are available only in selected sib pairs. One possible solution to this missing-data problem is to assign prior IBD probabilities (i.e., expected values) to the unselected sib pairs. The effect of this assignment in genetic covariance-structure modeling is investigated in the present paper. Two maximum-likelihood approaches to estimation are considered, the pi-hat approach and the IBD-mixture approach. In the simulations, sample size, selection criteria, QTL-increaser allele frequency, and gene action are manipulated. The results indicate that the assignment of prior IBD probabilities results in serious estimation bias in the pi-hat approach. Bias is also present in the IBD-mixture approach, although here the bias is generally much smaller. The null distribution of the log-likelihood ratio (i.e., in absence of any QTL effect) does not follow the expected null distribution in the pi-hat approach after selection. In the IBD-mixture approach, the null distribution does agree with expectation.

Alleles↗

Analysis strategies for serial multivariate ultrasonographic data that are incomplete.

Ultrasonographic measurement of intima-media thickness in the carotid artery has emerged as an important non-invasive means of assessing atherosclerosis, and has served to define primary outcome measures related to progression of arterial lesions in several large clinical trials and epidemiologic studies. It is characteristic that measurements often cannot be obtained from all sites during repeated examinations. This leads to incomplete multivariate serial data, for which the set and number of visualized sites may vary across time. We have contrasted several conditional and unconditional maximum likelihood analytical approaches, and have evaluated these with a simulation experiment based on characteristics of ultrasound measurements collected during the course of the Asymptomatic Carotid Artery Plaque Study. We examined analyses based on unweighted and generalized least squares regression in which we estimated cross-sectional summary statistics using raw means, unconditional maximum likelihood estimates and full maximum likelihood estimates. Since the genesis of missing data is not fully clear, and since the approaches we examined are based, to some degree, on the assumption that data are missing at random, we also examined the relative impact of deviations from such an assumption on each of the approaches considered. We found that maximum likelihood based approaches increased the expected efficiency of the analysis of serial ultrasound data over ignoring missing data by up to 21 per cent.

Arteriosclerosis↗

Applications of computer-intensive statistical methods to environmental research.

Conventional statistical approaches rely heavily on the properties of the central limit theorem to bridge the gap between the characteristics of a sample and some theoretical sampling distribution. Problems associated with nonrandom sampling, unknown population distributions, heterogeneous variances, small sample sizes, and missing data jeopardize the assumptions of such approaches and cast skepticism on conclusions. Conventional nonparametric alternatives offer freedom from distribution assumptions, but design limitations and loss of power can be serious drawbacks. With the data-processing capacity of today's computers, a new dimension of distribution-free statistical methods has evolved that addresses many of the limitations of conventional parametric and nonparametric methods. Computer-intensive statistical methods involve reshuffling, resampling, or simulating a data set thousands of times to empirically define a sampling distribution for a chosen test statistic. The only assumption necessary for valid results is the random assignment of experimental units to the test groups or treatments. Application to a real data set illustrates the advantages of these methods, including freedom from distribution assumptions without loss of power, complete choice over test statistics, easy adaptation to design complexities and missing data, and considerable intuitive appeal. The illustrations also reveal that computer-intensive methods can be more time consuming than conventional methods and the amount of computer code required to orchestrate reshuffling, resampling, or simulation procedures can be appreciable.

Analysis of Variance↗

The effect of question structure on self-reports of heavy drinking: closed-ended versus open-ended questions.

OBJECTIVE: We compared open-ended versus closed-ended questions on the frequency of consuming five or more drinks in a single sitting. METHOD: From a general population survey of Ontario adults (N = 2,022, 62% male), we analyzed a subsample of 649 respondents who reported drinking five or more drinks in a single sitting at least once in the past year. Differences in agreement between the two questions and rates of missing data were evaluated. RESULTS: For the most part, the two measures were not consistent, with the closed-ended question eliciting higher rates of heavier drinking. Rates of missing data were also higher for the open-ended question. CONCLUSIONS: Open-ended question may not necessarily be more suitable than closed-ended questions for estimating the frequency of heavy alcohol use.

Adult↗

Parametric and nonparametric linkage analysis: a unified multipoint approach.

In complex disease studies, it is crucial to perform multipoint linkage analysis with many markers and to use robust nonparametric methods that take account of all pedigree information. Currently available methods fall short in both regards. In this paper, we describe how to extract complete multipoint inheritance information from general pedigrees of moderate size. This information is captured in the multipoint inheritance distribution, which provides a framework for a unified approach to both parametric and nonparametric methods of linkage analysis. Specifically, the approach includes the following: (1) Rapid exact computation of multipoint LOD scores involving dozens of highly polymorphic markers, even in the presence of loops and missing data. (2) Non-parametric linkage (NPL) analysis, a powerful new approach to pedigree analysis. We show that NPL is robust to uncertainty about mode of inheritance, is much more powerful than commonly used nonparametric methods, and loses little power relative to parametric linkage analysis. NPL thus appears to be the method of choice for pedigree studies of complex traits. (3) Information-content mapping, which measures the fraction of the total inheritance information extracted by the available marker data and points out the regions in which typing additional markers is most useful. (4) Maximum-likelihood reconstruction of many-marker haplotypes, even in pedigrees with missing data. We have implemented NPL analysis, LOD-score computation, information-content mapping, and haplotype reconstruction in a new computer package, GENEHUNTER. The package allows efficient multipoint analysis of pedigree data to be performed rapidly in a single user-friendly environment.

Algorithms↗

Active life expectancy from annual follow-up data with missing responses.

Active life expectancy (ALE) at a given age is defined as the expected remaining years free of disability. In this study, three categories of health status are defined according to the ability to perform activities of daily living independently. Several studies have used increment-decrement life tables to estimate ALE, without error analysis, from only a baseline and one follow-up interview. The present work conducts an individual-level covariate analysis using a three-state Markov chain model for multiple follow-up data. Using a logistic link, the model estimates single-year transition probabilities among states of health, accounting for missing interviews. This approach has the advantages of smoothing subsequent estimates and increased power by using all follow-ups. We compute ALE and total life expectancy from these estimated single-year transition probabilities. Variance estimates are computed using the delta method. Data from the Iowa Established Population for the Epidemiologic Study of the Elderly are used to test the effects of smoking on ALE on all 5-year age groups past 65 years, controlling for sex and education.

Aged↗

A comparison of drugs versus placebo for the treatment of dysthymia.

OBJECTIVES: Dysthymia is a depressive disorder of chronic nature but of less severity than major depression, which depressive symptoms are more or less continuous for at least two years. The aim of this review was to conduct a systematic review of all RCTs comparing drugs and placebo for dysthymia. SEARCH STRATEGY: Electronic searches of Cochrane Library, EMBASE, MEDLINE, PsycLIT, Biological Abstracts and LILACS; reference searching; personal communication; conference abstracts; unpublished trials from the pharmaceutical industry; book chapters on the treatment of depression. SELECTION CRITERIA: The inclusion criteria for all randomised controlled trials were that they should focus on the use of drugs versus placebo for dysthymic patients. Exclusion criteria were: non randomised, mixed major depression/ dysthymia (trials not providing separate data) and depression secondary to other disorders (e.g. substance abuse). DATA COLLECTION AND ANALYSIS: The reviewers extracted the data independently. In order to achieve an intention-to-treat analysis, when trials failed to report it was assumed that people who died or dropped out had no improvement. Authors of relevant trials were contacted for additional and missing data. Absence of treatment response as defined by authors was the main measure of outcome used. Relative Risks (RR) and 95% confidence intervals (CI) of dichotomous data were calculated with the Random Effects Model. Where possible, number needed to treat (NNT) and number needed to harm (NNH) were estimated, taking the reciprocal of the absolute risk reduction. MAIN RESULTS: Currently the review includes 15 trials. Similar results were obtained in terms of efficacy for different groups of drugs, such as tricyclic (TCA), selective serotonin reuptake inhibitors (SSRI), monoamine oxidase inhibitors (MAOI) and other drugs (sulpiride, amineptine, and ritanserin). The pooled RR for absence of treatment response was 0. 68 (95% CI 0.59-0.78) for TCA and the NNT was 4.3 (95% CI 3.2-6.5). SSRIs showed similar RR for this outcome: 0.64 (95% CI 0.55-0.74), the NNT being 4.7 (95% CI 3.5-6.9). Concerning MAOIs, the RR was 0. 59 (95% CI 0.48-0.71) and the NNT was 2.9 (95% CI 2.2-4.3). Other drugs (amisulpride, amineptine and ritanserin) showed similar results in terms of absence of treatment response. Using more stringent criteria for improvement - full remission - the results were unchanged. Patients treated on TCA were more likely to report adverse events, compared with placebo. REVIEWER'S CONCLUSIONS: Drugs are effective in the treatment of dysthymia with no differences between and within class of drugs. Tricyclic antidepressants are more likely to cause adverse events and dropouts. As dysthymia is a chronic condition, there remains little information on quality of life and medium or long-term outcome.

Antidepressive Agents↗

Analysis of incomplete multivariate data from repeated measurement experiments.

This paper analyses two sets of data that consist of repeated measurements with missing data. The missing observations always occur at the end of the series of repeated measurements. The score test for multivariate normal data is used to compare treatment groups; if the original data are not multivariate normal they are replaced by expected normal scores.

Animals↗

Determinants of HIV infection among female commercial sex workers in northeastern Thailand: results from a longitudinal study.

Our objective was to estimate HIV seroconversion rates among commercial sex workers (CSWs) between 1990 and 1991 and to identify the behavioral, demographic, and reproductive determinants of these rates. This study has a prospective (n = 240 with 15 cases) and a cross-sectional component (n = 271 with 34 cases). In November 1990, HIV-negative female CSWs from 24 brothels in Khon Kaen city were interviewed and were followed prospectively for up to 1 year. In March, June, and September 1991, additional HIV-negative CSWs were enrolled and prospectively followed. HIV seroconversion rates were calculated, and the Cox regression model was used to estimate the relative risks of HIV seroconversion from demographic, sexual practice, and reproductive factors, adjusted for the effects of the others, among 232 of the 240 without missing data. Seroprevalence rates were also calculated for the 271 participants enrolled between March and December 1991, and relative risks of HIV seroprevalence were calculated for demographic, sexual practice, and reproductive risk factors among 184 of the 271 without missing data. The average seroprevalence was 12.5% (95% confidence interval 9.6-15.4%). With 1,947 person-months of observation obtained from 240 participants who were uninfected at baseline and seen at least twice during the course of the study, the cumulative incidence of HIV seroconversion between November 1990 and December 1991 was 9.4% (95% confidence interval 5.4-13.4%), and the average incidence rate of HIV seroconversion was 9.2 per 100 person-years (95% confidence interval 4.6-13.9 per 100 person-years). In the multivariate analysis, later date of enrollment into the study, having < 3 months experience as a CSW, and use of injectable contraceptives were the only risk factors that remained significant, with relative risks of 2.1 (95% confidence interval 1.2-3.7) for enrollment 3 months later, 3.8 (95% confidence interval 1.0-14.4) for < 3 months experience as a CSW versus > 3 months experience, and 3.9 (95% confidence interval 1.3-11.8) [corrected] for use of injectable contraceptives. In multivariate analysis of the cross-sectional data with 184 participants, of whom 21 were HIV seropositive, risk of HIV seropositivity increased significantly with current syphilis infection (odds ratio 5.8, 95% confidence interval 1.1-31.0). The results of this study will contribute to a better understanding of the risk factors of infection with HIV and thus allow for better targeting of group-specific interventions, particularly for CSWs and their clients. Further investigation of a possible association between injectable contraceptive use and HIV infection is needed.

Adolescent↗

What is meant by intention to treat analysis? Survey of published randomised controlled trials.

OBJECTIVES: To assess the methodological quality of intention to treat analysis as reported in randomised controlled trials in four large medical journals. DESIGN: Survey of all reports of randomised controlled trials published in 1997 in the BMJ, Lancet, JAMA, and New England Journal of Medicine. MAIN OUTCOME MEASURES: Methods of dealing with deviations from random allocation and missing data. RESULTS: 119 (48%) of the reports mentioned intention to treat analysis. Of these, 12 excluded any patients who did not start the allocated intervention and three did not analyse all randomised subjects as allocated. Five reports explicitly stated that there were no deviations from random allocation. The remaining 99 reports seemed to analyse according to random allocation, but only 34 of these explicitly stated this. 89 (75%) trials had some missing data on the primary outcome variable. The methods used to deal with this were generally inadequate, potentially leading to a biased treatment effect. 29 (24%) trials had more than 10% of responses missing for the primary outcome, the methods of handling the missing responses were similar in this subset. CONCLUSIONS: The intention to treat approach is often inadequately described and inadequately applied. Authors should explicitly describe the handling of deviations from randomised allocation and missing responses and discuss the potential effect of any missing response. Readers should critically assess the validity of reported intention to treat analyses.

Data Collection↗

Comparison of open and closed questionnaire formats in obtaining demographic information from Canadian general internists.

The objective of this study was to compare the impact of closed- versus open-ended question formats on the completeness and accuracy of demographic data collected in a mailed survey questionnaire. We surveyed general internists in five Canadian provinces to determine their career satisfaction. We randomized respondents to receive versions of the questionnaire in which 16 demographic questions were presented in a closed-ended or open-ended format. Two questions required respondents to make a relatively simple computation (ensuring that three or four categories of response added to 100%). The response rate was 1007/1192 physicians (80.0%). The proportion of respondents with no missing data for all 16 questions was 44.7% for open-ended and 67.0% for closed-ended formats (P < 0.001). The odds of having missing items remained higher for open-ended response options after adjusting for a number of respondent characteristics (2.67, 95% confidence interval 2.01 to 3.55). For the two questions requiring computations focused on professional activity and income, there were more missing data (P = 0.02, 0.02, respectively) but fewer inaccurate responses (P = 0.009, 0.20, respectively) for the open-ended compared to the closed-ended format. Investigators can achieve higher response rates for demographic items using closed format response options, but at the risk of increasing inaccuracy in response to questions requiring computation.

Adult↗

Two methods for recommending bat weights.

Baseball players swung very light and very heavy bats through our instrument and the speed of the bat was recorded. These data were used to make mathematical models for each person. Then these models were coupled with equations of physics for bat-ball collisions to compute the Ideal Bat Weight for each individual. However, these calculations required the use of a sophisticated instrument that is not conveniently available to most people. So, we tried to find items in our database that correlated with Ideal Bat Weight. However, because many cells in the database were empty, we could not use traditional statistical techniques or even neural networks. Therefore, three new methods were used to estimate the missing data: (i) a neural network was trained using subjects that had no empty cells, then that neural network was used to predict the missing data, (ii) the data patching facility of a commercial software package was used, and (iii) the empty cells were filled with random numbers. Then, using these fully populated databases, several simple models were derived for recommending bat weights.

Adolescent↗

Prevalence and patterns of same-gender sexual contact among men.

The prevalence and patterns of same-gender sexual contact among men are key components of models of the spread of HIV infection and AIDS in the U.S. population. Previous estimates by Kinsey et al. from data collected between 1938 and 1948 have been widely criticized for inadequacies of sample design. New lower-bound estimates of prevalence developed from data from a national sample survey conducted in 1970 indicate that minimums of 20.3 percent of adult men in the United States in 1970 had sexual contact to orgasm with another man at some time in life; 6.7 percent had such contact after age 19; and between 1.6 and 2.0 percent had such contact within the previous year. Although these estimates incorporate adjustments for missing data, the likelihood of underreporting suggests that these estimates might be lower bounds on the prevalence of same-gender sex among men. Two sets of alternative estimates are derived to assess the sensitivity of these estimates to the assumptions made in imputing values to missing data. Detailed estimates are presented by frequency of contact, age, education, and marital status; and supporting estimates are derived from a 1988 national survey. Data from both the 1970 and 1988 surveys indicate that never-married men are more likely than other men to have had same-gender sexual contacts within the last year. The 1970 survey also indicates, however, that approximately half the men estimated to have such contacts are found among the more numerous population of currently or previously married men.

Adult↗

Trends in trauma care in England and Wales 1989-97. UK Trauma Audit and Research Network.

BACKGROUND: In 1988, the Royal College of Surgeons reported major deficiencies in trauma care in UK hospitals. We investigated whether and how that care has changed in the last decade by use of data collected by the UK Trauma Audit and Research Network. METHODS: We analysed injury-severity, process, and outcome variables from 91602 patients' records on the database at the end of 1997, collected from 97 (49% of trauma-receiving) hospitals in England, Wales, and two in Ireland. We did longitudinal analyses of odds of death, process variables, and individual hospitals' performance. We took account of potential selection bias from missing data and recruitment of new hospitals. FINDINGS: The severity-adjusted odds of death after trauma declined gradually from 1989 (odds ratio 1997/1989 0.63 [95% CI [0.49-0.82]). In 1997, the reduction in odds of death was significant even after adjustment for missing data (ratio 1997/1989 0.72 [0.55-0.92]) and recruitment of new hospitals (0.64 [0.44-0.93]). There was significant variability in the proportion of survivors (adjusted for severity of injury and age) between the highest and lowest 10% of UK hospitals. The time between the call to the emergency services and arrival at hospital increased from 32 min in 1989 to 45 min in 1997, irrespective of injury severity. The proportion of severely injured patients seen first by senior doctors increased from 32% to 60%. INTERPRETATION: Hospital care has made a valuable but variable contribution to reductions in case fatality after injury in the UK in the past 10 years, though further improvement is possible.

Aged↗