PubMed HealthSearch

SEARCH · PubMed Health

Results for “missing data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Surrogate endpoints in clinical trials: cardiovascular diseases.

A surrogate endpoint in a cardiovascular clinical trial is defined as endpoint measured in lieu of some other so-called 'true' endpoint. A surrogate is especially useful if it is easily measured and highly correlated with the true endpoint. Often the 'true' endpoint is one with clinical importance to the patient, for example, mortality or a major clinical outcome, while a surrogate is one biologically closer to the process of disease, for example, ejection fraction. Use of the surrogate can often lead to dramatic reductions in sample size and much shorter studies than use of the true endpoint. We discuss several problems common in trials with surrogate endpoints. Most important is the effect of missing data, especially in the face of informative censoring. Possible solutions are the assignment of scores or formal penalties to missing data.

Cardiovascular Diseases

Evaluation of most frequent errors in daily compilation and use of a radiation treatment chart.

Between 1 March and 30 April (1994) we recorded the errors detected by the physician, the radiographer or the physicist during prescription, preparation and execution phases of 227 treatment plans. The radiation treatment modalities used were the following: (i) single or opposed fields, moulded or not; and (ii) multiple fields or kinetic techniques. The total number of sessions performed is 1613 with the cobalt unit and 2131 with the linear accelerator (total, 3744). The total number of wrong data is 155, consisting of 24/227 (10.5%) in compilation, 22/3744 (0.58%) in execution and 109/3744 (2.9%) in registration phases. The number of missing data is 140, consisting of 10/227 (4.4%) in compilation, 9/3744 (0.2%) in execution and 121/3744 (3.2%) in registration phases. Wrong data of compilation, even if in high rate (10.5%), were all found during the same compilation phase or at the first treatment, so that they did not alter the exactness of the treatment plan. Wrong and missing data, found in the registration phase (2.9% and 3.2%, respectively), depend on the repetition of daily treatment and on the registration of data on the chart after having digitized them on the display.

Cobalt Radioisotopes

Estimating the distribution of times from HIV seroconversion to AIDS using multiple imputation. Multicentre AIDS Cohort Study.

Multiple imputation is a model based technique for handling missing data problems. In this application we use the technique to estimate the distribution of times from HIV seroconversion to AIDS diagnosis with data from a cohort study of 4954 homosexual men with 4 years of follow-up. In this example the missing data are the dates of diagnosis with AIDS. The imputation procedure is performed in two stages. In the first stage, we estimate the residual AIDS-free time distribution as a function of covariates measured on the study participants with data provided by the participants who were seropositive at study entry. Specifically, we assume the residual AIDS-free times follow a log-normal regression model that depends on the covariates measured at enrolment on the seropositive participants. In the second stage we impute the date of AIDS diagnosis for the participants who seroconverted during the course of the study and are AIDS-free with use of the log-normal distribution estimated in the first stage and the covariates from each seroconverter's latest visit. The estimated proportions developing AIDS within 4 and within 7 years of seroconversion are 15 and 36 per cent respectively, with associated 95 per cent confidence intervals of (10, 21) and (26, 47) per cent. We discuss the Bayesian foundations of the multiple imputation technique and the statistical and scientific assumptions.

AIDS Serodiagnosis

Single-channel data and missed events: analysis of a two-state Markov model.

Patch-clamp recording permits investigation of the gating kinetics of single ion channels. Careful statistical analysis of kinetic data can yield clues as to the molecular events underlying channel gating. However, it is important that such analysis should take full account of the limitations that arise from the finite time resolution of patch-clamp recording techniques. Single-ion-channel data are generally interpreted in terms of Markov process models of channel gating mechanisms. Experimental channel records suffer from time interval omission, i.e. failure to detect brief channel openings and closings. This leads to an identifiability problem when analysing single-channel data, i.e. different gating mechanisms provide equally convincing descriptions of the same experimental data. We consider a two-state Markov model of receptor-channel gating in which the channel opening rate is proportional to the agonist concentration, C in equilibrium with OA. By using computer-simulated data, the approximate likelihood of the data is maximized to yield parameter estimates for the model. At a single agonist concentration there is an identifiability problem in that two pairs of parameter estimates are obtained. The 'true' parameter estimates cannot be distinguished from the 'false' ones. By considering data corresponding to a range of agonist concentrations one may identify the 'true' parameter estimates as those that do not change as the agonist concentration is increased. Alternatively, one may identify the 'true' parameter estimates directly by maximizing a global likelihood, the latter being obtained by simultaneous consideration of data obtained at several different agonist concentrations.(ABSTRACT TRUNCATED AT 250 WORDS)

Animals

Some conceptual and statistical issues in analysis of longitudinal psychiatric data. Application to the NIMH treatment of Depression Collaborative Research Program dataset.

Longitudinal studies have a prominent role in psychiatric research; however, statistical methods for analyzing these data are rarely commensurate with the effort involved in their acquisition. Frequently the majority of data are discarded and a simple end-point analysis is performed. In other cases, so called repeated-measures analysis of variance procedures are used with little regard to their restrictive and often unrealistic assumptions and the effect of missing data on the statistical properties of their estimates. We explored the unique features of longitudinal psychiatric data from both statistical and conceptual perspectives. We used a family of statistical models termed random regression models that provide a more realistic approach to analysis of longitudinal psychiatric data. Random regression models provide solutions to commonly observed problems of missing data, serial correlation, time-varying covariates, and irregular measurement occasions, and they accommodate systematic person-specific deviations from the average time trend. Properties of these models were compared with traditional approaches at a conceptual level. The approach was then illustrated in a new analysis of the National Institute of Mental Health Treatment of Depression Collaborative Research Program dataset, which investigated two forms of psychotherapy, pharmacotherapy with clinical management, and a placebo with clinical management control. Results indicated that both person-specific effects and serial correlation play major roles in the longitudinal psychiatric response process. Ignoring either of these effects produces misleading estimates of uncertainty that form the basis of statistical tests of hypotheses.

Analysis of Variance

Efficacy and safety of deucravacitinib, an oral, selective tyrosine kinase 2 inhibitor, in patients with active psoriatic arthritis: 52-week results from the randomised, double-blind, placebo-controlled phase 3 POETYK PsA-1 trial.

OBJECTIVES: The randomised, double-blind, placebo-controlled, phase 3 Program fOr Evaluation of TYK2 inhibitor Psoriatic Arthritis-1 (POETYK PsA-1) trial evaluated the efficacy, safety, and tolerability of deucravacitinib, an oral, selective tyrosine kinase 2 inhibitor, in patients with PsA na&#xef;ve to biologic disease-modifying antirheumatic drugs. METHODS: Adults with active PsA, high-sensitivity C-reactive protein concentration &#x2265; 3 mg/L, and &#x2265; 1 PsA-related hand and/or foot erosion detectable via radiograph were randomised 1:1 to oral deucravacitinib 6 mg once daily or placebo through week (W) 16. At W16, patients continued receiving deucravacitinib or switched from placebo to deucravacitinib through W52. The primary endpoint was American College of Rheumatology 20% improvement in response (ACR20) at W16. Nonresponder imputation was used for missing data. Efficacy and safety were evaluated through W52. Post hoc rank analysis of covariance was used to evaluate structural damage with no missing data imputation. RESULTS: In 670 patients, a significantly greater proportion of those receiving deucravacitinib vs placebo achieved ACR20 at W16 (54.2% vs 34.1%, P < .001). Responses with deucravacitinib were increased at W52. Patients who switched from placebo to deucravacitinib achieved improvements similar to those in patients who received continuous deucravacitinib. Inhibition of structural damage was observed at W16 and W52. At W16, incidences of serious adverse events (AEs) (deucravacitinib, 1.8%; placebo, 2.4%) and discontinuations due to AEs (2.4%; 1.8%) were low and remained low through W52, without imbalances in cardiovascular events, malignancies, or opportunistic infections. No new safety signals were detected; no deaths occurred. CONCLUSIONS: Deucravacitinib demonstrated superiority vs placebo for clinical responses, patient-reported outcomes, and structural damage inhibition in patients with PsA, with favourable tolerability and safety.

Humans

Analysis strategies for serial multivariate ultrasonographic data that are incomplete.

Ultrasonographic measurement of intima-media thickness in the carotid artery has emerged as an important non-invasive means of assessing atherosclerosis, and has served to define primary outcome measures related to progression of arterial lesions in several large clinical trials and epidemiologic studies. It is characteristic that measurements often cannot be obtained from all sites during repeated examinations. This leads to incomplete multivariate serial data, for which the set and number of visualized sites may vary across time. We have contrasted several conditional and unconditional maximum likelihood analytical approaches, and have evaluated these with a simulation experiment based on characteristics of ultrasound measurements collected during the course of the Asymptomatic Carotid Artery Plaque Study. We examined analyses based on unweighted and generalized least squares regression in which we estimated cross-sectional summary statistics using raw means, unconditional maximum likelihood estimates and full maximum likelihood estimates. Since the genesis of missing data is not fully clear, and since the approaches we examined are based, to some degree, on the assumption that data are missing at random, we also examined the relative impact of deviations from such an assumption on each of the approaches considered. We found that maximum likelihood based approaches increased the expected efficiency of the analysis of serial ultrasound data over ignoring missing data by up to 21 per cent.

Arteriosclerosis

Applications of computer-intensive statistical methods to environmental research.

Conventional statistical approaches rely heavily on the properties of the central limit theorem to bridge the gap between the characteristics of a sample and some theoretical sampling distribution. Problems associated with nonrandom sampling, unknown population distributions, heterogeneous variances, small sample sizes, and missing data jeopardize the assumptions of such approaches and cast skepticism on conclusions. Conventional nonparametric alternatives offer freedom from distribution assumptions, but design limitations and loss of power can be serious drawbacks. With the data-processing capacity of today's computers, a new dimension of distribution-free statistical methods has evolved that addresses many of the limitations of conventional parametric and nonparametric methods. Computer-intensive statistical methods involve reshuffling, resampling, or simulating a data set thousands of times to empirically define a sampling distribution for a chosen test statistic. The only assumption necessary for valid results is the random assignment of experimental units to the test groups or treatments. Application to a real data set illustrates the advantages of these methods, including freedom from distribution assumptions without loss of power, complete choice over test statistics, easy adaptation to design complexities and missing data, and considerable intuitive appeal. The illustrations also reveal that computer-intensive methods can be more time consuming than conventional methods and the amount of computer code required to orchestrate reshuffling, resampling, or simulation procedures can be appreciable.

Analysis of Variance

The effect of question structure on self-reports of heavy drinking: closed-ended versus open-ended questions.

OBJECTIVE: We compared open-ended versus closed-ended questions on the frequency of consuming five or more drinks in a single sitting. METHOD: From a general population survey of Ontario adults (N = 2,022, 62% male), we analyzed a subsample of 649 respondents who reported drinking five or more drinks in a single sitting at least once in the past year. Differences in agreement between the two questions and rates of missing data were evaluated. RESULTS: For the most part, the two measures were not consistent, with the closed-ended question eliciting higher rates of heavier drinking. Rates of missing data were also higher for the open-ended question. CONCLUSIONS: Open-ended question may not necessarily be more suitable than closed-ended questions for estimating the frequency of heavy alcohol use.

Adult

Parametric and nonparametric linkage analysis: a unified multipoint approach.

In complex disease studies, it is crucial to perform multipoint linkage analysis with many markers and to use robust nonparametric methods that take account of all pedigree information. Currently available methods fall short in both regards. In this paper, we describe how to extract complete multipoint inheritance information from general pedigrees of moderate size. This information is captured in the multipoint inheritance distribution, which provides a framework for a unified approach to both parametric and nonparametric methods of linkage analysis. Specifically, the approach includes the following: (1) Rapid exact computation of multipoint LOD scores involving dozens of highly polymorphic markers, even in the presence of loops and missing data. (2) Non-parametric linkage (NPL) analysis, a powerful new approach to pedigree analysis. We show that NPL is robust to uncertainty about mode of inheritance, is much more powerful than commonly used nonparametric methods, and loses little power relative to parametric linkage analysis. NPL thus appears to be the method of choice for pedigree studies of complex traits. (3) Information-content mapping, which measures the fraction of the total inheritance information extracted by the available marker data and points out the regions in which typing additional markers is most useful. (4) Maximum-likelihood reconstruction of many-marker haplotypes, even in pedigrees with missing data. We have implemented NPL analysis, LOD-score computation, information-content mapping, and haplotype reconstruction in a new computer package, GENEHUNTER. The package allows efficient multipoint analysis of pedigree data to be performed rapidly in a single user-friendly environment.

Algorithms

Analysis of incomplete multivariate data from repeated measurement experiments.

This paper analyses two sets of data that consist of repeated measurements with missing data. The missing observations always occur at the end of the series of repeated measurements. The score test for multivariate normal data is used to compare treatment groups; if the original data are not multivariate normal they are replaced by expected normal scores.

Animals

Determinants of HIV infection among female commercial sex workers in northeastern Thailand: results from a longitudinal study.

Our objective was to estimate HIV seroconversion rates among commercial sex workers (CSWs) between 1990 and 1991 and to identify the behavioral, demographic, and reproductive determinants of these rates. This study has a prospective (n = 240 with 15 cases) and a cross-sectional component (n = 271 with 34 cases). In November 1990, HIV-negative female CSWs from 24 brothels in Khon Kaen city were interviewed and were followed prospectively for up to 1 year. In March, June, and September 1991, additional HIV-negative CSWs were enrolled and prospectively followed. HIV seroconversion rates were calculated, and the Cox regression model was used to estimate the relative risks of HIV seroconversion from demographic, sexual practice, and reproductive factors, adjusted for the effects of the others, among 232 of the 240 without missing data. Seroprevalence rates were also calculated for the 271 participants enrolled between March and December 1991, and relative risks of HIV seroprevalence were calculated for demographic, sexual practice, and reproductive risk factors among 184 of the 271 without missing data. The average seroprevalence was 12.5% (95% confidence interval 9.6-15.4%). With 1,947 person-months of observation obtained from 240 participants who were uninfected at baseline and seen at least twice during the course of the study, the cumulative incidence of HIV seroconversion between November 1990 and December 1991 was 9.4% (95% confidence interval 5.4-13.4%), and the average incidence rate of HIV seroconversion was 9.2 per 100 person-years (95% confidence interval 4.6-13.9 per 100 person-years). In the multivariate analysis, later date of enrollment into the study, having < 3 months experience as a CSW, and use of injectable contraceptives were the only risk factors that remained significant, with relative risks of 2.1 (95% confidence interval 1.2-3.7) for enrollment 3 months later, 3.8 (95% confidence interval 1.0-14.4) for < 3 months experience as a CSW versus > 3 months experience, and 3.9 (95% confidence interval 1.3-11.8) [corrected] for use of injectable contraceptives. In multivariate analysis of the cross-sectional data with 184 participants, of whom 21 were HIV seropositive, risk of HIV seropositivity increased significantly with current syphilis infection (odds ratio 5.8, 95% confidence interval 1.1-31.0). The results of this study will contribute to a better understanding of the risk factors of infection with HIV and thus allow for better targeting of group-specific interventions, particularly for CSWs and their clients. Further investigation of a possible association between injectable contraceptive use and HIV infection is needed.

Adolescent

Two methods for recommending bat weights.

Baseball players swung very light and very heavy bats through our instrument and the speed of the bat was recorded. These data were used to make mathematical models for each person. Then these models were coupled with equations of physics for bat-ball collisions to compute the Ideal Bat Weight for each individual. However, these calculations required the use of a sophisticated instrument that is not conveniently available to most people. So, we tried to find items in our database that correlated with Ideal Bat Weight. However, because many cells in the database were empty, we could not use traditional statistical techniques or even neural networks. Therefore, three new methods were used to estimate the missing data: (i) a neural network was trained using subjects that had no empty cells, then that neural network was used to predict the missing data, (ii) the data patching facility of a commercial software package was used, and (iii) the empty cells were filled with random numbers. Then, using these fully populated databases, several simple models were derived for recommending bat weights.

Adolescent

Prevalence and patterns of same-gender sexual contact among men.

The prevalence and patterns of same-gender sexual contact among men are key components of models of the spread of HIV infection and AIDS in the U.S. population. Previous estimates by Kinsey et al. from data collected between 1938 and 1948 have been widely criticized for inadequacies of sample design. New lower-bound estimates of prevalence developed from data from a national sample survey conducted in 1970 indicate that minimums of 20.3 percent of adult men in the United States in 1970 had sexual contact to orgasm with another man at some time in life; 6.7 percent had such contact after age 19; and between 1.6 and 2.0 percent had such contact within the previous year. Although these estimates incorporate adjustments for missing data, the likelihood of underreporting suggests that these estimates might be lower bounds on the prevalence of same-gender sex among men. Two sets of alternative estimates are derived to assess the sensitivity of these estimates to the assumptions made in imputing values to missing data. Detailed estimates are presented by frequency of contact, age, education, and marital status; and supporting estimates are derived from a 1988 national survey. Data from both the 1970 and 1988 surveys indicate that never-married men are more likely than other men to have had same-gender sexual contacts within the last year. The 1970 survey also indicates, however, that approximately half the men estimated to have such contacts are found among the more numerous population of currently or previously married men.

Adult

Assessing inner-city patients' hospital experiences. A controlled trial of telephone interviews versus mailed surveys.

OBJECTIVES: Obtaining accurate and representative patient-centered data may be difficult among poor, inner-city patients because of changing addresses, variable access to telephones, and a higher prevalence of illiteracy than in the populations in which many survey instruments were developed and tested. Assumptions about the usefulness of mailed surveys versus telephone interviews may not hold for the urban poor. Therefore, identifying the most efficient mode of survey administration in this population becomes an important methodological question. METHODS: We conducted a randomized trial of patients discharged from the inpatient medicine service of an urban teaching hospital to compare telephone interview with mailed self-administration of a detailed instrument for measuring patients' experiences with hospital care. Our primary outcomes were response rate, missing data, and data collection costs. Patients were excluded if they were not discharged to home or were mentally or physically unable to complete mailed or telephone interviews. The research assistant contacted eligible patients while hospitalized, informed them of the postdischarge survey, and obtained current phone numbers and addresses. Patients then were randomized to receive a 116-item satisfaction survey via one of two survey methods: mail-first (mailed surveys with follow-up on nonrespondents by telephone) or telephone-first (telephone interviews with follow-up of nonrespondents by mail). RESULTS: Of the 252 patients enrolled, 130 were randomized to the mail-first and 122 to the telephone-first method. Response rates were higher with the telephone-first (73%) compared with the mail-first method (50%; P < 0.0001). Surveys obtained by the telephone-first method had fewer missing data (0.7 +/- 2.39) for those items not involved in skip patterns compared with the mail-first method (7.1 +/- 12.3; P < 0.001) and were 42% less expensive per completed survey ($26.32 versus $37.35; P < 0.0001). CONCLUSIONS: In this survey of patients served by an urban teaching hospital, a strategy of telephone interviews with mail follow-up proved less expensive and yielded a higher response rate with more complete data than using a method where mailed surveys were followed by back-up telephone interviews. In addition, we believe that the improved response rate for telephone interviews compared with those reported in the literature for similar populations is the result of informing inpatients of the survey and obtaining telephone numbers and addresses in the hospital.

Female

Digital Mindfulness Intervention for Pregnant Women With Affective Disorders and Acute Stress Reactions: Prespecified Secondary Analysis of a Randomized Controlled Trial.

BACKGROUND: Pregnant women with ICD-10 (International Statistical Classification of Diseases, Tenth Revision) affective or stress-related disorders face an elevated risk of perinatal depression and anxiety, yet evidence on digital nonpharmacologic interventions for this population remains limited. OBJECTIVE: This study evaluated the effectiveness of an 8-week digital mindfulness-based intervention (eMBI) compared with treatment as usual (TAU) among pregnant women with ICD-10 affective or stress-related disorders participating in a randomized controlled trial (RCT). METHODS: This prespecified secondary analysis was conducted within a multicenter RCT in Baden-W&#xfc;rttemberg, Germany. Pregnant women aged 18 years and older with elevated depressive symptoms (Edinburgh Postnatal Depression Scale [EPDS]>9) and ICD-10-diagnosed affective or stress-related disorders were randomized 1:1 to eMBI or TAU. The intervention consisted of 8 weekly app-based mindfulness sessions (45 min each) delivered during gestational weeks 29-36, with no direct therapist contact. The primary outcome was continuous depressive symptom severity measured with the EPDS at 4-6 weeks post partum. Secondary outcomes included the EPDS at 6 months post partum, generalized anxiety (State-Trait Anxiety Inventory-State [STAI-S], State-Trait Anxiety Inventory-Trait [STAI-T]), and Pregnancy-Related Anxiety Questionnaire-Revised (PRAQ-R). Analyses followed the intention-to-treat (ITT) principle, using mixed models for repeated measures and multiple imputation. RESULTS: Of the 5299 screened women, 147 met the inclusion criteria for this subgroup analysis (intervention group [IG] had n=73 women and control group had n=74 women). Groups were comparable at baseline. The IG showed significantly greater reductions in EPDS scores at gestational week 34 (&#x394;=-2.21, P=.01), week 36 (&#x394;=-3.25, P=.01), and 4-6 weeks post partum (&#x394;=-4.81, P=.007). Treatment effects remained robust under conservative missing-data assumptions. At 4-6 weeks post partum, a higher proportion of participants in the IG achieved clinically meaningful improvement (31/73, 42.5% vs 21/74, 28.4%; adjusted odds ratio 1.56, 95% CI 1.19-2.05; P=.001). Anxiety outcomes followed a similar pattern, whereas pregnancy-related anxiety did not differ between groups. CONCLUSIONS: In this prespecified subgroup of pregnant women with ICD-10 affective or stress-related disorders, the eMBI was associated with clinically meaningful reductions in depressive symptoms from late pregnancy to 4-6 weeks post partum. Effects at 6 months post partum were attenuated and less stable across missing-data assumptions. These findings support eMBIs as a scalable, nonpharmacological adjunct to perinatal mental health care for women with affective or stress-related disorders, while confirmation in adequately powered trials with strategies to reduce postpartum attrition is warranted.

Humans

Performance characteristics of a composite multivariate quality control system.

We present the results of an evaluation of the performance characteristics of a composite multivariate quality control (CMQC) system that incorporates quality control rules for univariate, multivariate, and correlation conditions. The CMQC system evaluated is designed to help analysts detect unacceptable trends and systematic error in one or more variables, unacceptable random error in one or more variables, and unacceptable changes in the correlation structure of any pair of variables. It is also designed to be tolerant of missing data, to allow analysts to reject as few as one or as many as all variables in a run, and to provide analysts with control statistics and graphics that logically relate to sources of analytical error. We show that the various components of the CMQC system have adequate statistical power to detect systematic errors, random errors, and correlation changes under the conditions likely to be encountered with multivariate analytical measurement systems: (1) a single variable with increased systematic or random error; (2) all variables or a subgroup of variables affected by a common problem that increases systematic or random error; and (3) missing data for one or more variables in a run. We also show that the power of the multivariate component of the CMQC system to detect systematic and random errors is higher than the power of an alternative multivariate test criterion.

Chemistry Techniques, Analytical

Data mining issues for improved birth outcomes.

Issues obstructing progress in data mining for improved health outcomes include data quality problems, data redundancy, data inconsistency, repeated measures, temporal (time-contextual) measures, and data volume. Related issues involve theoretical and technical problems involving uncertainty management, missing data and missing values, and matching appropriate data mining techniques to patient data sets. Results of data mining research in progress are reported for Duke University's perinatal database that contains nearly a decade of clinical patient data, 71,753 database (patient) records and 4-5000 variables per patient.

Artificial Intelligence