PubMed HealthSearch

SEARCH · PubMed Health

Results for “missing data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Analysis strategies for serial multivariate ultrasonographic data that are incomplete.

Ultrasonographic measurement of intima-media thickness in the carotid artery has emerged as an important non-invasive means of assessing atherosclerosis, and has served to define primary outcome measures related to progression of arterial lesions in several large clinical trials and epidemiologic studies. It is characteristic that measurements often cannot be obtained from all sites during repeated examinations. This leads to incomplete multivariate serial data, for which the set and number of visualized sites may vary across time. We have contrasted several conditional and unconditional maximum likelihood analytical approaches, and have evaluated these with a simulation experiment based on characteristics of ultrasound measurements collected during the course of the Asymptomatic Carotid Artery Plaque Study. We examined analyses based on unweighted and generalized least squares regression in which we estimated cross-sectional summary statistics using raw means, unconditional maximum likelihood estimates and full maximum likelihood estimates. Since the genesis of missing data is not fully clear, and since the approaches we examined are based, to some degree, on the assumption that data are missing at random, we also examined the relative impact of deviations from such an assumption on each of the approaches considered. We found that maximum likelihood based approaches increased the expected efficiency of the analysis of serial ultrasound data over ignoring missing data by up to 21 per cent.

Arteriosclerosis

Applications of computer-intensive statistical methods to environmental research.

Conventional statistical approaches rely heavily on the properties of the central limit theorem to bridge the gap between the characteristics of a sample and some theoretical sampling distribution. Problems associated with nonrandom sampling, unknown population distributions, heterogeneous variances, small sample sizes, and missing data jeopardize the assumptions of such approaches and cast skepticism on conclusions. Conventional nonparametric alternatives offer freedom from distribution assumptions, but design limitations and loss of power can be serious drawbacks. With the data-processing capacity of today's computers, a new dimension of distribution-free statistical methods has evolved that addresses many of the limitations of conventional parametric and nonparametric methods. Computer-intensive statistical methods involve reshuffling, resampling, or simulating a data set thousands of times to empirically define a sampling distribution for a chosen test statistic. The only assumption necessary for valid results is the random assignment of experimental units to the test groups or treatments. Application to a real data set illustrates the advantages of these methods, including freedom from distribution assumptions without loss of power, complete choice over test statistics, easy adaptation to design complexities and missing data, and considerable intuitive appeal. The illustrations also reveal that computer-intensive methods can be more time consuming than conventional methods and the amount of computer code required to orchestrate reshuffling, resampling, or simulation procedures can be appreciable.

Analysis of Variance

The effect of question structure on self-reports of heavy drinking: closed-ended versus open-ended questions.

OBJECTIVE: We compared open-ended versus closed-ended questions on the frequency of consuming five or more drinks in a single sitting. METHOD: From a general population survey of Ontario adults (N = 2,022, 62% male), we analyzed a subsample of 649 respondents who reported drinking five or more drinks in a single sitting at least once in the past year. Differences in agreement between the two questions and rates of missing data were evaluated. RESULTS: For the most part, the two measures were not consistent, with the closed-ended question eliciting higher rates of heavier drinking. Rates of missing data were also higher for the open-ended question. CONCLUSIONS: Open-ended question may not necessarily be more suitable than closed-ended questions for estimating the frequency of heavy alcohol use.

Adult

Parametric and nonparametric linkage analysis: a unified multipoint approach.

In complex disease studies, it is crucial to perform multipoint linkage analysis with many markers and to use robust nonparametric methods that take account of all pedigree information. Currently available methods fall short in both regards. In this paper, we describe how to extract complete multipoint inheritance information from general pedigrees of moderate size. This information is captured in the multipoint inheritance distribution, which provides a framework for a unified approach to both parametric and nonparametric methods of linkage analysis. Specifically, the approach includes the following: (1) Rapid exact computation of multipoint LOD scores involving dozens of highly polymorphic markers, even in the presence of loops and missing data. (2) Non-parametric linkage (NPL) analysis, a powerful new approach to pedigree analysis. We show that NPL is robust to uncertainty about mode of inheritance, is much more powerful than commonly used nonparametric methods, and loses little power relative to parametric linkage analysis. NPL thus appears to be the method of choice for pedigree studies of complex traits. (3) Information-content mapping, which measures the fraction of the total inheritance information extracted by the available marker data and points out the regions in which typing additional markers is most useful. (4) Maximum-likelihood reconstruction of many-marker haplotypes, even in pedigrees with missing data. We have implemented NPL analysis, LOD-score computation, information-content mapping, and haplotype reconstruction in a new computer package, GENEHUNTER. The package allows efficient multipoint analysis of pedigree data to be performed rapidly in a single user-friendly environment.

Algorithms

Analysis of incomplete multivariate data from repeated measurement experiments.

This paper analyses two sets of data that consist of repeated measurements with missing data. The missing observations always occur at the end of the series of repeated measurements. The score test for multivariate normal data is used to compare treatment groups; if the original data are not multivariate normal they are replaced by expected normal scores.

Animals

Determinants of HIV infection among female commercial sex workers in northeastern Thailand: results from a longitudinal study.

Our objective was to estimate HIV seroconversion rates among commercial sex workers (CSWs) between 1990 and 1991 and to identify the behavioral, demographic, and reproductive determinants of these rates. This study has a prospective (n = 240 with 15 cases) and a cross-sectional component (n = 271 with 34 cases). In November 1990, HIV-negative female CSWs from 24 brothels in Khon Kaen city were interviewed and were followed prospectively for up to 1 year. In March, June, and September 1991, additional HIV-negative CSWs were enrolled and prospectively followed. HIV seroconversion rates were calculated, and the Cox regression model was used to estimate the relative risks of HIV seroconversion from demographic, sexual practice, and reproductive factors, adjusted for the effects of the others, among 232 of the 240 without missing data. Seroprevalence rates were also calculated for the 271 participants enrolled between March and December 1991, and relative risks of HIV seroprevalence were calculated for demographic, sexual practice, and reproductive risk factors among 184 of the 271 without missing data. The average seroprevalence was 12.5% (95% confidence interval 9.6-15.4%). With 1,947 person-months of observation obtained from 240 participants who were uninfected at baseline and seen at least twice during the course of the study, the cumulative incidence of HIV seroconversion between November 1990 and December 1991 was 9.4% (95% confidence interval 5.4-13.4%), and the average incidence rate of HIV seroconversion was 9.2 per 100 person-years (95% confidence interval 4.6-13.9 per 100 person-years). In the multivariate analysis, later date of enrollment into the study, having < 3 months experience as a CSW, and use of injectable contraceptives were the only risk factors that remained significant, with relative risks of 2.1 (95% confidence interval 1.2-3.7) for enrollment 3 months later, 3.8 (95% confidence interval 1.0-14.4) for < 3 months experience as a CSW versus > 3 months experience, and 3.9 (95% confidence interval 1.3-11.8) [corrected] for use of injectable contraceptives. In multivariate analysis of the cross-sectional data with 184 participants, of whom 21 were HIV seropositive, risk of HIV seropositivity increased significantly with current syphilis infection (odds ratio 5.8, 95% confidence interval 1.1-31.0). The results of this study will contribute to a better understanding of the risk factors of infection with HIV and thus allow for better targeting of group-specific interventions, particularly for CSWs and their clients. Further investigation of a possible association between injectable contraceptive use and HIV infection is needed.

Adolescent

Two methods for recommending bat weights.

Baseball players swung very light and very heavy bats through our instrument and the speed of the bat was recorded. These data were used to make mathematical models for each person. Then these models were coupled with equations of physics for bat-ball collisions to compute the Ideal Bat Weight for each individual. However, these calculations required the use of a sophisticated instrument that is not conveniently available to most people. So, we tried to find items in our database that correlated with Ideal Bat Weight. However, because many cells in the database were empty, we could not use traditional statistical techniques or even neural networks. Therefore, three new methods were used to estimate the missing data: (i) a neural network was trained using subjects that had no empty cells, then that neural network was used to predict the missing data, (ii) the data patching facility of a commercial software package was used, and (iii) the empty cells were filled with random numbers. Then, using these fully populated databases, several simple models were derived for recommending bat weights.

Adolescent

Prevalence and patterns of same-gender sexual contact among men.

The prevalence and patterns of same-gender sexual contact among men are key components of models of the spread of HIV infection and AIDS in the U.S. population. Previous estimates by Kinsey et al. from data collected between 1938 and 1948 have been widely criticized for inadequacies of sample design. New lower-bound estimates of prevalence developed from data from a national sample survey conducted in 1970 indicate that minimums of 20.3 percent of adult men in the United States in 1970 had sexual contact to orgasm with another man at some time in life; 6.7 percent had such contact after age 19; and between 1.6 and 2.0 percent had such contact within the previous year. Although these estimates incorporate adjustments for missing data, the likelihood of underreporting suggests that these estimates might be lower bounds on the prevalence of same-gender sex among men. Two sets of alternative estimates are derived to assess the sensitivity of these estimates to the assumptions made in imputing values to missing data. Detailed estimates are presented by frequency of contact, age, education, and marital status; and supporting estimates are derived from a 1988 national survey. Data from both the 1970 and 1988 surveys indicate that never-married men are more likely than other men to have had same-gender sexual contacts within the last year. The 1970 survey also indicates, however, that approximately half the men estimated to have such contacts are found among the more numerous population of currently or previously married men.

Adult

Assessing inner-city patients' hospital experiences. A controlled trial of telephone interviews versus mailed surveys.

OBJECTIVES: Obtaining accurate and representative patient-centered data may be difficult among poor, inner-city patients because of changing addresses, variable access to telephones, and a higher prevalence of illiteracy than in the populations in which many survey instruments were developed and tested. Assumptions about the usefulness of mailed surveys versus telephone interviews may not hold for the urban poor. Therefore, identifying the most efficient mode of survey administration in this population becomes an important methodological question. METHODS: We conducted a randomized trial of patients discharged from the inpatient medicine service of an urban teaching hospital to compare telephone interview with mailed self-administration of a detailed instrument for measuring patients' experiences with hospital care. Our primary outcomes were response rate, missing data, and data collection costs. Patients were excluded if they were not discharged to home or were mentally or physically unable to complete mailed or telephone interviews. The research assistant contacted eligible patients while hospitalized, informed them of the postdischarge survey, and obtained current phone numbers and addresses. Patients then were randomized to receive a 116-item satisfaction survey via one of two survey methods: mail-first (mailed surveys with follow-up on nonrespondents by telephone) or telephone-first (telephone interviews with follow-up of nonrespondents by mail). RESULTS: Of the 252 patients enrolled, 130 were randomized to the mail-first and 122 to the telephone-first method. Response rates were higher with the telephone-first (73%) compared with the mail-first method (50%; P < 0.0001). Surveys obtained by the telephone-first method had fewer missing data (0.7 +/- 2.39) for those items not involved in skip patterns compared with the mail-first method (7.1 +/- 12.3; P < 0.001) and were 42% less expensive per completed survey ($26.32 versus $37.35; P < 0.0001). CONCLUSIONS: In this survey of patients served by an urban teaching hospital, a strategy of telephone interviews with mail follow-up proved less expensive and yielded a higher response rate with more complete data than using a method where mailed surveys were followed by back-up telephone interviews. In addition, we believe that the improved response rate for telephone interviews compared with those reported in the literature for similar populations is the result of informing inpatients of the survey and obtaining telephone numbers and addresses in the hospital.

Female

Digital Mindfulness Intervention for Pregnant Women With Affective Disorders and Acute Stress Reactions: Prespecified Secondary Analysis of a Randomized Controlled Trial.

BACKGROUND: Pregnant women with ICD-10 (International Statistical Classification of Diseases, Tenth Revision) affective or stress-related disorders face an elevated risk of perinatal depression and anxiety, yet evidence on digital nonpharmacologic interventions for this population remains limited. OBJECTIVE: This study evaluated the effectiveness of an 8-week digital mindfulness-based intervention (eMBI) compared with treatment as usual (TAU) among pregnant women with ICD-10 affective or stress-related disorders participating in a randomized controlled trial (RCT). METHODS: This prespecified secondary analysis was conducted within a multicenter RCT in Baden-W&#xfc;rttemberg, Germany. Pregnant women aged 18 years and older with elevated depressive symptoms (Edinburgh Postnatal Depression Scale [EPDS]>9) and ICD-10-diagnosed affective or stress-related disorders were randomized 1:1 to eMBI or TAU. The intervention consisted of 8 weekly app-based mindfulness sessions (45 min each) delivered during gestational weeks 29-36, with no direct therapist contact. The primary outcome was continuous depressive symptom severity measured with the EPDS at 4-6 weeks post partum. Secondary outcomes included the EPDS at 6 months post partum, generalized anxiety (State-Trait Anxiety Inventory-State [STAI-S], State-Trait Anxiety Inventory-Trait [STAI-T]), and Pregnancy-Related Anxiety Questionnaire-Revised (PRAQ-R). Analyses followed the intention-to-treat (ITT) principle, using mixed models for repeated measures and multiple imputation. RESULTS: Of the 5299 screened women, 147 met the inclusion criteria for this subgroup analysis (intervention group [IG] had n=73 women and control group had n=74 women). Groups were comparable at baseline. The IG showed significantly greater reductions in EPDS scores at gestational week 34 (&#x394;=-2.21, P=.01), week 36 (&#x394;=-3.25, P=.01), and 4-6 weeks post partum (&#x394;=-4.81, P=.007). Treatment effects remained robust under conservative missing-data assumptions. At 4-6 weeks post partum, a higher proportion of participants in the IG achieved clinically meaningful improvement (31/73, 42.5% vs 21/74, 28.4%; adjusted odds ratio 1.56, 95% CI 1.19-2.05; P=.001). Anxiety outcomes followed a similar pattern, whereas pregnancy-related anxiety did not differ between groups. CONCLUSIONS: In this prespecified subgroup of pregnant women with ICD-10 affective or stress-related disorders, the eMBI was associated with clinically meaningful reductions in depressive symptoms from late pregnancy to 4-6 weeks post partum. Effects at 6 months post partum were attenuated and less stable across missing-data assumptions. These findings support eMBIs as a scalable, nonpharmacological adjunct to perinatal mental health care for women with affective or stress-related disorders, while confirmation in adequately powered trials with strategies to reduce postpartum attrition is warranted.

Humans

Performance characteristics of a composite multivariate quality control system.

We present the results of an evaluation of the performance characteristics of a composite multivariate quality control (CMQC) system that incorporates quality control rules for univariate, multivariate, and correlation conditions. The CMQC system evaluated is designed to help analysts detect unacceptable trends and systematic error in one or more variables, unacceptable random error in one or more variables, and unacceptable changes in the correlation structure of any pair of variables. It is also designed to be tolerant of missing data, to allow analysts to reject as few as one or as many as all variables in a run, and to provide analysts with control statistics and graphics that logically relate to sources of analytical error. We show that the various components of the CMQC system have adequate statistical power to detect systematic errors, random errors, and correlation changes under the conditions likely to be encountered with multivariate analytical measurement systems: (1) a single variable with increased systematic or random error; (2) all variables or a subgroup of variables affected by a common problem that increases systematic or random error; and (3) missing data for one or more variables in a run. We also show that the power of the multivariate component of the CMQC system to detect systematic and random errors is higher than the power of an alternative multivariate test criterion.

Chemistry Techniques, Analytical

Data mining issues for improved birth outcomes.

Issues obstructing progress in data mining for improved health outcomes include data quality problems, data redundancy, data inconsistency, repeated measures, temporal (time-contextual) measures, and data volume. Related issues involve theoretical and technical problems involving uncertainty management, missing data and missing values, and matching appropriate data mining techniques to patient data sets. Results of data mining research in progress are reported for Duke University's perinatal database that contains nearly a decade of clinical patient data, 71,753 database (patient) records and 4-5000 variables per patient.

Artificial Intelligence

[Study on distribution form of mesiodistal crown diameter in large sample: Part II].

The purpose of this research was to examine the distribution of the tooth size in a large sample. The objective teeth were the left upper and lower fourteen teeth except the third molar. The tooth size of 1,000 dental casts from the Japanese female orthodontic patients was measured. On each of them, a histogram and a set of statistics (mean, standard deviation, coefficient of variation, skewness, kurtosis, Geary value) are given in order to examine the distribution. The findings are as follows: 1) Each tooth may be classified into the following four types of distribution except the congenitally missing data. TYPE I: A normal distribution was observed in the upper and lower central incisors, the lower lateral incisor, the lower canine, the upper and lower first premolars, the upper second premolar, the upper and lower first molars and the lower second molar. TYPE II: A positively skewed distribution was observed in the lower second premolar. TYPE III: A negatively skewed and leptokurtic distribution was observed in the upper canine and the upper second molar. TYPE IV: An extremely negatively skewed and leptokurtic distribution was observed in upper lateral incisor. 2) With the four teeth which were classified into TYPE II, TYPE III and TYPE IV, the distribution of the lower second premolar was concluded to be of normal distribution by logarithmic transformation. The distribution of the upper canine and the upper second molar was judged to be of lognormal distribution and the upper lateral incisor also was judged to be of three parameter lognormal distribution and four parameter lognormal distribution. 3) The distribution of thirteen teeth except the upper lateral incisor was judged to be of normal distribution, by considering the congenitally missing data and the outlier in statistical data of the tooth.

Asian People

Pre-natal blood lead levels and learning difficulties in children: an analysis of non-randomly missing categorical data.

This paper presents an analysis of categorical variables subject to non-response. We incorporate the incomplete data into the analysis by modelling the distribution of the variables of interest and the non-response mechanism. We discuss issues of model selection and interpretation and the effect of discarding incomplete observations. In addition, we describe how to perform all of the computations with standard statistical software. We discuss the problem of incomplete categorical data within the context of a study of the effect of lead exposure on learning difficulties in children. In this study, many of the children are not observed on some of the variables of interest. It is particularly important in this study to incorporate the incomplete data, since there is evidence that non-response is related to the variables of interest. We reach different conclusions when we incorporate the incomplete data into the analysis than we reach when we discard the incomplete data. We also examine the sensitivity of our conclusions to the choice of a model for the non-response mechanism.

Algorithms

Correcting single channel data for missed events.

Interpretation of currents recorded from single ion channels in cellular membranes or lipid bilayers is complicated by the necessarily limited time resolution of the recording and detection systems. All intervals less than a certain duration, depending on the frequency response of the system, are not detected. Such missed events produce increases in the durations of observed open and shut intervals. In order to obtain the true kinetic scheme and rate constants underlying the observed activity, it is necessary to take into account missed events. We develop methods to correct for missed events for models with two or more states, including models with multiple open and shut states, compound states, and loops. Our methods can be used in a forward direction to predict observed distributions of open and shut intervals for a given kinetic scheme and time resolution. They can also be used in a backwards direction with iterative methods to determine rate constants consistent with the observed distributions. While a given kinetic scheme with rate constants predicts unique observed distributions of open and shut intervals, rate constants determined from observed distributions are not necessarily unique. Using these correction methods, we examine the effects of missed events for a five-state model consistent with some properties of large conductance Ca-activated K channels.

Ion Channels

Malaria vaccine trials: the missing qualitative data.

Recent population-based efficacy trials of the synthetic malaria vaccine SPf66 have shown restricted, if any, clinical protection against Plasmodium falciparum infection. Despite the well-established role of antibodies in effector responses against asexual blood-stage malaria parasites, the titres of anti-SPf66 IgG antibodies do not correlate with the ability of sera from vaccine recipients to inhibit parasite growth in vitro nor with partial clinical protection which could be detected in some trials. Qualitative or functional parameters of SP66-induced antibody responses, such as IgG subclass composition and affinity, may be more predictive of clinical protection against malaria than quantitative estimates of antibody concentration or titre. Since these parameters are readily estimated by laboratory techniques currently available, and may be modulated by changes in vaccination protocols and by the use of different adjuvants, a better understanding of qualitative antibody responses induced by SPf66 and other asexual blood-stage malaria vaccine candidates, and of their relationship with clinical protection in vivo, is urgently needed for the improvement of currently used immunization schedules.

Animals

Research in physical medicine and rehabilitation. VIII. Preliminary data analysis.

This paper describes important aspects of preliminary data analysis to be taken after data are checked for clerical entry errors and before the primary statistical analysis is performed. These include description and graphic display of each variable, recoding categorical data, transforming continuous data into another continuous variable and recoding continuous to categorical data. Missing values and outlying data points are identified and several techniques are recommended to minimize mistakes in variable recoding. Related variables measured with different units may be combined by using the z transformation and converted back to one of the original units for ease of interpretation. Finally, both categorical and continuous variables are checked for reliability by using kappa or the intraclass R.

Data Collection