PubMed HealthSearch

SEARCH · PubMed Health

Results for “missing data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

vcfsim: flexible simulation of all-sites VCFs with missing data.

BACKGROUND |: VCFs are the most widely used data format for encoding genetic variation. By design, standard VCFs do not include data from sites where all individuals are homozygous for the reference allele ("invariant sites") and thus do not differentiate these from sites where data are completely missing. However, missing data are a key feature of biological datasets across all domains of genomics, and many recent studies have shown that missing data can introduce a variety of statistical biases in the estimation of key population genetic parameters. A solution to this limitation is to include invariant sites in a standard VCF, creating an "all-sites VCF", exposing missing and invariant sites explicitly. One hurdle to the wider adoption of all-sites VCFs is a reliable parameterized simulation framework for generating biologically realistic all-sites VCFs. RESULTS |: Here, we introduce an open-source command line tool, vcfsim, that interfaces with the popular coalescent simulation platform msprime and provides convenience functions for simulating all-sites VCFs with variable levels of ploidy and missing data. We show that the post-processed VCFs generated using vcfsim align precisely with population genetic expectations (i.e. are statistically identical to raw msprime output), accurately introduce missing data, and permit the simulation of data with varying ploidy levels, including the simulation of intraindividual ploidy variation (e.g. heterogametic sex chromosomes) and population structures. CONCLUSIONS |: Our results vcfsim is a useful and easy-to-use tool for the benchmarking of new software tools, performing population genetic inference, training of machine learning models, and the exploration of the effects of missing data in genomics data sets.

Benchmarking

A framework for block-wise missing data in multi-omics.

High-throughput technologies have generated vast amounts of omic data. It is a consensus that the integration of diverse omics sources improves predictive models and biomarker discovery. However, managing multiple omics data poses challenges such as data heterogeneity, noise, high-dimensionality and missing data, especially in block-wise patterns. This study addresses the challenges of high dimensionality and block-wise missing data through a regularization and constrained-based approach. The methodology is implemented in the R package bwm for binary and continuous response variables, and applied to breast cancer and exposome multi-omics datasets, achieving strong performance even in scenarios with missing data present in all omics. In binary classification task, our proposed model achieves accuracy in the range of 86% to 92%, and F1 in the range of 68% to 79%. And, in regression task the correlation between true and predicted responses is in the range of 72% to 76%. However, there is a slight decline in performance metrics as the percentage of missing data increases. In scenarios where block-wise missing data affects multiple omics, the model performance actually surpasses that of scenarios where missing data is present in only one omics. One possible explanation for this might be that the other scenarios introduce a greater diversity of observation profiles, leading to a more robust model. Depending on the specific omics being studied, there is greater consistency in feature selection when comparing block-wise missing data scenarios.

Humans

miss-SNF: a multimodal patient similarity network integration approach to handle completely missing data sources.

MOTIVATION: Precision medicine leverages patient-specific multimodal data to improve prevention, diagnosis, prognosis, and treatment of diseases. Advancing precision medicine requires the non-trivial integration of complex, heterogeneous, and potentially high-dimensional data sources, such as multi-omics and clinical data. In the literature, several approaches have been proposed to manage missing data, but are usually limited to the recovery of subsets of features for a subset of patients. A largely overlooked problem is the integration of multiple sources of data when one or more of them are completely missing for a subset of patients, a relatively common condition in clinical practice. RESULTS: We propose miss-Similarity Network Fusion (miss-SNF), a novel general-purpose data integration approach designed to manage completely missing data in the context of patient similarity networks. miss-SNF integrates incomplete unimodal patient similarity networks by leveraging a non-linear message-passing strategy borrowed from the SNF algorithm. miss-SNF is able to recover missing patient similarities and is "task agnostic", in the sense that can integrate partial data for both unsupervised and supervised prediction tasks. Experimental analyses on nine cancer datasets from The Cancer Genome Atlas (TCGA) demonstrate that miss-SNF achieves state-of-the-art results in recovering similarities and in identifying patients subgroups enriched in clinically relevant variables and having differential survival. Moreover, amputation experiments show that miss-SNF supervised prediction of cancer clinical outcomes and Alzheimer's disease diagnosis with completely missing data achieves results comparable to those obtained when all the data are available. AVAILABILITY AND IMPLEMENTATION: miss-SNF code, implemented in R, is available at https://github.com/AnacletoLAB/missSNF.

Humans

An application of multivariate ratio methods for the analysis of a longitudinal clinical trial with missing data.

This paper presents an analysis of a longitudinal multi-center clinical trial with missing data. It illustrates the application, the appropriateness, and the limitations of a straightforward ratio estimation procedure for dealing with multivariate situations in which missing data occur at random and with small probability. The parameter estimates are computed via matrix operators such as those used for the generalized least squares analysis of catetorical data. Thus, the estimates may be conveniently analyzed by asymptotic regression methods within the same computer program which computes the estimates, provided that the sample size is sufficiently computer program which computes the estimates, provided that the sample size is sufficiently large.

Clinical Trials as Topic

The reporting and handling of missing data in genetic epidemiological studies of mental health in childhood and adolescence: A systematic review.

BACKGROUND: Genetic epidemiological analyses of child and adolescent mental health often use data from prospective longitudinal cohorts. Missingness due to selective attrition is therefore an important potential source of bias in such analyses. Informatively reporting on missingness and taking appropriate steps to handle it in analyses can mitigate this potential bias. Here, we aim to systematically assess how researchers report and address missingness in genetic epidemiological studies of child and adolescent mental health-related outcomes using cohort data. METHODS: We systematically searched the Ovid Medline database for studies published between August 2012 and August 2025, reporting polygenic score, genome-wide association, or Mendelian randomization analyses, of data on children or adolescents participating in cohort studies. We extracted information from eligible studies based on criteria adapted from the strengthening and reporting of observational studies in epidemiology (STROBE) guidelines. RESULTS: A total of 133 eligible studies were included, of which 125 (93.98%) reported the number of complete cases in all waves, while 84 (63.16%) detailed the amount of missingness on all key variables. Most studies used complete case analysis, while 39 studies explicitly reported applying other methods to handle missingness, with multiple imputation (n = 20, 15.04%) being the most common, followed by full information maximum likelihood 10 (8.1%). Only 18 studies (13.53%) reported an assumed missing mechanism along with the method used to address missingness. Full reporting of both the extent and handling of missingness at the item level was rare, occurring in only 5 (3.76%) and 15 (11.28%) studies, respectively, among the 123 studies that used multi-item instruments. CONCLUSION: Best practice recommendations for reporting on missing data handling emphasize the importance of detailing the proportion of missingness, types of mechanisms underpinning missingness, and details of approaches used. Based on this review, these recommendations for proper reporting of missing data are rarely followed in full.

children and adolescents

Statistical analysis of longitudinal quality of life data with missing measurements.

The statistical analysis of longitudinal quality of life data in the presence of missing data is discussed. In cancer trials missing data are generated due to the fact that patients die, drop out, or are censored. These missing data are problematic in the monitoring of the quality of life during the trial. However, by means of assuming that the cause of the missing data lies in the observed history of the patients and not in their unobserved future, the missing data are ignorable. Consequently, all available data can be used to estimate quality of life change patterns with time. The computations that are required are illustrated with real quality of life data and three commonly used computer packages for statistical analysis.

Analysis of Variance

Fitting growth curve models to longitudinal data with missing observations.

We discuss the analysis of growth curve data with missing or incomplete information. The approach is to fit subject-specific models and then to carry out an analysis in terms of the estimated parameters. This achieves reduction of data and eliminates the need for special considerations for subjects with missing data. Although there is no perfect substitute for complete data, our approach provides a way to handle missing data using a straightforward application of well-known statistical methodology.

Cephalometry

Biased estimation of the odds ratio in case-control studies due to the use of ad hoc methods of correcting for missing values for confounding variables.

The effects of missing values for a confounding variable are investigated in the setting of case-control studies in which, for simplicity, the effect of one binary risk factor and one categoric confounding variable on disease risk is under investigation. Some ad hoc techniques with which to deal with missing values are examined under different assumptions about the missing-data mechanism. Examples are given to illustrate that the magnitude of the bias that is introduced by applying an inadequate procedure can be large under circumstances that occur frequently in empiric research. This is true even for so-called complete case analysis, i.e., when only data on subjects with complete information are used. Appropriate bias corrections are derived. Making use of data on those subjects who are neglected in complete case analysis by creating an additional category always results in biased estimation. An alternative is to allocate these subjects to the cells of the contingency table in an appropriate manner. This approach yields consistent estimates if the data are missing at random. Choosing an appropriate method for dealing with missing values always requires some knowledge of why the data are missing. This suggests that investigators should carry out validation studies to understand whether the missing values occur randomly across the study population or occur more frequently in specific subgroups.

Bias

Building phenotypic character matrices for phylogenetic inference: exploration of 35 years of practice.

Recent methodological development in phylogenetic inference has focused predominantly on molecular data. However, renewed interest in other data types, particularly morphological data, has followed from the increased recognition of the power of total evidence and tip-dating approaches, including fossil data, for inference of time-scaled trees and rates of evolution. However, attention has largely focused on the improvement of models of morphological evolution and other analytical tools with much less discussion about data acquisition itself. Here we review past and current practice for describing and collecting morphological data for phylogenetic inference. We present a systematic review of 164 phylogenetic analyses conducted over the last 35 years and focused on a diverse group of extinct arthropods: trilobites. Trends in increasing matrix size, data type, and coding strategy are evident. Where present, polymorphic characters have been predominantly derived from discretized continuous characters, although increasingly practitioners are utilizing alternative approaches for the treatment of quantitative characters. Not surprisingly, traditional indices that describe character consistency are highly correlated with matrix size but show surprising variation at different taxonomic scales. More recent attempts to describe data quality using information theory imply that characters can have high information content even if data are missing for many tips, providing support against the exclusion of characters because of missing data. In consideration of this, as well as advances in the study of developmental biology and variational complexity, we identify several avenues for increasing the quality and quantity of morphological data going forward.

Phylogeny

Methods for the analysis of informatively censored longitudinal data.

This paper describes the problem of informative censoring in longitudinal studies where the primary outcome is rate of change in a continuous variable. Standard approaches based on the linear random effects model are valid only when the data are missing in a non-ignorable fashion. Informative censoring, which is a special type of non-ignorably missing data, occurs when the probability of early termination is related to an individual subject's true rate of change. When present, informative censoring causes bias in standard likelihood-based analyses, as well as in weighted averages of individual least-squares slopes. This paper reviews several methods proposed by others for analysis of informatively censored longitudinal data, and outlines a new approach based on a log-normal survival model. Maximum likelihood estimates may be obtained via the EM algorithm. Advantages of this approach are that it allows general unbalanced data caused by staggered entry and unequally-timed visits, it utilizes all available data, including data from patients with only a single measurement, and it provides a unified method for estimating all model parameters. Issues related to study design when informative censoring may occur are also discussed.

Linear Models

Autoregressive spectral models of heart rate variability. Practical issues.

Autoregressive time series model-based spectral estimates of heart period sequences can provide a parsimonious and visually attractive representation of the dynamics of interbeat intervals. While a corollary to Wold's decomposition theorem implies that the discrete Fourier periodogram spectral estimate and the autoregressive spectral estimate converge asymptotically, there are practical differences between the two approaches when applied to short blocks of data. Autoregressive spectra can achieve good frequency resolution and excellent statistical stability on short segments of heart period data of sinus origin. However, the order of the autoregressive model (number of free parameters to be estimated) must be explicitly chosen, a decision that influences the trade-off of frequency resolution with statistical stability. Akaike's Information Criterion (AIC), an information-theoretic rule for picking the optimum order, is sensitive to the aggregate amount of data in the analysis. Thus, the best model order for estimating the spectrum of a 4-minute segment of data will generally be lower than the best order for estimating an hourly spectrum based on averaging 15 4-minute spectra. A major advantage of the autoregressive model approach to spectral analysis is the ease with which it can be extended to handle messy data frequently seen in heart rate variability studies. A number of autoregressive-based robust-resistant techniques are available for the analysis of heart period sequences that contain a high volume of nonsinus and other unusual beats intervals. A theoretically satisfying framework is also available for spectral analysis of unevenly sampled data and missing data.

Fourier Analysis

Inductive learning of thyroid functional states using the ID3 algorithm. The effect of poor examples on the learning result.

The ID3 algorithm for inductive learning was tested using preclassified material for patients suspected to have a thyroid illness. Classification followed a rule-based expert system for the diagnosis of thyroid function. Thus, the knowledge to be learned was limited to the rules existing in the knowledge base of that expert system. The learning capability of the ID3 algorithm was tested with an unselected learning material (with some inherent missing data) and with a selected learning material (no missing data). The selected learning material was a subgroup which formed a part of the unselected learning material. When the number of learning cases was increased, the accuracy of the program improved. When the learning material was large enough, an increase in the learning material did not improve the results further. A better learning result was achieved with the selected learning material not including missing data as compared to unselected learning material. With this material we demonstrate a weakness in the ID3 algorithm: it can not find available information from good example cases if we add poor examples to the data.

Algorithms

A computer program for multivariate ratio analysis (MISCAT).

Analysts must deal frequently with missing data in multivariate analysis. In such cases, estimating the covariance maxtrix V of the dependent variables usually involves initial estimation and iterative adjustment of imputed missing data values, and/or smoothing of an estimate V which is not necessarily positive semi-definite. This paper presents an alternative procedure for computing estimates of relevant multivariate parameters in situations where missing data occur at random and with small probability. MISCAT is a computer program which computes multivariate ratio estimates of the means and a corresponding positive semi-definite estimate of the covariance matrix. It is an extension of GENCAT, which is a program for the generalizaed least squares analysis of categorical data. Thus, one advantage of dealing with missing data in this manner is that variation among the ratio estimates may be conveniently analyzed within MISCAT using asymptotic regression methodology, provided that sample sizes are sufficiently large. An example is given to illustrate such analysis for longitudinal data from a multicenter clinical trial.

Computers

National Survey of Family Growth: design, estimation, and inference.

The purpose of this report is to document the procedures used in the 1988 National Survey of Family Growth (NSFG) to select the sample, weight the data to produce national estimates, impute missing data, and estimate sampling errors. Therefore, this report necessarily contains a great deal of technical detail. For readers who do not need this level of detail, this summary briefly describes the procedures used. The National Survey of Family Growth is conducted every few years by the National Center for Health Statistics (NCHS), a part of the U.S. Department of Health and Human Services. The purpose of the survey is to collect and publish data from a national sample of women on childbearing, factors affecting childbearing (such as contraception, sterilization, and infertility), and related aspects of maternal and infant health. Interviewing for Cycle IV of the survey was done in 1988 by Westat, Inc., under a contract with NCHS. Personal interviews were conducted between January and August of 1988 with a national sample of 8,450 women in the civilian noninstitutionalized population of the United States. Interviews were conducted in person by trained female interviewers and lasted an average of 70 minutes. The interview focused on the woman's pregnancies, if any; her use of contraception; her ability to bear children (fecundity and infertility); her use of medical services for family planning, infertility, and prenatal care; her marriage and cohabitation history, if any; and a wide range of demographic and economic characteristics. This report describes some of the main methodological aspects of the survey, including the sample design, weighting, sampling errors, and imputation of missing data. These topics will be described briefly and less technically in this summary. Each topic is discussed in more detail in the rest of the report.

Adolescent

Prophylactic antibiotics to prevent chest infections in children with neurological impairment: the PARROT RCT.

BACKGROUND: Improvements in neonatal and paediatric care in recent decades have increased the survival of children with non-progressive neurological impairment. Respiratory disease in children with neurological impairment is common, with symptoms difficult to manage and lower respiratory tract infection occurring frequently. To reduce these, prophylactic antibiotics are being increasingly used, but the type, duration and dose of antibiotics can vary considerably, and there is limited evidence about their effectiveness in children and young people. A joint United Kingdom and Australia multicentre, randomised, double-blind, placebo-controlled trial comparing 52 weeks of azithromycin to placebo in children and young people with neurological impairment at risk of lower respiratory tract infection (PARROT) was planned to address this gap. PARROT was a multicentre, parallel group, blinded, pragmatic randomised controlled trial of 52-week duration with a planned sample size of 500 (250 in each arm) participants with neurological impairment. The primary outcome was the proportion of children and young people hospitalised with lower respiratory tract infection over the 52-week period. RESULTS: In total, 90 children and young people (62 in Australia, 28 in the United Kingdom) aged 3-17 years, with a diagnosed non-progressive, non-neuromuscular neurological impairment, who had persistent respiratory symptoms were randomised (1 : 1) to receive azithromycin or placebo. Baseline demographic and clinical characteristics were relatively well balanced across the two treatment groups and countries. Overall, mean (standard deviation) age was 9.2 (4.4) years, with 64% of participants having cerebral palsy, 67% being non-ambulant and 54% being totally tube-fed. At baseline, mean (standard deviation) numbers of hospital admissions with lower respiratory tract infection in the preceding year were 1.8 (2.0)/year, and general practitioner attendances 3.3 (3.0)/year. The PARROT trial was closed early to recruitment due to challenges arising from the COVID-19 pandemic. Sixty-five (72%) participants (azithromycin n = 30, placebo n = 35) completed 52 weeks of treatment and were not withdrawn early from the trial. Regarding the primary outcome, 11 (36.7%) in the azithromycin group were hospitalised with lower respiratory tract infection and 9 (25.7%) in the placebo group [absolute risk reduction 0.11 (95% confidence interval -0.12 to 0.33), relative risk 1.43 (95% confidence interval 0.68 to 2.97)]. Analysis of secondary outcome data was limited by the number of missing data, but parent-reported quality of life for young person and parent, sleep amount/quality for young person and parent, and respiratory symptoms were similar between groups and countries. LIMITATIONS: As PARROT was stopped early and was consequently underpowered, it is not possible to say whether azithromycin prophylaxis is any more effective than placebo in reducing the proportion of children admitted to hospital with lower respiratory tract infection after a 52-week period. CONCLUSIONS AND FUTURE WORK: Although we cannot comment on the effectiveness of prophylactic antibiotics in this context, we can draw some useful conclusions from this trial. Thus, the importance placed by families on hospitalisation and its prevalence in both treatment groups, even during the pandemic, would suggest that this is an appropriate primary outcome measure for future trials in this high-risk group of children and young people. Furthermore, the high attrition rate and large numbers of missing data, specifically for questionnaire-based outcomes at later follow-up points, should encourage researchers to be mindful of minimising trial burden to families for any future trials wherever possible. FUNDING: This synopsis presents independent research funded by the National Institute for Health and Care Research (NIHR) Health Technology Assessment programme as award number 16/17/01.

Humans

Single-channel data and missed events: analysis of a two-state Markov model.

Patch-clamp recording permits investigation of the gating kinetics of single ion channels. Careful statistical analysis of kinetic data can yield clues as to the molecular events underlying channel gating. However, it is important that such analysis should take full account of the limitations that arise from the finite time resolution of patch-clamp recording techniques. Single-ion-channel data are generally interpreted in terms of Markov process models of channel gating mechanisms. Experimental channel records suffer from time interval omission, i.e. failure to detect brief channel openings and closings. This leads to an identifiability problem when analysing single-channel data, i.e. different gating mechanisms provide equally convincing descriptions of the same experimental data. We consider a two-state Markov model of receptor-channel gating in which the channel opening rate is proportional to the agonist concentration, C in equilibrium with OA. By using computer-simulated data, the approximate likelihood of the data is maximized to yield parameter estimates for the model. At a single agonist concentration there is an identifiability problem in that two pairs of parameter estimates are obtained. The 'true' parameter estimates cannot be distinguished from the 'false' ones. By considering data corresponding to a range of agonist concentrations one may identify the 'true' parameter estimates as those that do not change as the agonist concentration is increased. Alternatively, one may identify the 'true' parameter estimates directly by maximizing a global likelihood, the latter being obtained by simultaneous consideration of data obtained at several different agonist concentrations.(ABSTRACT TRUNCATED AT 250 WORDS)

Animals