PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data commons”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

The establishment of the photo management system for orthodontics and its clinical application.

It is necessary for orthodontists to collect and analyze the patients' photographic data. By conventional methods, these photo data were commonly saved on film, which were frail and often resulted in data loss. Furthermore, it is not convenient for the orthodontists to consult and study these data during clinical research. A critical problem thus arises in managing these photo data scientifically. The computer technique and picture processing method were employed in the present study to establish the Photo Management System for Orthodontics (PMSO), which makes the administration of patients' photo data more scientific, convenient, and effective than before. This system is characterized as follows: (1) Clinical orthodontists designed and programmed the system, which is "close to clinical reality and serves the clinic". Orthodontists can easily use it without any training course. (2) The images can be imported from many devices, such as digital cameras, scanners, and electronic data storage media such as floppy disks. (3) Even two groups of photos (four images) could be displayed in the same window for comparison and study. This function allows orthodontists to easily observe the differences in the same patient's photos at different stages, which is important to obtain comprehensive knowledge of the patient. (4) The network technique makes it possible for the photos to be shared. The orthodontists can get the patients' photos through the network without going to the radiology department or the imaging room. (5) It is helpful for orthodontic research since the photos can be exported easily in different ways. (6) This is shared software; orthodontists can use it for free.

Computer Communication Networks↗

SimpleMicrobiome: An integrated web-based platform for streamlined microbiome data analysis and visualization.

Microbiome studies require multiple analytical steps after initial sequence processing. These steps commonly include data harmonization, preprocessing, taxonomic profiling, diversity analysis, differential abundance testing, predictive modeling, network inference, and preparation of publication-ready outputs. Although robust packages are available for many of these tasks, routine use often depends on command-line workflows, repeated data reformatting, and method-specific scripting. These requirements can limit accessibility for experimental researchers and complicate consistent analysis across interdisciplinary teams. We developed SimpleMicrobiome, a web-based R Shiny platform that integrates established microbiome analysis methods into a single interactive downstream workflow. The application accepts standard abundance, taxonomy, and metadata tables, supports interactive preprocessing and sample filtering, and provides modules for taxa profile visualization, alpha and beta diversity analysis, ANCOM-BC2 and MaAsLin2 differential abundance testing, Random Forest modeling with SHAP-based interpretation, microbial association network inference using SparCC and SPIEC-EASI through NetCoMi, correlation heatmaps, and dbRDA/CAP-style association biplots. The platform is implemented as a modular Shiny application so that preprocessing choices are propagated across downstream analyses, results can be exported as figures and tables, and the same application can be run through the public server, source-code installation, or a Docker image. SimpleMicrobiome consolidates major downstream microbiome analysis tasks in an accessible browser-based environment while retaining links to established analytical frameworks. The platform may reduce technical barriers for non-programming users, improve consistency across exploratory and reporting-oriented analyses, and support collaborative microbiome research. The public application is available at https://simplemicrobiome.mglab.org, the source code is available at https://github.com/yjcho2252/SimpleMicrobiome, and a Docker image for local deployment is available at https://hub.docker.com/r/mglab2252/simplemicrobiome.

differential abundance↗

Who 'controls' quality control data analysis?

A common quality control tool is peer group comparison of data from commercial controls. While its real-time effectiveness is limited, inappropriate statistical management of the data can cause an individual lab's performance to be misrepresented. Here are two examples where vendor-directed data analysis contained flagrant errors. The finding that vendors use inappropriate algorithms to compare accuracy and precision of peer performance suggests a need, the author believes, to set rigorous standards of reporting for the protection of participating laboratories.

California↗

Possible effects of non-HLA antibodies in common typing sera on HLA antigen frequency data in leukemia.

Selective adsorption of several common "monospecific" HLA typing sera with HLA typed platelets, purified B lymphocytes, cultured lymphoid cells, or lymphocytes from patients with active chronic lymphocytic leukemia demonstrated that many of these sera contain antibodies to non-HLA antigens. Antibodies were detected to antigens present on peripheral blood B-lymphocytes, cultured lymphoid cells and leukemic cells from patients with both myelocytic and lymphocytic forms of leukemia but absent from T-lymphocytes and platelets. Since these kinds of antibodies appear to be present in a large proportion of common HLA typing sera, caution should be used in interpreting all data related to HLA antigen expression in leukemia.

Antibodies↗

Identification of significant host factors for HIV dynamics modelled by non-linear mixed-effects models.

Non-linear mixed-effects models are powerful tools for modelling HIV viral dynamics. In AIDS clinical trials, the viral load measurements for each subject are often sparse. In such cases, linearization procedures are usually used for inferences. Under such linearization procedures, however, standard covariate selection methods based on the approximate likelihood, such as the likelihood ratio test, may not be reliable. In order to identify significant host factors for HIV dynamics, in this paper we consider two alternative approaches for covariate selection: one is based on individual non-linear least square estimates and the other is based on individual empirical Bayes estimates. Our simulation study shows that, if the within-individual data are sparse and the between-individual variation is large, the two alternative covariate selection methods are more reliable than the likelihood ratio test, and the more powerful method based on individual empirical Bayes estimates is especially preferable. We also consider the missing data in covariates. The commonly used missing data methods may lead to misleading results. We recommend a multiple imputation method to handle missing covariates. A real data set from an AIDS clinical trial is analysed based on various covariate selection methods and missing data methods.

Acquired Immunodeficiency Syndrome↗

Marginal regression of multivariate event times based on linear transformation models.

Multivariate event time data are common in medical studies and have received much attention recently. In such data, each study subject may potentially experience several types of events or recurrences of the same type of event, or event times may be clustered. Marginal distributions are specified for the multivariate event times in multiple events and clustered events data, and for the gap times in recurrent events data, using the semiparametric linear transformation models while leaving the dependence structures for related events unspecified. We propose several estimating equations for simultaneous estimation of the regression parameters and the transformation function. It is shown that the resulting regression estimators are asymptotically normal, with variance-covariance matrix that has a closed form and can be consistently estimated by the usual plug-in method. Simulation studies show that the proposed approach is appropriate for practical use. An application to the well-known bladder cancer tumor recurrences data is also given to illustrate the methodology.

Biomedical Research↗

[Computer based documentation of ultrasound data].

Usually, hospitals and private doctors are well equipped with computers. However, the documentation of ultrasound data is commonly paper-based. The paper presents a computer-based ultrasound data recording and reporting, using the ultrasound documentation software Digisono.

Cost-Benefit Analysis↗

Statistical models for predicting response to interferon-alpha and spontaneous seroconversion in children with chronic hepatitis B.

To develop prognostic models for identifying children with hepatitis B who are likely to respond to interferon-alpha (IFN-alpha) or to spontaneously seroconvert, we evaluated results of a multinational controlled trial comprising 70 children with chronic hepatitis B who received IFN-alpha and 74 children who did not receive therapy. Prognostic models were developed using SMILES (similarity of least squares), which is a data analysis network that incorporates multidimensional relationships in the clinical data of complex diseases. Commonly collected clinical data included age, gender, serum aminotransferase (aspartate aminotransferase [AST] and alanine aminotransferase [ALT]) and hepatitis B virus (HBV) DNA levels, and IFN-alpha dose. Additional data included pretreatment directional information (e.g. increases or decreases in serum aminotransferase and HBV DNA levels), liver biopsy results, race and transmission mode. Using data available prior to initiation of treatment, the SMILES models achieved prospective predictions of 89% for responders, 96% for non-responders, 100% for seroconverters and 93% for non-seroconverters. Although not predictive by themselves, the variables that had the greatest impact on predictions for IFN-alpha response were HBV DNA pretreatment direction, baseline HBV DNA, IFN-alpha dose and gender. The variables that had the greatest impact on predictions for spontaneous seroconversion were ALT pretreatment direction, baseline HBV DNA level, age and AST pretreatment direction. Therefore, these models may be useful in determining, in children with hepatitis B, the likelihood of response to IFN-alpha and spontaneous seroconversion.

Adolescent↗

The MGED Ontology: a resource for semantics-based description of microarray experiments.

MOTIVATION: The generation of large amounts of microarray data and the need to share these data bring challenges for both data management and annotation and highlights the need for standards. MIAME specifies the minimum information needed to describe a microarray experiment and the Microarray Gene Expression Object Model (MAGE-OM) and resulting MAGE-ML provide a mechanism to standardize data representation for data exchange, however a common terminology for data annotation is needed to support these standards. RESULTS: Here we describe the MGED Ontology (MO) developed by the Ontology Working Group of the Microarray Gene Expression Data (MGED) Society. The MO provides terms for annotating all aspects of a microarray experiment from the design of the experiment and array layout, through to the preparation of the biological sample and the protocols used to hybridize the RNA and analyze the data. The MO was developed to provide terms for annotating experiments in line with the MIAME guidelines, i.e. to provide the semantics to describe a microarray experiment according to the concepts specified in MIAME. The MO does not attempt to incorporate terms from existing ontologies, e.g. those that deal with anatomical parts or developmental stages terms, but provides a framework to reference terms in other ontologies and therefore facilitates the use of ontologies in microarray data annotation. AVAILABILITY: The MGED Ontology version.1.2.0 is available as a file in both DAML and OWL formats at http://mged.sourceforge.net/ontologies/index.php. Release notes and annotation examples are provided. The MO is also provided via the NCICB's Enterprise Vocabulary System (http://nciterms.nci.nih.gov/NCIBrowser/Dictionary.do). CONTACT: Stoeckrt@pcbi.upenn.edu SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Computational Biology↗

Should aspirin be continued in patients started on warfarin?

BACKGROUND AND OBJECTIVE: Clinicians frequently face the decision of whether to continue aspirin when starting patients on warfarin. We performed a meta-analysis to characterize the tradeoffs involved in this common clinical dilemma. DATA SOURCES: Multiple computerized databases (1966 to 2003), reference lists of relevant articles, conference proceedings, and queries of primary authors. STUDY SELECTION: Randomized trials comparing warfarin plus aspirin versus warfarin alone. Studies with target international normalized ratios (INRs) <2 were excluded. DATA EXTRACTION: Two reviewers independently extracted baseline data and major outcomes: rates of thromboembolism, hemorrhage, and all-cause mortality. DATA SYNTHESIS: Nine studies met the inclusion criteria. Of the five that enrolled patients with mechanical heart valves, four used the same target INR in both groups, while one used a reduced target INR for the warfarin plus aspirin group. Pooling the results of the first four studies demonstrated that combination of warfarin plus aspirin significantly decreased thromboembolic events (relative risk [RR], 0.33; 95% confidence interval [CI], 0.19 to 0.58), increased major bleeding (RR, 1.58; 95% CI, 1.02 to 2.44), and decreased all-cause mortality (RR, 0.43; 95% CI, 0.23 to 0.81) compared to warfarin alone. The one valve trial using a reduced INR in the warfarin plus aspirin group reported no difference in thromboembolic outcomes but found decreased major bleeding and a significant mortality benefit with combination therapy. Of the remaining trials, three evaluated a warfarin indication not routinely used in the United States (post-myocardial infarction), and the only trial that considered atrial fibrillation was terminated early due to inadequate enrollment. CONCLUSIONS: For mechanical heart valve patients, the benefits of continuing aspirin when starting warfarin therapy are clear. For other routine warfarin indications, there are not adequate data to guide this common clinical decision.

Anticoagulants↗

Organization and representation of patient safety data: current status and issues around generalizability and scalability.

Recent reports have identified medical errors as a significant cause of morbidity and mortality among patients. A variety of approaches have been implemented to identify errors and their causes. These approaches include retrospective reporting and investigation of errors and adverse events and prospective analyses for identifying hazardous situations. The above approaches, along with other sources, contribute to data that are used to analyze patient safety risks. A variety of data structures and terminologies have been created to represent the information contained in these sources of patient safety data. Whereas many representations may be well suited to the particular safety application for which they were developed, such application-specific and often organization-specific representations limit the sharability of patient safety data. The result is that aggregation and comparison of safety data across organizations, practice domains, and applications is difficult at best. A common reference data model and a broadly applicable terminology for patient safety data are needed to aggregate safety data at the regional and national level and conduct large-scale studies of patient safety risks and interventions.

Humans↗

AAPI youth tobacco use: a comparative analysis of current cigarette use data from the Florida, Texas, and National Youth Tobacco Surveys.

OBJECTIVES: The objectives of this study were to compile data on Asian American and Pacific Islander (AAPI) youth tobacco use in Florida and conduct comparisons with state and national data. This research will contribute to reducing the gap in information regarding current smoking prevalence among AAPI youth in Florida and provide direct comparisons with another state (Texas) and National Youth Tobacco Survey (NYTS) data on AAPI youth. METHODS: Current cigarette use data from the Florida Youth Tobacco Surveys conducted in 1998, 1999, and 2000 were examined for trends in AAPI and state prevalence rates. AAPI data from Florida's baseline 1998 youth tobacco survey were compared to Texas data after applying a common set of data preparation edits. AAPI data from the NYTS were also compared to Florida's AAPI youth population. FINDINGS: Current cigarette use for AAPI students in Florida was generally below the overall prevalence rates among Florida's public middle or high school students. In 1998, current smoking prevalence among Texas AAPI middle and high school students was 18.7% compared to 19.4% among Florida students. Among high school students, the NYTS found that 21.2% of AAPI students were current cigarette smokers nationally in comparison to 21.7% of AAPI high school youth in Florida. In middle school, the current smoking prevalence among AAPI students was 5.5% in the NYTS as compared to 9.4% in Florida. The NYTS data in particular highlight the magnitude of the increasing trend of cigarette smoking among AAPI youth as they progress through the high school grades. CONCLUSIONS: Of all the racial/ethnic groups in Florida, only AAPIs did not have a significant decline in current cigarette use. While the Florida Tobacco Pilot Program has implemented many worthwhile initiatives, the anti-tobacco interventions do not appear to have exerted a noticeable effect on AAPI youth.

Adolescent↗

Tools for statistical analysis with missing data: application to a large medical database.

Missing data is a common feature of large data sets in general and medical data sets in particular. Depending on the goal of statistical analysis, various techniques can be used to tackle this problem. Imputation methods consist in substituting the missing values with plausible or predicted values so that the completed data can then be analysed with any chosen data mining procedure. In this work, we study imputation in the context of multivariate data and we evaluate a number of methods which can be used by today's standard statistical software packages. Imputation using multivariate classification, multiple imputation and imputation by factorial analysis are compared using simulated data and a large medical database (from the diabetes field) with numerous missing values. Our main result is to provide a control chart for assessing data quality after the imputation process. To this end, we developed an algorithm for which the input is a set of parameters describing the underlying data (e.g., covariance matrix, distribution) and the output is a chart which plots the change in the prediction error with respect to the proportion of missing values. The chart is built by means of an iterative algorithm involving four steps: (1) a sample of simulated data is drawn by using the input parameters; (2) missing values are randomly generated; (3) an imputation method is used to fill in the missing data and (4) the prediction error is computed. Steps 1 to 4 are repeated in order to estimate the distribution of the prediction error. The control chart was established for the 3 imputation methods studied here, assuming a multivariate normal distribution of data. The use of this tool on a large medical database was then investigated. We show how the control chart can be used to assess the quality of the imputation process in the pre-processing step upstream of data mining procedures.

Algorithms↗

Elemental microanalysis of biological specimens.

Although X-ray microanalysis in the electron microscope is the most common method for microanalysis of biological specimens, other methods of elemental microanalysis (electron energy loss spectroscopy, scanning Auger microanalysis, and proton, ion, and laser microprobe analysis) may provide important complementary information and help overcome some of the limitations of electron probe X-ray microanalysis. Despite differences in physical principles and instrumentation, the various microanalytical methods have much in common with regard to specimen preparation, quantitative analysis, and interpretation of analytical data. A common approach to microanalytical problems in the biological sciences, irrespective of the analytical techniques used, seems therefore indicated.

Electron Probe Microanalysis↗

Predictive model selection for repeated measures random effects models using Bayes factors.

The random effects model fit to repeated measures data is an extremely common model and data structure in current biostatistical practice. Modern data analysis often involves the selection of models within broad classes of prespecified models, but for models beyond the generalized linear model, few model-selection tools have been actively studied. In a Bayesian analysis, Bayes factors are the natural tool to use to explore these classes of models. In this paper, we develop a predictive approach for specifying the priors of a repeated measures random effects model with emphasis on selecting the fixed effects. The advantage of the predictive approach is that a single predictive specification is used to specify priors for all models considered. The methodology is applied to a pediatric pain data analysis.

Bayes Theorem↗

Risk of spontaneous preterm birth is associated with common proinflammatory cytokine polymorphisms.

BACKGROUND: Preliminary data suggest that common genetic variation in immune response genes can contribute to the risk for spontaneous preterm birth and possibly small-for-gestational age (SGA). METHODS: We investigated the relationship of polymorphisms in 6 cytokine genes associated with inflammation-interleukin (IL)1alpha, IL1beta, IL2, IL6, tumor necrosis factor (TNF), and lymphotoxin alpha (LTA)-with spontaneous preterm and SGA birth in a nested case-control study drawn from a prospective pregnancy cohort. Women were recruited between 24 and 29 weeks' gestation at the Wake County and University of North Carolina, Chapel Hill obstetric clinics between February 1996 and June 2000. We inferred haplotypes using the EM algorithm and the Bayesian method, PHASE. We then compared haplotype frequency distributions and implemented semi-Bayesian hierarchical logistic regression analyses to obtain odds ratio (OR) estimates and 95% confidence intervals (CIs) for each polymorphism. RESULTS: Two haplotypes spanning the TNF/LTA genes were associated with increased risk for spontaneous preterm birth in white subjects (for the AGG haplotype, OR = 1.5 [95% CI=0.8-2.6]; for the GAC haplotype, 1.6 [0.9-2.9]). Additionally, carriers of the GAG haplotype were found to have decreased risk of spontaneous preterm birth (0.6; 0.3-1.0). The TNF(-488)A and LTA(IVS1-82)C variants, constituents of the AGG and GAC haplotypes respectively, were also strongly associated with increased risk of spontaneous preterm birth. CONCLUSIONS: Our results suggest that common genetic variants in proinflammatory cytokine genes could influence the risk for spontaneous preterm birth. Selected TNF/LTA haplotypes were associated with spontaneous preterm birth in both African-American and white subjects. Our data do not support an inflammatory etiology for SGA.

Black or African American↗

Model-based clustering and data transformations for gene expression data.

MOTIVATION: Clustering is a useful exploratory technique for the analysis of gene expression data. Many different heuristic clustering algorithms have been proposed in this context. Clustering algorithms based on probability models offer a principled alternative to heuristic algorithms. In particular, model-based clustering assumes that the data is generated by a finite mixture of underlying probability distributions such as multivariate normal distributions. The issues of selecting a 'good' clustering method and determining the 'correct' number of clusters are reduced to model selection problems in the probability framework. Gaussian mixture models have been shown to be a powerful tool for clustering in many applications. RESULTS: We benchmarked the performance of model-based clustering on several synthetic and real gene expression data sets for which external evaluation criteria were available. The model-based approach has superior performance on our synthetic data sets, consistently selecting the correct model and the number of clusters. On real expression data, the model-based approach produced clusters of quality comparable to a leading heuristic clustering algorithm, but with the key advantage of suggesting the number of clusters and an appropriate model. We also explored the validity of the Gaussian mixture assumption on different transformations of real data. We also assessed the degree to which these real gene expression data sets fit multivariate Gaussian distributions both before and after subjecting them to commonly used data transformations. Suitably chosen transformations seem to result in reasonable fits. AVAILABILITY: MCLUST is available at http://www.stat.washington.edu/fraley/mclust. The software for the diagonal model is under development. CONTACT: kayee@cs.washington.edu. SUPPLEMENTARY INFORMATION: http://www.cs.washington.edu/homes/kayee/model.

Algorithms↗