PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

Multivariate data analysis of pollutant profiles: PCB levels across Europe.

It is not always recognised that standard multivariate analyses applied to pollution profile data (i.e. where data are relative amounts of pollutants expressed as proportions of their total) give rise to problems in the analysis and interpretation of results: a simple solution is to carry out analyses on log-ratios of proportions. However, while solving many problems, this approach is very sensitive to the issue of values below detection limits. These approaches have been applied to a dataset of the levels of 29 PCB congeners in ambient air samples across Europe during the summer of 2002. Multivariate descriptive methods (principal component analysis and cluster analysis) and inferential techniques (multivariate ANOVA, multiple linear and logistic regression) and graphical tools (2D and 3D plots, principal components plots, biplots and triangular diagrams) were used to analyse the proportions of five PCB homologues (tri-hepta). These established that there was considerable difference in the pollution profiles of the 71 samples: the greatest variation was between samples with differing ratios of tri-hexa and tri-hepta PCB homologues, and the samples showed little sign of consistent clusters. There was a significant difference between typical profiles from rural and urban areas such that urban samples (and those with high total PCBs) had higher proportions of tetra- and tri-PCBs compared to hexa- and hepta-PCBs.

Analysis of Variance↗

A microarray data analysis framework for postmortem tissues.

This paper will give a complete methodological approach to the processing of oligonucleotide microarray data from postmortem tissue, particularly brain matter. Attention will be drawn to each of the important stages in the process; specifically the quality control, gene expression value calculation, multiple hypothesis testing and correlation analyses. We shall initially discuss the theoretical foundations of each individual method and subsequently apply the ensemble to a sample data set to illustrate and visualise important points.

Algorithms↗

NanoASV: a snakemake workflow for reproducible field-based Nanopore full-length 16S metabarcoding amplicon data analysis.

SUMMARY: NanoASV is a conda environment and snakemake-based workflow using state-of-the-art bioinformatics software to process full-length SSU rRNA (16S/18S) amplicons acquired with Oxford Nanopore Sequencing technology. Its strength lies in reproducibility, portability, and the possibility to run offline, allowing in-field analysis. It can be installed on the Nanopore MK1C sequencing device and process data locally. AVAILABILITY AND IMPLEMENTATION: Source code and documentation are freely available at https://github.com/ImagoXV/NanoASV and Zenodo archive at https://doi.org/10.5281/zenodo.14730742.

Software↗

Microarray data analysis: from disarray to consolidation and consensus.

In just a few years, microarrays have gone from obscurity to being almost ubiquitous in biological research. At the same time, the statistical methodology for microarray analysis has progressed from simple visual assessments of results to a weekly deluge of papers that describe purportedly novel algorithms for analysing changes in gene expression. Although the many procedures that are available might be bewildering to biologists who wish to apply them, statistical geneticists are recognizing commonalities among the different methods. Many are special cases of more general models, and points of consensus are emerging about the general approaches that warrant use and elaboration.

Algorithms↗

Risk factors for failure of Helicobacter pylori therapy--results of an individual data analysis of 2751 patients.

AIM: To study risk factors for failure of Helicobacter pylori eradication treatment. METHODS: Individual data from 2751 patients included in 11 multicentre clinical trials carried out in France and using a triple therapy, were gathered in a unique database. The 27 treatment regimens were regrouped into four categories. RESULTS: The global failure rate was 25.8% [95% CI: 24-27]. There was a difference in failure rate between duodenal ulcer patients and non-ulcer dyspeptic patients, 21.9% and 33.7%, respectively (P < 10(-6)). In a random-effect model, the risk factors identified for eradication failure in duodenal ulcer patients (n = 1400) were: to be a smoker, and to have received the group 4 treatment, while to receive a 10 day treatment vs. 7 days protected from failure. In non-ulcer dyspeptic patients (n = 913), the group 2 treatment was associated with failure. In both groups, age over 60 was associated with successful H. pylori eradication. There were less strains resistant to clarithromycin in duodenal ulcer patients than in non-ulcer dyspeptic patients. Clarithromycin resistance predicted failure almost perfectly. CONCLUSION: Duodenal ulcer and non-ulcer dyspeptic patients should be managed differently in medical practice and considered independently in eradication trials.

Adolescent↗

Increasing the efficiency of fuzzy logic-based gene expression data analysis.

DNA microarray technology can accommodate a multifaceted analysis of the expression of genes in an organism. The wealth of spatiotemporal data generated by this technology allows researchers to potentially reverse engineer a particular genetic network. "Fuzzy logic" has been proposed as a method to analyze the relationships between genes and help decipher a genetic network. This method can identify interacting genes that fit a known "fuzzy" model of gene interaction by testing all combinations of gene expression profiles. This paper introduces improvements made over previous fuzzy gene regulatory models in terms of computation time and robustness to noise. Improvement in computation time is achieved by using a cluster analysis as a preprocessing method to reduce the total number of gene combinations analyzed. This approach speeds up the algorithm by a factor of 50% with minimal effect on the results. The model's sensitivity to noise is reduced by implementing appropriate methods of "fuzzy rule aggregation" and "conjunction" that produce reliable results in the face of minor changes in model input.

Animals↗

Problems in health data analysis: the Maryland permanent pacemaker experience in 1979 and 1980.

In a recent report the Maryland statewide health data base, which is derived from "face sheet" data, was used to determine the appropriateness of permanent pacemaker insertion. In the present study the same indications were utilized and both the complete medical records and the face sheet were reviewed for those patients who had been classified as having permanent pacemakers inserted for inappropriate or questionable reasons. In 32 hospitals, 75% of the records were reviewed (610 of 817 patients). Although coded as having received permanent pacemakers, 16% had received temporary pacemakers, battery change, and the like. Diagnoses justifying permanent pacemaker insertion had been omitted in 53% of the face sheet s, and coding errors were found in 39%. Although none of the 610 medical records reviewed had a valid indication for permanent pacemaker insertion listed on the face sheet, complete medical record review demonstrated valid indications in 95%. Inherent difficulties arise in attempting to list rigid indications for permanent pacemaker insertion. The face sheet does not provide adequate data for assessing the appropriateness of permanent pacemaker insertion.

Arrhythmias, Cardiac↗

Multi-target models and their application to data analysis of cellular mortality due to radiation exposure.

We consider multi-target models for use in analyzing data of the dose-response relationship. The target sizes we are concerned with here are both homogeneous, as assumed in the classical model, and heterogeneous, as simplified using geometric progression. We apply two models for establishing the multi-target models: a Poisson regression model constructed by assuming that the response variable Y follows Poisson distribution, and a gamma-frailty model as a Poisson mixture model derived by adding random common risks having a gamma distribution. Applying these models to experimental data relating the effects of miso fermentation-stages on the survival rate of cells of intestinal crypts of mice exposed to radiation yielded the result that there were substantial frailties associated with all miso fermentation-stages. Short-term and medium-term fermented miso provided similar effects, whereas long-term fermentation had the lowest-relative risk value, indicating a significant protection of the crypts against exposure effects. A gamma-frailty model based on heterogeneous target size was more suitably applied when there were at least 3 dead stem cells having 10 target genes.

Animals↗

Marginalized kernels for RNA sequence data analysis.

We present novel kernels that measure similarity of two RNA sequences, taking account of their secondary structures. Two types of kernels are presented. One is for RNA sequences with known secondary structures, the other for those without known secondary structures. The latter employs stochastic context-free grammar (SCFG) for estimating the secondary structure. We call the latter the marginalized count kernel (MCK). We show computational experiments for MCK using 74 sets of human tRNA sequence data: (i) kernel principal component analysis (PCA) for visualizing tRNA similarities, (ii) supervised classification with support vector machines (SVMs). Both types of experiment show promising results for MCKs.

Computational Biology↗

PEELing: an integrated and user-centric platform for spatially resolved proteomics data analysis.

SUMMARY: Molecular compartmentalization is vital for cellular physiology. Spatially resolved proteomics allows biologists to survey protein composition and dynamics with subcellular resolution. Here, we present PEELing, an integrated package and user-friendly web service for analyzing spatially resolved proteomics data. PEELing assesses data quality using curated or user-defined references, performs cutoff analysis to remove contaminants, connects to databases for functional annotation, and generates data visualizations-providing a streamlined and reproducible workflow to explore spatially resolved proteomics data. AVAILABILITY AND IMPLEMENTATION: PEELing and its tutorial are publicly available at https://peeling.janelia.org/ (Zenodo DOI: 10.5281/zenodo.15692517). A Python package of PEELing is available at https://github.com/JaneliaSciComp/peeling/ (Zenodo DOI: 10.5281/zenodo.15692434).

Proteomics↗

TaxTriage: an open-source metagenomic sequencing data analysis pipeline enabling putative pathogen detection.

MOTIVATION: TaxTriage is a comprehensive pathogen identification workflow designed for both short- and long-read untargeted DNA and RNA sequencing data. Combining read classification, mapping, and de novo assembly approaches, putative pathogens are identified through comparisons to curated pathogens and abundance expectations from healthy cohort data. Flexible installation options are enabled using Nextflow&#x2122; (NF), including cloud deployment via NF Tower (Seqera Platform) and local installation on a variety of systems, including standalone installations without external internet access. Final analysis summaries are compiled into an Organism Discovery Report, which lists likely pathogens and supporting data, including a custom confidence score. RESULTS: Evaluation of published in silico, clinical, and outbreak datasets identified performance comparable to alternative cloud-based processing pipelines for expected pathogen and co-infection detection with similar sensitivity and increased specificity. To support both public health and veterinary diagnostics communities, customization options have been incorporated to enable improved performance for host species of interest. AVAILABILITY AND IMPLEMENTATION: Source code for TaxTriage is freely available at https://github.com/jhuapl-bio/taxtriage. TaxTriage v2.1.1 has been archived on Zenodo at https://zenodo.org/records/17081354 to permit reproducible analysis as described in this manuscript.

Software↗

Blood contacts in the operating room after hospital-specific data analysis and action.

BACKGROUND: There has been much work recently to quantify risk of blood exposures among operating room personnel. Little has been done to show the outcome of preventive strategies. Three hospitals implemented a variety of changes after detailed feedback on blood contact data. This report follows those hospitals to document changes in blood contact rates. METHODS: Each hospital reviewed detailed data on blood exposures and developed a range of strategies to reduce contacts. In a second data collection period, uniformly trained circulating nurses sought information during surgical procedures on blood contacts among staff. Data were collected on all blood contacts and surgeries during which they occurred. These data were then compared with data from the previous study period, before changes in practices. RESULTS: All blood contacts combined and in each hospital decreased significantly in the second data collection period. Percutaneous exposures also consistently decreased, but did not reach statistical significance. The distribution of types of contact changed, with percutaneous exposures representing a larger proportion of contacts seen in the second period. Similar anatomic locations, devices, and characteristics of surgeries were associated with blood contacts in both periods. DISCUSSION: Specific data provided to operating room personnel motivated the development of specific strategies, although the influence of feedback alone versus specific interventions can not be separated. Analysis and generation of hospital-specific data on blood exposures among operating personnel may have a positive influence in lowering the risk of blood exposures in this population group.

Blood-Borne Pathogens↗

Regression splines for threshold selection in survival data analysis.

The Cox proportional hazards model restricts the hazard ratio to be linear in the covariates. A survival model based on data from a clinical trial is developed using spline functions with variable knots to estimate the log hazard function. Moreover, the main point of the method is that a knot, seen as free parameters for a piecewise linear spline, represents a break point in the log hazard function which may be interpreted as a threshold value. The likelihood ratio test is used to select the final model and to determine the threshold number for a covariate. Confidence intervals for these threshold values are computed by bootstrapping the data. Two examples illustrate the method.

Carcinoma, Small Cell↗

Pattern recognition techniques in microarray data analysis: a survey.

Recent development of technologies (e.g., microarray technology) that are capable of producing massive amounts of genetic data has highlighted the need for new pattern recognition techniques that can mine and discover biologically meaningful knowledge in large data sets. Many researchers have begun an endeavor in this direction to devise such data-mining techniques. As such, there is a need for survey articles that periodically review and summarize the work that has been done in the area. This article presents one such survey. The first portion of the paper is meant to provide the basic biology (mostly for non-biologists) that is required in such a project. This part is only meant to be a starting point for those experts in the technical fields who wish to embark on this new area of bioinformatics. The second portion of the paper is a survey of various data-mining techniques that have been used in mining microarray data for biological knowledge and information (such as sequence information). This survey is not meant to be treated as complete in any form, since the area is currently one of the most active, and the body of research is very large. Furthermore, the applications of the techniques mentioned here are not meant to be taken as the most significant applications of the techniques, but simply as examples among many.

Algorithms↗

Propagating uncertainty in microarray data analysis.

Microarray technology is associated with many sources of experimental uncertainty. In this review we discuss a number of approaches for dealing with this uncertainty in the processing of data from microarray experiments. We focus here on the analysis of high-density oligonucleotide arrays, such as the popular Affymetrix GeneChip array, which contain multiple probes for each target. This set of probes can be used to determine an estimate for the target concentration and can also be used to determine the experimental uncertainty associated with this measurement. This measurement uncertainty can then be propagated through the downstream analysis using probabilistic methods. We give examples showing how these credibility intervals can be used to help identify differential expression, to combine information from replicated experiments and to improve the performance of principal component analysis.

Computational Biology↗

Attenuation estimations using envelope echo data: analysis and simulations.

Previously we described a video signal analysis (VSA) method for measuring backscatter and attenuation from B-Mode image data. VSA computes depth-dependent ratios of the mean echo intensity from a sample to the mean echo intensity from a reference phantom imaged using identical scanner settings. The slope of a line-fit of this ratio (expressed in dB) versus depth is related to the attenuation of the sample. This paper investigates conditions for which the echo intensity ratio versus depth is independent of transducer pulsing characteristics and instrument settings, and depends only on the properties of the sample and the reference. A theoretical model is described for the echo signal power versus depth from a uniform medium containing scatterers. The model incorporates bandwidth, frequency and media attenuation. Results show that the sample-to-reference echo intensity ratio versus depth is a curve, the departure of which from a straight line is a function of the relative attenuation of the two media, the imaging system bandwidth and the initial frequency. The model also leads to a depth-dependent "effective frequency" determination in the VSA method. Model predictions are verified using RF signals computed by an acoustic pulse-echo simulation program.

Computer Simulation↗

[The height census of first grade schoolchildren: anthropometric data analysis].

The present study aims at identifying prior areas for nutritional programs considering growth scores of children entering the first grade of school. Data were obtained through the nutritional surveillance system in a Brazilian city called Osasco, São Paulo. The analysis was meant to determine the magnitude and distribution of growth retardation. In order to establish the nutritional status of children the indicator height/age expressed by standard deviation scores (z-score) was used. Values below -2 s.d. of the reference population median (NCHS) were considered height retarded. Children's growth was geographically distributed into groups of schools according to height deficit prevalence. Results showed marked differences among the schools with deficit prevalence varying between 0 and 16.1%. This led to the identification of communities with good health and nutrition levels, and communities exposed to different levels of malnutrition. The level of dissociation was such that it was possible to pinpoint areas and micro-areas where social programs and investments are most needed.

English Abstract↗

Local dimensionality reduction and supervised learning within natural clusters for biomedical data analysis.

Inductive learning systems were successfully applied in a number of medical domains. Nevertheless, the effective use of these systems often requires data preprocessing before applying a learning algorithm. This is especially important for multidimensional heterogeneous data presented by a large number of features of different types. Dimensionality reduction (DR) is one commonly applied approach. The goal of this paper is to study the impact of natural clustering--clustering according to expert domain knowledge--on DR for supervised learning (SL) in the area of antibiotic resistance. We compare several data-mining strategies that apply DR by means of feature extraction or feature selection with subsequent SL on microbiological data. The results of our study show that local DR within natural clusters may result in better representation for SL in comparison with the global DR on the whole data.

Algorithms↗