PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

A hierarchical, count-based model highlights challenges in scATAC-seq data analysis and points to opportunities to extract finer-resolution information.

BACKGROUND: Data from Single-cell Assay for Transposase Accessible Chromatin with Sequencing (scATAC-seq) is highly sparse. While current computational methods feature a range of transformation procedures to extract meaningful information, major challenges remain. RESULTS: Here, we discuss the major scATAC-seq data analysis challenges such as sequencing depth normalization and region-specific biases. We present a hierarchical count model that is motivated by the data generating process of scATAC-seq data. Our simulations show that current scATAC-seq data, while clearly containing physical single-cell resolution, are too sparse to infer true informational-level single-cell, single-region of chromatin accessibility states. CONCLUSIONS: While the broad utility of scATAC-seq at a cell type level is undeniable, describing it as fully resolving chromatin accessibility at single-cell resolution, particularly at individual locus level, may overstate the level of detail currently achievable. We conclude that chromatin accessibility profiling at true single-cell, single-region resolution is challenging with current data sensitivity, but that it may be achieved with promising developments in optimizing the efficiency of scATAC-seq assays.

Single-Cell Analysis↗

[Computer-assisted data analysis in a pediatric intensive care unit].

Computer assisted real time data analysis introduces a reasonable method of judgment into patient monitoring systems. From fast changing vital parameters discrete heart and respiration rate samples are immediately evaluated and presented as graphs near the bedside. Thus, statistical routines can increase the better understanding of instable clinical conditions and lend support to the decision making process. The early detection of a pathological trend in a patient whose ability to compensate is still present provides necessary time for diagnostic or preventive countermeasures in case of emergency.

Computers↗

Optical illusions from visual data analysis: example of the New Zealand asthma mortality epidemic.

The abundance of health-related statistics routinely collected worldwide invites their misuse from haphazard associations between secular trends of these data. This misuse is often compounded by assessing these associations simply on the basis of a visual inspection of the data. The visual approach to data analysis, known to have several pitfalls, is particularly tempting in the context of asthma where it has often been used. For example, the epidemic of asthma deaths that occurred in New Zealand during the last two decades has been imputed to fenoterol, a medication for asthma, on the basis of a visual assessment of ecological data. The simultaneity of time trends in the asthma death rate and fenoterol market share in that country formed an important part of the statistical basis of the evidence. We verified whether the results of such visual analyses are corroborated by more objective quantitative statistical methods of analysis. We reanalyzed these same data, namely the time trend data of New Zealand asthma death rates, fenoterol market share, sales of beta-agonists and inhaled corticosteroids, measured yearly for the 16-year span 1976-1991, using Poisson weighted loglinear regression. We found that the protective effect of inhaled corticosteroids (rate ratio 0.5 per canister per month; 95% confidence interval 0.4 to 0.7; p = 0.0001) was more closely associated with changes in asthma mortality than either fenoterol (RR 2.7 per canister per month; 95% CI: 0.9 to 7.5; p = 0.06) or all beta-agonists combined (RR 1.6; 95% CI: 0.8 to 3.0; p = .19). We conclude from this quantitative analysis that these ecological asthma mortality data provide evidence of a stronger association with inhaled corticosteroids, little used in New Zealand at the onset of the epidemic but used abundantly at its termination, than with fenoterol. This conclusion is diametrically opposite to that found by the visual approach. The quantitative analysis demonstrates that the visual approach to the analysis of ecological data, although seemingly convincing, can be misleading by creating an optical illusion. This purely visual approach to data analysis may thus have serious implications when the resulting scientific information is used to make vital public health and policy decisions.

Administration, Inhalation↗

RealSpot: software validating results from DNA microarray data analysis with spot images.

The spot images from DNA microarray highly affect the discovery of biological knowledge from gene expression data. However, results from quality analysis, normalization, differential expression, and cluster analysis are rarely validated with spot images in current data analysis methods or software packages. We designed RealSpot, a software package, to validate the results by directly associating spot quality and data with spot images in a spreadsheet table. RealSpot splits hybridization images into individual spots stored in a spreadsheet table. It subsequently associates microarray data with spot images and performs data validation through the standard table operation such as sorting, searching, and editing. RealSpot has several built-in functions to facilitate data validation, including spot quality analysis, data organization, one-way ANOVA, gene ontology association, verification, import, and export. We used RealSpot to evaluate 77 slides (30,000 features each) from real hybridization experiments and to validate results from each step of data analysis. It took approximately 10 min to validate results of spot quality after initial evaluation and correct approximately 0.3% of falsely assigned qualities of 10,000 spots. We validated 1,641 of 2,110 differentially expressed genes identified by SAM analysis in approximately 1/2 h by comparing each gene with its respective spot image. Furthermore, we found that 6 of 48 genes in one cluster from k-mean clustering method showed inconsistent trends of spot images. RealSpot is efficient for validating microarray results and thus helpful for improving the reliability of the whole microarray experiment for experimentalists.

Algorithms↗

Coronary plaque composition of nonculprit lesions, assessed by in vivo intracoronary ultrasound radio frequency data analysis, is related to clinical presentation.

BACKGROUND: Identification of subclinical high-risk plaques is potentially important because they may have greater likelihood of rupture and subsequent thrombosis. The purpose of this study was to assess the relationship between plaque composition determined by intravascular ultrasound (IVUS) radio frequency (RF) data analysis and clinical presentation. METHODS: In 55 patients, a nonculprit vessel with < 50% diameter stenosis was studied with IVUS. Tissue maps were reconstructed from RF data using IVUS-Virtual Histology software. RESULTS: Mean percentage of the different plaque components were 0.99% +/- 0.9%, calcium; 68.04% +/- 9.8%, fibrous; 19.31% +/- 7.3%, fibrolipidic; and 9.43% +/- 6.6%, lipid core. Mean lipid core percentage was significantly larger in patients with acute coronary syndrome (ACS) when compared with stable patients (12.26% +/- 7.0% vs 7.40% +/- 5.5%, P = .006). In addition, stable patients showed more fibrotic vessels (70.97% +/- 9.3% vs 63.96% +/- 9.1%, P = .007). There was no significant difference for either mean calcium (1.20% +/- 1.1% vs 0.83% +/- 0.7%, P = .124) or fibrolipidic (20.57% +/- 6.9% vs 18.40% +/- 7.6%, P = .281) percentages in ACS and stable patients, respectively. Vessel area obstruction did not differ between groups (46.49% +/- 10.9% vs 42.83% +/- 11.8%, P = .221). There was a significant, albeit weak, positive correlation between lipid core percentage and stenosis severity as determined by vessel area obstruction (r = 0.34, P = .015). CONCLUSIONS: In this study, plaque characterization of nonculprit vessels using spectral analysis of IVUS RF data analysis was significantly related to clinical presentation. Percentage of lipid core, a feature related to acute coronary events and worse prognosis, was significantly larger in patients with ACS. Conversely, stable patients showed more fibrotic content.

Aged↗

Simpleaffy: a BioConductor package for Affymetrix Quality Control and data analysis.

UNLABELLED: Quality Control is a fundamental aspect of successful microarray data analysis. Simpleaffy is a BioConductor package that provides access to a variety of QC metrics for assessing the quality of RNA samples and of the intermediate stages of sample preparation and hybridization. Simpleaffy also offers fast implementations of popular algorithms for generating expression summaries and detection calls. AVAILABILITY: Simpleaffy can be downloaded from http://www.bioconductor.org. SUPPLEMENTARY INFORMATION: Additional information can be found on the supplementary website located at http://bioinformatics.picr.man.ac.uk.

Computational Biology↗

The time-rescaling theorem and its application to neural spike train data analysis.

Measuring agreement between a statistical model and a spike train data series, that is, evaluating goodness of fit, is crucial for establishing the model's validity prior to using it to make inferences about a particular neural system. Assessing goodness-of-fit is a challenging problem for point process neural spike train models, especially for histogram-based models such as perstimulus time histograms (PSTH) and rate functions estimated by spike train smoothing. The time-rescaling theorem is a well-known result in probability theory, which states that any point process with an integrable conditional intensity function may be transformed into a Poisson process with unit rate. We describe how the theorem may be used to develop goodness-of-fit tests for both parametric and histogram-based point process models of neural spike trains. We apply these tests in two examples: a comparison of PSTH, inhomogeneous Poisson, and inhomogeneous Markov interval models of neural spike trains from the supplementary eye field of a macque monkey and a comparison of temporal and spatial smoothers, inhomogeneous Poisson, inhomogeneous gamma, and inhomogeneous inverse gaussian models of rat hippocampal place cell spiking activity. To help make the logic behind the time-rescaling theorem more accessible to researchers in neuroscience, we present a proof using only elementary probability theory arguments. We also show how the theorem may be used to simulate a general point process model of a spike train. Our paradigm makes it possible to compare parametric and histogram-based neural spike train models directly. These results suggest that the time-rescaling theorem can be a valuable tool for neural spike train data analysis.

Action Potentials↗

Who 'controls' quality control data analysis?

A common quality control tool is peer group comparison of data from commercial controls. While its real-time effectiveness is limited, inappropriate statistical management of the data can cause an individual lab's performance to be misrepresented. Here are two examples where vendor-directed data analysis contained flagrant errors. The finding that vendors use inappropriate algorithms to compare accuracy and precision of peer performance suggests a need, the author believes, to set rigorous standards of reporting for the protection of participating laboratories.

California↗

Data analysis now and then: significant changes in approaches and results.

Modern data analysis is one of the many prerequisites for telemedical applications. Classical statistical methods alone are no longer sufficient to fulfill the various demands of modern analytical procedures. Cluster and association analysis among others have filled this gap and are capable of producing more adequate and better suitable results as well as to provide information not detectable in the past.

Algorithms↗

An evaluation of five commercial immunoassay data analysis software systems.

An evaluation of five commercial software systems used for immunoassay data analysis revealed numerous deficiencies. Often, the utility of statistical output was compromised by poor documentation. Several data sets were run through each system using a four-parameter calibration function, and the results were compared to those from an independent method. Comparable results between systems were obtained, but often several attempts at analysis were necessary. The evaluation process revealed that it is difficult to monitor the numerous options available on these types of programs, and that incorrect results could easily be obtained if comparison analyses were not used. Recommendations for improved software functionality and for using the four-parameter calibration model are presented.

Data Interpretation, Statistical↗

Comparison of fMRI data analysis by SPM99 on different operating systems.

The hardware chosen for fMRI data analysis may depend on the platform already present in the laboratory or the supporting software. In this study, we ran SPM99 software on multiple platforms to examine whether we could analyze fMRI data by SPM99, and to compare their differences and limitations in processing fMRI data, which can be attributed to hardware capabilities. Six normal right-handed volunteers participated in a study of hand-grasping to obtain fMRI data. Each subject performed a run that consisted of 98 images. The run was measured using a gradient echo-type echo planar imaging sequence on a 1.5T apparatus with a head coil. We used several personal computer (PC), Unix and Linux machines to analyze the fMRI data. There were no differences in the results obtained on several PC, Unix and Linux machines. The only limitations in processing large amounts of the fMRI data were found using PC machines. This suggests that the results obtained with different machines were not affected by differences in hardware components, such as the CPU, memory and hard drive. Rather, it is likely that the limitations in analyzing a huge amount of the fMRI data were due to differences in the operating system (OS).

Adult↗

Functional data analysis in longitudinal settings using smoothing splines.

Data in many experiments arise as curves and therefore it is natural to use a curve as a basic unit in the analysis, which is termed functional data analysis (FDA). In longitudinal studies, recent developments in FDA have extended classical linear models and linear mixed effects models to functional linear models (also termed varying-coefficient models) and functional mixed effects models. In this paper we focus our review on the functional mixed effects models using smoothing splines, because functional linear models are special cases of this more general framework. Due to the connection between smoothing splines and linear mixed effects models, functional mixed effects models can be fitted using existing software such as SAS Proc Mixed. A case study is presented as an illustration.

Biomedical Research↗

Exploratory data analysis of hyperlipidemia on the Macintosh: software tools for analysis of biochemical, clinical, and genetic variables in 1677 consecutive lipid clinic patients.

Exploratory data analysis (EDA) software facilitates unstructured, iterative open exploration of complex datasets with the aid of multiple linked graphical displays. We are investigating relationships between plasma lipoproteins and coronary artery disease by retrospective analysis of 1677 consecutive UCSF Lipid Clinic patients. Our preliminary experience is with Data Deck 3.0 although several additional software programs (JMP 2.0, Systat 5.1, Minitab 8.0, StatView 4.0) are mentioned. Lipid diagnosis (751 women and 925 men) was 22% primary hypercholesterolemia, 19% combined hyperlipidemia, 3% dysbetalipoproteinemia, 15% endogenous lipemia, 4% mixed lipemia, 5% elevated Lp(a) and 32% with no major lipid abnormality. We found the Macintosh platform (68030) to be flexible and powerful for analysis of moderate size (less than 1 Mb) clinical datasets. High resolution color monitors (1024 x 768 pixels), fast hard disks (< 18 msec) and moderate amounts of system memory (8 + Mb) facilitate exploratory analysis.

Artificial Intelligence↗

Longitudinal data analysis in pedigree studies.

Longitudinal family studies provide a valuable resource for investigating genetic and environmental factors that influence long-term averages and changes over time in a complex trait. This paper summarizes 13 contributions to Genetic Analysis Workshop 13, which include a wide range of methods for genetic analysis of longitudinal data in families. The methods can be grouped into two basic approaches: 1) two-step modeling, in which repeated observations are first reduced to one summary statistic per subject (e.g., a mean or slope), after which this statistic is used in a standard genetic analysis, or 2) joint modeling, in which genetic and longitudinal model parameters are estimated simultaneously in a single analysis. In applications to Framingham Heart Study data, contributors collectively reported evidence for genes that affected trait mean on chromosomes 1, 2, 3, 5, 8, 9, 10, 13, and 17, but most did not find genes affecting slope. Applications to simulated data suggested that even for a gene that only affected slope, use of a mean-type statistic could provide greater power than a slope-type statistic for detecting that gene. We report on the results of a small experiment that sheds some light on this apparently paradoxical finding, and indicate how one might form a more powerful test for finding a slope-affecting gene. Several areas for future research are discussed.

Cardiovascular Diseases↗

WinXAS: a program for X-ray absorption spectroscopy data analysis under MS-Windows.

WinXAS is a new X-ray absorption spectroscopy (XAS) data-analysis program. It runs under the operating system MS-Windows 95/NT and offers several unique features. It has a user-friendly graphical environment and is capable of reading a variety of data formats. It contains a number of useful numerical algorithms beyond those used in conventional XAS analysis and offers a simple interface to the ab-initio theoretical code FEFF. The availability of fast macros in WinXAS makes it particularly useful for on-line data examination at synchrotron radiation facilities during XAS experiments as well as for the analysis of multiple-scan data such as those from time-resolved experiments.

Journal Article↗

Histopathological criteria for progressive dementia disorders: clinical-pathological correlation and classification by multivariate data analysis.

Autopsied brains from 55 patients with dementia between 59-95 years of age (mean age 77.9 +/- 8.1 years) and 19 non-demented individuals between 46-91 years of age (mean age 74.3 +/- 10.5 years) were examined to establish histopathological criteria for normal ageing, primary degenerative [Alzheimer's disease (AD)/senile dementia of Alzheimer type (SDAT)] and vascular (multi-infarct) dementia (MID) disorders. Senile/neuritic plaques, neurofibrillary tangles, microscopic infarcts and perivascular serum protein deposits were quantified in the frontal lobe (Brodmann area 10) and in the hippocampus. The demented patients were classified according to the DSM-III criteria into AD/SDAT and MID. Operationally defined histopathological criteria for dementias, based on the degree/amount of the histopathological changes seen in aged non-demented patients, were postulated. The demented patients were clearly separable into three histopathological types, namely AD/SDAT, MID and AD-MID, the dementia type where both the degenerative and the vascular changes are coexistent in greater extent than are seen in the non-demented individuals. Using general clinical, gross neuroanatomical and histopathological data three separate dementia classes, namely AD/SDAT, MID and AD-MID, were visualized in two-dimensional space by multivariate data analysis. This analysis revealed that the pathology in the AD-MID patients was not merely a linear combination of the pathology in AD/SDAT and MID, indicating that AD-MID might represent a dementia type of its own. The clinical diagnosis for AD/SDAT and MID was certain in only half of the AD/SDAT and one third of the MID cases when evaluated histopathologically and by multivariate data analysis. AD/SDAT, MID and AD-MID were histopathologically diagnosed in 49%, 24% and 27%, respectively, of all the dementia cases studied. Opposite correlation between the number of tangles, plaques and the patient age in non-demented and AD/SDAT cases were observed, indicating that the pathogenesis of tangles and plaques in the two groups of patients might be different and that AD/SDAT might not be a form of an exaggerated ageing process.

Aged↗

VANTED: a system for advanced data analysis and visualization in the context of biological networks.

BACKGROUND: Recent advances with high-throughput methods in life-science research have increased the need for automatized data analysis and visual exploration techniques. Sophisticated bioinformatics tools are essential to deduct biologically meaningful interpretations from the large amount of experimental data, and help to understand biological processes. RESULTS: We present VANTED, a tool for the visualization and analysis of networks with related experimental data. Data from large-scale biochemical experiments is uploaded into the software via a Microsoft Excel-based form. Then it can be mapped on a network that is either drawn with the tool itself, downloaded from the KEGG Pathway database, or imported using standard network exchange formats. Transcript, enzyme, and metabolite data can be presented in the context of their underlying networks, e. g. metabolic pathways or classification hierarchies. Visualization and navigation methods support the visual exploration of the data-enriched networks. Statistical methods allow analysis and comparison of multiple data sets such as different developmental stages or genetically different lines. Correlation networks can be automatically generated from the data and substances can be clustered according to similar behavior over time. As examples, metabolite profiling and enzyme activity data sets have been visualized in different metabolic maps, correlation networks have been generated and similar time patterns detected. Some relationships between different metabolites were discovered which are in close accordance with the literature. CONCLUSION: VANTED greatly helps researchers in the analysis and interpretation of biochemical data, and thus is a useful tool for modern biological research. VANTED as a Java Web Start Application including a user guide and example data sets is available free of charge at http://vanted.ipk-gatersleben.de.

Algorithms↗

New data analysis and mining approaches identify unique proteome and transcriptome markers of susceptibility to autoimmune diabetes.

Non-obese diabetic (NOD) mice spontaneously develop autoimmunity to the insulin producing beta cells leading to insulin-dependent diabetes. In this study we developed and used new data analysis and mining approaches on combined proteome and transcriptome (molecular phenotype) data to define pathways affected by abnormalities in peripheral leukocytes of young NOD female mice. Cells were collected before mice show signs of autoimmunity (age, 2-4 weeks). We extracted both protein and RNA from NOD and C57BL/6 control mice to conduct both proteome analysis by two-dimensional gel electrophoresis and transcriptome analysis on Affymetrix expression arrays. We developed a new approach to analyze the two-dimensional gel proteome data that included two-way analysis of variance, cluster analysis, and principal component analysis. Lists of differentially expressed proteins and transcripts were subjected to pathway analysis using a commercial service. From the list of 24 proteins differentially expressed between strains we identified two highly significant and interconnected networks centered around oncogenes (Myc and Mycn) and apoptosis-related genes (Bcl2 and Casp3). The 273 genes with significant strain differences in RNA expression levels created six interconnected networks with a significant over-representation of genes related to cancer, cell cycle, and cell death. They contained many of the same genes found in the proteome networks (including Myc and Mycn). The combination of the eight, highly significant networks created one large network of 272 genes of which 82 had differential expression between strains either at the protein or the RNA level. We conclude that new proteome data analysis strategies and combined information from proteome and transcriptome can enhance the insights gained from either type of data alone. The overall systems biology of prediabetic NOD mice points toward abnormalities in regulation of the opposing processes of cell renewal and cell death even before there are any clear signatures of immune system activation.

Analysis of Variance↗