PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Ambiguous results in functional neuroimaging data analysis due to covariate correlation.

In this note we draw attention to a source of potential ambiguity in functional neuroimaging results when data analysis is based on the resolution of a linear model. This ambiguity arises whenever there exists correlation between the model covariates. A single-subject PET activation experiment helps to illustrate to what extent correlation can affect statistical results interpretation, possibly leading to misinterpretation of part of the activation pattern. This note is intended to clarify this point and to suggest the use of a simple and well-known procedure to deal with these situations. In the Appendix, we suggest a convenient mathematical formulation for statistical tests particularly useful in such cases.

Brain↗

Linear models for microarray data analysis: hidden similarities and differences.

In the past several years many linear models have been proposed for analyzing two-color microarray data. As presented in the literature, many of these models appear dramatically different. However, many of these models are reformulations of the same basic approach to analyzing microarray data. This paper demonstrates the equivalence of some of these models. Attention is directed at choices in microarray data analysis that have a larger impact on the results than the choice of linear model.

Analysis of Variance↗

The effects of overhead transparency design on retention, recall, and application of data analysis content.

This experimental study tested the effects of overhead transparency design in conjunction with live lecture on retention, recall, and application of data analysis content over three occasions using a Solomon Four-Group, pretest-posttest design. Pretested subjects showed significant (p less than .001) gains in test scores from pre to posttest. No significant differences were found in pretest scores between control and experimental treatment groups or among the posttest scores of either the experimental or control groups.

Audiovisual Aids↗

EM-REML estimation of covariance parameters in Gaussian mixed models for longitudinal data analysis.

This paper presents procedures for implementing the EM algorithm to compute REML estimates of variance covariance components in Gaussian mixed models for longitudinal data analysis. The class of models considered includes random coefficient factors, stationary time processes and measurement errors. The EM algorithm allows separation of the computations pertaining to parameters involved in the random coefficient factors from those pertaining to the time processes and errors. The procedures are illustrated with Pothoff and Roy's data example on growth measurements taken on 11 girls and 16 boys at four ages. Several variants and extensions are discussed.

Journal Article↗

Experimental layout, data analysis, and thresholds in ELISA testing of maize for aphid-borne viruses.

Several aspects of enzyme-linked immunosorbent assay (ELISA) procedures and data analysis have been examined in an attempt to find a rapid and reliable method for discriminating between 'positive' and 'negative' results when testing a large number of samples. A layout of ELISA plates was designed to reduce uncontrolled variation and to optimize the number of negative and positive controls. A transformation using the fourth root (A(1/4)) of the optical density readings corrected for the blank (A) stabilized the variance of most ELISA data examined. Transformed A values were used to calculate the true limits, at a set protection level, for false positive (C) and false negative (D). Methods are discussed to reduce the number of undifferentiated samples, i.e. the samples with response falling between C and D. The whole procedure was set up for use with an electronic spreadsheet. With the addition of few instructions of the type 'if em leader then em leader else' in the spreadsheet, the ELISA results were obtained in the simple trichotomous form 'negative/undefined/positive'. This allowed rapid analysis of more than 1100 maize samples testing for the presence of seven aphid-borne viruses-in fact almost 8000 ELISA samples.

Animals↗

Non-homogeneous Markov processes for biomedical data analysis.

Some time ago, the Markov processes were introduced in biomedical sciences in order to study disease history events. Homogeneous and Non-homogeneous Markov processes are an important field of research into stochastic processes, especially when exact transition times are unknown and interval-censored observations are present in the analysis. Non-homogeneous Markov process should be used when the homogeneous assumption is too strong. However these sorts of models increase the complexity of the analysis and standard software is limited. In this paper, some methods for fitting non-homogeneous Markov models are reviewed and an algorithm is proposed for biomedical data analysis. The method has been applied to analyse breast cancer data. Specific software for this purpose has been implemented.

Algorithms↗

An artificial immune system for data analysis.

We present a simplified view of those parts of the human immune system which can be used to provide the basis for a data analysis tool. The motivation for and reasoning behind such a model is given and the desire for a 'transparent' model and meaningful visualization and interpretation techniques is noted. A minimalist formulation of an artificial immune system and some of its behaviour is described. A simple implementation and a suitable visualization technique are demonstrated using some trivial data and the famous 'iris' data set.

Antibody Formation↗

The application of new software tools to quantitative protein profiling via isotope-coded affinity tag (ICAT) and tandem mass spectrometry: II. Evaluation of tandem mass spectrometry methodologies for large-scale protein analysis, and the application of statistical tools for data analysis and interpretation.

Proteomic approaches to biological research that will prove the most useful and productive require robust, sensitive, and reproducible technologies for both the qualitative and quantitative analysis of complex protein mixtures. Here we applied the isotope-coded affinity tag (ICAT) approach to quantitative protein profiling, in this case proteins that copurified with lipid raft plasma membrane domains isolated from control and stimulated Jurkat human T cells. With the ICAT approach, cysteine residues of the two related protein isolates were covalently labeled with isotopically normal and heavy versions of the same reagent, respectively. Following proteolytic cleavage of combined labeled proteins, peptides were fractionated by multidimensional chromatography and subsequently analyzed via automated tandem mass spectrometry. Individual tandem mass spectrometry spectra were searched against a human sequence database, and a variety of recently developed, publicly available software applications were used to sort, filter, analyze, and compare the results of two repetitions of the same experiment. In particular, robust statistical modeling algorithms were used to assign measures of confidence to both peptide sequences and the proteins from which they were likely derived, identified via the database searches. We show that by applying such statistical tools to the identification of T cell lipid raft-associated proteins, we were able to estimate the accuracy of peptide and protein identifications made. These tools also allow for determination of the false positive rate as a function of user-defined data filtering parameters, thus giving the user significant control over and information about the final output of large-scale proteomic experiments. With the ability to assign probabilities to all identifications, the need for manual verification of results is substantially reduced, thus making the rapid evaluation of large proteomic datasets possible. Finally, by repeating the experiment, information relating to the general reproducibility and validity of this approach to large-scale proteomic analyses was also obtained.

Amino Acid Sequence↗

Variation in organochlorine bioaccumulation by a predatory fish; gender, geography, and data analysis methods.

SigmaPCB and p,p'-DDE levels within and among walleye (Stizostedion vitreum) populations were examined to determine how the method of data analysis could influence the interpretation of (i) gender differences and (ii) geographic variation. In the lower Great Lakes (Huron, Erie, and Ontario) whole-body burdens of both contaminants tended to increase with body mass at a faster rate in males than in females. Thus, males generally had higher burdens than females at large body sizes but not at small body sizes. This result was not strongly influenced by the method of expressing contaminant level (burden, wet mass concentration, or lipid mass concentration) but was influenced by the choice of covariate (body mass, body length, or age) in some cases. Mean sigmaPCB and p,p'-DDE concentrations of walleye muscle declined along a gradient from the lower Great Lakes to the Northwest Territories. Analyses using means adjusted for age yielded a stronger contrast between Great Lakes and non-Great Lakes populations than analyses using means adjusted for body length. The gender composition of fish samples and the type and level of covariate used in statistical analyses should be considered in studies of spatiotemporal variation in organochlorine bioaccumulation in fish.

Animals↗

Non-linear mapping for exploratory data analysis in functional genomics.

BACKGROUND: Several supervised and unsupervised learning tools are available to classify functional genomics data. However, relatively less attention has been given to exploratory, visualisation-driven approaches. Such approaches should satisfy the following factors: Support for intuitive cluster visualisation, user-friendly and robust application, computational efficiency and generation of biologically meaningful outcomes. This research assesses a relaxation method for non-linear mapping that addresses these concerns. Its applications to gene expression and protein-protein interaction data analyses are investigated. RESULTS: Publicly available expression data originating from leukaemia, round blue-cell tumours and Parkinson disease studies were analysed. The method distinguished relevant clusters and critical analysis areas. The system does not require assumptions about the inherent class structure of the data, its mapping process is controlled by only one parameter and the resulting transformations offer intuitive, meaningful visual displays. Comparisons with traditional mapping models are presented. As a way of promoting potential, alternative applications of the methodology presented, an example of exploratory data analysis of interactome networks is illustrated. Data from the C. elegans interactome were analysed. Results suggest that this method might represent an effective solution for detecting key network hubs and for clustering biologically meaningful groups of proteins. CONCLUSION: A relaxation method for non-linear mapping provided the basis for visualisation-driven analyses using different types of data. This study indicates that such a system may represent a user-friendly and robust approach to exploratory data analysis. It may allow users to gain better insights into the underlying data structure, detect potential outliers and assess assumptions about the cluster composition of the data.

Algorithms↗

Dispersion models and longitudinal data analysis.

Dispersion models provide a flexible class of non-normal distributions with many potential applications in biostatistics, accommodating a wide range of continuous, discrete and mixed data. Starting with Liang and Zeger's generalized estimating equation method, we review some recent applications of dispersion models in longitudinal data analysis, including state space models based on the Tweedie class of exponential dispersion models. In medical applications the latent process of a state space model may often be interpreted as an unobserved potential morbidity process, which is modelled as a function of time varying covariates. By allowing a multivariate response vector of 'symptoms', the model integrates several response variables of mixed types into a single model. For growth curve models, the latent process reflects the 'true' growth.

Air Pollutants↗

Multivariate data analysis in empirical research. A look on the bright side.

The interpretive benefits of employing multivariate analysis methods on experimental data with more than one dependent variable are described heuristically and illustrated on a set of data from a simply designed experiment in physiological psychology. Multivariate analysis of variance (MANOVA) is performed on the 9 dependent variables contained in the sample data and on the four composites derived from a principal components analysis (PCA) of the variability of the nine. A linear discriminant analysis (LDA) is conducted following both MANOVA results, and 5 methods of determining the "important" dependent variables in the experimental-control group difference are presented and discussed in terms of the data at hand.

Analysis of Variance↗

Detecting changes in natural resources using Land Condition Trend Analysis data.

The Land Condition Trend Analysis (LCTA) program is the US Army's standard for land inventory and monitoring, employing standardized methods of natural resources data collection, analyses, and reporting designed to meet multiple goals and objectives. Critical to using LCTA data in natural resources management decisions is the ability of the LCTA protocols to detect changes in natural resources. To quantify the ability of LCTA protocols to detect resource changes, power analysis techniques were used to estimate minimum detectable effect sizes (MDES) for selected primary and secondary management variables for three Army installations. MDES for a subset of primary variables were estimated using data from 27 installation LCTA programs. MDES for primary and secondary variables varied widely. However, LCTA programs implemented at larger installations with lower sampling intensities detected changes in installation resources as well as programs implemented at smaller more intensively sampled installations. As a national monitoring program that is implemented at individual installations, LCTA protocols provide relatively consistent monitoring data to detect changes in resources despite diverse resource characteristics and implementation constraints.

Conservation of Natural Resources↗

The importance of gel properties for mucoadhesion measurements: a multivariate data analysis approach.

In this study we used tensile strength measurements and a recently developed interpretation procedure to evaluate the mucoadhesive properties of a large set of gel preparations with diverse rheological properties. Multivariate data analysis in the form of principal component analysis (PCA) and partial least square projection to latent structures (PLS) was applied to extract useful information from the rather large quantities of data obtained. PCA showed that the selected series of gels was heterogeneous. Some groupings could be detected but none of the gels was identified as an outlier. By using PLS we investigated the relations between the rheological properties of a gel and the parameters defining the cohesiveness, as measured with the texture analyser used for the mucoadhesion measurements. The rheological properties proved to be important for the results of both the mucoadhesion and the cohesiveness measurements. Furthermore, by using PLS two different measurement configurations were evaluated and it was concluded that the combination of a relatively small volume of gel and two pieces of mucosa seems to be more appropriate than a large volume of gel in combination with one piece of mucosa.

Adhesiveness↗

MGraph: graphical models for microarray data analysis.

UNLABELLED: This paper introduces a MATLAB toolbox, MGraph, which applies graphical models as a natural environment to formulate and solve problems in microarray data analysis. MGraph with its graphical interface allows the user to predict genetic regulatory networks by a graphical gaussian model (GGM), and to quantify the effects of different experimental treatment conditions on gene expression profiles by a graphical log-linear model (GLM). The power of graphical models was explored and illustrated through two example applications. First, four MAPK pathways in yeast were meaningfully reconstructed through GGM. Second, GLM was used to quantify the contributions of sex, genotype and age to transcriptional variance in Drosophila melanogaster. This application may provide a valuable aid in the prediction of genetic regulatory networks, as well as in investigations of various experimental conditions that affect global gene expression profiles. AVAILABILITY: The MATLAB program MGraph is freely available at http://www.uio.no/~junbaiw/mgraph/mgraph.html for academics.

Algorithms↗

HLA-D typing with lymphoblastoid cell lines. VII. A computer program for data analysis.

When lymphoblastoid cell lines (LCL) are substituted for peripheral blood lymphocytes from human typing cell donors in HLA-D typing experiments, a data analysis program must be designed to distinguish the effect of allo-reactivity from those peculiar to LCL, mainly the "autologous-stimulation" effect. The computer program described in this report was created specifically for such an analysis. The rationale for the design of this program is presented in the preceding report (see this issue).

Cell Line↗

Diffusion-tensor MRI: theory, experimental design and data analysis - a technical review.

This article treats the theoretical underpinnings of diffusion-tensor magnetic resonance imaging (DT-MRI), as well as experimental design and data analysis issues. We review the mathematical model underlying DT-MRI, discuss the quantitative parameters that are derived from the measured effective diffusion tensor, and describe artifacts that arise in typical DT-MRI acquisitions. We also discuss difficulties in identifying appropriate models to describe water diffusion in heterogeneous tissues, as well as in interpreting experimental data obtained in such issues. Finally, we describe new statistical methods that have been developed to analyse DT-MRI data, and their potential uses in clinical and multi-site studies.

Anisotropy↗

Quantifying the pathology of neurodegenerative disorders: quantitative measurements, sampling strategies and data analysis.

The use of quantitative methods has become increasingly important in the study of neurodegenerative disease. Disorders such as Alzheimer's disease (AD) are characterized by the formation of discrete, microscopic, pathological lesions which play an important role in pathological diagnosis. This article reviews the advantages and limitations of the different methods of quantifying the abundance of pathological lesions in histological sections, including estimates of density, frequency, coverage, and the use of semiquantitative scores. The major sampling methods by which these quantitative measures can be obtained from histological sections, including plot or quadrat sampling, transect sampling, and point-quarter sampling, are also described. In addition, the data analysis methods commonly used to analyse quantitative data in neuropathology, including analyses of variance (anova) and principal components analysis (PCA), are discussed. These methods are illustrated with reference to particular problems in the pathological diagnosis of AD and dementia with Lewy bodies (DLB).

Alzheimer Disease↗