PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Explorative data analysis of two-dimensional electrophoresis gels.

Methods for classification of two-dimensional (2-DE) electrophoresis gels based on multivariate data analysis are demonstrated. Two-dimensional gels of ten wheat varieties are analyzed and it is demonstrated how to classify the wheat varieties in two qualities and a method for initial screening of gels is presented. First, an approach is demonstrated in which no prior knowledge of the separated proteins is used. Alignment of the gels followed by a simple transformation of data makes it possible to analyze the gels in an automated explorative manner by principal component analysis, to determine if the gels should be further analyzed. A more detailed approach is done by analyzing spot volume lists by principal components analysis and partial least square regression. The use of spot volume data offers a mean to investigate the spot pattern and link the classified protein patterns to distinct spots on the gels for further investigation. The explorative approach in analysis of 2-D gels makes it possible, in a fast and convenient way, to screen many gels in order to determine the protein patterns that form clusters and could be selected for further examination.

Diagnostic Imaging↗

An automated monitoring system for VOC ozone precursors in ambient air: development, implementation and data analysis.

An automated system for the monitoring of volatile organic compound (VOC) ozone precursors in ambient air is described. The measuring technique consists of subambient preconcentration on a cooled trap followed by thermal desorption and GC/FID analysis. First, the technical development, which permits detection limits below 0.05 ppbv to be reached, proceeded in two steps: (1). the determination of optimum sampling parameters (trap composition and conditioning, outlet split, desorption temperature); (2). the development of a reliable calibration method based on a highly accurate standard. Then, a 4-year field application of the hourly measuring chain was carried out at two urban sites. On the one hand, quality control procedures provided the best VOC identification (peak assignment) and quantification (reproducibility, blank system control). On the other hand, the success and performances of the routine experience (88% of the measurements covered more than 40 target compounds) indicated the high quality and suitability of the instrumentation which is actually applied in several French air quality monitoring networks. Finally, an example of data analysis is presented. Data handling identified important organic compound sources other than vehicle exhaust.

Journal Article↗

Proper multivariate conditional autoregressive models for spatial data analysis.

In the past decade conditional autoregressive modelling specifications have found considerable application for the analysis of spatial data. Nearly all of this work is done in the univariate case and employs an improper specification. Our contribution here is to move to multivariate conditional autoregressive models and to provide rich, flexible classes which yield proper distributions. Our approach is to introduce spatial autoregression parameters. We first clarify what classes can be developed from the family of Mardia (1988) and contrast with recent work of Kim et al. (2000). We then present a novel parametric linear transformation which provides an extension with attractive interpretation. We propose to employ these models as specifications for second-stage spatial effects in hierarchical models. Two applications are discussed; one for the two-dimensional case modelling spatial patterns of child growth, the other for a four-dimensional situation modelling spatial variation in HLA-B allele frequencies. In each case, full Bayesian inference is carried out using Markov chain Monte Carlo simulation.

Alleles↗

Incremental genetic K-means algorithm and its application in gene expression data analysis.

BACKGROUND: In recent years, clustering algorithms have been effectively applied in molecular biology for gene expression data analysis. With the help of clustering algorithms such as K-means, hierarchical clustering, SOM, etc, genes are partitioned into groups based on the similarity between their expression profiles. In this way, functionally related genes are identified. As the amount of laboratory data in molecular biology grows exponentially each year due to advanced technologies such as Microarray, new efficient and effective methods for clustering must be developed to process this growing amount of biological data. RESULTS: In this paper, we propose a new clustering algorithm, Incremental Genetic K-means Algorithm (IGKA). IGKA is an extension to our previously proposed clustering algorithm, the Fast Genetic K-means Algorithm (FGKA). IGKA outperforms FGKA when the mutation probability is small. The main idea of IGKA is to calculate the objective value Total Within-Cluster Variation (TWCV) and to cluster centroids incrementally whenever the mutation probability is small. IGKA inherits the salient feature of FGKA of always converging to the global optimum. C program is freely available at http://database.cs.wayne.edu/proj/FGKA/index.htm. CONCLUSIONS: Our experiments indicate that, while the IGKA algorithm has a convergence pattern similar to FGKA, it has a better time performance when the mutation probability decreases to some point. Finally, we used IGKA to cluster a yeast dataset and found that it increased the enrichment of genes of similar function within the cluster.

Algorithms↗

Late patency of the carotid artery after endarterectomy. Problems of definition, follow-up methodology, and data analysis.

To determine the relative incidence of recurrent carotid stenosis (RCS) and the effect of methodology on data analysis and interpretation, late results were obtained for 232 patients (270 procedures) from 1 to 51 months (mean 22 months) after carotid endarterectomy (group A). Patency of the carotid artery was confirmed by postoperative intravenous digital subtraction angiography (DSA) for most of the series, and a subset (subgroup A1) of 113 patients (129 procedures) also received DSA studies at later intervals of 4 to 49 months (mean 26 months). There were 23 late deaths and five late strokes. Only two of the strokes were ipsilateral to previous endarterectomy, and both of these patients had normal follow-up DSA studies. Late DSA imaging revealed either no RCS or only trivial defects (20% diameter or less) in 111 arteries, moderate (36% to 60%) RCS in nine, severe (70% to 90%) RCS requiring secondary procedures in eight, and internal carotid occlusion in one. Depending on the definition of RCS (secondary operation vs greater than or equal to 30% angiographic lesions), the cohort selected for analysis (group A vs subgroup A1), and the approach to calculations (crude vs cumulative), the incidence of recurrent stenosis after carotid reconstruction in this single study could be expressed within the extraordinary wide range of 3% to 32%. Although carotid endarterectomy was associated with uniformly low risk for late stroke, these results confirm that the reported recurrence rate may be substantially influenced by the method in which data are grouped and manipulated. Consistently presented data are essential to any comparisons concerning the surgical therapy for extracranial disease.

Actuarial Analysis↗

Component plane presentation integrated self-organizing map for microarray data analysis.

We describe a powerful approach, component plane presentation integrated self-organizing map (SOM), for the analysis of microarray data. This approach allows the display of multi-dimensional SOM outputs of microarray data in multiple sample specific presentations, providing distinct advantages in visual inspection of biological significances of genes clustered in each map unit with respect to each RNA sample. Beneficial potentials of the approach are highlighted by processing microarray data from yeast cells as well as human breast malignancies.

Gene Expression Profiling↗

A mixture model for duration data: analysis of second births in China.

In this paper we introduce a mixture model in which we combine logistic regression and piecewise proportional hazards models for analysis of duration data. The model allows simultaneous estimation of two sets of effects of covariates: one of the probability of an event and the other of the timing of the event. We illustrate the application of the model through an analysis of the effects of women's characteristics and of the acceptance of a one-child certificate on the birth of second children in China. Both factors affect the probability of having a second child, but only the acceptance of a one-child certificate has a significant and strong effect on the second-birth interval.

Algorithms↗

Graph-based normalization and whitening for non-linear data analysis.

In this paper we construct a graph-based normalization algorithm for non-linear data analysis. The principle of this algorithm is to get a spherical average neighborhood with unit radius. First we present a class of global dispersion measures used for "global normalization"; we then adapt these measures using a weighted graph to build a local normalization called "graph-based" normalization. Then we give details of the graph-based normalization algorithm and illustrate some results. In the second part we present a graph-based whitening algorithm built by analogy between the "global" and the "local" problem.

Algorithms↗

Constrained and restrained refinement in EXAFS data analysis with curved wave theory.

This paper describes methods of constrained and restrained refinement of EXAFS data which provide a means of substantially reducing the number of independent parameters compared to conventional least-squares methods commonly used. Constrained refinement allows a major reduction in the number of free parameters for a refinement of a structural model. In restrained refinement, additional structural information from well-characterized small molecules is used to provide additional observations in the data analysis. Even though these methods are of general application to the majority of complex systems, they are particularly valuable for biological molecules. The methods are of major advantage for ligands where significant multiple scattering is present, e.g., histidine, tyrosine, CO, CN, etc. The bases of these methods are described, and applications to some complex chemical and biological systems are given.

Fetal Hemoglobin↗

Application of neural networks to population pharmacokinetic data analysis.

This research examined the applicability of using a neural network approach to analyze population pharmacokinetic data. Such data were collected retrospectively from pediatric patients who had received tobramycin for the treatment of bacterial infection. The information collected included patient-related demographic variables (age, weight, gender, and other underlying illness), the individual's dosing regimens (dose and dosing interval), time of blood drawn, and the resulting tobramycin concentration. Neural networks were trained with this information to capture the relationships between the plasma tobramycin levels and the following factors: patient-related demographic factors, dosing regimens, and time of blood drawn. The data were also analyzed using a standard population pharmacokinetic modeling program, NON-MEM. The observed vs predicted concentration relationships obtained from the neural network approach were similar to those from NONMEM. The residuals of the predictions from neural network analyses showed a positive correlation with that from NONMEM. Average absolute errors were 33.9 and 37.3% for neural networks and 39.9% for NONMEM. Average prediction errors were found to be 2.59 and -5.01% for neural networks and 17.7% for NONMEM. We concluded that neural networks were capable of capturing the relationships between plasma drug levels and patient-related prognostic factors from routinely collected sparse within-patient pharmacokinetic data. Neural networks can therefore be considered to have potential to become a useful analytical tool for population pharmacokinetic data analysis.

Anti-Bacterial Agents↗

Data analysis in qualitative research: a plea for sharing the magic and the effort.

This discussion of data analysis in qualitative research addresses the question of how authors describe this aspect of their research. I suggest that the tradition of organization of research papers from quantitative research is not a good fit for writing qualitative research. I argue for less jargon and more detailed description, with the analytic process integrated into the findings and interpretation.

Humans↗

Data mining and structuring of executable data analysis reports: guideline development and implementation in a narrow sense.

In this paper we present a data mining scenario that supports development of automated web-based documentation of data analysis for diagnosis and treatment. The documents can be seen as guidelines in a narrow sense, and are designed to include executable modules for the corresponding decision support systems. Our aim is to discuss the possibilities of identifying certain types of diagnoses and treatments for which guidelines can be generated and computerised more systematically.

Artificial Intelligence↗

Analog processing of vestibular nystagmus for on-line cross- correlation data analysis.

An analog processing circuit is described which allow accurate measurement of the phase relationships between input angular acceleration and resulting eye velocity. Vestibular nystagmic data are processed via analog technics to yield slowphase eye velocity. The turntable velocity input is cross-correlated with the eye velocity output, using a Nicolet MED-80 minicomputer system. The resulting correlograms are further processed to obtain precise phase information. Test data analysis shows a system resolution within 1 degree. Data from human and animal subjects are portrayed.

Acceleration↗

Computer assisted data analysis in intensive care: the ICDEV project--development of a scientific database system for intensive care (Intensive Care Data Evaluation Project).

INTRODUCTION: Patient Data Management Systems (PDMS) for ICUs collect, present and store clinical data. Various intentions make analysis of those digitally stored data desirable, such as quality control or scientific purposes. The aim of the Intensive Care Data Evaluation project (ICDEV), was to provide a database tool for the analysis of data recorded at various ICUs at the University Clinics of Vienna. SETTINGS: General Hospital of Vienna, with two different PDMSs used: CareVue 9000 (Hewlett Packard, Andover, USA) at two ICUs (one medical ICU and one neonatal ICU) and PICIS Chart+ (PICIS, Paris, France) at one Cardiothoracic ICU. CONCEPT AND METHODS: Clinically oriented analysis of the data collected in a PDMS at an ICU was the beginning of the development. After defining the database structure we established a client-server based database system under Microsoft Windows NI and developed a user friendly data quering application using Microsoft Visual C++ and Visual Basic; RESULTS: ICDEV was successfully installed at three different ICUs, adjustment to the different PDMS configurations were done within a few days. The database structure developed by us enables a powerful query concept representing an 'EXPERT QUESTION COMPILER' which may help to answer almost any clinical questions. Several program modules facilitate queries at the patient, group and unit level. Results from ICDEV-queries are automatically transferred to Microsoft Excel for display (in form of configurable tables and graphs) and further processing. CONCLUSIONS: The ICDEV concept is configurable for adjustment to different intensive care information systems and can be used to support computerized quality control. However, as long as there exists no sufficient artifact recognition or data validation software for automatically recorded patient data, the reliability of these data and their usage for computer assisted quality control remain unclear and should be further studied.

Austria↗

Multiple peak alignment in sequential data analysis: a scale-space-based approach.

In this paper, we address the multiple peak alignment problem in sequential data analysis with an approach based on the Gaussian scale-space theory. We assume that multiple sets of detected peaks are the observed samples of a set of common peaks. We also assume that the locations of the observed peaks follow unimodal distributions (e.g., normal distribution) with their means equal to the corresponding locations of the common peaks and variances reflecting the extension of their variations. Under these assumptions, we convert the problem of estimating locations of the unknown number of common peaks from multiple sets of detected peaks into a much simpler problem of searching for local maxima in the scale-space representation. The optimization of the scale parameter is achieved using an energy minimization approach. We compare our approach with a hierarchical clustering method using both simulated data and real mass spectrometry data. We also demonstrate the merit of extending the binary peak detection method (i.e., a candidate is considered either as a peak or as a nonpeak) with a quantitative scoring measure-based approach (i.e., we assign to each candidate a possibility of being a peak).

Algorithms↗

Gene expression microarray data analysis of decidual and placental cell differentiation.

Gene expression analysis using DNA microarray approaches have provided new insights into the physiology and pathophysiology of many biological processes. These include identification of genetic programs and pathways that underlie cell and tissue differentiation and gene expression programs responsive to genetic perturbations, drugs, toxins, and infectious agents. In this chapter, we present methods for the analysis of microarray data using earlier investigations from our laboratory as examples of how gene expression patterns for cellular differentiation may be detected and analyzed for biological significance and how regulated genes may be classified into functional categories and pathways.

Animals↗

Clustering binary fingerprint vectors with missing values for DNA array data analysis.

Oligonucleotide fingerprinting is a powerful DNA array-based method to characterize cDNA and ribosomal RNA gene (rDNA) libraries and has many applications including gene expression profiling and DNA clone classification. We are especially interested in the latter application. A key step in the method is the cluster analysis of fingerprint data obtained from DNA array hybridization experiments. Most of the existing approaches to clustering use (normalized) real intensity values and thus do not treat positive and negative hybridization signals equally (positive signals are much more emphasized). In this paper, we consider a discrete approach. Fingerprint data are first normalized and binarized using control DNA clones. Because there may exist unresolved (or missing) values in this binarization process, we formulate the clustering of (binary) oligonucleotide fingerprints as a combinatorial optimization problem that attempts to identify clusters and resolve the missing values in the fingerprints simultaneously. We study the computational complexity of this clustering problem and a natural parameterized version and present an efficient greedy algorithm based on MINIMUM CLIQUE PARTITION on graphs. The algorithm takes advantage of some unique properties of the graphs considered here, which allow us to efficiently find the maximum cliques as well as some special maximal cliques. Our preliminary experimental results on simulated and real data demonstrate that the algorithm runs faster and performs better than some popular hierarchical and graph-based clustering methods. The results on real data from DNA clone classification also suggest that this discrete approach is more accurate than clustering methods based on real intensity values in terms of separating clones that have different characteristics with respect to the given oligonucleotide probes.

Algorithms↗

Longitudinal data analysis for linear Gaussian models with random disturbed-highest-derivative-polynomial subject effects.

For linear regression analysis of longitudinal data with Gaussian response, I propose a new model to generalize the traditional class of random effects models in which the random effects are deterministic polynomials with coefficients randomly distributed over subjects with mean zero. The generalization is accomplished by adding zero mean Gaussian 'disturbances' to the highest derivative of each random coefficient subject polynomial, independently at each observation time. The resulting random effects, which have mean zero at each observation time, are called disturbed highest derivative polynomials (DHDPs). The disturbances induce serial correlation and also allow the subject-specific DHDP time trends to be non-linear. I do not estimate the subject-specific DHDP time trends. Analysis is based on the marginal model, that is, the fixed effects or population model obtained by integrating the random polynomial coefficients and all disturbances out of the joint distribution of themselves and the response vector. This allows a 'population averaged' interpretation. One can select the DHDP order by an information criterion. When the population time trend is not correctly modelled, the optimal DHDP order will be larger than when it is correctly modelled. One can make the covariance matrix of the regression coefficients robust to errors in modelling the within-subject dependence. I describe the relationship of a DHDP to a smoothing polynomial spline, and show how to replace the DHDP model with a smoothing polynomial spline model for the within-subject dependence in the marginal model.

Bias↗