PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Data Analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Polychlorinated biphenyls, TEQs, children, and data analysis.

The data analyses in the publications reviewed in Kimbrough et al (1) and Kimbrough and Krouskas (2) are discussed. All of the studies involve infants exposed in utero to small amounts of polychlorinated biphenyls (PCBs) and/or other chlorinated hydrocarbons. Because of missing sampling data, attrition of the cohorts over time, the use of TEQvalues (toxic equivalence) for some congeners of PCBs (polychlorinated biphenyls), PCDDs (polychlorinated dibenzo-p-dioxins) and PCDFs (polychlorinated dibenzofurans), the correlations reported by the authors between chemical exposure and adverse outcomes are not convincing. For the most part the exposures in these cohorts in the US and in Europe were within the ranges observed in the general population, and in many biological samples the level of organochlorines were below the limit of detection of the analytical method. Appropriate evaluation of such data is difficult given the many existing confounders. Many comparisons were made in the different studies. The observed differences were not consistent across studies and most likely occurred by chance. None of the observed differences reported in the children represent clinical disease.

Analysis of Variance↗

Frequency of intravenous medication administration to hospitalised patients: secondary data-analysis of the Belgian nursing minimum data set.

The purpose of this study was to investigate the frequency of intravenous medication administration with Belgian hospitalised patients. Factors, which might influence this frequency of administration, were also studied. Research questions were investigated by secondary data-analysis of the Belgian Nursing Minimum Data Set. The randomised sample consisted of 1,035,681 observations on 421,530 patients. Results of this study demonstrate that one out of three (34%) hospitalised patients received intravenous medication. Medical diagnoses, for which most intravenous medications were administered, were oncological diseases: myeloid (77.9%) and lymphoid (69.4%) leukaemia. Elderly (6.7%) and female (31.2%) patients received significantly less intravenous medication than respectively young (32.9%) (chi(2) = 98411, df = 1, p<0.001) and male (38%) (chi(2) = 2033, df = 1, p<0.001) patients. Patients with intravenous medication administration were labour intensive for nursing staff.

Adolescent↗

Systems for data analysis.

A system for data analysis is the end product of study planning, form design, data entry, data verification, and statistical analysis. This article reviews these steps and considers the fundamental choices in software for data entry and analysis. The appendix includes a listing of general and specialized software for data management and statistical analysis.

Data Interpretation, Statistical↗

Temporal abstraction in intelligent clinical data analysis: a survey.

OBJECTIVE: Intelligent clinical data analysis systems require precise qualitative descriptions of data to enable effective and context sensitive interpretation to take place. Temporal abstraction (TA) provides the means to achieve such descriptions, which can then be used as input to a reasoning engine where they are evaluated against a knowledge base to arrive at possible clinical hypotheses. This paper surveys previous research into the development of intelligent clinical data analysis systems that incorporate TA mechanisms and presents research synergies and trends across the research reviewed, especially those associated with the multi-dimensional nature of real-time patient data streams. The motivation for this survey is case study based research into the development of an intelligent real-time, high-frequency patient monitoring system to provide detection of temporal patterns within multiple patient data streams. RESULTS: The survey was based on factors that are of importance to broaden research into temporal abstraction and on characteristics we believe will assume an increasing level of importance for future clinical IDA systems. These factors were: aspects of the data that is abstracted such as source domain and sample frequency, complexity available within abstracted patterns, dimensionality of the TA and data environment and the knowledge and reasoning underpinning TA processes. CONCLUSION: It is evident from the review that for intelligent clinical data analysis systems to progress into the future where clinical environments are becoming increasingly data-intensive, the ability for managing multi-dimensional aspects of data at high observation and sample frequencies must be provided. Also, the detection of complex patterns within patient data requires higher levels of TA than are presently available. The conflicting matters of computational tractability and temporal reasoning within a real-time environment present a non-trivial problem for investigation in regard to these matters. Finally, to be able to fully exploit the value of learning new knowledge from stored clinical data through data mining and enable its application to data abstraction, the fusion of data mining and TA processes becomes a necessity.

Artificial Intelligence↗

Modelling and simulation for metabolomics data analysis.

The advent of large data sets, such as those produced in metabolomics, presents a considerable challenge in terms of their interpretation. Several mathematical and statistical methods have been proposed to analyse these data, and new ones continue to appear. However, these methods often disagree in their analyses, and their results are hard to interpret. A major contributing factor for the difficulties in interpreting these data lies in the data analysis methods themselves, which have not been thoroughly studied under controlled conditions. We have been producing synthetic data sets by simulation of realistic biochemical network models with the purpose of comparing data analysis methods. Because we have full knowledge of the underlying 'biochemistry' of these models, we are better able to judge how well the analyses reflect true knowledge about the system. Another advantage is that the level of noise in these data is under our control and this allows for studying how the inferences are degraded by noise. Using such a framework, we have studied the extent to which correlation analysis of metabolomics data sets is capable of recovering features of the biochemical system. We were able to identify four major metabolic regulatory configurations that result in strong metabolite correlations. This example demonstrates the utility of biochemical simulation in the analysis of metabolomics data.

Algorithms↗

[GPS--good practice secondary data analysis. Working Group for the Survey and Utilization of Secondary Data (AGENS) of the German Society for Social Medicine and Prevention (DGSMP)].

The scientific use of secondary data, especially of claims data from health insurance funds, has continuously increased in the last years. Therefore the Working Group "Collection and Use of Secondary Data" (AGENS) of the German Society of Social Medicine and Prevention (DGSMP) took the initiative to define quality standards for secondary data analysis. Starting with a review of the Good Epidemiologic Practice (GEP) AGENS adapted the GEP to the specific requirements of secondary data analysis by a multi-stage consensus process. The guideline Good Practice Secondary Date Analysis (GPS) was adopted on January 15 (th), 2005. GPS consists of 10 guidelines which are divided in explaining comments and recommendations. The GPS are targeted to set up standards for secondary data analysis, and they may also be used as a foundation of contracts between data owners and scientists. They are addressed to scientists from health services research and social medicine. AGENS commits itself to revise GPS continuously.

Benchmarking↗

SED, a normalization free method for DNA microarray data analysis.

BACKGROUND: Analysis of DNA microarray data usually begins with a normalization step where intensities of different arrays are adjusted to the same scale so that the intensity levels from different arrays can be compared with one other. Both simple total array intensity-based as well as more complex "local intensity level" dependent normalization methods have been developed, some of which are widely used. Much less developed methods for microarray data analysis include those that bypass the normalization step and therefore yield results that are not confounded by potential normalization errors. RESULTS: Instead of focusing on the raw intensity levels, we developed a new method for microarray data analysis that maps each gene's expression intensity level to a high dimensional space of SEDs (Signs of Expression Difference), the signs of the expression intensity difference between a given gene and every other gene on the array. Since SED are unchanged under any monotonic transformation of intensity levels, the SED based method is normalization free. When tested on a multi-class tumor classification problem, simple Naive Bayes and Nearest Neighbor methods using the SED approach gave results comparable with normalized intensity-based algorithms. Furthermore, a high percentage of classifiers based on a single gene's SED gave good classification results, suggesting that SED does capture essential information from the intensity levels. CONCLUSION: The results of testing this new method on multi-class tumor classification problems suggests that the SED-based, normalization-free method of microarray data analysis is feasible and promising.

Central Nervous System Neoplasms↗

Data analysis: statistical analysis and use of historical control data.

Survival-adjusted methods for the statistical analysis of tumor data from long-term rodent carcinogenicity studies are described. Although most of these methods require knowledge of whether individual tumors are "fatal" or "incidental," such determinations may be difficult. Several methods for dealing with this and with other data analysis issues are discussed. Historical control tumor data may be useful in the interpretation of rodent carcinogenicity studies, particularly for rare tumors and for borderline effects. Although statistical methods are available for using historical control data in a formal testing framework, the primary difficulty is establishing a database that is truly comparable to the study under evaluation with respect to those factors known to influence tumor occurrence. Major sources of variability in tumor incidence include the animal room environment, dietary factors/body weight, gross necropsy and slide preparation procedures, and histopathology diagnosis. The National Toxicology Program's use of historical control data is briefly described and illustrated.

Adrenal Cortex Neoplasms↗

Early prediction of wheat quality: analysis during grain development using mass spectrometry and multivariate data analysis.

Matrix-assisted laser desorption/ionisation time-of-flight mass spectrometry and multivariate data analysis have been used for the determination of wheat quality at different stages of grain development. Wheat varieties with one of two different end-use qualities (i.e. suitable or not suitable for bread-making purposes) were investigated. The samples were collected from grains from 15 until 45 days post-anthesis (dpa). Gluten proteins from wheat grains were extracted and subsequently analysed by mass spectrometry. Discrimination partial least-squares regression and soft independent modelling of class analogy were used to determine the quality of new and unknown wheat samples. With these methods, we were able to predict correctly the end-use qualities at every stage investigated. This new fast technique, based on the rapidity of mass spectrometry combined with the objectivity of multivariate data analysis, offers a method that can replace the traditional rather time-consuming ones such as gel electrophoresis. This study focused on the determination of wheat quality at 15 dpa, when the grain is due for harvest 1 month later.

Edible Grain↗

Megavariate data analysis of mass spectrometric proteomics data using latent variable projection method.

There are many data mining techniques for processing and general learning of multivariate data. However, we believe the wavelet transformation and latent variable projection method are particularly useful for spectroscopic and chromatographic data. Projection based methods are designed to handle hugely multivariate nature of such data effectively. For the actual analysis of the data we have used latent variable projection methods such as principal component analysis (PCA) and partial least squares projection to latent structures based discriminant analysis (PLS-DA) to analyze the raw data presented to the participants of the First Duke Proteomics Data Mining Conference. PCA was used to solve problem #1 (clustering problem) and the PLS-DA was used to solve problem #2 (classification problem). The idea of internal and external cross-validation was used to validate the model obtained from the classification analysis. The simple two-component PLS-DA model obtained from the analysis performed well. The model has completely separated the two groups from all the data. The same model applied on two-thirds of the data showed good performance by external validation with independent test set of remaining 13 specimens obtained by setting aside the spectra of every third specimen (accuracy of 85%).

Artificial Intelligence↗

Single-trial variable model for event-related fMRI data analysis.

Most methods for fMRI data analysis assume that the hemodynamic responses (HRs) across similar experimental events are same. This assumption is not appropriate when HRs vary unpredictably from trial to trial. Here, we introduce a new method for fMRI data analysis. The main features of the proposed method are as follows: 1) The trial-to-trial variability is modeled as meaningful signal rather than assuming that the same HR is evoked in each trial; 2) Since the proposed method is a constrained optimization based general framework, it could be extended by utilizing prior knowledge of HR; 3) The traditional deconvolution method can be included into our method as a special case. A comparison of performance on simulated fMRI datasets is made using the general linear model, the deconvolution method and the proposed method with receiver operating characteristic (ROC) methodology. In addition, we examined the effectiveness and usefulness of our method on real experimental data.

Adult↗

Longitudinal data analysis. A comparison between generalized estimating equations and random coefficient analysis.

The analysis of data from longitudinal studies requires special techniques, which take into account the fact that the repeated measurements within one individual are correlated. In this paper, the two most commonly used techniques to analyze longitudinal data are compared: generalized estimating equations (GEE) and random coefficient analysis. Both techniques were used to analyze a longitudinal dataset with six measurements on 147 subjects. The purpose of the example was to analyze the relationship between serum cholesterol and four predictor variables, i.e., physical fitness at baseline, body fatness (measured by sum of the thickness of four skinfolds), smoking and gender. The results showed that for a continuous outcome variable, GEE and random coefficient analysis gave comparable results, i.e., GEE-analysis with an exchangeable correlation structure and random coefficient analysis with only a random intercept were the same. There was also no difference between both techniques in the analysis of a dataset with missing data, even when the missing data was highly selective on earlier observed data. For a dichotomous outcome variable, the magnitude of the regression coefficients and standard errors was higher when calculated with random coefficient analysis then when calculated with GEE-analysis. Analysis of a dataset with missing data with a dichotomous outcome variable showed unpredictable results for both GEE and random coefficient analysis. It can be concluded that for a continuous outcome variable, GEE and random coefficient analysis are comparable. Longitudinal data-analysis with dichotomous outcome variables should, however, be interpreted with caution, especially when there are missing data.

Adipose Tissue↗

Strong-association-rule mining for large-scale gene-expression data analysis: a case study on human SAGE data.

BACKGROUND: The association-rules discovery (ARD) technique has yet to be applied to gene-expression data analysis. Even in the absence of previous biological knowledge, it should identify sets of genes whose expression is correlated. The first association-rule miners appeared six years ago and proved efficient at dealing with sparse and weakly correlated data. A huge international research effort has led to new algorithms for tackling difficult contexts and these are particularly suited to analysis of large gene-expression matrices. To validate the ARD technique we have applied it to freely available human serial analysis of gene expression (SAGE) data. RESULTS: The approach described here enables us to designate sets of strong association rules. We normalized the SAGE data before applying our association rule miner. Depending on the discretization algorithm used, different properties of the data were highlighted. Both common and specific interpretations could be made from the extracted rules. In each and every case the extracted collections of rules indicated that a very strong co-regulation of mRNA encoding ribosomal proteins occurs in the dataset. Several rules associating proteins involved in signal transduction were obtained and analyzed, some pointing to yet-unexplored directions. Furthermore, by examining a subset of these rules, we were able both to reassign a wrongly labeled tag, and to propose a function for an expressed sequence tag encoding a protein of unknown function. CONCLUSIONS: We show that ARD is a promising technique that turns out to be complementary to existing gene-expression clustering techniques.

Algorithms↗

A computerized data analysis system for electrogastrogram.

A comprehensive computerized data analysis system for the electrogastrogram is presented in this paper. The electrogastrogram (EGG) is a cutaneous measurement of electrical activity of the stomach by positioning electrodes on the abdominal skin. Since the signal-to-noise ratio of the EGG is very low, visual analysis is impossible. The data analysis system presented in this paper contains a series of PC programs to perform: (a) data acquisition and real-time A/D conversion; (b) digital filter design and digital filtering; (c) adaptive cancellation of respiratory artifact; (d) smoothed power spectral analysis; (e) adaptive running spectral analysis; (f) two- or three-dimensional display of the EGG and analysis results. The basic principles of the system and sample results are presented.

Data Display↗

Analysis of bone scintigram data using speech recognition reporting system--data analysis with speech recognition system.

Five hundred eighty bone scintigram reports were stored using a voice pattern recognition system in a general-purpose, middle-sized computer (ACOS-650). Bone scintigraphy carried out in our institute was examined by analyzing these data. The results of the examination showed that the introduction of this system made it possible to analyze all the data quickly. Before the introduction of this system, the data able to be analyzed had been restricted because of their complexity. The results also showed that this system would be useful for understanding the examinations carried out in the whole hospital as well as for analyzing metastatic tumors and the number of patients receiving examinations. Furthermore, this system would be helpful in the logical analysis of reports prepared by doctors.

Bone and Bones↗

Effectiveness of an electronic medical record clinical quality alert prepared by off-line data analysis.

We tested whether off-line data analysis, instead of event monitoring, was a viable method for initiating a clinical quality alert. A cohort of patients eligible for an alert was identified by off-line data analysis and a flag was set in their ambulatory Electronic Medical Records. One hundred clinicians were randomly assigned either to a control group or to a group that received the alert when viewing the electronic medical record of eligible patients. Primarily due to actions of their clinicians, 315 of the 580 patients (54.3%) seen by alerted clinicians were no longer eligible for the alert at the end of the one month study, compared to 128 of the 496 patients (25.8%) seen by control clinicians (p<.001). When not alerted, Allied Health clinicians were less likely than physicians to prescribe aspirin, but they responded similarly to the alert. There were no differences in response by specialty or gender of the clinician. Off-line data analysis proved to be an effective method of initiating a clinical alert.

Ambulatory Care Information Systems↗

Flow cytometry data analysis: comparing large multivariate data sets using classification trees.

This paper describes a method to compare flow cytometry data sets, which typically contain 50,000 six-parameter measurements each. By this method, the data points in two such data sets are divided into subpopulations using a binary classification tree generated from the data. The chi 2 test is then used to establish the homogeneity of the two data sets based on how their data are distributed across these subpopulations. Preliminary results indicate that this comparison method is sufficiently sensitive to detect differences between flow cytometry data sets that are too subtle for human investigators to notice.

Animals↗