PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

The meaning of diagnostic test results: a spreadsheet for swift data analysis.

AIMS: To design a spreadsheet program to: (a) analyse rapidly diagnostic test result data produced in local research or reported in the literature; (b) correct reported predictive values for disease prevalence in any population; (c) estimate the post-test probability of disease in individual patients. MATERIALS AND METHODS: Microsoft Excel(TM)was used. Section A: a contingency (2 x 2) table was incorporated into the spreadsheet. Formulae for standard calculations [sample size, disease prevalence, sensitivity and specificity with 95% confidence intervals, predictive values and likelihood ratios (LRs)] were linked to this table. The results change automatically when the data in the true or false negative and positive cells are changed. Section B: this estimates predictive values in any population, compensating for altered disease prevalence. Sections C-F: Bayes' theorem was incorporated to generate individual post-test probabilities. The spreadsheet generates 95% confidence intervals, LRs and a table and graph of conditional probabilities once the sensitivity and specificity of the test are entered. The latter shows the expected post-test probability of disease for any pre-test probability when a test of known sensitivity and specificity is positive or negative. RESULTS: This spreadsheet can be used on desktop and palmtop computers. The MS Excel(TM)version can be downloaded via the Internet from the URL ftp://radiography.com/pub/Rad-data99.xls CONCLUSION: A spreadsheet is useful for contingency table data analysis and assessment of the clinical meaning of diagnostic test results.

Bayes Theorem↗

A new approach to near-infrared spectral data analysis using independent component analysis.

This paper presents a new approach to near-infrared spectral (NIR) data analysis that is based on independent component analysis (ICA). The main advantage of the new method is that it is able to separate the spectra of the constituent components from the spectra of their mixtures. The separation is a blind operation, since the constituent components of mixtures can be unknown. The ICA based method is therefore particularly useful in identifying the unknown components in a mixture as well as in estimating their concentrations. The approach is introduced by reference to case studies and compared to other techniques for NIR analysis including principal component regression (PCR), multiple linear regression (MLR), and partial least squares (PLS) as well as Fourier and wavelet transforms.

Adipose Tissue↗

Microarray data analysis: a practical approach for selecting differentially expressed genes.

BACKGROUND: The biomedical community is rapidly developing new methods of data analysis for microarray experiments, with the goal of establishing new standards to objectively process the massive datasets produced from functional genomic experiments. Each microarray experiment measures thousands of genes simultaneously producing an unprecedented amount of biological information across increasingly numerous experiments; however, in general, only a very small percentage of the genes present on any given array are identified as differentially regulated. The challenge then is to process this information objectively and efficiently in order to obtain knowledge of the biological system under study and by which to compare information gained across multiple experiments. In this context, systematic and objective mathematical approaches, which are simple to apply across a large number of experimental designs, become fundamental to correctly handle the mass of data and to understand the true complexity of the biological systems under study. RESULTS: The present report develops a method of extracting differentially expressed genes across any number of experimental samples by first evaluating the maximum fold change (FC) across all experimental parameters and across the entire range of absolute expression levels. The model developed works by first evaluating the FC across the entire range of absolute expression levels in any number of experimental conditions. The selection of those genes within the top X% of highest FCs observed within absolute expression bins was evaluated both with and without the use of replicates. Lastly, the FC model was validated by both real time polymerase chain reaction (RT-PCR) and variance data. Semi-quantitative RT-PCR analysis demonstrated 73% concordance with the microarray data from Mu11K Affymetrix GeneChips. Furthermore, 94.1% of those genes selected by the 5% FC model were found to lie above measurement variability using a SDwithin confidence level of 99.9%. CONCLUSION: As evidenced by the high rate of validation, the FC model has the potential to minimize the number of required replicates in expensive microarray experiments by extracting information on gene expression patterns (e.g. characterizing biological and/or measurement variance) within an experiment. The simplicity of the overall process allows the analyst to easily select model limits which best describe the data. The genes selected by this process can be compared between experiments and are shown to objectively extract information which is biologically & statistically significant.

Animals↗

Examination of dioxin fluxes recorded in dated aquatic-sediment cores in the Kanto region of Japan using multivariate data analysis.

Past dioxin (coplanar polychlorinated biphenyl (Co-PCB), 2,3.7,8-substituted polychlorinated dibenzo-p-dioxin (PCDD) and 2,3,7,8-substituted polychlorinated dibenzofuran (PCDF)) fluxes recorded in dated aquatic-sediment cores were analyzed using principal component analysis (PCA). The data set consisted of samples from four cores collected from the Kanto region of Japan. Time trends and spatial differences in the dioxin flux were examined, and the potential relationship to emission sources was investigated. Twenty-five compounds and 58 core slices, corresponding to the later half of the 20th century, were subjected to the analysis. The PCA of both log-transformed and maximum-value-standardized data successfully divided the dioxin compounds into a small number of groups, and three similar clusters of Co-PCBs. PCDDs and penta- to hepta-CDFs were identified. PCB formulations used in the past are judged to have been responsible for the major part of the Co-PCB flux recorded in the sediment cores. However, the relationship to emission sources needs further investigation. It is suggested that most 2,3,7,8-substituted PCDDs and PCDFs are different from Co-PCBs in their emission sources or movements in the environment. The subcore clusters obtained from the PCA of log-transformed data show that the cores from different sampling areas exhibited distinct dioxin fluxes and compositions. Common time trends among the cores were more effectively summarized by the PCA of maximum-value-standardized data focusing on relative time trends. PC scores show that recently the flux of each dioxin compound in the four cores has been generally declining after having reached a peak.

Benzofurans↗

Proficiency of the Tradescantia-micronucleus image analysis system for scoring micronucleus frequencies and data analysis.

The Tradescantia-micronucleus (Trad-MCN) bioassay is an efficient short-term test for genotoxicity of pollutants. In order to increase the efficiency and to standardize the micronucleus (MCN) scoring process, an automated scoring system was developed using the principle of image analysis in computer science. This assemblage is called the Tradescantia-micronucleus image analysis (Trad-MCNIA) system. The MCN frequencies scored by this system were compared with those scored by human observation for its proficiency. A set of low MCN frequency (around 5 MCN/100 tetrads) slides prepared from a control group, a set of medium MCN frequency (around 20 MCN/100 tetrads) slides prepared from sodium azide treated plant cuttings and a set of high MCN frequency (around 50 MCN/100 tetrads) slides prepared from X-ray treated materials were used for this study. In the low MCN frequency slides, the Trad-MCNIA system scored about the same value as human observation. In the medium and high frequency slides, MCN frequencies scored by the system were lower than those scored by human observers. This discrepancy was corrected by increasing the power of the objective of the microscope in the system. The MCN frequencies scored by the system attained 90% congruity with those scored by human observers after the correction. The scoring speed of the system was about 3.5 times as fast as that by human observers, and the data could be statistically analyzed immediately after the data scores were recorded. Further improvements can be made by upgrading the video camera and the computer speed.

Azides↗

Recruitment in NHLBI population-based studies and randomized clinical trials: data analysis and survey results.

Data on screening and recruitment from current and previous NHLBI population-based studies (PBSs) and randomized clinical trials (RCTs) were examined. In only two of the studies examined was the projected recruitment completed within the planned recruitment period. The shape of the graph of the relation between enrollment of participants and time varies by study. A single summary statistic for measuring the efficiency of recruitment in RCTs and PBSs is proposed and applied to the examined studies. In addition to providing summary data on recruitment for several studies, this article reports the survey results of a questionnaire sent to the coordinating centers of currently and previously funded National Heart, Lung, and Blood Institute and Veteran's Administration studies. The purpose was to ascertain the desirability of recommending that a generic core of information be collected on recruitment and screening in future studies. Most respondents believed that comparing data collected uniformly and prospectively might be helpful in designing further studies. The variables most respondents believed to be potentially useful are described.

Clinical Trials as Topic↗

Use of prior information to stabilize a population data analysis.

When modeling new data with a complex population pharmacokinetic/pharmacodynamic model, there may not be sufficient information to obtain estimates of all parameters. In this case information from previous studies can also be used to help stabilize estimation. Using simulated data, we explored three different ways to do this. (i) Some parameter values were fixed to estimates obtained from earlier data. (ii) The earlier data were combined with the current data. (iii) The objective function based on the current data was augmented by a penalty function expressing summary information obtained from the earlier data. This last method is similar to the use of a Bayesian prior. It may be particularly useful when either the combined data set of method (ii) is very large and leads to large computation times or when the early data are not readily available. With this method, two different types of penalty functions were used. With our examples, the three methods all resulted in stabilized estimation. Methods (ii) and (iii) gave similar results for parameter and standard error estimation, especially with respect to fixed effects parameters. For hypothesis testing, results obtained with method (i) are very problematic. There are also problems with the results obtained with method (iii), but they are much less severe, and when the design for the earlier data is known, they can be corrected by using a computer-intensive simulation test procedure.

Algorithms↗

A procedure for data analysis of the rodent micronucleus test involving a historical control.

No standard procedure of data analysis for rodent micronucleus tests involving historical controls has been established. In the present paper, under the presumption that the distribution of the historical control is stable and reliable, a procedure with three statistical steps is proposed to analyze the frequency of micronucleated polychromatic erythrocytes (MNPCEs). In the first step, the frequencies of MNPCEs in negative and positive control groups of a current experiment of the micronucleus test are compared with the distribution of historical negative and positive controls to examine the technical validity of the current experiment. In the second step, the frequency of MNPCEs in each treatment group is compared with the distribution of the historical negative control. In the third step, the dose-response relation is tested with the Cochran-Armitage trend test. A Monte Carlo stimulation study shows that the power of this procedure is acceptable and also this procedure is robust. An application of this procedure on real data reveals that it is effective in detecting clastogenic chemicals when the probability of a type I error is nearly .01.

Animals↗

Exploratory data analysis of evoked response single trials based on minimal spanning tree.

OBJECTIVE: An exploratory data analysis framework, based on minimal spanning tree, is proposed as a means to support the analysis of single trial (ST) electrophysiological signals. The core of this framework is the compact description of the input ST sample in a form of content-dependent ordered lists. Based on the established hierarchies, efficient ways to increase the SNR, extract prototypical responses, visualize possible self-organization trends in the sample and track the course of evoked response along the trial-to-trial dimension, are proposed. METHOD: Magnetoencephalographic auditory evoked responses were used for demonstrating and validating the introduced framework. RESULTS AND CONCLUSION: The results demonstrate the benefits, from this intelligent manipulation of STs, in understanding and enhancing the actual evoked signal. Specifically we find support for stimulus-induced phase-resetting hypothesis in the 3-20 Hz band, the existence of trials void of the prototypical evoked response, and an order across the single trial set hinting at an underlying process with long time scale.

Adult↗

Secondary data analysis: research method for the clinical nurse specialist.

This article presents a description of secondary data analysis and suggests that this type of research methodology may be helpful in facilitating research by the clinical nurse specialist (CNS). The article discusses the advantages and disadvantages of the use of this method specifically in relation to the CNS and offers suggestions for sources of data.

Data Collection↗

Techniques to identify clinical contexts during automated data analysis.

The interpretation of automatically collected data to produce intelligent alarms and identify particular conditions is nearly impossible without identifying the specific context in which the data are obtained. Shifts in clinical context occur because of changes in the patient's physiologic state, or due to the passage of time, or due to changes imposed by therapeutic intervention such as surgery. Techniques to identify such changes in clinical context are discussed with particular attention to the application of cluster analysis, discriminant analysis, and statistical predictors. An example of these analyses applied to EEG data is presented, showing an unexpected hysteresis of EEG behavior in response to an hypoxic challenge.

Cluster Analysis↗

Software for temporal gait data analysis.

This study presents a computer program, developed to support a low-cost, portable telemetry system that has been designed to assess footfall timing. This software frees the user from data processing and allows concentration on data analysis. The new technique has been applied with accuracy and reliability to the analysis of the gait of orthopedic patients, athletes, mountaineers, etc. The subroutines developed for data acquisition, storage and analysis are explained in detail, and an example is presented.

Data Interpretation, Statistical↗

Comparison of Ehrlich ascites tumour and mouse liver cells by analytical subcellular fractionation combined with a sensitive computational method for data analysis.

A simple method of analytical subcellular fractionation, combined with a sensitive computational method for data analysis and presentation, has been used to reinvestigate the distribution and relative amounts of several enzymes in the cytoplasmic and plasma membranes of two different cell types: one is a neoplastic, transformed cell type (Ehrlich ascites tumour cells), the other an untransformed, highly differentiated cell type (liver hepatocytes plus Kupffer and endothelial cells). In general the distribution of the enzymes in particular membranes is similar in the two cell types, however the relative amounts differ. Ehrlich ascites tumour cells have a higher specific activity of galactosyltransferase and ouabain-sensitive (Na,K)ATPase, while liver cells have higher glucose-6-phosphatase, 5'-nucleotidase and succinate dehydrogenase activity. These differences appear to be correlated with morphological and, in some cases, functional differences between the two cell types.

5'-Nucleotidase↗

Seeing the forest despite the trees. The benefit of exploratory data analysis to program evaluation research.

In the present article, it is argued that there is a benefit to applying techniques of exploratory data analysis (EDA) to program evaluation. To exemplify this, an evaluation of a rehabilitation program for people with rheumatoid arthritis is presented. The perceived health status of patients receiving intensive rehabilitation services from a major rehabilitation institute was compared with that of patients receiving customary office-based care over an 18-month period. The data were analyzed in a conventional way (analysis of variance) and then by way of EDA techniques (graphic display of medians and boxplots). The conventional analysis suggested that all patients improved over time and that intensive rehabilitation services provided no particular benefit or harm. The exploratory analysis showed that the distribution of the outcome variable was patently nonnormal, thus casting doubt on the validity of the conventional analysis. The EDA further showed that the rehabilitation group lagged behind the comparison group for a year, with a precipitious improvement at the 18-month period. This suggests that a selection factor was operating (i.e., those in the rehabilitation group could have been sicker) or that the patients in the rehabilitation group were made more aware of their condition by the intensive health services they received. The EDA provided an important insight.

Analysis of Variance↗

A simplified method of echocardiographic data analysis.

Rapid accurate analysis of echocardiographic data is accomplished using a sonic digitizer and programmable calculator. This method allows the echocardiographer to select technically optimal areas of the recording for analysis. The resolution of the measuring device is 0.1 mm. A hardcopy printout of both measurement and calculation is provided. Instead of expensive on-line computer, an inexpensive programmable calculator is used.

Computers↗

A consultation system constructor for medical data analysis.

MAD is a system that helps an expert data analyst in a specific application domain (like epidemiology or image analysis) to build reasoning models aimed at fulfilling specific tasks. These models may be subsequently used to guide doctors in the analysis of a set of data referring to a specific ground domain. Expert knowledge is represented at various levels: a general description of an application domain and various models that formalize the reasoning followed to perform specific tasks within a defined application domain. Reasoning models are represented as rules of propositional calculus, and a meta-knowledge permits to support knowledge acquisition. During the consultation, different external programs may be run when needed, without the doctor having to learn how to use them. MAD is written in Golden Common LISP and may be linked to any external software for data analysis, provided it runs under MS-DOS and does not require more than 192 Kb. Examples of application of the system to epidemiology and image analysis are given.

Computer Simulation↗

A pooled data analysis on the use of intermittent cyclical etidronate therapy for the prevention and treatment of corticosteroid induced bone loss.

OBJECTIVE: To conduct a pooled data analysis in a group of patients defined by sex, menopausal status, and underlying disease in order to examine the effect of intermittent cyclical etidronate in the prevention and treatment of corticosteroid induced osteoporosis. METHODS: We selected 5 randomized, placebo controlled studies that examined the efficacy of intermittent cyclical etidronate therapy in which the raw data were available for analysis. Three were prevention studies and 2 treatment studies. The primary outcome was the difference between treatment groups in the percentage change from baseline in lumbar spine bone density. Secondary outcomes included the difference between treatment groups in the percentage change from baseline in femoral neck and trochanter bone density, and vertebral fracture rates. RESULTS: Results are separately pooled for the prevention and treatment studies. The prevention studies had significant mean differences (95% CI) between groups in mean percentage change from baseline in lumbar spine, femoral neck, and trochanter bone density of 3.7 (2.6 to 4.7), 1.7 (0.4 to 2.9), and 2.8% (1.3 to 4.2) after one year of treatment, in favor of the etidronate group. The treatment studies displayed a mean difference between groups in mean percentage change from baseline in lumbar spine bone density of 4.8 (2.7 to 6.9) and 5.4% (2.5 to 8.4) after one and 2 years of therapy. In the prevention studies, a reduced fracture incidence was observed in the etidronate group compared with the placebo group (relative risk 0.50; CI 0.21 to 1.19). CONCLUSION: Etidronate therapy was effective in preventing bone loss in the prevention studies and in preventing or slightly increasing bone mass in the treatment studies. A fracture benefit was observed in postmenopausal women treated with etidronate in the prevention studies.

Adult↗

Proper multivariate conditional autoregressive models for spatial data analysis.

In the past decade conditional autoregressive modelling specifications have found considerable application for the analysis of spatial data. Nearly all of this work is done in the univariate case and employs an improper specification. Our contribution here is to move to multivariate conditional autoregressive models and to provide rich, flexible classes which yield proper distributions. Our approach is to introduce spatial autoregression parameters. We first clarify what classes can be developed from the family of Mardia (1988) and contrast with recent work of Kim et al. (2000). We then present a novel parametric linear transformation which provides an extension with attractive interpretation. We propose to employ these models as specifications for second-stage spatial effects in hierarchical models. Two applications are discussed; one for the two-dimensional case modelling spatial patterns of child growth, the other for a four-dimensional situation modelling spatial variation in HLA-B allele frequencies. In each case, full Bayesian inference is carried out using Markov chain Monte Carlo simulation.

Alleles↗