PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Is mixed effects modeling or naïve pooled data analysis preferred for the interpretation of single sample per subject toxicokinetic data?

The purpose of this study was to evaluate whether mixed effects modeling (MEM) performs better than either noncompartmental or compartmental naïve pooled data (NPD) analysis for the interpretation of single sample per subject pharmacokinetic (PK) data. Using PK parameters determined during a toxicokinetic study in rats, we simulated data sets that might emerge from similar experiments. Data sets were simulated with varying numbers of animals at each sampling time (4-48) and the number of samples taken (1-3) from each individual. Each data set was replicated 50 times and analyzed using several variations of MEM that differed in the assumptions made regarding intraindividual error, NPD, and a graphical noncompartmental method. These analyses attempted to retrieve the underlying parameter and covariate effect values. We compared these analysis methods with respect to how well the underlying values were retrieved. All analysis methods performed poorly with single sample per subject data but MEM gave less biased estimates under the simulated conditions used here. MEM performance increased when covariate effects were sought in the analysis compared with analyses seeking only PK parameters. Decreasing the number of animals used per sampling time from 48 to 16 did not influence the quality of parameter estimates but further reductions (< 16 animals per sampling time) resulted in a reduced proportion of acceptable estimates. Parameter estimate quality improved and worsened with MEM and NPD, respectively, when additional samples were obtained from each individual. Assumptions made regarding the magnitude of intraindividual error were unimportant with single sample per subject data but influenced parameter estimates if more samples were obtained from each individual. MEM is preferable to both NPD and noncompartmental approaches for the analysis of single sample per subject data but even with MEM estimates of clearance are often biased.

Animals↗

Data analysis in behavioral cerebral blood flow activation studies using xenon-133 clearance.

BACKGROUND AND PURPOSE: Three mainstream strategies exist to detect the responses of regional cerebral blood flow to functional activation. We tested the significance of changes in raw regional cerebral blood flow data, regional cerebral blood flow data normalized by division by global cerebral blood flow (dependent model of the regional-to-global cerebral blood flow relation), and regional cerebral blood flow data treating global cerebral blood flow as a covariate (independent model). Both latter models attempt to enhance regional sensitivity by removing global effects. We examined the sensitivity and pitfalls of these three strategies in behavioral activation studies. METHODS: These three strategies of data analysis were applied to changes in regional cerebral blood flow induced by a visuospatial problem-solving task in 38 healthy subjects as measured by the intravenous xenon-133 method with 32 stationary detectors. RESULTS: Mental activation increased blood flow in all regions of interest. Raw data were most sensitive and reliable to detect responses to mental stimulation. Both the independent and dependent models to remove global effects were less sensitive and falsely indicated deactivation in regions that were clearly stimulated. CONCLUSIONS: In behavioral activation paradigms, safe data analysis should be restricted to using raw regional cerebral blood flow increases without normalization or separation of global from regional effects. Studies using complex stimulation tasks should be scrutinized for global cerebral blood flow effects confounding regional responses.

Behavior↗

A visual data analysis system for the medical image processing.

We developed a visual data analysis system that can easily manage a large volume of medical imaging data. This system can analyze sets of imaging data using general image processing methods, so that various kinds of medical imaging data such as ECG charts, X ray image films, and MRI images, can be processed. The system has a graphical user interface (GUI). A physician who is novice at the system can manipulate the imaging data intuitively by pull down menus, pop up menus and buttons within the window system. The system can run on a standard UNIX workstation which is faster and more powerful than most personal computers. The system needs an X window system/Motif and C compiler. These are standard system programs already available on most UNIX workstations. The source code of the system can be retrieved from our anonymous ftp site via Internet.

Computer Graphics↗

Effects of resolution reduction on data analysis.

BACKGROUND: There is often a need in flow cytometry to display and analyze histograms at resolutions lower than those native to the data. It is common, for example, to analyze DNA histograms at 256-channel resolution, even though the data were acquired at 1,024 channels or more. The most common method for reducing resolution, referred to as the consecutive summation (CS) method, can introduce distortions into the shape of histograms. Peaks that were symmetric in the original data can become skewed in the reduced-resolution histogram. Data analysis can be negatively affected by the distortions produced by reducing the histogram resolution. An alternative technique for reducing histogram resolution, the unbiased summation (US) method, minimizes shape distortion. This paper describes the US method and examines the benefits it provides in the analysis of DNA histograms. METHODS: Reduced chi-square (RCS) was used to measure the response to three experimental variables in the least-squares analysis of simulated DNA histograms. For each variable (the percentage of coefficient of variation [%CV], number of events, and mean position of the G1 distribution), a test data set of 1,000 histograms was generated at 1,024-channel resolution. Histogram resolutions were reduced with each method and then analyzed with ModFit LT cell-cycle analysis software (Verity Software House, Topsham, ME). S-phase error and processor computation time of each method also were evaluated. A Monte Carlo experiment was performed to compare CS and US methods to theoretically correct reductions. RESULTS: CS method analysis results were negatively affected by changes in %CV, number of events, and G1 peak position. The US method produced consistently lower RCS values (more accurate results) within the tested ranges. The US method eliminated bias in S-phase error and had negligible impact on analysis processing speed. It improved RCS values 44.50% on average (P < 0.0002) with actual DNA histograms. Whereas the CS method became less accurate (chi-square test) as the amount of reduction increased, the US method was unaffected, producing consistently better results. CONCLUSIONS: The US method is recommended for reducing histogram resolution in modeling applications such as DNA cell-cycle analysis. It may have implications in other areas of flow cytometric data analysis.

Algorithms↗

Evaluation of replication studies, combined data analysis, and analytical methods in complex diseases.

Due to genetic heterogeneity, phenocopies, incomplete penetrance, misdiagnosis, and unknown mode of inheritance, linkage studies of most complex diseases are unlikely to provide conclusive findings with unambiguously high lod scores. Typically, several marginally significant lod scores or elevated lod scores are observed in a genome-wide screen. However, it is usually difficult to differentiate these findings from false positives (type I errors). Two approaches are commonly used to guard against false positives: replication studies in independent samples and combined data analysis. In the current paper, we evaluated these two common approaches using simulated data where data from multiple groups were available and locations of disease genes were known. We found replication studies and combined data analysis performed similarly in terms of their ability to identify true and false positive linkages. Both approaches confirmed two true linkages and did not confirm any false positive linkages. The results also indicated that it is not appropriate to apply the criteria proposed for confirming significant evidence for linkage to confirm regions with only suggestive evidence for linkage. The current results support previous findings that parametric analysis using an incorrect genetic model can still identify a true linkage.

Environment↗

CARMAweb: comprehensive R- and bioconductor-based web service for microarray data analysis.

CARMAweb (Comprehensive R-based Microarray Analysis web service) is a web application designed for the analysis of microarray data. CARMAweb performs data preprocessing (background correction, quality control and normalization), detection of differentially expressed genes, cluster analysis, dimension reduction and visualization, classification, and Gene Ontology-term analysis. This web application accepts raw data from a variety of imaging software tools for the most widely used microarray platforms: Affymetrix GeneChips, spotted two-color microarrays and Applied Biosystems (ABI) microarrays. R and packages from the Bioconductor project are used as an analytical engine in combination with the R function Sweave, which allows automatic generation of analysis reports. These report files contain all R commands used to perform the analysis and guarantee therefore a maximum transparency and reproducibility for each analysis. The web application is implemented in Java based on the latest J2EE (Java 2 Enterprise Edition) software technology. CARMAweb is freely available at https://carmaweb.genome.tugraz.at.

Cluster Analysis↗

Temperature data analysis for 22 patients with advanced cervical carcinoma treated in Rotterdam using radiotherapy, hyperthermia and chemotherapy: a reference point is needed.

INTRODUCTION: The growing interest and participation in multi-institutional trials involving deep hyperthermia treatment is an important step towards the further consolidation of hyperthermia as an oncological treatment modality. However, the differences in the clinical procedures of hyperthermia application also raises questions as how to compare the reported temperatures data obtained by the different institutes. In this study our recent developed approach, RHyThM (Rotterdam Hyperthermia Thermal Modulator), has been used for thermal data analysis to investigate the temperature dynamics behaviour of a series of deep hyperthermia treatments. PATIENTS AND METHODS: All 22 patients (104 hyperthermia treatments) with locally advanced cervical carcinoma who participated in a feasibility study for treatment with a three-modality therapy were selected. The patients received mega-voltage external beam radiotherapy to the pelvis in daily fractions of 2 Gy five times a week to a total dose of 46 Gy and additional brachytherapy, at least four courses of weekly cisplatin (40 mg m-2) and five sessions of weekly loco regional deep hyperthermia treatments with the BSD2000-3D with the Sigma 60 or the Sigma-eye applicators at frequencies 70-120 MHz. Using RHyThM tissue type was defined along the insertion length, based on the CT scan information in radiotherapy position, for each single treatment. A step change in the slope of the profile of the first temperature map was identified to verify the insertion length of the thermometry catheter and precise location of the transition between in- and outside the body. Data analysis was performed based on the temperature readout provided by RHyThM. RESULTS: The temperature and RF-power data of 97 treatments could be analysed. The intra-vaginal temperature indices were slightly lower than those for bladder and rectum. The average T50 (median temperature) in all lumens, i.e. bladder, vagina and rectum, was 40.4 +/- 0.6 degrees Celsius. The average vagina all lumen T50 was 40.0 +/- 0.8 degrees Celsius. The average bladder and rectum all lumen T50 was 40.6 +/- 0.7 degrees Celsius and 40.5 +/- 0.6, respectively. When the analysis was restricted to the deepest 5 cm of the vagina lumen, the average T50 was 39.8 +/- 0.9 degrees Celsius. Good correlation exists between the various temperature indices like T20, T50 and T90, for all lumen measurements in bladder, vagina and rectum. No correlation was found between temperature indices and treatment number. For the complete patient population, no relationship was found between T50 and net integrated RF-power applied. In an explorative analysis on individual patients a positive correlation coefficient or trend was found in 14 patients between normalized net integrated RF-power and vagina T50. CONCLUSION: Average all lumen T50 for bladder, vagina and rectum differ less than 1 degrees Celsius, indicating that a large volume was heated relatively homogeneously. The vagina T50 value depends on how many measurement points are included for the analysis. In this group of patients the vagina T50 of the first treatment is not a good measure to discriminate between patients with 'heatable' and 'non-heatable' tumours. In order to compare temperature data reported by different institutes dealing with the same group of patients, one needs a strict and clear agreement on which temperature measurements or reference point(s) that should be included in the analysis.

Adult↗

[Comparison of software programs for data analysis of complex surveys].

OBJECTIVE: To compare specific software programs for data analysis of complex surveys regarding the following characteristics: ease of application, computer efficiency and accuracy of the results. METHODS: Secondary data from the Pesquisa Nacional sobre Demografia e Saúde (National survey on demography and health) (1996) with a target population of women aged 15 to 49 years old were used. This was a probabilistic subsampling drawn in two stages, then stratified, with the probability proportional to size in the first stage. The northern and mid-western regions of the country were selected for the study. The parameters of interest were mean for the age variable, and the proportion for five other qualitative variables. The software programs used were Epi Info, Stata and WesVarPC. RESULTS: The programs have two common options for the files import: the dBASE and text type files. The number of steps previous to the execution of the analyses were twenty- one for Epi Info, eleven for Stata and nine for WesVarPC. Efficiency was high for all them, that is, less that three seconds. The standard errors estimated using Epi Info and Stata were the same, with approximation up to the third decimal; those for WesVarPC were generally higher. CONCLUSIONS: Epi Info is the most limited software program regarding the analyses currently performed; however it is easy to use and free. Stata and WesVarPC are far more complete, however the disadvantage is their cost. The choice of the software program will depend mainly on the user's specific needs.

Adolescent↗

Augmented kurtosis-based projection pursuit: a novel, advanced machine learning approach for multi-omics data analysis and integration.

Due to the heterogeneity of multi-omics data, exacting their maximum information potential remains a challenge. Whereas some solutions have been offered, most cannot overcome the large linear dynamic range associated with such data, while others require large biological effect sizes to produce meaningful models. Here, we (i)&#xa0;perform a comprehensive benchmarking of multi-omics data analysis tools, and (ii)&#xa0;introduce kurtosis-based projection pursuit analysis, augmented with classification and regression trees (kPPA-CART) as a robust, easy-to-implement alternative. Using ground truth data, we demonstrate that kPPA-CART exhibits superiority in inferring biological significance from low-intensity (low-count) features and studies with small biological effect sizes. Applying it to experimental breast cancer data from The Cancer Genome Atlas, we identify novel genes that cluster the samples into subtypes that mimic the canonical PAM50 classes with notable improvements. Validating with external metastatic breast cancer data from the AURORA US consortium, kPPA-CART identifies genes that are associated with poor event-free survival and additional clustering associated with increased tumor mutational burden. Finally, we provide an R package and an online implementation of kPPA-CART.

Humans↗

Estimating fertility potential via semen analysis data.

The aim of this study was to evaluate diagnostic profiles for the assessment of semen analysis data with respect to male fertility potential. Semen samples taken from 208 patients of known fertility and suspected infertility were studied by conventional semen analysis methods. The data throw doubt upon the validity of an approach based on the number of deviations from the normal standard values defined by the World Health Organization. The alternative approach of a specific semen characteristic (particularly morphology) as the major predictor of fertility produced no beneficial results. However, the semen analysis index based on semen volume, sperm count, percentage motility and normal forms resulted in a high accuracy of classification but for only 44% of the cases, with 3% false negatives and 10% false positives using cut-off indices of > or = 0.6 and < or = -1.0 for defining 'fertile' and 'infertile' zones, respectively. In conclusion, it is emphasized that there are a number of specific semen analysis variables, each expressing a different aspect of male fertility potential which, when combined in correct proportion, do provide the optimal evaluation of the male fertility status. However, in order to increase the prognostic potential of the semen sample, new and meaningful parameters must be discovered.

Adult↗

Risk behavior data analysis: ordinal or dichotomous the choice is yours.

OBJECTIVE: To demonstrate the differences of 2 approaches to data analysis. METHODS: Using the South Carolina YRBS data, study focused on contingency tables and ANOVA. Additive chi squares are utilized to illustrate information loss when collapsing a contingency table. Odds ratios are derived from contingency tables or logistic regression. Means are utilized in ANOVA. Five measures of life satisfaction were summed to create a pseudo-continuous response variable that was subsequently trichotomized. All predictors are dichotomized risk variables. RESULTS: Chi squares from subtables added exactly to that of the original table measuring lost information. ANOVA conveyed the same clinical message. CONCLUSION: Clinically relevant conclusions might be the same even when drawn from any of several different analyses of the same risk-behavior data.

Adolescent↗

A high-throughput urinalysis of abused drugs based on a SPE-LC-MS/MS method coupled with an in-house developed post-analysis data treatment system.

A rapid urinalysis system based on SPE-LC-MS/MS with an in-house post-analysis data management system has been developed for the simultaneous identification and semi-quantitation of opiates (morphine, codeine), methadone, amphetamines (amphetamine, methylamphetamine (MA), 3,4-methylenedioxyamphetamine (MDA) and 3,4-methylenedioxymethamphetamine (MDMA)), 11-benzodiazepines or their metabolites and ketamine. The urine samples are subjected to automated solid phase extraction prior to analysis by LC-MS (Finnigan Surveyor LC connected to a Finnigan LCQ Advantage) fitted with an Alltech Rocket Platinum EPS C-18 column. With a single point calibration at the cut-off concentration for each analyte, simultaneous identification and semi-quantitation for the above mentioned drugs can be achieved in a 10 min run per urine sample. A computer macro-program package was developed to automatically retrieve appropriate data from the analytical data files, compare results with preset values (such as cut-off concentrations, MS matching scores) of each drug being analyzed and generate user-defined Excel reports to indicate all positive and negative results in batch-wise manner for ease of checking. The final analytical results are automatically copied into an Access database for report generation purposes. Through the use of automation in sample preparation, simultaneous identification and semi-quantitation by LC-MS/MS and a tailored made post-analysis data management system, this new urinalysis system significantly improves the quality of results, reduces the post-data treatment time, error due to data transfer and is suitable for high-throughput laboratory in batch-wise operation.

Automation↗

Novel data analysis for synchronised spontaneous neuromagnetic activity.

A novel approach to neuromagnetic data analysis is presented. This technique is aimed at studying synchronised spontaneous activity (SSA) and has been used to resolve two different signals from one single evoked response, providing evidence for two possibly distinct sources. The data presented are consistent with a model that permits the generators of spontaneous activity to be synchronised by sensory stimuli.

Brain↗

Focus-group interview and data analysis.

In recent years focus-group interviews, as a means of qualitative data collection, have gained popularity amongst professionals within the health and social care arena. Despite this popularity, analysing qualitative data, particularly focus-group interviews, poses a challenge to most practitioner researchers. The present paper responds to the needs expressed by public health nutritionists, community dietitians and health development specialists following two training sessions organised collaboratively by the Health Development Agency, the Nutrition Society and the British Dietetic Association in 2003. The focus of the present paper is on the concepts and application of framework analysis, especially the use of Krueger's framework. It provides some practical steps for the analysis of individual data, as well as focus-group data using examples from the author's own research, in such a way as to assist the newcomer to qualitative research to engage with the methodology. Thus, it complements the papers by Draper (2004) and Fade (2004) that discuss in detail the complementary role of qualitative data in researching human behaviours, feelings and attitudes. Draper (2004) has provided theoretical and philosophical bases for qualitative data analysis. Fade (2004) has described interpretative phenomenology analysis as a method of analysing individual interview data. The present paper, using framework analysis concentrating on focus-group interviews, provides another approach to qualitative data analysis.

Data Collection↗

Distributed intelligent data analysis in diabetic patient management.

This paper outlines the methodologies that can be used to perform an intelligent analysis of diabetic patients' data, realized in a distributed management context. We present a decision-support system architecture based on two modules, a Patient Unit and a Medical Unit, connected by telecommunication services. We stress the necessity to resort to temporal abstraction techniques, combined with time series analysis, in order to provide useful advice to patients; finally, we outline how data analysis and interpretation can be cooperatively performed by the two modules.

Computer Communication Networks↗

Intelligent data analysis to interpret major risk factors for diabetic patients with and without ischemic stroke in a small population.

This study proposes an intelligent data analysis approach to investigate and interpret the distinctive factors of diabetes mellitus patients with and without ischemic (non-embolic type) stroke in a small population. The database consists of a total of 16 features collected from 44 diabetic patients. Features include age, gender, duration of diabetes, cholesterol, high density lipoprotein, triglyceride levels, neuropathy, nephropathy, retinopathy, peripheral vascular disease, myocardial infarction rate, glucose level, medication and blood pressure. Metric and non-metric features are distinguished. First, the mean and covariance of the data are estimated and the correlated components are observed. Second, major components are extracted by principal component analysis. Finally, as common examples of local and global classification approach, a k-nearest neighbor and a high-degree polynomial classifier such as multilayer perceptron are employed for classification with all the components and major components case. Macrovascular changes emerged as the principal distinctive factors of ischemic-stroke in diabetes mellitus. Microvascular changes were generally ineffective discriminators. Recommendations were made according to the rules of evidence-based medicine. Briefly, this case study, based on a small population, supports theories of stroke in diabetes mellitus patients and also concludes that the use of intelligent data analysis improves personalized preventive intervention.

Brain Infarction↗

Clustering binary fingerprint vectors with missing values for DNA array data analysis.

Oligonucleotide fingerprinting is a powerful DNA array based method to characterize cDNA and ribosomal RNA gene (rDNA) libraries and has many applications including gene expression profiling and DNA clone classification. We are especially interested in the latter application. A key step in the method is the cluster analysis of fingerprint data obtained from DNA array hybridization experiments. Most of the existing approaches to clustering use (normalized) real intensity values and thus do not treat positive and negative hybridization signals equally (positive signals are much more emphasized). In this paper, we consider a discrete approach. Fingerprint data are first normalized and binarized using control DNA clones. Because there may exist unresolved (or missing) values in this binarization process, we formulate the clustering of (binary) oligonucleotide fingerprints as a combinatorial optimization problem that attempts to identify clusters and resolve the missing values in the fingerprints simultaneously. We study the computational complexity of this clustering problem and a natural parameterized version, and present an efficient greedy algorithm based on MINIMUM CLIQUE PARTITION on graphs. The algorithm takes advantage of some unique properties of the graphs considered here, which allow us to efficiently find the maximum cliques as well as some special maximal cliques. Our experimental results on simulated and real data demonstrate that the algorithm runs faster and performs better than some popular hierarchical and graph-based clustering methods. The results on real data from DNA clone classification also suggest that this discrete approach is more accurate than clustering methods based on real intensity values, in terms of separating clones that have different characteristics with respect to the given oligonucleotide probes.

Algorithms↗

Making sense of qualitative data analysis: an introduction with illustrations from DIPEx (personal experiences of health and illness).

OBJECTIVES: This paper outlines an approach to analysing qualitative textual data from interviews and discusses how to ensure analytic procedures are appropriately rigorous. OVERVIEW: Qualitative data analysis should begin at an early stage in data collection and be highly systematic. It is important to identify issues that emerge during the data collection and analysis as well as those that the researcher may have anticipated (from reading or experience). Analysis is very time-consuming, but careful sampling, the collection of rich material and analytic depth mean that a relatively small number of cases can generate insights that apply well beyond the confines of the study. One particular approach to thematic analysis is introduced with examples from the DIPEx (personal experiences of health and illness) project, which collects video- and audio-taped interviews that are freely accessible through http://www.dipex.org. EVALUATION: Qualitative analysis of patients' perspectives of illness can illuminate numerous issues that are important for medical education, some of which are unlikely to arise in the clinical encounter. Qualitative studies can also cover a much broader range of experiences - of both common and rare disease - than clinicians will see in practice. The DIPEx website is based on qualitative analysis of collections of interviews, illustrated with hundreds of video and audio clips, and is an innovative resource for medical education.

Attitude to Health↗