PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Implementing a training resource for large-scale genomic data analysis in the All of Us Researcher Workbench.

A lack of representation in genomic research and limited access to computational training create barriers for many researchers seeking to analyze large-scale genetic datasets. The All of Us Research Program provides an unprecedented opportunity to address these gaps by offering genomic data from a broad range of participants, but its impact depends on equipping researchers with the necessary skills to use it effectively. The All of Us Biomedical Researcher (BR) Scholars Program at Baylor College of Medicine aims to break down these barriers by providing early-career researchers with hands-on training in computational genomics through the All of Us Evenings with Genetics Research Program. The year-long program begins with the faculty summit, an in-person computational boot camp that introduces scholars to foundational skills for using the All of Us dataset via a cloud-based research environment. The genomics tutorials focus on genome-wide association studies (GWASs), utilizing Jupyter Notebooks and the Hail computing framework to provide an accessible and scalable approach to large-scale data analysis. Scholars engage in hands-on exercises covering data preparation, quality control, association testing, and result interpretation. By the end of the summit, participants will have successfully conducted a GWAS, visualized key findings, and gained confidence in computational resource management. This initiative expands access to genomic research by equipping early-career researchers from a variety of backgrounds with the tools and knowledge to analyze All of Us data. By lowering barriers to entry and promoting the study of representative populations, the program fosters innovation in precision medicine and advances equity in genomic research.

Humans↗

Using an integrated software package for clinical data analysis on a microcomputer.

An integrated software package was used effectively for entering, organizing and analyzing clinical research data on a microcomputer. Both the database and the spreadsheet components of the package were used in the process. The database component enabled a form to be created for entering the data. The spreadsheet component was used in the organization and analysis of data. Macros were written within the spreadsheet environment for the statistical analysis of data. The purpose of this paper is to illustrate how an integrated software like Symphony could offer features beyond simply the use of a spreadsheet for the analysis of research data. Also highlighted are the other useful features of the integrated software that are not directly related to data analysis.

Adolescent↗

The meaning of diagnostic test results: a spreadsheet for swift data analysis.

AIMS: To design a spreadsheet program to: (a) analyse rapidly diagnostic test result data produced in local research or reported in the literature; (b) correct reported predictive values for disease prevalence in any population; (c) estimate the post-test probability of disease in individual patients. MATERIALS AND METHODS: Microsoft Excel(TM)was used. Section A: a contingency (2 x 2) table was incorporated into the spreadsheet. Formulae for standard calculations [sample size, disease prevalence, sensitivity and specificity with 95% confidence intervals, predictive values and likelihood ratios (LRs)] were linked to this table. The results change automatically when the data in the true or false negative and positive cells are changed. Section B: this estimates predictive values in any population, compensating for altered disease prevalence. Sections C-F: Bayes' theorem was incorporated to generate individual post-test probabilities. The spreadsheet generates 95% confidence intervals, LRs and a table and graph of conditional probabilities once the sensitivity and specificity of the test are entered. The latter shows the expected post-test probability of disease for any pre-test probability when a test of known sensitivity and specificity is positive or negative. RESULTS: This spreadsheet can be used on desktop and palmtop computers. The MS Excel(TM)version can be downloaded via the Internet from the URL ftp://radiography.com/pub/Rad-data99.xls CONCLUSION: A spreadsheet is useful for contingency table data analysis and assessment of the clinical meaning of diagnostic test results.

Bayes Theorem↗

A new approach to near-infrared spectral data analysis using independent component analysis.

This paper presents a new approach to near-infrared spectral (NIR) data analysis that is based on independent component analysis (ICA). The main advantage of the new method is that it is able to separate the spectra of the constituent components from the spectra of their mixtures. The separation is a blind operation, since the constituent components of mixtures can be unknown. The ICA based method is therefore particularly useful in identifying the unknown components in a mixture as well as in estimating their concentrations. The approach is introduced by reference to case studies and compared to other techniques for NIR analysis including principal component regression (PCR), multiple linear regression (MLR), and partial least squares (PLS) as well as Fourier and wavelet transforms.

Adipose Tissue↗

Microarray data analysis: a practical approach for selecting differentially expressed genes.

BACKGROUND: The biomedical community is rapidly developing new methods of data analysis for microarray experiments, with the goal of establishing new standards to objectively process the massive datasets produced from functional genomic experiments. Each microarray experiment measures thousands of genes simultaneously producing an unprecedented amount of biological information across increasingly numerous experiments; however, in general, only a very small percentage of the genes present on any given array are identified as differentially regulated. The challenge then is to process this information objectively and efficiently in order to obtain knowledge of the biological system under study and by which to compare information gained across multiple experiments. In this context, systematic and objective mathematical approaches, which are simple to apply across a large number of experimental designs, become fundamental to correctly handle the mass of data and to understand the true complexity of the biological systems under study. RESULTS: The present report develops a method of extracting differentially expressed genes across any number of experimental samples by first evaluating the maximum fold change (FC) across all experimental parameters and across the entire range of absolute expression levels. The model developed works by first evaluating the FC across the entire range of absolute expression levels in any number of experimental conditions. The selection of those genes within the top X% of highest FCs observed within absolute expression bins was evaluated both with and without the use of replicates. Lastly, the FC model was validated by both real time polymerase chain reaction (RT-PCR) and variance data. Semi-quantitative RT-PCR analysis demonstrated 73% concordance with the microarray data from Mu11K Affymetrix GeneChips. Furthermore, 94.1% of those genes selected by the 5% FC model were found to lie above measurement variability using a SDwithin confidence level of 99.9%. CONCLUSION: As evidenced by the high rate of validation, the FC model has the potential to minimize the number of required replicates in expensive microarray experiments by extracting information on gene expression patterns (e.g. characterizing biological and/or measurement variance) within an experiment. The simplicity of the overall process allows the analyst to easily select model limits which best describe the data. The genes selected by this process can be compared between experiments and are shown to objectively extract information which is biologically & statistically significant.

Animals↗

Examination of dioxin fluxes recorded in dated aquatic-sediment cores in the Kanto region of Japan using multivariate data analysis.

Past dioxin (coplanar polychlorinated biphenyl (Co-PCB), 2,3.7,8-substituted polychlorinated dibenzo-p-dioxin (PCDD) and 2,3,7,8-substituted polychlorinated dibenzofuran (PCDF)) fluxes recorded in dated aquatic-sediment cores were analyzed using principal component analysis (PCA). The data set consisted of samples from four cores collected from the Kanto region of Japan. Time trends and spatial differences in the dioxin flux were examined, and the potential relationship to emission sources was investigated. Twenty-five compounds and 58 core slices, corresponding to the later half of the 20th century, were subjected to the analysis. The PCA of both log-transformed and maximum-value-standardized data successfully divided the dioxin compounds into a small number of groups, and three similar clusters of Co-PCBs. PCDDs and penta- to hepta-CDFs were identified. PCB formulations used in the past are judged to have been responsible for the major part of the Co-PCB flux recorded in the sediment cores. However, the relationship to emission sources needs further investigation. It is suggested that most 2,3,7,8-substituted PCDDs and PCDFs are different from Co-PCBs in their emission sources or movements in the environment. The subcore clusters obtained from the PCA of log-transformed data show that the cores from different sampling areas exhibited distinct dioxin fluxes and compositions. Common time trends among the cores were more effectively summarized by the PCA of maximum-value-standardized data focusing on relative time trends. PC scores show that recently the flux of each dioxin compound in the four cores has been generally declining after having reached a peak.

Benzofurans↗

Proficiency of the Tradescantia-micronucleus image analysis system for scoring micronucleus frequencies and data analysis.

The Tradescantia-micronucleus (Trad-MCN) bioassay is an efficient short-term test for genotoxicity of pollutants. In order to increase the efficiency and to standardize the micronucleus (MCN) scoring process, an automated scoring system was developed using the principle of image analysis in computer science. This assemblage is called the Tradescantia-micronucleus image analysis (Trad-MCNIA) system. The MCN frequencies scored by this system were compared with those scored by human observation for its proficiency. A set of low MCN frequency (around 5 MCN/100 tetrads) slides prepared from a control group, a set of medium MCN frequency (around 20 MCN/100 tetrads) slides prepared from sodium azide treated plant cuttings and a set of high MCN frequency (around 50 MCN/100 tetrads) slides prepared from X-ray treated materials were used for this study. In the low MCN frequency slides, the Trad-MCNIA system scored about the same value as human observation. In the medium and high frequency slides, MCN frequencies scored by the system were lower than those scored by human observers. This discrepancy was corrected by increasing the power of the objective of the microscope in the system. The MCN frequencies scored by the system attained 90% congruity with those scored by human observers after the correction. The scoring speed of the system was about 3.5 times as fast as that by human observers, and the data could be statistically analyzed immediately after the data scores were recorded. Further improvements can be made by upgrading the video camera and the computer speed.

Azides↗

Recruitment in NHLBI population-based studies and randomized clinical trials: data analysis and survey results.

Data on screening and recruitment from current and previous NHLBI population-based studies (PBSs) and randomized clinical trials (RCTs) were examined. In only two of the studies examined was the projected recruitment completed within the planned recruitment period. The shape of the graph of the relation between enrollment of participants and time varies by study. A single summary statistic for measuring the efficiency of recruitment in RCTs and PBSs is proposed and applied to the examined studies. In addition to providing summary data on recruitment for several studies, this article reports the survey results of a questionnaire sent to the coordinating centers of currently and previously funded National Heart, Lung, and Blood Institute and Veteran's Administration studies. The purpose was to ascertain the desirability of recommending that a generic core of information be collected on recruitment and screening in future studies. Most respondents believed that comparing data collected uniformly and prospectively might be helpful in designing further studies. The variables most respondents believed to be potentially useful are described.

Clinical Trials as Topic↗

Use of prior information to stabilize a population data analysis.

When modeling new data with a complex population pharmacokinetic/pharmacodynamic model, there may not be sufficient information to obtain estimates of all parameters. In this case information from previous studies can also be used to help stabilize estimation. Using simulated data, we explored three different ways to do this. (i) Some parameter values were fixed to estimates obtained from earlier data. (ii) The earlier data were combined with the current data. (iii) The objective function based on the current data was augmented by a penalty function expressing summary information obtained from the earlier data. This last method is similar to the use of a Bayesian prior. It may be particularly useful when either the combined data set of method (ii) is very large and leads to large computation times or when the early data are not readily available. With this method, two different types of penalty functions were used. With our examples, the three methods all resulted in stabilized estimation. Methods (ii) and (iii) gave similar results for parameter and standard error estimation, especially with respect to fixed effects parameters. For hypothesis testing, results obtained with method (i) are very problematic. There are also problems with the results obtained with method (iii), but they are much less severe, and when the design for the earlier data is known, they can be corrected by using a computer-intensive simulation test procedure.

Algorithms↗

[Genetic diversity of carrion and jungle crows from RAPD-PCR analysis data].

RAPD-PCR analysis of the genetic diversity of the carrion crow (Corvus corone) and jungle crow (C. macrorhynchos) living in the continental parts of their species ranges and on some Russian and Japanese Far Eastern islands has been performed. Taxon-specific molecular markers have been found for each species. The genetic diversity of the carrion crow is considerably less than that of the jungle crow at the same genetic distance (P95 = 68.2%, DN = 0.27 and P95 = 88.4%, DN = 0.24, respectively). In both species, the genetic polymorphism of island samples is almost two times greater than that of continental samples (62 and 31.8%, respectively, for C. corone and 81.5 and 47.2%, respectively, for C. macrorhynchos). In addition, differences in genetic diversity between males and females (P95 = 55.1 and P95 = 72.1, respectively) has been found in the carrion crow but not in the jungle crow. The gene diversity of C. macrorhynchos is greater than that of C. corone: the mean numbers of alleles per locus are 2 and 1.81, effective numbers of alleles are 1.62 and 1.43, and the mean expected heterozygosities are 0.39 and 0.30, respectively. The phenograms and phylograms significantly segregate the clusters of the carrion and jungle crows. The clustering patterns of carrion crows corresponds to the intraspecies taxonomic and geographic differentiation: subspecies C. c. corone and C.c. orientalis living in the western and eastern parts of the species range, respectively, form different subclusters. The cluster of the jungle crow does not exhibit differentiation into subspecies C. m. mandshuricus and C. m. japonensis; molecular genetic differences between them are small.

Animals↗

A procedure for data analysis of the rodent micronucleus test involving a historical control.

No standard procedure of data analysis for rodent micronucleus tests involving historical controls has been established. In the present paper, under the presumption that the distribution of the historical control is stable and reliable, a procedure with three statistical steps is proposed to analyze the frequency of micronucleated polychromatic erythrocytes (MNPCEs). In the first step, the frequencies of MNPCEs in negative and positive control groups of a current experiment of the micronucleus test are compared with the distribution of historical negative and positive controls to examine the technical validity of the current experiment. In the second step, the frequency of MNPCEs in each treatment group is compared with the distribution of the historical negative control. In the third step, the dose-response relation is tested with the Cochran-Armitage trend test. A Monte Carlo stimulation study shows that the power of this procedure is acceptable and also this procedure is robust. An application of this procedure on real data reveals that it is effective in detecting clastogenic chemicals when the probability of a type I error is nearly .01.

Animals↗

Exploratory data analysis of evoked response single trials based on minimal spanning tree.

OBJECTIVE: An exploratory data analysis framework, based on minimal spanning tree, is proposed as a means to support the analysis of single trial (ST) electrophysiological signals. The core of this framework is the compact description of the input ST sample in a form of content-dependent ordered lists. Based on the established hierarchies, efficient ways to increase the SNR, extract prototypical responses, visualize possible self-organization trends in the sample and track the course of evoked response along the trial-to-trial dimension, are proposed. METHOD: Magnetoencephalographic auditory evoked responses were used for demonstrating and validating the introduced framework. RESULTS AND CONCLUSION: The results demonstrate the benefits, from this intelligent manipulation of STs, in understanding and enhancing the actual evoked signal. Specifically we find support for stimulus-induced phase-resetting hypothesis in the 3-20 Hz band, the existence of trials void of the prototypical evoked response, and an order across the single trial set hinting at an underlying process with long time scale.

Adult↗

Zherlock: an open source data analysis software.

Zherlock is an open source software that provides state-of-the-art data analysis tools to the user in an intuitive and flexible way. It is a front-end to different numerical "engines" to produce a seamless integration of algorithms written in different computer languages. Of particular interest is creating an interface to high-level scientific languages such as Octave (a Matlab clone) and R (an S-PLUS clone) to enable efficient porting of new data analytical methods. Zherlock uses advanced scientific visualization tools in 2-D and 3-D and has been extended to work on virtual reality (VR) systems. Central to Zherlock is a visual programming environment (VPE) which enables diagram based programming. These diagrams consist of nodes and connection lines where each node is an operator or a method and lines describe the flow of data between nodes. A VPE is chosen for Zherlock because it forms an effective way to control the processing pipeline in complex data analyses. The VPE is similar in functionality to other programs such as IRIS Explorer, AVS or LabVIEW.

Algorithms↗

Secondary data analysis: research method for the clinical nurse specialist.

This article presents a description of secondary data analysis and suggests that this type of research methodology may be helpful in facilitating research by the clinical nurse specialist (CNS). The article discusses the advantages and disadvantages of the use of this method specifically in relation to the CNS and offers suggestions for sources of data.

Data Collection↗

Fundamentals of cDNA microarray data analysis.

Microarray technology is a powerful approach for genomics research. The multi-step, data-intensive nature of this technology has created an unprecedented informatics and analytical challenge. It is important to understand the crucial steps that can affect the outcome of the analysis. In this review, we provide an overview of the contemporary trend on various main analysis steps in the microarray data analysis process, which includes experimental design, data standardization, image acquisition and analysis, normalization, statistical significance inference, exploratory data analysis, class prediction and pathway analysis, as well as various considerations relevant to their implementation.

Animals↗

Extreme value distribution based gene selection criteria for discriminant microarray data analysis using logistic regression.

One important issue commonly encountered in the analysis of microarray data is to decide which and how many genes should be selected for further studies. For discriminant microarray data analyses based on statistical models, such as the logistic regression models, gene selection can be accomplished by a comparison of the maximum likelihood of the model given the real data, L(D|M), and the expected maximum likelihood of the model given an ensemble of surrogate data with randomly permuted label, L(D(0)|M). Typically, the computational burden for obtaining L(D(0)M) is immense, often exceeding the limits of available computing resources by orders of magnitude. Here, we propose an approach that circumvents such heavy computations by mapping the simulation problem to an extreme-value problem. We present the derivation of an asymptotic distribution of the extreme-value as well as its mean, median, and variance. Using this distribution, we propose two gene selection criteria, and we apply them to two microarray datasets and three classification tasks for illustration.

Chromosome Mapping↗

Techniques to identify clinical contexts during automated data analysis.

The interpretation of automatically collected data to produce intelligent alarms and identify particular conditions is nearly impossible without identifying the specific context in which the data are obtained. Shifts in clinical context occur because of changes in the patient's physiologic state, or due to the passage of time, or due to changes imposed by therapeutic intervention such as surgery. Techniques to identify such changes in clinical context are discussed with particular attention to the application of cluster analysis, discriminant analysis, and statistical predictors. An example of these analyses applied to EEG data is presented, showing an unexpected hysteresis of EEG behavior in response to an hypoxic challenge.

Cluster Analysis↗

Software for temporal gait data analysis.

This study presents a computer program, developed to support a low-cost, portable telemetry system that has been designed to assess footfall timing. This software frees the user from data processing and allows concentration on data analysis. The new technique has been applied with accuracy and reliability to the analysis of the gait of orthopedic patients, athletes, mountaineers, etc. The subroutines developed for data acquisition, storage and analysis are explained in detail, and an example is presented.

Data Interpretation, Statistical↗