PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Microarray data analysis and mining.

DNA microarray is an innovative technology for obtaining information on gene function. Because it is a high-throughput method, computational tools are essential in data analysis and mining to extract the knowledge from experimental results. Filtering procedures and statistical approaches are frequently combined to identify differentially expressed genes. However, obtaining a list of differentially expressed genes is only the starting point because an important step is the integration of differential expression profiles in a biological context, which is a hot topic in data mining. In this chapter an integrated approach of filtering and statistical validation to select trustable differentially expressed genes is described together with a brief introduction on data mining focusing on the classification of co-regulated genes on the basis of their biological function.

Cluster Analysis↗

[The effect of smoking habit on aortic pulse wave velocity using a new method for data analysis].

We measured aortic pulse wave velocity (PWV) in 168 male adult cases of various arteriosclerotic diseases. In order to evaluate the effects of age, smoking habits, alcohol intake, and blood pressure, we applied the least median of squares (LMS) regression which was considered to be very useful for data analysis. The results showed that PWV level increased with age. Furthermore smoking was associated with increasing PWV level and this effect was also related to age. We concluded that the PWV was valuable as an index of arteriosclerosis, and instead of the classical least squares method, LMS regression was very useful for analysis of medical data.

Adult↗

Secondary structure determination of proteins in aqueous solution by infrared spectroscopy: a comparison of multivariate data analysis methods.

The accuracy of the secondary structure prediction from an infrared spectra data base of 39 proteins with known X-ray structure was investigated by different methods of multivariate data analysis. The best agreements with the secondary structure determined by X-ray crystallography are obtained if both the amide I and amide II bands are used for calibration. With optimized parameters the methods singular value decomposition, partial least squares, and ridge regression yield similar results. As judged by the standard error of prediction, the secondary structure elements helix and beta-sheet can be predicted with the highest accuracy. Small data sets of less than 20 protein spectra, which exhibit the variance in secondary structure content of the whole set, can pretend an increased prediction accuracy only if column cross-validation is used as reference; however, with these calibration sets the average secondary structure prediction of all 39 proteins is debased. The hydrogen-bonded turns or bridges are predicted with higher accuracy than the assigned secondary structure types helix and beta-sheet.

Multivariate Analysis↗

Functional data analysis for gait curves study in Parkinson's disease.

In Parkinson's disease, precise analysis of gait disorders remains essential for the diagnostic or the evaluation of treatments. During a gait analysis session, a series of successive dynamic gait trials are recorded and data involves a set of continuous curves for each patient. An important aspect of such data is the infinite dimension of the space data belong. Therefore, classical multivariate statistical analysis are inadequate. Recent methods known as functional data analysis allow to deal with this kind of data. In this paper, we present a functional data analysis approach for solving two problems encountered in clinical practice: (1) for a given patient, assessing the reliability of the gait curves corresponding to the different trials (2) performing intra individual curves comparisons for assessing the effect of a therapy. In a first step, each discretized curve was interpolated using cubic B-splines bases in order to ensure the continuous character of data. A cluster analysis was performed on the smoothed curves to assess the reliability and to identify a subset of representative curves for a given patient. Intra individual curves comparisons were carried out in the following way: (1) functional principal component analysis was performed to describe the temporal structure of data and to derive a finite number of reliable principal components. (2) These principal components were used in a linear discriminant analysis to point out the differences between the curves. This procedure was applied to compare the gait curves of 12 parkinsonian patients under 4 therapeutic conditions. This study allowed us to develop objective criteria for measuring the improvements in a subject's gait and comparing the effect of different treatments. The methods presented in this paper could be used in other medical domains when data consist in continuous curves.

France↗

Data analysis methods for detection of differential protein expression in two-dimensional gel electrophoresis.

The recent development of microarray technology has led statisticians and bioinformaticians to develop new statistical methodologies for comparing different biological samples. The objective is to identify a small number of differentially expressed genes from among thousands. In quantitative proteomics, analysis of protein expression using two-dimensional gel electrophoresis shows some similarities with transcriptomic studies. Thus, the goal of this study was to evaluate different data analysis methodologies widely used in array analysis using different proteomic data sets of hundreds of proteins. Even with few replications, the significance analysis of microarrays method appeared to be more powerful than the Student's t test in truly declaring differentially expressed proteins. This procedure will avoid wasting time due to false positives and losing information with false negatives.

Animals↗

Diagnostic accuracy of pancreatic enzymes evaluated by use of multivariate data analysis.

We analyzed pancreatic enzyme data from 508 patients with suspected pancreatitis by neural network analysis, by an Expert multirule generation protocol, and by receiver-operator characteristic (ROC) curve analysis of a single test result. Neural network analysis showed that use of lipase provided the best means for diagnosing pancreatitis. Diagnostic accuracies achieved by using amylase only, lipase only, and amylase and lipase in combination were 76%, 82%, and 84%, respectively. Use of the Expert rule generation protocol provided a diagnostic accuracy of 92% when rules for single and multiple samplings were combined. ROC curve analysis for initial enzyme activities showed the maximal diagnostic accuracy to be 82% and 85% for amylase and lipase, respectively; use of peak enzyme activities yielded accuracies of 81% and 88%, respectively. The evaluation of laboratory test data should include analysis of the diagnostic accuracy of laboratory tests by multivariate techniques such as neural network analysis or an Expert systems approach. Multivariate analysis should allow for a more realistic assessment of the diagnosis accuracy of laboratory tests because all the available data are included in the evaluation.

Amylases↗

Replica-exchange Monte Carlo scheme for bayesian data analysis.

We develop a sampling algorithm to explore the probability densities arising in Bayesian data analysis problems. Our algorithm is a multiparameter generalization of a replica-exchange Monte Carlo scheme. The strategy relies on gradual weighing of experimental data and on Tsallis generalized statistics. We demonstrate the effectiveness of the method on nuclear magnetic resonance data for a folded protein.

Algorithms↗

Implementing a training resource for large-scale genomic data analysis in the All of Us Researcher Workbench.

A lack of representation in genomic research and limited access to computational training create barriers for many researchers seeking to analyze large-scale genetic datasets. The All of Us Research Program provides an unprecedented opportunity to address these gaps by offering genomic data from a broad range of participants, but its impact depends on equipping researchers with the necessary skills to use it effectively. The All of Us Biomedical Researcher (BR) Scholars Program at Baylor College of Medicine aims to break down these barriers by providing early-career researchers with hands-on training in computational genomics through the All of Us Evenings with Genetics Research Program. The year-long program begins with the faculty summit, an in-person computational boot camp that introduces scholars to foundational skills for using the All of Us dataset via a cloud-based research environment. The genomics tutorials focus on genome-wide association studies (GWASs), utilizing Jupyter Notebooks and the Hail computing framework to provide an accessible and scalable approach to large-scale data analysis. Scholars engage in hands-on exercises covering data preparation, quality control, association testing, and result interpretation. By the end of the summit, participants will have successfully conducted a GWAS, visualized key findings, and gained confidence in computational resource management. This initiative expands access to genomic research by equipping early-career researchers from a variety of backgrounds with the tools and knowledge to analyze All of Us data. By lowering barriers to entry and promoting the study of representative populations, the program fosters innovation in precision medicine and advances equity in genomic research.

Humans↗

Using an integrated software package for clinical data analysis on a microcomputer.

An integrated software package was used effectively for entering, organizing and analyzing clinical research data on a microcomputer. Both the database and the spreadsheet components of the package were used in the process. The database component enabled a form to be created for entering the data. The spreadsheet component was used in the organization and analysis of data. Macros were written within the spreadsheet environment for the statistical analysis of data. The purpose of this paper is to illustrate how an integrated software like Symphony could offer features beyond simply the use of a spreadsheet for the analysis of research data. Also highlighted are the other useful features of the integrated software that are not directly related to data analysis.

Adolescent↗

Integrated software suite for magnetocardiographic data analysis--a proposal based on an interactive programming environment.

OBJECTIVES: This paper describes an integrated software suite (ISS) for the processing of magnetocardiographic (MCG) recordings obtained with super-conducting multi-channel systems having different characteristics. We aimed to develop a highly flexible suite including toolboxes for current MCG applications, organized consistently with an open architecture that allows function integrations and upgrades with minimal modifications; the suite was designed for the compliance not only of physicists and engineers but also of physicians, who have a different professional profile and are accustomed to retrieve information in different ways. METHODS: The MCG-ISS was designed to work with all common graphical user interface operative systems. MATLAB was chosen as the interactive programming environment (IPE), and the software was developed to achieve usability, interactivity, reliability, modularity, expansibility, interoperability, adaptability and graphics style tailoring. Three users, already experienced in MCG data analysis, have intensively tested MCG-ISS for six months. A great amount of MCG data on normal subjects and patients was used to assess software performances in terms of user compliance and confidence and total analysis time. RESULTS: The proposed suite is an all-in-one analysis tool that succeeded in speeding MCG data analysis up to about 55% with respect to standard reference routines; it consequently enhanced analysis performance and user compliance. CONCLUSIONS: Those results, together with the MCG-ISS advantage of being independent on the acquisition system, suggest that software suites like the proposed one could uphold a wider diffusion of MCG as a diagnostic tool in the clinical setting.

Diagnosis, Computer-Assisted↗

The meaning of diagnostic test results: a spreadsheet for swift data analysis.

AIMS: To design a spreadsheet program to: (a) analyse rapidly diagnostic test result data produced in local research or reported in the literature; (b) correct reported predictive values for disease prevalence in any population; (c) estimate the post-test probability of disease in individual patients. MATERIALS AND METHODS: Microsoft Excel(TM)was used. Section A: a contingency (2 x 2) table was incorporated into the spreadsheet. Formulae for standard calculations [sample size, disease prevalence, sensitivity and specificity with 95% confidence intervals, predictive values and likelihood ratios (LRs)] were linked to this table. The results change automatically when the data in the true or false negative and positive cells are changed. Section B: this estimates predictive values in any population, compensating for altered disease prevalence. Sections C-F: Bayes' theorem was incorporated to generate individual post-test probabilities. The spreadsheet generates 95% confidence intervals, LRs and a table and graph of conditional probabilities once the sensitivity and specificity of the test are entered. The latter shows the expected post-test probability of disease for any pre-test probability when a test of known sensitivity and specificity is positive or negative. RESULTS: This spreadsheet can be used on desktop and palmtop computers. The MS Excel(TM)version can be downloaded via the Internet from the URL ftp://radiography.com/pub/Rad-data99.xls CONCLUSION: A spreadsheet is useful for contingency table data analysis and assessment of the clinical meaning of diagnostic test results.

Bayes Theorem↗

A new approach to near-infrared spectral data analysis using independent component analysis.

This paper presents a new approach to near-infrared spectral (NIR) data analysis that is based on independent component analysis (ICA). The main advantage of the new method is that it is able to separate the spectra of the constituent components from the spectra of their mixtures. The separation is a blind operation, since the constituent components of mixtures can be unknown. The ICA based method is therefore particularly useful in identifying the unknown components in a mixture as well as in estimating their concentrations. The approach is introduced by reference to case studies and compared to other techniques for NIR analysis including principal component regression (PCR), multiple linear regression (MLR), and partial least squares (PLS) as well as Fourier and wavelet transforms.

Adipose Tissue↗

Microarray data analysis: a practical approach for selecting differentially expressed genes.

BACKGROUND: The biomedical community is rapidly developing new methods of data analysis for microarray experiments, with the goal of establishing new standards to objectively process the massive datasets produced from functional genomic experiments. Each microarray experiment measures thousands of genes simultaneously producing an unprecedented amount of biological information across increasingly numerous experiments; however, in general, only a very small percentage of the genes present on any given array are identified as differentially regulated. The challenge then is to process this information objectively and efficiently in order to obtain knowledge of the biological system under study and by which to compare information gained across multiple experiments. In this context, systematic and objective mathematical approaches, which are simple to apply across a large number of experimental designs, become fundamental to correctly handle the mass of data and to understand the true complexity of the biological systems under study. RESULTS: The present report develops a method of extracting differentially expressed genes across any number of experimental samples by first evaluating the maximum fold change (FC) across all experimental parameters and across the entire range of absolute expression levels. The model developed works by first evaluating the FC across the entire range of absolute expression levels in any number of experimental conditions. The selection of those genes within the top X% of highest FCs observed within absolute expression bins was evaluated both with and without the use of replicates. Lastly, the FC model was validated by both real time polymerase chain reaction (RT-PCR) and variance data. Semi-quantitative RT-PCR analysis demonstrated 73% concordance with the microarray data from Mu11K Affymetrix GeneChips. Furthermore, 94.1% of those genes selected by the 5% FC model were found to lie above measurement variability using a SDwithin confidence level of 99.9%. CONCLUSION: As evidenced by the high rate of validation, the FC model has the potential to minimize the number of required replicates in expensive microarray experiments by extracting information on gene expression patterns (e.g. characterizing biological and/or measurement variance) within an experiment. The simplicity of the overall process allows the analyst to easily select model limits which best describe the data. The genes selected by this process can be compared between experiments and are shown to objectively extract information which is biologically & statistically significant.

Animals↗

Examination of dioxin fluxes recorded in dated aquatic-sediment cores in the Kanto region of Japan using multivariate data analysis.

Past dioxin (coplanar polychlorinated biphenyl (Co-PCB), 2,3.7,8-substituted polychlorinated dibenzo-p-dioxin (PCDD) and 2,3,7,8-substituted polychlorinated dibenzofuran (PCDF)) fluxes recorded in dated aquatic-sediment cores were analyzed using principal component analysis (PCA). The data set consisted of samples from four cores collected from the Kanto region of Japan. Time trends and spatial differences in the dioxin flux were examined, and the potential relationship to emission sources was investigated. Twenty-five compounds and 58 core slices, corresponding to the later half of the 20th century, were subjected to the analysis. The PCA of both log-transformed and maximum-value-standardized data successfully divided the dioxin compounds into a small number of groups, and three similar clusters of Co-PCBs. PCDDs and penta- to hepta-CDFs were identified. PCB formulations used in the past are judged to have been responsible for the major part of the Co-PCB flux recorded in the sediment cores. However, the relationship to emission sources needs further investigation. It is suggested that most 2,3,7,8-substituted PCDDs and PCDFs are different from Co-PCBs in their emission sources or movements in the environment. The subcore clusters obtained from the PCA of log-transformed data show that the cores from different sampling areas exhibited distinct dioxin fluxes and compositions. Common time trends among the cores were more effectively summarized by the PCA of maximum-value-standardized data focusing on relative time trends. PC scores show that recently the flux of each dioxin compound in the four cores has been generally declining after having reached a peak.

Benzofurans↗

Surveillance system of vaccine adverse events and local data analysis--the experience in a middle-sized city in Brazil, 1999-2001.

We reviewed all vaccine adverse events (VAE) notified in a middle-sized Brazilian city (n=247) to the National Immunisation Program between January 1999 and December 2001. Vaccine doses used in that period were considered for rate estimates. Aspects of the surveillance system (SS) and their influences on collected data were considered, searching for contributions of local data analysis to investigation of VAE and to the monitoring of vaccine safety. Notification rates in our study were higher when compared to national data. Changes in the notification pattern were observed following vaccination campaign periods. An increase in aseptic meningitis cases temporally associated to yellow fever vaccine was detected. The analysis of local data provided information unperceived in national consolidated data. Through this analysis we detected: events related to application technique and handling; people's perception changes on VAE; and the local SS's ability to raise new hypothesis. We suggested changes to the notification form regarding data entry criteria and analysis.

Adolescent↗

Proficiency of the Tradescantia-micronucleus image analysis system for scoring micronucleus frequencies and data analysis.

The Tradescantia-micronucleus (Trad-MCN) bioassay is an efficient short-term test for genotoxicity of pollutants. In order to increase the efficiency and to standardize the micronucleus (MCN) scoring process, an automated scoring system was developed using the principle of image analysis in computer science. This assemblage is called the Tradescantia-micronucleus image analysis (Trad-MCNIA) system. The MCN frequencies scored by this system were compared with those scored by human observation for its proficiency. A set of low MCN frequency (around 5 MCN/100 tetrads) slides prepared from a control group, a set of medium MCN frequency (around 20 MCN/100 tetrads) slides prepared from sodium azide treated plant cuttings and a set of high MCN frequency (around 50 MCN/100 tetrads) slides prepared from X-ray treated materials were used for this study. In the low MCN frequency slides, the Trad-MCNIA system scored about the same value as human observation. In the medium and high frequency slides, MCN frequencies scored by the system were lower than those scored by human observers. This discrepancy was corrected by increasing the power of the objective of the microscope in the system. The MCN frequencies scored by the system attained 90% congruity with those scored by human observers after the correction. The scoring speed of the system was about 3.5 times as fast as that by human observers, and the data could be statistically analyzed immediately after the data scores were recorded. Further improvements can be made by upgrading the video camera and the computer speed.

Azides↗

Recruitment in NHLBI population-based studies and randomized clinical trials: data analysis and survey results.

Data on screening and recruitment from current and previous NHLBI population-based studies (PBSs) and randomized clinical trials (RCTs) were examined. In only two of the studies examined was the projected recruitment completed within the planned recruitment period. The shape of the graph of the relation between enrollment of participants and time varies by study. A single summary statistic for measuring the efficiency of recruitment in RCTs and PBSs is proposed and applied to the examined studies. In addition to providing summary data on recruitment for several studies, this article reports the survey results of a questionnaire sent to the coordinating centers of currently and previously funded National Heart, Lung, and Blood Institute and Veteran's Administration studies. The purpose was to ascertain the desirability of recommending that a generic core of information be collected on recruitment and screening in future studies. Most respondents believed that comparing data collected uniformly and prospectively might be helpful in designing further studies. The variables most respondents believed to be potentially useful are described.

Clinical Trials as Topic↗