PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Crash risks of older drivers: a panel data analysis.

Considerable progress has been made on understanding older drivers' safety issues. None the less, findings from previous research have been rather inconclusive. Differences in data and research methodology have been suggested as factors that contribute to the discrepancies in previous findings. One of the methodological limitations is the lack of considering temporal order between events (i.e. the time between onset of medical condition, symptom and crash). Without time-series data, a 'snap-shot' of medical conditions and driving patterns were often linked to more than 1 year--of crash data, hoping to accumulate enough data on crashes. The interpretation of the results from these studies is difficult in that one cannot explicitly attribute the increase in highway crash rates to medical conditions and/or physical limitations. This paper uses a panel data analysis to identify factors that place older drivers at greater crash risk. Our results show that factors that place female drivers at greater crash risk are different from those influencing male drivers. More risk factors were found to be significant in affecting older men's involvement in crashes than older women. When the analysis controlled for the amount of driving, women who live alone or who experience back pain were found to have a higher crash risk. Similarly, men who are employed, score low on word-recall tests, have a history of glaucoma, or use antidepressant drugs were found to have a higher crash risk. The most influential risk factors in men were the number of miles driven, and use of antidepressants.

Accidents, Traffic↗

Classification of atherosclerotic rabbit aorta samples with an infrared attenuated total reflection catheter and multivariate data analysis.

The strongly overlapping infrared absorption features of atherosclerotic and normal rabbit aorta samples as governed by their water, lipid, and protein content render the direct evaluation of molecular characteristics obtained from infrared (IR) spectroscopic measurements challenging for classification. We have successfully applied multivariate data analysis and classification techniques based on partial least squares regression (PLS), linear discriminant analysis (LDA), and principal component regression (PCR) to IR spectroscopic data obtained by using a recently developed infrared attenuated total reflectance (IR-ATR) catheter prototype for future in vivo diagnostic applications. Training data were collected ex vivo from atherosclerotic and normal rabbit aorta samples. The successful classification results on atherosclerotic and normal aorta samples utilizing the developed data evaluation routines reveals the potential of spectroscopy combined with multivariate classification strategies for the identification of normal and atherosclerotic aorta tissue for in vitro and, in the future, in vivo applications.

Algorithms↗

Predicting the subcellular localization of human proteins using machine learning and exploratory data analysis.

Identifying the subcellular localization of proteins is particularly helpful in the functional annotation of gene products. In this study, we use Machine Learning and Exploratory Data Analysis (EDA) techniques to examine and characterize amino acid sequences of human proteins localized in nine cellular compartments. A dataset of 3,749 protein sequences representing human proteins was extracted from the SWISS-PROT database. Feature vectors were created to capture specific amino acid sequence characteristics. Relative to a Support Vector Machine, a Multi-layer Perceptron, and a Naive Bayes classifier, the C4.5 Decision Tree algorithm was the most consistent performer across all nine compartments in reliably predicting the subcellular localization of proteins based on their amino acid sequences (average Precision=0.88; average Sensitivity=0.86). Furthermore, EDA graphics characterized essential features of proteins in each compartment. As examples, proteins localized on the plasma membrane had higher proportions of hydrophobic amino acids; cytoplasmic proteins had higher proportions of neutral amino acids; and mitochondrial proteins had higher proportions of neutral amino acids and lower proportions of polar amino acids. These data showed that the C4.5 classifier and EDA tools can be effective for characterizing and predicting the subcellular localization of human proteins based on their amino acid sequences.

Algorithms↗

[Planning and data analysis in prospective controlled clinical trials (author's transl)].

Planning of prospective controlled clinical trials in surgery requires the use of test and control groups, sufficiently frequent repetition of experiments, random allocation of patients to the groups (example), and balancing. The descriptive data analysis should be performed in a stepwise manner (list of new data, rank list, range, median, quartiles, histogram, mean value standard deviation). The advantages of the median-quartile-system and the prerequisites for application of various significance tests are pointed out. In the conduct of controlled clinical trials, the consultative role of experimental surgeons is proposed.

Clinical Trials as Topic↗

FM-test: a fuzzy-set-theory-based approach to differential gene expression data analysis.

BACKGROUND: Microarray techniques have revolutionized genomic research by making it possible to monitor the expression of thousands of genes in parallel. As the amount of microarray data being produced is increasing at an exponential rate, there is a great demand for efficient and effective expression data analysis tools. Comparison of gene expression profiles of patients against those of normal counterpart people will enhance our understanding of a disease and identify leads for therapeutic intervention. RESULTS: In this paper, we propose an innovative approach, fuzzy membership test (FM-test), based on fuzzy set theory to identify disease associated genes from microarray gene expression profiles. A new concept of FM d-value is defined to quantify the divergence of two sets of values. We further analyze the asymptotic property of FM-test, and then establish the relationship between FM d-value and p-value. We applied FM-test to a diabetes expression dataset and a lung cancer expression dataset, respectively. Within the 10 significant genes identified in diabetes dataset, six of them have been confirmed to be associated with diabetes in the literature and one has been suggested by other researchers. Within the 10 significantly overexpressed genes identified in lung cancer data, most (eight) of them have been confirmed by the literatures which are related to the lung cancer. CONCLUSION: Our experiments on synthetic datasets show that FM-test is effective and robust. The results in diabetes and lung cancer datasets validated the effectiveness of FM-test. FM-test is implemented as a Web-based application and is available for free at http://database.cs.wayne.edu/bioinformatics.

Algorithms↗

Horses for courses: facilitating postgraduate research students' choice of Computer Assisted Qualitative Data Analysis System (CAQDAS).

Supervisors of postgraduate students are increasingly likely to find themselves discussing whether or not the student should use a CAQDAS (computer assisted qualitative data analysis system) in their research. This paper discusses Weitzman and Miles (1995) framework for decision-making about CAQDAS and then reports the experiences of five postgraduate students, each of whom made a different decision. (These were variously: not to use a CAQDAS, using Atlas-Ti, Ethnograph, N. VIVO and N5). It explores the fit between Weitzman's and Miles' principles and the students' experiences then suggests some modifications of the principles and strategies for advising students.

Choice Behavior↗

Exploratory spatial data analysis for the identification of risk factors to birth defects.

BACKGROUND: Birth defects, which are the major cause of infant mortality and a leading cause of disability, refer to "Any anomaly, functional or structural, that presents in infancy or later in life and is caused by events preceding birth, whether inherited, or acquired (ICBDMS)". However, the risk factors associated with heredity and/or environment are very difficult to filter out accurately. This study selected an area with the highest ratio of neural-tube birth defect (NTBD) occurrences worldwide to identify the scale of environmental risk factors for birth defects using exploratory spatial data analysis methods. METHODS: By birth defect registers based on hospital records and investigation in villages, the number of birth defects cases within a four-year period was acquired and classified by organ system. The neural-tube birth defect ratio was calculated according to the number of births planned for each village in the study area, as the family planning policy is strictly adhered to in China. The Bayesian modeling method was used to estimate the ratio in order to remove the dependence of variance caused by different populations in each village. A recently developed statistical spatial method for detecting hotspots, Getis's 7, was used to detect the high-risk regions for neural-tube birth defects in the study area. RESULTS: After the Bayesian modeling method was used to calculate the ratio of neural-tube birth defects occurrences, Getis's statistics method was used in different distance scales. Two typical clustering phenomena were present in the study area. One was related to socioeconomic activities, and the other was related to soil type distributions. CONCLUSION: The fact that there were two typical hotspot clustering phenomena provides evidence that the risk for neural-tube birth defect exists on two different scales (a socioeconomic scale at 6.84 km and a soil type scale at 22.8 km) for the area studied. Although our study has limited spatial exploratory data for the analysis of the neural-tube birth defect occurrence ratio and for finding clues to risk factors, this result provides effective clues for further physical, chemical and even more molecular laboratory testing according to these two spatial scales.

Bayes Theorem↗

Histopathological classification of dementias by multivariate data analysis.

Autopsied brains from 55 demented patients, clinically classified according to DSM-III criteria into AD/SDAT and MID and 19 nondemented individuals were available for this study. Using general clinical, gross neuroanatomical and histopathological data three separate dementia classes, namely AD/SDAT, MID and AD-MID, were visualized in two-dimensional space by multivariate data analysis. This analysis revealed that the pathology in AD-MID patients were not merely a linear combination of the pathology in AD/SDAT and MID, indicating that AD-MID might represent a dementia type of its own.

Aged↗

Correlation of human jejunal permeability (in vivo) of drugs with experimentally and theoretically derived parameters. A multivariate data analysis approach.

The effective permeability (Peff) in the human jejunum (in vivo) of 22 structurally diverse compounds was correlated with both experimentally determined lipophilicity values and calculated molecular descriptors. The permeability data were previously obtained by using a regional in vivo perfusion system in the proximal jejunum in humans as part of constructing a biopharmaceutical classification system for oral immediate-release products. pKa, log P, and, where relevant, log Pion values were determined using the pH-metric technique. On the basis of these experiments, log D values were calculated at pH 5.5, 6.5, and 7.4. Multivariate data analysis was used to derive models that correlate passive intestinal permeability to physicochemical descriptors. The best model obtained, based on 13 passively transcellularly absorbed compounds, used the variables HBD (number of hydrogen bond donors), PSA (polar surface area), and either log D5.5 or log D6.5 (octanol/water distribution coefficient at pH 5.5 and 6.5, respectively). Statistically good models for prediciting human in vivo Peff values were also obtained by using only HBD and PSA or HBD, PSA, and CLOGP. These models can be used to predict passive intestinal membrane diffusion in humans for compounds that fit within the defined property space. We used one of the models obtained above to predict the log Peff values for an external validation set consisting of 34 compounds. A good correlation with the absorption data of these compounds was found.

Humans↗

Data analysis issues for protocols with overlapping enrollment.

Many persons with HIV require and take several medications. The efficacy and safety of many of these medications are uncertain. Usually limited data on drug interactions are available. Thus simultaneous and sequential enrolment of patients into multiple studies is desired for reasons of science and efficiency. This paper discusses the analysis of data arising from coenrolment in multiple studies sponsored by the Community Programs for Clinical Research on AIDS (CPCRA). Factorial designs and those in which patients are sequentially instead of simultaneously randomized are compared. Approaches to data analysis, based on intention-to-treat, for individual and pairs of trials are described. An antiretroviral trial and a trial for prophylaxis of Pneumocystis carinii pneumonia (PCP) are used for illustration. We conclude that such analyses may yield useful information on drug interactions and that a more vigorous coenrolment policy should be pursued in AIDS research.

AIDS-Related Opportunistic Infections↗

Simulation of DNA array hybridization experiments and evaluation of critical parameters during subsequent image and data analysis.

BACKGROUND: Gene expression analyses based on complex hybridization measurements have increased rapidly in recent years and have given rise to a huge amount of bioinformatic tools such as image analyses and cluster analyses. However, the amount of work done to integrate and evaluate these tools and the corresponding experimental procedures is not high. Although complex hybridization experiments are based on a data production pipeline that incorporates a significant amount of error parameters, the evaluation of these parameters has not been studied yet in sufficient detail. RESULTS: In this paper we present simulation studies on several error parameters arising in complex hybridization experiments. A general tool was developed that allows the design of exactly defined hybridization data incorporating, for example, variations of spot shapes, spot positions and local and global background noise. The simulation environment was used to judge the influence of these parameters on subsequent data analysis, for example image analysis and the detection of differentially expressed genes. As a guide for simulating expression data real experimental data were used and model parameters were adapted to these data. Our results show how measurement error can be balanced by the analysis tools. CONCLUSIONS: We describe an implemented model for the simulation of DNA-array experiments. This tool was used to judge the influence of critical parameters on the subsequent image analysis and differential expression analysis. Furthermore the tool can be used to guide future experiments and to improve performance by better experimental design. Series of simulated images varying specific parameters can be downloaded from our web-site: http://www.molgen.mpg.de/~lh_bioinf/projects/simulation/biotech/

Algorithms↗

ITTACA: a new database for integrated tumor transcriptome array and clinical data analysis.

Transcriptome microarrays have become one of the tools of choice for investigating the genes involved in tumorigenesis and tumor progression, as well as finding new biomarkers and gene expression signatures for the diagnosis and prognosis of cancer. Here, we describe a new database for Integrated Tumor Transcriptome Array and Clinical data Analysis (ITTACA). ITTACA centralizes public datasets containing both gene expression and clinical data. ITTACA currently focuses on the types of cancer that are of particular interest to research teams at Institut Curie: breast carcinoma, bladder carcinoma and uveal melanoma. A web interface allows users to carry out different class comparison analyses, including the comparison of expression distribution profiles, tests for differential expression and patient survival analyses. ITTACA is complementary to other databases, such as GEO and SMD, because it offers a better integration of clinical data and different functionalities. It also offers more options for class comparison analyses when compared with similar projects such as Oncomine. For example, users can define their own patient groups according to clinical data or gene expression levels. This added flexibility and the user-friendly web interface makes ITTACA especially useful for comparing personal results with the results in the existing literature. ITTACA is accessible online at http://bioinfo.curie.fr/ittaca.

Breast Neoplasms↗

Bioinformatics in mass spectrometry data analysis for proteomics studies.

Mass spectrometry is a technique widely employed for the identification and characterization of proteins. The role of bioinformatics is fundamental for the elaboration of mass spectrometry data due to the amount of data that this technique can produce. To process data efficiently, new software packages and algorithms are continuously being developed to improve protein identification and characterization in terms of high-throughput and statistical accuracy. However, many limitations exist concerning bioinformatics spectral data elaboration. This review aims to critically cover the recent and future developments of new bioinformatics approaches in mass spectrometry data analysis for proteomics studies.

Computational Biology↗

A grid computing infrastructure for MEG data analysis.

Magnetoencephalography (MEG) is widely used for studying brain functions, but clinical applications of MEG have been less prevalent. One reason is that only clinicians who have highly specialized knowledge can use MEG diagnostically, and such clinicians are found at only a few major hospitals. Another reason is that MEG data analysis is getting more and more complicated, and deals with a large amount of data, and thus requires high-performance computing. These problems can be solved by the collaboration of human and computing resources distributed in multiple facilities. A new computing infrastructure for brain scientists and clinicians in distant locations was therefore developed by the Grid technology, which provides virtual computing environments composed of geographically distributed computers and experimental devices. A prototype system connecting an MEG system at the AIST in Japan, a Grid environment composed of PC clusters at Osaka University in Japan and Nanyang Technological University in Singapore, and user terminals in Baltimore was developed. MEG data measured at the AIST were transferred in real-time through a 1-GB/s network to the PC clusters for processing by a wavelet cross-correlation method, and then monitored in Baltimore. The current system is the basic model for remote-access to MEG equipment and high-speed processing of MEG data.

Computing Methodologies↗

Thermodynamical model of indefinite mixed association of two components and NMR data analysis for caffeine-AMP interaction.

This paper describes the model used to estimate the parameters of caffeine-AMP interactions from corresponding 1H-NMR measurements and some methods of data analysis by which the applicability of the model has been checked. The model of mixed association is applicable to a mixture of any two substances A and C which exhibit indefinite aggregates in both self-association and mixed association. In aggregates, only nearest neighbour interaction is assumed. The model is described by three equilibrium constants: Kaa and Kcc (for self-association of A, or C, respectively), and Kac (for mixed association).

Journal Article↗

Maximum entropy and Bayesian data analysis: Entropic prior distributions.

The problem of assigning probability distributions which reflect the prior information available about experiments is one of the major stumbling blocks in the use of Bayesian methods of data analysis. In this paper the method of maximum (relative) entropy (ME) is used to translate the information contained in the known form of the likelihood into a prior distribution for Bayesian inference. The argument is inspired and guided by intuition gained from the successful use of ME methods in statistical mechanics. For experiments that cannot be repeated the resulting "entropic prior" is formally identical with the Einstein fluctuation formula. For repeatable experiments, however, the expected value of the entropy of the likelihood turns out to be relevant information that must be included in the analysis. The important case of a Gaussian likelihood is treated in detail.

Journal Article↗

Functional data analysis of knee joint kinematics in the vertical jump.

Understanding of the motor development process is usually based on descriptive studies using either cross-sectional or longitudinal designs. These data typically consist of sets of measurements on groups of individuals gathered during the performance of a single task. A natural approach is to represent the set of measurements for an individual as a single entity, a function. In practice, however, this approach is seldom applied. Typically, the analysis proceeds by reducing what are intrinsically functional responses to a single summary measurement and then using this to draw conclusions. As a result, many potentially informative data are ignored. Functional data analysis (FDA) is an emerging field in statistics that focuses on treating an entire sequence of measurements for an experimental unit as a single function. Therefore, functional data analysis appears to be inherently suitable for analysing kinematic data. In this paper, the key concepts and procedures of functional data analysis are introduced and illustrated using data obtained from a cross-sectional study on the development of the vertical jump.

Biomechanical Phenomena↗

Discussion on the choice of separated components in fMRI data analysis by spatial independent component analysis.

By measuring the changes of magnetic resonance signals during a stimulation, the functional magnetic resonance imaging (fMRI) is able to localize the neural activation in the brain. In this report, we discuss the fMRI application of the spatial independent component analysis (spatial ICA), which maximizes statistical independence over spatial images. Included simulations show the possibility of the spatial ICA on discriminating asynchronous activations or different response patterns in an fMRI data set. An in vivo visual stimulation fMRI test was conducted, and the result shows a proper sum of the separated components as the final image is better than a single component, using fMRI data analysis by spatial ICA. Our result means that spatial ICA is a useful tool for the detection of different response activations and suggests that a proper sum of the separated independent components should be used for the imaging result of fMRI data processing.

Algorithms↗