Exploratory data analysis. Part II: Steps in examining data.
Explore the source record for details and available documents.
SEARCH · PubMed Health
Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
The occurrence of acne in women with hyperandrogenemia is well known; a question remains, however, as to whether a further positive relationship can be detected between the intensity of acne and the levels of testosterone, androgen precursors and sex hormone binding globulin (SHBG). A procedure of interactive data analysis extracting relevant information from original data was applied. Exploratory data analysis (EDA) identifies basic statistical features and patterns of data using a variety of diagnostic displays. The need for this step is particularly acute in biochemical and clinical data, the distribution of which is mostly non-Gaussian and often corrupted by the outliers. The omission of EDA can lead to incorrect results and false conclusions. In the EDA (i) several graphical tools for summarizing data are applied, (ii) the peculiarities of a sample distribution are investigated, (iii) a construction of distribution is carried out, (iv) a graphical comparison of the sample distribution with selected theoretical distributions is employed. The proposed procedure is illustrated by typical case study in the evaluation of differences between mean values of serum levels of testosterone, androgen precursors and SHBG in a group of patients with mild and severe forms of acne. A knowledge of the interval estimate of the mean value in both groups enables their comparison at the chosen probability level. As will be apparent from the evaluation of inter-group SHBG differences, an incorrect approach to the determination of group mean values could result in a complete misinterpretation of the data. The results indicate that androgens are not significantly related to the intensity of acne, and that SHBG is higher in patients with more severe forms of acne.
In the present work a simple technique for fMRI data analysis is presented. Artifacts due to random and stimulus-correlated motions are corrected without image registration procedures. The first step of our procedure is the calculation of the raw activation map by correlation analysis. The task related motion artifacts arise at the tissue interfaces, including vessels: when image intensity gradient is calculated the high values correspond to interface regions. To eliminate stimulus-correlated motion artifacts the intensity gradient image, obtained from the fMRI data set, is compared to the raw activation map. Since small random motions decrease the value of the correlation coefficient (R) of the external pixels of the activation areas, in the last step of our analysis procedures the clusters are extended to connected pixels having R values smaller than the defined threshold. Each cluster is expanded until the R value of the cluster average intensity is kept constant. The procedure has been tested with both GRE and EPI studies. The presented approach is a fast and robust technique useful for preliminary or on-line analysis of fMRI data.
Two-dimensional difference gel electrophoresis (DIGE) in combination with univariate (Student's t-test) and multivariate data analysis, principal component analysis (PCA) and partial least squares discriminant analysis (PLS-DA) were used to study the anti-inflammatory effects of the beta(2)-adrenergic receptor (beta(2)-AR) agonist zilpaterol. U937 macrophages were exposed to the endotoxin lipopolysaccharide (LPS) to induce an inflammatory reaction, which was inhibited by the addition of zilpaterol (LZ). This inhibition was counteracted by addition of the beta(2)-AR antagonist propranolol (LZP). The extracellular proteome of the U937 cells induced by the three treatments were examined by DIGE. PCA was used as an explorative tool to investigate the clustering of the proteome dataset. Using this tool, the dataset obtained from cells treated with LPS and LZP were separated from those obtained from LZ treated cells. PLS-DA, a multivariate data analysis tool that also takes correlations between protein spots and class assignment into account, correctly classified the different extracellular proteomes and showed that many proteins were differentially expressed between the proteome of inflamed cells (LPS and LZP) and cells in which the inflammatory response was inhibited (LZ). The Student's t-test revealed 8 potential protein biomarkers, each of which was expressed at a similar level in the LPS and LZP treated cells, but differently expressed in the LZ treated cells. Two of the identified proteins, macrophage inflammatory protein-1beta (MIP-1beta) and macrophage inflammatory protein-1alpha (MIP-1alpha) are known secreted proteins. The inhibition of MIP-1beta by zilpaterol and the involvement of the beta(2)-AR and cAMP were confirmed using a specific immunoassay.
An analytical data analysis method is developed for slug tests in partially penetrating wells in confined or unconfined aquifers of high hydraulic conductivity. As adapted from the van der Kamp method, the determination of the hydraulic conductivity is based on the occurrence times and the displacements of the extreme points measured from the oscillatory data and their theoretical counterparts available in the literature. This method is applied to two sets of slug test response data presented by Butler et al.: one set shows slow damping with seven discernable extremities, and the other shows rapid damping with three extreme points. The estimates of the hydraulic conductivity obtained by the analytic method are in good agreement with those determined by an available curve-matching technique.
Hermetia illucens is an important insect resource. Studies have shown that exploring the effects of Cu2+-stressed on the growth and development of the Hermetia illucens genome holds significant scientific importance. There are three major challenges in the current studies of Hermetia illucens genomic data analysis: firstly, the lack of available genomic data which limits researchers in Hermetia illucens genomic data analysis. Secondly, to the best of our knowledge, there are no Artificial Intelligence (AI) feature selection models designed specifically for Hermetia illucens genome. Unlike human genomic data, noise in Hermetia illucens data is a more serious problem. Third, how to choose those genes located in the pathway enrichment region. Existing models assume that each gene probe has the same priori weight. However, researchers usually pay more attention to gene probes which are in the pathway enrichment region. Based on the above challenges, we initially construct experiments and establish a new Cu2+-stressed Hermetia illucens growth genome dataset. Subsequently, we propose AWGE-ESPCA: an edge Sparse PCA model based on adaptive noise elimination regularization and weighted gene network. The AWGE-ESPCA model innovatively proposes an adaptive noise elimination regularization method, effectively addressing the noise challenge in Hermetia illucens genomic data. We also integrate the known gene-pathway quantitative information into the Sparse PCA(SPCA) framework as a priori knowledge, which allows the model to filter out the gene probes in pathway-rich regions as much as possible. Ultimately, this study conducts five independent experiments and compared four latest Sparse PCA models as well as representative supervised and unsupervised baseline models to validate the model performance. The experimental results demonstrate the superior pathway and gene selection capabilities of the AWGE-ESPCA model. Ablation experiments validate the role of the adaptive regularizer and network weighting module. To summarize, this paper presents an innovative unsupervised model for Hermetia illucens genome analysis, which can effectively help researchers identify potential biomarkers. In addition, we also provide a working AWGE - ESPCA model code in the address: https://github.com/yhyresearcher/AWGE_ESPCA.
Multivariate analysis such as principal-components analysis (PCA) and partial-least-squares-discriminant analysis (PLS-DA) have been applied to peptidomics data from clinical urine samples subjected to LC/MS analysis. We show that it is possible to use these methods to get information from a complex set of clinical data. The aim of the work is to use this information as a first step in the further search for clinical biomarker data. It is possible to identify peptide-biomarker fingerprints related to disease diagnosis and progression. Further, we review clinical proteomics and pharmacogenomics data analyzed with the same multivariate approach.
A novel molecular dynamics (MD) analysis algorithm, DASH, is introduced in this paper. DASH has been developed to utilize the sequential nature of MD simulation data. By adjusting a set of parameters, the sensitivity of DASH can be controlled, allowing molecular motions of varying magnitudes to be detected or ignored as desired, with no knowledge of the number of conformations required being prerequisite. MD simulations of three synthetic ligands of the orphan nuclear receptor PPARgamma were generated in vacuo using Tripos's SYBYL and used as the training set for DASH. Two X-ray crystal structures of PPARgamma complexed with Rosiglitazone were compared to gain knowledge of the pharmacophoric conformation; this showed that the conformation of the ligand is significantly different between the two structures, indicating that there is no distinct conformation in which rosiglitazone binds to PPARgamma but multiple binding modes. An investigation into simulation length was carried out. A simulation of 5 ns was found to give highly variable results, whereas a simulation of 25 ns gave a representative window of motion for molecules of this size. DASH was compared with Ward's hierarchical cluster analysis method. The results show that DASH analysis is as good as Ward analysis in some areas (e.g. conformation identification) and is superior in others (e.g. speed and input size).
Two-dimensional (2-D) polyacrylamide gel electrophoresis can detect thousands of polypeptides, separating them by apparent molecular weight (Mr) and isoelectric point (pI). Thus it provides a more realistic and global view of cellular genetic expression than any other technique. This technique has been useful for finding sets of key proteins of biological significance. However, a typical experiment with more than a few gels often results in an unwiedly data management problem. In this paper, the GELLAB-II system is discussed with respect to how data reduction and exploratory data analysis can be aided by computer data management and statistical search techniques. By encoding the gel patterns in a "three-dimensional" (3-D) database, an exploratory data analysis can be carried out in an environment that might be called a "spread sheet for 2-D gel protein data". From such databases, complex parametric network models of protein expression during events such as differentiation might be constructed. For this, 2-D gel databases must be able to include data from other domains external to the gel itself. Because of the increasing complexity of such databases, new tools are required to help manage this complexity. Two such tools, object-oriented databases and expert-system rule-based analysis, are discussed in this context. Comparisons are made between GELLAB and other 2-D gel database analysis systems to illustrate some of the analysis paradigms common to these systems and where this technology may be heading.
"Given the crucial role played by census data in informing economic and social policies directed at the Aboriginal population in remote areas, some assessment of the quality of remote area data is required as these are derived from enumeration procedures which differ fundamentally from the standard approach employed in the census. This paper discusses the remote area census enumeration strategy employed by the Australian Bureau of Statistics (ABS), with a particular focus on the Northern Territory, and highlights possible implications for the interpretation of census counts and census characteristics."
Data envelopment analysis (DEA) is used to evaluate the relative technical efficiency and assist in the management of a chain of nursing homes. As with any DEA model, variables chosen are particularly important. The study looks at two possibly critical issues. The first is the appropriateness of models that include only financial and economic measures to evaluate administrators when quality care is an expected output. The second issue is the appropriateness of using noncontrollable variables, in this case operating income, to evaluate administrators. We show how efficiency scores differ when quality variables and/or operating income are included. We also demonstrate the usefulness of DEA information to both the home administrator and chain managers for improving operating efficiency.
The planned Australian National Cardiac Surgery Database is likely to have a number of positive outcomes, including increased patient satisfaction, improved quality assurance and increased economic efficiency. In relation to cardiac surgery, performance indicators associated with the commonly performed procedure of coronary artery bypass surgery will be used for peer review and to measure outcomes. Several different risk-adjusted models are available for analysing national databases. However, the potential weaknesses of database analysis are lack of both compliance and data validity. A number of other major issues, such as location of the data analysis centre, who will hold authority over data accuracy, and the security of and access to the Database, must also be considered when setting up the National Database. Overall, however, the benefits of a national database will be enormous. Cardiologists and cardiac surgeons will benefit from a disease-based registry with shared common definitions. In addition, the provision of such a database will represent a crucial step towards developing national strategies for treating heart disease.
Data Envelopment Analysis (DEA) identifies price and technical inefficiencies among decision-making units. With controls for differences in case-mix and standardized outcomes, DEA's "best practice" frontier can be interpreted as a "cost-effectiveness" frontier. This study illustrates the key concepts, identifies the decisions required to use the technique for medical care decision making, and presents an application to a system of nine hospitals that offer obstetric services.
For quantification of gene-specific mRNA, quantitative real-time RT-PCR has become one of the most frequently used methods over the last few years. This article focuses on the issue of real-time PCR data analysis and its mathematical background, offering a general concept for efficient, fast and precise data analysis superior to the commonly used comparative CT (DeltaDeltaCT) and the standard curve method, as it considers individual amplification efficiencies for every PCR. This concept is based on a novel formula for the calculation of relative gene expression ratios, termed GED (Gene Expression's CT Difference) formula. Prerequisites for this formula, such as real-time PCR kinetics, the concept of PCR efficiency and its determination, are discussed. Additionally, this article offers some technical considerations and information on statistical analysis of real-time PCR data.
Langerhans cells are dendritic cells situated in the mammalian epidermis. In human epidermis, the concentration is between 460 and 1000 mm(-2). Langerhans cells fulfill an essential role in skin immune responses. Numerous scientific reports on Langerhans cells have appeared, but with no systematic research on the pattern of the spatial distributions. On the contrary, in certain fields, a spatial distribution is an important theme, and spatial data analysis has a long history. We hypothesized that epidermal Langerhans cells were set in the best formation for their immuno-surveillance by a sophisticated mechanism. To prove this hypothesis, we have imported spatial data analysis into the study of epidermal Langerhans cells. Here, we show that the distribution is completely regular; the pattern of Voronoi divisions fits the territories; the random packing model simulates their bone marrow derivation; a repulsive interaction is demonstrated and a repulsive potential function is estimated. Spatial data analysis-based computer simulation will be a new method of Langerhans cell study. In addition, this procedure shows promise for future distribution research of certain cells.
Microarrays are one of the latest breakthroughs in experimental molecular biology, which allow monitoring of gene expression for tens of thousands of genes in parallel and are already producing huge amounts of valuable data. Analysis and handling of such data is becoming one of the major bottlenecks in the utilization of the technology. The raw microarray data are images, which have to be transformed into gene expression matrices--tables where rows represent genes, columns represent various samples such as tissues or experimental conditions, and numbers in each cell characterize the expression level of the particular gene in the particular sample. These matrices have to be analyzed further, if any knowledge about the underlying biological processes is to be extracted. In this paper we concentrate on discussing bioinformatics methods used for such analysis. We briefly discuss supervised and unsupervised data analysis and its applications, such as predicting gene function classes and cancer classification. Then we discuss how the gene expression matrix can be used to predict putative regulatory signals in the genome sequences. In conclusion we discuss some possible future directions.
Microarrays are one of the latest breakthroughs in experimental molecular biology, which allow monitoring of gene expression for tens of thousands of genes in parallel and are already producing huge amounts of valuable data. Analysis and handling of such data is becoming one of the major bottlenecks in the utilization of the technology. The raw microarray data are images, which have to be transformed into gene expression matrices, tables where rows represent genes, columns represent various samples such as tissues or experimental conditions, and numbers in each cell characterize the expression level of the particular gene in the particular sample. These matrices have to be analyzed further if any knowledge about the underlying biological processes is to be extracted. In this paper we concentrate on discussing bioinformatics methods used for such analysis. We briefly discuss supervised and unsupervised data analysis and its applications, such as predicting gene function classes and cancer classification as well as some possible future directions.
Computer programs are described that allow facile analysis of data from a protein sequencer and amino acid analyzer. The sequencer program provides automated sequence interpretation while requiring minimal user interaction. The program serves as a powerful aid in deciphering mixture sequences and allows routine monitoring of sequencer performance. The computer program for amino acid analysis data provides the following calculations: mole percent, protein concentration and residues per mole with comparison between theoretical and calculated values. A plot of molecular weight versus deviation from integer values is calculated providing a measure of peptide or protein purity.