PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Data Mining”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Gene Aging Nexus: a web database and data mining platform for microarray data on aging.

The recent development of microarray technology provided unprecedented opportunities to understand the genetic basis of aging. So far, many microarray studies have addressed aging-related expression patterns in multiple organisms and under different conditions. The number of relevant studies continues to increase rapidly. However, efficient exploitation of these vast data is frustrated by the lack of an integrated data mining platform or other unifying bioinformatic resource to enable convenient cross-laboratory searches of array signals. To facilitate the integrative analysis of microarray data on aging, we developed a web database and analysis platform 'Gene Aging Nexus' (GAN) that is freely accessible to the research community to query/analyze/visualize cross-platform and cross-species microarray data on aging. By providing the possibility of integrative microarray analysis, GAN should be useful in building the systems-biology understanding of aging. GAN is accessible at http://gan.usc.edu.

Aging↗

Case study: how to apply data mining techniques in a healthcare data warehouse.

Healthcare provider organizations are faced with a rising number of financial pressures. Both administrators and physicians need help analyzing large numbers of clinical and financial data when making decisions. To assist them, Rush-Presbyterian-St. Luke's Medical Center and Hitachi America, Ltd. (HAL), Inc., have partnered to build an enterprise data warehouse and perform a series of case study analyses. This article focuses on one analysis, which was performed by a team of physicians and computer science researchers, using a commercially available on-line analytical processing (OLAP) tool in conjunction with proprietary data mining techniques developed by HAL researchers. The initial objective of the analysis was to discover how to use data mining techniques to make business decisions that can influence cost, revenue, and operational efficiency while maintaining a high level of care. Another objective was to understand how to apply these techniques appropriately and to find a repeatable method for analyzing data and finding business insights. The process used to identify opportunities and effect changes is described.

Chicago↗

Data mining the protein data bank: automatic detection and assignment of carbohydrate structures.

Knowledge of the 3D structure of glycans is a prerequisite for a complete understanding of the biological processes glycoproteins are involved in. However, due to a lack of standardised nomenclature, carbohydrate compounds are difficult to locate within the Protein Data Bank (PDB). Using an algorithm that detects carbohydrate structures only requiring element types and atom coordinates, we were able to detect 1663 entries containing a total of 5647 carbohydrate chains. The majority of chains are found to be N-glycosidically bound. Noncovalently bound ligands are also frequent, while O-glycans form a minority. About 30% of all carbohydrate containing PDB entries comprise one or several errors. The automatic assignment of carbohydrate structures in PDB entries will improve the cross-linking of glycobiology resources with genomic and proteomic data collections, which will be an important issue of the upcoming glycomics projects. By aiding in detection of erroneous annotations and structures, the algorithm might also help to increase database quality.

Algorithms↗

Sequential projection pursuit using genetic algorithms for data mining of analytical data.

Sequential projection pursuit (SPP) is proposed to detect inhomogeneities (clusters) in high-dimensional analytical data. Such inhomogeneities indicate that there are groups of objects (samples) with different chemical characteristics. The method is compared with principal component analysis (PCA). PCA is generally applied to visually explore structure in high-dimensional data, but is not specifically used to find clustering tendency. Projection pursuit (PP) is specifically designed to find inhomogeneities, but the original method is computationally very intensive. SPP combines the advantages of both methods and overcomes most of their weak points. In this method, latent variables are obtained sequentially according to their importance measured by the entropy index. This involves an optimization step, which is achieved by using a genetic algorithm. The performance of the method is demonstrated and evaluated, first on simulated data sets, and then on near-infrared and gas chromatography data sets. It is shown that SPP indeed reveals more easily information about inhomogeneities than PCA.

Algorithms↗

Data mining of the GAW14 simulated data using rough set theory and tree-based methods.

Rough set theory and decision trees are data mining methods used for dealing with vagueness and uncertainty. They have been utilized to unearth hidden patterns in complicated datasets collected for industrial processes. The Genetic Analysis Workshop 14 simulated data were generated using a system that implemented multiple correlations among four consequential layers of genetic data (disease-related loci, endophenotypes, phenotypes, and one disease trait). When information of one layer was blocked and uncertainty was created in the correlations among these layers, the correlation between the first and last layers (susceptibility genes and the disease trait in this case), was not easily directly detected. In this study, we proposed a two-stage process that applied rough set theory and decision trees to identify genes susceptible to the disease trait. During the first stage, based on phenotypes of subjects and their parents, decision trees were built to predict trait values. Phenotypes retained in the decision trees were then advanced to the second stage, where rough set theory was applied to discover the minimal subsets of genes associated with the disease trait. For comparison, decision trees were also constructed to map susceptible genes during the second stage. Our results showed that the decision trees of the first stage had accuracy rates of about 99% in predicting the disease trait. The decision trees and rough set theory failed to identify the true disease-related loci.

Computer Simulation↗

Data mining in brain imaging.

Data mining in brain imaging is proving to be an effective methodology for disease prognosis and prevention. This, together with the rapid accumulation of massive heterogeneous data sets, motivates the need for efficient methods that filter, clarify, assess, correlate and cluster brain-related information. Here, we present data mining methods that have been or could be employed in the analysis of brain images. These methods address two types of brain imaging data: structural and functional. We introduce statistical methods that aid the discovery of interesting associations and patterns between brain images and other clinical data. We consider several applications of these methods, such as the analysis of task-activation, lesion-deficit, and structure morphological variability; the development of probabilistic atlases; and tumour analysis. We include examples of applications to real brain data. Several data mining issues, such as that of method validation or verification, are also discussed.

Algorithms↗

Data mining: qualitative analysis with health informatics data.

The new computational algorithms emerging in the data mining literature--in particular, the self-organizing map (SOM) and decision tree analysis (DTA)--offer qualitative researchers a unique set of tools for analyzing health informatics data. The uniqueness of these tools is that although they can be used to find meaningful patterns in large, complex quantitative databases, they are qualitative in orientation. To illustrate the utility of these tools, the authors review the two most popular: the SOM and DTA. They provide a basic definition of health informatics, focusing on how data mining assists this field, and apply the SOM and DTA to a hypothetical example to demonstrate what these tools are and how qualitative researchers can use them.

Algorithms↗

Myocardial infarction--pinpointing the key indicators in the 12-lead ECG using data mining.

In this paper we describe how data mining techniques were used in order to pinpoint the key indicators for myocardial infarction in the electrocardiogram (ECG) by determining existing trends in a large data set. In order to provide a test bed for the data mining techniques a data mining tool was developed so that the effectiveness of various data mining techniques could be determined. The material consisted of 2730 ECGs recorded at an emergency department. A total of 517 ECGs were recorded on patients suffering acute myocardial infarction. The remaining ECGs were defined as control ECGs. A subset of the material was used to train the data mining tool. After training, the data mining tool was able to pinpoint the key ECG indicators for myocardial infarction in the test set (duration and amplitude of the Q wave and R duration in lead V2) and successfully determine which patients had suffered a heart attack.

Algorithms↗

Data mining in pharmacovigilance: the need for a balanced perspective.

Data mining is receiving considerable attention as a tool for pharmacovigilance and is generating many perspectives on its uses. This paper presents four concepts that have appeared in various professional venues and represent potential sources of misunderstanding and/or entail extended discussions: (i) data mining algorithms are unvalidated; (ii) data mining algorithms allow data miners to objectively screen spontaneous report data; (iii) mathematically more complex Bayesian algorithms are superior to frequentist algorithms; and (iv) data mining algorithms are not just for hypothesis generation. Key points for a balanced perspective are that: (i) validation exercises have been done but lack a gold standard for comparison and are complicated by numerous nuances and pitfalls in the deployment of data mining algorithms. Their performance is likely to be highly situation dependent; (ii) the subjective nature of data mining is often underappreciated; (iii) simpler data mining models can be supplemented with 'clinical shrinkage', preserving sensitivity; and (iv) applications of data mining beyond hypothesis generation are risky, given the limitations of the data. These extended applications tend to 'creep', not pounce, into the public domain, leading to potential overconfidence in their results. Most importantly, in the enthusiasm generated by the promise of data mining tools, users must keep in mind the limitations of the data and the importance of clinical judgment and context, regardless of statistical arithmetic. In conclusion, we agree that contemporary data mining algorithms are promising additions to the pharmacovigilance toolkit, but the level of verification required should be commensurate with the nature and extent of the claimed applications.

Adverse Drug Reaction Reporting Systems↗

Data mining and structuring of executable data analysis reports: guideline development and implementation in a narrow sense.

In this paper we present a data mining scenario that supports development of automated web-based documentation of data analysis for diagnosis and treatment. The documents can be seen as guidelines in a narrow sense, and are designed to include executable modules for the corresponding decision support systems. Our aim is to discuss the possibilities of identifying certain types of diagnoses and treatments for which guidelines can be generated and computerised more systematically.

Artificial Intelligence↗

Automated detection of hereditary syndromes using data mining.

Computer-based data mining methodology applied to family history clinical data can algorithmically create highly accurate, clinically oriented hereditary disease pattern recognizers. For the example of hereditary colon cancer, the data mining's selection of relevant factors to assess for hereditary colon cancer was statistically significant (P < 0.05). All final recognizer-formulated patterns of hereditary colon cancer were independently confirmed by a clinical expert. Applied to previously analyzed family histories, the recognizer identified the definitive hereditary histories, correctly responded negatively to the putative hereditary histories, and correctly responded negatively to empirically elevated colon cancer risk situations. This capability facilitates patient selection for DNA studies in search of gene mutations. When genetic mutations are included as parameters in a patient database for a genetic disease, the process yields an expert system which characterizes variations in clinical disease presentations in terms of genetic mutations. Such information can greatly improve the efficiency of gene testing.

Adult↗

Data mining in child welfare.

Data mining is the sifting through of voluminous data to extract knowledge for decision making. This article illustrates the context, concepts, processes, techniques, and tools of data mining, using statistical and neural network analyses on a dataset concerning employee turnover. The resulting models and their predictive capability, advantages and disadvantages, and implications for decision support are highlighted.

Child↗

Appropriate medical data categorization for data mining classification techniques.

Some data mining (DM) methods, or software tools, require normalized data, others rely on categorized data, and some can accommodate multiple data scales. Each DM technique has a specific background theory; therefore, different results are expected when applying multiple methods. The purpose of this study is to find the data format appropriate for each DM classification technique for wider applications, and efficiently to obtain trustworthy results. Considering the nature of medical data, categorical variables are sometimes useful for making decisions and can make it easier to extrapolate knowledge. In this study, three mathematical data categorization methods (Fusinter, minimum description length principle [MDLPC] and Chi-merge) were applied to accommodate five data mining classification techniques (statistics discriminant analysis, supervised classification with Neural Networks, Decision trees, Genetic supervised clustering and Bayesian classification [probability neural networks; PNN]) using a heart disease database with four types of data (continuous data, binary data, nominal data, and ordinal data). Compared with original or normalized data, data categorized by the MDLPC categorization method was found to perform better in most of the DM classification techniques used in this study. Categorical data is good for most DM classification techniques (e.g. classification of disease and non-disease groups) and is relatively easy to use for extracting medical knowledge.

Bayes Theorem↗

Preliminary implementation of new data mining techniques for the analysis of simulation data from Genetic Analysis Workshop 12: problem 2.

We introduce a new data mining method applicable to complex disease genetics. Our approach is suited to a broad spectrum of diseases, identifying the noteworthy sharing of combinations of alleles in unrelated affected individuals. Furthermore, this approach may be extended to comprise the common types of genotype data, including single-nucleotide polymorphisms, candidate-gene sequences, etc. Using a method derived from data-mining computer algorithms, we analyze a data set of unrelated affected individuals chosen from the simulated pedigrees of problem 2 of the Genetics Analysis Workshop 12. We observe that most marker subsets containing a flanking marker for each of six or seven of the disease-gene loci yield significant numbers of individuals manifesting substantially similar genotypes. However, initial attempts (blind to the generating model) to identify the predisposing loci have not been successful. Refining our methods so that such loci may routinely be found and validated is underway.

Algorithms↗

Evolutionary optimization of radial basis function classifiers for data mining applications.

In many data mining applications that address classification problems, feature and model selection are considered as key tasks. That is, appropriate input features of the classifier must be selected from a given (and often large) set of possible features and structure parameters of the classifier must be adapted with respect to these features and a given data set. This paper describes an evolutionary algorithm (EA) that performs feature and model selection simultaneously for radial basis function (RBF) classifiers. In order to reduce the optimization effort, various techniques are integrated that accelerate and improve the EA significantly: hybrid training of RBF networks, lazy evaluation, consideration of soft constraints by means of penalty terms, and temperature-based adaptive control of the EA. The feasibility and the benefits of the approach are demonstrated by means of four data mining problems: intrusion detection in computer networks, biometric signature verification, customer acquisition with direct marketing methods, and optimization of chemical production processes. It is shown that, compared to earlier EA-based RBF optimization techniques, the runtime is reduced by up to 99% while error rates are lowered by up to 86%, depending on the application. The algorithm is independent of specific applications so that many ideas and solutions can be transferred to other classifier paradigms.

Algorithms↗

Increasing the diversity of medical data mining through distributed object technology.

Data mining has been recently experiencing a boom of interest from researchers and software producers. In medicine, however, its applications are still rather rare. In this paper, we argue that this is primarily due to the requirements of reproducibility of results and diversity of available data mining tools, both of which are crucially important for medical research. We propose to tackle the diversity requirement by means of distributed object technologies. The results presented here rely on our experience with medical data mining using the method GUHA and with further developing that method.

Artificial Intelligence↗

Medical data mining: knowledge discovery in a clinical data warehouse.

Clinical databases have accumulated large quantities of information about patients and their medical conditions. Relationships and patterns within this data could provide new medical knowledge. Unfortunately, few methodologies have been developed and applied to discover this hidden knowledge. In this study, the techniques of data mining (also known as Knowledge Discovery in Databases) were used to search for relationships in a large clinical database. Specifically, data accumulated on 3,902 obstetrical patients were evaluated for factors potentially contributing to preterm birth using exploratory factor analysis. Three factors were identified by the investigators for further exploration. This paper describes the processes involved in mining a clinical database including data warehousing, data query and cleaning, and data analysis.

Data Interpretation, Statistical↗

Data mining in spontaneous reports.

The increasing size of spontaneous report data sets and the increasing capability for screening such data due to increases in computational power has led to a recent increase in interest and use of data mining on such data. While data mining plays an important role in the analysis of spontaneous reports, there is general debate on how and when data mining should be best performed. While the cornerstone principles for data mining of spontaneous reports have been in place since the 1960s, several significant changes have occurred to make their use widespread. Superficially the Bayesian methods seem unnecessarily complex, particularly given the nature of the data, but in practice implementation in Bayesian framework gives clear benefits. There are difficulties evaluating the performance of the methods, but they work and save resources in managing large data sets. The use of neural networks allows more sophisticated pattern recognition to be performed.

Adverse Drug Reaction Reporting Systems↗