PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Data Mining”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Data mining approach to policy analysis in a health insurance domain.

This study examined the characteristics of the knowledge discovery and data mining algorithms to demonstrate how they can be used to predict health outcomes and provide policy information for hypertension management using the Korea Medical Insurance Corporation database. Specifically, this study validated the predictive power of data mining algorithms by comparing the performance of logistic regression and two decision tree algorithms, CHIAD (Chi-squared Automatic Interaction Detection) and C5.0 (a variant of C4.5) using the test set of 4588 beneficiaries and the training set of 13,689 beneficiaries. Contrary to the previous study, the CHIAD algorithm performed better than the logistic regression in predicting hypertension, and C5.0 had the lowest predictive power. In addition, the CHIAD algorithm and the association rule also provided the segment-specific information for the risk factors and target group that may be used in a policy analysis for hypertension management.

Algorithms↗

Using data mining to build a customer-focused organization.

Data mining is a new buzz word in managed care. More than simply a method of unlocking a vault of useful information in MCO data banks and warehouses, the author believes that it can help steer an organization through the reengineering process, leading to the health system's transformation toward a customer-focused organization.

Algorithms↗

SPINE: an integrated tracking database and data mining approach for identifying feasible targets in high-throughput structural proteomics.

High-throughput structural proteomics is expected to generate considerable amounts of data on the progress of structure determination for many proteins. For each protein this includes information about cloning, expression, purification, biophysical characterization and structure determination via NMR spectroscopy or X-ray crystallography. It will be essential to develop specifications and ontologies for standardizing this information to make it amenable to retrospective analysis. To this end we created the SPINE database and analysis system for the Northeast Structural Genomics Consortium. SPINE, which is available at bioinfo.mbb.yale.edu/nesg or nesg.org, is specifically designed to enable distributed scientific collaboration via the Internet. It was designed not just as an information repository but as an active vehicle to standardize proteomics data in a form that would enable systematic data mining. The system features an intuitive user interface for interactive retrieval and modification of expression construct data, query forms designed to track global project progress and external links to many other resources. Currently the database contains experimental data on 985 constructs, of which 740 are drawn from Methanobacterium thermoautotrophicum, 123 from Saccharomyces cerevisiae, 93 from Caenorhabditis elegans and the remainder from other organisms. We developed a comprehensive set of data mining features for each protein, including several related to experimental progress (e.g. expression level, solubility and crystallization) and 42 based on the underlying protein sequence (e.g. amino acid composition, secondary structure and occurrence of low complexity regions). We demonstrate in detail the application of a particular machine learning approach, decision trees, to the tasks of predicting a protein's solubility and propensity to crystallize based on sequence features. We are able to extract a number of key rules from our trees, in particular that soluble proteins tend to have significantly more acidic residues and fewer hydrophobic stretches than insoluble ones. One of the characteristics of proteomics data sets, currently and in the foreseeable future, is their intermediate size ( approximately 500-5000 data points). This creates a number of issues in relation to error estimation. Initially we estimate the overall error in our trees based on standard cross-validation. However, this leaves out a significant fraction of the data in model construction and does not give error estimates on individual rules. Therefore, we present alternative methods to estimate the error in particular rules.

Animals↗

A general 13C NMR spectrum predictor using data mining techniques.

A general-case neural network model for 13C NMR spectrum prediction (estimation) was built from more than 8,300 carbon atoms having various environments. Building the model from the data set required a few weeks' work using commercial software. Average deviation on test data is ca. 4 ppm. There is no limit on molecule complexity. Estimation error does not depend on molecule size or complexity. The emphasis is on the data, the method and the results, not on the processes that take place inside the modelling software. Advantages, disadvantages and peculiarities of neural network-based data modelling ("data mining") are described at length. The differences in data handling between the data mining approach and traditional statistical modelling techniques are discussed and illustrated in detail. The spectrum predictor is available from PMSI at no charge.

Carbon Isotopes↗

Comparison of data mining methodologies using Japanese spontaneous reports.

PURPOSE: Five data mining methodologies for detecting a possible signal from spontaneous reports on adverse drug reactions (ADRs) were compared. METHODS: The five methodologies, the Bayesian method using the Gamma Poissson Shrinker (GPS), the method employed in the UK Medicines Control Agency (MCA), the Bayesian Confidence Propagation Neural Network (BCPNN), the method using the 95% confidence interval (CI) for the reporting odds ratio (RORCI) and that using the 95% CI of the proportional reporting ratio (PRRCI) were compared using Japanese data obtained between 1998 and 2000. RESULTS: There were all in all 38,731 drug-ADR combinations. The count of drug-ADR pairs was equal to 1 or 2 for 31,230 combinations and none of them were identified as a possible signal with the MCA or BCPNN. Similarly, the GPS detected a possible signal in none of the combinations where the count was equal to 1 but in 7.5% of the combinations where the count was equal to 2. The RORCI and PRRCI detected a possible signal in more than half of the combinations where the count was equal to 1 or 2. When the pairwise agreement on whether or not a drug-ADR combination satisfied the criteria for a possible signal was assessed for the 38,731 combinations, the concordance measure kappa was greater than 0.9 between the MCA and BCPNN and between the RORCI and PRRCI. Kappa was around 0.6 between the GPS and MCA and between the GPS and BCPNN. Otherwise, kappa was smaller than 0.2. CONCLUSIONS: The drug-ADR combinations detected as a possible signal vary between different methodologies.

Adverse Drug Reaction Reporting Systems↗

Using fragment chemistry data mining and probabilistic neural networks in screening chemicals for acute toxicity to the fathead minnow.

The paper is illustrating how the general data mining methodology may be adapted to provide solutions to the problem of high throughput virtual screening of organic chemicals for possible acute toxicity to the fathead minnow fish. The present approach involves mining fragment information from chemical structures and is using probabilistic neural networks to model the relationship between structure and toxicity. Probabilistic neural networks implement a special class of multivariate non-linear Bayesian statistical models. The mathematical principles supporting their use for value prediction purposes are clarified and their peculiarities discussed. As part of the research phase of the data mining process, a dataset consisting of 800 structures and associated fathead minnow (Pimephales promelas) 96-h LC50 acute toxicity endpoint information is used for both the purpose of identifying an advantageous combination of fragment descriptors and for training the neural networks. As a result, two powerful models are generated. Model 1 implements the basic PNN with Gaussian kernel (statistical corrections included) while Model 2 implements the PNN with Gaussian kernel and separated variables. External validation is performed using a separate dataset consisting of 86 structures and associated toxicity information. Both learning and generalization capabilities of the two models are investigated and their limitations discussed.

Animals↗

Early postmarketing drug safety surveillance: data mining points to consider.

BACKGROUND: Computer-assisted data mining algorithms (DMAs) are being studied to screen spontaneous reporting databases for signals of novel adverse events. The performance characteristics and optimum deployment of these techniques remain to be established. OBJECTIVE: To explore issues in the practical evaluation and deployment of DMAs by comparing findings from an empirical Bayesian DMA with those from a traditional drug safety surveillance program. METHODS: Published findings from early postmarketing safety surveillance of thalidomide were compared with findings from an empirical Bayesian DMA. Differential results were used to explore practical issues in the evaluation and deployment of DMAs. RESULTS: Most adverse events highlighted by each method were compatible with the product labeling or natural history/complications of reported treatment indications. Traditional surveillance highlighted 4 potentially serious and unexpected adverse events (Stevens-Johnson syndrome, toxic epidermal necrolysis, seizures, skin ulcers) warranting labeling amendments or close monitoring. None of these adverse event terms generated a signal using the DMA. CONCLUSIONS: The DMA would not have enhanced early postmarketing surveillance in this particular setting. While the results cannot be used to draw inferences about the global performance of DMAs, they illustrate the following: (1) DMA performance may be highly situation dependent; (2) over-reliance on these methods may have deleterious consequences, especially with so-called "designated medical events"; and (3) the most appropriate selection of pharmacovigilance tools needs to be tailored to each situation, being mindful of the numerous factors that may influence comparative performance and incremental utility of DMAs.

Algorithms↗

Data mining applied to linkage disequilibrium mapping.

We introduce a new method for linkage disequilibrium mapping: haplotype pattern mining (HPM). The method, inspired by data mining methods, is based on discovery of recurrent patterns. We define a class of useful haplotype patterns in genetic case-control data and use the algorithm for finding disease-associated haplotypes. The haplotypes are ordered by their strength of association with the phenotype, and all haplotypes exceeding a given threshold level are used for prediction of disease susceptibility-gene location. The method is model-free, in the sense that it does not require (and is unable to utilize) any assumptions about the inheritance model of the disease. The statistical model is nonparametric. The haplotypes are allowed to contain gaps, which improves the method's robustness to mutations and to missing and erroneous data. Experimental studies with simulated microsatellite and SNP data show that the method has good localization power in data sets with large degrees of phenocopies and with lots of missing and erroneous data. The power of HPM is roughly identical for marker maps at a density of 3 single-nucleotide polymorphisms/cM or 1 microsatellite/cM. The capacity to handle high proportions of phenocopies makes the method promising for complex disease mapping. An example of correct disease susceptibility-gene localization with HPM is given with real marker data from families from the United Kingdom affected by type 1 diabetes. The method is extendable to include environmental covariates or phenotype measurements or to find several genes simultaneously.

Adolescent↗

[Data mining in diagnostic knowledge acquisition from patients with brain glioma].

In order to correctly predict the malignant degree of brain glioma, three data mining algorithms: multi-layer perceptron network(MLP), decision tree, and rule induction are adopted to acquire diagnostic knowledge from patients with brain glioma cases. Totally 280 cases are collected, and some of them contain missing values. Preprocessing is taken to make them applicable to all three algorithms. Performance comparisons are carried out with a 10-fold cross validation test. Although the result of MLP is hard to be understood and cannot be applied directly, its reliability and accuracy are the highest when only a few hidden nodes are involved. Unlike MLP, both decision tree and rule induction use attribute-value pairs to represent diagnostic knowledge derived from treated cases. These could improve both the understandability and applicability of their results. When compared with rule induction, the inherent restriction in structure makes decision tree more efficient in decision-making but meanwhile hurts its simplicity, accuracy, and reliability. For testing samples, results of all these algorithms can achieve accuracy rate over 80%, which satisfies the basic requirement of neuroradiologists. If diagnostic accuracy rate is the main factor to be considered, MLP with only a few hidden nodes is the best. If the result is expected to be further checked or evaluated, rule induction will be the best algorithm. This work proves that data mining techniques can be used to obtain valid diagnostic knowledge from brain glioma cases and make computer aided diagnosis system in this field feasible.

Algorithms↗

Term domain distribution analysis: a data mining tool for text databases.

In this paper, we give a case history illustrating the real-world application of a useful technique for data mining of text databases. The technique, which we call Term Domain Distribution Analysis (TDDA), consists of keeping track of term frequencies for specific finite domains and announcing significant differences from standard frequency distributions over these domains as a hypothesis. TDDA is part of a larger framework, the Digital Filter Model, for data mining of text documents. In the case study presented, the domain of terms was the pair {right, left}, over which we expected a uniform distribution. In analyzing term frequencies in a thoracic lung cancer database, the TDDA technique led to the surprising discovery that primary thoracic lung cancer tumors appear in the right lung more often than the left lung, with a ratio of 3:2. Treating the text discovery as a hypothesis, we verified this relationship against the medical literature in which primary lung tumor sites were reported, using a standard chi 2 statistic. We subsequently developed a working theoretical model of lung cancer that may explain the discovery. This discovery and our model may change how oncologists view the mechanisms of primary lung tumor location.

Humans↗

Association rules and data mining in hospital infection control and public health surveillance.

OBJECTIVES: The authors consider the problem of identifying new, unexpected, and interesting patterns in hospital infection control and public health surveillance data and present a new data analysis process and system based on association rules to address this problem. DESIGN: The authors first illustrate the need for automated pattern discovery and data mining in hospital infection control and public health surveillance. Next, they define association rules, explain how those rules can be used in surveillance, and present a novel process and system--the Data Mining Surveillance System (DMSS)--that utilize association rules to identify new and interesting patterns in surveillance data. RESULTS: Experimental results were obtained using DMSS to analyze Pseudomonas aeruginosa infection control data collected over one year (1996) at University of Alabama at Birmingham Hospital. Experiments using one-, three-, and six-month time partitions yielded 34, 57, and 28 statistically significant events, respectively. Although not all statistically significant events are clinically significant, a subset of events generated in each analysis indicated potentially significant shifts in the occurrence of infection or antimicrobial resistance patterns of P. aeruginosa. CONCLUSION: The new process and system are efficient and effective in identifying new, unexpected, and interesting patterns in surveillance data. The clinical relevance and utility of this process await the results of prospective studies currently in progress.

Data Interpretation, Statistical↗

Data mining of molecular dynamics trajectories of nucleic acids.

Analysis, storage, and transfer of molecular dynamic trajectories are becoming the bottleneck of computer simulations. In this paper we discuss different approaches for data mining and data processing of huge trajectory files generated from molecular dynamic simulations of nucleic acids.

Computer Simulation↗

The small-world dynamics of tree networks and data mining in phyloinformatics.

MOTIVATION: A noble and ultimate objective of phyloinformatic research is to assemble, synthesize, and explore the evolutionary history of life on earth. Data mining methods for performing these tasks are not yet well developed, but one avenue of research suggests that network connectivity dynamics will play an important role in future methods. Analysis of disordered networks, such as small-world networks, has applications as diverse as disease propagation, collaborative networks, and power grids. Here we apply similar analyses to networks of phylogenetic trees in order to understand how synthetic information can emerge from a database of phylogenies. RESULTS: Analyses of tree network connectivity in TreeBASE show that a collection of phylogenetic trees behaves as a small-world network-while on the one hand the trees are clustered, like a non-random lattice, on the other hand they have short characteristic path lengths, like a random graph. Tree connectivities follow a dual-scale power-law distribution (first power-law exponent approximately 1.87; second approximately 4.82). This unusual pattern is due, in part, to the presence of alternative tree topologies that enter the database with each published study. As expected, small collections of trees decrease connectivity as new trees are added, while large collections of trees increase connectivity. However, the inflection point is surprisingly low: after about 600 trees the network suddenly jumps to a higher level of coherence. More stringent definitions of 'neighbour' greatly delay the threshold whence a database achieves sufficient maturity for a coherent network to emerge. However, more stringent definitions of 'neighbour' would also likely show improved focus in data mining. AVAILABILITY: http://treebase.org

Algorithms↗

A data mining approach to the development of a diagnostic test for male infertility.

The paper presents a database of published Y chromosome deletions and the results of analyzing the database with data mining and other heuristic techniques with the goal of developing a diagnostic test for male infertility. The database describes 382 patients for which 177 markers were tested. Two data mining techniques, clustering and decision tree induction were used, as well as a heuristic set cover algorithm. Clustering was used to group markers according to their appearance across patients, while a heuristic set covering algorithm was used to select as small a set of markers that cover as many patients with deletions as possible. This algorithm created a diagnostic set of 13 markers that cover more than 90% of the patients with deletions. Finally, decision tree induction was used to relate deletion patterns to the severity of the clinical phenotype. A decision tree induced from the data uses 5 markers, all of which are also in the diagnostic set of 13 markers, to show relations between the severity of the clinical phenotype and deletion patterns which have not been known previously.

Algorithms↗

Using data mining and OLAP to discover patterns in a database of patients with Y-chromosome deletions.

The paper presents a database of published Y chromosome deletions and the results of analyzing the database with data mining techniques. The database describes 382 patients for which 177 different markers were tested: 364 of the 382 patients had deletions. Two data mining techniques, clustering and decision tree induction were used. Clustering was used to group patients according to the overall presence/absence of deletions at the tested markers. Decision trees and On-Line-Analytical-Processing (OLAP) were used to inspect the resulting clustering and look for correlations between deletion patterns, populations and the clinical picture of infertility. The results of the analysis indicate that there are correlations between deletion patterns and patient populations, as well as clinical phenotype severity.

Chromosome Deletion↗

Cancer surveillance using data warehousing, data mining, and decision support systems.

This article discusses how data warehousing, data mining, and decision support systems can reduce the national cancer burden or the oral complications of cancer therapies, especially as related to oral and pharyngeal cancers. An information system is presented that will deliver the necessary information technology to clinical, administrative, and policy researchers and analysts in an effective and efficient manner. The system will deliver the technology and knowledge that users need to readily: (1) organize relevant claims data, (2) detect cancer patterns in general and special populations, (3) formulate models that explain the patterns, and (4) evaluate the efficacy of specified treatments and interventions with the formulations. Such a system can be developed through a proven adaptive design strategy, and the implemented system can be tested on State of Maryland Medicaid data (which includes women, minorities, and children).

Database Management Systems↗

Extraction of substructures of proteins essential to their biological functions by a data mining technique.

Correlation between the sequential, structural, and functional features of proteins is one of the most important open questions in the field of molecular biology. To this problem, we apply a technique known as data mining for discovering associations across protein sequence, structure, and function. We were able to find various association rules on the substructures essential to some protein functions. Moreover, structure-structure associations were found between proteins having different functions. The results suggest that data mining might be a powerful tool in protein analysis.

Algorithms↗

Data mining tools for biological sequences.

We describe a methodology, as well as some related data mining tools, for analyzing sequence data. The methodology comprises three steps: (a) generating candidate features from the sequences, (b) selecting relevant features from the candidates, and (c) integrating the selected features to build a system to recognize specific properties in sequence data. We also give relevant techniques for each of these three steps. For generating candidate features, we present various types of features based on the idea of k-grams. For selecting relevant features, we discuss signal-to-noise, t-statistics, and entropy measures, as well as a correlation-based feature selection method. For integrating selected features, we use machine learning methods, including C4.5, SVM, and Naive Bayes. We illustrate this methodology on the problem of recognizing translation initiation sites. We discuss how to generate and select features that are useful for understanding the distinction between ATG sites that are translation initiation sites and those that are not. We also discuss how to use such features to build reliable systems for recognizing translation initiation sites in DNA sequences.

Artificial Intelligence↗