PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Data Mining”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Predicting survival causes after out of hospital cardiac arrest using data mining method.

BACKGROUND: The prognosis of life for patients with heart failure remains poor. By using data mining methods, the purpose of this study was to evaluate the most important criteria for predicting patient survival and to profile patients to estimate their survival chances together with the most appropriate technique for health care. METHODS: Five hundred and thirty three patients who had suffered from cardiac arrest were included in the analysis. We performed classical statistical analysis and data mining analysis using mainly Bayesian networks. RESULTS: The mean age of the 533 patients was 63 (+/- 17) and the sample was composed of 390 (73 %) men and 143 (27 %) women. Cardiac arrest was observed at home for 411 (77 %) patients, in a public place for 62 (12 %) patients and on a public highway for 60 (11 %) patients. The belief network of the variables showed that the probability of remaining alive after heart failure is directly associated to five variables: age, sex, the initial cardiac rhythm, the origin of the heart failure and specialized resuscitation techniques employed. CONCLUSIONS: Data mining methods could help clinicians to predict the survival of patients and then adapt their practices accordingly. This work could be carried out for each medical procedure or medical problem and it would become possible to build a decision tree rapidly with the data of a service or a physician. The comparison between classic analysis and data mining analysis showed us the contribution of the data mining method for sorting variables and quickly conclude on the importance or the impact of the data and variables on the criterion of the study. The main limit of the method is knowledge acquisition and the necessity to gather sufficient data to produce a relevant model.

Bayes Theorem↗

Data mining applications in the context of casemix.

In October 1999, the Singapore Government introduced casemix-based funding to public hospitals. The casemix approach to health care funding is expected to yield significant benefits, including equity and rationality in financing health care, the use of comparative casemix data for quality improvement activities, and the provision of information that enables hospitals to understand their cost behaviour and reinforces the drive for more cost-efficient services. However, there is some concern about the "quicker and sicker" syndrome (that is, the rapid discharge of patients with little regard for the quality of outcome). As it is likely that consequences of premature discharges will be reflected in the readmission data, an analysis of possible systematic patterns in readmission data can provide useful insight into the "quicker and sicker" syndrome. This paper explores potential data mining applications in the context of casemix by using readmission data as an illustration. In particular, it illustrates how data mining can be used to better understand readmission data and to detect systematic patterns, if any. From a technical perspective, data mining (which is capable of analysing complex non-linear and interaction relationships) supplements and complements traditional statistical methods in data analysis. From an applications perspective, data mining provides the technology and methodology to analyse mass volume of data to detect hidden patterns in data. Using readmission data as an illustrative data mining application, this paper explores potential data mining applications in the general casemix context.

Cluster Analysis↗

Using data mining to explore complex clinical decisions: A study of hospitalization after a suicide attempt.

BACKGROUND: Medical education is moving toward developing guidelines using the evidence-based approach; however, controlled data are missing for answering complex treatment decisions such as those made during suicide attempts. A new set of statistical techniques called data mining (or machine learning) is being used by different industries to explore complex databases and can be used to explore large clinical databases. METHOD: The study goal was to reanalyze, using data mining techniques, a published study of which variables predicted psychiatrists' decisions to hospitalize in 509 suicide attempters over the age of 18 years who were assessed in the emergency department. Patients were recruited for the study between 1996 and 1998. Traditional multivariate statistics were compared with data mining techniques to determine variables predicting hospitalization. RESULTS: Five analyses done by psychiatric researchers using traditional statistical techniques classified 72% to 88% of patients correctly. The model developed by researchers with no psychiatric knowledge and employing data mining techniques used 5 variables (drug consumption during the attempt, relief that the attempt was not effective, lack of family support, being a housewife, and family history of suicide attempts) and classified 99% of patients correctly (99% sensitivity and 100% specificity). CONCLUSIONS: This reanalysis of a published study fundamentally tries to make the point that these new multivariate techniques, called data mining, can be used to study large clinical databases in psychiatry. Data mining techniques may be used to explore important treatment questions and outcomes in large clinical databases and to help develop guidelines for problems where controlled data are difficult to obtain. New opportunities for good clinical research may be developed by using data mining analyses.

Adult↗

Perspectives on the use of data mining in pharmaco-vigilance.

In the last 5 years, regulatory agencies and drug monitoring centres have been developing computerised data-mining methods to better identify reporting relationships in spontaneous reporting databases that could signal possible adverse drug reactions. At present, there are no guidelines or standards for the use of these methods in routine pharmaco-vigilance. In 2003, a group of statisticians, pharmaco-epidemiologists and pharmaco-vigilance professionals from the pharmaceutical industry and the US FDA formed the Pharmaceutical Research and Manufacturers of America-FDA Collaborative Working Group on Safety Evaluation Tools to review best practices for the use of these methods.In this paper, we provide an overview of: (i) the statistical and operational attributes of several currently used methods and their strengths and limitations; (ii) information about the characteristics of various postmarketing safety databases with which these tools can be deployed; (iii) analytical considerations for using safety data-mining methods and interpreting the results; and (iv) points to consider in integration of safety data mining with traditional pharmaco-vigilance methods. Perspectives from both the FDA and the industry are provided. Data mining is a potentially useful adjunct to traditional pharmaco-vigilance methods. The results of data mining should be viewed as hypothesis generating and should be evaluated in the context of other relevant data. The availability of a publicly accessible global safety database, which is updated on a frequent basis, would further enhance detection and communication about safety issues.

Adverse Drug Reaction Reporting Systems↗

Application of an empiric Bayesian data mining algorithm to reports of pancreatitis associated with atypical antipsychotics.

STUDY OBJECTIVE: To compare the results from one frequently cited data mining algorithm with those from a study, which was published in a peer-reviewed journal, that examined the association of pancreatitis with selected atypical antipsychotics observed by traditional rule-based methods of signal detection. DESIGN: Retrospective pharmacovigilance study. INTERVENTION: The widely studied data mining algorithm known as the Multi-item Gamma Poisson Shrinker (MGPS) was applied to adverse-event reports from the United States Food and Drug Administration's Adverse Event Reporting System database through the first quarter of 2003 for clozapine, olanzapine, and risperidone to determine if a significant signal of pancreatitis would have been generated by this method in advance of their review or the addition of these events to the respective product labels. MEASUREMENTS AND MAIN RESULTS: Data mining was performed by using nine preferred terms relevant to drug-induced pancreatitis from the Medical Dictionary for Regulatory Activities (MedDRA). Results from a previous study on the antipsychotics were reviewed and analyzed. Physicians' Desk References (PDRs) starting from 1994 were manually reviewed to determine the first year that pancreatitis was listed as an adverse event in the product label for each antipsychotic. This information was used as a surrogate marker of the timing of initial signal detection by traditional criteria. Pancreatitis was listed as an adverse event in a PDR for all three atypical antipsychotics. Despite the presence of up to 88 reports/drug-event combination in the Food and Drug Administration's Adverse Event Reporting System database, the MGPS failed to generate a signal of disproportional reporting of pancreatitis associated with the three antipsychotics despite the signaling of these drug-event combinations by traditional rule-based methods, as reflected in product labeling and/or the literature. These discordant findings illustrate key principles in the application of data mining algorithms to drug safety surveillance. CONCLUSION: The optimal place of data mining algorithms in the pharmacovigilance tool kit remains to be determined, requires consideration of numerous factors that may affect their performance, and is highly situation dependent.

Adverse Drug Reaction Reporting Systems↗

Learning from the data: mining of large high-throughput screening databases.

High-throughput screening (HTS) campaigns in pharmaceutical companies have accumulated a large amount of data for several million compounds over a couple of hundred assays. Despite the general awareness that rich information is hidden inside the vast amount of data, little has been reported for a systematic data mining method that can reliably extract relevant knowledge of interest for chemists and biologists. We developed a data mining approach based on an algorithm called ontology-based pattern identification (OPI) and applied it to our in-house HTS database. We identified nearly 1500 scaffold families with statistically significant structure-HTS activity profile relationships. Among them, dozens of scaffolds were characterized as leading to artifactual results stemming from the screening technology employed, such as assay format and/or readout. Four types of compound scaffolds can be characterized based on this data mining effort: tumor cytotoxic, general toxic, potential reporter gene assay artifact, and target family specific. The OPI-based data mining approach can reliably identify compounds that are not only structurally similar but also share statistically significant biological activity profiles. Statistical tests such as Kruskal-Wallis test and analysis of variance (ANOVA) can then be applied to the discovered scaffolds for effective assignment of relevant biological information. The scaffolds identified by our HTS data mining efforts are an invaluable resource for designing SAR-robust diversity libraries, generating in silico biological annotations of compounds on a scaffold basis, and providing novel target family specific scaffolds for focused compound library design.

Algorithms↗

Data mining techniques applied to medical information.

Knowledge discovery from the dramatically increased data of an auto-stored medical information system is still in its infancy. The purpose of this study is to use widely available and easily operated techniques that can satisfy general users in extracting specific knowledge to make the medical information system more functional. Data mining techniques, including data visualisation, correlation analysis, discriminant analysis, and neural networks supervised classification, were applied to heart disease databases. These techniques can help to identify high risk patients, define the most important factors (variables) in heart disease, and build a multivariate relationship model to show the relationship between any two variables in a way that such relationships are easy to view. Simple visualization techniques were utilised to construct this model, which corresponds with current medical knowledge. Two nonparametric (distribution assumption free) classification tools were employed to identify high risk heart disease patients. Both the neural networks supervised classification methods and the discriminant analysis method produced reliable classification rates for heart disease patients. However, neural networks yielded a higher percentage of correct classifications (averaging 89%) than discriminant analysis (79%). Data visualisation and correlation analysis resulted in similar conclusions regarding the most important factors in heart disease. These data mining tools provide simple and effective methods of extracting knowledge from general medical information. The treatment of missing data is also discussed.

Computer Graphics↗

Data mining: a strategy for knowledge development and structure in nursing practice.

Data mining is an emerging technique used more widely by the business world than the world of nursing and health care. However, this strategy can be helpful for improving the quality of decision making by clinicians and health care administrators. This paper addresses the concepts and techniques of data mining that could be useful for practicing nurses as well as nurse administrators. Data mining can be an important tool for the development of nursing knowledge and knowledge structures. An example of the use of the technique in an inpatient setting is provided and insights from the process are discussed.

Cluster Analysis↗

Potential utility of data-mining algorithms for early detection of potentially fatal/disabling adverse drug reactions: a retrospective evaluation.

The objective of this study was to apply 2 data-mining algorithms to a drug safety database to determine if these methods would have flagged potentially fatal/disabling adverse drug reactions that triggered black box warnings/drug withdrawals in advance of initial identification via "traditional" methods. Relevant drug-event combinations were identified from a journal publication. Data-mining algorithms using commonly cited disproportionality thresholds were then applied to the US Food and Drug Administration database. Seventy drug-event combinations were considered sufficiently specific for retrospective data mining. In a minority of instances, potential signals of disproportionate reporting were provided clearly in advance of initial identification via traditional pharmacovigilance methods. Data-mining algorithms have the potential to improve pharmacovigilance screening; however, for the majority of drug-event combinations, there was no substantial benefit of either over traditional methods. They should be considered as potential supplements to, and not substitutes for, traditional pharmacovigilance strategies. More research and experience will be needed to optimize deployment of data-mining algorithms in pharmacovigilance.

Adverse Drug Reaction Reporting Systems↗

Using ontologies in PROTEUS for modeling proteomics data mining applications.

Bioinformatics applications are often characterized by a combination of (pre) processing of raw data representing biological elements, (e.g. sequence alignment, structure prediction), and an high level data mining analysis. Developing such applications needs knowledge of both data mining and bioinformatics domains, that can be effectively achieved by combining ontology about the application domain and ontology about the approaches and processes to solve the given problem. In this paper we talk about using ontologies to model proteomics in silico experiments. In particular data mining of mass spectrometry proteomics data is considered.

Computational Biology↗

Data-mining methods as useful tools for predicting individual drug response: application to CYP2D6 data.

OBJECTIVES: Selecting a maximally informative subset of polymorphisms to predict a clinical outcome, such as drug response, requires appropriate search methods due to the increased dimensionality associated with looking at multiple genotypes. In this study, we investigated the ability of several pattern recognition methods to identify the most informative markers in the CYP2D6 gene for the prediction of CYP2D6 metabolizer status. METHODS: Four data-mining tools were explored: decision trees, random forests, artificial neural networks, and the multifactor dimensionality reduction (MDR) method. Marker selection was performed separately in eight population samples of different ethnic origin to evaluate to what extent the most informative markers differ across ethnic groups. RESULTS: Our results show that the number of polymorphisms required to predict CYP2D6 metabolic phenotype with a high accuracy can be dramatically reduced owing to the strong haplotype block structure observed at CYP2D6. MDR and neural networks provided nearly identical results and performed the best. CONCLUSION: Data-mining methods, such as MDR and neural networks, appear as promising tools to improve the efficiency of genotyping tests in pharmacogenetics with the ultimate goal of pre-screening patients for individual therapy selection with minimum genotyping effort.

Cytochrome P-450 CYP2D6↗

A survey of data mining methods for linkage disequilibrium mapping.

Data mining methods are gaining more interest as potential tools in mapping and identification of complex disease loci. The methods are well suited to large numbers of genetic marker loci produced by high-throughput laboratory analyses, but also might be useful for clarifying the phenotype definitions prior to more traditional mapping analyses. Here, the current data mining-based methods for linkage disequilibrium mapping and phenotype analyses are reviewed.

Chromosome Mapping↗

Using questions and interests to guide data mining for medical quality management.

The healthcare sector is currently facing both the economic necessity and the technical opportunity of a data based approach to quality management. Against this background, a process model for such a data based medical quality management is proposed and intelligent data mining methods are applied to patient data. Intelligent data mining incorporates advantages of both knowledge acquisition from data and from experts. A controlled language for business questions is presented which abstracts from database and data mining terminology to allow high-level interaction. Objective and subjective interestingness of results is measured and used to filter and sort.

Artificial Intelligence↗

[Introduction to medical data mining].

Modern medicine generates a great deal of information stored in the medical database. Extracting useful knowledge and providing scientific decision-making for the diagnosis and treatment of disease from the database increasingly becomes necessary. Data mining in medicine can deal with this problem. It can also improve the management level of hospital information and promote the development of telemedicine and community medicine. Because the medical information is characteristic of redundancy, multi-attribution, incompletion and closely related with time, medical data mining differs from other one. In this paper we have discussed the key techniques of medical data mining involving pretreatment of medical data, fusion of different pattern and resource, fast and robust mining algorithms and reliability of mining results. The methods and applications of medical data mining based on computation intelligence such as artificial neural network, fuzzy system, evolutionary algorithms, rough set, and support vector machine have been introduced. The features and problems in data mining are summarized in the last section.

Algorithms↗

Dental data mining: potential pitfalls and practical issues.

Knowledge Discovery and Data Mining (KDD) have become popular buzzwords. But what exactly is data mining? What are its strengths and limitations? Classic regression, artificial neural network (ANN), and classification and regression tree (CART) models are common KDD tools. Some recent reports (e.g., Kattan et al., 1998) show that ANN and CART models can perform better than classic regression models: CART models excel at covariate interactions, while ANN models excel at nonlinear covariates. Model prediction performance is examined with the use of validation procedures and evaluating concordance, sensitivity, specificity, and likelihood ratio. To aid interpretation, various plots of predicted probabilities are utilized, such as lift charts, receiver operating characteristic curves, and cumulative captured-response plots. A dental caries study is used as an illustrative example. This paper compares the performance of logistic regression with KDD methods of CART and ANN in analyzing data from the Rochester caries study. With careful analysis, such as validation with sufficient sample size and the use of proper competitors, problems of naïve KDD analyses (Schwarzer et al., 2000) can be carefully avoided.

Child↗

Data mining.

Group 14 used data-mining strategies to evaluate a number of issues, including appropriate diagnosis, haplotype estimation, genetic linkage and association studies, and type I error. Methods ranged from exploratory analyses, to machine learning strategies (neural networks, supervised learning, and tree-based methods), to false discovery rate control of type I errors. The general motivations were to find the "story" in the data and to summarize information from a multitude of measures. Several methods illustrated strategies for better trait definition, using summarization of related traits. In the few studies that sought to identify genes for alcoholism, there was little agreement among the different strategies, likely reflecting the complexities of the disease. Nevertheless, Group 14 found that these methods offered strategies to gain a better understanding of the complex pathways by which disease develops.

Alcoholism↗

GeneKeyDB: a lightweight, gene-centric, relational database to support data mining environments.

BACKGROUND: The analysis of biological data is greatly enhanced by existing or emerging databases. Most existing databases, with few exceptions are not designed to easily support large scale computational analysis, but rather offer exclusively a web interface to the resource. We have recognized the growing need for a database which can be used successfully as a backend to computational analysis tools and pipelines. Such database should be sufficiently versatile to allow easy system integration. RESULTS: GeneKeyDB is a gene-centered relational database developed to enhance data mining in biological data sets. The system provides an underlying data layer for computational analysis tools and visualization tools. GeneKeyDB relies primarily on existing database identifiers derived from community databases (NCBI, GO, Ensembl, et al.) as well as the known relationships among those identifiers. It is a lightweight, portable, and extensible platform for integration with computational tools and analysis environments. CONCLUSION: GeneKeyDB can enable analysis tools and users to manipulate the intersections, unions, and differences among different data sets.

Algorithms↗

Accurate prediction of protein functional class from sequence in the Mycobacterium tuberculosis and Escherichia coli genomes using data mining.

The analysis of genomics data needs to become as automated as its generation. Here we present a novel data-mining approach to predicting protein functional class from sequence. This method is based on a combination of inductive logic programming clustering and rule learning. We demonstrate the effectiveness of this approach on the M. tuberculosis and E. coli genomes, and identify biologically interpretable rules which predict protein functional class from information only available from the sequence. These rules predict 65% of the ORFs with no assigned function in M. tuberculosis and 24% of those in E. coli, with an estimated accuracy of 60-80% (depending on the level of functional assignment). The rules are founded on a combination of detection of remote homology, convergent evolution and horizontal gene transfer. We identify rules that predict protein functional class even in the absence of detectable sequence or structural homology. These rules give insight into the evolutionary history of M. tuberculosis and E. coli.

Amino Acid Sequence↗