PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Data Mining”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Data mining parasite genomes.

The term 'data mining' can be used to describe any process where useful information is extracted from data with a large background of 'noise'. In the context of a genome project, several stages involve data mining. Amongst the sequence data, 'signals' need to be detected that indicate the presence of interesting features. Often this involves differentiating between transcribed and non-transcribed bases to predict coding regions. After detection, defining the roles of these sequences involves sifting through multiple lines of evidence. If these roles are accurately reflected in genome annotation, they can be used by researchers to frame queries and interrogate the data further.

Animals↗

Distributed data mining on grids: services, tools, and applications.

Data mining algorithms are widely used today for the analysis of large corporate and scientific datasets stored in databases and data archives. Industry, science, and commerce fields often need to analyze very large datasets maintained over geographically distributed sites by using the computational power of distributed and parallel systems. The grid can play a significant role in providing an effective computational support for distributed knowledge discovery applications. For the development of data mining applications on grids we designed a system called Knowledge Grid. This paper describes the Knowledge Grid framework and presents the toolset provided by the Knowledge Grid for implementing distributed knowledge discovery. The paper discusses how to design and implement data mining applications by using the Knowledge Grid tools starting from searching grid resources, composing software and data components, and executing the resulting data mining process on a grid. Some performance results are also discussed.

Algorithms↗

Application of data mining techniques in pharmacovigilance.

AIMS: To discuss the potential use of data mining and knowledge discovery in databases for detection of adverse drug events (ADE) in pharmacovigilance. METHODS: A literature search was conducted to identify articles, which contained details of data mining, signal generation or knowledge discovery in relation to adverse drug reactions or pharmacovigilance in medical databases. RESULTS: ADEs are common and result in significant mortality, and despite existing systems drugs have been withdrawn due to ADEs many years after licensing. Knowledge discovery in databases (KDD) is a technique which may be used to detect potential ADEs more efficiently. KDD involves the selection of data variables and databases, data preprocessing, data mining and data interpretation and utilization. Data mining encompasses a number of statistical techniques including cluster analysis, link analysis, deviation detection and disproportionality assessment which can be utilized to determine the presence of and to assess the strength of ADE signals. Currently the only data mining methods to be used in pharmacovigilance are those of disproportionality, such as the Proportional Reporting Ratio and Information Component, which have been used to analyse the UK Yellow Card Scheme spontaneous reporting database and the WHO Uppsala Monitoring Centre database. The association of pericarditis with practolol but not with other beta-blockers, the association of captopril and other angiotensin-converting enzymes with cough, and the association of terfenadine with heart rate and rhythm disorders could be identified by mining the WHO database. CONCLUSION: In view of the importance of ADEs and the development of massive data storage systems and powerful computer systems, the use of data mining techniques in knowledge discovery in medical databases is likely to be of increasing importance in the process of pharmacovigilance as they are likely to be able to detect signals earlier than using current methods.

Data Collection↗

Computation of the physio-chemical properties and data mining of large molecular collections.

Very large data sets of molecules screened against a broad range of targets have become available due to the advent of combinatorial chemistry. This information has led to the realization that ADME (absorption, distribution, metabolism, and excretion) and toxicity issues are important to consider prior to library synthesis. Furthermore, these large data sets provide a unique and important source of information regarding what types of molecular shapes may interact with specific receptor or target classes. Thus, the requirement for rapid and accurate data mining tools became paramount. To address these issues Pharmacopeia, Inc. formed a computational research group, The Center for Informatics and Drug Discovery (CIDD).* In this review we cover the work done by this group to address both in silico ADME modeling and data mining issues faced by Pharmacopeia because of the availability of a large and diverse collection (over 6 million discrete compounds) of drug-like molecules. In particular, in the data mining arena we discuss rapid docking tools and how we employ them, and we describe a novel data mining tool based on a ID representation of a molecule followed by a molecular sequence alignment step. For the ADME area we discuss the development and application of absorption, blood-brain barrier (BBB) and solubility models. Finally, we summarize the impact the tools and approaches might have on the drug discovery process.

Algorithms↗

Comparing data mining methods on the VAERS database.

PURPOSE: Data mining may enhance traditional surveillance of vaccine adverse events by identifying events that are reported more commonly after administering one vaccine than other vaccines. Data mining methods find signals as the proportion of times a condition or group of conditions is reported soon after the administration of a vaccine; thus it is a relative proportion compared across vaccines, and not an absolute rate for the condition. The Vaccine Adverse Event Reporting System (VAERS) contains approximately 150 000 reports of adverse events that are possibly associated with vaccine administration. METHODS: We studied four data mining techniques: empirical Bayes geometric mean (EBGM), lower-bound of the EBGM's 90% confidence interval (EB05), proportional reporting ratio (PRR), and screened PRR (SPRR). We applied these to the VAERS database and compared the agreement among methods and other performance properties, particularly focusing on the vaccine-event combinations with the highest numerical scores in the various methods. RESULTS: The vaccine-event combinations with the highest numerical scores varied substantially among the methods. Not all combinations representing known associations appeared in the top 100 vaccine-event pairs for all methods. CONCLUSIONS: The four methods differ in their ranking of vaccine-COSTART pairs. A given method may be superior in certain situations but inferior in others. This paper examines the statistical relationships among the four estimators. Determining which method is best for public health will require additional analysis that focuses on the true alarm and false alarm rates using known vaccine-event associations. Evaluating the properties of these data mining methods will help determine the value of such methods in vaccine safety surveillance.

Adverse Drug Reaction Reporting Systems↗

Application of data mining to intensive care unit microbiologic data.

We describe refinements to and new experimental applications of the Data Mining Surveillance System (DMSS), which uses a large electronic health-care database for monitoring emerging infections and antimicrobial resistance. For example, information from DMSS can indicate potentially important shifts in infection and antimicrobial resistance patterns in the intensive care units of a single health-care facility.

Alabama↗

Data mining in the US Vaccine Adverse Event Reporting System (VAERS): early detection of intussusception and other events after rotavirus vaccination.

The Vaccine Adverse Event Reporting System (VAERS) is the US passive surveillance system monitoring vaccine safety. A major limitation of VAERS is the lack of denominator data (number of doses of administered vaccine), an element necessary for calculating reporting rates. Empirical Bayesian data mining, a data analysis method, utilizes the number of events reported for each vaccine and statistically screens the database for higher than expected vaccine-event combinations signaling a potential vaccine-associated event. This is the first study of data mining in VAERS designed to test the utility of this method to detect retrospectively a known side effect of vaccination-intussusception following rotavirus (RV) vaccine. From October 1998 to December 1999, 112 cases of intussusception were reported. The data mining method was able to detect a signal for RV-intussusception in February 1999 when only four cases were reported. These results demonstrate the utility of data mining to detect significant vaccine-associated events at early date. Data mining appears to be an efficient and effective computer-based program that may enhance early detection of adverse events in passive surveillance systems.

Adverse Drug Reaction Reporting Systems↗

Knowledge discovery in nursing minimum data set using data mining.

PURPOSE: The purposes of this study were to apply data mining tool to nursing specific knowledge discovery process and to identify the utilization of data mining skill for clinical decision making. METHODS: Data mining based on rough set model was conducted on a large clinical data set containing NMDS elements. Randomized 1,000 patient data were selected from year 1998 database which had at least one of the five most frequently used nursing diagnoses. Patient characteristics and care service characteristics including nursing diagnoses, interventions and outcomes were analyzed to derive the meaningful decision rules. RESULTS: Number of comorbidity, marital status, nursing diagnosis related to risk for infection and nursing intervention related to infection protection, and discharge status were the predictors that could determine the length of stay. Four variables (age, impaired skin integrity, pain, and discharge status) were identified as valuable predictors for nursing outcome, relieved pain. Five variables (age, pain, potential for infection, marital status, and primary disease) were identified as important predictors for mortality. CONCLUSIONS: This study demonstrated the utilization of data mining method through a large data set with standardized language format to identify the contribution of nursing care to patient's health.

Adult↗

Uniqueness of medical data mining.

This article addresses the special features of data mining with medical data. Researchers in other fields may not be aware of the particular constraints and difficulties of the privacy-sensitive, heterogeneous, but voluminous data of medicine. Ethical and legal aspects of medical data mining are discussed, including data ownership, fear of lawsuits, expected benefits, and special administrative issues. The mathematical understanding of estimation and hypothesis formation in medical data may be fundamentally different than those from other data collection activities. Medicine is primarily directed at patient-care activity, and only secondarily as a research resource; almost the only justification for collecting medical data is to benefit the individual patient. Finally, medical data have a special status based upon their applicability to all people; their urgency (including life-or-death); and a moral obligation to be used for beneficial purposes.

Computer Security↗

Predictive data mining in clinical medicine: current issues and guidelines.

BACKGROUND: The widespread availability of new computational methods and tools for data analysis and predictive modeling requires medical informatics researchers and practitioners to systematically select the most appropriate strategy to cope with clinical prediction problems. In particular, the collection of methods known as 'data mining' offers methodological and technical solutions to deal with the analysis of medical data and construction of prediction models. A large variety of these methods requires general and simple guidelines that may help practitioners in the appropriate selection of data mining tools, construction and validation of predictive models, along with the dissemination of predictive models within clinical environments. PURPOSE: The goal of this review is to discuss the extent and role of the research area of predictive data mining and to propose a framework to cope with the problems of constructing, assessing and exploiting data mining models in clinical medicine. METHODS: We review the recent relevant work published in the area of predictive data mining in clinical medicine, highlighting critical issues and summarizing the approaches in a set of learned lessons. RESULTS: The paper provides a comprehensive review of the state of the art of predictive data mining in clinical medicine and gives guidelines to carry out data mining studies in this field. CONCLUSIONS: Predictive data mining is becoming an essential instrument for researchers and clinical practitioners in medicine. Understanding the main issues underlying these methods and the application of agreed and standardized procedures is mandatory for their deployment and the dissemination of results. Thanks to the integration of molecular and clinical data taking place within genomic medicine, the area has recently not only gained a fresh impulse but also a new set of complex problems it needs to address.

Clinical Medicine↗

Data mining issues for improved birth outcomes.

Issues obstructing progress in data mining for improved health outcomes include data quality problems, data redundancy, data inconsistency, repeated measures, temporal (time-contextual) measures, and data volume. Related issues involve theoretical and technical problems involving uncertainty management, missing data and missing values, and matching appropriate data mining techniques to patient data sets. Results of data mining research in progress are reported for Duke University's perinatal database that contains nearly a decade of clinical patient data, 71,753 database (patient) records and 4-5000 variables per patient.

Artificial Intelligence↗

Data mining methods for improving birth outcomes prediction.

Data mining is a research method that is increasingly being used to predict clinical outcomes, for example, cancer or AIDS survival, diagnostic accuracy in abdominal pain or brain tumors, and much more. In clinical practice, predicting which patients will deliver preterm versus full term remains a complex clinical problem for families and the healthcare system. Exploratory data mining was used for predicting birth outcomes in a racially diverse sample (n = 19,970). Duke University provided data (1622 variables) for data mining methods that found 7 demographic variables yielded .72 area under the curve for receiver operating characteristic (ROC) analyses, suggesting that a parsimonious set of preterm birth outcomes predictors may be possible. Improved prediction is needed for interventions to be appropriately targeted for improved birth outcomes management.

Databases, Factual↗

The use of Data Mining in the categorization of patients with Azoospermia.

OBJECTIVE: Data Mining is a relatively new field of Medical Informatics. The aim of this study was to compare Data Mining diagnosis with clinical diagnosis by applying a Data Miner (DM) to a clinical dataset of infertile men with azoospermia. DESIGN: One hundred and forty-seven azoospermic men were clinically classified into four groups: a) obstructive azoospermia (n=63), b) non-obstructive azoospermia (n=71), c) hypergonadotropic hypogonadism (n=2), and d) hypogonadotropic hypogonadism (n=11). The DM (IBM's DB2/Intelligent Miner for Data 6.1) was asked to reproduce a four-cluster model. RESULTS: DM formed four groups of patients: a) eugonadal men with normal testicular volume and normal FSH levels (n=86), b) eugonadal men with significantly reduced testicular volume (median 6.5 cm3) and very high FSH levels (n=29), c) eugonadal men with moderately reduced testicular volume (median 14.5 cm3) and raised FSH levels (n=20), and d) hypogonadal men (n=12). Overall DM concordance rate in hypogonadal men was 92%, in obstructive azoospermia 73%, and in non-obstructive azoospermia 69%. CONCLUSIONS: Data Mining produces clinically meaningful results but different from those of the clinical diagnosis. It is possible that the use of large sets of structured and formalised data and continuous evaluation of DM results will generate a useful methodology for the Clinician.

Cohort Studies↗

Using data mining to find fraud in HCFA health care claims.

Data mining can be/used to detect health care fraud and abuse through visualization of very large data sets to isolate new and unusual patterns of activity. Data mining has allowed better direction and use of health care fraud detection and investigative resources by recognizing and quantifying the underlying indicators of fraudulent claims, fraudulent providers, and fraudulent beneficiaries. A large amount of work must be performed prior to the actual data mining. These precursory tasks include: customer discussions, data extraction and cleaning, transformation of the database, and auditing (basic statistics and visualization of the information) of the data. This paper describes the tasks performed in support of a project for HCFA (Health Care Financing Administration).

Centers for Medicare and Medicaid Services, U.S.↗

Applying data mining techniques to library design, lead generation and lead optimization.

Many data mining techniques have been applied to activity and ADMET datasets and the resulting models are being used to understand quantitative structure-activity relationships and design new libraries. This review summarizes data mining concepts and discuss their application to library design, lead generation (particularly for sequential screening) and lead optimization (specifically for generating and interpreting QSAR models). Also, this review discusses recent comparative studies between data mining techniques and draws some conclusions about the patterns emerging in the drug discovery data mining field.

Algorithms↗

Using data mining to address heterogeneity in the Southampton data.

When analyzing complex traits such as asthma, heterogeneity needs to be assumed. With this in mind, to identify a more homogeneous group of asthmatic patients, we analyzed the Southampton data using the data mining technique known as the regression tree method and the two most inheritable quantitative phenotypes (LnIgE and RAST) as the target variables. Two-point and multipoint nonparametric linkage analyses were carried out using one of the subgroups as affected. In addition, we performed quantitative trait loci nonparametric linkage analysis using each phenotype as the outcome. The results from the affected-sibpairs method and quantitative linkage analysis were compared.

Adult↗

Development of quantitative structure-property relationship models for pseudoternary microemulsions formulated with nonionic surfactants and cosurfactants: application of data mining and molecular modeling.

Data mining, computer-aided molecular modeling, descriptor calculation and multiple linear regression techniques were utilized to produce statistically significant and predictive models for O/W and W/O microemulsions. The literature was scanned over the last 20 years, subsequently, 68 phase diagrams from eight different references were collected. Molecular modeling techniques were then applied on the components of the microemulsion systems to generate plausible 3-D structures. Subsequently, various physicochemical descriptors were calculated based on the resulting 3-D structures. The generated descriptors were correlated with microemulsion existence areas utilizing multiple linear regression analysis (MLR). The generated models were statistically cross-validated and were found to be of significant predictive power. Furthermore, the resulting models allowed better understanding of the process of microemulsion formation.

Drug Compounding↗

Data mining issues and opportunities for building nursing knowledge.

Health care information systems tend to capture data for nursing tasks, and have little basis in nursing knowledge. Opportunity lies in an important issue where the knowledge used by expert nurses (nursing knowledge workers) in caring for patients is undervalued in the health care system. The complexity of nursing's knowledge base remains poorly articulated and inadequately represented in contemporary information systems. There is opportunity for data mining methods to assist with discovering important linkages between clinical data, nursing interventions, and patient outcomes. Following a brief overview of relevant data mining techniques, a preterm risk prediction case study illustrates the opportunities and describes typical data mining issues in the nontrivial task of building knowledge. Building knowledge in nursing, using data mining or any other method, will make progress only if important data that capture expert nurses' contributions are available in clinical information systems configurations.

Artificial Intelligence↗