PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Data Mining”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Novel statistical classification model of type 2 diabetes mellitus patients for tailor-made prevention using data mining algorithm.

To estimate the usefulness of data mining algorithms for extracting risk predictors of diabetic vascular complications in proper order in the future, we tried applying the Classification and Regression Trees (CART) method to the prevalence data of 165 type 2 diabetic outpatients and already known risk factors. Among the 6 categorical and 15 continuous risk factors, age (cutoff: 65.4) was the best predictor for classifying patients into groups with and without macroangiopathy (p=0.000). Body weight (cutoff: 53.9) was the best predictor (p=0.006) in the older group (age >65.4), whereas systolic blood pressure (cutoff: 144.5) was the best predictor in the remaining group (p=0.002). Age (cutoff: 64.8) was also the best predictor for categorizing them into groups with and without microangiopathy (p=0.000). In the older group (age >64.8), BMI (cutoff: 21.5) was the best predictor (p=0.001), whereas morbidity term (cutoff: 15.5) was the best predictor in the other group (p=0.01 0). Because the orders and values of all risk factors and cutoff points mined were reasonable clinically, this method may have the potential to highlight predictors in order of importance to apply tailor-made prevention of diabetic vascular complications.

Algorithms↗

Online analytical processing (OLAP): a fast and effective data mining tool for gene expression databases.

Gene expression databases contain a wealth of information, but current data mining tools are limited in their speed and effectiveness in extracting meaningful biological knowledge from them. Online analytical processing (OLAP) can be used as a supplement to cluster analysis for fast and effective data mining of gene expression databases. We used Analysis Services 2000, a product that ships with SQLServer2000, to construct an OLAP cube that was used to mine a time series experiment designed to identify genes associated with resistance of soybean to the soybean cyst nematode, a devastating pest of soybean. The data for these experiments is stored in the soybean genomics and microarray database (SGMD). A number of candidate resistance genes and pathways were found. Compared to traditional cluster analysis of gene expression data, OLAP was more effective and faster in finding biologically meaningful information. OLAP is available from a number of vendors and can work with any relational database management system through OLE DB.

Journal Article↗

Application of information-theoretic data mining techniques in a national ambulatory practice outcomes research network.

The Medical Quality Improvement Consortium data warehouse contains de-identified data on more than 3.6 million patients including their problem lists, test results, procedures and medication lists. This study uses reconstructability analysis, an information-theoretic data mining technique, on the MQIC data warehouse to empirically identify risk factors for various complications of diabetes including myocardial infarction and microalbuminuria. The risk factors identified match those risk factors identified in the literature, demonstrating the utility of the MQIC data warehouse for outcomes research, and RA as a technique for mining clinical data warehouses.

Adult↗

Integration of text- and data-mining using ontologies successfully selects disease gene candidates.

Genome-wide techniques such as microarray analysis, Serial Analysis of Gene Expression (SAGE), Massively Parallel Signature Sequencing (MPSS), linkage analysis and association studies are used extensively in the search for genes that cause diseases, and often identify many hundreds of candidate disease genes. Selection of the most probable of these candidate disease genes for further empirical analysis is a significant challenge. Additionally, identifying the genes that cause complex diseases is problematic due to low penetrance of multiple contributing genes. Here, we describe a novel bioinformatic approach that selects candidate disease genes according to their expression profiles. We use the eVOC anatomical ontology to integrate text-mining of biomedical literature and data-mining of available human gene expression data. To demonstrate that our method is successful and widely applicable, we apply it to a database of 417 candidate genes containing 17 known disease genes. We successfully select the known disease gene for 15 out of 17 diseases and reduce the candidate gene set to 63.3% (+/-18.8%) of its original size. This approach facilitates direct association between genomic data describing gene expression and information from biomedical texts describing disease phenotype, and successfully prioritizes candidate genes according to their expression in disease-affected tissues.

Anatomy↗

[Estimation of a nationwide statistics of hernia operation applying data mining technique to the National Health Insurance Database].

OBJECTIVES: The aim of this study is to develop a methodology for estimating a nationwide statistic for hernia operations with using the claim database of the Korea Health Insurance Cooperation (KHIC). METHODS: According to the insurance claim procedures, the claim database was divided into the electronic data interchange database (EDI_DB) and the sheet database (Paper_DB). Although the EDI_DB has operation and management codes showing the facts and kinds of operations, the Paper_DB doesn't. Using the hernia matched management code in the EDI_DB, the cases of hernia surgery were extracted. For drawing the potential cases from the Paper_DB, which doesn't have the code, the predictive model was developed using the data mining technique called SEMMA. The claim sheets of the cases that showed a predictive probability of an operation over the threshold, as was decided by the ROC curve, were identified in order to get the positive predictive value as an index of usefulness for the predictive model. RESULTS: Of the claim databases in 2004, 14,386 cases had hernia related management codes with using the EDI system. For fitting the models with applying the data mining technique, logistic regression was chosen rather than the neural network method or the decision tree method. From the Paper_DB, 1,019 cases were extracted as potential cases. Direct review of the sheets of the extracted cases showed that the positive predictive value was 95.3%. CONCLUSIONS: The results suggested that applying the data mining technique to the claim database in the KHIC for estimating the nationwide surgical statistics would be useful from the aspect of execution and cost-effectiveness.

Adolescent↗

Predicting survival time for kidney dialysis patients: a data mining approach.

The cost for providing care for patients on hemodialysis due to end stage kidney disease is high. Finding ways to improve patient outcomes and reduce the cost of dialysis is important. Dialysis care is intricate and multiple factors may influence patient survival. Over 50 parameters may be monitored on a regular basis in providing kidney dialysis treatments. Understanding the collective role of these parameters in determining outcomes for an individual patient and administering individualized treatments allowing specific interventions is a challenge. Individual patient survival may depend on a complex interrelationship between multiple demographic and clinical parameters, medications, medical interventions, and the dialysis treatment prescription. In this research, data preprocessing, data transformations, and a data mining approach are used to elicit knowledge about the interaction between many of these measured parameters and patient survival. Two different data mining algorithms were employed for extracting knowledge in the form of decision rules. These rules were used by a decision-making algorithm, which predicts survival of new unseen patients. Important parameters identified by data mining are interpreted for their medical significance. The concepts introduced in this research have been applied and tested using data collected at four dialysis sites. The computational results are reported in the paper.

Algorithms↗

New Powder Diffraction File (PDF-4) in relational database format: advantages and data-mining capabilities.

The International Centre for Diffraction Data (ICDD) is responding to the changing needs in powder diffraction and materials analysis by developing the Powder Diffraction File (PDF) in a very flexible relational database (RDB) format. The PDF now contains 136,895 powder diffraction patterns. In this paper, an attempt is made to give an overview of the PDF-4, search/match methods and the advantages of having the PDF-4 in RDB format. Some case studies have been carried out to search for crystallization trends, properties, frequencies of space groups and prototype structures. These studies give a good understanding of the basic structural aspects of classes of compounds present in the database. The present paper also reports data-mining techniques and demonstrates the power of a relational database over the traditional (flat-file) database structures.

Journal Article↗

A hybrid model for improving response time in distributed data mining.

This paper presents a hybrid distributed data mining (DDM) model for optimization of response time. The model combines a mobile agent approach with client server strategies to reduce the overall response time. The hybrid model proposes and develops accurate a priori estimates of the computation and communication components of response time as the costing strategy to support optimization. Experimental evaluation of the hybrid model is presented.

Algorithms↗

A data mining system for infection control surveillance.

Nosocomial infections and antimicrobial resistance are problems of enormous magnitude that impact the morbidity and mortality of hospitalized patients as well as their cost of care. The Data Mining Surveillance System (DMSS) uses novel data mining techniques to discover unsuspected, useful patterns of nosocomial infections and antimicrobial resistance from the analysis of hospital laboratory data. This report details a mature version of DMSS as well as an experiment in which DMSS was used to analyze all inpatient culture data, collected over 15 months at the University of Alabama at Birmingham Hospital.

Alabama↗

Genomic research and data-mining technology: implications for personal privacy and informed consent.

This essay examines issues involving personal privacy and informed consent that arise at the intersection of information and communication technology (ICT) and population genomics research. I begin by briefly examining the ethical, legal, and social implications (ELSI) program requirements that were established to guide researchers working on the Human Genome Project (HGP). Next I consider a case illustration involving deCODE Genetics, a privately owned genetic company in Iceland, which raises some ethical concerns that are not clearly addressed in the current ELSI guidelines. The deCODE case also illustrates some ways in which an ICT technique known as data mining has both aided and posed special challenges for researchers working in the field of population genomics. On the one hand, data-mining tools have greatly assisted researchers in mapping the human genome and in identifying certain "disease genes" common in specific populations (which, in turn, has accelerated the process of finding cures for diseases tha affect those populations). On the other hand, this technology has significantly threatened the privacy of research subjects participating in population genomics studies, who may, unwittingly, contribute to the construction of new groups (based on arbitrary and non-obvious patterns and statistical correlations) that put those subjects at risk for discrimination and stigmatization. In the final section of this paper I examine some ways in which the use of data mining in the context of population genomics research poses a critical challenge for the principle of informed consent, which traditionally has played a central role in protecting the privacy interests of research subjects participating in epidemiological studies.

Computational Biology↗

Development of EST-SSR markers by data mining in three species of shrimp: Litopenaeus vannamei, Litopenaeus stylirostris, and Trachypenaeus birdy.

We report on the data mining of publicly available Litopenaeus vannamei expressed sequence tags (ESTs) to generate simple sequence repeat (SSRs) markers and on their transferability between related Penaeid shrimp species. Repeat motifs were found in 3.8% of the evaluated ESTs at a frequency of one repeat every 7.8 kb of sequence data. A total of 206 primer pairs were designed, and 112 loci were amplified with the highest success in L. vannamei. A high percentage (69%) of EST-SSRs were transferable within the genus Litopenaeus. More than half of the amplified products were polymorphic in a small testing panel of L. vannamei. Evaluation of those primers in a larger testing panel showed that 72% of the markers fit Hardy-Weinberg equilibrium, which shows their utility for population genetic analysis. Additionally, a set of 26 of the EST-SSRs were evaluated for Mendelian segregation. A high percentage of monomorphic markers (46%) proved to be polymorphic by singles-stranded conformational polymorphism analysis. Because of the high number of ESTs available in public databases, a data mining approach similar to the one outlined here might yield high numbers of SSR markers in many animal taxa.

Animals↗

Data mining in the US using the Vaccine Adverse Event Reporting System.

The US Vaccine Adverse Event Reporting System (VAERS), which is charged with vigilance for detecting vaccine-related safety issues, faces an increasingly complex immunisation environment. Since 1990, steady increases in vaccine licensing and distribution have resulted in increasing numbers of reports to VAERS. Prominent features of current reports include more routine vaccine co-administration and frequent reports of new postvaccination clinical syndromes. Data-mining methods, based on disproportionality analyses, are one strategy being pursued by VAERS researchers to increase the utility of its complex database. The types of analyses used include proportional reporting ratios, association rule discovery, and various 'historic limits' methods that compare observed versus expected event counts. The use of such strategies in VAERS has been primarily supplemental and retrospective. Signals for inactivated influenza, typhoid and tetanus toxoid-containing vaccines have been successfully identified. Concerns flagged through data mining should always be subject to clinical case review as a first evaluation step. Persistent issues should be subject to formal hypothesis testing in large linked databases or other controlled-study settings. Automated data-mining techniques for prospective use are currently undergoing development and evaluation within VAERS. Their use (as one signal-detection tool among many) by trained medical evaluators who are aware of system limitations is one legitimate approach to improving the ability of VAERS to generate vaccine-safety hypotheses. Such approaches are needed as more new vaccines continue to be licensed.

Adverse Drug Reaction Reporting Systems↗

Clinical and pharmacogenomic data mining: 3. Zeta theory as a general tactic for clinical bioinformatics.

A new approach, a Zeta Theory of observations, data, and data mining, is being forged from a theory of expected information into an even more cohesive and comprehensive form by the challenge of general genomic, pharmacogenomic, and proteomic data. In this paper, the focus is not on studies using the specific tool FANO (CliniMiner) but on extensions to a new broader theoretical approach, aspects of which can easily be implemented into, or otherwise support, excellent existing methods, such as forms of multivariate analysis and IBM's product Intelligent Miner. The theory should perhaps be distinguished from an existing purely number-theoretic area sometimes also known as Zeta Theory, which focuses on the Riemann Zeta Function and the ways in which it governs the distribution of prime numbers. However, Zeta Theory as used here overlaps heavily with it and actually makes use of these same matters. The distinction is that it enters from a Bayesian information theory and data representation perspective. It could thus be considered an application of the 'mathematician's version'. The application is by no means confined to areas of modern biomedicine, and indeed its generality, even merging into quantum mechanics, is a key feature. Other areas with some similar challenges as modern biology, and which have inspired data mining methods such as IBM's Intelligent Miner, include commerce. But for several reasons discussed, modern molecular biology and medicine seem particularly challenging, and this relates to the often irreducible high dimensionality of the data. This thus remains our main target.

Computational Biology↗

MedlineR: an open source library in R for Medline literature data mining.

SUMMARY: We describe an open source library written in the R programming language for Medline literature data mining. This MedlineR library includes programs to query Medline through the NCBI PubMed database; to construct the co-occurrence matrix; and to visualize the network topology of query terms. The open source nature of this library allows users to extend it freely in the statistical programming language of R. To demonstrate its utility, we have built an application to analyze term-association by using only 10 lines of code. We provide MedlineR as a library foundation for bioinformaticians and statisticians to build more sophisticated literature data mining applications. AVAILABILITY: The library is available from http://dbsr.duke.edu/pub/MedlineR.

Abstracting and Indexing↗

Automatic rule generation for protein annotation with the C4.5 data mining algorithm applied on SWISS-PROT.

MOTIVATION: The gap between the amount of newly submitted protein data and reliable functional annotation in public databases is growing. Traditional manual annotation by literature curation and sequence analysis tools without the use of automated annotation systems is not able to keep up with the ever increasing quantity of data that is submitted. Automated supplements to manually curated databases such as TrEMBL or GenPept cover raw data but provide only limited annotation. To improve this situation automatic tools are needed that support manual annotation, automatically increase the amount of reliable information and help to detect inconsistencies in manually generated annotations. RESULTS: A standard data mining algorithm was successfully applied to gain knowledge about the Keyword annotation in SWISS-PROT. 11 306 rules were generated, which are provided in a database and can be applied to yet unannotated protein sequences and viewed using a web browser. They rely on the taxonomy of the organism, in which the protein was found and on signature matches of its sequence. The statistical evaluation of the generated rules by cross-validation suggests that by applying them on arbitrary proteins 33% of their keyword annotation can be generated with an error rate of 1.5%. The coverage rate of the keyword annotation can be increased to 60% by tolerating a higher error rate of 5%. AVAILABILITY: The results of the automatic data mining process can be browsed on http://golgi.ebi.ac.uk:8080/Spearmint/ Source code is available upon request. CONTACT: kretsch@ebi.ac.uk.

Algorithms↗

Bioinformatics and data mining in proteomics.

Proteomic studies involve the identification as well as qualitative and quantitative comparison of proteins expressed under different conditions, and elucidation of their properties and functions, usually in a large-scale, high-throughput format. The high dimensionality of data generated from these studies will require the development of improved bioinformatics tools and data-mining approaches for efficient and accurate data analysis of biological specimens from healthy and diseased individuals. Mining large proteomics data sets provides a better understanding of the complexities between the normal and abnormal cell proteome of various biological systems, including environmental hazards, infectious agents (bioterrorism) and cancers. This review will shed light on recent developments in bioinformatics and data-mining approaches, and their limitations when applied to proteomics data sets, in order to strengthen the interdependence between proteomic technologies and bioinformatics tools.

Computational Biology↗

Data mining and visualization for decision support and modeling of public health-care resources.

This paper proposes an innovative use of data mining and visualization techniques for decision support in planning and regional-level management of Slovenian public health-care. Data mining and statistical techniques were used to analyze databases collected by a regional Public Heath Institute. We also studied organizational aspects of public health resources in the selected Celje region with the objective to identify the areas that are atypical in terms of availability and accessibility of public health services for the population. The most important step was the detection of outliers and the analysis of availability and accessibility deviations. The results are applicable to health-care planning and support in decision making by local and regional health-care authorities. In addition to the practical results, which are directly useful for decision making in planning of the regional health-care system, the main methodological contribution of the paper are the developed visualization methods that can be used to facilitate knowledge management and decision making processes.

Database Management Systems↗

Microarray analysis of gene expression: considerations in data mining and statistical treatment.

DNA microarray represents a powerful tool in biomedical discoveries. Harnessing the potential of this technology depends on the development and appropriate use of data mining and statistical tools. Significant current advances have made microarray data mining more versatile. Researchers are no longer limited to default choices that generate suboptimal results. Conflicting results in repeated experiments can be resolved through attention to the statistical details. In the current dynamic environment, there are many choices and potential pitfalls for researchers who intend to incorporate microarrays as a research tool. This review is intended to provide a simple framework to understand the choices and identify the pitfalls. Specifically, this review article discusses the choice of microarray platform, preprocessing raw data, differential expression and validation, clustering, annotation and functional characterization of genes, and pathway construction in light of emergent concepts and tools.

Cluster Analysis↗