PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Data Mining”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

QSPR modeling of pseudoternary microemulsions formulated employing lecithin surfactants: application of data mining, molecular and statistical modeling.

Data mining, computer aided molecular modeling, descriptor calculation, genetic algorithm and multiple linear regression analysis techniques were combined together to generate predictive quantitative structure property relationship (QSPR) models explaining the formation of lecithin-based W/O microemulsions. Ninety-four microemulsion phase diagrams were collected from five different references published over the past few years. Computer-based molecular modeling techniques were then applied on the components of the collected microemulsion systems to generate corresponding plausible three-dimensional (3D) structures. The resulting 3D models were utilized to calculate a group of molecular physicochemical descriptors. Thereafter, genetic algorithm and backward stepwise regression analysis were separately assessed as means for selecting optimal descriptor sets for statistical modeling. The selected descriptors were correlated with microemulsion existence areas employing multiple linear regression analysis. The resulting W/O models were statistically validated and found to be of significant predictive power. The models allowed better understanding of the process of microemulsion formation. Unfortunately, all QSPR modeling efforts directed towards O/W microemulsions failed completely.

Emulsions↗

Accomplishments and challenges in literature data mining for biology.

We review recent results in literature data mining for biology and discuss the need and the steps for a challenge evaluation for this field. Literature data mining has progressed from simple recognition of terms to extraction of interaction relationships from complex sentences, and has broadened from recognition of protein interactions to a range of problems such as improving homology search, identifying cellular location, and so on. To encourage participation and accelerate progress in this expanding field, we propose creating challenge evaluations, and we describe two specific applications in this context.

Benchmarking↗

A novel data mining approach to the identification of effective drugs or combinations for targeted endpoints--application to chronic heart failure as a new form of evidence-based medicine.

BACKGROUND: Data mining is a technique for discovering useful information hidden in a database, which has recently been used by the chemical, financial, pharmaceutical, and insurance industries. It may enable us to detect the interesting and hidden data on useful drugs especially in the field of cardiovascular disease. METHODS AND RESULTS: We evaluated the current treatments for chronic heart failure (CHF) in our institute using a decision tree method of data mining and compared the results with those of large-scale clinical trials. We enrolled 1,100 patients with CHF (NYHA classes II-IV and EF < 40%) who were hospitalized at the National Cardiovascular Center during the past 31 months. Drugs prescribed at discharge were extracted from the clinical database. Both echocardiograms and plasma BNP level at 6-12 months after discharge were determined prospectively. It was found that beta-blockers, angiotensin converting enzyme inhibitors, and angiotensin II receptor antagonists independently improve both the plasma BNP level and %fractional shortening (FS), while oral inotropic agents increased the plasma BNP level and decreased %FS. These findings agree with evidence accumulated from several large-scale trials. Interestingly, statins, histamine receptor blockers, and alpha-glucosidase inhibitors also attenuated the severity of CHF, suggesting the possibility of new treatment of CHF. CONCLUSION: Clinical data mining using Japanese CHF patients yielded almost identical data to the results of large-scale trials, and also suggested novel and unexpected candidates for CHF therapy. Further validation of the data mining approved in the cardiovascular field is warranted.

Chronic Disease↗

Data mining of gene expression changes in Alzheimer brain.

Genome-wide transcription profiling is a powerful technique for studying the enormous complexity of cellular states. Moreover, when applied to disease tissue it may reveal quantitative and qualitative alterations in gene expression that give information on the context or underlying basis for the disease and may provide a new diagnostic approach. However, the data obtained from high-density microarrays is highly complex and poses considerable challenges in data mining. The data requires care in both pre-processing and the application of data mining techniques. This paper addresses the problem of dealing with microarray data that come from two known classes (Alzheimer and normal). We have applied three separate techniques to discover genes associated with Alzheimer disease (AD). The 67 genes identified in this study included a total of 17 genes that are already known to be associated with Alzheimer's or other neurological diseases. This is higher than any of the previously published Alzheimer's studies. Twenty known genes, not previously associated with the disease, have been identified as well as 30 uncharacterized expressed sequence tags (ESTs). Given the success in identifying genes already associated with AD, we can have some confidence in the involvement of the latter genes and ESTs. From these studies we can attempt to define therapeutic strategies that would prevent the loss of specific components of neuronal function in susceptible patients or be in a position to stimulate the replacement of lost cellular function in damaged neurons. Although our study is based on a relatively small number of patients (four AD and five normal), we think our approach sets the stage for a major step in using gene expression data for disease modeling (i.e. classification and diagnosis). It can also contribute to the future of gene function identification, pathology, toxicogenomics, and pharmacogenomics.

Alzheimer Disease↗

Data mining and healthcare informatics.

OBJECTIVE: To acquaint members of the Academy with a relatively recent development in the area of data exploration and statistical analysis. METHODS: A review of the concepts and methods inherent in data-mining with a special emphasis on the those methods applicable to predictive modeling. RESULTS: Data-mining is demonstrated to be a useful tool for researchers in those circumstances where large amounts of information are available. CONCLUSIONS: With the advent and proliferation of on-line data collection, truly massive databases are now available to health care researchers. In that situation, data-mining methods yield some unique opportunities to researchers who wish to develop prediction models and to establish associations.

Database Management Systems↗

Data mining: sophisticated forms of managed care modeling through artificial intelligence.

Data mining is a recent development in computer science that combines artificial intelligence algorithms and relational databases to discover patterns automatically, without the use of traditional statistical methods. Work with data mining tools in health care is in a developmental stage that holds great promise, given the combination of demographic and diagnostic information.

Algorithms↗

Data mining and computationally intensive methods: summary of Group 7 contributions to Genetic Analysis Workshop 13.

The Framingham Heart Study data, as well as a related simulated data set, were generously provided to the participants of the Genetic Analysis Workshop 13 in order that newly developed and emerging statistical methodologies could be tested on that well-characterized data set. The impetus driving the development of novel methods is to elucidate the contributions of genes, environment, and interactions between and among them, as well as to allow comparison between and validation of methods. The seven papers that comprise this group used data-mining methodologies (tree-based methods, neural networks, discriminant analysis, and Bayesian variable selection) in an attempt to identify the underlying genetics of cardiovascular disease and related traits in the presence of environmental and genetic covariates. Data-mining strategies are gaining popularity because they are extremely flexible and may have greater efficiency and potential in identifying the factors involved in complex disorders. While the methods grouped together here constitute a diverse collection, some papers asked similar questions with very different methods, while others used the same underlying methodology to ask very different questions. This paper briefly describes the data-mining methodologies applied to the Genetic Analysis Workshop 13 data sets and the results of those investigations.

Bayes Theorem↗

Autonomous decision-making: a data mining approach.

The researchers and practitioners of today create models, algorithms, functions, and other constructs defined in abstract spaces. The research of the future will likely be data driven. Symbolic and numeric data that are becoming available in large volumes will define the need for new data analysis techniques and tools. Data mining is an emerging area of computational intelligence that offers new theories, techniques, and tools for analysis of large data sets. In this paper, a novel approach for autonomous decision-making is developed based on the rough set theory of data mining. The approach has been tested on a medical data set for patients with lung abnormalities referred to as solitary pulmonary nodules (SPNs). The two independent algorithms developed in this paper either generate an accurate diagnosis or make no decision. The methodolgy discussed in the paper depart from the developments in data mining as well as current medical literature, thus creating a variable approach for autonomous decision-making.

Algorithms↗

Application of data mining for examining polypharmacy and adverse effects in cardiology patients.

This article comments upon the use of data mining tools to examine clinical data. Many cardiovascular patients have co-morbid diseases that put them at risk for polypharmacy, or severe adverse reactions from the interactions of multiple medications. Clinical trials typically use too few patients with stringent inclusion/exclusion criteria that prevent an examination of the issue of polypharmacy. However, clinical data collected in the course of patient treatment can be used in conjunction with data mining to find meaningful results.

Clinical Trials as Topic↗

Data mining techniques for cancer detection using serum proteomic profiling.

OBJECTIVE: Pathological changes in an organ or tissue may be reflected in proteomic patterns in serum. It is possible that unique serum proteomic patterns could be used to discriminate cancer samples from non-cancer ones. Due to the complexity of proteomic profiling, a higher order analysis such as data mining is needed to uncover the differences in complex proteomic patterns. The objectives of this paper are (1) to briefly review the application of data mining techniques in proteomics for cancer detection/diagnosis; (2) to explore a novel analytic method with different feature selection methods; (3) to compare the results obtained on different datasets and that reported by Petricoin et al. in terms of detection performance and selected proteomic patterns. METHODS AND MATERIAL: Three serum SELDI MS data sets were used in this research to identify serum proteomic patterns that distinguish the serum of ovarian cancer cases from non-cancer controls. A support vector machine-based method is applied in this study, in which statistical testing and genetic algorithm-based methods are used for feature selection respectively. Leave-one-out cross validation with receiver operating characteristic (ROC) curve is used for evaluation and comparison of cancer detection performance. RESULTS AND CONCLUSIONS: The results showed that (1) data mining techniques can be successfully applied to ovarian cancer detection with a reasonably high performance; (2) the classification using features selected by the genetic algorithm consistently outperformed those selected by statistical testing in terms of accuracy and robustness; (3) the discriminatory features (proteomic patterns) can be very different from one selection method to another. In other words, the pattern selection and its classification efficiency are highly classifier dependent. Therefore, when using data mining techniques, the discrimination of cancer from normal does not depend solely upon the identity and origination of cancer-related proteins.

Biomarkers, Tumor↗

Crystallization data mining in structural genomics: using positive and negative results to optimize protein crystallization screens.

Recent efforts to collect and mine crystallization data from structural genomics (SG) consortia have led to the identification of minimal screens and novel screening strategies that can be used to streamline the crystallization process. Two groups, the Joint Center for Structural Genomics and the University of Toronto, carried out large-scale crystallization trials on different sets of bacterial targets (539, JCSG and 755, Toronto), using different sample processing and crystallization methods, and then analyzed their results to identify the smallest subset of conditions that would have crystallized the maximum number of protein targets. The JCSG Core Screen contains 67 conditions (from 480) while the Toronto Minimal Screen contains 6 (from 48). While the exact conditions included in the two screens do not overlap, the major precipitants of the conditions are similar and thus both screens can be used to determine if a protein has a natural propensity to crystallize. In addition, studies from other groups including the University of Queensland, the Mycobacterium tuberculosis SG group, the Southeast Collaboratory for SG, and the York Structural Biology Laboratory indicate that alternative crystallization strategies may be more successful at identifying initial crystallization conditions than typical sparse matrix screens. These minimal screens and alternative screening strategies are already being used to optimize the crystallization processes within large SG efforts. The differences between these results, however, demonstrate that additional studies which examine the influence of protein biophysical properties and sample preparation methods on crystal formation must also be carried out before more robust screens can be identified.

Chemistry Techniques, Analytical↗

The importance of scaling in data mining for toxicity prediction.

While mining a data set of 554 chemicals in order to extract information on their toxicity value, we faced the problem of scaling all the data. There are numerous different approaches to this procedure, and in most cases the choice greatly influences the results. The aim of this paper is 2-fold. First, we propose a universal scaling procedure for acute toxicity in fish according to the Directive 92/32/EEC. Second, we look at how expert preprocessing of the data effects the performance of qualitative structure-activity relationship (QSAR) approach to toxicity prediction.

Animals↗

Deriving knowledge through data mining high-throughput screening data.

Deriving general knowledge from high-throughput screening data is made difficult by the significant amount of noise, arising primarily from false positives, in the data. The paradigm established for screening an encoded combinatorial library on polymeric support, an ECLiPS library, has a significant amount of built-in redundancy. Because of this redundancy, the resulting data can be interpreted through a rigorous statistical analysis procedure, thereby significantly reducing the number of false positives. Here, we develop the statistical models used to analyze data from high-throughput screens of ECLiPS libraries to derive unbiased true hit rates. These hit rates can also be calculated on subsets of the collection such as those compounds containing a carboxylic acid or those with molecular weight below 350 Da. The relative value of the hit rate on the subset of the collection can then be compared to the overall hit rate to determine the effect of the substructure or physical property on the likelihood of a molecule having biological activity. Here, we show the effects that various functional groups and the standard physical properties, molecular weight, hydrogen bond donors, hydrogen bond acceptors, log P, and rotatable bonds, have on the likelihood of a compound being biologically active. To our knowledge this is the first published account of the use of high-throughput screening data to elucidate the effects of physical properties and substructures on the likelihood of compounds showing biological activity over a broad range of pharmaceutically relevant targets.

Algorithms↗

Applying data mining in healthcare: an info-structure for delivering 'data-driven' strategic services.

Presently, there is a growing demand from the healthcare community to leverage upon and transform the vast quantities of healthcare data into value-added, 'decision-quality' knowledge, vis-à-vis, strategic knowledge services oriented towards healthcare management and planning. To meet this end, we present a Strategic Knowledge Services Info-structure that leverages on existing healthcare knowledge/data bases to derive decision-quality knowledge-knowledge that is extracted from healthcare data through services akin to knowledge discovery in databases and data mining.

Artificial Intelligence↗

Data mining for signals in spontaneous reporting databases: proceed with caution.

PURPOSE: To provide commentary and points of caution to consider before incorporating data mining as a routine component of any Pharmacovigilance program, and to stimulate further research aimed at better defining the predictive value of these new tools as well as their incremental value as an adjunct to traditional methods of post-marketing surveillance. METHODS/RESULTS: Commentary includes review of current data mining methodologies employed and their limitations, caveats to consider in the use of spontaneous reporting databases and caution against over-confidence in the results of data mining. CONCLUSIONS: Future research should focus on more clearly delineating the limitations of the various quantitative approaches as well as the incremental value that they bring to traditional methods of pharmacovigilance.

Adverse Drug Reaction Reporting Systems↗

Knowledge discovery and data mining in toxicology.

Knowledge discovery and data mining tools are gaining increasing importance for the analysis of toxicological databases. This paper gives a survey of algorithms, capable to derive interpretable models from toxicological data, and presents the most important application areas. The majority of techniques in this area were derived from symbolic machine learning, one commercial product was developed especially for toxicological applications. The main application area is presently the detection of structure-activity relationships, very few authors have used these techniques to solve problems in epidemiological and clinical toxicology. Although the discussed algorithms are very flexible and powerful, further research is required to adopt the algorithms to the specific learning problems in this area, to develop improved representations of chemical and biological data and to enhance the interpretability of the derived models for toxicological experts.

Algorithms↗

Analysis of patient flows via data mining.

The paper presents DoMiner, a data mining tool for the analysis of patient flows among public hospitals and care units. DoMiner is based on the theory of rough sets, and allows for the extraction of association rules from data base tables. The paper describes both the clustering and rule extraction algorithm of DoMiner, and illustrates an introductory analysis example.

Algorithms↗

Investigation of protein functions through data-mining on integrated human transcriptome database, H-Invitational database (H-InvDB).

H-Invitational Database (H-InvDB; ) is a human transcriptome database, containing integrative annotation of 41,118 full-length cDNA clones originated from 21,037 loci. H-InvDB is a product of the H-Invitational project, an international collaboration to systematically and functionally validate human genes by analysis of a unique set of high quality full-length cDNA clones using automatic annotation and human curation under unified criteria. Here, 19,574 proteins encoded by these cDNAs were classified into 11,709 function-known and 7865 function-unknown hypothetical proteins by similarity with protein databases and motif prediction (InterProScan). The proportion of "hypothetical proteins" in H-InvDB was as high as 40.4%. In this study, we thus conducted data-mining in H-InvDB with the aim of assigning advanced functional annotations to those hypothetical proteins. First, by data-mining in the H-InvDB version of GTOP, we identified 337 SCOP domains within 7865 H-Inv hypothetical proteins. Second, by data-mining of predicted subcellular localization by SOSUI and TMHMM in H-InvDB, we found 1032 transmembrane proteins within H-Inv hypothetical proteins. These results clearly demonstrate that structural prediction is effective for functional annotation of proteins with unknown functions. All the data in H-InvDB are shown in two main views, the cDNA view and the Locus view, and five auxiliary databases with web-based viewers; DiseaseInfo Viewer, H-ANGEL, Clustering Viewer, G-integra and TOPO Viewer; the data also are provided as flat files and XML files. The data consists of descriptions of their gene structures, novel alternative splicing isoforms, functional RNAs, functional domains, subcellular localizations, metabolic pathways, predictions of protein 3D structure, mapping of SNPs and microsatellite repeat motifs in relation with orphan diseases, gene expression profiling, and comparisons with mouse full-length cDNAs in the context of molecular evolution. This unique integrative platform for conducting in silico data-mining represents a substantial contribution to resources required for the exploration of human biology and pathology.

Amino Acid Sequence↗