PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data mining”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

A data mining technique for discovering distinct patterns of hand signs: implications in user training and computer interface design.

Hand signs are considered as one of the important ways to enter information into computers for certain tasks. Computers receive sensor data of hand signs for recognition. When using hand signs as computer inputs, we need to (1) train computer users in the sign language so that their hand signs can be easily recognized by computers, and (2) design the computer interface to avoid the use of confusing signs for improving user input performance and user satisfaction. For user training and computer interface design, it is important to have a knowledge of which signs can be easily recognized by computers and which signs are not distinguishable by computers. This paper presents a data mining technique to discover distinct patterns of hand signs from sensor data. Based on these patterns, we derive a group of indistinguishable signs by computers. Such information can in turn assist in user training and computer interface design.

Algorithms↗

A Consensus Data Mining secondary structure prediction by combining GOR V and Fragment Database Mining.

The major aim of tertiary structure prediction is to obtain protein models with the highest possible accuracy. Fold recognition, homology modeling, and de novo prediction methods typically use predicted secondary structures as input, and all of these methods may significantly benefit from more accurate secondary structure predictions. Although there are many different secondary structure prediction methods available in the literature, their cross-validated prediction accuracy is generally <80%. In order to increase the prediction accuracy, we developed a novel hybrid algorithm called Consensus Data Mining (CDM) that combines our two previous successful methods: (1) Fragment Database Mining (FDM), which exploits the Protein Data Bank structures, and (2) GOR V, which is based on information theory, Bayesian statistics, and multiple sequence alignments (MSA). In CDM, the target sequence is dissected into smaller fragments that are compared with fragments obtained from related sequences in the PDB. For fragments with a sequence identity above a certain sequence identity threshold, the FDM method is applied for the prediction. The remainder of the fragments are predicted by GOR V. The results of the CDM are provided as a function of the upper sequence identities of aligned fragments and the sequence identity threshold. We observe that the value 50% is the optimum sequence identity threshold, and that the accuracy of the CDM method measured by Q(3) ranges from 67.5% to 93.2%, depending on the availability of known structural fragments with sufficiently high sequence identity. As the Protein Data Bank grows, it is anticipated that this consensus method will improve because it will rely more upon the structural fragments.

Algorithms↗

GeneMerge--post-genomic analysis, data mining, and hypothesis testing.

SUMMARY: GeneMerge is a web-based and standalone program written in PERL that returns a range of functional and genomic data for a given set of study genes and provides statistical rank scores for over-representation of particular functions or categories in the data set. Functional or categorical data of all kinds can be analyzed with GeneMerge, facilitating regulatory and metabolic pathway analysis, tests of population genetic hypotheses, cross-experiment comparisons, and tests of chromosomal clustering, among others. GeneMerge can perform analyses on a wide variety of genomic data quickly and easily and facilitates both data mining and hypothesis testing. AVAILABILITY: GeneMerge is available free of charge for academic use over the web and for download from: http://www.oeb.harvard.edu/hartl/lab/publications/GeneMerge.html.

Algorithms↗

Protecting patient privacy in clinical data mining.

This paper investigates whether HIPAA de-identification requirements--as well as proposed AAMC de-identification standards--were met in a large clinical data mining study (1997-2001) conducted at Duke University prior to the publication of the final rule. While HIPAA has improved de-identification standards, the study also shows that privacy issues may persist even in de-identified large clinical databases.

Biomedical Research↗

Revealing transforming growth factor-beta signaling transduction in human kidney by gene expression data mining.

Tumor growth factor-beta (TGF-beta) is a key mediator of glomerular and tubulointerstitial pathobiology in chronic kidney disease. Its signaling transduction controls a diverse number of biological processes in a dynamic and context-dependent manner. We applied a data mining strategy to deconvolute gene expression patterns across hundreds of microarray data sets to reveal members of the TGF-beta signaling network in human kidney. This strategy is composed of three major steps: (i) select genes known to be involved and expressionally regulated in TGF-beta signaling as "bait"; (ii) select microarray data sets in which the bait genes are strongly co-regulated; (iii) identify (or "fish") additional TGF-beta signaling genes by a non-parametric statistic-based gene scoring system (NP score). The 40 genes with highest NP scores and significant permutation p values were selected for in silico validation, and used to identify a network, in which 35 of these genes were found to be connected by literature- derived relationships. Transcription factors were found to be enriched in the top list. Among them, activated transcription factor 3 (ATF3) had the highest NP score, and was proposed to play a pivotal role in TGF-beta signaling in human kidney. Finally, we implemented a non-parametric pathway ranking (NPPR) tool (Mootha et al., 2003) to rank pathways and identified canonical biological pathways associated with the down-stream of TGF-beta signaling.

Computational Biology↗

Harnessing data mining to explore incident databases.

Large numbers of incident related databases have been established in the last three decades. The majority of attempts to explore these data marts were trials to identify patterns via first glance into the datasets. This study investigated a subset of incidents from fixed facilities in Harris County, TX, extracted from the National Response Center database. By classifying the information into groups and using data mining techniques, interesting patterns of incidents according to characteristics such as type of equipment involved, type of chemical released and causes involved were revealed and further these were used to modify the annual failure probabilities of equipments.

Chemical Industry↗

A probabilistic Classifier System and its application in data mining.

The article is about a new Classifier System framework for classification tasks called BYP-CS (for BaYesian Predictive Classifier System). The proposed CS approach abandons the focus on high accuracy and addresses a well-posed Data Mining goal, namely, that of uncovering the low-uncertainty patterns of dependence that manifest often in the data. To attain this goal, BYP-CS uses a fair amount of probabilistic machinery, which brings its representation language closer to other related methods of interest in statistics and machine learning. On the practical side, the new algorithm is seen to yield stable learning of compact populations, and these still maintain a respectable amount of predictive power. Furthermore, the emerging rules self-organize in interesting ways, sometimes providing unexpected solutions to certain benchmark problems.

Algorithms↗

Automatic detection of interictal spikes using data mining models.

A prospective candidate for epilepsy surgery is studied both the ictal and interictal spikes (IS) to determine the localization of the epileptogenic zone. In this work, data mining (DM) classification techniques were utilized to build an automatic detection model. The selected DM algorithms are: Decision Trees (J 4.8), and Statistical Bayesian Classifier (naïve model). The main objective was the detection of IS, isolating them from the EEG's base activity. On the other hand, DM has an attractive advantage in such applications, in that the recognition of epileptic discharges does not need a clear definition of spike morphology. Furthermore, previously 'unseen' patterns could be recognized by the DM with proper 'training'. The results obtained showed that the efficacy of the selected DM algorithms is comparable to the current visual analysis used by the experts. Moreover, DM is faster than the time required for the visual analysis of the EEG. So this tool can assist the experts by facilitating the analysis of a patient's information, and reducing the time and effort required in the process.

Action Potentials↗

A preprocessing method for improving data mining techniques. Application to a large medical diabetes database.

The Knowledge Discovery in Databases (KDD) methodology seems to be attractive on the analyze of large clinical databases. In the KDD process, the preprocessing step (data cleaning and handling of missing values) is paramount since it conditions the quality of the results obtained by data mining procedures and represents about 80% of the whole project time. The aims of the present study were to analyze this step and provide tools to handle inconsistent data and missing values. We have broken down the process into 3 main stages: data cleaning--explanatory study of missing values--choice of the procedure used for handling missing values. The data cleaning stage was based on a system of logical rules to correct mistakes and on cluster analysis to discard the poorly filled files. The missing-data mechanism was analyzed by means of multivariate statistical procedures. Two methods to deal with missing values were compared: imputation by the most common value (mode) and imputation using decision trees. This study was performed on a large medical diabetes database (23,601 patients) including numerous missing values. A system of logical rules allowed to correct mistakes on essential parameters (for example, the type of diabetes). Cluster analysis allowed to identify 10% of poorly filled files. After multivariate analysis, the missing-data mechanism could be considered as random. For variables with low number of missing values (< 10%) and categories (< 4), imputation using decision trees provided better results than imputation by mode.

Data Interpretation, Statistical↗

Data mining and knowledge discovery in predictive toxicology.

This article describes the knowledge discovery process in predictive toxicology. This process consists of five major steps (i) feature calculation, (ii) feature selection, (iii) model induction, (iv) model validation and (v) interpretation of predictions and models. Data mining is a part of the knowledge discovery process and consists of the application of data analysis and discovery algorithms, which can be useful in all of the above steps. A brief review of suitable algorithms and their advantages and disadvantages is given for each knowledge discovery step, followed by a more detailed description of a problem-specific implementation of the lazar prediction system.

Algorithms↗

Automatic categorization of medical images for content-based retrieval and data mining.

Categorization of medical images means selecting the appropriate class for a given image out of a set of pre-defined categories. This is an important step for data mining and content-based image retrieval (CBIR). So far, published approaches are capable to distinguish up to 10 categories. In this paper, we evaluate automatic categorization into more than 80 categories describing the imaging modality and direction as well as the body part and biological system examined. Based on 6231 reference images from hospital routine, 85.5% correctness is obtained combining global texture features with scaled images. With a frequency of 97.7%, the correct class is within the best ten matches, which is sufficient for medical CBIR applications.

Automation↗

Evaluation of text data mining for database curation: lessons learned from the KDD Challenge Cup.

MOTIVATION: The biological literature is a major repository of knowledge. Many biological databases draw much of their content from a careful curation of this literature. However, as the volume of literature increases, the burden of curation increases. Text mining may provide useful tools to assist in the curation process. To date, the lack of standards has made it impossible to determine whether text mining techniques are sufficiently mature to be useful. RESULTS: We report on a Challenge Evaluation task that we created for the Knowledge Discovery and Data Mining (KDD) Challenge Cup. We provided a training corpus of 862 articles consisting of journal articles curated in FlyBase, along with the associated lists of genes and gene products, as well as the relevant data fields from FlyBase. For the test, we provided a corpus of 213 new ('blind') articles; the 18 participating groups provided systems that flagged articles for curation, based on whether the article contained experimental evidence for gene expression products. We report on the evaluation results and describe the techniques used by the top performing groups.

Abstracting and Indexing↗

Data mining of sequences and 3D structures of allergenic proteins.

MOTIVATION: Many sequences, and in some cases structures, of proteins that induce an allergic response in atopic individuals have been determined in recent years. This data indicates that allergens, regardless of source, fall into discreet protein families. Similarities in the sequence may explain clinically observed cross-reactivities between different biological triggers. However, previously available allergy databases group allergens according to their biological sources, or observed clinical cross-reactivities, without providing data about the proteins. A computer-aided data mining system is needed to compare the sequential and structural details of known allergens. This information will aid in predicting allergenic cross-responses and eventually in determining possible common characteristics of IgE recognition. RESULTS: The new web-based Structural Database of Allergenic Proteins (SDAP) permits the user to quickly compare the sequence and structure of allergenic proteins. Data from literature sources and previously existing lists of allergens are combined in a MySQL interactive database with a wide selection of bioinformatics applications. SDAP can be used to rapidly determine the relationship between allergens and to screen novel proteins for the presence of IgE or T-cell epitopes they may share with known allergens. Further, our novel similarity search method, based on five dimensional descriptors of amino acid properties, can be used to scan the SDAP entries with a peptide sequence. For example, when a known IgE binding epitope from shrimp tropomyosin was used as a query, the method rapidly identified a similar sequence in known shellfish and insect allergens. This prediction of cross-reactivity between allergens is consistent with clinical observations. AVAILABILITY: SDAP is available on the web at http://fermi.utmb.edu/SDAP/index.html

Allergens↗

Sight: automating genomic data-mining without programming skills.

SUMMARY: We created and tested Sight, a Java-based package that provides a user-friendly interface to generate and connect agents for automatic genomic data-mining for individual requirements without requiring programming skills from the user. AVAILABILITY: http://physiologie.uni-ulm.de//Seiten/Arbeitsgruppe/Jurkat-Rott/Jurkat-Rott.htm. The system does not require additional components and runs on IBM PCs under Windows (NT 4.0, 2000 and XP) or Linux (Phat 4.0 and Mandrake 9.0).

Algorithms↗

yMGV: a cross-species expression data mining tool.

The yeast Microarray Global Viewer (yMGV @ http://transcriptome.ens.fr/ymgv) was created 3 years ago as a database that houses a collection of Saccharomyces cerevisiae and Schizosaccharo myces pombe microarray data sets published in 82 different articles. yMGV couples data mining tools with a user-friendly web interface so that, with a few mouse clicks, one can identify the conditions that affect the expression of a gene or list of genes regulated in a set of experiments. One of the major new features we present here is a set of tools that allows for inter-organism comparisons. This should enable the fission yeast community to take advantage of the large amount of available information on budding yeast transcriptome. New tools and ongoing developments are also presented here.

Computational Biology↗

Delay-coordinates embeddings as a data mining tool for denoising speech signals.

In this paper, we utilize techniques from the theory of nonlinear dynamical systems to define a notion of embedding estimators. More specifically, we use delay-coordinates embeddings of sets of coefficients of the measured signal (in some chosen frame) as a data mining tool to separate structures that are likely to be generated by signals belonging to some predetermined data set. We implement the embedding estimator in a windowed Fourier frame, and we apply it to speech signals heavily corrupted by white noise. Our experimental work suggests that, after training on the data sets of interest, these estimators perform well for a variety of white noise processes and noise intensity levels.

Algorithms↗

A prediction model for patient classification according to nursing need: Using data mining techniques.

The purpose of this study was to construct a prediction model for patient classification according to nursing need. The results were assessed from the classification of the hospitalized cancer patients by three different data mining techniques: logistic regression, decision tree and neural network. Among these three techniques, neural network showed the best prediction power in ROC curve verification. The prediction model for patient classification developed by neural network based on nurse needs produced a prediction accuracy of 84.06%.

Adult↗

Application of a data-mining method based on Bayesian networks to lesion-deficit analysis.

Although lesion-deficit analysis (LDA) has provided extensive information about structure-function associations in the human brain, LDA has suffered from the difficulties inherent to the analysis of spatial data, i.e., there are many more variables than subjects, and data may be difficult to model using standard distributions, such as the normal distribution. We herein describe a Bayesian method for LDA; this method is based on data-mining techniques that employ Bayesian networks to represent structure-function associations. These methods are computationally tractable, and can represent complex, nonlinear structure-function associations. When applied to the evaluation of data obtained from a study of the psychiatric sequelae of traumatic brain injury in children, this method generates a Bayesian network that demonstrates complex, nonlinear associations among lesions in the left caudate, right globus pallidus, right side of the corpus callosum, right caudate, and left thalamus, and subsequent development of attention-deficit hyperactivity disorder, confirming and extending our previous statistical analysis of these data. Furthermore, analysis of simulated data indicates that methods based on Bayesian networks may be more sensitive and specific for detecting associations among categorical variables than methods based on chi-square and Fisher exact statistics.

Algorithms↗