PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Data Mining”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Data mining and electroencephalography.

An overview of data mining (DM) and its application to the analysis of DM and electroencephalography (EEG) is given by: (i) presenting a working definition of DM, (ii) motivating why EEG analysis is a challenging field of application for DM technology and (iii) by reviewing exemplary work on DM applied to EEG analysis. The current status of work on DM and EEG is discussed and some general conclusions are drawn.

Algorithms↗

Improving sensitivity in shotgun proteomics using a peptide-centric database with reduced complexity: protease cleavage and SCX elution rules from data mining of MS/MS spectra.

Correct identification of a peptide sequence from MS/MS data is still a challenging research problem, particularly in proteomic analyses of higher eukaryotes where protein databases are large. The scoring methods of search programs often generate cases where incorrect peptide sequences score higher than correct peptide sequences (referred to as distraction). Because smaller databases yield less distraction and better discrimination between correct and incorrect assignments, we developed a method for editing a peptide-centric database (PC-DB) to remove unlikely sequences and strategies for enabling search programs to utilize this peptide database. Rules for unlikely missed cleavage and nontryptic proteolysis products were identified by data mining 11 849 high-confidence peptide assignments. We also evaluated ion exchange chromatographic behavior as an editing criterion to generate subset databases. When used to search a well-annotated test data set of MS/MS spectra, we found no loss of critical information using PC-DBs, validating the methods for generating and searching against the databases. On the other hand, improved confidence in peptide assignments was achieved for tryptic peptides, measured by changes in DeltaCN and RSP. Decreased distraction was also achieved, consistent with the 3-9-fold decrease in database size. Data mining identified a major class of common nonspecific proteolytic products corresponding to leucine aminopeptidase (LAP) cleavages. Large improvements in identifying LAP products were achieved using the PC-DB approach when compared with conventional searches against protein databases. These results demonstrate that peptide properties can be used to reduce database size, yielding improved accuracy and information capture due to reduced distraction, but with little loss of information compared to conventional protein database searches.

Amino Acid Sequence↗

Applying data mining techniques to the mapping of complex disease genes.

The simulated sequence data for the Genetic Analysis Workshop 12 were analyzed using data mining techniques provided by SAS ENTERPRISE MINER Release 4.0 in addition to traditional statistical tests for linkage and association of genetic markers with disease status. We examined two ways of combining these approaches to make use of the covariate data along with the genotypic data. The result of incorporating data mining techniques with more classical methods is an improvement in the analysis, both by correctly classifying the affection status of more individuals and by locating more single nucleotide polymorphisms related to the disease, relative to analyses that use classical methods alone.

Chromosome Mapping↗

A data mining approach in home healthcare: outcomes and service use.

BACKGROUND: The purpose of this research is to understand the performance of home healthcare practice in the US. The relationships between home healthcare patient factors and agency characteristics are not well understood. In particular, discharge destination and length of stay have not been studied using a data mining approach which may provide insights not obtained through traditional statistical analyses. METHODS: The data were obtained from the 2000 National Home and Hospice Care Survey data for three specific conditions (chronic obstructive pulmonary disease, heart failure and hip replacement), representing nearly 580 patients from across the US. The data mining approach used was CART (Classification and Regression Trees). Our aim was twofold: 1) determining the drivers of home healthcare service outcomes (discharge destination and length of stay) and 2) examining the applicability of induction through data mining to home healthcare data. RESULTS: Patient age (85 and older) was a driving force in discharge destination and length of stay for all three conditions. There were also impacts from the type of agency, type of payment, and ethnicity. CONCLUSION: Patients over 85 years of age experience differential outcomes depending on the condition. There are also differential effects related to agency type by condition although length of stay was generally lower for hospital-based agencies. The CART procedure was sufficiently accurate in correctly classifying patients in all three conditions which suggests continuing utility in home health care.

Aftercare↗

Assessment of hierarchical clustering methodologies for proteomic data mining.

Hierarchical clustering methodology is a powerful data mining approach for a first exploration of proteomic data. It enables samples or proteins to be grouped blindly according to their expression profiles. Nevertheless, the clustering results depend on parameters such as data preprocessing, between-profile similarity measurement, and the dendrogram construction procedure. We assessed several clustering strategies by calculating the F-measure, a widely used quality metric. The combination, on logged matrix, of Pearson correlation and Ward's methods for data aggregation is among the best clustering strategies, at least with the data sets we studied. This study was carried out using PermutMatrix, a freely available software derived from transcriptomics.

Algorithms↗

Linear data mining the Wichita clinical matrix suggests sleep and allostatic load involvement in chronic fatigue syndrome.

OBJECTIVES: To provide a mathematical introduction to the Wichita (KS, USA) clinical dataset, which is all of the nongenetic data (no microarray or single nucleotide polymorphism data) from the 2-day clinical evaluation, and show the preliminary findings and limitations, of popular, matrix algebra-based data mining techniques. METHODS: An initial matrix of 440 variables by 227 human subjects was reduced to 183 variables by 164 subjects. Variables were excluded that strongly correlated with chronic fatigue syndrome (CFS) case classification by design (for example, the multidimensional fatigue inventory [MFI] data), that were otherwise self reporting in nature and also tended to correlate strongly with CFS classification, or were sparse or nonvarying between case and control. Subjects were excluded if they did not clearly fall into well-defined CFS classifications, had comorbid depression with melancholic features, or other medical or psychiatric exclusions. The popular data mining techniques, principle components analysis (PCA) and linear discriminant analysis (LDA), were used to determine how well the data separated into groups. Two different feature selection methods helped identify the most discriminating parameters. RESULTS: Although purely biological features (variables) were found to separate CFS cases from controls, including many allostatic load and sleep-related variables, most parameters were not statistically significant individually. However, biological correlates of CFS, such as heart rate and heart rate variability, require further investigation. CONCLUSIONS: Feature selection of a limited number of variables from the purely biological dataset produced better separation between groups than a PCA of the entire dataset. Feature selection highlighted the importance of many of the allostatic load variables studied in more detail by Maloney and colleagues in this issue [1] , as well as some sleep-related variables. Nonetheless, matrix linear algebra-based data mining approaches appeared to be of limited utility when compared with more sophisticated nonlinear analyses on richer data types, such as those found in Maloney and colleagues [1] and Goertzel and colleagues [2] in this issue.

Adult↗

Application of data mining approaches to drug delivery.

Computational approaches play a key role in all areas of the pharmaceutical industry from data mining, experimental and clinical data capture to pharmacoeconomics and adverse events monitoring. They will likely continue to be indispensable assets along with a growing library of software applications. This is primarily due to the increasingly massive amount of biology, chemistry and clinical data, which is now entering the public domain mainly as a result of NIH and commercially funded projects. We are therefore in need of new methods for mining this mountain of data in order to enable new hypothesis generation. The computational approaches include, but are not limited to, database compilation, quantitative structure activity relationships (QSAR), pharmacophores, network visualization models, decision trees, machine learning algorithms and multidimensional data visualization software that could be used to improve drug delivery after mining public and/or proprietary data. We will discuss some areas of unmet needs in the area of data mining for drug delivery that can be addressed with new software tools or databases of relevance to future pharmaceutical projects.

Computer Simulation↗

Large scale data mining approach for gene-specific standardization of microarray gene expression data.

MOTIVATION: The identification of the change of gene expression in multifactorial diseases, such as breast cancer is a major goal of DNA microarray experiments. Here we present a new data mining strategy to better analyze the marginal difference in gene expression between microarray samples. The idea is based on the notion that the consideration of gene's behavior in a wide variety of experiments can improve the statistical reliability on identifying genes with moderate changes between samples. RESULTS: The availability of a large collection of array samples sharing the same platform in public databases, such as NCBI GEO, enabled us to re-standardize the expression intensity of a gene using its mean and variation in the wide variety of experimental conditions. This approach was evaluated via the re-identification of breast cancer-specific gene expression. It successfully prioritized several genes associated with breast tumor, for which the expression difference between normal and breast cancer cells was marginal and thus would have been difficult to recognize using conventional analysis methods. Maximizing the utility of microarray data in the public database, it provides a valuable tool particularly for the identification of previously unrecognized disease-related genes. AVAILABILITY: A user friendly web-interface (http://compbio.sookmyung.ac.kr/~lage/) was constructed to provide the present large-scale approach for the analysis of GEO microarray data (GS-LAGE server).

Algorithms↗

An intelligent data mining model approach for adverse effects of hormone replacement therapy.

OBJECTIVES: A number of controversial studies have been reported on the potential risk of breast cancer caused by hormone replacement therapy (HRT). Some studies showed a positive relationship between HRT and breast cancer onset, but other studies have not confirmed these results. To clarify the contradictory outcomes in the relationships between HRT and the onset of breast cancer, we have designed an intelligent data mining model (IDM), which is able to find proper prognostic factors for cancer onset and provides alternate measures in interpretation of outcome of clinical data through hierarchies of attributes. METHODS: Based on the selection criteria, we selected 22 sets of random and case-control studies of the last 15 years, which identified any involvements of HRT with breast cancer. We analyzed the relationship between HRT and breast cancer using an IDM model consisting of data mining algorithms and public domain data mining tools. Prognostic factors which underline the major etiological dispositions of breast cancer were identified. RESULTS: The variables which are closely associated with cancer onset to some degree are age 60-69, age at menopause 40-49, parity 0, age 40-49, and types of menopause oophorectomy. An implementation of IDM model on overall pooled data indicated that there is no significant relationship between breast cancer onset and HRT. It is suggested that HRT patients with specific physiological and pathological conditions related with the higher ranks of prognostic factors may have a greater chance to get breast cancer. CONCLUSION: The results of this study may guide biomedical research directed at establishing the causal relationships between various medications and their complications, allowing an accurate assessment of efficacy and side effects of new therapeutic treatment in clinical trials without reliance on a large control population.

Adult↗

Data mining as a tool for research and knowledge development in nursing.

The ability to collect and store data has grown at a dramatic rate in all disciplines over the past two decades. Healthcare has been no exception. The shift toward evidence-based practice and outcomes research presents significant opportunities and challenges to extract meaningful information from massive amounts of clinical data to transform it into the best available knowledge to guide nursing practice. Data mining, a step in the process of Knowledge Discovery in Databases, is a method of unearthing information from large data sets. Built upon statistical analysis, artificial intelligence, and machine learning technologies, data mining can analyze massive amounts of data and provide useful and interesting information about patterns and relationships that exist within the data that might otherwise be missed. As domain experts, nurse researchers are in ideal positions to use this proven technology to transform the information that is available in existing data repositories into useful and understandable knowledge to guide nursing practice and for active interdisciplinary collaboration and research.

Algorithms↗

[DNA chip data mining].

DNA chip data routinely contain gene expression levels of thousands of genes and the analysis should be supported by various computational tools. To be brief, the analysis procedure consists of four steps including image scanning, image processing, mathematical interpretation and biological interpretation. In image processing step, we should detect the spots and measure the signals of the spots and the background. In mathematical interpretation step, first of all we should massage the measured signals to make them appropriate for further mathematical analysis. The massaged data could be analyzed by various computational methods especially when the data were generated for multiple samples comparisons. The clustering techniques including hierarchical clustering, k-means clustering, SOTA, SOM are the most popular methods in this step. Various other multivariate statistics and related machine learning techniques are being introduced and applied to DNA chip data analysis recently. And finally the most important step we should tackle is the biological interpretation task. Although the depth of the domain knowledge about the biological situation under which the data were generated is the most important factor to elucidate the biological context, it could be supported by various bioinformatics tools including MEDLINE abstract processing by NLP techniques or genetic network models constructed by Boolean networks algorithms.

Algorithms↗

Canonical correlation analysis for data reduction in data mining applied to predictive models for breast cancer recurrence.

Data mining methods can be used for extracting specific medical knowledge such as important predictors for recurrence of breast cancer in pertinent data material. However, when there is a huge quantity of variables in the data material it is first necessary to identify and select important variables. In this study we present a preprocessing method for selecting important variables in a dataset prior to building a predictive model.In the dataset, data from 5787 female patients were analysed. To cover more predictors and obtain a better assessment of the outcomes, data were retrieved from three different registers: the regional breast cancer, tumour markers, and cause of death registers. After retrieving information about selected predictors and outcomes from the different registers, the raw data were cleaned by running different logical rules. Thereafter, domain experts selected predictors assumed to be important regarding recurrence of breast cancer. After that, Canonical Correlation Analysis (CCA) was applied as a dimension reduction technique to preserve the character of the original data.Artificial Neural Network (ANN) was applied to the resulting dataset for two different analyses with the same settings. Performance of the predictive models was confirmed by ten-fold cross validation. The results showed an increase in the accuracy of the prediction and reduction of the mean absolute error.

Breast Neoplasms↗

WormBase: methods for data mining and comparative genomics.

WormBase is a comprehensive repository for information on Caenorhabditis elegans and related nematodes. Although the primary web-based interface of WormBase (http:// www.wormbase.org/) is familiar to most C. elegans researchers, WormBase also offers powerful data-mining features for addressing questions of comparative genomics, genome structure, and evolution. In this chapter, we focus on data mining at WormBase through the use of flexible web interfaces, custom queries, and scripts. The intended audience includes users wishing to query the database beyond the confines of the web interface or fetch data en masse. No knowledge of programming is necessary or assumed, although users with intermediate skills in the Perl scripting language will be able to utilize additional data-mining approaches.

Animals↗

A knowledge-driven agent-centred framework for data mining in EMG.

In this paper, we present a multi-agent framework for data mining in electromyography. This application, based on a web interface, provides a set of functionalities allowing to manipulate 1000 medical cases and more than 25,000 neurological tests stored in a medical database. The aim is to extract medical information using data mining algorithms and to supply a knowledge base with pertinent information. The multi-agent platform gives the possibility to distribute the data management process between several autonomous entities. This framework provides a parallel and flexible data manipulation.

Data Collection↗

Temporal data mining for the quality assessment of hemodialysis services.

OBJECTIVE: This paper describes the temporal data mining aspects of a research project that deals with the definition of methods and tools for the assessment of the clinical performance of hemodialysis (HD) services, on the basis of the time series automatically collected during hemodialysis sessions. METHODS: Intelligent data analysis and temporal data mining techniques are applied to gain insight and to discover knowledge on the causes of unsatisfactory clinical results. In particular, two new methods for association rule discovery and temporal rule discovery are applied to the time series. Such methods exploit several pre-processing techniques, comprising data reduction, multi-scale filtering and temporal abstractions. RESULTS: We have analyzed the data of more than 5800 dialysis sessions coming from 43 different patients monitored for 19 months. The qualitative rules associating the outcome parameters and the measured variables were examined by the domain experts, which were able to distinguish between rules confirming available background knowledge and unexpected but plausible rules. CONCLUSION: The new methods proposed in the paper are suitable tools for knowledge discovery in clinical time series. Their use in the context of an auditing system for dialysis management helped clinicians to improve their understanding of the patients' behavior.

Algorithms↗

ARTMAP neural networks for information fusion and data mining: map production and target recognition methodologies.

The Sensor Exploitation Group of MIT Lincoln Laboratory incorporated an early version of the ARTMAP neural network as the recognition engine of a hierarchical system for fusion and data mining of registered geospatial images. The Lincoln Lab system has been successfully fielded, but is limited to target/non-target identifications and does not produce whole maps. Procedures defined here extend these capabilities by means of a mapping method that learns to identify and distribute arbitrarily many target classes. This new spatial data mining system is designed particularly to cope with the highly skewed class distributions of typical mapping problems. Specification of canonical algorithms and a benchmark testbed has enabled the evaluation of candidate recognition networks as well as pre- and post-processing and feature selection options. The resulting mapping methodology sets a standard for a variety of spatial data mining tasks. In particular, training pixels are drawn from a region that is spatially distinct from the mapped region, which could feature an output class mix that is substantially different from that of the training set. The system recognition component, default ARTMAP, with its fully specified set of canonical parameter values, has become the a priori system of choice among this family of neural networks for a wide variety of applications.

Computer Simulation↗

Data mining methods find demographic predictors of preterm birth.

BACKGROUND: Preterm births in the United States increased from 11.0% to 11.4% between 1996 and 1997; they continue to be a complex healthcare problem in the United States. OBJECTIVE: The objective of this research was to compare traditional statistical methods with emerging new methods called data mining or knowledge discovery in databases in identifying accurate predictors of preterm births. METHOD: An ethnically diverse sample (N = 19,970) of pregnant women provided data (1,622 variables) for new methods of analysis. Preterm birth predictors were evaluated using traditional statistical and newer data mining analyses. RESULTS: Seven demographic variables (maternal age and binary coding for county of residence, education, marital status, payer source, race, and religion) yielded a .72 area under the curve using Receiving Operating Characteristic curves to test predictive accuracy. The addition of hundreds of other variables added only a .03 to the area under the curve. CONCLUSION: Similar results across data mining methods suggest that results are data-driven and not method-dependent, and that demographic variables offer a small set of parsimonious variables with reasonable accuracy in predicting preterm birth outcomes in a racially diverse population.

Data Collection↗

Prediction of cardiovascular risk in hemodialysis patients by data mining.

OBJECTIVES: The objective of this work was to contribute to the development, validation and application of data mining methods for prediction in decision support systems in medicine. The particular focus was on the prediction of cardiovascular risk factors in hemodialysis patients, specifically the interventricular septum (IVS) thickness of the heart of individual patients as an important quantitative indicator to diagnose left ventricular hypertrophy. The work was based on data from 63 long-term hemodialysis patients of the KfH Dialysis Centre in Jena, Germany. METHODS: The approach applied is based on data mining methods and involves four major steps: data based clustering, cluster based rule extraction, rulebase construction and cluster and rule based prediction. The methods employed include crisp and fuzzy algorithms. At each step, logical and medical validation of results was carried out. Different sets of randomly selected patient data were used to train, test and optimize the clusterbases and rulebases for prediction. RESULTS: Using the best clusterbase/rulebase combination designed, the IVS thickness cluster ('small' or 'large') was predicted correctly for 30 of the 35 patients with known IVS values in the training data set; no patient was predicted incorrectly and 5 were parity predicted. For the test data set, 4 of the 6 patients with known IVS values were predicted correctly, no patient incorrectly and 2 parity. These results did not substantially differ from those obtained using the second best clusterbase/rulebase combination which was finally recommended for use based on further performance criteria. The prediction of the IVS thickness clusters of the 22 patients with unknown IVS values also yielded good results that were (and could only be) validated by a medical individual risk assessment of these patients. CONCLUSIONS: The approach applied proved successful for the cluster and rule based prediction of a quantitative variable, such as IVS thickness, for individual patients from other variables relevant to the problem. The results obtained demonstrate the high potential of the approach and the methods developed and validated to support decision-making in hemodialysis and other fields of medicine by individual risk prediction.

Algorithms↗