PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Data Mining”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Report generation and data mining in the domain of thoracic surgery.

As a part of AssistMe system, the reporting system has been developed for the thoracic surgery domain. Reporting System is defined as software for dynamic report generation purpose and based on the data-mining techniques. The target users of the future reporting system-physicians, administrative staff, and patients-have been identified. Two major types of clinical reports have been found: predefined and customized. The decision of splitting reports into groups has been taken mainly because users were heterogeneous and had different access rights to the sensitive information. Data-mining process in the reporting system is based on descriptive statistics. It allows dynamically mined AssistMe databases and generates statistical reports about patient's morbidity, mortality, and comorbidity. Information is visualized in the chart way and can be also observed in tabular form. User interaction is also supported by the system.

Heart-Assist Devices↗

Application of a data-mining technique to analyze coprescription patterns for antacids in Taiwan.

BACKGROUND: Although antacids are popular drugs with a long history of use, their true utilization patterns-including over-the-counter use-have rarely been documented. Because all antacids are reimbursed under the National Health Insurance program in Taiwan, it is possible to access and analyze nationwide data for these drugs. OBJECTIVES: The purposes of this study were to estimate the scale of antacid prescribing in Taiwan using the national insurance claims for outpatient services and to analyze coprescribing patterns of antacids using modern data-mining techniques. METHODS: The National Health Insurance Research Database in Taiwan supplied the visit-based sampling data sets, which had a sampling ratio of 0.2% for all claims for outpatient medical services in the year 2000. In addition to the plain statistics (ie, data from simple calculations) for antacid prescriptions, we also analyzed relationships between prescriptions for antacids and nonantacid drugs. A data-mining technique-association rule mining-was applied to identify the drugs prescribed in combination with antacids. RESULTS: Among a total of 409,049 eligible prescriptions for 1,704,595 drug items to be administered orally, antacids were present in 213,494 prescriptions (52.2%). Antacid users were generally older than nonusers (mean [SD] age, 39.9 [23.4] years vs 32.4 [25.7] years). In all, 88.8% of antacid items (189,531/213,494) were prescribed without claims diagnoses of gastrointestinal disorders. Using association rule mining with a 1.0% minimum support factor, there were 36 strong association rules between prescriptions for antacids and other drug subgroups at the third level of Anatomical Therapeutic Chemical classification. Nonsteroidal anti-inflammatory drugs and drugs for treating upper respiratory infections played dominant roles in the associations with antacid prescriptions; vitamin B complex and antivertigo preparations were also strongly associated with antacids. CONCLUSIONS: Antacid coprescriptions were common in Taiwan in the year 2000. Further study should investigate whether antacid prescribing patterns are influenced by Taiwanese perceptions that Western drugs injure the stomach.

Adolescent↗

Data mining approach to model the diagnostic service management.

Korea has National Health Insurance Program operated by the government-owned National Health Insurance Corporation, and diagnostic services are provided every two year for the insured and their family members. Developing a customer relationship management (CRM) system using data mining technology would be useful to improve the performance of diagnostic service programs. Under these circumstances, this study developed a model for diagnostic service management taking into account the characteristics of subjects using a data mining approach. This study could be further used to develop an automated CRM system contributing to the increase in the rate of receiving diagnostic services.

Databases, Factual↗

Pathology information systems: data mining leads to knowledge discovery.

Information systems in pathology provide opportunities for pathologists and clinical laboratory scientists to impact both clinical care and modern research agendas. The paradigm shift in health care from individualized care to population-based and standardized delivery systems has created both of these opportunities. In research, pathology information systems can provide key databases for health services research and new informatics-based approaches to database research. The latter is characterized by utilization of pathology databases for data mining to discover new patterns that provide new knowledge. The multidisciplinary knowledge discovery and data mining program at the University of Alabama at Birmingham focuses on this health care application, which has the potential to make a major impact on health care research and delivery.

Artificial Intelligence↗

Confirmation of data mining based predictions of protein function.

MOTIVATION: A central problem in bioinformatics is the assignment of function to sequenced open reading frames (ORFs). The most common approach is based on inferred homology using a statistically based sequence similarity (SIM) method, e.g. PSI-BLAST. Alternative non-SIM based bioinformatic methods are becoming popular. One such method is Data Mining Prediction (DMP). This is based on combining evidence from amino-acid attributes, predicted structure and phylogenic patterns; and uses a combination of Inductive Logic Programming data mining, and decision trees to produce prediction rules for functional class. DMP predictions are more general than is possible using homology. In 2000/1, DMP was used to make public predictions of the function of 1309 Escherichia coli ORFs. Since then biological knowledge has advanced allowing us to test our predictions. RESULTS: We examined the updated (20.02.02) Riley group genome annotation, and examined the scientific literature for direct experimental derivations of ORF function. Both tests confirmed the DMP predictions. Accuracy varied between rules, and with the detail of prediction, but they were generally significantly better than random. For voting rules, accuracies of 75-100% were obtained. Twenty-one of these DMP predictions have been confirmed by direct experimentation. The DMP rules also have interesting biological explanations. DMP is, to the best of our knowledge, the first non-SIM based prediction method to have been tested directly on new data. AVAILABILITY: We have designed the "Genepredictions" database for protein functional predictions. This is intended to act as an open repository for predictions for any organism and can be accessed at http://www.genepredictions.org

Abstracting and Indexing↗

Safety related drug-labelling changes: findings from two data mining algorithms.

INTRODUCTION: With increasing volumes of postmarketing safety surveillance data, data mining algorithms (DMAs) have been developed to search large spontaneous reporting system (SRS) databases for disproportional statistical dependencies between drugs and events. A crucial question is the proper deployment of such techniques within the universe of methods historically used for signal detection. One question of interest is comparative performance of algorithms based on simple forms of disproportionality analysis versus those incorporating Bayesian modelling. A potential benefit of Bayesian methods is a reduced volume of signals, including false-positive signals. OBJECTIVE: To compare performance of two well described DMAs (proportional reporting ratios [PRRs] and an empirical Bayesian algorithm known as multi-item gamma Poisson shrinker [MGPS]) using commonly recommended thresholds on a diverse data set of adverse events that triggered drug labelling changes. METHODS: PRRs and MGPS were retrospectively applied to a diverse sample of drug-event combinations (DECs) identified on a government Internet site for a 7-month period. Metrics for this comparative analysis included the number and proportion of these DECs that generated signals of disproportionate reporting with PRRs, MGPS, both or neither method, differential timing of signal generation between the two methods, and clinical nature of events that generated signals with only one, both or neither method. RESULTS: There were 136 relevant DECs that triggered safety-related labelling changes for 39 drugs during a 7-month period. PRRs generated a signal of disproportionate reporting with almost twice as many DECs as MGPS (77 vs 40). No DECs were flagged by MGPS only. PRRs highlighted DECs in advance of MGPS (1-15 years) and a label change (1-30 years). For 59 DECs, there was no signal with either DMA. DECs generating signals of disproportionate reporting with only PRRs were both medically serious and non-serious. DISCUSSION/CONCLUSION: In most instances in which a DEC generated a signal of disproportionate reporting with both DMAs (almost twice as many with PRRs), the signal was generated using PRRs in advance of MGPS. No medically important events were signalled only by MGPS. It is likely that the incremental utility of DMAs are highly situation-dependent. It is clear, however, that the volume of signals generated by itself is an inadequate criterion for comparison and that clinical nature of signalled events and differential timing of signals needs to be considered. Accepting commonly recommended threshold criteria for DMAs examined in this study as universal benchmarks for signal detection is not justified.

Algorithms↗

Genomic comparison using data mining techniques based on a possibilistic fuzzy sets model.

Current copiousness of genomic information stored in biological databases [Mar Albà, M., Lee, M., Pearl, D., Shepherd, F.M.G., Martin, A.J., Orengo, N., Kellam, C.A., 2001. P. VIDA: a virus database system for the organisation of virus genome open reading frames. Nuleic Acids Res. 133-136] makes ultimately feasible the proposal for an application of knowledge management aimed to discover general rules in subcellular phenomena. The goal of this work is primarily to discover relationships between genes by microarray analysis. The tools exploited come from clustering techniques and are mainly based on Knowledge Discovery in Databases (KDD) concepts [Fayyad, U., Piatetsky-Shapiro, G., Smyth, P., 1996. From data mining to knowledge discovery in databases. AI Magazine 17(3), 37-54]. Starting from a data set, each element can be represented by a characteristic matrix, which sums up all data attributes. In this case data mining is oriented to perform a Pattern Recognition of related sequences, hidden in databases [Hand, D.J., Nicholas, A., 2005. Heard finding groups in gene expression data. J. Biomed. Biotechnol. 215-225]. Following a bottom up approach, the next refinement is to compare retrieved data to gather similar features, by dedicated clustering algorithms [Kaufman, L., Rousseeuw, P.J., 1990. Finding groups in data. An Introduction to Cluster Analysis. John Wiley & Sons, New York; Forman, G., Zhang, B., 2000. Distributed Data clustering can be efficient and exact HP. Laboratories Palo Alto HPL-2000, p. 158], driven by fuzzy logic, allowing us to perceive by intuition a common denominator for various genomic families and to anticipate likely future developments.

Algorithms↗

From data collection to knowledge data discovery: a medical application of data mining.

Prison inmates are exposed to a variety of major risk factors (psychiatric disorders, suicide attempts, illicit drug use). From 1986 to 1996, the USA prison population more than doubled while in France, it increased from 35655 in 1980 to 51623 in 1995. In spite of these findings, very little information concerning the inmates population is available. At the present time, there is a desire to adopt a policy based on the prevention of recidivism, on adequate release planning and referrals to community-based services. The aim of the RAPPEL project was to build an information system for assessing the social and health status of prison inmates. The pilot project was set up at the prison of Loos and allowed the collection and analysis of nearly 15000 records. The aim of this paper is to present the extension of the project consisting in developing a regional network grouping 11 jails. Information locally available will serve as the basis for the information system of regional jails. Data mining techniques will provide solutions for the extraction of new information. Three data mining tools were experimented : association rules, classification trees and clustering. Further extension consists in a distributed approach allowing direct access to the information system by WEB tools.

Classification↗

Novel approaches to visualization and data mining reveals diagnostic information in the low amplitude region of serum mass spectra from ovarian cancer patients.

The ability to identify patterns of diagnostic signatures in proteomic data generated by high throughput mass spectrometry (MS) based serum analysis has recently generated much excitement and interest from the scientific community. These data sets can be very large, with high-resolution MS instrumentation producing 1-2 million data points per sample. Approaches to analyze mass spectral data using unsupervised and supervised data mining operations would greatly benefit from tools that effectively allow for data reduction without losing important diagnostic information. In the past, investigators have proposed approaches where data reduction is performed by a priori "peak picking" and alignment/warping/smoothing components using rule-based signal-to-noise measurements. Unfortunately, while this type of system has been employed for gene microarray analysis, it is unclear whether it will be effective in the analysis of mass spectral data, which unlike microarray data, is comprised of continuous measurement operations. Moreover, it is unclear where true signal begins and noise ends. Therefore, we have developed an approach to MS data analysis using new types of data visualization and mining operations in which data reduction is accomplished by culling via the intensity of the peaks themselves instead of by location. Applying this new analysis method on a large study set of high resolution mass spectra from healthy and ovarian cancer patients, shows that all of the diagnostic information is contained within the very lowest amplitude regions of the mass spectra. This region can then be selected and studied to identify the exact location and amplitude of the diagnostic biomarkers.

Blood Proteins↗

Using data mining to characterize DNA mutations by patient clinical features.

In most hereditary cancer syndromes, finding a correspondence between various genetic mutations within a gene (genotype) and a patient's clinical cancer history (phenotype) is challenging; to date there are few clinically meaningful correlations between specific DNA intragenic mutations and corresponding cancer types. To define possible genotype and phenotype correlations, we evaluated the application of data mining methodology whereby the clinical cancer histories of gene-mutation-positive patients were used to define valid or "true" patterns for a specific DNA intragenic mutation. The clinical histories of patients with their corresponding detailed attributes without the same oncologic intragenic mutation were labeled incorrect or "false" patterns. The results of data mining technology yielded characterizing rules for the true cases that constituted clinical features which predicted the intragenic mutation. Some of the initial results derived correlations already independently known in the literature, adding to the confidence of using this methodological approach.

Age of Onset↗

Whole-body gene expression by data mining.

To date, a comprehensive survey of the expression of lysyl oxidase (LOX), lysyl oxidase-like 1 (LOXL1), and lysyl oxidase-like 2 (LOXL2) has yet to be performed. The use of in vitro strategies to accomplish this task would prove daunting as it is both time-consuming and costly. We present a new in silico data mining strategy that directly addresses these limitations. Sequences corresponding to the 3' untranslated regions of LOX, LOXL1, and LOXL2 were individually queried against the human expressed sequence tag database (dbEST). In this manner, the entire tissue repertoire available in the dbEST was surveyed. This provided an estimate of the levels of mRNA transcripts in a variety of adult and fetal tissues. We have also employed this strategy to determine the pattern of expression and levels of a newly discovered gene, CGI-15. The veracity of this technique has been independently assessed by semiquantitative PCR analysis. The application of this technology is bounded only by the ever-growing information available in the GenBank, UniGene, and human EST databases. The utility of our data mining strategy to establish relative transcript levels in numerous tissues is presented.

3' Untranslated Regions↗

Prediction of blood glucose level of type 1 diabetics using response surface methodology and data mining.

In order to improve the accuracy of predicting blood glucose levels, it is necessary to obtain details about the lifestyle and to optimize the input variables dependent on diabetics. In this study, using four subjects who are type 1 diabetics, the fasting blood glucose level (FBG), metabolic rate, food intake, and physical condition are recorded for more than 5 months as a preliminary study. Then, using data mining, an estimation model of FBG is obtained, and subsequently, the trend in fluctuations in the next morning's glucose level is predicted. The subject's physical condition is self-assessed on a scale from positive (1) to negative (5), and the values are set as the physical condition variable. By adding the physical condition variable to the input variables for the data mining, the accuracy of the FBG prediction is improved. In order to determine more appropriate input variables from the biological information reflecting on the subject's glucose metabolism, response surface methodology (RSM) is employed. As a result, using the variables exhibiting positive correlations with the FBG in the RSM, the accuracy of the FBG prediction improved. Conditions could be found such that the accuracy of the predicting trends in fluctuations in blood glucose level reached around 80%. The prediction method of the trend in fluctuations in the next morning's glucose levels might be useful to improve the quality of life of type 1 diabetics through insulin treatment, and to prevent hypoglycemia.

Adult↗

Data mining tools for the Saccharomyces cerevisiae morphological database.

For comprehensive understanding of precise morphological changes resulting from loss-of-function mutagenesis, a large collection of 1,899,247 cell images was assembled from 91,71 micrographs of 4782 budding yeast disruptants of non-lethal genes. All the cell images were processed computationally to measure approximately 500 morphological parameters in individual mutants. We have recently made this morphological quantitative data available to the public through the Saccharomyces cerevisiae Morphological Database (SCMD). Inspecting the significance of morphological discrepancies between the wild type and the mutants is expected to provide clues to uncover genes that are relevant to the biological processes producing a particular morphology. To facilitate such intensive data mining, a suite of new software tools for visualizing parameter value distributions was developed to present mutants with significant changes in easily understandable forms. In addition, for a given group of mutants associated with a particular function, the system automatically identifies a combination of multiple morphological parameters that discriminates a mutant group from others significantly, thereby characterizing the function effectively. These data mining functions are available through the World Wide Web at http://scmd.gi.k.u-tokyo.ac.jp/.

Computer Graphics↗

Knowledge representation forms for data mining methodologies as applied in thoracic surgery.

Typical ways of disseminating and using results of clinical research are scientific journals and reports. Presentation forms are condensed and comprehensible mainly to the experts following the specific topics. A vast amount of information remains unutilized due to the complex form of presenting the knowledge. Subject of this research is to explore possibilities of representation and also visualization of the results obtained using data mining methodologies. The intention is to formulate more than scientific ways to communicate facts that are of interest for the clinicians, medical students and even patients. Internet technologies as already widely established media support knowledge representation forms such as hypertext documents and structured knowledge components. The "Assist Me" decision support system for surgical treatment of cardiac patients integrates several forms of data mining and representation methodologies. We are showing a feasibility study in which scientific outcomes were forwarded to a broad group of potential users.

Artificial Intelligence↗

Predicting breast cancer survivability: a comparison of three data mining methods.

OBJECTIVE: The prediction of breast cancer survivability has been a challenging research problem for many researchers. Since the early dates of the related research, much advancement has been recorded in several related fields. For instance, thanks to innovative biomedical technologies, better explanatory prognostic factors are being measured and recorded; thanks to low cost computer hardware and software technologies, high volume better quality data is being collected and stored automatically; and finally thanks to better analytical methods, those voluminous data is being processed effectively and efficiently. Therefore, the main objective of this manuscript is to report on a research project where we took advantage of those available technological advancements to develop prediction models for breast cancer survivability. METHODS AND MATERIAL: We used two popular data mining algorithms (artificial neural networks and decision trees) along with a most commonly used statistical method (logistic regression) to develop the prediction models using a large dataset (more than 200,000 cases). We also used 10-fold cross-validation methods to measure the unbiased estimate of the three prediction models for performance comparison purposes. RESULTS: The results indicated that the decision tree (C5) is the best predictor with 93.6% accuracy on the holdout sample (this prediction accuracy is better than any reported in the literature), artificial neural networks came out to be the second with 91.2% accuracy and the logistic regression models came out to be the worst of the three with 89.2% accuracy. CONCLUSION: The comparative study of multiple prediction models for breast cancer survivability using a large dataset along with a 10-fold cross-validation provided us with an insight into the relative prediction ability of different data mining methods. Using sensitivity analysis on neural network models provided us with the prioritized importance of the prognostic factors used in the study.

Breast Neoplasms↗

Reviewing mobile phases used on Chiralcel OD through an application of data mining tools to CHIRBASE database.

During the past decade, thousands of compounds have been resolved on Chiralcel OD (a cellulose-based chiral stationary phase) under diverse eluting conditions. Many researches have documented the effects of mobile phase on enantioselectivity for a given family of samples but today no comprehensive study aimed at identifying the associations between the structural features present on solute and appropriate mobile phase conditions has yet been proposed. In this review of mobile phases used on Chiralcel OD, we try to go far beyond a simple enumeration of eluting conditions and an effort is made to explore the utility of data mining tools for assessing the knowledge contained in CHIRBASE database. We have extracted from CHIRBASE the chemical features of 2363 chiral compounds separated on Chiralcel OD and their corresponding mobile phases. This data set was submitted to data mining programs for molecular pattern recognition and mobile phase predictions for new cases. Some substructural characteristics of solutes were related to the efficient use of some specific mobile phases. For example, the application of CH3CN/salt buffer at pH 6-7 was found convenient for reversed-phase separation of compounds bearing a tertiary amine functional group. Furthermore, a cluster analysis allowed the arrangement of the mobile phases according to similarity found in molecular patterns of solutes. A decision tree, which may lead to a more rational choice of the mobile phase under reversed-phase conditions, is also proposed.

Carbamates↗

Data mining of Mycobacterium tuberculosis complex genotyping results using mycobacterial interspersed repetitive units validates the clonal structure of spoligotyping-defined families.

Recently, a combination of spoligotyping and bioinformatics was proposed as a potential tool for defining major circulating clades of tuberculosis bacilli. In the present study, we attempted to validate the above mentioned classification using a new high-throughput marker, named mycobacterial interspersed repetitive units (MIRUs). Using 12 MIRU loci and spoligotyping, we performed data mining of results on clinical isolates of the Mycobacterium tuberculosis complex representative of global mycobacterial allelic diversity. Knowledge rules permitting automatic labeling of major M. tuberculosis families were defined. Using this strategy, MIRU 24 appeared to be most appropriate for classifying our dataset. The Bovis family was shown to be perfectly classified by a maximum of 3 MIRUs, followed by Africanum and East African Indian (EAI) families by 4 MIRUs, the Beijing family by 6 MIRUs, Haarlem and X families by 8 MIRUs, the T family by 9, and the Latin-American and Mediterranean (LAM) family by 10 MIRUs. Considering the hierarchy of family divergence, our results corroborate a recent suggestion that EAI is the ancestral family followed by Africanum and Bovis. On the other hand, T, X, LAM and Haarlem families appear to be of more recent evolution. These results indicate that data mining of MIRUs is a valuable new tool for analyzing the evolutionary dynamics of the M. tuberculosis complex, and for monitoring an infectious disease such as tuberculosis.

Bacterial Typing Techniques↗