PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data mining”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Can the US minimum data set be used for predicting admissions to acute care facilities?

This paper is intended to give an overview of Knowledge Discovery in Large Datasets (KDD) and data mining applications in healthcare particularly as related to the Minimum Data Set, a resident assessment tool which is used in US long-term care facilities. The US Health Care Finance Administration, which mandates the use of this tool, has accumulated massive warehouses of MDS data. The pressure in healthcare to increase efficiency and effectiveness while improving patient outcomes requires that we find new ways to harness these vast resources. The intent of this preliminary study design paper is to discuss the development of an approach which utilizes the MDS, in conjunction with KDD and classification algorithms, in an attempt to predict admission from a long-term care facility to an acute care facility. The use of acute care services by long term care residents is a negative outcome, potentially avoidable, and expensive. The value of the MDS warehouse can be realized by the use of the stored data in ways that can improve patient outcomes and avoid the use of expensive acute care services. This study, when completed, will test whether the MDS warehouse can be used to describe patient outcomes and possibly be of predictive value.

Algorithms↗

[Differences in ethnicity and emergency department visits in the Negev].

The population of the Negev consists mainly of Jews and Bedouin, who have very different life styles. Patients of both ethnic groups use our emergency department exclusively, providing a unique opportunity to study comparative patient habits. In gathering and processing the information we used Data Mining technology, which allows search for unique patterns in large data bases. We examined demographic data on some 64,000 emergency department visits during 1997-8, mostly medical and surgical cases, but not trauma cases. Many more were by Bedouin than Jews, and between the ages of 25 and 44, more by women than men. There were changes in trends in comparison with an arrival survey conducted some 11 years before.

Adult↗

Gene expression informatics--it's all in your mine.

Technologies for whole-genome RNA expression studies are becoming increasingly reliable and accessible. However, universal standards to make the data more suitable for comparative analysis and for inter-operability with other information resources have yet to emerge. Improved access to large electronic data sets, reliable and consistent annotation and effective tools for 'data mining' are critical. Analysis methods that exploit large data warehouses of gene expression experiments will be necessary to realize the full potential of this technology.

Animals↗

Visualization and interactive analysis of blood parameters with InfoZoom.

This paper describes the application of the data analysis tool InfoZoom to a database containing the results of blood examinations for about 400 patients with a suspect of thrombosis. The main goal was to find correlations between the measurements and the occurrence of a thrombosis. No automatic method for data mining is used. Instead, InfoZoom uses a novel technique to display data sets as highly compressed tables which always fit completely onto the screen. The user can interactively explore animated tabular views of the data. In this way, the user gets a feeling of the data, detects interesting knowledge, and gains a deep understanding of the data set.

Databases, Factual↗

Advanced database methodology for the Collation of Connectivity data on the Macaque brain (CoCoMac).

The need to integrate massively increasing amounts of data on the mammalian brain has driven several ambitious neuroscientific database projects that were started during the last decade. Databasing the brain's anatomical connectivity as delivered by tracing studies is of particular importance as these data characterize fundamental structural constraints of the complex and poorly understood functional interactions between the components of real neural systems. Previous connectivity databases have been crucial for analysing anatomical brain circuitry in various species and have opened exciting new ways to interpret functional data, both from electrophysiological and from functional imaging studies. The eventual impact and success of connectivity databases, however, will require the resolution of several methodological problems that currently limit their use. These problems comprise four main points: (i) objective representation of coordinate-free, parcellation-based data, (ii) assessment of the reliability and precision of individual data, especially in the presence of contradictory reports, (iii) data mining and integration of large sets of partially redundant and contradictory data, and (iv) automatic and reproducible transformation of data between incongruent brain maps. Here, we present the specific implementation of the 'collation of connectivity data on the macaque brain' (CoCoMac) database (http://www.cocomac.org). The design of this database addresses the methodological challenges listed above, and focuses on experimental and computational neuroscientists' needs to flexibly analyse and process the large amount of published experimental data from tracing studies. In this article, we explain step-by-step the conceptual rationale and methodology of CoCoMac and demonstrate its practical use by an analysis of connectivity in the prefrontal cortex.

Animals↗

Quantitative collagen as a golden standard in differential diagnosing of fibrotic changes in liver tissue.

Determining a presence and degree of liver fibrosis provides means for diagnosing disease related processes. We have used two data mining methods, discriminant and regression analyses, to acquire knowledge from the data of 211 patients. We have shown and discussed that quantitative collagen has a distinguished discriminating power and can serve as a golden standard. We have additionally succeeded to obtain a formula consisting of standardised blood tests that can replace quantitative collagen. Practical implications of this is a non-invasive and cost efficient patient examination. All the results are now left for clinical evaluation and so is the current way of histopathological classifications.

Biopsy↗

Liver guide for monitoring of chronic hepatitis C.

The severity of chronic hepatitis C infection in the individual patient is monitored using blood laboratory findings and liver biopsy. If blood test results could be shown to provide sufficient information concerning the disease, the invasive procedure of liver biopsy could perhaps be avoided in some instances. This study assessed the clinical relevance of blood laboratory tests for detecting disease-related changes in the liver. Histopathological classification was used to assign class membership of the patients and data mining operations were performed in an elaborate way on 19 different data sets. Disease activity could be detected by a small set of blood tests. Extended sets could identify more severe changes, but failed to distinguish them. The extracted rules are implemented as a part of the knowledge base of a corresponding decision support system aimed at specialists and general practitioners.

Analysis of Variance↗

Comparisons of gene colinearity in genomes using GeneOrder2.0.

Comparative genomics is enhanced by data mining the rapidly expanding DNA sequence databases. Because of the immense amount of data, computational tools and methods are needed to augment traditional manual visualizations and manipulations of these data. GeneOrder2.0, a Java-based interactive software programme, organizes genome sequence data into tabular and graphical visualizations of the extent of colinearity of genes between any two chromosome genomes of < or =250 kilobases. Both GenBank and proprietary data can be analyzed with this tool.

Computational Biology↗

Using data warehousing and OLAP in public health care.

The paper describes the possibilities of using data warehousing and OLAP technologies in public health care in general and then our own experience with these technologies gained during the implementation of a data warehouse of outpatient data at the national level. Such a data warehouse serves as a basis for advanced decision support systems based on statistical, OLAP or data mining methods. We used OLAP to enable interactive exploration and analysis of the data. We found out that data warehousing and OLAP are suitable for the domain of public health and that they enable new analytical possibilities in addition to the traditional statistical approaches.

Ambulatory Care↗

Applications of qualitative multi-attribute decision models in health care.

Hierarchical decision models are a general decision support methodology aimed at the classification or evaluation of options that occur in decision-making processes. They are also important for the analysis, simulation and explanation of options. Decision models are typically developed through the decomposition of complex decision problems into smaller and less complex subproblems; the result of such decomposition is a hierarchical structure that consists of attributes and utility functions. This article presents an approach to the development and application of qualitative hierarchical decision models that is based on DEX, an expert system shell for multi-attribute decision support. The distinguishing characteristics of DEX are the use of qualitative (symbolic) attributes, and 'if-then' decision rules. Also, DEX provides a number of methods for the analysis of models and options, such as selective explanation and what-if analysis. We demonstrate the applicability and flexibility of the approach presenting four real-life applications of DEX in health care: assessment of breast cancer risk, assessment of basic living activities in community nursing, risk assessment in diabetic foot care, and technical analysis of radiogram errors. In particular, we highlight and justify the importance of knowledge presentation and option analysis methods for practical decision-making. We further show that, using a recently developed data mining method called HINT, such hierarchical decision models can be discovered from retrospective patient data.

Breast Neoplasms↗

Systematic functional evaluation of CNGA1 missense variants associated with retinitis pigmentosa.

BACKGROUND: Missense variants are frequently classified as variants of uncertain significance (VUS) according to the guidelines of the American College of Medical Genetics and Genomics and the Association of Molecular Pathology (ACMG/AMP). Consequently, disease relevance remains elusive, impeding molecular genetic diagnostics, patients` and family genetic counseling, and identification of patients eligible for clinical trials. Functional studies are critical for resolving the clinical significance of VUS. CNGA1 encodes the main subunit of the rod cyclic nucleotide-gated (CNG) channel, a vital component of the phototransduction cascade. Variants in CNGA1 are a rare cause of autosomal recessive retinitis pigmentosa and a phase I/II gene augmentation trial (NCT06291935) is currently ongoing highlighting the necessity to differentiate benign from pathogenic variants. METHODS: CNGA1 missense variants compiled from retinal disease patient cohorts, public databases and literature were functionally investigated using a medium-throughput aequorin-based assay and in vitro minigene splice assays for predicted exonic spliceogenic variants. Functional data were correlated with the in silico prediction of five variant effect predictors (VEPs) and applied to support or revise variants' ACMG/AMP classification. RESULTS: Data mining revealed 86 missense CNGA1 variants - including three novel - most of them lacking functional data; 65.1% of the variants were initially classified as VUS. The aequorin-based assay showed that 72.1% of tested variants significantly impaired CNG channel function and were classified as functionally abnormal, while 23.3% were functionally normal and 5% remained functionally uncertain. Correlation of the functional data with in silico predictions identified AlphaMissense and CPT-1 to be the most suitable tools for assessing CNGA1 missense variants. Using in vitro minigene splice assays, two putative missense variants were shown to induce missplicing. Based on the functional findings, 62.1% of the variants initially classified as VUS were re-categorized as likely pathogenic or likely benign. Furthermore, 93.3% of the variants initially classified as likely pathogenic showed an effect on CNGA1 channel function, confirming their disease relevance and supporting their reclassification as pathogenic. CONCLUSION: This study represents the first comprehensive functional assessment of disease-associated CNGA1 missense variants, thus significantly advancing the understanding of their disease relevance and improving molecular genetic diagnostics in patients.

Humans↗

Information Management System for Site Remediation Efforts.

/ Environmental regulatory agencies are responsible for protecting human health and the environment in their constituencies. Their responsibilities include the identification, evaluation, and cleanup of contaminated sites. Leaking underground storage tanks (USTs) constitute a major source of subsurface and groundwater contamination. A significant portion of a regulatory body's efforts may be directed toward the management of UST-contaminated sites. In order to manage remedial sites effectively, vast quantities of information must be maintained, including analytical dataon chemical contaminants, remedial design features, and performance details. Currently, most regulatory agencies maintain such information manually. This makes it difficult to manage the data effectively. Some agencies have introduced automated record-keeping systems. However, the ad hoc approach in these endeavors makes it difficult to efficiently analyze, disseminate, and utilize the data. This paper identifies the information requirements for UST-contaminated site management at the Waste Cleanup Section of the Department of Environmental Resources Management in Dade County, Florida. It presents a viable design for an information management system to meet these requirements. The proposed solution is based on a back-end relational database management system with relevant tools for sophisticated data analysis and data mining. The database is designed with all tables in the third normal form to ensure data integrity, flexible access, and efficient query processing. In addition to all standard reports required by the agency, the system provides answers to ad hoc queries that are typically difficult to answer under the existing system. The database also serves as a repository of information for a decision support system to aid engineering design and risk analysis. The system may be integrated with a geographic information system for effective presentation and dissemination of spatial data.

Journal Article↗

Extended SQL for manipulating clinical warehouse data.

Health care institutions are beginning to collect large amounts of clinical data through patient care applications. Clinical data warehouses make these data available for complex analysis across patient records, benefiting administrative reporting, patient care and clinical research. Data gathered for patient care purposes are difficult to manipulate for analytic tasks; the schema presents conceptual difficulties for the analyst, and many queries perform poorly. An extension to SQL is presented that enables the analyst to designate groups of rows. These groups can then be manipulated and aggregated in various ways to solve a number of useful analytic problems. The extended SQL is concise and runs in linear time, while standard SQL requires multiple statements with polynomial performance. The extensions are extremely powerful for performing aggregations on large amounts of data, which is useful in clinical data mining applications.

Clinical Laboratory Techniques↗

Metabolite profiling for plant functional genomics.

Multiparallel analyses of mRNA and proteins are central to today's functional genomics initiatives. We describe here the use of metabolite profiling as a new tool for a comparative display of gene function. It has the potential not only to provide deeper insight into complex regulatory processes but also to determine phenotype directly. Using gas chromatography/mass spectrometry (GC/MS), we automatically quantified 326 distinct compounds from Arabidopsis thaliana leaf extracts. It was possible to assign a chemical structure to approximately half of these compounds. Comparison of four Arabidopsis genotypes (two homozygous ecotypes and a mutant of each ecotype) showed that each genotype possesses a distinct metabolic profile. Data mining tools such as principal component analysis enabled the assignment of "metabolic phenotypes" using these large data sets. The metabolic phenotypes of the two ecotypes were more divergent than were the metabolic phenotypes of the single-loci mutant and their parental ecotypes. These results demonstrate the use of metabolite profiling as a tool to significantly extend and enhance the power of existing functional genomics approaches.

Arabidopsis↗

Implementing a data warehouse at Inglis Innovative Services.

Data warehouses, data marts, and data mining have been hot topics in the 1990s, offering the promise of a vault of corporate data ripe for decision making. As is true with all promising technologies, the key issue is how to get started. Implementation of a corporate data warehouse involves a lot more than spending a huge amount of money on hardware, software, and consultants. Successful implementation of a data warehouse involves a corporate treasure hunt--identifying and cataloging data. It involves data ownership, data integrity, and business process analysis to determine what the data are, who owns them, how reliable they are, and how they are processed. Finally, implementation of the warehouse drives the issue of how good the decisions are that are based on the information in the warehouse. This article presents a case study of how one healthcare facility dealt with the challenges of implementing a data warehouse.

Computer Communication Networks↗

What is bioinformatics? A proposed definition and overview of the field.

BACKGROUND: The recent flood of data from genome sequences and functional genomics has given rise to new field, bioinformatics, which combines elements of biology and computer science. OBJECTIVES: Here we propose a definition for this new field and review some of the research that is being pursued, particularly in relation to transcriptional regulatory systems. METHODS: Our definition is as follows: Bioinformatics is conceptualizing biology in terms of macromolecules (in the sense of physical-chemistry) and then applying "informatics" techniques (derived from disciplines such as applied maths, computer science, and statistics) to understand and organize the information associated with these molecules, on a large-scale. RESULTS AND CONCLUSIONS: Analyses in bioinformatics predominantly focus on three types of large datasets available in molecular biology: macromolecular structures, genome sequences, and the results of functional genomics experiments (e.g. expression data). Additional information includes the text of scientific papers and "relationship data" from metabolic pathways, taxonomy trees, and protein-protein interaction networks. Bioinformatics employs a wide range of computational techniques including sequence and structural alignment, database design and data mining, macromolecular geometry, phylogenetic tree construction, prediction of protein structure and function, gene finding, and expression data clustering. The emphasis is on approaches integrating a variety of computational methods and heterogeneous data sources. Finally, bioinformatics is a practical discipline. We survey some representative applications, such as finding homologues, designing drugs, and performing large-scale censuses. Additional information pertinent to the review is available over the web at http://bioinfo.mbb.yale.edu/what-is-it.

Computational Biology↗

Building manageable rough set classifiers.

An interesting aspect of techniques for data mining and knowledge discovery is their potential for generating hypotheses by discovering underlying relationships buried in the data. However, the set of possible hypotheses is often very large and the extracted models may become prohibitively complex. It is therefore typically desirable to only consider the "strongest" hypotheses, so that smaller models can be obtained that also retain good classificatory capabilities. This paper outlines how rule-based classifiers based on rough set theory and Boolean reasoning that are both small and perform well can be developed. Applied to a real-world medical dataset, the final models are shown to exhibit good performance using only a subset of the available information. Furthermore, the number of resulting rules is low and enables practical a posteriori inspection and interpretation of the models.

Classification↗

Application of a novel and fast information-theoretic method to the discovery of higher-order correlations in protein databases.

We present a fast, discrete data-mining approach to the problem of finding kappa-tuples of correlated amino acid residues in protein sequence data. When sets of sequence-distant sites display high mutual information, they may bespeak important structural or functional features. Our novel methodology overcomes the limitations of previous methods which examined only single-residue features or pairwise interactions.

AIDS Vaccines↗