PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data mining”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Surgeon and type of anesthesia predict variability in surgical procedure times.

BACKGROUND: Variability in surgical procedure times increases the cost of healthcare delivery by increasing both the underutilization and overutilization of expensive surgical resources. To reduce variability in surgical procedure times, we must identify and study its sources. METHODS: Our data set consisted of all surgeries performed over a 7-yr period at a large teaching hospital, resulting in 46,322 surgical cases. To study factors associated with variability in surgical procedure times, data mining techniques were used to segment and focus the data so that the analyses would be both technically and intellectually feasible. The data were subdivided into 40 representative segments of manageable size and variability based on headers adopted from the common procedural terminology classification. Each data segment was then analyzed using a main-effects linear model to identify and quantify specific sources of variability in surgical procedure times. RESULTS: The single most important source of variability in surgical procedure times was surgeon effect. Type of anesthesia, age, gender, and American Society of Anesthesiologists risk class were additional sources of variability. Intrinsic case-specific variability, unexplained by any of the preceding factors, was found to be highest for shorter surgeries relative to longer procedures. Variability in procedure times among surgeons was a multiplicative function (proportionate to time) of surgical time and total procedure time, such that as procedure times increased, variability in surgeons' surgical time increased proportionately. CONCLUSIONS: Surgeon-specific variability should be considered when building scheduling heuristics for longer surgeries. Results concerning variability in surgical procedure times due to factors such as type of anesthesia, age, gender, and American Society of Anesthesiologists risk class may be extrapolated to scheduling in other institutions, although specifics on individual surgeons may not. This research identifies factors associated with variability in surgical procedure times, knowledge of which may ultimately be used to improve surgical scheduling and operating room utilization.

Adolescent↗

The Human Genome Project and the role of genetics in health care.

The Human Genome Project, the mapping of our 100,000 genes and the sequencing of all of our DNA, will have major impact on biomedical research and the therapeutic and preventive health care. The tracing of genetic diseases to their molecular causes is rapidly expanding diagnostic and preventive options, while the increased insights into molecular pathways open tremendous perspectives for pharmacological and genetic therapies. The design of animal model systems for the functional study of disease and development of bioinformatics and biostatistics to improve our pattern recognition abilities are greatly accelerating progress. However, the optimal value from the current explosion of 'data mining' possibilities will only be gained when the basic data are made and kept publicly accessible, at the same time preventing the jeopardisation of the protection of intellectual property, arising from downstream inventions. This is one of the goals of HUGO, the international Human Genome Organisation, established 9 years ago to assist coordinating data acquisition and exchange and societal implementation of the genome project. Additional points of major importance in this historic endeavour are the safeguarding of a worldwide balance in the contribution and benefits to countries and population, the prevention of stigmatisation and discrimination of individuals and groups and the maintenance of respect for the priceless diversity of our world's cultures and traditions.

Delivery of Health Care↗

Classification algorithms applied to narrative reports.

Narrative text reports represent a significant source of clinical data. However, the information stored in these reports is inaccessible to many automated decision support systems. Data mining techniques can assist in extracting information from narrative data. Multiple classification methods, such as rule generation, decision trees, Bayesian classifiers, and information retrieval were used to classify a set of 200 chest X-ray reports according to 6 clinical conditions indicated. A general-purpose natural language processor was used to convert the narrative text into a coded form that could be used by the classification algorithms. Significant differences in performance were found between algorithms. The best performing algorithm applied to the processor output was significantly better than information retrieval applied to raw text. Predictor variables from the coded processor output were limited to avoid overfitting. Methods that limited by domain knowledge performed significantly better than those that limited by conditional probabilities of the variables in the training set. Algorithms were also shown to be dependent on training set size.

Algorithms↗

Mining association rules with improved semantics in medical databases.

The discovery of new knowledge by mining medical databases is crucial in order to make an effective use of stored data, enhancing patient management tasks. One of the main objectives of data mining methods is to provide a clear and understandable description of patterns held in data. We introduce a new approach to find association rules among quantitative values in relational databases. The semantics of such rules are improved by introducing imprecise terms in both the antecedent and the consequent, as these terms are the most commonly used in human conversation and reasoning. The terms are modeled by means of fuzzy sets defined in the appropriate domains. However, the mining task is performed on the precise data. These "fuzzy association rules" are more informative than rules relating precise values. We also introduce a new measure of accuracy, based on Shortliffe and Buchanan's certainty factors [Shortliffe E, Buchanan B. Math Biosci 1975;23:351-79]. Also, the semantics of the usual measure of usefulness of an association rule, called support are discussed and some new criteria are introduced. Our new measures have been shown to be more understandable and appropriate than ordinary ones. Several experiments on large medical databases show that our new approach can provide useful knowledge with better semantics in this field.

Databases, Factual↗

Can the US minimum data set be used for predicting admissions to acute care facilities?

This paper is intended to give an overview of Knowledge Discovery in Large Datasets (KDD) and data mining applications in healthcare particularly as related to the Minimum Data Set, a resident assessment tool which is used in US long-term care facilities. The US Health Care Finance Administration, which mandates the use of this tool, has accumulated massive warehouses of MDS data. The pressure in healthcare to increase efficiency and effectiveness while improving patient outcomes requires that we find new ways to harness these vast resources. The intent of this preliminary study design paper is to discuss the development of an approach which utilizes the MDS, in conjunction with KDD and classification algorithms, in an attempt to predict admission from a long-term care facility to an acute care facility. The use of acute care services by long term care residents is a negative outcome, potentially avoidable, and expensive. The value of the MDS warehouse can be realized by the use of the stored data in ways that can improve patient outcomes and avoid the use of expensive acute care services. This study, when completed, will test whether the MDS warehouse can be used to describe patient outcomes and possibly be of predictive value.

Algorithms↗

[Differences in ethnicity and emergency department visits in the Negev].

The population of the Negev consists mainly of Jews and Bedouin, who have very different life styles. Patients of both ethnic groups use our emergency department exclusively, providing a unique opportunity to study comparative patient habits. In gathering and processing the information we used Data Mining technology, which allows search for unique patterns in large data bases. We examined demographic data on some 64,000 emergency department visits during 1997-8, mostly medical and surgical cases, but not trauma cases. Many more were by Bedouin than Jews, and between the ages of 25 and 44, more by women than men. There were changes in trends in comparison with an arrival survey conducted some 11 years before.

Adult↗

Gene expression informatics--it's all in your mine.

Technologies for whole-genome RNA expression studies are becoming increasingly reliable and accessible. However, universal standards to make the data more suitable for comparative analysis and for inter-operability with other information resources have yet to emerge. Improved access to large electronic data sets, reliable and consistent annotation and effective tools for 'data mining' are critical. Analysis methods that exploit large data warehouses of gene expression experiments will be necessary to realize the full potential of this technology.

Animals↗

Quantitative collagen as a golden standard in differential diagnosing of fibrotic changes in liver tissue.

Determining a presence and degree of liver fibrosis provides means for diagnosing disease related processes. We have used two data mining methods, discriminant and regression analyses, to acquire knowledge from the data of 211 patients. We have shown and discussed that quantitative collagen has a distinguished discriminating power and can serve as a golden standard. We have additionally succeeded to obtain a formula consisting of standardised blood tests that can replace quantitative collagen. Practical implications of this is a non-invasive and cost efficient patient examination. All the results are now left for clinical evaluation and so is the current way of histopathological classifications.

Biopsy↗

Liver guide for monitoring of chronic hepatitis C.

The severity of chronic hepatitis C infection in the individual patient is monitored using blood laboratory findings and liver biopsy. If blood test results could be shown to provide sufficient information concerning the disease, the invasive procedure of liver biopsy could perhaps be avoided in some instances. This study assessed the clinical relevance of blood laboratory tests for detecting disease-related changes in the liver. Histopathological classification was used to assign class membership of the patients and data mining operations were performed in an elaborate way on 19 different data sets. Disease activity could be detected by a small set of blood tests. Extended sets could identify more severe changes, but failed to distinguish them. The extracted rules are implemented as a part of the knowledge base of a corresponding decision support system aimed at specialists and general practitioners.

Analysis of Variance↗

Using data warehousing and OLAP in public health care.

The paper describes the possibilities of using data warehousing and OLAP technologies in public health care in general and then our own experience with these technologies gained during the implementation of a data warehouse of outpatient data at the national level. Such a data warehouse serves as a basis for advanced decision support systems based on statistical, OLAP or data mining methods. We used OLAP to enable interactive exploration and analysis of the data. We found out that data warehousing and OLAP are suitable for the domain of public health and that they enable new analytical possibilities in addition to the traditional statistical approaches.

Ambulatory Care↗

Applications of qualitative multi-attribute decision models in health care.

Hierarchical decision models are a general decision support methodology aimed at the classification or evaluation of options that occur in decision-making processes. They are also important for the analysis, simulation and explanation of options. Decision models are typically developed through the decomposition of complex decision problems into smaller and less complex subproblems; the result of such decomposition is a hierarchical structure that consists of attributes and utility functions. This article presents an approach to the development and application of qualitative hierarchical decision models that is based on DEX, an expert system shell for multi-attribute decision support. The distinguishing characteristics of DEX are the use of qualitative (symbolic) attributes, and 'if-then' decision rules. Also, DEX provides a number of methods for the analysis of models and options, such as selective explanation and what-if analysis. We demonstrate the applicability and flexibility of the approach presenting four real-life applications of DEX in health care: assessment of breast cancer risk, assessment of basic living activities in community nursing, risk assessment in diabetic foot care, and technical analysis of radiogram errors. In particular, we highlight and justify the importance of knowledge presentation and option analysis methods for practical decision-making. We further show that, using a recently developed data mining method called HINT, such hierarchical decision models can be discovered from retrospective patient data.

Breast Neoplasms↗

Systematic functional evaluation of CNGA1 missense variants associated with retinitis pigmentosa.

BACKGROUND: Missense variants are frequently classified as variants of uncertain significance (VUS) according to the guidelines of the American College of Medical Genetics and Genomics and the Association of Molecular Pathology (ACMG/AMP). Consequently, disease relevance remains elusive, impeding molecular genetic diagnostics, patients` and family genetic counseling, and identification of patients eligible for clinical trials. Functional studies are critical for resolving the clinical significance of VUS. CNGA1 encodes the main subunit of the rod cyclic nucleotide-gated (CNG) channel, a vital component of the phototransduction cascade. Variants in CNGA1 are a rare cause of autosomal recessive retinitis pigmentosa and a phase I/II gene augmentation trial (NCT06291935) is currently ongoing highlighting the necessity to differentiate benign from pathogenic variants. METHODS: CNGA1 missense variants compiled from retinal disease patient cohorts, public databases and literature were functionally investigated using a medium-throughput aequorin-based assay and in vitro minigene splice assays for predicted exonic spliceogenic variants. Functional data were correlated with the in silico prediction of five variant effect predictors (VEPs) and applied to support or revise variants' ACMG/AMP classification. RESULTS: Data mining revealed 86 missense CNGA1 variants - including three novel - most of them lacking functional data; 65.1% of the variants were initially classified as VUS. The aequorin-based assay showed that 72.1% of tested variants significantly impaired CNG channel function and were classified as functionally abnormal, while 23.3% were functionally normal and 5% remained functionally uncertain. Correlation of the functional data with in silico predictions identified AlphaMissense and CPT-1 to be the most suitable tools for assessing CNGA1 missense variants. Using in vitro minigene splice assays, two putative missense variants were shown to induce missplicing. Based on the functional findings, 62.1% of the variants initially classified as VUS were re-categorized as likely pathogenic or likely benign. Furthermore, 93.3% of the variants initially classified as likely pathogenic showed an effect on CNGA1 channel function, confirming their disease relevance and supporting their reclassification as pathogenic. CONCLUSION: This study represents the first comprehensive functional assessment of disease-associated CNGA1 missense variants, thus significantly advancing the understanding of their disease relevance and improving molecular genetic diagnostics in patients.

Humans↗

Information Management System for Site Remediation Efforts.

/ Environmental regulatory agencies are responsible for protecting human health and the environment in their constituencies. Their responsibilities include the identification, evaluation, and cleanup of contaminated sites. Leaking underground storage tanks (USTs) constitute a major source of subsurface and groundwater contamination. A significant portion of a regulatory body's efforts may be directed toward the management of UST-contaminated sites. In order to manage remedial sites effectively, vast quantities of information must be maintained, including analytical dataon chemical contaminants, remedial design features, and performance details. Currently, most regulatory agencies maintain such information manually. This makes it difficult to manage the data effectively. Some agencies have introduced automated record-keeping systems. However, the ad hoc approach in these endeavors makes it difficult to efficiently analyze, disseminate, and utilize the data. This paper identifies the information requirements for UST-contaminated site management at the Waste Cleanup Section of the Department of Environmental Resources Management in Dade County, Florida. It presents a viable design for an information management system to meet these requirements. The proposed solution is based on a back-end relational database management system with relevant tools for sophisticated data analysis and data mining. The database is designed with all tables in the third normal form to ensure data integrity, flexible access, and efficient query processing. In addition to all standard reports required by the agency, the system provides answers to ad hoc queries that are typically difficult to answer under the existing system. The database also serves as a repository of information for a decision support system to aid engineering design and risk analysis. The system may be integrated with a geographic information system for effective presentation and dissemination of spatial data.

Journal Article↗

Extended SQL for manipulating clinical warehouse data.

Health care institutions are beginning to collect large amounts of clinical data through patient care applications. Clinical data warehouses make these data available for complex analysis across patient records, benefiting administrative reporting, patient care and clinical research. Data gathered for patient care purposes are difficult to manipulate for analytic tasks; the schema presents conceptual difficulties for the analyst, and many queries perform poorly. An extension to SQL is presented that enables the analyst to designate groups of rows. These groups can then be manipulated and aggregated in various ways to solve a number of useful analytic problems. The extended SQL is concise and runs in linear time, while standard SQL requires multiple statements with polynomial performance. The extensions are extremely powerful for performing aggregations on large amounts of data, which is useful in clinical data mining applications.

Clinical Laboratory Techniques↗

Metabolite profiling for plant functional genomics.

Multiparallel analyses of mRNA and proteins are central to today's functional genomics initiatives. We describe here the use of metabolite profiling as a new tool for a comparative display of gene function. It has the potential not only to provide deeper insight into complex regulatory processes but also to determine phenotype directly. Using gas chromatography/mass spectrometry (GC/MS), we automatically quantified 326 distinct compounds from Arabidopsis thaliana leaf extracts. It was possible to assign a chemical structure to approximately half of these compounds. Comparison of four Arabidopsis genotypes (two homozygous ecotypes and a mutant of each ecotype) showed that each genotype possesses a distinct metabolic profile. Data mining tools such as principal component analysis enabled the assignment of "metabolic phenotypes" using these large data sets. The metabolic phenotypes of the two ecotypes were more divergent than were the metabolic phenotypes of the single-loci mutant and their parental ecotypes. These results demonstrate the use of metabolite profiling as a tool to significantly extend and enhance the power of existing functional genomics approaches.

Arabidopsis↗

Building manageable rough set classifiers.

An interesting aspect of techniques for data mining and knowledge discovery is their potential for generating hypotheses by discovering underlying relationships buried in the data. However, the set of possible hypotheses is often very large and the extracted models may become prohibitively complex. It is therefore typically desirable to only consider the "strongest" hypotheses, so that smaller models can be obtained that also retain good classificatory capabilities. This paper outlines how rule-based classifiers based on rough set theory and Boolean reasoning that are both small and perform well can be developed. Applied to a real-world medical dataset, the final models are shown to exhibit good performance using only a subset of the available information. Furthermore, the number of resulting rules is low and enables practical a posteriori inspection and interpretation of the models.

Classification↗

Application of a novel and fast information-theoretic method to the discovery of higher-order correlations in protein databases.

We present a fast, discrete data-mining approach to the problem of finding kappa-tuples of correlated amino acid residues in protein sequence data. When sets of sequence-distant sites display high mutual information, they may bespeak important structural or functional features. Our novel methodology overcomes the limitations of previous methods which examined only single-residue features or pairwise interactions.

AIDS Vaccines↗

SPECT electronic collimation resolution enhancement using chi-square minimization.

An electronic collimation technique is developed which utilizes the chi-square goodness-of-fit measure to filter scattered gammas incident upon a medical imaging detector. In this data mining technique, Compton kinematic expressions are used as the chi-square fitting templates for measured energy-deposition data involving multiple-interaction scatter sequences. Fit optimization is conducted using the Davidon variable metric minimization algorithm to simultaneously determine the best-fit gamma scatter angles and their associated uncertainties, with the uncertainty associated with the first scatter angle corresponding to the angular resolution precision for the source. The methodology requires no knowledge of materials and geometry. This pattern recognition application enhances the ability to select those gammas that will provide the best resolution for input to reconstruction software. Illustrative computational results are presented for a conceptual truncated-ellipsoid polystyrene position-sensitive fibre head-detector Monte Carlo model using a triple Compton scatter gamma sequence assessment for a 99mTc point source. A filtration rate of 94.3% is obtained, resulting in an estimated sensitivity approximately three orders of magnitude greater than a high-resolution mechanically collimated device. The technique improves the nominal single-scatter angular resolution by up to approximately 24 per cent as compared with the conventional analytic electronic collimation measure.

Algorithms↗