PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data mining”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Gene expression informatics--it's all in your mine.

Technologies for whole-genome RNA expression studies are becoming increasingly reliable and accessible. However, universal standards to make the data more suitable for comparative analysis and for inter-operability with other information resources have yet to emerge. Improved access to large electronic data sets, reliable and consistent annotation and effective tools for 'data mining' are critical. Analysis methods that exploit large data warehouses of gene expression experiments will be necessary to realize the full potential of this technology.

Animals↗

Visualization and interactive analysis of blood parameters with InfoZoom.

This paper describes the application of the data analysis tool InfoZoom to a database containing the results of blood examinations for about 400 patients with a suspect of thrombosis. The main goal was to find correlations between the measurements and the occurrence of a thrombosis. No automatic method for data mining is used. Instead, InfoZoom uses a novel technique to display data sets as highly compressed tables which always fit completely onto the screen. The user can interactively explore animated tabular views of the data. In this way, the user gets a feeling of the data, detects interesting knowledge, and gains a deep understanding of the data set.

Databases, Factual↗

Quantitative collagen as a golden standard in differential diagnosing of fibrotic changes in liver tissue.

Determining a presence and degree of liver fibrosis provides means for diagnosing disease related processes. We have used two data mining methods, discriminant and regression analyses, to acquire knowledge from the data of 211 patients. We have shown and discussed that quantitative collagen has a distinguished discriminating power and can serve as a golden standard. We have additionally succeeded to obtain a formula consisting of standardised blood tests that can replace quantitative collagen. Practical implications of this is a non-invasive and cost efficient patient examination. All the results are now left for clinical evaluation and so is the current way of histopathological classifications.

Biopsy↗

Liver guide for monitoring of chronic hepatitis C.

The severity of chronic hepatitis C infection in the individual patient is monitored using blood laboratory findings and liver biopsy. If blood test results could be shown to provide sufficient information concerning the disease, the invasive procedure of liver biopsy could perhaps be avoided in some instances. This study assessed the clinical relevance of blood laboratory tests for detecting disease-related changes in the liver. Histopathological classification was used to assign class membership of the patients and data mining operations were performed in an elaborate way on 19 different data sets. Disease activity could be detected by a small set of blood tests. Extended sets could identify more severe changes, but failed to distinguish them. The extracted rules are implemented as a part of the knowledge base of a corresponding decision support system aimed at specialists and general practitioners.

Analysis of Variance↗

Using data warehousing and OLAP in public health care.

The paper describes the possibilities of using data warehousing and OLAP technologies in public health care in general and then our own experience with these technologies gained during the implementation of a data warehouse of outpatient data at the national level. Such a data warehouse serves as a basis for advanced decision support systems based on statistical, OLAP or data mining methods. We used OLAP to enable interactive exploration and analysis of the data. We found out that data warehousing and OLAP are suitable for the domain of public health and that they enable new analytical possibilities in addition to the traditional statistical approaches.

Ambulatory Care↗

Applications of qualitative multi-attribute decision models in health care.

Hierarchical decision models are a general decision support methodology aimed at the classification or evaluation of options that occur in decision-making processes. They are also important for the analysis, simulation and explanation of options. Decision models are typically developed through the decomposition of complex decision problems into smaller and less complex subproblems; the result of such decomposition is a hierarchical structure that consists of attributes and utility functions. This article presents an approach to the development and application of qualitative hierarchical decision models that is based on DEX, an expert system shell for multi-attribute decision support. The distinguishing characteristics of DEX are the use of qualitative (symbolic) attributes, and 'if-then' decision rules. Also, DEX provides a number of methods for the analysis of models and options, such as selective explanation and what-if analysis. We demonstrate the applicability and flexibility of the approach presenting four real-life applications of DEX in health care: assessment of breast cancer risk, assessment of basic living activities in community nursing, risk assessment in diabetic foot care, and technical analysis of radiogram errors. In particular, we highlight and justify the importance of knowledge presentation and option analysis methods for practical decision-making. We further show that, using a recently developed data mining method called HINT, such hierarchical decision models can be discovered from retrospective patient data.

Breast Neoplasms↗

Systematic functional evaluation of CNGA1 missense variants associated with retinitis pigmentosa.

BACKGROUND: Missense variants are frequently classified as variants of uncertain significance (VUS) according to the guidelines of the American College of Medical Genetics and Genomics and the Association of Molecular Pathology (ACMG/AMP). Consequently, disease relevance remains elusive, impeding molecular genetic diagnostics, patients` and family genetic counseling, and identification of patients eligible for clinical trials. Functional studies are critical for resolving the clinical significance of VUS. CNGA1 encodes the main subunit of the rod cyclic nucleotide-gated (CNG) channel, a vital component of the phototransduction cascade. Variants in CNGA1 are a rare cause of autosomal recessive retinitis pigmentosa and a phase I/II gene augmentation trial (NCT06291935) is currently ongoing highlighting the necessity to differentiate benign from pathogenic variants. METHODS: CNGA1 missense variants compiled from retinal disease patient cohorts, public databases and literature were functionally investigated using a medium-throughput aequorin-based assay and in vitro minigene splice assays for predicted exonic spliceogenic variants. Functional data were correlated with the in silico prediction of five variant effect predictors (VEPs) and applied to support or revise variants' ACMG/AMP classification. RESULTS: Data mining revealed 86 missense CNGA1 variants - including three novel - most of them lacking functional data; 65.1% of the variants were initially classified as VUS. The aequorin-based assay showed that 72.1% of tested variants significantly impaired CNG channel function and were classified as functionally abnormal, while 23.3% were functionally normal and 5% remained functionally uncertain. Correlation of the functional data with in silico predictions identified AlphaMissense and CPT-1 to be the most suitable tools for assessing CNGA1 missense variants. Using in vitro minigene splice assays, two putative missense variants were shown to induce missplicing. Based on the functional findings, 62.1% of the variants initially classified as VUS were re-categorized as likely pathogenic or likely benign. Furthermore, 93.3% of the variants initially classified as likely pathogenic showed an effect on CNGA1 channel function, confirming their disease relevance and supporting their reclassification as pathogenic. CONCLUSION: This study represents the first comprehensive functional assessment of disease-associated CNGA1 missense variants, thus significantly advancing the understanding of their disease relevance and improving molecular genetic diagnostics in patients.

Humans↗

Information Management System for Site Remediation Efforts.

/ Environmental regulatory agencies are responsible for protecting human health and the environment in their constituencies. Their responsibilities include the identification, evaluation, and cleanup of contaminated sites. Leaking underground storage tanks (USTs) constitute a major source of subsurface and groundwater contamination. A significant portion of a regulatory body's efforts may be directed toward the management of UST-contaminated sites. In order to manage remedial sites effectively, vast quantities of information must be maintained, including analytical dataon chemical contaminants, remedial design features, and performance details. Currently, most regulatory agencies maintain such information manually. This makes it difficult to manage the data effectively. Some agencies have introduced automated record-keeping systems. However, the ad hoc approach in these endeavors makes it difficult to efficiently analyze, disseminate, and utilize the data. This paper identifies the information requirements for UST-contaminated site management at the Waste Cleanup Section of the Department of Environmental Resources Management in Dade County, Florida. It presents a viable design for an information management system to meet these requirements. The proposed solution is based on a back-end relational database management system with relevant tools for sophisticated data analysis and data mining. The database is designed with all tables in the third normal form to ensure data integrity, flexible access, and efficient query processing. In addition to all standard reports required by the agency, the system provides answers to ad hoc queries that are typically difficult to answer under the existing system. The database also serves as a repository of information for a decision support system to aid engineering design and risk analysis. The system may be integrated with a geographic information system for effective presentation and dissemination of spatial data.

Journal Article↗

Extended SQL for manipulating clinical warehouse data.

Health care institutions are beginning to collect large amounts of clinical data through patient care applications. Clinical data warehouses make these data available for complex analysis across patient records, benefiting administrative reporting, patient care and clinical research. Data gathered for patient care purposes are difficult to manipulate for analytic tasks; the schema presents conceptual difficulties for the analyst, and many queries perform poorly. An extension to SQL is presented that enables the analyst to designate groups of rows. These groups can then be manipulated and aggregated in various ways to solve a number of useful analytic problems. The extended SQL is concise and runs in linear time, while standard SQL requires multiple statements with polynomial performance. The extensions are extremely powerful for performing aggregations on large amounts of data, which is useful in clinical data mining applications.

Clinical Laboratory Techniques↗

Metabolite profiling for plant functional genomics.

Multiparallel analyses of mRNA and proteins are central to today's functional genomics initiatives. We describe here the use of metabolite profiling as a new tool for a comparative display of gene function. It has the potential not only to provide deeper insight into complex regulatory processes but also to determine phenotype directly. Using gas chromatography/mass spectrometry (GC/MS), we automatically quantified 326 distinct compounds from Arabidopsis thaliana leaf extracts. It was possible to assign a chemical structure to approximately half of these compounds. Comparison of four Arabidopsis genotypes (two homozygous ecotypes and a mutant of each ecotype) showed that each genotype possesses a distinct metabolic profile. Data mining tools such as principal component analysis enabled the assignment of "metabolic phenotypes" using these large data sets. The metabolic phenotypes of the two ecotypes were more divergent than were the metabolic phenotypes of the single-loci mutant and their parental ecotypes. These results demonstrate the use of metabolite profiling as a tool to significantly extend and enhance the power of existing functional genomics approaches.

Arabidopsis↗

Building manageable rough set classifiers.

An interesting aspect of techniques for data mining and knowledge discovery is their potential for generating hypotheses by discovering underlying relationships buried in the data. However, the set of possible hypotheses is often very large and the extracted models may become prohibitively complex. It is therefore typically desirable to only consider the "strongest" hypotheses, so that smaller models can be obtained that also retain good classificatory capabilities. This paper outlines how rule-based classifiers based on rough set theory and Boolean reasoning that are both small and perform well can be developed. Applied to a real-world medical dataset, the final models are shown to exhibit good performance using only a subset of the available information. Furthermore, the number of resulting rules is low and enables practical a posteriori inspection and interpretation of the models.

Classification↗

Application of a novel and fast information-theoretic method to the discovery of higher-order correlations in protein databases.

We present a fast, discrete data-mining approach to the problem of finding kappa-tuples of correlated amino acid residues in protein sequence data. When sets of sequence-distant sites display high mutual information, they may bespeak important structural or functional features. Our novel methodology overcomes the limitations of previous methods which examined only single-residue features or pairwise interactions.

AIDS Vaccines↗

SPECT electronic collimation resolution enhancement using chi-square minimization.

An electronic collimation technique is developed which utilizes the chi-square goodness-of-fit measure to filter scattered gammas incident upon a medical imaging detector. In this data mining technique, Compton kinematic expressions are used as the chi-square fitting templates for measured energy-deposition data involving multiple-interaction scatter sequences. Fit optimization is conducted using the Davidon variable metric minimization algorithm to simultaneously determine the best-fit gamma scatter angles and their associated uncertainties, with the uncertainty associated with the first scatter angle corresponding to the angular resolution precision for the source. The methodology requires no knowledge of materials and geometry. This pattern recognition application enhances the ability to select those gammas that will provide the best resolution for input to reconstruction software. Illustrative computational results are presented for a conceptual truncated-ellipsoid polystyrene position-sensitive fibre head-detector Monte Carlo model using a triple Compton scatter gamma sequence assessment for a 99mTc point source. A filtration rate of 94.3% is obtained, resulting in an estimated sensitivity approximately three orders of magnitude greater than a high-resolution mechanically collimated device. The technique improves the nominal single-scatter angular resolution by up to approximately 24 per cent as compared with the conventional analytic electronic collimation measure.

Algorithms↗

Rough sets: a knowledge discovery technique for multifactorial medical outcomes.

Rough sets is a fairly new and promising technique for data mining and knowledge discovery from databases. This tutorial article presents the fundamentals of rough set theory in a nontechnical manner and outlines how the technique can be used to extract minimal if-then rules from tables of empirical data that either fully or approximately describe given example classifications. An example application for prediction of ambulation for patients with spinal cord injury is given. Because such rules are readily interpretable, they can be inspected to yield possible new insight into how various contributing factors interact and, thus, serve as hypothesis generators for further research. Additionally, the set of mined rules may function as a classifier of new, unseen cases.

Data Interpretation, Statistical↗

A Bayesian neural network method for adverse drug reaction signal generation.

OBJECTIVE: The database of adverse drug reactions (ADRs) held by the Uppsala Monitoring Centre on behalf of the 47 countries of the World Health Organization (WHO) Collaborating Programme for International Drug Monitoring contains nearly two million reports. It is the largest database of this sort in the world, and about 35,000 new reports are added quarterly. The task of trying to find new drug-ADR signals has been carried out by an expert panel, but with such a large volume of material the task is daunting. We have developed a flexible, automated procedure to find new signals with known probability difference from the background data. METHOD: Data mining, using various computational approaches, has been applied in a variety of disciplines. A Bayesian confidence propagation neural network (BCPNN) has been developed which can manage large data sets, is robust in handling incomplete data, and may be used with complex variables. Using information theory, such a tool is ideal for finding drug-ADR combinations with other variables, which are highly associated compared to the generality of the stored data, or a section of the stored data. The method is transparent for easy checking and flexible for different kinds of search. RESULTS: Using the BCPNN, some time scan examples are given which show the power of the technique to find signals early (captopril-coughing) and to avoid false positives where a common drug and ADRs occur in the database (digoxin-acne; digoxin-rash). A routine application of the BCPNN to a quarterly update is also tested, showing that 1004 suspected drug-ADR combinations reached the 97.5% confidence level of difference from the generality. Of these, 307 were potentially serious ADRs, and of these 53 related to new drugs. Twelve of the latter were not recorded in the CD editions of The physician's Desk Reference or Martindale's Extra Pharmacopoea and did not appear in Reactions Weekly online. CONCLUSION: The results indicate that the BCPNN can be used in the detection of significant signals from the data set of the WHO Programme on International Drug Monitoring. The BCPNN will be an extremely useful adjunct to the expert assessment of very large numbers of spontaneously reported ADRs.

Adverse Drug Reaction Reporting Systems↗

Future of benchmarking: more data, more sharing, and better patient care.

Automated systems that provide whatever regulatory information is needed when it is needed; sharing of data to improve quality; data mined for specific groups of patients: Those are just a few of the trends predicted by health care experts asked to comment on the future of benchmarking and data strategies. Such improvements are needed; many hospitals continually run into problems when it comes to finding the right data sets for targeted patient groups.

Benchmarking↗

Logan: Planetary-Scale Genome Assembly Surveys Life's Diversity.

The breadth of life's diversity is unfathomable, but public nucleic acid sequencing data offers a window into the dispersion and evolution of genetic diversity across Earth. However the rapid growth and accumulation of sequence data have outpaced efficient analysis capabilities. The largest collection of freely available sequencing data is the Sequence Read Archive (SRA), comprising 27.3 million datasets or 5 × 1016 basepairs. To realize the potential of the SRA, we constructed Logan, a massive sequence assembly transforming short reads into long contigs and compressing the data over 100-fold, enabling highly efficient petabase-scale analysis. We created Logan-Search, a k-mer index of Logan for free planetary-scale sequence search, returning matches in minutes. We used Logan contigs to identify >200 million plastic-degrading enzyme homologs, and validate novel enzymes with catalytic activities exceeding current reference standards. Further, we vastly expand the known diversity of proteins (30-fold over UniRef50), plasmids (22-fold over PLSDB), P4 satellites (4.5-fold), and the recently described Obelisk RNA elements (3.7-fold). Logan also enables ecological and biomedical data mining, such as global tracking of antimicrobial resistance genes and the characterization of viral reactivation across millions of human BioSamples. By transforming the SRA, Logan democratizes access to the world's public genetic data and opens frontiers in biotechnology, molecular ecology, and global health.

Journal Article↗

Molecular anatomy of an intracranial aneurysm: coordinated expression of genes involved in wound healing and tissue remodeling.

BACKGROUND AND PURPOSE: Approximately 6% of human beings harbor an unruptured intracranial aneurysm. Each year in the United States, >30 000 people suffer a ruptured intracranial aneurysm, resulting in subarachnoid hemorrhage. Despite the high incidence and catastrophic consequences of a ruptured intracranial aneurysm and the fact that there is considerable evidence that predisposition to intracranial aneurysm has a strong genetic component, very little is understood with regard to the pathology and pathogenesis of this disease. METHODS: To begin characterizing the molecular pathology of intracranial aneurysm, we used a global gene expression analysis approach (SAGE-Lite) in combination with a novel data-mining approach to perform a high-resolution transcript analysis of a single intracranial aneurysm, obtained from a 3-year-old girl. RESULTS: SAGE-Lite provides a detailed molecular snapshot of a single intracranial aneurysm. These data suggest that, at least in this specific case, aneurysmal dilation results in a highly dynamic cellular environment in which extensive wound healing and tissue/extracellular matrix remodeling are taking place. Specifically, we observed significant overexpression of genes encoding extracellular matrix components (eg, COL3A1, COL1A1, COL1A2, COL6A1, COL6A2, elastin) and genes involved in extracellular matrix turnover (TIMP-3, OSF-2), cell adhesion and antiadhesion (SPARC, hevin), cytokinesis (PNUTL2), and cell migration (tetraspanin-5). CONCLUSIONS: Although these are preliminary data, representing analysis of only one individual, we present a unique first insight into the molecular basis of aneurysmal disease and define numerous candidate markers for future biochemical, physiological, and genetic studies of intracranial aneurysm. Products of these genes will be the focus of future studies in wider sample sets.

Calcium-Binding Proteins↗