PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data mining”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Bayesian data mining of protein domains gives an efficient predictive algorithm and new insight.

Identification of structural domains in uncharacterized protein sequences is important in the prediction of protein tertiary folds and functional sites, and hence in designing biologically active molecules. We present a new predictive computational method of classifying a protein into single, two continuous or two discontinuous domains using Bayesian Data Mining. The algorithm requires only the primary sequence and computer-predicted secondary structure. It incorporates correlation patterns between certain 3-dimensional motifs and some local helical folds found conserved in the vicinity of protein domains with high statistical confidence. The prediction of domain-class by this computationally simple and fast method shows good accuracy of prediction-average accuracies 83.3% for single domain, 60% for two continuous and 65.7% for two discontinuous domain proteins. Experiments on the large validation sample show its performance to be significantly better than that of DGS and DomSSEA. Computations of Bayesian probabilities show important features in terms of correlation of certain conserved patterns of secondary folds and tertiary motifs and give new insight. Applications for improved accuracy of predicting domain boundary points relevant to protein structural and functional modeling are also highlighted.

Algorithms↗

In silico identification and expression of SLC30 family genes: an expressed sequence tag data mining strategy for the characterization of zinc transporters' tissue expression.

BACKGROUND: Intracellular zinc concentration and localization are strictly regulated by two main protein components, metallothioneins and membrane transporters. In mammalian cells, two membrane transporters family are involved in intracellular zinc homeostasis: the uptake transporters called SLC39 or Zip family and the efflux transporters called SLC30 or ZnT family. ZnT proteins are members of the cation diffusion facilitator (CDF) family of metal ion transporters. RESULTS: From genomic databanks analysis, we identified the full-length sequences of two novel SLC30 genes, SLC30A8 and SLC30A10, extending the SLC30 family to ten members. We used an expressed sequence tag (EST) data mining strategy to determine the pattern of ZnT genes expression in tissues. In silico results obtained for already studied ZnT sequences were compared to experimental data, previously published. We determined an overall good correlation with expression pattern obtained by RT-PCR or immunomethods, particularly for highly tissue specific genes. CONCLUSION: The method presented herein provides a useful tool to complete gene families from sequencing programs and to produce preliminary expression data to select the proper biological samples for laboratory experimentation.

Amino Acid Sequence↗

Data mining of tree-based models to analyze freeway accident frequency.

INTRODUCTION: Statistical models, such as Poisson or negative binomial regression models, have been employed to analyze vehicle accident frequency for many years. However, these models have their own model assumptions and pre-defined underlying relationship between dependent and independent variables. If these assumptions are violated, the model could lead to erroneous estimation of accident likelihood. Classification and Regression Tree (CART), one of the most widely applied data mining techniques, has been commonly employed in business administration, industry, and engineering. CART does not require any pre-defined underlying relationship between target (dependent) variable and predictors (independent variables) and has been shown to be a powerful tool, particularly for dealing with prediction and classification problems. METHOD: This study collected the 2001-2002 accident data of National Freeway 1 in Taiwan. A CART model and a negative binomial regression model were developed to establish the empirical relationship between traffic accidents and highway geometric variables, traffic characteristics, and environmental factors. RESULTS: The CART findings indicated that the average daily traffic volume and precipitation variables were the key determinants for freeway accident frequencies. By comparing the prediction performance between the CART and the negative binomial regression models, this study demonstrates that CART is a good alternative method for analyzing freeway accident frequencies. IMPACT ON INDUSTRY: By comparing the prediction performance between the CART and the negative binomial regression models, this study demonstrates that CART is a good alternative method for analyzing freeway accident frequencies.

Accidents, Traffic↗

Data mining to support simulation modeling of patient flow in hospitals.

Spiraling health care costs in the United States are driving institutions to continually address the challenge of optimizing the use of scarce resources. One of the first steps towards optimizing resources is to utilize capacity effectively. For hospital capacity planning problems such as allocation of inpatient beds, computer simulation is often the method of choice. One of the more difficult aspects of using simulation models for such studies is the creation of a manageable set of patient types to include in the model. The objective of this paper is to demonstrate the potential of using data mining techniques, specifically clustering techniques such as K-means, to help guide the development of patient type definitions for purposes of building computer simulation or analytical models of patient flow in hospitals. Using data from a hospital in the Midwest this study brings forth several important issues that researchers need to address when applying clustering techniques in general and specifically to hospital data.

Algorithms↗

MtDB: a database for personalized data mining of the model legume Medicago truncatula transcriptome.

In order to identify the genes and gene functions that underlie key aspects of legume biology, researchers have selected the cool season legume Medicago truncatula (Mt) as a model system for legume research. A set of >170 000 Mt ESTs has been assembled based on in-depth sampling from various developmental stages and pathogen-challenged tissues. MtDB is a relational database that integrates Mt transcriptome data and provides a wide range of user-defined data mining options. The database is interrogated through a series of interfaces with 58 options grouped into two filters. In addition, the user can select and compare unigene sets generated by different assemblers: Phrap, Cap3 and Cap4. Sequence identifiers from all public Mt sites (e.g. IDs from GenBank, CCGB, TIGR, NCGR, INRA) are fully cross-referenced to facilitate comparisons between different sites, and hypertext links to the appropriate database records are provided for all queries' results. MtDB's goal is to provide researchers with the means to quickly and independently identify sequences that match specific research interests based on user-defined criteria. The underlying database and query software have been designed for ease of updates and portability to other model organisms. Public access to the database is at http://www.medicago.org/MtDB.

Chromosome Mapping↗

The role of data mining in pharmacovigilance.

A principle concern of pharmacovigilance is the timely detection of adverse drug reactions that are novel by virtue of their clinical nature, severity and/or frequency. The cornerstone of this process is the scientific acumen of the pharmacovigilance domain expert. There is understandably an interest in developing database screening tools to assist human reviewers in identifying associations worthy of further investigation (i.e., signals) embedded within a database consisting largely of background 'noise' containing reports of no substantial public health significance. Data mining algorithms are, therefore, being developed, tested and/or used by health authorities, pharmaceutical companies and academic researchers. After a focused review of postapproval drug safety signal detection, the authors explain how the currently used algorithms work and address key questions related to their validation, comparative performance, deployment in naturalistic pharmacovigilance settings, limitations and potential for misuse. Suggestions for further research and development are offered.

Adverse Drug Reaction Reporting Systems↗

Data mining of NCI's anticancer screening database reveals mitochondrial complex I inhibitors cytotoxic to leukemia cell lines.

Mitochondria are principal mediators of apoptosis and thus can be considered molecular targets for new chemotherapeutic agents in the treatment of cancer. Inhibitors of mitochondrial complex I of the electron transport chain have been shown to induce apoptosis and exhibit antitumor activity. In an effort to find novel complex I inhibitors which exhibited anticancer activity in the NCI's tumor cell line screen, we examined organized tumor cytotoxicity screening data available as SOM (self-organized maps) (http://www.spheroid.ncifcrf.gov) at the developmental therapeutics program (DTP) of the National Cancer Institute (NCI). Our analysis focused on an SOM cluster comprised of compounds which included a number of known mitochondrial complex I (NADH:CoQ oxidoreductase) inhibitors. From these clusters 10 compounds whose mechanism of action was unknown were tested for inhibition of complex I activity in bovine heart sub-mitochondrial particles (SMP) resulting in the discovery that 5 of the 10 compounds demonstrated significant inhibition with IC50's in the nM range for three of the five. Examination of screening profiles of the five inhibitors toward the NCI's tumor cell lines revealed that they were cytotoxic to the leukemia subpanel (particularly K562 cells). Oxygen consumption experiments with permeabilized K562 cells revealed that the five most active compounds inhibited complex I activity in these cells in the same rank order and similar potency as determined with bovine heart SMP. Our findings thus fortify the appeal of mitochondrial complex I as a possible anticancer molecular target and provide a data mining strategy for selecting candidate inhibitors for further testing.

Animals↗

Databases and data mining for computational vaccinology.

Drugs and vaccines are keys to the effective fight against disease. While the pharmaceutical industry has developed an awesome array of real and virtual approaches to rational drug discovery, the complexity of the immune system hampers attempts to design and develop vaccines in a rational manner. The goal of immunoinformatics (the application of informatics techniques to immunological macromolecules), an emergent sub-discipline of bioinformatics, is to develop computational vaccinology as a potent tool in the quest for new vaccines. Databases and data mining, the two principal weapons at the disposal of the in silico vaccinologist, will be presented in the light of current developments.

Computational Biology↗

Decision algorithm based on data mining for coagulant type and dosage in water treatment systems.

Water shortages are gradually accelerating because higher standards of living are required and water resources are more heavily utilised. Therefore, effective water treatment is necessary in order to retain the required quality and amount of water. General treatment includes coagulation, flocculation, filtering and disinfection. Coagulation, flocculation and disinfection are major components of water treatment processes. In this paper, a new automatic decision algorithm is proposed for coagulation. The proposed method shows how to determine the coagulant type and amount using data mining techniques.

Algorithms↗

Combining data mining tools with health care models for improved understanding of health processes and resource utilisation.

Variability and uncertainty are inherent characteristics of most health care processes. Patient pathways and dwelling times even within the same process typically vary from patient to patient, such as the flow of patients through a particular health care provider or patient progression through the natural history of a given disease. The challenge for the OR modeller is to adequately handle and capture the stochastic features within developed models. This paper will discuss the benefits of combining patient classification tools (data mining techniques) with developed OR models, such as simulation tools, to more accurately capture patient outcomes, risks and resource needs. Illustrative applications will demonstrate the approach.

Decision Support Systems, Management↗

Data mining crystallization databases: knowledge-based approaches to optimize protein crystal screens.

Protein crystallization is a major bottleneck in protein X-ray crystallography, the workhorse of most structural proteomics projects. Because the principles that govern protein crystallization are too poorly understood to allow them to be used in a strongly predictive sense, the most common crystallization strategy entails screening a wide variety of solution conditions to identify the small subset that will support crystal nucleation and growth. We tested the hypothesis that more efficient crystallization strategies could be formulated by extracting useful patterns and correlations from the large data sets of crystallization trials created in structural proteomics projects. A database of crystallization conditions was constructed for 755 different proteins purified and crystallized under uniform conditions. Forty-five percent of the proteins formed crystals. Data mining identified the conditions that crystallize the most proteins, revealed that many conditions are highly correlated in their behavior, and showed that the crystallization success rate is markedly dependent on the organism from which proteins derive. Of the proteins that crystallized in a 48-condition experiment, 60% could be crystallized in as few as 6 conditions and 94% in 24 conditions. Consideration of the full range of information coming from crystal screening trials allows one to design screens that are maximally productive while consuming minimal resources, and also suggests further useful conditions for extending existing screens.

Archaeal Proteins↗

Data mining and multiparameter analysis of lung surfactant protein genes in bronchopulmonary dysplasia.

Bronchopulmonary dysplasia (BPD), the most common chronic lung disease in infancy, is influenced by a number of antenatal and postnatal risk factors and is mostly preceded by respiratory distress syndrome (RDS) in the newborn. Surfactant protein (SP-A, -B, -C and -D) gene variations may play a role in both BPD and RDS. An association study between these candidate genes and BPD was performed. A total of 365 preterm Finnish infants in a high-risk population with gestational age <or=32 weeks were genotyped for all SP genes. A multiparameter analysis was performed using Agrawal's algorithm based data mining and conventional methods of statistical allelic association. In singletons and presenting multiples, the frequency of SP-B intron 4 deletion variant allele was increased in BPD versus controls (P=0.008, OR=2.0, 95%CI 1.2-3.4). The presence of the SP-B intron 4 deletion variant was a risk factor for BPD even when essential external confounding factors were included in the analyses. No other SP polymorphisms associated with BPD, and the SP-B intron 4 variation did not associate with RDS. Transcription Element Search Software predicted allele-specific differences at several putative transcription factor binding sites that may be important in SP-B regulation. The present multiparameter analysis demonstrates the presumable direct involvement of the SP-B intron 4 deletion variant allele as a genetic risk factor to BPD. We propose that two separate SP-B gene polymorphisms have a phenotypic significance via separate molecular mechanisms: the intron 4 length variation affecting transcriptional regulation, and the exonic Ile131Thr variation affecting post-translationally.

Bronchopulmonary Dysplasia↗

Fullerene data mining using bibliometrics and database tomography

Database tomography (DT) is a textual database analysis system consisting of two major components: (1) algorithms for extracting multiword phrase frequencies and phrase proximities (physical closeness of the multiword technical phrases) from any type of large textual database, to augment (2) interpretative capabilities of the expert human analyst. DT was used to derive technical intelligence from a fullerenes database derived from the Science Citation Index and the Engineering Compendex. Phrase frequency analysis by the technical domain experts provided the pervasive technical themes of the fullerenes database, and phrase proximity analysis provided the relationships among the pervasive technical themes. Bibliometric analysis of the fullerenes literature supplemented the DT results with author/journal/institution publication and citation data. Comparisons of fullerenes results with past analyses of similarly structured near-earth space, chemistry, hypersonic/supersonic flow, aircraft, and ship hydrodynamics databases are made. One important finding is that many of the normalized bibliometric distribution functions are extremely consistent across these diverse technical domains and could reasonably be expected to apply to broader chemical topics than fullerenes that span multiple structural classes. Finally, lessons learned about integrating the technical domain experts with the data mining tools are presented.

Journal Article↗

Automated gamma knife radiosurgery treatment planning with image registration, data-mining, and Nelder-Mead simplex optimization.

Gamma knife treatments are usually planned manually, requiring much expertise and time. We describe a new, fully automatic method of treatment planning. The treatment volume to be planned is first compared with a database of past treatments to find volumes closely matching in size and shape. The treatment parameters of the closest matches are used as starting points for the new treatment plan. Further optimization is performed with the Nelder-Mead simplex method: the coordinates and weight of the isocenters are allowed to vary until a maximally conformal plan specific to the new treatment volume is found. The method was tested on a randomly selected set of 10 acoustic neuromas and 10 meningiomas. Typically, matching a new volume took under 30 seconds. The time for simplex optimization, on a 3 GHz Xeon processor, ranged from under a minute for small volumes (<1000 cubic mm, 2-3 isocenters), to several tens of hours for large volumes (>30,000 cubic mm, >20 isocenters). In 8/10 acoustic neuromas and 8/10 meningiomas, the automatic method found plans with conformation number equal or better than that of the manual plan. In 4/10 acoustic neuromas and 5/10 meningiomas, both overtreatment and undertreatment ratios were equal or better in automated plans. In conclusion, data-mining of past treatments can be used to derive starting parameters for treatment planning. These parameters can then be computer optimized to give good plans automatically.

Automation↗

RNA expression profiles and data mining of sugarcane response to low temperature.

Tropical and subtropical plants are generally sensitive to cold and can show appreciable variation in their response to cold stress when exposed to low positive temperatures. Using nylon filter arrays, we analyzed the expression profile of 1,536 expressed sequence tags (ESTs) of sugarcane (Saccharum sp. cv SP80-3280) exposed to cold for 3 to 48 h. Thirty-four cold-inducible ESTs were identified, of which 20 were novel cold-responsive genes that had not previously been reported as being cold inducible, including cellulose synthase, ABI3-interacting protein 2, a negative transcription regulator, phosphate transporter, and others, as well as several unknown genes. In addition, 25 ESTs were identified as being down-regulated during cold exposure. Using a database of cold-regulated proteins reported for other plants, we searched for homologs in the sugarcane EST project database (SUCEST), which contains 263,000 ESTs. Thirty-three homologous putative cold-regulated proteins were identified in the SUCEST database. On the basis of the expression profiles of the cold-inducible genes and the data-mining results, we propose a molecular model for the sugarcane response to low temperature.

Cold Temperature↗

Data mining for regulatory elements in yeast genome.

We have examined methods and developed a general software tool for finding and analyzing combinations of transcription factor binding sites that occur relatively often in gene upstream regions (putative promoter regions) in the yeast genome. Such frequently occurring combinations may be essential parts of possible promoter classes. The regions upstream to all genes were first isolated from the yeast genome database MIPS using the information in the annotation files of the database. The ones that do not overlap with coding regions were chosen for further studies. Next, all occurrences of the yeast transcription factor binding sites, as given in the IMD database, were located in the genome and in the selected regions in particular. Finally, by using a general purpose data mining software in combination with our own software, which parametrizes the search, we can find the combinations of binding sites that occur in the upstream regions more frequently than would be expected on the basis of the frequency of individual sites. The procedure also finds so-called association rules present in such combinations. The developed tool is available for use through the WWW.

Binding Sites↗

Predicting post-synaptic activity in proteins with data mining.

The bioinformatics problem being addressed in this paper is to predict whether or not a protein has post-synaptic activity. This problem is of great intrinsic interest because proteins with post-synaptic activities are connected with functioning of the nervous system. Indeed, many proteins having post-synaptic activity have been functionally characterized by biochemical, immunological and proteomic exercises. They represent a wide variety of proteins with functions in extracellular signal reception and propagation through intracellular apparatuses, cell adhesion molecules and scaffolding proteins that link them in a web. The challenge is to automatically discover features of the primary sequences of proteins that typically occur in proteins with post-synaptic activity but rarely (or never) occur in proteins without post-synaptic activity, and vice-versa. In this context, we used data mining to automatically discover classification rules that predict whether or not a protein has post-synaptic activity. The discovered rules were analysed with respect to their predictive accuracy (generalization ability) and with respect to their interestingness to biologists (in the sense of representing novel, unexpected knowledge).

Database Management Systems↗

Data-mining analyses of pharmacovigilance signals in relation to relevant comparison drugs.

OBJECTIVE: The aim of this paper is to demonstrate the usefulness of the Bayesian Confidence Propagation Neural Network (BCPNN) in the detection of drug-specific and drug-group effects in the database of adverse drug reactions of the World Health Organization Programme for International Drug Monitoring. METHODS: Examples of drug-adverse reaction combinations highlighted by the BCPNN as quantitative associations were selected. The anatomical therapeutic chemical (ATC) group to which the drug belonged was then identified, and the information component (IC) was calculated for this ATC group and the adverse drug reaction (ADR). The IC of the ATC group with the ADR was then compared with the IC of the drug-ADR by plotting the change in IC and its 95% confidence limit over time for both. RESULTS: The chosen examples show that the BCPNN data-mining approach can identify drug-specific as well as group effects. In the known examples that served as test cases, beta-blocking agents other than practolol are not associated with sclerosing peritonitis, but all angiotensin-converting enzyme inhibitors are associated with coughing, as are antihistamines with heart-rhythm disorders and antipsychotics with myocarditis. The recently identified association between antipsychotics and myocarditis remains even after consideration of concomitant medication. CONCLUSION: The BCPNN can be used to improve the ability of a signal detection system to highlight group and drug-specific effects.

Adverse Drug Reaction Reporting Systems↗