PubMed HealthSearch

PubMed · 8768768

Neural network studies. 2. Variable selection.

Abstract

Quantitative structure-activity relationship (QSAR) studies usually require an estimation of the relevance of a very large set of initial variables. Determination of the most important variables allows theoretically a better generalization by all pattern recognition methods. This study introduces and investigates five pruning algorithms designed to estimate the importance of input variables in feed-forward artificial neural network trained by back propagation algorithm (ANN) applications and to prune nonrelevant ones in a statistically reliable way. The analyzed algorithms performed similar variable estimations for simulated data sets, but differences were detected for real QSAR examples. Improvement of ANN prediction ability was shown after the pruning of redundant input variables. The statistical coefficients computed by ANNs for QSAR examples were better than those of multiple linear regression. Restrictions of the proposed algorithms and the potential use of ANNs are discussed.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

I V Tetko, A E Villa, D J Livingstone. Neural network studies. 2. Variable selection.. https://doi.org/10.1021/ci950204c

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Gene therapy and medical genetics on Internet.

In this report we consider the development of the Internet, from its origins as a military invention in the times of the cold war to its present day role, together with the World Wide Web, as a means of global communication which plays a key role in medical research and particularly in medical genetics. A few of the major genetics related research projects and gene research centers are introduced and their aims are briefly discussed. Detailed information about chromosome and gene mapping, together with sequence and structure databases, can be easily and rapidly accessed through the Internet. A variety of web-sites are briefly described and then listed at the end of the report, which will serve as a useful starting point from which the interested reader can access an almost endless source of genetics related information on the Internet. Finally, some of the ethical, legal and social implications of the links between gene therapy and the Intemet are considered.

Databases, Factual

A scoring scheme for discriminating between drugs and nondrugs.

A scoring scheme for the rapid and automatic classification of molecules into drugs and nondrugs was developed. The method is a valuable new tool that can aid in the selection and prioritization of compounds from large compound collections for purchase or biological testing and that can replace a considerable amount of laborious manual work by a more unbiased approach. It is based on the extraction of knowledge from large databases of drugs and nondrugs. The method was set up by using atom type descriptors for encoding the molecular structures and by training a feedforward neural network for classifying the molecules. It was parametrized and validated by using large databases of drugs and nondrugs (169 331 molecules from the Available Chemicals Directory, ACD, and 38 416 molecules from the World Drug Index, WDI). The method revealed features in the molecular descriptors that either qualify or disqualify a molecule for being a drug and classified 83% of the ACD and 77% of the WDI adequately.

Databases, Factual