PubMed · 7584373
Protein structure prediction: selecting salient features from large candidate pools.
Abstract
We introduce a parallel approach, "DT-SELECT," for selecting features used by inductive learning algorithms to predict protein secondary structure. DT-SELECT is able to rapidly choose small, nonredundant feature sets from pools containing hundreds of thousands of potentially useful features. It does this by building a decision tree, using features from the pool, that classifies a set of training examples. The features included in the tree provide a compact description of the training data and are thus suitable for use as inputs to other inductive learning algorithms. Empirical experiments in the protein secondary-structure task, in which sets of complex features chosen by DT-SELECT are used to augment a standard artificial neural network representation, yield surprisingly little performance gain, even though features are selected from very large feature pools. We discuss some possible reasons for this result.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
K J Cherkauer, J W Shavlik. 1993. Protein structure prediction: selecting salient features from large candidate pools.. https://pubmed.ncbi.nlm.nih.gov/7584373/
Cite the original work for its findings. Save a collection to share your selection of sources.