PubMed Health⌕ Search

Biomedical subjects

AK Dunker

Publications and source records attributed to AK Dunker.

4 recordsLinked to original sources

Predicting Protein Disorder for N-, C-, and Internal Regions.

Logistic regression (LR), discriminant analysis (DA), and neural networks (NN) were used to predict ordered and disordered regions in proteins. Training data were from a set of non-redundant X-ray crystal structures, with the data being partitioned into N-terminal, C-terminal and internal (I) regions. The DA and LR methods gave almost identical 5-cross validation accuracies that averaged to the following values: 75.9 +/- 3.1% (N-regions), 70.7 +/- 1.5% (I-regions), and 74.6 +/- 4.4% (C-regions). NN predictions gave slightly higher scores: 78.8 +/- 1.2% (N-regions), 72.5 +/- 1.2% (I-regions), and 75.3 +/- 3.3% (C-regions). Predictions improved with length of the disordered regions. Averaged over the three methods, values ranged from 52% to 78% for length = 9-14 to >/= 21, respectively, for I-regions, from 72% to 81% for length = 5 to 12-15, respectively, for N-regions, and from 70% to 80% for length = 5 to 12-15, respectively, for C-regions. These data support the hypothesis that disorder is encoded by the amino acid sequence.

Journal Article↗

Predicting Binding Regions within Disordered Proteins.

Disordered regions are sequences within proteins that fail to fold into a fixed tertiary structure and have been shown to be involved in a variety of biological functions. We recently applied neural network predictors of disorder developed from X-ray data to several protein sequences characterized as disordered by NMR (Garner, Cannon, Romero, Obradovic and Dunker, Genome Informatics, 9:201-213, 1998). A few predictions on the NMR-characterized disordered regions were noted to contain false negative indications of order that correlated with regions of function. These and additional examples are examined in more detail here. Overall, 8 of 9 functional segments in 5 disordered proteins were identified or partially identified by this approach. The functions of these regions appear to involve binding to DNA, RNA, and proteins. These regions are known to undergo disorder-to-order transitions upon binding. This apparent ability of the predictors to identify functional regions in disordered proteins could be due to the existence of different flavors, or sub-classes of disorder, originating from the sequence of the disordered regions and perhaps owing to local inclinations toward order. These different flavors may be a characteristic that could be used to identify binding regions within proteins that are difficult to characterize structurally.

Journal Article↗

The Sequence Attribute Method for Determining Relationships Between Sequence and Protein Disorder.

The conditional probability, P(s|x), is a statement of the probability that the event, s, will occur given prior knowledge for the value of x. If x is given and if s is randomly distributed, then an empirical approximation of the true conditional probability can be computed by the application of Bayes' Theorem. Here s represents one of two structural classes, either ordered, s (o), or disordered, s (d), and x represents an attribute value calculated over a window of 21 amino acids. Plots of P(s|x) versus x provide information about the correlation between the given sequence attribute and disorder or order. These conditional probability plots allow quantitative comparisons between individual attributes for their ability to discriminate between order and disorder states. Using such quantitative comparisons, 38 different sequence attributes have been rank-ordered. Attributes based on cysteine, the aromatics, flexible tendencies, and charge were found to be the best attributes for distinguishing order and disorder among those tested so far.

Journal Article↗

Predicting Disordered Regions from Amino Acid Sequence: Common Themes Despite Differing Structural Characterization.

Using ordered and disordered regions identified either by X-ray crystallography or by NMR spectroscopy, we trained neural networks to predict order and disorder from amino acid sequence. Although the NMR-based predictor initially appeared to be much better than the one based on the X-ray data, both predictors yielded similar overall accuracies when tested on each other's training sets, and indicated similar regions of disorder upon each sequence. The predictors trained with X-ray data showed similar results for a 5-cross validation experiment and for the out-of-sample predictions on the NMR characterized data. In contrast, the predictor trained with NMR data gave substantially worse accuracies on the out-of-sample X-ray data as compared to the accuracies displayed by the 5-cross validation during the network training. Overall, the results from the two predictors suggest that disordered regions comprise a sequence-dependant category distinct from that of ordered protein structure.

Journal Article↗