PubMed Health⌕ Search

Biomedical subjects

Chin-Hsien Tai

Publications and source records attributed to Chin-Hsien Tai.

4 recordsLinked to original sources

ROC and confusion analysis of structure comparison methods identify the main causes of divergence from manual protein classification.

BACKGROUND: Current classification of protein folds are based, ultimately, on visual inspection of similarities. Previous attempts to use computerized structure comparison methods show only partial agreement with curated databases, but have failed to provide detailed statistical and structural analysis of the causes of these divergences. RESULTS: We construct a map of similarities/dissimilarities among manually defined protein folds, using a score cutoff value determined by means of the Receiver Operating Characteristics curve. It identifies folds which appear to overlap or to be "confused" with each other by two distinct similarity measures. It also identifies folds which appear inhomogeneous in that they contain apparently dissimilar domains, as measured by both similarity measures. At a low (1%) false positive rate, 25 to 38% of domain pairs in the same SCOP folds do not appear similar. Our results suggest either that some of these folds are defined using criteria other than purely structural consideration or that the similarity measures used do not recognize some relevant aspects of structural similarity in certain cases. Specifically, variations of the "common core" of some folds are severe enough to defeat attempts to automatically detect structural similarity and/or to lead to false detection of similarity between domains in distinct folds. Structures in some folds vary greatly in size because they contain varying numbers of a repeating unit, while similarity scores are quite sensitive to size differences. Structures in different folds may contain similar substructures, which produce false positives. Finally, the common core within a structure may be too small relative to the entire structure, to be recognized as the basis of similarity to another. CONCLUSION: A detailed analysis of the entire available protein fold space by two automated similarity methods reveals the extent and the nature of the divergence between the automatically determined similarity/dissimilarity and the manual fold type classifications. Some of the observed divergences can probably be addressed with better structure comparison methods and better automatic, intelligent classification procedures. Others may be intrinsic to the problem, suggesting a continuous rather than discrete protein fold space.

Algorithms↗

Domain definition and target classification for CASP6.

Assessment of structure predictions in CASP6 was based on single domains isolated from experimentally determined structures, which were categorized into comparative modeling, fold recognition, and new fold targets. Domain definitions were defined upon visual examination of the structures with the aid of automated domain-parsing programs. Domain categorization was determined by comparison of the target structures with those in the Protein Data Bank at the time each target expired and a variety of sequence and structure-based methods to determine potential homologous relationships.

Amino Acid Sequence↗

Assessment of CASP6 predictions for new and nearly new fold targets.

This is a report of the assessment of the predictions made for the CASP6 protein structure prediction experiment conducted in 2004 in the New Fold (NF) category. There were nine protein domains that were judged to have new folds (NF) and 16 for which a similar structure was known but the sequence similarity was judged to be too low for them to be easily recognized (FR/A). We selected all NF targets and eight of the 16 FR/A targets judged to be at the borderline between NF and FR/A for evaluation in the NF category. A total of 165 prediction groups submitted over 7400 structural models for these targets. The quality of these models was evaluated using the GDT_TS scores of the structural similarity detection program LGA and by visual inspection of the top-scoring models. The best models submitted bore an overall similarity to the target structure for three or four of the nine NF targets and for all but one of the FR/A targets. High-scoring models for the NF targets were submitted by several different groups. When both the NF and FR/A targets were considered, Baker group dominated by submitting best models for seven of the 17 targets, but 14 other groups also managed to submit best models for one or more targets.

Algorithms↗

Evaluation of domain prediction in CASP6.

We present an analysis of the domain boundary prediction, a new category, in the sixth community-wide experiment on the Critical Assessment of Techniques for Protein Structure Prediction (CASP6). There were 1011 predictions submitted for 63 targets. Each prediction was compared to the set of domains defined manually by visual inspection of the experimental structure. The comparison was scored using a new domain prediction scoring scheme. As the definition of a domain is subjective, many targets were assigned alternate definitions. For such targets, each prediction was compared with all different definitions and the best score was chosen. The predictors found it difficult to accurately predict domain boundaries when the target protein contained many domains or domains made of multiple sequence segments. The CBRC-DR (P0536) and Sternberg (P0237) groups were the most successful among human experts, while Baker-Rossettadom (P0353) and Baker-Robetta-Ginzu (P0421) did well among servers.

Algorithms↗