PubMed Health⌕ Search

Biomedical subjects

Cornelius G Hunter

Publications and source records attributed to Cornelius G Hunter.

3 recordsLinked to original sources

Protein fragment clustering and canonical local shapes.

A novel clustering method is used to cluster protein fragments by shape. The centroids (mean fragments from each cluster) form a basis set of structural motifs. A database of 156,643 seven-residue fragments is used, and eight different basis sets with varying levels of resolution are generated. Coarse basis sets contain tens of centroids and provide meaningful local shapes, which are more detailed than the traditional secondary structure categories. High-resolution basis sets contain thousands of centroids and can be used to model tertiary structure of longer segments. The basis sets generated fit nontraining set proteins with the expected accuracy.

Animals↗

Protein local structure prediction from sequence.

A basis set of protein canonical fragments, or centroids, represents the range of local structure found in globular proteins. We develop a methodology to predict centroids from the amino acid sequence. The predictor gives the probability of each centroid in the basis set, at each loci along the backbone. The predictor selects the best-fit centroid at about 40% of the loci. The predicted probabilities are accurate and can be used to judge the confidence of each centroid prediction. For example, when filtering out centroids with <0.50 probability, the predictor is 65% accurate, although such high-probability centroids occur at only 28% of the loci. Centroids with high probability can be interpreted as segments that are highly influenced by the amino acid sequence, whereas centroids with low probability can be interpreted as segments that are more likely influenced by tertiary contacts. Low-resolution, starting point structures, can be generated by fitting the predicted centroids together.

Animals↗

Natural coordinate representation for the protein backbone structure.

A new model for describing the geometry of the C(alpha) backbone atoms in protein molecules is derived. This model uses one continuous variable per amino acid. This is half the number of degrees-of-freedom used in traditional backbone models. The new model was tested on 721 PDB structures and its average accuracy was determined to be 1.14 A cRMSD. This model can be used as a description of local structure that provides higher resolution than the traditional secondary structure categories. Also, because this structure description is one-dimensional, it can be used to align structures with the same efficiency and convergence properties available in the popular sequence alignment tools. Furthermore, the 1:1 correspondence with the amino acid sequence has implications for combined sequence/structure alignment. Conventional secondary structure prediction was used to further reduce the number of degrees-of-freedom in 16 test proteins. In those cases, the average cRMSD degraded from 0.96 to 2.33 A while the number of degrees-of-freedom improved (reduced) by more than 30%.

Amino Acids↗