PubMed Health⌕ Search

Biomedical subjects

Igor B Kuznetsov

Publications and source records attributed to Igor B Kuznetsov.

7 recordsLinked to original sources

Using evolutionary and structural information to predict DNA-binding sites on DNA-binding proteins.

Proteins that interact with DNA are involved in a number of fundamental biological activities such as DNA replication, transcription, and repair. A reliable identification of DNA-binding sites in DNA-binding proteins is important for functional annotation, site-directed mutagenesis, and modeling protein-DNA interactions. We apply Support Vector Machine (SVM), a supervised pattern recognition method, to predict DNA-binding sites in DNA-binding proteins using the following features: amino acid sequence, profile of evolutionary conservation of sequence positions, and low-resolution structural information. We use a rigorous statistical approach to study the performance of predictors that utilize different combinations of features and how this performance is affected by structural and sequence properties of proteins. Our results indicate that an SVM predictor based on a properly scaled profile of evolutionary conservation in the form of a position specific scoring matrix (PSSM) significantly outperforms a PSSM-based neural network predictor. The highest accuracy is achieved by SVM predictor that combines the profile of evolutionary conservation with low-resolution structural information. Our results also show that knowledge-based predictors of DNA-binding sites perform significantly better on proteins from mainly-alpha structural class and that the performance of these predictors is significantly correlated with certain structural and sequence properties of proteins. These observations suggest that it may be possible to assign a reliability index to the overall accuracy of the prediction of DNA-binding sites in any given protein using its sequence and structural properties. A web-server implementation of the predictors is freely available online at http://lcg.rit.albany.edu/dp-bind/.

Amino Acid Sequence↗

A novel sensitive method for the detection of user-defined compositional bias in biological sequences.

MOTIVATION: Most biological sequences contain compositionally biased segments in which one or more residue types are significantly overrepresented. The function and evolution of these segments are poorly understood. Usually, all types of compositionally biased segments are masked and ignored during sequence analysis. However, it has been shown for a number of proteins that biased segments that contain amino acids with similar chemical properties are involved in a variety of molecular functions and human diseases. A detailed large-scale analysis of the functional implications and evolutionary conservation of different compositionally biased segments requires a sensitive method capable of detecting user-specified types of compositional bias. RESULTS: We present BIAS, a novel sensitive method for the detection of compositionally biased segments composed of a user-specified set of residue types. BIAS uses the discrete scan statistics that provides a highly accurate correction for multiple tests to compute analytical estimates of the significance of each compositionally biased segment. The method can take into account global compositional bias when computing analytical estimates of the significance of local clusters. BIAS is benchmarked against SEG, SAPS and CAST programs. We also use BIAS to show that groups of proteins with the same biological function are significantly associated with particular types of compositionally biased segments.

Algorithms↗

Class-specific correlations between protein folding rate, structure-derived, and sequence-derived descriptors.

Small single-domain proteins that fold by simple two-state kinetics have been shown to exhibit a wide variation in their folding rates. It has been proposed that folding mechanisms in these proteins are largely determined by the native-state topology, and a significant correlation between folding rate and measures of the average topological complexity, such as relative contact order (RCO), has been reported. We perform a statistical analysis of folding rate and RCO in all three major structural classes (alpha, beta, and alpha/beta) of small two-state proteins and of RCO in groups of analogous and homologous small single-domain proteins with the same topology. We also study correlation between folding rate and the average physicochemical properties of amino acid sequences in two-state proteins. Our results indicate that 1) helical proteins have statistically distinguishable, class-specific folding rates; 2) RCO accounts for essentially all the variation of folding rate in helical proteins, but for only a part of the variation in beta-sheet-containing proteins; and 3) only a small fraction of the protein topologies studied show a topology-specific RCO. We also report a highly significant correlation between the folding rate and average intrinsic structural propensities of protein sequences. These results suggest that intrinsic structural propensities may be an important determinant of the rate of folding in small two-state proteins.

Databases, Protein↗

Comparative computational analysis of prion proteins reveals two fragments with unusual structural properties and a pattern of increase in hydrophobicity associated with disease-promoting mutations.

Prion diseases are a group of neurodegenerative disorders associated with conversion of a normal prion protein, PrPC, into a pathogenic conformation, PrPSc. The PrPSc is thought to promote the conversion of PrPC. The structure and stability of PrPC are well characterized, whereas little is known about the structure of PrPSc, what parts of PrPC undergo conformational transition, or how mutations facilitate this transition. We use a computational knowledge-based approach to analyze the intrinsic structural propensities of the C-terminal domain of PrP and gain insights into possible mechanisms of structural conversion. We compare the properties of PrP sequences to those of a PrP paralog, Doppel, and to the distributions of structural propensities observed in known protein structures from the Protein Data Bank. We show that the prion protein contains at least two sequence fragments with highly unusual intrinsic propensities, PrP(114-125) and helix B. No segments with unusual properties were found in Doppel protein, which is topologically identical to PrP but does not undergo structural rearrangements. Known disease-promoting PrP mutations form a statistically significant cluster in the region comprising helices B and C. Due to their unusual properties, PrP(114-125) and the C terminus of helix B may be considered as primary candidates for sites involved in conformational transition from PrPC to PrPSc. The results of our study also show that most PrP mutations associated with neurodegenerative disorders increase local hydrophobicity. We suggest that the observed increase in hydrophobicity may facilitate PrP-to-PrP or/and PrP-to-cofactor interactions, and thus promote structural conversion.

Amino Acid Sequence↗

Similarity between the C-terminal domain of the prion protein and chimpanzee cytomegalovirus glycoprotein UL9.

Prion diseases are a group of fatal neurodegenerative disorders associated with structural conversion of a normal, mostly alpha-helical cellular prion protein, PrP(C), into a pathogenic beta-sheet-rich conformation, PrP(Sc). The structure of PrP(C) is well studied, whereas the insolubility of PrP(Sc) makes the characterization of its structure problematic. No proteins similar to PrP, except for its paralog with the same fold, PrP-Doppel, are known. However, PrP-Doppel does not undergo a structural transition into a beta-sheet-rich conformation. Structural information from proteins that share a weak but significant sequence similarity with PrP may be used to gain additional insights into the conformation of PrP(Sc). We construct a sequence profile corresponding to the structured domain of PrP and use this profile to search the SWISS-PROT and TrEMBL databases. We identify a significant sequence similarity between PrP and chimpanzee cytomegalovirus glycoprotein UL9. This glycoprotein scores higher than all PrP-Doppel sequences. Fold recognition methods assign a mainly-beta fold to UL9. Owing to the observed sequence similarity with PrP and a putative mainly-beta fold, the UL9 glycoprotein may represent a potential target for experimental structure determination aimed at obtaining a structural template for PrP(Sc) modeling.

Amino Acid Sequence↗

On the properties and sequence context of structurally ambivalent fragments in proteins.

The goal of this work is to characterize structurally ambivalent fragments in proteins. We have searched the Protein Data Bank and identified all structurally ambivalent peptides (SAPs) of length five or greater that exist in two different backbone conformations. The SAPs were classified in five distinct categories based on their structure. We propose a novel index that provides a quantitative measure of conformational variability of a sequence fragment. It measures the context-dependent width of the distribution of (phi,xi) dihedral angles associated with each amino acid type. This index was used to analyze the local structural propensity of both SAPs and the sequence fragments contiguous to them. We also analyzed type-specific amino acid composition, solvent accessibility, and overall structural properties of SAPs and their sequence context. We show that each type of SAP has an unusual, type-specific amino acid composition and, as a result, simultaneous intrinsic preferences for two distinct types of backbone conformation. All types of SAPs have lower sequence complexity than average. Fragments that adopt helical conformation in one protein and sheet conformation in another have the lowest sequence complexity and are sampled from a relatively limited repertoire of possible residue combinations. A statistically significant difference between two distinct conformations of the same SAP is observed not only in the overall structural properties of proteins harboring the SAP but also in the properties of its flanking regions and in the pattern of solvent accessibility. These results have implications for protein design and structure prediction.

Amino Acid Sequence↗

Discriminative ability with respect to amino acid types: assessing the performance of knowledge-based potentials without threading.

We present a novel method designed to analyze the discriminative ability of knowledge-based potentials with respect to the 20 residue types. The method is based on the preference of amino acids for specific types of protein environment, and uses a virtual mutagenesis experiment to estimate how much information a given potential can provide about environments of each amino acid type. This allows one to test and optimize the performance of real potentials at the level of individual amino acids, using actual data on residue environments from a dataset of known protein structures. We have applied our method to long-range and medium-range pairwise distance-dependent potentials. The results of our study indicate that these potentials are only able to discriminate between a very limited number of residue types, and that discriminative ability is extremely sensitive to the choice of parameters used to construct the potentials, and even to the size of the training dataset. We also show that different types of pairwise distance potentials are dominated by different types of interactions. These dominant interactions strongly depend on the type of approximation used to define residue position. For each potential, our methodology is able to identify a potential-specific amino acid distance matrix and a reduced amino acid alphabet of any specified size, which may have implications for sequence alignment and multibody models.

Amino Acid Substitution↗