PubMed Health⌕ Search

Biomedical subjects

I I Vaisman

Publications and source records attributed to I I Vaisman.

6 recordsLinked to original sources

SECOST: sequence-conformation-structure database for amino acid residues in proteins.

The sequence-conformation-structure database for amino acid residues contains information on 114 828 individual residues derived from the spatial structures of 473 high-quality non-homologous proteins. The information in the database is obtained using a variety of different methods and can be used in various protein modeling applications.

Computational Biology↗

A new approach to protein fold recognition based on Delaunay tessellation of protein structure.

We propose new algorithms for sequence-structure compatibility (fold recognition) searches in multi-dimensional sequence-structure space. Individual amino acid residues in protein structures are represented by their C alpha atoms; thus each protein is described as a collection of points in three-dimensional space. Delaunay tessellation of a protein generates an aggregate of space-filling, irregular tetrahedra, or Delaunay simplices. Statistical analysis of quadruplet residue compositions of all Delaunay simplices in a representative dataset of protein structures leads to a novel four body contact residue potential expressed as log likelihood factor q. The q factors are calculated for native 20 letter amino acid alphabet and several reduced alphabets. Two sequence-structure compatibility functions are computed as (i) the sum of q factors for all Delaunay simplices in a given protein, or (ii) 3D-1D Delaunay tessellation profiles where the individual residue profile value is calculated as the sum of q factors for all simplices that share this vertex residue. Both threading functions have been implemented in structure-recognizes-sequence and sequence-recognizes-structure protocols for protein fold recognition. We find that both profile and total score based threading functions can distinguish both the native fold from incorrect folds for a sequence, and the native sequence from non-native sequences for a fold.

Amino Acid Sequence↗

Delaunay tessellation of proteins: four body nearest-neighbor propensities of amino acid residues.

Delaunay tessellation is applied for the first time in the analysis of protein structure. By representing amino acid residues in protein chains by C alpha atoms, the protein is described as a set of points in three-dimensional space. Delaunay tessellation of a protein structure generates an aggregate of space-filling irregular tetrahedra, or Delaunay simplices. The vertices of each simplex define objectively four nearest neighbor C alpha atoms, i.e., four nearest-neighbor residues. A simplex classification scheme is introduced in which simplices are divided into five classes based on the relative positions of vertex residues in protein primary sequence. Statistical analysis of the residue composition of Delaunay simplices reveals nonrandom preferences for certain quadruplets of amino acids to be clustered together. This nonrandom preference may be used to develop a four-body potential that can be used in evaluating sequence-structure compatibility for the purpose of inverted structure prediction.

Amino Acids↗

Statistical geometry analysis of proteins: implications for inverted structure prediction.

The topology of folded proteins from the representative dataset of well-defined three-dimensional protein structures is studied using a statistical geometry approach. Amino acid residues in protein chains are represented by C alpha atoms, thus reducing the protein three-dimensional structure to a set of points in three dimensional space. The Delaunay tessellation of a protein structure generates an aggregate of space-filling irregular tetrahedra, or Delaunay simplices. Each simplex objectively defines four nearest neighbor C alpha atoms, i.e. four nearest neighbor residues. The statistical analysis of residue composition of Delaunay simplices reveals nonrandom preferences for certain quadruplets of amino acids. These nonrandom preferences are used to develop a fitness function that evaluates sequence-structure compatibility. Using this fitness function, several tested native proteins score the highest among 100,000 random sequences with average protein amino acid composition. The statistical geometry approach, based solely on first principles, provides a unique means for protein structure analysis and has direct implications for inverted protein structure prediction.

Amino Acid Sequence↗

Pseudotorsional OCCO backbone angle as a single descriptor of protein secondary structure.

Protein secondary structure is conventionally identified using characteristic ranges of two backbone torsional angles phi and psi. We suggest that the secondary structure can be adequately characterized by a single descriptor, the Oi-1Ci-1CiOi (where i is the residue number) pseudotorsional backbone angle. A set of 102 structurally distinct protein chains from the Protein Data Bank was used to evaluate the adequacy of this descriptor. We find that a specific range of OCCO angles corresponds to each major secondary structure. The complete range of OCCO angles (-180 degrees to 179 degrees) was broken into 18 consecutive subranges of 20 degrees each, and each subrange was assigned a letter. Thus, the OCCO profiles for each protein in the database were "translated" into a sequence of letters. The Needleman-Wunsch primary sequence alignment algorithm was then used for secondary/tertiary structure comparison and alignment. Preliminary results indicate that this new approach has a significant potential for rapid identification of fold families in the Protein Data Bank.

Amino Acid Sequence↗

Rapid protein structure classification using one-dimensional structure profiles on the bioSCAN parallel computer.

Rapid growth of protein structures database in recent years requires an effective approach for objective comparison and classification of deposited protein structures. We describe a novel method for structure comparison and classification based on the alignment of one-dimensional structure profiles. These profiles are obtained by calculating the OCCO pseudodihedral angles (formed by O-C-C-O atoms of carbonyl groups of consecutive amino acid residues) from protein three-dimensional coordinates. These angle measurements are then converted into a 24 letter alphabet, and the protein structures are represented by sequences of letter from this alphabet. The BioSCAN parallel computer, designed for primary sequence alignment, is used to rapidly align and classify these one-dimensional structure profiles. We have developed and implemented weighted scoring matrix to identify structural classes based on commonly found structural motifs. The results of our experiments are in good agreement with the traditional protein structure classification schemes. One-dimensional structure profiles significantly improve efficiency of structure comparison and classification.

Algorithms↗