PubMed HealthSearch

Biomedical subjects

A S Kolaskar

Publications and source records attributed to A S Kolaskar.

16 recordsLinked to original sources

Sequence alignment approach to pick up conformationally similar protein fragments.

Crystal structure data of globular proteins were used to prepare (phi, psi) probability maps of 20 proteinous amino acids. These maps were compared grid-wise with each other and a conformational similarity index was calculated for each pair of amino acids. A weight matrix, called Conformational Similarity Weight (CSW) matrix, was prepared using the conformational similarity index. This weight matrix was used to align sequences of 21 pairs of proteins whose crystal structures are known. The aligned regions with more than seven contiguous amino acids were further analysed by plotting average weight (W) values of overlapping hepatapeptides in these regions and carrying out curve fitting by Fourier series having TEN harmonics. The protein fragments corresponding to the half-linewidth of peaks were predicted as fragments having similar conformation in the protein pair under consideration. Such an approach allows us to pick up conformationally similar protein fragments with more than 67% accuracy.

Amino Acids

Computerization of virus data and its usefulness in virus classification.

Data on 537 Arboviruses and 180 other viruses have been collected and coded in two different formats. These data include information not only regarding the taxonomy and history of isolation, but also regarding the properties of biomacromolecules, proteins and nucleic acids. Information on antigenic relationships, histopathology and experimental viremia is also included. This information is stored in formats which allow the manipulation and analysis of data by dBASE III PLUS and MICRO-IS. A set of programs was written for interconversion and editing purposes. Transmission electron micrographs are scanned and stored. This stored information can be used in viral classification as shown by carrying out analysis of data on the Bunyaviridae family.

Arboviruses

Analysis of inverted repeats in primary structure of proteins.

A computer program has been developed to locate exact inverted repeating subsequences present anywhere in the given primary structure of proteins or nucleic acids. The output is amenable to protein sequence/nucleic acid query (PSQ/NAQ) packages. Our analysis has shown that there is a large number of proteins which have inverted repeats of more than four amino acid residues in length. However, the number is small when conditions such as the existence of more than 20 inverted repeats in given sequence or the existence of inverted repeats having more than five different types of amino acids are applied.

Amino Acid Sequence

A protein secondary structure database (PSS).

A protein secondary structure database (PSS) has been designed to correlate the Protein Sequence Database of the PIR-International with the atomic coordinates and bond connectivities database of the Protein Data Bank in the Brookhaven National Laboratory. The present database includes secondary structures determined by X-ray diffraction analysis, but not predicted structures. The database currently contains data from both the Protein Sequence Database and the Protein Data Bank Database, and will encompass the NMR database in the future. The main characteristics of the database are as follows: (1) the secondary structures, sites, regions and domains of structural interest are displayed together with protein primary structures; and (2) the secondary structure of a desired length of peptide fragment is displayed upon request, as are the peptide fragment(s) that correspond to a defined secondary structure. This database also has software to indicate amino acid pairs having hydrogen bonds and to count the occurrence frequency of each pair as well as the conformational parameters widely used in semi-empirical methods of secondary structure prediction.

Amino Acid Sequence

A semi-empirical method for prediction of antigenic determinants on protein antigens.

Analysis of data from experimentally determined antigenic sites on proteins has revealed that the hydrophobic residues Cys, Leu and Val, if they occur on the surface of a protein, are more likely to be a part of antigenic sites. A semi-empirical method which makes use of physicochemical properties of amino acid residues and their frequencies of occurrence in experimentally known segmental epitopes was developed to predict antigenic determinants on proteins. Application of this method to a large number of proteins has shown that our method can predict antigenic determinants with about 75% accuracy which is better than most of the known methods. This method is based on a single parameter and thus very simple to use.

Algorithms

An extension of the graph theoretical approach to predict the secondary structure of large RNAs: the complex of 16S and 23S rRNAs from E. coli as a case study.

An algorithm using the graph theoretical approach to predict secondary structures of large nucleic acids is discussed. Reliability of prediction can be improved by incorporating available experimental data and sequence homology information. As a case study, this algorithm is applied to predict the secondary structure of the 16S-23S rRNA complex from E. coli. It was found that several structures of the complex can coexist. The computer program developed to predict the secondary structure of large RNAs can be run on IBM PC/AT compatible systems.

Algorithms

Prediction of the recognition sites on 16S and 23S rRNAs from E. coli for the formation of 16S-23S rRNA complex.

Interactions between RNA molecules have been postulated to play an important role in the assembly of ribosomes. Using the sequence analysis and the search of continuous complementary regions on 16S rRNA and 23S rRNA, the recognition sites involved in the formation of ribosome of E. coli are postulated. The number of postulated sites was narrowed down by taking available experimental data. The suggestive evidence for correct postulation is obtained from sequence comparison studies of 16S and 23S rRNAs from various species. The sites 891-899 and 1195-1203 on 16S rRNA along with the corresponding complementary sites 1904-1912 and 760-768 on 23S rRNA are predicted to be the most probable candidates for the sites of recognition between 16S and 23S rRNAs. The possibility of the involvement of the additional site 630-638 on 16S rRNA with its complementary site 2031-2039 on 23S rRNA cannot be ruled out.

Computer Simulation

Contextual constraints on codon pair usage: structural and biological implications.

Complementary DNA sequence data of 278 protein coding genes from prokaryotic systems have been analysed at the level of near neighbour codon pairs. Our analysis points out that constraints exist even at the level of near neighbour codon pairs. These constraints are in addition to those which arise due to relative levels of tRNA. Codon pairs, which in the data base have different occurrence values from their expected values, neither have common secondary structure nor do have better stabilization due to high base stacking. Our study points out that there are strong interaction between constituent codons in these codon pairs. These strongly interacting codon pairs, we suggest, are involved in the formation of three dimensional structural elements of cDNA/mRNA and interact with ribosome and thus modulate translation.

Base Sequence

Analysis of repeating oligonucleotide sequences in ribonucleic acids using an Apple II microcomputer.

A simple computer program has been developed to locate repeating subsequences of all possible lengths in a given nucleic acid. The observed number of repeats of subsequences was compared with the expected number of such repeats in several RNAs. The analysis showed that, in the case of rRNAs, there are no constraints in the choice of the fourth and the higher order nucleotides, while the selection is maximum at the level of nearest neighbour. This is, however, not true for RNAs coding for proteins, where the constraints are also found at the level of nucleotides containing five or more bases.

Algorithms

A method to locate protein coding sequences in DNA of prokaryotic systems.

cDNA sequence data from E. coli phages, for which complete genome sequences are known, have been analysed, From this analysis thirteen triplets have been identified as markers to distinguish protein-coding frames from fortuitous open reading frames. The region of -18 to +18 nucleotides around ATG/GTG, has been analysed and used to identify initiator codons from internal ATG/GTG. With the aid of criteria defined above a method has been developed to locate protein coding sequences by a combination of 'gene search by signal' and 'gene search by content' approaches. Application of this method to prokaryotic systems including those which were not part of our data base indicates that it is quite accurate and general in nature.

Bacterial Proteins

Empirical torsional potential functions from protein structure data. Phi- and psi-potentials for non-glycyl amino acid residues.

The torsional potential functions Vt(phi) and Vt(psi) around single bonds N--C alpha and C alpha--C, which can be used in conformational studies of oligopeptides, polypeptides and proteins, have been derived, using crystal structure data of 22 globular proteins, fitting the observed distribution in the (phi, psi)-plane with the value of Vtot(phi, psi), using the Boltzmann distribution. The averaged torsional potential functions, obtained from various amino acid residues in L-configuration, are Vt(phi) = 1.0 cos (phi + 60 degrees); Vt(psi) = 0.5 cos (psi + 60 degrees) - 1.0 cos (2 psi + 30 degrees) - 0.5 cos (3 psi + 30 degrees). The dipeptide energy maps Vtot(phi, psi) obtained using these functions, instead of the normally accepted torsional functions, were found to explain various observations, such as the absence of the left-handed alpha helix and the C7 conformation, and the relatively high density of points near the line psi = 0 degrees. These functions derived from observational data on protein structures, will, it is hoped, explain various previously unexplained facts in polypeptide conformation.

Amino Acid Sequence

Recognition of helper T cell epitopes in envelope (E) glycoprotein of Japanese encephalitis, west Nile and Dengue viruses.

Helper T (Th) cell antigenic sites were predicted from the primary amino acid sequence (approximately 500 in length) of the envelope (E) glycoprotein (gp) of Japanese encephalitis (JE), West Nile (WN) and Dengue (DEN) I-IV flaviviruses. Prediction of Th epitopes was done by analyzing the occurrence of amphipathic segments, Rothbard-Taylor tetra/pentamer motifs and presence of alpha helix-preferring amino acids. The simultaneous occurrence of all these parameters in segments of E gp were used as criteria for prediction as Th epitopes. Only one cross reactive epitope was predicted in the C-terminal region of the E gp predicted segments of all flaviviruses analyzed. This region is one of the longest amphipathic stretch (approximately from 420 to 455) and also has a fairly large amphipathic score. Based on the predicted findings three selected peptides were synthesized and analyzed for their ability to induce in vitro T cell proliferative response in different inbred strains of mice (Balb/c, C57BL6, C3H/HeJ). Synthetic peptide I and II prepared from C-terminal region gave a cross reactive response to JE, WN and Den-II in Balb/c and C3H/HeJ mice. Synthetic peptide III prepared from N-terminal region gave a proliferative response to DEN-II in Balb/c strain only, indicating differential antigen presentation.

Algorithms