PubMed Health⌕ Search

Biomedical subjects

M K Sakharkar

Publications and source records attributed to M K Sakharkar.

8 recordsLinked to original sources

Towards the MHC-peptide combinatorics.

The exponentially increased sequence information on major histocompatibility complex (MHC) alleles points to the existence of a high degree of polymorphism within them. To understand the functional consequences of MHC alleles, 36 nonredundant MHC-peptide complexes in the protein data bank (PDB) were examined. Induced fit molecular recognition patterns such as those in MHC-peptide complexes are governed by numerous rules. The 36 complexes were clustered into 19 subgroups based on allele specificity and peptide length. The subgroups were further analyzed for identifying common features in MHC-peptide binding pattern. The four major observations made during the investigation were: (1) the positional preference of peptide residues defined by percentage burial upon complex formation is shown for all the 19 subgroups and the burial profiles within entries in a given subgroup are found to be similar; (2) in class I specific 8- and 9-mer peptides, the fourth residue is consistently solvent exposed, however this observation is not consistent in class I specific 10-mer peptides; (3) an anchor-shift in positional preference is observed towards the C terminal as the peptide length increases in class II specific peptides; and (4) peptide backbone atoms are proportionately dominant at the MHC-peptide interface.

Animals↗

Phylogenetic relationships of the seven coat protein subunits of the coatomer complex, and comparative sequence analysis of murine xenin and proxenin.

The coatomer complex is involved in intracellular protein transport and comprises an assembly of seven polypeptide subunits designated alpha, beta, beta', gamma, delta, epsilon, and zeta COP. Rooted phylogenetic trees constructed from the full-length cDNA and amino acid sequences of 49 COP entities in different eukaryotes from yeast to man generally revealed striking conservation of each subunit through evolution. Both nucleotide and protein trees displayed close relationships between alpha and beta' subunits, between beta and gamma subunits, and between delta and zeta subunits, implying evolution from common ancestors as well as functional similarity. Interestingly, although 6 out of 7 epsilon-COP genes appeared to be grouped and related to the beta-COP genes, 4 out of 7 epsilon

Amino Acid Sequence↗

Generation of a database containing discordant intron positions in eukaryotic genes (MIDB).

MOTIVATION: Intron sliding is the relocation of intron-exon boundaries over short distances and is often also referred to as intron slippage or intron migration or intron drift. We have generated a database containing discordant intron positions in homologous genes (MIDB--Mismatched Intron DataBase). Discordant intron positions are those that are either closely located in homologous genes (within a window of 10 nucleotides) or an intron position that is present in one gene but not in any of its homologs. The MIDB database aims at systematically collecting information about mismatched introns in the genes from GenBank and organizing it into a form useful for understanding the genomics and dynamics of introns thereby helping understand the evolution of genes. RESULTS: Intron displacement or sliding is critically important for explaining the present distribution of introns among orthologous and paralogous genes. MIDB allows examining of intron movements and allows mapping of intron positions from homologous proteins onto a single sequence. The database is of potential use for molecular biologists in general and for researchers who are interested in gene evolution and eukaryotic gene structure. Partial analysis of this database allowed us to identify a few putative cases of intron sliding. AVAILABILITY: http://intron.bic.nus.edu.sg/midb/midb.html

Amino Acid Sequence↗

Knowledge-based grouping of modeled HLA peptide complexes.

Human leukocyte antigens are the most polymorphic of human genes and multiple sequence alignment shows that such polymorphisms are clustered in the functional peptide binding domains. Because of such polymorphism among the peptide binding residues, the prediction of peptides that bind to specific HLA molecules is very difficult. In recent years two different types of computer based prediction methods have been developed and both the methods have their own advantages and disadvantages. The nonavailability of allele specific binding data restricts the use of knowledge-based prediction methods for a wide range of HLA alleles. Alternatively, the modeling scheme appears to be a promising predictive tool for the selection of peptides that bind to specific HLA molecules. The scoring of the modeled HLA-peptide complexes is a major concern. The use of knowledge based rules (van der Waals clashes and solvent exposed hydrophobic residues) to distinguish binders from nonbinders is applied in the present study. The rules based on (1) number of observed atomic clashes between the modeled peptide and the HLA structure, and (2) number of solvent exposed hydrophobic residues on the modeled peptide effectively discriminate experimentally known binders from poor/nonbinders. Solved crystal complexes show no vdW Clash (vdWC) in 95% cases and no solvent exposed hydrophobic peptide residues (SEHPR) were seen in 86% cases. In our attempt to compare experimental binding data with the predicted scores by this scoring scheme, 77% of the peptides are correctly grouped as good binders with a sensitivity of 71%.

Alleles↗

IE-Kb: intron exon knowledge base.

SUMMARY: IE-Kb (Intron Exon-Knowledge base) illustrates the intron-exon dynamics in eukaryotic genes. We have developed three different knowledge sets, namely 'Non-redundant ExInt', 'Non-redundant Pfam-ExInt complement' and 'Non-redundant GenBank eukaryotic subdivisional sets' to understand this phenomenon. Statistical analysis is performed on each knowledge set and the results are made available online. The entries in knowledge sets are ranked based on their intron length, exon length and protein length with relational hyper-links to the corresponding intron phase, intron position, intron sequence, gene definition and parent GenBank entry.

Artificial Intelligence↗

TRES: comparative promoter sequence analysis.

Comparative promoter analysis is a promising strategy for elucidation of common regulatory modules conserved in evolutionarily related sequences or in genes showing common expression profiles. To facilitate such analysis, we have developed a software tool that detects conserved transcription factor binding sites, cis-elements, palindromes and k-tuples simultaneously in a set of sequences, and thus helps to identify putative motifs for designing further experiments.

Animals↗

Integration of bioInformatics tools at the National University of Singapore (NUS).

In the past decade "Big Science" such as the Genome Project has generated an enormous amount of data in the life sciences. Concurrently, the synergy of this project with existing research has quickened the pace of biological discovery. But the major drawback that is beginning to be felt worldwide is the primitive level of organisation in the data accumulated. Without a proper framework or knowledge scaffold to hang and interconnect the various bits of data and information, the national knowledge-to-data ratio is declining rapidly. We are trying to serve a solution to this enigma by providing a World Wide Web (WWW) interface to Biosoftware and at the same time have come up with a database integration tool that can query heterogeneous, geographically scattered and disparate databases simultaneously. In this report we will talk about BioInformatics in general with specific reference to BioInformatics Centre (BIC) at the National University of Singapore.

Computational Biology↗

Development of software tools at BioInformatics Centre (BIC) at the National University of Singapore (NUS).

There is burgeoning volume of information and data arising from the rapid research and unprecedented progress in molecular biology. This has been particularly affected by the Human Genome Project which is trying to completely sequence three billion nucleotides of the human genome (1),(1a). Other genome sequencing projects are also contributing substantially to this exponential growth in the number of DNA nucleotides and proteins sequenced. The number of journals, reports and research papers and tools required for the analysis of these sequences has also increased. For this the life sciences today needs tools in information technology and computation to prevent degeneration of this data into an inchoate accretion of unconnected facts and figures. The recently formed BioInformatics Centre (BIC) at the National University of Singapore (NUS) provides access to various commonly used computational tools available over the World Wide Web (WWW)--using a uniform interface and easy access. We have also come up with a new database tool. BioKleisli, which allows you to interact with various geographically scattered, heterogeneous, structurally complex and constantly evolving data sources. This paper summarises the importance of network access and database integration to biomedical research and gives a glimpse of current research conducted at BIC.

Animals↗