PubMed Health⌕ Search

Biomedical subjects

Debasis Dash

Publications and source records attributed to Debasis Dash.

12 recordsLinked to original sources

Role of intrinsic disorder in transient interactions of hub proteins.

Hubs in the protein-protein interaction network have been classified as "party" hubs, which are highly correlated in their mRNA expression with their partners while "date" hubs show lesser correlation. In this study, we explored the role of intrinsic disorder in date and party hub interactions. The data reveals that intrinsic disorder is significantly enriched in date hub proteins when compared with party hub proteins. Intrinsic disorder has been largely implicated in transient binding interactions. The disorder to order transition, which occurs during binding interactions in disordered regions, renders the interaction highly reversible while maintaining the high specificity. The enrichment of intrinsic disorder in date hubs may facilitate transient interactions, which might be required for date hubs to interact with different partners at different times.

Computational Biology↗

Conformational flexibility may explain multiple cellular roles of PEST motifs.

PEST sequences are one of the major motifs that serve as signal for the protein degradation and are also involved in various cellular processes such as phosphorylation and protein-protein interaction. In our earlier study, we found that these motifs contribute largely to eukaryotic protein disorder. This observation led us to evaluate their conformational variability in the nonredundant Protein Data Bank (PDB) structures. For this purpose, crystallographic temperature factors, structural alignment of multiple NMR models, and dihedral angle order parameters have been used in this study. The study has revealed the hypermobility of PEST motifs as compared to other regions of the protein. Conformational flexibility may allow them to participate in number of molecular interactions under different conditions. This analysis may explain the role of protein backbone flexibility in bringing about multiple cellular roles of PEST motifs.

Amino Acid Motifs↗

Intrinsic unstructuredness and abundance of PEST motifs in eukaryotic proteomes.

The study of unfolded protein regions has gained importance because of their prevalence and important roles in various cellular functions. These regions have characteristically high net charge and low hydrophobicity. The amino acid sequence determines the intrinsic unstructuredness of a region and, therefore, efforts are ongoing to delineate the sequence motifs, which might contribute to protein disorder. We find that PEST motifs are enriched in the characterized disordered regions as compared with globular ones. Analysis of representative PDB chains revealed very few structures containing PEST sequences and the majority of them lacked regular secondary structure. A proteome-wide study in completely sequenced eukaryotes with predicted unfolded and folded proteins shows that PEST proteins make up a large fraction of unfolded dataset as compared with the folded proteins. Our data also reveal the prevalence of PEST proteins in eukaryotic proteomes (approximately 25%). Functional classification of the PEST-containing proteins shows an over- and under-representation in proteins involved in regulation and metabolism, respectively. Furthermore, our analysis shows that predicted PEST regions do not exhibit any preference to be localized in the C terminals of proteins, as reported earlier.

Amino Acid Sequence↗

Conformational analysis of invariant peptide sequences in bacterial genomes.

The functional significance of evolutionarily conserved motifs/patterns of short regions in proteins is well documented. Although a large number of sequences are conserved, only a small fraction of these are invariant across several organisms. Here, we have examined the structural features of the functionally important peptide sequences, which have been found invariant across diverse bacterial genera. Ramachandran angles (phi,psi) have been used to analyze the conformation, folding patterns and geometrical location (buried/exposed) of these invariant peptides in different crystal structures harboring these sequences. The analysis indicates that the peptides preferred a single conformation in different protein structures, with the exception of only a few longer peptides that exhibited some conformational variability. In addition, it is noticed that the variability of conformation occurs mainly due to flipping of peptide units about the virtual C(alpha)...C(alpha) bond. However, for a given invariant peptide, the folding patterns are found to be similar in almost all the cases. Over and above, such peptides are found to be buried in the protein core. Thus, we can safely conclude that these invariant peptides are structurally important for the proteins, since they acquire unique structures across different proteins and can act as structural determinants (SD) of the proteins. The location of these SD peptides on the protein chain indicated that most of them are clustered towards the N-terminal and middle region of the protein with the C-terminal region exhibiting low preference. Another feature that emerges out of this study is that some of these SD peptides can also play the roles of "fold boundaries" or "hinge nucleus" in the protein structure. The study indicates that these SD peptides may act as chain-reversal signatures, guiding the proteins to adopt appropriate folds. In some cases the invariant signature peptides may also act as folding nuclei (FN) of the proteins.

Amino Acid Sequence↗

In silico characterization of the INO80 subfamily of SWI2/SNF2 chromatin remodeling proteins.

Proteins belonging to SNF2 family of DNA dependent ATPases are important members of the chromatin remodeling complexes that are implicated in epigenetic control of gene expression. The yeast Ino80, the catalytic ATPase subunit of the INO80 complex, is the most recently described member of the SNF2 family. Outside the conserved ATPase domain, it has very little similarity with other well-characterized SNF2 proteins hence it is believed to represent a new subfamily. We have identified new members of this subfamily in different organisms and have detected characteristic features of this subfamily. Using various data mining tools we have identified a new, previously undetected domain in all members of this subfamily. This domain designated DBINO is characteristic of the INO80 subfamily and is predicted to have DNA-binding function. The presence of this domain in all the INO80 subfamily proteins from different organisms suggests its conserved function in evolution.

Amino Acid Sequence↗

Comparative analysis of protein unfoldedness in human housekeeping and non-housekeeping proteins.

Absence of any regular structure is increasingly being observed in structural studies of proteins. These disordered regions or random coils, which have been observed under physiological conditions, are indicators of protein plasticity. The wide variety of interactions possible due to the flexibility of these 'natively disordered' regions confers functional advantage to the protein and the organism in general. This concept is underscored by the increasing proportion of intrinsically unstructured proteins seen with the ascension in the complexity of the organisms. The 'natively unfolded/disordered' state of the protein can be predicted utilizing Uversky's or Dunker's algorithm. We utilized Uversky's prediction scheme and based on the unique position of a protein in the charge-hydrophobicity plot, a derived net score was used to predict the overall disorder of the human housekeeping and non-housekeeping proteins. Substantial numbers of proteins in both the classes were predicted to be unfolded. However, comparative genomic analysis of predicted unfolded Homo sapiens proteins with homologues in Caenorhabditis elegans, Drosophila melanogaster and Mus musculus revealed significant increase in unfoldedness in non-housekeeping proteins in comparison with housekeeping proteins. Our analysis in the evolutionary context suggests addition or substitution of amino acid residues which favour unfoldedness in non-housekeeping proteins compared to housekeeping proteins.

Algorithms↗

CoPS: Comprehensive Peptide Signature database.

UNLABELLED: We present the development of a Comprehensive database of 12 076 invariant Peptide Signatures (CoPS) derived from 52 bacterial genomes with a minimum occurrence in at least seven organisms. These peptides were observed in functionally similar proteins and are distributed over nearly 1250 different functional proteins. The database provides function, structure and occurrence in biochemical pathways of the proteins containing these signature peptides. It houses additional information on the signature peptides, such as identical match in other motif/pattern (e.g. PROSITE, BLOCKS, PRINTS and Pfam) databases and the database of interacting proteins, human proteome and mutation effect on these signature peptides. There is a wide applicability of this database in the identification of critical functional residues in proteins. The database also facilitates the identification of folding nucleus/structural determinants in proteins and functional assignment to yet unknown proteins. We demonstrate functional assignment to 2605 hypothetical proteins in bacterial genomes and 112 unknown proteins in human using this database. AVAILABILITY: The database can be freely accessed through the following URL: http://203.195.151.46/copsv2/index.html or http://203.90.127.70/copsv2/index.html

Bacterial Proteins↗

Recognition and analysis of protein-coding genes in severe acute respiratory syndrome associated coronavirus.

MOTIVATION: The recent outbreak of severe acute respiratory syndrome (SARS) caused by SARS coronavirus (SARS-CoV) has necessitated an in-depth molecular understanding of the virus to identify new drug targets. The availability of complete genome sequence of several strains of SARS virus provides the possibility of identification of protein-coding genes and defining their functions. Computational approach to identify protein-coding genes and their putative functions will help in designing experimental protocols. RESULTS: In this paper, a novel analysis of SARS genome using gene prediction method GeneDecipher developed in our laboratory has been presented. Each of the 18 newly sequenced SARS-CoV genomes has been analyzed using GeneDecipher. In addition to polyprotein 1ab(1), polyprotein 1a and the four genes coding for major structural proteins spike (S), small envelope (E), membrane (M) and nucleocapsid (N), six to eight additional proteins have been predicted depending upon the strain analyzed. Their lengths range between 61 and 274 amino acids. Our method also suggests that polyprotein 1ab, polyprotein 1a, S, M and N are proteins of viral origin and others are of prokaryotic. Putative functions of all predicted protein-coding genes have been suggested using conserved peptides present in their open reading frames. AVAILABILITY: Detailed results of GeneDecipher analysis of all the 18 strains of SARS-CoV genomes are available at http://www.igib.res.in/sarsanalysis.html

Algorithms↗

A novel complexity measure for comparative analysis of protein sequences from complete genomes.

Analysis of sequence complexities of proteins is an important step in the characterization and classification of new genomes. A new measure has been proposed to compute sequence complexity in protein sequences based on linguistic complexity. The algorithm requires a single parameter, is computationally simple and provides a framework for comparative genomic analysis. Protein sequences were classified into groups of high or low complexity based on a quantitative measure termed F(c), which is proportional to the fraction of low complexity sequence present in the protein. The algorithm was tested on sequences of 196 non-homologous proteins whose crystal structures are available at </=2.0 A resolution. Protein sequences of high complexity had 'globular' structures (95% agreement), whereas those of low complexity had non-globular structures (80% agreement). Application of this measure to proteins of unknown structure/function from different genomes revealed that the sequences of high complexity constitute the majority in all genomes (about 90% in Archaea, about 93% in Eubacteria, 89% in Saccharomyces cerevisiae and 90% in Caenorhabditis elegans). Aeropyrum pernix among Archaeae and Deinococcus radiodurans among Eubacteria have the lowest fraction of high complexity proteins (75% and 80% respectively). Further, it was observed that a few bacterial pathogens (Mycobacterium tuberculosis, Pseudomonas aeruginosa) have high fraction of low complexity proteins. The program ScanCom is available from the authors as a PERL script (UNIX system).

Algorithms↗

Role of histidine interruption in mitigating the pathological effects of long polyglutamine stretches in SCA1: A molecular approach.

Polyglutamine expansions, leading to aggregation, have been implicated in various neurodegenerative disorders. The range of repeats observed in normal individuals in most of these diseases is 19-36, whereas mutant proteins carry 40-81 repeats. In one such disorder, spinocerebellar ataxia (SCA1), it has been reported that certain individuals with expanded polyglutamine repeats in the disease range (Q(12)HQHQ(12)HQHQ(14/15)) but with histidine interruptions were found to be phenotypically normal. To establish the role of histidine, a comparative study of conformational properties of model peptide sequences with (Q(12)HQHQ(12)HQHQ(12)) and without (Q(42)) interruptions is presented here. Q(12)HQHQ(12)HQHQ(12) displays greater solubility and lesser aggregation propensity compared to uninterrupted Q(42) as well as much shorter Q(22). The solvent and temperature-driven conformational transitions (beta structure <--> random coil --> alpha helix) displayed by these model polyQ stretches is also discussed in the present report. The study strengthens our earlier hypothesis of the importance of histidine interruptions in mitigating the pathogenicity of expanded polyglutamine tract at the SCA1 locus. The relatively lower propensity for aggregation observed in case of histidine interrupted stretches even in the disease range suggests that at a very low concentration, the protein aggregation in normal cells, is possibly not initiated at all or the disease onset is significantly delayed. Our present study also reveals that besides histidine interruption, proline interruption in polyglutamine stretches can lower their aggregation propensity.

Amino Acid Sequence↗

Spectrum of beta-thalassemia mutations and their association with allelic sequence polymorphisms at the beta-globin gene cluster in an Eastern Indian population.

In this report, the spectrum of beta-thalassemia mutations and genotype-to-phenotype correlations were defined in large number of patients (beta-thalassemia carriers and major) with varying disease severity in an Eastern Indian population mainly from the state of West Bengal. The five most common beta-thalassemia mutations were detected, which included IVS1-5 (G-->C), codon 15 (G-->A), codon 26 (G-->A), codon 30 (G-->C), and codon 41/42 (-TCTT). These accounted for 85% in 80 beta-thalassemic alleles deciphered from 56 patients, including beta-thalassemia major and carriers, and 15% of alleles remained uncharacterized in these patients. Expression of the human beta-globin gene is regulated by an array of cis-acting DNA elements, including five DNase I hypersensitive sites (HSs) in the locus control region (LCR), promoters that incorporate certain silencer elements, and enhancers at 3' of the beta-globin gene. For detailed studies and to understand the molecular basis of beta-thalassemia, we studied two groups of subjects: a group of 12 patients from four families having beta-thalassemia major and carrier phenotype and a control group of 26 healthy individuals. In these two groups, we examined portions of the beta-globin gene locus control region HSs 1, 2, 3, and 4, which included the (CA)(x)(TA)(y) repeat motif, the (AT)(x)N(y)(AT)(z) repeat motif, the inverted repeat sequence TGGGGACCCCA, the promoter region of the (G)gamma-globin gene, an (AT)(x)(T)(y) repeat 5' of the silencer region, and the beta-globin gene and its 3' flanking region. We investigated the allelic sequence polymorphisms in these regions and their association with the beta-thalassemia mutations to know the possible genotype-phenotype relationship in beta-thalassemia patients. An analysis of cis-acting regulatory regions showed varied sequence haplotypes associated with some frequent beta-thalassemia mutations in this Eastern Indian population.

Adult↗

Origin and instability of GAA repeats: insights from Alu elements.

Expansion of GAA repeats in the intron of the frataxin gene is involved in the autosomal recessive Friedreich's ataxia (FRDA). The GAA repeats arise from a stretch of adenine residues of an Alu element. These repeats have a size ranging from 7- 38 in the normal population, and expand to thousands in the affected individuals. The mechanism of origin of GAA repeats, their polymorphism and stability are not well understood. In this study, we have carried out an extensive analysis of GAA repeats at several loci in the humans. This analysis indicates the association of a majority of GAA repeats with the 3' end of an "A" stretch present in the Alu repeats. Further, the prevalence of GAA repeats correlates with the evolutionary age of Alu subfamilies as well as with their relative frequency in the genome. Our study on GAA repeat polymorphism at some loci in the normal population reveals that the length of the GAA repeats is determined by the relative length of the flanking A stretch. Based on these observations, a possible mechanism for origin of GAA repeats and modulatory effects of flanking sequences on repeat instability mediated by DNA triplex is proposed.

3' Flanking Region↗