PubMed Health⌕ Search

Biomedical subjects

Meena Kishore Sakharkar

Publications and source records attributed to Meena Kishore Sakharkar.

18 recordsLinked to original sources

Functional and evolutionary analyses on expressed intronless genes in the mouse genome.

Using computational approaches we have identified 2017 expressed intronless genes in the mouse genome. Evolutionary analysis reveals that 56 intronless genes are conserved among the three domains of life--bacteria, archea and eukaryotes. These highly conserved intronless genes were found to be involved in essential housekeeping functions. About 80% of expressed mouse intronless genes have orthologs in eukaryotic genomes only, and thus are specific to eukaryotic organisms. 608 of these genes have intronless human orthologs and 302 of these orthologs have a match in OMIM database. Investigation into these mouse genes will be important in generating mouse models for understanding human diseases.

Animals↗

Huge proteins in the human proteome and their participation in hereditary diseases.

Protein lengths vary considerably from a few to thousands of amino acids and length variations are documented to have multiple effects. A computational approach to investigate the functional impact of protein length variation in genetic disorders is presented. The genes for huge proteins are found to have more introns. Our analysis also shows greater involvement of huge proteins in hereditary diseases.

Exons↗

Intron position conservation across eukaryotic lineages in tubulin genes.

A compilation of intron positions obtained from a large number of eukaryotic genomes across orthologous tubulins is explored for molecular evolution. Comparison of intron positions for 41 alpha, 80 beta, and 30 gamma tubulin genomic sequences indicates that the putative ancestral tubulin gene contained at least 19, 33, and 52 intron positions distributed at different sites in the coding regions for alpha, beta, and gamma tubulins, respectively. Many intron positions are old and are conserved across different eukaryotic lineages and intron distribution patterns are consistent with 'introns-early' hypothesis.

Animals↗

Human genome -- from pieces to patterns.

A profile of exon-intron lengths in genes shows a normal distribution. This observation suggests that different genes may have portions of their total exon and/or intron lengths in common. In order to explore the common exon-intron structural patterns that may arise due to common lengths across genes, we compared the exon-intron length patterns of annotated human genes. We discovered 1762278 conserved arrangements of exon-intron length across the otherwise unrelated and diverse genomic landscape. The existence of common exon-intron length patterns across unrelated genes suggests for their role of in gene assemblage and human genome design and architecture.

Base Sequence↗

Insights to metabolic network evolution by fusion proteins.

Human fusion proteins consisting of two or more fusion partners of prokaryotic origin exhibit accreted function. Recent studies have elucidated the importance of fusion proteins in complex regulatory networks. The significance of fusion proteins in cellular networks and their evolutionary mechanism is largely unknown. Here, we discuss the association of six fusion proteins with the citric acid cycle. We define possible gene fusion scenarios and show that they produce metabolites with high connectivity for complex networking. Complex networking of metabolites requires proteins with incremental structural architectures and functional capabilities. Such higher order functionality is frequently provided by fusion proteins. Therefore, evolution of fusion proteins capable of producing metabolites with greater connectivity for enhanced cross-talk between pathways is critical for the selection of multiple trajectories in maintaining a stoichiometric balance during regulation. The association of six fusion proteins with the citric acid cycle and their capability to produce metabolites with high connectivity index is intriguing. This suggests that fusion gene products and their evolution have had a key role in the selection of complex multifaceted networks. In addition, we propose that fusion proteins have gained additive biochemical function for a balanced regulation of metabolic networks.

Biological Evolution↗

Generation of a dataset for studying ligand effect on homodimer interface.

Protein dimer interfaces (homodimer - same polypeptide and heterodimer - different polypeptide) display geometric and chemical properties that give the non-covalent assembly its stability and specificity. Therefore, it is important to understand the molecular principles of dimer interaction. Several studies on homodimer interaction are available. However, a study on the effect of ligands (i.e. non-peptide compounds) on subunit interactions is not available. Hence, we generated a dataset of 62 identical homodimer pairs (one structure determined with an interface ligand and the other without an interface ligand) and analyzed the effect of interface ligands on dimer interface. The analysis suggests that homodimer interfaces having ligands are less hydrophobic with small interface area compared to those without ligands. We also found that ligands occupying = 7% interface area have negligible effect on dimer interaction.

Binding Sites↗

Identification of critical heterodimer protein interface parameters by multi-dimensional scaling in euclidian space.

Protein subunit dimers are either homodimers (consisting of identical polypeptides) or heterodimers (consisting of different polypeptides). Protein dimers are involved in several cellular processes and an understanding of their molecular principle in complexations (subunit-subunit interaction) is essential. This is generally studied using 3D structures of homodimers and heterodimers determined by X-ray crystallography. However, the current knowledge on subunit interaction is limited due to lack of sufficient 3D dimer structures. It is our interest to study heterodimers using 3D structures to identify interaction parameters that would help in the development of a model to predict heterodimer interaction sites just from protein sequences. The efficiency of such models depends on the weighted contribution of numerous parameters characterizing heterodimer interfaces. Therefore, we studied the salient features of 111 interface parameters in 65 heterodimer structures. In this study, we applied multi-dimensional scaling for dimensionality reduction on these parameters to select the most critical ones that best characterize heterodimer interfaces. The significance of these parameters in subunit interaction is discussed.

Computational Biology↗

A framework to sub-type HLA supertypes.

The human leukocyte antigen (HLA) alleles are extremely polymorphic among ethnic population and the peptide binding specificity varies for different alleles in a combinatorial manner. However, it has been suggested that majority of alleles can be covered within few HLA supertypes, where different members of a supertype bind similar peptides, yet exhibiting distinct repertoires. Since the overlap between different members of a supertype appears to be extensive, it is crucial to develop a framework for grouping alleles into supertypes just from sequence information. In this report, we define sub supertypes, where members show functional overlap with identical repertoire, and describe a strategy to group HLA-A, B and C alleles into different categories of sub supertypes. The strategy grouped 47% of 295 A alleles, 44% of 540 B alleles and 35% of 156 C alleles to just 36, 71 and 18 groups, respectively. The grouping is moderately validated using available binding data. However, the validation is limited due to lack of binding data. Hence, the data presented in this article serve as a framework to test specific functional overlap between alleles. The grouping of HLA alleles into different categories of sub supertypes has profound use in the understanding of antigenic peptide selection, degeneration and discrimination during T-cell mediated immune response. A complete knowledge of this phenomenon finds utility in epitope design for the development of HLA based vaccines and immuno-therapeutics.

Alleles↗

An analysis on gene architecture in human and mouse genomes.

A comparative genome analysis on exon-intron distribution profiles is performed for human and mouse genomes to deduce similarities and differences between them. Interestingly, both in human and mouse genomes, the total length in introns and intergenic DNA on each chromosome is significantly correlated to the chromosome size. The results presented provide a framework for understanding the nature and patterns of exon-intron length distributions, the constraints on them and their role in genome design and evolution.

Animals↗

u-Genome: a database on genome design in unicellular genomes.

Unicellular eukaryotes were among the first ones to be selected for complete genome sequencing because of the small size of their genomes and their interactions with humans and a broad range of animals and plants. Currently, ten completely sequenced unicellular genome sequences have been publicly released and as the number of available unicellular genomes increases, comparative genomics analysis within this group of organisms becomes more and more instructive. However, such an analysis is difficult to carry out without a suitable platform gathering not only the original annotations but also relevant information available in public databases or obtained by applying common bioinformatics methods. With the aim of solving these difficulties, we have developed a web-accessible database named u-Genome, the unicellular genome design database. The database is unique in featuring three datasets namely (1) orthologous proteins (2) paralogous proteins and (3) statistical distributions on exons, introns, intergenic DNA and correlations between them. A tool, Uniview, designed to visualize the gene structures for individual genes in the genome is also integrated. This database is of importance in understanding unicellular genome design and architecture and evolution related studies. The database is available through a web interface at http://sege.ntu.edu.sg/wester/ugenome.

Animals↗

Alternatively spliced human genes by exon skipping--a database (ASHESdb).

UNLABELLED: Alternative splicing of mRNA allows many gene products with different functions to be produced from a single coding sequence. Exon skipping is the most commonly known alternative splicing mechanism. A comprehensive database of alternative splicing by exon skipping is made available for the human genome data. 1,229 human genes are identified to exhibit alternative splicing by exon skipping. AVAILABILITY: http://sege.ntu.edu.sg/wester/ashes/.

Alternative Splicing↗

Can ends justify the means? Digging deep for human fusion genes of prokaryotic origin.

Gene fusion has been described as an important evolutionary phenomenon. This report focuses on identifying, analyzing, and tabulating human fusion proteins of prokaryotic origin. These fusion proteins are found to mimic operons, simulate protein-protein interfaces in prokaryotes, exhibiting multiple functions and alternative splicing in humans. The accredited biological functions for each of these proteins is made available as a database at http://sege.ntu.edu.sg/wester/fusion/

Alternative Splicing↗

A report on single exon genes (SEG) in eukaryotes.

Single exon genes (SEG) are archetypical of prokaryotes. Hence, their presence in intron-rich, multi-cellular eukaryotic genomes is perplexing. Consequently, a study on SEG origin and evolution is important. Towards this goal, we took the first initiative of identifying and counting SEG in nine completely sequenced eukaryotic organisms--four of which are unicellular (E. cuniculi, S. cerevisiae, S. pombe, P. falciparum) and five of which are multi-cellular (C. elegans, A. thaliana, D. melanogaster, M. musculus, H. sapiens). This exercise enabled us to compare their proportion in unicellular and multi-cellular genomes. The comparison suggests that the SEG fraction decreases with gene count (r = -0.80) and increases with gene density (r = 0.88) in these genomes. We also examined the distribution patterns of their protein lengths in different genomes.

Animals↗

Genome SEGE: a database for 'intronless' genes in eukaryotic genomes.

BACKGROUND: A number of completely sequenced eukaryotic genome data are available in the public domain. Eukaryotic genes are either 'intron containing' or 'intronless'. Eukaryotic 'intronless' genes are interesting datasets for comparative genomics and evolutionary studies. The SEGE database containing a collection of eukaryotic single exon genes is available. However, SEGE is derived using GenBank. The redundant, incomplete and heterogeneous qualities of GenBank data are a bottleneck for biological investigation in comparative genomics and evolutionary studies. Such studies often require representative gene sets from each genome and this is possible only by deriving specific datasets from completely sequenced genome data. Thus Genome SEGE, a database for 'intronless' genes in completely sequenced eukaryotic genomes, has been constructed. AVAILABILITY: http://sege.ntu.edu.sg/wester/intronless DESCRIPTION: Eukaryotic 'intronless' genes are extracted from nine completely sequenced genomes (four of which are unicellular and five of which are multi-cellular). The complete dataset is available for download. Data subsets are also available for 'intronless' pseudo-genes. The database provides information on the distribution of 'intronless' genes in different genomes together with their length distributions in each genome. Additionally, the search tool provides pre-computed PROSITE motifs for each sequence in the database with appropriate hyperlinks to InterPro. A search facility is also available through the web server. CONCLUSIONS: The unique features that distinguish Genome SEGE from SEGE is the service providing representative 'intronless' datasets for completely sequenced genomes. 'Intronless' gene sets available in this database will be of use for subsequent bio-computational analysis in comparative genomics and evolutionary studies. Such analysis may help to revisit the original genome data for re-examination and re-annotation.

Databases, Genetic↗

Distributions of exons and introns in the human genome.

The human genome is revisited using exon and intron distribution profiles. The 26,564 annotated genes in the human genome (build October, 2003) contain 233,785 exons and 207,344 introns. On average, there are 8.8 exons and 7.8 introns per gene. About 80% of the exons on each chromosome are < 200 bp in length. < 0.01% of the introns are < 20 bp in length and < 10% of introns are more than 11,000 bp in length. These results suggest constraints on the splicing machinery to splice out very long or very short introns and provide insight to optimal intron length selection. Interestingly, the total length in introns and intergenic DNA on each chromosome is significantly correlated to the determined chromosome size with a coefficient of correlation r = 0.95 and r = 0.97, respectively. These results suggest their implication in genome design.

Alternative Splicing↗

A novel MHCp binding prediction model.

Many statistical and molecular mechanics models have been developed and tested for major histocompatibility complex peptide (MHCp) binding predictions during the last decade. The statistical model prediction using pooled peptide sequence data and three-dimensional modeling prediction by molecular mechanics calculations have been assessed for efficiency and human leukocyte antigen diversity coverage. We describe a novel predictive model using information gleaned from 29 human MHCp crystal structures. The validation for the new model is performed using four different sets of data: (1) MHCp crystal structures, (2) peptides with known IC(50) binding values, (3) peptides tested positive by tetramer staining, (4) peptides with known binding information at the MHCBN database. The model produces high prediction efficiencies (average 60 %) with good sensitivity (approximately 50%-73%) and specificity (52%-58%) values. The average positive predictive value of the model is 89%, while the average negative predictive value is only 18%. The efficiency is very high in predicting binders and very low in predicting nonbinders. This model is superior to many existing methods because of its potential application to any given MHC allele whose sequence is clearly defined.

Amino Acid Sequence↗

Compression of functional space in HLA-A sequence diversity.

The major histocompatibility complex (MHC) is highly polymorphic and more than 1500 human MHC alleles are known to date. These alleles do not bind to a given peptide with identical affinity. Although MHC alleles are functionally related, it is difficult to quantify the functional variation between them. Three-dimensional structures of known MHC-peptide (MHCp) complexes suggest that specific peptide residues bind selectively to functional pockets in the binding groove. From a set of known MHCp structures we identified 21 critical polymorphic functional residue positions (CPFRP) that significantly reduced functional pocket variability to just 189 among 212 HLA-A alleles. Interestingly 101 HLA-A alleles clustered into 29 clusters such that the six functional pockets formed by the CPFRPs are identical within the cluster.

Alleles↗

Types of inter-atomic interactions at the MHC-peptide interface: identifying commonality from accumulated data.

BACKGROUND: Quantitative information on the types of inter-atomic interactions at the MHC-peptide interface will provide insights to backbone/sidechain atom preference during binding. Qualitative descriptions of such interactions in each complex have been documented by protein crystallographers. However, no comprehensive report is available to account for the common types of inter-atomic interactions in a set of MHC-peptide complexes characterized by variation in MHC allele and peptide sequence. The available x-ray crystallography data for these complexes in the Protein Databank (PDB) provides an opportunity to identify the prevalent types of such interactions at the binding interface. RESULTS: We calculated the percentage distributions of four types of interactions at varying inter-atomic distances. The mean percentage distribution for these interactions and their standard deviation about the mean distribution is presented. The prevalence of SS and SB interactions at the MHC-peptide interface is shown in this study. SB is clearly dominant at an inter-atomic distance of 3A. CONCLUSION: The prevalently dominant SB interactions at the interface suggest the importance of peptide backbone conformation during MHC-peptide binding. Currently, available algorithms are developed for protein sidechain prediction upon fixed backbone template. This study shows the preference of backbone atoms in MHC-peptide binding and hence emphasizes the need for accurate peptide backbone prediction in quantitative MHC-peptide binding calculations.

Binding Sites↗