PubMed Health⌕ Search

Biomedical subjects

Natarajan Kannan

Publications and source records attributed to Natarajan Kannan.

8 recordsLinked to original sources

CLASPP: A unified model for predicting post-translational modifications.

Post-Translational Modifications (PTMs) are a fundamental mechanism for regulating cellular pathways and increasing the functional diversity of the proteome. Accurately predicting the PTM types that are likely to occur at a given site in the primary sequence is a key challenge in functional proteomics. Existing PTM prediction models predominantly focus on either single PTM types or employ ensemble methods that combine multiple models to predict different PTM types. This fragmentation is largely driven by the vast imbalance in data availability across PTM types, making it difficult to predict multiple PTM types with a single model. To address this limitation, we present the Contrastively Learned Attention-based Stratified PTM Predictor (CLASPP), a unified PTM prediction model. CLASPP addresses imbalance challenges by leveraging unsupervised clustering-based undersampling and a novel contrastive learning framework tailored to PTM data. Additionally, our hierarchical data organization and curation are shown to improve CLASPP's performance by balancing the representation of individual PTM types and provides a standardized dataset to train and validate future model designs. Drawing inspiration from advancements in image and natural language processing, the CLASPP model employs a multi-stage training strategy and a high-quality, curated training dataset to improve PTM prediction performance. To uncover what is learned during the contrastive learning stage, the CLASPP model is shown to distinguish known protein kinase substrate specificity profiles as a form of explainability. Finally, we evaluate the application of CLASPP in predicting PTMs in different model organisms and experimentally validated ubiquitination sites in the understudied DCLK3 kinase. Overall, CLASPP represents a unified model for PTM prediction that addresses key bottlenecks in data imbalance and offers new strategies for biological data curation, thereby improving PTM-type prediction performance across diverse organisms.

Protein Processing, Post-Translational↗

Identification and classification of ion-channels across the tree of life provide functional insights into understudied CALHM channels.

The ion channel (IC) genes encoded in the human genome play fundamental roles in cellular functions and disease and are one of the largest classes of druggable proteins. However, limited knowledge of the diverse molecular and cellular functions carried out by ICs presents a major bottleneck in developing selective chemical probes for modulating their functions in disease states. The wealth of sequence data available on ICs from diverse organisms provides a valuable source of untapped information for illuminating the unique modes of channel regulation and functional specialization. However, the extensive diversification of IC sequences and the lack of a unified resource present a challenge in effectively using existing data for IC research. Here, we perform integrative mining of available sequence, structure, and functional data on 419 human ICs across disparate sources, including extensive literature mining by leveraging advances in large language models to annotate and curate the full complement of the "channelome". We employ a well-established orthology inference approach to identify and extend the IC orthologs across diverse organisms to above 48,000. We show that the depth of conservation and taxonomic representation of IC sequences can further be translated to functional similarities by clustering them into functionally relevant groups, which can be used for downstream functional prediction on understudied members. We demonstrate this by delineating co-conserved patterns characteristic of the understudied family of the Calcium Homeostasis Modulator (CALHM) family of ICs. Through mutational analysis of co-conserved residues altered in human diseases and electrophysiological studies, we show that these evolutionarily-constrained residues play an important role in channel gating functions. Thus, by providing new tools and resources for performing large comparative analyses on ICs, this study addresses the unique needs of the IC community and provides the groundwork for accelerating the functional characterization of dark channels for therapeutic intervention.

CALHM1↗

The hallmark of AGC kinase functional divergence is its C-terminal tail, a cis-acting regulatory module.

The catalytic activities of eukaryotic protein kinases (EPKs) are regulated by movement of the C-helix, movement of the N and C lobes upon ATP binding, and movement of the activation loop upon phosphorylation. Statistical analysis of the selective constraints associated with AGC kinase functional divergence reveals conserved interactions between these regulatory regions and three regions of the C-terminal tail (C-tail): the N-lobe tether (NLT), the active-site tether (AST), and the C-lobe tether (CLT). The NLT serves as a docking site for an upstream kinase PDK1 and, upon activation, positions the C-helix within the ATP binding pocket. The AST directly interacts with the ATP binding pocket, and the CLT interacts with the interlobe linker and the alphaC-beta4 loop, which appears to serve as a hinge for C-helix movement. The C-tail is a hallmark of AGC functional divergence inasmuch as most of the conserved core residues that distinguish AGC kinases from other EPKs are associated with the NLT, AST, or CLT. Moreover, several AGC catalytic core conserved residues that interact with the C-tail strikingly diverge from the canonical residues observed at corresponding positions in nearly all other EPKs, suggesting that the catalytic core may have coevolved with the C-tail in AGC kinases. These observations, along with the fact that the C-tail is needed for catalytic activity suggests that the C-tail is a cis-acting regulatory module that can also serve as a regulatory "handle," to which trans-acting cellular components can bind to modulate activity.

Amino Acid Sequence↗

Did protein kinase regulatory mechanisms evolve through elaboration of a simple structural component?

Statistical analysis of the functional constraints acting on eukaryotic protein kinases (EPKs) and on distantly related kinases suggests that EPK regulatory mechanisms evolved around an ancient structural component whose most distinctive features include the HxD-motif adjoining the catalytic loop, the F-helix, an F-helix aspartate, and the DFG-motif adjoined to the activation loop. The HxD-histidine constitutes a convergence point for signal integration, as conserved interactions link it to key catalytic residues, to the F-helix aspartate, and to both ends of the DFG-motif. These and other conserved features appear to be associated with DFG conformational changes and with coordinated movements possibly associated with phosphate transfer and ADP release. The EPKs have acquired structural features that link this core component to likely substrate-interacting regions at either end of the F-helix (most notably involving an F-helix tryptophan) and to three regions undergoing conformational changes upon kinase activation: the activation segment, the C-helix, and the nucleotide-binding pocket.

Adenosine Diphosphate↗

Computational analysis of protein tyrosine phosphatases: practical guide to bioinformatics and data resources.

The exponential growth of sequence data has become a challenge to database curators and end-users alike and biologists seeking to utilize the data effectively are faced with numerous analysis methods. Here, with practical examples from our bioinformatics analysis of the protein tyrosine phosphatases (PTPs), we show how computational analysis can be exploited to fuel hypothesis-driven experimental research through the exploration of online databases. We cover the following elements: (i) similarity searches and strategies to collect a non-redundant database of tyrosine-specific PTP domains; (ii) utilization of this database to classify human, fly, and worm PTPs (based on alignments and phylogenetic analysis); (iii) three-dimensional structural analysis to identify conserved regions (structure-function) and non-conserved selectivity-determining regions (substrate specificity); and (iv) genomic analysis, including mapping of exon structure, identification of pseudogenes, and exploration of disease databases. We discuss the importance of manual curation, illustrating examples in which pseudogenes give rise to predicted proteins in GenBank and note that domain servers, such as PFAM and SMART, erroneously include dual-specificity and lipid phosphatases in their collection of tyrosine-specific PTPs. To capitalize on our annotated set of 402 PTP domains (from 47 species and five phyla), we identify sequence conservation across taxonomic categories and explore structure-function relationships among tandem domain receptor-like PTPs. We define three Src homology 2 domain-containing PTP genes in stingray, zebrafish, and fugu and speculate on their evolutionary relationship with human pseudogenes. Our annotated sequences, along with a web service for phylogenetic classification of PTP domains, are available online (http://ptp.cshl.edu and http://science.novonordisk.com/ptp).

Amino Acid Sequence↗

Crystal structure of the E230Q mutant of cAMP-dependent protein kinase reveals an unexpected apoenzyme conformation and an extended N-terminal A helix.

Glu230, one of the acidic residues that cluster around the active site of the catalytic subunit of cAMP-dependent protein kinase, plays an important role in substrate recognition. Specifically, its side chain forms a direct salt-bridge interaction with the substrate's P-2 Arg. Previous studies showed that mutation of Glu230 to Gln (E230Q) caused significant decreases not only in substrate binding but also in the rate of phosphoryl transfer. To better understand the importance of Glu230 for structure and function, we solved the crystal structure of the E230Q mutant at 2.8 A resolution. Surprisingly, the mutant preferred an open conformation with no bound ligands observed, even though the crystals were grown in the presence of MgATP and the inhibitor peptide, IP20. This is in contrast to the wild-type protein that, under the same conditions, prefers the closed conformation of a ternary complex. The structure highlights the importance of the electrostatic surface not only for substrate binding and catalysis, but also for the mechanism for closing the active site cleft. This surface mutation clearly disrupts the recognition and binding of substrate peptide so that the enzyme prefers an open conformation that cannot trap ATP. This is consistent with the reinforcing concepts of conformational dynamics and the synergistic binding of ATP and substrate peptide. Another unusual feature of the structure is the observation of the entire N terminus (Gly1-Thr32) assumes an extended alpha-helix conformation. Finally, based on temperature factors, this mutant structure is more stable than the wild-type C-subunit in the apo state.

Adenosine Triphosphate↗

Evolutionary constraints associated with functional specificity of the CMGC protein kinases MAPK, CDK, GSK, SRPK, DYRK, and CK2alpha.

Amino acid residues associated with functional specificity of cyclin-dependent kinases (CDKs), mitogen-activated protein kinases (MAPKs), glycogen synthase kinases (GSKs), and CDK-like kinases (CLKs), which are collectively termed the CMGC group, were identified by categorizing and quantifying the selective constraints acting upon these proteins during evolution. Many constraints specific to CMGC kinases correspond to residues between the N-terminal end of the activation segment and a CMGC-conserved insert segment associated with coprotein binding. The strongest such constraint is imposed on a "CMGC-arginine" near the substrate phosphorylation site with a side chain that plays a role both in substrate recognition and in kinase activation. Two nearby buried waters, which are also present in non-CMGC kinases, typically position the main chain of this arginine relative to the catalytic loop. These and other CMGC-specific features suggest a structural linkage between coprotein binding, substrate recognition, and kinase activation. Constraints specific to individual subfamilies point to mechanisms for CMGC kinase specialization. Within casein kinase 2alpha (CK2alpha), for example, the binding of one of the buried waters appears prohibited by the side chain of a leucine that is highly conserved within CK2alpha and that, along with substitution of lysine for the CMGC-arginine, may contribute to the broad substrate specificity of CK2alpha by relaxing characteristically conserved, precise interactions near the active site. This leucine is replaced by a conserved isoleucine or valine in other CMGC kinases, thereby illustrating the potential functional significance of subtle amino acid substitutions. Analysis of other CMGC kinases similarly suggests candidate family-specific residues for experimental follow-up.

Amino Acid Sequence↗

Ran's C-terminal, basic patch, and nucleotide exchange mechanisms in light of a canonical structure for Rab, Rho, Ras, and Ran GTPases.

Proteins comprising the core of the eukaryotic cellular machinery are often highly conserved, presumably due to selective constraints maintaining important structural features. We have developed statistical procedures to decompose these constraints into distinct categories and to pinpoint critical structural features within each category. When applied to P-loop GTPases, this revealed within Rab, Rho, Ras, and Ran a canonical network of molecular interactions centered on bound nucleotide. This network presumably performs a crucial structural and/or mechanistic role considering that it has persisted for more than a billion years after the divergence of these families. We call these 'FY-pivot' GTPases after their most distinguishing feature, a phenylalanine or tyrosine that functions as a pivot within this network. Specific families deviate somewhat from canonical features in interesting ways, presumably reflecting their functional specialization during evolution. We illustrate this here for Ran GTPases, within which two highly conserved histidines, His30 and His139, strikingly diverge from their canonical counterparts. These, along with other residues specifically conserved in Ran, such as Tyr98, Lys99, and Phe138, appear to work in conjunction with FY-pivot canonical residues to facilitate alternative conformations in which these histidines are strategically positioned to couple Ran's basic patch and C-terminal switch to nucleotide exchange and effector binding. Other core components of the cellular machinery are likewise amenable to this approach, which we term Contrast Hierarchical Alignment and Interaction Network (CHAIN) analysis.

Amino Acid Sequence↗