PubMed Health⌕ Search

Biomedical subjects

Artem Cherkasov

Publications and source records attributed to Artem Cherkasov.

At least 19 recordsLinked to original sources

Structure-based discovery of inhibitors of Mac1 domain of nonstructural protein-3 of SARS-CoV-2 by machine learning-augmented screening of chemical space.

Significant efforts have been recently dedicated to the discovery of small molecule inhibitors against the Macrodomain 1 (Mac1) of nonstructural protein 3 (NSP3) as potential antivirals for SARS-CoV-2. Thus, Mac1 has also been selected as the target for the Critical Assessment of Hit-finding Experiments (CACHE) challenge #3. As contestants in that challenge, we developed a computational strategy that ranked on the top among all 23 participants in the competition and resulted in the discovery of a novel chemical series of non-charged Mac1 inhibitors. Those have been identified through the combination of machine learning-accelerated virtual screening of Enamine REAL Diversity Subset of approximately 25 million compounds and consequent hit expansion into the entire Enamine REAL Space library. In particular, the initially identified hit compound CACHE3-HI_1706_56 (KD = 20 μM) was explored by probing 17 close analogues from a library of 44 billion molecules from the Enamine REAL. All those analogues effectively displaced the Mac1-binding ADP-ribose peptide, and 12 were confirmed to engage with Mac1 by the Surface Plasmon Resonance experiments, revealing a new chemical series of compounds for hit-to-lead optimization. The structure of the CACHE3-HI_1706_56-Mac1 complex was further determined at high resolution with crystallography, confirming initial computational predictions. Our results illustrate the effectiveness of ML-accelerated docking to rapidly identify novel chemical series and provide a strong foundation for the development of SARS-CoV-2 NSP3 Mac1 inhibitors.

CACHE challenge↗

Progressive docking: a hybrid QSAR/docking approach for accelerating in silico high throughput screening.

A combination of protein-ligand docking and ligand-based QSAR approaches has been elaborated, aiming to speed-up the process of virtual screening. In particular, this approach utilizes docking scores generated for already processed compounds to build predictive QSAR models that, in turn, assess hypothetical target binding affinities for yet undocked entries. The "progressive docking" has been tested on drug-like substances from the NCI database that have been docked into several unrelated targets, including human sex hormone binding globulin (SHBG), carbonic anhydrase, corticosteroid-binding globulin, SARS 3C-like protease, and HIV1 reverse transcriptase. We demonstrate that progressive docking can reduce the amount of computations 1.2- to 2.6-fold (when compared to traditional docking), while maintaining 80-99% hit recovery rates. This progressive-docking procedure, therefore, substantially accelerates high throughput screening, especially when using high accuracy (slower) docking approaches and large-sized datasets, and has allowed us to identify several novel potent nonsteroidal SHBG ligands.

Binding Sites↗

A phosphorylation site in the Toll-like receptor 5 TIR domain is required for inflammatory signalling in response to flagellin.

Flagellin, the major structural subunit of bacterial flagella, potently induces inflammatory responses in mammalian cells by activating Toll-like receptor (TLR) 5. Like other TLRs, TLR5 recruits signalling molecules to its intracellular TIR domain, leading to inflammatory responses. Phosphatidylinositol 3-kinase (PI3K) has been reported to play a role in early TLR signalling. We identified a putative binding site for PI3K at tyrosine 798 in the TLR5 TIR domain, at a site analogous to the PI3K recruitment domain in the interleukin-1 receptor. Mutation of this residue did not affect homodimerization, but prevented inflammatory responses to flagellin. While we did not detect direct interaction of PI3K with TLR5, we demonstrated by mass spectrometry that Y798 is phosphorylated in flagellin-treated HEK 293T cells. Together, these results suggest that phosphorylation of Y798 in TLR5 is required for signalling, but not for TLR5 dimerization.

Amino Acid Sequence↗

Distance based algorithms for small biomolecule classification and structural similarity search.

MOTIVATION: Structural similarity search among small molecules is a standard tool used in molecular classification and in-silico drug discovery. The effectiveness of this general approach depends on how well the following problems are addressed. The notion of similarity should be chosen for providing the highest level of discrimination of compounds wrt the bioactivity of interest. The data structure for performing search should be very efficient as the molecular databases of interest include several millions of compounds. RESULTS: In this paper we focus on the k-nearest-neighbor search method, which, until recently was not considered for small molecule classification. The few recent applications of k-nn to compound classification focus on selecting the most relevant set of chemical descriptors which are then compared under standard Minkowski distance L(p). Here we show how to computationally design the optimal weighted Minkowski distance wL(p) for maximizing the discrimination between active and inactive compounds wrt bioactivities of interest. We then show how to construct pruning based k-nn search data structures for any wL(p) distance that minimizes similarity search time. The accuracy achieved by our classifier is better than the alternative LDA and MLR approaches and is comparable to the ANN methods. In terms of running time, our classifier is considerably faster than the ANN approach especially when large data sets are used. Furthermore, our classifier quantifies the level of bioactivity rather than returning a binary decision and thus is more informative than the ANN approach.

Algorithms↗

Large-scale survey for potentially targetable indels in bacterial and protozoan proteins.

Our previous results demonstrated that some essential, housekeeping proteins from pathogenic microorganisms may contain sizable insertions-deletions in their sequences (compared to close human homologs) that can be responsible for unexpected virulence properties. For example, we found that indel-bearing elongation factor-1alpha from several pathogenic protozoa can activate a human tyrosine phosphatase SHP-1 leading to deactivation of macrophages. On the one hand, these findings allowed development of a strategy for targeting some indel-containing pathogen proteins that have similar human counterparts. On the other hand, the results raised numerous questions regarding the nature and implications of sequence indels in pathogen proteins. In the present study, we conducted a large-scale survey of indels in proteins from 136 bacterial and protozoan genomes. It has been established that sizable insertions and deletions occur in approximately 5-10% of bacterial proteins with close human homologs, while proteins from the protozoan pathogens such as Trypanosoma cruzi, Plasmodium falciparum, and Leishmania donovani exhibit elevated indel content that can reach up to 25%. The finding suggested that the occurrence of sequence indels may be involved in the evolution of pathogenic mechanisms in these protozoa.

Animals↗

Selective targeting of indel-inferred differences in spatial structures of homologous proteins.

The eukaryotic pathogen Leishmania donovani possesses a housekeeping protein Elongation-Factor-1alpha (EF-1alpha) which has been found to be unexpectedly involved in the pathogen's virulence. Because it is associated with virulence and essential for cell survival, this protein is an attractive choice for drug targeting; however, its sequence is highly similar (> 80% sequence identity) to that of its human homolog, rendering it a risky choice for a drug target. The chief difference between these two proteins has been found to be a 12 amino acid sequence present in human EF-1alpha but absent from leishmania EF-1alpha. Furthermore, it has been shown that this 12 amino acid insert in the human sequence corresponds to a hairpin loop on the surface of the protein. In this study, we searched for those spatial features in leishmania EF-1alpha that are impacted or obscured by the extra hairpin loop in the human counterpart. We have also conducted a large-scale in silico screening for small molecules that could plausibly bind to these protein features. While experimental evidence is required to verify our results, our findings thus far appear to support this approach as a new strategy for the development of antagonists against pathogenic targets having close human homologs.

Amino Acid Sequence↗

Successful in silico discovery of novel nonsteroidal ligands for human sex hormone binding globulin.

Using "in silico" drug design methodologies, we have discovered several nonsteroidal compounds of natural origin that bind to human sex hormone binding globulin (SHBG) with affinity constants of 0.1 x 10(6) to 1.2 x 10(6) M(-1). The computational solutions we developed involved pharmacophore-aided database search, virtual protein-ligand docking, and structure-activity modeling with "inductive" QSAR descriptors. By screening 23 836 natural substance structures, we identified 29 potential SHBG ligands, and eight of these bound the protein in vitro. These nonsteroidal ligands belong to four classes of molecular scaffolds with several available substitution positions that could allow chemical modification to enhance SHBG-binding activity. Interestingly, one of these compounds is structurally similar to a dicyclohexane derivative that binds to rat SHBG and causes azospermia when administered to male rats. Taken together, the in silico strategy we have developed will aid in the discovery of nonsteroidal ligands of SHBG with novel pharmacological properties.

Algorithms↗

Selective targeting of indel-inferred differences in spatial structures of highly homologous proteins.

Recent findings have shown that the protein elongation factor-1alpha (EF-1alpha) from the eukaryotic pathogen Leishmania donovani possesses virulence properties. This was unexpected, since it has greater than 80% sequence identity with its human homologue. Given that EF-1alpha is essential for cell survival, in principle, it can be considered an attractive drug target. However, the challenge is to be able to selectively target the protein so as not to affect function of the human homologue. While a limited number of discrete differences were scattered throughout the sequence, most of the difference between these 2 homologues could be attributed to a 12-amino acid insert present in human EF-1alpha and absent from the leishmania sequence. In the present study, we modeled the spatial differences in structures of human and L. donovani EF-1alpha's inferred by this insertion-deletion (or "indel"). The protein models were used to develop antibodies directed specifically toward the deletion region of the pathogen protein. The strategy described allowed successful selective targeting of this putative leishmania virulence factor while avoiding recognition of the highly similar human EF-1alpha homologue. These findings may establish a new strategy for the development of antagonists directed against certain pathogenic targets having close human homologues.

Amino Acid Sequence↗

Modeling of cell signaling pathways in macrophages by semantic networks.

BACKGROUND: Substantial amounts of data on cell signaling, metabolic, gene regulatory and other biological pathways have been accumulated in literature and electronic databases. Conventionally, this information is stored in the form of pathway diagrams and can be characterized as highly "compartmental" (i.e. individual pathways are not connected into more general networks). Current approaches for representing pathways are limited in their capacity to model molecular interactions in their spatial and temporal context. Moreover, the critical knowledge of cause-effect relationships among signaling events is not reflected by most conventional approaches for manipulating pathways. RESULTS: We have applied a semantic network (SN) approach to develop and implement a model for cell signaling pathways. The semantic model has mapped biological concepts to a set of semantic agents and relationships, and characterized cell signaling events and their participants in the hierarchical and spatial context. In particular, the available information on the behaviors and interactions of the PI3K enzyme family has been integrated into the SN environment and a cell signaling network in human macrophages has been constructed. A SN-application has been developed to manipulate the locations and the states of molecules and to observe their actions under different biological scenarios. The approach allowed qualitative simulation of cell signaling events involving PI3Ks and identified pathways of molecular interactions that led to known cellular responses as well as other potential responses during bacterial invasions in macrophages. CONCLUSIONS: We concluded from our results that the semantic network is an effective method to model cell signaling pathways. The semantic model allows proper representation and integration of information on biological structures and their interactions at different levels. The reconstruction of the cell signaling network in the macrophage allowed detailed investigation of connections among various essential molecules and reflected the cause-effect relationships among signaling events. The simulation demonstrated the dynamics of the semantic network, where a change of states on a molecule can alter its function and potentially cause a chain-reaction effect in the system.

Computer Simulation↗

Structural characterization of genomes by large scale sequence-structure threading: application of reliability analysis in structural genomics.

BACKGROUND: We establish that the occurrence of protein folds among genomes can be accurately described with a Weibull function. Systems which exhibit Weibull character can be interpreted with reliability theory commonly used in engineering analysis. For instance, Weibull distributions are widely used in reliability, maintainability and safety work to model time-to-failure of mechanical devices, mechanisms, building constructions and equipment. RESULTS: We have found that the Weibull function describes protein fold distribution within and among genomes more accurately than conventional power functions which have been used in a number of structural genomic studies reported to date. It has also been found that the Weibull reliability parameter beta for protein fold distributions varies between genomes and may reflect differences in rates of gene duplication in evolutionary history of organisms. CONCLUSIONS: The results of this work demonstrate that reliability analysis can provide useful insights and testable predictions in the fields of comparative and structural genomics.

Computational Biology↗

An approach to large scale identification of non-obvious structural similarities between proteins.

BACKGROUND: A new sequence independent bioinformatics approach allowing genome-wide search for proteins with similar three dimensional structures has been developed. By utilizing the numerical output of the sequence threading it establishes putative non-obvious structural similarities between proteins. When applied to the testing set of proteins with known three dimensional structures the developed approach was able to recognize structurally similar proteins with high accuracy. RESULTS: The method has been developed to identify pathogenic proteins with low sequence identity and high structural similarity to host analogues. Such protein structure relationships would be hypothesized to arise through convergent evolution or through ancient horizontal gene transfer events, now undetectable using current sequence alignment techniques. The pathogen proteins, which could mimic or interfere with host activities, would represent candidate virulence factors. The developed approach utilizes the numerical outputs from the sequence-structure threading. It identifies the potential structural similarity between a pair of proteins by correlating the threading scores of the corresponding two primary sequences against the library of the standard folds. This approach allowed up to 64% sensitivity and 99.9% specificity in distinguishing protein pairs with high structural similarity. CONCLUSION: Preliminary results obtained by comparison of the genomes of Homo sapiens and several strains of Chlamydia trachomatis have demonstrated the potential usefulness of the method in the identification of bacterial proteins with known or potential roles in virulence.

Bacterial Proteins↗

Structural characterization of genomes by large scale sequence-structure threading.

BACKGROUND: Using sequence-structure threading we have conducted structural characterization of complete proteomes of 37 archaeal, bacterial and eukaryotic organisms (including worm, fly, mouse and human) totaling 167,888 genes. RESULTS: The reported data represent first rather general evaluation of performance of full sequence-structure threading on multiple genomes providing opportunity to evaluate its general applicability for large scale studies. According to the estimated results the sequence-structure threading has assigned protein folds to more then 60% of eukaryotic, 68% of archaeal and 70% of bacterial proteomes.The repertoires of protein classes, architectures, topologies and homologous superfamilies (according to the CATH 2.4 classification) have been established for distant organisms and superkingdoms. It has been found that the average abundance of CATH classes decreases from "alpha and beta" to "mainly beta", followed by "mainly alpha" and "few secondary structures".3-Layer (aba) Sandwich has been characterized as the most abundant protein architecture and Rossman fold as the most common topology. CONCLUSION: The analysis of genomic occurrences of CATH 2.4 protein homologous superfamilies and topologies has revealed the power-law character of their distributions. The corresponding double logarithmic "frequency - genomic occurrence" dependences characteristic of scale-free systems have been established for individual organisms and for three superkingdoms.

Animals↗

Molecular cloning, biochemical and structural analysis of elongation factor-1 alpha from Leishmania donovani: comparison with the mammalian homologue.

The Src-homology 2 domain containing protein tyrosine phosphatase-1 (SHP-1) is involved in the pathogenesis of infection with Leishmania. Recently, we identified elongation factor-1 alpha (EF-1 alpha) from Leishmania donovani as a SHP-1 binding and activating protein [J. Biol. Chem. 277 (2002) 50190]. To characterize this apparent Leishmania virulence factor further, the cDNA encoding L. donovani EF-1 alpha was cloned and sequenced. Whereas nearly complete sequence conservation was observed amongst EF-1 alpha proteins from trypanosomatids, the deduced amino acid sequence of EF-1 alpha of L. donovani when compared to mammalian EF-1 alpha sequences showed a number of significant changes. Protein structure modeling-based upon the known crystal structure of EF-1 alpha for Saccharomyces cerevisiae-identified a hairpin loop present in mammalian EF-1 alpha and absent from the Leishmania protein which corresponded to a 12 amino acid deletion. Consistent with these structural differences, the sub-cellular distributions of L. donovani EF-1 alpha and host EF-1 alpha were strikingly different. Interestingly, infection of macrophages with L. donovani caused redistribution of host as well as pathogen EF-1 alpha. Since EF-1 alpha is essential for survival, the distinct biochemical and structural properties of Leishmania EF-1 alpha may provide a novel target for drug development.

Amino Acid Sequence↗

Molecular analysis of the multiple GroEL proteins of Chlamydiae.

Genome sequencing revealed that all six chlamydiae genomes contain three groEL-like genes (groEL1, groEL2, and groEL3). Phylogenetic analysis of groEL1, groEL2, and groEL3 indicates that these genes are likely to have been present in chlamydiae since the beginning of the lineage. Comparison of deduced amino acid sequences of the three groEL genes with those of other organisms showed high homology only for groEL1, although comparison of critical amino acid residues that are required for polypeptide binding of the Escherichia coli chaperonin GroEL revealed substantial conservation in all three chlamydial GroELs. This was further supported by three-dimensional structural predictions. All three genes are expressed constitutively throughout the developmental cycle of Chlamydia trachomatis, although groEL1 is expressed at much higher levels than are groEL2 and groEL3. Transcription of groEL1, but not groEL2 and groEL3, was elevated when HeLa cells infected with C. trachomatis were subjected to heat shock. Western blot analysis with polyclonal antibodies raised against recombinant GroEL1, GroEL2, and GroEL3 demonstrated the presence of the three proteins in C. trachomatis elementary bodies, with GroEL1 being present in the largest amount. Only C. trachomatis groEL1 and groES together complemented a temperature-sensitive E. coli groEL mutant. Complementation did not occur with groEL2 or groEL3 alone or together with groES. The role for each of the three GroELs in the chlamydial developmental cycle and in disease pathogenesis requires further study.

Amino Acid Sequence↗

Evidence that plant-like genes in Chlamydia species reflect an ancestral relationship between Chlamydiaceae, cyanobacteria, and the chloroplast.

An unusually high proportion of proteins encoded in Chlamydia genomes are most similar to plant proteins, leading to proposals that a Chlamydia ancestor obtained genes from a plant or plant-like host organism by horizontal gene transfer. However, during an analysis of bacterial-eukaryotic protein similarities, we found that the vast majority of plant-like sequences in Chlamydia are most similar to plant proteins that are targeted to the chloroplast, an organelle derived from a cyanobacterium. We present further evidence suggesting that plant-like genes in Chlamydia, and other Chlamydiaceae, are likely a reflection of an unappreciated evolutionary relationship between the Chlamydiaceae and the cyanobacteria-chloroplast lineage. Further analyses of bacterial and eukaryotic genomes indicates the importance of evaluating organellar ancestry of eukaryotic proteins when identifying bacteria-eukaryote homologs or horizontal gene transfer and supports the proposal that Chlamydiaceae, which are obligate intracellular bacterial pathogens of animals, are not likely exchanging DNA with their hosts.

Animals↗

Inductive electronegativity scale. Iterative calculation of inductive partial charges.

A number of novel QSAR descriptors have been introduced on the basis of the previously elaborated models for steric and inductive effects. The developed "inductive" parameters include absolute and effective electronegativity, atomic partial charges, and local and global chemical hardness and softness. Being based on traditional inductive and steric substituent constants these 3D descriptors provide a valuable insight into intramolecular steric and electronic interactions and can find broad application in structure-activity studies. Possible interpretation of physical meaning of the inductive descriptors has been suggested by considering a neutral molecule as an electrical capacitor formed by charged atomic spheres. This approximation relates inductive chemical softness and hardness of bound atom(s) with the total area of the facings of electrical capacitor formed by the atom(s) and the rest of the molecule. The derived full electronegativity equalization scheme allows iterative calculation of inductive partial charges on the basis of atomic electronegativities, covalent radii, and intramolecular distances. A range of inductive descriptors has been computed for a variety of organic compounds. The calculated inductive charges in the studied molecules have been validated by experimental C-1s Electron Core Binding Energies and molecular dipole moments. Several semiempirical chemical rules, such as equalized electronegativity's arithmetic mean, principle of maximum hardness, and principle of hardness borrowing could be explicitly illustrated in the framework of the developed approach.

Journal Article↗

'Inductive' charges on atoms in proteins: comparative docking with the extended steroid benchmark set and discovery of a novel SHBG ligand.

We have developed a novel iterative approach for calculation of partial charges in proteins within the framework of the 'molecular capacitance' model. The method operates by an effective 'inductive' electronegativity scale derived from a number of the conventional charge systems including CHARMM, AMBER, MMFF, OPLS, and PEOE among others. Our novel 'inductive' electronegativity equalization procedure allows rapid and conformation sensitive computation of adequate partial charges in proteins. Accuracy of the 'inductive' values was confirmed by their correlation with DFT-computed partial charges in common amino acids. A comparative docking study with an extended steroid data set not only illustrated the adequacy of 'inductive' protein charges but also demonstrated their superior performance compared to several conventional protein charging systems. Subsequent docking with 'inductive' charges resulted in identification of five potential leads as human Sex Hormone Binding Globulin (SHBG) ligands from a commercial library of natural compounds. When the selected substances were evaluated for their ability to bind SHBG in vitro, three of them displaced testosterone from the SHBG steroid-binding site, and with one compound this was achieved at micromolar concentrations.

Algorithms↗

Can 'Bacterial-Metabolite-Likeness' model improve odds of 'in silico' antibiotic discovery?

'Inductive' QSAR descriptors have been used to develop the series of QSAR models enabling 'in silico' distinguishing between antimicrobial compounds, conventional drugs, and druglike substances. The constructed neural network-based models operating by 30 'inductive' parameters have been validated on an extensive set of 2686 chemical structures and resulted in up to 97% accurate separation of the three types of molecular activities. The demonstrated ability of 'inductive' parameters to adequately capture molecular features determining 'antibiotic-like' and 'druglike' potentials have been further utilized to construct a model of 'Bacterial-Metabolite-Likeness' (BML). The same 'inductive' descriptors have been used to train a neural network that could very accurately recognize substances involved into bacterial metabolism (that have been experimentally identified). When the developed model has been applied to the mixed set of antimicrobials, drugs, and druglike chemicals (not used for training the BML model), it exhibited a 2-5-fold recognition preference toward antimicrobial compounds compared to general drugs and an 18- to 45-fold preference when compared to a druglike substance (depending on the model stringency). These results illustrate immanent similarity between conventional antimicrobials and native bacterial metabolites and suggest that the developed BML model can be an effective classification tool for 'in silico' antibiotic studies.

Anti-Bacterial Agents↗