PubMed Health⌕ Search

Biomedical subjects

Olivier Lichtarge

Publications and source records attributed to Olivier Lichtarge.

At least 19 recordsLinked to original sources

Meta-evolutionary exome analysis identifies novel type 2 diabetes mellitus genes in the UK Biobank and all of us.

Type 2 diabetes mellitus (T2DM) risk is heavily influenced by genetics, yet current association tests have explained only parts of its heritability. We developed MEVA (Meta-Evolutionary Action), a meta-analytic framework that integrates three complementary methods-EAML, Sigma-Diff, and GeneEMBED-to assess the functional burden of protein-coding variants using evolutionary data. MEVA was applied to exome data from 28,115 T2DM cases and 28,115 controls in the UK Biobank (UKB), identifying 101 genes (p&#x2009;<&#x2009;1e-5). MEVA outperformed its component methods, each of which substantially outperformed a conventional burden test (MAGMA), in recovering known T2DM genes (AUROC&#x2009;=&#x2009;0.925) and maintaining robustness in progressively smaller cohorts (AUROC&#x2009;=&#x2009;0.917). MEVA showed significant enrichment for T2DM-related loci (p&#x2009;=&#x2009;6.8e-10, p&#x2009;=&#x2009;2.0e-34), protein interactions (z&#x2009;=&#x2009;4.6, z&#x2009;=&#x2009;4.2), pathways (p&#x2009;=&#x2009;1.3e-6, z&#x2009;=&#x2009;2.0), phenotypes (p&#x2009;=&#x2009;1.3e-21, z&#x2009;=&#x2009;9.1), and literature mentions (z&#x2009;=&#x2009;7.2). Replication in 16,915 T2DM cases and 16,915 controls from All of Us (AoU) yielded 99 genes (p&#x2009;<&#x2009;1e-5), 23 of which were also recovered in the UKB cohort - far exceeding random chance. These included established genes (SLC30A8, WFS1, HNF1A) and less-characterized candidates (NRIP1, ADAM30, CALCOCO2, TUBB1, ZFP36L2, WDR90). Notably, NRIP1 loss-of-function variants were associated with increased T2DM risk in both the UKB (OR = 1.09, FDR&#x2009;=&#x2009;5.4e-4) and AoU (OR = 1.09, FDR&#x2009;=&#x2009;0.046), and TUBB1 and CALCOCO2 gain-of-function variants showed consistent risk effects (FDR&#x2009;<&#x2009;0.05). Pathway analyses revealed convergence on endoplasmic reticulum chaperone complexes (FDR&#x2009;=&#x2009;0.02) and Hippo signaling (FDR&#x2009;=&#x2009;8.5e-4). Finally, all 177 candidate genes were functionally prioritized using ten orthogonal criteria to guide experimental follow-up. These results demonstrate that combining complementary, impact-aware association tests increases sensitivity, improves replication, and expands the catalog of genetic risk factors for T2DM.

Humans↗

Essential helix interactions in the anion transporter domain of prestin revealed by evolutionary trace analysis.

Prestin, a member of the SLC26A family of anion transporters, is a polytopic membrane protein found in outer hair cells (OHCs) of the mammalian cochlea. Prestin is an essential component of the membrane-based motor that enhances electromotility of OHCs and contributes to frequency sensitivity and selectivity in mammalian hearing. Mammalian cells expressing prestin display a nonlinear capacitance (NLC), widely accepted as the electrical signature of electromotility. The associated charge movement requires intracellular anions reflecting the membership of prestin in the SLC26A family. We used the computational approach of evolutionary trace analysis to identify candidate functional (trace) residues in prestin for mutational studies. We created a panel of mutations at each trace residue and determined membrane expression and nonlinear capacitance associated with each mutant. We observe that several residue substitutions near the conserved sulfate transporter domain of prestin either greatly reduce or eliminate NLC, and the effect is dependent on the size of the substituted residue. These data suggest that packing of helices and interactions between residues surrounding the "sulfate transporter motif" is essential for normal prestin activity.

Amino Acid Sequence↗

Rapid detection of similarity in protein structure and function through contact metric distances.

The characterization of biological function among newly determined protein structures is a central challenge in structural genomics. One class of computational solutions to this problem is based on the similarity of protein structure. Here, we implement a simple yet efficient measure of protein structure similarity, the contact metric. Even though its computation avoids structural alignments and is therefore nearly instantaneous, we find that small values correlate with geometrical root mean square deviations obtained from structural alignments. To test whether the contact metric detects functional similarity, as defined by Gene Ontology (GO) terms, it was compared in large-scale computational experiments to four other measures of structural similarity, including alignment algorithms as well as alignment independent approaches. The contact metric was the fastest method and its sensitivity, at any given specificity level, was a close second only to Fast Alignment and Search Tool--a structural alignment method that is slower by three orders of magnitude. Critically, nearly 40% of correct functional inferences by the contact metric were not identified by any other approach, which shows that the contact metric is complementary and computationally efficient in detecting functional relationships between proteins. A public 'Contact Metric Internet Server' is provided.

Algorithms↗

Rank information: a structure-independent measure of evolutionary trace quality that improves identification of protein functional sites.

Protein functional sites are key targets for drug design and protein engineering, but their large-scale experimental characterization remains difficult. The evolutionary trace (ET) is a computational approach to this problem that has been useful in a variety of case studies, but its proteomic scale application is partially hindered because automated retrieval of input sequences from databases often includes some with errors that degrade functional site identification. To recognize and purge these sequences, this study introduces a novel and structure-free measure of ET quality called rank information (RI). It is shown that RI decreases in response to errors in sequences, alignments, or functional classifications. Conversely, an automated procedure to increase RI by selectively removing sequences improves functional site identification so as to nearly match manually curated traces in kinases and in a test set of 79 diverse proteins. Thus we conclude that RI partially reflects the evolutionary consistency of sequence, structure, and function. In practice, as the size of the proteome continues to grow exponentially, it provides a novel and structure-free measure of ET quality that increases its accuracy for large-scale automated annotation of protein functional sites.

Algorithms↗

Evolutionary identification of a subtype specific functional site in the ligand binding domain of steroid receptors.

Nuclear receptors are ubiquitous eukaryotic ligand-activated transcription factors that modulate gene expression through varied interactions. However, the highly conserved functional sites known today seem insufficient to explain receptor specific recruitment of different coactivator and corepressor proteins and regulation of transcription. To search for new receptor-subtype specific functional sites, we applied difference evolutionary trace (difference ET) analysis to the ligand binding domain of steroid receptors, a subgroup of the nuclear receptor (NR) family. This computational approach identified a new functional site located on a surface opposite to currently known protein-protein interaction sites and distinct from the ligand binding pocket. Strikingly, the literature shows that in vivo variations at residues in the new site are linked to androgen resistance and leukemia, and our own targeted mutations to this site lower but do not eradicate transcriptional activation by estrogen receptor alpha (ERalpha), with reduced ligand binding affinity and SRC-1 interaction. Thus, these data demonstrate that this evolutionary important surface can function as an allosteric site that modulates some but not all receptor binding interactions. Evolutionary analysis further shows that this allosteric regulatory site is shared among all NRs from groups 2 (HNF4-like) and 4 (NGFIB-like), suggesting a role among many nuclear receptors. Its concave structure, hydrophobic composition, and residue variability among nuclear receptors further suggest that it would be amenable for specific drug design. This highlights the power of evolutionary information for the identification of new functional sites even in a protein family as well studied as NRs.

Allosteric Site↗

ET viewer: an application for predicting and visualizing functional sites in protein structures.

SUMMARY: The Evolutionary Trace Viewer (ETV) provides a one-stop environment in which to run, visualize and interpret Evolutionary Trace (ET) predictions of functional sites in protein structures. ETV is implemented using Java to run across different operating systems using Java Web Start technology. AVAILABILITY: The ETV is available for download from our website at http://mammoth.bcm.tmc.edu/traceview/index.html. This webpage also links to sample trace results and a user manual that describes ET Viewer functions in detail.

Amino Acid Sequence↗

Recurrent use of evolutionary importance for functional annotation of proteins based on local structural similarity.

The annotation of protein function has not kept pace with the exponential growth of raw sequence and structure data. An emerging solution to this problem is to identify 3D motifs or templates in protein structures that are necessary and sufficient determinants of function. Here, we demonstrate the recurrent use of evolutionary trace information to construct such 3D templates for enzymes, search for them in other structures, and distinguish true from spurious matches. Serine protease templates built from evolutionarily important residues distinguish between proteases and other proteins nearly as well as the classic Ser-His-Asp catalytic triad. In 53 enzymes spanning 33 distinct functions, an automated pipeline identifies functionally related proteins with an average positive predictive power of 62%, including correct matches to proteins with the same function but with low sequence identity (the average identity for some templates is only 17%). Although these template building, searching, and match classification strategies are not yet optimized, their sequential implementation demonstrates a functional annotation pipeline which does not require experimental information, but only local molecular mimicry among a small number of evolutionarily important residues.

Algorithms↗

Role of transmembrane domain/transmembrane domain interfaces of P-glycoprotein (ABCB1) in solute transport. Convergent information from photoaffinity labeling, site directed mutagenesis and in silico importance prediction.

Human P-glycoprotein (P-gp, ABCB1) plays an important role in the development of resistance to anticancer therapy. This ABC-transporter (ATP-binding cassette transporter) intercepts drugs at the level of the plasma membrane and effluxes them before they are able to reach their intracellular target structures. Inhibition of P-gp by low molecular weight compounds has been advocated as a concept for resensitization of cells to anticancer agents and several clinical studies in oncological patients have advanced to phase III. Even more importantly, P-glycoprotein also represents an antitarget. Its expression in cells lining the intestinal tract, the canalicular side of hepatocytes, renal tubuli and the blood brain barrier lead to interference with pharmacokinetics of compounds that are recognized as pump substrates. An early prediction of ADMET (Absorption-Distribution-Metabolism-Excretion-Toxicity) properties is important during drug development, since interference of a compound with P-gp might compromise its future development into a drug. Despite considerable efforts, the mechanism by which P-gp binds and transports its solutes remains unclear. Generation of homology models of the protein allowed integration of data obtained by photoaffinity labeling, in silico prediction of functional importance by evolutionary tracing and site directed mutagenesis. An integral view of data indicates that these three lines of evidence converge to indicate two pseudosymmetric P-gp drug binding pockets located at the two transmembrane domain interfaces.

ATP Binding Cassette Transporter, Subfamily B, Mem↗

beta-arrestin-dependent, G protein-independent ERK1/2 activation by the beta2 adrenergic receptor.

Physiological effects of beta adrenergic receptor (beta2AR) stimulation have been classically shown to result from G(s)-dependent adenylyl cyclase activation. Here we demonstrate a novel signaling mechanism wherein beta-arrestins mediate beta2AR signaling to extracellular-signal regulated kinases 1/2 (ERK 1/2) independent of G protein activation. Activation of ERK1/2 by the beta2AR expressed in HEK-293 cells was resolved into two components dependent, respectively, on G(s)-G(i)/protein kinase A (PKA) or beta-arrestins. G protein-dependent activity was rapid, peaking within 2-5 min, was quite transient, was blocked by pertussis toxin (G(i) inhibitor) and H-89 (PKA inhibitor), and was insensitive to depletion of endogenous beta-arrestins by siRNA. beta-Arrestin-dependent activation was slower in onset (peak 5-10 min), less robust, but more sustained and showed little decrement over 30 min. It was insensitive to pertussis toxin and H-89 and sensitive to depletion of either beta-arrestin1 or -2 by small interfering RNA. In G(s) knock-out mouse embryonic fibroblasts, wild-type beta2AR recruited beta-arrestin2-green fluorescent protein and activated pertussis toxin-insensitive ERK1/2. Furthermore, a novel beta2AR mutant (beta2AR(T68F,Y132G,Y219A) or beta2AR(TYY)), rationally designed based on Evolutionary Trace analysis, was incapable of G protein activation but could recruit beta-arrestins, undergo beta-arrestin-dependent internalization, and activate beta-arrestin-dependent ERK. Interestingly, overexpression of GRK5 or -6 increased mutant receptor phosphorylation and beta-arrestin recruitment, led to the formation of stable receptor-beta-arrestin complexes on endosomes, and increased agonist-stimulated phospho-ERK1/2. In contrast, GRK2, membrane translocation of which requires Gbetagamma release upon G protein activation, was ineffective unless it was constitutively targeted to the plasma membrane by a prenylation signal (CAAX). These findings demonstrate that the beta2AR can signal to ERK via a GRK5/6-beta-arrestin-dependent pathway, which is independent of G protein coupling.

Amino Acid Sequence↗

Correlated evolutionary pressure at interacting transcription factors and DNA response elements can guide the rational engineering of DNA binding specificity.

Understanding the molecular mechanisms of the specific interaction between transcription factor proteins and DNA is key to comprehend the regulation of gene expression and to develop technologies to engineer transcription factors. Thus far, although there have been several attempts to elucidate protein-DNA interaction through amino acid-base recognition codes, sequence based profiles, or physical models of interaction, the greatest successes in engineering DNA binding specificity remain experimental. Here we present the first systematic evidence of correlated evolutionary pressure at interacting amino acid residues and DNA base-pairs in transcription factors, and show that it can be used to rationally engineer DNA binding specificity. The correlation is between the relative evolutionary importance of protein residues and DNA bases, measured, respectively, in terms of the Evolutionary Trace (ET) rank and information entropy. The evolutionarily most important residues interact with the most conserved base-pairs within the response element while residues of least importance interact with the most variable base-pairs. The correlation averages 0.74 over 12 unrelated families of transcriptional regulators, including nuclear hormone receptors, basic helix-loop-helix, ETS- and homeo-domain family. To test the predictive power of this correlation, we targeted a mutational swap of top-ranked ET residues in a transcription factor, LRH-1. This redirects LRH-1 binding as predicted and showed that, in this case, evolutionary importance and binding specificity are coupled sufficiently strongly for the Evolutionary Trace to guide the computational design of DNA binding specificity. This establishes the existence of evolutionary importance correlation at protein-DNA interfaces, and demonstrates that it is a useful principle for the rational engineering of binding specificity.

Animals↗

Evolutionary trace-based peptides identify a novel asymmetric interaction that mediates oligomerization in nuclear receptors.

Germ cell nuclear factor (GCNF) is an orphan nuclear receptor that plays important roles in development and reproduction, by repressing the expression of essential genes such as Oct4, GDF9, and BMP15, through binding to DR0 elements. Surprisingly, whereas recombinant GCNF binds to DR0 sequences as a homodimer, endogenous GCNF does not exist as a homodimer but rather as part of a large complex termed the transiently retinoid-induced factor (TRIF). Here, we use evolutionary trace (ET) analysis to design mutations and peptides that probe the molecular basis for the formation of this unusual complex. We find that GCNF homodimerization and TRIF complex formation are DNA-dependent, and ET suggests that dimerization involves key functional sites on both helix 3 and helix 11, which are located on opposing surfaces of the ligand binding domain. Targeted mutations in either helix of GCNF disrupt the formation of both the homodimer and the endogenous TRIF complex. Moreover, peptide mimetics of both of these ET-determined sites inhibit dimerization and TRIF complex formation. This suggests that a novel helix 3-helix 11 heterotypic interaction mediates GCNF interaction and would facilitate oligomerization. Indeed, it was determined that the endogenous TRIF complex is composed of a GCNF oligomer. These findings shed light on an evolutionarily selected mechanism that reveals the unusual DNA-binding, dimerization, and oligomerization properties of GCNF.

Adaptor Proteins, Vesicular Transport↗

Character and evolution of protein-protein interfaces.

Protein-protein interactions create the macromolecular assemblies and sequential signaling pathways essential for cell function. Their number far exceeds the number of proteins themselves and their experimental characterization, while improving, remains relatively slow. For these reasons, novel computational methods have important roles to play in understanding the physical basis of protein interactions, and in constraining the molecular basis of their specificity. This paper discusses methods based on multiple sequence alignments of protein homologues and phylogenetic trees.

Artificial Intelligence↗

Algorithms for structural comparison and statistical analysis of 3D protein motifs.

The comparison of structural subsites in proteins is increasingly relevant to the prediction of their biological function. To address this problem, we present the Match Augmentation algorithm (MA). Given a structural motif of interest, such as a functional site, MA searches a target protein structure for a match: the set of atoms with the greatest geometric and chemical similarity. MA is extremely efficient because it exploits the fact that the amino acids in a structural motif are not equally important to function. Using motif residues ranked on functional significance via the Evolutionary Trace (ET), MA prioritizes its search by initially forming matches with functionally significant residues, then, guided by ET, it augments this partial match stepwise until the whole motif is found. With this hierarchical strategy, MA runs considerably faster than other methods, and almost always identifies matches in homologs known to have cognate functional sites. Second, in order to interpret matches, we further introduce a statistical method using nonparametric density estimation of the frequency distribution of structural matches. Our results show that the hierarchy of functional importance within structural motifs speeds up the search within targets, and points to a new method to score their statistical significance.

Amino Acid Sequence↗

Computational and biochemical identification of a nuclear pore complex binding site on the nuclear transport carrier NTF2.

Nuclear transport carriers interact with proteins of the nuclear pore complex (NPC) to transport their cargo across the nuclear envelope. One such carrier is nuclear transport factor 2 (NTF2), whose import cargo is the small GTPase Ran. A domain highly homologous to the small NTF2 protein (14kDa) is also found in a number of additional proteins, which together make up the NTF2 domain containing superfamily of proteins. Using structural, computational and biochemical analysis we have identified a functional site that is present throughout this superfamily, and our results indicate that this site functions as an NPC binding site in NTF2. Previously we showed that a D23A mutant of NTF2 exhibits increased affinity for the NPC. The mechanism of this mutation, however, was unknown as this region of NTF2 had not been implicated in binding to NPC proteins. Here we show that the D23A mutation in NTF2 does not result in gross structural changes affecting other known NPC binding sites. Instead, the D23 residue is located in an evolutionarily important region in the NTF2 domain containing superfamily, that in NTF2, is involved in binding to the NPC.

Active Transport, Cell Nucleus↗

Evolution of neural precursor selection: functional divergence of proneural proteins.

How conserved pathways are differentially regulated to produce diverse outcomes is a fundamental question of developmental and evolutionary biology. The conserved process of neural precursor cell (NPC) selection by basic helix-loop-helix (bHLH) proneural transcription factors in the peripheral nervous system (PNS) by atonal related proteins (ARPs) presents an excellent model in which to address this issue. Proneural ARPs belong to two highly related groups: the ATONAL (ATO) group and the NEUROGENIN (NGN) group. We used a cross-species approach to demonstrate that the genetic and molecular mechanisms by which ATO proteins and NGN proteins select NPCs are different. Specifically, ATO group genes efficiently induce neurogenesis in Drosophila but very weakly in Xenopus, while the reverse is true for NGN group proteins. This divergence in proneural activity is encoded by three residues in the basic domain of ATO proteins. In NGN proteins, proneural capacity is encoded by the equivalent three residues in the basic domain and a novel motif in the second Helix (H2) domain. Differential interactions with different types of zinc (Zn)-finger proteins mediate the divergence of ATO and NGN activities: Senseless is required for ATO group activity, whereas MyT1 is required for NGN group function. These data suggest an evolutionary divergence in the mechanisms of NPC selection between protostomes and deuterostomes.

Animals↗

Evolutionary trace of G protein-coupled receptors reveals clusters of residues that determine global and class-specific functions.

G protein-coupled receptor (GPCR) activation mediated by ligand-induced structural reorganization of its helices is poorly understood. To determine the universal elements of this conformational switch, we used evolutionary tracing (ET) to identify residue positions commonly important in diverse GPCRs. When mapped onto the rhodopsin structure, these trace residues cluster into a network of contacts from the retinal binding site to the G protein-coupling loops. Their roles in a generic transduction mechanism were verified by 211 of 239 published mutations that caused functional defects. When grouped according to the nature of the defects, these residues sub-divided into three striking sub-clusters: a trigger region, where mutations mostly affect ligand binding, a coupling region near the cytoplasmic interface to the G protein, where mutations affect G protein activation, and a linking core in between where mutations cause constitutive activity and other defects. Differential ET analysis of the opsin family revealed an additional set of opsin-specific residues, several of which form part of the retinal binding pocket, and are known to cause functional defects upon mutation. To test the predictive power of ET, we introduced novel mutations in bovine rhodopsin at a globally important position, Leu-79, and at an opsin-specific position, Trp-175. Both were functionally critical, causing constitutive G protein activation of the mutants and rapid loss of regeneration after photobleaching. These results define in GPCRs a canonical signal transduction mechanism where ligand binding induces conformational changes propagated through adjacent trigger, linking core, and coupling regions.

Amino Acid Sequence↗

An accurate, sensitive, and scalable method to identify functional sites in protein structures.

Functional sites determine the activity and interactions of proteins and as such constitute the targets of most drugs. However, the exponential growth of sequence and structure data far exceeds the ability of experimental techniques to identify their locations and key amino acids. To fill this gap we developed a computational Evolutionary Trace method that ranks the evolutionary importance of amino acids in protein sequences. Studies show that the best-ranked residues form fewer and larger structural clusters than expected by chance and overlap with functional sites, but until now the significance of this overlap has remained qualitative. Here, we use 86 diverse protein structures, including 20 determined by the structural genomics initiative, to show that this overlap is a recurrent and statistically significant feature. An automated ET correctly identifies seven of ten functional sites by the least favorable statistical measure, and nine of ten by the most favorable one. These results quantitatively demonstrate that a large fraction of functional sites in the proteome may be accurately identified from sequence and structure. This should help focus structure-function studies, rational drug design, protein engineering, and functional annotation to the relevant regions of a protein.

Amino Acid Motifs↗

Conserved motifs in somatostatin, D2-dopamine, and alpha 2B-adrenergic receptors for inhibiting the Na-H exchanger, NHE1.

Receptor subtypes within families of G protein-coupled receptors that are activated by similar ligands can regulate distinct intracellular effectors. We identified conserved motifs within intracellular domains 2 and 3 of selective subtypes of several G protein-coupled receptor families that confer coupling to the Na-H exchanger, NHE1. A T(s,p)V motif within intracellular domain 2 and a QQ(r) motif within intracellular domain 3 are shared by the somatostatin receptor subtypes SSTR1, -3, and -4, which couple to the inhibition of NHE1, but not by SSTR2 and -5, which do not signal to NHE1. Only the collective substitution of cognate SSTR2 residues with these two motifs conferred the ability of mutant SSTR2 to inhibit NHE1. Both motifs are present in D(2)-dopamine receptors, which inhibit NHE1, and in alpha(2B)-adrenergic receptors, which couple to the inhibition of NHE1, but not in alpha(2A)-adrenergic receptors, which do not regulate NHE1. These findings indicate that motifs shared by different subfamilies of G protein-coupled receptors, but not necessarily by receptor subtypes within a subfamily, can confer coupling to a common effector.

Amino Acid Motifs↗