PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Intrinsically Disordered Proteins”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

DescribePROT Database of Residue-Level Protein Structure and Function Annotations.

DescribePROT is a freely available online database of structural and functional descriptors of proteins at the amino acid level. It provides access to 13 diverse descriptors that include sequence conservation, putative secondary structure, solvent accessibility, intrinsic disorder, and signal peptides, and putative annotations of residues that interact with proteins, peptides and nucleic acids. These data can be used to elucidate protein functions, to support efforts to develop therapeutics, and to develop and evaluate future predictors of protein structure and function. DescribePROT includes 7.8 billion predictions for 1.4 million proteins from 83 complete proteomes of popular model organisms. This information can be downloaded at multiple levels of scope (entire database, specific organisms, and individual proteins) and can be interacted with using a graphical interface that simultaneously displays data on multiple descriptors. We describe the contents of this resource, provide directions on how to use its interface, and offer instructions on how to obtain and interact with the underlying data. Moreover, we briefly discuss plans for a future expansion of this database. DescribePROT is available at http://biomine.cs.vcu.edu/servers/DESCRIBEPROT/ .

Databases, Protein↗

Primary contact sites in intrinsically unstructured proteins: the case of calpastatin and microtubule-associated protein 2.

Intrinsically unstructured proteins (IUPs) exist in a disordered conformational state, often considered to be equivalent with the random-coil structure. We challenge this simplifying view by limited proteolysis, circular dichroism (CD) spectroscopy, and solid-state (1)H NMR, to show short- and long-range structural organization in two IUPs, the first inhibitory domain of calpastatin (CSD1) and microtubule-associated protein 2c (MAP2c). Proteases of either narrow (trypsin, chymotrypsin, and plasmin) or broad (subtilisin and proteinase K) substrate specificity, applied at very low concentrations, preferentially cleaved both proteins in regions, i.e., subdomains A, B, and C in CSD1 and the proline-rich region (PRR) in MAP2c, that are destined to form contacts with their targets. For CSD1, nonadditivity of the CD spectra of its two halves and suboptimal hydration of the full-length protein measured by solid-state NMR demonstrate that long-range tertiary interactions provide the structural background of this structural feature. In MAP2c, such tertiary interactions are absent, which points to the importance of local structural constraints. In fact, urea and temperature dependence of the CD spectrum of its PRR reveals the presence of the extended and rather stiff polyproline II helix conformation that keeps the interaction site exposed. These data suggest that functionally significant residual structure exists in both of these IUPs. This structure, manifest as either transient local and/or global organization, ensures the spatial exposure of short contact segments on the surface. Pertinent data from other IUPs suggest that the presence of such recognition motifs may be a general feature of disordered proteins. To emphasize the possible importance of this structural trait, we propose that these motifs be called primary contact sites in IUPs.

Amino Acid Sequence↗

NMR relaxation studies on the hydrate layer of intrinsically unstructured proteins.

Intrinsically unstructured/disordered proteins (IUPs) exist in a disordered and largely solvent-exposed, still functional, structural state under physiological conditions. As their function is often directly linked with structural disorder, understanding their structure-function relationship in detail is a great challenge to structural biology. In particular, their hydration and residual structure, both closely linked with their mechanism of action, require close attention. Here we demonstrate that the hydration of IUPs can be adequately approached by a technique so far unexplored with respect to IUPs, solid-state NMR relaxation measurements. This technique provides quantitative information on various features of hydrate water bound to these proteins. By freezing nonhydrate (bulk) water out, we have been able to measure free induction decays pertaining to protons of bound water from which the amount of hydrate water, its activation energy, and correlation times could be calculated. Thus, for three IUPs, the first inhibitory domain of calpastatin, microtubule-associated protein 2c, and plant dehydrin early responsive to dehydration 10, we demonstrate that they bind a significantly larger amount of water than globular proteins, whereas their suboptimal hydration and relaxation parameters are correlated with their differing modes of function. The theoretical treatment and experimental approach presented in this article may have general utility in characterizing proteins that belong to this novel structural class.

Arabidopsis Proteins↗

Intrinsic proteins and their effect upon lipid hydrocarbon chain order.

We present evidence that at temperatures greater than their main transition temperature, phospholipid molecules that are trapped within clusters of intrinsic molecules such as polypeptides or proteins have the ends of their hydrocarbon chains more statically disordered than those of lipid molecules far from such intrinsic molecules. We have constructed a model in which the lipids are divided into three populations: (i) those that are not adjacent to any protein ("free" lipids), (ii) those that are adjacent to only one protein ("adjacent" lipids), and (iii) those that are "trapped" between two or three proteins. We applied this model to study deuterium nuclear magnetic resonance of dimyristoyl-3-sn-phosphatidylcholine (DMPC) bilayers containing gramicidin A' or cytochrome oxidase and found that while the methyl groups of adjacent lipids are slightly more statically ordered than those of free lipids, the methyl groups of trapped lipids are more statically disordered than those of free lipids. We propose a physical explanation for this and show that phosphorus-31 nuclear magnetic resonance data for DMPC-cytochrome oxidase bilayers can be understood as a consequence of changes in the polar region of only trapped lipids.

Dimyristoylphosphatidylcholine↗

SNAREs prefer liquid-disordered over "raft" (liquid-ordered) domains when reconstituted into giant unilamellar vesicles.

Membrane domains ("rafts") have received great attention as potential platforms for proteins in signaling and trafficking. Because rafts are believed to form by cooperative lipid interactions but are not directly accessible in vivo, artificial phase-separating lipid bilayers are useful model systems. Giant unilamellar vesicles (GUVs) offer large free-standing bilayers, but suitable methods for incorporating proteins are still scarce. Here we report the reconstitution of two water-insoluble SNARE proteins into GUVs without fusogenic additives. Following reconstitution, protein functionality was assayed by confocal imaging and fluorescence auto- and cross-correlation spectroscopy. Incorporation into GUVs containing phase-separating lipids revealed that, in the absence of other cellular factors, both proteins exhibit an intrinsic preference for the liquid-disordered phase. Although the picture from detergent resistance assays on whole cells is ambiguous, reconstitutions of components of the exocytic machinery into GUVs by this new approach should yield insight into the dynamics of protein complex associations with hypothesized liquid-ordered phase microdomains, the correspondence between detergent-resistant membranes and liquid-ordered phase, and the mechanism of SNARE-mediated membrane fusion.

Exocytosis↗

Intrinsically disordered C-terminal segments of voltage-activated potassium channels: a possible fishing rod-like mechanism for channel binding to scaffold proteins.

Membrane-embedded voltage-activated potassium channels (Kv) bind intracellular scaffold proteins, such as the Post Synaptic Density 95 (PSD-95) protein, using a conserved PDZ-binding motif located at the channels' C-terminal tip. This interaction underlies Kv-channel clustering, and is important for the proper assembly and functioning of the synapse. Here we demonstrate that the C-terminal segments of Kv channels adjacent to the PDZ-binding motif are intrinsically disordered. Phylogenetic analysis of the Kv channel family reveals a cluster of channel sequences belonging to three out of the four main channel families, for which an association is demonstrated between the presence of the consensus terminal PDZ-binding motif and the intrinsically disordered nature of the immediately adjacent C-terminal segment. Our observations, combined with a structural analogy to the N-terminal intra-molecular ball-and-chain mechanism for Kv channel inactivation, suggest that the C-terminal disordered segments of these channel families encode an inter-molecular fishing rod-like mechanism for K(+) channel binding to scaffold proteins.

Amino Acid Motifs↗

The pairwise energy content estimated from amino acid composition discriminates between folded and intrinsically unstructured proteins.

The structural stability of a protein requires a large number of interresidue interactions. The energetic contribution of these can be approximated by low-resolution force fields extracted from known structures, based on observed amino acid pairing frequencies. The summation of such energies, however, cannot be carried out for proteins whose structure is not known or for intrinsically unstructured proteins. To overcome these limitations, we present a novel method for estimating the total pairwise interaction energy, based on a quadratic form in the amino acid composition of the protein. This approach is validated by the good correlation of the estimated and actual energies of proteins of known structure and by a clear separation of folded and disordered proteins in the energy space it defines. As the novel algorithm has not been trained on unstructured proteins, it substantiates the concept of protein disorder, i.e. that the inability to form a well-defined 3D structure is an intrinsic property of many proteins and protein domains. This property is encoded in their sequence, because their biased amino acid composition does not allow sufficient stabilizing interactions to form. By limiting the calculation to a predefined sequential neighborhood, the algorithm was turned into a position-specific scoring scheme that characterizes the tendency of a given amino acid to fall into an ordered or disordered region. This application we term IUPred and compare its performance with three generally accepted predictors, PONDR VL3H, DISOPRED2 and GlobPlot on a database of disordered proteins.

Amino Acids↗

Prevalent structural disorder in E. coli and S. cerevisiae proteomes.

Intrinsically unstructured proteins, which exist without a well-defined 3D structure, carry out essential functions and occur with high frequency, as predicted for genomes. The generality of this phenomenon, however, is questioned by the uncertainty of what fraction of genomes actually encodes for expressed proteins. Here, we used two independent bioinformatic predictors, PONDR VSL1, and IUPred, to demonstrate that disorder prevails in the recently characterized proteomes and essential proteins of E. coli and S. cerevisiae, at levels exceeding that estimated from the genomes. The S. cerevisiae proteome contains three times as much disorder as that of E. coli, with 50-60% of proteins containing at least one long (>30 residues) disordered segment. This evolutionary advance can be explained by the observation that disorder is much higher in Gene Ontology categories related to regulatory, as opposed to metabolic, functions, and also in categories unique to yeast. Thus, protein disorder is a widespread and functionally important phenomenon, which needs to be characterized in full detail for understanding complex organisms at the molecular level.

Algorithms↗

Predicting intrinsic disorder from amino acid sequence.

Blind predictions of intrinsic order and disorder were made on 42 proteins subsequently revealed to contain 9,044 ordered residues, 284 disordered residues in 26 segments of length 30 residues or less, and 281 disordered residues in 2 disordered segments of length greater than 30 residues. The accuracies of the six predictors used in this experiment ranged from 77% to 91% for the ordered regions and from 56% to 78% for the disordered segments. The average of the order and disorder predictions ranged from 73% to 77%. The prediction of disorder in the shorter segments was poor, from 25% to 66% correct, while the prediction of disorder in the longer segments was better, from 75% to 95% correct. Four of the predictors were composed of ensembles of neural networks. This enabled them to deal more efficiently with the large asymmetry in the training data through diversified sampling from the significantly larger ordered set and achieve better accuracy on ordered and long disordered regions. The exclusive use of long disordered regions for predictor training likely contributed to the disparity of the predictions on long versus short disordered regions, while averaging the output values over 61-residue windows to eliminate short predictions of order or disorder probably contributed to the even greater disparity for three of the predictors. This experiment supports the predictability of intrinsic disorder from amino acid sequence.

Amino Acids↗

Exploiting heterogeneous sequence properties improves prediction of protein disorder.

During the past few years we have investigated methods to improve predictors of intrinsically disordered regions longer than 30 consecutive residues. Experimental evidence, however, showed that these predictors were less successful on short disordered regions, as observed two years ago during the fifth Critical Assessment of Techniques for Protein Structure Prediction (CASP5). To address this shortcoming, we developed a two-level model called VSL1 (CASP6 id: 193-1). At the first level, VSL1 consists of two specialized predictors, one of which was optimized for long disordered regions (>30 residues) and the other for short disordered regions (< or =30 residues). At the second level, a meta-predictor was built to assign weights for combining the two first-level predictors. As the results of the CASP6 experiment showed, this new predictor has achieved the highest accuracy yet and significantly improved performance on short disordered regions, while maintaining high performance on long disordered regions.

Algorithms↗

Intrinsically disordered structure of Bacillus pasteurii UreG as revealed by steady-state and time-resolved fluorescence spectroscopy.

UreG is an essential protein for the in vivo activation of urease. In a previous study, UreG from Bacillus pasteurii was shown to behave as an intrinsically unstructured dimeric protein. Here, intrinsic and extrinsic fluorescence experiments were performed, in the absence and presence of denaturant, to provide information about the form (fully folded, molten globule, premolten globule, or random coil) that the native state of BpUreG assumes in solution. The features of the emission band of the unique tryptophan residue (W192) located on the C-terminal helix, as well as the rate of bimolecular quenching by potassium iodide, indicated that, in the native state, W192 is protected from the aqueous polar solvent, while upon addition of denaturant, a conformational change occurs that causes solvent exposure of the indole side chain. This structural change, mainly affecting the C-terminal helix, is associated with the release of static quenching, as shown by resolution of the decay-associated spectra. The exposure of protein hydrophobic sites, monitored using the fluorescent probe bis-ANS, indicated that the native dimeric state of BpUreG is disordered even though it maintains a significant amount of tertiary structure. ANS fluorescence also indicated that, upon addition of a small amount of GuHCl, a transition to a molten globule state occurs, followed by formation of a pre-molten globule state at a higher denaturant concentration. The latter form is resistant to full unfolding, as also revealed by far-UV circular dichroism spectroscopy. The hydrodynamic parameters obtained by time-resolved fluorescence anisotropy at maximal denaturant concentrations (3 M GuHCl) confirmed the existence of a disordered but stable dimeric protein core. The nature of the forces holding together the two monomers of BpUreG was investigated. Determination of free thiols in native or denaturant conditions, as well as light scattering experiments in the absence and presence of dithiothreitol as a reducing agent, under native or denaturing conditions, indicates that a disulfide bond, involving the unique conserved cysteine C68, is present under native conditions and maintained upon addition of denaturant. This covalent bond is therefore important for the stabilization of the dimer under native conditions. The intrinsically disordered structure of UreG is discussed with respect to the role of this protein as a chaperone in the urease assembly system.

Bacillus↗

Modular organization of SARS coronavirus nucleocapsid protein.

The SARS-CoV nucleocapsid (N) protein is a major antigen in severe acute respiratory syndrome. It binds to the viral RNA genome and forms the ribonucleoprotein core. The SARS-CoV N protein has also been suggested to be involved in other important functions in the viral life cycle. Here we show that the N protein consists of two non-interacting structural domains, the N-terminal RNA-binding domain (RBD) (residues 45-181) and the C-terminal dimerization domain (residues 248-365) (DD), surrounded by flexible linkers. The C-terminal domain exists exclusively as a dimer in solution. The flexible linkers are intrinsically disordered and represent potential interaction sites with other protein and protein-RNA partners. Bioinformatics reveal that other coronavirus N proteins could share the same modular organization. This study provides information on the domain structure partition of SARS-CoV N protein and insights into the differing roles of structured and disordered regions in coronavirus nucleocapsid proteins.

Amino Acid Sequence↗

Abundance of intrinsically unstructured proteins in P. falciparum and other apicomplexan parasite proteomes.

Preliminary sequence analysis of Plasmodium falciparum has shown that the proteome of this organism is enriched in intrinsically unstructured proteins (IUPs), which are either completely disordered or contain large disordered regions. IUPs have been characterized as a unique class of proteins that plays an important role in biology and disease. In this study, the IUP contents in the proteomes of apicomplexan parasites, especially the proteome of P. falciparum and its various life cycle stages, have been evaluated with DisEMBL-1.4. Compared with other proteomes, apicomplexan species are extremely abundant in proteins containing long disordered regions, and the IUP contents in mammalian Plasmodium species are higher than in most other apicomplexan parasites. The proteome of the P. falciparum sporozoite appears to be distinct from the other life cycle stages in having an even higher content of disordered proteins. The abundance of IUPs in the P. falciparum proteome correlates with its enrichment in repetitive sequences. The structural plasticity of IUPs, which allows promiscuous binding interactions, may favour parasite survival both by inhibiting the generation of effective high affinity antibody responses and by facilitating the interactions with host molecules necessary for attachment and invasion of host cells.

Animals↗

Natively unfolded proteins: a point where biology waits for physics.

The experimental material accumulated in the literature on the conformational behavior of intrinsically unstructured (natively unfolded) proteins was analyzed. Results of this analysis showed that these proteins do not possess uniform structural properties, as expected for members of a single thermodynamic entity. Rather, these proteins may be divided into two structurally different groups: intrinsic coils, and premolten globules. Proteins from the first group have hydrodynamic dimensions typical of random coils in poor solvent and do not possess any (or almost any) ordered secondary structure. Proteins from the second group are essentially more compact, exhibiting some amount of residual secondary structure, although they are still less dense than native or molten globule proteins. An important feature of the intrinsically unstructured proteins is that they undergo disorder-order transition during or prior to their biological function. In this respect, the Protein Quartet model, with function arising from four specific conformations (ordered forms, molten globules, premolten globules, and random coils) and transitions between any two of the states, is discussed.

Databases, Factual↗

Prediction of unfolded segments in a protein sequence based on amino acid composition.

MOTIVATION: Partially and wholly unstructured proteins have now been identified in all kingdoms of life--more commonly in eukaryotic organisms. This intrinsic disorder is related to certain critical functions. Apart from their fundamental interest, unstructured regions in proteins may prevent crystallization. Therefore, the prediction of disordered regions is an important aspect for the understanding of protein function, but may also help to devise genetic constructs. RESULTS: In this paper we present a computational tool for the detection of unstructured regions in proteins based on two properties of unfolded fragments: (1) disordered regions have a biased composition and (2) they usually contain either small or no hydrophobic clusters. In order to quantify these two facts we first calculate the amino acid distributions in structured and unstructured regions. Using this distribution, we calculate for a given sequence fragment the probability to be part of either a structured or an unstructured region. For each amino acid, the distance to the nearest hydrophobic cluster is also computed. Using these three values along a protein sequence allows us to predict unstructured regions, with very simple rules. This method requires only the primary sequence, and no multiple alignment, which makes it an adequate method for orphan proteins. AVAILABILITY: http://genomics.eu.org/

Algorithms↗

Comprehensive evaluation of AlphaFold/OpenFold prediction of experimentally unresolved proteins through novel metrics.

Predicting accurate protein structures is essential for understanding molecular mechanisms, interpreting the impact of sequence variation, and supporting translational applications ranging from drug discovery to clinical genomics. Recent advances in deep-learning-based predictors such as AlphaFold2, OpenFold, and AlphaFold3 have transformed structural biology, enabling routine in silico modeling even for challenging or previously uncharacterized proteins. However, systematic benchmarking of these tools-especially for novel targets and single amino acid variants-remains limited. Conventional global metrics often fail to capture biologically meaningful discrepancies. By evaluating multiple implementations of AlphaFold2 and OpenFold, together with ColabFold and the AlphaFold3 server, across 10 different proteins and 222 single amino acid protein variants encompassing a wide range of sizes, structures, and functions, we show that although widely used global indicators-like mean pLDDT, pTM-score, and RMSD-frequently suggest comparable performance, substantial local-level differences remain elusive. To address this gap, we introduce a comparative framework leveraging Bland-Altman agreement analysis, to evaluate per-residue C&#x3b1;-confidence differences and Per-Residue profiles (PRPs), complemented by Uniform Manifold Approximation and Projection (UMAP). This approach reveals marked localized divergences, particularly within flexible or intrinsically disordered regions, where both predictor choice and single-residue substitutions trigger the largest conformational shifts. We further demonstrate that using reduced homology databases has minimal impact on predicted structural quality, offering computationally efficient alternatives. Collectively, our findings underscore the importance of integrating global and residue-specific evaluations to more accurately assess robustness, agreement, and practical usability across contemporary protein structure prediction methods.

Proteins↗

Characterizing residual structure in disordered protein States using nuclear magnetic resonance.

The importance of disordered protein states in biology is gaining recognition, and can be attributed in part to the participation of unfolded and partially folded states of globular proteins in normal and abnormal biological functions, such as protein translation, protein translocation, protein degradation, protein assembly, and protein aggregation (1-5). There is also a growing awareness that a significant fraction of gene products from various genomes, including the human genome, fall into a category that includes low complexity, low globularity, or intrinsically unstructured proteins (6-9). Unlike native states of globular proteins, disordered protein states, by definition, do not adopt a fixed structure that can be determined using classical high-resolution methods. Nevertheless, there has long been evidence that many disordered states contain detectable and significant residual or nascent structure (10-16). This structure has been found to be important for nucleating local structure, as well as mediating long range contacts upon either intramolecular folding to the native state (17-21) or intermolecular folding with specific binding partners (22-24), and is also predicted to influence intermolecular folding into structured aggregates (25,26). The primary tool for the characterization of such structure is high-resolution solution state nuclear magnetic resonance (NMR) spectroscopy. Advances in NMR instrumentation and methods have greatly facilitated this task and in principle can now be accomplished by those without extensive prior experience in NMR spectroscopy. This chapter describes how this can be accomplished.

Nuclear Magnetic Resonance, Biomolecular↗