PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “intrinsically disordered regions”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Spritz: a server for the prediction of intrinsically disordered regions in protein sequences using kernel machines.

Intrinsically disordered proteins have long stretches of their polypeptide chain, which do not adopt a single native structure composed of stable secondary and tertiary structure in the absence of binding partners. The prediction of intrinsically disordered regions in proteins from sequence is increasingly becoming of interest, as the presence of many such regions in the complete genome sequences are discovered and important functional roles are associated with them. We have developed a machine learning approach based on two support vector machines (SVM) to discriminate disordered regions from sequence. The SVM are trained and benchmarked on two sets, representing long and short disordered regions. A preliminary version of Spritz was shown to perform consistently well at the recent biannual CASP-6 experiment [Critical Assessment of Techniques for Protein Structure Prediction (CASP), 2004]. The fully developed Spritz method is freely available as a web server at http://distill.ucd.ie/spritz/ and http://protein.cribi.unipd.it/spritz/.

Artificial Intelligence↗

The 1H, 15N and 13C backbone resonance assignments of an intrinsically disordered region (467-696) of breast cancer type 1 susceptibility protein (BRCA1).

The tumor suppressor protein breast cancer type 1 susceptibility protein (BRCA1) plays a central role in maintaining genome stability through its involvement in DNA damage repair, transcriptional regulation, and cell-cycle control. BRCA1 functions as an obligate heterodimer with its binding partner, the BRCA1-associated RING domain protein 1 (BARD1), to coordinate accurate DNA repair. While the structured N- and C-terminal domains of BRCA1 have been well-characterized, the large central region encoded largely by exon 11 that comprises ~ 80% of the protein, is intrinsically disordered, and remains poorly structurally characterized. This intrinsically disordered region (IDR) harbors critical interaction interfaces for key proteins involved in genome maintenance, including RAD50, RAD51, MYC, and RB. Here, we report the backbone resonance assignments of a BRCA1 IDR construct spanning residues 467-696, providing a foundation for future studies aimed at understanding how the disordered central region of BRCA1 contributes to homologous recombination, interactions with BARD1, and overall BRCA1 tumor suppressor function.

BRCA1 Protein↗

Super-enhancer trapping by the nuclear pore via intrinsically disordered regions of proteins in squamous cell carcinoma cells.

Master transcription factors such as TP63 establish super-enhancers (SEs) to drive core transcriptional networks in cancer cells, yet the spatiotemporal regulation of SEs within the nucleus remains unknown. The nuclear pore complex (NPC) may tether SEs to the nuclear pore where RNA export rates are maximal. Here, we report that NUP153, a component of the NPC, anchors SEs to the NPC and enhances TP63 expression by maximizing mRNA export. This anchoring is mediated through protein-protein interaction between the intrinsically disordered regions (IDRs) of NUP153 and the coactivator BRD4. Silencing of NUP153 excludes SEs from the nuclear periphery, decreases TP63 expression, impairs cellular growth, and induces epidermal differentiation of squamous cell carcinoma. Overall, this work reveals the critical roles of NUP153 IDRs in the regulation of SE localization, thus providing insights into a new layer of gene regulation at the epigenomic and spatial level.

Humans↗

The1H, 15N and13C backbone resonance assignments of an intrinsically disordered region (124-270) of BRCA1 associated RING domain 1 (BARD1).

The BRCA1-associated RING domain protein 1 (BARD1) is the obligate binding partner of the tumor suppressor breast cancer type 1 susceptibility protein (BRCA1) and plays a critical role in maintaining genome integrity. BARD1 contains structured N- and C-terminal domains that mediate heterodimerization with BRCA1, recognition of chromatin marks, and DNA repair functions. Approximately 40% of BARD1 is intrinsically disordered, particularly in the central region of the protein. This intrinsically disordered region (IDR) engages DNA and key repair proteins such as RAD51, BLM, and WRN. DNA binding through the BARD1 IDR facilitates H2A ubiquitination by the BRCA1-BARD1 complex and is essential for stimulating long-range DNA end resection during homologous recombination, underscoring its role in accurate DNA repair. Despite these insights, structural characterization of the IDR remains limited, leaving questions regarding its functional interplay with BRCA1 and other repair factors unresolved. Here, we report the backbone resonance assignments of a BARD1 IDR construct spanning residues 124-270, providing a foundation for future studies aimed at understanding how the disordered regions of BARD1 interact with various binding partners, and cooperates with itself and BRCA1 to regulate genome stability.

Nuclear Magnetic Resonance, Biomolecular↗

Human transcription factors contain a high fraction of intrinsically disordered regions essential for transcriptional regulation.

Human transcriptional regulation factors, such as activators, repressors, and enhancer-binding factors are quite different from their prokaryotic counterparts in two respects: the average sequence in human is more than twice as long as that in prokaryotes, while the fraction of sequence aligned to domains of known structure is 31% in human transcription factors (TFs), less than half of that in bacterial TFs (72%). Intrinsically disordered (ID) regions were identified by a disorder-prediction program, and were found to be in good agreement with available experimental data. Analysis of 401 human TFs with experimental evidence from the Swiss-Prot database showed that as high as 49% of the entire sequence of human TFs is occupied by ID regions. More than half of the human TFs consist of a small DNA binding domain (DBD) and long ID regions frequently sandwiching unassigned regions. The remaining TFs have structural domains in addition to DBDs and ID regions. Experimental studies, particularly those with NMR, revealed that the transactivation domains in unbound TFs are usually unstructured, but become structured upon binding to their partners. The sequences of human and mouse TF orthologues are 90.5% identical despite a high incidence of ID regions, probably reflecting important functional roles played by ID regions. In general ID regions occupy a high fraction in TFs of eukaryotes, but not in prokaryotes. Implications of this dichotomy are discussed in connection with their functional roles in transcriptional regulation and evolution.

Animals↗

ATRX Condensates as Candidate Organizers of Enhancer-Centered Nuclear Microenvironments in Neural Progenitors: A Hypothesis for Enhancer-Associated ATRX Function in Neural Progenitors.

Neural progenitor cells (NPCs) must preserve lineage identity while remaining responsive to developmental cues. Here, we discuss the hypothesis that ATRX condensates help organize enhancer-centered nuclear microenvironments in NPCs. ATRX has long been studied in heterochromatin maintenance, histone variant deposition, and chromatin remodeling; earlier work has also shown that ATRX can occupy euchromatic and active regulatory regions and contribute to transcriptional regulation. Recent evidence in human NPCs indicates that ATRX forms nuclear puncta with condensate-like properties, associates with neurogenic enhancer-rich regions, and incorporates regulatory factors such as CHD7 and p300. Perturbation of ATRX condensate formation is associated with changes in enhancer-associated ATRX occupancy, neural gene-expression programs, and neuroepithelial organization, suggesting a regulatory mode that may complement canonical heterochromatin-associated functions. We propose a dual-mode model in which folded domains contribute to chromatin anchoring at repressive regions, whereas intrinsically disordered regions support condensate-associated organization at active developmental enhancers. We emphasize that whether ATRX condensates activate enhancers de novo, stabilize pre-existing enhancer states, buffer transcriptional variability, or primarily organize cofactor localization remains unresolved. We also discuss limitations of the current evidence and outline acute, locus-specific experiments needed to test the model.

X-linked Nuclear Protein↗

Conservation of intrinsic disorder in protein domains and families: I. A database of conserved predicted disordered regions.

Many protein regions have been shown to be intrinsically disordered, lacking unique structure under physiological conditions. These intrinsically disordered regions are not only very common in proteomes, but also crucial to the function of many proteins, especially those involved in signaling, recognition, and regulation. The goal of this work was to identify the prevalence, characteristics, and functions of conserved disordered regions within protein domains and families. A database was created to store the amino acid sequences of nearly one million proteins and their domain matches from the InterPro database, a resource integrating eight different protein family and domain databases. Disorder prediction was performed on these protein sequences. Regions of sequence corresponding to domains were aligned using a multiple sequence alignment tool. From this initial information, regions of conserved predicted disorder were found within the domains. The methodology for this search consisted of finding regions of consecutive positions in the multiple sequence alignments in which a 90% or more of the sequences were predicted to be disordered. This procedure was constrained to find such regions of conserved disorder prediction that were at least 20 amino acids in length. The results of this work included 3,653 regions of conserved disorder prediction, found within 2,898 distinct InterPro entries. Most regions of conserved predicted disorder detected were short, with less than 10% of those found exceeding 30 residues in length.

Amino Acid Sequence↗

Enrichment of G-to-U Substitution in SARS-CoV-2 Functional Regions and Its Compensation via Concurrent Mutations.

We surveyed single nucleotide variant (SNV) patterns from 5 903 647 complete SARS-CoV-2 genomes. Among 10 012 SNVs, APOBEC-mediated C-to-U (C > U) deamination was the most prevalent, followed by G > U and other RNA editing-related substitutions including (A > G, U > C, G > A). However, C > U mutations were less frequent in functional regions, for example, S protein, intrinsic disordered regions, and nonsynonymous mutations, where G > U were over-represented. Notably, G-loss substitutions rarely appeared together. Instead, G-gain mutations tended to more frequently co-occur with others, with a marked preference in the S protein, suggesting a compensatory mechanism for G loss in G > U mutations. The temporal patterns revealed C > U frequency declined until late 2021 then resurged in early 2022. Conversely, G > U steadily decreased, with a pronounced drop in January 2022, coinciding with reduced COVID-19 severity. Vaccinated individuals exhibited a slightly but significantly higher C > U frequency and a notably lower G > U frequency compared to the unvaccinated group. Additionally, cancer patients had higher G > U frequency than general patients during the same period. Interestingly, none of the C > U SNVs were uniquely identified in 2724 environmental samples. These findings suggest novel functional roles of G > U in COVID-19 symptoms, potentially linked to oxidative stress and reactive oxygen species, while C > U remains the dominant substitution, likely driven by host immune-mediated RNA editing.

SARS-CoV-2↗

Intrinsic disorder in the Protein Data Bank.

The Protein Data Bank (PDB) is the preeminent source of protein structural information. PDB contains over 32,500 experimentally determined 3-D structures solved using X-ray crystallography or nuclear magnetic resonance spectroscopy. Intrinsically disordered regions fail to form a fixed 3-D structure under physiological conditions. In this study, we compare the amino-acid sequences of proteins whose structures are determined by X-ray crystallography with the corresponding sequences from the Swiss-Prot database. The analyzed dataset includes 16,370 structures, which represent 18,101 PDB chains and 5,434 different proteins from 910 different organisms (2,793 eukaryotic, 2,109 bacterial, 288 viral, and 244 archaeal). In this dataset, on average, each Swiss-Prot protein is represented by 7 PDB chains with 76% of the crystallized regions being represented by more than one structure. Intriguingly, the complete sequences of only approximately 7% of proteins are observed in the corresponding PDB structures, and only approximately 25% of the total dataset have >95% of their lengths observed in the corresponding PDB structures. This suggests that the vast majority of PDB proteins is shorter than their corresponding Swiss-Prot sequences and/or contain numerous residues, which are not observed in maps of electron density. To determine the prevalence of disordered regions in PDB, the residues in the Swiss-Prot sequences were grouped into four general categories, "Observed" (which correspond to structured regions), "Not observed" (regions with missing electron density, potentially disordered), "Uncharacterized," and "Ambiguous," depending on their appearance in the corresponding PDB entries. This non-redundant set of residues can be viewed as a 'fragment' or empirical domain database that contains a set of experimentally determined structured regions or domains and a set of experimentally verified disordered regions or domains. We studied the propensities and properties of residues in these four categories and analyzed their relations to the predictions of disorder using several algorithms. "Non-observed," "Ambiguous," and "Uncharacterized" regions were shown to possess the amino acid compositional biases typical of intrinsically disordered proteins. The application of four different disorder predictors (PONDR(R) VL-XT, VL3-BA, VSL1P, and IUPred) revealed that the vast majority of residues in the "Observed" dataset are ordered, and that the "Not observed" regions are mostly disordered. The "Uncharacterized" regions possess some tendency toward order, whereas the predictions for the short "Ambiguous" regions are really ambiguous. Long "Ambiguous" regions (>70 amino acid residues) are mostly predicted to be ordered, suggesting that they are likely to be "wobbly" domains. Overall, we showed that completely ordered proteins are not highly abundant in PDB and many PDB sequences have disordered regions. In fact, in the analyzed dataset approximately 10% of the PDB proteins contain regions of consecutive missing or ambiguous residues longer than 30 amino-acids and approximately 40% of the proteins possess short regions (> or =10 and < 30 amino-acid long) of missing and ambiguous residues.

Algorithms↗

Comprehensive evaluation of AlphaFold/OpenFold prediction of experimentally unresolved proteins through novel metrics.

Predicting accurate protein structures is essential for understanding molecular mechanisms, interpreting the impact of sequence variation, and supporting translational applications ranging from drug discovery to clinical genomics. Recent advances in deep-learning-based predictors such as AlphaFold2, OpenFold, and AlphaFold3 have transformed structural biology, enabling routine in silico modeling even for challenging or previously uncharacterized proteins. However, systematic benchmarking of these tools-especially for novel targets and single amino acid variants-remains limited. Conventional global metrics often fail to capture biologically meaningful discrepancies. By evaluating multiple implementations of AlphaFold2 and OpenFold, together with ColabFold and the AlphaFold3 server, across 10 different proteins and 222 single amino acid protein variants encompassing a wide range of sizes, structures, and functions, we show that although widely used global indicators-like mean pLDDT, pTM-score, and RMSD-frequently suggest comparable performance, substantial local-level differences remain elusive. To address this gap, we introduce a comparative framework leveraging Bland-Altman agreement analysis, to evaluate per-residue C&#x3b1;-confidence differences and Per-Residue profiles (PRPs), complemented by Uniform Manifold Approximation and Projection (UMAP). This approach reveals marked localized divergences, particularly within flexible or intrinsically disordered regions, where both predictor choice and single-residue substitutions trigger the largest conformational shifts. We further demonstrate that using reduced homology databases has minimal impact on predicted structural quality, offering computationally efficient alternatives. Collectively, our findings underscore the importance of integrating global and residue-specific evaluations to more accurately assess robustness, agreement, and practical usability across contemporary protein structure prediction methods.

Proteins↗

An Intrinsically Disordered RNA Binding Protein Modulates mRNA Translation and Storage.

Proteins with intrinsically disordered regions (IDR) play diverse functions in regulating gene expression in the cell. Many of these proteins interact with cytoplasmic ribosomes. However, the molecular functions related to the interactions are largely unclear. In this study, using an abundant RNA-binding protein, Sbp1, with a structurally well-defined RNA recognition motif and an intrinsically disordered RGG domain as a model system, we investigated how an RNA binding protein with IDR modulates mRNA storage and translation. Using genomic and molecular approaches, we show that Sbp1 slows ribosome movement on cellular mRNAs and promotes polysome stacking or aggregation. Sbp1-associated polysomes display a ring-shaped structure in addition to a beads-on-string morphology visualized under the electron microscope, likely to be an intermediate slow translation state between actively translating polysomes and the translation-sequestered RNA granule. Moreover, the binding of Sbp1 to the 5'UTRs of mRNAs represses both cap-dependent and cap-independent translation initiation of proteins, many are functionally important for general protein synthesis in the cell. Finally, post-translational modifications at the arginine in the RGG motif change the Sbp1 protein interactome and play important roles in directing cellular mRNAs to either translation or storage. Taken together, our study demonstrates that under physiological conditions, intrinsically disordered RNA binding proteins promote polysome aggregation and regulate mRNA translation and storage using multiple distinctive mechanisms. This research also establishes a framework with which functions of other IDR-containing proteins can be investigated and defined.

RNA-Binding Proteins↗

pkaPS: prediction of protein kinase A phosphorylation sites with the simplified kinase-substrate binding model.

BACKGROUND: Protein kinase A (cAMP-dependent kinase, PKA) is a serine/threonine kinase, for which ca. 150 substrate proteins are known. Based on a refinement of the recognition motif using the available experimental data, we wished to apply the simplified substrate protein binding model for accurate prediction of PKA phosphorylation sites, an approach that was previously successful for the prediction of lipid posttranslational modifications and of the PTS1 peroxisomal translocation signal. RESULTS: Approximately 20 sequence positions flanking the phosphorylated residue on both sides have been found to be restricted in their sequence variability (region -18...+23 with the site at position 0). The conserved physical pattern can be rationalized in terms of a qualitative binding model with the catalytic cleft of the protein kinase A. Positions -6...+4 surrounding the phosphorylation site are influenced by direct interaction with the kinase in a varying degree. This sequence stretch is embedded in an intrinsically disordered region composed preferentially of hydrophilic residues with flexible backbone and small side chain. This knowledge has been incorporated into a simplified analytical model of productive binding of substrate proteins with PKA. CONCLUSION: The scoring function of the pkaPS predictor can confidently discriminate PKA phosphorylation sites from serines/threonines with non-permissive sequence environments (sensitivity of appoximately 96% at a specificity of approximately 94%). The tool "pkaPS" has been applied on the whole human proteome. Among new predicted PKA targets, there are entirely uncharacterized protein groups as well as apparently well-known families such as those of the ribosomal proteins L21e, L22 and L6. AVAILABILITY: The supplementary data as well as the prediction tool as WWW server are available at http://mendel.imp.univie.ac.at/sat/pkaPS. REVIEWERS: Erik van Nimwegen (Biozentrum, University of Basel, Switzerland), Sandor Pongor (International Centre for Genetic Engineering and Biotechnology, Trieste, Italy), Igor Zhulin (University of Tennessee, Oak Ridge National Laboratory, USA).

Journal Article↗

Viral replication through phase separation: Cytosolic and nuclear condensates.

Replication of many RNA and DNA viruses occurs within specialized intracellular hubs organized as membraneless biomolecular condensates (BCs) driven by liquid-liquid phase separation. As obligate intracellular parasites, viruses depend on the host cell machinery to complete their replication cycles and therefore actively remodel the intracellular environment to favor viral genome replication, transcription, and assembly. Cytosolic and nuclear phase-separated replication compartments (RC) provide concentrated and dynamic platforms that promote efficient interactions between viral genomes and viral or host proteins essential for infection. The formation of viral replication BCs is typically facilitated by viral proteins enriched in intrinsically disordered regions and low-complexity domains, which enable multivalent interactions with viral nucleic acids and cellular factors. These interactions are mediated by diverse biophysical forces, including hydrophobic and &#x3c0; interactions, hydrogen bonding, molecular crowding, and osmotic effects. Throughout infection, viral BCs remain highly dynamic, allowing continuous exchange of components and functional maturation of replication hubs. Their properties and activities are further regulated by post-translational modifications of viral and host proteins, such as phosphorylation, acetylation, and methylation. In this review, we summarize current evidence supporting liquid-liquid phase separation as a central organizing principle of viral RCs. We focus on representative RNA and DNA viruses that replicate in the cytosol or nucleus, highlighting virus-specific strategies, conserved mechanisms, and the consequences of BC formation for viral replication efficiency, host antiviral responses, and therapeutic intervention.

Phase Separation↗

MobiDB-lite 4.0: faster prediction of intrinsic protein disorder and structural compactness.

MOTIVATION: In recent years, many disorder predictors have been developed to identify intrinsically disordered regions (IDRs) in proteins, achieving high accuracy. However, it may be difficult to interpret differences in predictions across methods. Consensus methods offer a simple solution, highlighting reliable predictions while filtering out uncertain positions. Here, we present a new version of MobiDB-lite, a consensus method designed to predict long IDRs and classify them based on compositional biases and conformational properties. RESULTS: MobiDB-lite 4.0 pipeline was optimized to be ten times faster than the previous version. It now provides compactness annotations based on predicted apparent scaling exponent. The newly added features and disorder subclassifications allow the users to get a comprehensive insight into the protein's function and characteristics. MobiDB-lite 4.0 is integrated into the MobiDB and DisProt databases. A version without the compactness predictor is integrated into InterProScan, propagating MobiDB-lite annotations to UniProtKB. AVAILABILITY AND IMPLEMENTATION: The MobiDB-lite 4.0 source code and a Docker container are available from the GitHub repository: https://github.com/BioComputingUP/MobiDB-lite.

Intrinsically Disordered Proteins↗

Exploiting heterogeneous sequence properties improves prediction of protein disorder.

During the past few years we have investigated methods to improve predictors of intrinsically disordered regions longer than 30 consecutive residues. Experimental evidence, however, showed that these predictors were less successful on short disordered regions, as observed two years ago during the fifth Critical Assessment of Techniques for Protein Structure Prediction (CASP5). To address this shortcoming, we developed a two-level model called VSL1 (CASP6 id: 193-1). At the first level, VSL1 consists of two specialized predictors, one of which was optimized for long disordered regions (>30 residues) and the other for short disordered regions (< or =30 residues). At the second level, a meta-predictor was built to assign weights for combining the two first-level predictors. As the results of the CASP6 experiment showed, this new predictor has achieved the highest accuracy yet and significantly improved performance on short disordered regions, while maintaining high performance on long disordered regions.

Algorithms↗

Small GTPase RAN-driven PNET2 oligomerization and phase separation at the nuclear lamina promote nuclear envelope integrity in plants.

The nuclear envelope is a fundamental organizer of eukaryotic cells, yet how plants regulate its architecture and integrity remains poorly understood. In this study, we identified the plant inner nuclear membrane protein PLANT NUCLEAR ENVELOPE TRANSMEMBRANE 2 (PNET2) as a scaffold that maintains nuclear envelope integrity and genome stability. Loss of PNET2 function compromises nuclear membrane structure and sensitizes cells to DNA damage, whereas overexpression drives aberrant nuclear membrane expansion. Biochemically, PNET2 cooperates with the nuclear lamin protein KAKU4 and CROWDED NUCLEI 1 within the nuclear lamina to promote nuclear membrane remodeling, a process driven by biomolecular condensate formation via their intrinsically disordered regions. We further uncovered a direct interaction between PNET2 and the small GTPase RAN. Structural modeling and biochemical analyses revealed that its active GTP-bound form stimulates PNET2 oligomerization, potentially promoting its phase separation to drive membrane expansion. Genetic analyses showed that PNET2 and RAN function in a shared pathway essential for nuclear membrane integrity. Together, our findings define a regulatory module that orchestrates GTPase signaling to sustain nuclear membrane homeostasis in plants, positioning PNET2 as a nexus linking membrane dynamics, nuclear lamina organization, and genome protection.

PNET2↗

DisP-seq reveals the genome-wide functional organization of DNA-associated disordered proteins.

Intrinsically disordered regions (IDRs) in DNA-associated proteins are known to influence gene regulation, but their distribution and cooperative functions in genome-wide regulatory programs remain poorly understood. Here we describe DisP-seq (disordered protein precipitation followed by DNA sequencing), an antibody-independent chemical precipitation assay that can simultaneously map endogenous DNA-associated disordered proteins genome-wide through a combination of biotinylated isoxazole precipitation and next-generation sequencing. DisP-seq profiles are composed of thousands of peaks that are associated with diverse chromatin states, are enriched for disordered transcription factors (TFs) and are often arranged in large lineage-specific clusters with high local concentrations of disordered proteins and different combinations of histone modifications linked to regulatory potential. We use DisP-seq to analyze cancer cells and reveal how disordered protein-associated islands enable IDR-dependent mechanisms that control the binding and function of disordered TFs, including oncogene-dependent sequestration of TFs through long-range interactions and the reactivation of differentiation pathways upon loss of oncogenic stimuli in Ewing sarcoma.

DNA↗