PubMed Health⌕ Search

Biomedical subjects

Christopher J Oldfield

Publications and source records attributed to Christopher J Oldfield.

14 recordsLinked to original sources

Intrinsic disorder and functional proteomics.

The recent advances in the prediction of intrinsically disordered proteins and the use of protein disorder prediction in the fields of molecular biology and bioinformatics are reviewed here, especially with regard to protein function. First, a close look is taken at intrinsically disordered proteins and then at the methods used for their experimental characterization. Next, the major statistical properties of disordered regions are summarized, and prediction models developed thus far are described, including their numerous applications in functional proteomics. The future of the prediction of protein disorder and the future uses of such predictions in functional proteomics comprise the last section of this article.

Algorithms↗

Abundance of intrinsic disorder in protein associated with cardiovascular disease.

Evidence that many protein regions and even entire proteins lacking stable tertiary and/or secondary structure in solution (i.e., intrinsically disordered proteins) might be involved in protein-protein interactions, regulation, recognition, and signal transduction is rapidly accumulating. These signaling proteins play a crucial role in the development of several pathological conditions, including cancer. To test a hypothesis that intrinsic disorder is also abundant in cardiovascular disease (CVD), a data set of 487 CVD-related proteins was extracted from SWISS-PROT. CVD-related proteins are depleted in major order-promoting residues (Trp, Phe, Tyr, Ile, and Val) and enriched in some disorder-promoting residues (Arg, Gln, Ser, Pro, and Glu). The application of a neural network predictor of natural disordered regions (PONDR VL-XT) together with cumulative distribution function (CDF) analysis, charge-hydropathy plot (CH plot) analysis, and alpha-helical molecular recognition feature (alpha-MoRF) indicator revealed that CVD-related proteins are enriched in intrinsic disorder. In fact, the percentage of proteins with 30 or more consecutive residues predicted by PONDR VL-XT to be disordered was 57 +/- 4% for CVD-associated proteins. This value is close that described earlier for signaling proteins (66 +/- 6%) and is significantly larger than the content of intrinsic disorder in eukaryotic proteins from SWISS-PROT (47 +/- 4%) and in nonhomologous protein segments with a well-defined three-dimensional structure (13 +/- 4%). Furthermore, CDF and CH-plot analyses revealed that 120 and 36 CVD-related proteins, respectively, are wholly disordered. This high level of intrinsic disorder could be important for the function of CVD-related proteins and for the control and regulation of processes associated with cardiovascular disease. In agreement with this hypothesis, 198 alpha-MoRFs were predicted in 101 proteins from the CVD data set. A comparison of disorder predictions with the experimental structural and functional data for a subset of the CVD-associated proteins indicated good agreement between predictions and observations. Thus, our data suggest that intrinsically disordered proteins might play key roles in cardiovascular disease.

3',5'-Cyclic-AMP Phosphodiesterases↗

Analysis of molecular recognition features (MoRFs).

Several proteomic studies in the last decade revealed that many proteins are either completely disordered or possess long structurally flexible regions. Many such regions were shown to be of functional importance, often allowing a protein to interact with a large number of diverse partners. Parallel to these findings, during the last five years structural bioinformatics has produced an explosion of results regarding protein-protein interactions and their importance for cell signaling. We studied the occurrence of relatively short (10-70 residues), loosely structured protein regions within longer, largely disordered sequences that were characterized as bound to larger proteins. We call these regions molecular recognition features (MoRFs, also known as molecular recognition elements, MoREs). Interestingly, upon binding to their partner(s), MoRFs undergo disorder-to-order transitions. Thus, in our interpretation, MoRFs represent a class of disordered region that exhibits molecular recognition and binding functions. This work extends previous research showing the importance of flexibility and disorder for molecular recognition. We describe the development of a database of MoRFs derived from the RCSB Protein Data Bank and present preliminary results of bioinformatics analyses of these sequences. Based on the structure adopted upon binding, at least three basic types of MoRFs are found: alpha-MoRFs, beta-MoRFs, and iota-MoRFs, which form alpha-helices, beta-strands, and irregular secondary structure when bound, respectively. Our data suggest that functionally significant residual structure can exist in MoRF regions prior to the actual binding event. The contribution of intrinsic protein disorder to the nature and function of MoRFs has also been addressed. The results of this study will advance the understanding of protein-protein interactions and help towards the future development of useful protein-protein binding site predictors.

Algorithms↗

Rational drug design via intrinsically disordered protein.

Despite substantial increases in research funding by the pharmaceutical industry, drug discovery rates seem to have reached a plateau or perhaps are even declining, suggesting the need for new strategies. Protein-protein interactions have long been thought to provide interesting drug discovery targets, but the development of small molecules that modulate such interactions has so far achieved a low success rate. In contrast to this historic trend, a few recent successes raise hopes for routinely identifying druggable protein-protein interactions. In this Opinion article, we point out the importance of coupled binding and folding for protein-protein signalling interactions generally, and from this and associated observations, we develop a new strategy for identifying protein-protein interactions that would be particularly promising targets for modulation by small molecules. This novel strategy, based on intrinsically disordered protein, has the potential to increase significantly the discovery rate for new molecule entities.

Chemistry, Pharmaceutical↗

Intrinsic disorder is a common feature of hub proteins from four eukaryotic interactomes.

Recent proteome-wide screening approaches have provided a wealth of information about interacting proteins in various organisms. To test for a potential association between protein connectivity and the amount of predicted structural disorder, the disorder propensities of proteins with various numbers of interacting partners from four eukaryotic organisms (Caenorhabditis elegans, Saccharomyces cerevisiae, Drosophila melanogaster, and Homo sapiens) were investigated. The results of PONDR VL-XT disorder analysis show that for all four studied organisms, hub proteins, defined here as those that interact with > or = 10 partners, are significantly more disordered than end proteins, defined here as those that interact with just one partner. The proportion of predicted disordered residues, the average disorder score, and the number of predicted disordered regions of various lengths were higher overall in hubs than in ends. A binary classification of hubs and ends into ordered and disordered subclasses using the consensus prediction method showed a significant enrichment of wholly disordered proteins and a significant depletion of wholly ordered proteins in hubs relative to ends in worm, fly, and human. The functional annotation of yeast hubs and ends using GO categories and the correlation of these annotations with disorder predictions demonstrate that proteins with regulation, transcription, and development annotations are enriched in disorder, whereas proteins with catalytic activity, transport, and membrane localization annotations are depleted in disorder. The results of this study demonstrate that intrinsic structural disorder is a distinctive and common characteristic of eukaryotic hub proteins, and that disorder may serve as a determinant of protein interactivity.

Amino Acids↗

Intrinsic disorder in transcription factors.

Intrinsic disorder (ID) is highly abundant in eukaryotes, which reflect the greater need for disorder-associated signaling and transcriptional regulation in nucleated cells. Although several well-characterized examples of intrinsically disordered proteins in transcriptional regulation have been reported, no systematic analysis has been reported so far. To test for the general prevalence of intrinsic disorder in transcriptional regulation, we used the predictor of natural disorder regions (PONDR) to analyze the abundance of intrinsic disorder in three transcription factor datasets and two control sets. This analysis revealed that from 94.13 to 82.63% of transcription factors possess extended regions of intrinsic disorder, relative to 54.51 and 18.64% of the proteins in two control datasets, which indicates the significant prevalence of intrinsic disorder in transcription factors. This propensity of transcription factors to intrinsic disorder was confirmed by cumulative distribution function analysis and charge-hydropathy plots. The amino acid composition analysis showed that all three transcription factor datasets were substantially depleted in order-promoting residues and significantly enriched in disorder-promoting residues. Our analysis of the distribution of disorder within the transcription factor datasets revealed that (a) the AT-hooks and basic regions of transcription factor DNA-binding domains are highly disordered; (b) the degree of disorder in transcription factor activation regions is much higher than that in DNA-binding domains; (c) the degree of disorder is significantly higher in eukaryotic transcription factors than in prokaryotic transcription factors; and (d) the level of alpha-MoRF (molecular recognition feature) prediction is much higher in transcription factors. Overall, our data reflected the fact that eukaryotes with well-developed gene transcription machinery require transcription factor flexibility to be more efficient.

Amino Acid Sequence↗

Alternative splicing in concert with protein intrinsic disorder enables increased functional diversity in multicellular organisms.

Alternative splicing of pre-mRNA generates two or more protein isoforms from a single gene, thereby contributing to protein diversity. Despite intensive efforts, an understanding of the protein structure-function implications of alternative splicing is still lacking. Intrinsic disorder, which is a lack of equilibrium 3D structure under physiological conditions, may provide this understanding. Intrinsic disorder is a common phenomenon, particularly in multicellular eukaryotes, and is responsible for important protein functions including regulation and signaling. We hypothesize that polypeptide segments affected by alternative splicing are most often intrinsically disordered such that alternative splicing enables functional and regulatory diversity while avoiding structural complications. We analyzed a set of 46 differentially spliced genes encoding experimentally characterized human proteins containing both structured and intrinsically disordered amino acid segments. We show that 81% of 75 alternatively spliced fragments in these proteins were associated with fully (57%) or partially (24%) disordered protein regions. Regions affected by alternative splicing were significantly biased toward encoding disordered residues, with a vanishingly small P value. A larger data set composed of 558 SwissProt proteins with known isoforms produced by 1,266 alternatively spliced fragments was characterized by applying the pondr vsl1 disorder predictor. Results from prediction data are consistent with those obtained from experimental data, further supporting the proposed hypothesis. Associating alternative splicing with protein disorder enables the time- and tissue-specific modulation of protein function needed for cell differentiation and the evolution of multicellular organisms.

Alternative Splicing↗

Protein intrinsic disorder and human papillomaviruses: increased amount of disorder in E6 and E7 oncoproteins from high risk HPVs.

It is recognized now that many functional proteins or their long segments are devoid of stable secondary and/or tertiary structure and exist instead as very dynamic ensembles of conformations. They are known by different names including natively unfolded, intrinsically disordered, intrinsically unstructured, rheomorphic, pliable, and different combinations thereof. Many important functions and activities have been associated with these intrinsically disordered proteins (IDPs), including molecular recognition, signaling, and regulation. It is also believed that disorder of these proteins allows function to be readily modified through phosphorylation, acetylation, ubiquitination, hydroxylation, and proteolysis. Bioinformatics analysis revealed that IDPs comprise a large fraction of different proteomes. Furthermore, it is established that the intrinsic disorder is relatively abundant among cancer-related and other disease-related proteins and IDPs play a number of key roles in oncogenesis. There are more than 100 different types of human papillomaviruses (HPVs), which are the causative agents of benign papillomas/warts, and cofactors in the development of carcinomas of the genital tract, head and neck, and epidermis. With respect to their association with cancer, HPVs are grouped into two classes, known as low (e.g., HPV-6 and HPV-11) and high-risk (e.g., HPV-16 and HPV-18) types. The entire proteome of HPV includes six nonstructural proteins [E1, E2, E4, E5, E6, and E7 (the latter two are known to function as oncoproteins in the high-risk HPVs)] and two structural proteins (L1 and L2). To understand whether intrinsic disorder plays a role in the oncogenic potential of different HPV types, we have performed a detailed bioinformatics analysis of proteomes of high-risk and low-risk HPVs with the major focus on E6 and E7 oncoproteins. The results of this analysis are consistent with the conclusion that high-risk HPVs are characterized by the increased amount of intrinsic disorder in transforming proteins E6 and E7.

Algorithms↗

Coupled folding and binding with alpha-helix-forming molecular recognition elements.

Many protein-protein and protein-nucleic acid interactions involve coupled folding and binding of at least one of the partners. Here, we propose a protein structural element or feature that mediates the binding events of initially disordered regions. This element consists of a short region that undergoes coupled binding and folding within a longer region of disorder. We call these features "molecular recognition elements" (MoREs). Examples of MoREs bound to their partners can be found in the alpha-helix, beta-strand, polyproline II helix, or irregular secondary structure conformations, and in various mixtures of the four structural forms. Here we describe an algorithm that identifies regions having propensities to become alpha-helix-forming molecular recognition elements (alpha-MoREs) based on a discriminant function that indicates such regions while giving a low false-positive error rate on a large collection of structured proteins. Application of this algorithm to databases of genomics and functionally annotated proteins indicates that alpha-MoREs are likely to play important roles protein-protein interactions involved in signaling events.

Binding Sites↗

Addressing the intrinsic disorder bottleneck in structural proteomics.

The Center for Eukaryotic Structural Genomics (CESG), as part of the Protein Structure Initiative (PSI), has established a high-throughput structure determination pipeline focused on eukaryotic proteins. NMR spectroscopy is an integral part of this pipeline, both as a method for structure determinations and as a means for screening proteins for stable structure. Because computational approaches have estimated that many eukaryotic proteins are highly disordered, about 1 year into the project, CESG began to use an algorithm (the Predictor of Naturally Disordered Regions, PONDR to avoid proteins that were likely to be disordered. We report a retrospective analysis of the effect of this filtering on the yield of viable structure determination candidates. In addition, we have used our current database of results on 70 protein targets from Arabidopsis thaliana and 1 from Caenorhabditis elegans, which were labeled uniformly with nitrogen-15 and screened for disorder by NMR spectroscopy, to compare the original algorithm with 13 other approaches for predicting disorder from sequence. Our study indicates that the efficiency of structural proteomics of eukaryotes can be improved significantly by removing targets predicted to be disordered by an algorithm chosen to provide optimal performance.

Algorithms↗

Comparing and combining predictors of mostly disordered proteins.

Intrinsically disordered proteins and regions carry out varied and vital cellular functions. Proteins with disordered regions are especially common in eukaryotic cells, with a subset of these proteins being mostly disordered, e.g., with more disordered than ordered residues. Two distinct methods have been previously described for using amino acid sequences to predict which proteins are likely to be mostly disordered. These methods are based on the net charge-hydropathy distribution and disorder prediction score distribution. Each of these methods is reexamined, and the prediction results are compared herein. A new prediction method based on consensus is described. Application of the consensus method to whole genomes reveals that approximately 4.5% of Yersinia pestis, 5% of Escherichia coli K12, 6% of Archaeoglobus fulgidus, 8% of Methanobacterium thermoautotrophicum, 23% of Arabidopsis thaliana, and 28% of Mus musculus proteins are mostly disordered. The unexpectedly high frequency of mostly disordered proteins in eukaryotes has important implications both for large-scale, high-throughput projects and also for focused experiments aimed at determination of protein structure and function.

Algorithms↗

The C-terminal domain of measles virus nucleoprotein belongs to the class of intrinsically disordered proteins that fold upon binding to their physiological partner.

The nucleoprotein of measles virus consists of an N-terminal domain, N(CORE) (aa 1-400), resistant to proteolysis, and a C-terminal domain, N(TAIL) (aa 401-525), hypersensitive to proteolysis and not visible by electron microscopy. Using two complementary computational approaches, we predict that N(TAIL) belongs to the class of natively unfolded proteins. Using different biochemical and biophysical approaches, we show that N(TAIL) is indeed unstructured in solution. In particular, the spectroscopic and hydrodynamic properties of N(TAIL) indicate that this protein domain belongs to the premolten globule subfamily within the class of intrinsically disordered proteins. The isolated N(TAIL) domain was shown to be able to bind to its physiological partner, the phosphoprotein (P), and to undergo an induced folding upon binding to the C-terminal moiety of P [J. Biol. Chem. 278 (2003) 18638]. Using a computational analysis, we have identified within N(TAIL) a putative alpha-helical molecular recognition element (alpha-MoRE, aa 488-499), which could be involved in binding to P via induced folding. We report the bacterial expression and purification of a truncated form of N(TAIL) (N(TAIL2), aa 401-488) devoid of the alpha-MoRE. We show that N(TAIL2) has lost the ability to bind to P, thus supporting the hypothesis that the alpha-MoRE may play a role in binding to P. We have further analyzed the alpha-helical propensities of N(TAIL2) and N(TAIL) using circular dichroism in the presence of 2,2,2-trifluoroethanol. We show that N(TAIL2) has a lower alpha-helical potential compared to N(TAIL), thus suggesting that the alpha-MoRE may be indeed involved in the induced folding of N(TAIL).

Circular Dichroism↗

Evolutionary rate heterogeneity in proteins with long disordered regions.

The dominant view in protein science is that a three-dimensional (3-D) structure is a prerequisite for protein function. In contrast to this dominant view, there are many counterexample proteins that fail to fold into a 3-D structure, or that have local regions that fail to fold, and yet carry out function. Protein without fixed 3-D structure is called intrinsically disordered. Motivated by anecdotal accounts of higher rates of sequence evolution in disordered protein than in ordered protein we are exploring the molecular evolution of disordered proteins. To test whether disordered protein evolves more rapidly than ordered protein, pairwise genetic distances were compared between the ordered and the disordered regions of 26 protein families having at least one member with a structurally characterized region of disorder of 30 or more consecutive residues. For five families, there were no significant differences in pairwise genetic distances between ordered and disordered sequences. The disordered region evolved significantly more rapidly than the ordered region for 19 of the 26 families. The functions of these disordered regions are diverse, including binding sites for protein, DNA, or RNA and also including flexible linkers. The functions of some of these regions are unknown. The disordered regions evolved significantly more slowly than the ordered regions for the two remaining families. The functions of these more slowly evolving disordered regions include sites for DNA binding. More work is needed to understand the underlying causes of the variability in the evolutionary rates of intrinsically ordered and disordered protein.

Amino Acid Sequence↗

Showing your ID: intrinsic disorder as an ID for recognition, regulation and cell signaling.

Regulation, recognition and cell signaling involve the coordinated actions of many players. To achieve this coordination, each participant must have a valid identification (ID) that is easily recognized by the others. For proteins, these IDs are often within intrinsically disordered (also ID) regions. The functions of a set of well-characterized ID regions from a diversity of proteins are presented herein to support this view. These examples include both more recently described signaling proteins, such as p53, alpha-synuclein, HMGA, the Rieske protein, estrogen receptor alpha, chaperones, GCN4, Arf, Hdm2, FlgM, measles virus nucleoprotein, RNase E, glycogen synthase kinase 3beta, p21(Waf1/Cip1/Sdi1), caldesmon, calmodulin, BRCA1 and several other intriguing proteins, as well as historical prototypes for signaling, regulation, control and molecular recognition, such as the lac repressor, the voltage gated potassium channel, RNA polymerase and the S15 peptide associating with the RNA polymerase S-protein. The frequent occurrence and the common use of ID regions in important protein functions raise the possibility that the relationship between amino acid sequence, disordered ensemble and function might be the dominant paradigm for the molecular recognition that serves as the basis for signaling and regulation by protein molecules.

Animals↗