PubMed Health⌕ Search

Biomedical subjects

Roland L Dunbrack

Publications and source records attributed to Roland L Dunbrack.

At least 19 recordsLinked to original sources

A germline KDM3C polymorphism impairs DNA repair and sensitizes to chemoradiotherapy.

Chemoradiotherapy (CRT) is the standard-of-care therapy for many solid malignancies, yet predictive biomarkers of treatment response remain limited. We identified a germline single nucleotide polymorphism (SNP) in an intrinsically disordered region of the lysine demethylase KDM3C/JMJD1C (p.S464T) that is associated with CRT outcomes in locally advanced rectal cancers (LARC) and head and neck squamous cell carcinoma (LA-HNSCC). In silico modeling with AlphaFold predicted S464T substitution influenced interaction between phosphorylated KDM3C and RNF8 FHA domain. In cellular models, conversion of S464 to T464 increased sensitivity to DNA-damaging agents. S464T substitution impaired damage-induced MDC1-RAP80 signaling and downstream RAP80-BRCA1 colocalization. SNP carrying cells impaired DNA repair causing genotoxic stress that is associated with increased cGAS-cGAMP innate immune signaling and increased apoptosis. Population analyses with the SNP highlighted an increase incidence of UV-induced skin and other cancers, linking inherited variation in the chromatin regulatory gene KDM3C to genome instability, cancer risk, and therapeutic vulnerability.

Journal Article↗

Statistical and conformational analysis of the electron density of protein side chains.

Protein side chains make most of the specific contacts between proteins and other molecules, and their conformational properties have been studied for many years. These properties have been analyzed primarily in the form of rotamer libraries, which cluster the observed conformations into groups and provide frequencies and average dihedral angles for these groups. In recent years, these libraries have improved with higher resolution structures and using various criteria such as high thermal factors to eliminate side chains that may be misplaced within the crystallographic model coordinates. Many of these side chains have highly non-rotameric dihedral angles. The origin of side chains with high B-factors and/or with non-rotameric dihedral angles is of interest in the determination of protein structures and in assessing the prediction of side chain conformations. In this paper, using a statistical analysis of the electron density of a large set of proteins, it is shown that: (1) most non-rotameric side chains have low electron density compared to rotameric side chains; (2) up to 15% of chi1 non-rotameric side chains in PDB models can clearly be fit to density at a single rotameric conformation and in some cases multiple rotameric conformations; (3) a further 47% of non-rotameric side chains have highly dispersed electron density, indicating potentially interconverting rotameric conformations; (4) the entropy of these side chains is close to that of side chains annotated as having more than one chi(1) rotamer in the crystallographic model; (5) many rotameric side chains with high entropy clearly show multiple conformations that are not annotated in the crystallographic model. These results indicate that modeling of side chains alternating between rotamers in the electron density is important and needs further improvement, both in structure determination and in structure prediction.

Algorithms↗

ProtBuD: a database of biological unit structures of protein families and superfamilies.

MOTIVATION: Modeling of protein interactions is often possible from known structures of related complexes. It is often time-consuming to find the most appropriate template. Hypothesized biological units (BUs) often differ from the asymmetric units and it is usually preferable to model from the BUs. RESULTS: ProtBuD is a database of BUs for all structures in the Protein Data Bank (PDB). We use both the PDBs BUs and those from the Protein Quaternary Server. ProtBuD is searchable by PDB entry, the Structural Classification of Proteins (SCOP) designation or pairs of SCOP designations. The database provides the asymmetric and BU contents of related proteins in the PDB as identified in SCOP and Position-Specific Iterated BLAST (PSI-BLAST). The asymmetric unit is different from PDB and/or Protein Quaternary Server (PQS) BUs for 52% of X-ray structures, and the PDB and PQS BUs disagree on 18% of entries. AVAILABILITY: The database is provided as a standalone program and a web server from http://dunbrack.fccc.edu/ProtBuD.php.

Amino Acid Sequence↗

Sequence comparison and protein structure prediction.

Sequence comparison is a major step in the prediction of protein structure from existing templates in the Protein Data Bank. The identification of potentially remote homologues to be used as templates for modeling target sequences of unknown structure and their accurate alignment remain challenges, despite many years of study. The most recent advances have been in combining as many sources of information as possible--including amino acid variation in the form of profiles or hidden Markov models for both the target and template families, known and predicted secondary structures of the template and target, respectively, the combination of structure alignment for distant homologues and sequence alignment for close homologues to build better profiles, and the anchoring of certain regions of the alignment based on existing biological data. Newer technologies have been applied to the problem, including the use of support vector machines to tackle the fold classification problem for a target sequence and the alignment of hidden Markov models. Finally, using the consensus of many fold recognition methods, whether based on profile-profile alignments, threading or other approaches, continues to be one of the most successful strategies for both recognition and alignment of remote homologues. Although there is still room for improvement in identification and alignment methods, additional progress may come from model building and refinement methods that can compensate for large structural changes between remotely related targets and templates, as well as for regions of misalignment.

Computational Biology↗

PISCES: recent improvements to a PDB sequence culling server.

PISCES is a database server for producing lists of sequences from the Protein Data Bank (PDB) using a number of entry- and chain-specific criteria and mutual sequence identity. Our goal in culling the PDB is to provide the longest list possible of the highest resolution structures that fulfill the sequence identity and structural quality cut-offs. The new PISCES server uses a combination of PSI-BLAST and structure-based alignments to determine sequence identities. Structure alignment produces more complete alignments and therefore more accurate sequence identities than PSI-BLAST. PISCES now allows a user to cull the PDB by-entry in addition to the standard culling by individual chains. In this scenario, a list will contain only entries that do not have a chain that has a sequence identity to any chain in any other entry in the list over the sequence identity cut-off. PISCES also provides fully annotated sequences including gene name and species. The server allows a user to cull an input list of entries or chains, so that other criteria, such as function, can be used. Results from a search on the re-engineered RCSB's site for the PDB can be entered into the PISCES server by a single click, combining the powerful searching abilities of the PDB with PISCES's utilities for sequence culling. The server's data are updated weekly. The server is available at http://dunbrack.fccc.edu/pisces.

Databases, Protein↗

MollDE: a homology modeling framework you can click with.

UNLABELLED: Molecular Integrated Development Environment (MolIDE) is an integrated application designed to provide homology modeling tools and protocols under a uniform, user-friendly graphical interface. Its main purpose is to combine the most frequent modeling steps in a semi-automatic, interactive way, guiding the user from the target protein sequence to the final three-dimensional protein structure. The typical basic homology modeling process is composed of building sequence profiles of the target sequence family, secondary structure prediction, sequence alignment with PDB structures, assisted alignment editing, side-chain prediction and loop building. All of these steps are available through a graphical user interface. MolIDE's user-friendly and streamlined interactive modeling protocol allows the user to focus on the important modeling questions, hiding from the user the raw data generation and conversion steps. MolIDE was designed from the ground up as an open-source, cross-platform, extensible framework. This allows developers to integrate additional third-party programs to MolIDE. AVAILABILITY: http://dunbrack.fccc.edu/molide/molide.php CONTACT: rl_dunbrack@fccc.edu.

Algorithms↗

Domain definition and target classification for CASP6.

Assessment of structure predictions in CASP6 was based on single domains isolated from experimentally determined structures, which were categorized into comparative modeling, fold recognition, and new fold targets. Domain definitions were defined upon visual examination of the structures with the aid of automated domain-parsing programs. Domain categorization was determined by comparison of the target structures with those in the Protein Data Bank at the time each target expired and a variety of sequence and structure-based methods to determine potential homologous relationships.

Amino Acid Sequence↗

Assessment of fold recognition predictions in CASP6.

The Sixth Community Wide Experiment on the Critical Assessment of Techniques for Protein Structure Prediction (CASP6) held in December 2004 focused on the prediction of the structures of 90 protein domains from 64 targets. Thirty-eight of these were classified as "fold recognition," defined as being similar in fold to proteins of known structure at the time of submission of the predictions. Only the "first" predictions and those longer than 20 amino acids for each domain were assessed, resulting in 4527 predictions from 165 groups. The assessment was accomplished by the use of six structure alignment programs and three scoring measures based on these alignments. The use of a variety of measures resulted in scoring insensitive to the peculiarities of any one alignment method. The top-ranked methods in the prediction of structures that were clearly homologous to proteins in the Protein Data Bank primarily used servers and other programs based on achieving a consensus of many remote homology detection and fold recognition methods. The top-ranked methods in prediction of structures less clearly related or unrelated to proteins of known structures used fragment building methods in addition to the fold recognition meta methods.

Algorithms↗

Assessment of disorder predictions in CASP6.

Natively disordered proteins or protein segments are those without stable secondary or tertiary structure in the absence of binding partners. Such disordered regions often are important functional sites in many biological processes, especially those involved in transcription, translation, and cell signaling. The prediction of such regions is therefore of great importance in focusing experimental efforts on regions of proteins that may be critical for function. In CASP6, held in 2004, twenty research groups participated in the prediction of disordered regions. Both binary predictions (ordered or disordered) and assigned scores for disorder were assessed. Several groups performed quite well in predicting regions of disorder in the X-ray and NMR structures available to the assessors. The best of these groups performed better than the best groups in CASP5, held in 2002.

Algorithms↗

Formation of MacroH2A-containing senescence-associated heterochromatin foci and senescence driven by ASF1a and HIRA.

In senescent cells, specialized domains of transcriptionally silent senescence-associated heterochromatic foci (SAHF), containing heterochromatin proteins such as HP1, are thought to repress expression of proliferation-promoting genes. We have investigated the composition and mode of assembly of SAHF and its contribution to cell cycle exit. SAHF is enriched in a transcription-silencing histone H2A variant, macroH2A. As cells approach senescence, a known chromatin regulator, HIRA, enters PML nuclear bodies, where it transiently colocalizes with HP1 proteins, prior to incorporation of HP1 proteins into SAHF. A physical complex containing HIRA and another chromatin regulator, ASF1a, is rate limiting for formation of SAHF and onset of senescence, and ASF1a is required for formation of SAHF and efficient senescence-associated cell cycle exit. These data indicate that HIRA and ASF1a drive formation of macroH2A-containing SAHF and senescence-associated cell cycle exit, via a pathway that appears to depend on flux of heterochromatic proteins through PML bodies.

Amino Acid Sequence↗

Scoring profile-to-profile sequence alignments.

Sequence alignment profiles have been shown to be very powerful in creating accurate sequence alignments. Profiles are often used to search a sequence database with a local alignment algorithm. More accurate and longer alignments have been obtained with profile-to-profile comparison. There are several steps that must be performed in creating profile-profile alignments, and each involves choices in parameters and algorithms. These steps include (1) what sequences to include in a multiple alignment used to build each profile, (2) how to weight similar sequences in the multiple alignment and how to determine amino acid frequencies from the weighted alignment, (3) how to score a column from one profile aligned to a column of the other profile, (4) how to score gaps in the profile-profile alignment, and (5) how to include structural information. Large-scale benchmarks consisting of pairs of homologous proteins with structurally determined sequence alignments are necessary for evaluating the efficacy of each scoring scheme. With such a benchmark, we have investigated the properties of profile-profile alignments and found that (1) with optimized gap penalties, most column-column scoring functions behave similarly to one another in alignment accuracy; (2) some functions, however, have much higher search sensitivity and specificity; (3) position-specific weighting schemes in determining amino acid counts in columns of multiple sequence alignments are better than sequence-specific schemes; (4) removing positions in the profile with gaps in the query sequence results in better alignments; and (5) adding predicted and known secondary structure information improves alignments.

Algorithms↗

PISCES: a protein sequence culling server.

PISCES is a public server for culling sets of protein sequences from the Protein Data Bank (PDB) by sequence identity and structural quality criteria. PISCES can provide lists culled from the entire PDB or from lists of PDB entries or chains provided by the user. The sequence identities are obtained from PSI-BLAST alignments with position-specific substitution matrices derived from the non-redundant protein sequence database. PISCES therefore provides better lists than servers that use BLAST, which is unable to identify many relationships below 40% sequence identity and often overestimates sequence identity by aligning only well-conserved fragments. PDB sequences are updated weekly. PISCES can also cull non-PDB sequences provided by the user as a list of GenBank identifiers, a FASTA format file, or BLAST/PSI-BLAST output.

Algorithms↗

A structural basis for half-of-the-sites metal binding revealed in Drosophila melanogaster porphobilinogen synthase.

Porphobilinogen synthase (PBGS) proteins fall into several distinct groups with different metal ion requirements. Drosophila melanogaster porphobilinogen synthase (DmPBGS) is the first non-mammalian metazoan PBGS to be characterized. The sequence shows the determinants for two zinc binding sites known to be present in both mammalian and yeast PBGS, proteins that differ in the exhibition of half-of-the-sites metal binding. The pH-dependent activity of DmPBGS is uniquely affected by zinc. A tight binding catalytic zinc binds at 0.5/subunit with a Kd well below microm. A second inhibitory zinc exhibits a Kd of approximately 5 microm and appears to bind at a stoichiometry of 1/subunit. A molecular model of DmPBGS suggests that the inhibitory zinc is located at a subunit interface using Cys-219 and His-10 as ligands. Zinc binding to this previously unknown inhibitory site is proposed to inhibit opening of the active site lid. As predicted, the DmPBGS mutant H10F is active but is not inhibited by zinc. H10F binds a catalytic zinc at 0.5/subunit and binds a second nonessential and noninhibitory zinc at 0.5/subunit. This result reveals a structural basis for half-of-the-sites metal binding that is consistent with a reciprocating motion model for function of oligomeric PBGS.

Animals↗

CAFASP3: the third critical assessment of fully automated structure prediction methods.

We present the results of the fully automated CAFASP3 experiment, which was carried out in parallel with CASP5, using the same set of prediction targets. CAFASP participation is restricted to fully automatic structure prediction servers. The servers' performance is evaluated by using previously announced, objective, reproducible and fully automated evaluation methods. More than 60 servers participated in CAFASP3, covering all categories of structure prediction. As in the previous CAFASP2 experiment, it was possible to identify a group of 5-10 top performing independent servers. This group of top performing independent servers produced relatively accurate models for all the 32 "Homology Modeling" targets, and for up to 43% of the 30 "Fold Recognition" targets. One of the most important results of CAFASP3 was the realization of the value of all the independent servers as a group, as evidenced by the superior performance of "meta-predictors" (defined here as predictors that make use of the output of other CAFASP servers). The performance of the best automated meta-predictors was roughly 30% higher than that of the best independent server. More significantly, the performance of the best automated meta-predictors was comparable with that of the best 5-10 human CASP predictors. This result shows that significant progress has been achieved in automatic structure prediction and has important implications to the prospects of automated structure modeling in the context of structural genomics.

Computational Biology↗

Cyclic coordinate descent: A robotics algorithm for protein loop closure.

In protein structure prediction, it is often the case that a protein segment must be adjusted to connect two fixed segments. This occurs during loop structure prediction in homology modeling as well as in ab initio structure prediction. Several algorithms for this purpose are based on the inverse Jacobian of the distance constraints with respect to dihedral angle degrees of freedom. These algorithms are sometimes unstable and fail to converge. We present an algorithm developed originally for inverse kinematics applications in robotics. In robotics, an end effector in the form of a robot hand must reach for an object in space by altering adjustable joint angles and arm lengths. In loop prediction, dihedral angles must be adjusted to move the C-terminal residue of a segment to superimpose on a fixed anchor residue in the protein structure. The algorithm, referred to as cyclic coordinate descent or CCD, involves adjusting one dihedral angle at a time to minimize the sum of the squared distances between three backbone atoms of the moving C-terminal anchor and the corresponding atoms in the fixed C-terminal anchor. The result is an equation in one variable for the proposed change in each dihedral. The algorithm proceeds iteratively through all of the adjustable dihedral angles from the N-terminal to the C-terminal end of the loop. CCD is suitable as a component of loop prediction methods that generate large numbers of trial structures. It succeeds in closing loops in a large test set 99.79% of the time, and fails occasionally only for short, highly extended loops. It is very fast, closing loops of length 8 in 0.037 sec on average.

Algorithms↗

A graph-theory algorithm for rapid protein side-chain prediction.

Fast and accurate side-chain conformation prediction is important for homology modeling, ab initio protein structure prediction, and protein design applications. Many methods have been presented, although only a few computer programs are publicly available. The SCWRL program is one such method and is widely used because of its speed, accuracy, and ease of use. A new algorithm for SCWRL is presented that uses results from graph theory to solve the combinatorial problem encountered in the side-chain prediction problem. In this method, side chains are represented as vertices in an undirected graph. Any two residues that have rotamers with nonzero interaction energies are considered to have an edge in the graph. The resulting graph can be partitioned into connected subgraphs with no edges between them. These subgraphs can in turn be broken into biconnected components, which are graphs that cannot be disconnected by removal of a single vertex. The combinatorial problem is reduced to finding the minimum energy of these small biconnected components and combining the results to identify the global minimum energy conformation. This algorithm is able to complete predictions on a set of 180 proteins with 34342 side chains in <7 min of computer time. The total chi(1) and chi(1 + 2) dihedral angle accuracies are 82.6% and 73.7% using a simple energy function based on the backbone-dependent rotamer library and a linear repulsive steric energy. The new algorithm will allow for use of SCWRL in more demanding applications such as sequence design and ab initio structure prediction, as well addition of a more complex energy function and conformational flexibility, leading to increased accuracy.

Algorithms↗

Promotion of tumor growth by murine fibroblast activation protein, a serine protease, in an animal model.

Fibroblast activation protein (FAP) is a type II integral membrane glycoprotein belonging to the serine protease family. Human FAP is selectively expressed by tumor stromal fibroblasts in epithelial carcinomas, but not by epithelial carcinoma cells, normal fibroblasts, or other normal tissues. FAP has been shown to have both in vitro dipeptidyl peptidase and collagenase activity, but its biological function in the tumor microenvironment is unknown. The modeled structure of murine FAP consists of a short cytoplasmic tail, a single hydrophobic transmembrane region, and a large extracellular domain. A seven-bladed beta-propeller domain is situated on top of the catalytic triad and may serve as a "gate" to selectively filter protein access to the catalytic site. HEK293 cells transfected to constitutively express murine FAP, when xenografted into scid mice, were 2-4 times more likely to develop s.c. tumors and showed a 10-40-fold enhancement of tumor growth compared with mock-transfected HEK293 cells. Rabbits immunized with recombinant murine FAP developed polyclonal anti-FAP antibodies that significantly inhibited murine FAP dipeptidyl peptidase activity in vitro. HT-29 xenografts treated with these inhibitory anti-FAP antisera exhibited attenuated growth compared with tumors treated with preimmunization rabbit antisera. These data demonstrate the ability of FAP to potentiate tumor growth in an animal model. Moreover, tumor growth is attenuated by antibodies that inhibit the proteolytic activity of FAP. These findings suggest a possible therapeutic role for functional inhibition of FAP activity.

Animals↗