PubMed Health⌕ Search

Biomedical subjects

William W Chen

Publications and source records attributed to William W Chen.

6 recordsLinked to original sources

A knowledge-based move set for protein folding.

The free energy landscape of protein folding is rugged, occasionally characterized by compact, intermediate states of low free energy. In computational folding, this landscape leads to trapped, compact states with incorrect secondary structure. We devised a residue-specific, protein backbone move set for efficient sampling of protein-like conformations in computational folding simulations. The move set is based on the selection of a small set of backbone dihedral angles, derived from clustering dihedral angles sampled from experimental structures. We show in both simulated annealing and replica exchange Monte Carlo (REMC) simulations that the knowledge-based move set, when compared with a conventional move set, shows statistically significant improved ability at overcoming kinetic barriers, reaching deeper energy minima, and achieving correspondingly lower RMSDs to native structures. The new move set is also more efficient, being able to reach low energy states considerably faster. Use of this move set in determining the energy minimum state and for calculating thermodynamic quantities is discussed.

Glycine↗

All-atom ab initio folding of a diverse set of proteins.

Natural proteins fold to a unique, thermodynamically dominant state. Modeling of the folding process and prediction of the native fold of proteins are two major unsolved problems in biophysics. Here, we show successful all-atom ab initio folding of a representative diverse set of proteins by using a minimalist transferable-energy model that consists of two-body atom-atom interactions, hydrogen bonding, and a local sequence-energy term that models sequence-specific chain stiffness. Starting from a random coil, the native-like structure was observed during replica exchange Monte Carlo (REMC) simulation for most proteins regardless of their structural classes; the lowest energy structure was close to native-in the range of 2-6 A root-mean-square deviation (rmsd). Our results demonstrate that the successful folding of a protein chain to its native state is governed by only a few crucial energetic terms.

Models, Molecular↗

Entropic stabilization of proteins and its proteomic consequences.

Evolutionary traces of thermophilic adaptation are manifest, on the whole-genome level, in compositional biases toward certain types of amino acids. However, it is sometimes difficult to discern their causes without a clear understanding of underlying physical mechanisms of thermal stabilization of proteins. For example, it is well-known that hyperthermophiles feature a greater proportion of charged residues, but, surprisingly, the excess of positively charged residues is almost entirely due to lysines but not arginines in the majority of hyperthermophilic genomes. All-atom simulations show that lysines have a much greater number of accessible rotamers than arginines of similar degree of burial in folded states of proteins. This finding suggests that lysines would preferentially entropically stabilize the native state. Indeed, we show in computational experiments that arginine-to-lysine amino acid substitutions result in noticeable stabilization of proteins. We then hypothesize that if evolution uses this physical mechanism as a complement to electrostatic stabilization in its strategies of thermophilic adaptation, then hyperthermostable organisms would have much greater content of lysines in their proteomes than comparably sized and similarly charged arginines. Consistent with that, high-throughput comparative analysis of complete proteomes shows extremely strong bias toward arginine-to-lysine replacement in hyperthermophilic organisms and overall much greater content of lysines than arginines in hyperthermophiles. This finding cannot be explained by genomic GC compositional biases or by the universal trend of amino acid gain and loss in protein evolution. We discovered here a novel entropic mechanism of protein thermostability due to residual dynamics of rotamer isomerization in native state and demonstrated its immediate proteomic implications. Our study provides an example of how analysis of a fundamental physical mechanism of thermostability helps to resolve a puzzle in comparative genomics as to why amino acid compositions of hyperthermophilic proteomes are significantly biased toward lysines but not similarly charged arginines.

Aminopeptidases↗

Lessons from the design of a novel atomic potential for protein folding.

We investigate all-atom potentials of mean force for estimating free energies in protein folding and fold recognition. We search through the space potentials and design novel atomic potentials with a random mixing approximation and a contact-correlated Gaussian approximation of decoy states. We show that the two derived potentials are highly correlated, supporting the use of the random energy model as an accurate statistical description of protein conformational states. The novel atomic potentials perform well in a Z-score and fold decoy recognition test. Furthermore, the designed atomic potential performs slightly and significantly better than atomic potentials derived under a quasi-chemical assumption. While accounting for connectivity correlations between atom types does not improve the performance of the designed potential, we show these correlations lead to ambiguities in the distribution of energetic contributions for atoms on the same residue. Within the confines of the model then, many potentials may exist which stabilize all native folds in subtly different ways. Comparison of different protein conformations under the various atomic potentials reveals both a remarkable degree of correspondence in the estimated free energies and a remarkable degree of correspondence in the identity of the contacts types that make the dominant contributions to the estimated free energies. This consistency may be interpreted as a sign that the design procedure is extracting physically meaningful quantities.

Amino Acids↗

Recognition of nucleic acid bases and base-pairs by hydrogen bonding to amino acid side-chains.

Sequence-specific protein-nucleic acid recognition is determined, in part, by hydrogen bonding interactions between amino acid side-chains and nucleotide bases. To examine the repertoire of possible interactions, we have calculated geometrically plausible arrangements in which amino acids hydrogen bond to unpaired bases, such as those found in RNA bulges and loops, or to the 53 possible RNA base-pairs. We find 32 possible interactions that involve two or more hydrogen bonds to the six unpaired bases (including protonated A and C), 17 of which have been observed. We find 186 "spanning" interactions to base-pairs in which the amino acid hydrogen bonds to both bases, in principle allowing particular base-pairs to be selectively targeted, and nine of these have been observed. Four calculated interactions span the Watson-Crick pairs and 15 span the G:U wobble pair, including two interesting arrangements with three hydrogen bonds to the Arg guanidinum group that have not yet been observed. The inherent donor-acceptor arrangements of the bases support many possible interactions to Asn (or Gln) and Ser (or Thr or Tyr), few interactions to Asp (or Glu) even though several already have been observed, and interactions to U (or T) only if the base is in an unpaired context, as also observed in several cases. This study highlights how complementary arrangements of donors and acceptors can contribute to base-specific recognition of RNA, predicts interactions not yet observed, and provides tools to analyze proposed contacts or design novel interactions.

Amino Acids↗