Biochemistry. Proteins in a small world.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to Todd O Yeates.
Explore the source record for details and available documents.
Previous studies of symmetry preferences in protein crystals suggest that symmetric proteins, such as homodimers, might crystallize more readily on average than asymmetric, monomeric proteins. Proteins that are naturally monomeric can be made homodimeric artificially by forming disulfide bonds between individual cysteine residues introduced by mutagenesis. Furthermore, by creating a variety of single-cysteine mutants, a series of distinct synthetic dimers can be generated for a given protein of interest, with each expected to gain advantage from its added symmetry and to exhibit a crystallization behavior distinct from the other constructs. This strategy was tested on phage T4 lysozyme, a protein whose crystallization as a monomer has been studied exhaustively. Experiments on three single-cysteine mutants, each prepared in dimeric form, yielded numerous novel crystal forms that cannot be realized by monomeric lysozyme. Six new crystal forms have been characterized. The results suggest that synthetic symmetrization may be a useful approach for enlarging the search space for crystallizing proteins.
In a natively folded protein of moderate or larger size, the protein backbone may weave through itself in complex ways, raising questions about what sequence of events might have to occur in order for the protein to reach its native configuration from the unfolded state. A mathematical framework is presented here for describing the notion of a topological folding barrier, which occurs when a protein chain must pass through a hole or opening, formed by other regions of the protein structure. Different folding pathways encounter different numbers of such barriers and therefore different degrees of frustration. A dynamic programming algorithm finds the optimal theoretical folding path and minimal degree of frustration for a protein based on its natively folded configuration. Calculations over a database of protein structures provide insights into questions such as whether the path of minimal frustration might tend to favor folding from one or from many sites of folding nucleation, or whether proteins favor folding around the N terminus, thereby providing support for the hypothesis that proteins fold co-translationally. The computational methods are applied to a multi-disulfide bonded protein, with computational findings that are consistent with the experimentally observed folding pathway. Attention is drawn to certain complex protein folds for which the computational method suggests there may be a preferred site of nucleation or where folding is likely to proceed through a relatively well-defined pathway or intermediate. The computational analyses lead to testable models for protein folding.
CsoSCA (formerly CsoS3) is a bacterial carbonic anhydrase localized in the shell of a cellular microcompartment called the carboxysome, where it converts HCO(3)(-) to CO(2) for use in carbon fixation by ribulose-bisphosphate carboxylase/oxygenase (RuBisCO). CsoSCA lacks significant sequence similarity to any of the four known classes of carbonic anhydrase (alpha, beta, gamma, or delta), and so it was initially classified as belonging to a new class, epsilon. The crystal structure of CsoSCA from Halothiobacillus neapolitanus reveals that it is actually a representative member of a new subclass of beta-carbonic anhydrases, distinguished by a lack of active site pairing. Whereas a typical beta-carbonic anhydrase maintains a pair of active sites organized within a two-fold symmetric homodimer or pair of fused, homologous domains, the two domains in CsoSCA have diverged to the point that only one domain in the pair retains a viable active site. We suggest that this defunct and somewhat diminished domain has evolved a new function, specific to its carboxysomal environment. Despite the level of sequence divergence that separates CsoSCA from the other two subclasses of beta-carbonic anhydrases, there is a remarkable level of structural similarity among active site regions, which suggests a common catalytic mechanism for the interconversion of HCO(3)(-) and CO(2). Crystal packing analysis suggests that CsoSCA exists within the carboxysome shell either as a homodimer or as extended filaments.
Numerous diseases are characterized by the formation of insoluble, amyloid protein fibrils. Intensive investigations are beginning to unravel the detailed molecular and structural principles that underlie the spontaneous formation of these fibrils. The amyloid protein transthyretin serves as an excellent system for dissecting the conformational changes and ensuing subunit-subunit associations that lead to amyloid. One working model for tranthyretin amyloid involves the exposure of an "unprotected" edge beta strand, followed by symmetric assembly of subunits to give head-to-head and tail-to-tail protofibrils. The models and principles emerging from studies on transthyretin lead to connections to other amyloid systems.
The 2.5-A resolution crystal structure is reported for an actin dimer, composed of two protomers cross-linked along the longitudinal (or vertical) direction of the F-actin filament. The crystal structure provides an atomic resolution view of a molecular interface between actin protomers, which we argue represents a near-native interaction in the F-actin filament. The interaction involves subdomains 3 and 4 from distinct protomers. The atomic positions in the interface visualized differ by 5-10 A from those suggested by previous models of F-actin. Such differences fall within the range of uncertainties allowed by the fiber diffraction and electron microscopy methods on which previous models have been based. In the crystal, the translational arrangement of protomers lacks the slow twist found in native filaments. A plausible model of F-actin can be constructed by reintroducing the known filament twist, without disturbing significantly the interface observed in the actin dimer crystal.
In several natural settings, the standard genetic code is expanded to incorporate two additional amino acids with distinct functionality, selenocysteine and pyrrolysine. These rare amino acids can be overlooked inadvertently, however, as they arise by recoding at certain stop codons. We report a method for such recoding prediction from genomic data, using read-through similarity evaluation. A survey across a set of microbial genomes identifies almost all the known cases as well as a number of novel candidate proteins.
Thermophilic organisms flourish in varied high-temperature environmental niches that are deadly to other organisms. Recently, genomic evidence has implicated a critical role for disulfide bonds in the structural stabilization of intracellular proteins from certain of these organisms, contrary to the conventional view that structural disulfide bonds are exclusively extracellular. Here both computational and structural data are presented to explore the occurrence of disulfide bonds as a protein-stabilization method across many thermophilic prokaryotes. Based on computational studies, disulfide-bond richness is found to be widespread, with thermophiles containing the highest levels. Interestingly, only a distinct subset of thermophiles exhibit this property. A computational search for proteins matching this target phylogenetic profile singles out a specific protein, known as protein disulfide oxidoreductase, as a potential key player in thermophilic intracellular disulfide-bond formation. Finally, biochemical support in the form of a new crystal structure of a thermophilic protein with three disulfide bonds is presented together with a survey of known structures from the literature. Together, the results provide insight into biochemical specialization and the diversity of methods employed by organisms to stabilize their proteins in exotic environments. The findings also motivate continued efforts to sequence genomes from divergent organisms.
Bacterial microcompartments are primitive organelles composed entirely of protein subunits. Genomic sequence databases reveal the widespread occurrence of microcompartments across diverse microbes. The prototypical bacterial microcompartment is the carboxysome, a protein shell for sequestering carbon fixation reactions. We report three-dimensional crystal structures of multiple carboxysome shell proteins, revealing a hexameric unit as the basic microcompartment building block and showing how these hexamers assemble to form flat facets of the polyhedral shell. The structures suggest how molecular transport across the shell may be controlled and how structural variations might govern the assembly and architecture of these subcellular compartments.
The wealth of available genomic data has spawned a corresponding interest in computational methods that can impart biological meaning and context to these experiments. Traditional computational methods have drawn relationships between pairs of proteins or genes based on notions of equality or similarity between their patterns of occurrence or behavior. For example, two genes displaying similar variation in expression, over a number of experiments, may be predicted to be functionally related. We have introduced a natural extension of these approaches, instead identifying logical relationships involving triplets of proteins. Triplets provide for various discrete kinds of logic relationships, leading to detailed inferences about biological associations. For instance, a protein C might be encoded within an organism if, and only if, two other proteins A and B are also both encoded within the organism, thus suggesting that gene C is functionally related to genes A and B. The method has been applied fruitfully to both phylogenetic and microarray expression data, and has been used to associate logical combinations of protein activity with disease state phenotypes, revealing previously unknown ternary relationships among proteins, and illustrating the inherent complexities that arise in biological data.
A major focus of genome research is to decipher the networks of molecular interactions that underlie cellular function. We describe a computational approach for identifying detailed relationships between proteins on the basis of genomic data. Logic analysis of phylogenetic profiles identifies triplets of proteins whose presence or absence obey certain logic relationships. For example, protein C may be present in a genome only if proteins A and B are both present. The method reveals many previously unidentified higher order relationships. These relationships illustrate the complexities that arise in cellular networks because of branching and alternate pathways, and they also facilitate assignment of cellular functions to uncharacterized proteins.
The Genomic Disulfide Analysis Program (GDAP) provides web access to computationally predicted protein disulfide bonds for over one hundred microbial genomes, including both bacterial and achaeal species. In the GDAP process, sequences of unknown structure are mapped, when possible, to known homologous Protein Data Bank (PDB) structures, after which specific distance criteria are applied to predict disulfide bonds. GDAP also accepts user-supplied protein sequences and subsequently queries the PDB sequence database for the best matches, scans for possible disulfide bonds and returns the results to the client. These predictions are useful for a variety of applications and have previously been used to show a dramatic preference in certain thermophilic archaea and bacteria for disulfide bonds within intracellular proteins. Given the central role these stabilizing, covalent bonds play in such organisms, the predictions available from GDAP provide a rich data source for designing site-directed mutants with more stable thermal profiles. The GDAP web application is a gateway to this information and can be used to understand the role disulfide bonds play in protein stability both in these unusual organisms and in sequences of interest to the individual researcher. The prediction server can be accessed at http://www.doe-mbi.ucla.edu/Services/GDAP.
The advent of whole-genome sequencing has led to methods that infer protein function and linkages. We have combined four such algorithms (phylogenetic profile, Rosetta Stone, gene neighbor and gene cluster) in a single database--Prolinks--that spans 83 organisms and includes 10 million high-confidence links. The Proteome Navigator tool allows users to browse predicted linkage networks interactively, providing accompanying annotation from public databases. The Prolinks database and the Proteome Navigator tool are available for use online at http://dip.doe-mbi.ucla.edu/pronav.
The three-dimensional structure of the RNA-modifying enzyme, psi55 tRNA pseudouridine synthase from Mycobacterium tuberculosis, is reported. The 1.9-A resolution crystal structure reveals the enzyme, free of substrate, in two distinct conformations. The structure depicts an interesting mode of protein flexibility involving a hinged bending in the central beta-sheet of the catalytic module. Key parts of the active site cleft are also found to be disordered in the substrate-free form of the enzyme. The hinge bending appears to act as a clamp to position the substrate. Our structural data furthers the previously proposed mechanism of tRNA recognition. The present crystal structure emphasizes the significant role that protein dynamics must play in tRNA recognition, base flipping, and modification.
Genome-wide functional linkages among proteins in cellular complexes and metabolic pathways can be inferred from high throughput experimentation, such as DNA microarrays, or from bioinformatic analyses. Here we describe a method for the visualization and interpretation of genome-wide functional linkages inferred by the Rosetta Stone, Phylogenetic Profile, Operon and Conserved Gene Neighbor computational methods. This method involves the construction of a genome-wide functional linkage map, where each significant functional linkage between a pair of proteins is displayed on a two-dimensional scatter-plot, organized according to the order of genes along the chromosome. Subsequent hierarchical clustering of the map reveals clusters of genes with similar functional linkage profiles and facilitates the inference of protein function and the discovery of functionally linked gene clusters throughout the genome. We illustrate this method by applying it to the genome of the pathogenic bacterium Mycobacterium tuberculosis, assigning cellular functions to previously uncharacterized proteins involved in cell wall biosynthesis, signal transduction, chaperone activity, energy metabolism and polysaccharide biosynthesis.
A new approach to analyzing macromolecular single-crystal X-ray diffraction intensity statistics is presented. Instead of considering reflections in resolution shells, differences between local pairs of reflection intensities are taken and normalized separately. When the two reflections to be compared (having intensities I(1) and I(2), respectively) are chosen appropriately, the behavior of the parameter L = (I(1) - I(2))/(I(1) + I(2)) is insensitive to phenomena that tend to confound traditional intensity statistics, such as anisotropic diffraction and pseudo-centering. The distributions and expected values for L take simple forms when the intensity data are from ordinary crystals or from perfectly twinned specimens. The robustness of the approach is demonstrated with examples using real proteins whose diffraction data appear aberrant by other methods of intensity analysis. The new statistic is better suited than other available methods for diagnosing perfect hemihedral twinning.
The iron-containing superoxide dismutase (FeSOD) from the thermophilic cyanobacterium Thermosynechococcus elongatus has been isolated. The protein crystallizes readily and we have determined the structure to 1.6 A resolution. This is the first structural characterization of an FeSOD isolated from a cyanobacterium and one of the highest resolution FeSOD structures determined to date. The activity of the T. elongatus FeSOD has been measured both at 25 degrees C and 50 degrees C and it has been spectroscopically characterized. The T. elongatus FeSOD EPR spectra at pH 5.1, 7.5 and 10.0 are similar. This indicates that no change in the geometry of the Fe(III) site occurs over a wide range of pH. This is in contrast to the other FeSODs described in the literature.
The crystal structure at 1.54 A resolution of a double mutant of interleukin-1beta (F42W/W120F), a cytokine secreted by macrophages, was determined by multiple-wavelength anomalous dispersion (MAD) using data from highly twinned selenomethionine-modified crystals. The space group is P4(3), with unit-cell parameters a = b = 53.9, c = 77.4 A. Self-rotation function analysis and various intensity statistics revealed the presence of merohedral twinning in crystals of both the native (twinning fraction alpha approximately 0.35) and SeMet (alpha approximately 0.40) forms. Structure determination and refinement are discussed with emphasis on the possible reasons for successful phasing using untreated twinned MAD data.