PubMed HealthSearch

Biomedical subjects

C Sander

Publications and source records attributed to C Sander.

At least 55 records · Page 3Linked to original sources

The FSSP database: fold classification based on structure-structure alignment of proteins.

The FSSP database presents a continuously updated classification of 3-D protein folds based on an all-against-all comparison of structures currently in the Protein Data Bank (PDB) [Bernstein et al. (1977) J. Mol. Biol., 112, 535- 542]. The database currently contains an extended structural family for each of 600 representative protein chains which have <25% mutual sequence identity. The results of the exhaustive pairwise structure comparisons are reported in the form of a fold tree generated by hierarchical clustering and as a series of structurally representative sets of folds at varying levels of uniqueness. For each query structure from the representative set, there is a database entry containing structure-structure alignments with its structural neighbours in the representative set and its sequence homologs in the PDB. All alignments are based purely on the 3-D co-ordinates of the proteins and are derived by an automatic structure comparison program (Dali). The FSSP database is accessible electronically on the World Wide Web and by anonymous ftp.

Amino Acid Sequence

Sequencing and analysis of a 35.4 kb region on the right [corrected] arm of chromosome IV from Saccharomyces cerevisiae reveal 23 open reading frames.

The complete DNA sequence of cosmid clone 31A5 containing a 35 452 bp segment from the right [corrected] arm of chromosome IV from Saccharomyces cerevisiae, was determined from an ordered set of subclones in combination with primer walking on the cosmid. The sequence contains 23 open reading frames (ORFs) of more than 100 amino acid residues and the tRNA-Va12a gene. Five ORFs corresponded to the known yeast genes SNQ2, SES1, GCV1, RPL2B and RPS18A. The DNA sequence for RPS18A is interrupted by an intron. One ORF corresponded to a part of the yeast gene HEX2 at the end of the cosmid insert. Four ORFs encoded putative proteins which showed strong homologies to other previously known proteins, three of yeast origin and one of non-yeast origin. Two ORFs were classified as having borderline homologies: one had similarity to two protein families and another to two protein products of unknown function from other species. The remaining 11 ORFs bore no significant similarity to any published protein.

Amino Acid Sequence

Positioning hydrogen atoms by optimizing hydrogen-bond networks in protein structures.

A method is presented that positions polar hydrogen atoms in protein structures by optimizing the total hydrogen bond energy. For this goal, an empirical hydrogen bond force field was derived from small molecule crystal structures. Bifurcated hydrogen bonds are taken into account. The procedure also predicts ionization states of His, Asp, and Glu residues. During optimization, side-chain conformations of His, Gln, and Asn residues are allowed to change their last chi angle by 180 degrees to compensate for crystallographic misassignments. Crystal structure symmetry is taken into account where appropriate. The results can have significant implications for molecular dynamics simulations, protein engineering, and docking studies. The largest impact, however, is in protein structure verification: over 85% of protein structures tested can be improved by using our procedure.

Amino Acids

Computational comparisons of model genomes.

Complete genomes from model organisms provide new challenges for computational molecular biology. Novel questions emerge from the genome data obtained from the functional prediction of thousands of gene products. In this review, we present some approaches to the computational comparison of genomes, based on sequence and text analysis, and comparisons of genome composition and gene order.

Biotechnology

A comparison of structural and dynamic properties of different simulation methods applied to SH3.

The dynamic and static properties of molecular dynamics simulations using various methods for treating solvent were compared. The SH3 protein domain was chosen as a test case because of its small size and high surface-to-volume ratio. The simulations were analyzed in structural terms by examining crystal packing, distribution of polar residues, and conservation of secondary structure. In addition, the "essential dynamics" method was applied to compare each of the molecular dynamics trajectories with a full solvent simulation. This method proved to be a powerful tool for the comparison of large concerted atomic motions in SH3. It identified methods of simulation that yielded significantly different dynamic properties compared to the full solvent simulation. Simulating SH3 using the stochastic dynamics algorithm with a vacuum (reduced charge) force field produced properties close to those of the full solvent simulation. The application of a recently described solvation term did not improve the dynamic properties. The large concerted atomic motions in the full solvent simulation as revealed by the essential dynamics method were analyzed for possible biological implications. Two loops, which have been shown to be involved in ligand binding, were seen to move in concert to open and close the ligand-binding site.

Algorithms

The PDBFINDER database: a summary of PDB, DSSP and HSSP information with added value.

MOTIVATION: The Protein Data Bank currently contains more than 4700 protein coordinate sets. It is often desirable to make a selection from these files based on a criterion like R-factor, experimental method, length of the amino acid sequence, or the number of homologous sequences in SWISSPROT. Doing this using the distributed form of the Protein Data Bank can be a tedious task, because (1) this requires reading one file for every single entry, and (2) not all of the information is present in a consistent computer readable way in all of the entries. RESULTS: The PDBFINDER database provides an easy to interpret file containing summary information about all Protein Data Bank files. Summary information from the DSSP (Definition of Secondary Structure of Proteins) and HSSP (Homology derived Secondary Structure of Proteins) databases is also included. Furthermore, where essential data were missing from the Protein Data Bank file, this information has been retrieved from the original literature. AVAILABILITY: The latest version of the PDBFINDER database can be downloaded by anonymous ftp from swift.embl-heidelberg.de, directory:/pdbfinder. CONTACT: E-mail address hooft@embl-heidelberg.de.

Amino Acid Sequence

The prediction of protein contacts from multiple sequence alignments.

We have studied the question of how much extra predictive power the correlated mutational behaviour of pairs of amino acid residues separated along a sequence has concerning the likelihood of those residues being in contact in the folded protein. The mutational behaviour is deduced from multiple sequence alignments. Our findings are that there is, indeed, some valuable information available from this source and that it is sufficient to make a significant improvement in our ability to predict contacts, when compared with earlier methods that do not take into account the correlations between the mutations. This improvement is approximately twice as large as can be obtained by the more economical method of simply averaging pair preferences over the same sequence alignment. Even when using a method based on pair preferences, a further significant improvement can be made by penalizing more variable regions (on the reasonable assumption that invariant residues are relatively more likely to be in contact), though we have found no way of improving the pair preference method to the extent that it matches the method based on correlated behaviour. Our new method is thought to be the best data-based method of contact prediction developed so far, achieving, on average, an improvement over a random (i.e. information-free) prediction of a factor of five when the number of contacts predicted is chosen to match the number that actually occur.

Algorithms

Bridging the protein sequence-structure gap by structure predictions.

The problem of accurately predicting protein three-dimensional structure from sequence has yet to be solved. Recently, several new and promising methods that work in one, two, or three dimensions have invigorated the field. Modeling by homology can yield fairly accurate three-dimensional structures for approximately 25% of the currently known protein sequences. Techniques for cooperatively fitting sequences into known three-dimensional folds, called threading methods, can increase this rate by detecting very remote homologies in favorable cases. Prediction of protein structure in two dimensions, i.e. prediction of interresidue contacts, is in its infancy. Prediction tools that work in one dimension are both mature and generally applicable; they predict secondary structure, residue solvent accessibility, and the location of transmembrane helices with reasonable accuracy. These and other prediction methods have gained immensely from the rapid increase of information in publicly accessible databases. Growing databases will lead to further improvements of prediction methods and, thus, to narrowing the gap between the number of known protein sequences and known protein structures.

Amino Acid Sequence

Mutation of the ras genes is a rare genetic event in the histologic transformation of follicular lymphoma.

The role of ras gene mutations in the progression of follicular lymphoma has been ascertained by SSCP-PCR and sequencing. A total of 40 transformed lymphomas were studied, 16 of which had a matched preceding low-grade biopsy. Only one transformed lymphoma was found to have a missense mutation at codon 12 of N-ras, resulting in an amino acid change of glycine to serine. We conclude that mutation within the ras gene family is a rare event in the transformation of follicular lymphoma.

Base Sequence

Macromolecular structure information and databases. The EU BRIDGE Database Project Consortium.

The current status and future outlook of macromolecular structure databases and information handling, with particular reference to European databases, are reviewed. Issues concerning the efficiency with which data are represented, validated, archived and accessed are discussed in view of the fast growing body of information on structures of biological macromolecules.

Databases, Factual

A sequence property approach to searching protein databases.

Currently available sequence alignment programs are generally not capable of detecting functional and structural homologs in the twilight zone of sequence similarity, i.e. when the sequence identity falls below about 25%. Here we attempt to detect such weak similarities using an approach based on a notion of protein sequence similarity radically different from that used in sequential alignment. The approach defines protein sequence dissimilarity (or distance) as a weighted sum of differences of compositional properties such as singlet and doublet amino acid composition, molecular weight, isoelectric point (protein property search or PropSearch). With PropSearch, either single sequences can be used for a database query, or multiple sequences can be merged into an "average" sequence reflecting the average composition of a protein family. First, we show that members of structural protein families have a low mutual PropSearch distance when the weights are optimized to discriminate maximally between structural families. Second, we demonstrate the results of database searches using the PropSearch method. Such searches are very rapid when scanning a preprocessed database and do not require alignments. In cases in which conventional alignment tools fail to detect similarities, PropSearch can be used to generate hypotheses about possible structural or functional relationships between a new sequence and sequences in the database.

Algorithms

Investigating the structural determinants of the p21-like triphosphate and Mg2+ binding site.

Amongst the superfamily of nucleotide binding proteins, the classical mononucleotide binding fold (CMBF), is the one that has been best characterized structurally. The common denominator of all the members is the triphosphate/Mg2+ binding site, whose signature has been recognized as two structurally conserved stretches of residues: the Kinase 1 and 2 motifs that participate in triphosphate and Mg2+ binding, respectively. The Kinase 1 motif is borne by a loop (the P-loop), whose structure is conserved throughout the whole CMBF family. The low sequence similarity between the different members raises questions about which interactions are responsible for the active structure of the P-loop. What are the minimal requirements for the active structure of the P-loop? Why is the P-loop structure conserved despite the diverse environments in which it is found? To address this question, we have engineered the Kinase 1 and 2 motifs into a protein that has the CMBF and no nucleotide binding activity, the chemotactic protein from Escherichia coli, CheY. The mutant does not exhibit any triphosphate/Mg2+ binding activity. The crystal structure of the mutant reveals that the engineered P-loop is in a different conformation than that found in the CMBF. This demonstrates that the native structure of the P-loop requires external interactions with the rest of the protein. On the basis of an analysis of the conserved tertiary contacts of the P-loop in the mononucleotide binding superfamily, we propose a set of residues that could play an important role in the acquisition of the active structure of the P-loop.

Amino Acid Sequence

Evolutionary link between glycogen phosphorylase and a DNA modifying enzyme.

We report here an unexpected similarity in three-dimensional structure between glucosyltransferases involved in very different biochemical pathways, with interesting evolutionary and functional implications. One is the DNA modifying enzyme beta-glucosyltransferase from bacteriophage T4, alias UDP-glucose:5-hydroxymethyl-cytosine beta-glucosyltransferase. The other is the metabolic enzyme glycogen phosphorylase, alias 1.4-alpha-D-glucan:orthophosphate alpha-glucosyltransferase. Structural alignment revealed that the entire structure of beta-glucosyltransferase is topographically equivalent to the catalytic core of the much larger glycogen phosphorylase. The match includes two domains in similar relative orientation and connecting helices, with a positional root-mean-square deviation of only 3.4 A for 256 C alpha atoms. An interdomain rotation seen in the R- to T-state transition of glycogen phosphorylase is similar to that observed in beta-glucosyltransferase on substrate binding. Although not a single functional residue is identical, there are striking similarities in the spatial arrangement and in the chemical nature of the substrates. The functional analogies are (beta-glucosyltransferase-glycogen phosphorylase): ribose ring of UDP-pyridoxal ring of pyridoxal phosphate co-enzyme; phosphates of UDP-phosphate of co-enzyme and reactive orthophosphate; glucose unit transferred to DNA-terminal glucose unit extracted from glycogen. We anticipate the discovery of additional structurally conserved members of the emerging glucosyltransferase superfamily derived from a common ancient evolutionary ancestor of the two enzymes.

Amino Acid Sequence