PubMed Health⌕ Search

Biomedical subjects

Andrej Sali

Publications and source records attributed to Andrej Sali.

11 recordsLinked to original sources

From words to literature in structural proteomics.

Technical advances on several frontiers have expanded the applicability of existing methods in structural biology and helped close the resolution gaps between them. As a result, we are now poised to integrate structural information gathered at multiple levels of the biological hierarchy - from atoms to cells - into a common framework. The goal is a comprehensive description of the multitude of interactions between molecular entities, which in turn is a prerequisite for the discovery of general structural principles that underlie all cellular processes.

Animals↗

ModView, visualization of multiple protein sequences and structures.

SUMMARY: We describe ModView, a web application for visualization of multiple protein sequences and structures. ModView integrates a multiple structure viewer, a multiple sequence alignment editor, and a database querying engine. It is possible to interactively manipulate hundreds of proteins, to visualize conservative and variable residues, active and binding sites, fragments, and domains in protein families, as well as to display large macromolecular complexes such as ribosomes or viruses. As a Netscape plug-in, ModView can be included in HTML pages along with text and figures, which makes it useful for teaching and presentations. ModView is also suitable as a graphical interface to various databases because it can be controlled through JavaScript commands and called from CGI scripts. AVAILABILITY: ModView is available at http://guitar.rockefeller.edu/modview.

Database Management Systems↗

Use of single point mutations in domain I of beta 2-glycoprotein I to determine fine antigenic specificity of antiphospholipid autoantibodies.

Autoantibodies against beta(2)-glycoprotein I (beta(2)GPI) appear to be a critical feature of the antiphospholipid syndrome (APS). As determined using domain deletion mutants, human autoantibodies bind to the first of five domains present in beta(2)GPI. In this study the fine detail of the domain I epitope has been examined using 10 selected mutants of whole beta(2)GPI containing single point mutations in the first domain. The binding to beta(2)GPI was significantly affected by a number of single point mutations in domain I, particularly by mutations in the region of aa 40-43. Molecular modeling predicted these mutations to affect the surface shape and electrostatic charge of a facet of domain I. Mutation K19E also had an effect, albeit one less severe and involving fewer patients. Similar results were obtained in two different laboratories using affinity-purified anti-beta(2)GPI in a competitive inhibition ELISA and with whole serum in a direct binding ELISA. This study confirms that anti-beta(2)GPI autoantibodies bind to domain I, and that the charged surface patch defined by residues 40-43 contributes to a dominant target epitope.

Amino Acid Substitution↗

Probing the specificity of a trypanosomal aromatic alpha-hydroxy acid dehydrogenase by site-directed mutagenesis.

The aromatic l-alpha-hydroxy acid dehydrogenase (AHDAH) from Trypanosoma cruzi has over 50% sequence identity with cytosolic malate dehydrogenases (cMDHs), yet it is unable to reduce oxaloacetate. Molecular modeling of the three-dimensional structure of AHADH using the pig cMDH as template directed the construction of several mutants. AHADH shares with MDHs the essential catalytic residues H195 and R171 (using Eventoff's numbering). The AHADH A102R mutant became able to reduce oxaloacetate, while remaining fully active towards aromatic alpha-oxoacids. The Y237G mutant diminished its affinity for all of the natural substrates, whereas the double mutant A102R/Y237G was more active than Y237G and had similar activity with oxaloacetate and with aromatic substrates. The present results reinforce our proposal that AHADH arose by a moderate number of point mutations from a cMDH no longer present in the parasite.

Alcohol Oxidoreductases↗

RasGRP4, a new mast cell-restricted Ras guanine nucleotide-releasing protein with calcium- and diacylglycerol-binding motifs. Identification of defective variants of this signaling protein in asthma, mastocytosis, and mast cell leukemia patients and demonstration of the importance of RasGRP4 in mast cell development and function.

A cDNA was isolated from interleukin 3-developed, mouse bone marrow-derived mast cells (MCs) that contained an insert (designated mRasGRP4) that had not been identified in any species at the gene, mRNA, or protein level. By using a homology-based cloning approach, the approximately 2.6-kb hRasGRP4 transcript was also isolated from the mononuclear progenitors residing in the peripheral blood of normal individuals. This transcript information was then used to locate the RasGRP4 gene in the mouse and human genomes, to deduce its exon/intron organization, and then to identify 10 single nucleotide polymorphisms in the human gene that result in 5 amino acid differences. The >15-kb hRasGRP4 gene consists of 18 exons and resides on a region of chromosome 19q13.1 that had not been sequenced by the Human Genome Project. Human and mouse MCs and their progenitors selectively express RasGRP4, and this new intracellular protein contains all of the domains present in the RasGRP family of guanine nucleotide exchange factors even though it is <50% identical to its closest homolog. Recombinant RasGRP4 can activate H-Ras in a cation-dependent manner. Transfection experiments also suggest that RasGRP4 is a diacylglycerol/phorbol ester receptor. Transcript analysis of an asthma patient, a mastocytosis patient, and the HMC-1 cell line derived from a MC leukemia patient revealed the presence of substantial amounts of non-functional forms of hRasGRP4 due to an inability to remove intron 5 in the precursor transcript. Because only abnormal forms of hRasGRP4 were identified in the HMC-1 cell line, this immature MC progenitor was used to address the function of RasGRP4 in MCs. HMC-1 leukemia cells differentiated and underwent granule maturation when induced to express a normal form of RasGRP4. Thus, RasGRP4 plays an important role in the final stages of MC development.

Amino Acid Sequence↗

MODBASE, a database of annotated comparative protein structure models.

MODBASE (http://guitar.rockefeller.edu/modbase) is a relational database of annotated comparative protein structure models for all available protein sequences matched to at least one known protein structure. The models are calculated by MODPIPE, an automated modeling pipeline that relies on PSI-BLAST, IMPALA and MODELLER. MODBASE uses the MySQL relational database management system for flexible and efficient querying, and the MODVIEW Netscape plugin for viewing and manipulating multiple sequences and structures. It is updated regularly to reflect the growth of the protein sequence and structure databases, as well as improvements in the software for calculating the models. For ease of access, MODBASE is organized into different datasets. The largest dataset contains models for domains in 304 517 out of 539 171 unique protein sequences in the complete TrEMBL database (23 March 2001); only models based on significant alignments (PSI-BLAST E-value < 10(-4)) and models assessed to have the correct fold are included. Other datasets include models for target selection and structure-based annotation by the New York Structural Genomics Research Consortium, models for prediction of genes in the Drosophila melanogaster genome, models for structure determination of several ribosomal particles and models calculated by the MODWEB comparative modeling web server.

Animals↗

Statistical potentials for fold assessment.

A protein structure model generally needs to be evaluated to assess whether or not it has the correct fold. To improve fold assessment, four types of a residue-level statistical potential were optimized, including distance-dependent, contact, Phi/Psi dihedral angle, and accessible surface statistical potentials. Approximately 10,000 test models with the correct and incorrect folds were built by automated comparative modeling of protein sequences of known structure. The criterion used to discriminate between the correct and incorrect models was the Z-score of the model energy. The performance of a Z-score was determined as a function of many variables in the derivation and use of the corresponding statistical potential. The performance was measured by the fractions of the correctly and incorrectly assessed test models. The most discriminating combination of any one of the four tested potentials is the sum of the normalized distance-dependent and accessible surface potentials. The distance-dependent potential that is optimal for assessing models of all sizes uses both C(alpha) and C(beta) atoms as interaction centers, distinguishes between all 20 standard residue types, has the distance range of 30 A, and is derived and used by taking into account the sequence separation of the interacting atom pairs. The terms for the sequentially local interactions are significantly less informative than those for the sequentially nonlocal interactions. The accessible surface potential that is optimal for assessing models of all sizes uses C(beta) atoms as interaction centers and distinguishes between all 20 standard residue types. The performance of the tested statistical potentials is not likely to improve significantly with an increase in the number of known protein structures used in their derivation. The parameters of fold assessment whose optimal values vary significantly with model size include the size of the known protein structures used to derive the potential and the distance range of the accessible surface potential. Fold assessment by statistical potentials is most difficult for the very small models. This difficulty presents a challenge to fold assessment in large-scale comparative modeling, which produces many small and incomplete models. The results described in this study provide a basis for an optimal use of statistical potentials in fold assessment.

Algorithms↗

Reliability of assessment of protein structure prediction methods.

The reliability of ranking of protein structure modeling methods is assessed. The assessment is based on the parametric Student's t test and the nonparametric Wilcox signed rank test of statistical significance of the difference between paired samples. The approach is applied to the ranking of the comparative modeling methods tested at the fourth meeting on Critical Assessment of Techniques for Protein Structure Prediction (CASP). It is shown that the 14 CASP4 test sequences may not be sufficient to reliably distinguish between the top eight methods, given the model quality differences and their standard deviations. We suggest that CASP needs to be supplemented by an assessment of protein structure prediction methods that is automated, continuous in time, based on several criteria applied to a large number of models, and with quantitative statistical reliability assigned to each characterization.

Computer Simulation↗

Evolution and physics in comparative protein structure modeling.

From a physical perspective, the native structure of a protein is a consequence of physical forces acting on the protein and solvent atoms during the folding process. From a biological perspective, the native structure of proteins is a result of evolution over millions of years. Correspondingly, there are two types of protein structure prediction methods, de novo prediction and comparative modeling. We review comparative protein structure modeling and discuss the incorporation of physical considerations into the modeling process. A good starting point for achieving this aim is provided by comparative modeling by satisfaction of spatial restraints. Incorporation of physical considerations is illustrated by an inclusion of solvation effects into the modeling of loops.

Biophysical Phenomena↗

LigBase: a database of families of aligned ligand binding sites in known protein sequences and structures.

A database comprising all ligand-binding sites of known structure aligned with all related protein sequences and structures is described. Currently, the database contains approximately 50000 ligand-binding sites for small molecules found in the Protein Data Bank (PDB). The structure-structure alignments are obtained by the Combinatorial Extension (CE) program (Shindyalov and Bourne, Protein Eng., 11, 739-747, 1998) and sequence-structure alignments are extracted from the ModBase database of comparative protein structure models for all known protein sequences (Sanchez et al., Nucleic Acids Res., 28, 250-253, 2000). It is possible to search for binding sites in LigBase by a variety of criteria. LigBase reports summarize ligand data including relevant structural information from the PDB file, such as ligand type and size, and contain links to all related protein sequences in the TrEMBL database. Residues in the binding sites are graphically depicted for comparison with other structurally defined family members. LigBase provides a resource for the analysis of families of related binding sites.

Binding Sites↗