PubMed HealthSearch

Biomedical subjects

P Argos

Publications and source records attributed to P Argos.

At least 19 recordsLinked to original sources

Recognition of distantly related protein sequences using conserved motifs and neural networks.

A sensitive technique for protein sequence motif recognition based on neural networks has been developed. It involves three major steps. (1) At each appropriate alignment position of a set of N matched sequences, a set of N aligned oligopeptides is specified with preselected window length. N neural nets are subsequently and successively trained on N-1 amino acid spans after eliminating each ith oligopeptide. A test for recognition of each of the ith spans is performed. The average neural net recognition over N such trials is used as a measure of conservation for the particular windowed region of the multiple alignment. This process is repeated for all possible spans of given length in the multiple alignment. (2) The M most conserved regions are regarded as motifs and the oligopeptides within each are used to train intensively M individual neural networks. (3) The M networks are then applied in a search for related primary structures in a databank of known protein sequences. The oligopeptide spans in the database sequence with strongest neural net output for each of the M networks are saved and then scored according to the output signals and the proper combination that follows the expected N- to C-terminal sequence order. The motifs from the database with highest similarity scores can then be used to retrain the M neural nets, which can be subsequently utilized for further searches in the databank, thus providing even greater sensitivity to recognize distant familial proteins. This technique was successfully applied to the integrase, DNA-polymerase and immunoglobulin families.

Aldehyde Dehydrogenase

Anatomy and evolution of proteins displaying the viral capsid jellyroll topology.

In this paper the anatomy of 25 structures containing a jellyroll motif, consisting of eight antiparallel beta-strands forming a so-called beta-barrel, was investigated. This involved performing a careful structural alignment based on hydrogen bonds for the equivalent regions of the tertiary folds and a subsequent analysis of conserved amino acids, equivalenced residue-residue contacts, and various parameters describing the size, shape and other geometrical characteristics of these regions. It was found that the jellyroll motif is best viewed as a two-sheet wedge structure rather than a barrel. The more conserved parameters are discussed. A model of evolutionary development for the jellyroll fold in the various protein and viral structures is proposed.

Amino Acid Sequence

Prediction of protein folding pathways.

Recent 1H nuclear magnetic resonance (n.m.r.) hydrogen exchange experiments on five different proteins have delineated the secondary structures formed in trapped, partially folded intermediates. The early forming structural elements are identifiable through a technique described in this work to predict folding pathways. The method assumes that the sequential selection of structural fragments such as alpha-helices and beta-strands involved in the folding process is founded upon the maximal burial of solvent accessible surface from both the formation of internal structure and substructure association. The substructural elements were defined objectively by major changes in main-chain direction. The predicted folding pathways are in complete correspondence with the n.m.r. results in that the formed structural fragments found in the folding intermediates are those predicted earliest in the pathways. The technique was also applied to proteins of known tertiary structure and with fold similar to one of the five proteins examined by 1H n.m.r. The pathways for these structures also showed general consistency with the n.m.r. observations, suggesting conservation of a secondary structural framework or molten globule about which folding nucleates and proceeds.

Amino Acid Sequence

Optimal protocol and trajectory visualization for conformational searches of peptides and proteins.

Conformational searches by molecular dynamics and different types of Monte Carlo or build-up methods usually aim to find the lowest-energy conformation. However, this is often misleading, as the energy functions used in conformational calculations are imprecise. For instance, though positions of local minima defined by the repulsive part of the Lennard-Jones potential are usually altered only slightly by functional modification, the relative depths of the minima could change significantly. Thus, the purpose of conformational searches and, correspondingly, performance criteria should be reformulated and appropriate methods found to extract different local minima from the search trajectory and allow visualization in the search space. Attempts at convergence to the lowest-energy structure should be replaced with efforts to visit a maximum number of different local energy minima with energies within a certain range. We use this quantitative criterion consistently to evaluate performances of different search procedures. To utilize information generated in the course of simulation, a "stack" of low energy conformations is created and stored. It keeps track of variables and visit numbers for the best representatives of different conformational families. To visualize the search, projection of multidimensional walks onto a principal plane defined by a set of reference structures is used. With Met-enkephalin as a structural example and a Monte Carlo procedure combined with energy minimization (MCM) as a basic search method, we analyzed the influence on search efficiency of different characteristics as temperature schedules, the step size for variable modification, constrained random step and response mechanisms to search difficulties. Simulated annealing MCM had comparable efficiency with MCM at constant and elevated temperature (about 600 K). Constraining the randomized choice of side-chain chi angles to optimal values (rotamers) on every MCM step did not improve, but rather worsened, the search efficiency. Two low-energy Met-enkephalin conformations with parallel Tyr1 and Phe4 rings, a gamma-turn around the Gly2 residue, and Phe4 and Met5 side-chains forming together a compact hydrophobic cluster were found and are suggested as possible structural candidates for interaction with a receptor or a membrane.

Amino Acid Sequence

Analysis of insertions/deletions in protein structures.

An analysis of insertions and deletions (indels) occurring in a databank of multiple sequence alignments based on protein tertiary structure is reported. Indels prefer to be short (1 to 5 residues). The average intervening sequence length between them versus the percentage of residue identity in pairwise alignments shows an exponential behaviour, suggesting a stochastic process such that nearly every loop in an ancestral structure is a possible target for indels during evolution. The results also suggest a limit to the average size of indels accommodated by protein structures. The preferred indel conformations are reverse turn and coil as are the preferred conformations at the indel edges (N- and C-terminal sides). Interruptions in helices and strands were observed as very rare events.

Amino Acid Sequence

Overseer: a nucleotide sequence searching tool.

Overseer is a computer program that searches databases of nucleic acid sequences for objects of interest to the user. Such objects may consist of any number of simpler building blocks such as repeats, palindromes or stem-loops, strings of particular bases with or without mismatches, etc. Written in standard Pascal, this program runs under Unix and VMS and should also run under other operating systems. A simple interface allows the user to generate interactively a file containing a description of the target to be found. The searching program runs non-interactively, processing the information from the file and searching the sequences. The results are output to a file. Search capabilities are quite flexible and the code is designed to be modified. Since the framework of the program is simple, adding new modules to search for new target types as the need arises is possible.

Algorithms

Searching for distantly related protein sequences in large databases by parallel processing on a transputer machine.

AliMac is an implementation of a sensitive sequence alignment algorithm on a parallel computer. The method achieves reliable alignments for very distantly related sequences from a combined use of amino acid exchange weights and physicochemical characteristics. The algorithm is computing intensive and its usage on conventional computers is limited to a relatively small number of sequences. The parallel implementation uses a Macintosh IIcx host computer and 21 transputers and achieves 22 times the speed of a VAX 8650 at a fraction of the cost. This paper describes the AliMac hardware and software and discusses problems and peculiarities of parallel implementations, especially with transputers. Finally, several popular sequence alignment algorithms are compared in their ability to detect distantly related sequences in searching large databases.

Algorithms

OBSTRUCT: a program to obtain largest cliques from a protein sequence set according to structural resolution and sequence similarity.

A program OBSTRUCT has been developed to obtain the largest possible subset according to specific constraints from a set of protein sequences whose tertiary structures have been determined crystallographically. The user can request a range in sequence similarity level and/or structural resolution. The program optionally includes sequences with known three-dimensional folds elicited from NMR data.

Protein Conformation

A data bank merging related protein structures and sequences.

A data collection which merges protein structural and sequence information is described. Structural superpositions amongst proteins with similar main-chain fold were performed or collected from the literature. Sequences taken from the protein primary structure databases were associated with the multiple structural alignments providing they were at least 50% homologous in residue identity to one of the structural sequences and at least 50% of the structural sequence residues were alignable. Such restrictions allow reasonable confidence that the primary sequences share the conformation of the tertiary structural templates, except in the less conserved loop regions. Multiple structural superpositions were collected for 38 familial groups containing a total of 209 tertiary structures; 45 structures had no superposable mates and were used individually. Other information is also provided as main-chain and side-chain conformational angles, secondary structural assignments and the like. Wedding the primary and tertiary structural data resulted in an 8-fold increase of data bank sequence entries over those associated with the known three-dimensional architectures alone.

Amino Acid Sequence

Potential of genetic algorithms in protein folding and protein engineering simulations.

Genetic algorithms are very efficient search mechanisms which mutate, recombine and select amongst tentative solutions to a problem until a near optimal one is achieved. We introduce them as a new tool to study proteins. The identification and motivation for different fitness functions is discussed. The evolution of the zinc finger sequence motif from a random start is modelled. User specified changes of the lambda repressor structure were simulated and critical sites and exchanges for mutagenesis identified. Vast conformational spaces are efficiently searched as illustrated by the ab initio folding of a model protein of a four beta strand bundle. The genetic algorithm simulation which mimicked important folding constraints as overall hydrophobic packaging and a propensity of the betaphilic residues for trans positions achieved a unique fold. Cooperativity in the beta strand regions and a length of 3-5 for the interconnecting loops was critical. Specific interaction sites were considerably less effective in driving the fold.

Algorithms

Identification of proteins in sequence databases from amino acid composition data.

Having obtained the amino acid composition of a protein, chemists and molecular biologists may wish to identify the protein from this data alone. In general such data will have errors associated with them and the length of the protein may be known only approximately or not at all. In this paper a method is described which enables searching of protein sequence databases for sequences or fragments of sequences which have a composition similar to the one being sought. Such searches are generally quite discriminating as shown by the examples provided. This method has been implemented as part of the computer program Scrutineer and is being freely distributed. It is simple to use.

Algorithms

Side-chain clusters in protein structures and their role in protein folding.

A method has been developed to detect dense clusters of residue side-chains in proteins, where contact is based upon the percentage of the maximum possible for a given residue type. The clusters represent protein sites with the highest degree of interaction amongst their member residues, while contacts with the environment surrounding the cluster are lower in number. The method has been applied to three distinct structural sets of proteins to check for consistency: mixed alpha-helical/beta-sheet proteins, all beta-strand proteins, and all alpha-helical proteins. A number of cluster features generated from these sets are of general interest for protein folding. (1) A majority of the clusters, comprising three to four residues on average, are localized near the protein surfaces and not within the protein cores. (2) The clusters have preferences for the N- and C-terminal ends of alpha-helices and beta-strands in alpha/beta and alpha-proteins, while beta-proteins utilize the middle strand regions more often. A number of clusters connect three or more beta-strands and/or alpha-helices. (3) More than half of the clusters display residue pairs with oppositely charged atoms within 4.5 A of each other. (4) The residue composition of the clusters does not show correlation with hydrophobicity measures but rather with side-chain volume and surface. The highly preferred cluster residues are (in order of decreasing preference) Trp, His, Arg, Tyr, Glu, Gln and Phe. Clusters with extensive internal contacts in related haemoglobin and immunoglobulin tertiary structures show respective conservation. Several examples illustrate "strategic" folding positions in proteins that often bring together a number of sheets and/or helices, suggesting a folding model in which largely preformed secondary structures are joined together in a cluster induced collapse. Alternatively, the clusters may form at some stage in the folding process to reduce considerably the searchable conformational space and help maintain the proper folding pathway. The clusters also provide hints for site-directed mutagenesis and protein engineering experiments as they are also suggested to be important for structural stability.

Amino Acid Sequence

Homology between IRE-BP, a regulatory RNA-binding protein, aconitase, and isopropylmalate isomerase.

Iron-responsive elements (IREs) are regulatory RNA elements which serve as specific binding sites for the IRE-binding protein (IRE-BP). Interaction between IREs and IRE-BP induces repression of ferritin mRNA translation and transferrin receptor mRNA stabilization. We describe the identification of extensive amino acid sequence homology between IRE-BP and two known isomerases, aconitase and isopropylmalate (IPM) isomerase. We discuss the implications of this observation with regard to structure/function relationships of IRE-BP. The structural conservation between a regulatory RNA-binding protein and two enzymes involved in intermediary metabolism provides a surprising example of the functional flexibility in biological structures.

Aconitate Hydratase

Motif recognition and alignment for many sequences by comparison of dot-matrices.

Calculation of dot-matrices is a widespread tool in the search for sequence similarities. When sequences are distant, even this approach may fail to point out common regions. If several plots calculated for all members of a sequence set consistently displayed a similarity between them, this would increase its credibility. We present an algorithm to delineate dot-plot agreement. A novel procedure based on matrix multiplication is developed to identify common patterns and reliably aligned regions in a set of distantly related sequences. The algorithm finds motifs independent of input sequence lengths and reduces the dependence on gap penalties. When sequences share greater similarity, the same approach converts to a multiple sequence alignment procedure.

Algorithms

Suggestions for "safe" residue substitutions in site-directed mutagenesis.

The conserved topological structure observed in various molecular families such as globins or cytochromes c allows structural equivalencing of residues in every homologous structure and defines in a coherent way a global alignment in each sequence family. A search was performed for equivalent residue pairs in various topological families that were buried in protein cores or exposed at the protein surface and that had mutated but maintained similar unmutated environments. Amino acid residues with atoms in contact with the mutated residue pairs defined the environment. Matrices of preferred amino acid exchanges were then constructed and preferred or avoided amino acid substitutions deduced. Given the conserved atomic neighborhoods, such natural in vivo substitutions are subject to similar constrains as point mutations performed in site-directed mutagenesis experiments. The exchange matrices should provide guidelines for "safe" amino acid substitutions least likely to disturb the protein structure, either locally or in its overall folding pathway, and most likely to allow probing the structural and functional significance of the substituted site.

Amino Acid Sequence

Beta-COP, a 110 kd protein associated with non-clathrin-coated vesicles and the Golgi complex, shows homology to beta-adaptin.

We have cloned and sequenced beta-COP, a peripheral 110 kd Golgi membrane protein. beta-COP shows significant homology to beta-adaptin. It is present in a membrane-bound form and in a cytosolic complex of 13-14S, with a Stokes radius of approximately 10 nm and an estimated Mr of approximately 550,000. By immunofluorescence labeling, beta-COP is associated with the structures of the Golgi complex. Immunoelectron microscopy has localized beta-COP to non-clathrin-coated vesicles and cisternae of the Golgi complex. These coated vesicles accumulate in rat liver Golgi fractions treated with GTP gamma S and strongly label for beta-COP. Our data suggest that beta-COP is a component of a coat associated with vesicles and cisternae of the Golgi complex.

Amino Acid Sequence

Automated protein sequence pattern handling and PROSITE searching.

The protein sequence searching program Scrutineer has been modified to search for targets from a file. We are distributing a reformatted file of PROSITES which can be read by Scrutineer. In addition, Scrutineer still accepts targets typed in interactively but can now write them out in the format required as input. Since the input format is the same as the output format, target management and re-use is simple.

Amino Acid Sequence