PubMed HealthSearch

Biomedical subjects

P Bork

Publications and source records attributed to P Bork.

16 recordsLinked to original sources

Cloning of a gene from Bacillus cereus with homology to the mreB gene from Escherichia coli.

We have cloned and sequenced a gene coding for a putative shape-determining protein (MreB) highly homologous to the mreB gene product of Escherichia coli. The amino acid (aa) identity was 53% and the similarity 72%. The gene is expressed early in the logarithmic phase. The aa sequence comparison showed that the protein, like the E. coli MreB, has structural similarity to actin and heat-shock protein Hsc70 encoded by a new super-gene family.

Amino Acid Sequence

Proposed acquisition of an animal protein domain by bacteria.

A systematic screen of a protein sequence data base confirms that the fibronectin type III (Fn3) domain is widely distributed among animal proteins and occurs also in several bacterial carbohydrate-splitting enzymes. The motif has yet to be identified in proteins from plants or fungi. All indications are that the bacterial sequences are much too similar to the animal type to be the result of conventional vertical descent. Rather, it is likely that the bacterial units were initially acquired from an animal source and are being spread further by horizontal transfers between distantly related bacteria.

Amino Acid Sequence

An ATPase domain common to prokaryotic cell cycle proteins, sugar kinases, actin, and hsp70 heat shock proteins.

The functionally diverse actin, hexokinase, and hsp70 protein families have in common an ATPase domain of known three-dimensional structure. Optimal superposition of the three structures and alignment of many sequences in each of the three families has revealed a set of common conserved residues, distributed in five sequence motifs, which are involved in ATP binding and in a putative interdomain hinge. From the multiple sequence alignment in these motifs a pattern of amino acid properties required at each position is defined. The discriminatory power of the pattern is in part due to the use of several known three-dimensional structures and many sequences and in part to the "property" method of generalizing from observed amino acid frequencies to amino acid fitness at each sequence position. A sequence data base search with the pattern significantly matches sugar kinases, such as fuco-, glucono-, xylulo-, ribulo-, and glycerokinase, as well as the prokaryotic cell cycle proteins MreB, FtsA, and StbA. These are predicted to have subdomains with the same tertiary structure as the ATPase subdomains Ia and IIa of hexokinase, actin, and Hsc70, a very similar ATP binding pocket, and the capacity for interdomain hinge motion accompanying functional state changes. A common evolutionary origin for all of the proteins in this class is proposed.

Actins

The modular architecture of vertebrate collagens.

Collagens are typical mosaic proteins containing a number of shuffled domains. These domains have been classified by sequence similarity in order to characterize their structural and functional relationships to other proteins. This analysis provides an overview of homologies of collagen domains. It also reveals two new relationships: (i) a module common to type V, IX, XI, and XII collagens was found to be homologous to the heparin binding domain of thrombospondin; (ii) the modular architecture of a human type VII collagen fragment was identified. Its N-terminal globular domain contains fibronectin type III repeats located adjacent to a Von Willebrand factor type A module. The proposed structural similarities point to analogous subfunctions of the respective domains in otherwise distinct proteins.

Amino Acid Sequence

A large domain common to sperm receptors (Zp2 and Zp3) and TGF-beta type III receptor.

A new family of mosaic proteins is defined by sequence analysis. The family is characterized by a 260 residue domain common to proteins of apparently diverse function and tissue specificity: sperm receptors Zp2 and Zp3, betaglycan (also called TGF-beta type III receptor), uromodulin, as well as the major zymogen granule membrane protein (GP-2). The location of the common domain is similar with respect to putative transmembrane regions. The results lead to the hypothesis that this type of domain has a common tertiary structure and that there is a functional similarity in the recognition mechanism of the sperm receptor system and the TGF-beta receptor complex.

Amino Acid Sequence

Comprehensive sequence analysis of the 182 predicted open reading frames of yeast chromosome III.

With the completion of the first phase of the European yeast genome sequencing project, the complete DNA sequence of chromosome III of Saccharomyces cerevisiae has become available (Oliver, S. G., et al., 1992, Nature 357, 38-46). We have tested the predictive power of computer sequence analysis of the 176 probable protein products of this chromosome, after exclusion of six problem cases. When the results of database similarity searches are pooled with prior knowledge, a likely function can be assigned to 42% of the proteins, and a predicted three-dimensional structure to a third of these (14% of the total). The function of the remaining 58% remains to be determined. Of these, about one-third have one or more probable transmembrane segments. Among the most interesting proteins with predicted functions are a new member of the type X polymerase family, a transcription factor with an N-terminal DNA-binding domain related to GAL4, a "fork head" DNA-binding domain previously known only in Drosophila and in mammals, and a putative methyltransferase. Our analysis increased the number of known significant sequence similarities on chromosome III by 13, to now 67. Although the near 40% success rate of identifying unknown protein function by sequence analysis is surprisingly high, the information gap between known protein sequences and unknown function is expected to widen and become a major bottleneck of genome projects in the near future. Based on the experience gained in this test study, we suggest that the development of an automated computer workbench for protein sequence analysis must be an important item in genome projects.

Acetolactate Synthase

On alpha-helices terminated by glycine. 1. Identification of common structural features.

About one third of all helices is terminated by residues with a positive torsion angle phi. 74% of them are glycines. This strong propensity can be explained by typical bifurcated three-center hydrogen bonds which are only compatible with a positive torsion angle phi, causing helix termination. An algorithm was developed to identify these structural features in alpha-helices. 158 out of 456 helices in 79 different well-refined protein structures examined in our analysis were found to have a glycine with this special conformation which have been conserved remarkably during evolution.

Algorithms

On alpha-helices terminated by glycine. 2. Recognition by sequence patterns.

Consensus sequence patterns were constructed to describe helix ends with a characteristic conformation caused by specific three-center hydrogen bonds. This special type of hydrogen bond pattern comprises about one third of all helices and mostly contains glycine with a positive torsion angle phi at the helix ends. After a simple clustering procedure 6 resulting consensus sequence patterns were able to identify 501 out of 575 helix ends in the Brookhaven Protein Data Bank, showing the above-mentioned features. The patterns did not detect any false segment, but numerous sequence segments not identified by structural criteria were recognized. It is likely that they are indeed helices terminated by glycine with a positive torsion angle phi.

Amino Acid Sequence

Shuffled domains in extracellular proteins.

A comprehensive list of domains in extracellular mosaic proteins is presented. About 40 domains were distinguished by consensus patterns. A subsequent sequence database search recognized these domains in more than 200 extracellular proteins. The results point to a structural network, which may also represent the molecular basis for a complex coordination of various functions within the world of extracellular proteins.

Amino Acid Sequence

Complement components C1r/C1s, bone morphogenic protein 1 and Xenopus laevis developmentally regulated protein UVS.2 share common repeats.

Property patterns were constructed, based on an alignment of related domains in human complement subcomponents C1r and C1s as well as in the sea urchin protein uEGF. This kind of consensus pattern was able to identify similar domains in a human bone morphogenic protein, in a Xenopus laevis embryonal protein involved in dorsoanterior development and in a calcium-dependent serine protease secreted from malignant hamster embryo fibroblast cells. Because of the high level of overall sequence homology this protease may be the hamsters' equivalent of the human complement subcomponent C1s. The resulting multiple alignment of all studied domains suggests functionally and structurally important regions.

Amino Acid Sequence

Sequence similarities between tryptophan synthase beta subunit and other pyridoxal-phosphate-dependent enzymes.

On the basis of 8 tryptophan synthase beta subunits (EC 4.2.1.20) consensus patterns were constructed comprising two conserved motifs. Screening of the SWISSPROT protein sequence database with these patterns indicates similarities with O-acetylserine sulfhydrolases (EC 4.2.99.8), threonine synthases (EC 4.2.99.2), L- and D-serine dehydratases (EC 4.2.1.13/EC 4.2.1.14) and threonine dehydratases (EC 4.2.1.16). Using multiple alignment procedures the similar regions could be extended. In connection with their pyridoxal-phosphate-binding-capacity and their positions in biochemical pathways evolutionary relationships among these enzymes are discussed.

Amino Acid Sequence

Recognition of different nucleotide-binding sites in primary structures using a property-pattern approach.

Consensus sequence patterns for beta-alpha-beta folds binding FAD, NAD and GTP were constructed on the basis of 11 steric and physicochemical properties. These property patterns permit detection and distinction of the respective nucleotide-binding sites on the basis of amino acid sequence analysis alone. The SWISS-PROT database (release 9) was screened with the three calculated patterns, and nucleotide-binding sites identified are presented. They correspond to existing structure data (if known). For the detected sequence segments we are able to predict the beta-alpha-beta motif as well as the respective binding sites. For some of the proteins so detected a nucleotide-binding capacity has not previously been reported.

Algorithms

Recognition of functional regions in primary structures using a set of property patterns.

32 consensus patterns for a set of functional regions and structural motifs in protein sequences were constructed. The pattern definition is heuristic and based on 11 selected steric and physicochemical properties. By comparison with these patterns, it was possible to identify, without false detection, 1532 sites in 8702 protein sequences of SWISSPROT. Screening against such a pattern library offers a considerable chance to identify functional regions or structural motifs in proteins from which only the sequence is known.

Amino Acid Sequence