PubMed HealthSearch

Biomedical subjects

M J Sternberg

Publications and source records attributed to M J Sternberg.

At least 19 recordsLinked to original sources

Drug design by machine learning: the use of inductive logic programming to model the structure-activity relationships of trimethoprim analogues binding to dihydrofolate reductase.

The machine learning program GOLEM from the field of inductive logic programming was applied to the drug design problem of modeling structure-activity relationships. The training data for the program were 44 trimethoprim analogues and their observed inhibition of Escherichia coli dihydrofolate reductase. A further 11 compounds were used as unseen test data. GOLEM obtained rules that were statistically more accurate on the training data and also better on the test data than a Hansch linear regression model. Importantly machine learning yields understandable rules that characterized the chemistry of favored inhibitors in terms of polarity, flexibility, and hydrogen-bonding character. These rules agree with the stereochemistry of the interaction observed crystallographically.

Artificial Intelligence

Evaluation of the sequence template method for protein structure prediction. Discrimination of the (beta/alpha)8-barrel fold.

A multiple alignment of five (beta/alpha)8-barrel enzymes has been derived from their structure. The eight beta-strands and eight alpha-helices of the (beta/alpha)8-barrel are correctly aligned and the equivalenced residues in these regions fulfil similar structural roles. Each beta-strand has a central core of usually four residues, two residues contribute side-chains to the barrel core and the other two residues are involved in beta-strand/alpha-helix contacts. However, the fold imposes no constraints on the volumes of the residues at either a local or global level: the volume of the beta-barrel core varies between 1088 A3 in glycolate oxidase and 1571 A3 in taka-amylase. Sequence motifs derived from the multiple alignment were scanned against a database of 124 protein sequences, including 17 (beta/alpha)8-barrel enzymes. The results were evaluated in terms of the discrimination of (beta/alpha)8-barrel sequences and the quality of the alignments obtained. One motif was able to identify the top 12% of high scoring sequences as forming (beta/alpha)8-barrels with 50% accuracy and the bottom 50% of sequences as not being (beta/alpha)8-barrel proteins with 100% accuracy. However, in most instances the alignments were poor. The reasons for this are discussed with reference to the (beta/alpha)8-barrel proteins and the sequence motif method in general.

Alcohol Oxidoreductases

New algorithm to model protein-protein recognition based on surface complementarity. Applications to antibody-antigen docking.

A novel algorithm is presented which models protein-protein interactions using surface complementarity. The method is applied to antibody-antigen docking. A steric scoring scheme, based upon a soft potential, is used to assess complementarity, and a simple electrostatic model is then used to remove infeasible interactions. The soft potential allows for structural changes that occur during docking. Biochemical knowledge is necessary to reduce the number of docking orientations produced by the method to a manageable size. The information used includes the known epitope residues and a single loose distance constraint. The method is applied to all three crystallographically determined antibody-lysozyme complexes, HyHEL-10, D1.3 and HyHEL-5. For the first time, a predicted antibody structure (that of D1.3) is used as a docking target. In the four systems modelled, the method identifies between 15 and 40 possible docking orientations. The root-mean-square (r.m.s.) deviation between these orientations and the relevant crystallographic complex is measured in the interface region. For all four complexes an orientation is found with r.m.s. deviation in the range 1.9 A and 4.8 A. The algorithm is implemented on a single instruction/multiple datastream (SI/MD) architecture computer. The use of a parallel architecture computer ensures detailed coverage of the search space, whilst still maintaining a search time of two days.

Algorithms

Prediction of structural and functional features of protein and nucleic acid sequences by artificial neural networks.

The applications of artificial neural networks to the prediction of structural and functional features of protein and nucleic acid sequences are reviewed. A brief introduction to neural networks is given, including a discussion of learning algorithms and sequence encoding. The protein applications mostly involve the prediction of secondary and tertiary structure from sequence. The problems in nucleic acid analysis tackled by neural networks are the prediction of translation initiation sites in Escherichia coli, the recognition of splice junctions in human mRNA, and the prediction of promoter sites in E. coli. The performance of the approach is compared with other current statistical methods.

Algorithms

A predicted three-dimensional structure for the carcinoembryonic antigen (CEA).

A three-dimensional model for the carcinoembryonic antigen (CEA) has been constructed by knowledge-based computer modelling. Each of the seven extracellular domains of CEA are expected to have immunoglobulin folds. The N-terminal domain of CEA was modelled using the first domain of the recently solved NMR structure of rat CD2, as well as the first domain of the X-ray crystal structure of human CD4 and an immunoglobulin variable domain REI as templates. The remaining domains were modelled from the first and second domains of CD4 and REI. Link conformations between the domains were taken from the elbow region of antibodies. A possible packing model between each of the seven domains is proposed. Each residue of the model is labelled as to its suitability for site-directed mutagenesis.

Amino Acid Sequence

The binding site on ICAM-1 for Plasmodium falciparum-infected erythrocytes overlaps, but is distinct from, the LFA-1-binding site.

The intercellular adhesion molecule-1 (ICAM-1, CD54) is one of three putative endothelial receptors that mediate in vitro cytoadherence of P. falciparum-infected erythrocytes. Since cytoadherence to postcapillary venular endothelium is thought to be a major factor in the virulence of P. falciparum malaria, we have examined the interaction between ICAM-1 and the P. falciparum-infected cell, and have compared it with the interaction to the physiological counter receptor, the leukocyte integrin LFA-1. Our results demonstrate that the malaria-binding site resides in the first two domains of the ICAM-1 molecule and overlaps, but is distinct from, the LFA-1 site.

Amino Acid Sequence

Three dimensional structure of the transmembrane region of the proto-oncogenic and oncogenic forms of the neu protein.

The neu proto-oncogene may be converted into a dominantly transforming oncogene by a single point mutation. Substitution of a valine residue at position 664 in the transmembrane region with glutamic acid activates the tyrosine kinase of the molecule and is associated with increased receptor dimerization. Previously we have proposed a model in which the glutamic acid side chain stabilizes receptor dimerization by hydrogen bonding. Other models have been proposed in which the mutation leads to a conformational change in the transmembrane region mimicking that assumed to occur following binding of a natural ligand. Synthetic peptides representing part of the transmembrane region were prepared. Some residues were replaced with serine in order to improve peptide solubility to allow purification and analysis. Both the peptides containing valine and glutamic acid dissolved in water and in an artificial lipid monolayer. The structures of the peptides were determined by NMR spectroscopy to be alpha-helical. No significant difference in conformation was observed between the two peptides. This result does not support the model proposing a conformational change. The receptor structures determined experimentally do allow alternative models involving receptor transmembrane region packing.

Amino Acid Sequence

Modelling the structure and function of enzymes by machine learning.

A machine learning program, GOLEM, has been applied to two problems: (1) the prediction of protein secondary structure from sequence and (2) modelling a quantitative structure-activity relationship in drug design. GOLEM takes as input observations and combines them with background knowledge of chemistry to yield rules expressed as stereochemical principles for prediction. The secondary structure prediction was explored on the alpha/alpha class of proteins; on an unrelated test set it yielded 81% accuracy. The rules from GOLEM defined patterns of residues forming alpha-helices. The system studied for drug design was the activities of trimethoprim analogues binding to E. coli dihydrofolate reductase. The GOLEM rules were a better model than standard regression approaches. More importantly, these rules described the chemical properties of the enzyme-binding site that were in broad agreement with the crystallographic structure.

Amino Acid Sequence

Towards an automatic method of predicting protein structure by homology: an evaluation of suboptimal sequence alignments.

A major problem in predicting protein structure by homology modelling is that the sequence alignment from which the model is built may not be the best one in terms of the correct equivalencing of residues assessed by structural or functional criteria. A useful strategy is to generate and examine a number of suboptimal alignments as better alignments can often be found away from the optimal. A procedure to filter rapidly suboptimal alignments based on measurement of core volumes and packing pair potentials is investigated. The approach is benchmarked on three pairs of sequences which are non-trivial to align correctly, namely two immunoglobulin domains, plastocyanin with azurin and two distant globin sequences. It is shown to be useful to reduce a large ensemble of possible alignments down to a few which correspond more closely to the correct (structure based) alignment.

Algorithms

Protein secondary structure prediction using logic-based machine learning.

Many attempts have been made to solve the problem of predicting protein secondary structure from the primary sequence but the best performance results are still disappointing. In this paper, the use of a machine learning algorithm which allows relational descriptions is shown to lead to improved performance. The Inductive Logic Programming computer program, Golem, was applied to learning secondary structure prediction rules for alpha/alpha domain type proteins. The input to the program consisted of 12 non-homologous proteins (1612 residues) of known structure, together with a background knowledge describing the chemical and physical properties of the residues. Golem learned a small set of rules that predict which residues are part of the alpha-helices--based on their positional relationships and chemical and physical properties. The rules were tested on four independent non-homologous proteins (416 residues) giving an accuracy of 81% (+/- 2%). This is an improvement, on identical data, over the previously reported result of 73% by King and Sternberg (1990, J. Mol. Biol., 216, 441-457) using the machine learning program PROMIS, and of 72% using the standard Garnier-Osguthorpe-Robson method. The best previously reported result in the literature for the alpha/alpha domain type is 76%, achieved using a neural net approach. Machine learning also has the advantage over neural network and statistical methods in producing more understandable results.

Amino Acid Sequence

A simple method to generate non-trivial alternate alignments of protein sequences.

A major problem in sequence alignments based on the standard dynamic programming method is that the optimal path does not necessarily yield the best equivalencing of residues assessed by structural or functional criteria. An algorithm is presented that finds suboptimal alignments of protein sequences by a simple modification to the standard dynamic programming method. The standard pairwise weight matrix elements are modified in order to penalize, but not eliminate, the equivalencing of residues obtained from previous alignments. The algorithm thereby yields a limited set of alternate alignments that can differ considerably from the optimal. The approach is benchmarked on the alignments of immunoglobulin domains. Without a prior knowledge of the optimal choice of gap penalty, one of the suboptimal alignments is shown to be more accurate than the optimal.

Algorithms

PROMOT: a FORTRAN program to scan protein sequences against a library of known motifs.

Information about the three-dimensional structure or function of a newly determined protein sequence can be obtained if the protein is found to contain a characterized motif or pattern of residues. Recently a database (PROSITE) has been established that contains 337 known motifs encoded as a list of allowed residue types at specific positions along the sequence. PROMOT is a FORTRAN computer program that takes a protein sequence and examines if it contains any of the motifs in PROSITE. The program also extends the definitions of patterns beyond those used in PROSITE to provide a simple, yet flexible, method to scan either a PROSITE or a user-defined pattern against a protein sequence database.

Amino Acid Sequence

A three-dimensional molecular template for substrates of human cytochrome P450 involved in debrisoquine 4-hydroxylation.

A three-dimensional molecular template has been generated for substrates of human debrisoquine 4-hydroxylase cytochrome P450 (CYP2D6). This template defines the stereochemical requirements for CYP2D6 substrates in terms of the volume occupied and positions of key atoms. The modelling was based on the X-ray crystallographic coordinates of the location of the attacked C5 atom of camphor in relation to the haem in cytochrome P450 cam. Interactive molecular graphics combined with energy calculations were used to identify allowed conformers to superpose known CYP2D6 substrates to yield a molecular template. This model takes into account the site of attack of the known substrates and the requirement for a protonated nitrogen atom to interact with an anion site of the protein. A nitrogen-anion distance of between 2.5 and 4.5 A was allowed for the interaction. The substrates modelled were cardiovascular drugs (debrisoquine, sparteine, guanoxan and perhexiline), beta-adrenergic blocking agents (bufuralol and propranolol), tricyclic anti-depressants (desipramine, amitriptyline and nortriptyline) and other miscellaneous compounds (phenformin, methoxy-amphetamine, codeine and dextromethorphan). The template generated in this manner was then used to determine the likelihood that certain other compounds were substrates for CYP2D6. A carcinogenic protein pyrolysate product, 2-amino-1-methyl-6-phenylimidazo[4,5-b]pyridine (PhIP), did not fit the template and is therefore unlikely to be activated by this enzyme. A potent carcinogen in tobacco smoke, 4-(N-methyl-N-nitrosamino)-1-(3-pyridyl)-1-butanone (NNK), fitted the template but could not be modelled to form a favourable nitrogen-anion interaction. Experimental substrate competition studies also showed that NNK is unlikely to be a CYP2D6 substrate. It was also shown that the widely used drug for treatment of breast cancer, trans-1-(4-beta-dimethylaminoethoxyphenyl)-,2-diphenyl-1-ene (tamoxifen), did not fit the molecular template and is unlikely to be metabolized by CYP2D6. Coordinates of the template are available.

Camphor

A predicted three-dimensional structure of human cytochrome P450: implications for substrate specificity.

A three-dimensional structure for human cytochrome P450IA1 was predicted based on the crystal coordinates of cytochrome P450cam from Pseudomonas putida. As there was only 15% residue identity between the two enzymes, additional information was used to establish an accurate sequence alignment that is a prerequisite for model building. Twelve representative eukaryotic sequences were aligned and a net prediction of secondary structure was matched against the known alpha-helices and beta-sheets of P450cam. The cam secondary structure provided a fixed main-chain framework onto which loops of appropriate length from the human P450IA1 structure were added. The model-built structure of the human cytochrome conformed to the requirements for the segregation of polar and nonpolar residues between the core and the surface. The first 44 residues of human cytochrome P450 could not be built into the model and sequence analysis suggested that residues 1-26 formed a single membrane-spanning segment. Examination of the sequences of cytochrome P450s from distinct gene families suggested specific residues that could account for the differences in substrate specificity. A major substrate for P450IA1, 3-methyl-cholanthrene, was fitted into the proposed active site and this planar aromatic molecule could be accommodated into the available cavity. Residues that are likely to interact with the haem were identified. The sequence similarity between 59 eukaryotic enzymes was represented as a dendrogram that in general clustered according to gene family. Until a crystallographic structure is available, this model-building study identifies potential residues in cytochrome P450s important in the function of these enzymes and these residues are candidates for site-directed mutagenesis.

Amino Acid Sequence

Prediction of ATP/GTP-binding motif: a comparison of a perceptron type neural network and a consensus sequence method [corrected].

Neural networks have been applied to a number of protein structure problems. In some applications their success has not been substantiated by a comparison with the performance of a suitable alternative statistical method on the same data. In this paper, a two-layer feed-forward neural network has been trained to recognize ATP/GTP-binding [corrected] local sequence motifs. The neural network correctly classified 78% of the 349 sequences used. This was much better than a simple motif-searching program. A more sophisticated statistical method was developed, however, which performed marginally better (80% correct classification) than the neural network. The neural network and the statistical method performed similarly on sequences of varying degrees of homology. These results do not imply that neural networks, especially those with hidden layers, are not useful tools, but they do suggest that two-layer networks in particular should be carefully tested against other statistical methods.

Adenosine Triphosphate