PubMed Health⌕ Search

Biomedical subjects

Jay H Lee

Publications and source records attributed to Jay H Lee.

4 recordsLinked to original sources

Identifying the interacting positions of a protein using Boolean learning and support vector machines.

It is known that in the three-dimensional structure of a protein, certain amino acids can interact with each other in order to provide structural integrity or aid in its catalytic function. If these positions are mutated the loss of this interaction usually leads to a non-functional protein. Directed evolution experiments, which probe the sequence space of a protein through mutations in search for an improved variant, frequently result in such inactive sequences. In this work, we address the use of machine learning algorithms, Boolean learning and support vector machines (SVMs), to find such pairs of amino acid positions. The recombination method of imparting mutations was simulated to create in silico sequences that were used as training data for the algorithms. The two algorithms were combined together to develop an approach that weighs the structural risk as well as the empirical risk to solve the problem. This strategy was adapted to a multi-round framework of experiments where the data generated in the present round is used to design experiments for the next round to improve the generated library, as well as the estimation of the interacting positions. It is observed that this strategy can greatly improve the number of functional variants that are generated as well as the average number of mutations that can be made in the library.

Algorithms↗

Simulation modeling of pooling for combinatorial protein engineering.

Pooling in directed-evolution experiments will greatly increase the throughput of screening systems, but important parameters such as the number of good mutants created and the activity level increase of the good mutants will depend highly on the protein being engineered. The authors developed and validated a Monte Carlo simulation model of pooling that allows the testing of various scenarios in silico before starting experimentation. Using a simplified test system of 2 enzymes, betagalactosidase (supermutant, or greatly improved enzyme) and beta-glucuronidase (dud, or enzyme with ancestral level of activity), the model accurately predicted the number of supermutants detected in experiments within a factor of 2. Additional simulations using more complex activity distributions show the versatility of the model. Pooling is most suited to cases such as the directed evolution of new function in a protein, where the background level of activity is minimized, making it easier to detect small increases in activity level. Pooling is most successful when a sensitive assay is employed. Using the model will increase the throughput of screening procedures for directed-evolution experiments and thus lead to speedier engineering of proteins.

Cells, Cultured↗

Support vector machines for learning to identify the critical positions of a protein.

A method for identifying the positions in the amino acid sequence, which are critical for the catalytic activity of a protein using support vector machines (SVMs) is introduced and analysed. SVMs are supported by an efficient learning algorithm and can utilize some prior knowledge about the structure of the problem. The amino acid sequences of the variants of a protein, created by inducing mutations, along with their fitness are required as input data by the method to predict its critical positions. To investigate the performance of this algorithm, variants of the beta-lactamase enzyme were created in silico using simulations of both mutagenesis and recombination protocols. Results from literature on beta-lactamase were used to test the accuracy of this method. It was also compared with the results from a simple search algorithm. The algorithm was also shown to be able to predict critical positions that can tolerate two different amino acids and retain function.

Algorithms↗

Pooling for improved screening of combinatorial libraries for directed evolution.

Following diversity generation in combinatorial protein engineering, a significant amount of effort is expended in screening the library for improved variants. Pooling, or combining multiple cells into the same assay well when screening, is a means to increase throughput and screen a larger portion of the library with less time and effort. We have developed and validated a Monte Carlo simulation model of pooling and used it to screen a library of beta-galactosidase mutants randomized in the active site to increase their activity toward fucosides. Here, we show that our model can successfully predict the number of highly improved mutants obtained via pooling and that pooling does increase the number of good mutants obtained. In unpooled conditions, we found a total of three mutants with higher activity toward p-nitrophenyl-beta-D-fucoside than that of the wild-type beta-galactosidase, whereas when pooling 10 cells per well we found a total of approximately 10 improved mutants. In addition, the number of "supermutants", those with the highest activity increase, was also higher when pooling was used. Pooling is a useful tool for increasing the efficiency of screening combinatorial protein engineering libraries.

Binding Sites↗