PubMed Health⌕ Search

Biomedical subjects

Douglas Brutlag

Publications and source records attributed to Douglas Brutlag.

5 recordsLinked to original sources

Development and validation of a consistency based multiple structure alignment algorithm.

SUMMARY: We introduce an algorithm that uses the information gained from simultaneous consideration of an entire group of related proteins to create multiple structure alignments (MSTAs). Consistency-based alignment (CBA) first harnesses the information contained within regions that are consistently aligned among a set of pairwise superpositions in order to realign pairs of proteins through both global and local refinement methods. It then constructs a multiple alignment that is maximally consistent with the improved pairwise alignments. We validate CBA's alignments by assessing their accuracy in regions where at least two of the aligned structures contain the same conserved sequence motif. RESULTS: CBA correctly aligns well over 90% of motif residues in superpositions of proteins belonging to the same family or superfamily, and it outperforms a number of previously reported MSTA algorithms.

Algorithms↗

FoldMiner and LOCK 2: protein structure comparison and motif discovery on the web.

The FoldMiner web server (http://foldminer.stanford.edu/) provides remote access to methods for protein structure alignment and unsupervised motif discovery. FoldMiner is unique among such algorithms in that it improves both the motif definition and the sensitivity of a structural similarity search by combining the search and motif discovery methods and using information from each process to enhance the other. In a typical run, a query structure is aligned to all structures in one of several databases of single domain targets in order to identify its structural neighbors and to discover a motif that is the basis for the similarity among the query and statistically significant targets. This process is fully automated, but options for manual refinement of the results are available as well. The server uses the Chime plugin and customized controls to allow for visualization of the motif and of structural superpositions. In addition, we provide an interface to the LOCK 2 algorithm for rapid alignments of a query structure to smaller numbers of user-specified targets.

Algorithms↗

FoldMiner: structural motif discovery using an improved superposition algorithm.

We report an unsupervised structural motif discovery algorithm, FoldMiner, which is able to detect global and local motifs in a database of proteins without the need for multiple structure or sequence alignments and without relying on prior classification of proteins into families. Motifs, which are discovered from pairwise superpositions of a query structure to a database of targets, are described probabilistically in terms of the conservation of each secondary structure element's position and are used to improve detection of distant structural relationships. During each iteration of the algorithm, the motif is defined from the current set of homologs and is used both to recruit additional homologous structures and to discard false positives. FoldMiner thus achieves high specificity and sensitivity by distinguishing between homologous and nonhomologous structures by the regions of the query to which they align. We find that when two proteins of the same fold are aligned, highly conserved secondary structure elements in one protein tend to align to highly conserved elements in the second protein, suggesting that FoldMiner consistently identifies the same motif in members of a fold. Structural alignments are performed by an improved superposition algorithm, LOCK 2, which detects distant structural relationships by placing increased emphasis on the alignment of secondary structure elements. LOCK 2 obeys several properties essential in automated analysis of protein structure: It is symmetric, its alignments of secondary structure elements are transitive, its alignments of residues display a high degree of transitivity, and its scoring system is empirically found to behave as a metric.

Algorithms↗

Remote homology detection: a motif based approach.

MOTIVATION: Remote homology detection is the problem of detecting homology in cases of low sequence similarity. It is a hard computational problem with no approach that works well in all cases. RESULTS: We present a method for detecting remote homology that is based on the presence of discrete sequence motifs. The motif content of a pair of sequences is used to define a similarity that is used as a kernel for a Support Vector Machine (SVM) classifier. We test the method on two remote homology detection tasks: prediction of a previously unseen SCOP family and prediction of an enzyme class given other enzymes that have a similar function on other substrates. We find that it performs significantly better than an SVM method that uses BLAST or Smith-Waterman similarity scores as features.

Algorithms↗

Using robotics to fold proteins and dock ligands.

The problems of protein folding and ligand docking have been explored largely using molecular dynamics or Monte Carlo methods. These methods are very compute intensive because they often explore a much wider range of energies, conformations and time than necessary. In addition, Monte Carlo methods often get trapped in local minima. We initially showed that robotic motion planning permitted one to determine the energy of binding and dissociation of ligands from protein binding sites (Singh et al., 1999). The robotic motion planning method maps complicated three-dimensional conformational states into a much simpler, but higher dimensional space in which conformational rearrangements can be represented as linear paths. The dimensionality of the conformation space is of the same order as the number of degrees of conformational freedom in three-dimensional space. We were able to determine the relative energy of association and dissociation of a ligand to a protein by calculating the energetics of interaction for a few thousand conformational states in the vicinity of the protein and choosing the best path from the roadmap. More recently, we have applied roadmap planning to the problem of protein folding (Apaydin et al., 2002a). We represented multiple conformations of a protein as nodes in a compact graph with the edges representing the probability of moving between neighboring states. Instead of using Monte Carlo simulation to simulate thousands of possible paths through various conformational states, we were able to use Markov methods to calculate the steady state occupancy of each conformation, needing to calculate the energy of each conformation only once. We referred to this Markov method of representing multiple conformations and transitions as stochastic roadmap simulation or SRS. We demonstrated that the distribution of conformational states calculated with exhaustive Monte Carlo simulations asymptotically approached the Markov steady state if the same Boltzman energy distribution was used in both methods. SRS permits one to calculate contributions from all possible paths simultaneously with far fewer energy calculations than Monte Carlo or molecular dynamics methods. The SRS method also permits one to represent multiple unfolded starting states and multiple, near-native, folded states and all possible paths between them simultaneously. The SRS method is also independent of the function used to calculate the energy of the various conformational states. In a paper to be presented at this conference (Apaydin et al., 2002b) we have also applied SRS to ligand docking in which we calculate the dynamics of ligand-protein association and dissociation in the region of various binding sites on a number of proteins. SRS permits us to determine the relative times of association to and dissociation from various catalytic and non-catalytic binding sites on protein surfaces. Instead of just following the best path in a roadmap, we can calculate the contribution of all the possible binding or dissociation paths and their relative probabilities and energies simultaneously.

Algorithms↗