PubMed Health⌕ Search

Biomedical subjects

Julian Mintseris

Publications and source records attributed to Julian Mintseris.

10 recordsLinked to original sources

Design of a combinatorial DNA microarray for protein-DNA interaction studies.

BACKGROUND: Discovery of precise specificity of transcription factors is an important step on the way to understanding the complex mechanisms of gene regulation in eukaryotes. Recently, double-stranded protein-binding microarrays were developed as a potentially scalable approach to tackle transcription factor binding site identification. RESULTS: Here we present an algorithmic approach to experimental design of a microarray that allows for testing full specificity of a transcription factor binding to all possible DNA binding sites of a given length, with optimally efficient use of the array. This design is universal, works for any factor that binds a sequence motif and is not species-specific. Furthermore, simulation results show that data produced with the designed arrays is easier to analyze and would result in more precise identification of binding sites. CONCLUSION: In this study, we present a design of a double stranded DNA microarray for protein-DNA interaction studies and show that our algorithm allows optimally efficient use of the arrays for this purpose. We believe such a design will prove useful for transcription factor binding site identification and other biological problems.

3' Flanking Region↗

ZDOCK and RDOCK performance in CAPRI rounds 3, 4, and 5.

We present an evaluation of the results of our ZDOCK and RDOCK algorithms in Rounds 3, 4, and 5 of the protein docking challenge CAPRI. ZDOCK is a Fast Fourier Transform (FFT)-based, initial-stage rigid-body docking algorithm, and RDOCK is an energy minimization algorithm for refining and reranking ZDOCK results. Of the 9 targets for which we submitted predictions, we attained at least acceptable accuracy for 7, at least medium accuracy for 6, and high accuracy for 3. These results are evidence that ZDOCK in combination with RDOCK is capable of making accurate predictions on a diverse set of protein complexes.

Algorithms↗

Protein-Protein Docking Benchmark 2.0: an update.

We present a new version of the Protein-Protein Docking Benchmark, reconstructed from the bottom up to include more complexes, particularly focusing on more unbound-unbound test cases. SCOP (Structural Classification of Proteins) was used to assess redundancy between the complexes in this version. The new benchmark consists of 72 unbound-unbound cases, with 52 rigid-body cases, 13 medium-difficulty cases, and 7 high-difficulty cases with substantial conformational change. In addition, we retained 12 antibody-antigen test cases with the antibody structure in the bound form. The new benchmark provides a platform for evaluating the progress of docking methods on a wide variety of targets. The new version of the benchmark is available to the public at http://zlab.bu.edu/benchmark2.

Algorithms↗

Structure, function, and evolution of transient and obligate protein-protein interactions.

Recent analyses of high-throughput protein interaction data coupled with large-scale investigations of evolutionary properties of interaction networks have left some unanswered questions. To what extent do protein interactions act as constraints during evolution of the protein sequence? How does the type of interaction, specifically transient or obligate, play into these constraints? Are the mutations in the binding site of an interacting protein correlated with mutations in the binding site of its partner? We address these and other questions by relying on a carefully curated dataset of protein complex structures. Results point to the importance of distinguishing between transient and obligate interactions. We conclude that residues in the interfaces of obligate complexes tend to evolve at a relatively slower rate, allowing them to coevolve with their interacting partners. In contrast, the plasticity inherent in transient interactions leads to an increased rate of substitution for the interface residues and leaves little or no evidence of correlated mutations across the interface.

Amino Acid Substitution↗

Are protein-protein interfaces more conserved in sequence than the rest of the protein surface?

Protein interfaces are thought to be distinguishable from the rest of the protein surface by their greater degree of residue conservation. We test the validity of this approach on an expanded set of 64 protein-protein interfaces using conservation scores derived from two multiple sequence alignment types, one of close homologs/orthologs and one of diverse homologs/paralogs. Overall, we find that the interface is slightly more conserved than the rest of the protein surface when using either alignment type, with alignments of diverse homologs showing marginally better discrimination. However, using a novel surface-patch definition, we find that the interface is rarely significantly more conserved than other surface patches when using either alignment type. When an interface is among the most conserved surface patches, it tends to be part of an enzyme active site. The most conserved surface patch overlaps with 39% (+/- 28%) and 36% (+/- 28%) of the actual interface for diverse and close homologs, respectively. Contrary to results obtained from smaller data sets, this work indicates that residue conservation is rarely sufficient for complete and accurate prediction of protein interfaces. Finally, we find that obligate interfaces differ from transient interfaces in that the former have significantly fewer alignment gaps at the interface than the rest of the protein surface, as well as having buried interface residues that are more conserved than partially buried interface residues.

Amino Acid Sequence↗

Optimizing protein representations with information theory.

The problem of describing a protein representation by breaking up the amino acids atoms into functionally similar atom groups has been addressed by many researchers in the past 25 years. They have used a variety of physical, chemical and biological criteria of varying degrees of rigor to essentially impose our understanding of protein structures onto various atom-typing schemes used in studies of protein folding, protein-protein and protein-ligand interactions, and others. Here, instead, we have chosen to rely primarily on the data and use information-theoretic techniques to dissect it. We show that we can obtain an optimized protein representation for a given alphabet size from protein monomers or protein interface datasets that are in agreement with general concepts of protein energetics. Closer inspection of the atom partitions led to interesting observations pointing to the greater importance of the hydrophobic interactions in protein monomers compared to interfaces and, conversely, greater importance of polar/charged interaction in protein interfaces. Comparing the atom partitions from the two datasets we show that the two are strikingly similar at alphabet size of five, proving that despite some differences, the general energetic concepts are very similar for folding and binding. Implications for further structural studies are discussed.

Binding Sites↗

Atomic contact vectors in protein-protein recognition.

The ability to analyze and compare protein-protein interactions on the structural level is critical to our understanding of various aspects of molecular recognition and the functional interplay of components of biochemical networks. In this study, we introduce atomic contact vectors (ACVs) as an intuitive way to represent the physico-chemical characteristics of a protein-protein interface as well as a way to compare interfaces to each other. We test the utility of ACVs in classification by using them to distinguish between homodimers and crystal contacts. Our results compare favorably with those reported by other authors. We then apply ACVs to mine the PDB for all known protein-protein complexes and separate transient recognition complexes from permanent oligomeric ones. Getting at the basis of this difference is important for our understanding of recognition and we achieved a success rate of 91% for distinguishing these two classes of complexes. Although accessible surface area of the interface is a major discriminating feature, we also show that there are distinct differences in the contact preferences between the two kinds of complexes. Illustrating the superiority of ACVs as a basic comparison measure over a sequence-based approach, we derive a general rule of thumb to determine whether two protein-protein interfaces are redundant. With this method, we arrive at a nonredundant set of 209 recognition complexes--the largest set reported so far.

Crystallization↗

ZDOCK predictions for the CAPRI challenge.

The CAPRI Challenge is a blind test of protein-protein-docking algorithms that predict the complex structure from the crystal structures of the interacting proteins. We participated in both rounds of this blind test and submitted predictions for all seven targets, relying mainly on our Fast Fourier Transform based algorithm ZDOCK that combines shape complementarity, desolvation, and electrostatics. Our group made good predictions for three targets and had at least some success with three others. Implications of the treatment of prior biological information as well as contributions of manual inspection to docking predictions are also discussed.

Algorithms↗

A protein-protein docking benchmark.

We have developed a nonredundant benchmark for testing protein-protein docking algorithms. Currently it contains 59 test cases: 22 enzyme-inhibitor complexes, 19 antibody-antigen complexes, 11 other complexes, and 7 difficult test cases. Thirty-one of the test cases, for which the unbound structures of both the receptor and ligand are available, are classified as follows: 16 enzyme-inhibitor, 5 antibody-antigen, 5 others, and 5 difficult. Such a centralized resource should benefit the docking community not only as a large curated test set but also as a common ground for comparing different algorithms. The benchmark is available at (http://zlab.bu.edu/~rong/dock/benchmark.shtml).

Algorithms↗

Predictome: a database of putative functional links between proteins.

The current deluge of genomic sequences has spawned the creation of tools capable of making sense of the data. Computational and high-throughput experimental methods for generating links between proteins have recently been emerging. These methods effectively act as hypothesis machines, allowing researchers to screen large sets of data to detect interesting patterns that can then be studied in greater detail. Although the potential use of these putative links in predicting gene function has been demonstrated, a central repository for all such links for many genomes would maximize their usefulness. Here we present Predictome, a database of predicted links between the proteins of 44 genomes based on the implementation of three computational methods--chromosomal proximity, phylogenetic profiling and domain fusion--and large-scale experimental screenings of protein-protein interaction data. The combination of data from various predictive methods in one database allows for their comparison with each other, as well as visualization of their correlation with known pathway information. As a repository for such data, Predictome is an ongoing resource for the community, providing functional relationships among proteins as new genomic data emerges. Predictome is available at http://predictome.bu.edu.

Animals↗