PubMed Health⌕ Search

Biomedical subjects

S C Pegg

Publications and source records attributed to S C Pegg.

3 recordsLinked to original sources

Functional assignment of the 20 S proteasome from Trypanosoma brucei using mass spectrometry and new bioinformatics approaches.

As experimental technologies for characterization of proteomes emerge, bioinformatic analysis of the data becomes essential. Separation and identification technologies currently based on two-dimensional gels/mass spectrometry provide the inherent analytical power required. This strategy involves protein spot digestion and accurate mass mapping together with computational interrogation of available data bases for protein functional identification. When either no exact match is found or when the possible matches only partially account for molecular weights actually observed, peptide sequencing by tandem mass spectrometry has emerged as the methodology of choice to provide the basic additional information required. To evaluate the capabilities of bioinformatics methods employed for identifying homologs of a protein of interest, we attempted to identify the major proteins from the 20 S proteasome of Trypanosoma brucei using sequence information determined using mass spectrometry. The results suggest that neither the traditional query engines, BLAST and FASTA, nor specialized software developed for analysis of sequence information obtained by mass spectrometry are able to identify even closely related sequences at statistically significant scores. To address this deficit, new bioinformatics approaches were developed for concomitant use of the multiple fragments of short sequence typically available from methods of tandem mass spectrometry. These approaches rely on the occurrence of congruence across searches of multiple fragments from a single protein. This method resulted in sharply better statistical significance values for correct hits in the data base output relative to that achieved for independent searches using single sequence fragments.

Algorithms↗

A genetic algorithm for structure-based de novo design.

Genetic algorithms have properties which make them attractive in de novo drug design. Like other de novo design programs, genetic algorithms require a method to reduce the enormous search space of possible compounds. Most often this is done using information from known ligands. We have developed the ADAPT program, a genetic algorithm which uses molecular interactions evaluated with docking calculations as a fitness function to reduce the search space. ADAPT does not require information about known ligands. The program takes an initial set of compounds and iteratively builds new compounds based on the fitness scores of the previous set of compounds. We describe the particulars of the ADAPT algorithm and its application to three well-studied target systems. We also show that the strategies of enhanced local sampling and re-introducing diversity to the compound population during the design cycle provide better results than conventional genetic algorithm protocols.

Algorithms↗

Shotgun: getting more from sequence similarity searches.

MOTIVATION: As genomic sequencing reveals the range of structural classes generated through the evolution of proteins, analysis of the superfamilies to which they belong can contribute important insights for understanding their structure-function relationships. Current database search techniques fall short of identifying the majority of distant sequence relationships at statistically significant levels. We developed the Shotgun program in an effort to enhance the sensitivity and utility of current database search output. RESULTS: We have developed and used the Shotgun program to identify both new superfamily members and to reconstruct several known enzyme superfamilies using BLAST database searches. An analysis of the false-positive rates generated in the analysis and other control experiments provides evidence that high Shotgun scores indicate real evolutionary relationships. Shotgun is also a useful tool for identifying subgroup relationships within superfamilies and for testing hypotheses about related protein families. AVAILABILITY: By request from the Babbitt lab homepage: http://mako.cgl.ucsf. edu/babbittlab/ CONTACT: babbitt@cgl.ucsf.edu

Algorithms↗