PubMed Health⌕ Search

Biomedical subjects

Francis D Gibbons

Publications and source records attributed to Francis D Gibbons.

7 recordsLinked to original sources

Chipper: discovering transcription-factor targets from chromatin immunoprecipitation microarrays using variance stabilization.

Chromatin immunoprecipitation combined with microarray technology (Chip2) allows genome-wide determination of protein-DNA binding sites. The current standard method for analyzing Chip2 data requires additional control experiments that are subject to systematic error. We developed methods to assess significance using variance stabilization, learning error-model parameters without external control experiments. The method was validated experimentally, shows greater sensitivity than the current standard method, and incorporates false-discovery rate analysis. The corresponding software ('Chipper') is freely available. The method described here should help reveal an organism's transcription-regulatory 'wiring diagram'.

Analysis of Variance↗

Towards a proteome-scale map of the human protein-protein interaction network.

Systematic mapping of protein-protein interactions, or 'interactome' mapping, was initiated in model organisms, starting with defined biological processes and then expanding to the scale of the proteome. Although far from complete, such maps have revealed global topological and dynamic features of interactome networks that relate to known biological properties, suggesting that a human interactome map will provide insight into development and disease mechanisms at a systems level. Here we describe an initial version of a proteome-scale map of human binary protein-protein interactions. Using a stringent, high-throughput yeast two-hybrid system, we tested pairwise interactions among the products of approximately 8,100 currently available Gateway-cloned open reading frames and detected approximately 2,800 interactions. This data set, called CCSB-HI1, has a verification rate of approximately 78% as revealed by an independent co-affinity purification assay, and correlates significantly with other biological attributes. The CCSB-HI1 data set increases by approximately 70% the set of available binary interactions within the tested space and reveals more than 300 new connections to over 100 disease-associated proteins. This work represents an important step towards a systematic and comprehensive human interactome project.

Cloning, Molecular↗

Genomewide identification of Sko1 target promoters reveals a regulatory network that operates in response to osmotic stress in Saccharomyces cerevisiae.

In Saccharomyces cerevisiae, the ATF/CREB transcription factor Sko1 (Acr1) regulates the expression of genes induced by osmotic stress under the control of the high osmolarity glycerol (HOG) mitogen-activated protein kinase pathway. By combining chromatin immunoprecipitation and microarrays containing essentially all intergenic regions, we estimate that yeast cells contain approximately 40 Sko1 target promoters in vivo; 20 Sko1 target promoters were validated by direct analysis of individual loci. The ATF/CREB consensus sequence is not statistically overrepresented in confirmed Sko1 target promoters, although some sites are evolutionarily conserved among related yeast species, suggesting that they are functionally important in vivo. These observations suggest that Sko1 association in vivo is affected by factors beyond the protein-DNA interaction defined in vitro. Sko1 binds a number of promoters for genes directly involved in defense functions that relieve osmotic stress. In addition, Sko1 binds to the promoters of genes encoding transcription factors, including Msn2, Mot3, Rox1, Mga1, and Gat2. Stress-induced expression of MSN2, MOT3, and MGA1 is diminished in sko1 mutant cells, while transcriptional regulation of ROX1 seems to be unaffected. Lastly, Sko1 targets PTP3, which encodes a phosphatase that negatively regulates Hog1 kinase activity, and Sko1 is required for osmotic induction of PTP3 expression. Taken together our results suggest that Sko1 operates a transcriptional network upon osmotic stress, which involves other specific transcription factors and a phosphatase that regulates the key component of the signal transduction pathway.

Base Sequence↗

Predicting protein complex membership using probabilistic network reliability.

Evidence for specific protein-protein interactions is increasingly available from both small- and large-scale studies, and can be viewed as a network. It has previously been noted that errors are frequent among large-scale studies, and that error frequency depends on the large-scale method used. Despite knowledge of the error-prone nature of interaction evidence, edges (connections) in this network are typically viewed as either present or absent. However, use of a probabilistic network that considers quantity and quality of supporting evidence should improve inference derived from protein networks. Here we demonstrate inference of membership in a partially known protein complex by using a probabilistic network model and an algorithm previously used to evaluate reliability in communication networks.

Fungal Proteins↗

Intensity-based protein identification by machine learning from a library of tandem mass spectra.

Tandem mass spectrometry (MS/MS) has emerged as a cornerstone of proteomics owing in part to robust spectral interpretation algorithms. Widely used algorithms do not fully exploit the intensity patterns present in mass spectra. Here, we demonstrate that intensity pattern modeling improves peptide and protein identification from MS/MS spectra. We modeled fragment ion intensities using a machine-learning approach that estimates the likelihood of observed intensities given peptide and fragment attributes. From 1,000,000 spectra, we chose 27,000 with high-quality, nonredundant matches as training data. Using the same 27,000 spectra, intensity was similarly modeled with mismatched peptides. We used these two probabilistic models to compute the relative likelihood of an observed spectrum given that a candidate peptide is matched or mismatched. We used a 'decoy' proteome approach to estimate incorrect match frequency, and demonstrated that an intensity-based method reduces peptide identification error by 50-96% without any loss in sensitivity.

Algorithms↗

SILVER helps assign peptides to tandem mass spectra using intensity-based scoring.

Tandem mass spectrometry is commonly used to identify peptides (and thereby proteins) that are present in complex mixtures. Peptide identification from tandem mass spectra is partially automated, but still requires human curation to resolve "borderline" peptide-spectrum matches (PSMs). SILVER is web-based software that assists manual curation of tandem mass spectra, using a recently developed intensity-based machine-learning approach to scoring PSMs, Elias et al. In this method, a large training set of peptide, fragment, and peak-intensity properties for both matched and mismatched PSMs was used to develop a score measuring consistency between each predicted fragment ion of a candidate peptide and its corresponding observed spectral peak intensity. The SILVER interface provides a visual representation of match quality between each candidate fragment ion and the observed spectrum, thereby expediting manual curation of tandem mass spectra. SILVER is available online at http://llama.med.harvard.edu/Software.html.

Amino Acid Sequence↗

Judging the quality of gene expression-based clustering methods using gene annotation.

We compare several commonly used expression-based gene clustering algorithms using a figure of merit based on the mutual information between cluster membership and known gene attributes. By studying various publicly available expression data sets we conclude that enrichment of clusters for biological function is, in general, highest at rather low cluster numbers. As a measure of dissimilarity between the expression patterns of two genes, no method outperforms Euclidean distance for ratio-based measurements, or Pearson distance for non-ratio-based measurements at the optimal choice of cluster number. We show the self-organized-map approach to be best for both measurement types at higher numbers of clusters. Clusters of genes derived from single- and average-linkage hierarchical clustering tend to produce worse-than-random results.

Algorithms↗