PubMed Health⌕ Search

Biomedical subjects

Frederick P Roth

Publications and source records attributed to Frederick P Roth.

35 records · Page 2Linked to original sources

Predicting protein complex membership using probabilistic network reliability.

Evidence for specific protein-protein interactions is increasingly available from both small- and large-scale studies, and can be viewed as a network. It has previously been noted that errors are frequent among large-scale studies, and that error frequency depends on the large-scale method used. Despite knowledge of the error-prone nature of interaction evidence, edges (connections) in this network are typically viewed as either present or absent. However, use of a probabilistic network that considers quantity and quality of supporting evidence should improve inference derived from protein networks. Here we demonstrate inference of membership in a partially known protein complex by using a probabilistic network model and an algorithm previously used to evaluate reliability in communication networks.

Fungal Proteins↗

Predicting co-complexed protein pairs using genomic and proteomic data integration.

BACKGROUND: Identifying all protein-protein interactions in an organism is a major objective of proteomics. A related goal is to know which protein pairs are present in the same protein complex. High-throughput methods such as yeast two-hybrid (Y2H) and affinity purification coupled with mass spectrometry (APMS) have been used to detect interacting proteins on a genomic scale. However, both Y2H and APMS methods have substantial false-positive rates. Aside from high-throughput interaction screens, other gene- or protein-pair characteristics may also be informative of physical interaction. Therefore it is desirable to integrate multiple datasets and utilize their different predictive value for more accurate prediction of co-complexed relationship. RESULTS: Using a supervised machine learning approach--probabilistic decision tree, we integrated high-throughput protein interaction datasets and other gene- and protein-pair characteristics to predict co-complexed pairs (CCP) of proteins. Our predictions proved more sensitive and specific than predictions based on Y2H or APMS methods alone or in combination. Among the top predictions not annotated as CCPs in our reference set (obtained from the MIPS complex catalogue), a significant fraction was found to physically interact according to a separate database (YPD, Yeast Proteome Database), and the remaining predictions may potentially represent unknown CCPs. CONCLUSIONS: We demonstrated that the probabilistic decision tree approach can be successfully used to predict co-complexed protein (CCP) pairs from other characteristics. Our top-scoring CCP predictions provide testable hypotheses for experimental validation.

Computational Biology↗

Global mapping of the yeast genetic interaction network.

A genetic interaction network containing approximately 1000 genes and approximately 4000 interactions was mapped by crossing mutations in 132 different query genes into a set of approximately 4700 viable gene yeast deletion mutants and scoring the double mutant progeny for fitness defects. Network connectivity was predictive of function because interactions often occurred among functionally related genes, and similar patterns of interactions tended to identify components of the same pathway. The genetic network exhibited dense local neighborhoods; therefore, the position of a gene on a partially mapped network is predictive of other genetic interactions. Because digenic interactions are common in yeast, similar networks may underlie the complex genetics associated with inherited phenotypes in other organisms.

Amino Acid Sequence↗

Intensity-based protein identification by machine learning from a library of tandem mass spectra.

Tandem mass spectrometry (MS/MS) has emerged as a cornerstone of proteomics owing in part to robust spectral interpretation algorithms. Widely used algorithms do not fully exploit the intensity patterns present in mass spectra. Here, we demonstrate that intensity pattern modeling improves peptide and protein identification from MS/MS spectra. We modeled fragment ion intensities using a machine-learning approach that estimates the likelihood of observed intensities given peptide and fragment attributes. From 1,000,000 spectra, we chose 27,000 with high-quality, nonredundant matches as training data. Using the same 27,000 spectra, intensity was similarly modeled with mismatched peptides. We used these two probabilistic models to compute the relative likelihood of an observed spectrum given that a candidate peptide is matched or mismatched. We used a 'decoy' proteome approach to estimate incorrect match frequency, and demonstrated that an intensity-based method reduces peptide identification error by 50-96% without any loss in sensitivity.

Algorithms↗

A map of the interactome network of the metazoan C. elegans.

To initiate studies on how protein-protein interaction (or "interactome") networks relate to multicellular functions, we have mapped a large fraction of the Caenorhabditis elegans interactome network. Starting with a subset of metazoan-specific proteins, more than 4000 interactions were identified from high-throughput, yeast two-hybrid (HT=Y2H) screens. Independent coaffinity purification assays experimentally validated the overall quality of this Y2H data set. Together with already described Y2H interactions and interologs predicted in silico, the current version of the Worm Interactome (WI5) map contains approximately 5500 interactions. Topological and biological features of this interactome network, as well as its integration with phenome and transcriptome data sets, lead to numerous biological hypotheses.

Animals↗

SILVER helps assign peptides to tandem mass spectra using intensity-based scoring.

Tandem mass spectrometry is commonly used to identify peptides (and thereby proteins) that are present in complex mixtures. Peptide identification from tandem mass spectra is partially automated, but still requires human curation to resolve "borderline" peptide-spectrum matches (PSMs). SILVER is web-based software that assists manual curation of tandem mass spectra, using a recently developed intensity-based machine-learning approach to scoring PSMs, Elias et al. In this method, a large training set of peptide, fragment, and peak-intensity properties for both matched and mismatched PSMs was used to develop a score measuring consistency between each predicted fragment ion of a candidate peptide and its corresponding observed spectral peak intensity. The SILVER interface provides a visual representation of match quality between each candidate fragment ion and the observed spectrum, thereby expediting manual curation of tandem mass spectra. SILVER is available online at http://llama.med.harvard.edu/Software.html.

Amino Acid Sequence↗

Characterizing gene sets with FuncAssociate.

SUMMARY: FuncAssociate is a web-based tool to help researchers use Gene Ontology attributes to characterize large sets of genes derived from experiment. Distinguishing features of FuncAssociate include the ability to handle ranked input lists, and a Monte Carlo simulation approach that is more appropriate to determine significance than other methods, such as Bonferroni or idák p-value correction. FuncAssociate currently supports 10 organisms (Vibrio cholerae, Shewanella oneidensis, Saccharomyces cerevisiae, Schizosaccharomyces pombe, Arabidopsis thaliana, Caenorhaebditis elegans, Drosophila melanogaster, Mus musculus, Rattus norvegicus and Homo sapiens). AVAILABILITY: FuncAssociate is freely accessible at http://llama.med.harvard.edu/Software.html. Source code (in Perl and C) is freely available to academic users 'as is'.

Algorithms↗

A non-parametric model for transcription factor binding sites.

We introduce a non-parametric representation of transcription factor binding sites which can model arbitrary dependencies between positions. As two parameters are varied, this representation smoothly interpolates between the empirical distribution of binding sites and the standard position-specific scoring matrix (PSSM). In a test of generalization to unseen binding sites using 10-fold cross-validation on known binding sites for 95 TRANSFAC transcription factors, this representation outperforms PSSMs on between 65 and 89 of the 95 transcription factors, depending on the choice of the two adjustable parameters. We also discuss how the non- parametric representation may be incorporated into frameworks for finding binding sites given only a collection of unaligned promoter regions.

Base Sequence↗

Predicting gene function from patterns of annotation.

The Gene Ontology (GO) Consortium has produced a controlled vocabulary for annotation of gene function that is used in many organism-specific gene annotation databases. This allows the prediction of gene function based on patterns of annotation. For example, if annotations for two attributes tend to occur together in a database, then a gene holding one attribute is likely to hold the other as well. We modeled the relationships among GO attributes with decision trees and Bayesian networks, using the annotations in the Saccharomyces Genome Database (SGD) and in FlyBase as training data. We tested the models using cross-validation, and we manually assessed 100 gene-attribute associations that were predicted by the models but that were not present in the SGD or FlyBase databases. Of the 100 manually assessed associations, 41 were judged to be true, and another 42 were judged to be plausible.

Animals↗

GoFish finds genes with combinations of Gene Ontology attributes.

SUMMARY: GoFish is a Java application that allows users to search for gene products with particular gene ontology (GO) attributes, or combinations of attributes. GoFish ranks gene products by the degree to which they satisfy a Boolean query. Four organisms are currently supported: Saccaromyces cerevisiae, Caenorhabditis elegans, Drosophila melanogaster, and M.musculus.

Amino Acid Sequence↗

Assessing experimentally derived interactions in a small world.

Experimentally determined networks are susceptible to errors, yet important inferences can still be drawn from them. Many real networks have also been shown to have the small-world network properties of cohesive neighborhoods and short average distances between vertices. Although much analysis has been done on small-world networks, small-world properties have not previously been used to improve our understanding of individual edges in experimentally derived graphs. Here we focus on a small-world network derived from high-throughput (and error-prone) protein-protein interaction experiments. We exploit the neighborhood cohesiveness property of small-world networks to assess confidence for individual protein-protein interactions. By ascertaining how well each protein-protein interaction (edge) fits the pattern of a small-world network, we stratify even those edges with identical experimental evidence. This result promises to improve the quality of inference from protein-protein interaction networks in particular and small-world networks in general.

Cluster Analysis↗

Predicting phenotype from patterns of annotation.

MOTIVATION: Predicting the outcome of specific experiments (such as the growth of a particular mutant strain in a particular medium) has the potential to allow researchers to devote resources to experiments with higher expected numbers of 'hits'. RESULTS: We use decision trees to predict phenotypes associated with Saccharomyces cerevisiae genes on the basis of Gene Ontology (GO) functional annotations from the Saccharomyces Genome Database (SGD) and other phenotypic annotations from the Yeast Phenotype Catalog at the Munich Information Center for Protein Sequences (MIPS). We assess the methodology in three ways: (1) we use cross-validation on the phenotypic annotations listed in MIPS, and show ROC curves indicating the tradeoff between true-positive rate and false-positive rate; (2) we do a literature-search for 100 of the predicted gene-phenotype associations that are not listed in MIPS, and find evidence for 43 of them; (3) we use deletion strains to experimentally assess 61 predicted gene-phenotype associations not listed in MIPS; significantly more of these deletion strains show abnormal growth than would be expected by chance.

Algorithms↗

Regulating general mutation rates: examination of the hypermutable state model for Cairnsian adaptive mutation.

In the lac adaptive mutation system of Cairns, selected mutant colonies but not unselected mutant types appear to arise from a nongrowing population of Escherichia coli. The general mutagenesis suffered by the selected mutants has been interpreted as support for the idea that E. coli possesses an evolved (and therefore beneficial) mechanism that increases the mutation rate in response to stress (the hypermutable state model, HSM). This mechanism is proposed to allow faster genetic adaptation to stressful conditions and to explain why mutations appear directed to useful sites. Analysis of the HSM reveals that it requires implausibly intense mutagenesis (10(5) times the unselected rate) and even then cannot account for the behavior of the Cairns system. The assumptions of the HSM predict that selected revertants will carry an average of eight deleterious null mutations and thus seem unlikely to be successful in long-term evolution. The experimentally observed 35-fold increase in the level of general mutagenesis cannot account for even one Lac(+) revertant from a mutagenized subpopulation of 10(5) cells (the number proposed to enter the hypermutable state). We conclude that temporary general mutagenesis during stress is unlikely to provide a long-term selective advantage in this or any similar genetic system.

Adaptation, Biological↗

Latent herpes simplex virus infection of sensory neurons alters neuronal gene expression.

The persistence of herpes simplex virus (HSV) and the diseases that it causes in the human population can be attributed to the maintenance of a latent infection within neurons in sensory ganglia. Little is known about the effects of latent infection on the host neuron. We have addressed the question of whether latent HSV infection affects neuronal gene expression by using microarray transcript profiling of host gene expression in ganglia from latently infected versus mock-infected mouse trigeminal ganglia. (33)P-labeled cDNA probes from pooled ganglia harvested at 30 days postinfection or post-mock infection were hybridized to nylon arrays printed with 2,556 mouse genes. Signal intensities were acquired by phosphorimager. Mean intensities (n = 4 replicates in each of three independent experiments) of signals from mock-infected versus latently infected ganglia were compared by using a variant of Student's t test. We identified significant changes in the expression of mouse neuronal genes, including several with roles in gene expression, such as the Clk2 gene, and neurotransmission, such as genes encoding potassium voltage-gated channels and a muscarinic acetylcholine receptor. We confirmed the neuronal localization of some of these transcripts by using in situ hybridization. To validate the microarray results, we performed real-time reverse transcriptase PCR analyses for a selection of the genes. These studies demonstrate that latent HSV infection can alter neuronal gene expression and might provide a new mechanism for how persistent viral infection can cause chronic disease.

Animals↗

The genome-wide localization of Rsc9, a component of the RSC chromatin-remodeling complex, changes in response to stress.

The cellular response to environmental changes includes widespread modifications in gene expression. Here we report the identification and characterization of Rsc9, a member of the RSC chromatin-remodeling complex in yeast. The genome-wide localization of Rsc9 indicated a relationship between genes targeted by Rsc9 and genes regulated by stress; treatment with hydrogen peroxide or rapamycin, which inhibits TOR signaling, resulted in genome-wide changes in Rsc9 occupancy. We further show that Rsc9 is involved in both repression and activation of mRNAs regulated by TOR as well as the synthesis of rRNA. Our results illustrate the response of a chromatin-remodeling factor to signaling cascades and suggest that changes in the activity of chromatin-remodeling factors are reflected in changes in their localization in the genome.

Amino Acid Sequence↗

Judging the quality of gene expression-based clustering methods using gene annotation.

We compare several commonly used expression-based gene clustering algorithms using a figure of merit based on the mutual information between cluster membership and known gene attributes. By studying various publicly available expression data sets we conclude that enrichment of clusters for biological function is, in general, highest at rather low cluster numbers. As a measure of dissimilarity between the expression patterns of two genes, no method outperforms Euclidean distance for ratio-based measurements, or Pearson distance for non-ratio-based measurements at the optimal choice of cluster number. We show the self-organized-map approach to be best for both measurement types at higher numbers of clusters. Clusters of genes derived from single- and average-linkage hierarchical clustering tend to produce worse-than-random results.

Algorithms↗

Using high-throughput screening data to discriminate compounds with single-target effects from those with side effects.

The most desirable compound leads from high-throughput assays are those with novel biological activities resulting from their action on a single biological target. Valuable resources can be wasted on compound leads with significant 'side effects' on additional biological targets; therefore, technical refinements to identify compounds that primarily have effects resulting from a single target are needed. This study explores the use of multiple assays of a chemical library and a statistic based on entropy to identify lead compound classes that have patterns of assay activity resulting primarily from small molecule action on a single target. This statistic, called the coincidence score, discriminates with 88% accuracy compound classes known to act primarily on a single target from compound classes with significant side effects on nonhomologous targets. Furthermore, a significant number of the compound classes predicted to have primarily single-target effects contain known bioactive compounds. We also show that a compound's known biological target or mechanism of action can often be suggested by its pattern of activities in multiple assays.

Drug-Related Side Effects and Adverse Reactions↗