PubMed Health⌕ Search

Biomedical subjects

Frederick P Roth

Publications and source records attributed to Frederick P Roth.

At least 19 recordsLinked to original sources

'Truthsets' for clinical validation of large-scale functional assays: Practice recommendations from Cancer Variant Interpretation Group UK (CanVIG-UK).

BACKGROUND: Large-scale functional assays, including multiplex assays of variant effect, have substantial potential to resolve variants of uncertain significance (VUS), particularly for rare missense variants where clinical and population evidence are limited. The ClinGen assay-level clinical validation framework described by Brnich et al provided baseline guidance for the use of functional data for variant classification. However, clear consensus regarding construction of variant 'truthsets' by which to clinically validate functional data remains lacking. METHODS: CanVIG-UK developed consensus recommendations for truthset construction through an iterative national consultation process involving the CanVIG Steering Advisory Group (CStAG), wider CanVIG-UK membership, and engagement with international functional genomics experts. Consultation was based on previous analyses of 2,120 truthset constructions examining the impact of truthset composition on evidence point allocation within the ClinGen assay-level clinical validation framework. RESULTS: Across several consultations, CanVIG-UK established nine guiding principles and seven best-practice recommendations for assay-level clinical validation, using the assumed context of an assay for a cancer susceptibility gene where loss-of-function is the mechanism of pathogenicity. The principal recommendation stipulates, where assays are intended for use in interpretation of largely missense variants, the truthset used to validate should comprise only missense variants. Rather than mixtures of different variant types which may serve to over-estimate assay performance. Additional recommendations support option for relaxation of truthset stringency to improve power, augmentation of benign missense truthsets with systematically derived 'proxy-clinical' benign variants, independent clinical validation separate from assayist-defined validation, and careful evaluation of missense score distributions against that of protein-truncating and synonymous variants. Guidance is also provided for scenarios with limited pathogenic truthset availability and for assays reporting multiple deleterious zones or readouts. CONCLUSIONS: The CanVIG-UK principles and recommendations for truthset construction upon the ClinGen assay-level clinical validation framework, while aiming to form a baseline for future discussion regarding other functional and disease contexts and helping to address the gap between publication of new data and routine clinical implementation.

Journal Article↗

BOGO: A Proteome-Wide Gene Overexpression Platform for Discovering Rational Cancer Combination Therapies.

Cancer drug resistance remains a major barrier to durable treatment success, often leading to relapse despite advances in precision oncology. While combination therapies are being increasingly investigated, such as chemotherapy with small molecule inhibitors, predicting drug response and identifying rational drug combinations based on resistance mechanisms remain major challenges. Therefore, a proteome-wide, single-gene overexpression screening platform is essential for guiding rational therapy selection. Here, we present BOGO (Bxb1-landing pad human ORFeome-integrated system for a proteome-wide Gene Overexpression), a robust, scalable, and reproducible screening platform that enables single-copy, site-specific integration and overexpression of ~19,000 human open across cancer cell models. Using BOGO, we identified drug-specific response drivers for 16 chemotherapeutic agents and integrated clinical datasets to uncover proliferation and resistance-associated genes with prognostic potential. Drug response similarity networks revealed both shared and unique mechanisms, highlighting key pathways such as autophagy, apoptosis, and Wnt signaling, and notable resistance-associated genes including BCL2, POLD2, and TRADD. In particular, we proposed a synergistic combination of the BCL2 family inhibitor ABT-263 (Navitoclax®) and the DNA analog TAS-102 (Lonsurf®), which revealed that lysosomal modulation is a key mechanism driving DNA analog resistance. This combination therapy selectively enhanced cytotoxicity in colorectal and pancreatic cancer cells in vitro, and demonstrated therapeutic benefit in vivo in both cell line-derived xenograft (CDX) and patient-derived xenograft (PDX) models. Together, these findings establish BOGO as a powerful gene overexpression perturbation platform for systematically identifying chemoresistance and chemosensitization drivers, and for discovering rational combination therapies. Its scalability and reproducibility position BOGO as a broadly applicable tool for functional genomics and therapeutic discovery beyond cancer resistance.

Journal Article↗

Creating an atlas of variant effects to resolve variants of uncertain significance and guide cardiovascular medicine.

Cardiovascular diseases are leading global causes of death and disability, often presenting as interrelated phenotypes of atherosclerotic vascular disease, heart failure and arrhythmias. Cardiovascular diseases arise from interactions between environmental factors and predisposing genotypes and include common Mendelian lipid disorders, cardiomyopathies and arrhythmia syndromes. The identification of a pathogenic variant through genetic testing can inform disease diagnosis, risk prediction, treatment and family screening. However, a major roadblock in genomic medicine is that for many variants, especially missense variants, we lack sufficient evidence to enable a definitive classification, and therefore these variants are deemed as 'variants of uncertain significance'. In this Review, we describe how multiplexed assays of variant effects can enable the functional assessment of nearly all coding variants in a target sequence, potentially offering a proactive approach to identifying the functional significance of gene variants that are observed later in a patient. We discuss validation, including the role of in silico variant effect predictors, and how multiplexed experimental methods are informing cardiovascular disease biology and ultimately resolving the problem of variants of uncertain significance at scale.

Humans↗

Combining biological networks to predict genetic interactions.

Genetic interactions define overlapping functions and compensatory pathways. In particular, synthetic sick or lethal (SSL) genetic interactions are important for understanding how an organism tolerates random mutation, i.e., genetic robustness. Comprehensive identification of SSL relationships remains far from complete in any organism, because mapping these networks is highly labor intensive. The ability to predict SSL interactions, however, could efficiently guide further SSL discovery. Toward this end, we predicted pairs of SSL genes in Saccharomyces cerevisiae by using probabilistic decision trees to integrate multiple types of data, including localization, mRNA expression, physical interaction, protein function, and characteristics of network topology. Experimental evidence demonstrated the reliability of this strategy, which, when extended to human SSL interactions, may prove valuable in discovering drug targets for cancer therapy and in identifying genes responsible for multigenic diseases.

Animals↗

Evidence for dynamically organized modularity in the yeast protein-protein interaction network.

In apparently scale-free protein-protein interaction networks, or 'interactome' networks, most proteins interact with few partners, whereas a small but significant proportion of proteins, the 'hubs', interact with many partners. Both biological and non-biological scale-free networks are particularly resistant to random node removal but are extremely sensitive to the targeted removal of hubs. A link between the potential scale-free topology of interactome networks and genetic robustness seems to exist, because knockouts of yeast genes encoding hubs are approximately threefold more likely to confer lethality than those of non-hubs. Here we investigate how hubs might contribute to robustness and other cellular properties for protein-protein interactions dynamically regulated both in time and in space. We uncovered two types of hub: 'party' hubs, which interact with most of their partners simultaneously, and 'date' hubs, which bind their different partners at different times or locations. Both in silico studies of network connectivity and genetic interactions described in vivo support a model of organized modularity in which date hubs organize the proteome, connecting biological processes--or modules--to each other, whereas party hubs function inside modules.

Computer Simulation↗

Prediction of similarly acting cis-regulatory modules by subsequence profiling and comparative genomics in Drosophila melanogaster and D.pseudoobscura.

MOTIVATION: To date, computational searches for cis-regulatory modules (CRMs) have relied on two methods. The first, phylogenetic footprinting, has been used to find CRMs in non-coding sequence, but does not directly link DNA sequence with spatio-temporal patterns of expression. The second, based on searches for combinations of transcription factor (TF) binding motifs, has been employed in genome-wide discovery of similarly acting enhancers, but requires prior knowledge of the set of TFs acting at the CRM and the TFs' binding motifs. RESULTS: We propose a method for CRM discovery that combines aspects of both approaches in an effort to overcome their individual limitations. By treating phylogenetically footprinted non-coding regions (PFRs) as proxies for CRMs, we endeavor to find PFRs near co-regulated genes that are comprised of similar short, conserved sequences. Using Markov chains as a convenient formulation to assess similarity, we develop a sampling algorithm to search a large group of PFRs for the most similar subset. When starting with a set of genes involved in Drosophila early blastoderm development and using phylogenetic comparisons of Drosophila melanogaster and D.pseudoobscura genomes, we show here that our algorithm successfully detects known CRMs. Further, we use our similarity metric, based on Markov chain discrimination, in a genome-wide search, and uncover additional known and many candidate early blastoderm CRMs. AVAILABILITY: Software is available via http://arep.med.harvard.edu/enhancer

Algorithms↗

Predicting protein complex membership using probabilistic network reliability.

Evidence for specific protein-protein interactions is increasingly available from both small- and large-scale studies, and can be viewed as a network. It has previously been noted that errors are frequent among large-scale studies, and that error frequency depends on the large-scale method used. Despite knowledge of the error-prone nature of interaction evidence, edges (connections) in this network are typically viewed as either present or absent. However, use of a probabilistic network that considers quantity and quality of supporting evidence should improve inference derived from protein networks. Here we demonstrate inference of membership in a partially known protein complex by using a probabilistic network model and an algorithm previously used to evaluate reliability in communication networks.

Fungal Proteins↗

Predicting co-complexed protein pairs using genomic and proteomic data integration.

BACKGROUND: Identifying all protein-protein interactions in an organism is a major objective of proteomics. A related goal is to know which protein pairs are present in the same protein complex. High-throughput methods such as yeast two-hybrid (Y2H) and affinity purification coupled with mass spectrometry (APMS) have been used to detect interacting proteins on a genomic scale. However, both Y2H and APMS methods have substantial false-positive rates. Aside from high-throughput interaction screens, other gene- or protein-pair characteristics may also be informative of physical interaction. Therefore it is desirable to integrate multiple datasets and utilize their different predictive value for more accurate prediction of co-complexed relationship. RESULTS: Using a supervised machine learning approach--probabilistic decision tree, we integrated high-throughput protein interaction datasets and other gene- and protein-pair characteristics to predict co-complexed pairs (CCP) of proteins. Our predictions proved more sensitive and specific than predictions based on Y2H or APMS methods alone or in combination. Among the top predictions not annotated as CCPs in our reference set (obtained from the MIPS complex catalogue), a significant fraction was found to physically interact according to a separate database (YPD, Yeast Proteome Database), and the remaining predictions may potentially represent unknown CCPs. CONCLUSIONS: We demonstrated that the probabilistic decision tree approach can be successfully used to predict co-complexed protein (CCP) pairs from other characteristics. Our top-scoring CCP predictions provide testable hypotheses for experimental validation.

Computational Biology↗

Global mapping of the yeast genetic interaction network.

A genetic interaction network containing approximately 1000 genes and approximately 4000 interactions was mapped by crossing mutations in 132 different query genes into a set of approximately 4700 viable gene yeast deletion mutants and scoring the double mutant progeny for fitness defects. Network connectivity was predictive of function because interactions often occurred among functionally related genes, and similar patterns of interactions tended to identify components of the same pathway. The genetic network exhibited dense local neighborhoods; therefore, the position of a gene on a partially mapped network is predictive of other genetic interactions. Because digenic interactions are common in yeast, similar networks may underlie the complex genetics associated with inherited phenotypes in other organisms.

Amino Acid Sequence↗

Intensity-based protein identification by machine learning from a library of tandem mass spectra.

Tandem mass spectrometry (MS/MS) has emerged as a cornerstone of proteomics owing in part to robust spectral interpretation algorithms. Widely used algorithms do not fully exploit the intensity patterns present in mass spectra. Here, we demonstrate that intensity pattern modeling improves peptide and protein identification from MS/MS spectra. We modeled fragment ion intensities using a machine-learning approach that estimates the likelihood of observed intensities given peptide and fragment attributes. From 1,000,000 spectra, we chose 27,000 with high-quality, nonredundant matches as training data. Using the same 27,000 spectra, intensity was similarly modeled with mismatched peptides. We used these two probabilistic models to compute the relative likelihood of an observed spectrum given that a candidate peptide is matched or mismatched. We used a 'decoy' proteome approach to estimate incorrect match frequency, and demonstrated that an intensity-based method reduces peptide identification error by 50-96% without any loss in sensitivity.

Algorithms↗

A map of the interactome network of the metazoan C. elegans.

To initiate studies on how protein-protein interaction (or "interactome") networks relate to multicellular functions, we have mapped a large fraction of the Caenorhabditis elegans interactome network. Starting with a subset of metazoan-specific proteins, more than 4000 interactions were identified from high-throughput, yeast two-hybrid (HT=Y2H) screens. Independent coaffinity purification assays experimentally validated the overall quality of this Y2H data set. Together with already described Y2H interactions and interologs predicted in silico, the current version of the Worm Interactome (WI5) map contains approximately 5500 interactions. Topological and biological features of this interactome network, as well as its integration with phenome and transcriptome data sets, lead to numerous biological hypotheses.

Animals↗

SILVER helps assign peptides to tandem mass spectra using intensity-based scoring.

Tandem mass spectrometry is commonly used to identify peptides (and thereby proteins) that are present in complex mixtures. Peptide identification from tandem mass spectra is partially automated, but still requires human curation to resolve "borderline" peptide-spectrum matches (PSMs). SILVER is web-based software that assists manual curation of tandem mass spectra, using a recently developed intensity-based machine-learning approach to scoring PSMs, Elias et al. In this method, a large training set of peptide, fragment, and peak-intensity properties for both matched and mismatched PSMs was used to develop a score measuring consistency between each predicted fragment ion of a candidate peptide and its corresponding observed spectral peak intensity. The SILVER interface provides a visual representation of match quality between each candidate fragment ion and the observed spectrum, thereby expediting manual curation of tandem mass spectra. SILVER is available online at http://llama.med.harvard.edu/Software.html.

Amino Acid Sequence↗

Characterizing gene sets with FuncAssociate.

SUMMARY: FuncAssociate is a web-based tool to help researchers use Gene Ontology attributes to characterize large sets of genes derived from experiment. Distinguishing features of FuncAssociate include the ability to handle ranked input lists, and a Monte Carlo simulation approach that is more appropriate to determine significance than other methods, such as Bonferroni or idák p-value correction. FuncAssociate currently supports 10 organisms (Vibrio cholerae, Shewanella oneidensis, Saccharomyces cerevisiae, Schizosaccharomyces pombe, Arabidopsis thaliana, Caenorhaebditis elegans, Drosophila melanogaster, Mus musculus, Rattus norvegicus and Homo sapiens). AVAILABILITY: FuncAssociate is freely accessible at http://llama.med.harvard.edu/Software.html. Source code (in Perl and C) is freely available to academic users 'as is'.

Algorithms↗

A non-parametric model for transcription factor binding sites.

We introduce a non-parametric representation of transcription factor binding sites which can model arbitrary dependencies between positions. As two parameters are varied, this representation smoothly interpolates between the empirical distribution of binding sites and the standard position-specific scoring matrix (PSSM). In a test of generalization to unseen binding sites using 10-fold cross-validation on known binding sites for 95 TRANSFAC transcription factors, this representation outperforms PSSMs on between 65 and 89 of the 95 transcription factors, depending on the choice of the two adjustable parameters. We also discuss how the non- parametric representation may be incorporated into frameworks for finding binding sites given only a collection of unaligned promoter regions.

Base Sequence↗

Predicting gene function from patterns of annotation.

The Gene Ontology (GO) Consortium has produced a controlled vocabulary for annotation of gene function that is used in many organism-specific gene annotation databases. This allows the prediction of gene function based on patterns of annotation. For example, if annotations for two attributes tend to occur together in a database, then a gene holding one attribute is likely to hold the other as well. We modeled the relationships among GO attributes with decision trees and Bayesian networks, using the annotations in the Saccharomyces Genome Database (SGD) and in FlyBase as training data. We tested the models using cross-validation, and we manually assessed 100 gene-attribute associations that were predicted by the models but that were not present in the SGD or FlyBase databases. Of the 100 manually assessed associations, 41 were judged to be true, and another 42 were judged to be plausible.

Animals↗

GoFish finds genes with combinations of Gene Ontology attributes.

SUMMARY: GoFish is a Java application that allows users to search for gene products with particular gene ontology (GO) attributes, or combinations of attributes. GoFish ranks gene products by the degree to which they satisfy a Boolean query. Four organisms are currently supported: Saccaromyces cerevisiae, Caenorhabditis elegans, Drosophila melanogaster, and M.musculus.

Amino Acid Sequence↗

Assessing experimentally derived interactions in a small world.

Experimentally determined networks are susceptible to errors, yet important inferences can still be drawn from them. Many real networks have also been shown to have the small-world network properties of cohesive neighborhoods and short average distances between vertices. Although much analysis has been done on small-world networks, small-world properties have not previously been used to improve our understanding of individual edges in experimentally derived graphs. Here we focus on a small-world network derived from high-throughput (and error-prone) protein-protein interaction experiments. We exploit the neighborhood cohesiveness property of small-world networks to assess confidence for individual protein-protein interactions. By ascertaining how well each protein-protein interaction (edge) fits the pattern of a small-world network, we stratify even those edges with identical experimental evidence. This result promises to improve the quality of inference from protein-protein interaction networks in particular and small-world networks in general.

Cluster Analysis↗

Predicting phenotype from patterns of annotation.

MOTIVATION: Predicting the outcome of specific experiments (such as the growth of a particular mutant strain in a particular medium) has the potential to allow researchers to devote resources to experiments with higher expected numbers of 'hits'. RESULTS: We use decision trees to predict phenotypes associated with Saccharomyces cerevisiae genes on the basis of Gene Ontology (GO) functional annotations from the Saccharomyces Genome Database (SGD) and other phenotypic annotations from the Yeast Phenotype Catalog at the Munich Information Center for Protein Sequences (MIPS). We assess the methodology in three ways: (1) we use cross-validation on the phenotypic annotations listed in MIPS, and show ROC curves indicating the tradeoff between true-positive rate and false-positive rate; (2) we do a literature-search for 100 of the predicted gene-phenotype associations that are not listed in MIPS, and find evidence for 43 of them; (3) we use deletion strains to experimentally assess 61 predicted gene-phenotype associations not listed in MIPS; significantly more of these deletion strains show abnormal growth than would be expected by chance.

Algorithms↗