Publisher Correction: Translating genomic data into healthcare practice with the Singapore National Precision Medicine program.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to Nicolas Bertin.
Explore the source record for details and available documents.
Although numerous fundamental aspects of development have been uncovered through the study of individual genes and proteins, system-level models are still missing for most developmental processes. The first two cell divisions of Caenorhabditis elegans embryogenesis constitute an ideal test bed for a system-level approach. Early embryogenesis, including processes such as cell division and establishment of cellular polarity, is readily amenable to large-scale functional analysis. A first step toward a system-level understanding is to provide 'first-draft' models both of the molecular assemblies involved and of the functional connections between them. Here we show that such models can be derived from an integrated gene/protein network generated from three different types of functional relationship: protein interaction, expression profiling similarity and phenotypic profiling similarity, as estimated from detailed early embryonic RNA interference phenotypes systematically recorded for hundreds of early embryogenesis genes. The topology of the integrated network suggests that C. elegans early embryogenesis is achieved through coordination of a limited set of molecular machines. We assessed the overall predictive value of such molecular machine models by dynamic localization of ten previously uncharacterized proteins within the living embryo.
RNA interference (RNAi) of target genes is triggered by double-stranded RNAs (dsRNAs) processed by conserved nucleases and accessory factors. To identify the genetic components required for RNAi, we performed a genome-wide screen using an engineered RNAi sensor strain of Caenorhabditis elegans. The RNAi screen identified 90 genes. These included Piwi/PAZ proteins, DEAH helicases, RNA binding/processing factors, chromatin-associated factors, DNA recombination proteins, nuclear import/export factors, and 11 known components of the RNAi machinery. We demonstrate that some of these genes are also required for germline and somatic transgene silencing. Moreover, the physical interactions among these potential RNAi factors suggest links to other RNA-dependent gene regulatory pathways.
Currently available protein-protein interaction (PPI) network or 'interactome' maps, obtained with the yeast two-hybrid (Y2H) assay or by co-affinity purification followed by mass spectrometry (co-AP/MS), only cover a fraction of the complete PPI networks. These partial networks display scale-free topologies--most proteins participate in only a few interactions whereas a few proteins have many interaction partners. Here we analyze whether the scale-free topologies of the partial networks obtained from Y2H assays can be used to accurately infer the topology of complete interactomes. We generated four theoretical interaction networks of different topologies (random, exponential, power law, truncated normal). Partial sampling of these networks resulted in sub-networks with topological characteristics that were virtually indistinguishable from those of currently available Y2H-derived partial interactome maps. We conclude that given the current limited coverage levels, the observed scale-free topology of existing interactome maps cannot be confidently extrapolated to complete interactomes.
In apparently scale-free protein-protein interaction networks, or 'interactome' networks, most proteins interact with few partners, whereas a small but significant proportion of proteins, the 'hubs', interact with many partners. Both biological and non-biological scale-free networks are particularly resistant to random node removal but are extremely sensitive to the targeted removal of hubs. A link between the potential scale-free topology of interactome networks and genetic robustness seems to exist, because knockouts of yeast genes encoding hubs are approximately threefold more likely to confer lethality than those of non-hubs. Here we investigate how hubs might contribute to robustness and other cellular properties for protein-protein interactions dynamically regulated both in time and in space. We uncovered two types of hub: 'party' hubs, which interact with most of their partners simultaneously, and 'date' hubs, which bind their different partners at different times or locations. Both in silico studies of network connectivity and genetic interactions described in vivo support a model of organized modularity in which date hubs organize the proteome, connecting biological processes--or modules--to each other, whereas party hubs function inside modules.
To initiate a system-level analysis of C. elegans DAF-7/TGF-beta signaling, we combined interactome mapping with single and double genetic perturbations. Yeast two-hybrid (Y2H) screens starting with known DAF-7/TGF-beta pathway components defined a network of 71 interactions among 59 proteins. Coaffinity purification (co-AP) assays in mammalian cells confirmed the overall quality of this network. Systematic perturbations of the network using RNAi, both in wild-type and daf-7/TGF-beta pathway mutant animals, identified nine DAF-7/TGF-beta signaling modifiers, seven of which are conserved in humans. We show that one of these has functional homology to human SNO/SKI oncoproteins and that mutations at the corresponding genetic locus daf-5 confer defects in DAF-7/TGF-beta signaling. Our results reveal substantial molecular complexity in DAF-7/TGF-beta signal transduction. Integrating interactome maps with systematic genetic perturbations may be useful for developing a systems biology approach to this and other signaling modules.
To initiate studies on how protein-protein interaction (or "interactome") networks relate to multicellular functions, we have mapped a large fraction of the Caenorhabditis elegans interactome network. Starting with a subset of metazoan-specific proteins, more than 4000 interactions were identified from high-throughput, yeast two-hybrid (HT=Y2H) screens. Independent coaffinity purification assays experimentally validated the overall quality of this Y2H data set. Together with already described Y2H interactions and interologs predicted in silico, the current version of the Worm Interactome (WI5) map contains approximately 5500 interactions. Topological and biological features of this interactome network, as well as its integration with phenome and transcriptome data sets, lead to numerous biological hypotheses.
Proteins function mainly through interactions, especially with DNA and other proteins. While some large-scale interaction networks are now available for a number of model organisms, their experimental generation remains difficult. Consequently, interolog mapping--the transfer of interaction annotation from one organism to another using comparative genomics--is of significant value. Here we quantitatively assess the degree to which interologs can be reliably transferred between species as a function of the sequence similarity of the corresponding interacting proteins. Using interaction information from Saccharomyces cerevisiae, Caenorhabditis elegans, Drosophila melanogaster, and Helicobacter pylori, we find that protein-protein interactions can be transferred when a pair of proteins has a joint sequence identity >80% or a joint E-value <10(-70). (These "joint" quantities are the geometric means of the identities or E-values for the two pairs of interacting proteins.) We generalize our interolog analysis to protein-DNA binding, finding such interactions are conserved at specific thresholds between 30% and 60% sequence identity depending on the protein family. Furthermore, we introduce the concept of a "regulog"--a conserved regulatory relationship between proteins across different species. We map interologs and regulogs from yeast to a number of genomes with limited experimental annotation (e.g., Arabidopsis thaliana) and make these available through an online database at http://interolog.gersteinlab.org. Specifically, we are able to transfer approximately 90,000 potential protein-protein interactions to the worm. We test a number of these in two-hybrid experiments and are able to verify 45 overlaps, which we show to be statistically significant.
The bacteria of the Brucella genus are responsible for a worldwide zoonosis called brucellosis. They belong to the alpha-proteobacteria group, as many other bacteria that live in close association with a eukaryotic host. Importantly, the Brucellae are mainly intracellular pathogens, and the molecular mechanisms of their virulence are still poorly understood. Using the complete genome sequence of Brucella melitensis, we generated a database of protein-coding open reading frames (ORFs) and constructed an ORFeome library of 3091 Gateway Entry clones, each containing a defined ORF. This first version of the Brucella ORFeome (v1.1) provides the coding sequences in a user-friendly format amenable to high-throughput functional genomic and proteomic experiments, as the ORFs are conveniently transferable from the Entry clones to various Expression vectors by recombinational cloning. The cloning of the Brucella ORFeome v1.1 should help to provide a better understanding of the molecular mechanisms of virulence, including the identification of bacterial protein-protein interactions, but also interactions between bacterial effectors and their host's targets.
The advent of systems biology necessitates the cloning of nearly entire sets of protein-encoding open reading frames (ORFs), or ORFeomes, to allow functional studies of the corresponding proteomes. Here, we describe the generation of a first version of the human ORFeome using a newly improved Gateway recombinational cloning approach. Using the Mammalian Gene Collection (MGC) resource as a starting point, we report the successful cloning of 8076 human ORFs, representing at least 7263 human genes, as mini-pools of PCR-amplified products. These were assembled into the human ORFeome version 1.1 (hORFeome v1.1) collection. After assessing the overall quality of this version, we describe the use of hORFeome v1.1 for heterologous protein expression in two different expression systems at proteome scale. The hORFeome v1.1 represents a central resource for the cloning of large sets of human ORFs in various settings for functional proteomics of many types, and will serve as the foundation for subsequent improved versions of the human ORFeome.
To verify the genome annotation and to create a resource to functionally characterize the proteome, we attempted to Gateway-clone all predicted protein-encoding open reading frames (ORFs), or the 'ORFeome,' of Caenorhabditis elegans. We successfully cloned approximately 12,000 ORFs (ORFeome 1.1), of which roughly 4,000 correspond to genes that are untouched by any cDNA or expressed-sequence tag (EST). More than 50% of predicted genes needed corrections in their intron-exon structures. Notably, approximately 11,000 C. elegans proteins can now be expressed under many conditions and characterized using various high-throughput strategies, including large-scale interactome mapping. We suggest that similar ORFeome projects will be valuable for other organisms, including humans.
By integrating functional genomic and proteomic mapping approaches, biological hypotheses should be formulated with increasing levels of confidence. For example, yeast interactome and transcriptome data can be correlated in biologically meaningful ways. Here, we combine interactome mapping data generated for a multicellular organism with data from both large-scale phenotypic analysis ("phenome mapping") and transcriptome profiling. First, we generated a two-hybrid interactome map of the Caenorhabditis elegans germline by using 600 transcripts enriched in this tissue. We compared this map to a phenome map of the germline obtained by RNA interference (RNAi) and to a transcriptome map obtained by clustering worm genes across 553 expression profiling experiments. In this dataset, we find that essential proteins have a tendency to interact with each other, that pairs of genes encoding interacting proteins tend to exhibit similar expression profiles, and that, for approximately 24% of germline interactions, both partners show overlapping embryonic lethal or high incidence of males RNAi phenotypes and similar expression profiles. We propose that these interactions are most likely to be relevant to germline biology. Similar integration of interactome, phenome, and transcriptome data should be possible for other biological processes in the nematode and for other organisms, including humans.