PubMed Health⌕ Search

Biomedical subjects

Marc Vidal

Publications and source records attributed to Marc Vidal.

At least 19 recordsLinked to original sources

Towards a proteome-scale map of the human protein-protein interaction network.

Systematic mapping of protein-protein interactions, or 'interactome' mapping, was initiated in model organisms, starting with defined biological processes and then expanding to the scale of the proteome. Although far from complete, such maps have revealed global topological and dynamic features of interactome networks that relate to known biological properties, suggesting that a human interactome map will provide insight into development and disease mechanisms at a systems level. Here we describe an initial version of a proteome-scale map of human binary protein-protein interactions. Using a stringent, high-throughput yeast two-hybrid system, we tested pairwise interactions among the products of approximately 8,100 currently available Gateway-cloned open reading frames and detected approximately 2,800 interactions. This data set, called CCSB-HI1, has a verification rate of approximately 78% as revealed by an independent co-affinity purification assay, and correlates significantly with other biological attributes. The CCSB-HI1 data set increases by approximately 70% the set of available binary interactions within the tested space and reveals more than 300 new connections to over 100 disease-associated proteins. This work represents an important step towards a systematic and comprehensive human interactome project.

Cloning, Molecular↗

Interactome: gateway into systems biology.

Protein-protein interactions are fundamental to all biological processes, and a comprehensive determination of all protein-protein interactions that can take place in an organism provides a framework for understanding biology as an integrated system. The availability of genome-scale sets of cloned open reading frames has facilitated systematic efforts at creating proteome-scale data sets of protein-protein interactions, which are represented as complex networks or 'interactome' maps. Protein-protein interaction mapping projects that follow stringent criteria, coupled with experimental validation in orthogonal systems, provide high-confidence data sets immanently useful for interrogating developmental and disease mechanisms at a system level as well as elucidating individual protein function and interactome network topology. Although far from complete, currently available maps provide insight into how biochemical properties of proteins and protein complexes are integrated into biological systems. Such maps are also a useful resource to predict the function(s) of thousands of genes.

Animals↗

Pooled ORF expression technology (POET): using proteomics to screen pools of open reading frames for protein expression.

We have developed a pooled ORF expression technology, POET, that uses recombinational cloning and proteomic methods (two-dimensional gel electrophoresis and mass spectrometry) to identify ORFs that when expressed are likely to yield high levels of soluble, purified protein. Because the method works on pools of ORFs, the procedures needed to subclone, express, purify, and assay protein expression for hundreds of clones are greatly simplified. Small scale expression and purification of 12 positive clones identified by POET from a pool of 688 Caenorhabditis elegans ORFs expressed in Escherichia coli yielded on average 6 times as much protein as 12 negative clones. Larger scale expression and purification of six of the positive clones yielded 47-374 mg of purified protein/liter. Using POET, pools of ORFs can be constructed, and the pools of the resulting proteins can be analyzed and manipulated to rapidly acquire information about the attributes of hundreds of proteins simultaneously.

Animals↗

Predictive models of molecular machines involved in Caenorhabditis elegans early embryogenesis.

Although numerous fundamental aspects of development have been uncovered through the study of individual genes and proteins, system-level models are still missing for most developmental processes. The first two cell divisions of Caenorhabditis elegans embryogenesis constitute an ideal test bed for a system-level approach. Early embryogenesis, including processes such as cell division and establishment of cellular polarity, is readily amenable to large-scale functional analysis. A first step toward a system-level understanding is to provide 'first-draft' models both of the molecular assemblies involved and of the functional connections between them. Here we show that such models can be derived from an integrated gene/protein network generated from three different types of functional relationship: protein interaction, expression profiling similarity and phenotypic profiling similarity, as estimated from detailed early embryonic RNA interference phenotypes systematically recorded for hundreds of early embryogenesis genes. The topology of the integrated network suggests that C. elegans early embryogenesis is achieved through coordination of a limited set of molecular machines. We assessed the overall predictive value of such molecular machine models by dynamic localization of ten previously uncharacterized proteins within the living embryo.

Algorithms↗

Systematic analysis of genes required for synapse structure and function.

Chemical synapses are complex structures that mediate rapid intercellular signalling in the nervous system. Proteomic studies suggest that several hundred proteins will be found at synaptic specializations. Here we describe a systematic screen to identify genes required for the function or development of Caenorhabditis elegans neuromuscular junctions. A total of 185 genes were identified in an RNA interference screen for decreased acetylcholine secretion; 132 of these genes had not previously been implicated in synaptic transmission. Functional profiles for these genes were determined by comparing secretion defects observed after RNA interference under a variety of conditions. Hierarchical clustering identified groups of functionally related genes, including those involved in the synaptic vesicle cycle, neuropeptide signalling and responsiveness to phorbol esters. Twenty-four genes encoded proteins that were localized to presynaptic specializations. Loss-of-function mutations in 12 genes caused defects in presynaptic structure.

Aldicarb↗

lin-8, which antagonizes Caenorhabditis elegans Ras-mediated vulval induction, encodes a novel nuclear protein that interacts with the LIN-35 Rb protein.

Ras-mediated vulval development in C. elegans is inhibited by the functionally redundant sets of class A, B, and C synthetic Multivulva (synMuv) genes. Three of the class B synMuv genes encode an Rb/DP/E2F complex that, by analogy with its mammalian and Drosophila counterparts, has been proposed to silence genes required for vulval specification through chromatin modification and remodeling. Two class A synMuv genes, lin-15A and lin-56, encode novel nuclear proteins that appear to function as a complex. We show that a third class A synMuv gene, lin-8, is the defining member of a novel C. elegans gene family. The LIN-8 protein is nuclear and can interact physically with the product of the class B synMuv gene lin-35, the C. elegans homolog of mammalian Rb. LIN-8 likely acts with the synMuv A proteins LIN-15A and LIN-56 in the nucleus, possibly in a protein complex with the synMuv B protein LIN-35 Rb. Other LIN-8 family members may function in similar complexes in different cells or at different stages. The nuclear localization of LIN-15A, LIN-56, and LIN-8, as well as our observation of a direct physical interaction between class A and class B synMuv proteins, supports the hypothesis that the class A synMuv genes control vulval induction through the transcriptional regulation of gene expression.

Amino Acid Sequence↗

Local modeling of global interactome networks.

MOTIVATION: Systems biology requires accurate models of protein complexes, including physical interactions that assemble and regulate these molecular machines. Yeast two-hybrid (Y2H) and affinity-purification/mass-spectrometry (AP-MS) technologies measure different protein-protein relationships, and issues of completeness, sensitivity and specificity fuel debate over which is best for high-throughput 'interactome' data collection. Static graphs currently used to model Y2H and AP-MS data neglect dynamic and spatial aspects of macromolecular complexes and pleiotropic protein function. RESULTS: We apply the local modeling methodology proposed by Scholtens and Gentleman (2004) to two publicly available datasets and demonstrate its uses, interpretation and limitations. Specifically, we use this technology to address four major issues pertaining to protein-protein networks. (1) We motivate the need to move from static global interactome graphs to local protein complex models. (2) We formally show that accurate local interactome models require both Y2H and AP-MS data, even in idealized situations. (3) We briefly discuss experimental design issues and how bait selection affects interpretability of results. (4) We point to the implications of local modeling for systems biology including functional annotation, new complex prediction, pathway interactivity and coordination with gene-expression data. AVAILABILITY: The local modeling algorithm and all protein complex estimates reported here can be found in the R package apComplex, available at http://www.bioconductor.org CONTACT: dscholtens@northwestern.edu SUPPLEMENTARY INFORMATION: http://daisy.prevmed.northwestern.edu/~denise/pubs/LocalModeling

Algorithms↗

Biochemical clustering of monomeric GTPases of the Ras superfamily.

To date phylogeny has been used to compare entire families of proteins based on their nucleotide or amino acid sequence. Here we developed a novel analytical platform allowing a systematic comparison of protein families based on their biochemical properties. This approach was validated on the Rho subfamily of GTPases. We used two high throughput methods, referred to as AlphaScreen and FlashPlate, to measure nucleotide binding capacity, exchange, and hydrolysis activities of small monomeric GTPases. These two technologies have the characteristics to be very sensitive and to allow homogenous and high throughput assays. To analyze and integrate the data obtained, we developed an algorithm that allows the classification of GTPases according to their enzymatic activities. Integration and hierarchical clustering of these results revealed unexpected features of the small Rho GTPases when compared with primary sequence-based trees. Hence we propose a novel phylobiochemical classification of the Ras superfamily of GTPases.

Algorithms↗

Functional genomic analysis of RNA interference in C. elegans.

RNA interference (RNAi) of target genes is triggered by double-stranded RNAs (dsRNAs) processed by conserved nucleases and accessory factors. To identify the genetic components required for RNAi, we performed a genome-wide screen using an engineered RNAi sensor strain of Caenorhabditis elegans. The RNAi screen identified 90 genes. These included Piwi/PAZ proteins, DEAH helicases, RNA binding/processing factors, chromatin-associated factors, DNA recombination proteins, nuclear import/export factors, and 11 known components of the RNAi machinery. We demonstrate that some of these genes are also required for germline and somatic transgene silencing. Moreover, the physical interactions among these potential RNAi factors suggest links to other RNA-dependent gene regulatory pathways.

Amino Acid Motifs↗

Interactome modeling.

A long-term goal of the field of interactome modeling is to understand how global and local properties of complex macromolecular networks impact on observable biological properties, and how changes in such properties can lead to human diseases. The information available at this stage of development of the field provides strong evidence for the existence of such interesting global and local properties, but also demonstrates that many more datasets will be needed to provide accurate models with increasingly predictive capacity. This review focuses on an early attempt at mapping a multicellular interactome network and on the lessons learned from that attempt.

Animals↗

A gene expression fingerprint of C. elegans embryonic motor neurons.

BACKGROUND: Differential gene expression specifies the highly diverse cell types that constitute the nervous system. With its sequenced genome and simple, well-defined neuroanatomy, the nematode C. elegans is a useful model system in which to correlate gene expression with neuron identity. The UNC-4 transcription factor is expressed in thirteen embryonic motor neurons where it specifies axonal morphology and synaptic function. These cells can be marked with an unc-4::GFP reporter transgene. Here we describe a powerful strategy, Micro-Array Profiling of C. elegans cells (MAPCeL), and confirm that this approach provides a comprehensive gene expression profile of unc-4::GFP motor neurons in vivo. RESULTS: Fluorescence Activated Cell Sorting (FACS) was used to isolate unc-4::GFP neurons from primary cultures of C. elegans embryonic cells. Microarray experiments detected 6,217 unique transcripts of which approximately 1,000 are enriched in unc-4::GFP neurons relative to the average nematode embryonic cell. The reliability of these data was validated by the detection of known cell-specific transcripts and by expression in UNC-4 motor neurons of GFP reporters derived from the enriched data set. In addition to genes involved in neurotransmitter packaging and release, the microarray data include transcripts for receptors to a remarkably wide variety of signaling molecules. The added presence of a robust array of G-protein pathway components is indicative of complex and highly integrated mechanisms for modulating motor neuron activity. Over half of the enriched genes (537) have human homologs, a finding that could reflect substantial overlap with the gene expression repertoire of mammalian motor neurons. CONCLUSION: We have described a microarray-based method, MAPCeL, for profiling gene expression in specific C. elegans motor neurons and provide evidence that this approach can reveal candidate genes for key roles in the differentiation and function of these cells. These methods can now be applied to generate a gene expression map of the C. elegans nervous system.

Animals↗

Effect of sampling on topology predictions of protein-protein interaction networks.

Currently available protein-protein interaction (PPI) network or 'interactome' maps, obtained with the yeast two-hybrid (Y2H) assay or by co-affinity purification followed by mass spectrometry (co-AP/MS), only cover a fraction of the complete PPI networks. These partial networks display scale-free topologies--most proteins participate in only a few interactions whereas a few proteins have many interaction partners. Here we analyze whether the scale-free topologies of the partial networks obtained from Y2H assays can be used to accurately infer the topology of complete interactomes. We generated four theoretical interaction networks of different topologies (random, exponential, power law, truncated normal). Partial sampling of these networks resulted in sub-networks with topological characteristics that were virtually indistinguishable from those of currently available Y2H-derived partial interactome maps. We conclude that given the current limited coverage levels, the observed scale-free topology of existing interactome maps cannot be confidently extrapolated to complete interactomes.

Animals↗

New genes with roles in the C. elegans embryo revealed using RNAi of ovary-enriched ORFeome clones.

Several RNA interference (RNAi)-based functional genomic projects have been performed in Caenorhabditis elegans to identify genes required during embryogenesis. These studies have demonstrated that the ovary is enriched for transcripts essential for the first cell divisions. However, comparing RNAi results suggests that many genes involved in embryogenesis have yet to be identified, especially those eliciting partially penetrant phenotypes. To discover additional genes required for C. elegans embryonic development, we tested by RNAi 1123 ORFeome clones selected to represent ovary-enriched genes not associated with an embryonic phenotype. We discovered 155 new ovary-enriched genes with roles during embryogenesis, of which 69% show partial penetrance lethality. Time-lapse microscopy revealed specific phenotypes during early embryogenesis for genes giving rise to high penetrance lethality. Together with previous studies, we now have evidence that 1843 C. elegans genes have roles in embryogenesis, and that many more remain to be found. Using all available RNAi phenotypic data for the ovary-enriched genes, we re-examined the distribution of genes by chromosomal location, functional class, ovary enrichment, and conservation and found that trends are driven almost exclusively by genes eliciting high-penetrance phenotypes. Furthermore, we discovered a striking direct relationship between phylogenetic distribution and the penetrance level of embryonic lethality elicited by RNAi.

Animals↗

Closing in on the C. elegans ORFeome by cloning TWINSCAN predictions.

The genome of Caenorhabditis elegans was the first animal genome to be sequenced. Although considerable effort has been devoted to annotating it, the standard WormBase annotation contains thousands of predicted genes for which there is no cDNA or EST evidence. We hypothesized that a more complete experimental annotation could be obtained by creating a more accurate gene-prediction program and then amplifying and sequencing predicted genes. Our approach was to adapt the TWINSCAN gene prediction system to C. elegans and C. briggsae and to improve its splice site and intron-length models. The resulting system has 60% sensitivity and 58% specificity in exact prediction of open reading frames (ORFs), and hence, proteins-the best results we are aware of any multicellular organism. We then attempted to amplify, clone, and sequence 265 TWINSCAN-predicted ORFs that did not overlap WormBase gene annotations. The success rate was 55%, adding 146 genes that were completely absent from WormBase to the ORF clone collection (ORFeome). The same procedure had a 7% success rate on 90 Worm Base "predicted" genes that do not overlap TWINSCAN predictions. These results indicate that the accuracy of WormBase could be significantly increased by replacing its partially curated predicted genes with TWINSCAN predictions. The technology described in this study will continue to drive the C. elegans ORFeome toward completion and contribute to the annotation of the three Caenorhabditis species currently being sequenced. The results also suggest that this technology can significantly improve our knowledge of the "parts list" for even the best-studied model organisms.

Animals↗

Combining biological networks to predict genetic interactions.

Genetic interactions define overlapping functions and compensatory pathways. In particular, synthetic sick or lethal (SSL) genetic interactions are important for understanding how an organism tolerates random mutation, i.e., genetic robustness. Comprehensive identification of SSL relationships remains far from complete in any organism, because mapping these networks is highly labor intensive. The ability to predict SSL interactions, however, could efficiently guide further SSL discovery. Toward this end, we predicted pairs of SSL genes in Saccharomyces cerevisiae by using probabilistic decision trees to integrate multiple types of data, including localization, mRNA expression, physical interaction, protein function, and characteristics of network topology. Experimental evidence demonstrated the reliability of this strategy, which, when extended to human SSL interactions, may prove valuable in discovering drug targets for cancer therapy and in identifying genes responsible for multigenic diseases.

Animals↗

Evidence for dynamically organized modularity in the yeast protein-protein interaction network.

In apparently scale-free protein-protein interaction networks, or 'interactome' networks, most proteins interact with few partners, whereas a small but significant proportion of proteins, the 'hubs', interact with many partners. Both biological and non-biological scale-free networks are particularly resistant to random node removal but are extremely sensitive to the targeted removal of hubs. A link between the potential scale-free topology of interactome networks and genetic robustness seems to exist, because knockouts of yeast genes encoding hubs are approximately threefold more likely to confer lethality than those of non-hubs. Here we investigate how hubs might contribute to robustness and other cellular properties for protein-protein interactions dynamically regulated both in time and in space. We uncovered two types of hub: 'party' hubs, which interact with most of their partners simultaneously, and 'date' hubs, which bind their different partners at different times or locations. Both in silico studies of network connectivity and genetic interactions described in vivo support a model of organized modularity in which date hubs organize the proteome, connecting biological processes--or modules--to each other, whereas party hubs function inside modules.

Computer Simulation↗

Systematic interactome mapping and genetic perturbation analysis of a C. elegans TGF-beta signaling network.

To initiate a system-level analysis of C. elegans DAF-7/TGF-beta signaling, we combined interactome mapping with single and double genetic perturbations. Yeast two-hybrid (Y2H) screens starting with known DAF-7/TGF-beta pathway components defined a network of 71 interactions among 59 proteins. Coaffinity purification (co-AP) assays in mammalian cells confirmed the overall quality of this network. Systematic perturbations of the network using RNAi, both in wild-type and daf-7/TGF-beta pathway mutant animals, identified nine DAF-7/TGF-beta signaling modifiers, seven of which are conserved in humans. We show that one of these has functional homology to human SNO/SKI oncoproteins and that mutations at the corresponding genetic locus daf-5 confer defects in DAF-7/TGF-beta signaling. Our results reveal substantial molecular complexity in DAF-7/TGF-beta signal transduction. Integrating interactome maps with systematic genetic perturbations may be useful for developing a systems biology approach to this and other signaling modules.

Animals↗