PubMed Health⌕ Search

Biomedical subjects

David E Hill

Publications and source records attributed to David E Hill.

At least 19 recordsLinked to original sources

BOGO: A Proteome-Wide Gene Overexpression Platform for Discovering Rational Cancer Combination Therapies.

Cancer drug resistance remains a major barrier to durable treatment success, often leading to relapse despite advances in precision oncology. While combination therapies are being increasingly investigated, such as chemotherapy with small molecule inhibitors, predicting drug response and identifying rational drug combinations based on resistance mechanisms remain major challenges. Therefore, a proteome-wide, single-gene overexpression screening platform is essential for guiding rational therapy selection. Here, we present BOGO (Bxb1-landing pad human ORFeome-integrated system for a proteome-wide Gene Overexpression), a robust, scalable, and reproducible screening platform that enables single-copy, site-specific integration and overexpression of ~19,000 human open across cancer cell models. Using BOGO, we identified drug-specific response drivers for 16 chemotherapeutic agents and integrated clinical datasets to uncover proliferation and resistance-associated genes with prognostic potential. Drug response similarity networks revealed both shared and unique mechanisms, highlighting key pathways such as autophagy, apoptosis, and Wnt signaling, and notable resistance-associated genes including BCL2, POLD2, and TRADD. In particular, we proposed a synergistic combination of the BCL2 family inhibitor ABT-263 (Navitoclax®) and the DNA analog TAS-102 (Lonsurf®), which revealed that lysosomal modulation is a key mechanism driving DNA analog resistance. This combination therapy selectively enhanced cytotoxicity in colorectal and pancreatic cancer cells in vitro, and demonstrated therapeutic benefit in vivo in both cell line-derived xenograft (CDX) and patient-derived xenograft (PDX) models. Together, these findings establish BOGO as a powerful gene overexpression perturbation platform for systematically identifying chemoresistance and chemosensitization drivers, and for discovering rational combination therapies. Its scalability and reproducibility position BOGO as a broadly applicable tool for functional genomics and therapeutic discovery beyond cancer resistance.

Journal Article↗

hORFeome v3.1: a resource of human open reading frames representing over 10,000 human genes.

Complete sets of cloned protein-encoding open reading frames (ORFs), or ORFeomes, are essential tools for large-scale proteomics and systems biology studies. Here we describe human ORFeome version 3.1 (hORFeome v3.1), currently the largest publicly available resource of full-length human ORFs (available at ). Generated by Gateway recombinational cloning, this collection contains 12,212 ORFs, representing 10,214 human genes, and corresponds to a 51% expansion of the original hORFeome v1.1. An online human ORFeome database, hORFDB, was built and serves as the central repository for all cloned human ORFs (http://horfdb.dfci.harvard.edu). This expansion of the original ORFeome resource greatly increases the potential experimental search space for large-scale proteomics studies, which will lead to the generation of more comprehensive datasets.

Animals↗

A protein-protein interaction network for human inherited ataxias and disorders of Purkinje cell degeneration.

Many human inherited neurodegenerative disorders are characterized by loss of balance due to cerebellar Purkinje cell (PC) degeneration. Although the disease-causing mutations have been identified for a number of these disorders, the normal functions of the proteins involved remain, in many cases, unknown. To gain insight into the function of proteins involved in PC degeneration, we developed an interaction network for 54 proteins involved in 23 inherited ataxias and expanded the network by incorporating literature-curated and evolutionarily conserved interactions. We identified 770 mostly novel protein-protein interactions using a stringent yeast two-hybrid screen; of 75 pairs tested, 83% of the interactions were verified in mammalian cells. Many ataxia-causing proteins share interacting partners, a subset of which have been found to modify neurodegeneration in animal models. This interactome thus provides a tool for understanding pathogenic mechanisms common for this class of neurodegenerative disorders and for identifying candidate genes for inherited ataxias.

Animals↗

From genome to proteome: developing expression clone resources for the human genome.

cDNA clones have long been valuable reagents for studying the structure and function of proteins. With recent access to the entire human genome sequence, it has become possible and highly productive to compare the sequences of mRNAs to their genes, in order to validate the sequences and protein-coding annotations of each (1,2). Thus, well-characterized collections of human cDNAs are now playing an essential role in defining the structure and function of human genes and proteins. In this review, we will summarize the major collections of human cDNA clones, discuss some limitations common to most of these collections and describe several noteworthy proteomics applications, focusing on the detection and analysis of protein-protein interactions (PPI). These human cDNA collections contain principally two types of cDNA clones. The largest collections comprise cDNAs with full-length protein coding sequences (FL-CDS). Some but not all of these cDNA clones may represent the entire mRNA sequence, but many are missing considerable non-coding UTR sequence, usually at the 5' end. A second type of cDNA clone, a 'full-ORF' (F-ORF) expression clone, is one where the annotated protein-coding sequence, excised of 5' UTR and 3' UTR sequence, has been transferred to a vector designed to facilitate transfer to other vectors for protein expression.

Cloning, Molecular↗

Towards a proteome-scale map of the human protein-protein interaction network.

Systematic mapping of protein-protein interactions, or 'interactome' mapping, was initiated in model organisms, starting with defined biological processes and then expanding to the scale of the proteome. Although far from complete, such maps have revealed global topological and dynamic features of interactome networks that relate to known biological properties, suggesting that a human interactome map will provide insight into development and disease mechanisms at a systems level. Here we describe an initial version of a proteome-scale map of human binary protein-protein interactions. Using a stringent, high-throughput yeast two-hybrid system, we tested pairwise interactions among the products of approximately 8,100 currently available Gateway-cloned open reading frames and detected approximately 2,800 interactions. This data set, called CCSB-HI1, has a verification rate of approximately 78% as revealed by an independent co-affinity purification assay, and correlates significantly with other biological attributes. The CCSB-HI1 data set increases by approximately 70% the set of available binary interactions within the tested space and reveals more than 300 new connections to over 100 disease-associated proteins. This work represents an important step towards a systematic and comprehensive human interactome project.

Cloning, Molecular↗

Interactome: gateway into systems biology.

Protein-protein interactions are fundamental to all biological processes, and a comprehensive determination of all protein-protein interactions that can take place in an organism provides a framework for understanding biology as an integrated system. The availability of genome-scale sets of cloned open reading frames has facilitated systematic efforts at creating proteome-scale data sets of protein-protein interactions, which are represented as complex networks or 'interactome' maps. Protein-protein interaction mapping projects that follow stringent criteria, coupled with experimental validation in orthogonal systems, provide high-confidence data sets immanently useful for interrogating developmental and disease mechanisms at a system level as well as elucidating individual protein function and interactome network topology. Although far from complete, currently available maps provide insight into how biochemical properties of proteins and protein complexes are integrated into biological systems. Such maps are also a useful resource to predict the function(s) of thousands of genes.

Animals↗

Pooled ORF expression technology (POET): using proteomics to screen pools of open reading frames for protein expression.

We have developed a pooled ORF expression technology, POET, that uses recombinational cloning and proteomic methods (two-dimensional gel electrophoresis and mass spectrometry) to identify ORFs that when expressed are likely to yield high levels of soluble, purified protein. Because the method works on pools of ORFs, the procedures needed to subclone, express, purify, and assay protein expression for hundreds of clones are greatly simplified. Small scale expression and purification of 12 positive clones identified by POET from a pool of 688 Caenorhabditis elegans ORFs expressed in Escherichia coli yielded on average 6 times as much protein as 12 negative clones. Larger scale expression and purification of six of the positive clones yielded 47-374 mg of purified protein/liter. Using POET, pools of ORFs can be constructed, and the pools of the resulting proteins can be analyzed and manipulated to rapidly acquire information about the attributes of hundreds of proteins simultaneously.

Animals↗

Systematic analysis of genes required for synapse structure and function.

Chemical synapses are complex structures that mediate rapid intercellular signalling in the nervous system. Proteomic studies suggest that several hundred proteins will be found at synaptic specializations. Here we describe a systematic screen to identify genes required for the function or development of Caenorhabditis elegans neuromuscular junctions. A total of 185 genes were identified in an RNA interference screen for decreased acetylcholine secretion; 132 of these genes had not previously been implicated in synaptic transmission. Functional profiles for these genes were determined by comparing secretion defects observed after RNA interference under a variety of conditions. Hierarchical clustering identified groups of functionally related genes, including those involved in the synaptic vesicle cycle, neuropeptide signalling and responsiveness to phorbol esters. Twenty-four genes encoded proteins that were localized to presynaptic specializations. Loss-of-function mutations in 12 genes caused defects in presynaptic structure.

Aldicarb↗

Biochemical clustering of monomeric GTPases of the Ras superfamily.

To date phylogeny has been used to compare entire families of proteins based on their nucleotide or amino acid sequence. Here we developed a novel analytical platform allowing a systematic comparison of protein families based on their biochemical properties. This approach was validated on the Rho subfamily of GTPases. We used two high throughput methods, referred to as AlphaScreen and FlashPlate, to measure nucleotide binding capacity, exchange, and hydrolysis activities of small monomeric GTPases. These two technologies have the characteristics to be very sensitive and to allow homogenous and high throughput assays. To analyze and integrate the data obtained, we developed an algorithm that allows the classification of GTPases according to their enzymatic activities. Integration and hierarchical clustering of these results revealed unexpected features of the small Rho GTPases when compared with primary sequence-based trees. Hence we propose a novel phylobiochemical classification of the Ras superfamily of GTPases.

Algorithms↗

New genes with roles in the C. elegans embryo revealed using RNAi of ovary-enriched ORFeome clones.

Several RNA interference (RNAi)-based functional genomic projects have been performed in Caenorhabditis elegans to identify genes required during embryogenesis. These studies have demonstrated that the ovary is enriched for transcripts essential for the first cell divisions. However, comparing RNAi results suggests that many genes involved in embryogenesis have yet to be identified, especially those eliciting partially penetrant phenotypes. To discover additional genes required for C. elegans embryonic development, we tested by RNAi 1123 ORFeome clones selected to represent ovary-enriched genes not associated with an embryonic phenotype. We discovered 155 new ovary-enriched genes with roles during embryogenesis, of which 69% show partial penetrance lethality. Time-lapse microscopy revealed specific phenotypes during early embryogenesis for genes giving rise to high penetrance lethality. Together with previous studies, we now have evidence that 1843 C. elegans genes have roles in embryogenesis, and that many more remain to be found. Using all available RNAi phenotypic data for the ovary-enriched genes, we re-examined the distribution of genes by chromosomal location, functional class, ovary enrichment, and conservation and found that trends are driven almost exclusively by genes eliciting high-penetrance phenotypes. Furthermore, we discovered a striking direct relationship between phylogenetic distribution and the penetrance level of embryonic lethality elicited by RNAi.

Animals↗

BRCA1/BARD1 orthologs required for DNA repair in Caenorhabditis elegans.

Inherited germline mutations in the tumor suppressor gene BRCA1 predispose individuals to early onset breast and ovarian cancer. BRCA1 together with its structurally related partner BARD1 is required for homologous recombination and DNA double-strand break repair, but how they perform these functions remains elusive. As part of a comprehensive search for DNA repair genes in C. elegans, we identified a BARD1 ortholog. In protein interaction screens, Ce-BRD-1 was found to interact with components of the sumoylation pathway, the TACC domain protein TAC-1, and most importantly, a homolog of mammalian BRCA1. We show that animals depleted for either Ce-brc-1 or Ce-brd-1 display similar abnormalities, including a high incidence of males, elevated levels of p53-dependent germ cell death before and after irradiation, and impaired progeny survival and chromosome fragmentation after irradiation. Furthermore, depletion of ubc-9 and tac-1 leads to radiation sensitivity and a high incidence of males, respectively, potentially linking these genes to the C. elegans BRCA1 pathway. Our findings support a shared role for Ce-BRC-1 and Ce-BRD-1 in C. elegans DNA repair processes, and this role will permit studies of the BRCA1 pathway in an organism amenable to rapid genetic and biochemical analysis.

Amino Acid Sequence↗

A map of the interactome network of the metazoan C. elegans.

To initiate studies on how protein-protein interaction (or "interactome") networks relate to multicellular functions, we have mapped a large fraction of the Caenorhabditis elegans interactome network. Starting with a subset of metazoan-specific proteins, more than 4000 interactions were identified from high-throughput, yeast two-hybrid (HT=Y2H) screens. Independent coaffinity purification assays experimentally validated the overall quality of this Y2H data set. Together with already described Y2H interactions and interologs predicted in silico, the current version of the Worm Interactome (WI5) map contains approximately 5500 interactions. Topological and biological features of this interactome network, as well as its integration with phenome and transcriptome data sets, lead to numerous biological hypotheses.

Animals↗

ORFeome projects: gateway between genomics and omics.

The availability of entire genome sequences is expected to revolutionize the way in which biology and medicine are conducted for years to come. However, achieving this promise still requires significant effort in the areas of gene annotation, cloning and expression of thousands of known and heretofore unknown protein-encoding genes. Traditional technologies of manipulating genes are too cumbersome and inefficient when one is dealing with more than a few genes at a time. Entire libraries composed of all protein-encoding open reading frames (ORFs) cloned in highly flexible vectors will be needed to take full advantage of the information found in any genome sequence. The creation of such ORFeome resources using novel technologies for cloning and expressing entire proteomes constitutes an effective gateway from whole genome sequencing efforts to downstream 'omics' applications.

Animals↗

Generation of the Brucella melitensis ORFeome version 1.1.

The bacteria of the Brucella genus are responsible for a worldwide zoonosis called brucellosis. They belong to the alpha-proteobacteria group, as many other bacteria that live in close association with a eukaryotic host. Importantly, the Brucellae are mainly intracellular pathogens, and the molecular mechanisms of their virulence are still poorly understood. Using the complete genome sequence of Brucella melitensis, we generated a database of protein-coding open reading frames (ORFs) and constructed an ORFeome library of 3091 Gateway Entry clones, each containing a defined ORF. This first version of the Brucella ORFeome (v1.1) provides the coding sequences in a user-friendly format amenable to high-throughput functional genomic and proteomic experiments, as the ORFs are conveniently transferable from the Entry clones to various Expression vectors by recombinational cloning. The cloning of the Brucella ORFeome v1.1 should help to provide a better understanding of the molecular mechanisms of virulence, including the identification of bacterial protein-protein interactions, but also interactions between bacterial effectors and their host's targets.

Bacterial Proteins↗

C. elegans ORFeome version 3.1: increasing the coverage of ORFeome resources with improved gene predictions.

The first version of the Caenorhabditis elegans ORFeome cloning project, based on release WS9 of Wormbase (August 1999), provided experimental verifications for approximately 55% of predicted protein-encoding open reading frames (ORFs). The remaining 45% of predicted ORFs could not be cloned, possibly as a result of mispredicted gene boundaries. Since the release of WS9, gene predictions have improved continuously. To test the accuracy of evolving predictions, we attempted to PCR-amplify from a highly representative worm cDNA library and Gateway-clone approximately 4200 ORFs missed earlier and for which new predictions are available in WS100 (May 2003). In this set we successfully cloned 63% of ORFs with supporting experimental data ("touched" ORFs), and 42% of ORFs with no supporting experimental evidence ("untouched" ORFs). Approximately 2000 full-length ORFs were cloned in-frame, 13% of which were corrected in their exon/intron structure relative to WS100 predictions. In total, approximately 12,500 C. elegans ORFs are now available as Gateway Entry clones for various reverse proteomics (ORFeome v3.1). This work illustrates why the cloning of a complete C. elegans ORFeome, and likely the ORFeomes of other multicellular organisms, needs to be an iterative process that requires multiple rounds of experimental validation together with gradually improving gene predictions.

Animals↗

A first version of the Caenorhabditis elegans Promoterome.

An important aspect of the development of systems biology approaches in metazoans is the characterization of expression patterns of nearly all genes predicted from genome sequences. Such "localizome" maps should provide information on where (in what cells or tissues) and when (at what stage of development or under what conditions) genes are expressed. They should also indicate in what cellular compartments the corresponding proteins are localized. Caenorhabditis elegans is particularly suited for the development of a localizome map since all its 959 adult somatic cells can be visualized by microscopy, and its cell lineage has been completely described. Here we address one of the challenges of C. elegans localizome mapping projects: that of obtaining a genome-wide resource of C. elegans promoters needed to generate transgenic animals expressing localization markers such as the green fluorescent protein (GFP). To ensure high flexibility for future uses, we utilized the newly developed MultiSite Gateway system. We generated and validated "version 1.1" of the Promoterome: a resource of approximately 6000 C. elegans promoters. These promoters can be transferred easily into various Gateway Destination vectors to drive expression of markers such as GFP, alone (promoter::GFP constructs), or in fusion with protein-encoding open reading frames available in ORFeome resources (promoter::ORF::GFP).

Animals↗

Toward improving Caenorhabditis elegans phenome mapping with an ORFeome-based RNAi library.

The recently completed Caenorhabditis elegans genome sequence allows application of high-throughput (HT) approaches for phenotypic analyses using RNA interference (RNAi). As large phenotypic data sets become available, "phenoclustering" strategies can be used to begin understanding the complex molecular networks involved in development and other biological processes. The current HT-RNAi resources represent a great asset for phenotypic profiling but are limited by lack of flexibility. For instance, existing resources do not take advantage of the latest improvements in RNAi technology, such as inducible hairpin RNAi. Here we show that a C. elegans ORFeome resource, generated with the Gateway cloning system, can be used as a starting point to generate alternative HT-RNAi resources with enhanced flexibility. The versatility inherent to the Gateway system suggests that additional HT-RNAi libraries can now be readily generated to perform gene knockdowns under various conditions, increasing the possibilities for phenome mapping in C. elegans.

Animals↗

High-throughput expression of C. elegans proteins.

Proteome-scale studies of protein three-dimensional structures should provide valuable information for both investigating basic biology and developing therapeutics. Critical for these endeavors is the expression of recombinant proteins. We selected Caenorhabditis elegans as our model organism in a structural proteomics initiative because of the high quality of its genome sequence and the availability of its ORFeome, protein-encoding open reading frames (ORFs), in a flexible recombinational cloning format. We developed a robotic pipeline for recombinant protein expression, applying the Gateway cloning/expression technology and utilizing a stepwise automation strategy on an integrated robotic platform. Using the pipeline, we have carried out heterologous protein expression experiments on 10,167 ORFs of C. elegans. With one expression vector and one Escherichia coli strain, protein expression was observed for 4854 ORFs, and 1536 were soluble. Bioinformatics analysis of the data indicates that protein hydrophobicity is a key determining factor for an ORF to yield a soluble expression product. This protein expression effort has investigated the largest number of genes in any organism to date. The pipeline described here is applicable to high-throughput expression of recombinant proteins for other species, both prokaryotic and eukaryotic, provided that ORFeome resources become available.

Animals↗