PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Cell type annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

The Eukaryotic Promoter Database EPD: the impact of in silico primer extension.

The Eukaryotic Promoter Database (EPD) is an annotated non-redundant collection of eukaryotic POL II promoters, experimentally defined by a transcription start site (TSS). There may be multiple promoter entries for a single gene. The underlying experimental evidence comes from journal articles and, starting from release 73, from 5' ESTs of full-length cDNA clones used for so-called in silico primer extension. Access to promoter sequences is provided by pointers to TSS positions in nucleotide sequence entries. The annotation part of an EPD entry includes a description of the type and source of the initiation site mapping data, links to other biological databases and bibliographic references. EPD is structured in a way that facilitates dynamic extraction of biologically meaningful promoter subsets for comparative sequence analysis. Web-based interfaces have been developed that enable the user to view EPD entries in different formats, to select and extract promoter sequences according to a variety of criteria and to navigate to related databases exploiting different cross-references. Tools for analysing sequence motifs around TSSs defined in EPD are provided by the signal search analysis server. EPD can be accessed at http://www.epd. isb-sib.ch.

Animals↗

Gene discovery and annotation using LCM-454 transcriptome sequencing.

454 DNA sequencing technology achieves significant throughput relative to traditional approaches. More than 261,000 ESTs were generated by 454 Life Sciences from cDNA isolated using laser capture microdissection (LCM) from the developmentally important shoot apical meristem (SAM) of maize (Zea mays L.). This single sequencing run annotated >25,000 maize genomic sequences and also captured approximately 400 expressed transcripts for which homologous sequences have not yet been identified in other species. Approximately 70% of the ESTs generated in this study had not been captured during a previous EST project conducted using a cDNA library constructed from hand-dissected apex tissue that is highly enriched for SAMs. In addition, at least 30% of the 454-ESTs do not align to any of the approximately 648,000 extant maize ESTs using conservative alignment criteria. These results indicate that the combination of LCM and the deep sequencing possible with 454 technology enriches for SAM transcripts not present in current EST collections. RT-PCR was used to validate the expression of 27 genes whose expression had been detected in the SAM via LCM-454 technology, but that lacked orthologs in GenBank. Significantly, transcripts from approximately 74% (20/27) of these validated SAM-expressed "orphans" were not detected in meristem-rich immature ears. We conclude that the coupling of LCM and 454 sequencing technologies facilitates the discovery of rare, possibly cell-type-specific transcripts.

Base Sequence↗

A sodium bicarbonate transporter from sea urchin spermatozoa.

Bicarbonate (HCO3-) transporters play crucial roles in cell-signaling pathways and are essential for cell viability. Here we describe the first cloning and localization of a HCO3- transporter from sperm of the sea urchin, Strongylocentrotus purpuratus. The deduced protein is 1214 amino acids and has a calculated molecular mass of 135 kDa. The annotated protein coding region of the transporter gene consists of 24 exons. The most similar human protein is the Na+/HCO3- cotransporter-2 (NBC2), which has 53% identity and 68% similarity to the sea urchin protein. The sea urchin protein shares the major structural features of HCO3- transporters, including 13 transmembrane segments, a DIDS (4,4-diiodothiocyanatostilbene-2, 2-disulfonic acid) binding motif and N-linked glycosylation sites. It has longer N- and C-terminal cytoplasmic domains compared to human HCO3- transporters. The sea urchin protein possesses a relatively long 3rd extracellular loop with four conserved cysteine residues. This is characteristic for Na+/HCO3- cotransporters, but not for anion exchangers, suggesting that the sea urchin protein is a Na+/HCO3- cotransporter. It is therefore designated as Sp-NBC. A neighbor-joining tree shows that Sp-NBC branches closer to the electroneutral type of HCO3- transporters. Western immunoblots and immunoflourescence show that Sp-NBC is concentrated in the flagellar plasma membrane, suggesting a role in motility regulation.

Amino Acid Sequence↗

Towards precise classification of cancers based on robust gene functional expression profiles.

BACKGROUND: Development of robust and efficient methods for analyzing and interpreting high dimension gene expression profiles continues to be a focus in computational biology. The accumulated experiment evidence supports the assumption that genes express and perform their functions in modular fashions in cells. Therefore, there is an open space for development of the timely and relevant computational algorithms that use robust functional expression profiles towards precise classification of complex human diseases at the modular level. RESULTS: Inspired by the insight that genes act as a module to carry out a highly integrated cellular function, we thus define a low dimension functional expression profile for data reduction. After annotating each individual gene to functional categories defined in a proper gene function classification system such as Gene Ontology applied in this study, we identify those functional categories enriched with differentially expressed genes. For each functional category or functional module, we compute a summary measure (s) for the raw expression values of the annotated genes to capture the overall activity level of the module. In this way, we can treat the gene expressions within a functional module as an integrative data point to replace the multiple values of individual genes. We compare the classification performance of decision trees based on functional expression profiles with the conventional gene expression profiles using four publicly available datasets, which indicates that precise classification of tumour types and improved interpretation can be achieved with the reduced functional expression profiles. CONCLUSION: This modular approach is demonstrated to be a powerful alternative approach to analyzing high dimension microarray data and is robust to high measurement noise and intrinsic biological variance inherent in microarray data. Furthermore, efficient integration with current biological knowledge has facilitated the interpretation of the underlying molecular mechanisms for complex human diseases at the modular level.

Algorithms↗

The Helicobacter pylori chemotaxis receptor TlpB (HP0103) is required for pH taxis and for colonization of the gastric mucosa.

The location of Helicobacter pylori in the gastric mucosa of mammals is defined by natural pH gradients within the gastric mucus, which are more alkaline proximal to the mucosal epithelial cells and more acidic toward the lumen. We have used a microscope slide-based pH gradient assay and video data collection system to document pH-tactic behavior. In response to hydrochloric acid (HCl), H. pylori changes its swimming pattern from straight-line random swimming to arcing or circular patterns that move the motile population away from the strong acid. Bacteria in more-alkaline regions did not swim toward the acid, suggesting the pH taxis is a form of negative chemotaxis. To identify the chemoreceptor(s) responsible for the transduction of pH-tactic signals, a vector-free allelic replacement strategy was used to construct mutations in each of the four annotated chemoreceptor genes (tlpA, tlpB, tlpC, and tlpD) in H. pylori strain SS1 and a motile variant of strain KE26695. All deletion mutants were motile and displayed normal chemotaxis in brucella soft agar, but only tlpB mutants were defective for pH taxis. tlpD mutants exhibited more tumbling and arcing swimming, while tlpC mutants were hypermotile and responsive to acid. While tlpA, tlpC, and tlpD mutants colonized mice to near wild-type levels, tlpB mutants were defective for colonization of highly permissive C57BL/6 interleukin-12 (IL-12) (p40-/-)-deficient mice. Complementation of the tlpB mutant (tlpB expressed from the rdxA locus) restored pH taxis and infectivity for mice. pH taxis, like motility and urease activity, is essential for colonization and persistence in the gastric mucosa, and thus TlpB function might represent a novel target in the development of therapeutics that blind tactic behavior.

Animals↗

The Arabidopsis phenylalanine ammonia lyase gene family: kinetic characterization of the four PAL isoforms.

In Arabidopsis thaliana, four genes have been annotated as provisionally encoding PAL. In this study, recombinant native AtPAL1, 2, and 4 were demonstrated to be catalytically competent for l-phenylalanine deamination, whereas AtPAL3, obtained as a N-terminal His-tagged protein, was of very low activity and only detectable at high substrate concentrations. All four PALs displayed similar pH optima, but not temperature optima; AtPAL3 had a lower temperature optimum than the other three isoforms. AtPAL1, 2 and 4 had similar K(m) values (64-71 microM) for l-Phe, with AtPAL2 apparently being slightly more catalytically efficacious due to decreased K(m) and higher k(cat) values, relative to the others. As anticipated, PAL activities with l-tyrosine were either low (AtPAL1, 2, and 4) or undetectable (AtPAL3), thereby establishing that l-Phe is the true physiological substrate. This detailed knowledge of the kinetic and functional properties of the various PAL isoforms now provides the necessary biochemical foundation required for the systematic investigation and dissection of the organization of the PAL metabolic network/gene circuitry involved in numerous aspects of phenylpropanoid metabolism in A. thaliana spanning various cell types, tissues and organs.

Arabidopsis↗

Automatic generation of gene finders for eukaryotic species.

BACKGROUND: The number of sequenced eukaryotic genomes is rapidly increasing. This means that over time it will be hard to keep supplying customised gene finders for each genome. This calls for procedures to automatically generate species-specific gene finders and to re-train them as the quantity and quality of reliable gene annotation grows. RESULTS: We present a procedure, Agene, that automatically generates a species-specific gene predictor from a set of reliable mRNA sequences and a genome. We apply a Hidden Markov model (HMM) that implements explicit length distribution modelling for all gene structure blocks using acyclic discrete phase type distributions. The state structure of the each HMM is generated dynamically from an array of sub-models to include only gene features represented in the training set. CONCLUSION: Acyclic discrete phase type distributions are well suited to model sequence length distributions. The performance of each individual gene predictor on each individual genome is comparable to the best of the manually optimised species-specific gene finders. It is shown that species-specific gene finders are superior to gene finders trained on other species.

Algorithms↗

A bile salt hydrolase of Brucella abortus contributes to the establishment of a successful infection through the oral route in mice.

Choloylglycine hydrolase (CGH), a bile salt hydrolase, has been annotated in all the available genomes of Brucella species. We obtained the Brucella CGH in recombinant form and demonstrated in vitro its capacity to cleave glycocholate into glycine and cholate. Brucella abortus 2308 (wild type) and its isogenic Deltacgh deletion mutant exhibited similar growth rates in tryptic soy broth in the absence of bile. In contrast, the growth of the Deltacgh mutant was notably impaired by both 5% and 10% bile. The bile resistance of the complemented mutant was similar to that of the wild-type strain. In mice infected through the intragastric or the intraperitoneal route, splenic infection was significantly lower at 10 and 20 days postinfection in animals infected with the Deltacgh mutant than in those infected with the wild-type strain. For both routes, no differences in spleen CFU were found between animals infected with the wild-type strain and those infected with the complemented mutant. Mice immunized intragastrically with recombinant CGH mixed with cholera toxin (CGH+CT) developed a specific mucosal humoral (immunoglobulin G [IgG] and IgA) and cellular (interleukin-2) immune responses. Fifteen days after challenge by the same route with live B. abortus 2308 cells, splenic CFU counts were 10-fold lower in mice immunized with CGH+CT than in mice immunized with CT or phosphate-buffered saline. This study shows that CGH confers on Brucella the ability to resist the antimicrobial action of bile salts. The results also suggest that CGH may contribute to the ability of Brucella to infect the host through the oral route.

Amidohydrolases↗

Functional annotation of mouse mutations in embryonic stem cells by use of expression profiling.

Expression profiling offers a potential high-throughput phenotype screen for mutant mouse embryonic stem (ES) cells. We have assessed the ability of expression arrays to distinguish among heterozygous mutant ES cell lines and to accurately reflect the normal function of the mutated genes. Two ES cell lines hemizygous for overlapping regions of mouse Chromosome (Chr) 5 differed substantially from the wildtype parental line and from each other. Expression differences included frequent downregulation of hemizygous genes and downstream effects on genes mapping to other chromosomes. Some genes were affected similarly in each deletion line, consistent with the overlap of the deletions. To determine whether such downstream effects reveal pathways impacted by a mutation, we examined ES cell lines heterozygous for mutations in either of two well-characterized genes. A heterozygous mutation in the gene encoding the cell cycle regulator, cyclin D kinase 4 ( Cdk4), affected expression of many genes involved in cell growth and proliferation. A heterozygous mutation in the ATP binding cassette transporter family A, member 1 ( Abca1) gene, altered genes associated with lipid homeostasis, the cytoskeleton, and vesicle trafficking. Heterozygous Abca1 mutation had similar effects in liver, indicating that ES cell expression profile reflects changes in fundamental processes relevant to mutant gene function in multiple cell types.

ATP Binding Cassette Transporter 1↗

The Drosophila melanogaster LEM-domain protein MAN1.

Here we describe the Drosophila melanogaster LEM-domain protein encoded by the annotated gene CG3167 which is the putative ortholog to vertebrate MAN1. MAN1 of Drosophila (dMAN1) and vertebrates have the following properties in common. Firstly, both molecules are integral membrane proteins of the inner nuclear membrane (INM) and share the same structural organization comprising an N-terminally located LEM motif, two transmembrane domains in the middle of the molecule, and a conserved RNA recognition motif in the C-terminal region. Secondly, dMAN1 has similar targeting domains as it has been reported for the human protein. Thirdly, immunoprecipitations with dMAN1-specific antibodies revealed that this Drosophila LEM-domain protein is contained in protein complexes together with lamins Dm0 and C. It has been previously shown that human MAN1 binds to A- and B-type lamins in vitro. During embryogenesis and early larval development LEM-domain proteins dMAN1 and otefin show the same expression pattern and are much more abundant in eggs and the first larval instar than in later larval stages and young pupae whereas the LEM-domain protein Bocksbeutel is uniformly expressed in all developmental stages. dMAN1 is detectable in the nuclear envelope of embryonic cells including the pole cells. In mitotic cells of embryos at metaphase and anaphase, LEM-domain proteins dMAN1, otefin and Bocksbeutel were predominantly localized in the region of the two spindle poles whereas the lamin B receptor and lamin Dm0 were more homogeneously distributed. Downregulation of dMAN1 by RNA interference (RNAi) in Drosophila cultured Kc167 cells has no obvious effect on nuclear architecture, viability of RNAi-treated cells and the intracellular distribution of the LEM-domain proteins Bocksbeutel and otefin. In contrast, the localization of dMAN1, Bocksbeutel and otefin at the INM is supported by lamin Dm0. We conclude that the dMAN1 protein is not a limiting component of the nuclear architecture in Drosophila cultured cells.

Amino Acid Motifs↗

Modeling Lactococcus lactis using a genome-scale flux model.

BACKGROUND: Genome-scale flux models are useful tools to represent and analyze microbial metabolism. In this work we reconstructed the metabolic network of the lactic acid bacteria Lactococcus lactis and developed a genome-scale flux model able to simulate and analyze network capabilities and whole-cell function under aerobic and anaerobic continuous cultures. Flux balance analysis (FBA) and minimization of metabolic adjustment (MOMA) were used as modeling frameworks. RESULTS: The metabolic network was reconstructed using the annotated genome sequence from L. lactis ssp. lactis IL1403 together with physiological and biochemical information. The established network comprised a total of 621 reactions and 509 metabolites, representing the overall metabolism of L. lactis. Experimental data reported in the literature was used to fit the model to phenotypic observations. Regulatory constraints had to be included to simulate certain metabolic features, such as the shift from homo to heterolactic fermentation. A minimal medium for in silico growth was identified, indicating the requirement of four amino acids in addition to a sugar. Remarkably, de novo biosynthesis of four other amino acids was observed even when all amino acids were supplied, which is in good agreement with experimental observations. Additionally, enhanced metabolic engineering strategies for improved diacetyl producing strains were designed. CONCLUSION: The L. lactis metabolic network can now be used for a better understanding of lactococcal metabolic capabilities and potential, for the design of enhanced metabolic engineering strategies and for integration with other types of 'omic' data, to assist in finding new information on cellular organization and function.

Bacteria, Anaerobic↗

Integrative analysis of cancer-related data using CAP.

The development of human cancer is a highly complex process and can be considered the result of several combined events, such as genetic alterations, disturbance of signal transduction, or failure of immunological surveillance. Cancer-related databases usually focus on specific fields of research, e.g., cancer genetics or cancer immunology, whereas the complexity of cancer genesis requires an integrated analysis of heterogeneous data from several sources. Here we present the cancer-associated protein database (CAP), a novel analysis system for cancer-related data. CAP integrates data from multiple external databases, augments these data with functional annotations, and offers tools for statistical analysis of these data. We have employed CAP to analyze genes that have been found to cause an autoimmune response in cancer. In particular, we explored the connection between the autoimmune response, mutations, and overexpression of these genes. Our preliminary results suggest that mutations are not significant contributors to raising an antibody response against tumor antigens, whereas overexpression seems to play a more important role. We hereby demonstrate how different types of data can be integrated and analyzed successfully, providing interesting results. As the amount of available data is growing rapidly, a combined analysis will play an important role in exploring the genetic and immunological basis of cancer. CAP is freely available at the following web site: http://www.bioinf.uni-sb.de/CAP/.

Autoimmunity↗

Meta-analysis of microarray data on pancreatic cancer defines a set of commonly dysregulated genes.

Pancreatic ductal adenocarcinoma is the eighth most common cancer with the lowest overall 5-year relative survival rate of any tumor type today. Expression profiling using microarrays has been widely used to identify genes associated with pancreatic cancer development. To extract maximum value from the available gene expression data, we applied a meta-analysis to search for commonly differentially expressed genes in pancreatic ductal adenocarcinoma. We obtained data sets from four different gene expression studies on pancreatic cancer. We selected a consensus set of 2984 genes measured in all four studies and applied a meta-analysis approach to evaluate the combined data. Of the genes identified as differentially expressed, several were validated using RT-PCR and immunohistochemistry. Additionally, we used a class discovery algorithm to identify a gene expression signature. Our meta-analysis revealed that the pancreatic cancer gene expression data sets shared a significant number of up- and downregulated genes, independent of the technology used. This interstudy crossvalidation approach generated a set of 568 genes that were consistently and significantly dysregulated in pancreatic cancer. Of these, 364 (64.1%) were upregulated and 204 (35.9%) were downregulated in pancreatic cancer. Only 127 (22%) were described in the published individual analyses. Functional annotation of the genes revealed that genes presumably associated with the cell adhesion-mediated drug resistance pathway are frequently overexpressed in pancreatic cancer. Meta-analysis is an important tool for the identification and validation of differentially expressed genes. These could represent good candidates for novel diagnostic and therapeutic approaches to pancreatic cancer.

Adenocarcinoma↗

Transcriptomic footprints disclose specificity of reactive oxygen species signaling in Arabidopsis.

Reactive oxygen species (ROS) are key players in the regulation of plant development, stress responses, and programmed cell death. Previous studies indicated that depending on the type of ROS (hydrogen peroxide, superoxide, or singlet oxygen) or its subcellular production site (plastidic, cytosolic, peroxisomal, or apoplastic), a different physiological, biochemical, and molecular response is provoked. We used transcriptome data generated from ROS-related microarray experiments to assess the specificity of ROS-driven transcript expression. Data sets obtained by exogenous application of oxidative stress-causing agents (methyl viologen, Alternaria alternata toxin, 3-aminotriazole, and ozone) and from a mutant (fluorescent) and transgenic plants, in which the activity of an individual antioxidant enzyme was perturbed (catalase, cytosolic ascorbate peroxidase, and copper/zinc superoxide dismutase), were compared. In total, the abundance of nearly 26,000 transcripts of Arabidopsis (Arabidopsis thaliana) was monitored in response to different ROS. Overall, 8,056, 5,312, and 3,925 transcripts showed at least a 3-, 4-, or 5-fold change in expression, respectively. In addition to marker transcripts that were specifically regulated by hydrogen peroxide, superoxide, or singlet oxygen, several transcripts were identified as general oxidative stress response markers because their steady-state levels were at least 5-fold elevated in most experiments. We also assessed the expression characteristics of all annotated transcription factors and inferred new candidate regulatory transcripts that could be responsible for orchestrating the specific transcriptomic signatures triggered by different ROS. Our analysis provides a framework that will assist future efforts to address the impact of ROS signals within environmental stress conditions and elucidate the molecular mechanisms of the oxidative stress response in plants.

Arabidopsis↗

Influence of cyclical mechanical strain on extracellular matrix gene expression in human lamina cribrosa cells in vitro.

PURPOSE: The mechanical effect of raised intraocular pressure is a recognised stimulus for optic neuropathy in primary open angle glaucoma (POAG). Characteristic extracellular matrix (ECM) remodelling accompanies axonal damage in the lamina cribrosa (LC) of the optic nerve head in POAG. Glial cells in the lamina cribrosa may play a role in this process but the precise cellular responses to mechanical forces in this region are unknown. The authors examined global gene expression profiles in lamina cribrosa cells exposed to cyclical mechanical stretch, with an emphasis on ECM genes. METHODS: Glial fibrillary acid protein negative primary LC cells were generated from the optic nerve head tissue of three normal human donors. Confluent cell passages (n=4) were exposed to 15% stretch at 1 Hz or static conditions for 24 h using the Flexercell system. Gene expression was assessed using Affymetrix U133A microarrays with pooled RNA. Expression levels were normalized using robust multi-chip average (RMA). Expression data was annotated using NIH DAVID software. ECM-related gene expression was validated in an independent experiment using quantitative real-time PCR and protein synthesis was measured using ELISA and immunohistochemistry. RESULTS: Compared with static controls, 805 genes were upregulated and 644 were downregulated by +/-1.5 fold in stretched LC cells. Gene ontologies included ECM, cell proliferation, growth factor activity, and signal transduction. Differentially expressed ECM genes included elastin, collagens (IV, VI, VIII, IX), thrombospondin 1, perlecan, and lysl oxidase. Quantitative PCR demonstrated that the expression of TGF-beta2, BMP-7, elastin, collagen VI, biglycan, versican, EMMPRIN, VEGF, and thrombomodulin were reproducible and consistent with the microarray data. VEGF and TGF-beta2 protein levels were also significantly (p<0.05) increased in stretched cell media supernatants. Immunohistochemistry demonstrated increased EMMPRIN (an extracellular matrix metalloproteinase inducer) protein in human POAG optic nerve head tissue compared to nonglaucomatous controls. CONCLUSIONS: These findings demonstrate that LC cells respond to mechanical stimuli in vitro by transcription of several components and modulators of the ECM. Some of the upregulated ECM genes identified are novel in the context of glaucomatous optic neuropathy (biglycan, versican, EMMPRIN, and BMP-7). The LC cell may represent both an important pro-fibrotic cell type in the optic nerve head and an attractive target for novel therapeutic intervention in POAG.

Cells, Cultured↗

Real-time RT-PCR profiling of over 1400 Arabidopsis transcription factors: unprecedented sensitivity reveals novel root- and shoot-specific genes.

Summary To overcome the detection limits inherent to DNA array-based methods of transcriptome analysis, we developed a real-time reverse transcription (RT)-PCR-based resource for quantitative measurement of transcripts for 1465 Arabidopsis transcription factors (TFs). Using closely spaced gene-specific primer pairs and SYBR Green to monitor amplification of double-stranded DNA (dsDNA), transcript levels of 83% of all target genes could be measured in roots or shoots of young Arabidopsis wild-type plants. Only 4% of reactions produced non-specific PCR products. The amplification efficiency of each PCR was determined from the log slope of SYBR Green fluorescence versus cycle number in the exponential phase, and was used to correct the readout for each primer pair and run. Measurements of transcript abundance were quantitative over six orders of magnitude, with a detection limit equivalent to one transcript molecule in 1000 cells. Transcript levels for different TF genes ranged between 0.001 and 100 copies per cell. Only 13% of TF transcripts were undetectable in these organs. For comparison, 22K Arabidopsis Affymetrix chips detected less than 55% of TF transcripts in the same samples, the range of transcript levels was compressed by a factor more than 100, and the data were less accurate especially in the lower part of the response range. Real-time RT-PCR revealed 35 root-specific and 52 shoot-specific TF genes, most of which have not been identified as organ-specific previously. Finally, many of the TF transcripts detected by RT-PCR are not represented in Arabidopsis EST (expressed sequence tag) or Massively Parallel Signature Sequencing (MPSS) databases. These genes can now be annotated as expressed.

Arabidopsis↗

The Eukaryotic Promoter Database, EPD: new entry types and links to gene expression data.

The Eukaryotic Promoter Database (EPD) is an annotated, non-redundant collection of eukaryotic Pol II promoters, for which the transcription start site has been determined experimentally. Access to promoter sequences is provided by pointers to positions in nucleotide sequence entries. The annotation part of an entry includes a description of the initiation site mapping data, exhaustive cross-references to the EMBL nucleotide sequence database, SWISS-PROT, TRANSFAC and other databases, as well as bibliographic references. EPD is structured in a way that facilitates dynamic extraction of biologically meaningful promoter subsets for comparative sequence analysis. World Wide Web-based interfaces have been developed which enable the user to view EPD entries in different formats, to select and extract promoter sequences according to a variety of criteria, and to navigate to related databases exploiting different cross-references. The EPD web site also features yearly updated base frequency matrices for major eukaryotic promoter elements. EPD can be accessed at http://www.epd.isb-sib.ch.

Animals↗

A wiring of the human nucleolus.

Recent proteomic efforts have created an extensive inventory of the human nucleolar proteome. However, approximately 30% of the identified proteins lack functional annotation. We present an approach of assigning function to uncharacterized nucleolar proteins by data integration coupled to a machine-learning method. By assembling protein complexes, we present a first draft of the human ribosome biogenesis pathway encompassing 74 proteins and hereby assign function to 49 previously uncharacterized proteins. Moreover, the functional diversity of the nucleolus is underlined by the identification of a number of protein complexes with functions beyond ribosome biogenesis. Finally, we were able to obtain experimental evidence of nucleolar localization of 11 proteins, which were predicted by our platform to be associates of nucleolar complexes. We believe other biological organelles or systems could be "wired" in a similar fashion, integrating different types of data with high-throughput proteomics, followed by a detailed biological analysis and experimental validation.

Artificial Intelligence↗