PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Cell type annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Protein complexes and functional modules in molecular networks.

Proteins, nucleic acids, and small molecules form a dense network of molecular interactions in a cell. Molecules are nodes of this network, and the interactions between them are edges. The architecture of molecular networks can reveal important principles of cellular organization and function, similarly to the way that protein structure tells us about the function and organization of a protein. Computational analysis of molecular networks has been primarily concerned with node degree [Wagner, A. & Fell, D. A. (2001) Proc. R. Soc. London Ser. B 268, 1803-1810; Jeong, H., Tombor, B., Albert, R., Oltvai, Z. N. & Barabasi, A. L. (2000) Nature 407, 651-654] or degree correlation [Maslov, S. & Sneppen, K. (2002) Science 296, 910-913], and hence focused on single/two-body properties of these networks. Here, by analyzing the multibody structure of the network of protein-protein interactions, we discovered molecular modules that are densely connected within themselves but sparsely connected with the rest of the network. Comparison with experimental data and functional annotation of genes showed two types of modules: (i) protein complexes (splicing machinery, transcription factors, etc.) and (ii) dynamic functional units (signaling cascades, cell-cycle regulation, etc.). Discovered modules are highly statistically significant, as is evident from comparison with random graphs, and are robust to noise in the data. Our results provide strong support for the network modularity principle introduced by Hartwell et al. [Hartwell, L. H., Hopfield, J. J., Leibler, S. & Murray, A. W. (1999) Nature 402, C47-C52], suggesting that found modules constitute the "building blocks" of molecular networks.

Biophysical Phenomena↗

Integrative bioinformatics for functional genome annotation: trawling for G protein-coupled receptors.

G protein-coupled receptors (GPCR) are amongst the best studied and most functionally diverse types of cell-surface protein. The importance of GPCRs as mediates or cell function and organismal developmental underlies their involvement in key physiological roles and their prominence as targets for pharmacological therapeutics. In this review, we highlight the requirement for integrated protocols which underline the different perspectives offered by different sequence analysis methods. BLAST and FastA offer broad brush strokes. Motif-based search methods add the fine detail. Structural modelling offers another perspective which allows us to elucidate the physicochemical properties that underlie ligand binding. Together, these different views provide a more informative and a more detailed picture of GPCR structure and function. Many GPCRs remain orphan receptors with no identified ligand, yet as computer-driven functional genomics starts to elaborate their functions, a new understanding of their roles in cell and developmental biology will follow.

Algorithms↗

Automating genomic data mining via a sequence-based matrix format and associative rule set.

There is an enormous amount of information encoded in each genome--enough to create living, responsive and adaptive organisms. Raw sequence data alone is not enough to understand function, mechanisms or interactions. Changes in a single base pair can lead to disease, such as sickle-cell anemia, while some large megabase deletions have no apparent phenotypic effect. Genomic features are varied in their data types and annotation of these features is spread across multiple databases. Herein, we develop a method to automate exploration of genomes by iteratively exploring sequence data for correlations and building upon them. First, to integrate and compare different annotation sources, a sequence matrix (SM) is developed to contain position-dependant information. Second, a classification tree is developed for matrix row types, specifying how each data type is to be treated with respect to other data types for analysis purposes. Third, correlative analyses are developed to analyze features of each matrix row in terms of the other rows, guided by the classification tree as to which analyses are appropriate. A prototype was developed and successful in detecting coinciding genomic features among genes, exons, repetitive elements and CpG islands.

Base Sequence↗

Comparative profiling of the sense and antisense transcriptome of maize lines.

BACKGROUND: There are thousands of maize lines with distinctive normal as well as mutant phenotypes. To determine the validity of comparisons among mutants in different lines, we first address the question of how similar the transcriptomes are in three standard lines at four developmental stages. RESULTS: Four tissues (leaves, 1 mm anthers, 1.5 mm anthers, pollen) from one hybrid and one inbred maize line were hybridized with the W23 inbred on Agilent oligonucleotide microarrays with 21,000 elements. Tissue-specific gene expression patterns were documented, with leaves having the most tissue-specific transcripts. Haploid pollen expresses about half as many genes as the other samples. High overlap of gene expression was found between leaves and anthers. Anther and pollen transcript expression showed high conservation among the three lines while leaves had more divergence. Antisense transcripts represented about 6 to 14 percent of total transcriptome by tissue type but were similar across lines. Gene Ontology (GO) annotations were assigned and tabulated. Enrichment in GO terms related to cell-cycle functions was found for the identified antisense transcripts. Microarray results were validated via quantitative real-time PCR and by hybridization to a second oligonucleotide microarray platform. CONCLUSION: Despite high polymorphisms and structural differences among maize inbred lines, the transcriptomes of the three lines displayed remarkable similarities, especially in both reproductive samples (anther and pollen). We also identified potential stage markers for maize anther development. A large number of antisense transcripts were detected and implicated in important biological functions given the enrichment of particular GO classes.

Expressed Sequence Tags↗

CD70 (TNFSF7) is expressed at high prevalence in renal cell carcinomas and is rapidly internalised on antibody binding.

In order to identify potential markers of renal cancer, the plasma membrane protein content of renal cell carcinoma (RCC)-derived cell lines was annotated using a proteomics process. One unusual protein identified at high levels in A498 and 786-O cells was CD70 (TNFSF7), a type II transmembrane receptor normally expressed on a subset of B, T and NK cells, where it plays a costimulatory role in immune cell activation. Immunohistochemical analysis of CD70 expression in multiple carcinoma types demonstrated strong CD70 staining in RCC tissues. Metastatic tissues from eight of 11 patients with clear cell RCC were positive for CD70 expression. Immunocytochemical analysis demonstrated that binding of an anti-CD70 antibody to CD70 endogenously expressed on the surface of A498 and 786-O cell lines resulted in the rapid internalisation of the antibody-receptor complex. Coincubation of the internalising anti-CD70 antibody with a saporin-conjugated secondary antibody before addition to A498 cells resulted in 50% cell killing. These data indicate that CD70 represents a potential target antigen for toxin-conjugated therapeutic antibody treatment of RCC.

Antibodies↗

[Cancer genome or the development of molecular portraits of tumors].

The rapid development of cancer genomics is due to important progresses in oncogenesis, human genome sequencing and emergence of new technologies in genome and transcriptome analysis. In this context, the aim of the French program 'Cartes d'Identites des Tumeurs--Molecular Portraits of Tumors' is to build a public data base containing a pan genome assessment of genome and transcriptome alterations in the major types of tumors as well as in relevant normal cells and experimental models. Data mining is done in the context of genome annotations and clinical and biological informations attached to the enrolled samples. The goal of the program is to define new tests useful for diagnostic procedures in clinical laboratories and new targets for biological treatments of tumors.

France↗

Gene profiling and bioinformatic analysis of Schwann cell embryonic development and myelination.

To elucidate the molecular mechanisms involved in Schwann cell development, we profiled gene expression in the developing and injured rat sciatic nerve. The genes that showed significant changes in expression in developing and dedifferentiated nerve were validated with RT-PCR, in situ hybridisation, Western blot and immunofluorescence. A comprehensive approach to annotating micro-array probes and their associated transcripts was performed using Biopendium, a database of sequence and structural annotation. This approach significantly increased the number of genes for which a functional insight could be found. The analysis implicates agrin and two members of the collapsin response-mediated protein (CRMP) family in the switch from precursors to Schwann cells, and synuclein-1 and alphaB-crystallin in peripheral nerve myelination. We also identified a group of genes typically related to chondrogenesis and cartilage/bone development, including type II collagen, that were expressed in a manner similar to that of myelin-associated genes. The comprehensive function annotation also identified, among the genes regulated during nerve development or after nerve injury, proteins belonging to high-interest families, such as cytokines and kinases, and should therefore provide a uniquely valuable resource for future research.

Agrin↗

Enzyme-Metabolite Network Analysis of Endometrial Cancer-Derived Extracellular Vesicles Through Integrated Proteomics and Metabolomics.

Endometrial cancer (EC) is the most common gynecological malignancy in high-income countries. Extracellular vesicles (EVs) are key mediators of intercellular communication and metabolic reprogramming, but their molecular cargo in EC remains poorly characterized. EVs were isolated from four EC cell lines representing Type I and Type II subtypes (AN3CA, ISHIKAWA, HEC1A, and KLE). Untargeted metabolomics was performed by HILIC-LC-MS/MS, proteomics by data-independent acquisition (DIA) mass spectrometry, and multi-omics integration using MetaboAnalyst and OmicsNet. Metabolomic profiling identified 1463 annotated features and revealed significant differences among EC cell lines (PERMANOVA, p = 0.002). Twenty-eight differentially abundant metabolites, including lactic acid, succinic acid, and uric acid, were identified. Proteomic analysis quantified 8513 proteins with subtype-specific expression patterns. Integrated analysis revealed seven significantly enriched pathways, including glycolysis/gluconeogenesis, central carbon metabolism in cancer, and the pentose phosphate pathway. Increased LDHA abundance in metastatic AN3CA-derived EVs was confirmed by Western blot (p = 0.047). EC-derived EVs display subtype- and metastatic-status-specific metabolo-proteomic signatures, with glycolysis, TCA cycle remodeling, and central carbon metabolism as convergent pathway signatures of molecular reprogramming. These findings establish a multi-omics framework for characterizing EV cargo in EC and identify candidate enzyme-metabolite nodes for future biomarker validation in patient-derived specimens.

Female↗

ChromBERT-tools: a versatile toolkit for context-specific regulatory representations of transcription regulators across different cell types.

SUMMARY: Representations that encode the genome-wide regulatory behavior of transcription regulators provide a foundation for flexible transcription modeling and in silico regulatory analysis. Existing regulator representations are commonly derived from gene co-expression, motif annotations, or static protein features, which capture useful but limited aspects of regulator identity but do not directly model how regulators participate in region-specific regulatory programs across the genome. ChromBERT addresses this gap by learning context-aware regulatory representations from large-scale ChIP-seq data. However, routine bioinformatics applications require lightweight, accessible, and modular tools for generating, adapting, and interpreting these representations in user-defined biological contexts. Here, we present ChromBERT-tools, a user-oriented toolkit built upon ChromBERT that converts its regulatory representation framework into practical workflows for customizable analysis across cellular contexts. ChromBERT-tools provides command-line interfaces and Python APIs organized into three functional layers: representation generation, predictive modeling, and regulatory interpretation. The representation generation layer produces representations of genomic regions and transcription regulators. The predictive modeling layer fine-tunes ChromBERT for genome-wide regulatory activity prediction through classification or regression tasks, with optimized implementation to reduce running time and computational resource requirements. The regulatory interpretation layer supports inference of the context-specific roles of cis-regulatory elements and transcription regulators. These modules can be used independently or integrated into end-to-end workflows, enabling flexible analyses across diverse datasets. ChromBERT-tools lowers the barrier to applying context-specific regulatory representations in routine genomic analyses. AVAILABILITY AND IMPLEMENTATION: ChromBERT-tools is freely available at https://github.com/TongjiZhanglab/ChromBERT-tools, with documentation at https://chrombert-tools.readthedocs.io/en/latest/. A frozen archival snapshot is available on Zenodo under DOI: 10.5281/zenodo.20094206.

Software↗

Analysis of bovine mammary gland EST and functional annotation of the Bos taurus gene index.

Functional genomic studies of the mammary gland require an appropriate collection of cDNA sequences to assess gene expression patterns from the different developmental and operational states of underlying cell types. To better capture the range of gene expression, a normalized cDNA library was constructed from pooled bovine mammary tissues, and 23,202 expressed sequence tags (EST) were produced and deposited into GenBank. Assembly of these EST with sequences in the Bos taurus Gene Index (BtGI) helped to form 5751 of the current 23,883 tentative consensus (TC) sequences. The majority (87%) of these 5751 assemblies contained only one to three mammary-derived EST. In contrast, 18% of the mammary EST assembled with TC sequences corresponding to 12 genes. These results suggest library normalization was only partially effective, because the reduction in EST for genes abundantly transcribed during lactation could be attributed to pooling. For better assessment of novel content in the mammary library and to add to existing annotation of all bovine sequence elements, gene ontology assignments, and comparative sequence analyses against human genome sequence, human and rodent gene indices, and an index of orthologous alignments of genes across eukaryotes (TOGA) were performed, and results were added to existing BtGI annotation. Over 35,000 of the bovine elements significantly matched human genome sequence, and the positions of some alignments (3%) were unique relative to those using human expressed sequences. Because 3445 TC sequences had no significant match with any data set, mammary-derived cDNA clones representing 23 of these elements were analyzed further for expression and novelty. Only one clone met criteria suggesting the corresponding gene was a divergent ortholog or expressed sequence unique to cattle. These results demonstrate that bovine sequence expression data serve as a resource for characterizing mammalian transcriptomes and identifying those genes potentially unique to ruminants.

Animals↗

Simultaneous modelling of metabolic, genetic and product-interaction networks.

The creation of cell models from annotated genome information, as well as additional data from other databases, requires both a format and medium for its distribution. Standards are described for the representation of the data in the form of Document Type Definitions (DTDs) for XML files. Separate DTDs are detailed for genetic, metabolic and gene product-interaction networks, which can be used to hold information on individual subsystems, or which may be combined to create a whole cell DTD. In the execution of this work, a fifth DTD was also created for a metabolite thesaurus, which allows incorporation of metabolite synonyms and generic nomenclature data into the models. A gene-regulation classification scheme was also created, to facilitate incorporation of gene regulatory information in an efficient manner. The work is described with particular reference to the metabolic network of Escherichia coli, which contains 808 individual enzymes. The assignment of confidence levels to these data, through the use of Gene Ontology evidence codes, is highlighted. In silico investigations may now be performed using the mathematical simulation workbench, DBsolve, which incorporates the facility to introduce data directly from XML.

Computational Biology↗

Synergistic computational and experimental proteomics approaches for more accurate detection of active serine hydrolases in yeast.

An analysis of the structurally and catalytically diverse serine hydrolase protein family in the Saccharomyces cerevisiae proteome was undertaken using two independent but complementary, large-scale approaches. The first approach is based on computational analysis of serine hydrolase active site structures; the second utilizes the chemical reactivity of the serine hydrolase active site in complex mixtures. These proteomics approaches share the ability to fractionate the complex proteome into functional subsets. Each method identified a significant number of sequences, but 15 proteins were identified by both methods. Eight of these were unannotated in the Saccharomyces Genome Database at the time of this study and are thus novel serine hydrolase identifications. Three of the previously uncharacterized proteins are members of a eukaryotic serine hydrolase family, designated as Fsh (family of serine hydrolase), identified here for the first time. OVCA2, a potential human tumor suppressor, and DYR-SCHPO, a dihydrofolate reductase from Schizosaccharomyces pombe, are members of this family. Comparing the combined results to results of other proteomic methods showed that only four of the 15 proteins were identified in a recent large-scale, "shotgun" proteomic analysis and eight were identified using a related, but similar, approach (neither identifies function). Only 10 of the 15 were annotated using alternate motif-based computational tools. The results demonstrate the precision derived from combining complementary, function-based approaches to extract biological information from complex proteomes. The chemical proteomics technology indicates that a functional protein is being expressed in the cell, while the computational proteomics technology adds details about the specific type of function and residue that is likely being labeled. The combination of synergistic methods facilitates analysis, enriches true positive results, and increases confidence in novel identifications. This work also highlights the risks inherent in annotation transfer and the use of scoring functions for determination of correct annotations.

Amino Acid Sequence↗

GenomeRNAi: a database for cell-based RNAi phenotypes.

RNA interference (RNAi) has emerged as a powerful tool to generate loss-of-function phenotypes in a variety of organisms. Combined with the sequence information of almost completely annotated genomes, RNAi technologies have opened new avenues to conduct systematic genetic screens for every annotated gene in the genome. As increasing large datasets of RNAi-induced phenotypes become available, an important challenge remains the systematic integration and annotation of functional information. Genome-wide RNAi screens have been performed both in Caenorhabditis elegans and Drosophila for a variety of phenotypes and several RNAi libraries have become available to assess phenotypes for almost every gene in the genome. These screens were performed using different types of assays from visible phenotypes to focused transcriptional readouts and provide a rich data source for functional annotation across different species. The GenomeRNAi database provides access to published RNAi phenotypes obtained from cell-based screens and maps them to their genomic locus, including possible non-specific regions. The database also gives access to sequence information of RNAi probes used in various screens. It can be searched by phenotype, by gene, by RNAi probe or by sequence and is accessible at http://rnai.dkfz.de.

Animals↗

GlycoSuiteDB: a new curated relational database of glycoprotein glycan structures and their biological sources.

GlycoSuiteDB is a relational database that curates information from the scientific literature on glyco-protein derived glycan structures, their biological sources, the references in which the glycan was described and the methods used to determine the glycan structure. To date, the database includes most published O:-linked oligosaccharides from the last 50 years and most N:-linked oligosaccharides that were published in the 1990s. For each structure, information is available concerning the glycan type, linkage and anomeric configuration, mass and composition. Detailed information is also provided on native and recombinant sources, including tissue and/or cell type, cell line, strain and disease state. Where known, the proteins to which the glycan structures are attached are reported, and cross-references to the SWISS-PROT/TrEMBL protein sequence databases are given if applicable. The GlycoSuiteDB annotations include literature references which are linked to PubMed, and detailed information on the methods used to determine each glycan structure are noted to help the user assess the quality of the structural assignment. GlycoSuiteDB has a user-friendly web interface which allows the researcher to query the database using mono-isotopic or average mass, monosaccharide composition, glycosylation linkages (e.g. N:- or O:-linked), reducing terminal sugar, attached protein, taxonomy, tissue or cell type and GlycoSuiteDB accession number. Advanced queries using combinations of these parameters are also possible. GlycoSuiteDB can be accessed on the web at http://www.glycosuite.com.

Animals↗

Profiling of generic anti-phosphopeptide antibodies and kinases with peptide microarrays using radioactive and fluorescence-based assays.

Kinases represent one of the largest enzyme families and key regulatory proteins in the cell. Only a small subset of these enzymes has been characterised so far. We have prepared different types of phosphopeptide and peptide microarrays displaying peptides deduced from annotated human phosphorylation sites and cytoplasmic domains of all annotated human membrane proteins. This approach was enabled by fully-automated high throughput micro-scale synthesis of peptides by the SPOT technology combined with chemo-selective immobilisation on modified glass slides. The phosphopeptide microarrays displaying 2923 peptides in total have been used for the characterisation of commercially available generic anti-phosphopeptide antibodies. This enabled us to detect Abl kinase activity on a microarray with anti-phosphotyrosine antibodies yielding results comparable to those obtained from a radioactive assay. More than 13 000 peptides deposited on six glass slides were used to profile casein kinase 2 (CK2) using a radioactive assay, since no generic antibody for the reliable detection of serine or threonine phosphorylation could be identified. All previously identified substrates were detected in the microarray experiment. In order to confirm whether substrates on the microarray are substrates in solution phase assays, more than 700 peptides were synthesised and tested with CK2 in a solution phase assay. All substrates identified in the solution phase assay were also detected on the microarray.

Amino Acid Sequence↗

High-quality peptide evidence for annotating non-canonical open reading frames as human proteins.

A major scientific drive is to characterize the protein-coding genome as it provides the primary basis for the study of human health. But the fundamental question remains: what has been missed in prior genomic analyses? Over the past decade, the translation of non-canonical open reading frames (ncORFs) has been observed across human cell types and disease states, with major implications for proteomics, genomics, and clinical science. However, the impact of ncORFs has been limited by the absence of a large-scale understanding of their contribution to the human proteome. Here, we report the collaborative efforts of stakeholders in proteomics, immunopeptidomics, Ribo-seq ORF discovery, and gene annotation, to produce a consensus landscape of protein-level evidence for ncORFs. We show that at least 25% of a set of 7,264 ncORFs give rise to translated gene products, yielding over 3,000 peptides in a pan-proteome analysis encompassing 3.8 billion mass spectra from 95,520 experiments. With these data, we developed an annotation framework for ncORFs and created public tools for researchers through GENCODE and PeptideAtlas. This work will provide a platform to advance ncORF-derived proteins in biomedical discovery and, beyond humans, diverse animals and plants where ncORFs are similarly observed.

GENCODE↗

Cross-species annotation of basic leucine zipper factor interactions: Insight into the evolution of closed interaction networks.

Dimeric basic leucine zipper (bZIP) factors constitute one of the most important classes of enhancer-type transcription factors. In vertebrates, bZIP factors are involved in many cellular processes, including cell survival, learning and memory, cancer progression, lipid metabolism, and a variety of developmental processes. These factors have the ability to homodimerize and heterodimerize in a specific and predictable manner, resulting in hundreds of dimers with unique effects on transcription. In recent years, several studies have described dimerization preferences for bZIP factors from different species, including Homo sapiens, Drosophila melanogaster, Arabidopsis thaliana, and Saccharomyces cerevisiae. Here, these findings are summarized as novel, graphical representations of closed, interacting protein networks. These representations combine phylogenetic information, DNA-binding properties, and dimerization preference. Beyond summarizing bZIP dimerization preferences within selected species, we have included annotation for a solitary bZIP factor found in the primitive eukaryote, Giardia lamblia, a possible evolutionary precursor to the complex networks of bZIP factors encoded by other genomes. Finally, we discuss the fundamental similarities and differences between dimerization networks within the context of bZIP factor evolution.

Amino Acid Sequence↗

Methods for the functional genomic analysis of ubiquitin ligases.

Ubiquitin ligases (E3s) are critical components of the ubiquitin-proteasome system as they are the major determinants of specificity in ubiquitin conjugation. The number of predicted E3s in the mammalian genome is exceeding 400 and is represented by two major subfamilies: HECT domain-containing E3s and RING finger-type E3s. Given the size of this protein family and lack of knowledge on the functions of most of these 400 proteins, their functional annotation should benefit from modern genomic tools. This article presents a methodology consisting of the use of a cDNA expression library to identify suppressors of polyglutamine (polyQ)-mediated protein aggregate formation in cells, as an example of a genomic approach to assign functions to E3s. In this screen, we identified novel RING finger-type E3s exhibiting suppressor activity among >50% of all the potential E3s in the mouse and human genomes. This method could be adapted easily to identify E3s that function in other processes and signaling pathways.

Algorithms↗