PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Cell type annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Diversity and relatedness among the type I interferons.

Type I interferons (IFNs) include the IFN-alpha family of subtypes, IFN-beta, IFN-omega, IFN-tau, IFN-kappa, IFN-lambda, and IFN-zeta. IFN genes lack introns and encode secretory signal peptide sequences that are proteolytically cleaved prior to secretion from the cell. In contrast to the approximately 50% amino acid sequence identity among the human IFN-alpha subtypes, human IFN-alphas share approximately 22% identity with human IFN-beta and 37% identity with human IFN-omega. Many of the conserved residues among the type I IFNs are implicated in receptor recognition and structural integrity. This report provides an update on the gene annotations for the mouse and human IFN gene clusters on chromosome 4 and 9, respectively, with accompanying amino acid sequence alignments. Based on sequence identities, a phylogenic tree analysis for the different mammalian Type I IFNs is also presented, showing the high degree of relatedness among these IFNs. Notably, sequence alignment of the different human and mouse IFN promoter regions reveals different signature patterns for transcription factor binding sites, implying different inducers might differentially activate the transcription of the different IFNs.

Amino Acid Sequence↗

Protein interaction mapping on a functional shotgun sequence of Rickettsia sibirica.

Protein interaction maps can reveal novel pathways and functional complexes, allowing 'guilt by association' annotation of uncharacterized proteins. To address the need for large-scale protein interaction analyses, a bacterial two-hybrid system was coupled with a whole genome shotgun sequencing approach for microbial genome analysis. We report the first large-scale proteomics study using this system, integrating de novo genome sequencing with functional interaction mapping and annotation in a high-throughput format. We apply the approach by shotgun sequencing and annotating the genome of Rickettsia sibirica strain 246, an obligate intracellular human pathogen among the Spotted Fever Group rickettsiae. The bacteria invade endothelial cells and cause lysis after large amounts of progeny have accumulated. Little is known about specific Rickettsial virulence factors and their mode of pathogenicity. Analysis of the combined genomic sequence and protein-protein interaction data for a set of virulence related Type IV secretion system (T4SS) proteins revealed over 250 interactions and will provide insight into the mechanism of Rickettsial pathogenicity.

Bacterial Proteins↗

Positional clustering of differentially expressed genes on human chromosomes 20, 21 and 22.

BACKGROUND: Clusters of genes co-expressed are known in prokaryotes (operons) and were recently described in several eukaryote organisms, including Human. According to some studies, these clusters consist of housekeeping genes, whereas other studies suggest that these clustered genes exhibit similar tissue specificity. Here we further explore the relationship between co-expression and chromosomal co-localization in the human genome by analyzing the expression status of the genes along the best-annotated chromosomes 20, 21 and 22. METHODS: Gene expression levels were estimated according to their publicly available ESTs and gene differential expressions were assessed using a previously described and validated statistical test. Gene sequences for chromosomes 20, 21 and 22 were taken from the Ensembl annotation. RESULTS: We identified clusters of genes specifically expressed in similar tissues along chromosomes 20, 21 and 22. These co-expression clusters occurred more frequently than expected by chance and may thus be biologically significant. CONCLUSION: The co-expression of co-localized genes might be due to higher chromatin structures influencing the gene availability for transcription in a given tissue or cell type.

Chromosomes, Human, Pair 20↗

Decoding the fine-scale structure of a breast cancer genome and transcriptome.

A comprehensive understanding of cancer is predicated upon knowledge of the structure of malignant genomes underlying its many variant forms and the molecular mechanisms giving rise to them. It is well established that solid tumor genomes accumulate a large number of genome rearrangements during tumorigenesis. End Sequence Profiling (ESP) maps and clones genome breakpoints associated with all types of genome rearrangements elucidating the structural organization of tumor genomes. Here we extend the ESP methodology in several directions using the breast cancer cell line MCF-7. First, targeted ESP is applied to multiple amplified loci, revealing a complex process of rearrangement and co-amplification in these regions reminiscent of breakage/fusion/bridge cycles. Second, genome breakpoints identified by ESP are confirmed using a combination of DNA sequencing and PCR. Third, in vitro functional studies assign biological function to a rearranged tumor BAC clone, demonstrating that it encodes anti-apoptotic activity. Finally, ESP is extended to the transcriptome identifying four novel fusion transcripts and providing evidence that expression of fusion genes may be common in tumors. These results demonstrate the distinct advantages of ESP including: (1) the ability to detect all types of rearrangements and copy number changes; (2) straightforward integration of ESP data with the annotated genome sequence; (3) immortalization of the genome; (4) ability to generate tumor-specific reagents for in vitro and in vivo functional studies. Given these properties, ESP could play an important role in a tumor genome project.

Breast Neoplasms↗

A genetic screen implicates miRNA-372 and miRNA-373 as oncogenes in testicular germ cell tumors.

Endogenous small RNAs (miRNAs) regulate gene expression by mechanisms conserved across metazoans. While the number of verified human miRNAs is still expanding, only few have been functionally annotated. To perform genetic screens for novel functions of miRNAs, we developed a library of vectors expressing the majority of cloned human miRNAs and created corresponding DNA barcode arrays. In a screen for miRNAs that cooperate with oncogenes in cellular transformation, we identified miR-372 and miR-373, each permitting proliferation and tumorigenesis of primary human cells that harbor both oncogenic RAS and active wild-type p53. These miRNAs neutralize p53-mediated CDK inhibition, possibly through direct inhibition of the expression of the tumor-suppressor LATS2. We provide evidence that these miRNAs are potential novel oncogenes participating in the development of human testicular germ cell tumors by numbing the p53 pathway, thus allowing tumorigenic growth in the presence of wild-type p53.

Cells, Cultured↗

MapID-based quantitative mapping of chemical modifications and expression of human transfer RNA.

Detection and quantification of tRNA chemical modifications are critical for understanding their regulatory functions in biology and diseases. However, tRNA-seq-based methods for modification mapping encountered challenges both experimentally (poor processivity of heavily modified tRNAs during reverse transcription or RT) and bioinformatically (frequent reads misalignment to highly similar tRNA genes). Here, we report "MapID-tRNA-seq" where we deployed an evolved reverse transcriptase (RT-1306) into tRNA-seq and developed "MapIDs" that reduce redundancy of the human tRNA genome and explicitly annotate genetic variances. RT-1306 generated robust mutations against m1A and m3C, and RT stops against multiple bulky roadblock modifications. MapID-assisted data processing enabled systematic exclusion of false-positive discoveries of modifications which arise from reads misalignment onto similar genes. We applied MapID-tRNA-seq into mapping m1A, m3C and expression levels of tRNAs in three mammary cell lines, which revealed cell-type dependent modification sites and potential translational regulation of the reduced mitochondrial activities in breast cancer.

Humans↗

Transcriptomic and proteomic analyses of the pMOL30-encoded copper resistance in Cupriavidus metallidurans strain CH34.

The four replicons of Cupriavidus metallidurans CH34 (the genome sequence was provided by the US Department of Energy-University of California Joint Genome Institute) contain two gene clusters putatively encoding periplasmic resistance to copper, with an arrangement of genes resembling that of the copSRABCD locus on the 2.1 Mb megaplasmid (MPL) of Ralstonia solanacearum, a closely related plant pathogen. One of the copSRABCD clusters was located on the 2.6 Mb MPL, while the second was found on the pMOL30 (234 kb) plasmid as part of a larger group of genes involved in copper resistance, spanning 17 857 bp in total. In this region, 19 ORFs (copVTMKNSRABCDIJGFLQHE) were identified based on the sequencing of a fragment cloned in an IncW vector, on the preliminary annotation by the Joint Genome Institute, and by using transcriptomic and proteomic data. When introduced into plasmid-cured derivatives of C. metallidurans CH34, the cop locus was able to restore the wild-type MIC, albeit with a biphasic survival curve, with respect to applied Cu(II) concentration. Quantitative-PCR data showed that the 19 ORFs were induced from 2- to 1159-fold when cells were challenged with elevated Cu(II) concentrations. Microarray data showed that the genes that were most induced after a Cu(II) challenge of 0.1 mM belonged to the pMOL30 cop cluster. Megaplasmidic cop genes were also induced, but at a much lower level, with the exception of the highly expressed MPL copD. Proteomic data allowed direct observation on two-dimensional gel electrophoresis, and via mass spectrometry, of pMOL30 CopK, CopR, CopS, CopA, CopB and CopC proteins. Individual cop gene expression depended on both the Cu(II) concentration and the exposure time, suggesting a sequential scheme in the resistance process, involving genes such as copK and copT in an initial phase, while other genes, such as copH, seem to be involved in a late response phase. A concentration of 0.4 mM Cu(II) was the highest to induce maximal expression of most cop genes.

Amino Acid Sequence↗

A single-cell atlas of multiple myeloma defines malignant archetypes and proliferative states.

Multiple myeloma (MM) is a plasma-cell malignancy with extensive genomic and transcriptional heterogeneity, limiting disease classification and precision therapy. Here we generated a clinically annotated, population-scale, single-cell atlas of MM from 341 individuals spanning the disease and treatment continuum. We identified five recurrent malignant transcriptional archetypes and an orthogonal proliferative program associated with genomic features, therapeutic resistance and clinical outcomes. Validation in the independent CoMMpass cohort demonstrated robustness, prognostic relevance and portability across platforms. We developed a single-cell, target-discovery pipeline prioritizing malignant enrichment, cell-type specificity and tissue restriction, identifying FCRL2 as a plasma-restricted or B cell-lineage-restricted surface target expressed by malignant plasma cells. FCRL2-targeted chimeric antigen receptor T cells demonstrated antigen-specific activity in vitro and survival benefit in vivo. Together, these data provide a clinically actionable blueprint for patient stratification and precision target nomination in plasma-cell malignancies.

Multiple Myeloma↗

[Transcriptomes for serial analysis of gene expression].

The availability of the sequences for whole genomes is changing our understanding of cell biology. Functional genomics refers to the comprehensive analysis, at the protein level (proteome) and at the mRNA level (transcriptome) of all events associated with the expression of whole sets of genes. New methods have been developed for transcriptome analysis. Serial Analysis of Gene Expression (SAGE) is based on the massive sequential analysis of short cDNA sequence tags. Each tag is derived from a defined position within a transcript. Its size (14 bp) is sufficient to identify the corresponding gene and the number of times each tag is observed provides an accurate measurement of its expression level. Since tag populations can be widely amplified without altering their relative proportions, SAGE may be performed with minute amounts of biological extract. Dealing with the mass of data generated by SAGE necessitates computer analysis. A software is required to automatically detect and count tags from sequence files. Criterias allowing to assess the quality of experimental data can be included at this stage. To identify the corresponding genes, a database is created registering all virtual tags susceptible to be observed, based on the present status of the genome knowledge. By using currently available database functions, it is easy to match experimental and virtual tags, thus generating a new database registering identified tags, together with their expression levels. As an open system, SAGE is able to reveal new, yet unknown, transcripts. Their identification will become increasingly easier with the progress of genome annotation. However, their direct characterization can be attempted, since tag information may be sufficient to design primers allowing to extend unknown sequences. A major advantage of SAGE is that, by measuring expression levels without reference to an arbitrary standard, data are definitively acquired and cumulative. All publicly available data can thus be stored in a unique database, facilitating whole-genome analysis of differential expression between cell types, normal and diseased samples, or samples with and without drug treatment. SAGE data are readily amenable to statistical comparisons, allowing to determine the level of confidence of the observed variations. A major limitation of SAGE is that, because each analysis is obligatory performed on the whole set of expressed genes, it can hardly be performed on multiple samples, for example in kinetics studies or to compare the effects of large numbers of drugs. To overcome this limitation, high-throughput detection of a subset of mRNAs is more rapidly performed by parallel hybridization of mRNAs on arrays of nucleic acids immobilized on solid supports. From this point of view, a SAGE platform is a powerful instrument for selecting the most informative subset of genes, assembling them to design microarrays dedicated to a specific problem and calibrating measurement by comparison with a standard cell model for which SAGE data are available. This approach is an attractive alternative to strategies based exclusively on pangenomic arrays. A very large amount of SAGE data are already available and the problem is now to extract their biological meaning. Knowledge on metabolic pathways is already organized so that its successful integration in a SAGE platform can be undertaken. For other cell components and pathways, the problem lies on the lack of controlled vocabulary to describe gene activities, starting form a clear definition of the concept of biological function itself. Progress in gene and cell ontology is expected to facilitate computer-based extraction of biological knowledge from existing and forthcoming SAGE data.

Animals↗

IMGT/LIGM-DB, the IMGT comprehensive database of immunoglobulin and T cell receptor nucleotide sequences.

IMGT/LIGM-DB is the IMGT comprehensive database of immunoglobulin (IG) and T cell receptor (TR) nucleotide sequences from human and other vertebrate species. It was created in 1989 by LIGM, Montpellier, France and is the oldest and the largest database of IMGT. IMGT/LIGM-DB includes all germline (non-rearranged) and rearranged IG and TR genomic DNA (gDNA) and complementary DNA (cDNA) sequences published in generalist databases. IMGT/LIGM-DB allows searches from the Web interface according to biological and immunogenetic criteria through five distinct modules depending on the user interest. For a given entry, nine types of display are available including the IMGT flat file, the translation of the coding regions and the analysis by the IMGT/V-QUEST tool. IMGT/LIGM-DB distributes expertly annotated sequences. The annotations hugely enhance the quality and the accuracy of the distributed detailed information. They include the sequence identification, the gene and allele classification, the constitutive and specific motif description, the codon and amino acid numbering, and the sequence obtaining information, according to the main concepts of IMGT-ONTOLOGY. They represent the main source of IG and TR gene and allele knowledge stored in IMGT/GENE-DB and in the IMGT reference directory. IMGT/LIGM-DB is freely available at http://imgt.cines.fr.

Animals↗

Detection of novel gene expression in paraffin-embedded tissues by isotopic in situ hybridization in tissue microarrays.

Correlating altered gene expression patterns with particular disease states is a critical step in understanding disease processes and developing treatment strategies. Many thousands of novel gene sequences have recently been annotated in public and private databases and are now available for analysis. Tissue-specific expression patterns of these sequences can be evaluated physically on DNA arrays and other high throughput assays, or virtually by bioinformatics mining of expressed sequence tag (EST) databases. As a secondary screening tool, in situ hybridisation (ISH) not only confirms tissue specificity, but also reveals what is often valuable information about cell-type expression patterns of nov16l sequences. Due to their availability and long-term stability at room temperature, formalin-fixed paraffin-embedded clinical specimens provide an invaluable resource for evaluating expression patterns of novel human genes. We describe a high-throughput approach for identifying and quantifying the expression of novel genes in paraffin-embedded human tissues using isotopic in situ hybridisation and tissue microarrays (TMA).

Blotting, Northern↗

From gene networks to brain networks.

The brain's structural organization is so complex that 2,500 years of analysis leaves pervasive uncertainty about (i) the identity of its basic parts (regions with their neuronal cell types and pathways interconnecting them), (ii) nomenclature, (iii) systematic classification of the parts with respect to topographic relationships and functional systems and (iv) the reliability of the connectional data itself. Here we present a prototype knowledge management system (http://brancusi.usc.edu/bkms/) for analyzing the architecture of brain networks in a systematic, interactive and extendable way. It supports alternative interpretations and models, is based on fully referenced and annotated data and can interact with genomic and functional knowledge management systems through web services protocols.

Animals↗

Temperature-regulated transcription in the pathogenic fungus Cryptococcus neoformans.

The basidiomycete fungus Cryptococcus neoformans is an opportunistic pathogen of worldwide importance that causes meningitis, leading to death in immunocompromised individuals. Unlike many basidiomycete fungi, C. neoformans is thermotolerant, and its ability to grow at 37 degrees C is considered to be a virulence factor. We used serial analysis of gene expression (SAGE) to characterize the transcriptomes of C. neoformans strains that represent two varieties with different polysaccharide capsule serotypes. These include a serotype D strain of the C. neoformans variety neoformans and a serotype A strain of variety grubii. In this report, we describe the construction and characterization of SAGE libraries from each strain grown at 25 degrees C and 37 degrees C. The SAGE data reveal transcriptome differences between the two strains, even at this early stage of analysis, and identify sets of genes with higher transcript levels at 25 degrees C or 37 degrees C. Notably, growth at the lower temperature increased transcript levels for histone genes, indicating a general influence of temperature on chromatin structure. At 37 degrees C, we noted elevated transcript levels for several genes encoding heat shock proteins and translation machinery. Some of these genes may play a role in temperature-regulated phenotypes in C. neoformans, such as the adaptation of the fungus to growth in the host and the dimorphic transition between budding and filamentous growth. Overall, this work provides the most comprehensive gene expression data available for C. neoformans; this information will be a critical resource both for gene discovery and genome annotation in this pathogen.

Blotting, Northern↗

Functional Annotation of the Major Histocompatibility Complex Locus.

The human major histocompatibility complex (MHC) locus has the greatest density of disease-associations in the human genome, including links to over 100 polygenic disorders. Its complex haplotype structure, rich gene density, and high degree of linkage disequilibrium combine to make deciphering the gene regulatory logic of the MHC locus extremely challenging. Employing complementary high-throughput CRISPR interference (CRISPRi) and activation (CRISPRa) epigenetic screens coupled with single-cell transcriptome profiling across three distinct human cell types, we identified hundreds of new connections between cis -regulatory elements (CREs) and their target genes in this locus. These CRE-gene links are largely cell type-specific and act as enhancers. Additionally, some CREs have complex features, including harboring both active and repressive histone marks, lacking chromatin accessibility, targeting multiple genes, or acting as silencers. Computational methods fail to predict a majority of these CRE-gene connections. These findings emphasize the potential for functional perturbation experiments to dissect complex loci and reveal shared and cell type-specific regulatory mechanisms relevant to genomics of complex diseases. Collectively, this study provides a unique resource for understanding the complex regulatory landscape within the MHC locus and supports the need for creating new models that encompass CRE-gene interactions, cell type-specific gene expression, and disease genetics in the noncoding genome.

Journal Article↗

IMGT unique numbering for immunoglobulin and T cell receptor variable domains and Ig superfamily V-like domains.

IMGT, the international ImMunoGeneTics database (http://imgt.cines.fr) is a high quality integrated information system specializing in immunoglobulins (IG), T cell receptors (TR) and major histocompatibility complex (MHC) of human and other vertebrates. IMGT provides a common access to expertly annotated data on the genome, proteome, genetics and structure of the IG and TR, based on the IMGT Scientific chart and IMGT-ONTOLOGY. The IMGT unique numbering defined for the IG and TR variable regions and domains of all jawed vertebrates has allowed a redefinition of the limits of the framework (FR-IMGT) and complementarity determining regions (CDR-IMGT), leading, for the first time, to a standardized description of mutations, allelic polymorphisms, 2D representations (Colliers de Perles) and 3D structures, whatever the antigen receptor, the chain type, or the species. The IMGT numbering has been extended to the V-like domain and is, therefore, highly valuable for comparative analysis and evolution studies of proteins belonging to the IG superfamily.

Amino Acid Sequence↗

Comparative organization of cattle chromosome 5 revealed by comparative mapping by annotation and sequence similarity and radiation hybrid mapping.

A whole genome cattle-hamster radiation hybrid cell panel was used to construct a map of 54 markers located on bovine chromosome 5 (BTA5). Of the 54 markers, 34 are microsatellites selected from the cattle linkage map and 20 are genes. Among the 20 mapped genes, 10 are new assignments that were made by using the comparative mapping by annotation and sequence similarity strategy. A LOD-3 radiation hybrid framework map consisting of 21 markers was constructed. The relatively low retention frequency of markers on this chromosome (19%) prevented unambiguous ordering of the other 33 markers. The length of the map is 398.7 cR, corresponding to a ratio of approximately 2.8 cR(5,000)/cM. Type I genes were binned for comparison of gene order among cattle, humans, and mice. Multiple internal rearrangements within conserved syntenic groups were apparent upon comparison of gene order on BTA5 and HSA12 and HSA22. A similarly high number of rearrangements were observed between BTA5 and MMU6, MMU10, and MMU15. The detailed comparative map of BTA5 should facilitate identification of genes affecting economically important traits that have been mapped to this chromosome and should contribute to our understanding of mammalian chromosome evolution.

Animals↗

DeepPlaque: a scalable multimodal platform for Aβ pathology and cell analysis in Alzheimer's disease.

Histological analysis is essential for understanding disease pathology and the microenvironment, particularly in Alzheimer's disease (AD), characterized by beta-amyloid (Aβ) plaques that exist as diffuse, fibrillar, and core species, with distinct toxicity levels. However, accurate classification of Aβ plaque types in postmortem brain tissues and profiling of surrounding cells present significant challenges. To address these challenges, we developed "DeepPlaque", an integrated system featuring "PlaqueNet", a deep learning model for automated classification of Aβ plaque species from diverse imaging platforms. DeepPlaque includes automated workflows for cellular phenotyping and proteomic profiling through targeted laser microdissection. PlaqueNet achieves expert-level accuracy (AUC > 90%) in classifying the 3 major Aβ plaque species, supporting consistent and large-scale annotation. By integrating spatial cellular phenotyping with laser microdissection, DeepPlaque enables high-throughput proteomic analysis of Aβ plaque niches, revealing that microglia are more abundant around core and fibrillar Aβ plaques, with increased expression of apolipoprotein E and amyloid precursor protein in core Aβ plaques. This customizable platform enhances the molecular and cellular characterization of Aβ plaque-associated environments, providing critical insights into AD pathology.

Alzheimer Disease↗

Gene expression profile induced by 17 alpha-ethynyl estradiol in the prepubertal female reproductive system of the rat.

The profound effects of 17beta-estradiol on cell growth, differentiation, and general homeostasis of the reproductive and other systems, are mediated mostly by regulation of temporal and cell type-specific expression of different genes. In order to understand better the molecular events associated with the activation of the estrogen receptor (ER), we have used microarray technology to determine the transcriptional program and dose-response characteristics of exposure to a potent synthetic estrogen, 17 alpha-ethynyl estradiol (EE), during prepubertal development. Changes in patterns of gene expression were determined in the immature uterus and ovaries of Sprague-Dawley rats on postnatal day (PND) 24, 24 h after exposure to EE, at 0.001, 0.01, 0.1, 1 and 10 micro g EE/kg/day (sc), for four days (dosing from PND 20 to 23). The transcript profiles were compared between treatment groups and controls using oligonucleotide arrays to determine the expression level of approximately 7000 annotated rat genes and over 1740 expressed sequence tags (ESTs). Quantification of the number of genes whose expression was modified by the treatment, for each of the various doses of EE tested, showed clear evidence of a dose-dependent treatment effect that follows a monotonic response, concordant with the dose-response pattern of uterine wet-weight gain and luminal epithelial cell height. The number of genes whose expression is affected by EE exposure increases according to dose. At the highest dose tested of EE, we determined that the expression level of over 300 genes was modified significantly (p < or = 0.0001). A dose-dependent analysis of the transcript profile revealed a set of 88 genes whose expression is significantly and reproducibly modified (increased or decreased) by EE exposure (p < or = 0.0001). The results of this study demonstrate that, exposure to a potent estrogenic chemical during prepubertal maturation changes the gene expression profile of estrogen-sensitive tissues. Furthermore, the products of the EE-regulated genes identified in these tissues have a physiological role in different intracellular pathways, information that will be valuable to determine the mechanism of action of estrogens. Moreover, those genes could be used as biomarkers to identify chemicals with estrogenic activity.

Animals↗