PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Cell type annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Methods for the functional genomic analysis of ubiquitin ligases.

Ubiquitin ligases (E3s) are critical components of the ubiquitin-proteasome system as they are the major determinants of specificity in ubiquitin conjugation. The number of predicted E3s in the mammalian genome is exceeding 400 and is represented by two major subfamilies: HECT domain-containing E3s and RING finger-type E3s. Given the size of this protein family and lack of knowledge on the functions of most of these 400 proteins, their functional annotation should benefit from modern genomic tools. This article presents a methodology consisting of the use of a cDNA expression library to identify suppressors of polyglutamine (polyQ)-mediated protein aggregate formation in cells, as an example of a genomic approach to assign functions to E3s. In this screen, we identified novel RING finger-type E3s exhibiting suppressor activity among >50% of all the potential E3s in the mouse and human genomes. This method could be adapted easily to identify E3s that function in other processes and signaling pathways.

Algorithms↗

Cell-type signatures of Alzheimer's disease shared across population groups.

Genomic studies at single-cell resolution have identified several cell types associated with clinical and pathological traits in Alzheimer's disease1-9, but have not examined associations that are shared across populations. To bridge this gap, here we use single-nucleus RNA sequencing and assay for transposase-accessible chromatin with sequencing to profile cortical and subcortical regions in post-mortem brain-tissue samples from Latin, white (excluding Latin) and African American (excluding Latin) individuals. Using discrete and continuous dissections of molecular programs, we identify cell-type-specific clusters associated with Alzheimer's disease in a region-specific manner across all three population groups, including microglial (GPNMB+ and CD74+ subgroups), astrocytic (SERPINH1+, CD44+ and WIF1+ subgroups) and neuronal (SST+ GABAergic and superficial-layer glutamatergic) signatures. We also report continuous gene-expression factors in astrocytes and oligodendrocytes that are not captured by discrete cluster assignments, but which show strong associations with disease phenotypes; these factors are enriched for genes associated with annotated functions such as lipid processing and neurotransmitter reuptake. Finally, we find that molecular programs reveal six distinct subgroups of individuals with cognitive impairment that span all three populations, are not captured by neuropathology, and are instead distinguished by molecular signatures that are not universally present but are nonetheless associated with ante-mortem impairment. Overall, our study identifies key cell types and gene programs implicated in Alzheimer's disease that are shared across population groups, and underscores how representative sampling can capture both shared signatures and disease heterogeneity, thereby enabling better prioritization of key cell types for further investigation.

Female↗

Rare variant contribution to the heritability of coronary artery disease.

Whole genome sequences (WGS) enable discovery of rare variants which may contribute to missing heritability of coronary artery disease (CAD). To measure their contribution, we apply the GREML-LDMS-I approach to WGS of 4949 cases and 17,494 controls of European ancestry from the NHLBI TOPMed program. We estimate CAD heritability at 34.3% assuming a prevalence of 8.2%. Ultra-rare (minor allele frequency ≤ 0.1%) variants with low linkage disequilibrium (LD) score contribute ~50% of the heritability. We also investigate CAD heritability enrichment using a diverse set of functional annotations: i) constraint; ii) predicted protein-altering impact; iii) cis-regulatory elements from a cell-specific chromatin atlas of the human coronary; and iv) annotation principal components representing a wide range of functional processes. We observe marked enrichment of CAD heritability for most functional annotations. These results reveal the predominant role of ultra-rare variants in low LD on the heritability of CAD. Moreover, they highlight several functional processes including cell type-specific regulatory mechanisms as key drivers of CAD genetic risk.

Humans↗

CLN3 transcript complexity revealed by long-read RNA sequencing analysis.

BACKGROUND: Batten disease is a group of rare inherited neurodegenerative diseases. Juvenile CLN3 disease is the most prevalent type, and the most common pathogenic variant shared by most patients is the "1-kb" deletion which removes two internal coding exons (7 and 8) in CLN3. Previously, we identified two transcripts in patient fibroblasts homozygous for the 1-kb deletion: the 'major' and 'minor' transcripts. To understand the full variety of disease transcripts and their role in disease pathogenesis, it is necessary to first investigate CLN3 transcription in "healthy" samples without juvenile CLN3 disease. METHODS: We leveraged PacBio long-read RNA sequencing datasets from ENCODE to investigate the full range of CLN3 transcripts across various tissues and cell types in human control samples. Then we sought to validate their existence using data from different sources. RESULTS: We found that a readthrough gene affects the quantification and annotation of CLN3. After taking this into account, we detected over 100 novel CLN3 transcripts, with no dominantly expressed CLN3 transcript. The most abundant transcript has median usage of 42.9%. Surprisingly, the known disease-associated 'major' transcripts are detected. Together, they have median usage of 1.5% across 22 samples. Furthermore, we identified 48 CLN3 ORFs, of which 26 are novel. The predominant ORF that encodes the canonical CLN3 protein isoform has median usage of 66.7%, meaning around one-third of CLN3 transcripts encode protein isoforms with different stretches of amino acids. The same ORFs could be found with alternative UTRs. Moreover, we were able to validate the translational potential of certain transcripts using public mass spectrometry data. CONCLUSION: Overall, these findings provide valuable insights into the complexity of CLN3 transcription, highlighting the importance of studying both canonical and non-canonical CLN3 protein isoforms as well as the regulatory role of UTRs to fully comprehend the regulation and function(s) of CLN3. This knowledge is essential for investigating the impact of the 1-kb deletion and rare pathogenic variants on CLN3 transcription and disease pathogenesis.

Humans↗

T1DBase: integration and presentation of complex data for type 1 diabetes research.

T1DBase (http://T1DBase.org) [Smink et al. (2005) Nucleic Acids Res., 33, D544-D549; Burren et al. (2004) Hum. Genomics, 1, 98-109] is a public website and database that supports the type 1 diabetes (T1D) research community. T1DBase provides a consolidated T1D-oriented view of the complex data world that now confronts medical researchers and enables scientists to navigate from information they know to information that is new to them. Overview pages for genes and markers summarize information for these elements. The Gene Dossier summarizes information for a list of genes. GBrowse [Stein et al. (2002) Genome Res., 10, 1599-1610] displays genes and other features in their genomic context, and Cytoscape [Shannon et al. (2003) Genome Res., 13, 2498-2504] shows genes in the context of interacting proteins and genes. The Beta Cell Gene Atlas shows gene expression in beta cells, islets, and related cell types and lines, and the Tissue Expression Viewer shows expression across other tissues. The Microarray Viewer shows expression from more than 20 array experiments. The Beta Cell Gene Expression Bank contains manually curated gene and pathway annotations for genes expressed in beta cells. T1DMart is a query tool for markers and genotypes. PosterPages are 'home pages' about specific topics or datasets. The key challenge, now and in the future, is to provide powerful informatics capabilities to T1D scientists in a form they can use to enhance their research.

Animals↗

Altered lung gene expression in CCSP-null mice suggests immunoregulatory roles for Clara cells.

Clara cell secretory protein (CCSP) is one of the most abundant proteins present in airway lining fluid of mammals. In an effort to elucidate the function of CCSP, we established CCSP-null [CCSP(-/-)] mice and demonstrated altered sensitivity to various environmental agents including oxidant pollutants and microorganisms. Although CCSP deficiency itself may be central to the observed changes in environmental susceptibility, altered lung gene expression associated with CCSP deficiency may contribute to the observed phenotype. To determine whether CCSP deficiency results in altered lung gene expression, high-density cDNA microarrays were used to profile gene expression in the total lung RNA of wild-type and CCSP(-/-) mice. Genes that were differentially expressed between wild-type and CCSP(-/-) mice included a previously non-annotated expressed sequence tag (EST W82219) and immunoglobulin A (IgA), both of which were elevated with CCSP deficiency. mRNA expression of EST W82219 and IgA was localized in the lungs of wild-type and CCSP(-/-) mice to airway Clara cells and peribronchial lymphoid tissues, respectively. We conclude that CCSP deficiency is associated with 1) altered gene expression in Clara cells of the conducting airway epithelium and 2) alterations to peribronchial B lymphocytes. These findings identify new roles for Clara cells and their secretions in airway homeostasis.

Animals↗

Targeted overexpression of the Escherichia coli MinC protein in higher plants results in abnormal chloroplasts.

Higher plant chloroplast division involves some of the same types of proteins that are required in prokaryotic cell division. These include two of the three Min proteins, MinD and MinE, encoded by the min operon in bacteria. Noticeably absent from annotated sequences from higher plants is a MinC homologue. A higher plant functional MinC homologue that would interfere with FtsZ polymerization, has yet to be identified. We sought to determine whether expression of the bacterial MinC in higher plants could affect chloroplast division. The Escherichia coli minC (EcMinC) gene was isolated and inserted behind the Arabidopsis thaliana RbcS transit peptide sequence for chloroplast targeting. This TP-EcMinC gene driven by the CaMV 35S(2) constitutive promoter was then transformed into tobacco (Nicotiana tabacum L.). Abnormally large chloroplasts were observed in the transgenic plants suggesting that overexpression of the E. coli MinC perturbed higher plant chloroplast division.

Chloroplasts↗

HemoPDB: Hematopoiesis Promoter Database, an information resource of transcriptional regulation in blood cell development.

Hematopoiesis describes the process of the normal formation and development of blood cells, involving both proliferation and differentiation from stem cells. Abnormalities in this developmental program yield blood cell diseases, such as leukemia. Although, in recent years, extensive molecular research in normal hematopoietic development has characterized transcription factors and their binding sites in the target gene promoters, the information generated is highly fragmented. In order to integrate this important regulatory information with the corresponding genomic sequences, we have developed a new database called Hematopoiesis Promoter Database (HemoPDB). HemoPDB is a comprehensive resource focused on transcriptional regulation during hematopoietic development and associated aberrances that result in malignancy. HemoPDB (version 1.0) contains 246 promoter sequences and 604 experimentally known cis-regulatory elements of 187 different transcription factors, with links to published references. Orthologous promoters from different species are linked with each other and displayed in the same database record, accompanied by a visual image of the promoters and corresponding annotations of cis-regulatory elements. HemoPDB may be searched for the promoter of a specific gene, transcription factors and target genes, and genes that are expressed in a certain cell type or lineage, through a user-friendly web interface at http://bioinformatics.med.ohio-state.edu/HemoPDB. Links to the documentation and other technical details are provided on this website.

Animals↗

Effect of ionizing irradiation on human esophageal cancer cell lines by cDNA microarray gene expression analysis.

To provide new insights into the molecular mechanisms underlying the effect of irradiation on esophageal squamous cell carcinomas (ESCCs), we used a cDNA microarray screening of more than 4,000 genes with known functions to identify genes involved in the early response to ionizing irradiation. Two human ESCC cell lines, one each of well (TE-1) and poorly (TE-2) differentiated phenotypes were screened. Subconfluent cells of each phenotype were treated with single doses of 2.0 Gy or 8.0 Gy irradiations. After a 15 min incubation time-point, the cells were collected and analyzed. Compared with non-irradiated cells, many genes revealed at least 2-fold upregulation or downregulation at both doses in well or poorly differentiated ESCC cells. The common upregulated genes in well and poorly differentiated cell types at both irradiation doses included SCYA5, CYP51, SMARCD2, COX6C, MAPK8, FOS, UBE2M, RPL6, PDGFRL, TRAF2, TNFAIP6, ITGB4, GSTM3, and SP3 and common downregulated genes involved NFIL3, SMARCA2, CAPZA1, MetAP2, CITED2, DAP3, MGAT2, ATRX, CIAO1, and STAT6. Several of these genes were novel and not previously known to be associated with irradiation. Functional annotations of the modulated genes suggested that at the molecular level, irradiation appears to induce a regularizing balance in ESCC cell function. The genes modulated in the early response to irradiation may be useful in our understanding of the molecular basis of radiotherapy and in developing strategies to augment its effect or establish novel less hazardous alternative adjuvant therapies.

Carcinoma, Squamous Cell↗

ChickGCE: a novel germ cell EST database for studying the early developmental stage in chickens.

We established a database to study germ cells during the early developmental stage in the chicken. The ChickGCE database provides integrated expressed sequence tag (EST) data from chicken testis, ovary, embryonic gonads, and primordial germ cells. We gathered data on 10,294 ESTs from approximately 1000 embryonic gonads, and we experimentally determined 10,851 ESTs from primordial germ cells purified from 7955 embryonic gonads by magnetically activated cell sorting. The EST testis and ovary datasets were retrieved from the public database of The Institute for Genomic Research (TIGR). The EST data were clustered and assembled into unique sequences, contigs, and singletons. The ChickGCE database provides functional annotation, identification, and putative embryonic germ-cell-specific novel transcripts based on the Gene Ontology database, as well as statistical analyses of expression patterns and pair-wise comparisons of two types of tissue- and germ-cell-specific alternative splicing events in the chicken. The new database is accessible online and queries can be answered using several search options, including tissue database searches, keywords, clone IDs, expected values, and BLAST search scores.

Animals↗

Mining gene expression data by interpreting principal components.

BACKGROUND: There are many methods for analyzing microarray data that group together genes having similar patterns of expression over all conditions tested. However, in many instances the biologically important goal is to identify relatively small sets of genes that share coherent expression across only some conditions, rather than all or most conditions as required in traditional clustering; e.g. genes that are highly up-regulated and/or down-regulated similarly across only a subset of conditions. Equally important is the need to learn which conditions are the decisive ones in forming such gene sets of interest, and how they relate to diverse conditional covariates, such as disease diagnosis or prognosis. RESULTS: We present a method for automatically identifying such candidate sets of biologically relevant genes using a combination of principal components analysis and information theoretic metrics. To enable easy use of our methods, we have developed a data analysis package that facilitates visualization and subsequent data mining of the independent sources of significant variation present in gene microarray expression datasets (or in any other similarly structured high-dimensional dataset). We applied these tools to two public datasets, and highlight sets of genes most affected by specific subsets of conditions (e.g. tissues, treatments, samples, etc.). Statistically significant associations for highlighted gene sets were shown via global analysis for Gene Ontology term enrichment. Together with covariate associations, the tool provides a basis for building testable hypotheses about the biological or experimental causes of observed variation. CONCLUSION: We provide an unsupervised data mining technique for diverse microarray expression datasets that is distinct from major methods now in routine use. In test uses, this method, based on publicly available gene annotations, appears to identify numerous sets of biologically relevant genes. It has proven especially valuable in instances where there are many diverse conditions (10's to hundreds of different tissues or cell types), a situation in which many clustering and ordering algorithms become problematic. This approach also shows promise in other topic domains such as multi-spectral imaging datasets.

Algorithms↗

Cell-specific DNA methylation in human alpha and beta cells regulates gene expression in type 2 diabetes.

Epigenome-wide studies of pancreatic islets provide valuable insights into type 2 diabetes (T2D) but lack methylomes from individual cell types. Here we show changes to alpha and beta cell-specific methylomes and transcriptomes from people with or without T2D, using whole-genome bisulfite sequencing and RNA sequencing. We discover 22,544 differentially methylated regions annotated to 7,975 genes in alpha versus beta cells, such as INS, GCG, PDX1 and PCSK1, with ~50% showing differential expression. CRISPR-dCas9-DNMT3A-based epigenetic editing increases INS and TH DNA methylation, while CRISPR-dCas9-TET1-based editing decreases GCG methylation, each altering INS, TH or GCG expression and content in beta cells. Pre-T2D/T2D-associated differentially methylated regions in alpha and beta cells overlap 12-18% of T2D-associated genome-wide association study candidates. Additionally, ONECUT2 is epigenetically upregulated in beta cells from people with pre-T2D/T2D and elevated in male Goto-Kakizaki rat islets. ONECUT2 overexpression in beta cells/islets downregulates gene sets impacting insulin secretion and glucose homeostasis, and reduces mitochondrial activity, ATP/ADP ratio and insulin secretion. We also provide 'alpha-beta-methylome' ( https://alpha-beta-methylome.serve.scilifelab.se/app/alpha-beta-methylome/ ), a resource exploring T2D, age and sex associations on methylation, highlighting cell-specific epigenetic regulation and dysfunctions contributing to T2D.

Humans↗

Dynamic covariation between gene expression and proteome characteristics.

BACKGROUND: Cells react to changing intra- and extracellular signals by dynamically modulating complex biochemical networks. Cellular responses to extracellular signals lead to changes in gene and protein expression. Since the majority of genes encode proteins, we investigated possible correlations between protein parameters and gene expression patterns to identify proteome-wide characteristics indicative of trends common to expressed proteins. RESULTS: Numerous bioinformatics methods were used to filter and merge information regarding gene and protein annotations. A new statistical time point-oriented analysis was developed for the study of dynamic correlations in large time series data. The method was applied to investigate microarray datasets for different cell types, organisms and processes, including human B and T cell stimulation, Drosophila melanogaster life span, and Saccharomyces cerevisiae cell cycle. CONCLUSION: We show that the properties of proteins synthesized correlate dynamically with the gene expression profile, indicating that not only is the actual identity and function of expressed proteins important for cellular responses but that several physicochemical and other protein properties correlate with gene expression as well. Gene expression correlates strongly with amino acid composition, composition- and sequence-derived variables, functional, structural, localization and gene ontology parameters. Thus, our results suggest that a dynamic relationship exists between proteome properties and gene expression in many biological systems, and therefore this relationship is fundamental to understanding cellular mechanisms in health and disease.

Animals↗

Identifying active transcription factors and kinases from expression data using pathway queries.

MOTIVATION: Although progress has been made identifying regulatory relationships from expression data in general, only few methods have focused on detecting biological mechanisms like active pathways using a single measurement. This is of particular importance when only few measurements are available, e.g. if special cell types or conditions are under investigation. Here we present a method to test user specified hypotheses (pathway queries) on expression data where prior knowledge is given in the form of networks and functional annotations. Based on this method, we develop a scoring function to identify active transcription factors or kinases, thus making a first step toward explaining the measured expression data. RESULTS: We apply the algorithm to the Rosetta Yeast Compendium dataset, finding that in many cases the results are in concordance with biological knowledge. We were able to confirm that transcription factors and to a lesser degree, kinases identified by our method play an important role in the biological processes affected by the respective knockouts. Furthermore, we show that correlation of inferred activities can provide evidence for a physical interaction or cooperation of transcription factors where correlation of plain expression data fails to do so.

Algorithms↗

Iron-related transcriptomic variations in CaCo-2 cells, an in vitro model of intestinal absorptive cells.

Regulation of iron absorption by duodenal enterocytes is essential for the maintenance of homeostasis by preventing iron deficiency or overload. Despite the identification of a number of genes implicated in iron absorption and its regulation, it is likely that further factors remain to be identified. For that purpose, we used a global transcriptomic approach, using the CaCo-2 cell line as an in vitro model of intestinal absorptive cells. Pangenomic screening for variations in gene expression correlating with intracellular iron content allowed us to identify 171 genes. One hundred nine of these genes are clustered into five types of expression profile. This is the first time that most of these genes have been associated with iron metabolism. Functional annotation of these five clusters indicates potential links between the immune response, proteolysis processes, and iron depletion. In contrast, iron overload is associated with cellular metabolism, especially that of lipids and glutathione involving redox function and electron transfer.

Caco-2 Cells↗

CMAtlas: a comprehensive DNA methylation atlas for exploring epigenetic alterations in 34 human cancer types.

MOTIVATION: Aberrant DNA methylation is a fundamental epigenetic hallmark of cancer. However, existing resources often lack technological diversity and comprehensive cancer coverage. Furthermore, most platforms fail to achieve deep multi-omics integration and tend to ignore cancer-type-specific methylation features, limiting their utility in precision oncology and drug discovery. RESULTS: We developed Cancer Methylation Atlas (CMAtlas), a comprehensive platform integrating 13 753 samples across 34 cancer types. By applying technology-tailored pipelines to data from various profiling technologies, we identified 830 725 tumor-specific differentially methylated elements (DMEs) and 1 480 098 differentially methylated regions (DMRs), alongside 1 154 256 cancer-type-specific DMEs and 329 154 DMRs. The platform demonstrates high cross-platform consistency and strong concordance between tumor tissues and cell lines, ensuring the robustness of our findings. All DMEs and DMRs are annotated with multi-omics data (RNA expression, somatic mutations, and chromatin accessibility) and clinical relevance (survival associations and cell-free DNA profiling). We further demonstrate the utility of CMAtlas by identifying prognostic aberrant methylation in colorectal cancer driver genes. AVAILABILITY AND IMPLEMENTATION: CMAtlas is freely accessible at {{https://cmatlas.renlab.cn/}}. The platform offers an intuitive web interface supporting gene-centric and cancer-centric queries, alongside customizable analysis modules designed to facilitate user-specific research needs.

Humans↗

Assessment of differential gene expression in vestibular epithelial cell types using microarray analysis.

Current global gene expression techniques allow the evaluation and comparison of the expression of thousands of genes in a single experiment, providing a tremendous amount of information. However, the data generated by these techniques are context-dependent, and minor differences in the individual biological samples, methodologies for RNA acquisition, amplification, hybridization protocol and gene chip preparation, as well as hardware and analysis software, lead to poor correlation between the results. One of the significant difficulties presently faced is the standardization of the protocols for the meaningful comparison of results. In the inner ear, the acquisition of RNA from individual cell populations remains a challenge due to the high density of the different cell types and the paucity of tissue. Consequently, laser capture microdissection was used to selectively collect individual cells and regions of cells from cristae ampullares followed by extraction of total RNA and amplification to amounts sufficient for high throughput analysis. To demonstrate hair cell-specific gene expression, myosin VIIA, calmodulin and alpha9 nicotinic acetylcholine receptor subunit mRNAs were amplified using reverse transcription-polymerase chain reaction (RT-PCR). To demonstrate supporting cell-specific gene expression, cyclin-dependent kinase inhibitor p27kip1 mRNA was amplified using RT-PCR. Subsequent experiments with alpha9 RT-PCR demonstrated phenotypic differences between type I and type II hair cells, with expression only in type II hair cells. Using the laser capture microdissection technique, microarray expression profiling demonstrated 408 genes with more than a five-fold difference in expression between the hair cells and supporting cells, of these 175 were well annotated. There were 97 annotated genes with greater than a five-fold expression difference in the hair cells relative to the supporting cells, and 78 annotated genes with greater than a five-fold expression difference in the supporting cells relative to the hair cells.

Acoustic Maculae↗

Gene discovery and gene expression in the rice blast fungus, Magnaporthe grisea: analysis of expressed sequence tags.

Over 28,000 expressed sequence tags (ESTs) were produced from cDNA libraries representing a variety of growth conditions and cell types. Several Magnaporthe grisea strains were used to produce the libraries, including a nonpathogenic strain bearing a mutation in the PMK1 mitogen-activated protein kinase. Approximately 23,000 of the ESTs could be clustered into 3,050 contigs, leaving 5,127 singleton sequences. The estimate of 8,177 unique sequences indicates that over half of the genes of the fungus are represented in the ESTs. Analysis of EST frequency reveals growth and cell type-specific patterns of gene expression. This analysis establishes criteria for identification of fungal genes involved in pathogenesis. A large fraction of the genes represented by ESTs have no known function or described homologs. Manual annotation of the most abundant cDNAs with no known homologs allowed us to identify a family of metallothionein proteins present in M. grisea, Neurospora crassa, and Fusarium graminearum. In addition, multiply represented ESTs permitted the identification of alternatively spliced mRNA species. Alternative splicing was rare, and in most cases, the alternate mRNA forms were unspliced, although alternative 5' splice sites were also observed.

Expressed Sequence Tags↗