PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Cell type annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Primary and secondary metabolism, and post-translational protein modifications, as portrayed by proteomic analysis of Streptomyces coelicolor.

The newly sequenced genome of Streptomyces coelicolor is estimated to encode 7825 theoretical proteins. We have mapped approximately 10% of the theoretical proteome experimentally using two-dimensional gel electrophoresis and matrix-assisted laser desorption ionization time-of-flight (MALDI-TOF) mass spectrometry. Products from 770 different genes were identified, and the types of proteins represented are discussed in terms of their annotated functional classes. An average of 1.2 proteins per gene was observed, indicating extensive post-translational regulation. Examples of modification by N-acetylation, adenylylation and proteolytic processing were characterized using mass spectrometry. Proteins from both primary and certain secondary metabolic pathways are strongly represented on the map, and a number of these enzymes were identified at more than one two-dimensional gel location. Post-translational modification mechanisms may therefore play a significant role in the regulation of these pathways. Unexpectedly, one of the enzymes for synthesis of the actinorhodin polyketide antibiotic appears to be located outside the cytoplasmic compartment, within the cell wall matrix. Of 20 gene clusters encoding enzymes characteristic of secondary metabolism, eight are represented on the proteome map, including three that specify the production of novel metabolites. This information will be valuable in the characterization of the new metabolites.

Acetylation↗

Whole-genome analysis of Brevibacterium sanguinis AZMABM HM27: a bacterial isolate from the sea anemone Radianthus magnifica and exhibiting promising multi-therapeutic properties.

BACKGROUND: The marine anemone Radianthus magnifica harbors symbiotic microbes with promising biomedical potential, yet their diversity and therapeutic properties remain underexplored. This study aimed to characterize a symbiotic bacterium isolated from R. magnifica collected from Samalona Island, Indonesia, and to evaluate its multi-therapeutic potential. METHODS: Strain AZMABM HM27 was characterized using whole-genome sequencing, functional annotation, biosynthetic gene cluster prediction, molecular docking, and in vitro bioactivity assays. RESULTS: Phylogenetic and genome-based analyses confirmed AZMABM HM27 as Brevibacterium sanguinis, with an OrthoANI value of 97.37% and a dDDH value of 76.50% against the type strain. The genome comprises a 3,834,082 bp chromosome encoding 3,362 protein-coding genes, including 95 genes involved in secondary metabolite biosynthesis. Five biosynthetic gene clusters were predicted, including those associated with ectoine, terpene, and siderophore production. The crude extract demonstrated antioxidant activity (IC₅₀ = 0.87 mg/mL), anti-inflammatory activity (up to 60% inhibition), antidiabetic activity through α-glucosidase inhibition (up to 40% inhibition), and dose-dependent antiproliferative activity against MCF-7 breast cancer cells (74.10% viability at 1 mg/mL). Molecular docking identified a lead compound, 8,9,9,10,10,11-hexafluoro-4,4-dimethyl-3,5-dioxatetracyclo [5.4.1.0(2,6)0.0(8,11)] dodecane, with strong binding affinities to selected therapeutic targets. CONCLUSIONS: B. sanguinis AZMABM HM27 represents a marine symbiotic strain associated with R. magnifica and a promising source of bioactive compounds with antioxidant, anti-inflammatory, antidiabetic, and antiproliferative potential. Further purification, structural elucidation, and in vivo studies are warranted to validate its therapeutic potential.

Animals↗

Proteome analysis of serovars Typhimurium and Pullorum of Salmonella enterica subspecies I.

BACKGROUND: Salmonella enterica subspecies I includes several closely related serovars which differ in host ranges and ability to cause disease. The basis for the diversity in host range and pathogenic potential of the serovars is not well understood, and it is not known how host-restricted variants appeared and what factors were lost or acquired during adaptations to a specific environment. Differences apparent from the genomic data do not necessarily correspond to functional proteins and more importantly differential regulation of otherwise identical gene content may play a role in the diverse phenotypes of the serovars of Salmonella. RESULTS: In this study a comparative analysis of the cytosolic proteins of serovars Typhimurium and Pullorum was performed using two-dimensional gel electrophoresis and the proteins of interest were identified using mass spectrometry. An annotated reference map was created for serovar Typhimurium containing 233 entries, which included many metabolic enzymes, ribosomal proteins, chaperones and many other proteins characteristic for the growing cell. The comparative analysis of the two serovars revealed a high degree of variation amongst isolates obtained from different sources and, in some cases, the variation was greater between isolates of the same serovar than between isolates with different sero-specificity. However, several serovar-specific proteins, including intermediates in sulphate utilisation and cysteine synthesis, were also found despite the fact that the genes encoding those proteins are present in the genomes of both serovars. CONCLUSION: Current microbial proteomics are generally based on the use of a single reference or type strain of a species. This study has shown the importance of incorporating a large number of strains of a species, as the diversity of the proteome in the microbial population appears to be significantly greater than expected. The characterisation of a diverse selection of strains revealed parts of the proteome of S. enterica that alter their expression while others remain stable and allowed for the identification of serovar-specific factors that have so far remained undetected by other methods.

Bacterial Proteins↗

Comparative analysis of eukaryotic-type protein phosphatases in two streptomycete genomes.

Inspection of the genomes of Streptomyces coelicolor A3(2) and Streptomyces avermitilis reveals that each contains 55 putative eukaryotic-type protein phosphatases (PPs), the largest number ever identified from any single prokaryotic organism. Unlike most other prokaryotic genomes that have only one or two superfamilies of eukaryotic-type PPs, the streptomycete genomes possess the eukaryotic-type PPs that belong to four superfamilies: 2 phosphoprotein phosphatases and 2 low-molecular-mass protein tyrosine phosphatases in each species, 49 Mg(2+)- or Mn(2+)-dependent protein phosphatases (PPMs) and 2 conventional protein tyrosine phosphatases (CPTPs) in S. coelicolor A3(2), and 48 PPMs and 3 CPTPs in S. avermitilis. Sixty-four percent of the PPs found in S. coelicolor A3(2) have orthologues in S. avermitilis, indicating that they originated from a common ancestor and might be involved in the regulation of more conserved metabolic activities. The genes of eukaryotic-type PP unique to each surveyed streptomycete genome are mainly located in two arms of the linear chromosomes and their evolution might be involved in gene acquisition or duplication to adapt to the extremely variable soil environments where these organisms live. In addition, 56 % of the PPs from S. coelicolor A3(2) and 65 % of the PPs from S. avermitilis possess at least one additional domain having a putative biological function. These include the domains involved in the detection of redox potential, the binding of cyclic nucleotides, mRNA, DNA and ATP, and the catalysis of phosphorylation reactions. Because they contained multiple functional domains, most of them were assigned functions other than PPs in previous annotations. Although few studies have been conducted on the physiological functions of the PPs in streptomycetes, the existence of large numbers of putative PPs in these two streptomycete genomes strongly suggests that eukaryotic-type PPs play important regulatory roles in primary or secondary metabolic pathways. The identification and analysis of such a large number of putative eukaryotic-type PPs from S. coelicolor A3(2) and S. avermitilis constitute a basis for further exploration of the signal transduction pathways mediated by these phosphatases in industrially important strains of streptomycetes.

Computational Biology↗

Reannotation of the CELO genome characterizes a set of previously unassigned open reading frames and points to novel modes of host interaction in avian adenoviruses.

BACKGROUND: The genome of the avian adenovirus Chicken Embryo Lethal Orphan (CELO) has two terminal regions without detectable homology in mammalian adenoviruses that are left without annotation in the initial analysis. Since adenoviruses have been a rich source of new insights into molecular cell biology and practical applications of CELO as gene a delivery vector are being considered, this genome appeared worth revisiting. We conducted a systematic reannotation and in-depth sequence analysis of the CELO genome. RESULTS: We describe a strongly diverged paralogous cluster including ORF-2, ORF-12, ORF-13, and ORF-14 with an ATPase/helicase domain most likely acquired from adeno-associated parvoviruses. None of these ORFs appear to have retained ATPase/helicase function and alternative functions (e.g. modulation of gene expression during the early life-cycle) must be considered in an adenoviral context. Further, we identified a cluster of three putative type-1-transmembrane glycoproteins with IG-like domains (ORF-9, ORF-10, ORF-11) which are good candidates to substitute for the missing immunomodulatory functions of mammalian adenoviruses. ORF-16 (located directly adjacent) displays distant homology to vertebrate mono-ADP-ribosyltransferases. Members of this family are known to be involved in immuno-regulation and similiar functions during CELO life cycle can be considered for this ORF. Finally, we describe a putative triglyceride lipase (merged ORF-18/19) with additional domains, which can be expected to have specific roles during the infection of birds, since they are unique to avian adenoviruses and Marek's disease-like viruses, a group of pathogenic avian herpesviruses. CONCLUSIONS: We could characterize most of the previously unassigned ORFs pointing to functions in host-virus interaction. The results provide new directives for rationally designed experiments.

ADP Ribose Transferases↗

Cloning and characterization of multiple glycosyl hydrolase genes from Trichoderma virens.

Trichoderma virens is a widely distributed soil fungus that is parasitic on other soil fungi. The mycoparasitic activity of T. virens is correlated with the production of numerous antifungal activities, including the secretion of a considerable repertoire of fungal cell wall-degrading enzymes. Here, we report the characterization of a diverse set of chitinase and glucanase genes from T. virens. In each case, full-length genomic clones were isolated and characterized, while sequencing of the corresponding cDNA clones and manual annotation provided a basis for establishing gene structure. Based on homology of the deduced amino acid sequences, we have identified three members of the 42Kd endochitinase gene family, two 33Kd exochitinases, two exochitinases with homology to N-acetylglucosaminidases, and three glucanase genes predicted to encode beta-1,3- and beta-1,6-proteins. The majority of these genes appear to encode signal peptides, suggesting an extracellular location for the corresponding proteins. Despite their overall similarity, the 42Kd class of chitinases can be subdivided, based on the presence of distinct N-terminal domains, suggesting that the proteins may have distinct cellular roles, while Northern blot analysis confirms that these genes possess distinct patterns of gene regulation. Similarly, one of the 33Kd chitinase genes is unique, because it is predicted to encode a protein C-terminus with high homology to the conserved family I cellulose-binding domain. The expression patterns of the chitinase genes were analyzed in both a wild-type strain and a strain disrupted for the major 42Kd chitinase gene of T. virens. The results of these transcript analyses, together with enzymatic assay of the extracellular proteins, suggest interdependent regulation of this important gene family in T. virens.

Acetylglucosaminidase↗

Purification and characterization of ferredoxin-NADP+ reductase encoded by Bacillus subtilis yumC.

From Bacillus subtilis cell extracts, ferredoxin-NADP+ reductase (FNR) was purified to homogeneity and found to be the yumC gene product by N-terminal amino acid sequencing. YumC is a approximately 94-kDa homodimeric protein with one molecule of non-covalently bound FAD per subunit. In a diaphorase assay with 2,6-dichlorophenol-indophenol as electron acceptor, the affinity for NADPH was much higher than that for NADH, with Km values of 0.57 microM vs >200 microM. Kcat values of YumC with NADPH were 22.7 s(-1) and 35.4 s(-1) in diaphorase and in a ferredoxin-dependent NADPH-cytochrome c reduction assay, respectively. The cell extracts contained another diaphorase-active enzyme, the yfkO gene product, but its affinity for ferredoxin was very low. The deduced YumC amino acid sequence has high identity to that of the recently identified Chlorobium tepidum FNR. A genomic database search indicated that there are more than 20 genes encoding proteins that share a high level of amino acid sequence identity with YumC and which have been annotated variously as NADH oxidase, thioredoxin reductase, thioredoxin reductase-like protein, etc. These genes are found notably in gram-positive bacteria, except Clostridia, and less frequently in archaea and proteobacteria. We propose that YumC and C. tepidum FNR constitute a new group of FNR that should be added to the already established plant-type, bacteria-type, and mitochondria-type FNR groups.

Bacillus subtilis↗

Detection of a luxS-signaling molecule in Bacillus anthracis.

Quorum-sensing regulation of density-dependent genes has been described for numerous bacterial species. The partially annotated genome sequence of Bacillus anthracis contains an open reading frame (BA5047) predicted to encode an ortholog of luxS, required for synthesis of the quorum-sensing signaling molecule autoinducer-2 (AI-2). To determine whether B. anthracis produces AI-2, the Vibrio harveyi luminescence bioassay was used. Cell-free conditioned media from vaccine (Sterne) strain 34F(2) induced luminescence in V. harveyi reporter strain BB170, indicating its production of AI-2. Cloned BA5047, expressed in Escherichia coli DH5 alpha cells, restored AI-2 activity to these cells. To evaluate whether BA5047 is essential for AI-2 synthesis, it was deleted through allelic exchange with marker rescue; the resulting mutant had no functional luxS activity and had reduced growth in vitro. In the wild-type strain, AI-2 activity was greatest during the exponential phase of growth. In total, these data indicate that BA5047 is a functional luxS ortholog in B. anthracis necessary for growth-phase-specific AI-2 expression. Thus, B. anthracis may utilize extracellular signaling molecules to regulate density-dependent gene expression.

Bacillus anthracis↗

A single-cell transcriptomic atlas of the pigtail macaque placenta in late gestation.

The placenta is a complex organ with multiple immune and non-immune cell types that promote fetal tolerance and facilitate the transfer of nutrients and oxygen. The nonhuman primate (NHP) is a key experimental model for studying human pregnancy complications, in part due to similarities in placental structure, which makes it essential to understand how single-cell populations compare across the human and NHP maternal-fetal interface. We constructed a single-cell RNA-Seq (scRNA-Seq) atlas of the placenta from the pigtail macaque ( Macaca nemestrina ) in the third trimester, comprising three different tissues at the maternal-fetal interface: the chorionic villi (placental disc), chorioamniotic membranes, and the maternal decidua. Each tissue was separately dissociated into single cells and processed through the 10X Genomics and Seurat pipeline, followed by aggregation, unsupervised clustering, and cluster annotation. Next, we determined the maternal-fetal origins of cell populations and analyzed single-cell RNA trajectory, Gene Ontology enrichment, and cell-cell communication. Single-cell populations in the pigtail macaque were strikingly similar in their identity and frequency to those found in the human placenta, including cells from trophoblast, stromal cell, immune, and macrophage lineages. An advantage of our approach was the deep sequencing of three tissues at the maternal-fetal interface, which yielded a rich diversity of common and rare single-cell populations. The third-trimester pigtail macaque single-cell atlas enables the identification of cellular subclusters analogous to those in humans and provides a powerful resource for understanding experimental perturbations on the NHP placenta.

Journal Article↗

Characterization of AtCHX17, a member of the cation/H+ exchangers, CHX family, from Arabidopsis thaliana suggests a role in K+ homeostasis.

The Arabidopsis genome contains many sequences annotated as encoding H(+)-coupled cotransporters. Among those are the members of the cation:proton antiporter-2 (CPA2) family (or CHX family), predicted to encode Na(+),K(+)/H(+) antiporters. AtCHX17, a member of the CPA2 family, was selected for expression studies, and phenotypic analysis of knockout mutants was performed. AtCHX17 expression was only detected in roots. The gene was strongly induced by salt stress, potassium starvation, abscisic acid (ABA) and external acidic pH. Using the beta-glucuronidase reporter gene strategy and in situ RT-PCR experiments, we have found that AtCHX17 was expressed preferentially in epidermal and cortical cells of the mature root zones. Knockout mutants accumulated less K(+) in roots in response to salt stress and potassium starvation compared with the wild type. These data support the hypothesis that AtCHX17 is involved in K(+) acquisition and homeostasis.

Amino Acid Sequence↗

Analysis of the interaction of extracellular matrix and phenotype of bladder cancer cells.

BACKGROUND: The extracellular matrix has a major effect upon the malignant properties of bladder cancer cells both in vitro in 3-dimensional culture and in vivo. Comparing gene expression of several bladder cancer cells lines grown under permissive and suppressive conditions in 3-dimensional growth on cancer-derived and normal-derived basement membrane gels respectively and on plastic in conventional tissue culture provides a model system for investigating the interaction of malignancy and extracellular matrix. Understanding how the extracellular matrix affects the phenotype of bladder cancer cells may provide important clues to identify new markers or targets for therapy. METHODS: Five bladder cancer cell lines and one immortalized, but non-tumorigenic, urothelial line were grown on Matrigel, a cancer-derived ECM, on SISgel, a normal-derived ECM, and on plastic, where the only ECM is derived from the cells themselves. The transcriptomes were analyzed on an array of 1186 well-annotated cancer derived cDNAs containing most of the major pathways for malignancy. Hypervariable genes expressing more variability across cell lines than a set expressing technical variability were analyzed further. Expression values were clustered, and to identify genes most likely to represent biological factors, statistically over-represented ontologies and transcriptional regulatory elements were identified. RESULTS: Approximately 400 of the 1186 total genes were expressed 2 SD above background. Approximately 100 genes were hypervariable in cells grown on each ECM, but the pattern was different in each case. A core of 20 were identified as hypervariable under all 3 growth conditions, and 33 were hypervariable on both SISgel and Matrigel, but not on plastic. Clustering of the hypervariable genes showed very different patterns for the same 6 cell types on the different ECM. Even when loss of cell cycle regulation was identified, different genes were involved, depending on the ECM. Under the most permissive conditions of growth where the malignant phenotype was fully expressed, activation of AKT was noted. TGFbeta1 signaling played a major role in the response of bladder cancer cells to ECM. Identification of TREs on genes that clustered together suggested some clustering was driven by specific transcription factors. CONCLUSION: The extracellular matrix on which cancer cells are grown has a major effect on gene expression. A core of 20 malignancy-related genes were not affected by matrix, and 33 were differentially expressed on 3-dimensional culture as opposed to plastic. Other than these genes, the patterns of expression were very different in cells grown on SISgel than on Matrigel or even plastic, supporting the hypothesis that growth of bladder cancer cells on normal matrix suppresses some malignant functions. Unique underlying regulatory networks were driving gene expression and could be identified by the approach outlined here.

Cell Line, Tumor↗

Causal circuit tracing reveals distinct computational architectures in single-cell foundation models: inhibitory dominance, biological coherence, and cross-model convergence.

MOTIVATION: Sparse autoencoders (SAEs) decompose foundation-model activations into interpretable features, but the model-internal causal interactions between those features (i.e. what ablating one feature does to the others, as distinct from the biological causal structure of the underlying cells)-and how those model-internal relationships relate to biological structure-are uncharacterized in single-cell foundation models. RESULTS: We introduce model-internal causal circuit tracing-zeroing one SAE feature at a source layer and measuring the resulting change in all downstream SAE features, for each of 120 source features-and apply it to Geneformer V2-316M and scGPT whole-human across four conditions (96&#xa0;892 ablation-derived edges, 80&#xa0;191 forward passes). On annotation-selected source features, edges share GO/KEGG/Reactome/STRING/TRRUST ontology terms at 50.9%-68.5%, a 2.9-6.2&#xd7; enrichment over a configuration-preserving permutation null (P<.002); on 20 randomly sampled source features this attenuates to 21.5%-26.3%-still 2.5-3.1&#xd7; above null-quantifying the annotation-selection contribution. Inhibitory dominance (fraction of ablation edges with d<0, i.e. source activation supports downstream target) is 65.5%-89.4%. scGPT produces larger raw per-edge effects (mean |d|=1.40 versus 1.05); after feature-share normalization, Geneformer is stronger (paired gene-pair ratio 0.64 on 33&#xa0;301 shared pairs). Cross-model consensus yields 1142 architecture-invariant domain pairs (ordered pairs of GO biological-process categories "A&#x2192;B" each connected by at least one ablation edge in both models; 10.6&#xd7; enrichment over permutation null; P<.001). Circuit edge magnitude explains <1% of the variance in marginal driver-gene coexpression on the same cells (R2=0.010, n=31&#xa0;176): the graph encodes structure beyond bivariate correlation. Against a matched-cell-type ENCODE ChIP-seq prior, circuit-predicted transcription factor (TF)&#x2192;target pairs are enriched 2.06&#xd7; (Fisher OR 5.84), markedly higher than 1.12&#xd7; against TRRUST; direct ChIP-seq-supported target pairs show 10-30&#xd7; larger CRISPRi sign-bias-corrected excess than indirect pairs. Gene-level CRISPRi validation on Replogle K562 and the noncancer RPE1 arm (and a true primary-T-cell control from Shifrut E, Carnevale J, Tobin V et&#xa0;al. Genome-wide CRISPR screens in primary human T cells reveal key regulators of immune function. Cell 2018; 175: 1958-71.e15) after sign-bias correction shows excess over baseline of +0.03 and +0.35 percentage points on K562 and RPE1, respectively (baseline already 52%-56% from sign marginals); effect-magnitude Spearman correlations &#x3c1;&#x2248;0. Bootstrap and per-cell-type stability (N&#x2208;{50,100,200}; B cell, CD4&#xa0;+ T, macrophage) give Pearson r&#x2265;0.97 on shared edges with 100% sign agreement; edge Jaccard grows monotonically with sample size. The circuit graph is therefore highly reproducible as an effect-size map, cell type specific in edge identity, consistent with coexpression encoding, and weakly but detectably enriched for ChIP-seq-supported direct regulatory edges. AVAILABILITY AND IMPLEMENTATION: https://github.com/Biodyn-AI/bio-sae-circuits (Python). Archival DOI: 10.5281/zenodo.19,633,166 (Zenodo).

Humans↗

A focused microarray approach to functional glycomics: transcriptional regulation of the glycome.

Glycosylation is the most common posttranslational modification of proteins, yet genes relevant to the synthesis of glycan structures and function are incompletely represented and poorly annotated on the commercially available arrays. To fill the need for expression analysis of such genes, we employed the Affymetrix technology to develop a focused and highly annotated glycogene-chip representing human and murine glycogenes, including glycosyltransferases, nucleotide sugar transporters, glycosidases, proteoglycans, and glycan-binding proteins. In this report, the array has been used to generate glycogene-expression profiles of nine murine tissues. Global analysis with a hierarchical clustering algorithm reveals that expression profiles in immune tissues (thymus [THY], spleen [SPL], lymph node, and bone marrow [BM]) are more closely related, relative to those of nonimmune tissues (kidney [KID], liver [LIV], brain [BRN], and testes [TES]). Of the biosynthetic enzymes, those responsible for synthesis of the core regions of N- and O-linked oligosaccharides are ubiquitously expressed, whereas glycosyltransferases that elaborate terminal structures are expressed in a highly tissue-specific manner, accounting for tissue and ultimately cell-type-specific glycosylation. Comparison of gene expression profiles with matrix-assisted laser desorption ionization-time of flight (MALDI-TOF) profiling of N-linked oligosaccharides suggested that the alpha1-3 fucosyltransferase 9, Fut9, is the enzyme responsible for terminal fucosylation in KID and BRN, a finding validated by analysis of Fut9 knockout mice. Two families of glycan-binding proteins, C-type lectins and Siglecs, are predominately expressed in the immune tissues, consistent with their emerging functions in both innate and acquired immunity. The glycogene chip reported in this study is available to the scientific community through the Consortium for Functional Glycomics (CFG) (http://www.functionalglycomics.org).

Animals↗

Complexities in ETS-domain transcription factor function and regulation: lessons from the TCF (ternary complex factor) subfamily. The Colworth Medal Lecture.

The ETS-domain transcription factor family can be divided into a series of subfamilies. Elk-1 represents the founding member of the ternary complex factor (TCF) subfamily. By focusing on the TCF subfamily, we can demonstrate the complexities that exist in the function and regulation of ETS-domain transcription factors. This article focuses on Elk-1 in detail and summarizes the functions of other TCFs. The key themes covered include the domain structure of the TCFs, the mechanisms of complex formation with serum response factor, regulation of TCFs by mitogen-activated protein kinase cascades, and transcriptional regulatory properties of the TCFs. Finally, the emerging role of the TCFs in vivo is discussed. A picture is developing indicating that, while these proteins exhibit significant sequence and functional conservation, key differences in their structure and regulation are being identified which may relate to unique functions of these proteins in vivo.

Amino Acid Sequence↗

Annotation of glycoproteins in the SWISS-PROT database.

SWISS-PROT is a protein sequence database, which aims to be nonredundant, fully annotated and highly cross-referenced. Most eukaryotic gene products undergo co- and/or post-translational modifications, and these need to be included in the database in order to describe the mature protein. SWISS-PROT includes information on many types of different protein modifications. As glycosylation is the most common type of post-translational protein modification, we are currently placing an emphasis on annotation of protein glycosylation in SWISS-PROT. Information on the position of the sugar within the polypeptide chain, the reducing terminal linkage as well as additional information on biological function of the sugar is included in the database. In this paper we describe how we account for the different types of protein glycosylation, namely N-linked glycosylation, O-linked glycosylation, proteoglycans, C-linked glycosylation and the attachment of glycosyl-phosphatidylinosital anchors to proteins.

Amino Acid Sequence↗

Relationship of gene expression and chromosomal abnormalities in colorectal cancer.

Several studies have verified the existence of multiple chromosomal abnormalities in colon cancer. However, the relationships between DNA copy number and gene expression have not been adequately explored nor globally monitored during the progression of the disease. In this work, three types of array-generated data (expression, single nucleotide polymorphism, and comparative genomic hybridization) were collected from a large set of colon cancer patients at various stages of the disease. Probes were annotated to specific chromosomal locations and coordinated alterations in DNA copy number and transcription levels were revealed at specific positions. We show that across many large regions of the genome, changes in expression level are correlated with alterations in DNA content. Often, large chromosomal segments, containing multiple genes, are transcriptionally affected in a coordinated way, and we show that the underlying mechanism is a corresponding change in DNA content. This implies that whereas specific chromosomal abnormalities may arise stochastically, the associated changes in expression of some or all of the affected genes are responsible for selecting cells bearing these abnormalities for clonal expansion. Indeed, particular chromosomal regions are frequently gained and overexpressed (e.g., 7p, 8q, 13q, and 20q) or lost and underexpressed (e.g., 1p, 4, 5q, 8p, 14q, 15q, and 18) in primary colon tumors, making it likely that these changes favor tumorigenicity. Furthermore, we show that these aberrations are absent in normal colon mucosa, appear in benign adenomas (albeit only in a small fraction of the samples), become more frequent as disease advances, and are found in the majority of metastatic samples.

Adenoma↗

Shared genetic architecture and therapeutic targets across paediatric immune-mediated diseases.

OBJECTIVES: Paediatric-onset immune-mediated inflammatory diseases (IMIDs), including juvenile idiopathic arthritis and related rheumatic diseases, remain genetically undercharacterised. We aimed to define shared and category-specific genetic architecture across paediatric IMIDs, compare signals with adult IMIDs, and identify therapeutic opportunities. METHODS: We analysed 24 paediatric IMIDs classified as autoimmune, polygenic-autoinflammatory, mixed-pattern, or allergic. Genome-wide association analyses included 18,086 cases and 131,019 controls of European ancestry. We estimated single nucleotide polymorphism (SNP)-based heritability, genetic correlations, and polygenic overlap; performed subset-based meta-analysis; and conducted functional annotation, gene prioritisation, pathway and protein network analyses, adult-IMID comparison, and drug-target prioritisation. RESULTS: SNP-based heritability ranged from 28.9% for allergic IMIDs to 61.9% for autoimmune IMIDs. Genetic correlation and polygenic modelling supported partial sharing across categories with category-specific components. Meta-analysis identified 39 genome-wide significant loci outside the Major Histocompatibility Complex (MHC) region, including 15 previously unreported loci; 19 loci were shared between categories. Gene-prioritisation and protein interaction analyses identified a core MHC-centred antigen-presentation network, with category-enriched modules involving complement, innate/barrier pathways, epithelial biology, and type 2 immunity. Enriched pathways included nuclear factor &#x3ba;B signalling, T helper 17 related pathways, Janus kinase-signal transducer and activator of transcription signalling, programmed cell death protein 1/programmed death&#x2011;ligand 1, cytotoxic T&#x2011;lymphocyte associated protein 4 regulation, and osteoclast differentiation, several of which are relevant to rheumatic diseases. Paediatric IMIDs shared broad polygenic architecture with adult IMIDs, whereas top-ranked genes converged strongly with adult rheumatic diseases. Priority Index analysis identified 178 high-scoring genes, including 43 approved or investigational IMID drug targets. CONCLUSIONS: Paediatric-onset IMIDs share core pathways with adult forms but exhibit distinct genetic architecture shaped by age-specific immune and neurodevelopmental biology. These findings provide a genomic framework for paediatric precision medicine, guiding classification, risk prediction, and therapeutic development.

Humans↗

Comprehensive genome sequence analysis of a breast cancer amplicon.

Gene amplification occurs in most solid tumors and is associated with poor prognosis. Amplification of 20q13.2 is common to several tumor types including breast cancer. The 1 Mb of sequence spanning the 20q13.2 breast cancer amplicon is one of the most exhaustively studied segments of the human genome. These studies have included amplicon mapping by comparative genomic hybridization (CGH), fluorescent in-situ hybridization (FISH), array-CGH, quantitative microsatellite analysis (QUMA), and functional genomic studies. Together these studies revealed a complex amplicon structure suggesting the presence of at least two driver genes in some tumors. One of these, ZNF217, is capable of immortalizing human mammary epithelial cells (HMEC) when overexpressed. In addition, we now report the sequencing of this region in human and mouse, and on quantitative expression studies in tumors. Amplicon localization now is straightforward and the availability of human and mouse genomic sequence facilitates their functional analysis. However, comprehensive annotation of megabase-scale regions requires integration of vast amounts of information. We present a system for integrative analysis and demonstrate its utility on 1.2 Mb of sequence spanning the 20q13.2 breast cancer amplicon and 865 kb of syntenic murine sequence. We integrate tumor genome copy number measurements with exhaustive genome landscape mapping, showing that amplicon boundaries are associated with maxima in repetitive element density and a region of evolutionary instability. This integration of comprehensive sequence annotation, quantitative expression analysis, and tumor amplicon boundaries provide evidence for an additional driver gene prefoldin 4 (PFDN4), coregulated genes, conserved noncoding regions, and associate repetitive elements with regions of genomic instability at this locus.

Animals↗