PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Cell type annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

ProteoformDB: A Built-In Application to Generate Proteoform Database.

Proteins play essential functions through their complex regulations on cell-type-specific expression, localization, and molecular complexes. Protein complexity is further enhanced by proteoforms, which are the diverse molecular forms that each gene can produce through genomic alterations, transcriptional variations, translational regulations, and protein modifications. Profiling of proteoforms is a promising method for gaining a deeper understanding of the role of proteins in biological pathways and disease mechanisms. Here, we developed ProteoformDB, an application tool for generating proteoform databases, and we cataloged a total of over one million unique single-site human proteoforms. We showed that ProteoformDB can serve as a valuable resource to document the experimentally identified proteoforms in a database, supporting protein characterization in quantitative proteomics for both total protein abundances and modified protein forms.

Humans↗

A novel beta-glucanase gene from Bacillus halodurans C-125.

A novel endo-beta-1,3(4)-D-glucanase gene was found in the complete genome sequence of Bacillus halodurans C-125. The gene was previously annotated as an "unknown" protein and assigned an incorrect open reading frame (ORF). However, determining the biochemical characteristics has elucidated the function and correct ORF of the gene. The gene encodes 231 amino acids, and its calculated molecular mass was estimated to be 26743.16 Da. The amino acid sequence alignment showed that the highest sequence identity was only 28% with that of the beta-1,3-1,4-glucanase from Bacillus subtilis. Moreover, the nucleotide sequence did not match any other known Bacillus beta-glucanase gene. The member of the gene cluster that includes this novel gene was apparently different from that of the gene cluster including the putative beta-glucanase genes (bh3231 and bh3232) from B. halodurans C-125. Therefore, the novel gene is not a copy of either of these genes, and in B. halodurans cells, the putative role of the encoded protein may differ from that of bh3231 and bh3232. To examine the activity of the gene product, the gene was cloned as a His-tagged protein and expressed in Escherichia coli. The purified enzyme showed activity against lichenan, barley beta-glucan, laminarin, and carboxymethyl curdlan. Thin-layer chromatography showed that the enzyme hydrolyzes substrates in an endo-type manner. When beta-glucan was used as a substrate, the pH optimum was between 6 and 8, and the temperature optimum was 60 degrees C. After 2 h incubation at 50 and 60 degrees C, the residual activity remained 100% and 50%, respectively. The enzymatic activity was abolished after 30 min incubation at 70 degrees C. Based on the results, the gene encodes an endo-type beta-1,3(4)-D-glucanase (E.C. 3.2.1.6).

Amino Acid Sequence↗

Expression profiling in transformed human B cells: influence of Btk mutations and comparison to B cell lymphomas using filter and oligonucleotide arrays.

We have used both Clontech Atlas Human Hematology/Immunology cDNA microarrays, containing 588 genes, and Affymetrix oligonucleotide U95Av2 human array complementary to more than 12,500 genes to get a global view of genes expressed in Epstein-Barr virus (EBV)-transformed B cells and genes regulated by Bruton's tyrosine kinase (Btk). We compared EBV-transformed wild-type (WT) B cells from a healthy individual, WT1 and an X-linked agammaglobulinemia (XLA) patient cell line, XLA1, using the Clontech filters arrays. Eleven genes were > or =1.9-fold induced in absence of functional Btk. Furthermore, we analyzed a second patient cell line, XLA2, and compared this to two WT cell lines using oligonucleotide arrays. A total of 391 genes were found to be differentially expressed, including kinases and transcriptions factors. Furthermore, one expressed sequence tag and eight complementary DNA clones with unknown function were down-regulated in XLA2, indicating their biological role. Higher-fold inductions, Fyn (39.5), Hck (15.5) and Cyp1B1 (5.8), were observed using oligonucleotide array and were confirmed using real-time PCR for Fyn (20.8), Hck (6.7) and Cyp1B1 (10). Two genes, B cell translocation gene1 (BTG1) and B cell-specific OCT binding factor-1 (OBF-1) were induced > or =1.9-fold in both XLA1 and XLA2 analyzed by Atlas filter arrays andAffymetrix chips, respectively. Data from both filter and oligonucleotide arrays were compared to the gene clusters of a previously published lymphoma expression profile by linking to the UniGene transcript database. Our findings demonstrate for the first time the use of microarray to study the influence of Btk mutations and the use of functional annotation and validation of expression data by comparison of microarray analyses.

Agammaglobulinaemia Tyrosine Kinase↗

N-terminal N-myristoylation of proteins: prediction of substrate proteins from amino acid sequence.

Myristoylation by the myristoyl-CoA:protein N-myristoyltransferase (NMT) is an important lipid anchor modification of eukaryotic and viral proteins. Automated prediction of N-terminal N-myristoylation from the substrate protein sequence alone is necessary for large-scale sequence annotation projects but it requires a low rate of false positive hits in addition to a sufficient sensitivity. Our previous analysis of substrate protein sequence variability, NMT sequences and 3D structures has revealed motif properties in addition to the known PROSITE motif that are utilized in a new predictor described here. The composite prediction function (with separate ad hoc parameterization (a) for queries from non-fungal eukaryotes and their viruses and (b) for sequences from fungal species) consists of terms evaluating amino acid type preferences at sequences positions close to the N terminus as well as terms penalizing deviations from the physical property pattern of amino acid side-chains encoded in multi-residue correlation within the motif sequence. The algorithm has been validated with a self-consistency and two jack-knife tests for the learning set as well as with kinetic data for model substrates. The sensitivity in recognizing documented NMT substrates is above 95 % for both taxon-specific versions. The corresponding rate of false positive prediction (for sequences with an N-terminal glycine residue) is close to 0.5 %; thus, the technique is applicable for large-scale automated sequence database annotation. The predictor is available as public WWW-server with the URL http://mendel.imp.univie.ac.at/myristate/. Additionally, we propose a version of the predictor that identifies a number of proteolytic protein processing sites at internal glycine residues and that evaluates possible N-terminal myristoylation of the protein fragments.A scan of public protein databases revealed new potential NMT targets for which the myristoyl modification may be of critical importance for biological function. Among others, the list includes kinases, phosphatases, proteasomal regulatory subunit 4, kinase interacting proteins KIP1/KIP2, protozoan flagellar proteins, homologues of mitochondrial translocase TOM40, of the neuronal calcium sensor NCS-1 and of the cytochrome c-type heme lyase CCHL. Analyses of complete eukaryote genomes indicate that about 0.5 % of all encoded proteins are apparent NMT substrates except for a higher fraction in Arabidopsis thaliana ( approximately 0.8 %).

Acyltransferases↗

Annotated draft genomic sequence from a Streptococcus pneumoniae type 19F clinical isolate.

The public availability of numerous microbial genomes is enabling the analysis of bacterial biology in great detail and with an unprecedented, organism-wide and taxon-wide, broad scope. Streptococcus pneumoniae is one of the most important bacterial pathogens throughout the world. We present here sequences and functional annotations for 2.1-Mbp of pneumococcal DNA, covering more than 90% of the total estimated size of the genome. The sequenced strain is a clinical isolate resistant to macrolides and tetracycline. It carries a type 19F capsular locus, but multilocus sequence typing for several conserved genetic loci suggests that the strain sequenced belongs to a pneumococcal lineage that most often expresses a serotype 15 capsular polysaccharide. A total of 2,046 putative open reading frames (ORFs) longer than 100 amino acids were identified (average of 1,009 bp per ORF), including all described two-component systems and aminoacyl tRNA synthetases. Comparisons to other complete, or nearly complete, bacterial genomes were made and are presented in a graphical form for all the predicted proteins.

DNA, Bacterial↗

Molecular weight assessment of proteins in total proteome profiles using 1D-PAGE and LC/MS/MS.

BACKGROUND: The observed molecular weight of a protein on a 1D polyacrylamide gel can provide meaningful insight into its biological function. Differences between a protein's observed molecular weight and that predicted by its full length amino acid sequence can be the result of different types of post-translational events, such as alternative splicing (AS), endoproteolytic processing (EPP), and post-translational modifications (PTMs). The characterization of these events is one of the important goals of total proteome profiling (TPP). LC/MS/MS has emerged as one of the primary tools for TPP, but since this method identifies tryptic fragments of proteins, it has not generally been used for large-scale determination of the molecular weight of intact proteins in complex mixtures. RESULTS: We have developed a set of computational tools for extracting molecular weight information of intact proteins from total proteome profiles in a high throughput manner using 1D-PAGE and LC/MS/MS. We have applied this technology to the proteome profile of a human lymphoblastoid cell line under standard culture conditions. From a total of 1 x 10(7) cells, we identified 821 proteins by at least two tryptic peptides. Additionally, these 821 proteins are well-localized on the 1D-SDS gel. 656 proteins (80%) occur in gel slices in which the observed molecular weight of the protein is consistent with its predicted full-length sequence. A total of 165 proteins (20%) are observed to have molecular weights that differ from their predicted full-length sequence. We explore these molecular-weight differences based on existing protein annotation. CONCLUSION: We demonstrate that the determination of intact protein molecular weight can be achieved in a high-throughput manner using 1D-PAGE and LC/MS/MS. The ability to determine the molecular weight of intact proteins represents a further step in our ability to characterize gene expression at the protein level. The identification of 165 proteins whose observed molecular weight differs from the molecular weight of the predicted full-length sequence provides another entry point into the high-throughput characterization of protein modification.

Journal Article↗

Predicting eukaryotic protein subcellular location by fusing optimized evidence-theoretic K-Nearest Neighbor classifiers.

Facing the explosion of newly generated protein sequences in the post genomic era, we are challenged to develop an automated method for fast and reliably annotating their subcellular locations. Knowledge of subcellular locations of proteins can provide useful hints for revealing their functions and understanding how they interact with each other in cellular networking. Unfortunately, it is both expensive and time-consuming to determine the localization of an uncharacterized protein in a living cell purely based on experiments. To tackle the challenge, a novel hybridization classifier was developed by fusing many basic individual classifiers through a voting system. The "engine" of these basic classifiers was operated by the OET-KNN (Optimized Evidence-Theoretic K-Nearest Neighbor) rule. As a demonstration, predictions were performed with the fusion classifier for proteins among the following 16 localizations: (1) cell wall, (2) centriole, (3) chloroplast, (4) cyanelle, (5) cytoplasm, (6) cytoskeleton, (7) endoplasmic reticulum, (8) extracell, (9) Golgi apparatus, (10) lysosome, (11) mitochondria, (12) nucleus, (13) peroxisome, (14) plasma membrane, (15) plastid, and (16) vacuole. To get rid of redundancy and homology bias, none of the proteins investigated here had >/=25% sequence identity to any other in a same subcellular location. The overall success rates thus obtained via the jack-knife cross-validation test and independent dataset test were 81.6% and 83.7%, respectively, which were 46 approximately 63% higher than those performed by the other existing methods on the same benchmark datasets. Also, it is clearly elucidated that the overwhelmingly high success rates obtained by the fusion classifier is by no means a trivial utilization of the GO annotations as prone to be misinterpreted because there is a huge number of proteins with given accession numbers and the corresponding GO numbers, but their subcellular locations are still unknown, and that the percentage of proteins with GO annotations indicating their subcellular components is even less than the percentage of proteins with known subcellular location annotation in the Swiss-Prot database. It is anticipated that the powerful fusion classifier may also become a very useful high throughput tool in characterizing other attributes of proteins according to their sequences, such as enzyme class, membrane protein type, and nuclear receptor subfamily, among many others. A web server, called "Euk-OET-PLoc", has been designed at http://202.120.37.186/bioinf/euk-oet for public to predict subcellular locations of eukaryotic proteins by the fusion OET-KNN classifier.

Amino Acids↗

Differentially expressed genes in pancreatic ductal adenocarcinomas identified through serial analysis of gene expression.

Serial analysis of gene expression (SAGE) is a powerful tool for the discovery of novel tumor markers. The publicly available online SAGE libraries of normal and neoplastic tissues (http://www.ncbi.nlm.nih.gov/SAGE/) have recently been expanded; in addition, a more complete annotation of the human genome and better biocomputational techniques have substantially improved the assignment of differentially expressed SAGE "tags" to human genes. These improvements have provided us with an opportunity to re-evaluate global gene expression in pancreatic cancer using existing SAGE libraries. SAGE libraries generated from six pancreatic cancers were compared to SAGE libraries generated from 11 non-neoplastic tissues. Compared to normal tissue libraries, we identified 453 SAGE tags as differentially expressed in pancreatic cancer, including 395 that mapped to known genes and 58 "uncharacterized" tags. Of the 395 SAGE tags assigned to known genes, 223 were overexpressed in pancreatic cancer, and 172 were underexpressed. In order to map the 58 uncharacterized differentially expressed SAGE tags to genes, we used a newly developed resource called TAGmapper (http://tagmapper.ibioinformatics.org), to identify 16 additional differentially expressed genes. The differential expression of seven genes, involved in multiple cellular processes such as signal transduction (MIC-1), differentiation (DMBT1 and Neugrin), immune response (CD74), inflammation (CXCL2), cell cycle (CEB1) and enzymatic activity (Kallikrein 6), was confirmed by either immunohistochemical labeling of tissue microarrays (Kallikrein 6, CD74 and DMBT1) or by RT-PCR (CEB1, Neugrin, MIC1 and CXCL2). Of note, Neugrin was one of the genes whose previously uncharacterized SAGE tag was correctly assigned using TAGmapper, validating the utility of this program. Novel differentially expressed genes in a cancer type can be identified by revisiting updated and expanded SAGE databases. TAGmapper should prove to be a powerful tool for the discovery of novel tumor markers through assignment of uncharacterized SAGE tags.

Adenocarcinoma↗

Systematic determination of patterns of gene expression during Drosophila embryogenesis.

BACKGROUND: Cell-fate specification and tissue differentiation during development are largely achieved by the regulation of gene transcription. RESULTS: As a first step to creating a comprehensive atlas of gene-expression patterns during Drosophila embryogenesis, we examined 2,179 genes by in situ hybridization to fixed Drosophila embryos. Of the genes assayed, 63.7% displayed dynamic expression patterns that were documented with 25,690 digital photomicrographs of individual embryos. The photomicrographs were annotated using controlled vocabularies for anatomical structures that are organized into a developmental hierarchy. We also generated a detailed time course of gene expression during embryogenesis using microarrays to provide an independent corroboration of the in situ hybridization results. All image, annotation and microarray data are stored in publicly available database. We found that the RNA transcripts of about 1% of genes show clear subcellular localization. Nearly all the annotated expression patterns are distinct. We present an approach for organizing the data by hierarchical clustering of annotation terms that allows us to group tissues that express similar sets of genes as well as genes displaying similar expression patterns. CONCLUSIONS: Analyzing gene-expression patterns by in situ hybridization to whole-mount embryos provides an extremely rich dataset that can be used to identify genes involved in developmental processes that have been missed by traditional genetic analysis. Systematic analysis of rigorously annotated patterns of gene expression will complement and extend the types of analyses carried out using expression microarrays.

Animals↗

Prediction of potential GPI-modification sites in proprotein sequences.

Glycosylphosphatidylinositol (GPI) lipid anchoring is a common posttranslational modification known mainly from extracellular eukaryotic proteins. Attachment of the GPI moiety to the carboxyl terminus (omega-site) of the polypeptide follows after proteolytic cleavage of a C-terminal propeptide. For the first time, a new prediction technique locating potential GPI-modification sites in precursor sequences has been applied for large-scale protein sequence database searches. The composite prediction function (with separate parametrisation for metazoan and protozoan proteins) consists of terms evaluating both amino acid type preferences at sequence positions near a supposed omega-site as well as the concordance with general physical properties encoded in multi-residue correlation within the motif sequence. The latter terms are especially successful in rejecting non-appropriate sequences from consideration. The algorithm has been validated with a self-consistency and two jack-knife tests for the learning set of fully annotated sequences from the SWISS-PROT database as well as with a newly created database "big-Pi" (more than 300 GPI-motif mutations extracted from original literature sources). The accuracy of predicting the effect of mutations in the GPI sequence motif was above 83 %. Lists of potential precursor proteins which are non-annotated in SWISS-PROT and SPTrEMBL are presented on the WWW-page http://www.embl-heidelberg.de/beisenha/gpi/gpi_p rediction. html The algorithm has been implemented in the prototype software "big-Pi predictor" which may find application as a genome annotation and target selection tool.

Algorithms↗

Identification and validation of novel ERBB2 (HER2, NEU) targets including genes involved in angiogenesis.

V-erb-b2 erythroblastic leukemia viral oncogene homolog 2 (ERBB2; synonyms HER2, NEU) encodes a transmembrane glycoprotein with tyrosine kinase-specific activity that acts as a major switch in different signal-transduction processes. ERBB2 amplification and overexpression have been found in a number of human cancers, including breast, ovary and kidney carcinoma. Our aim was to detect ERBB2-regulated target genes that contribute to its tumorigenic effect on a genomewide scale. The differential gene expression profile of ERBB2-transfected and wild-type mouse fibroblasts was monitored employing DNA microarrays. Regulated expression of selected genes was verified by RT-PCR and validated by Western blot analysis. Genome wide gene expression profiling identified (i) known targets of ERBB2 signaling, (ii) genes implicated in tumorigenesis but so far not associated with ERBB2 signaling as well as (iii) genes not yet associated with oncogenic transformation, including novel genes without functional annotation. We also found that at least a fraction of coexpressed genes are closely linked on the genome. ERBB2 overexpression suppresses the transcription of antiangiogenic factors (e.g., Sparc, Timp3, Serpinf1) but induces expression of angiogenic factors (e.g., Klf5, Tnfaip2, Sema3c). Profiling of ERBB2-dependent gene regulation revealed a compendium of potential diagnostic markers and putative therapeutic targets. Identification of coexpressed genes that colocalize in the genome may indicate gene regulatory mechanisms that require further study to evaluate functional coregulation. (Supplementary material for this article can be found on the International Journal of Cancer website at http://www.interscience.wiley.com/jpages/0020-7136/suppmat/index.html.)

Animals↗

The consensus coding sequences of human breast and colorectal cancers.

The elucidation of the human genome sequence has made it possible to identify genetic alterations in cancers in unprecedented detail. To begin a systematic analysis of such alterations, we determined the sequence of well-annotated human protein-coding genes in two common tumor types. Analysis of 13,023 genes in 11 breast and 11 colorectal cancers revealed that individual tumors accumulate an average of approximately 90 mutant genes but that only a subset of these contribute to the neoplastic process. Using stringent criteria to delineate this subset, we identified 189 genes (average of 11 per tumor) that were mutated at significant frequency. The vast majority of these genes were not known to be genetically altered in tumors and are predicted to affect a wide range of cellular functions, including transcription, adhesion, and invasion. These data define the genetic landscape of two human cancer types, provide new targets for diagnostic and therapeutic intervention, and open fertile avenues for basic research in tumor biology.

Amino Acid Substitution↗

PubMatrix: a tool for multiplex literature mining.

BACKGROUND: Molecular experiments using multiplex strategies such as cDNA microarrays or proteomic approaches generate large datasets requiring biological interpretation. Text based data mining tools have recently been developed to query large biological datasets of this type of data. PubMatrix is a web-based tool that allows simple text based mining of the NCBI literature search service PubMed using any two lists of keywords terms, resulting in a frequency matrix of term co-occurrence. RESULTS: For example, a simple term selection procedure allows automatic pair-wise comparisons of approximately 1-100 search terms versus approximately 1-10 modifier terms, resulting in up to 1,000 pair wise comparisons. The matrix table of pair-wise comparisons can then be surveyed, queried individually, and archived. Lists of keywords can include any terms currently capable of being searched in PubMed. In the context of cDNA microarray studies, this may be used for the annotation of gene lists from clusters of genes that are expressed coordinately. An associated PubMatrix public archive provides previous searches using common useful lists of keyword terms. CONCLUSIONS: In this way, lists of terms, such as gene names, or functional assignments can be assigned genetic, biological, or clinical relevance in a rapid flexible systematic fashion. http://pubmatrix.grc.nia.nih.gov/

Cell Line, Tumor↗

Bioinformatic analysis of an unusual gene-enzyme relationship in the arginine biosynthetic pathway among marine gamma proteobacteria: implications concerning the formation of N-acetylated intermediates in prokaryotes.

BACKGROUND: The N-acetylation of L-glutamate is regarded as a universal metabolic strategy to commit glutamate towards arginine biosynthesis. Until recently, this reaction was thought to be catalyzed by either of two enzymes: (i) the classical N-acetylglutamate synthase (NAGS, gene argA) first characterized in Escherichia coli and Pseudomonas aeruginosa several decades ago and also present in vertebrates, or (ii) the bifunctional version of ornithine acetyltransferase (OAT, gene argJ) present in Bacteria, Archaea and many Eukaryotes. This paper focuses on a new and surprising aspect of glutamate acetylation. We recently showed that in Moritella abyssi and M. profunda, two marine gamma proteobacteria, the gene for the last enzyme in arginine biosynthesis (argH) is fused to a short sequence that corresponds to the C-terminal, N-acetyltransferase-encoding domain of NAGS and is able to complement an argA mutant of E. coli. Very recently, other authors identified in Mycobacterium tuberculosis an independent gene corresponding to this short C-terminal domain and coding for a new type of NAGS. We have investigated the two prokaryotic Domains for patterns of gene-enzyme relationships in the first committed step of arginine biosynthesis. RESULTS: The argH-A fusion, designated argH(A), and discovered in Moritella was found to be present in (and confined to) marine gamma proteobacteria of the Alteromonas- and Vibrio-like group. Most of them have a classical NAGS with the exception of Idiomarina loihiensis and Pseudoalteromonas haloplanktis which nevertheless can grow in the absence of arginine and therefore appear to rely on the arg(A) sequence for arginine biosynthesis. Screening prokaryotic genomes for virtual argH-X 'fusions' where X stands for a homologue of arg(A), we retrieved a large number of Bacteria and several Archaea, all of them devoid of a classical NAGS. In the case of Thermus thermophilus and Deinococcus radiodurans, the arg(A)-like sequence clusters with argH in an operon-like fashion. In this group of sequences, we find the short novel NAGS of the type identified in M. tuberculosis. Among these organisms, at least Thermus, Mycobacterium and Streptomyces species appear to rely on this short NAGS version for arginine biosynthesis. CONCLUSION: The gene-enzyme relationship for the first committed step of arginine biosynthesis should now be considered in a new perspective. In addition to bifunctional OAT, nature appears to implement at least three alternatives for the acetylation of glutamate. It is possible to propose evolutionary relationships between them starting from the same ancestral N-acetyltransferase domain. In M. tuberculosis and many other bacteria, this domain evolved as an independent enzyme, whereas it fused either with a carbamate kinase fold to give the classical NAGS (as in E. coli) or with argH as in marine gamma proteobacteria. Moreover, there is an urgent need to clarify the current nomenclature since the same gene name argA has been used to designate structurally different entities. Clarifying the confusion would help to prevent erroneous genomic annotation.

Acetylation↗

Assessing the performance of different high-density tiling microarray strategies for mapping transcribed regions of the human genome.

Genomic tiling microarrays have become a popular tool for interrogating the transcriptional activity of large regions of the genome in an unbiased fashion. There are several key parameters associated with each tiling experiment (e.g., experimental protocols and genomic tiling density). Here, we assess the role of these parameters as they are manifest in different tiling-array platforms used for transcription mapping. First, we analyze how a number of published tiling-array experiments agree with established gene annotation on human chromosome 22. We observe that the transcription detected from high-density arrays correlates substantially better with annotation than that from other array types. Next, we analyze the transcription-mapping performance of the two main high-density oligonucleotide array platforms in the ENCODE regions of the human genome. We hybridize identical biological samples and develop several ways of scoring the arrays and segmenting the genome into transcribed and nontranscribed regions, with the aim of making the platforms most comparable to each other. Finally, we develop a platform comparison approach based on agreement with known annotation. Overall, we find that the performance improves with more data points per locus, coupled with statistical scoring approaches that properly take advantage of this, where this larger number of data points arises from higher genomic tiling density and the use of replicate arrays and mismatches. While we do find significant differences in the performance of the two high-density platforms, we also find that they complement each other to some extent. Finally, our experiments reveal a significant amount of novel transcription outside of known genes, and an appreciable sample of this was validated by independent experiments.

Cell Line↗

Functional replacement of the FabA and FabB proteins of Escherichia coli fatty acid synthesis by Enterococcus faecalis FabZ and FabF homologues.

The anaerobic unsaturated fatty acid synthetic pathway of Escherichia coli requires two specialized proteins, FabA and FabB. However, the fabA and fabB genes are found only in the Gram-negative alpha- and gamma-proteobacteria, and thus other anaerobic bacteria must synthesize these acids using different enzymes. We report that the Gram-positive bacterium Enterococcus faecalis encodes a protein, annotated as FabZ1, that functionally replaces the E. coli FabA protein, although the sequence of this protein aligns much more closely with E. coli FabZ, a protein that plays no specific role in unsaturated fatty acid synthesis. Therefore E. faecalis FabZ1 is a bifunctional dehydratase/isomerase, an enzyme activity heretofore confined to a group of Gram-negative bacteria. The FabZ2 protein is unable to replace the function of E. coli FabZ, although FabZ2, a second E. faecalis FabZ homologue, has this ability. Moreover, an E. faecalis FabF homologue (FabF1) was found to replace the function of E. coli FabB, whereas a second FabF homologue was inactive. From these data it is clear that bacterial fatty acid biosynthetic pathways cannot be deduced solely by sequence comparisons.

3-Oxoacyl-(Acyl-Carrier-Protein) Synthase↗

iMTSS: an integrated framework for biology- and patient-driven prognosis in myelofibrosis undergoing transplantation.

BACKGROUND: Allogeneic hematopoietic cell transplantation is the only curative treatment for myelofibrosis, but failure occurs by two mechanistically distinct routes: relapse of the neoplasm, which reflects its underlying genetics, and non-relapse mortality, which reflects whether the patient and graft tolerate the procedure. Established prognostic systems either lack molecular granularity or were derived in the non-transplant setting, and all collapse these two routes into a single survival estimate. None can indicate why an individual patient is at risk, or which class of intervention might reduce that risk. OBJECTIVE: To determine why an individual patient is at risk and to develop and validate an integrated framework that quantifies biology- and patient-driven prognosis. STUDY DESIGN: We analyzed 1,550 adults undergoing first allogeneic transplantation for primary or secondary myelofibrosis across international centers, the largest genomically annotated transplant cohort in this disease. The cohort was split into development (n=930) and validation (n=620) sets. Overall survival was modeled by Cox regression; relapse and non-relapse mortality were modeled as competing events by Fine-Gray subdistribution-hazard regression at 2 years. Discrimination was assessed by the concordance index with bootstrap confidence intervals. The molecular contribution was quantified by variance decomposition of, and robustness to the analytic choices was examined by resampling. RESULTS: A genetically defined disease-intrinsic axis, including TP53 allelic state, RAS pathway mutations, ASXL1 and driver genotype, blasts and blood counts, predicted 2 year relapse incidence (validation concordance 0.69, 95% CI 0.63 to 0.74), whereas a non-overlapping host and structural axis, including portal vein thrombosis, donor type, patients' performance status, and age predicted 2-year non-relapse mortality (0.63, 95% CI 0.59 to 0.68). The two scores shared only 3.4% of their variance, indicating that a patient's disease genetics carried almost no information about non-relapse mortality. Variance decomposition showed that TP53 allelic state alone accounted for 30% of the relapse score. Recombined, the framework discriminated overall survival (concordance 0.640, 95% CI 0.616 to 0.662) better than every established prognostic system. For proof of concept, 3 risk groups separated in the validation cohort, with 5 year survival of 72%, 58%, and 39% (P<0.001), and the models were well calibrated. CONCLUSIONS: Relapse and non-relapse mortality after transplantation for myelofibrosis are governed by distinct dimensions. Estimating both outcomes independently with genetic and clinical information, in addition to overall survival, establishes an individualized basis for transplant decision-making. The calculator is openly available (https://imtss-calculator.com).

mortality↗

The 630-kb lung cancer homozygous deletion region on human chromosome 3p21.3: identification and evaluation of the resident candidate tumor suppressor genes. The International Lung Cancer Chromosome 3p21.3 Tumor Suppressor Gene Consortium.

We used overlapping and nested homozygous deletions, contig building, genomic sequencing, and physical and transcript mapping to further define a approximately 630-kb lung cancer homozygous deletion region harboring one or more tumor suppressor genes (TSGs) on chromosome 3p21.3. This location was identified through somatic genetic mapping in tumors, cancer cell lines, and premalignant lesions of the lung and breast, including the discovery of several homozygous deletions. The combination of molecular manual methods and computational predictions permitted us to detect, isolate, characterize, and annotate a set of 25 genes that likely constitute the complete set of protein-coding genes residing in this approximately 630-kb sequence. A subset of 19 of these genes was found within the deleted overlap region of approximately 370-kb. This region was further subdivided by a nesting 200-kb breast cancer homozygous deletion into two gene sets: 8 genes lying in the proximal approximately 120-kb segment and 11 genes lying in the distal approximately 250-kb segment. These 19 genes were analyzed extensively by computational methods and were tested by manual methods for loss of expression and mutations in lung cancers to identify candidate TSGs from within this group. Four genes showed loss-of-expression or reduced mRNA levels in non-small cell lung cancer (CACNA2D2/alpha2delta-2, SEMA3B [formerly SEMA(V), BLU, and HYAL1] or small cell lung cancer (SEMA3B, BLU, and HYAL1) cell lines. We found six of the genes to have two or more amino acid sequence-altering mutations including BLU, NPRL2/Gene21, FUS1, HYAL1, FUS2, and SEMA3B. However, none of the 19 genes tested for mutation showed a frequent (>10%) mutation rate in lung cancer samples. This led us to exclude several of the genes in the region as classical tumor suppressors for sporadic lung cancer. On the other hand, the putative lung cancer TSG in this location may either be inactivated by tumor-acquired promoter hypermethylation or belong to the novel class of haploinsufficient genes that predispose to cancer in a hemizygous (+/-) state but do not show a second mutation in the remaining wild-type allele in the tumor. We discuss the data in the context of novel and classic cancer gene models as applied to lung carcinogenesis. Further functional testing of the critical genes by gene transfer and gene disruption strategies should permit the identification of the putative lung cancer TSG(s), LUCA, Analysis of the approximately 630-kb sequence also provides an opportunity to probe and understand the genomic structure, evolution, and functional organization of this relatively gene-rich region.

Carcinoma, Non-Small-Cell Lung↗