PubMed Health⌕ Search

PubMed · 10966805

A Bayesian system integrating expression data with sequence patterns for localizing proteins: comprehensive application to the yeast genome.

Abstract

We develop a probabilistic system for predicting the subcellular localization of proteins and estimating the relative population of the various compartments in yeast. Our system employs a Bayesian approach, updating a protein's probability of being in a compartment, based on a diverse range of 30 features. These range from specific motifs (e.g. signal sequences or the HDEL motif) to overall properties of a sequence (e.g. surface composition or isoelectric point) to whole-genome data (e.g. absolute mRNA expression levels or their fluctuations). The strength of our approach is the easy integration of many features, particularly the whole-genome expression data. We construct a training and testing set of approximately 1300 yeast proteins with an experimentally known localization from merging, filtering, and standardizing the annotation in the MIPS, Swiss-Prot and YPD databases, and we achieve 75 % accuracy on individual protein predictions using this dataset. Moreover, we are able to estimate the relative protein population of the various compartments without requiring a definite localization for every protein. This approach, which is based on an analogy to formalism in quantum mechanics, gives better accuracy in determining relative compartment populations than that obtained by simply tallying the localization predictions for individual proteins (on the yeast proteins with known localization, 92% versus 74%). Our training and testing also highlights which of the 30 features are informative and which are redundant (19 being particularly useful). After developing our system, we apply it to the 4700 yeast proteins with currently unknown localization and estimate the relative population of the various compartments in the entire yeast genome. An unbiased prior is essential to this extrapolated estimate; for this, we use the MIPS localization catalogue, and adapt recent results on the localization of yeast proteins obtained by Snyder and colleagues using a minitransposon system. Our final localizations for all approximately 6000 proteins in the yeast genome are available over the web at: http://bioinfo.mbb.yale. edu/genome/localize.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

A Drawid, M Gerstein. 2000-08-25. A Bayesian system integrating expression data with sequence patterns for localizing proteins: comprehensive application to the yeast genome.. https://doi.org/10.1006/jmbi.2000.3968

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

GAMMA: gap-aware motif mining under incomplete labeling with applications to MHC motifs.

MOTIVATION: Sequence motif identification is crucial for understanding molecular recognition, particularly in immune responses involving peptide binding to major histocompatibility complex (MHC) Class I molecules for antigen presentation to T cells. Traditionally, MHC Class I binding motifs are assumed to be contiguous and span nine amino acids. However, structural evidence suggests that binding may involve nonadjacent residues, challenging the assumptions of existing methods. RESULTS: In this study, we propose Gap-Aware Motif Mining Algorithm (GAMMA), a probabilistic framework designed to identify noncontiguous motifs under conditions of incomplete labeling. GAMMA employs Bayesian inference with Markov chain Monte Carlo sampling to jointly estimate motif parameters, binding locations, and the relative spacing between binding positions. Through extensive simulations and real-world applications to MHC Class I peptide datasets, GAMMA outperforms existing motif discovery tools such as GLAM2 in accurately localizing binding residues and identifying the underlying motifs. Notably, our results suggest that the true number of binding residues may be eight, fewer than the commonly assumed nine. In addition, for longer peptides, the model captures increased flexibility in the central region, consistent with structural observations that peptides may bulge in the middle. AVAILABILITY AND IMPLEMENTATION: The raw data and the source codes are available on GitHub (https://github.com/RanLIUaca/GAMMAmotif).

Amino Acid Motifs↗

Signalling thresholds and negative B-cell selection in acute lymphoblastic leukaemia.

B cells are selected for an intermediate level of B-cell antigen receptor (BCR) signalling strength: attenuation below minimum (for example, non-functional BCR) or hyperactivation above maximum (for example, self-reactive BCR) thresholds of signalling strength causes negative selection. In ∼25% of cases, acute lymphoblastic leukaemia (ALL) cells carry the oncogenic BCR-ABL1 tyrosine kinase (Philadelphia chromosome positive), which mimics constitutively active pre-BCR signalling. Current therapeutic approaches are largely focused on the development of more potent tyrosine kinase inhibitors to suppress oncogenic signalling below a minimum threshold for survival. We tested the hypothesis that targeted hyperactivation--above a maximum threshold--will engage a deletional checkpoint for removal of self-reactive B cells and selectively kill ALL cells. Here we find, by testing various components of proximal pre-BCR signalling in mouse BCR-ABL1 cells, that an incremental increase of Syk tyrosine kinase activity was required and sufficient to induce cell death. Hyperactive Syk was functionally equivalent to acute activation of a self-reactive BCR on ALL cells. Despite oncogenic transformation, this basic mechanism of negative selection was still functional in ALL cells. Unlike normal pre-B cells, patient-derived ALL cells express the inhibitory receptors PECAM1, CD300A and LAIR1 at high levels. Genetic studies revealed that Pecam1, Cd300a and Lair1 are critical to calibrate oncogenic signalling strength through recruitment of the inhibitory phosphatases Ptpn6 (ref. 7) and Inpp5d (ref. 8). Using a novel small-molecule inhibitor of INPP5D (also known as SHIP1), we demonstrated that pharmacological hyperactivation of SYK and engagement of negative B-cell selection represents a promising new strategy to overcome drug resistance in human ALL.

Amino Acid Motifs↗

The ETS domain transcription factor Elk-1 contains a novel class of repression domain.

The ETS domain transcription factor Elk-1 serves as an integration point for different mitogen-activated protein (MAP) kinase pathways. Phosphorylation of Elk-1 by MAP kinases triggers its activation. However, while the activation process is well understood, its downregulation-inactivation is less well characterized. The ETS DNA-binding domain plays a role in the downregulation of Elk-dependent promoter activity following mitogenic activation by recruiting the mSin3A-HDAC complex. Here we have identified a novel evolutionarily conserved repression domain in Elk-1, termed the R motif, which serves to reduce the basal transcriptional activity of Elk-1 and dampen its response to mitogenic signals. This domain is highly potent and portable and can repress transcription in trans. The R motif is related to the CRD1 repression domain in p300 and can functionally replace this domain and confer p21(waf1/cip1) inducibility on p300. However, the R motif acts in a context-dependent manner and is not p21(waf1/cip1) responsive in Elk-1. Thus, the Elk-1 R motif and the p300 CRD1 motif represent a new class of repression domains that are regulated in a context-dependent manner.

Amino Acid Motifs↗