PubMed Health⌕ Search

PubMed · 12217908

Identification of regulatory elements using a feature selection method.

Abstract

MOTIVATION: Many methods have been described to identify regulatory motifs in the transcription control regions of genes that exhibit similar patterns of gene expression across a variety of experimental conditions. Here we focus on a single experimental condition, and utilize gene expression data to identify sequence motifs associated with genes that are activated under this experimental condition. We use a linear model with two-way interactions to model gene expression as a function of sequence features (words) present in presumptive transcription control regions. The most relevant features are selected by a feature selection method called stepwise selection with monte carlo cross validation. We apply this method to a publicly available dataset of the yeast Saccharomyces cerevisiae, focussing on the 800 basepairs immediately upstream of each gene's translation start site (the upstream control region (UCR)). RESULTS: We successfully identify regulatory motifs that are known to be active under the experimental conditions analyzed, and find additional significant sequences that may represent novel regulatory motifs. We also discuss a complementary method that utilizes gene expression data from a single microarray experiment and allows averaging over variety of experimental conditions as an alternative to motif finding methods that act on clusters of co-expressed genes. AVAILABILITY: The software is available upon request from the first author or may be downloaded from http://www.stat.berkeley.edu/~sunduz. CONTACT: keles@stat.berkeley.edu

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Sündüz Keleş, Mark van der Laan, Michael B Eisen. 2002. Identification of regulatory elements using a feature selection method.. https://doi.org/10.1093/bioinformatics%2F18.9.1167

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

GAMMA: gap-aware motif mining under incomplete labeling with applications to MHC motifs.

MOTIVATION: Sequence motif identification is crucial for understanding molecular recognition, particularly in immune responses involving peptide binding to major histocompatibility complex (MHC) Class I molecules for antigen presentation to T cells. Traditionally, MHC Class I binding motifs are assumed to be contiguous and span nine amino acids. However, structural evidence suggests that binding may involve nonadjacent residues, challenging the assumptions of existing methods. RESULTS: In this study, we propose Gap-Aware Motif Mining Algorithm (GAMMA), a probabilistic framework designed to identify noncontiguous motifs under conditions of incomplete labeling. GAMMA employs Bayesian inference with Markov chain Monte Carlo sampling to jointly estimate motif parameters, binding locations, and the relative spacing between binding positions. Through extensive simulations and real-world applications to MHC Class I peptide datasets, GAMMA outperforms existing motif discovery tools such as GLAM2 in accurately localizing binding residues and identifying the underlying motifs. Notably, our results suggest that the true number of binding residues may be eight, fewer than the commonly assumed nine. In addition, for longer peptides, the model captures increased flexibility in the central region, consistent with structural observations that peptides may bulge in the middle. AVAILABILITY AND IMPLEMENTATION: The raw data and the source codes are available on GitHub (https://github.com/RanLIUaca/GAMMAmotif).

Amino Acid Motifs↗

Signalling thresholds and negative B-cell selection in acute lymphoblastic leukaemia.

B cells are selected for an intermediate level of B-cell antigen receptor (BCR) signalling strength: attenuation below minimum (for example, non-functional BCR) or hyperactivation above maximum (for example, self-reactive BCR) thresholds of signalling strength causes negative selection. In ∼25% of cases, acute lymphoblastic leukaemia (ALL) cells carry the oncogenic BCR-ABL1 tyrosine kinase (Philadelphia chromosome positive), which mimics constitutively active pre-BCR signalling. Current therapeutic approaches are largely focused on the development of more potent tyrosine kinase inhibitors to suppress oncogenic signalling below a minimum threshold for survival. We tested the hypothesis that targeted hyperactivation--above a maximum threshold--will engage a deletional checkpoint for removal of self-reactive B cells and selectively kill ALL cells. Here we find, by testing various components of proximal pre-BCR signalling in mouse BCR-ABL1 cells, that an incremental increase of Syk tyrosine kinase activity was required and sufficient to induce cell death. Hyperactive Syk was functionally equivalent to acute activation of a self-reactive BCR on ALL cells. Despite oncogenic transformation, this basic mechanism of negative selection was still functional in ALL cells. Unlike normal pre-B cells, patient-derived ALL cells express the inhibitory receptors PECAM1, CD300A and LAIR1 at high levels. Genetic studies revealed that Pecam1, Cd300a and Lair1 are critical to calibrate oncogenic signalling strength through recruitment of the inhibitory phosphatases Ptpn6 (ref. 7) and Inpp5d (ref. 8). Using a novel small-molecule inhibitor of INPP5D (also known as SHIP1), we demonstrated that pharmacological hyperactivation of SYK and engagement of negative B-cell selection represents a promising new strategy to overcome drug resistance in human ALL.

Amino Acid Motifs↗

The tumour-suppressor function of PTEN requires an N-terminal lipid-binding motif.

The PTEN (phosphatase and tensin homologue deleted on chromosome 10) tumour-suppressor protein is a phosphoinositide 3-phosphatase which antagonizes phosphoinositide 3-kinase-dependent signalling by dephosphorylating PtdIns(3,4,5)P3. Most tumour-derived point mutations of PTEN induce a loss of function, which correlates with profoundly reduced catalytic activity. However, here we characterize a point mutation at the N-terminus of PTEN, K13E from a human glioblastoma, which displayed wild-type activity when assayed in vitro. This mutation occurs within a conserved polybasic motif, a putative PtdIns(4,5)P2-binding site that may participate in membrane targeting of PTEN. We found that catalytic activity against lipid substrates and vesicle binding of wild-type PTEN, but not of PTEN K13E, were greatly stimulated by anionic lipids, especially PtdIns(4,5)P2. The K13E mutation also greatly reduces the efficiency with which anionic lipids inhibit PTEN activity against soluble substrates, supporting the hypothesis that non-catalytic membrane binding orientates the active site to favour lipid substrates. Significantly, in contrast to the wild-type enzyme, PTEN K13E failed either to prevent protein kinase B/Akt phosphorylation, or inhibit cell proliferation when expressed in PTEN-null U87MG cells. The cellular functioning of K13E PTEN was recovered by targeting to the plasma membrane through inclusion of a myristoylation site. Our results establish a requirement for the conserved N-terminal motif of PTEN for correct membrane orientation, cellular activity and tumour-suppressor function.

Amino Acid Motifs↗