PubMed Health⌕ Search

PubMed · 16639720

Predicting protein subcellular location by fusing multiple classifiers.

Abstract

One of the fundamental goals in cell biology and proteomics is to identify the functions of proteins in the context of compartments that organize them in the cellular environment. Knowledge of subcellular locations of proteins can provide key hints for revealing their functions and understanding how they interact with each other in cellular networking. Unfortunately, it is both time-consuming and expensive to determine the localization of an uncharacterized protein in a living cell purely based on experiments. With the avalanche of newly found protein sequences emerging in the post genomic era, we are facing a critical challenge, that is, how to develop an automated method to fast and reliably identify their subcellular locations so as to be able to timely use them for basic research and drug discovery. In view of this, an ensemble classifier was developed by the approach of fusing many basic individual classifiers through a voting system. Each of these basic classifiers was trained in a different dimension of the amphiphilic pseudo amino acid composition (Chou [2005] Bioinformatics 21: 10-19). As a demonstration, predictions were performed with the fusion classifier for proteins among the following 14 localizations: (1) cell wall, (2) centriole, (3) chloroplast, (4) cytoplasm, (5) cytoskeleton, (6) endoplasmic reticulum, (7) extracellular, (8) Golgi apparatus, (9) lysosome, (10) mitochondria, (11) nucleus, (12) peroxisome, (13) plasma membrane, and (14) vacuole. The overall success rates thus obtained via the resubstitution test, jackknife test, and independent dataset test were all significantly higher than those by the existing classifiers. It is anticipated that the novel ensemble classifier may also become a very useful vehicle in classifying other attributes of proteins according to their sequences, such as membrane protein type, enzyme family/sub-family, G-protein coupled receptor (GPCR) type, and structural class, among many others. The fusion ensemble classifier will be available at www.pami.sjtu.edu.cn/people/hbshen.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Kuo-Chen Chou, Hong-Bin Shen. 2006-10-01. Predicting protein subcellular location by fusing multiple classifiers.. https://doi.org/10.1002/jcb.20879

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Chimeric structural isomer fragments as cost-efficient internal standards for amino acid quantification by mass spectrometry.

Amino acid (AA) profiles from body fluids such as blood and urine are clinical indicators for diagnosing metabolic and hepatic diseases. Current quantitative methods, such as liquid chromatography-mass spectrometry (LC-MS) with isotopically labelled internal standards (ISs), are costly and technically demanding. This study proposes a cost-efficient alternative using structural isomers as ISs in a direct liquid infusion (DLI) tandem mass spectrometry (MS/MS) approach. The method leverages chimeric spectra and fragment intensity ratios to quantify AAs, demonstrating high linearity and precision even with a 3D ion trap mass analyser. This approach offers a viable strategy for AA quantification in preventive medicine, particularly for screening metabolic diseases such as phenylketonuria, diabetes, and liver dysfunction.

Amino Acids↗

MODEL-molecular descriptor lab: a web-based server for computing structural and physicochemical features of compounds.

Molecular descriptors represent structural and physicochemical features of compounds. They have been extensively used for developing statistical models, such as quantitative structure activity relationship (QSAR) and artificial neural networks (NN), for computer prediction of the pharmacodynamic, pharmacokinetic, or toxicological properties of compounds from their structure. While computer programs have been developed for computing molecular descriptors, there is a lack of a freely accessible one. We have developed a web-based server, MODEL (Molecular Descriptor Lab), for computing a comprehensive set of 3,778 molecular descriptors, which is significantly more than the approximately 1,600 molecular descriptors computed by other software. Our computational algorithms have been extensively tested and the computed molecular descriptors have been used in a number of published works of statistical models for predicting variety of pharmacodynamic, pharmacokinetic, and toxicological properties of compounds. Several testing studies on the computed molecular descriptors are discussed. MODEL is accessible at http://jing.cz3.nus.edu.sg/cgi-bin/model/model.cgi free of charge for academic use.

Amino Acids↗

The penicillin G acylase production by B. megaterium is amino acid consumption dependent.

Aiming at to enhance the production of penicillin G acylase (PGA) by Bacillus megaterium, we have performed flasks experiments using different medium composition. Using 51 g/L of casein hydrolyzed with Alcalase and 2.7 g/L of phenylacetic acid (PhAc), the following carbon substrates were tested, individually and combined: glucose, glycerol, and lactose (present in cheese whey). Glycerol and glucose showed to be effective nutrients for the microorganism growth but delayed the PGA production. Cheese whey always increased enzyme production and cell mass. However, lactose (present in cheese whey) was not a significant carbon source for B. megaterium. PhAc, amino acids, and small peptides present in the hydrolyzed casein were the actual carbon sources for enzyme production. Replacement of hydrolyzed casein by free amino acids, 10.0 g/L, led to a significant increase in enzyme production (app. 150%), with a preferential consumption of alanine, aspartic acid, glycine, serine, arginine, threonine, lysine, and glutamic acid. A decrease of the enzyme production was observed when 20.0 g/L of amino acids were used. Using the single omission technique, it was shown that none of the 18 tested amino acids was essential for enzyme production. The use of a medium containing eight of the preferentially consumed amino acids lead to similar enzyme production level obtained when using 18 amino acids. PhAc, up to 2.7 g/L, did not inhibit enzyme production, even if added at the beginning of the cultivation.

Amino Acids↗