PubMed Health⌕ Search

Biomedical subjects

Songfeng Wu

Publications and source records attributed to Songfeng Wu.

7 recordsLinked to original sources

A dataset of human fetal liver proteome identified by subcellular fractionation and multiple protein separation and identification technology.

A high throughput process including subcellular fractionation and multiple protein separation and identification technology allowed us to establish the protein expression profile of human fetal liver, which was composed of at least 2,495 distinct proteins and 568 non-isoform groups identified from 64,960 peptides and 24,454 distinct peptides. In addition to the basic protein identification mentioned above, the MS data were used for complementary identification and novel protein mining. By doing the analysis with integrated protein, expressed sequence tag, and genome datasets, 223 proteins and 15 peptides were complementarily identified with high quality MS/MS data.

Cell Membrane↗

Multi-modality of pI distribution in whole proteome.

Multi-modality of pI distribution is a common feature in different whole proteomes. Some researchers considered it relate to the proteins with different subcellular locations, indicating the result of natural selection. We explored the pI distribution of predicted proteomes (including animals, plants, bacterium, archaeans) and random proteome [random protein sequences constructed according to the special amino acid composition and molecular weight (MW) distribution of human predicted proteome]. Our results suggest that the multi-modality is the result of discrete pK(R) values for different amino acids. Amino acid composition and MW distribution of a proteome also contributes to the specific pI distribution. Although protein subcellular location was related to pI value, our analyses revealed that comparing with the random proteome, neither the multi-modality phenomenon nor the distribution bias of pI values is caused by subcellular location. It seems that the multi-modality distribution is just a mathematical fun. The blank region near the neutral pI was caused by the absence of amino acids with neutral pK(R), and suggests that the selection of amino acids with ionizable side chain might be restricted by the requirement for a special pH environment during the origin of life. From this point of view, the special distribution was the result of natural selection.

Animals↗

Protein interaction networks of Saccharomyces cerevisiae, Caenorhabditis elegans and Drosophila melanogaster: large-scale organization and robustness.

High-throughput screens have begun to reveal protein interaction networks in several organisms. To understand the general properties of these protein interaction networks, a systematic analysis of topological structure and robustness was performed on the protein interaction networks of Saccharomyces cerevisiae, Caenorhabditis elegans and Drosophila melanogaster. It shows that the three protein interaction networks have a scale-free and high-degree clustering nature as the consequence of their hierarchical organization. It also shows that they have the small-world property with similar diameter at 4-5. Evaluation of the consequences of random removal of both proteins and interactions from the protein interaction networks suggests their high degree of robustness. Simulation of a protein's removal shows that the protein interaction network's error tolerance is accompanied by attack vulnerability. These fundamental analyses of the networks might serve as a starting point for further exploring complex biological networks and the coming research of "systems biology".

Algorithms↗

Protein probabilities in shotgun proteomics: evaluating different estimation methods using a semi-random sampling model.

The calculation of protein probabilities is one of the most intractable problems in large-scale proteomic research. Current available estimating methods, for example, ProteinProphet, PROT_PROBE, Poisson model and two-peptide hits, employ different models trying to resolve this problem. Until now, no efficient method is used for comparative evaluation of the above methods in large-scale datasets. In order to evaluate these various methods, we developed a semi-random sampling model to simulate large-scale proteomic data. In this model, the identified peptides were sampled from the designed proteins and their cross-correlation scores were simulated according to the results from reverse database searching. The simulated result of 18 control proteins was consistent with the experimental one, demonstrating the efficiency of our model. According to the simulated results of human liver sample, ProteinProphet returned slightly higher probabilities and lower specificity than real cases. PROT_PROBE was a more efficient method with higher specificity. Predicted results from a Poisson model roughly coincide with real datasets, and the method of two-peptide hits seems solid but imprecise. However, the probabilities of identified proteins are strongly correlated with several experimental factors including spectra number, database size and protein abundance distribution.

Chromatography, Liquid↗

Comparison of alternative analytical techniques for the characterisation of the human serum proteome in HUPO Plasma Proteome Project.

Based on the same HUPO reference specimen (C1-serum) with the six proteins of highest abundance depleted by immunoaffinity chromatography, we have compared five proteomics approaches, which were (1) intact protein fractionation by anion-exchange chromatography followed by 2-DE-MALDI-TOF-MS/MS for protein identification (2-DE strategy); (2) intact protein fractionation by 2-D HPLC followed by tryptic digestion of each fraction and microcapillary RP-HPLC/microESI-MS/MS identification (protein 2-D HPLC fractionation strategy); (3) protein digestion followed by automated online microcapillary 2-D HPLC (strong cation-exchange chromatography (SCX)-RPC) with IT microESI-MS/MS; (online shotgun strategy); (4) same as (3) with the SCX step performed offline (offline shotgun strategy) and (5) same as (4) with the SCX fractions reanalysed by optimised nanoRP-HPLC-nanoESI-MS/MS (offline shotgun-nanospray strategy). All five approaches yielded complementary sets of protein identifications. The total number of unique proteins identified by each of these five approaches was (1) 78, (2) 179, (3) 131, (4) 224 and (5) 330 respectively. In all, 560 unique proteins were identified. One hundred and sixty-five proteins were identified through two or more peptides, which could be considered a high-confidence identification. Only 37 proteins were identified by all five approaches. The 2-DE approach yielded more information on the pI-altered isoforms of some serum proteins and the relative abundance of identified proteins. The protein prefractionation strategy slightly improved the capacity to detect proteins of lower abundance. Optimising the separation at the peptide level and improving the detection sensitivity of ESI-MS/MS were more effective than fractionation of intact proteins in increasing the total number of proteins identified. Overall, electrophoresis and chromatography, coupled respectively with MALDI-TOF/TOF-MS and ESI-MS/MS, identified complementary sets of serum proteins.

Blood Proteins↗

Characterization of Ceap-11 and Ceap-16, two novel splicing-variant-proteins, associated with centrosome, microtubule aggregation and cell proliferation.

A novel human gene, encoding two polypeptide-isoforms, has been identified from human fetal liver cDNA library. These two alternatively spliced polypeptide-variants are associated with centrosomes, and are designated Ceap-11 and Ceap-16, respectively, according to the acronym Ceap for centrosomal-associated protein and the approximate relative molecular mass. The high degree of sequence similarity between Ceap proteins of divergent species indicates that the Ceap homologous genes are significantly conserved in evolution and constitute a new gene family without any functional information until now. Human Ceap gene is mapped on 10q24.2. These two Ceap cDNA isoforms are generated by RNA alternative splicing on the 5' terminus of the Ceap gene, and are composed of four and five exons, respectively. Ceap-11 and Ceap-16 are co-immunoprecipitated and co-located with gamma-tubulin; ectopic overexpression of these two proteins in NIH3T3 cells induces microtubule aggregation and cell proliferation; the protein level of Ceap in certain tumors is significantly higher than that in corresponding normal tissues. Taken together, our data provide the first evidence for the function of Ceap-11 and Ceap-16, the two novel human proteins, namely, association with centrosome, microtubule aggregation and cell proliferation.

Alternative Splicing↗

Proteomic analysis on structural proteins of Severe Acute Respiratory Syndrome coronavirus.

Recently, a new coronavirus was isolated from the lung tissue of autopsy sample and nasal/throat swabs of the patients with Severe Acute Respiratory Syndrome (SARS) and the causative association with SARS was determined. To reveal further the characteristics of the virus and to provide insight about the molecular mechanism of SARS etiology, a proteomic strategy was utilized to identify the structural proteins of SARS coronavirus (SARS-CoV) isolated from Vero E6 cells infected with the BJ-01 strain of the virus. At first, Western blotting with the convalescent sera from SARS patients demonstrated that there were various structural proteins of SARS-CoV in the cultured supernatant of virus infected-Vero E6 cells and that nucleocaspid (N) protein had a prominent immunogenicity to the convalescent sera from the patients with SARS, while the immune response of spike (S) protein probably binding with membrane (M) glycoprotein was much weaker. Then, sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE) was used to separate the complex protein constituents, and the strategy of continuous slicing from loading well to the bottom of the gels was utilized to search thoroughly the structural proteins of the virus. The proteins in sliced slots were trypsinized in-gel and identified by mass spectrometry. Three structural proteins named S, N and M proteins of SARS-CoV were uncovered with the sequence coverage of 38.9, 93.1 and 28.1% respectively. Glycosylation modification in S protein was also analyzed and four glycosylation sites were discovered by comparing the mass spectra before and after deglycosylation of the peptides with PNGase F digestion. Matrix-assisted laser desorption/ionization-mass spectrometry determination showed that relative molecular weight of intact N protein is 45 929 Da, which is very close to its theoretically calculated molecular weight 45 935 Da based on the amino acid sequence deduced from the genome with the first amino acid methionine at the N-terminus depleted and second, serine, acetylated, indicating that phosphorylation does not happen at all in the predicted phosphorylation sites within infected cells nor in virus particles. Intriguingly, a series of shorter isoforms of N protein was observed by SDS-PAGE and identified by mass spectrometry characterization. For further confirmation of this phenomenon and its related mechanism, recombinant N protein of SARS-CoV was cleaved in vitro by caspase-3 and -6 respectively. The results demonstrated that these shorter isoforms could be the products from cleavage of caspase-3 rather than that of caspase-6. Further, the relationship between the caspase cleavage and the viral infection to the host cell is discussed.

Amino Acid Sequence↗