PubMed Health⌕ Search

Biomedical subjects

Pan Du

Publications and source records attributed to Pan Du.

3 recordsLinked to original sources

Improved peak detection in mass spectrum by incorporating continuous wavelet transform-based pattern matching.

MOTIVATION: A major problem for current peak detection algorithms is that noise in mass spectrometry (MS) spectra gives rise to a high rate of false positives. The false positive rate is especially problematic in detecting peaks with low amplitudes. Usually, various baseline correction algorithms and smoothing methods are applied before attempting peak detection. This approach is very sensitive to the amount of smoothing and aggressiveness of the baseline correction, which contribute to making peak detection results inconsistent between runs, instrumentation and analysis methods. RESULTS: Most peak detection algorithms simply identify peaks based on amplitude, ignoring the additional information present in the shape of the peaks in a spectrum. In our experience, 'true' peaks have characteristic shapes, and providing a shape-matching function that provides a 'goodness of fit' coefficient should provide a more robust peak identification method. Based on these observations, a continuous wavelet transform (CWT)-based peak detection algorithm has been devised that identifies peaks with different scales and amplitudes. By transforming the spectrum into wavelet space, the pattern-matching problem is simplified and in addition provides a powerful technique for identifying and separating the signal from the spike noise and colored noise. This transformation, with the additional information provided by the 2D CWT coefficients can greatly enhance the effective signal-to-noise ratio. Furthermore, with this technique no baseline removal or peak smoothing preprocessing steps are required before peak detection, and this improves the robustness of peak detection under a variety of conditions. The algorithm was evaluated with SELDI-TOF spectra with known polypeptide positions. Comparisons with two other popular algorithms were performed. The results show the CWT-based algorithm can identify both strong and weak peaks while keeping false positive rate low. AVAILABILITY: The algorithm is implemented in R and will be included as an open source module in the Bioconductor project.

Algorithms↗

High-level expression and purification of Escherichia coli oligopeptidase B.

Oligopeptidase B (OpdB) of Escherichia coli, previously called protease II, has a trypsin-like specificity, cleaving peptides at lysine and arginine residues and belongs to the prolyl oligopeptidase family of new serine peptidases. In this study, we report the fusion expression of E. coli oligopeptidase B with an N-terminal histidine tag using pET28a as the expression vector. Although most of the recombinant OpdB was produced as inclusion bodies, the solubility of the recombinant protease increased significantly when the expression temperature shifted from 37 to 30 degrees C. Recombinant OpdB (approximately 10 mg) could be purified from the soluble fraction of the crude extract of 1L log-phase E. coli culture containing 1.5 g wet bacterial cells. The purified OpdB has a molecular weight of approximately 80 kDa and a specific activity of 4.8 x 10(4) U/mg. OpdB could also be purified from the inclusion bodies with a lower yield. The recombinant enzyme was very stable under 40 degrees C. By comparison of the substrate specificity of the purified OpdB with that of OpdA, another trypsin-like protease in E. coli, we found that Boc-Glu-Lys-Lys-MCA is a specific substrate for E. coli OpdB. We also found that compared to OpdA, OpdB is much more sensitive to GMCHA-OPh(t)Bu, a synthetic trypsin inhibitor that can retard the growth of E. coli.

Cyclohexanecarboxylic Acids↗

Modeling gene expression networks using fuzzy logic.

Gene regulatory networks model regulation in living organisms. Fuzzy logic can effectively model gene regulation and interaction to accurately reflect the underlying biology. A new multiscale fuzzy clustering method allows genes to interact between regulatory pathways and across different conditions at different levels of detail. Fuzzy cluster centers can be used to quickly discover causal relationships between groups of coregulated genes. Fuzzy measures weight expert knowledge and help quantify uncertainty about the functions of genes using annotations and the gene ontology database to confirm some of the interactions. The method is illustrated using gene expression data from an experiment on carbohydrate metabolism in the model plant Arabidopsis thaliana. Key gene regulatory relationships were evaluated using information from the gene ontology database. A new regulatory relationship concerning trehalose regulation of carbohydrate metabolism was also discovered in the extracted network.

Animals↗