PubMed Health⌕ Search

Biomedical subjects

Charles X Ling

Publications and source records attributed to Charles X Ling.

4 recordsLinked to original sources

IndexToolkit: an open source toolbox to index protein databases for high-throughput proteomics.

UNLABELLED: A software package, IndexToolkit, aimed at overcoming the disadvantage of FASTA-format databases for frequent searching, is developed to utilize an indexing strategy to substantially accelerate sequence queries. IndexToolkit includes user-friendly tools and an Application Programming Interface (API) to facilitate indexing, storage and retrieval of protein sequence databases. As open source, it provides a sequence-retrieval developing framework, which is easily extensible for high-speed-request proteomic applications, such as database searching or modification discovering. We applied IndexToolkit to database searching engine pFind to demonstrate its effect. Experimental studies show that IndexToolkit is able to support significantly faster searches of protein database. AVAILABILITY: The IndexToolkit is free to use under the open source GNU GPL license. The source code and the compiled binary can be freely accessed through the website http://pfind.jdl.ac.cn/IndexToolkit. In this website, the more detailed information including screenshots and documentations for users and developers is also available.

Database Management Systems↗

Customized generalization of support patterns for classification.

We propose a novel classification learning method called customized support pattern learner (CSPL). Given an instance to be classified, CSPL explores and discovers support patterns (SPs), which are essentially attribute value subsets of the instance to be classified. The final prediction of the class label is performed by combining some statistics of the discovered useful SPs. One advantage of the CSPL method is that it can explore a richer hypothesis space and discover useful classification patterns involving attribute values with almost indistinguishable information gain. The customized learning characteristic also allows that the target class can vary for different instances to be classified. It facilitates extremely easy training instance maintenance and updates. We have evaluated our method with real-world problems and benchmark data sets. The results demonstrate that CSPL can achieve good performance and high reliability.

Journal Article↗

pFind: a novel database-searching software system for automated peptide and protein identification via tandem mass spectrometry.

SUMMARY: Research in proteomics requires powerful database-searching software to automatically identify protein sequences in a complex protein mixture via tandem mass spectrometry. In this paper, we describe a novel database-searching software system called pFind (peptide/protein Finder), which employs an effective peptide-scoring algorithm that we reported earlier. The pFind server is implemented with the C++ STL, .Net and XML technologies. As a result, high speed and good usability of the software are achieved.

Algorithms↗

Exploiting the kernel trick to correlate fragment ions for peptide identification via tandem mass spectrometry.

MOTIVATION: The correlation among fragment ions in a tandem mass spectrum is crucial in reducing stochastic mismatches for peptide identification by database searching. Until now, an efficient scoring algorithm that considers the correlative information in a tunable and comprehensive manner has been lacking. RESULTS: This paper provides a promising approach to utilizing the correlative information for improving the peptide identification accuracy. The kernel trick, rooted in the statistical learning theory, is exploited to address this issue with low computational effort. The common scoring method, the tandem mass spectral dot product (SDP), is extended to the kernel SDP (KSDP). Experiments on a dataset reported previously demonstrate the effectiveness of the KSDP. The implementation on consecutive fragments shows a decrease of 10% in the error rate compared with the SDP. Our software tool, pFind, using a simple scoring function based on the KSDP, outperforms two SDP-based software tools, SEQUEST and Sonar MS/MS, in terms of identification accuracy. SUPPLEMENTARY INFORMATION: http://www.jdl.ac.cn/user/yfu/pfind/index.html

Algorithms↗