PubMed · 17018537
A Novel algorithm for identifying low-complexity regions in a protein sequence.
Abstract
MOTIVATION: We consider the problem of identifying low-complexity regions (LCRs) in a protein sequence. LCRs are regions of biased composition, normally consisting of different kinds of repeats. RESULTS: We define new complexity measures to compute the complexity of a sequence based on a given scoring matrix, such as BLOSUM 62. Our complexity measures also consider the order of amino acids in the sequence and the sequence length. We develop a novel graph-based algorithm called GBA to identify LCRs in a protein sequence. In the graph constructed for the sequence, each vertex corresponds to a pair of similar amino acids. Each edge connects two pairs of amino acids that can be grouped together to form a longer repeat. GBA finds short subsequences as LCR candidates by traversing this graph. It then extends them to find longer subsequences that may contain full repeats with low complexities. Extended subsequences are then post-processed to refine repeats to LCRs. Our experiments on real data show that GBA has significantly higher recall compared to existing algorithms, including 0j.py, CARD, and SEG. AVAILABILITY: The program is available on request.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xuehui Li, Tamer Kahveci. 2006-10-02. A Novel algorithm for identifying low-complexity regions in a protein sequence.. https://doi.org/10.1093/bioinformatics%2Fbtl495
Cite the original work for its findings. Save a collection to share your selection of sources.