PubMed HealthSearch

Biomedical subjects

J Gracy

Publications and source records attributed to J Gracy.

6 recordsLinked to original sources

Automated protein sequence database classification. I. Integration of compositional similarity search, local similarity search, and multiple sequence alignment.

MOTIVATION: Genome sequencing projects require the periodic application of analysis tools that can classify and multiply align related protein sequence domains. Full automation of this task requires an efficient integration of similarity and alignment techniques. RESULTS: We have developed a fully automated process that classifies entire protein sequence databases, resulting in alignment of the homologous sequences. The successive steps of the procedure are based on compositional and local sequence similarity searches followed by multiple sequence alignments. Global similarities are detected from the pairwise comparison of amino acid and dipeptide compositions of each protein. After the elimination of all but one sequence from each detected cluster of closely related proteins, the remaining sequences are compiled in a suffix tree which is self-compared to detect local sequence similarities. Sets of proteins which share similar sequence segments are then weighted according to their closeness and multiply aligned using a fast hierarchical dynamic programming algorithm. Computational strategies were devised to minimize computer processing time and memory space requirements. The accuracy of the sequence classifications has been evaluated for 12 462 primary structures distributed over 341 known families. The percentage of sequences with missed or incorrect family assignments was 6.8% on the test set. This low error level is only twice that of the manually constructed PROSITE database ( 3.4% ) and is substantially better than that found for the automatically built PRODOM database ( 34.9% ). AVAILABILITY: The resulting database, called DOMO, is available through database search routine SRS at Infobiogen (http://www.infobiogen.fr/srs5/), EBI (http://srs.ebi.ac.uk:5000/) and EMBL (http://www.embl-heidelberg.de/srs5/) World Wide Web sites. CONTACT: gracy@infobiogen.fr

Algorithms

Automated protein sequence database classification. II. Delineation Of domain boundaries from sequence similarities.

MOTIVATION: Decomposing each protein into modular domains is a basic prerequisite to classify accurately structural units in biological molecules. Boundaries between domains are indicated by two similar amino acid sequence segments located within the same protein (repeats) or within homologous proteins at notably different distances from their respective N- or C-termini. RESULTS: We have developed an automated method that combines such positional constraints derived from various detected pairwise sequence similarities to delineate the modular organization of proteins. The procedure has been applied to a non-redundant data set of 26 990 proteins whose sequences were taken from the PIR and SWISS-PROT databanks and shared <60% sequence identity amongst pairs. The resultant clustering, delineation and multiple alignment of 24 380 sequence fragments yielded a new database of 4364 domain families. Comparison of the domain collection with that of PRODOM indicates a clear improvement in the number and size of domain families, domain boundaries and multiple sequence alignments. The accuracy and sensitivity of the method are illustrated by results obtained for ankyrin-like repeats and EGF-like modules. AVAILABILITY: The resulting database, called DOMO, is available through the database search routine SRS at Infobiogen (http://www.infobiogen.fr/srs5/), EBI (http://srs.ebi.ac.uk:5000/) and EMBL (http://www.embl-heidelberg.de/srs5/) World Wide Web sites. CONTACT: gracy@infobiogen.fr

Algorithms

Learning and alignment methods applied to protein structure prediction.

Learning techniques are able to extract structural knowledge specific to a selected set of proteins. We describe two algorithms that optimize scores expressing the propensity of a polypeptide sequence to adopt a local fold. The first algorithm generates secondary structure prediction rules based on a dictionary of geometrical patterns frequently found in the learning database. The second algorithm leads to scores that indicate the fit between an amino acid and a given local structural environment. Dynamic programming is then used to align structural information profiles by modifying the local mutation cost with the above learned functions. The main features of the system are exemplified on the structural prediction of the N-terminal domain of the CD4 antigen. Then the usefulness of additional 3-D information in the alignment is benchmarked on eight pairs of weakly homologous proteins.

Algorithms

Improved alignment of weakly homologous protein sequences using structural information.

Protein sequence alignments can be improved when at least one of the proteins to be aligned has a known 3-D structure. In this work, geometrical constraints extracted from the target fold are evaluated in independent units that deal with complementary structural features. This information is used to set up mutation tables specific to the locally observed structural environments. The resulting partial evaluations are then combined linearly into a global function which is optimized by dynamic programming. Eventually, a score based on tertiary interactions can be used as a selection criterion to discriminate among a set of suboptimal alignments. The relevance of the scores given by each unit is tested on a representative set of protein families. Finally, a method for combining the different scores is described and its efficiency is evaluated on a few pairs of weakly homologous proteins.

Algorithms

A modular learning environment for protein modeling.

We propose in this paper a modular learning environment for protein modeling. In this system, the protein modeling problem is tackled in two successive phases. First, partial structural informations are determined via numerical learning techniques. Then, in the second phase, the multiple available informations are combined in pattern matching searches via dynamic programming. It is shown on real problems that various protein structure predictions can be improved in this way, such as secondary structure prediction, alignment of weakly homologous protein sequences or protein model evaluations.

Amino Acid Sequence