PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49Linked to original sources

Prediction of beta-turns in proteins from multiple alignment using neural network.

A neural network-based method has been developed for the prediction of beta-turns in proteins by using multiple sequence alignment. Two feed-forward back-propagation networks with a single hidden layer are used where the first-sequence structure network is trained with the multiple sequence alignment in the form of PSI-BLAST-generated position-specific scoring matrices. The initial predictions from the first network and PSIPRED-predicted secondary structure are used as input to the second structure-structure network to refine the predictions obtained from the first net. A significant improvement in prediction accuracy has been achieved by using evolutionary information contained in the multiple sequence alignment. The final network yields an overall prediction accuracy of 75.5% when tested by sevenfold cross-validation on a set of 426 nonhomologous protein chains. The corresponding Q(pred), Q(obs), and Matthews correlation coefficient values are 49.8%, 72.3%, and 0.43, respectively, and are the best among all the previously published beta-turn prediction methods. The Web server BetaTPred2 (http://www.imtech.res.in/raghava/betatpred2/) has been developed based on this approach.

Algorithms↗

The estimation of statistical parameters for local alignment score distributions.

The distribution of optimal local alignment scores of random sequences plays a vital role in evaluating the statistical significance of sequence alignments. These scores can be well described by an extreme-value distribution. The distribution's parameters depend upon the scoring system employed and the random letter frequencies; in general they cannot be derived analytically, but must be estimated by curve fitting. For obtaining accurate parameter estimates, a form of the recently described 'island' method has several advantages. We describe this method in detail, and use it to investigate the functional dependence of these parameters on finite-length edge effects.

Algorithms↗

Expression of chicken hepatic type I and type III iodothyronine deiodinases during embryonic development.

In embryonic chicken liver (ECL) two types of iodothyronine deiodinases are expressed: D1 and D3. D1 catalyzes the activation as well as the inactivation of thyroid hormone by outer and inner ring deiodination, respectively. D3 only catalyzes inner ring deiodination. D1 and D3 have been cloned from mammals and amphibians and shown to contain a selenocysteine (Sec) residue. We characterized chicken D1 and D3 complementary DNAs (cDNAs) and studied the expression of hepatic D1 and D3 messenger RNAs (mRNAs) during embryonic development. Oligonucleotides based on two amino acid sequences strongly conserved in the different deiodinases (NFGSCTSecP and YIEEAH) were used for reverse transcription-PCR of poly(A+) RNA isolated from embryonic day 17 (E17) chicken liver, resulting in the amplification of two 117-bp DNA fragments. Screening of an E17 chicken liver cDNA library with these probes led to the isolation of two cDNA clones, ECL1711 and ECL1715. The ECL1711 clone was 1360 bp long and lacked a translation start site. Sequence alignment showed that it shared highest sequence identity with D1s from other vertebrates and that the coding sequence probably lacked the first five nucleotides. An ATG start codon was engineered by site-directed mutagenesis, generating a mutant (ECL1711M) with four additional codons (coding for MGTR). The open reading frame of ECL1711M coded for a 249-amino acid protein showing 58-62% identity with mammalian D1s. An in-frame TGA codon was located at position 127, which is translated as Sec in the presence ofa Sec insertion sequence (SECIS) identified in the 3'-untranslated region. Enzyme activity expressed in COS-1 cells by transfection with ECL1711M showed the same catalytic, substrate, and inhibitor specificities as native chicken D1. The ECL1715 clone was 1366 bp long and also lacked a translation start site. Sequence alignment showed that it was most homologous with D3 from other species and that the coding sequence lacked approximately the first 46 nucleotides. The deduced amino acid sequence showed 62-72% identity with the D3 sequences from other species, including a putative Sec residue at a corresponding position. The 3'-untranslated region of ECL1715 also contained a SECIS element. These results indicate that ECL1711 and ECL1715 are near-full-length cDNA clones for chicken D1 and D3 selenoproteins, respectively. The ontogeny of D1 and D3 expression in chicken liver was studied between E14 and 1 day after hatching (C1). D1 activity showed a gradual increase from E14 until C1, whereas D1 mRNA level remained relatively constant. D3 activity and mRNA level were highly significantly correlated, showing an increase from E14 to E17 and a strong decrease thereafter. These results suggest that the regulation of chicken hepatic D3 expression during embryonic development occurs predominantly at the pretranslational level.

Amino Acid Sequence↗

Reduced space hidden Markov model training.

MOTIVATION: Complete forward-backward (Baum-Welch) hidden Markov model training cannot take advantage of the linear space, divide-and-conquer sequence alignment algorithms because of the examination of all possible paths rather than the single best path. RESULTS: This paper discusses the implementation and performance of checkpoint-based reduced space sequence alignment in the SAM hidden Markov modeling package. Implementation of the checkpoint algorithm reduced memory usage from O(mn) to O (m square root n) with only a 10% slowdown for small m and n, and vast speed-up for the larger values, such as m = n = 2000, that cause excessive paging on a 96 Mbyte workstation. The results are applicable to other types of dynamic programming. AVAILABILITY: A World-Wide Web server, as well as information on obtaining the Sequence Alignment and Modeling (SAM) software suite, can be found at http://www.cse.ucsc. edu/research/compbio/sam.html. CONTACT: rph@cse.ucsc.edu

Algorithms↗

Crystal structure of the dihaem cytochrome c4 from Pseudomonas stutzeri determined at 2.2A resolution.

BACKGROUND: . Cytochromes c4 are dihaem cytochromes c found in a variety of bacteria. They are assumed to take part in the electron-transport systems associated with both aerobic and anaerobic respiration. The cytochrome c4 proteins are located in the periplasm, predominantly bound to the inner membrane, and are able to transfer electrons between membrane-bound reduction systems and terminal oxidases. Alignment of cytochrome c4 sequences from three bacteria, Pseudomonas aeruginosa, Pseudomonas stutzeri and Azotobacter vinelandii, suggests that these dihaem proteins are composed of two similar domains. Two distinctly different redox potentials have been measured for the Ps. stutzeri cytochrome c4, however. RESULTS: . The crystal structure of the dihaem cytochrome c4 from Ps. stutzeri has been determined to 2.2A resolution by isomorphous replacement. The model, consisting of two entire cytochrome c4 molecules and 138 water molecules in the asymmetric unit, was refined to an R value of 20.1% for all observations in the resolution range 8-2.2A. The molecule is organized in two cytochrome c-like domains that are related by a pseudo-twofold axis. The symmetry is virtually perfectly close to the twofold axis, which passes through a short hydrogen bond between the two haem propionic acid groups, connecting the redox centre of each domain. This haem-haem interaction is further stabilized by an extensive symmetrical hydrogen-bond network. The twofold symmetry is not present further away from the axis, however, and the cytochrome c4 molecule can be considered to be a dipole with charged residues unevenly distributed between the two domains. The haem environment in the two domains show pronounced differences, mainly on the methionine side of the haem group. CONCLUSIONS: . The structure, in conjunction with sequence alignment, suggests that the cytochrome protein has evolved by duplication of a cytochrome c gene. The difference in charge distribution around each haem group in the two domains allows the haem group in the N-terminal domain to be associated with the lower redox potential of 241 mV and the C-terminal haem group with the higher potential of 328 mV. The molecular dipole characteristic of cytochrome c4 is important for its interaction with, and recognition of, its redox partners. In cytochrome c4, the hydrogen-bond network (between residues that are conserved in all known cytochrome c4 subspecies) seems to provide an efficient pathway for an intramolecular electron transfer that can ensure cooperativity between the two redox centres. The C-pyrrole corners of the haem edges are potential sites for external electron exchange.

Amino Acid Sequence↗

Definition of the tempo of sequence diversity across an alignment and automatic identification of sequence motifs: Application to protein homologous families and superfamilies.

It is often possible to identify sequence motifs that characterize a protein family in terms of its fold and/or function from aligned protein sequences. Such motifs can be used to search for new family members. Partitioning of sequence alignments into regions of similar amino acid variability is usually done by hand. Here, I present a completely automatic method for this purpose: one that is guaranteed to produce globally optimal solutions at all levels of partition granularity. The method is used to compare the tempo of sequence diversity across reliable three-dimensional (3D) structure-based alignments of 209 protein families (HOMSTRAD) and that for 69 superfamilies (CAMPASS). (The mean alignment length for HOMSTRAD and CAMPASS are very similar.) Surprisingly, the optimal segmentation distributions for the closely related proteins and distantly related ones are found to be very similar. Also, optimal segmentation identifies an unusual protein superfamily. Finally, protein 3D structure clues from the tempo of sequence diversity across alignments are examined. The method is general, and could be applied to any area of comparative biological sequence and 3D structure analysis where the constraint of the inherent linear organization of the data imposes an ordering on the set of objects to be clustered.

Amino Acid Motifs↗

Classification and phylogenetic analysis of the cAMP-dependent protein kinase regulatory subunit family.

The members of the PKA regulatory subunit family (PKA-R family) were analyzed by multiple sequence alignment and clustering based on phylogenetic tree construction. According to the phylogenetic trees generated from multiple sequence alignment of the complete sequences, the PKA-R family was divided into four subfamilies (types I to IV). Members of each subfamily were exclusively from animals (types I and II), fungi (type III), and alveolates (type IV). Application of the same methodology to the cAMP-binding domains, and subsequently to the region delimited by beta-strands 6 and 7 of the crystal structures of bovine RIalpha and rat RIIbeta (the phosphate-binding cassette; PBC), proved that this highly conserved region was enough to classify unequivocally the members of the PKA-R family. A single signature sequence, F-G-E-[LIV]-A-L-[LIMV]-x(3)-[PV]-R-[ANQV]-A, corresponding to the PBC was identified which is characteristic of the PKA-R family and is sufficient to distinguish it from other members of the cyclic nucleotide-binding protein superfamily. Specific determinants for the A and B domains of each R-subunit type were also identified. Conserved residues defining the signature motif are important for interaction with cAMP or for positioning the residues that directly interact with cAMP. Conversely, residues that define subfamilies or domain types are not conserved and are mostly located on the loop that connects alpha-helix B' and beta strand 7.

Amino Acid Sequence↗

Three-dimensional cluster analysis identifies interfaces and functional residue clusters in proteins.

Three-dimensional cluster analysis offers a method for the prediction of functional residue clusters in proteins. This method requires a representative structure and a multiple sequence alignment as input data. Individual residues are represented in terms of regional alignments that reflect both their structural environment and their evolutionary variation, as defined by the alignment of homologous sequences. From the overall (global) and the residue-specific (regional) alignments, we calculate the global and regional similarity matrices, containing scores for all pairwise sequence comparisons in the respective alignments. Comparing the matrices yields two scores for each residue. The regional conservation score (C(R)(x)) defines the conservation of each residue x and its neighbors in 3D space relative to the protein as a whole. The similarity deviation score (S(x)) detects residue clusters with sequence similarities that deviate from the similarities suggested by the full-length sequences. We evaluated 3D cluster analysis on a set of 35 families of proteins with available cocrystal structures, showing small ligand interfaces, nucleic acid interfaces and two types of protein-protein interfaces (transient and stable). We present two examples in detail: fructose-1,6-bisphosphate aldolase and the mitogen-activated protein kinase ERK2. We found that the regional conservation score (C(R)(x)) identifies functional residue clusters better than a scoring scheme that does not take 3D information into account. C(R)(x) is particularly useful for the prediction of poorly conserved, transient protein-protein interfaces. Many of the proteins studied contained residue clusters with elevated similarity deviation scores. These residue clusters correlate with specificity-conferring regions: 3D cluster analysis therefore represents an easily applied method for the prediction of functionally relevant spatial clusters of residues in proteins.

Adenosine Triphosphate↗

VIR: a computational tool for analysis of immunoglobulin sequences.

In this paper a microcomputer software named VIR (Variable domains of the Immune Receptors) is reported. This package can be used in sequence studies of immunoglobulin variable domains. The main features of the VIR software in the sequences management are: (1) ease of information recovery/extraction from amino acid sequences; and (2) its capability to obtain multiple sequence alignments with predefined characteristics (i.e. specie and/or specificity). As an analytical tool, the VIR package employs such multiple sequence alignments to compute: (1) tables showing amino acid frequencies; (2) three variability indexes; (3) identity matrices; (4) random samples; and (5) sequences with possible canonical structures. Thus the software reported here is proposed as a useful tool to carry out detailed studies of immunoglobulin variable domains.

Amino Acid Sequence↗

Choosing the best heuristic for seeded alignment of DNA sequences.

BACKGROUND: Seeded alignment is an important component of algorithms for fast, large-scale DNA similarity search. A good seed matching heuristic can reduce the execution time of genomic-scale sequence comparison without degrading sensitivity. Recently, many types of seed have been proposed to improve on the performance of traditional contiguous seeds as used in, e.g., NCBI BLASTN. Choosing among these seed types, particularly those that use information besides the presence or absence of matching residue pairs, requires practical guidance based on a rigorous comparison, including assessment of sensitivity, specificity, and computational efficiency. This work performs such a comparison, focusing on alignments in DNA outside widely studied coding regions. RESULTS: We compare seeds of several types, including those allowing transition mutations rather than matches at fixed positions, those allowing transitions at arbitrary positions ("BLASTZ" seeds), and those using a more general scoring matrix. For each seed type, we use an extended version of our Mandala seed design software to choose seeds with optimized sensitivity for various levels of specificity. Our results show that, on a test set biased toward alignments of noncoding DNA, transition information significantly improves seed performance, while finer distinctions between different types of mismatches do not. BLASTZ seeds perform especially well. These results depend on properties of our test set that are not shared by EST-based test sets with a strong bias toward coding DNA. CONCLUSION: Practical seed design requires careful attention to the properties of the alignments being sought. For noncoding DNA sequences, seeds that use transition information, especially BLASTZ-style seeds, are particularly useful. The Mandala seed design software can be found at http://www.cse.wustl.edu/~yanni/mandala/.

Algorithms↗

Indel seeds for homology search.

We are interested in detecting homologous genomic DNA sequences with the goal of locating approximate inverted, interspersed, and tandem repeats. Standard search techniques start by detecting small matching parts, called seeds, between a query sequence and database sequences. Contiguous seed models have existed for many years. Recently, spaced seeds were shown to be more sensitive than contiguous seeds without increasing the random hit rate. To determine the superiority of one seed model over another, a model of homologous sequence alignment must be chosen. Previous studies evaluating spaced and contiguous seeds have assumed that matches and mismatches occur within these alignments, but not insertions and deletions (indels). This is perhaps appropriate when searching for protein coding sequences (<5% of the human genome), but is inappropriate when looking for repeats in the majority of genomic sequence where indels are common. In this paper, we assume a model of homologous sequence alignment which includes indels and we describe a new seed model, called indel seeds, which explicitly allows indels. We present a waiting time formula for computing the sensitivity of an indel seed and show that indel seeds significantly outperform contiguous and spaced seeds when homologies include indels. We discuss the practical aspect of using indel seeds and finally we present results from a search for inverted repeats in the dog genome using both indel and spaced seeds.

Algorithms↗

Functional mapping of cannabinoid receptor homologs in mammals, other vertebrates, and invertebrates.

Over the past decade, several putative homologs of cannabinoid receptors (CBRs) have been identified by homology screening. Homology screening utilizes sequence alignment search engines to recognize homologs. We investigated these putative CBR homologs further by 'functional mapping' of their deduced amino acid sequences. The entire pharmacophore of a CBR has not yet been elucidated, but point-mutation studies have identified over 20 amino acid residues that impart CBR specificity for ligand recognition and/or signal transduction. Twenty point-mutation studies were used to construct a CBR functionality matrix. Sixteen putative CBR homologs were then mapped over the matrix. Several putative homologs did not hold up to this analysis: human GPR3, GPR6, GPR12, and Caenorhabditis elegans C02H7.2 expressed a series of crippling substitutions in the matrix, strongly suggesting they do not encode functional CBRs. Mapping the contested leech (Hirudo medicinalis) CBR sequence suggests that it encodes a functional CB1; it expresses fewer substitutions than the sea squirt (Ciona intestinalis) CB1 sequence. Mapping a putative CB2 ortholog in the puffer fish (Fugu rubripes T012234) suggests it may encode a CBR other than CB2. These findings are consistent with the lack of experimental data proving these putative CBRs have affinity for cannabinoid ligands. Matrix analysis also reveals that SR144528, a 'CB2-specific' synthetic antagonist, has affinity for non-mammalian CB1 receptors, and that L3.45 appears to be CB2-specific, its cognate in CB1 receptors is F3.45. In conclusion, functional mapping, utilizing point-mutation studies, may improve the specificity of homology screening performed by sequence alignment search engines.

Amino Acid Sequence↗

Utilization of the relative complexity measure to construct a phylogenetic tree for fungi.

The relative complexity measure (RCM) is a new approach to evaluate relatedness of DNA sequences which eliminates the requirement to align sequences prior to analysis, a step required with standard reference methods. The value of the RCM approach to generate distance matrices for use in phylogenetic analysis of organisms has not been determined. This study compared RCM with the algorithmic and tree searching reference methods for phylogenetic analysis using fungal sequences. Sequences of the cytochrome b gene and the 18S rDNA gene were obtained from the GenBank database to determine feasibility of this method for phylogenetic relatedness. The RCM approach was also used to construct a phylogenetic tree using internal transcribed spacer (ITS) sequences from 23 medically relevant fungal species. The robustness of the RCM and reference approaches was determined by comparing the topology of seven medically relevant fungi within the phylogenetic trees generated after progressive removal of 10, 20, 30, 40 and 50% of the nucleotide bases from either the 5' or 3' end of the three genomic target sequences. The results demonstrated that the RCM method was equivalent to the reference methods for construction of phylogenetic trees from cytochrome b and 18S rDNA gene sequences. The phylogenetic tree constructed using the ITS sequence generated no contradictory topology. The RCM generated trees retained the appropriate topology after removal of up to 50% of the cytochrome b sequence, 40% of the ITS sequence, and 30% of the 18S gene target sequence. Comparatively, the reference methods failed to maintain topology after only a 10% sequence deletion for each genomic target. The results showed the RCM to be a reliable and robust computational approach for use in the construction of fungal phylogenetic trees without the requirement for prior sequence alignment.

Algorithms↗

PRIMEGENS: robust and efficient design of gene-specific probes for microarray analysis.

MOTIVATION: DNA microarray is a powerful high-throughput tool for studying gene function and regulatory networks. Due to the problem of potential cross hybridization, using full-length genes for microarray construction is not appropriate in some situations. A bioinformatic tool, PRIMEGENS, has recently been developed for the automatic design of PCR primers using DNA fragments that are specific to individual open reading frames (ORFs). RESULTS: PRIMEGENS first carries out a BLAST search for each target ORF against all other ORFs of the genome to quickly identify possible homologous sequences. Then it performs optimal sequence alignment between the target ORF and each of its homologous ORFs using dynamic programming. PRIMEGENS uses the sequence alignments to select gene- specific fragments, and then feeds the fragments to the Primer3 program to design primer pairs for PCR amplification. PRIMEGENS can be run from the command line on Unix/Linux platforms as a stand-alone package or it can be used from a Web interface. The program runs efficiently, and it takes a few seconds per sequence on a typical workstation. PCR primers specific to individual ORFs from Shewanella oneidensis MR-1 and Deinococcus radiodurans R1 have been designed. The PCR amplification results indicate that this method is very efficient and reliable for designing specific probes for microarray analysis.

Algorithms↗

Fold-recognition and comparative modeling of human alpha2,3-sialyltransferases reveal their sequence and structural similarities to CstII from Campylobacter jejuni.

BACKGROUND: The 3-D structure of none of the eukaryotic sialyltransferases (SiaTs) has been determined so far. Sequence alignment algorithms such as BLAST and PSI-BLAST could not detect a homolog of these enzymes from the protein databank. SiaTs, thus, belong to the hard/medium target category in the CASP experiments. The objective of the current work is to model the 3-D structures of human SiaTs which transfer the sialic acid in alpha2,3-linkage viz., ST3Gal I, II, III, IV, V, and VI, using fold-recognition and comparative modeling methods. The pair-wise sequence similarity among these six enzymes ranges from 41 to 63%. RESULTS: Unlike the sequence similarity servers, fold-recognition servers identified CstII, a alpha2,3/8 dual-activity SiaT from Campylobacter jejuni as the homolog of all the six ST3Gals; the level of sequence similarity between CstII and ST3Gals is only 15-20% and the similarity is restricted to well-characterized motif regions of ST3Gals. Deriving template-target sequence alignments for the entire ST3Gal sequence was not straightforward: the fold-recognition servers could not find a template for the region preceding the L-motif and that between the L- and S-motifs. Multiple structural templates were identified to model these regions and template identification-modeling-evaluation had to be performed iteratively to choose the most appropriate templates. The modeled structures have acceptable stereochemical properties and are also able to provide qualitative rationalizations for some of the site-directed mutagenesis results reported in literature. Apart from the predicted models, an unexpected but valuable finding from this study is the sequential and structural relatedness of family GT42 and family GT29 SiaTs. CONCLUSION: The modeled 3-D structures can be used for docking and other modeling studies and for the rational identification of residues to be mutated to impart desired properties such as altered stability, substrate specificity, etc. Several studies in literature have focused on the development of tools and/or servers for the large-scale/automated modeling of 3-D structures of proteins. In contrast, the present study focuses on modeling the 3-D structure of a specific protein of interest to a biochemist and illustrates the associated difficulties. It is also able to establish a sequence/structure relationship between sialyltransferases of two distinct families.

Campylobacter jejuni↗

Hidden Markov models that use predicted secondary structures for fold recognition.

There are many proteins that share the same fold but have no clear sequence similarity. To predict the structure of these proteins, so called "protein fold recognition methods" have been developed. During the last few years, improvements of protein fold recognition methods have been achieved through the use of predicted secondary structures (Rice and Eisenberg, J Mol Biol 1997;267:1026-1038), as well as by using multiple sequence alignments in the form of hidden Markov models (HMM) (Karplus et al., Proteins Suppl 1997;1:134-139). To test the performance of different fold recognition methods, we have developed a rigorous benchmark where representatives for all proteins of known structure are matched against each other. Using this benchmark, we have compared the performance of automatically-created hidden Markov models with standard-sequence-search methods. Further, we combine the use of predicted secondary structures and multiple sequence alignments into a combined method that performs better than methods that do not use this combination of information. Using only single sequences, the correct fold of a protein was detected for 10% of the test cases in our benchmark. Including multiple sequence information increased this number to 16%, and when predicted secondary structure information was included as well, the fold was correctly identified in 20% of the cases. Moreover, if the correct secondary structure was used, 27% of the proteins could be correctly matched to a fold. For comparison, blast2, fasta, and ssearch identifies the fold correctly in 13-17% of the cases. Thus, standard pairwise sequence search methods perform almost as well as hidden Markov models in our benchmark. This is probably because the automatically-created multiple sequence alignments used in this study do not contain enough diversity and because the current generation of hidden Markov models do not perform very well when built from a few sequences.

Amino Acid Sequence↗

Protein structure prediction using a combination of sequence-based alignment, constrained energy minimization, and structural alignment.

We present a novel approach to protein structure prediction in which fold recognition techniques are combined with ab initio folding methods. Based on the predicted secondary structure, one of two different protocols is followed. For mostly alpha proteins, global optimization and sampling of a statistical energy function is used to generate many low-energy structures; these structures are then screened against a fold library. Any structural matches are then selected for further refinement. For proteins predicted to have significant beta-content, sequence and secondary structure-based alignment is used to identify candidate templates; spatial constraints are then extracted from these templates and used, along with the statistical energy function, in the global sampling and optimization program. Successes and failures of both protocols are discussed.

Algorithms↗

Protein homology detection by HMM-HMM comparison.

MOTIVATION: Protein homology detection and sequence alignment are at the basis of protein structure prediction, function prediction and evolution. RESULTS: We have generalized the alignment of protein sequences with a profile hidden Markov model (HMM) to the case of pairwise alignment of profile HMMs. We present a method for detecting distant homologous relationships between proteins based on this approach. The method (HHsearch) is benchmarked together with BLAST, PSI-BLAST, HMMER and the profile-profile comparison tools PROF_SIM and COMPASS, in an all-against-all comparison of a database of 3691 protein domains from SCOP 1.63 with pairwise sequence identities below 20%.Sensitivity: When the predicted secondary structure is included in the HMMs, HHsearch is able to detect between 2.7 and 4.2 times more homologs than PSI-BLAST or HMMER and between 1.44 and 1.9 times more than COMPASS or PROF_SIM for a rate of false positives of 10%. Approximately half of the improvement over the profile-profile comparison methods is attributable to the use of profile HMMs in place of simple profiles. Alignment quality: Higher sensitivity is mirrored by an increased alignment quality. HHsearch produced 1.2, 1.7 and 3.3 times more good alignments ('balanced' score >0.3) than the next best method (COMPASS), and 1.6, 2.9 and 9.4 times more than PSI-BLAST, at the family, superfamily and fold level, respectively.Speed: HHsearch scans a query of 200 residues against 3691 domains in 33 s on an AMD64 2GHz PC. This is 10 times faster than PROF_SIM and 17 times faster than COMPASS.

Algorithms↗