PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Analysis of comparative modeling predictions for CASP2 targets 1, 3, 9, and 17.

Comparative modeling targets 1, 3, 9 and 17 were predicted by alignment of multiple sequences and structures, when available, followed by minimization using the program AMMP. The minimization used improved potentials, and distance restraints for regions of common structure. New prediction procedures were evaluated. Three tested solvent corrections did not significantly improve the predictions. Target 17 had 85.3% sequence identity with the parent and no insertions or deletions. The prediction had a root-mean-square deviation from target 17 of 0.56 A on C alpha atoms, and 0.59 A for the ligand atoms, which verified the accuracy of the minimization. Targets 1, 3, and 9 had 36.4%, 46.7%, and 33.3% identity with the parent sequences, and predictions resulted in root-mean-square deviations for 79-85% of C alpha atoms of 1.49, 1.11, and 1.24 A, respectively. Conformational differences between parent and target crystal structures were difficult to predict. The use of distance restraints and multiple structures improved the positioning of gaps in sequence alignment. Distance restraints did not overcome errors in sequence alignment or ambiguities due to conformational variation in proteins. Predictions for targets 3 and 9 successfully reduced large deviations between parent and target structures.

Animals↗

Automatic assessment of alignment quality.

Multiple sequence alignments play a central role in the annotation of novel genomes. Given the biological and computational complexity of this task, the automatic generation of high-quality alignments remains challenging. Since multiple alignments are usually employed at the very start of data analysis pipelines, it is crucial to ensure high alignment quality. We describe a simple, yet elegant, solution to assess the biological accuracy of alignments automatically. Our approach is based on the comparison of several alignments of the same sequences. We introduce two functions to compare alignments: the average overlap score and the multiple overlap score. The former identifies difficult alignment cases by expressing the similarity among several alignments, while the latter estimates the biological correctness of individual alignments. We implemented both functions in the MUMSA program and demonstrate the overall robustness and accuracy of both functions on three large benchmark sets.

Algorithms↗

Fast protein fold recognition via sequence to structure alignment and contact capacity potentials.

We propose new empirical scoring potentials and associated alignment procedures for optimally aligning protein sequences to protein structures. The method has two main applications: first, the recognition of a plausible fold for a protein sequence of unknown structure out of a database of representative protein structures and, second, the improvement of sequence alignments by using structural information in order to find a better starting point for homology based modelling. The empirical scoring function is derived from an analysis of a nonredundant database of known structures by converting relative frequencies into pseudoenergies using a normalization according to the inverse Bolzmann law. These-so called contact capacity-potentials turn out to be discriminative enough to detect structural folds in the absence of significant sequence similarity and at the same time simple enough to allow for a very fast optimization in an alignment procedure.

Algorithms↗

Sequence diversification of the FK506-binding proteins in several different genomes.

Sequences of FK506-binding proteins (FKBPs) from four genomes of the following organisms were compared: the prokaryote Escherichia coli, the lower eukaryote Saccharomyces cerevisiae, the plant Arabidopsis thaliana, the nematode Caenorhabditis elegans and a composite of 14 unique FKBPs from two mammalian organisms Homo sapiens (man) and Mus musculus (domestic mouse). A singular FK506-like binding domain (FKBD) has about 12 kDa and occurs in the form of archetypal FKBP-12 and as a part of different proteins ranging in size from 13 to 135 kDa. Some organisms may contain a variable number of proteins which consist from two to four consecutively fused FKBDs. In the 12-kDa subgroup of archetypal FKBPs sequence identity (ID) varies from 100 to 83% (mammalian FKBPs-12), 75-50% in mammalian vs. invertebrate FKBPs-12, and fall to about 30% for pairwise sequence comparisons of mammalian and bacterial FKBPs-12 which suggests that their sequences are divergent. Multiple sequence alignment of FKBPs from the four genomes and a set of unique mammalian FKBPs does not contain any explicit consensus sequence but certain sequence positions have conserved physico-chemical characteristics. Variations of hydrophobicity and bulkiness in the multiple sequence alignment are nonsymmetrical because the physico-chemical properties of the aligned sequences changed during evolution. These variations at the sequence positions which are crucial for binding the immunosuppressive macrolide FK506 and peptidyl-prolyl cis/trans isomerase (PPIase) activity are small.

Amino Acid Sequence↗

Database on the structure of large subunit ribosomal RNA.

The Antwerp database on large subunit ribosomal RNA now contains 607 complete or nearly complete aligned sequences. The alignment incorporates secondary structure information for each sequence. Other information about the sequences, such as literature references, accession numbers and taxonomic information is also available. Information from the database can be downloaded or searched on the rRNA WWW Server at URL http://rrna.uia.ac.be/

Animals↗

Use of residue pairs in protein sequence-sequence and sequence-structure alignments.

Two new sets of scoring matrices are introduced: H2 for the protein sequence comparison and T2 for the protein sequence-structure correlation. Each element of H2 or T2 measures the frequency with which a pair of amino acid types in one protein, k-residues apart in the sequence, is aligned with another pair of residues, of given amino acid types (for H2) or in given structural states (for T2), in other structurally homologous proteins. There are four types, corresponding to the k-values of 1 to 4, for both H2 and T2. These matrices were set up using a large number of structurally homologous protein pairs, with little sequence homology between the pair, that were recently generated using the structure comparison program SHEBA. The two scoring matrices were incorporated into the main body of the sequence alignment program SSEARCH in the FASTA package and tested in a fold recognition setting in which a set of 107 test sequences were aligned to each of a panel of 3,539 domains that represent all known protein structures. Six procedures were tested; the straight Smith-Waterman (SW) and FASTA procedures, which used the Blosum62 single residue type substitution matrix; BLAST and PSI-BLAST procedures, which also used the Blosum62 matrix; PASH, which used Blosum62 and H2 matrices; and PASSC, which used Blosum62, H2, and T2 matrices. All procedures gave similar results when the probe and target sequences had greater than 30% sequence identity. However, when the sequence identity was below 30%, a similar structure could be found for more sequences using PASSC than using any other procedure. PASH and PSI-BLAST gave the next best results.

Algorithms↗

Probabilistic divergence measures for detecting interspecies recombination.

This paper proposes a graphical method for detecting interspecies recombination in multiple alignments of DNA sequences. A fixed-size window is moved along a given DNA sequence alignment. For every position, the marginal posterior probability over tree topologies is determined by means of a Markov chain Monte Carlo simulation. Two probabilistic divergence measures are plotted along the alignment, and are used to identify recombinant regions. The method is compared with established detection methods on a set of synthetic benchmark sequences and two real-world DNA sequence alignments.

Computational Biology↗

enoLOGOS: a versatile web tool for energy normalized sequence logos.

enoLOGOS is a web-based tool that generates sequence logos from various input sources. Sequence logos have become a popular way to graphically represent DNA and amino acid sequence patterns from a set of aligned sequences. Each position of the alignment is represented by a column of stacked symbols with its total height reflecting the information content in this position. Currently, the available web servers are able to create logo images from a set of aligned sequences, but none of them generates weighted sequence logos directly from energy measurements or other sources. With the advent of high-throughput technologies for estimating the contact energy of different DNA sequences, tools that can create logos directly from binding affinity data are useful to researchers. enoLOGOS generates sequence logos from a variety of input data, including energy measurements, probability matrices, alignment matrices, count matrices and aligned sequences. Furthermore, enoLOGOS can represent the mutual information of different positions of the consensus sequence, a unique feature of this tool. Another web interface for our software, C2H2-enoLOGOS, generates logos for the DNA-binding preferences of the C2H2 zinc-finger transcription factor family members. enoLOGOS and C2H2-enoLOGOS are accessible over the web at http://biodev.hgen.pitt.edu/enologos/.

Amino Acids↗

Total variation in the penA gene of Neisseria meningitidis: correlation between susceptibility to beta-lactam antibiotics and penA gene heterogeneity.

In recent decades, the prevalence of Neisseria meningitidis isolates with reduced susceptibility to penicillins has increased. The intermediate resistance to penicillin (Pen(i)) for most strains is due mainly to mosaic structures in the penA gene, encoding penicillin-binding protein 2. In this study, susceptibility to beta-lactam antibiotics was determined for 60 Swedish clinical N. meningitidis isolates and 19 reference strains. The penA gene was sequenced and compared to 237 penA sequences from GenBank in order to explore the total identified variation of penA. The divergent mosaic alleles differed by 3% to 24% compared to those of the designated wild-type penA gene. By studying the final 1,143 to 1,149 bp of penA in a sequence alignment, 130 sequence variants were identified. In a 402-bp alignment of the most variable regions, 84 variants were recognized. Good correlation between elevated MICs and the presence of penA mosaic structures was found especially for penicillin G and ampicillin. The Pen(i) isolates comprised an MIC of >0.094 microg/ml for penicillin G and an MIC of >0.064 microg/ml for ampicillin. Ampicillin was the best antibiotic for precise categorization as Pen(s) or Pen(i). In comparison with the wild-type penA sequence, two specific Pen(i) sites were altered in all except two mosaic penA sequences, which were published in GenBank and no MICs of the corresponding isolates were described. In conclusion, monitoring the relationship between penA sequences and MICs to penicillins is crucial for developing fast and objective methods for susceptibility determination. By studying the penA gene, genotypical determination of susceptibility in culture-negative cases can also be accomplished.

Amino Acid Sequence↗

Theseus: fast and optimal affine-gap sequence-to-graph alignment.

MOTIVATION: Sequence-to-graph alignment is a central problem in bioinformatics, with applications in multiple sequence alignment (MSA) and pangenome analysis, among others. However, current algorithms for optimal affine-gap alignment impose high memory and computational requirements, limiting their scalability to aligning long sequences to complex graphs. Practical solutions partially address this problem using heuristic strategies that ultimately trade off optimality for speed. RESULTS: This work presents Theseus, a novel, fast, and optimal affine-gap sequence-to-graph alignment algorithm. Theseus leverages similarities between genomic sequences to accelerate the alignment computation and reduces the overall memory requirements without compromising optimality. To that end, Theseus processes only a subset of the dynamic programming cells, using a sparse-data strategy that enables efficient sequence-to-graph alignment. Moreover, our algorithm supports optimal affine-gap alignment on arbitrary directed graphs, including those with cycles. We evaluate Theseus on two key problems: MSA and pangenome read mapping. For MSA, we compare it against SPOA, abPOA, and POASTA. Theseus is 1.6× to 17.6× faster than POASTA, and 7.3× faster, on average, than SPOA, both optimal aligners. Compared with abPOA, Theseus ensures optimality and scales to the largest problems. For pangenome read mapping, we benchmark Theseus against the alignment stage of the mapping tool vg map, along with the alignment kernels of SPOA, abPOA, and POASTA. Theseus outperforms the other methods, showing a 1.9× to 16.9× speedup on short reads. Moreover, Theseus is 1.5× to 36.3× faster than vg when aligning against synthetic cyclic graphs. AVAILABILITY AND IMPLEMENTATION: Theseus code and documentation are publicly available at https://github.com/albertjimenezbl/theseus-lib.

Algorithms↗

Diversity and relatedness among the type I interferons.

Type I interferons (IFNs) include the IFN-alpha family of subtypes, IFN-beta, IFN-omega, IFN-tau, IFN-kappa, IFN-lambda, and IFN-zeta. IFN genes lack introns and encode secretory signal peptide sequences that are proteolytically cleaved prior to secretion from the cell. In contrast to the approximately 50% amino acid sequence identity among the human IFN-alpha subtypes, human IFN-alphas share approximately 22% identity with human IFN-beta and 37% identity with human IFN-omega. Many of the conserved residues among the type I IFNs are implicated in receptor recognition and structural integrity. This report provides an update on the gene annotations for the mouse and human IFN gene clusters on chromosome 4 and 9, respectively, with accompanying amino acid sequence alignments. Based on sequence identities, a phylogenic tree analysis for the different mammalian Type I IFNs is also presented, showing the high degree of relatedness among these IFNs. Notably, sequence alignment of the different human and mouse IFN promoter regions reveals different signature patterns for transcription factor binding sites, implying different inducers might differentially activate the transcription of the different IFNs.

Amino Acid Sequence↗

CHOP: visualization of 'wobbling' and isolation of highly conserved regions from aligned DNA sequences.

The web software CHOP was developed to visualize the 'wobbling' in the third codon position of aligned DNA sequences. The simple features of this tool allow users to easily find regions suspected of containing coding sequences (CDSs). The program also allows visualization of the nucleotide diversity between two genomic or gene sequences by graphically plotting the percentage identity between the two sequences. CHOP can also isolate highly conserved regions within both CDSs and non-CDSs. Highly conserved regions within CDSs include the regions with lower rates of synonymous substitution in which nucleotide sequences are expected to be under strong selective pressure. CHOP is available at http://bunsei2.med.u-tokai.ac.jp:8080/~ohtsuka/cds_finding.html.

Animals↗

Statistical distributions of optimal global alignment scores of random protein sequences.

BACKGROUND: The inference of homology from statistically significant sequence similarity is a central issue in sequence alignments. So far the statistical distribution function underlying the optimal global alignments has not been completely determined. RESULTS: In this study, random and real but unrelated sequences prepared in six different ways were selected as reference datasets to obtain their respective statistical distributions of global alignment scores. All alignments were carried out with the Needleman-Wunsch algorithm and optimal scores were fitted to the Gumbel, normal and gamma distributions respectively. The three-parameter gamma distribution performs the best as the theoretical distribution function of global alignment scores, as it agrees perfectly well with the distribution of alignment scores. The normal distribution also agrees well with the score distribution frequencies when the shape parameter of the gamma distribution is sufficiently large, for this is the scenario when the normal distribution can be viewed as an approximation of the gamma distribution. CONCLUSION: We have shown that the optimal global alignment scores of random protein sequences fit the three-parameter gamma distribution function. This would be useful for the inference of homology between sequences whose relationship is unknown, through the evaluation of gamma distribution significance between sequences.

Algorithms↗

Protein structure prediction by threading methods: evaluation of current techniques.

This paper evaluates the results of a protein structure prediction contest. The predictions were made using threading procedures, which employ techniques for aligning sequences with 3D structures to select the correct fold of a given sequence from a set of alternatives. Nine different teams submitted 86 predictions, on a total of 21 target proteins with little or no sequence homology to proteins of known structure. The 3D structures of these proteins were newly determined by experimental methods, but not yet published or otherwise available to the predictors. The predictions, made from the amino acid sequence alone, thus represent a genuine test of the current performance of threading methods. Only a subset of all the predictions is evaluated here. It corresponds to the 44 predictions submitted for the 11 target proteins seen to adopt known folds. The predictions for the remaining 10 proteins were not analyzed, although weak similarities with known folds may also exist in these proteins. We find that threading methods are capable of identifying the correct fold in many cases, but not reliably enough as yet. Every team predicts correctly a different set of targets, with virtually all targets predicted correctly by at least one team. Also, common folds such as TIM barrels are recognized more readily than folds with only a few known examples. However, quite surprisingly, the quality of the sequence-structure alignments, corresponding to correctly recognized folds, is generally very poor, as judged by comparison with the corresponding 3D structure alignments. Thus, threading can presently not be relied upon to derive a detailed 3D model from the amino acid sequence. This raises a very intriguing question: how is fold recognition achieved? Our analysis suggests that it may be achieved because threading procedures maximize hydrophobic interactions in the protein core, and are reasonably good at recognizing local secondary structure.

Amino Acid Sequence↗

DISTREE: a tool for estimating genetic distances between aligned DNA sequences.

MOTIVATION: Substitution rates estimated from aligned DNA data can be used as genetic distances to investigate the phylogenetic relationship of those sequences. For this purpose, a Markov model of nucleotide substitution has to be assumed that describes this process most adequately. RESULTS: A program is presented that estimates substitution rates and their standard errors for a variety of Markov models. The model introduced by Hasegawa et al. (J. Mol. Evol., 22, 160-174, 1985) is the only one for which distances and standard deviations need to be calculated numerically, since analytical formulae cannot be derived. Each model is implemented in two different variants: (i) assuming rate homogeneity or (ii) starting from Gamma-distributed substitution rates across sequence sites. The estimation of heterogeneous substitution rates is based on a method suggested by Tamura and Nei (Mol. Biol. Evol., 10, 512-526, 1993). All required parameters are estimated from sequence data, hence the user is not asked to supply any additional input. One goal of the program is to support the user when choosing a particular model that describes most adequately the evolution of the given data set. For this purpose, a more detailed analysis of this model fit is provided. Phylogenetic trees reconstructed from the inferred distances using the neighbor-joining algorithm are also available.

Algorithms↗

A catalytic triad is required by the non-heme haloperoxidases to perform halogenation.

The bacterial non-heme haloperoxidases are highly related to an esterase from Pseudomonas fluorescens, at structural and functional levels. Both types of enzymes displayed brominating activity and esterase activity. The presence of the serine-hydrolase motif Gly-X-Ser-X-Gly, in the esterase as well as in all aligned haloperoxidase sequences, strongly suggested that they belong to the serine-hydrolase family. Sequence alignment with several serine-hydrolases and secondary structure superimposition revealed the striking conservation of structural features characterising the alpha/beta-hydrolase fold structure in all haloperoxidases. These structural predictions allowed us to identify a potential catalytic triad in haloperoxidases, perfectly matching the triad of all aligned serine-hydrolases. The structurally equivalent triad in the chloroperoxidase CPO-P comprised the amino acids Serine 97, Aspartic acid 229 and Histidine 258. The involvement of this catalytic triad in halogenation was further assessed by inhibition studies and site-directed mutagenesis. Inactivation of CPO-P by PMSF and DEPC strongly suggested that the serine residue from the serine-hydrolase motif and an histidine residue are essential for halogenation, similar to that demonstrated for typical serine-hydrolases. By site-directed mutagenesis of CPO-P, Ser-97 was exchanged against alanine or cysteine, Asp-229 against alanine and His-258 against glutamine. Western blot analysis indicated that each mutant gene was efficiently expressed. Whereas the mutant S97C conserved a very low residual activity, each other mutant S97A, D229A or H258Q was totally inactive. This study gives the direct demonstration of the requirement of a catalytic triad in the halogenation mechanism.

Amino Acid Sequence↗

Identification of important functional environs in protein tertiary structures from the analysis of residue variation in 3-D: application to cytochromes c and carboxypeptidases A and B.

A simple methodology is described to apply to aligned protein sequence sets for which at least one representative 3-D C alpha structure is known. The evolutionary variation observed at each residue position in the sequence alignment is qualified by taking into account the residue variation that has occurred at other positions located within 7 A (according to the probable chain fold). This expresses the evolutionary behaviour of any residue position in the more appropriate context of its immediate surroundings and distinguishes between invariant residues on the basis of the variation of their environment. The highest mechanistic significance is attached to conserved residues in conserved surroundings, but the quantitative nature of the analysis means that all residue vicinities can be ranked and merged according to the degree of conservation that they exhibit and the residue positions that comprise them. Therefore, with the aid of the chain fold, contour maps can be constructed that show graded foci of evolutionary conservation in the underlying superstructure of the protein type, and the irregular shapes and extents of large conserved areas. To test the methodology, it was applied to cytochromes c and the carboxypeptidases A and B.

Biological Evolution↗

Function-dependent clustering of orthologues and paralogues of cyclophilins.

The 18 kDa archetypal cyclosporin-A binding protein, cyclophilin-A, has multiple paralogues in the human genome. Only 18 of those paralogues have been detected as mRNAs or proteins whose masses vary from 18 to 354 kDa, whereas the functional significance of the open reading frames (ORFs) encoding other paralogues of cyclophilin-A remains unknown. The genomes of Drosophila melanogaster, Caenorhabditis elegans, Arabidopsis thaliana, Schizosaccharomyces pombe, and Saccharomyces cerevisiae encode different numbers of the cyclophilin paralogues, some of which are orthologous to the human cyclophilins. A library of novel algorithms was developed and used for computation of the conservation levels for hydrophobicity and bulkiness profiles, and amino acid compositions (AACs) of 303 aligned sequences of cyclophilins. The majority of the paralogues and orthologues encoded in these 6 genomes differ considerably from each other. Some of the orthologues and paralogues have high correlation coefficients (CCFs) for pairwise compared hydrophobicity and bulkiness profiles, and whose AACs differ to a low degree. Convergence of these three properties of the polypeptide chain and apparent conservation of the typical sequence hallmarks and parameters allowed for the clustering of the functionally related orthologues and paralogues of the cyclophilins. The clustering method allowed for sorting out the cyclophilins into several distinct classes. Analyses of the overlapping clusters of sequences permitted delineation of some hypothetical pathways that might have led to the creation of certain paralogues of cyclophilins in the eukaryotic genomes.

Amino Acid Sequence↗