PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Sequence alignment between vWF and peptides inhibiting the vWF-collagen interaction does not result in the identification of a collagen-binding site in vWF.

We previously found that two peptides (N- and Q-peptide) selected by phage display for binding to an anti-vWF antibody, were able to inhibit vWF-binding to collagen (1). The sequence of those peptides could be aligned with the sequence in vWF at position 1129-1136 just outside the A3-domain. As the peptides represent an epitope or mimotope of vWF for binding to collagen we next wanted to study whether the alignment resulted in the identification of a new collagen binding site in vWF. We mutated the 1129-1136 VWTLPDQC sequence in vWF to VATAPAAC. Expressing this mutant vWF (7.8-vWF) in a fur-BHK cell line resulted in well processed 7.8-vWF containing a normal distribution of molecular weight multimers. However, binding studies of this mutant vWF to rat tail, human and calf skin collagens type 1, to human collagen types III and VI, revealed no decrease in vWF-binding to any of these collagens. Thus, although the N- and Q-peptides did inhibit the vWF-collagen interaction, the resulting alignment with the vWF sequence did not identify a collagen binding site, pointing out that alignments (although with a high percentage of identity) do not always result in identification of binding epitopes. However, suprisingly removal of the A3-domain or changing the vWF sequence at position 1129-1136 resulted in an increase of vWF-binding to human collagen type V1 and to rat tail collagen type 1, implying that these changes result in a different conformation of vWF with an increased binding to these collagens as a consequence.

Animals↗

PIR-ALN: a database of protein sequence alignments.

MOTIVATION: The Protein Information Resource (PIR) maintains a database of annotated and curated alignments in order to visually represent interrelationships among sequences in the PIR-International Protein Sequence Database, to spread and standardize protein names, features and keywords among members of a family or superfamily, and to aid us in classifying sequences, in identifying conserved regions, and in defining new homology domains. RESULTS: Release 22.0, (December 1998), of the PIR-ALN database contains a total of 3806 alignments, including 1303 superfamily, 2131 family and 372 homology domain alignments. This is an appropriate dataset to develop and extract patterns, test profiles, train neural networks or build Hidden Markov Models (HMMs). These alignments can be used to standardize and spread annotation to newer members by homology, as well as to understand the modular architecture of multidomain proteins. PIR-ALN includes 529 alignments that can be used to develop patterns not represented in PROSITE, Blocks, PRINTS and Pfam databases. The ATLAS information retrieval system can be used to browse and query the PIR-ALN alignments. AVAILABILITY: PIR-ALN is currently being distributed as a single ASCII text file along with the title, member, species, superfamily and keyword indexes. The quarterly and weekly updates can be accessed via the WWW at pir.georgetown.edu. The quarterly updates can also be obtained by anonymous FTP from the PIR FTP site at NBRF.Georgetown.edu, directory [ANONYMOUS.PIR.ALIGNMENT].

Amino Acid Sequence↗

Sequence alignment: an approximation law for the Z-value with applications to databank scanning.

The Z-value is an attempt to estimate the statistical significance of a Smith and Waterman dynamic programming alignment score (H-score) through the use of a Monte-Carlo procedure. In this paper, we give an approximation for the Z-value law deduced from the Poisson clumping heuristic developed by Waterman and Vingron (Stat. Sci. 9 (1994) 367) in the case of independent and identically distributed sequences comparison. As for non-gapped alignment scores, our approximation is of Gumbel type but with parameters that are sequence independent. This result makes clear the related experimental results mentioned by Comet et al. (Comput. Chem. 23 (1999) 317). Using 'quasi-real' sequences (i.e. randomly shuffled sequences of the same length and amino acid composition as the real ones) we investigate the relevance of our approximation result. Since the Monte-Carlo approach we use generates a bias for the Gumbel decay parameter estimation, a correction procedure is proposed. Applications to real sequences are considered and we show how our results can be used to detect the potential biological relationships between real sequences.

Computing Methodologies↗

The Rieske protein: a case study on the pitfalls of multiple sequence alignments and phylogenetic reconstruction.

Previously published phylogenetic trees reconstructed on "Rieske protein" sequences frequently are at odds with each other, with those of other subunits of the parent enzymes and with small-subunit rRNA trees. These differences are shown to be at least partially if not completely due to problems in the reconstruction procedures. A major source of erroneous Rieske protein trees lies in the presence of a large, poorly conserved domain prone to accommodate very long insertions in well-defined structural hot spots substantially hampering multiple alignments. The remaining smaller domain, in contrast, is too conserved to allow distant phylogenies to be deduced with sufficient confidence. Three-dimensional structures of representatives from this protein family are now available from phylogenetically distant species and from diverse enzymes. Multiple alignments can thus be refined on the basis of these structures. We show that structurally guided alignments of Rieske proteins from Rieske-cytochrome b complexes and arsenite oxidases strongly reduce conflicts between resulting trees and those obtained on their companion enzyme subunits. Further problems encountered during this work, mainly consisting in database errors such as wrong annotations and frameshifts, are described. The obtained results are discussed against the background of hypotheses stipulating pervasive lateral gene transfer in prokaryotes.

Algorithms↗

Improvement of the A(*) Algorithm for Multiple Sequence Alignment.

The alignment problem of DNA or protein sequences is very applicable and important in various fields of molecular biology. This problem can be reduced to the shortest path problem and Ikeda and Imai (Genome Informatics 5: 90-99, 1994) showed that the A(*) algorithm works efficiently with the estimator utilizing all 2-dimensional sub-alignments. In this paper we present new powerful estimators utilizing k >/= 3 dimensional sub-alignments, and propose a new bounding technique using V (Delta), a set of vertices in the paths whose lengths are at most Delta longer than the shortest path. We also extend our algorithm to a recursive-estimate version. These algorithms become more efficient when the number of sequences increase, or the similarity among sequences is lower.

Journal Article↗

Determination of reliable regions in protein sequence alignments.

Judging the significance of alignments is still a major problem in sequence comparison. We present a method to delineate reliable regions within an alignment. This differs from standard approaches in that it does not attempt to attribute one significance value to the alignment as a whole, but assesses alignment quality locally. An algorithm is provided that predicts which residue pairs in an alignment are likely to be correctly matched. The predictions are evaluated by comparison with alignments taken from tertiary structural superpositions.

Algorithms↗

Approximate p-values for local sequence alignments: numerical studies.

Siegmund and Yakir (2000) have given an approximate p-value when two independent, identically distributed sequences from a finite alphabet are optimally aligned based on a scoring system that rewards similarities according to a general scoring matrix and penalizes gaps (insertions and deletions). The approximation involves an infinite sequence of difficult-to-compute parameters. In this paper, it is shown by numerical studies that these reduce to essentially two numerically distinct parameters, which can be computed as one-dimensional numerical integrals. For an arbitrary scoring matrix and affine gap penalty, this modified approximation is easily evaluated. Comparison with published numerical results show that it is reasonably accurate.

Models, Molecular↗

Pfam: multiple sequence alignments and HMM-profiles of protein domains.

Pfam contains multiple alignments and hidden Markov model based profiles (HMM-profiles) of complete protein domains. The definition of domain boundaries, family members and alignment is done semi-automatically based on expert knowledge, sequence similarity, other protein family databases and the ability of HMM-profiles to correctly identify and align the members. Release 2.0 of Pfam contains 527 manually verified families which are available for browsing and on-line searching via the World Wide Web in the UK at http://www.sanger.ac.uk/Pfam/ and in the US at http://genome.wustl. edu/Pfam/ Pfam 2.0 matches one or more domains in 50% of Swissprot-34 sequences, and 25% of a large sample of predicted proteins from the Caenorhabditis elegans genome.

Amino Acid Sequence↗

FFAS03: a server for profile--profile sequence alignments.

The FFAS03 server provides a web interface to the third generation of the profile-profile alignment and fold-recognition algorithm of fold and function assignment system (FFAS) [L. Rychlewski, L. Jaroszewski, W. Li and A. Godzik (2000), Protein Sci., 9, 232-241]. Profile-profile algorithms use information present in sequences of homologous proteins to amplify the patterns defining the family. As a result, they enable detection of remote homologies beyond the reach of other methods. FFAS, initially developed in 2000, is consistently one of the best ranked fold prediction methods in the CAFASP and LiveBench competitions. It is also used by several fold-recognition consensus methods and meta-servers. The FFAS03 server accepts a user supplied protein sequence and automatically generates a profile, which is then compared with several sets of sequence profiles of proteins from PDB, COG, PFAM and SCOP. The profile databases used by the server are automatically updated with the latest structural and sequence information. The server provides access to the alignment analysis, multiple alignment, and comparative modeling tools. Access to the server is open for both academic and commercial researchers. The FFAS03 server is available at http://ffas.burnham.org.

Algorithms↗

Estimation and reliability of molecular sequence alignments.

The problem of estimating the relatedness of a pair of biological sequences is addressed. A stochastic model of sequence evolution is described that allows insertion and deletion as well as replacement of amino acid residues (or substitution of nucleotides) over time. An expectation-maximization (EM) algorithm that obtains maximum likelihood estimates of the model parameters is introduced. The method assumes that the sequences are related by descent from a common ancestor but the alignment (i.e., the precise evolutionary correspondence between residues in each sequence) is unknown. Results from the E-step of the EM algorithm are used to assess the likelihood that any two residues are related by direct descent from a common ancestor.

Algorithms↗