PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Microsatellite evolution inferred from human-chimpanzee genomic sequence alignments.

Most studies of microsatellite evolution utilize long, highly mutable loci, which are unrepresentative of the majority of simple repeats in the human genome. Here we use an unbiased sample of 2,467 microsatellite loci derived from alignments of 5.1 Mb of genomic sequence from human and chimpanzee to investigate the mutation process of tandemly repetitive DNA. The results indicate that the process of microsatellite evolution is highly heterogeneous, exhibiting differences between loci of different lengths and motif sizes and between species. We find a highly significant tendency for human dinucleotide repeats to be longer than their orthologues in chimpanzees, whereas the opposite trend is observed in mononucleotide repeat arrays. Furthermore, the rate of divergence between orthologues is significantly higher at longer loci, which also show significantly greater mutability per repeat number. These observations have important consequences for understanding the molecular mechanisms of microsatellite mutation and for the development of improved measures of genetic distance.

Animals↗

Sequence alignment and structural modelling of the LamB glycoporin family.

lamB gene segments were obtained from Yersinia enterocolitica and Vibrio parahaemolyticus by the PCR and the DNA sequence determined. The deduced polypeptide sequences showed high similarity to six other LamB-related proteins and all contained typical signature sequences present in all members of the family but not other proteins. The aligned amino acid sequences permitted derivation of a model of LamB folding across the bacterial outer membrane using an approach successfully applied in the identification of structural features in other porins (Ferenci,T. (1994) Mol. Microbiol. 14:188-189). The alignment-based model differs from previous LamB structure predictions and is also more complex than that found for OmpF-related porins; more than 16 conserved stretches of amino acid sequence potentially corresponded to membrane-spanning segments.

Amino Acid Sequence↗

Sequence alignments of the H(+)-dependent oligopeptide transporter family PTR: inferences on structure and function of the intestinal PET1 transporter.

PURPOSE: To study the structure and function of the intestinal H+/ peptide transporter PET1, we compared its amino acid sequence with those of related transporters belonging to the oligopeptide transporter family PTR, and with more distant transporter families. METHODS: We have developed a new approach to the sequence analysis of proteins with multiple transmembrane domains (TMDs) which takes into account the repeated TMD-loop topology. In addition to conventional analyses of the entire sequence, each TMD and its adjacent loop residues (= TMD segments) were analyzed separately as independent structural units. In combination with hydropathy analysis, this approach reveals any changes in the order of the TMD segments in the primary structure and permits TMD alignments among divergent structures even if rearrangements of the order of TMD segments have occurred in the course of evolution. RESULTS: Alignments of TMD segments indicate that the TMD order in PTR transporters may have changed in the process of evolution. Consideration of such changes permits the alignment of homologous TMD segments from PTR transporters belonging to distant akaryotic and eukaryotic phyla. Multiple alignments of TMDs reveal several highly conserved regions that may play a role in transporter function. In comparing the PTR transporters with other transporter gene families, alignment scores using the entire primary structure are too low to support a finding of probable homology. However, statistically significant alignments were observed among individual TMD segments if one disregards the order in which they occur in the primary structure. CONCLUSIONS: Our results support the hypothesis that the PTR transporters may have evolved by rearrangement, duplication, or insertions and deletions of TMD segments as independent modules. This modular structure suggests new alignment strategies for determining functional domains and testing relationship among distant transporter families.

Amino Acid Sequence↗

Normalization of affine gap costs used in optimal sequence alignment.

It is shown how to normalize the costs of an alignment algorithm that employs affine or linear gap costs. The normalized costs are interpreted as the -log probabilities of the instructions of a finite-state edit-machine. This gives an explicit model relating sequences that can be linked to processes of mutation and evolution.

Animals↗

Phylogeny of prokaryotes and chloroplasts revealed by a simple composition approach on all protein sequences from complete genomes without sequence alignment.

The complete genomes of living organisms have provided much information on their phylogenetic relationships. Similarly, the complete genomes of chloroplasts have helped to resolve the evolution of this organelle in photosynthetic eukaryotes. In this paper we propose an alternative method of phylogenetic analysis using compositional statistics for all protein sequences from complete genomes. This new method is conceptually simpler than and computationally as fast as the one proposed by Qi et al. (2004b) and Chu et al. (2004). The same data sets used in Qi et al. (2004b) and Chu et al. (2004) are analyzed using the new method. Our distance-based phylogenic tree of the 109 prokaryotes and eukaryotes agrees with the biologists "tree of life" based on 16S rRNA comparison in a predominant majority of basic branching and most lower taxa. Our phylogenetic analysis also shows that the chloroplast genomes are separated to two major clades corresponding to chlorophytes s.l. and rhodophytes s.l. The interrelationships among the chloroplasts are largely in agreement with the current understanding on chloroplast evolution.

Chloroplasts↗

TEXshade: shading and labeling of multiple sequence alignments using LATEX2 epsilon.

MOTIVATION: Typesetting, shading and labeling of nucleotide and peptide alignments using standard word processing or graphics software is time consuming. Available automatic sequence shading programs usually do not allow manual application of additional shadings or labels. Hence, a flexible alignment shading package was designed for both calculated and manual shading, using the macro language of the scientific typesetting software LATEX2 epsilon. RESULTS: TEXshade is the first TEX-based alignment shading software featuring, in addition to standard identity and similarity shading, special modes for the display of functional aspects such as charge, hydropathy or solvent accessibility. A plenitude of commands for manual shading, graphical labels, re-arrangements of the sequence order, numbering, legends etc. is implemented. Further, TEXshade allows the inclusion and display of secondary structure predictions in the DSSP-, STRIDE- and PHD-format. AVAILABILITY: From http://homepages.uni-tuebingen.de/beitz/tse.h tml (macro package and on-line documentation) CONTACT: eric.beitz@uni-tuebingen.de

Algorithms↗

Empirical analysis of protein insertions and deletions determining parameters for the correct placement of gaps in protein sequence alignments.

To understand how protein segments are inserted and deleted during divergent evolution, a set of pairwise alignments contained exactly one gap, and therefore arising from the first insertion-deletion (indel) event in the time separating the homologs, was examined. The alignments showed that "structure breaking" amino acids (PGDNS) were preferred within and flanking gapped regions, as are two residues with hydrophilic side-chains (QE) that frequently occur at the surface of protein folds. Conversely, hydrophobic residues (FMILYVW) occur infrequently within and flanking the gapped region. These preferences are modestly different in protein pairs separated by an episode of adaptive evolution, than in pairs diverging under strong functional constraints. Surprisingly, regions near an indel have not evolved more rapidly than the sequence pair overall, showing no evidence that an indel event must be compensated by local amino acid replacement. The gap-lengths are best approximated by a Zipfian distribution, with the probability of a gap of length L decreasing as a function of L(-1.8). These features are largely independent of the length of the gap and the extent of divergence (measured by both silent and non-silent sequence changes) separating the two proteins. Surprisingly, amino acid repeats were discovered in more than a third of the polypeptide segments in and around the gap. These correspond to repeats in the DNA sequence. This suggests that a signature of the mechanism by which indels occur in the DNA sequence remains in the encoded protein sequences. These data suggest specific tools to score gap placement in an alignment. They also suggest tools that distinguish true indels from gaps created by mistaken gene finding, including under-predicted and over-predicted introns. By providing mechanisms to identify errors, the tools will enhance the value of genome sequence databases in support of integrated paleogenomics strategies used to extract functional information in a post-genomic environment.

Amino Acid Sequence↗

Phylogenies from amino acid sequences aligned with gaps: the problem of gap weighting.

The common but generally overlooked problem of how best to construct phylogenies from orthologous amino acid sequences, when their alignment requires the placement therein of gaps denoting insertions/deletions in the evolutionary history of their genes since their common ancestor, has been studied. Three diverse methods were examined: 1. each missing residue in a gap is weighted as equivalent to the average number of minimum nucleotide replacements in known conjugate amino acid pairs of those same two sequences, which weight necessarily differs for each pair of sequences; 2. each missing residue in a gap is weighted as equivalent to a fixed number of nucleotide replacements; and 3. each gap, regardless of length, is weighted as equivalent to a fixed number of nucleotide replacements. For the flavodoxins, each method yielded a different best tree and suggests that the choice of method may be crucial. For the plant ferredoxins, all methods give results inconsistent with botanical classification and suggests the sequences may not all be orthologous. For the bacterial ferredoxins, the method was less germane than the actual weight used, five different best trees being obtained depending upon the weight. The best tree for all ferredoxins (prokaryotic plus eukaryotic) combined proved to be greatly dependent upon the gap locations with several reasonable aligments yielding different best trees. They also suggest that functional equivalence may well prove to be a poor guide to which residues have a common ancestral codon. The rubredoxin sequences show that a partial internal gene duplication occurred in the Pseudomonas line, probably very soon after its divergence from the other genera. Together, the results clearly indicate that the phylogenetic answer one gets may greatly depend upon how one treats the gaps but they fail to indicate what treatment may be best.

Amino Acid Sequence↗

Engineering proteins for thermostability: the use of sequence alignments versus rational design and directed evolution.

With the advent of directed evolution techniques, protein engineering has received a fresh impetus. Engineering proteins for thermostability is a particularly exciting and challenging field, as it is crucial for broadening the industrial use of recombinant proteins. In addition to directed evolution, a variety of partially successful rational concepts for engineering thermostability have been developed in the past. Recent results suggest that amino acid sequence comparisons of mesophilic proteins alone can be used efficiently to engineer thermostable proteins. The potential benefits of the underlying, semirational 'consensus concept' are compared with those of rational design and directed evolution approaches.

Directed Molecular Evolution↗

Sequence alignment of the G-protein coupled receptor superfamily.

The multitude of G-protein coupled receptor (GPR) superfamily cDNAs recently isolated has exceeded the number of receptor subtypes anticipated by pharmacological studies. Analysis of the sequence similarities and unique features of the members of this family is valuable for designing strategies to isolate related cDNAs, for developing hypotheses concerning substrate-ligand and receptor-effector interactions, and for understanding the evolution of these genes. We have compiled and aligned the 74 unique amino acid sequences published to date and review the present understanding of the structural motifs contributing to ligand binding and G-protein coupling.

Amino Acid Sequence↗

Sequence alignment and penalty choice. Review of concepts, case studies and implications.

Alignment algorithms to compare DNA or amino acid sequences are widely used tools in molecular biology. The algorithms depend on the setting of various parameters, most notably gap penalties. The effect that such parameters have on the resulting alignments is still poorly understood. This paper begins by reviewing two recent advances in algorithms and probability that enable us to take a new approach to this question. The first tool we introduce is a newly developed method to delineate efficiently all optimal alignments arising under all choices of parameters. The second tool comprises insights into the statistical behavior of optimal alignment scores. From this we gain a better understanding of the dependence of alignments on parameters in general. We propose novel criteria to detect biologically good alignments and highlight some specific features about the interaction between similarity matrices and gap penalties. To illustrate our analysis we present a detailed study of the comparison of two immunoglobulin sequences.

Algorithms↗

SCOP, Structural Classification of Proteins database: applications to evaluation of the effectiveness of sequence alignment methods and statistics of protein structural data.

The Structural Classification of Proteins (SCOP) database provides a detailed and comprehensive description of the relationships of all known protein structures. The classification is on hierarchical levels: the first two levels, family and superfamily, describe near and far evolutionary relationships; the third, fold, describes geometrical relationships. The distinction between evolutionary relationships and those that arise from the physics and chemistry of proteins is a feature that is unique to this database, so far. The database can be used as a source of data to calibrate sequence search algorithms and for the generation of population statistics on protein structures. The database and its associated files are freely accessible from a number of WWW sites mirrored from URL http://scop. mrc-lmb.cam.ac.uk/scop/.

Algorithms↗

ParAlign: a parallel sequence alignment algorithm for rapid and sensitive database searches.

There is a need for faster and more sensitive algorithms for sequence similarity searching in view of the rapidly increasing amounts of genomic sequence data available. Parallel processing capabilities in the form of the single instruction, multiple data (SIMD) technology are now available in common microprocessors and enable a single microprocessor to perform many operations in parallel. The ParAlign algorithm has been specifically designed to take advantage of this technology. The new algorithm initially exploits parallelism to perform a very rapid computation of the exact optimal ungapped alignment score for all diagonals in the alignment matrix. Then, a novel heuristic is employed to compute an approximate score of a gapped alignment by combining the scores of several diagonals. This approximate score is used to select the most interesting database sequences for a subsequent Smith-Waterman alignment, which is also parallelised. The resulting method represents a substantial improvement compared to existing heuristics. The sensitivity and specificity of ParAlign was found to be as good as Smith-Waterman implementations when the same method for computing the statistical significance of the matches was used. In terms of speed, only the significantly less sensitive NCBI BLAST 2 program was found to outperform the new approach. Online searches are available at http://dna.uio.no/search/

Algorithms↗