PubMed HealthSearch

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Analysis of sequence variation among legume lectins. A ring of hypervariable residues forms the perimeter of the carbohydrate-binding site.

Twelve plant lectins from the Papilionoideae subfamily were selected to represent a range of carbohydrate specificities, and their sequences were aligned. Two variability indices were applied to the aligned sequences and the results were analysed using the three-dimensional structures of concanavalin A and the pea lectin. The areas of greatest variability were located in the carbohydrate-binding site region, forming a perimeter around a well-conserved core. These residues are inferred to be specificity determining, in the manner of antibodies, and the most variable position corresponded to Tyr100 in concanavalin A, a known ligand contact residue. In addition to the five peptide loops known to form the binding site from crystallographic studies, a sixth segment with variable residues was located in the binding-site region, and this may contribute to oligosaccharide specificity. In their overall composition, the lectin sites resemble those of the sugar-transport proteins rather than antibodies. The prospects for modelling lectin binding sites by the methods used for antibodies were also assessed.

Amino Acid Sequence

Prediction of surface loops of protein-folds from multiple alignments of homologous sequences.

Multiple alignments of distantly related homologous sequences may be used for the construction of consensus sequences that identify conserved motifs, variable segments and regions that tolerate gap events. It is suggested that such consensus sequences may be used for the prediction of key features of protein-folds. The validity of the proposed approach is illustrated in the case of the alpha 2 mu globulin superfamily: the consensus sequence derived from the multiple alignment of sequences succeeded in identifying conserved structural motifs and in predicting the location of surface loops that connect these motifs.

Amino Acid Sequence

Molecular modeling of the 3-D structure of cytochrome P-450scc.

Sequence-alignment studies of the bovine mitochondrial cholesterol side-chain cleavage enzyme cytochrome P-450scc with the bacterial cytochrome P-450cam (camphor hydroxylating enzyme) have been undertaken. Our novel alignment of the sequences revealed 69 identical residues and many highly conserved regions. The results of the sequence alignment studies were used to model the 3-D structure of P-450scc based on the available crystal structure of P-450cam. The major insertions in the sequence are found mainly on four external-loop regions of the molecule, while the core structure of P-450cam is retained with subtle internal modifications. The most hydrophobic of these four external loops is proposed as a candidate for membrane attachment.

Amino Acid Sequence

Theseus: fast and optimal affine-gap sequence-to-graph alignment.

MOTIVATION: Sequence-to-graph alignment is a central problem in bioinformatics, with applications in multiple sequence alignment (MSA) and pangenome analysis, among others. However, current algorithms for optimal affine-gap alignment impose high memory and computational requirements, limiting their scalability to aligning long sequences to complex graphs. Practical solutions partially address this problem using heuristic strategies that ultimately trade off optimality for speed. RESULTS: This work presents Theseus, a novel, fast, and optimal affine-gap sequence-to-graph alignment algorithm. Theseus leverages similarities between genomic sequences to accelerate the alignment computation and reduces the overall memory requirements without compromising optimality. To that end, Theseus processes only a subset of the dynamic programming cells, using a sparse-data strategy that enables efficient sequence-to-graph alignment. Moreover, our algorithm supports optimal affine-gap alignment on arbitrary directed graphs, including those with cycles. We evaluate Theseus on two key problems: MSA and pangenome read mapping. For MSA, we compare it against SPOA, abPOA, and POASTA. Theseus is 1.6× to 17.6× faster than POASTA, and 7.3× faster, on average, than SPOA, both optimal aligners. Compared with abPOA, Theseus ensures optimality and scales to the largest problems. For pangenome read mapping, we benchmark Theseus against the alignment stage of the mapping tool vg map, along with the alignment kernels of SPOA, abPOA, and POASTA. Theseus outperforms the other methods, showing a 1.9× to 16.9× speedup on short reads. Moreover, Theseus is 1.5× to 36.3× faster than vg when aligning against synthetic cyclic graphs. AVAILABILITY AND IMPLEMENTATION: Theseus code and documentation are publicly available at https://github.com/albertjimenezbl/theseus-lib.

Algorithms

SequenceEditingAligner: a multiple sequence editor and aligner.

Here we present the SequenceEditingAligner system for editing multiple, aligned genetic sequences. This is an interactive multi-window color system that displays more than 3500 nucleotides or amino acids. The system handles nucleic acid or protein sequences with or without secondary structure data. More than 300 sequences, each more than 1500 elements in length, may be analyzed together. With the system scientists can classify elements, align sequences, edit them, find consensus patterns, and simultaneously generate oligomer frequency histograms and other statistics.

Algorithms

Aligning two sequences within a specified diagonal band.

We describe an algorithm for aligning two sequences within a diagonal band that requires only O(NW) computation time and O(N) space, where N is the length of the shorter of the two sequences and W is the width of the band. The basic algorithm can be used to calculate either local or global alignment scores. Local alignments are produced by finding the beginning and end of a best local alignment in the band, and then applying the global alignment algorithm between those points. This algorithm has been incorporated into the FASTA program package, where it has decreased the amount of memory required to calculate local alignments from O(NW) to O(N) and decreased the time required to calculate optimized scores for every sequence in a protein sequence database by 40%. On computers with limited memory, such as the IBM-PC, this improvement both allows longer sequences to be aligned and allows optimization within wider bands, which can include longer gaps.

Algorithms

Locating a nucleotide-binding site in the thymidine kinase of vaccinia virus and of herpes simplex virus by scoring triply aligned protein sequences.

Computer techniques were used to locate related segments of amino acid sequences in the thymidine kinases of vaccinia virus and of herpes simplex virus type 1 and in porcine adenylate kinase. As determined by a procedure that evaluates triply aligned sequences, the probability that the similarities among the segments described here arose by chance was no greater than 0.001. Because the sequence in porcine adenylate kinase is a nucleotide phosphate-binding site it is concluded that the segments in the vaccinia virus and herpes simplex virus thymidine kinases perform similar functions. The segments are residues 16-23 in porcine adenylate kinase, 11-19 in vaccinia virus thymidine kinase, and 56-64 in herpes simplex virus thymidine kinase.

Adenylate Kinase

Searching for distantly related protein sequences in large databases by parallel processing on a transputer machine.

AliMac is an implementation of a sensitive sequence alignment algorithm on a parallel computer. The method achieves reliable alignments for very distantly related sequences from a combined use of amino acid exchange weights and physicochemical characteristics. The algorithm is computing intensive and its usage on conventional computers is limited to a relatively small number of sequences. The parallel implementation uses a Macintosh IIcx host computer and 21 transputers and achieves 22 times the speed of a VAX 8650 at a fraction of the cost. This paper describes the AliMac hardware and software and discusses problems and peculiarities of parallel implementations, especially with transputers. Finally, several popular sequence alignment algorithms are compared in their ability to detect distantly related sequences in searching large databases.

Algorithms

Envelope protein sequences of dengue virus isolates TH-36 and TH-Sman, and identification of a type-specific genetic marker for dengue and tick-borne flaviviruses.

Complementary DNAs were synthesized from the envelope protein genes of two isolates of dengue virus (TH-36 and TH-Sman, previously suggested as possible dengue virus type 5 and dengue virus type 6 respectively) and amplified by the polymerase chain reaction using sense and antisense primers designed from conserved dengue virus gene sequences. The amplified cDNA clones were sequenced in both directions by double-stranded dideoxynucleotide sequencing. Alignment with published dengue virus sequences enabled us to assign these viruses accurately to classified serotypes, confirming that TH-36 and TH-Sman are strains of dengue virus type 2 and dengue virus type 1 respectively. Amino acid changes between the proteins encoded by these two isolates and strains of their respective serotypes may account for the significant antigenic differences observed during previous serological typing of these viruses. Moreover, sequence alignment of flavivirus envelope proteins revealed a hypervariable region, within which members of the dengue and tick-borne virus antigenic complexes show unique peptide sequences. This type-specific hypervariable domain may be useful as a genetic marker for typing dengue and tick-borne flaviviruses.

Amino Acid Sequence

A novel randomized iterative strategy for aligning multiple protein sequences.

The rigorous alignment of multiple protein sequences becomes impractical even with a modest number of sequences, since computer memory and time requirements increase as the product of the lengths of the sequences. We have devised a strategy to approach such an optimal alignment, which modifies the intensive computer storage and time requirements of dynamic programming. Our algorithm randomly divides a group of unaligned sequences into two subgroups, between which an optimal alignment is then obtained by a Needleman-Wunsch style of algorithm. Our algorithm uses a matrix with dimensions corresponding to the lengths of the two aligned sequence subgroups. The pairwise alignment process is repeated using different random divisions of the whole group into two subgroups. Compared with the rigorous approach of solving the n-dimensional lattice by dynamic programming, our iterative algorithm results in alignments that match or are close to the optimal solution, on a limited set of test problems. We have implemented this algorithm in a computer program that runs on the IBM PC class of machines, together with a user-friendly environment for interactively selecting sequences or groups of sequences to be aligned either simultaneously or progressively.

Amino Acid Sequence

A complete amino acid sequence for the basic subunit of crotoxin.

The complete amino acid sequence of the basic subunit of crotoxin from the venom of Crotalus durissus terrificus has been determined. Fragmentation of the protein was achieved by using cyanogen bromide and arginine- and lysine-specific endoproteases. Sixteen Glx and Asx residues reported by Fraenkel-Conrat et al. (1980) in Natural Toxins (D. Eaker and T. Wadstrom, eds.), pp. 561-567, Pergamon, Oxford.) have been resolved as Glu or Gln and Asp or Asn residues, respectively. Most of the remaining sequence is identical to that reported by the foregoing authors although several significant differences were evident in our protein. Tyr-61 was not present; thus the correct sequence is Lys-60, Trp-61. The latter sequence aligns with sequences of all other known viperid and crotalid phospholipases A2 (S. D. Aird, I. I. Kaiser, R. V. Lewis, and W. G. Kruggel (1985) Biochemistry 24, 7054-7058). Other differences include Asx-99, which is Ser, and Asx-105, which is Tyr. Some positions display allelic variation. In some lots of venom Glx-33 is Gln, while in others it is Arg. Positions 37 and 69 occur as mixtures of both Lys and Arg. Amino acid sequence comparisons between the basic and acidic subunits of crotoxin and between the basic subunit and other phospholipase A2 molecules indicate that the basic subunit is structurally most similar to the monomers of nontoxic, dimeric phospholipases A2 from the venoms of Crotalus adamanteus, Crotalus atrox, and Trimeresurus okinavensis, and to the toxic monomeric phospholipase A2 from the venom of Bitis caudalis.

Amino Acid Sequence

DNA sequence homology between attB-related sites of Corynebacterium diphtheriae, Corynebacterium ulcerans, Corynebacterium glutamicum, and the attP site of gamma-corynephage.

Chromosomal restriction fragments of Corynebacterium ulcerans and C. diphtheriae, containing an integration site for corynephages of the beta family, show homology on Southern blots. Homologous DNA in also found in the soil isolate C. glutamicum, although this strain is not susceptible to beta-corynephages. Three of these DNA fragments, one for each bacterial strain, and a fragment of gamma-corynephage DNA previously shown to contain the phage integration site, were cloned and sequenced. Alignment of the 3 bacterial sequences shows a very high degree of homology in a stretch of ca 120 nucleotides, whereas the rest of the sequences is generally non-homologous. Within this common bacterial portion, a segment of ca. 96 nucleotides (core sequence) is also highly homologous to the phage sequence. The first half (ca. 50 bp) of the core sequence is identical in all aligned sequences whereas the second half, which is largely occupied by a stem-and-loop structure, contains point mutations peculiar to each clone. The described sequences are likely to be involved in phage integration/excision processes.

Attachment Sites, Microbiological

Comparison of the periplasmic receptors for L-arabinose, D-glucose/D-galactose, and D-ribose. Structural and Functional Similarity.

The primary sequence of the receptor for L-arabinose or Ara-binding protein (ABP) composed of 306 residues is very different from the D-glucose/D-galactose-binding protein (GGBP) which consists of 309 residues. Nevertheless, superimpositioning of the well-refined high resolution structures of ABP in complex with D-galactose and the GGBP in complex with D-glucose shows very similar structures; 220 of the residues (or about 70%) have a root mean square deviation of 2.0 A. From the superpositioning, nine pairs of continuous segments (consisting of 8-51 residues), mainly alpha-helices and beta-strands that form the core of the two lobes of the bilobate proteins were found to exhibit strong sequence homology. The equivalenced structures and aligned sequences show that many of the polar, as well as aromatic residues, in the sugar-binding sites located in the cleft between the two lobes are highly conserved. Surprisingly, however, the exact mode of binding of the D-galactose in ABP is totally different from that of the D-glucose in GGBP. Using the structurally aligned sequences of the ABP and GGBP as a template, we have matched the sequence of the ribose-binding protein (RBP) which consists of 271 residues with the ABP/GGBP pair. Although the nine aligned segments of all three proteins show little sequence identity, they have significant homology. Four additional segments of RBP were matched only with GGBP, leading to the alignment of about 90% of the RBP sequence with the GGBP sequence. Many of the conserved residues in the binding sites of ABP and GGBP matched with similar residues in RBP. Additional observations indicate that the GGBP/RBP pair is more closely related than the ABP/RBP or ABP/GGBP pair. All three binding proteins, which may have diverged from a common ancestor, serve as primary receptors for bacterial high affinity active transport systems. Moreover, GGBP and RBP, but not ABP, also act as receptors for chemotaxis. An exposed site located in one domain, which includes Gly74, for interacting with the trg transmembrane signal transducer that is involved in triggering chemotaxis has been located in the structure of GGBP (Vyas, N.K., Vyas, M.N., and Quiocho, F.A. (1988) Science 242, 1290-1295). Whereas the site is absent in the structure of ABP, it is strongly predicted to be present in RBP which shares the same trg transducer with GGBP. The knowledge-based alignment of RBP further revealed two possible additional peripheral chemotactic sites that show high structural and sequence similarity between GGBP and RBP only. At least one of these sites, together with the one proven to exist in the other domain, could be used by the signal transducer with which both binding proteins interact in a way which the substrate-loaded "closed cleft" structure could be discriminated from the unliganded "open cleft" form by the transducer.

Amino Acid Sequence

A workbench for multiple alignment construction and analysis.

Multiple sequence alignment can be a useful technique for studying molecular evolution, as well as for analyzing relationships between structure or function and primary sequence. We have developed for this purpose an interactive program, MACAW (Multiple Alignment Construction and Analysis Workbench), that allows the user to construct multiple alignments by locating, analyzing, editing, and combining "blocks" of aligned sequence segments. MACAW incorporates several novel features. (1) Regions of local similarity are located by a new search algorithm that avoids many of the limitations of previous techniques. (2) The statistical significance of blocks of similarity is evaluated using a recently developed mathematical theory. (3) Candidate blocks may be evaluated for potential inclusion in a multiple alignment using a variety of visualization tools. (4) A user interface permits each block to be edited by moving its boundaries or by eliminating particular segments, and blocks may be linked to form a composite multiple alignment. No completely automatic program is likely to deal effectively with all the complexities of the multiple alignment problem; by combining a powerful similarity search algorithm with flexible editing, analysis and display tools, MACAW allows the alignment strategy to be tailored to the problem at hand.

Algorithms

Leveraging basecaller's move table to generate a lightweight k-mer model for nanopore sequencing analysis.

MOTIVATION: Nanopore sequencing by Oxford Nanopore Technologies (ONT) enables direct analysis of DNA and RNA by capturing raw electrical signals. Different nanopore chemistries have varied k-mer lengths, current levels, and standard deviations, which are stored in "k-mer models." In cases where official models are lacking or unsuitable for specific sequencing conditions, tailored k-mer models are crucial to ensure precise signal-to-sequence alignment, analysis and interpretation. The process of transforming raw signal data into nucleotide sequences, known as basecalling, is a fundamental step in nanopore sequencing. RESULTS: In this study, we leverage the move table produced by ONT's basecalling software to create a lightweight de novo k-mer model for RNA004 chemistry. We demonstrate the validity of our custom k-mer model by using it to guide signal-to-sequence alignment analysis, achieving high alignment rates (97.48%) compared to larger default models. Additionally, our 5-mer model exhibits similar performance as the default 9-mer models another analysis, such as detection of m6A RNA modifications. We provide our method, termed Poregen, as a generalizable approach for creation of custom, de novo k-mer models for nanopore signal data analysis. AVAILABILITY AND IMPLEMENTATION: Poregen is an open source package under an MIT license: https://github.com/hiruna72/poregen.

Nanopore Sequencing

A fast homology program for aligning biological sequences.

The algorithm of Gotoh computes in two passes of MN steps the alignment of a pair of sequences of lengths M and N, subject to a constraint on the form of the gap weighting function. This compares with the previous algorithm of Waterman et al. which runs in M2N steps. Gotoh also gave a method using two passes of (L+2)MN steps in the case where gap weights remain constant for gaps of length greater than L. Here we describe a procedure for computing the alignment (evolutionary distance and optimal path) in a single pass of MN steps for both cases.

Amino Acid Sequence

Statistical analysis of DNA sequences.

Developments in the statistical analysis of DNA sequence data since 1984 are reviewed. Mathematical methods employing dynamic programming or incorporating Markov chain theory have been developed to search sequences for regions of similarity and to align sequences. When the biological forces of mutation and genetic drift are included in models, distances between aligned sequences allow the construction of evolutionary trees. Theory based on models may lead to estimates of variation of parameter estimates and so give a means of assessing the statistical significance of observed patterns and relationships. The complexity of DNA sequences, however, suggests that most statistical inferences will rest on random permutations of sequences.

Base Sequence

Maximum likelihood alignment of DNA sequences.

The optimal alignment problem for pairs of molecular sequences under a probabilistic model of evolutionary change is equivalent to the problem of estimating the maximum likelihood time required to transform one sequence to the other. When this time has been estimated, various alignments of high posterior probability may be written down. A simple model with two parameters is presented and a method is described by which the likelihood may be computed. Maximum likelihood estimates for some pairs of tRNA genes illustrate the method and allow us to obtain the best alignments under the model.

Animals