PubMed HealthSearch

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

An analysis of the periodicity of conserved residues in sequence alignments of G-protein coupled receptors. Implications for the three-dimensional structure.

Twenty-three sequences from the family of G-protein coupled receptors have been aligned according to the 'historical alignment' procedure of Feng and Doolittle. Fourier transform analysis of this reveals that parts of five of the seven putative membrane-spanning regions exhibit a periodicity of conserved/nonconserved residues which is compatible with the periodicity of the alpha-helix. This would place the conserved residues on one side of the helix, which may face the inside of the proposed seven membered helical bundle.

Amino Acid Sequence

Mast cell tryptases: examination of unusual characteristics by multiple sequence alignment and molecular modeling.

Tryptases are trypsin-like serine proteinases found in the granules of mast cells. Although they show 40% sequence identity with trypsin and contain only 20 or 21 additional residues, tryptases display several unusual features. Unlike trypsin, the tryptases only make limited cleavages in a few proteins and are not inhibited by natural trypsin inhibitors, they form tetramers, bind heparin, and their activity on synthetic substrates is progressively inhibited as the concentration of salt increases above 0.2 M. Unique sequence features of seven tryptases were identified by comparison to other serine proteinases. The three-dimensional structures of the tryptases were then predicted by molecular modeling based on the crystal structure of bovine trypsin. The models show two large insertions to lie on either side of the active-site cleft, suggesting an explanation for the limited activity of tryptases on protein substrates and the lack of inhibition by natural inhibitors. A group of conserved Trp residues and a unique proline-rich region make two surface hydrophobic patches that may account for the formation of tetramers and/or inhibition with increasing salt. Although they contain no consensus heparin-binding sequence, the tryptases have 10-13 more His residues than trypsin, and these are positioned on the surface of the model. In addition, clustering of Arg and Lys residues may also contribute to heparin binding. Putative Asn-linked glycosylation sites are found on the opposite side of the model from the active site. The model provides structural explanations for some to the unusual characteristics of the tryptases and a rational basis for future experiments, such as site-directed mutagenesis.

Amino Acid Sequence

Sequence alignment and evolutionary comparison of the L10 equivalent and L12 equivalent ribosomal proteins from archaebacteria, eubacteria, and eucaryotes.

The genes corresponding to the L10 and L12 equivalent ribosomal proteins (L10e and L12e) of Escherichia coli have been cloned and sequenced from two widely divergent species of archaebacteria, Halobacterium cutirubrum and Sulfolobus solfataricus. The deduced amino acid sequences of the L10e and L12e proteins have been compared to each other and to available eubacterial and eucaryotic sequences. We have identified the human P0 protein as the eucaryotic L10e. The L10e proteins from the three kingdoms were found to be colinear. The eubacterial L10e protein is much shorter than the archaebacterial-eucaryotic proteins because of two large deletions, one internal and one at the carboxy terminus. The archaebacterial and eucaryotic L12e proteins were also colinear; the eubacterial protein is homologous to the archaebacterial and eucaryotic L12e proteins, but has suffered rearrangement through what appear to be gene fusion events. Intraspecies comparisons between L10e and L12e sequences indicate the archaebacterial and eucaryotic L10e proteins contain a partial copy of the L12e protein fused to their carboxy terminus. In the eubacteria most of this fusion has been removed by the carboxy terminal deletion. Within the L12e-derived region, a 26-amino acid-long internal modular sequence reiterated thrice in the archaebacterial L10e, twice in the eucaryotic L10e, and once in the eubacterial L10e was discovered. This modular sequence also appears to be present as a single copy in all L12e proteins and may play a role in L12e dimerization, L10e-L12e complex formation, and the function of L10e-L12e complex in translation.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence

The bacterial porin superfamily: sequence alignment and structure prediction.

The porins of Gram-negative bacteria are responsible for the 'molecular sieve' properties of the outer membrane. They form large water-filled channels which allow the diffusion of hydrophilic molecules into the periplasmic space. Owing to the strong hydrophilicity of their amino acid sequence and the nature of their secondary structure (beta strands), conventional hydropathy methods for predicting membrane topology are useless for this class of protein. The large number of available porin amino acid sequences was exploited to improve the accuracy of the prediction in combination with tools detecting amphipathicity of secondary structure. Using the constraints of beta-sheet structure these porins are predicted to contain 16 membrane-spanning strands, 14 of which are common to the two (enteric and the neisserial) porin subfamilies.

Amino Acid Sequence

Evolutionary divergence plots of homologous proteins.

A simple and efficient method is described for analyzing quantitatively multiple protein sequence alignments and finding the most conserved blocks as well as the maxima of divergence within the set of aligned sequences. It consists of calculating the mean distance and the root-mean-square distance in each column of the multiple alignment, averaging the values in a window of defined length and plotting the results as a function of the position of the window. Due attention is paid to the presence of gaps in the columns. Several examples are provided, using the sequences of several cytochromes c, serine proteases, lysozymes and globins. Two distance matrices are compared, namely the matrix derived by Gribskov and Burgess from the Dayhoff matrix, and the Risler Structural Superposition Matrix. In each case, the divergence plots effectively point to the specific residues which are known to be essential for the catalytic activity of the proteins. In addition, the regions of maximum divergence are clearly delineated. Interestingly, they are generally observed in positions immediately flanking the most conserved blocks. The method should therefore be useful for delineating the peptide segments which will be good candidates for site-directed mutagenesis and for visualizing the evolutionary constraints along homologous polypeptide chains.

Amino Acid Sequence

Simultaneous and multivariate alignment of protein sequences: correspondence between physicochemical profiles and structurally conserved regions (SCR).

A general protein sequence alignment methodology for detecting a priori unknown common structural and functional regions is described. The method proposed in this paper is based on two basic requirements for a meaningful alignment. First, each sequence or segment of a sequence is characterized by a multivariate physicochemical profile. Second, the alignment is performed by considering all the sequences simultaneously, and the algorithm detects those regions that form a set of similar profiles. In order to test the structural meaning of the alignment obtained from the sequences, quantitative comparisons are performed with structurally conserved regions (SCR) determined from the X-ray structures of three serine proteases. Results suggest that the limits of the SCR may be predicted from the similarities between the physicochemical profiles of the sequences. The procedures are not completely automated. The final step requires a visual screening of alternative pathways in order to determine an optimal alignment.

Algorithms

Modelling of binding sites of the nicotinic acetylcholine receptor and their relation to models of the whole receptor.

Models for the acetylcholine (ACh)-binding site of the nicotinic acetylcholine receptor (nAChR) are proposed. These models have been developed by using the concept of the ligand-gated ion-channel (LGIC) superfamily of receptors that have evolved from a common ancestor. An initial component of the binding site was identified as a highly conserved 15-residue stretch of primary structure in the N-terminal extracellular region of all known LGIC subunits, based on aligned sequence data of LGICs. This subregion, termed the Cys-loop, was modelled as an amphiphilic beta-hairpin and we propose that it forms a major determinant of the binding cleft for agonists. This initial, partial binding-site model has been extended to include residues biochemically identified as spatially adjacent to the binding cleft. A recently developed technique for rapidly scanning the known protein structural database for 'non-homologous similarity' using just sequence information identified the known structure of the enzyme pyrophosphatase (PPase) as a candidate scaffold for the N-terminal domain of the nAChR. This similarity was investigated further using sequence alignments. A framework model of the full N-terminal domain in which the position of the Cys-loop and other binding-site determinants, as well as the main immunogenic region (MIR), have been mapped on to the PPase structure.

Amino Acid Sequence

A simple method to generate non-trivial alternate alignments of protein sequences.

A major problem in sequence alignments based on the standard dynamic programming method is that the optimal path does not necessarily yield the best equivalencing of residues assessed by structural or functional criteria. An algorithm is presented that finds suboptimal alignments of protein sequences by a simple modification to the standard dynamic programming method. The standard pairwise weight matrix elements are modified in order to penalize, but not eliminate, the equivalencing of residues obtained from previous alignments. The algorithm thereby yields a limited set of alternate alignments that can differ considerably from the optimal. The approach is benchmarked on the alignments of immunoglobulin domains. Without a prior knowledge of the optimal choice of gap penalty, one of the suboptimal alignments is shown to be more accurate than the optimal.

Algorithms

Molecular cloning of a human thyrotropin receptor cDNA fragment. Use of highly degenerate, inosine containing primers derived from aligned amino acid sequences of a homologous family of glycoprotein hormone receptors.

Autoantibodies to the thyrotropin (TSH) hormone receptor (TSH-R) are present in the sera of patients with thyroid autoimmune disease which are pathogenetic leading to hyperthyroidism of Graves' disease. Considerable interest has been focused on the cloning of the human TSH-R, which has until very recently, proven exceedingly difficult due to the very low receptor level expression on thyroid cells. We have used polymerase chain reaction and highly degenerate, inosine containing oligonucleotides derived from sequence alignments of the transmembrane regions 2 and 7 of a number of G-binding protein receptors including the lutropin/choriogonadotropin (LH/CG) receptors to amplify various cDNAs from human thyroid cDNA. Sequencing analysis of 27 different clones revealed that they fall into eight different groups. The very recent publication of the complete nucleotide sequence of the human TSH-R revealed that one of the groups (GT1) containing seven clones which had been sequenced belong to the human TSH-receptor. The sequence of all 7 GT1 clones was identical and in complete concordance with transmembrane regions 2 and 7 of the published TSH-R sequence. Our results show that by designing oligonucleotides to common transmembrane regions of G-binding proteins where the primers are biased in their sequence to the LH/CG receptors it is possible to amplify the TSH-R receptor sequence.

Amino Acid Sequence

Preferred positions of AA and TT dinucleotides in aligned nucleosomal DNA sequences.

Multiple alignment of 118 nucleosomal DNA sequences by maximizing simultaneously match of AA dinucleotides and match of TT dinucleotides results in a pattern of the dinucleotide distributions which is characteristic of the nucleosomal DNA sequences. The AA dinucleotides are found to be distributed symmetrically relative to the TT dinucleotide distribution, around the middle point of the nucleosomal DNA sequence. The distances between major peaks of the distributions are multiples of about 10.4 bases. The peaks of the TT distribution are shifted by 6 bases downstream from the peaks of the AA distribution.

Adenine

Analysis of sequence variation among legume lectins. A ring of hypervariable residues forms the perimeter of the carbohydrate-binding site.

Twelve plant lectins from the Papilionoideae subfamily were selected to represent a range of carbohydrate specificities, and their sequences were aligned. Two variability indices were applied to the aligned sequences and the results were analysed using the three-dimensional structures of concanavalin A and the pea lectin. The areas of greatest variability were located in the carbohydrate-binding site region, forming a perimeter around a well-conserved core. These residues are inferred to be specificity determining, in the manner of antibodies, and the most variable position corresponded to Tyr100 in concanavalin A, a known ligand contact residue. In addition to the five peptide loops known to form the binding site from crystallographic studies, a sixth segment with variable residues was located in the binding-site region, and this may contribute to oligosaccharide specificity. In their overall composition, the lectin sites resemble those of the sugar-transport proteins rather than antibodies. The prospects for modelling lectin binding sites by the methods used for antibodies were also assessed.

Amino Acid Sequence

Prediction of surface loops of protein-folds from multiple alignments of homologous sequences.

Multiple alignments of distantly related homologous sequences may be used for the construction of consensus sequences that identify conserved motifs, variable segments and regions that tolerate gap events. It is suggested that such consensus sequences may be used for the prediction of key features of protein-folds. The validity of the proposed approach is illustrated in the case of the alpha 2 mu globulin superfamily: the consensus sequence derived from the multiple alignment of sequences succeeded in identifying conserved structural motifs and in predicting the location of surface loops that connect these motifs.

Amino Acid Sequence

Molecular modeling of the 3-D structure of cytochrome P-450scc.

Sequence-alignment studies of the bovine mitochondrial cholesterol side-chain cleavage enzyme cytochrome P-450scc with the bacterial cytochrome P-450cam (camphor hydroxylating enzyme) have been undertaken. Our novel alignment of the sequences revealed 69 identical residues and many highly conserved regions. The results of the sequence alignment studies were used to model the 3-D structure of P-450scc based on the available crystal structure of P-450cam. The major insertions in the sequence are found mainly on four external-loop regions of the molecule, while the core structure of P-450cam is retained with subtle internal modifications. The most hydrophobic of these four external loops is proposed as a candidate for membrane attachment.

Amino Acid Sequence

Theseus: fast and optimal affine-gap sequence-to-graph alignment.

MOTIVATION: Sequence-to-graph alignment is a central problem in bioinformatics, with applications in multiple sequence alignment (MSA) and pangenome analysis, among others. However, current algorithms for optimal affine-gap alignment impose high memory and computational requirements, limiting their scalability to aligning long sequences to complex graphs. Practical solutions partially address this problem using heuristic strategies that ultimately trade off optimality for speed. RESULTS: This work presents Theseus, a novel, fast, and optimal affine-gap sequence-to-graph alignment algorithm. Theseus leverages similarities between genomic sequences to accelerate the alignment computation and reduces the overall memory requirements without compromising optimality. To that end, Theseus processes only a subset of the dynamic programming cells, using a sparse-data strategy that enables efficient sequence-to-graph alignment. Moreover, our algorithm supports optimal affine-gap alignment on arbitrary directed graphs, including those with cycles. We evaluate Theseus on two key problems: MSA and pangenome read mapping. For MSA, we compare it against SPOA, abPOA, and POASTA. Theseus is 1.6× to 17.6× faster than POASTA, and 7.3× faster, on average, than SPOA, both optimal aligners. Compared with abPOA, Theseus ensures optimality and scales to the largest problems. For pangenome read mapping, we benchmark Theseus against the alignment stage of the mapping tool vg map, along with the alignment kernels of SPOA, abPOA, and POASTA. Theseus outperforms the other methods, showing a 1.9× to 16.9× speedup on short reads. Moreover, Theseus is 1.5× to 36.3× faster than vg when aligning against synthetic cyclic graphs. AVAILABILITY AND IMPLEMENTATION: Theseus code and documentation are publicly available at https://github.com/albertjimenezbl/theseus-lib.

Algorithms

SequenceEditingAligner: a multiple sequence editor and aligner.

Here we present the SequenceEditingAligner system for editing multiple, aligned genetic sequences. This is an interactive multi-window color system that displays more than 3500 nucleotides or amino acids. The system handles nucleic acid or protein sequences with or without secondary structure data. More than 300 sequences, each more than 1500 elements in length, may be analyzed together. With the system scientists can classify elements, align sequences, edit them, find consensus patterns, and simultaneously generate oligomer frequency histograms and other statistics.

Algorithms

Aligning two sequences within a specified diagonal band.

We describe an algorithm for aligning two sequences within a diagonal band that requires only O(NW) computation time and O(N) space, where N is the length of the shorter of the two sequences and W is the width of the band. The basic algorithm can be used to calculate either local or global alignment scores. Local alignments are produced by finding the beginning and end of a best local alignment in the band, and then applying the global alignment algorithm between those points. This algorithm has been incorporated into the FASTA program package, where it has decreased the amount of memory required to calculate local alignments from O(NW) to O(N) and decreased the time required to calculate optimized scores for every sequence in a protein sequence database by 40%. On computers with limited memory, such as the IBM-PC, this improvement both allows longer sequences to be aligned and allows optimization within wider bands, which can include longer gaps.

Algorithms