PubMed HealthSearch

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Consistency of optimal sequence alignments.

Pairwise optimal alignments between three or more sequences are not necessarily consistent as a whole, but consistent and inconsistent residues are usually distributed in clusters. An efficient method has been developed for locating consistent regions when each pairwise alignment is given in the form of a "skeletal representation" (Bull. math. Biol. 52, 359-373). This method is further extended so that the combination of pairwise alignments that gives the greatest consistency is found when possibly many alignments are equally optimal for each pairwise comparison. A method for acceleration of simultaneous multiple sequence alignment is proposed in which consistent regions serve as "anchor points" limiting application of direct multi-way alignment to the rest of "inconsistent" regions.

Algorithms

Database of homology-derived protein structures and the structural meaning of sequence alignment.

The database of known protein three-dimensional structures can be significantly increased by the use of sequence homology, based on the following observations. (1) The database of known sequences, currently at more than 12,000 proteins, is two orders of magnitude larger than the database of known structures. (2) The currently most powerful method of predicting protein structures is model building by homology. (3) Structural homology can be inferred from the level of sequence similarity. (4) The threshold of sequence similarity sufficient for structural homology depends strongly on the length of the alignment. Here, we first quantify the relation between sequence similarity, structure similarity, and alignment length by an exhaustive survey of alignments between proteins of known structure and report a homology threshold curve as a function of alignment length. We then produce a database of homology-derived secondary structure of proteins (HSSP) by aligning to each protein of known structure all sequences deemed homologous on the basis of the threshold curve. For each known protein structure, the derived database contains the aligned sequences, secondary structure, sequence variability, and sequence profile. Tertiary structures of the aligned sequences are implied, but not modeled explicitly. The database effectively increases the number of known protein structures by a factor of five to more than 1800. The results may be useful in assessing the structural significance of matches in sequence database searches, in deriving preferences and patterns for structure prediction, in elucidating the structural role of conserved residues, and in modeling three-dimensional detail by homology.

Amino Acid Sequence

Amino acid sequence alignment of bacterial and mammalian pancreatic serine proteases based on topological equivalences.

The three-dimensional structures of the bacterial serine proteases SGPA, SGPB, and alpha-lytic protease have been compared with those of the pancreatic enzymes alpha-chymotrypsin and elastase. This comparison shows that approximately 60% (55-64%) of the alpha-carbon atom positions of the bacterial serine proteases are topologically equivalent to the alpha-carbon atom positions of the pancreatic enzymes. The corresponding value for a comparison of the bacterial enzymes among themselves is approximately 84%. The results of these topological comparisons have been used to deduce an experimentally sound sequence alignment for these several enzymes. This alignment shows that there is extensive tertiary structural homology among the bacteria and pancreatic enzymes without significant primary sequence identity (less than 21%). The acquisition of a zymogen function by the pancreatic enzymes is accompanied by two major changes to the bacterial enzymes' architecture: an insertion of 9 residues to increase the length of the N-terminal loop, and one of 12 residues to a loop near the activation salt bridge. In addition, in these two enzyme families, the methionine loop (residues 164-182) adopts very different comformations which are associated with their altered substrate specificities.

Amino Acid Sequence

Amino acid similarity coefficients for protein modeling and sequence alignment derived from main-chain folding angles.

A set of "similarity-parameters" was calculated that reflects the influence of the proteinogenic amino acids on the structure of the protein backbone. The parameters were derived from a detailed analysis of the amino acid specific main-chain torsion angle distributions as they are found in proteins (highly resolved protein structures from the Brookhaven Protein Data Bank). The purpose of these parameters is threefold: (1) they should help in estimating the structural effect of an amino acid substitution during the design of new mutants in protein-engineering; (2) in modeling by homology they should mark places in the protein where changes in the folding are expected; and (3) they should form a scoring matrix in protein sequence alignment superior to identity scoring. The usability of the "structure derived correlation matrix (SCM)" for these purposes is assessed and demonstrated for some examples in the paper.

Amino Acid Sequence

theBIGbam: compression and interactive exploration of large-scale sequencing alignments with circular mapping support.

SUMMARY: theBIGbam (github.com/bhagavadgitadu22/theBIGbam) is a genome browser and alignment viewer designed for massive metagenomic and metatranscriptomic datasets. The tool takes BAM files containing read alignments, together with genome assemblies in FASTA format or annotated genome sequences in GenBank format. Alternatively, it can start from raw FASTQ reads and generate alignments using a modified mapper that supports circular genomes, enabling seamless read mapping across genome ends. theBIGbam can compress hundreds of gigabytes of input files 10- to 100-fold into dedicated databases while retaining key per-position information, including coverage depth and recurrent mismatches, insertions, and deletions between reads and the reference. These databases can be served to a local web browser, enabling interactive exploration of any contig in any sample using DNAFeaturesViewer for genome maps and Bokeh for mapping-derived features. Contig-sample pairs available for visualization can be filtered using a range of summary metrics calculated per contig, per sample, and per contig-sample pair to guide users toward the most relevant signals. Through its interactive visualization, theBIGbam facilitates the exploration of complex datasets, while its integrated database-combining assembly features, annotated features, and mapping-derived features-provides the information needed to investigate biological hypotheses systematically. Designed to complement existing browsing tools like IGV and Anvi'o, theBIGbam is particularly suited for examining misassemblies, subpopulations, microdiversity, and contig topology in large-scale datasets. AVAILABILITY AND IMPLEMENTATION: theBIGbam is an open-source Rust/Python package that can be installed from Bioconda or PyPI. The source code and documentation are available on GitHub (github.com/bhagavadgitadu22/theBIGbam).

Software

Progressive sequence alignment and molecular evolution of the Zn-containing alcohol dehydrogenase family.

Sequences of 47 members of the Zn-containing alcohol dehydrogenase (ADH) family were aligned progressively, and an evolutionary tree with detailed branch order and branch lengths was produced. The alignment shows that only 9 amino acid residues (of 374 in the horse liver ADH sequence) are conserved in this family; these include eight Gly and one Val with structural roles. Three residues that bind the catalytic Zn and modulate its electrostatic environment are conserved in 45 members. Asp 223, which determines specificity for NAD, is found in all but the two NADP-dependent enzymes, which have Gly or Ala. Ser or Thr 48, which makes a hydrogen bond to the substrate, is present in 46 members. The four Cys ligands for the structural zinc are conserved except in zeta-crystallin, the sorbitol dehydrogenases, and two bacterial enzymes. Analysis of the evolutionary tree gives estimates of the times of divergence for different animal ADHs. The human class II (pi) and class III (chi) ADHs probably diverged about 630 million years ago, and the newly identified human ADH6 appeared about 520 million years ago, implying that these classes of enzymes may exist or have existed in all vertebrates. The human class I ADH isoenzymes (alpha, beta, and gamma) diverged about 80 million years ago, suggesting that these isoenzymes may exist or have existed in all primates. Analysis of branch lengths shows that these plant ADHs are more conserved than the animal ones and that class III ADHs are more conserved than class I ADHs. The rate of acceptance of point mutations (PAM units) shows that selection pressure has existed for ADHs, implying that these enzymes play definite metabolic roles.

Alcohol Dehydrogenase

Vertebrate protamine gene evolution I. Sequence alignments and gene structure.

The availability of the amino acid sequence for nine different mammalian P1 family protamines and the revised amino acid sequence of the chicken protamine galline (Oliva and Dixon 1989) reveals a much close relationship between mammalian and avian protamines than was previously thought (Nakano et al. 1976). Dot matrix analysis of all protamine genes for which genomic DNA or cDNA sequence is available reveals both marked sequence similarities in the mammalian protamine gene family and internal repeated sequences in the chicken protamine gene. The detailed alignments of the cis-acting regulatory DNA sequences shows several consensus sequence patterns, particularly the conservation of a cAMP response element (CRE) in all the protamine genes and of the regions flanking the TATA box, CAP site, N-terminal coding region, and polyadenylation signal. In addition we have found a high frequency of the CA dinucleotide immediately adjacent to the CRE element of both the protamine genes and the testis transition proteins, a feature not present in other genes, which suggests the existence of an extended CRE motif involved in the coordinate expression of protamine and transition protein genes during spermatogenesis. Overall these findings suggest the existence of an avian-mammalian P1 protamine gene line and are discussed in the context of different hypotheses for protamine gene evolution and regulation.

Amino Acid Sequence

A local algorithm for DNA sequence alignment with inversions.

A dynamic programming algorithm to find all optimal alignments of DNA subsequences is described. The alignments use not only substitutions, insertions and deletions of nucleotides but also inversions (reversed complements) of substrings of the sequences. The inversion alignments themselves contain substitutions, insertions and deletions of nucleotides. We study the problem of alignment with non-intersecting inversions. To provide a computationally efficient algorithm we restrict candidate inversions to the K highest scoring inversions. An algorithm to find the J best non-intersecting alignments with inversions is also described. The new algorithm is applied to the regions of mitochondrial DNA of Drosophila yakuba and mouse coding for URF6 and cytochrome b and the inversion of the URF6 gene is found. The open problem of intersecting inversions is discussed.

Algorithms

Phylogenies from amino acid sequences aligned with gaps: the problem of gap weighting.

The common but generally overlooked problem of how best to construct phylogenies from orthologous amino acid sequences, when their alignment requires the placement therein of gaps denoting insertions/deletions in the evolutionary history of their genes since their common ancestor, has been studied. Three diverse methods were examined: 1. each missing residue in a gap is weighted as equivalent to the average number of minimum nucleotide replacements in known conjugate amino acid pairs of those same two sequences, which weight necessarily differs for each pair of sequences; 2. each missing residue in a gap is weighted as equivalent to a fixed number of nucleotide replacements; and 3. each gap, regardless of length, is weighted as equivalent to a fixed number of nucleotide replacements. For the flavodoxins, each method yielded a different best tree and suggests that the choice of method may be crucial. For the plant ferredoxins, all methods give results inconsistent with botanical classification and suggests the sequences may not all be orthologous. For the bacterial ferredoxins, the method was less germane than the actual weight used, five different best trees being obtained depending upon the weight. The best tree for all ferredoxins (prokaryotic plus eukaryotic) combined proved to be greatly dependent upon the gap locations with several reasonable aligments yielding different best trees. They also suggest that functional equivalence may well prove to be a poor guide to which residues have a common ancestral codon. The rubredoxin sequences show that a partial internal gene duplication occurred in the Pseudomonas line, probably very soon after its divergence from the other genera. Together, the results clearly indicate that the phylogenetic answer one gets may greatly depend upon how one treats the gaps but they fail to indicate what treatment may be best.

Amino Acid Sequence

Sequence alignment of the G-protein coupled receptor superfamily.

The multitude of G-protein coupled receptor (GPR) superfamily cDNAs recently isolated has exceeded the number of receptor subtypes anticipated by pharmacological studies. Analysis of the sequence similarities and unique features of the members of this family is valuable for designing strategies to isolate related cDNAs, for developing hypotheses concerning substrate-ligand and receptor-effector interactions, and for understanding the evolution of these genes. We have compiled and aligned the 74 unique amino acid sequences published to date and review the present understanding of the structural motifs contributing to ligand binding and G-protein coupling.

Amino Acid Sequence

Mast cell tryptases: examination of unusual characteristics by multiple sequence alignment and molecular modeling.

Tryptases are trypsin-like serine proteinases found in the granules of mast cells. Although they show 40% sequence identity with trypsin and contain only 20 or 21 additional residues, tryptases display several unusual features. Unlike trypsin, the tryptases only make limited cleavages in a few proteins and are not inhibited by natural trypsin inhibitors, they form tetramers, bind heparin, and their activity on synthetic substrates is progressively inhibited as the concentration of salt increases above 0.2 M. Unique sequence features of seven tryptases were identified by comparison to other serine proteinases. The three-dimensional structures of the tryptases were then predicted by molecular modeling based on the crystal structure of bovine trypsin. The models show two large insertions to lie on either side of the active-site cleft, suggesting an explanation for the limited activity of tryptases on protein substrates and the lack of inhibition by natural inhibitors. A group of conserved Trp residues and a unique proline-rich region make two surface hydrophobic patches that may account for the formation of tetramers and/or inhibition with increasing salt. Although they contain no consensus heparin-binding sequence, the tryptases have 10-13 more His residues than trypsin, and these are positioned on the surface of the model. In addition, clustering of Arg and Lys residues may also contribute to heparin binding. Putative Asn-linked glycosylation sites are found on the opposite side of the model from the active site. The model provides structural explanations for some to the unusual characteristics of the tryptases and a rational basis for future experiments, such as site-directed mutagenesis.

Amino Acid Sequence

The bacterial porin superfamily: sequence alignment and structure prediction.

The porins of Gram-negative bacteria are responsible for the 'molecular sieve' properties of the outer membrane. They form large water-filled channels which allow the diffusion of hydrophilic molecules into the periplasmic space. Owing to the strong hydrophilicity of their amino acid sequence and the nature of their secondary structure (beta strands), conventional hydropathy methods for predicting membrane topology are useless for this class of protein. The large number of available porin amino acid sequences was exploited to improve the accuracy of the prediction in combination with tools detecting amphipathicity of secondary structure. Using the constraints of beta-sheet structure these porins are predicted to contain 16 membrane-spanning strands, 14 of which are common to the two (enteric and the neisserial) porin subfamilies.

Amino Acid Sequence

Evolutionary divergence plots of homologous proteins.

A simple and efficient method is described for analyzing quantitatively multiple protein sequence alignments and finding the most conserved blocks as well as the maxima of divergence within the set of aligned sequences. It consists of calculating the mean distance and the root-mean-square distance in each column of the multiple alignment, averaging the values in a window of defined length and plotting the results as a function of the position of the window. Due attention is paid to the presence of gaps in the columns. Several examples are provided, using the sequences of several cytochromes c, serine proteases, lysozymes and globins. Two distance matrices are compared, namely the matrix derived by Gribskov and Burgess from the Dayhoff matrix, and the Risler Structural Superposition Matrix. In each case, the divergence plots effectively point to the specific residues which are known to be essential for the catalytic activity of the proteins. In addition, the regions of maximum divergence are clearly delineated. Interestingly, they are generally observed in positions immediately flanking the most conserved blocks. The method should therefore be useful for delineating the peptide segments which will be good candidates for site-directed mutagenesis and for visualizing the evolutionary constraints along homologous polypeptide chains.

Amino Acid Sequence

Simultaneous and multivariate alignment of protein sequences: correspondence between physicochemical profiles and structurally conserved regions (SCR).

A general protein sequence alignment methodology for detecting a priori unknown common structural and functional regions is described. The method proposed in this paper is based on two basic requirements for a meaningful alignment. First, each sequence or segment of a sequence is characterized by a multivariate physicochemical profile. Second, the alignment is performed by considering all the sequences simultaneously, and the algorithm detects those regions that form a set of similar profiles. In order to test the structural meaning of the alignment obtained from the sequences, quantitative comparisons are performed with structurally conserved regions (SCR) determined from the X-ray structures of three serine proteases. Results suggest that the limits of the SCR may be predicted from the similarities between the physicochemical profiles of the sequences. The procedures are not completely automated. The final step requires a visual screening of alternative pathways in order to determine an optimal alignment.

Algorithms

Modelling of binding sites of the nicotinic acetylcholine receptor and their relation to models of the whole receptor.

Models for the acetylcholine (ACh)-binding site of the nicotinic acetylcholine receptor (nAChR) are proposed. These models have been developed by using the concept of the ligand-gated ion-channel (LGIC) superfamily of receptors that have evolved from a common ancestor. An initial component of the binding site was identified as a highly conserved 15-residue stretch of primary structure in the N-terminal extracellular region of all known LGIC subunits, based on aligned sequence data of LGICs. This subregion, termed the Cys-loop, was modelled as an amphiphilic beta-hairpin and we propose that it forms a major determinant of the binding cleft for agonists. This initial, partial binding-site model has been extended to include residues biochemically identified as spatially adjacent to the binding cleft. A recently developed technique for rapidly scanning the known protein structural database for 'non-homologous similarity' using just sequence information identified the known structure of the enzyme pyrophosphatase (PPase) as a candidate scaffold for the N-terminal domain of the nAChR. This similarity was investigated further using sequence alignments. A framework model of the full N-terminal domain in which the position of the Cys-loop and other binding-site determinants, as well as the main immunogenic region (MIR), have been mapped on to the PPase structure.

Amino Acid Sequence

A simple method to generate non-trivial alternate alignments of protein sequences.

A major problem in sequence alignments based on the standard dynamic programming method is that the optimal path does not necessarily yield the best equivalencing of residues assessed by structural or functional criteria. An algorithm is presented that finds suboptimal alignments of protein sequences by a simple modification to the standard dynamic programming method. The standard pairwise weight matrix elements are modified in order to penalize, but not eliminate, the equivalencing of residues obtained from previous alignments. The algorithm thereby yields a limited set of alternate alignments that can differ considerably from the optimal. The approach is benchmarked on the alignments of immunoglobulin domains. Without a prior knowledge of the optimal choice of gap penalty, one of the suboptimal alignments is shown to be more accurate than the optimal.

Algorithms