PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Amino acid sequence alignment of bacterial and mammalian pancreatic serine proteases based on topological equivalences.

The three-dimensional structures of the bacterial serine proteases SGPA, SGPB, and alpha-lytic protease have been compared with those of the pancreatic enzymes alpha-chymotrypsin and elastase. This comparison shows that approximately 60% (55-64%) of the alpha-carbon atom positions of the bacterial serine proteases are topologically equivalent to the alpha-carbon atom positions of the pancreatic enzymes. The corresponding value for a comparison of the bacterial enzymes among themselves is approximately 84%. The results of these topological comparisons have been used to deduce an experimentally sound sequence alignment for these several enzymes. This alignment shows that there is extensive tertiary structural homology among the bacteria and pancreatic enzymes without significant primary sequence identity (less than 21%). The acquisition of a zymogen function by the pancreatic enzymes is accompanied by two major changes to the bacterial enzymes' architecture: an insertion of 9 residues to increase the length of the N-terminal loop, and one of 12 residues to a loop near the activation salt bridge. In addition, in these two enzyme families, the methionine loop (residues 164-182) adopts very different comformations which are associated with their altered substrate specificities.

Amino Acid Sequence↗

Background rareness-based iterative multiple sequence alignment algorithm for regulatory element detection.

MOTIVATION: Experimental methods capable of generating sets of co-regulated genes have become commonplace, however, recognizing the regulatory motifs responsible for this regulation remains difficult. As a result, computational detection of transcription factor binding sites in such data sets has been an active area of research. Most approaches have utilized either Gibbs sampling or greedy strategies to identify such elements in sets of sequences. These existing methods have varying degrees of success depending on the strength and length of the signals and the number of available sequences. We present a new deterministic iterative algorithm for regulatory element detection based on a Markov chain background. As in other methods, sequences in the entire genome and the training set are taken into account in order to discriminate against commonly occurring signals and produce patterns, which are significant in the training set. RESULTS: The results of the algorithm compare favorably with existing tools on previously known and newly compiled data sets. The iteration based search appears rather rigorous, not only finding the binding sites, but also showing how the binding site stands out from genomic background. The approach used to score the results is critical and a discussion of various scoring schemes and options is also presented. Benchmarking of several methods shows that while most tools are good at detecting strong signals, Gibbs sampling algorithms give inconsistent results when the regulatory element signal becomes weak. A Markov chain based background model alleviates the drawbacks of MAP (maximum a posteriori log likelihood) scores. AVAILABILITY: Available on request from the authors. SUPPLEMENTARY INFORMATION: Data and the results presented in this paper are available on the web at http://compbio.ornl.gov/mira/index.html

Algorithms↗

Sequence alignment and homology threading reveals prokaryotic and eukaryotic proteins similar to lactose permease.

Certain prokaryotic transport proteins similar to the lactose permease of Escherichia coli (LacY) have been identified by BLAST searches from available genomic databanks. These proteins exhibit conservation of amino acid residues that participate in sugar binding and H(+) translocation in LacY. Homology threading of prokaryotic transporters based on the X-ray structure of LacY (PDB ID: 1PV7) and sequence similarities reveals a common overall fold for sugar transporters belonging to the Major Facilitator Superfamily (MFS) and suggest new targets for study. Evolution-based searches for sequence similarities also identify eukaryotic proteins bearing striking resemblance to MFS sugar transporters. Like LacY, the eukaryotic proteins are predicted to have 12 transmembrane domains (TMDs), and many of the irreplaceable residues for sugar binding and H(+) translocation in LacY appear to be largely conserved. The overall size of the eukaryotic homologs is about twice that of prokaryotic permeases with longer N and C termini and loops between TMDs III-IV and VI-VII. The human gene encoding protein FLJ20160 consists of six exons located on more than 60,000 bp of DNA sequences and requires splicing to produce mature mRNA. Cellular localization predictions suggest membrane insertion with possible proteolysis at the N terminus, and expression studies with the human protein FJL20160 demonstrate membrane insertion in both E.coli and Pichia pastoris. Widespread expression of the eukaryotic sugar transport candidates suggests an important role in cellular metabolism, particularly in brain and tumors. Homology is observed in the TMDs of both the eukaryotic and prokaryotic proteins that contain residues involved in sugar binding and H(+) translocation in LacY.

Amino Acid Sequence↗

Analysis of covariation in an SH3 domain sequence alignment: applications in tertiary contact prediction and the design of compensating hydrophobic core substitutions.

We have analyzed sequence covariation in an alignment of 266 non-redundant SH3 domain sequences using chi-squared statistical methods. Artifactual covariations arising from close evolutionary relationships among certain sequence subgroups were eliminated using empirically derived sequence diversity thresholds. This covariation detection method was able to predict residue-residue contacts (side-chain centres of mass within 8 A) in the structure of the SH3 domain with an accuracy of 85 %, which is greater than that achieved in many previous covariation studies. In examining the positions involved most frequently in covariations, we discovered a dramatic over-representation of a subset of five hydrophobic core positions. This covariation information was used to design second and third site substitutions that could compensate for highly destabilizing hydrophobic core substitutions in the Fyn SH3 domain, thus providing experimental data to validate the covariation analysis. The testing of our covariation detection method on 15 other alignments showed that the accuracy of contact prediction is highly variable depending on which sequence alignment is used, and useful levels of prediction accuracy were obtained with only approximately one-third of alignments. The results presented here provide insight into the difficulties inherent in covariation analysis, and suggest that it may have limited usefulness in tertiary structure prediction. On the other hand, our ability to use covariation analysis to design stabilizing combinations of hydrophobic core substitutions attests to its potential utility for gaining deeper insight into the stability determinants and functional mechanisms of proteins with known three-dimensional structures.

Algorithms↗

Natural selection for kinetic stability is a likely origin of correlations between mutational effects on protein energetics and frequencies of amino acid occurrences in sequence alignments.

It appears plausible that natural selection constrains, to some extent at least, the stability in many natural proteins. If, during protein evolution, stability fluctuates within a comparatively narrow range, then mutations are expected to be fixed with frequencies that reflect mutational effects on stability. Indeed, we recently reported a robust correlation between the effect of 27 conservative mutations on the thermodynamic stability (unfolding free energy) of Escherichia coli thioredoxin and the frequencies of residues occurrences in sequence alignments. We show here that this correlation likely implies a lower limit to thermodynamic stability of only a few kJ/mol below the unfolding free energy of the wild-type (WT) protein. We suggest, therefore, that the correlation does not reflect natural selection of thermodynamic stability by itself, but of some other factor which is linked to thermodynamic stability for the mutations under study. We propose that this other factor is the kinetic stability of thioredoxin in vivo, since( i) kinetic stability relates to irreversible denaturation, (ii) the rate of irreversible denaturation in a crowded cellular environment (or in a harsh extracellular environment) is probably determined by the rate of unfolding, and (iii) the half-life for unfolding changes in an exponential manner with activation free energy and, consequently, comparatively small free energy effects can have deleterious consequences for kinetic stability. This proposal is supported by the results of a kinetic study of the WT form and the 27 single-mutant variants of E. coli thioredoxin based on the global analyses of chevron plots and equilibrium unfolding profiles determined from double-jump unfolding assays. This kinetic study suggests, furthermore, one of the factors that may contribute to the high activation free energy for unfolding in thioredoxin (required for kinetic stability), namely the energetic optimization of native-state residue environments in regions, which become disrupted in the transition state for unfolding.

Amino Acid Sequence↗

Amino acid similarity coefficients for protein modeling and sequence alignment derived from main-chain folding angles.

A set of "similarity-parameters" was calculated that reflects the influence of the proteinogenic amino acids on the structure of the protein backbone. The parameters were derived from a detailed analysis of the amino acid specific main-chain torsion angle distributions as they are found in proteins (highly resolved protein structures from the Brookhaven Protein Data Bank). The purpose of these parameters is threefold: (1) they should help in estimating the structural effect of an amino acid substitution during the design of new mutants in protein-engineering; (2) in modeling by homology they should mark places in the protein where changes in the folding are expected; and (3) they should form a scoring matrix in protein sequence alignment superior to identity scoring. The usability of the "structure derived correlation matrix (SCM)" for these purposes is assessed and demonstrated for some examples in the paper.

Amino Acid Sequence↗

Further improvement in methods of group-to-group sequence alignment with generalized profile operations.

It has previously been shown that rigorous optimization of alignment between two groups of sequences in the sense of minimal sum of pairs (SP) score with a linear gap-weighting function can be achieved by an extended version of the dynamic programming algorithm. The major drawback of this algorithm was that the computation time grows in proportion to the product of the numbers (M and N) of sequences comprising the two groups. A new algorithm presented in this paper achieves the same rigorous alignment in a time complexity much less dependent on the sizes of the groups. Examinations with many groups of sequences indicated that the new algorithm runs faster than the earlier one when M x N > 6-10, approximately 10 times faster when M x N approximately 200, and > 100 times faster when M x N > 2500. This computational acceleration facilitates application of the algorithm to alignment of large groups, especially in the framework of iterative refinement strategies.

Algorithms↗

LumberJack: a heuristic tool for sequence alignment exploration and phylogenetic inference.

SUMMARY: LumberJack is a phylogenetic tool intended to serve two purposes: to facilitate sampling treespace to find likely tree topologies quickly, and to map phylogenetic signal onto regions of an alignment in a revealing way. LumberJack creates non-random jackknifed alignments by progressively sliding a window of omission along the alignment. A neighbor-joining tree is built from the full alignment and from each jackknifed alignment, and then the likelihood for each topology (given the original full alignment) is calculated. To determine whether any of the topologies generated is significantly more likely than the others, Kishino-Hasegawa, Shimodaira-Hasegawa and ELW tests are implemented. Availability and SUPPLEMENTARY INFORMATION: http://www.plantbio.uga.edu/~russell/software.html

Algorithms↗

Internal transcribed spacer rRNA gene-based phylogenetic reconstruction using algorithms with local and global sequence alignment for black yeasts and their relatives.

Sequences of rRNA gene internal transcribed spacer (ITS) of a standard set of black yeast-like fungal pathogens were compared using two methods: local and global alignments. The latter is based on DNA-walk divergence analysis. This method has become recently available as an algorithm (DNAWD program) which converts sequences into three-dimensional walks. The walks are compared with, or fit to, each other generating global alignments. The DNA-walk geometry defines a proper metric used to create a distance matrix appropriated for phylogenetic reconstruction. In this work, the analyses were carried out for species currently classified in Capronia, Cladophialophora, Exophiala, Fonsecaea, Phialophora, and Ramichloridium. Main groups were verified by small-subunit rRNA gene data. DNAWD applied to ITS2 alone enabled species recognition as well as phylogenetic reconstruction reflecting clades discriminated in small-subunit rRNA gene phylogeny, which was not possible with any other algorithm using local alignment for the same data set. It is concluded that DNAWD provides rapid insight into broader relationships between groups using genes that otherwise would be hardly usable for this purpose.

Algorithms↗

Progressive sequence alignment as a prerequisite to correct phylogenetic trees.

A progressive alignment method is described that utilizes the Needleman and Wunsch pairwise alignment algorithm iteratively to achieve the multiple alignment of a set of protein sequences and to construct an evolutionary tree depicting their relationship. The sequences are assumed a priori to share a common ancestor, and the trees are constructed from difference matrices derived directly from the multiple alignment. The thrust of the method involves putting more trust in the comparison of recently diverged sequences than in those evolved in the distant past. In particular, this rule is followed: "once a gap, always a gap." The method has been applied to three sets of protein sequences: 7 superoxide dismutases, 11 globins, and 9 tyrosine kinase-like sequences. Multiple alignments and phylogenetic trees for these sets of sequences were determined and compared with trees derived by conventional pairwise treatments. In several instances, the progressive method led to trees that appeared to be more in line with biological expectations than were trees obtained by more commonly used methods.

Algorithms↗

Pyruvate: ferredoxin oxidoreductase from the sulfate-reducing Archaeoglobus fulgidus: molecular composition, catalytic properties, and sequence alignments.

Archaeoglobus fulgidus is a hyperthermophilic sulfate-reducing archaeon. In this communication we describe the purification and properties of pyruvate: ferredoxin oxidoreductase from this organism. The catabolic enzyme was purified 250-fold to apparent homogeneity with a yield of 16%. The native enzyme had an apparent molecular mass of 120 kDa and was composed of four different subunits of apparent molecular masses of 45, 33, 25, and 13 kDa, indicating an alpha beta gamma delta structure. Per mol, the enzyme contained 0.8 mol thiamine pyrophosphate, 9 mol non-heme iron, and 8 mol acid-labile sulfur. FAD, FMN, lipoic acid, and copper were not found. The purified enzyme showed an apparent Km for coenzyme A of 0.02 mM, for pyruvate of 0.3 mM, and for clostridial ferredoxin of 0.01 mM, an apparent Vmax of 64 U/mg (at 65 degrees C) with a pH optimum near 7.5 and an Arrhenius activation energy of 75 kJ/mol (between 30 and 70 degrees C). The temperature optimum was above 90 degrees C. At 90 degrees C, the enzyme lost 50% activity within 60 min in the presence of 2 M KCl. The enzyme did not catalyze the oxidation of 2-oxoglutarate, indolepyruvate, phenylpyruvate, glyoxylate, and hydroxypyruvate. The N-terminal amino acid sequences of the four subunits were determined. The sequence of the alpha-subunit had similarities to the N-terminal amino acid sequence of the alpha-subunit of the heterotetrameric pyruvate: ferredoxin oxidoreductase from Pyrococcus furiosus and from Thermotoga maritima, and unexpectedly, to the N-terminal amino acid sequence of the homodimeric pyruvate:ferredoxin oxidoreductase from proteobacteria and from cyanobacteria. No sequence similarities were found, however, between the alpha-subunits of the enzyme from A. fulgidus and the heterodimeric pyruvate:ferredoxin oxidoreductase from Halobacterium halobium.

Amino Acid Sequence↗

High-throughput modeling of human G-protein coupled receptors: amino acid sequence alignment, three-dimensional model building, and receptor library screening.

The current study describes the development of a computer package (GPCRmod) aimed at the high-throughput modeling of the therapeutically important family of human G-protein coupled receptors (GPCRs). GPCRmod first proposes a reliable alignment of the seven transmembrane domains (7 TMs) of most druggable human GPCRs based on pattern/motif recognition for each of the 7 TMs that are considered independently. It then converts the alignment into knowledge-based three-dimensional (3-D) models starting from a set of 3-D backbone templates and two separate rotamer libraries for side chain positioning. The 7 TMs of 277 human GPCRs have been accurately aligned, unambiguously clustered in three different classes (rhodopsin-like, secretin-like, metabotropic glutamate-like), and converted into high-quality 3-D models at a remarkable throughput (ca. 3s/model). A 3-D GPCR target library of 277 receptors has consequently been setup. Its utility for "in silico" inverse screening purpose has been demonstrated by recovering among top scorers the receptor of a selective GPCR antagonist as well as the receptors of a promiscuous antagonist. The current GPCR target library thus constitutes a 3-D database of choice to address as soon as possible the "virtual selectivity" profile of any GPCR antagonist or inverse agonist in an early hit optimization process.

Amino Acid Sequence↗

A multiple sequence alignment program.

A program is described for simultaneously aligning two or more molecular sequences which is based on first finding common segments above a specified length and then piecing these together to maximize an alignment scoring function. Optimal as well as near-optimal alignments are found, and there is also provided a means for randomizing the given sequences for testing the statistical significance of an alignment. Alignments may be made in the original alphabets of the sequences or in user-specified alternate ones to take advantage of chemical similarities (such as hydrophobic-hydrophilic).

Amino Acid Sequence↗

The practical use of the A* algorithm for exact multiple sequence alignment.

Multiple alignment is an important problem in computational biology. It is well known that it can be solved exactly by a dynamic programming algorithm which in turn can be interpreted as a shortest path computation in a directed acyclic graph. The A* algorithm (or goal-directed unidirectional search) is a technique that speeds up the computation of a shortest path by transforming the edge lengths without losing the optimality of the shortest path. We implemented the A* algorithm in a computer program similar to MSA (Gupta et al., 1995) and FMA (Shibuya and Imai, 1997). We incorporated in this program new bounding strategies for both lower and upper bounds and show that the A* algorithm, together with our improvements, can speed up computations considerably. Additionally, we show that the A* algorithm together with a standard bounding technique is superior to the well-known Carrillo-Lipman bounding since it excludes more nodes from consideration.

Algorithms↗

Estimating statistical significance of sequence alignments.

Algorithms that compare two proteins or DNA sequences and produce an alignment of the best matching segments are widely used in molecular biology. These algorithms produce scores that when comparing random sequences of length n grow proportional to n or to log(n) depending on the algorithm parameters. The Azuma-Hoeffding inequality gives an upper bound on the probability of large deviations of the score from its mean in the linear case. Poisson approximation can be applied in the logarithmic case.

Algorithms↗

Improvements to GALA and dbERGE II: databases featuring genomic sequence alignment, annotation and experimental results.

We describe improvements to two databases that give access to information on genomic sequence similarities, functional elements in DNA and experimental results that demonstrate those functions. GALA, the database of Genome ALignments and Annotations, is now a set of interlinked relational databases for five vertebrate species, human, chimpanzee, mouse, rat and chicken. For each species, GALA records pairwise and multiple sequence alignments, scores derived from those alignments that reflect the likelihood of being under purifying selection or being a regulatory element, and extensive annotations such as genes, gene expression patterns and transcription factor binding sites. The user interface supports simple and complex queries, including operations such as subtraction and intersections as well as clustering and finding elements in proximity to features. dbERGE II, the database of Experimental Results on Gene Expression, contains experimental data from a variety of functional assays. Both databases are now run on the DB2 database management system. Improved hardware and tuning has reduced response times and increased querying capacity, while simplified query interfaces will help direct new users through the querying process. Links are available at http://www.bx.psu.edu/.

Animals↗