PubMed HealthSearch

SEARCH · PubMed Health

Results for “Multiple sequence alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

A new family of carbon-nitrogen hydrolases.

Using computer methods for database search and multiple alignment, statistically significant sequence similarities were identified between several nitrilases with distinct substrate specificity, cyanide hydratases, aliphatic amidases, beta-alanine synthase, and a few other proteins with unknown molecular function. All these proteins appear to be involved in the reduction of organic nitrogen compounds and ammonia production. Sequence conservation over the entire length, as well as the similarity in the reactions catalyzed by the known enzymes in this family, points to a common catalytic mechanism. The new family of enzymes is characterized by several conserved motifs, one of which contains an invariant cysteine that is part of the catalytic site in nitrilases. Another highly conserved motif includes an invariant glutamic acid that might also be involved in catalysis.

Amidohydrolases

Prediction of structurally conserved regions of D-specific hydroxy acid dehydrogenases by multiple alignment with formate dehydrogenase.

We propose a multiple alignment of the sequence of formate dehydrogenase with the D-specific 2-hydroxy acid dehydrogenases family. Structurally conserved regions are predicted for those sequences corresponding to important regions of the catalytic and the coenzyme binding domains defined from the known three-dimensional structure of the formate dehydrogenase, namely the nicotinamide binding site (beta D to beta F) and the beta A-loop-alpha B region containing the typical glycine pattern of the adenosine binding site, the catalytic histidine/aspartic acid pair and an arginine probably involved in the interaction with the carboxyl group of the substrate.

Alcohol Oxidoreductases

Characterization of the nuclear gene encoding mitochondrial aconitase in the marine red alga Gracilaria verrucosa.

We have cloned a nuclear gene from the marine red alga Gracilaria verrucosa that encodes the complete 779 amino-acid mitochondrial aconitase (m-ACN), the first characterized from a photosynthetic organism. The N-terminal 28 deduced amino acids are predicted to constitute the mitochondrial transit peptide, the first described from a red alga. Putative transcriptional cis-acting elements were identified in the upstream untranslated region. The G. verrucosa m-ACN gene (m-ACN) is present in a single copy and is located ca. 1.5 kb upstream from the single-copy polyubiquitin gene. The single spliceosomal intron is located near the 5' end of the region encoding the mature m-ACN in precisely the same location and phase as intron 2 in Caenorhabditis elegans m-ACN; sequences at its 3' and 5' splice junctions and at the predicted lariat branch point conform well to the eukaryote consensus sequences. Multiple protein-sequence alignment of m-ACN, bacterial aconitase (b-ACN) and iron-responsive element-binding protein (IRE-BP), and phylogenetic analyses, revealed that m-ACN does not share a recent common ancestry with either b-ACN or IRE-BP.

Aconitate Hydratase

Expansion of the mammalian 3 beta-hydroxysteroid dehydrogenase/plant dihydroflavonol reductase superfamily to include a bacterial cholesterol dehydrogenase, a bacterial UDP-galactose-4-epimerase, and open reading frames in vaccinia virus and fish lymphocystis disease virus.

Mammalian 3 beta-hydroxysteroid dehydrogenase and plant dihydroflavonol reductases are descended from a common ancestor. Here we present evidence that Nocardia cholesterol dehydrogenase, E. coli UDP-galactose-4 epimerase, and open reading frames in vaccinia virus and fish lymphocystis disease virus are homologous to 3 beta-hydroxysteroid dehydrogenase and dihydroflavonol reductase. Analysis of a multiple alignment of these sequences indicates that viral ORFs are most closely related to the mammalian 3 beta-hydroxysteroid dehydrogenases. The ancestral protein of this superfamily is likely to be one that metabolized sugar nucleotides. The sequence similarity between 3 beta-hydroxysteroid dehydrogenase and the viral ORFs is sufficient to suggest that these ORFs have an activity that is similar to 3 beta-hydroxysteroid dehydrogenase or cholesterol dehydrogenase, although the putative substrates are not yet known.

3-Hydroxysteroid Dehydrogenases

Computer-assisted dissection of rolling circle DNA replication.

A comparative analysis of the proteins involved in initiation and termination of rolling circle replication (RCR) was performed using computer-assisted methods of data based screening, motif search and multiple amino acid sequence alignment. Two vast classes of such proteins were delineated, one of these being associated with RCR proper, and the other with mobilization (conjugal transfer) of plasmid DNA. The common denominator of the two classes was found to be a conserved amino acid motif that consists of the sequence HisUHisUUU (U--bulky hydrophobic residue; hereafter HUH motif). Based on analogies with metalloenzymes, it is hypothesized that the two conserved His residues this motif may be involved in metal ion coordination required for the activity of the RCR and mobilization proteins. The proteins of the replication (Rep) class contained two additional conserved motifs, with the motif around the Tyr residue(s) forming the covalent link with nicked DNA being located C-proximally of the HUH motif. This class further split into two large superfamilies and several smaller families, with the proteins belonging to a single but not to different (super)families demonstrating statistically significant similarity to each other. Superfamily I, prototyped by the gene A proteins of small isometric single-stranded (ss) DNA bacteriophages, included also Rep proteins of P2-related double-stranded (ds) DNA bacteriophages, the small phage-plasmid hybrid phasyl, and several cyanobacterial and archaebacterial plasmids. These proteins contained two invariant Tyr residues separated by three partially conserved amino acids, suggesting that they all may share the cleavage-ligation mechanism proposed for phi X174 A protein and involving alternate covalent binding of both tyrosines to DNA (Van Mansfeld, A.D., Van Teeffelen, H.A., Baas, P.D., Jansz, H.S., 1986. Nucl. Acids Res. 14, 4229-4238). Superfamily II included Rep proteins of a number of ssDNA plasmids replicating mainly in gram-positive bacteria that unexpectedly were shown to be related to the Rep proteins of plant geminiviruses. Conservation of the "HUH" motif and a motif around the putative DNA-linking Tyr residue was observed also in the Rep proteins of animal parvoviruses containing linear ssDNA with a terminal hairpin and replicating via the rolling hairpin mechanism. The class of plasmid mobilization (Mob) proteins was characterized by the opposite orientation of the conserved motifs, with the (putative) DNA-linking Tyr being located N-proximally of the "HUH" motif.(ABSTRACT TRUNCATED AT 400 WORDS)

Amino Acid Sequence

Mutagenesis and the molecular modeling of the rat angiotensin II receptor (AT1).

The molecular interaction involved in the ligand binding of the rat angiotensin II receptor (AT1A) was studied by site-directed mutagenesis and receptor model building. The three-dimensional structure of AT1A was constructed on the basis of a multiple amino acid sequence alignment of seven transmembrane domain receptors and angiotensin II receptors and after the beta 2 adrenergic receptor model built on the template of the bacteriorhodopsin structure. These data indicated that there are conserved residues that are actively involved in the receptor-ligand interaction. Eleven conserved residues in AT1, His166, Arg167, Glu173, His183, Glu185, Lys199, Trp253, His256, Phe259, Thr260, and Asp263, were targeted individually for site-directed mutation to Ala. Using COS-7 cells transiently expressing these mutated receptors, we found that the binding of angiotensin II was not affected in three of the mutations in the second extracellular loop, whereas the ligand binding affinity was greatly reduced in mutants Lys199-->Ala, Trp253-->Ala, Phe259-->Ala, Asp263-->Ala, and Arg167-->Ala. These amino acid residues appeared to provide binding sites for Ang II. The molecular modeling provided useful structural information for the peptide hormone receptor AT1A. Binding of EXP985, a nonpeptide angiotensin II antagonist, was found to be involved with Arg167 but not Lys199.

Amino Acid Sequence

RNA sequence of potato virus X strain HB.

The genomic RNA of the potato virus X (PVX) strain HB, isolated in Bolivia and able to overcome all known resistance genes, has been cloned and sequenced. The PVXHB RNA sequence is 6432 nucleotides long and contains, similarly to the RNAs of other PVX strains, five open reading frames encoding proteins of M(r)s 165.1K, 24.5K, 12.4K, 7.6K and 25.1K (coat protein), respectively. Multiple amino acid sequence alignments of the coat proteins of four PVX strains identified eight amino acid residues unique for PVXHB. Structural prediction comparisons of the coat proteins of PVXHB and of the other strains suggest a general structural similarity. However, two of the eight amino acid residues unique for strain HB gave rise to a change in the predicted coat protein structure, suggesting a possible involvement in the resistance-breaking activity of PVXHB.

Amino Acid Sequence

Molecular dissection of the beta subunit of F1-ATPase into peptide fragments.

Partial digestion of the native beta subunit of F1-ATPase from the thermophilic Bacillus strain PS3 by three different proteases produced a limited number of peptide fragments. In most cases, the peptides remained associated, and the gross structure of the beta subunit was not destroyed. Furthermore, most peptides were able to reassociate into the form of the beta subunit after denaturating urea treatment. Therefore, the cleaved sites are most likely located in water-exposed loop regions in the tertiary structure of the protein. Almost all peptides were analyzed, and 17 cleaved sites were determined. From the analysis of the distribution of cleaved sites and deletions or insertions in the multiple amino acid sequence alignment of proteins homologous to the beta subunit, locations of five loops and four candidate loops in the beta subunit are suggested. There are two large loops in the central region of the beta subunit sequence, and dicyclohexylcarbodiimide-reactive Glu190 is located in one of them. Tyr341, involved in putative catalytic ATP binding, is also found in one of the loops. Then, taking cleaved sites as a reference, two kinds of expression plasmids, each of which carried genes of two complementary peptide fragments, 1-193 and 198-473 or 1-284 and 285-473, were constructed and expressed in Escherichia coli. For each plasmid, two peptides were coexpressed, associated into a stable beta subunit form in E. coli cells, and purified without dissociation. When these beta subunits were denatured by urea and applied to polyacrylamide gel without denaturant, a protein band with the same mobility as that of the beta subunit appeared, indicating that reassociation of peptide fragments into the form of the beta subunit occurred upon removal of urea. These beta subunits retained the ability to reconstitute the alpha 3 beta 3 gamma complexes even though the efficiency of reconstitution and the recovered ATPase activities were decreased. These complexes were stable at high or low temperature, and ATPase activities were sensitive to inhibition by N3-.

Amino Acid Sequence

fRagmentomics: an R package for integrating cell-free DNA fragment features with mutational status to support liquid biopsy interpretation.

SUMMARY: Liquid biopsy offers a non-invasive approach to study tumor-derived genetic material circulating in plasma. Beyond genetic alterations, the fragmentomic features of cell-free DNA-such as fragment size, genomic position, and end-motifs-provide valuable insights into the biological and clinical context of DNA release. fRagmentomics is a user-friendly R package designed to characterize cfDNA fragments overlapping one or multiple small mutations of any type, starting from an aligned sequencing file (BAM). It supports multiple mutation input formats, accommodates one-based and zero-based genomic conventions, resolves mutation representation ambiguities, and accepts any reference file in FASTA format. For each fragment overlapping a mutation of interest, fRagmentomics outputs fragment-level features including its fragment size, end-motifs, and mutational status, along with additional fragment-level or read-level information. The package implements an indel-aware and optionally soft-clip-preserving fragment size computation that improves accuracy over conventional size estimates based solely on aligned positions. AVAILABILITY AND IMPLEMENTATION: fRagmentomics is licensed under GNU General Public License v3.0 and available at https://github.com/ElsaB-Lab/fRagmentomics, https://anaconda.org/elsab-lab/r-fragmentomics and https://bioconductor.org/packages/fRagmentomics, with documentation and a tutorial. CONTACT: yoann.pradat@gustaveroussy.fr, elsa.bernard@gustaveroussy.fr. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Software

Playing with blocks: some pitfalls of forcing multiple alignments.

Block alignments of multiple amino acid sequences are useful representations of regions thought to share common ancestry and function. Often the block alignments are motivated by the expectation that a protein of interest is similar in function to members of a family of proteins. However, when alignments are forced by using ad hoc methods, it is often difficult to decide whether the proposed relationship is valid. Visual examination can be deceptive, especially when alignments are not carried out in the context of controls subjected to similar procedures. Even computer-aided methods can be misleading when biases are introduced. To illustrate some of the problems that can arise, a few examples from the literature are analyzed. It is concluded that when standard methods fail to find an interesting block alignment unaided by human intervention, then the result should be regarded with caution.

Amino Acid Sequence

Flexible protein sequence patterns. A sensitive method to detect weak structural similarities.

The concept of a flexible protein sequence pattern is defined. In contrast to conventional pattern matching, template or sequence alignment methods, flexible patterns allow residue patterns typical of a complete protein fold to be developed in terms of residue positions (elements), separated by gaps of defined range. An efficient dynamic programming algorithm is presented to enable the best alignment(s) of a pattern with a sequence to be identified. The flexible pattern method is evaluated in detail by reference to the globin protein family, and by comparison to alignment techniques that exploit single sequence, multiple sequence and secondary structural information. A flexible pattern derived from seven globins aligned on structural criteria successfully discriminates all 345 globins from non-globins in the Protein Identification Resource database. Furthermore, a pattern that uses helical regions from just human alpha-haemoglobin identified 337 globins compared to 318 for the best non-pattern global alignment method. Patterns derived from successively fewer, yet more highly conserved positions in a structural alignment of seven globins show that as few as 38 residue positions (25 buried hydrophobic, 4 exposed and 9 others) may be used to uniquely identify the globin fold. The study suggests that flexible patterns gain discriminating power both by discarding regions known to vary within the protein family, and by defining gaps within specific ranges. Flexible patterns therefore provide a convenient and powerful bridge between regular expression pattern matching techniques and more conventional local and global sequence comparison algorithms.

Amino Acid Sequence

Use of long sequence alignments to study the evolution and regulation of mammalian globin gene clusters.

The determination of long segments of DNA sequences encompassing the beta- and alpha-globin gene clusters has provided an unprecedented data base for analysis of genome evolution and regulation of gene clusters. A newly developed computer tool kit generates local alignments between such long sequences in a space-efficient manner, helps the user analyze the alignments effectively, and finds consistently aligning blocks of sequences in multiple pairwise comparisons. Such sequence analyses among the beta-like globin gene clusters of human, galago, rabbit, and mouse have revealed the general patterns of evolution of this gene cluster. Alignments in the flanking regions are very useful in assigning orthologous relationships. Investigation of such matches between the mouse and human beta-like globin gene clusters has led to a reassessment of some orthologous assignments in mouse and to a revision of the proposed pathway for evolution of this gene cluster. In general, the interspersed repetitive elements have inserted independently, presumably via a retrotransposition mechanism, in the different mammalian lineages. However, some examples of ancient L1 repeats are found, including one between the epsilon- and gamma-globin genes that appears to have been in the ancestral eutherian gene cluster. Prominent matching sequences are found in a long region 5' to the epsilon-globin gene, the locus control region (LCR) that is a positive regulator of the entire gene cluster. Three-way alignments among the human, goat, and rabbit sequences can extend for > or = 3 kb in part of the LCR (DNase hypersensitive site 3), indicating that the cis-acting components of this complex regulatory region cover a long segment of DNA. In contrast to the beta-like globin gene clusters, the alpha-like globin gene clusters of many mammals occur in very G+C-rich isochores and contain prominent CpG islands. The regions between the alpha-like globin genes are evolving faster than the intergenic regions of the beta-like globin gene clusters. The contrasts between the two gene clusters can be attributed to differences in DNA metabolism in the isochore. The proximal control elements of the rabbit alpha-globin gene are located both 5' to and within the gene. All of this region is part of a prominent CpG island that may be acting as an extended, enhancer-independent promoter. One can hypothesize that the analogue to the LCR in the alpha-globin gene cluster may interface with the distinctive alpha-globin promoter in ways different from the interaction between the beta LCR and the promoters of beta-like globin genes.(ABSTRACT TRUNCATED AT 400 WORDS)

Animals

Characterization of the active site and thermostability regions of endoxylanase from Thermoanaerobacterium saccharolyticum B6A-RI.

Deletion mutants were constructed from pZEP12, which contained the intact Thermoanaerobacterium saccharolyticum endoxylanase gene (xynA). Deletion of 1.75 kb from the N-terminal end of xynA resulted in a mutant enzyme that retained activity but lost thermostability. Deletion of 1.05 kb from the C terminus did not alter thermostability or activity. The deduced amino acid sequence of T. saccharolyticum B6A-RI endoxylanase XynA was aligned with five other family F beta-glycanases by using the PILEUP program of the Genetics Computer Group package. This multiple alignment of amino acid sequences revealed six highly conserved motifs which included the consensus sequence consisting of a hydrophobic amino acid, Ser or Thr, Glu, a hydrophobic amino acid, Asp, and a hydrophobic amino acid in the catalytic domain. Endoxylanase was inhibited by EDAC [1-(3-dimethylamino propenyl)-3-ethylcarbodiimide hydrochloride], suggesting that Asp and/or Glu was involved in catalysis. Three aspartic acids, two glutamic acids, and one histidine were conserved in all six enzymes aligned. Hydrophobic cluster analysis revealed that two Asp and one Glu occur in the same hydrophobic clusters in T. saccharolyticum B6A-RI endoxylanase and two other enzymes belonging to family F beta-glycanases and suggests their involvement in a catalytic triad. These two Asp and one Glu in XynA from T. saccharolyticum were targeted for analysis by site-specific mutagenesis. Substitution of Asp-537 and Asp-602 by Asn and Glu-600 by Gln completely destroyed endoxylanase activity. These results suggest that these three amino acids form a catalytic triad that functions in a general acid catalysis mechanism.

Amino Acid Sequence

Building multiple alignments from pairwise alignments.

Given a family of related sequences, one can first determine alignments between various pairs of those sequences, then construct a simultaneous alignment of all the sequences that is determined in a natural manner by the set of pairwise alignments. This approach is sometimes effective for exposing the existence and locations of conserved regions, which can then be aligned by more sensitive multiple-alignment methods. This paper presents an efficient algorithm for constructing a multiple alignment from a set of pairwise alignments.

Algorithms

A-liner: linear alignment visualizer for genome comparisons.

SUMMARY: A-liner is a flexible command-line tool for linear visualization of genome-scale sequence alignments, supporting outputs from multiple aligners and integrated visualization of annotations, highlights, quantitative tracks, and coordinate scales. It is applicable to a wide range of organisms, from bacteria to large eukaryotic genomes, and facilitates efficient generation of publication-ready comparative genome visualizations. AVAILABILITY AND IMPLEMENTATION: The source code and example output files for a-liner are available in the GitHub repository: https://github.com/mokuno3430/a-liner. A-liner v1.1.0 has been archived on Zenodo at https://doi.org/10.5281/zenodo.19702001.

Software

ANTHEPROT 2.0: a three-dimensional module fully coupled with protein sequence analysis methods.

ANTHEPROT is a fully interactive graphics program devoted to the analysis of the sequences and structures of proteins. This program, originally developed to facilitate the protein sequence analysis coupled with multiple alignments and predicted secondary structures of proteins, now comprises a powerful 3D module to display and handle macromolecular structures. All the methods that were previously integrated into ANTHEPROT are now directly coupled with a 3D window that provides the user all the classic features of a molecular modeling package. Indeed, it allows real-time rotation and translation of 3D structures with many kinds of models in depth-cueing mode (space filling, backbone, wire models, main chain, and ribbons), selections (atom type, residue type, segments, and chain), color-coding systems (amino acid properties, predicted or observed secondary structures, temperature B factor, and subunits), geometric calculations (Ramachandran plot, distances, and angles), and fitting molecules. Stereo views are possible as well as HPGL standard files. A module specifically devoted to the determination of 3D structures using nuclear magnetic resonance is also available. This major release of our program for IBM rs6000 workstations is available by anonymous ftp to ibcp.fr for academic institutions.

Antigens

Similarities between alanine dehydrogenase and the N-terminal part of pyridine nucleotide transhydrogenase and their possible implication in the virulence mechanism of Mycobacterium tuberculosis.

Recent developments in simultaneous multiple alignment methods of protein sequences allow prediction of structural similarity in related proteins. Alanine dehydrogenase and the N-terminal sequence of pyridine nucleotide transhydrogenase were compared for their sequences. High similarities of sequences were observed especially in their NAD(H)-binding sites. These similarities suggest that antibodies which recognized the alanine dehydrogenase of Mycobacterium tuberculosis can also be directed against the membrane bound pyridine nucleotide transhydrogenase. If this is the case, the virulent property of this pathogen could be linked to its higher synthesis of NADPH necessary for its anabolism.

Alanine Dehydrogenase

Human T-cell receptor variable gene segment families.

Multiple DNA and protein sequence alignments have been constructed for the human T-cell receptor alpha/delta, beta, and gamma (TCRA/D, B, and G) variable (V) gene segments. The traditional classification into subfamilies was confirmed using a much larger pool of sequences. For each sequence, a name was derived which complies with the standard nomenclature. The traditional numbering of V gene segments in the order of their discovery was continued and changed when in conflict with names of other segments. By discriminating between alleles at the same locus versus genes from different loci, we were able to reduce the number of more than 150 different TCRBV sequences in the database to a repertoire of only 47 functional TCRBV gene segments. An extension of this analysis to the over 100 TCRAV sequences results in a predicted repertoire of 42 functional TCRAV gene segments. Our alignment revealed two residues that distinguish between the highly homologous V delta and V alpha, one at a site that in VH contacts the constant region, the other at the interface between immunoglobulin VH and VL. This site may be responsible for restricted pairing between certain V delta and V gamma chains. On the other hand, V beta and V gamma appear to be related by the fact that their CDR2 length is increased by four residues as compared with that of V alpha/delta peptides.

Alleles