PubMed HealthSearch

SEARCH · PubMed Health

Results for “Multiple sequence alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Calculating percent identity between protein or DNA sequences with a word processor.

Two macros, to calculate percentage identity between protein or DNA sequences using the Microsoft Word word processor, are described. The user prepares an alignment file of multiple sequences which is used by the macros to calculate number of matches, number of mismatches, total number of compared positions, and the percent identity. The macros are especially useful when alignment of multiple sequences is possible only by eye.

Algorithms

Multiple alignment and hierarchical clustering of conserved amino acid sequences in the replication-associated proteins of plant RNA viruses.

We have used multiple alignment computer programs to align and hierarchically cluster the conserved amino acid "signature" sequences found in the replication-associated proteins of all plant RNA viruses sequenced so far. These regions, called "polymerase", "nucleotide-binding" and "N-terminal" are well conserved even between viruses which are only distantly related, and are thus very well suited for this type of analysis. Our results show that the clusterings obtained using these very short amino acid sequences are very robust to computing parameters and are surprisingly well matched with the taxonomic grouping of RNA plant viruses. The possibility of using this system as a new taxonomic criterion is discussed.

Amino Acid Sequence

Analysis of conserved domains and sequence motifs in cellular regulatory proteins and locus control regions using new software tools for multiple alignment and visualization.

With the tremendous expansion of molecular sequence data in recent years, multiple alignment is arguably one of the two most important analytic techniques (the other being fast database searching). A number of useful approaches to this problem have previously been developed, but often they are limited to only a subset of multiple-alignment applications and cannot easily deal with the complex structural organization seen in an increasing number of sequences. For example, a single sequence may contain several domains of different evolutionary origins, and the multiplicities and relative ordering of these domains may be quite different among related sequences. Here we describe an integrated set of interactive Unix tools that combines several multiple-alignment techniques with traditional "dot-plot" visualization to provide a flexible environment for approaching complex sequence analysis problems. We apply these tools to the identification and characterization of "catalytic" domains in ras and rho/rac GTPase-activating proteins, to "Src homology" (SH2, SH3) domains in cytoplasmic signaling proteins, to repetitive sequence motifs in the alpha and beta subunits of protein prenyltransferases, and to regulatory DNA sequences in the locus control region of the beta-globin gene cluster.

Alkyl and Aryl Transferases

Evolutionary relationship between the TonB-dependent outer membrane transport proteins: nucleotide and amino acid sequences of the Escherichia coli colicin I receptor gene.

The nucleotide sequence of the Escherichia coli colicin I receptor gene (cir) has been determined. The predicted mature protein consists of 599 amino acids and has a molecular weight of 67,169. Several previously noted characteristics of other E. coli outer membrane protein sequences were also identified in the sequence of Cir. These include an overall acidic nature, the absence of long hydrophobic stretches of amino acids, and a lack of predicted alpha-helical secondary structure. Because two classes of outer membrane proteins (the TonB-dependent transport proteins and the porins) share some structural features, protein sequences from both of these groups were aligned pairwise and scored for sequence similarity. Statistical evidence suggested that the porins were not related to the proteins in the TonB-dependent group; however, there was a significant relationship between the proteins in the TonB-dependent group. On the basis of the multiple progressive sequence alignment and the similarity scores derived from it, a tree representing evolutionary distance between five TonB-dependent outer membrane transport proteins was generated.

Amino Acid Sequence

The inference of evolutionary trees from molecular data.

1. Procedures for multiple alignment of sequence data, subsequent phylogenetic inference, and testing of the trees derived are presented. 2. The assumptions underlying different approaches and the extent to which they are valid are discussed.

Amino Acid Sequence

Molecular evolution of class A beta-lactamases: phylogeny and patterns of sequence conservation.

We present a multiple alignment of the amino acid sequences of eight class A beta-lactamases and utilized it to propose a phylogeny, based on the nucleotide sequences of their corresponding genes. We have also used the alignment, together with the alpha-carbon co-ordinates of the Staphylococcus aureus protein, to search systematically for neighbouring residues that share the same pattern of conservation among the different members of the protein family. The distribution of invariant residues and of groups of residues with co-ordinate changes map, predominantly, at the region of the active site and at interfaces between structural elements, respectively. We have also contrasted the distribution of conserved residues with the positions which are known to differ in mutants and variants of class A beta-lactamases.

Amino Acid Sequence

Eukaryotic DNA polymerase amino acid sequence required for 3'----5' exonuclease activity.

We have identified an amino-proximal sequence motif, Phe-Asp-Ile-Glu-Thr, in Saccharomyces cerevisiae DNA polymerase II that is almost identical to a sequence comprising part of the 3'----5' exonuclease active site of Escherichia coli DNA polymerase I. Similar motifs were identified by amino acid sequence alignment in related, aphidicolin-sensitive DNA polymerases possessing 3'----5' proofreading exonuclease activity. Substitution of Ala for the Asp and Glu residues in the motif reduced the exonuclease activity of partially purified DNA polymerase II at least 100-fold while preserving the polymerase activity. Yeast strains expressing the exonuclease-deficient DNA polymerase II had on average about a 22-fold increase in spontaneous mutation rate, consistent with a presumed proofreading role in vivo. In multiple amino acid sequence alignments of this and two other conserved motifs described previously, five residues of the 3'----5' exonuclease active site of E. coli DNA polymerase I appeared to be invariant in aphidicolin-sensitive DNA polymerases known to possess 3'----5' proofreading exonuclease activity. None of these residues, however, appeared to be identifiable in the catalytic subunits of human, yeast, or Drosophila alpha DNA polymerases.

Amino Acid Sequence

Selection of circularization sites in a group I IVS RNA requires multiple alignments of an internal template-like sequence.

Circularization and reverse circularization of the Tetrahymena thermophila rRNA intervening sequence resemble the first and second steps in splicing, respectively. However, site-specific base substitutions show that different nucleotides are involved in selection of the 5' splice site and the circularization sites. Furthermore, a substitution at the major circularization site that prevents circularization can be suppressed by second substitutions at two different nucleotide positions. A model is proposed in which adjacent and overlapping sequences can function as a binding site, forming a short duplex with the sequence at the circularization site and thus directing circularization and reverse circularization. Because the 5' exon-binding site and three potential circularization binding sites fall within a contiguous eight nucleotide region, this sequence may translocate relative to the catalytic core of the ribozyme in a template-like manner.

Animals

KCFtools: rapid alignment-free method for introgression screening and GWAS using k-mer profiles.

MOTIVATION: In the era of multiple genome references, researchers often align sequencing reads against distinct assemblies or even multiple references simultaneously. This enables applications such as the detection of introgressed segments or highly variable genomic regions, which are especially prevalent in large-genome crop species such as lettuce or wheat. However, these applications come at the cost of increased computational burden, inconsistencies in mapping methods, and reduced reproducibility across studies. To address these limitations, we developed KCFtools, a Java-based toolkit that identifies the presence and absence of k-mers in nonoverlapping genomic or transcriptomic windows by comparing query and reference genomes. This alignment-free approach enables the efficient computation of an identity score for each window, thereby facilitating robust detection of introgressed or variable regions across genomes. RESULTS: We systematically evaluated the performance and accuracy of the k-mer-based method implemented in KCFtools, benchmarking it against conventional single nucleotide variation-based introgression detection pipelines. Our results demonstrate that KCFtools effectively captures introgressed segments and structurally diverse regions, even in species with fragmented or highly divergent reference genomes. In addition, we extended KCFtools to generate genotype matrices from k-mer variation tables. These matrices are compatible with genome-wide association studies software and allow the identification of loci associated with phenotypic traits. We showcase the utility of this approach by detecting known and novel associations for downy mildew resistance in lettuce, underscoring the pipeline's potential for high-resolution, reference-agnostic population genetic analysis. AVAILABILITY AND IMPLEMENTATION: https://github.com/sivasubramanics/kcftools.

Software

Herpesviral deoxythymidine kinases contain a site analogous to the phosphoryl-binding arginine-rich region of porcine adenylate kinase; comparison of secondary structure predictions and conservation.

Twelve herpesviral deoxythymidine kinases were examined for regions of sequence similarity by multiple alignment. Six highly conserved sites were observed. Site 1 corresponded to a glycine-rich loop that forms part of the ATP-binding pocket in porcine adenylate kinase (PAK), and site 5 corresponded to a region in PAK, located on one lobe of the cleft, that contains arginine residues that bind substrate phosphoryl groups. Site 3, consisting of the motif -DRH-, is thought to be involved in thymine/deoxythymidine recognition; site 4, which is nearby, probably participates in this function as well. The functions of sites 2 and 6 have not been identified. Secondary structure predictions were made by the Garnier method and averaged for each position in the multiple alignment. The structure predicted for all six sites was typically a short flexible region (turn or coil) at or adjacent to the site, flanked by rigid structures (helix or sheet) on either side.

Adenylate Kinase

Novel GACG-hairpin pair motif in the 5' untranslated region of type C retroviruses related to murine leukemia virus.

We searched for the presence of common RNA structural motifs in mammalian type C retroviruses related to murine leukemia viruses and the closely related avian spleen necrosis virus. A novel motif consisting of a pair of hairpins, called hairpin pair motif, was detected in the 5' untranslated regions of the genomes of these retroviruses. A combination of computational analyses that included the assessment of phylogenetic sequence conservation by multiple alignment, the search for regions with unusual RNA folding properties, and the analysis of RNA secondary structure by suboptimal free-energy calculations highlighted the significance of this hairpin pair motif. The hairpin pair motif encompasses 70 to 80 nucleotides between the splice donor site and the gag translational initiation codon of these viruses. The motif is composed of two adjacent hairpins both with a perfectly conserved GACG tetraloop. We propose that the novel GACG-hairpin pair motif described here constitutes an essential component of the regulatory machinery in these type C retroviruses.

Base Sequence

Expansion of the mammalian 3 beta-hydroxysteroid dehydrogenase/plant dihydroflavonol reductase superfamily to include a bacterial cholesterol dehydrogenase, a bacterial UDP-galactose-4-epimerase, and open reading frames in vaccinia virus and fish lymphocystis disease virus.

Mammalian 3 beta-hydroxysteroid dehydrogenase and plant dihydroflavonol reductases are descended from a common ancestor. Here we present evidence that Nocardia cholesterol dehydrogenase, E. coli UDP-galactose-4 epimerase, and open reading frames in vaccinia virus and fish lymphocystis disease virus are homologous to 3 beta-hydroxysteroid dehydrogenase and dihydroflavonol reductase. Analysis of a multiple alignment of these sequences indicates that viral ORFs are most closely related to the mammalian 3 beta-hydroxysteroid dehydrogenases. The ancestral protein of this superfamily is likely to be one that metabolized sugar nucleotides. The sequence similarity between 3 beta-hydroxysteroid dehydrogenase and the viral ORFs is sufficient to suggest that these ORFs have an activity that is similar to 3 beta-hydroxysteroid dehydrogenase or cholesterol dehydrogenase, although the putative substrates are not yet known.

3-Hydroxysteroid Dehydrogenases

Molecular dissection of the beta subunit of F1-ATPase into peptide fragments.

Partial digestion of the native beta subunit of F1-ATPase from the thermophilic Bacillus strain PS3 by three different proteases produced a limited number of peptide fragments. In most cases, the peptides remained associated, and the gross structure of the beta subunit was not destroyed. Furthermore, most peptides were able to reassociate into the form of the beta subunit after denaturating urea treatment. Therefore, the cleaved sites are most likely located in water-exposed loop regions in the tertiary structure of the protein. Almost all peptides were analyzed, and 17 cleaved sites were determined. From the analysis of the distribution of cleaved sites and deletions or insertions in the multiple amino acid sequence alignment of proteins homologous to the beta subunit, locations of five loops and four candidate loops in the beta subunit are suggested. There are two large loops in the central region of the beta subunit sequence, and dicyclohexylcarbodiimide-reactive Glu190 is located in one of them. Tyr341, involved in putative catalytic ATP binding, is also found in one of the loops. Then, taking cleaved sites as a reference, two kinds of expression plasmids, each of which carried genes of two complementary peptide fragments, 1-193 and 198-473 or 1-284 and 285-473, were constructed and expressed in Escherichia coli. For each plasmid, two peptides were coexpressed, associated into a stable beta subunit form in E. coli cells, and purified without dissociation. When these beta subunits were denatured by urea and applied to polyacrylamide gel without denaturant, a protein band with the same mobility as that of the beta subunit appeared, indicating that reassociation of peptide fragments into the form of the beta subunit occurred upon removal of urea. These beta subunits retained the ability to reconstitute the alpha 3 beta 3 gamma complexes even though the efficiency of reconstitution and the recovered ATPase activities were decreased. These complexes were stable at high or low temperature, and ATPase activities were sensitive to inhibition by N3-.

Amino Acid Sequence

fRagmentomics: an R package for integrating cell-free DNA fragment features with mutational status to support liquid biopsy interpretation.

SUMMARY: Liquid biopsy offers a non-invasive approach to study tumor-derived genetic material circulating in plasma. Beyond genetic alterations, the fragmentomic features of cell-free DNA-such as fragment size, genomic position, and end-motifs-provide valuable insights into the biological and clinical context of DNA release. fRagmentomics is a user-friendly R package designed to characterize cfDNA fragments overlapping one or multiple small mutations of any type, starting from an aligned sequencing file (BAM). It supports multiple mutation input formats, accommodates one-based and zero-based genomic conventions, resolves mutation representation ambiguities, and accepts any reference file in FASTA format. For each fragment overlapping a mutation of interest, fRagmentomics outputs fragment-level features including its fragment size, end-motifs, and mutational status, along with additional fragment-level or read-level information. The package implements an indel-aware and optionally soft-clip-preserving fragment size computation that improves accuracy over conventional size estimates based solely on aligned positions. AVAILABILITY AND IMPLEMENTATION: fRagmentomics is licensed under GNU General Public License v3.0 and available at https://github.com/ElsaB-Lab/fRagmentomics, https://anaconda.org/elsab-lab/r-fragmentomics and https://bioconductor.org/packages/fRagmentomics, with documentation and a tutorial. CONTACT: yoann.pradat@gustaveroussy.fr, elsa.bernard@gustaveroussy.fr. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Software

Playing with blocks: some pitfalls of forcing multiple alignments.

Block alignments of multiple amino acid sequences are useful representations of regions thought to share common ancestry and function. Often the block alignments are motivated by the expectation that a protein of interest is similar in function to members of a family of proteins. However, when alignments are forced by using ad hoc methods, it is often difficult to decide whether the proposed relationship is valid. Visual examination can be deceptive, especially when alignments are not carried out in the context of controls subjected to similar procedures. Even computer-aided methods can be misleading when biases are introduced. To illustrate some of the problems that can arise, a few examples from the literature are analyzed. It is concluded that when standard methods fail to find an interesting block alignment unaided by human intervention, then the result should be regarded with caution.

Amino Acid Sequence

Flexible protein sequence patterns. A sensitive method to detect weak structural similarities.

The concept of a flexible protein sequence pattern is defined. In contrast to conventional pattern matching, template or sequence alignment methods, flexible patterns allow residue patterns typical of a complete protein fold to be developed in terms of residue positions (elements), separated by gaps of defined range. An efficient dynamic programming algorithm is presented to enable the best alignment(s) of a pattern with a sequence to be identified. The flexible pattern method is evaluated in detail by reference to the globin protein family, and by comparison to alignment techniques that exploit single sequence, multiple sequence and secondary structural information. A flexible pattern derived from seven globins aligned on structural criteria successfully discriminates all 345 globins from non-globins in the Protein Identification Resource database. Furthermore, a pattern that uses helical regions from just human alpha-haemoglobin identified 337 globins compared to 318 for the best non-pattern global alignment method. Patterns derived from successively fewer, yet more highly conserved positions in a structural alignment of seven globins show that as few as 38 residue positions (25 buried hydrophobic, 4 exposed and 9 others) may be used to uniquely identify the globin fold. The study suggests that flexible patterns gain discriminating power both by discarding regions known to vary within the protein family, and by defining gaps within specific ranges. Flexible patterns therefore provide a convenient and powerful bridge between regular expression pattern matching techniques and more conventional local and global sequence comparison algorithms.

Amino Acid Sequence

A-liner: linear alignment visualizer for genome comparisons.

SUMMARY: A-liner is a flexible command-line tool for linear visualization of genome-scale sequence alignments, supporting outputs from multiple aligners and integrated visualization of annotations, highlights, quantitative tracks, and coordinate scales. It is applicable to a wide range of organisms, from bacteria to large eukaryotic genomes, and facilitates efficient generation of publication-ready comparative genome visualizations. AVAILABILITY AND IMPLEMENTATION: The source code and example output files for a-liner are available in the GitHub repository: https://github.com/mokuno3430/a-liner. A-liner v1.1.0 has been archived on Zenodo at https://doi.org/10.5281/zenodo.19702001.

Software