PubMed HealthSearch

SEARCH · PubMed Health

Results for “Multiple Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Molecular evolution of class A beta-lactamases: phylogeny and patterns of sequence conservation.

We present a multiple alignment of the amino acid sequences of eight class A beta-lactamases and utilized it to propose a phylogeny, based on the nucleotide sequences of their corresponding genes. We have also used the alignment, together with the alpha-carbon co-ordinates of the Staphylococcus aureus protein, to search systematically for neighbouring residues that share the same pattern of conservation among the different members of the protein family. The distribution of invariant residues and of groups of residues with co-ordinate changes map, predominantly, at the region of the active site and at interfaces between structural elements, respectively. We have also contrasted the distribution of conserved residues with the positions which are known to differ in mutants and variants of class A beta-lactamases.

Amino Acid Sequence

Eukaryotic DNA polymerase amino acid sequence required for 3'----5' exonuclease activity.

We have identified an amino-proximal sequence motif, Phe-Asp-Ile-Glu-Thr, in Saccharomyces cerevisiae DNA polymerase II that is almost identical to a sequence comprising part of the 3'----5' exonuclease active site of Escherichia coli DNA polymerase I. Similar motifs were identified by amino acid sequence alignment in related, aphidicolin-sensitive DNA polymerases possessing 3'----5' proofreading exonuclease activity. Substitution of Ala for the Asp and Glu residues in the motif reduced the exonuclease activity of partially purified DNA polymerase II at least 100-fold while preserving the polymerase activity. Yeast strains expressing the exonuclease-deficient DNA polymerase II had on average about a 22-fold increase in spontaneous mutation rate, consistent with a presumed proofreading role in vivo. In multiple amino acid sequence alignments of this and two other conserved motifs described previously, five residues of the 3'----5' exonuclease active site of E. coli DNA polymerase I appeared to be invariant in aphidicolin-sensitive DNA polymerases known to possess 3'----5' proofreading exonuclease activity. None of these residues, however, appeared to be identifiable in the catalytic subunits of human, yeast, or Drosophila alpha DNA polymerases.

Amino Acid Sequence

Selection of circularization sites in a group I IVS RNA requires multiple alignments of an internal template-like sequence.

Circularization and reverse circularization of the Tetrahymena thermophila rRNA intervening sequence resemble the first and second steps in splicing, respectively. However, site-specific base substitutions show that different nucleotides are involved in selection of the 5' splice site and the circularization sites. Furthermore, a substitution at the major circularization site that prevents circularization can be suppressed by second substitutions at two different nucleotide positions. A model is proposed in which adjacent and overlapping sequences can function as a binding site, forming a short duplex with the sequence at the circularization site and thus directing circularization and reverse circularization. Because the 5' exon-binding site and three potential circularization binding sites fall within a contiguous eight nucleotide region, this sequence may translocate relative to the catalytic core of the ribozyme in a template-like manner.

Animals

KCFtools: rapid alignment-free method for introgression screening and GWAS using k-mer profiles.

MOTIVATION: In the era of multiple genome references, researchers often align sequencing reads against distinct assemblies or even multiple references simultaneously. This enables applications such as the detection of introgressed segments or highly variable genomic regions, which are especially prevalent in large-genome crop species such as lettuce or wheat. However, these applications come at the cost of increased computational burden, inconsistencies in mapping methods, and reduced reproducibility across studies. To address these limitations, we developed KCFtools, a Java-based toolkit that identifies the presence and absence of k-mers in nonoverlapping genomic or transcriptomic windows by comparing query and reference genomes. This alignment-free approach enables the efficient computation of an identity score for each window, thereby facilitating robust detection of introgressed or variable regions across genomes. RESULTS: We systematically evaluated the performance and accuracy of the k-mer-based method implemented in KCFtools, benchmarking it against conventional single nucleotide variation-based introgression detection pipelines. Our results demonstrate that KCFtools effectively captures introgressed segments and structurally diverse regions, even in species with fragmented or highly divergent reference genomes. In addition, we extended KCFtools to generate genotype matrices from k-mer variation tables. These matrices are compatible with genome-wide association studies software and allow the identification of loci associated with phenotypic traits. We showcase the utility of this approach by detecting known and novel associations for downy mildew resistance in lettuce, underscoring the pipeline's potential for high-resolution, reference-agnostic population genetic analysis. AVAILABILITY AND IMPLEMENTATION: https://github.com/sivasubramanics/kcftools.

Software

Herpesviral deoxythymidine kinases contain a site analogous to the phosphoryl-binding arginine-rich region of porcine adenylate kinase; comparison of secondary structure predictions and conservation.

Twelve herpesviral deoxythymidine kinases were examined for regions of sequence similarity by multiple alignment. Six highly conserved sites were observed. Site 1 corresponded to a glycine-rich loop that forms part of the ATP-binding pocket in porcine adenylate kinase (PAK), and site 5 corresponded to a region in PAK, located on one lobe of the cleft, that contains arginine residues that bind substrate phosphoryl groups. Site 3, consisting of the motif -DRH-, is thought to be involved in thymine/deoxythymidine recognition; site 4, which is nearby, probably participates in this function as well. The functions of sites 2 and 6 have not been identified. Secondary structure predictions were made by the Garnier method and averaged for each position in the multiple alignment. The structure predicted for all six sites was typically a short flexible region (turn or coil) at or adjacent to the site, flanked by rigid structures (helix or sheet) on either side.

Adenylate Kinase

Novel GACG-hairpin pair motif in the 5' untranslated region of type C retroviruses related to murine leukemia virus.

We searched for the presence of common RNA structural motifs in mammalian type C retroviruses related to murine leukemia viruses and the closely related avian spleen necrosis virus. A novel motif consisting of a pair of hairpins, called hairpin pair motif, was detected in the 5' untranslated regions of the genomes of these retroviruses. A combination of computational analyses that included the assessment of phylogenetic sequence conservation by multiple alignment, the search for regions with unusual RNA folding properties, and the analysis of RNA secondary structure by suboptimal free-energy calculations highlighted the significance of this hairpin pair motif. The hairpin pair motif encompasses 70 to 80 nucleotides between the splice donor site and the gag translational initiation codon of these viruses. The motif is composed of two adjacent hairpins both with a perfectly conserved GACG tetraloop. We propose that the novel GACG-hairpin pair motif described here constitutes an essential component of the regulatory machinery in these type C retroviruses.

Base Sequence

Characterization of the nuclear gene encoding mitochondrial aconitase in the marine red alga Gracilaria verrucosa.

We have cloned a nuclear gene from the marine red alga Gracilaria verrucosa that encodes the complete 779 amino-acid mitochondrial aconitase (m-ACN), the first characterized from a photosynthetic organism. The N-terminal 28 deduced amino acids are predicted to constitute the mitochondrial transit peptide, the first described from a red alga. Putative transcriptional cis-acting elements were identified in the upstream untranslated region. The G. verrucosa m-ACN gene (m-ACN) is present in a single copy and is located ca. 1.5 kb upstream from the single-copy polyubiquitin gene. The single spliceosomal intron is located near the 5' end of the region encoding the mature m-ACN in precisely the same location and phase as intron 2 in Caenorhabditis elegans m-ACN; sequences at its 3' and 5' splice junctions and at the predicted lariat branch point conform well to the eukaryote consensus sequences. Multiple protein-sequence alignment of m-ACN, bacterial aconitase (b-ACN) and iron-responsive element-binding protein (IRE-BP), and phylogenetic analyses, revealed that m-ACN does not share a recent common ancestry with either b-ACN or IRE-BP.

Aconitate Hydratase

Expansion of the mammalian 3 beta-hydroxysteroid dehydrogenase/plant dihydroflavonol reductase superfamily to include a bacterial cholesterol dehydrogenase, a bacterial UDP-galactose-4-epimerase, and open reading frames in vaccinia virus and fish lymphocystis disease virus.

Mammalian 3 beta-hydroxysteroid dehydrogenase and plant dihydroflavonol reductases are descended from a common ancestor. Here we present evidence that Nocardia cholesterol dehydrogenase, E. coli UDP-galactose-4 epimerase, and open reading frames in vaccinia virus and fish lymphocystis disease virus are homologous to 3 beta-hydroxysteroid dehydrogenase and dihydroflavonol reductase. Analysis of a multiple alignment of these sequences indicates that viral ORFs are most closely related to the mammalian 3 beta-hydroxysteroid dehydrogenases. The ancestral protein of this superfamily is likely to be one that metabolized sugar nucleotides. The sequence similarity between 3 beta-hydroxysteroid dehydrogenase and the viral ORFs is sufficient to suggest that these ORFs have an activity that is similar to 3 beta-hydroxysteroid dehydrogenase or cholesterol dehydrogenase, although the putative substrates are not yet known.

3-Hydroxysteroid Dehydrogenases

Mutagenesis and the molecular modeling of the rat angiotensin II receptor (AT1).

The molecular interaction involved in the ligand binding of the rat angiotensin II receptor (AT1A) was studied by site-directed mutagenesis and receptor model building. The three-dimensional structure of AT1A was constructed on the basis of a multiple amino acid sequence alignment of seven transmembrane domain receptors and angiotensin II receptors and after the beta 2 adrenergic receptor model built on the template of the bacteriorhodopsin structure. These data indicated that there are conserved residues that are actively involved in the receptor-ligand interaction. Eleven conserved residues in AT1, His166, Arg167, Glu173, His183, Glu185, Lys199, Trp253, His256, Phe259, Thr260, and Asp263, were targeted individually for site-directed mutation to Ala. Using COS-7 cells transiently expressing these mutated receptors, we found that the binding of angiotensin II was not affected in three of the mutations in the second extracellular loop, whereas the ligand binding affinity was greatly reduced in mutants Lys199-->Ala, Trp253-->Ala, Phe259-->Ala, Asp263-->Ala, and Arg167-->Ala. These amino acid residues appeared to provide binding sites for Ang II. The molecular modeling provided useful structural information for the peptide hormone receptor AT1A. Binding of EXP985, a nonpeptide angiotensin II antagonist, was found to be involved with Arg167 but not Lys199.

Amino Acid Sequence

Molecular dissection of the beta subunit of F1-ATPase into peptide fragments.

Partial digestion of the native beta subunit of F1-ATPase from the thermophilic Bacillus strain PS3 by three different proteases produced a limited number of peptide fragments. In most cases, the peptides remained associated, and the gross structure of the beta subunit was not destroyed. Furthermore, most peptides were able to reassociate into the form of the beta subunit after denaturating urea treatment. Therefore, the cleaved sites are most likely located in water-exposed loop regions in the tertiary structure of the protein. Almost all peptides were analyzed, and 17 cleaved sites were determined. From the analysis of the distribution of cleaved sites and deletions or insertions in the multiple amino acid sequence alignment of proteins homologous to the beta subunit, locations of five loops and four candidate loops in the beta subunit are suggested. There are two large loops in the central region of the beta subunit sequence, and dicyclohexylcarbodiimide-reactive Glu190 is located in one of them. Tyr341, involved in putative catalytic ATP binding, is also found in one of the loops. Then, taking cleaved sites as a reference, two kinds of expression plasmids, each of which carried genes of two complementary peptide fragments, 1-193 and 198-473 or 1-284 and 285-473, were constructed and expressed in Escherichia coli. For each plasmid, two peptides were coexpressed, associated into a stable beta subunit form in E. coli cells, and purified without dissociation. When these beta subunits were denatured by urea and applied to polyacrylamide gel without denaturant, a protein band with the same mobility as that of the beta subunit appeared, indicating that reassociation of peptide fragments into the form of the beta subunit occurred upon removal of urea. These beta subunits retained the ability to reconstitute the alpha 3 beta 3 gamma complexes even though the efficiency of reconstitution and the recovered ATPase activities were decreased. These complexes were stable at high or low temperature, and ATPase activities were sensitive to inhibition by N3-.

Amino Acid Sequence

fRagmentomics: an R package for integrating cell-free DNA fragment features with mutational status to support liquid biopsy interpretation.

SUMMARY: Liquid biopsy offers a non-invasive approach to study tumor-derived genetic material circulating in plasma. Beyond genetic alterations, the fragmentomic features of cell-free DNA-such as fragment size, genomic position, and end-motifs-provide valuable insights into the biological and clinical context of DNA release. fRagmentomics is a user-friendly R package designed to characterize cfDNA fragments overlapping one or multiple small mutations of any type, starting from an aligned sequencing file (BAM). It supports multiple mutation input formats, accommodates one-based and zero-based genomic conventions, resolves mutation representation ambiguities, and accepts any reference file in FASTA format. For each fragment overlapping a mutation of interest, fRagmentomics outputs fragment-level features including its fragment size, end-motifs, and mutational status, along with additional fragment-level or read-level information. The package implements an indel-aware and optionally soft-clip-preserving fragment size computation that improves accuracy over conventional size estimates based solely on aligned positions. AVAILABILITY AND IMPLEMENTATION: fRagmentomics is licensed under GNU General Public License v3.0 and available at https://github.com/ElsaB-Lab/fRagmentomics, https://anaconda.org/elsab-lab/r-fragmentomics and https://bioconductor.org/packages/fRagmentomics, with documentation and a tutorial. CONTACT: yoann.pradat@gustaveroussy.fr, elsa.bernard@gustaveroussy.fr. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Software

Playing with blocks: some pitfalls of forcing multiple alignments.

Block alignments of multiple amino acid sequences are useful representations of regions thought to share common ancestry and function. Often the block alignments are motivated by the expectation that a protein of interest is similar in function to members of a family of proteins. However, when alignments are forced by using ad hoc methods, it is often difficult to decide whether the proposed relationship is valid. Visual examination can be deceptive, especially when alignments are not carried out in the context of controls subjected to similar procedures. Even computer-aided methods can be misleading when biases are introduced. To illustrate some of the problems that can arise, a few examples from the literature are analyzed. It is concluded that when standard methods fail to find an interesting block alignment unaided by human intervention, then the result should be regarded with caution.

Amino Acid Sequence

Flexible protein sequence patterns. A sensitive method to detect weak structural similarities.

The concept of a flexible protein sequence pattern is defined. In contrast to conventional pattern matching, template or sequence alignment methods, flexible patterns allow residue patterns typical of a complete protein fold to be developed in terms of residue positions (elements), separated by gaps of defined range. An efficient dynamic programming algorithm is presented to enable the best alignment(s) of a pattern with a sequence to be identified. The flexible pattern method is evaluated in detail by reference to the globin protein family, and by comparison to alignment techniques that exploit single sequence, multiple sequence and secondary structural information. A flexible pattern derived from seven globins aligned on structural criteria successfully discriminates all 345 globins from non-globins in the Protein Identification Resource database. Furthermore, a pattern that uses helical regions from just human alpha-haemoglobin identified 337 globins compared to 318 for the best non-pattern global alignment method. Patterns derived from successively fewer, yet more highly conserved positions in a structural alignment of seven globins show that as few as 38 residue positions (25 buried hydrophobic, 4 exposed and 9 others) may be used to uniquely identify the globin fold. The study suggests that flexible patterns gain discriminating power both by discarding regions known to vary within the protein family, and by defining gaps within specific ranges. Flexible patterns therefore provide a convenient and powerful bridge between regular expression pattern matching techniques and more conventional local and global sequence comparison algorithms.

Amino Acid Sequence

Protein database searches for multiple alignments.

Protein database searches frequently can reveal biologically significant sequence relationships useful in understanding structure and function. Weak but meaningful sequence patterns can be obscured, however, by other similarities due only to chance. By searching a database for multiple as opposed to pairwise alignments, distant relationships are much more easily distinguished from background noise. Recent statistical results permit the power of this approach to be analyzed. Given a typical query sequence, an algorithm described here permits the current protein database to be searched for three-sequence alignments in less than 4 min. Such searches have revealed a variety of subtle relationships that pairwise search methods would be unable to detect.

Algorithms

A-liner: linear alignment visualizer for genome comparisons.

SUMMARY: A-liner is a flexible command-line tool for linear visualization of genome-scale sequence alignments, supporting outputs from multiple aligners and integrated visualization of annotations, highlights, quantitative tracks, and coordinate scales. It is applicable to a wide range of organisms, from bacteria to large eukaryotic genomes, and facilitates efficient generation of publication-ready comparative genome visualizations. AVAILABILITY AND IMPLEMENTATION: The source code and example output files for a-liner are available in the GitHub repository: https://github.com/mokuno3430/a-liner. A-liner v1.1.0 has been archived on Zenodo at https://doi.org/10.5281/zenodo.19702001.

Software

ANTHEPROT 2.0: a three-dimensional module fully coupled with protein sequence analysis methods.

ANTHEPROT is a fully interactive graphics program devoted to the analysis of the sequences and structures of proteins. This program, originally developed to facilitate the protein sequence analysis coupled with multiple alignments and predicted secondary structures of proteins, now comprises a powerful 3D module to display and handle macromolecular structures. All the methods that were previously integrated into ANTHEPROT are now directly coupled with a 3D window that provides the user all the classic features of a molecular modeling package. Indeed, it allows real-time rotation and translation of 3D structures with many kinds of models in depth-cueing mode (space filling, backbone, wire models, main chain, and ribbons), selections (atom type, residue type, segments, and chain), color-coding systems (amino acid properties, predicted or observed secondary structures, temperature B factor, and subunits), geometric calculations (Ramachandran plot, distances, and angles), and fitting molecules. Stereo views are possible as well as HPGL standard files. A module specifically devoted to the determination of 3D structures using nuclear magnetic resonance is also available. This major release of our program for IBM rs6000 workstations is available by anonymous ftp to ibcp.fr for academic institutions.

Antigens

aPhyloGeo: a Python application for correlating genetic and climatic conditions.

MOTIVATION: Environmental variation and its influence on genetic diversity is a central topic in evolutionary biology and phylogeography. Accurate correlations between genetic and climatic datasets to understand the genetic adaptations of different species to specific environments. It requires integrated and reproducible workflows. RESULTS: We developed aPhyloGeo, an open-source and multiplatform application implemented in Python, for investigating correlations between genetic variation and environmental data within a phylogenetic framework. The workflow integrates multiple analytical steps, including sequence alignment, sliding window phylogenetic inference, and statistical approaches such as the Mantel test and the Procrustean randomization test. These analyses enable the identification of mutation hotspots that exhibit strong associations with environmental variables. In addition, aPhyloGeo supports multicore data processing and provides a fully reproducible pipeline for evaluating localized relationships between genomic variation and climatic distributions. AVAILABILITY AND IMPLEMENTATION: aPhyloGeo is freely available on GitHub at: https://github.com/tahiri-lab/aPhyloGeo, as both a PyPI package and as Python scripts for Linux, macOS, and Windows.

Software