PubMed HealthSearch

SEARCH · PubMed Health

Results for “Multiple Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Rapidly evolving aphid gall effector proteins exhibit saposin-like folds.

Many insects manipulate plants by injecting effector proteins. In one extreme example of this molecular "hijacking", Hormaphis cornu aphids inject bicycle proteins into Hamamelis virginiana (Witch Hazel), contributing to the development of novel organs called galls. Bicycle proteins share no amino acid sequence similarity with proteins of known function. Here, we report the crystal structures of two divergent bicycle proteins. Both proteins contain saposin-like folds: one with multiple disulfide bonds exhibits a helix swap; the other has no disulfide bonds and possesses two tandem domains. To explore the structural evolution of bicycle proteins, we predicted bicycle protein structures with Alphafold2 (AF2). While AF2 did not recover the two experimental structures using existing databases, it succeeded after we provided multiple sequence alignments (MSAs) containing protein sequences encoded in new genome sequences from closely related aphid species. Using this customized approach at scale, we generated 2400 high-confidence predictions for bicycle proteins from seven aphid species. This dataset revealed that bicycle proteins without cysteines are outliers in fold space and appear to have evolved from ancestral proteins with disulfide-bonded saposin-like folds. While all bicycle proteins contain predicted saposin-like folds, they display a vast diversity of structural and physicochemical properties. While this diversity thwarts prediction of conserved functions encoded in structure, it suggests that bicycle proteins have evolved to target diverse plant processes and/or to evade plant immune surveillance.

AlphaFold predictions

KCFtools: rapid alignment-free method for introgression screening and GWAS using k-mer profiles.

MOTIVATION: In the era of multiple genome references, researchers often align sequencing reads against distinct assemblies or even multiple references simultaneously. This enables applications such as the detection of introgressed segments or highly variable genomic regions, which are especially prevalent in large-genome crop species such as lettuce or wheat. However, these applications come at the cost of increased computational burden, inconsistencies in mapping methods, and reduced reproducibility across studies. To address these limitations, we developed KCFtools, a Java-based toolkit that identifies the presence and absence of k-mers in nonoverlapping genomic or transcriptomic windows by comparing query and reference genomes. This alignment-free approach enables the efficient computation of an identity score for each window, thereby facilitating robust detection of introgressed or variable regions across genomes. RESULTS: We systematically evaluated the performance and accuracy of the k-mer-based method implemented in KCFtools, benchmarking it against conventional single nucleotide variation-based introgression detection pipelines. Our results demonstrate that KCFtools effectively captures introgressed segments and structurally diverse regions, even in species with fragmented or highly divergent reference genomes. In addition, we extended KCFtools to generate genotype matrices from k-mer variation tables. These matrices are compatible with genome-wide association studies software and allow the identification of loci associated with phenotypic traits. We showcase the utility of this approach by detecting known and novel associations for downy mildew resistance in lettuce, underscoring the pipeline's potential for high-resolution, reference-agnostic population genetic analysis. AVAILABILITY AND IMPLEMENTATION: https://github.com/sivasubramanics/kcftools.

Software

Novel GACG-hairpin pair motif in the 5' untranslated region of type C retroviruses related to murine leukemia virus.

We searched for the presence of common RNA structural motifs in mammalian type C retroviruses related to murine leukemia viruses and the closely related avian spleen necrosis virus. A novel motif consisting of a pair of hairpins, called hairpin pair motif, was detected in the 5' untranslated regions of the genomes of these retroviruses. A combination of computational analyses that included the assessment of phylogenetic sequence conservation by multiple alignment, the search for regions with unusual RNA folding properties, and the analysis of RNA secondary structure by suboptimal free-energy calculations highlighted the significance of this hairpin pair motif. The hairpin pair motif encompasses 70 to 80 nucleotides between the splice donor site and the gag translational initiation codon of these viruses. The motif is composed of two adjacent hairpins both with a perfectly conserved GACG tetraloop. We propose that the novel GACG-hairpin pair motif described here constitutes an essential component of the regulatory machinery in these type C retroviruses.

Base Sequence

fRagmentomics: an R package for integrating cell-free DNA fragment features with mutational status to support liquid biopsy interpretation.

SUMMARY: Liquid biopsy offers a non-invasive approach to study tumor-derived genetic material circulating in plasma. Beyond genetic alterations, the fragmentomic features of cell-free DNA-such as fragment size, genomic position, and end-motifs-provide valuable insights into the biological and clinical context of DNA release. fRagmentomics is a user-friendly R package designed to characterize cfDNA fragments overlapping one or multiple small mutations of any type, starting from an aligned sequencing file (BAM). It supports multiple mutation input formats, accommodates one-based and zero-based genomic conventions, resolves mutation representation ambiguities, and accepts any reference file in FASTA format. For each fragment overlapping a mutation of interest, fRagmentomics outputs fragment-level features including its fragment size, end-motifs, and mutational status, along with additional fragment-level or read-level information. The package implements an indel-aware and optionally soft-clip-preserving fragment size computation that improves accuracy over conventional size estimates based solely on aligned positions. AVAILABILITY AND IMPLEMENTATION: fRagmentomics is licensed under GNU General Public License v3.0 and available at https://github.com/ElsaB-Lab/fRagmentomics, https://anaconda.org/elsab-lab/r-fragmentomics and https://bioconductor.org/packages/fRagmentomics, with documentation and a tutorial. CONTACT: yoann.pradat@gustaveroussy.fr, elsa.bernard@gustaveroussy.fr. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Software

A-liner: linear alignment visualizer for genome comparisons.

SUMMARY: A-liner is a flexible command-line tool for linear visualization of genome-scale sequence alignments, supporting outputs from multiple aligners and integrated visualization of annotations, highlights, quantitative tracks, and coordinate scales. It is applicable to a wide range of organisms, from bacteria to large eukaryotic genomes, and facilitates efficient generation of publication-ready comparative genome visualizations. AVAILABILITY AND IMPLEMENTATION: The source code and example output files for a-liner are available in the GitHub repository: https://github.com/mokuno3430/a-liner. A-liner v1.1.0 has been archived on Zenodo at https://doi.org/10.5281/zenodo.19702001.

Software

aPhyloGeo: a Python application for correlating genetic and climatic conditions.

MOTIVATION: Environmental variation and its influence on genetic diversity is a central topic in evolutionary biology and phylogeography. Accurate correlations between genetic and climatic datasets to understand the genetic adaptations of different species to specific environments. It requires integrated and reproducible workflows. RESULTS: We developed aPhyloGeo, an open-source and multiplatform application implemented in Python, for investigating correlations between genetic variation and environmental data within a phylogenetic framework. The workflow integrates multiple analytical steps, including sequence alignment, sliding window phylogenetic inference, and statistical approaches such as the Mantel test and the Procrustean randomization test. These analyses enable the identification of mutation hotspots that exhibit strong associations with environmental variables. In addition, aPhyloGeo supports multicore data processing and provides a fully reproducible pipeline for evaluating localized relationships between genomic variation and climatic distributions. AVAILABILITY AND IMPLEMENTATION: aPhyloGeo is freely available on GitHub at: https://github.com/tahiri-lab/aPhyloGeo, as both a PyPI package and as Python scripts for Linux, macOS, and Windows.

Software

Distribution and molecular characterization of integron classes from Escherichia coli and Klebsiella pneumoniae isolates in Sulaymaniyah province of Iraq.

UNLABELLED: The environmental pollution from the misuse of antimicrobial drugs is fueling selection pressure in bacteria, thereby exacerbating the threat to global health. In Iraq, the situation is made worse by the poor implementation of the World Health Organization's Global Antimicrobial Resistance and Use Surveillance System (WHO-GLASS). Consequently, this study aimed to increase surveillance of the spread of antimicrobial resistance in Sulaymaniyah, Iraq. A total of 296 Enterobacteriaceae comprising 147 Klebsiella pneumoniae and 149 Escherichia coli were isolated from humans, poultry, and dairy farms. The isolates were screened using multiplex PCR to assess the prevalence of the clinically important integron integrase (intI) classes and antimicrobial resistance genes (ARGs) of commonly used antibiotics. Remarkably, 81.14% of the isolates carried at least 2 ARGs, 10.47% intI1, and 3.72% intI2. No intI3 was detected. A total of 663 ARGs were identified using multiplex PCR in the two Enterobacteriaceae: beta-lactamase genes were 43%, tetracycline resistance genes 25.20%, sulfonamide resistance gene 16.10%, quinolone resistance gene 10.2%, and aminoglycoside resistance genes 5.7%. K. pneumoniae harbored more integrons and ARGs than E. coli, thus posing a higher antimicrobial resistance threat in this province. This study underscores the importance of implementing more stringent WHO-GLASS and antibiotic stewardship to end the multidrug resistance crisis in Iraq. IMPORTANCE: These data are about the prevalence of integrons and resistance genes, helping to fill a significant gap in global surveillance efforts. Results can be used by global health authorities and the World Health Organization to develop national and international antimicrobial resistance (AMR) control strategies. The study is important because integrons are key genetic platforms that capture and disseminate antibiotic resistance genes among bacteria. In addition, Escherichia coli and Klebsiella spp. are among the top causes of hospital- and community-acquired infections, especially urinary tract infections, bloodstream infections, and pneumonia. Therefore, it will be riskier when these bacteria have a high rate of integrons and resistance genes because it impedes treatments during infection. Another importance of this study is that the study was carried out in Iraq. Iraq, like many low- and middle-income countries, faces challenges with unregulated antibiotic use, leading to high rates of AMR.

Escherichia coli

Cooperation of transposable elements to endow global networks of initiators of hybrid assembly pathways of endogenous multiprotein complexes.

Mechanisms governing initiation steps of the assembly of endogenous multi-protein complexes (EMC) remain incompletely understood. Here, multiple lines of observations are reported describing the function-aligned initiation sequence of hybrid assembly pathways (HAP) of EMC. The first step of HAP-guided chain reactions of protein-protein interactions (PPI) of EMC assemblies constitutes the creation of cell type-specific pools of hetero and homo dimers. The molecular anatomy of HAP was elucidated by defining qualitative and quantitative characteristics of protein binding to a compendium of 200,393 distinct genomic regulatory elements (GRE), including 49,667 sequences representing control sets of genomic loci as well as 150,726 GRE of different evolutionary origins. The consensus sequence of HAP actions consists of: a) Initiation on genomic DNA of the formation of metastable hetero- and homodimers of EMCs' protein constituents; b) Release of dimers from DNA templates for delivery to the EMC assembly compartments; c) Assembly of defined EMC by sequential on demand addition of proteins to preformed dimers serving as attractors of EMC-specific ensembles of monomers. Chromosome-naïve DNA scaffolds facilitating creation of intracellular dimer pools engage networks of ~700 transcription factors (TFs), 534 of which manifest region-specific patterns of significantly enriched expression in 1358 brain regions. HAP initiators appear to operate within nucleosome-depleted islands of transposable elements (TE) - derived sequences within heterochromatin. PPI assembly lines of EMCs operate in 2 concurrent modes: TF-TF PPI cascade and PPI HUB protein cascade. Regardless of the number of DNA-bound initiator TFs (ranging from one to 716 TFs), both modes of operations reached the equilibrium at the PPI constituents saturation levels of ~245 proteins for TF-TF PPI modes and of ~351 proteins for PPI HUB protein modes. Distinct panels of DNA-bound initiator TFs and proteins of PPI cascade ensembles are enriched in either defined sets of neuroanatomical structures (TF-TF mode) or among structural-functional constituents of synapses (HUB proteins mode). Thus, these bifurcated cascades appear biologically congruent: TF-TF constituents map to transcriptional signatures of hundreds of brain regions, whereas HUB constituents map to synaptogenesis and synaptic structures, suggesting the unified logic of genomic functions coordinating region identity and connectivity. Evidence-supported examples of default operations of PPI-guided assemblies of hetero- and homodimers of Yamanaka factors, neurogenesis constituents, and protein components of postsynaptic density of excitatory and inhibitory synaptogenesis are reported with detailed analytical focus on human Claustrum. The foundational set of observations reported in this contribution should facilitate experimental and theoretical explorations of TE-seeded genomic codes for initiators of PPI chain reactions of protein dimerization creating pools of attractors to guide and accelerate the EMC assemblies.

Humans

Genome-wide SNP data support species boundaries in sympatric Polylepis Ruiz & Pav. (Rosaceae) species from Bolivia and Ecuador.

Species delimitation in the South American genus Polylepis is notoriously challenging due to high morphological similarity and phenotypic plasticity, likely driven by hybridization and gene flow. Previous phylogenetic studies suggested that genetic structure aligns more strongly with geography than with taxonomy, questioning existing species concepts and hampering conservation efforts. We used double-digest RAD sequencing (ddRADseq) to generate genome-wide SNP data for 11 Polylepis species sampled across multiple localities in Bolivia and Ecuador. Population genetic analyses, phylogenetic inference, and network approaches were combined to assess whether genetic structure aligns more closely with taxonomy or geography. Morphologically defined species formed largely cohesive genetic lineages across regions, with species identity explaining substantially more genetic variation than locality. While localized admixture and reticulation were detected among closely related taxa, widespread species showed strong genetic cohesion and clear separation from congeners. Our results indicate that the sampled Polylepis species from Bolivia and Ecuador maintain distinct genetic identities despite localized signals consistent with gene flow. This genome-wide support for current taxonomy highlights Polylepis as a valuable model for studying speciation under gene flow and indicates that multiple geographic sampling will be essential in reconstructing a robust phylogeny of the genus, with important implications for conservation planning in Andean montane forests.

Bolivia

Mitochondrial Impostors: Prevalence and Impacts of NUMTs on Genetic and Evolutionary Studies in Carnivora.

Nuclear mitochondrial pseudogenes are mitochondria-derived DNA sequences integrated into the nuclear genome, which can introduce errors in species identification, phylogenetic inference, and population genetics. Although nuclear mitochondrial pseudogene contamination has been reported in some Carnivora species, a systematic investigation into the prevalence and impacts of nuclear mitochondrial pseudogenes across an order is still lacking. In this study, 22,102 mitochondrial DNA sequences of 80 Carnivora species from 14 families and 54 genera were retrieved from the public National Center for Biotechnology Information database and further analyzed. Using alignment-based methods, 158 problematic sequences/sequence groups were identified and categorized into four types: nuclear mitochondrial pseudogenes, species misidentification or mislabeling, sequence errors, and anomalous sites. Among families, Felidae exhibited the highest rate of nuclear mitochondrial pseudogene contamination, particularly in species of the genus Panthera. In contrast, no nuclear mitochondrial pseudogene contamination was detected in members of Ursidae and Ailuridae. Phylogenetic analysis revealed multiple independent origins of nuclear mitochondrial pseudogene, with some tracing back to the common ancestor of Carnivora. To mitigate nuclear mitochondrial pseudogene-related errors, rigorous sequence verification strategies, such as sequence alignment and phylogenetic validation, should be implemented. In conclusion, our findings highlight the necessity of nuclear mitochondrial pseudogene awareness in genetic and evolutionary studies of Carnivora and other taxa.

Animals

High-accuracy SNV calling for bacterial isolates using deep learning with AccuSNV.

Accurate detection of mutations within bacterial species is critical for fundamental studies of microbial evolution, reconstruction of transmission events, and identification of antimicrobial resistance mutations. Although many tools have been developed to identify single-nucleotide variants (SNVs) from whole-genome sequencing, they often suffer from high false-positive rates owing to the complexity of bacterial genomes and the need for different filtering cutoffs across sample types and sequencing depths. As data sets increase in size, the manual filtering required for high accuracy presents a significant obstacle. Here, we present AccuSNV, a novel deep learning-based tool for high-precision and automated bacterial SNV calling. Unlike traditional methods that process one sample at a time, AccuSNV leverages a convolutional neural network (CNN) that integrates alignment information across multiple samples, enhancing precision through learned across-sample patterns. We evaluate AccuSNV against seven popular SNV-calling tools using simulated data from six bacterial species with varied sequencing depths, numbers of isolates, mutations, and divergence levels. To further validate its real-world utility, we test AccuSNV on multiple curated bacterial data sets containing reported SNVs. In both simulated and real-world scenarios, AccuSNV consistently achieves the best performance. Moreover, AccuSNV provides comprehensive user-friendly downstream analysis modules and outputs, including mutation annotation information, phylogenetic inference, d N/d S calculations, and optional manual filtering. Together with the automated deep learning-based calling, these features make AccuSNV broadly accessible to users with different levels of computational expertise.

Deep Learning

Ulysses transposable element of Drosophila shows high structural similarities to functional domains of retroviruses.

We have determined the DNA structure of the Ulysses transposable element of Drosophila virilis and found that this transposon is 10,653 bp and is flanked by two unusually large direct repeats 2136 bp long. Ulysses shows the characteristic organization of LTR-containing retrotransposons, with matrix and capsid protein domains encoded in the first open reading frame. In addition, Ulysses contains protease, reverse transcriptase, RNase H and integrase domains encoded in the second open reading frame. Ulysses lacks a third open reading frame present in some retrotransposons that could encode an env-like protein. A dendrogram analysis based on multiple alignments of the protease, reverse transcriptase, RNase H, integrase and tRNA primer binding site of all known Drosophila LTR-containing retrotransposon sequences establishes a phylogenetic relationship of Ulysses to other retrotransposons and suggests that Ulysses belongs to a new family of this type of elements.

Amino Acid Sequence

Alignment-free integration of single-nucleus ATAC-seq across species with sPYce.

Changes in gene regulation largely contribute to differences in cellular identities and phenotypes between species. Single-nucleus assays for transposase-accessible chromatin with sequencing (snATAC-seq) are an efficient strategy to identify putative gene regulatory elements and provide new insight into evolutionary divergence of regulatory programmes. However, no dedicated framework exists to integrate and compare snATAC-seq data across species, while methods designed for single-cell gene expression data have serious limitations. Here we present sPYce, a cross-species snATAC-seq integration method that relies on sequence composition similarities through k-mer histograms of regulatory regions, removing the need for genome alignments to anchor data from different species. sPYce can embed datasets from multiple species into the same mathematical space and permits further downstream analysis steps. We benchmarked sPYce against existing approaches on two publicly available datasets spanning more than 160 myr of evolution, showing that it successfully uncovers conserved cellular programmes while preserving biologically relevant species-specific differences. By comparing cerebellar development in mice and opossums, sPYce identifies regulatory divergence in granule cell differentiation programmes, particularly driven by nuclear factor 1. As an easy-to-use, alignment-free cross-species snATAC-seq integration approach, sPYce opens new perspectives to compare gene regulatory evolution across species.

Animals

MAFin: motif detection in multiple alignment files.

MOTIVATION: Whole Genome and Proteome Alignments, represented by the multiple alignment file format, have become a standard approach in comparative genomics and proteomics. These often require identifying conserved motifs, which is crucial for understanding functional and evolutionary relationships. However, current approaches lack a direct method for motif detection within MAF files. We present MAFin, a novel tool that enables efficient motif detection and conservation analysis in MAF files to address this gap, streamlining genomic and proteomic research. RESULTS: We developed MAFin, the first motif detection tool for Multiple Alignment Format files. MAFin enables the multithreaded search of conserved motifs using three approaches: (i) using user-specified k-mers to search the sequences. (ii) with regular expressions, in which case one or more patterns are searched, and (iii) with predefined Position Weight Matrices. Once the motif has been found, MAFin detects the motif instances and calculates the conservation across the aligned sequences. MAFin also calculates a conservation percentage, which provides information about the conservation levels of each motif across the aligned sequences, based on the number of matches relative to the length of the motif. A set of statistics enables the interpretation of each motif's conservation level, and the detected motifs are exported in JSON and CSV files for downstream analyses. AVAILABILITY AND IMPLEMENTATION: MAFin is offered as a Python package under the GPL license as a multi-platform application and is available at: https://github.com/Georgakopoulos-Soares-lab/MAFin.

Software

Detection of Caenorhabditis transposon homologs in diverse organisms.

Although transposons that move via DNA intermediates are common in bacteria, invertebrates, and plants, none have been clearly documented in vertebrates and certain other classes of organisms. One such family of transposons includes invertebrate elements related to Caenorhabditis elegans Tc1. Blocks of aligned protein segments derived from this family were used to search a nucleotide sequence databank. Among the relatives detected were known bacterial insertion elements, revealing the ancient origin of the family. Furthermore, a Tc1-like homolog was detected in a catfish, raising the possibility that this valuable tool of C. elegans genetics can be used with vertebrate genomes. This study illustrates the use of multiple protein blocks for detection and evaluation of distant relationships.

Amino Acid Sequence

PanDelos-plus: A parallel algorithm for computing sequence homology in pangenomic analysis.

The identification of homologous gene families across multiple genomes is a central task in bacterial pangenomics traditionally requiring computationally demanding all-against-all comparisons. PanDelos addresses this challenge with an alignment-free and parameter-free approach based on k-mer profiles, combining high speed, ease of use, and competitive accuracy with state-of-the-art methods. However, the increasing availability of genomic data requires tools that can scale efficiently to larger datasets. To address this need, we present PanDelos-plus, a fully parallel, gene-centric redesign of PanDelos. The algorithm parallelizes the most computationally intensive phases (Best Hit detection and Bidirectional Best Hit extraction) through data decomposition and a thread pool strategy, while employing lightweight data structures to reduce memory usage. Benchmarks on synthetic datasets show that PanDelos-plus achieves up to 14x faster execution and reduces memory usage by up to 96%, while maintaining consistency with the original algorithm. These improvements allow the PanDelos methodology to be applied to population-scale comparative genomics, thus enabling more precise characterisation of pangenome structure and dynamics. PanDelos-plus is available at github.com/synbionics/PanDelos-plus.

Journal Article

A database of protein structure families with common folding motifs.

The availability of fast and robust algorithms for protein structure comparison provides an opportunity to produce a database of three-dimensional comparisons, called families of structurally similar proteins (FSSP). The database currently contains an extended structural family for each of 154 representative (below 30% sequence identity) protein chains. Each data set contains: the search structure; all its relatives with 70-30% sequence identity, aligned structurally; and all other proteins from the representative set that contain substructures significantly similar to the search structure. Very close relatives (above 70% sequence identity) rarely have significant structural differences and are excluded. The alignments of remote relatives are the result of pairwise all-against-all structural comparisons in the set of 154 representative protein chains. The comparisons were carried out with each of three novel automatic algorithms that cover different aspects of protein structure similarity. The user of the database has the choice between strict rigid-body comparisons and comparisons that take into account interdomain motion or geometrical distortions; and, between comparisons that require strictly sequential ordering of segments and comparisons, which allow altered topology of loop connections or chain reversals. The data sets report the structurally equivalent residues in the form of a multiple alignment and as a list of matching fragments to facilitate inspection by three-dimensional graphics. If substructures are ignored, the result is a database of structure alignments of full-length proteins, including those in the twilight zone of sequence similarity.(ABSTRACT TRUNCATED AT 250 WORDS)

Algorithms

DNA sequence determinants for binding of transformed Ah receptor to a dioxin-responsive enhancer.

We have utilized gel retardation analysis and DNA mutagenesis to examine the specific interaction of transformed guinea pig hepatic cytosolic TCDD.AhR complex with a dioxin-responsive element (DRE). Sequence alignment of the mouse CYPIA1 upstream DREs has identified a common invariant "core" consensus sequence of TNGCGTG flanked by several variable nucleotides. Competitive gel retardation analysis using a series of DRE oligonucleotides containing single or multiple base substitutions has allowed identification of those nucleotides important for TCDD.AhR.DRE complex formation. A putative TCDD.AhR DNA-binding consensus sequence of GCGTGNNA/TNNNC/G has been derived. The four core nucleotides, CGTG, appear to be critical for TCDD-inducible protein-DNA complex formation since their substitution decreased AhR binding affinity by 100-800-fold; the remaining conserved bases are also important, albeit to a lesser degree (3-5-fold). The 5'-ward thymine, present in the invariant core sequence of all the DREs identified to date, appears not to be involved in DNA binding of the AhR. The results obtained here indicate that although the primary interaction of the TCDD.AhR complex with the DRE occurs with the conserved "core" sequence, nucleotides flanking the core also contribute to the specificity of DRE binding.

Animals