PubMed HealthSearch

SEARCH · PubMed Health

Results for “Sequence motif”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A Sequence Motif Enables Widespread Use of Non-Canonical Redox Cofactors in Natural Enzymes.

Non-canonical redox cofactors (NRCs) are promising alternatives to nicotinamide adenine dinucleotide (phosphate) (NAD(P)+) for biomanufacturing due to low cost and exquisite electron delivery control, yet their adoption is limited by the scarcity of compatible enzymes. Here, we screened the aldehyde dehydrogenase (ALDH) protein family and identified a conserved RH/QxxR sequence motif that enables widespread NRC activity among natural enzymes. Bos taurus ALDH3a1 and Pseudanabaena biceps ALDH exhibit unprecedented turnover with nicotinamide mononucleotide (NMN+), with kcat values matching or exceeding that of NAD+ and surpassing most engineered NRC-active enzymes by 10 to 105-fold, based on the relative NRC to native activity. Structural and dynamic analyses reveal this motif reinforces cofactor positioning and pre-organizes the active site without dependence on the adenosine monophosphate moiety of NAD+. When introduced into diverse ALDH scaffolds, the RH/QxxR motif enhances NMN+ activity up to 60-fold. In addition to NMN+, this motif also supports activity across multiple non-nucleotide, simple synthetic NRCs such as 1-(2-carbamoylmethyl)nicotinamide (AmNA+). These findings elucidate Nature's solution to the engineering challenge of obtaining NRC-active enzymes and offers a blueprint to mine latent evolutionary plasticity in natural enzymes that serve as superior engineering starting points.

Active site pre-organization

A sequence motif enables widespread use of noncanonical redox cofactors in natural enzymes.

Noncanonical redox cofactors (NRCs) are low-cost alternatives to the natural redox cofactors nicotinamide adenine dinucleotide (NAD+) and nicotinamide adenine dinucleotide phosphate (NADP+) for biomanufacturing, offering exquisite electron-delivery control, yet their adoption is limited by the scarcity of compatible enzymes. Screening the aldehyde dehydrogenase (ALDH) family, we identified a conserved RH/QxxR motif that enables widespread NRC activity among natural enzymes. Bos taurus ALDH3a1 exhibits unprecedented turnover with nicotinamide mononucleotide (NMN+), with kcat values exceeding NAD+ and surpassing most engineered NRC-active enzymes by 10-105-fold. Structural analyses reveal that this motif reinforces cofactor positioning and preorganizes the active site independently of the NAD+ adenosine monophosphate moiety. This motif supports activity across simple-synthetic NRCs such as 1-(2-carbamoylmethyl)nicotinamide and, when introduced into diverse ALDH scaffolds, enhances NMN+ activity up to 60-fold. These findings elucidate nature's solution to engineering NRC-active enzymes and offer a blueprint to mine latent evolutionary plasticity in natural enzymes that serve as superior engineering starting points.

Journal Article

GAMMA: gap-aware motif mining under incomplete labeling with applications to MHC motifs.

MOTIVATION: Sequence motif identification is crucial for understanding molecular recognition, particularly in immune responses involving peptide binding to major histocompatibility complex (MHC) Class I molecules for antigen presentation to T cells. Traditionally, MHC Class I binding motifs are assumed to be contiguous and span nine amino acids. However, structural evidence suggests that binding may involve nonadjacent residues, challenging the assumptions of existing methods. RESULTS: In this study, we propose Gap-Aware Motif Mining Algorithm (GAMMA), a probabilistic framework designed to identify noncontiguous motifs under conditions of incomplete labeling. GAMMA employs Bayesian inference with Markov chain Monte Carlo sampling to jointly estimate motif parameters, binding locations, and the relative spacing between binding positions. Through extensive simulations and real-world applications to MHC Class I peptide datasets, GAMMA outperforms existing motif discovery tools such as GLAM2 in accurately localizing binding residues and identifying the underlying motifs. Notably, our results suggest that the true number of binding residues may be eight, fewer than the commonly assumed nine. In addition, for longer peptides, the model captures increased flexibility in the central region, consistent with structural observations that peptides may bulge in the middle. AVAILABILITY AND IMPLEMENTATION: The raw data and the source codes are available on GitHub (https://github.com/RanLIUaca/GAMMAmotif).

Amino Acid Motifs

An interpretable deep learning framework uncovers features governing CRISPR-Cas9 genome-editing efficiency.

MOTIVATION: CRISPR-Cas9 genome-editing efficiency is strongly influenced by the sequence composition and positional context of single-guide RNAs (sgRNAs). Although numerous deep learning-based models have been developed to predict Cas9 efficiency from sgRNA sequences, most operate as black boxes, offering limited insight into the sequence determinants underlying Cas9 activity. In addition, previous studies often overlook how the positional context of sequence motifs within sgRNAs influences their effects on Cas9 binding or cleavage. RESULTS: We introduce DeepCC9, an interpretable machine learning framework that combines explicit sequence feature extraction with a residual block-based deep architecture to improve interpretability and identify composition- and position-based motifs governing Cas9 genome-editing efficiency. We applied this method to multiple Cas9 variant datasets, achieving superior predictive performance compared with existing methods while enabling direct interpretation of sequence motifs and their positional effects. Our analysis uncovered 74 sequence motifs enriched or depleted at specific positions within sgRNAs and strongly associated with Cas9 efficiency, providing mechanistic insight into sequence features that influence guide performance. Together, these results establish DeepCC9 as a generalizable and interpretable framework for modeling sequence-function relationships and advancing the understanding of the sequence determinants underlying CRISPR-Cas9 genome editing. AVAILABILITY AND IMPLEMENTATION: The authors have implemented their algorithm in the Python programming language (version 3.X), which is accessible using (https://zenodo.org/records/20073890).

Deep Learning

Properties Governing Native State Entanglements and Relationships to Protein Function.

Non-covalent lasso entanglements are structural motifs found in a majority of globular proteins, and their misfolding has been linked to a range of biological consequences. Here, we characterize these motifs' structural and physicochemical properties, sequence biases, functional site correlations, and universal features across E. coli, S. cerevisiae, and H. sapiens. We find that the crossing residues, which pierce the plane of the entanglement loop, are 11-times more likely to be a β-strand than an α-helix or random coil, and that around this position the protein sequence is 2.5-times more likely to be composed of a stretch of all hydrophobic residues (most often Val, Ile, or Phe) compared to other sequence motifs. Functionally, crossing residues are enriched at enzyme active sites in S. cerevisiae and small molecule binding residues across all species to degrees greater than expected by random chance. Metal binding residues are enriched in these entanglements in H. sapiens. Increasing statistical power by pooling together these species data, we find RNA-binding residues are enriched in these entanglement components. On the other hand, there is a spatial depletion of crossing residues at sites involved in protein binding. Using machine learning, we identified eight robust features predictive of these entanglements, achieving AUROC scores of 0.8 across species. These results are significant because they suggest a direct role for components of native entanglements in particular protein functions, as well as identifying strong secondary structure and sequence preferences in native entanglements.

Humans

An atlas of non-redundant sequences and structures of transcription factor assemblies across domains of life.

Transcription factors (TFs) regulate gene expression by controlling the recruitment of transcriptional machinery to regulatory regions of the genome. Nearly 10% of the human genome encodes TFs, making them one of the largest protein families. Despite their central roles in gene regulation, TFs are historically considered challenging therapeutic targets due to their complex interactions with DNA, RNA and associated proteins. Although recent progress in studying TFs both at molecular and structural level excels our understanding on their function, yet a universal rule decoding their recognition process remains elusive. Here, we present a curated non-redundant dataset of TFs with 3570 sequences and 377 structures. We further characterize "unique interfaces" by quantifying interface identity across interacting chains in TF assemblies. Surprisingly, our data shows that the "unique interfaces" have optimal size ranging from 2000 Å2 to 4000 Å2 irrespective of their quaternary assembly. To understand the functional diversity, we integrate sequence motifs, structural domains, subcellular localization and functional enrichment of TFs. We have also catalogued association of TFs with various human diseases. Our dataset provides a comprehensive platform to perform large scale analysis of TF-assemblies and aid in computational methods for their prediction across domains of life.

Gene regulation

An expanded codebook of human transcription factor DNA-binding specificity.

Gene expression is regulated by transcription factors (TFs), which recognize specific DNA sequence motifs. Several hundred putative human TFs, identified mainly by an apparent DNA-binding domain, lack known binding motifs1. Furthermore, even for well-characterized TFs, it remains controversial the degree to which motifs accurately reflect binding sites in living cells2. Here we describe a systematic effort ('Codebook') to determine the sequence specificity of 332 putative and poorly characterized human TFs. More than 4,000 independent experiments, encompassing multiple in vitro and in vivo assays, produced motifs for just over half (177; 53%) of the TFs, of which most are associated with only a single protein. These results extend the vocabulary of sequence recognition encoded by human TFs by around 130 distinct motifs. Moreover, binding motifs identified in vitro are strongly enriched in cellular binding sites. Collectively, the data reveal tens of thousands of previously unknown, conserved and direct TF-binding sites across the human genome. These sites are concentrated in promoter regions and are predictive of gene expression. In summary, this new codebook provides an important step forward in decoding the human genome.

Humans

Evolutionary history and recombination in the mitochondrial carrier SLC25 superfamily analyzed by similarities in the exon and transmembrane α-helix sequences.

Mitochondrial carriers (MCs), which constitute a superfamily also called the solute carrier family 25 (SLC25), are characterized by conserved signature motif sequences and a six-transmembrane α-helical transporter domain. They transport a wide variety of substrates ranging from protons, inorganic ions, citric acid cycle intermediates, and amino acids to nucleotides and cofactors. The superfamily members can be divided into subfamilies, each with a distinct substrate specificity. In an attempt to understand how different subfamilies have evolved, we analyzed the protein sequences of the exons (with conserved boundaries) and the six transmembrane α-helices of MCs from highly diverged organisms. The results show that some MC subfamilies have all exons and transmembrane α-helices most similar to a closely related subfamily, which is consistent with a scenario of gene duplication and mutational divergence from a last common ancestor. However, several MC subfamilies appear to be mosaics of exons and transmembrane α-helices most similar to different and distant subfamilies, which in some cases could be explained by recombination between the superfamily genes during evolution. It seems that this latter mechanism could have played a role in the formation of new subfamilies with different substrate specificities by the combination of MC transporter domain segments that had been optimized previously for binding specific portions of the substrates. This study presents novel evolutionary relationships between MC subfamilies and may provide clues for how protein superfamilies have expanded and how to investigate their evolution.

Evolution, Molecular

The early injected genomic region determines sensitivity to Type I restriction-modification defence against Autographiviridae phages.

Bacteriophages must evade bacterial defences to establish successful infections. Type I restriction-modification (RM) systems recognize specific DNA motifs and degrade unmethylated foreign DNA, restricting phage replication. In this study, we detected that Marinomonas mediterranea MMB-2 uses a Type I RM system (Mme2I) to protect against several new phages in the Murciavirus genus within the Autographiviridae family. Whole-genome sequencing and methylation analysis revealed a DNA sequence motif methylated in M. mediterranea MMB-2, which is also present in the phages. Phages lacking the motif within the leading, first injected, region of their genomes, either natural isolates or escape mutants of sensitive phages, successfully infect M. mediterranea MMB-2, despite the presence of the recognition motif elsewhere in their genomes. These results highlight the importance of considering RM motif locations when predicting avoidance of restriction sites as escape mechanisms from RM systems. Additionally, our findings indicate an important role for RM systems in specifically influencing the organization of the leading injected regions of phage genomes, which are highly variable and often encode diverse anti-defence systems.

Genome, Viral

Flexible use of conserved motifs constrains genome access in cell type evolution.

Cell types can be organized into related families, but the regulatory mechanisms that define and maintain these families across deep evolutionary time remain unknown. Here, combining single-nucleus multi-omic sequencing with deep learning to analyse the accessible genomes of two groups of vastly divergent animals including flatworms and vertebrates, we find that hundreds of accessibility-dictating sequence motifs partition into distinct yet conserved sets, or 'vocabularies', each associated with a specific cell type family. However, combinatorial relationships among these motifs preferred by individual cell types are largely species specific. Deep-learning models trained on one species accurately predict family-level chromatin accessibility in distantly related species, albeit frequently rely on different motifs from shared vocabularies to reach convergent predictions. By contrast, models trained on individual cell types within a family lose cross-species predictive power, indicating that the regulatory syntax governing cell type-level identity evolves rapidly. We propose a 'collective maintenance' model in which motif vocabularies defining cell type families are evolutionarily stable, while recombination of these motifs generates cell type-specific regulatory programmes. This suggests that family identity is maintained collectively by large, conserved pools of regulatory factors, analogous to the logic of developmental homology, where character identity persists through network-level conservation despite extensive rewiring.

Journal Article

Genome-wide screening and functional validation of methylation barriers near promoters.

CpG islands near promoters are normally unmethylated despite being surrounded by densely methylated regions. Aberrant hypermethylation of these CpG islands has been associated with the development of various human diseases. Although local genetic elements have been speculated to play a role in protecting promoters from methylation, only a limited number of methylation barriers have been identified. In this study, we conducted an integrated computational and experimental investigation of colorectal cancer methylomes. Our study revealed 610 genes with disrupted methylation barriers. Genomic sequences of these barriers shared a common 41-bp sequence motif (MB-41) that displayed homology to the chicken HS4 methylation barrier. Using the CDKN2A (P16) tumor suppressor gene promoter, we validated the protective function of MB-41 and showed that loss of such protection led to aberrant hypermethylation. Our findings highlight a novel sequence signature of cis-acting methylation barriers in the human genome that safeguard promoters from silencing.

Animals

Diversity at the HYP1 locus in potato cyst nematodes does not result from developmentally-programmed somatic mutations.

Most genetic diversity stems from spontaneous mutations, that is, errors in DNA repair or replication. But for dozens of organisms across the tree of life, mutations at specific loci are not spontaneous but developmentally programmed: effectively, some organisms edit their own DNA sequences. This is perhaps most common among pathogens and parasites, many of which use editing to diversify genes that produce important antigens. Plant-parasitic potato cyst nematodes are damaging agricultural pests that establish a lifelong feeding site inside the root of their host plant. We previously observed extensive diversity of rare alleles at HYP1, the most highly expressed gene that encodes a protein secreted by potato cyst nematodes during parasitism. Importantly, HYP1 alleles differ from each other by complex, in-frame rearrangements of short repeated sequence motifs within a single exon. Combining several lines of evidence, we previously hypothesized that potato cyst nematodes use developmentally-programmed mutations, or editing, to diversify HYP1 alleles in the soma. In the current work, we now test this hypothesis. We employ highly accurate long-read DNA sequencing of a simplified genetic system to identify potential rare edited alleles, we use a transgenic yeast system to describe large de novo mutations at HYP1, and we interpret our findings in light of key population genetic parameters as well as the genetic diversity surrounding HYP1 and across the genome.

Animals

A token-pruning framework enables efficient representation of the human genome for RNA modification analysis.

MOTIVATION: Modelling long genomic sequences remains challenging due to extreme sequence length, high redundancy, and the need for biological interpretability. Although Transformer-based architectures have achieved strong performance across genomic tasks, their high computational cost and reliance on fixed tokenization strategies limit their scalability and ability to focus on biologically informative regions. RESULTS: We propose ATSFormer, a token-pruning Transformer framework for efficient and biologically informed genomic sequence modelling. ATSFormer incorporates an attention-guided and parameter-free Adaptive Token Sampling (ATS) module into Transformer layers. Guided by attention-derived importance scores, ATS dynamically retains informative tokens while probabilistically discarding redundant ones, thereby reducing sequence length, FLOPs, and memory usage without introducing additional learnable parameters or extra training procedures. Importantly, the retained tokens correspond to key contributors to model predictions, enabling ATSFormer to highlight biologically meaningful sites and sequence motifs. We evaluated ATSFormer on four benchmark RNA modification datasets derived from RMVar 2.0, covering A-to-I, m1A, m5C, and m7G. Experimental results show that ATSFormer consistently outperforms existing state-of-the-art methods while achieving substantial computational savings. Furthermore, structural analysis using AlphaFold3 supports the biological relevance of the motifs identified by ATSFormer. AVAILABILITY AND IMPLEMENTATION: The source data and code are freely available at GitHub (https://github.com/1gao2/ATSFormer) and Zenodo (https://doi.org/10.5281/zenodo.21813541).

Humans

Novel avian calicivirus genome with type IV internal ribosomal entry site (IRES) in black-headed gull (Chroicocephalus ridibundus) in Hungary.

In this study, a taxonomically novel avian calicivirus detected and characterized by next generation sequencing, RT-PCR and Sanger sequencing methods in faecal specimen collected from black-headed gull (Chroicocephalus ridibundus) in Hungary. The complete genome length of the calicivirus strain gull/HA15097/HUN/2018 (PZ810127) is remarkably long, 8,845 nucleotides, which had type IV internal ribosomal entry site (IRES) at the 5', and a stem-loop-II-like (s2m) sequence motif at the 3' untranslated regions. The VP1 capsid protein had less than 26% aa identity to the members of the known calicivirus genera. Caliciviruses appear to be widespread not only in mammals including humans but also in various bird species.

Animals

EvoSNR-Prom: Predicting promoters at single-nucleotide resolution with label-aware transfer learning of the pretrained EVO model.

The precise identification of promoters is crucial for understanding gene regulation. Deep learning methods have achieved considerable success in promoter prediction, yet most operate at the sequence level with coarse-grained labels. This means they label an entire DNA segment as either a "promoter" or "non-promoter," which results in a lack of the nucleotide-level resolution in prediction. In this study, we propose EvoSNR-Prom, a model designed for promoter prediction at single-nucleotide resolution. EvoSNR-Prom is built on the Evo foundation model and formulates promoter identification as a token-level sequence labeling problem, analogous to named entity recognition in natural language processing. To address the limited contextual information available in single-nucleotide tokenization, we introduce a lexicon-enhanced embedding strategy that incorporates biologically meaningful DNA lexicons, enriching contextual representations and improving the model's ability to capture complex sequence motifs. Furthermore, to enhance predictive performance on small size datasets, we integrate a label-aware transfer learning framework to leverage knowledge from well-annotated source species to a target organism. The results across various prokaryotic datasets show that EvoSNR-Prom achieves excellent performance. This work provides a valuable computational framework for the high-precision analysis of gene regulatory elements, contributing to the advancement of promoter prediction at single-nucleotide resolution.

Promoter Regions, Genetic

Human dopamine β-hydroxylase promoter variant alters transcription in chromaffin cells, enzyme secretion, and blood pressure.

BACKGROUND: Dopamine β-hydroxylase (DBH) plays an indispensable role in catecholamine synthesis by converting dopamine into norepinephrine. Here, we characterized a DBH promoter polymorphism (C-2073T; rs1989787; minor allele frequency ~16%) that influences not only gene transcription but also enzyme secretion and blood pressure (BP) in vivo. METHODS: Plasma DBH activity was measured spectrophotometrically. DBH genetic effects on BP were tested in subjects with the most extreme BP values in a large primary care population. Functional effects of promoter variants were studied by site-directed mutagenesis in DBH promoter haplotype/luciferase reporter plasmids transfected into chromaffin cells. Sequence motifs were predicted from position weight matrices, and endogenous transcription factor binding was probed by Chromatin ImmunoPrecipitation (ChIP). RESULTS: The T-allele of common promoter variant C-2073T was contained in a promoter haplotype that associated with plasma DBH activity, a trait also predicted by that variant itself. Promoter haplotypes including C-2073T predicted BP in the population, and the effect was also referable to C-2073T itself. Computationally, C-2073 disrupted a predicted match for transcription factor c-FOS. Site-directed mutagenesis at C-2073T altered not only basal promoter activity, but also transactivation by c-FOS, as well as the chromaffin cell secretory stimuli nicotine or pituitary adenylate cyclase-activating polypeptide (PACAP). Endogenous c-FOS bound to the motif in chromatin. CONCLUSIONS: These results suggest that DBH promoter variant C-2073T is functional in vivo: this promoter variant seems to initiate a cascade of transcriptional and biochemical changes including augmented DBH secretion, eventuating in elevation of basal BP, and hence cardiovascular risk. The observations suggest new strategies for probing the pathophysiology, risk, and treatment of hypertension.

Animals

Pervasive phosphorylation by phage T7 kinase disarms bacterial defences.

Bacteria and bacteriophages are in a constant arms race to develop defence and anti-defence systems, respectively. Currently known phage-encoded anti-defence systems are specific to the activity of the targeted bacterial defence system. Here we identify a mechanism by which the T7 bacteriophage broadly counteracts bacterial defences using protein phosphorylation. Its kinase (T7K), which has been reported to redirect the function of a few host proteins1-5, is actually a hyperpromiscuous dual-specificity kinase that phosphorylates nearly all host and phage proteins during infection. The scale of phosphorylation vastly exceeds known phosphosites in Escherichia coli, has no sequence motif specificity and results in a higher proteome-wide phosphorylation density than mammalian cells with around 500 kinases. Stoichiometry analysis of phosphorylation sites revealed strong bias in T7K activity towards nucleic-acid-binding substrates mediated by its C-terminal DNA-binding domain. This highly stoichiometric phosphorylation enables the deactivation of DNA-targeting or DNA-containing bacterial defence systems. We provide mechanistic insights into how T7K weakens DNA-containing Retron-Eco9 through specific phosphorylation events, with single phosphomimetic mutations in key sites of the toxin abolishing defence. Moreover, by screening a large collection of E. coli strains, we provide evidence of broad anti-defence abilities of T7K in nature, as counteracted strains contain diverse bacterial defence systems. T7K homologues are found almost exclusively in phages, with hyperpromiscuous kinase activity probably being enabled by a divergent DFG-like motif in the catalytic centre.

Journal Article

Abundant mRNA m1A modification in dinoflagellates: a new layer of gene regulation.

Dinoflagellates, a class of unicellular eukaryotic phytoplankton, exhibit minimal transcriptional regulation, representing a unique model for exploring gene expression. The biosynthesis, distribution, regulation, and function of mRNA N1-methyladenosine (m1A) remain controversial due to its limited presence in typical eukaryotic mRNA. This study provides a comprehensive map of m1A in dinoflagellate mRNA and shows that m1A, rather than N6-methyladenosine (m6A), is the most prevalent internal mRNA modification in various dinoflagellate species, with an asymmetric distribution along mature transcripts. In Amphidinium carterae, we identify 6549 m1A sites characterized by a non-tRNA T-loop-like sequence motif within the transcripts of 3196 genes, many of which are involved in regulating carbon and nitrogen metabolism. Enriched within 3'UTRs, dinoflagellate mRNA m1A levels negatively correlate with translation efficiency. Nitrogen depletion further decreases mRNA m1A levels. Our data suggest that distinctive patterns of m1A modification might influence the expression of metabolism-related genes through translational control.

Dinoflagellida