PubMed HealthSearch

SEARCH · PubMed Health

Results for “Nucleotide Motifs”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Disruption of a six-nucleotide miRNA motif improves PKD1 dosage and ameliorates polycystic kidney disease.

Disrupting microRNA interactions to restore protein expression from haploinsufficient genes offers a promising precision-therapy strategy for monogenic disorders. PKD1 heterozygosity underlies autosomal dominant polycystic kidney disease (ADPKD), a disorder affecting nearly 12 million people worldwide, where reduced PKD1 dosage drives progressive cyst formation and kidney failure. We previously identified a 55-bp cis-repressive element in the PKD1 3'UTR. Here, we define a six-nucleotide miR-17 seed match within this element that is sufficient to reproduce PKD1 repression. In vivo base substitution of this motif stabilizes Pkd1 messenger RNA and increases polycystin-1 (PC1) protein levels, producing a robust reduction in cyst growth and preservation of kidney function in mouse models. To therapeutically recapitulate this effect, we developed a steric-blocking oligonucleotide that occludes the motif, stabilizes PKD1 transcript levels, increases PC1 expression, and mitigates cyst-pathogenic events in both murine and patient-derived ADPKD cells. Together, these findings establish a minimal, targetable cis-regulatory motif and provide proof of concept for oligonucleotide-mediated PKD1 derepression, while offering a potentially generalizable strategy to restore other haploinsufficient genes.

Animals

A new tool for engineering Phaeodactylum tricornutum: the METE promoter drives both high expression and B12-tuneable regulation of transgenes.

For advanced metabolic engineering strategies, it is crucial to be able to regulate transgene expression, to prevent potential deleterious effects in the host organism during growth and allow optimisation of production levels. Here, we identified vitamin B12 (cobalamin)-responsive promoters in the diatom Phaeodactylum tricornutum, a promising biotechnological chassis that readily absorbs this metabolite with minimal physiological impact. Using promoter-reporter constructs, the promoters of the cobalamin acquisition protein 1 (CBA1) and the B12-independent form of methionine synthase (METE) were shown to regulate transgene expression in a B12-dependent manner. Further characterisation of the METE promoter (PMETE) demonstrated that it exhibited significantly higher expression levels than several previously characterised promoters, but could be repressed by nanomolar amounts of B12, with a dynamic range >100-fold. Tight regulation was demonstrated by the suppression of the lethal ribonuclease, barnase at 1 μg L-1 B12. Reporter expression was doubled when PMETE was paired with its cognate terminator, compared with the widely used FCPA terminator. Promoter truncations resulted in decreased expression, but no loss of B12 regulation. A 14 nucleotide motif, present in four copies in PMETE, was found to be necessary for expression, and when fused to the constitutive FCPA promoter, enhanced expression levels. Transgenic lines expressing the heterologous diterpenoid enzyme, casbene synthase, produced casbene titres of approximately 2 mg L-1 and this was tuneable by B12. This demonstrates the utility of PMETE in efforts to establish P. tricornutum as an industrial biotechnology production platform.

Promoter Regions, Genetic

Structural Features of DNA in TATA-Containing and TATA-Less Core Promoters of RNA Polymerase II Differ.

Nucleotide motifs in the core promoters of eukaryotic protein-coding genes transcribed by RNA polymerase II (Pol II) play an important role in the transcription process. We analyzed the role of an octanucleotide located in the TATA box position. Depending on whether this octanucleotide can form a complex with the TATA-binding protein (TBP), the promoter is classified as either TATA-containing or TATA-less. We analyzed the differences in the primary and spatial structures, as well as their dynamics, in TATA-containing and TATA-less promoters of mammals and plants. We divided the complete promoter sets of six organisms (H. sapiens, M. musculus, C. familiaris, A. thaliana, Z. mays, and H. vulgare) from the EPDnew database into TATA-containing and TATA-less fractions. The sizes of the TATA-containing promoter fractions are significantly smaller than those of the TATA-less fractions in all studied organisms, except in A. thaliana, where the sizes of both fractions are approximately equal. We characterized promoter architecture using variation profiles of various base-pair step parameters, minor-groove width, and the conformational dynamics of native DNA. The architectures of TATA-containing and TATA-less promoters differ significantly. The possible mechanistic influence of DNA structural features on the formation of the pre-initiation complex (PIC) in both types of promoters is discussed.

Promoter Regions, Genetic

An expanded realm of anti-CRISPR-associated proteins and regulatory mechanisms.

Many bacteriophages encode anti-CRISPR (Acr) proteins that inhibit bacterial CRISPR-Cas immune systems. Rapid acr gene expression upon phage entry enables CRISPR-Cas neutralization but can impact phage fitness if unregulated. Therefore, Acr production is often controlled by distinct families of co-encoded anti-CRISPR-associated (Aca) proteins, which are usually helix-turn-helix (HTH) regulators that bind DNA within acr-aca operon promoters. Previously, we demonstrated that the Aca2 family additionally represses Acr production translationally by binding structured RNA motifs within the 5' untranslated region (UTR) of the acr-aca mRNA. Here, through systematic bioinformatic analyses, we provide evidence of structured RNA motifs in the 5' UTRs of operons encoding members of other Aca families and show that Aca1 also specifically binds its cognate RNA motif. Additionally, many Aca proteins are predicted to regulate not only their own but also adjacent operons with potential anti-defence genes. Indeed, we show that Aca14, newly identified in this study, represses two predicted anti-defence operons. Aca14 is a ribbon-helix-helix domain protein, revealing regulatory diversity beyond the canonical HTH Aca family members. Collectively, our findings expand our understanding of acr regulation in mobile genetic elements and reveal novel mechanisms by which phages fine-tune anti-defence gene expression.

5' Untranslated Regions

Machine learning analysis of the human initiator region reveals key features of different types of core promoters.

The initiator (Inr) is the starting point for the transcription of many genes. Here, we generated highly predictive machine learning models of the human Inr region, and determined that the Inr is present in ∼60% of focused human promoters, identified a novel TATA-specific Inr, and detected the overlapping but functionally distinct TCT motif. Quantitative genome-wide analyses revealed a strict and synergistic interaction between the Inr and DPR, an inverse relationship between the TATA and DPR, a flexible and sometimes independent function of the TATA box in relation to the Inr, and different properties of the TCT motif in humans versus Drosophila.

Humans

Motif-Cluster: Motif driven prioritization of transcription factor binding clusters.

Genome-wide analyses of transcription factor (TF) motif binding sites have largely emphasized individual high-affinity sites, while overlooking the regulatory importance of locally repetitive motif clusters. Such clusters, including combinations of weak and strong binding sites, can collectively enhance TF occupancy and regulatory activity. Here we present Motif-Cluster, an open-source framework for motif-driven prioritization and visualization of TF binding clusters using sequence information alone. Motif-Cluster integrates a density-based clustering strategy with flexible modeling of binding-site gaps and affinity signals, enabling the identification and ranking of candidate regulatory regions without requiring experimental binding data. Through simulations and multiple real-data analyses, we show that combining gap distributions with binding affinity effectively balances cluster size and signal strength while reducing noise from weak sites. Application to ZNF410 successfully recovers the previously characterized binding clusters in the CHD4 promoter, which are conserved between human and mouse. Additional case studies involving PHB1, TWIST1, and EGR1 further demonstrate the general applicability of the method across diverse transcription factors. Motif-Cluster also provides intuitive visualization and reproducible workflows to facilitate interpretation of spatially dense motif patterns. Overall, Motif-Cluster offers a robust and flexible approach for prioritizing transcription factor regulatory regions from genome-wide motif scans, enabling biological discovery and guiding experimental design, particularly in settings where direct genome-wide binding assays are unavailable.

Transcription Factors

Heat-responsive ONSEN long terminal repeats integrate heat shock factor motifs, DNA methylation and natural sequence variation in Arabidopsis.

ONSEN is a heat-activated Ty1/copia retrotransposon in Arabidopsis thaliana controlled by heat shock factors (HSFs) and epigenetic silencing. Heat shock element (HSE)-like sequences in ONSEN long terminal repeats (LTRs) contribute to heat responsiveness, but relationships among sequence architecture, basal DNA methylation and natural variation remain unclear. We combined transcription-factor motif prediction, transposable-element comparisons, methylome and RNA sequencing (RNA-seq) data, and Arabidopsis genome assemblies. In silico disruption of five HSE cores eliminated HSF-family motif compatibility in the selected design and all 5119 exact-guanine-cytosine (GC) alternatives. Across 16 curated Columbia-0 terminal windows, ONSEN contained 33-49 non-redundant HSF motif-coordinate placements per 800 bp window and was strongly enriched relative to 1930 non-ONSEN transposable elements across score thresholds and continuous metrics. Direct comparison with 779 non-ONSEN LTR retrotransposons showed selectively elevated basal CHH methylation (where H = A, C or T) at ONSEN termini. Genome-wide RNA-seq analysis revealed broad heat-responsive gene and transposable-element changes, including strong ONSEN induction, whereas candidate-window analysis distinguished ONSEN from most HSF-rich non-ONSEN outliers. ONSEN-like variants across eight accessions generally retained HSF-compatible motifs while altering predicted DNA binding with one finger-family motif composition. Together, these findings define ONSEN terminal regions as HSF-rich regulatory sequences that retain heat-responsive potential within a methylated chromatin context and identify candidates for functional analysis.

DNA Methylation

Motif-centered analyses reveal universal and tissue-specific mutagenic mechanisms operating in the human body.

Somatic mutations are inevitable in human genomes and can lead to cancer initiation and tumor progression. Although many mutagenic processes have been linked to cancer, their activities in normal tissues before malignant transformation remain poorly characterized. Here, we analyzed the mutation profiles of 10,625 normal samples across 25 tissues obtained from whole-genome and whole-exome sequencing datasets. We applied stringent statistical hypothesis for detecting enrichment and enrichment-adjusted Minimal Estimate of Mutation Load in trinucleotide motifs preferred by known mutagenic processes. We found several cancer-associated mutational motifs in cancer-free tissues. Samples enriched with C→T mutations in nCg motif associated with clock-like spontaneous meCpG deamination were detected across all tissues. We also identified a second clock-like motif, T→C substitutions in aTn motif associated with exposure to small epoxides and other SN2 electrophiles, in several tissues. Motifs associated with other environmental and chemical mutagens showed sporadic and tissue-specific mutagenesis. APOBEC-induced C→T and C→G mutations in tCw motif were enriched in bladder, lung, small intestine, liver, and breast with preference for APOBEC3A-like mutagenesis in most tissues. Together, our analyses elucidated several cancer-associated mutagenic processes in normal tissues and provided a robust analytical framework for quantifying mutagenic activities from somatic mutation catalogs.

Humans

Chromosome-specific centromeric patterns define the centeny map of the human genome.

Centromeres are epigenetically specified by distinct chromatin, whereas their DNA varies between species and individuals. This extensive sequence divergence makes comparative analyses between centromeres challenging. In this study, we identified a chromosome-specific architectural pattern across the human genome, defined by the conserved spacing of a functionally relevant centromeric DNA motif. The distribution of these sites along chromosome arms constitutes the human "centeny map." By using a custom Genomic Centromere Profiling (GCP) pipeline, we leveraged the motif's position, orientation, and organization to construct structural models that enable reclassification of human chromosomal clusters, detection of centromere expansion, and identification of structural variants and misassembled regions. The high-resolution maps derived from this pattern not only provide a framework for comparative analysis of centromeres across evolution and disease but also offer a new dimension for chromosome annotation, assembly, and characterization.

Humans

The early injected genomic region determines sensitivity to Type I restriction-modification defence against Autographiviridae phages.

Bacteriophages must evade bacterial defences to establish successful infections. Type I restriction-modification (RM) systems recognize specific DNA motifs and degrade unmethylated foreign DNA, restricting phage replication. In this study, we detected that Marinomonas mediterranea MMB-2 uses a Type I RM system (Mme2I) to protect against several new phages in the Murciavirus genus within the Autographiviridae family. Whole-genome sequencing and methylation analysis revealed a DNA sequence motif methylated in M. mediterranea MMB-2, which is also present in the phages. Phages lacking the motif within the leading, first injected, region of their genomes, either natural isolates or escape mutants of sensitive phages, successfully infect M. mediterranea MMB-2, despite the presence of the recognition motif elsewhere in their genomes. These results highlight the importance of considering RM motif locations when predicting avoidance of restriction sites as escape mechanisms from RM systems. Additionally, our findings indicate an important role for RM systems in specifically influencing the organization of the leading injected regions of phage genomes, which are highly variable and often encode diverse anti-defence systems.

Genome, Viral

Genome-wide screening and functional validation of methylation barriers near promoters.

CpG islands near promoters are normally unmethylated despite being surrounded by densely methylated regions. Aberrant hypermethylation of these CpG islands has been associated with the development of various human diseases. Although local genetic elements have been speculated to play a role in protecting promoters from methylation, only a limited number of methylation barriers have been identified. In this study, we conducted an integrated computational and experimental investigation of colorectal cancer methylomes. Our study revealed 610 genes with disrupted methylation barriers. Genomic sequences of these barriers shared a common 41-bp sequence motif (MB-41) that displayed homology to the chicken HS4 methylation barrier. Using the CDKN2A (P16) tumor suppressor gene promoter, we validated the protective function of MB-41 and showed that loss of such protection led to aberrant hypermethylation. Our findings highlight a novel sequence signature of cis-acting methylation barriers in the human genome that safeguard promoters from silencing.

Animals

SLE-Associated rs2295613(A) Allele Strengthens a Predicted c-MYC Motif and Enhances SLAMF1 Promoter Reporter Activity in B Cells.

SLAMF1 encodes CD150, an immunoregulatory receptor involved in lymphocyte activation, T-B-cell interactions, and humoral immune responses. The SLAMF1 promoter polymorphism rs2295613(G>A) was previously associated with systemic lupus erythematosus (SLE) susceptibility in a Chinese case-control cohort. Here, we investigated the regulatory activity of rs2295613 in the transformed B-cell lines Raji and MP1 and in primary human CD19+ B cells. The rs2295613(A)-containing reporter showed higher promoter activity than the rs2295613(G)-containing reporter in all three cellular systems. Bioinformatic analysis predicted that the G to A substitution strengthens a pre-existing MYC-compatible motif. Substitutions disrupting the motif-containing region attenuated the rs2295613(A)-associated increase in reporter activity and reduced enrichment of the promoter fragment in anti-c-MYC DNA pull-down assays. Partial siRNA-mediated reduction in MYC mRNA also decreased the activity of the rs2295613(A)-containing reporter in Raji cells. Together, these findings identify rs2295613 as a functional SLAMF1 promoter variant in B-cell reporter systems and support a contribution of c-MYC-associated regulation to the enhanced activity of the rs2295613(A)-containing promoter.

Humans

An expanded codebook of human transcription factor DNA-binding specificity.

Gene expression is regulated by transcription factors (TFs), which recognize specific DNA sequence motifs. Several hundred putative human TFs, identified mainly by an apparent DNA-binding domain, lack known binding motifs1. Furthermore, even for well-characterized TFs, it remains controversial the degree to which motifs accurately reflect binding sites in living cells2. Here we describe a systematic effort ('Codebook') to determine the sequence specificity of 332 putative and poorly characterized human TFs. More than 4,000 independent experiments, encompassing multiple in vitro and in vivo assays, produced motifs for just over half (177; 53%) of the TFs, of which most are associated with only a single protein. These results extend the vocabulary of sequence recognition encoded by human TFs by around 130 distinct motifs. Moreover, binding motifs identified in vitro are strongly enriched in cellular binding sites. Collectively, the data reveal tens of thousands of previously unknown, conserved and direct TF-binding sites across the human genome. These sites are concentrated in promoter regions and are predictive of gene expression. In summary, this new codebook provides an important step forward in decoding the human genome.

Humans

Genome-wide computational analysis reveals cardiomyocyte-specific transcriptional Cis-regulatory motifs that enable efficient cardiac gene therapy.

Gene therapy is a promising emerging therapeutic modality for the treatment of cardiovascular diseases and hereditary diseases that afflict the heart. Hence, there is a need to develop robust cardiac-specific expression modules that allow for stable expression of the gene of interest in cardiomyocytes. We therefore explored a new approach based on a genome-wide bioinformatics strategy that revealed novel cardiac-specific cis-acting regulatory modules (CS-CRMs). These transcriptional modules contained evolutionary-conserved clusters of putative transcription factor binding sites that correspond to a "molecular signature" associated with robust gene expression in the heart. We then validated these CS-CRMs in vivo using an adeno-associated viral vector serotype 9 that drives a reporter gene from a quintessential cardiac-specific α-myosin heavy chain promoter. Most de novo designed CS-CRMs resulted in a >10-fold increase in cardiac gene expression. The most robust CRMs enhanced cardiac-specific transcription 70- to 100-fold. Expression was sustained and restricted to cardiomyocytes. We then combined the most potent CS-CRM4 with a synthetic heart and muscle-specific promoter (SPc5-12) and obtained a significant 20-fold increase in cardiac gene expression compared to the cytomegalovirus promoter. This study underscores the potential of rational vector design to improve the robustness of cardiac gene therapy.

Animals

Discovering human transcription factor physical interactions with genetic variants, novel DNA motifs, and repetitive elements using enhanced yeast one-hybrid assays.

Identifying transcription factor (TF) binding to noncoding variants, uncharacterized DNA motifs, and repetitive genomic elements has been technically and computationally challenging. Current experimental methods, such as chromatin immunoprecipitation, generally test one TF at a time, and computational motif algorithms often lead to false-positive and -negative predictions. To address these limitations, we developed an experimental approach based on enhanced yeast one-hybrid assays. The first variation of this approach interrogates the binding of >1000 human TFs to repetitive DNA elements, while the second evaluates TF binding to single nucleotide variants, short insertions and deletions (indels), and novel DNA motifs. Using this approach, we detected the binding of 75 TFs, including several nuclear hormone receptors and ETS factors, to the highly repetitive Alu elements. Further, we identified cancer-associated changes in TF binding, including gain of interactions involving ETS TFs and loss of interactions involving KLF TFs to different mutations in the TERT promoter, and gain of a MYB interaction with an 18-bp indel in the TAL1 superenhancer. Additionally, we identified TFs that bind to three uncharacterized DNA motifs identified in DNase footprinting assays. We anticipate that these enhanced yeast one-hybrid approaches will expand our capabilities to study genetic variation and undercharacterized genomic regions.

Algorithms

Molecular Characterization of the ClpC AAA+ ATPase in the Biology of Chlamydia trachomatis.

Bacterial AAA+ unfoldases are crucial for bacterial physiology by recognizing specific substrates and, typically, unfolding them for degradation by a proteolytic component. The caseinolytic protease (Clp) system is one example where a hexameric unfoldase (e.g., ClpC) interacts with the tetradecameric proteolytic core ClpP. Unfoldases can have both ClpP-dependent and ClpP-independent roles in protein homeostasis, development, virulence, and cell differentiation. ClpC is an unfoldase predominantly found in Gram-positive bacteria and mycobacteria. Intriguingly, the obligate intracellular Gram-negative pathogen Chlamydia, an organism with a highly reduced genome, also encodes a ClpC ortholog, implying an important function for ClpC in chlamydial physiology. Here, we used a combination of in vitro and cell culture approaches to gain insight into the function of chlamydial ClpC. ClpC exhibits intrinsic ATPase and chaperone activities, with a primary role for the Walker B motif in the first nucleotide binding domain (NBD1). Furthermore, ClpC binds ClpP1P2 complexes via ClpP2 to form the functional protease ClpCP2P1 in vitro, which degraded arginine-phosphorylated β-casein. Cell culture experiments confirmed that higher order complexes of ClpC are present in chlamydial cells. Importantly, these data further revealed severe negative effects of both overexpression and depletion of ClpC in Chlamydia as revealed by a significant reduction in chlamydial growth. Here, again, NBD1 was critical for ClpC function. Hence, we provide the first mechanistic insight into the molecular and cellular function of chlamydial ClpC, which supports its essentiality in Chlamydia. ClpC is, therefore, a potential novel target for the development of antichlamydial agents. IMPORTANCE Chlamydia trachomatis is an obligate intracellular pathogen and the world's leading cause of preventable infectious blindness and bacterial sexually transmitted infections. Due to the high prevalence of chlamydial infections along with negative effects of current broad-spectrum treatment strategies, new antichlamydial agents with novel targets are desperately needed. In this context, bacterial Clp proteases have emerged as promising new antibiotic targets, since they often play central roles in bacterial physiology and, for some bacterial species, are even essential for survival. Here, we report on the chlamydial AAA+ unfoldase ClpC, its functional reconstitution and characterization, individually and as part of the ClpCP2P1 protease, and establish an essential role for ClpC in chlamydial growth and intracellular development, thereby identifying ClpC as a potential target for antichlamydial compounds.

Humans

Interpreting the CTCF-mediated sequence grammar of genome folding with AkitaV2.

Interphase mammalian genomes are folded in 3D with complex locus-specific patterns that impact gene regulation. CTCF (CCCTC-binding factor) is a key architectural protein that binds specific DNA sites, halts cohesin-mediated loop extrusion, and enables long-range chromatin interactions. There are hundreds of thousands of annotated CTCF-binding sites in mammalian genomes; disruptions of some result in distinct phenotypes, while others have no visible effect. Despite their importance, the determinants of which CTCF sites are necessary for genome folding and gene regulation remain unclear. Here, we update and utilize Akita, a convolutional neural network model, to extract the sequence preferences and grammar of CTCF contributing to genome folding. Our analyses of individual CTCF sites reveal four predictions: (i) only a small fraction of genomic sites are impactful; (ii) impact is highly dependent on sequences flanking the core CTCF binding motif; (iii) core and flanking nucleotides contribute largely additively to the overall impact of a site; (iv) sites created as combinations of different core and flanking sequences have impacts proportional to the product of their average impacts, i.e. they are broadly compatible. Our analysis of collections of CTCF sites make two predictions for multi-motif grammar: (i) insulation strength depends on the number of CTCF sites within a cluster, and (ii) pattern formation is governed by the orientation and spacing of these sites, rather than any inherent specialization of the CTCF motifs themselves. In sum, we present a framework for using neural network models to probe the sequences instructing genome folding and provide a number of predictions to guide future experimental inquiries.

CCCTC-Binding Factor

Di-, tri-, and tetranucleotide frequencies covary with lifespan and genome size across protostome invertebrates.

Animal lifespans span orders of magnitude, yet how genome sequence covaries with lifespan remains poorly characterized outside vertebrates. Although promoter CpG density has been linked to vertebrate longevity due to its gene-regulatory function through DNA methylation, it is unclear whether such patterns are promoter- and CpG-specific, or if they reflect broader sequence evolution. We curated maximum lifespan estimates for 466 protostome species spanning eight phyla with available genome assemblies and quantified mono-, di-, tri-, and tetranucleotide composition across whole genomes, intergenic regions, and six gene-associated regions (two upstream regions, exons, introns, and two downstream regions) defined using Benchmarking Universal Single-Copy Orthologs. Dinucleotide observed/expected ratios showed significant associations with lifespan and genome size in different ways. Lifespan-associated motifs were most pronounced in gene-associated non-coding regions, especially in introns and downstream regions, whereas genome-size effects were strongest in whole-genome and intergenic sequence. Tri- and tetranucleotide observed/expected ratios broadly recapitulated this regional organization. In contrast, GC content was not associated with lifespan across regions, indicating that the observed signals are not explained by mononucleotide composition but instead by how those nucleotides are arranged into short sequence motifs. These results suggest that lifespan and genome size show distinct but overlapping associations with regional sequence composition across invertebrate species and that lifespan-associated motif evolution extends beyond vertebrate promoter methylation architectures.

CpG density