PubMed HealthSearch

SEARCH · PubMed Health

Results for “GC-rich sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

8 recordsLinked to original sources

Using the DNA language model, GROVER, to parse effects of sequence, chromatin and regulatory features on genome stability.

MOTIVATION: Genome stability is shaped by DNA sequence and chromatin context, but their relative contributions to double-strand break (DSB) sensitivity remain unclear. RESULTS: We show that the DNA language model, GROVER, can infer DSB location based on sequence. DSB hotspots tend to contain GC-rich sequences that belong to promoters, genes and short interspersed nuclear elements (SINEs). Additionally, we identified several specific short sequences (tokens) that are associated with modulating DSB sensitivity. Another model using chromatin and genome regulatory features outperforms the sequence-only model, highlighting complementary and cell-type specific information. Integrating sequence and genome biological features yields the best performance, demonstrating their synergy. Analyzing this model revealed that, dependent on the sample, genome stability information encoded in H3K36me3 and DNase-seq can be learned from the sequence, but not H3K27ac or H3K9me3. Embedding chromatin data directly into the GROVER architecture enabled cell-type specific modeling with performance matching the full chromatin feature model. Our results suggest that while chromatin and regulatory context provides important information, such as cell-type specificity, much of the information shaping DSB patterns is already encoded in the DNA sequence itself. Our integrative modeling approach not only reveals DSB patterns but also provides a generalizable strategy for tracing predictions in genomic data. AVAILABILITY: Data, models, and a tutorial are available on Zenodo.

Chromatin

ELYS associates with distinct DNA sequence environments during post-mitotic nuclear pore reassembly.

Nuclear pore complexes (NPCs) contribute to genome organization and cell identity, yet how post-mitotic NPC assembly is coordinated with chromatin architecture remains unclear. Here, we show that the nucleoporin ELYS preferentially associates with chromatin regions displaying distinct intrinsic DNA sequence features that are not explained by the repressive histone marks examined here. ELYS-bound regions are enriched for AT-rich sequences, whereas ELYS binding at super-enhancer-associated loci shift toward GC-rich sequence composition, revealing distinct sequence environments. These findings indicate that ELYS localization is associated with distinct intrinsic DNA sequence features and suggest a mechanism by which nuclear pore-associated architecture restores transcriptional programs after mitosis.

Journal Article

In vitro reconstitution of chromatin replication recapitulates symmetric histone recycling.

Symmetric histone recycling is vital for maintaining epigenetic inheritance upon eukaryotic DNA replication. Recent genome-wide studies have uncovered key determinants of this process, but how these factors collectively support parental histone transfer remains incompletely understood. Here, we successfully reconstitute histone recycling with 24 purified proteins and analyze the products digested by Micrococcal nuclease with Repli-pore-seq, the newly developed pipeline combining nanopore sequencing and deep-learning-based classification. As a result, we identify histones symmetrically recycled as tetrasomes or hexasomes on nucleosome-favorable sequences. We also observe the discordance of the recycled position between lagging and leading strands on the GC-rich DNA sequences. Moreover, removal of Pol δ, Pol32, Dpb3/4, Ctf4, Csm3/Tof1, or Mrc1 disrupts the balance of histone recycling between the two daughter strands, whereas removal of Ctf4, Csm3/Tof1, or Mrc1 additionally alters the positions at which histones were recycled. Furthermore, addition of the lagging-strand maturation factors Fen1 and Cdc9 enhances histone recycling to the lagging strand. These findings provide critical insights into the molecular players and mechanisms underlying symmetric histone recycling.

Histones

Genome-wide etiology analysis of autoimmune hypothyroidism supports somatic mutations of at-risk DNA as the underlying cause.

Autoimmune hypothyroidism (AIHT) is the most common autoimmune disease. Through an unidentified mechanism, the immune system attacks the thyroid gland, destroys thyroid follicular cells, and causes hypothyroidism. A new theory poses that all DNA is continuously damaged and, as a result, is exposed to somatic mutations at a constant rate. Based on this theory, several assumptions related to epidemiology and DNA sequence can be made. These have been summarized as a method called genome-wide etiology analysis (GWEA) to facilitate the interpretation of GWAS results of autoimmune diseases. Here, GWEA is applied to AIHT. The results show that existing epidemiological and genomic data of AIHT adhere to the principles of GWEA. Therefore, AIHT appears to be the result of somatic mutations in people at risk for the disease. AIHT develops once sufficient mutations create a new "autoimmune pathway" driven by non-self-signal and supported by neopeptide formation and signal amplification. Given the random nature of somatic mutations throughout life, the new theory explains why some people with AIHT develop additional autoimmune diseases, why family members may develop a range of non-AIHT autoimmune diseases, why the age of onset cannot be predicted, and why AIHT is transferred to the following generations through dominant inheritance with delayed, incomplete penetrance.

Humans

Analysis of targeted and whole genome sequencing of PacBio HiFi reads for a comprehensive genotyping of gene-proximal and phenotype-associated Variable Number Tandem Repeats.

Variable Number Tandem repeats (VNTRs) refer to repeating motifs of size greater than five bp. VNTRs are an important source of genetic variation, and have been associated with multiple Mendelian and complex phenotypes. However, the highly repetitive structures require reads to span the region for accurate genotyping. Pacific Biosciences HiFi sequencing spans large regions and is highly accurate but relatively expensive. Therefore, targeted sequencing approaches coupled with long-read sequencing have been proposed to improve efficiency and throughput. In this paper, we systematically explored the trade-off between targeted and whole genome HiFi sequencing for genotyping VNTRs. We curated a set of 10&#xa0;,&#xa0;787 gene-proximal (G-)VNTRs, and 48 phenotype-associated (P-)VNTRs of interest. Illumina reads only spanned 46% of the G-VNTRs and 71% of P-VNTRs, motivating the use of HiFi sequencing. We performed targeted sequencing with hybridization by designing custom probes for 9,999 VNTRs and sequenced 8 samples using HiFi and Illumina sequencing, followed by adVNTR genotyping. We compared these results against HiFi whole genome sequencing (WGS) data from 28 samples in the Human Pangenome Reference Consortium (HPRC). With the targeted approach only 4,091 (41%) G-VNTRs and only 4 (8%) of P-VNTRs were spanned with at least 15 reads. A smaller subset of 3,579 (36%) G-VNTRs had higher median coverage of at least 63 spanning reads. The spanning behavior was consistent across all 8 samples. Among 5,638 VNTRs with low-coverage (&#xa0;<&#xa0;15), 67% were located within GC-rich regions (&#xa0;>&#xa0;60%). In contrast, the 40X WGS HiFi dataset spanned 98% of all VNTRs and 49 (98%) of P-VNTRs with at least 15 spanning reads, albeit with lower coverage. Spanning reads were sufficient for accurate genotyping in both cases. Our findings demonstrate that targeted sequencing provides consistently high coverage for a small subset of low-GC VNTRs, but WGS is more effective for broad and sufficient sampling of a large number of VNTRs.

Minisatellite Repeats

Codon Composition in Human Oocytes Reveals Age-Associated Defects in mRNA Decay.

Oocytes from women of advanced reproductive age exhibit diminished developmental potential, but the underlying mechanisms remain incompletely defined. Oocyte maturation depends on translational control of maternal mRNA synthesized during growth. We performed a computational analysis on human oocytes from women <30 versus &#x2265;40 years and observed that mRNA GC content correlates negatively with half-life in oocytes from young (<30 yr) but positively with oocytes from aged (>40 yr) women. In young oocytes, longer mRNA half-life is associated with lower protein abundance, whereas in aged oocytes GC content correlates positively with protein abundance. During the GV-to-MII transition, codon composition stratifies stability: codons that support rapid translation (optimal) stabilize mRNA, while slow-translating codons (non-optimal) promote decay. With reproductive aging, GC-containing codons become more optimal and align with increased protein abundance. These findings indicate that reproductive aging remodels codon-optimality-linked, translation-coupled mRNA decay, stabilizing a subset of GC-rich maternal mRNA that may be prone to excess translation during maturation. Our analysis is explicitly within human reproductive aging; it does not revisit cross-species stability rules. Instead, it shows that sequence-stability relations are reprogrammed with age within human oocytes, including an inversion of the GC-stability association during GV-to-MII transition. Disruption of the normal mRNA clearance program in aged oocytes may compromise oocyte competence and alter maternal mRNA dosage, with downstream consequences for early embryonic development.

Humans

Improving spliced alignment by modeling splice sites with deep learning.

MOTIVATION: Spliced alignment refers to the alignment of messenger RNA (mRNA) or protein sequences to eukaryotic genomes. It plays a critical role in gene annotation and the study of gene functions. Accurate spliced alignment demands sophisticated modeling of splice sites, but current aligners use simple models, which may affect their accuracy given dissimilar sequences. RESULTS: We implemented minisplice to learn splice signals with a one-dimensional convolutional neural network (1D-CNN) and trained a model with 7,026 parameters for vertebrate and insect genomes. It captures conserved splice signals across phyla and reveals GC-rich introns specific to mammals and birds. We used this model to estimate the empirical splicing probability for every GT and AG in genomes, and modified minimap2 and miniprot to leverage pre-computed splicing probability during alignment. Evaluation on human long-read RNA-seq data and cross-species protein datasets showed our method greatly improves the junction accuracy especially for noisy long RNA-seq reads and proteins of distant homology. AVAILABILITY AND IMPLEMENTATION: https://github.com/lh3/minisplice.

Journal Article

TAp73beta and DNp73beta activate the expression of the pro-survival caspase-2S.

p73, the p53 homologue, exists as a transactivation-domain-proficient TAp73 or deficient deltaN(DN)p73 form. Expectedly, the oncogenic DNp73 that is capable of inactivating both TAp73 and p53 function, is over-expressed in cancers. However, the role of TAp73, which exhibits tumour-suppressive properties in gain or loss of function models, in human cancers where it is hyper-expressed is unclear. We demonstrate here that both TAp73 and DNp73 are able to specifically transactivate the expression of the anti-apoptotic member of the caspase family, caspase-2(S). Neither p53 nor TAp63 has this property, and only the p73beta form, but not the p73alpha form, has this competency. Caspase-2 promoter analysis revealed that a non-canonical, 18 bp GC-rich Sp-1-binding site-containing region is essential for p73beta-mediated activation. However, mutating the Sp-1-binding site or silencing Sp-1 expression did not affect p73beta's transactivation ability. In vitro DNA binding and in vivo chromatin immunoprecipitation assays indicated that p73beta is capable of directly binding to this region, and consistently, DNA binding p73 mutant was unable to transactivate caspase-2(S). Finally, DNp73beta over-expression in neuroblastoma cells led to resistance to cell death, and concomitantly to elevated levels of caspase-2(S.) Silencing p73 expression in these cells led to reduction of caspase-2(S) expression and increased cell death. Together, the data identifies caspase-2(S) as a novel transcriptional target common to both TAp73 and DNp73, and raises the possibility that TAp73 may be over-expressed in cancers to promote survival.

Binding Sites