PubMed HealthSearch

SEARCH · PubMed Health

Results for “repetitive genomic elements”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Discovering human transcription factor physical interactions with genetic variants, novel DNA motifs, and repetitive elements using enhanced yeast one-hybrid assays.

Identifying transcription factor (TF) binding to noncoding variants, uncharacterized DNA motifs, and repetitive genomic elements has been technically and computationally challenging. Current experimental methods, such as chromatin immunoprecipitation, generally test one TF at a time, and computational motif algorithms often lead to false-positive and -negative predictions. To address these limitations, we developed an experimental approach based on enhanced yeast one-hybrid assays. The first variation of this approach interrogates the binding of >1000 human TFs to repetitive DNA elements, while the second evaluates TF binding to single nucleotide variants, short insertions and deletions (indels), and novel DNA motifs. Using this approach, we detected the binding of 75 TFs, including several nuclear hormone receptors and ETS factors, to the highly repetitive Alu elements. Further, we identified cancer-associated changes in TF binding, including gain of interactions involving ETS TFs and loss of interactions involving KLF TFs to different mutations in the TERT promoter, and gain of a MYB interaction with an 18-bp indel in the TAL1 superenhancer. Additionally, we identified TFs that bind to three uncharacterized DNA motifs identified in DNase footprinting assays. We anticipate that these enhanced yeast one-hybrid approaches will expand our capabilities to study genetic variation and undercharacterized genomic regions.

Algorithms

Neotelomeres and telomere-spanning chromosomal arm fusions in cancer genomes revealed by long-read sequencing.

Alterations in the structure and location of telomeres are pivotal in cancer genome evolution. Here, we applied both long-read and short-read genome sequencing to assess telomere repeat-containing structures in cancers and cancer cell lines. Using long-read genome sequences that span telomeric repeats, we defined four types of telomere repeat variations in cancer cells: neotelomeres where telomere addition heals chromosome breaks, chromosomal arm fusions spanning telomere repeats, fusions of neotelomeres, and peri-centromeric fusions with adjoined telomere and centromere repeats. These results provide a framework for the systematic study of telomeric repeats in cancer genomes, which could serve as a model for understanding the somatic evolution of other repetitive genomic elements.

Humans

Neotelomeres and Telomere-Spanning Chromosomal Arm Fusions in Cancer Genomes Revealed by Long-Read Sequencing.

Alterations in the structure and location of telomeres are key events in cancer genome evolution. However, previous genomic approaches, unable to span long telomeric repeat arrays, could not characterize the nature of these alterations. Here, we applied both long-read and short-read genome sequencing to assess telomere repeat-containing structures in cancers and cancer cell lines. Using long-read genome sequences that span telomeric repeat arrays, we defined four types of telomere repeat variations in cancer cells: neotelomeres where telomere addition heals chromosome breaks, chromosomal arm fusions spanning telomere repeats, fusions of neotelomeres, and peri-centromeric fusions with adjoined telomere and centromere repeats. Analysis of lung adenocarcinoma genome sequences identified somatic neotelomere and telomere-spanning fusion alterations. These results provide a framework for systematic study of telomeric repeat arrays in cancer genomes, that could serve as a model for understanding the somatic evolution of other repetitive genomic elements.

Telomere

Reference-Guided Chromosome-Scale Genome Assembly With Insights on Population Genomics of the Atlantic Goliath Grouper (Epinephelus itajara), Islas del Rosario, Colombia.

Epinephelus itajara, commonly known as the Atlantic Goliath grouper, is the largest species among the western North Atlantic groupers and is critically endangered. This species plays a crucial ecological, cultural, and economic role and has been the focus of captive breeding efforts at the Oceanario of the Rosario Islands, Colombia. However, despite its ecological and conservation importance, genomic resources and population genomic data for E. itajara remain scarce, particularly in the Colombian Caribbean. This study presents a reference-guided chromosome-scale genome assembly and an analysis of the population genomic structure of E. itajara using PacBio HiFi sequencing and Illumina technologies. The assembled genome has a total size of 1.12 Gb, with a contig N50 of 42.69 Mb and a scaffold N50 of 46.30 Mb. A total of 22,692 protein-coding genes were identified after masking 46% of the genome, which consists of repetitive elements. Comparative genomic analyses revealed a high degree of collinearity with closely related Epinephelus species and identified E. lanceolatus as the closest relative, supporting recent divergence and conserved genome architecture within the genus. Additionally, a population genomics analysis was conducted using 7706 high-quality SNPs to assess the genomic structure of captive populations. The results revealed four distinct genomic lineages, with moderate genetic differentiation among the sampled individuals. In the Colombian Caribbean, two unique lineages were identified, associated with the localities of Bahía Cispatá and Bahía Barbacoas, suggesting possible geographic isolation. These genomic resources provide valuable tools and new opportunities to better understand the genomic diversity, evolutionary history, and reproductive mechanisms of E. itajara. Moreover, they serve as a foundation for conservation strategies, including selective breeding programs aimed at increasing genomic diversity in captive populations and guiding restoration efforts in its natural habitat.

Epinephelus itajara

Chromosome-level Genome Assembly of the Halophytic Turfgrass Zoysia macrostachya.

Zoysia macrostachya Franch. & Sav. is a halophytic perennial turfgrass in the Poaceae family, commonly found in the coastal regions of Korea, Japan, and East Asia. Z. macrostachya thrives in high-salinity environments, making it an excellent model for studying abiotic stress resilience. In this study, we present a chromosome-level genome assembly of Z. macrostachya, constructed using Oxford Nanopore long reads, Illumina short reads, and Omni-C sequencing data. The assembly spans 329.78 Mb across 20 chromosomes, with a scaffold N50 of 19.24 Mb, and includes complete telomeric sequences at both ends. The assembly showed 97.8% complete BUSCOs, indicating high genome completeness. Repeat element and gene annotation identified 44.03% of the genome as repetitive elements and 33,474 protein-coding genes. The gene annotation showed 97.1% complete BUSCOs and 86.92% functionally characterized genes. Macrosynteny analysis highlighted highly collinear relationships with related species, providing a foundational understanding of the Z. macrostachya genomic structure. This high-quality genome serves as a valuable resource for advancing salinity tolerance research and improving the genetic diversity of Zoysia species.

Genome, Plant

nf-core/pacsomatic: a scalable somatic analytic pipeline using PacBio HiFi data.

MOTIVATION: Pacific Biosciences (PacBio) HiFi long-read sequencing enables robust characterization of complex genomic regions, repetitive elements, and structural variants (SVs) that are often inaccessible to short-read technologies. To fully leverage HiFi reads to advance cancer genomics and epigenetics, researchers require an end-to-end, scalable and optimized bioinformatics workflow. The nf-core framework meets this need by providing rigorously tested, community-curated pipelines that ensure reproducibility, transparency, and broad compatibility across computational environments. RESULTS: We present nf-core/pacsomatic, an automated Nextflow DSL2 pipeline designed for comprehensive paired tumor-normal somatic analysis using PacBio HiFi data. The workflow includes steps for read alignments against reference genome, somatic SNV/indel, SV, and CNV calling, CpG methylation profiling and differential methylation region (DMR) detection. Additional downstream modules support functional annotation, mutational signature analysis, tumor purity and ploidy estimation, and homologous recombination deficiency (HRD) assessment. Utilizing nf-core's modular design and containerized execution, nf-core/pacsomatic provides a stable framework for the reproducible discovery of biological insights. AVAILABILITY: nf-core/pacsomatic is available under the MIT License at nf-core (https://nf-co.re/pacsomatic) and github (https://github.com/nf-core/pacsomatic).

Software

Diversity of ribosomes at the level of rRNA variation associated with human health and disease.

With hundreds of copies of rDNA, it is unknown whether they possess sequence variations that form different types of ribosomes. Here, we developed an algorithm for long-read variant calling, termed RGA, which revealed that variations in human rDNA loci are predominantly insertion-deletion (indel) variants. We developed full-length rRNA sequencing (RIBO-RT) and in situ sequencing (SWITCH-seq), which showed that translating ribosomes possess variation in rRNA. Over 1,000 variants are lowly expressed. However, tens of variants are abundant and form distinct rRNA subtypes with different structures near indels as revealed by long-read rRNA structure probing coupled to dimethyl sulfate sequencing. rRNA subtypes show differential expression in endoderm/ectoderm-derived tissues, and in cancer, low-abundance rRNA variants can become highly expressed. Together, this study identifies the diversity of ribosomes at the level of rRNA variants, their chromosomal location, and unique structure as well as the association of ribosome variation with tissue-specific biology and cancer.

Humans

The dark genome in cardiovascular medicine.

Only ∼1%-2% of the human genome directly codes for proteins. The remainder consists of non-coding DNA, often referred to as the 'dark genome'. This includes regulatory elements, transposable and repetitive sequences, structural genomic features, pseudogenes, intronic and intergenic regions, and non-coding RNA (ncRNA) genes. These components are increasingly recognized as major regulators of gene expression, cell identity, and disease susceptibility. Currently, dark genome elements, particularly ncRNAs are increasingly recognized as important regulators of cardiovascular health and disease. Advances in genome analysis technologies have greatly improved our understanding of these non-coding regions and revealed clearer connections between the dark genome and cardiovascular traits. This review highlights major parts of the dark genome involved in cardiovascular disease, with emphasis on those for which mechanistic understanding and translational relevance are beginning to emerge. As mechanistic insight into individual and collective components of the dark genome advances, it increasingly enables the development of new opportunities for targeted therapeutics for cardiovascular prevention and disease management.

Humans

A chromosome-level assembly of the alpine snow alga Chloromonas typhlos.

Chloromonas typhlos is a cosmopolitan alpine snow alga distributed across continents, and its blooming accelerates snow melting by decreasing the amount of snow albedo. To elucidate the genetic traits underlying the adaptation of C. typhlos to the alpine habitat, we combined PacBio sequencing and Hi-C to generate a high-quality chromosome-level genome assembly (contig N50: 1.29 Mb; scaffold N50: 7.23 Mb) with 31 chromosomes and a genome size of 200.86 Mb. Repetitive elements constituted 11.05% of the genome, and 16,133 protein-coding genes were predicted, of which 82% were functionally annotated. This study provides a set of omics resources both for snow algae and the genus Chloromonas.

Snow

The clustered and scrambled arrangement of moderately repetitive elements in Drosophila DNA.

An examination of cloned Drosophila DNA has revealed large clusters of densely spaced, short (less than or equal to 1 kb), moderately repetitive elements. Different clusters have many of the same repetitive elements, but these elements are arranged differently in each cluster. It is improbable that this clustered arrangement can be detected by conventional reassociation kinetic and electron microscopic techniques, but it can be detected and features of its fine structure can be determined by a two-dimensional version of Southern's blotting technique. The genomic organization of these clustered repetitive elements was investigated by hybridizing restriction fragments of cloned DNA to polytene chromosomes, to filter-bound recombinant DNA clones and to Southern blots of total Drosophila DNA. These studies demonstrated that clusters occur in euchromatic regions of the chromosomes and that at least one of the clusters has the same repetitive element organization in cloned and in chromosomal DNA. These studies also demonstrated that copies of the elements from one cluster are scattered in at least 1000 chromosomal regions. These regions appear to have differing concentrations of repetitive DNA, but together they account for a large fraction of Drosophila's moderately repetitive DNA. Aside from indicating the genomic organization of cluster elements, this work has identified cluster elements throughout a 9 kb region neighboring one of the heat shock genes, throughout the intron of the major rDNA repeat and within the apparently transposable element, 412.

Animals

ECHO: a nanopore sequencing-based workflow for (epi)genetic profiling of the human repeatome.

SUMMARY: The human genome is dominated by repetitive DNA, whose genetic and epigenetic variation plays a key role in gene regulation, genome stability, and disease. Recent advances in long-read sequencing now enable large-scale, haplotype-resolved, and DNA methylation-informative analysis of the human genome, including on previously inaccessible complex and repetitive regions. However, the comprehensive, simultaneous characterisation of the "human repeatome" remains challenging, largely due to the lack of comprehensive tools integrated in a single pipeline that can capture the full spectrum of variation across diverse types of DNA repeats. Here, we present ECHO, a user-friendly, Snakemake-based pipeline for the "(Epi)genomic Characterisation of Human Repetitive Elements using Oxford Nanopore Sequencing." ECHO provides a reproducible and scalable framework for end-to-end analysis of whole-genome nanopore sequencing data, enabling integrative but also tailored (epi)genetic analyses of the human repeatome. AVAILABILITY AND IMPLEMENTATION: ECHO is freely available at Github: https://github.com/leenput/ECHO-pipeline, with the archived version at Zenodo: https://zenodo.org/records/19068468.

Humans

Genome analysis of Amphioxus and speculation as to the origin of contrasting vertebrate genome organization patterns.

1. The genome of Amphioxus was investigated by DNA reassociation techniques for the amount of repetitive and non-repetitive sequences and its pattern of organization. 2. A comparison of the amount of non-repetitive DNA between Amphioxus and the tunicate Ciona intestinalis does not support the hypothesis that the Cephalochordates have arisen from the Tunicates by polyploidy. 3. In the Amphioxus genome repetitive and non-repetitive elements are predominantly arranged in a short period interspersion pattern. Conclusions are presented as to the evolution of contrasting genome organization patterns among vertebrates.

Animals

Chromosome-level genome assembly of Elaeocarpus petiolatus (Elaeocarpaceae).

Elaeocarpus petiolatus is an ecologically and economically important species in tropical and subtropical forests. Despite its significance, the lack of genomic resources has hindered research on the genetic diversity and adaptive traits of E. petiolatus. To address this gap, we present a comprehensive chromosome-level genome assembly of E. petiolatus generated using advanced PacBio high-fidelity (HiFi) long-read sequencing and Hi-C technology. The assembly spans 322.45 Mb, with a scaffold N50 of 20.58 Mb, indicating that 37.11% of the genome is composed of repetitive elements. We identified 25,295 protein-coding genes, of which 96.74% were functionally annotated. This high-quality genome provides a critical resource for understanding the genetic mechanisms underlying environmental adaptability and biosynthesis of bioactive compounds in E. petiolatus, thereby supporting conservation efforts and sustainable forest management. The assembled genome and associated sequencing data are publicly available, facilitating further evolutionary and functional studies on the Elaeocarpaceae family.

Chromosomes, Plant

Chromosome-level genome assembly of Ceroplastes pseudoceriferus Green, 1935 (Hemiptera: Coccidae).

Soft scales (Hemiptera: Coccidae) are significant polyphagous pests and majority of which are invasive species. The 364.14 Mb chromosome-level genome of Ceroplastes pseudoceriferus was assembled in this work, with a contig N50 length of 6.16 Mb and scafold N50 length of 21.24 Mb. Approximately 99.89% of assembled sequences were anchored into 18 chromosomes with the assistance of Hi-C reads. Furthermore, approximately 53.98% of the genome was composed of repetitive elements. In total, 10,475 protein-coding genes were predicted, of which 9503 (90.72%) genes were functionally annotated. The BUSCO analysis demonstrated the completeness of the genome annotation is 92.54%. This genome represents first high-quality chromosome level assembly of Coccidae, thereby advancing our knowledge of Coccidae insects and developing effective management strategies that protect crops, forests, and natural ecosystems.

Animals

Chromosome-level genome assembly of the ornamental plant Alcea rosea.

Alcea rosea, a member of the Malvaceae family, is celebrated for its rich floral palette and global horticultural significance. Here, we present a high-quality reference genome for A. rosea, achieving a genome assembly size of 1.01 Gbp, with a Contig N50 length of 36.61 Mbp. The genome sequence was successfully mapped to 21 chromosomes, and the scaffold N50 length reached 52.57 Mbp, with a scaffold genome completeness of 99.6%. A total of 565.84 Mbp (comprising 56% of the genome) of repetitive sequences were identified, with transposable elements being predominant, particularly long terminal repeat (LTR) elements, which accounted for 48.44% of the genome. 51,436 genes were annotated. Among these predicted genes, the average gene length and coding sequence (CDS) length were 2739.92 bp and 1242.54 bp, respectively.

Genome, Plant

Chromosome-level genome assembly of the large carpenter bee Xylocopa dejeanii Lepeletier, 1841 (Hymenoptera: Apidae).

Xylocopinae, a diverse bee subfamily comprising over 1,000 bee species, and also a major model system for studying the pollination and evolution of sociality. The lack of chromosome-level genome assembly resources for the Xylocopinae limits our research of their biology and evolution. Here, we provided the first pseudo-chromosomes genome assembly of the Xylocopa dejeanii combined PacBio CLR long reads, Illumina sequences, and Hi-C data. The final genome is 194.44 Mb located in 16 chromosomes. Our assembly includes 141 scaffolds, with a scaffold N50 length of 13.15 Mb. BUSCO analysis revealed 99.00% completeness. Genome annotation identified 28.27 Mb of repetitive elements, 10,970 protein-coding genes, and 432 ncRNAs. This high-quality X. dejeanii assembly advances our understanding of Xylocopinae genomics and provides new insights into bee evolution.

Animals

Chromosome-level genome assembly of an Arctic fish species pale eelpout (Lycodes pallidus).

Eelpouts (Zoarcidae) are known for their bipolar distributions and distinctive biogeographic histories. However, limited genomic data have hindered our understanding of their adaptive evolution. In this study, we present a thoroughly annotated chromosome-level genome assembly of pale eelpout (Lycodes pallidus) generated through the integration of Illumina, PacBio circular consensus, and Hi-C sequencing techniques. The final assembly spans 753.4 Mb, with its high quality confirmed by a scaffold N50 of 28.6 Mb and a Benchmarking Universal Single-Copy Ortholog (BUSCO) completeness of 99.3%. In comparison to other eelpouts and related fishes, the L. pallidus genome is larger and exhibits greater repetitive element content, accounting for approximately 45% of its total length. We annotated 21,419 protein-coding genes, a significant proportion of which are involved in signal transduction mechanisms and transcription. These findings provide valuable genetic resources for elucidating the evolutionary mechanisms underlying polar fish adaptation.

Animals

Loss of maternal PADI6 disrupts DNA methylation and genomic imprinting maintenance in late preimplantation mouse embryos.

BACKGROUND: The maternal-effect protein PADI6, which is part of the subcortical maternal complex, is involved in proper spindle assembly, organelle distribution, ribosome storage, and cytoplasmic lattice organization in mouse oocytes. In humans, variants of PADI6 are associated with female infertility and multilocus imprinting disturbance in offspring. Recently, it was demonstrated that PADI6 plays a role in the storage and cytoplasmic localization of epigenetic factors, including UHRF1 and DNMT1. Moreover, maternal PADI6 depletion leads to defective epigenetic reprogramming and zygotic genome activation but not to an imprinting defect in two-cell mouse embryos. These findings raise the possibility that imprinting disturbances arise later in development. RESULTS: By employing combined single-blastocyst RNA-seq/BS-seq and immunostaining validation in the embryos derived from Padi6P620A-mutant oocytes, we investigated the role of Padi6 in late preimplantation development. We demonstrated that embryos that overcame the two-cell stage block had a dramatic reduction in UHRF1 and DNMT1 protein levels, a decrease in H3K9me3, and whole-genome hypomethylation, including most imprinted loci and repetitive elements, at the blastocyst stage. Furthermore, these maternal mutant embryos showed deregulation of inner cell mass markers and defective blastocyst implantation, but no effect on trophoblast differentiation. CONCLUSION: Our results demonstrate that maternal PADI6 is a key regulator of the stability of epigenetic factors required to maintain repressive marks in late preimplantation mouse embryos. Its deficiency results in genomic imprinting defects that closely resemble those found in human patients and provide a mechanistic explanation for MLID caused by maternal PADI6 variants. Furthermore, the impairment of blastocyst implantation capacity, likely due to dysregulation of inner cell mass differentiation, provides new mechanistic insights into the control of female fertility and embryo development exerted by PADI6.

DNA Methylation