PubMed HealthSearch

SEARCH · PubMed Health

Results for “centromere evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Haplotype-resolved telomere-to-telomere genome assembly of Populus lasiocarpa unveils retrotransposon-driven centromere evolution.

Centromeres, essential for chromosome segregation, exhibit remarkable evolutionary dynamism in sequence composition and structural organization. Here, we report the first haplotype-resolved, telomere-to-telomere genome assembly of Populus lasiocarpa (PLAS) and precisely map all 38 functional centromeres through CENH3 ChIP-Seq. Unlike classical satellite-rich centromeres in model plants, PLAS centromeres lack abundant satellite arrays but are dominated by retrotransposons, particularly RLG and RIL elements, which form intricate nested TE arrays within the functional centromeric regions, disrupting their structural integrity and driving their evolution. Comparative analysis with P. trichocarpa reveals a conserved retrotransposon-dominated architecture, despite minimal sequence conservation. We propose a cyclic model of centromere evolution in which autonomous retrotransposons destabilize functional centromeres through epigenetic erosion, triggering neocentromere formation at pericentromeric sites enriched in transposable elements (TEs) and tandem repeats (TRs). These neocentromeres either succumb to recurrent retrotransposon invasions or stabilize through KARMA-mediated TR expansion, ultimately giving rise to satellite-rich centromeres. Our work redefines centromeres as dynamic, epigenetically plastic domains shaped by retrotransposon-TR antagonism, challenging the satellite-centric paradigm and offering novel insights into plant genome evolution.

Retroelements

Distinct evolutionary trajectories of subgenomic centromeres in polyploid wheat.

BACKGROUND: Centromeres are crucial for precise chromosome segregation and maintaining genome stability during cell division. However, their evolutionary dynamics, particularly in polyploid organisms with complex genomic architectures, remain largely enigmatic. Allopolyploid wheat, with its well-defined hierarchical ploidy series and recent polyploidization history, serves as an excellent model to explore centromere evolution. RESULTS: In this study, we perform a systematic comparative analysis of centromeres in common wheat and its corresponding ancestral species, utilizing the latest comprehensive reference genome assembly available. Our findings reveal that wheat centromeres predominantly consist of five types of centromeric-specific retrotransposon elements (CRWs), with CRW1 and CRW2 being the most prevalent. We identify distinct evolutionary trajectories in the functional centromeres of each subgenome, characterized by variations in copy number, insertion age, and CRW composition. By utilizing CENH3-ChIP data across various ploidy levels, we uncover a series of CRW invasion events that have shaped the evolution of AA subgenome centromeres. Conversely, the evolutionary process of the DD subgenome centromeres involves their expansion from diploid to hexaploid wheat, facilitating adaptation to a larger genomic context. Integration of complete einkorn centromere assemblies and Aegilops tauschii pan-genomes further revealed subgenome-specific centromere evolutionary trajectories. By inclusion of synthetic hexaploid from S2-S3 generations, alongside 2x/6 × natural accessions, we demonstrate that DD subgenome centromere expansion represents a gradual evolutionary process rather than an immediate response to polyploidization. CONCLUSIONS: Our study provides a comprehensive landscape of centromere adaptation, evolution, and maturation, along with insights into how retrotransposon invasions drive centromere evolution in polyploid wheat.

Centromere

Two CENH3 paralogs in the green alga Chlamydomonas reinhardtii have a redundantly essential function and associate with ZeppL-LINE1 elements.

Centromeres in eukaryotes are defined by the presence of histone H3 variant CENP-A/CENH3. Chlamydomonas encodes two predicted CENH3 paralogs, CENH3.1 and CENH3.2, that have not been previously characterized. We generated peptide antibodies to unique N-terminal epitopes for each of the two predicted Chlamydomonas CENH3 paralogs as well as an antibody against a shared CENH3 epitope. All three CENH3 antibodies recognized proteins of the expected size on immunoblots and had punctate nuclear immunofluorescence staining patterns. These results are consistent with both paralogs being expressed and localized to centromeres. CRISPR-Cas9-mediated insertional mutagenesis was used to generate predicted null mutations in either CENH3.1 or CENH3.2. Single mutants were viable but cenh3.1 cenh3.2 double mutants were not recovered, confirming that the function of CENH3 is essential. We sequenced and assembled two chromosome-scale Chlamydomonas genomes from strains CC-400 and UL-1690 (a derivative of CC-1690) with complete centromere sequences for 17/17 and 14/17 chromosomes respectively, enabling us to compare centromere evolution across four isolates with near complete assemblies. These data revealed significant changes across isolates between homologous centromeres including mobility and degeneration of ZeppL-LINE1 (ZeppL) transposons that comprise the major centromere repeat sequence in Chlamydomonas. We used cleavage under targets and tagmentation (CUT&Tag) to purify and map CENH3-bound genomic sequences and found enrichment of CENH3-binding almost exclusively at predicted centromere regions. An interesting exception was chromosome 2 in UL-1690, which had enrichment at its genetically mapped centromere repeat region as well as a second, distal location, centered around a single recently acquired ZeppL insertion. The CENH3-bound regions of the 17 Chlamydomonas centromeres ranged from 63.5 kb (average lower estimate) to 175 kb (average upper estimate). The relatively small size of its centromeres suggests that Chlamydomonas may be a useful organism for testing and deploying artificial chromosome technologies.

Chlamydomonas reinhardtii

The centromere landscapes of four karyotypically diverse Papaver species provide insights into chromosome evolution and speciation.

Understanding the roles played by centromeres in chromosome evolution and speciation is complicated by the fact that centromeres comprise large arrays of tandemly repeated satellite DNA, which hinders high-quality assembly. Here, we used long-read sequencing to generate nearly complete genome assemblies for four karyotypically diverse Papaver species, P. setigerum (2n = 44), P. somniferum (2n = 22), P. rhoeas (2n = 14), and P. bracteatum (2n = 14), collectively representing 45 gapless centromeres. We identified four centromere satellite (cenSat) families and experimentally validated two representatives. For the two allopolyploid genomes (P. somniferum and P. setigerum), we characterized the subgenomic distribution of each satellite and identified a "homogenizing" phase of centromere evolution in the aftermath of hybridization. An interspecies comparison of the peri-centromeric regions further revealed extensive centromere-mediated chromosome rearrangements. Taking these results together, we propose a model for studying cenSat competition after hybridization and shed further light on the complex role of the centromere in speciation.

Centromere

The dynamic centromere.

Centromeres are fundamental chromosomal structures that ensure accurate chromosome segregation during cell division. Despite their conserved and essential role in maintaining genomic stability, centromeres are subject to rapid evolutionary change. At the heart of centromere identity is the histone H3 variant CENP-A, an epigenetic mark that defines and propagates active centromeres and is essential for their function. Recent evidence supports a rapid evolution of centromere DNA sequences but also suggests a certain degree of flexibility in CENP-A deposition and propagation. The phenomenon of centromere drift, recently observed in humans, highlights how the dynamic repositioning of CENP-A and associated epigenetic environment over time maintains a regulated equilibrium, ensuring centromere function despite positional variation. Understanding these processes is crucial for unraveling centromere dynamics and their broader implications for genome stability and evolution.

Centromere

Chromosome-specific centromeric patterns define the centeny map of the human genome.

Centromeres are epigenetically specified by distinct chromatin, whereas their DNA varies between species and individuals. This extensive sequence divergence makes comparative analyses between centromeres challenging. In this study, we identified a chromosome-specific architectural pattern across the human genome, defined by the conserved spacing of a functionally relevant centromeric DNA motif. The distribution of these sites along chromosome arms constitutes the human "centeny map." By using a custom Genomic Centromere Profiling (GCP) pipeline, we leveraged the motif's position, orientation, and organization to construct structural models that enable reclassification of human chromosomal clusters, detection of centromere expansion, and identification of structural variants and misassembled regions. The high-resolution maps derived from this pattern not only provide a framework for comparative analysis of centromeres across evolution and disease but also offer a new dimension for chromosome annotation, assembly, and characterization.

Humans

Germline-restricted chromosome of songbirds has different centromere compared to regular chromosomes.

Centromeres are an important part of chromosomes which direct chromosome segregation during cell division. Their modifications can therefore explain the unusual mitotic and meiotic behaviour of certain chromosomes, such as the germline-restricted chromosome (GRC) of songbirds. This chromosome is eliminated from somatic cells during early embryogenesis and later also from male germ cells during spermatogenesis. Although the mechanism of elimination is not yet known, it is possible that it involves a modification of the centromeric sequence on the GRC, resulting in problems with the attachment of this chromosome to the mitotic or meiotic spindle and its lagging during anaphase, which eventually leads to its elimination from the nucleus. However, the repetitive nature and rapid evolution of centromeres make their identification and comparative analysis across species and chromosomes challenging. Here, we used a combination of cytogenetic and genomic approaches to identify the centromeric sequences of two closely related songbird species, the common nightingale (Luscinia megarhynchos) and the thrush nightingale (L. luscinia). We found a 436-bp satellite repeat present in the centromeric regions of all regular chromosomes (i.e., autosomes and sex chromosomes), making it a strong candidate for the centromeric repeat. This centromeric repeat was highly similar between the two nightingale species. Interestingly, hybridization of the probe to this satellite repeat on meiotic spreads suggested that this repeat is missing on the GRC. Our results indicate that the change of the centromeric sequence may underlie the unusual inheritance and programmed DNA elimination of the GRC in songbirds.

Animals

Evolution and domestication-trait associations of ultra-long centromere haplotypes in pepper plants.

Centromeric and pericentromeric regions of most eukaryotic genomes are highly repetitive and strongly recombination-suppressed, confounding efforts to resolve genetic variation, population structure and phenotypic associations. Pepper (Capsicum annuum) centromeres are nearly devoid of satellite repeats, facilitating assembly and population-level comparison of centromeric regions. Here we integrate 9 near-complete genome assemblies, CENH3 ChIP-seq profiles from 26 diverse accessions, and resequencing and phenotypic data from ~400 cultivated and wild accessions to investigate population-level diversity and phenotypic relevance of pepper peri/centromeric regions. Functional centromere positions are largely fixed on 8 of 12 chromosomes, whereas the remaining 4 carry distinct centromeric epialleles shaped mainly by centromere repositioning and pericentromeric inversions. Pepper centromeres are embedded within ultra-long centromere-spanning haplotype (cenhap) blocks, ranging from 29.8 to 112.9 Mb and collectively covering 23.96% of the genome; each block contains only 1-4 major haplotypes. Some cenhaps may act as supergene-like units and are strongly associated with fruit traits, probably because recombination-suppressed intervals harbour multiple fruit-related genes, including OFP and F-box genes. F2 segregation assays further reveal transmission distortion of chromosomes carrying alternative cenhaps. Together, these findings highlight peri/centromeric regions as underrecognized reservoirs of agronomically important variation.

Centromere

Satellite DNAs in Drosophila koepferae (repleta group) reveal patterns of origin, chromosomal organization, transcription, and turnover in the buzzatii cluster.

Satellite DNAs (satDNAs) are non-coding tandem repeats that can comprise more than 20% of eukaryotic genomes. They contribute to structural and regulatory processes in the genome and often evolve rapidly, shaping early stages of genetic differentiation between populations and species. Although Drosophila has long served as a model for studying satDNA biology, little is known about satDNAs in non-model Drosophila species, particularly within the repleta group, one of the most species-rich lineages in the genus. To reduce such bias, several studies have focused on the buzzatii cluster (repleta group). However, D. koepferae remained the only species lacking comprehensive satDNA data, limiting comparative analyses. Here, we used publicly available genomic sequencing data from two D. koepferae populations (Argentina and Bolivia) to characterize their satDNA content. Both populations share the same set of five satDNAs (CDSTR8, CDSTR138, CDSTR230, DBC-150 and CDSTR177), which together account for ~ 0,9% of the genomic DNA. We show that CDSTR177 originated through amplification of an internal segment of the Galileo transposable element, an event restricted to D. koepferae. All satDNAs localize to heterochromatic regions, with CDSTR138 most likely associated to the centromeres of most chromosomes. Transcripts from all satDNAs were detected, although at low levels. Our results provide new insights into the origin, genomic contribution, expression and evolution of satDNAs in the buzzatii cluster, support incipient differentiation between Argentinean and Bolivian populations of D. koepferae and contribute to clarifying the phylogenetic position of this species within the buzzatii cluster.

Animals

A complete genome for the common marmoset.

The common marmoset is a New World monkey widely used to study primate evolution and human disease. We present a telomere-to-telomere (T2T) reference assembly for the species, plus three near-T2T haplotypes. These resolve previously inaccessible regions, including the centromeres, sex chromosomes, subterminal satellites, acrocentric chromosomes, and the major histocompatibility complex (MHC). We find marmoset centromeres carry dimeric alpha satellites with chromosomal specificity, flanked by inactive layers interpreted as ancestral centromere remnants. We assemble gene-poor, satellite-rich short arms of the acrocentrics and find that most can harbor rDNA and all share pseudo-homolog regions (PHRs). PHR-sharing chromosomes also share closely related centromeric satellites, consistent with a model of ongoing rDNA-facilitated recombinational exchange between heterologous chromosomes. We further identify over 500 marmoset-lineage-specific transcribed genes with previously unknown transcript models or expansions. These resources, along with a preliminary pangenome, improve the utility of the marmoset as a model organism and address gaps in primate genome evolution.

Animals

Genetic Differentiation is Constrained to Chromosomal Inversions and Putative Centromeres in Locally Adapted Populations With Higher Gene Flow.

The impact of genome structure on adaptation is a growing focus in evolutionary biology, revealing an important role for structural variation and recombination landscapes in shaping genetic diversity across genomes and among populations. This is particularly relevant when local adaptation occurs despite gene flow, where clustering of differentiated loci can maintain locally adapted variants by reducing recombination between them. However, the limited genomic resources for nonmodel species, including reference genomes and recombination maps, have constrained our understanding of these patterns. In this study, we leverage the Atlantic silverside-a nonmodel fish with extensive local adaptation across a steep latitudinal gradient-as an ideal system to explore how genome structure influences adaptation under varying levels of gene flow, using a newly available reference genome and multiple recombination maps. Analyzing 168 genomes from four populations, we found a continuum of genome-wide differentiation increasing from south to north, reflecting higher connectivity among southern populations and reduced gene flow at northern latitudes. With increasing gene flow, the number and clustering of FST outlier loci also increased, with differentiated loci found exclusively within large haploblocks harboring inversions and smaller peaks overlapping putative centromeric regions. Notably, sequence divergence was only evident in inversions, supporting their role in adaptive divergence with gene flow, whereas centromeric regions appeared differentiated because of low recombination and diversity, with no indication of elevated divergence. Our results support the hypothesis that clustered genomic architectures evolve with high gene flow and enhance our understanding of how inversions and centromeres are linked to different evolutionary processes.

Gene Flow

Evolutionary fingerprints of epithelial-to-mesenchymal transition.

Mesenchymal plasticity has been extensively described in advanced epithelial cancers; however, its functional role in malignant progression is controversial1-5. The function of epithelial-to-mesenchymal transition (EMT) and cell plasticity in tumour heterogeneity and clonal evolution is poorly understood. Here we clarify the contribution of EMT to malignant progression in pancreatic cancer. We used somatic mosaic genome engineering technologies to trace and ablate malignant mesenchymal lineages along the EMT continuum. The experimental evidence clarifies the essential contribution of mesenchymal lineages to pancreatic cancer evolution. Spatial genomic analysis, single-cell transcriptomic and epigenomic profiling of EMT clarifies its contribution to the emergence of genomic instability, including events of chromothripsis. Genetic ablation of mesenchymal lineages robustly abolished these mutational processes and evolutionary patterns, as confirmed by cross-species analysis of pancreatic and other human solid tumours. Mechanistically, we identified that malignant cells with mesenchymal features display increased chromatin accessibility, particularly in the pericentromeric and centromeric regions, in turn resulting in delayed mitosis and catastrophic cell division. Thus, EMT favours the emergence of genomic-unstable, highly fit tumour cells, which strongly supports the concept of cell-state-restricted patterns of evolution, whereby cancer cell speciation is propagated to progeny within restricted functional compartments. Restraining the evolutionary routes through ablation of clones capable of mesenchymal plasticity, and extinction of the derived lineages, halts the malignant potential of one of the most aggressive forms of human cancer.

Animals

Satellite DNA evolution in Tytonidae (Aves: Strigiformes): dynamic repeat landscapes despite conserved karyotypes.

The elevated chromosome numbers observed in Tytonidae relative to the putative ancestral avian karyotype suggest that lineage-specific chromosomal fissions may have played an important role in the evolutionary history of this family. Here, we provide the first cytogenetic characterization of the American barn owl (Tyto furcata) and performs a comparative repeatome analysis across members of the Tytonidae, including other two species, the Western barn owl (Tyto alba), and the Oriental bay owl (Phodilus badius). The karyotype of T. furcata showed a 2n = 92, closely resembling that previously described for T. alba, indicating a high degree of chromosomal conservation within Tytonidae. Although T. furcata and T. alba exhibit similar karyotypic organization, comparative repeatomic analyses revealed differences in their composition, including variation in satellite DNA (satDNA) repertoires and abundance. Eight satDNA families were identified in T. furcata, nine in T. alba, and 28 in P. badius, highlighting the dynamic evolution of repetitive sequences. Several satDNA families were shared between T. furcata and T. alba, whereas some appeared species-specific, supporting the library hypothesis of satDNA evolution. In P. badius, multiple satDNAs exhibited similarity to transposable elements, suggesting that mobile elements contributed to their diversification. Cytogenetic analyses demonstrated centromeric heterochromatin distribution in T. furcata, as well as a large heterochromatic W chromosome enriched in DNA repeats. The localization of satDNAs in centromeric regions and the apparent accumulation of repeats on the W chromosome reinforce the role of repetitive sequences in chromosome organization and sex chromosome differentiation. Together, these findings reveal repeatome diversification despite conserved macrochromosomal structure and provide new insights into genome evolution and chromosomal dynamics in birds.

Animals

Neotelomeres and telomere-spanning chromosomal arm fusions in cancer genomes revealed by long-read sequencing.

Alterations in the structure and location of telomeres are pivotal in cancer genome evolution. Here, we applied both long-read and short-read genome sequencing to assess telomere repeat-containing structures in cancers and cancer cell lines. Using long-read genome sequences that span telomeric repeats, we defined four types of telomere repeat variations in cancer cells: neotelomeres where telomere addition heals chromosome breaks, chromosomal arm fusions spanning telomere repeats, fusions of neotelomeres, and peri-centromeric fusions with adjoined telomere and centromere repeats. These results provide a framework for the systematic study of telomeric repeats in cancer genomes, which could serve as a model for understanding the somatic evolution of other repetitive genomic elements.

Humans

Neotelomeres and Telomere-Spanning Chromosomal Arm Fusions in Cancer Genomes Revealed by Long-Read Sequencing.

Alterations in the structure and location of telomeres are key events in cancer genome evolution. However, previous genomic approaches, unable to span long telomeric repeat arrays, could not characterize the nature of these alterations. Here, we applied both long-read and short-read genome sequencing to assess telomere repeat-containing structures in cancers and cancer cell lines. Using long-read genome sequences that span telomeric repeat arrays, we defined four types of telomere repeat variations in cancer cells: neotelomeres where telomere addition heals chromosome breaks, chromosomal arm fusions spanning telomere repeats, fusions of neotelomeres, and peri-centromeric fusions with adjoined telomere and centromere repeats. Analysis of lung adenocarcinoma genome sequences identified somatic neotelomere and telomere-spanning fusion alterations. These results provide a framework for systematic study of telomeric repeat arrays in cancer genomes, that could serve as a model for understanding the somatic evolution of other repetitive genomic elements.

Telomere

Complete chromosome 21 centromere sequencing of families with Down syndrome reveals centromere size asymmetry.

Down syndrome, the most common form of human intellectual disability, is caused by nondisjunction and chromosome 21 trisomy (T21). Small centromeres have been hypothesized to contribute to its aetiology and studies on mammals suggest that larger centromeres are more efficiently transmitted, yet complete sequencing of chromosome 21 (chr21) centromeres has been particularly challenging. Using long-read sequencing, we sequenced and assembled the centromeres from eight families that include a child with free T21 (1 trio, 6 child-mother duos, and 1 singleton) all resulting from maternal meiosis I errors. Two of these families carry the smallest chr21 centromeres (143 and 181 kbp) observed in female individuals to date, exhibiting a ~10.7- and ~19.4-fold centromeric α-satellite higher-order repeat array size difference between the maternally inherited homologs, respectively. In both cases, the longer centromere harbors a poorly defined centromere dip region, marked by DNA hypomethylation, in the proband but not in the mother. A comparison of all proband chr21 centromeres (n=24) to those of controls (n=261) shows that small centromeres are not enriched in families with T21 (p-value=0.73); contrarily, chr21 extreme centromere size asymmetry (>10-fold) is unique of T21 (p-value=0.003), suggesting that this feature may represent a genetic risk factor for a subset of families with free T21. Additionally, phylogenetic reconstruction reveals that human chr21 has been particularly prone to such variation with some of the biggest size differences occurring over the last ~17 thousand years of human evolution.

Down syndrome

The genetic control of rapid genome content divergence in Arabidopsis thaliana.

Genome evolution in eukaryotes is predominantly driven by the dynamics of repetitive sequences, which vary widely in both copy number and sequence composition. Rates of repeat evolution differ between and within species and are likely modulated by both genetics and environment. To uncover factors shaping the rate of genome content evolution, we analyzed 1,142 resequenced Arabidopsis thaliana genomes using a novel K-mer based approach to characterize genome content variation and identify hypervariable regions underlying differences in repeat abundance. We next treated repeat abundance as a quantitative trait and performed genome-wide association analyses across more than 400 repeat families to identify the genetic basis of copy number variation. Integrating these results through a meta-GWAS approach revealed both cis-acting variants and more than 50 trans-acting loci that regulate repeat abundance genome-wide. Cis-acting variation was predominantly localized to pericentromeric and centromeric regions, whereas trans-acting loci were enriched for candidate genes involved in DNA replication, DNA repair, DNA methylation regulation. Finally, we found evidence that purifying selection acts against mutations that accelerate genome content divergence, favoring alleles that constrain repeat expansion. Together, these findings provide new insights into the genetic architecture and evolutionary forces shaping genome evolution in A. thaliana and establish a framework for investigating these processes in other plant species.

Journal Article

The genetic control of rapid genome content divergence in Arabidopsis thaliana.

Genome evolution in eukaryotes is predominantly driven by the dynamics of repetitive sequences, which vary widely in both copy number and sequence composition. Rates of repeat evolution differ between and within species and are likely modulated by both genetics and environment. To uncover factors shaping the rate of genome content evolution, we analyzed 1043 resequenced Arabidopsis thaliana genomes using a novel K-mer-based approach to characterize genome content variation and identify hypervariable regions underlying differences in repeat abundance. We next treated repeat abundance as a quantitative trait and performed genome-wide association analyses across more than 400 repeat families to identify the genetic basis of copy number variation. Integrating these results through a meta-GWAS approach revealed both cis-acting variants and more than 50 candidate trans-acting loci associated with repeat abundance genome-wide. Cis-acting variation was predominantly localized to pericentromeric and centromeric regions, whereas trans-acting loci were enriched for candidate genes involved in DNA replication, DNA repair, and DNA methylation regulation. The results are consistent with purifying selection acting against mutations that accelerate genome content divergence, favoring alleles that constrain repeat expansion. Together, these findings provide new insights into the genetic architecture and evolutionary forces shaping genome evolution in A. thaliana and establish a framework for investigating these processes in other plant species.

Arabidopsis