PubMed HealthSearch

SEARCH · PubMed Health

Results for “haplotype-resolved genome”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A chromosome-level, haplotype-resolved genome assembly for the barn owl, Tyto alba.

Recent advances in long-read sequencing have enabled near telomere-to-telomere (T2T) assemblies across diverse taxa. However, avian genomes remain challenging due to numerous microchromosomes, small, typically < 20Mb, DNA molecules that are gene-, GC-, and repeat-rich. As a consequence, microchromosomes are often missing from genome assemblies. Here, we present a chromosome-level, haplotype-resolved genome assembly for the Western barn owl (Tyto alba). Using a trio-binning strategy with Illumina parental reads combined with PacBio HiFi and Oxford Nanopore Technologies data, we generated two phased contig sets. These were scaffolded into 40 linkage groups using a linkage map. Comparative analyses identified unplaced HiFi scaffolds corresponding to microchromosomes, which we integrated into six additional microchromosomes using long reads information. The two assemblies present 46 chromosomes, matching the karyotype of the species. They exhibit strong synteny between parental haplotypes, except for a &#x223c;38 Mb complex region on chromosome 7 containing nested inversions. This high-quality reference provides a haplotype-resolved and chromosome-level genome for Strigiformes, enabling fine-scale studies of structural variation and avian genome evolution.

Tyto alba

Haplotype-resolved genome assembly and implementation of VitExpress, an open interactive transcriptomic platform for grapevine.

Haplotype-resolved genome assemblies were produced for Chasselas and Ugni Blanc, two heterozygous Vitis vinifera cultivars by combining high-fidelity long-read sequencing and high-throughput chromosome conformation capture (Hi-C). The telomere-to-telomere full coverage of the chromosomes allowed us to assemble separately the two haplo-genomes of both cultivars and revealed structural variations between the two haplotypes of a given cultivar. The deletions/insertions, inversions, translocations, and duplications provide insight into the evolutionary history and parental relationship among grape varieties. Integration of de novo single long-read sequencing of full-length transcript isoforms (Iso-Seq) yielded a highly improved genome annotation. Given its higher contiguity, and the robustness of the IsoSeq-based annotation, the Chasselas assembly meets the standard to become the annotated reference genome for V. vinifera. Building on these resources, we developed VitExpress, an open interactive transcriptomic platform, that provides a genome browser and integrated web tools for expression profiling, and a set of statistical tools (StatTools) for the identification of highly correlated genes. Implementation of the correlation finder tool for MybA1, a major regulator of the anthocyanin pathway, identified candidate genes associated with anthocyanin metabolism, whose expression patterns were experimentally validated as discriminating between black and white grapes. These resources and innovative tools for mining genome-related data are anticipated to foster advances in several areas of grapevine research.

Vitis

Haplotype-resolved genome of Forsythia suspensa reveals the reticulate evolution in Oleaceae and a novel gene cluster regulating stamen development.

The olive family (Oleaceae) comprises numerous species of economic, horticultural, and medicinal importance. Despite its significance, the evolutionary history of this complex family remains enigmatic. Here, we generated a high-quality haplotype-resolved genome of Forsythia suspensa, a distylous species that occupies a key phylogenetic position in Oleaceae. The 2 haplotypes exhibit significant allelic divergence with potential allele-specific regulation. We reconstructed the polyploidization history of Oleaceae by confirming and precisely dating a shared whole-genome triplication and an independent whole-genome duplication event. We revealed a complex reticulate evolution that gave rise to the tribe Oleeae: an initial hybridization between Forsythieae (&#x2642;) and Jasmineae (&#x2640;), a subsequent backcrossing event, and a final whole-genome duplication. We identified a novel tandemly duplicated pectin methylesterase inhibitor gene cluster that regulates filament length and pollen size via restricting cell elongation in the long-styled morph. Dosage augmentation via stepwise cluster formation (0.99 to 3.83&#x2005;Mya) may contribute to maintaining stamen traits of the long-styled morph. These FsPMEIs are co-expressed with many cell wall-related genes, suggesting a functional link in cell wall modification. Our study reveals the reticulate evolution in Oleaceae and a novel gene cluster controlling stamen development in F. suspensa and provides valuable haplotype-resolved genomic resources for heterostylous species, offering novel framework and molecular pathways to understand plant adaptive evolution.

Forsythia

De novo haplotype-resolved genome assembly of the endemic kiwifruit Actinidia hubeiensis.

The genus Actinidia, which encompasses the widely cultivated kiwifruit, is characterized by its rich species diversity. Wild Actinidia species serve as invaluable germplasm reservoirs for crop improvement. As an important kiwifruit species, Actinidia hubeiensis represents a unique taxonomic group endemic to Hubei Province, contributing valuable genetic diversity to the genus Actinidia. Here, we present a haplotype-resolved genome assembly for A. hubeiensis. The two haplotype assemblies (Hap1 and Hap2) spanned 658.03&#x2009;Mb (N50&#x2009;=&#x2009;23.16&#x2009;Mb) and 597.19&#x2009;Mb (N50&#x2009;=&#x2009;20.89&#x2009;Mb), encoding 35,741 and 36,647 high-confidence protein-coding genes, respectively. Based on comprehensive assessments, both haplotypes demonstrated high completeness (BUSCO completeness&#x2009;>&#x2009;99%), excellent continuity (LAI up to 21.67), low base-error rates (QV&#x2009;>&#x2009;40), and nearly complete read mapping rates (>&#xa0;98%). This genome assembly provides crucial genomic resources for the genus, enriching our understanding of kiwifruit biodiversity and offering new insights into the genetic background and evolutionary characteristics of this distinctive species.

Actinidia

Chromosome-level haplotype-resolved genome assembly of the giant honeycomb oyster, Hyotissa hyotis.

The giant honeycomb oyster, Hyotissa hyotis, a common bivalve inhabitant of tropical and subtropical coastal waters, holds significant ecological and economic importance due to its shell characteristics, rapid growth, and high-quality adductor muscle. However, the lack of high-quality genome has impeded the genetic study and artificial breeding of this species. In this study, we provided the first chromosomal-level haplotype-resolved assembly for the H. hyotis (2n&#x2009;=&#x2009;20) by combining PacBio HiFi long-read and Hi-C sequencing. We obtained a haplotype-resolved assembly of 3.39&#x2009;Gb in size, of which 96.69% were anchored to 20 chromosomes. The haplotype A and B genome (HapA and HapB) was 1,639.90 and 1,643.23&#x2009;Mb in size, respectively. Accordingly, a total of 28,720 and 29,003 protein-coding genes were annotated from HapA and HapB. Through the BUSCO evaluation, the assembly and annotation results exhibited the completeness value of 94.65% and 94.03% for HapA, while 94.13% and 92.98% for HapB. This high-quality genome assembly provides valuable resource for further genetic studies and genetic improvement of the group of oysters.

Animals

Haplotype-resolved 3D genome maps reveal RNAPII-mediated allelic regulation in hybrid rice.

To understand how the two parental genomes coordinate transcription in hybrids, chromatin architecture must be resolved at the haplotype level. Here, using phased Bridge-Linker Hi-C, we reconstructed a haplotype-resolved three-dimensional (3D) genome of the elite hybrid rice (Oryza sativa) line Shanyou 63 (SY63). We identified extensive allele-specific chromatin conformations. Furthermore, we generated allele-resolved RNAPII ChIA-PET maps and phased transcriptomes to explore how chromatin interactions contribute to allelic regulation. Although maternal and paternal homologs share broadly similar chromatin features, we detected widespread haplotype-biased RNAPII binding and chromatin looping at high resolution. These allele-specific RNAPII-mediated contacts were significantly associated with biased expression. Stronger RNAPII binding on one haplotype promoted the formation of long-range regulatory loops with distal genes, thereby contributing to allele-biased transcription at a subset of loci, even when promoter-proximal RNAPII occupancy was comparable between alleles. These results demonstrate that subtle differences in RNAPII engagement and 3D regulatory wiring between parental haplotypes can reshape transcriptional output in hybrids, providing new insights into the mechanisms underlying the allelic regulation of gene expression.

Allele-specific chromatin interactions

Comparative genomic analysis of Artemisia argyi reveals asymmetric expansion of terpene synthases and conservation of artemisinin biosynthesis.

Artemisia argyi, a perennial herb of the Asteraceae family, possesses significant therapeutic and economic value. We present a 7.88&#x2009;Gb chromosome-level haplotype-resolved genome assembly, revealing its unique evolutionary trajectory. The karyotype (2n&#x2009;=&#x2009;34) of A. argyi is that of an autotetraploid, which underwent gametic chromosome fusion prior to species-specific whole-genome duplication (WGD-3). The genome exhibits pronounced multivalent chromosome pairing and frequent recombination among homologous groups. Asymmetrical evolution following WGD-3 is a hallmark feature, evidenced by imbalanced allelic gene loss and widespread neofunctionalization. The terpene synthase (TPS) gene family exemplifies this pattern, having expanded through four duplication events in A. argyi. Recent tandem duplications and allelic functional differentiation have generated substantial gene functional diversity. Notably, we identified a tandem-duplicated six-copy ADS homolog (AarADS)-a key TPS gene in the artemisinin biosynthetic pathway of Artemisia annua (AanADS)-localized exclusively to a single chromosome in A. argyi. Unlike AanADS, which converts farnesyl pyrophosphate (FPP) to amorpha-4,11-diene, AarADS catalyzes FPP to &#x3b1;-bisabolol. Evolutionary analysis suggested that AanADS acquired its specialized function via a derived mutation in the A. annua lineage. This study elucidates the genomic evolution underpinning A. argyi's distinctive medicinal properties.

Alkyl and Aryl Transferases

Genomic and evolutionary basis of parthenogenesis in a disease-vector tick species.

Haemaphysalis longicornis is an important tick species and pathogen vector characterized by the co-circulation of triploid parthenogenetic and diploid bisexual strains. However, the evolutionary basis of parthenogenesis in this species is unclear. Here we report reference-quality, haplotype-resolved genome assemblies of the parthenogenetic strain and two reference-quality genomes of the bisexual strains. Comparative genomic analysis revealed high collinearity between the parthenogenetic and bisexual genomes, with a stable chromosomal architecture maintained among the three haplotypes of the parthenogenetic strain. The parthenogenetic H. longicornis genome exhibited a major expansion in cell cycle-related gene families, including the inhibitor of apoptosis protein (IAP) family, but was characterized by a contraction in other gene families. Population resequencing of 179 individuals revealed two distinct subpopulations, with chromosome&#x2009;7 harbouring high genetic differentiation and several candidate genes probably associated with parthenogenesis. Functional experiments showed that knockdown of the BIRC5 gene, a member of the IAP family, suppressed oviposition in both strains, with the parthenogenetic strain exhibiting milder adverse effects probably due to a stronger transcriptional response. Overall, our results reveal the genomic and evolutionary features associated with polyploid parthenogenesis in H. longicornis.

Animals

Integrated genomic, transcriptomic, and metabolomic analyses of Chrysanthemum aromaticum provide insights into the volatile terpene biosynthesis.

Chrysanthemum aromaticum is renowned for its uniformly emitted strong and attractive scent, primarily attributed to volatile terpenes. Despite its commercial and horticultural significance, the molecular mechanisms underlying volatile terpene production in C. aromaticum remain largely unexplored. Here, we present the haplotype-resolved genome assembly of C. aromaticum, with a total size of 3.10&#x2009;Gb, comprising nine anchored chromosomes with a contig N50 of 30.66&#x2009;Mb and a scaffold N50 of 350.58&#x2009;Mb. Phylogenetic analyses revealed a distant relationship between C. aromaticum and C. indicum, suggesting that C. aromaticum likely represents a distinct species rather than a variety of C. indicum. Through integrated genomic, transcriptomic, metabolomic, and biochemical analyses, we identified seven TPS involved in monoterpene biosynthesis and six TPS for sesquiterpene biosynthesis. Notably, comparative genomic analysis revealed a gene cluster for &#x3b1;-bisabolol biosynthesis in C. aromaticum, which has specifically expanded in Chrysanthemum species through tandem gene duplications, contributing to the elevated accumulation of &#x3b1;-bisabolol in the leaves of C. aromaticum. Our study provides important insights into the biosynthesis of volatile terpenes, highlighting the genetic basis for C. aromaticum's unique aromatic profile.

Chrysanthemum

Haplotype-resolved telomere-to-telomere genome assembly of Populus lasiocarpa unveils retrotransposon-driven centromere evolution.

Centromeres, essential for chromosome segregation, exhibit remarkable evolutionary dynamism in sequence composition and structural organization. Here, we report the first haplotype-resolved, telomere-to-telomere genome assembly of Populus lasiocarpa (PLAS) and precisely map all 38 functional centromeres through CENH3 ChIP-Seq. Unlike classical satellite-rich centromeres in model plants, PLAS centromeres lack abundant satellite arrays but are dominated by retrotransposons, particularly RLG and RIL elements, which form intricate nested TE arrays within the functional centromeric regions, disrupting their structural integrity and driving their evolution. Comparative analysis with P. trichocarpa reveals a conserved retrotransposon-dominated architecture, despite minimal sequence conservation. We propose a cyclic model of centromere evolution in which autonomous retrotransposons destabilize functional centromeres through epigenetic erosion, triggering neocentromere formation at pericentromeric sites enriched in transposable elements (TEs) and tandem repeats (TRs). These neocentromeres either succumb to recurrent retrotransposon invasions or stabilize through KARMA-mediated TR expansion, ultimately giving rise to satellite-rich centromeres. Our work redefines centromeres as dynamic, epigenetically plastic domains shaped by retrotransposon-TR antagonism, challenging the satellite-centric paradigm and offering novel insights into plant genome evolution.

Retroelements

Biosynthesis and heterologous production of the &#x3b1;-agarofuran scaffold of Celangulin V from Celastrus angulatus.

Celangulin V is a widely used biopesticide derived from Celastrus angulatus, and features antifeedant and insecticidal properties as a dihydro-&#x3b2;-agarofuran (DH&#x3b2;AF) sesquiterpenoid. Its biosynthesis remains largely unexplored. Here, we assemble a chromosome-level and haplotype-resolved reference genome of C. angulatus, with each haplotype assembled into 23 pseudochromosomes and achieving scaffold N50 of 14.31 and 14.01&#x2009;Mb, respectively. This high-quality genome reveals that a recent &#x3b2; whole-genome triplication (&#x3b2;-WGT) event occurred ~34.3 million years ago, and that the expansion of sesquiterpene synthases and cytochrome P450s from the CYP71BE family results from whole-genome duplication (WGD) event and tandem duplication, respectively. We identify CaTPS16 as a &#x3b3;-eudesmol synthase, and show that CYP71BE416 further catalyzes the &#x3b3;-eudesmol to tetrahydrofuran ring &#x3b1;-agarofuran for Celangulin V biosynthesis. We further achieve the de novo synthesis of &#x3b1;-agarofuran in Saccharomyces cerevisiae through combined coexpression of these genes. This study has significantly increases the available genomic resources of the Celastraceae family, improves our understanding of the biosynthetic origins and evolution of the tetrahydrofuran ring in DH&#x3b2;AF sesquiterpenoids, and enables its heterologous bioproduction in microbial chassis.

Celastrus

Haplotype-specific expression of a terpene synthase underlies linalool variation in the grapevine cultivar Riesling.

Grapevine cultivars vary widely in monoterpenoid content, yet the genetic and regulatory mechanisms underlying this variation remain poorly characterized beyond highly aromatic Muscat types. We profiled free volatiles and monoterpenoid glycosides in a Riesling &#xd7; Cabernet Sauvignon F1 mapping population, revealing extensive variation and transgressive segregation consistent with multigenic control. QTL mapping identified 70 significant loci associated with 48 volatile compounds and monoterpene glycosides, including two major QTLs explaining 33.6% and 33.4% of phenotypic variance in (3S)-linalool accumulation. Integration of haplotype-resolved transcriptomics with metabolite data, enabled by a chromosome-scale diploid Riesling genome assembly, resolved a (3S)-linalool/nerolidol synthase cluster on chromosome 10 and identified VviTPS54 as the strongest candidate underlying linalool variation. VviTPS54 exhibited haplotype-specific expression strongly correlated with (3S)-linalool accumulation across genotypes, while no QTL was detected at the 1-deoxy-D-xylulose-5-phosphate synthase 1 (VviDXS1) locus previously identified in Muscat cultivars. In addition, VviDXS1 expression was not correlated with terpene levels, indicating that regulatory variation within terpene synthase clusters, rather than methylerythritol phosphate (MEP) pathway flux, drives monoterpenoid composition in this population. These results establish regulatory variation of terpene synthases as a key mechanism underlying monoterpenoid diversity in grapevine and demonstrate that resolving such variation requires haplotype-phased genome assemblies coupled with haplotype-resolved transcriptomics to detect allele-specific expression differences at complex, heterozygous loci.

Grapevine

Phased telomere-to-telomere reference genome and pangenome reveal an expansion of resistance genes during apple domestication.

The cultivated apple (Malus domestica Borkh.) is a cross-pollinated perennial fruit tree of great economic importance. Earlier versions of apple reference genomes were unphased, fragmented, and lacked comprehensive insights into the apple's highly heterozygous genome, which impeded advances in genetic studies and breeding programs. In this study, we assembled a haplotype-resolved telomere-to-telomere (T2T) reference genome for the diploid apple cultivar Golden Delicious. Subsequently, we constructed a pangenome based on 12 assemblies from wild and cultivated species to investigate the dynamic changes of functional genes. Our results revealed the gene gain and loss events during apple domestication. Compared with cultivated species, more gene families in wild species were significantly enriched in oxidative phosphorylation, pentose metabolic process, responses to salt, and abscisic acid biosynthesis process. Our analyses also demonstrated a higher prevalence of different types of resistance gene analogs (RGAs) in cultivars than their wild relatives, partially attributed to segmental and tandem duplication events in certain RGAs classes. Structural variations, mainly deletions and insertions, have affected the presence and absence of TIR-NB-ARC-LRR, NB-ARC-LRR, and CC-NB-ARC-LRR genes. Additionally, hybridization/introgression from wild species has also contributed to the expansion of resistance genes in domesticated apples. Our haplotype-resolved T2T genome and pangenome provide important resources for genetic studies of apples, emphasizing the need to study the evolutionary mechanisms of resistance genes in apple breeding.

Malus

Comparative genomics reveals lineage-associated structural variation and diversification in a barley fungal pathogen.

Leaf rust, caused by Puccinia hordei, is a major barley disease worldwide. Despite repeated shifts in virulence, contrasting reproductive histories, and emerging fungicide insensitivity, the genomic basis of its diversification and adaptation remains poorly understood. In this study, we generated haplotype-resolved, chromosome-level genome assemblies for two isolates with contrasting virulence and analyzed 41 Australian isolates collected over 54&#x2009;yr (1966-2020), integrating comparative and population genomics, mating-type gene phylogenies, chromosome-specific k-mer profiling, genome-wide copy-number variation (CNV) analysis, and gene-expression analysis. We identified a structurally dynamic chromosome characterized by repeat-associated rearrangements, structural variation, and lineage-associated CNV, representing the first evidence in a rust fungus of chromosome-scale structural diversification of this extent. Population analyses distinguished clonally expanded lineages from recombination-associated lineages, with mating-type gene phylogenies providing further support for lineage differentiation. More recently collected isolates showed increased duplication-associated variation, and CNV boundaries were associated with structural-variant breakpoints. We also identified lineage-associated amplification of Cyp51, with increased copy number associated with higher transcript abundance, supporting a potential role in fungicide adaptation. Overall, our findings highlight structural variation, contrasting reproductive histories, and lineage-associated CNV as important contributors to diversification in P. hordei, providing insights for future rust pathogen surveillance and management strategies.

Cyp51 gene

A chromosome-level genome of the Nicobar pigeon, Caloenas nicobarica.

The Nicobar pigeon (Caloenas nicobarica), the closest living relative of the extinct Dodo (Raphus cucullatus), is endemic to Southeast Asia with a fragmented distribution across numerous small islands. It suffers from habitat loss, hunting, and predation from invasive species, resulting in its classification as Near Threatened by the International Union for the Conservation of Nature. We have generated a haplotype-resolved and chromosome-level genome assembly of the Nicobar pigeon using a combination of PacBio HiFi long-read sequencing and Arima Hi-C chromatin interaction mapping. This assembly includes two haplotypes, each spanning approximately 1.2 Gb. Haplotype 1 has a contig N50 of 25.2&#xa0;Mb and a scaffold N50 of 79.7&#xa0;Mb, whereas haplotype 2 has a contig N50 of 24.7&#xa0;Mb and a scaffold N50 of 107.9&#xa0;Mb. As the first high-quality genome assembly of any bird in the Columbidae Indo-Pacific clade, this resource provides valuable insights for phylogenetic studies. Furthermore, the phylogenetic proximity of the Nicobar pigeon to the Dodo (R. cucullatus) and the Rodrigues Solitaire (Pezophaps solitaria) offers a unique opportunity to study these extinct species, making this assembly a critical resource for evolutionary studies. It also offers a unique model for studying genetic diversity, adaptation, and speciation in island environments. This genomic resource will not only enhance our understanding of the evolutionary history of the Nicobar pigeon but also serve as a valuable tool for future conservation efforts aimed at preserving this unique species and its fragile island ecosystem.

Animals

ECHO: a nanopore sequencing-based workflow for (epi)genetic profiling of the human repeatome.

SUMMARY: The human genome is dominated by repetitive DNA, whose genetic and epigenetic variation plays a key role in gene regulation, genome stability, and disease. Recent advances in long-read sequencing now enable large-scale, haplotype-resolved, and DNA methylation-informative analysis of the human genome, including on previously inaccessible complex and repetitive regions. However, the comprehensive, simultaneous characterisation of the "human repeatome" remains challenging, largely due to the lack of comprehensive tools integrated in a single pipeline that can capture the full spectrum of variation across diverse types of DNA repeats. Here, we present ECHO, a user-friendly, Snakemake-based pipeline for the "(Epi)genomic Characterisation of Human Repetitive Elements using Oxford Nanopore Sequencing." ECHO provides a reproducible and scalable framework for end-to-end analysis of whole-genome nanopore sequencing data, enabling integrative but also tailored (epi)genetic analyses of the human repeatome. AVAILABILITY AND IMPLEMENTATION: ECHO is freely available at Github: https://github.com/leenput/ECHO-pipeline, with the archived version at Zenodo: https://zenodo.org/records/19068468.

Humans

Haplotype-resolved reconstruction and functional interrogation of cancer karyotypes.

Complex karyotype changes are widespread in cancer genomes. A major gap in cancer genome characterization is the resolution of rearranged chromosomes with chromosome-length continuity. Here, we describe a two-tiered approach to determine the segmental composition of rearranged chromosomes with haplotype resolution. First, we present refLinker, a bioinformatic method for robust determination of chromosomal haplotypes using cancer Hi-C data. By contrast with existing methods, refLinker is insensitive to the presence of large-scale DNA deletions, duplications, and high-level amplification in cancer genomes. Second, we demonstrate a computational strategy to determine the segmental structure of rearranged chromosomes using haplotype-specific Hi-C contacts. We apply these methods to breast cancer genomes and provide direct evidence for long-range transcriptional changes associated with rearrangements of the inactive X chromosome. Together, these results highlight refLinker's broad utility for studying the functional consequences of chromosomal rearrangements.

Humans