PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “haplotype-resolved genome assembly”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A chromosome-level, haplotype-resolved genome assembly for the barn owl, Tyto alba.

Recent advances in long-read sequencing have enabled near telomere-to-telomere (T2T) assemblies across diverse taxa. However, avian genomes remain challenging due to numerous microchromosomes, small, typically < 20Mb, DNA molecules that are gene-, GC-, and repeat-rich. As a consequence, microchromosomes are often missing from genome assemblies. Here, we present a chromosome-level, haplotype-resolved genome assembly for the Western barn owl (Tyto alba). Using a trio-binning strategy with Illumina parental reads combined with PacBio HiFi and Oxford Nanopore Technologies data, we generated two phased contig sets. These were scaffolded into 40 linkage groups using a linkage map. Comparative analyses identified unplaced HiFi scaffolds corresponding to microchromosomes, which we integrated into six additional microchromosomes using long reads information. The two assemblies present 46 chromosomes, matching the karyotype of the species. They exhibit strong synteny between parental haplotypes, except for a &#x223c;38 Mb complex region on chromosome 7 containing nested inversions. This high-quality reference provides a haplotype-resolved and chromosome-level genome for Strigiformes, enabling fine-scale studies of structural variation and avian genome evolution.

Tyto alba↗

Haplotype-resolved genome assembly and implementation of VitExpress, an open interactive transcriptomic platform for grapevine.

Haplotype-resolved genome assemblies were produced for Chasselas and Ugni Blanc, two heterozygous Vitis vinifera cultivars by combining high-fidelity long-read sequencing and high-throughput chromosome conformation capture (Hi-C). The telomere-to-telomere full coverage of the chromosomes allowed us to assemble separately the two haplo-genomes of both cultivars and revealed structural variations between the two haplotypes of a given cultivar. The deletions/insertions, inversions, translocations, and duplications provide insight into the evolutionary history and parental relationship among grape varieties. Integration of de novo single long-read sequencing of full-length transcript isoforms (Iso-Seq) yielded a highly improved genome annotation. Given its higher contiguity, and the robustness of the IsoSeq-based annotation, the Chasselas assembly meets the standard to become the annotated reference genome for V. vinifera. Building on these resources, we developed VitExpress, an open interactive transcriptomic platform, that provides a genome browser and integrated web tools for expression profiling, and a set of statistical tools (StatTools) for the identification of highly correlated genes. Implementation of the correlation finder tool for MybA1, a major regulator of the anthocyanin pathway, identified candidate genes associated with anthocyanin metabolism, whose expression patterns were experimentally validated as discriminating between black and white grapes. These resources and innovative tools for mining genome-related data are anticipated to foster advances in several areas of grapevine research.

Vitis↗

De novo haplotype-resolved genome assembly of the endemic kiwifruit Actinidia hubeiensis.

The genus Actinidia, which encompasses the widely cultivated kiwifruit, is characterized by its rich species diversity. Wild Actinidia species serve as invaluable germplasm reservoirs for crop improvement. As an important kiwifruit species, Actinidia hubeiensis represents a unique taxonomic group endemic to Hubei Province, contributing valuable genetic diversity to the genus Actinidia. Here, we present a haplotype-resolved genome assembly for A. hubeiensis. The two haplotype assemblies (Hap1 and Hap2) spanned 658.03&#x2009;Mb (N50&#x2009;=&#x2009;23.16&#x2009;Mb) and 597.19&#x2009;Mb (N50&#x2009;=&#x2009;20.89&#x2009;Mb), encoding 35,741 and 36,647 high-confidence protein-coding genes, respectively. Based on comprehensive assessments, both haplotypes demonstrated high completeness (BUSCO completeness&#x2009;>&#x2009;99%), excellent continuity (LAI up to 21.67), low base-error rates (QV&#x2009;>&#x2009;40), and nearly complete read mapping rates (>&#xa0;98%). This genome assembly provides crucial genomic resources for the genus, enriching our understanding of kiwifruit biodiversity and offering new insights into the genetic background and evolutionary characteristics of this distinctive species.

Actinidia↗

Chromosome-level haplotype-resolved genome assembly of the giant honeycomb oyster, Hyotissa hyotis.

The giant honeycomb oyster, Hyotissa hyotis, a common bivalve inhabitant of tropical and subtropical coastal waters, holds significant ecological and economic importance due to its shell characteristics, rapid growth, and high-quality adductor muscle. However, the lack of high-quality genome has impeded the genetic study and artificial breeding of this species. In this study, we provided the first chromosomal-level haplotype-resolved assembly for the H. hyotis (2n&#x2009;=&#x2009;20) by combining PacBio HiFi long-read and Hi-C sequencing. We obtained a haplotype-resolved assembly of 3.39&#x2009;Gb in size, of which 96.69% were anchored to 20 chromosomes. The haplotype A and B genome (HapA and HapB) was 1,639.90 and 1,643.23&#x2009;Mb in size, respectively. Accordingly, a total of 28,720 and 29,003 protein-coding genes were annotated from HapA and HapB. Through the BUSCO evaluation, the assembly and annotation results exhibited the completeness value of 94.65% and 94.03% for HapA, while 94.13% and 92.98% for HapB. This high-quality genome assembly provides valuable resource for further genetic studies and genetic improvement of the group of oysters.

Animals↗

Genomic and evolutionary basis of parthenogenesis in a disease-vector tick species.

Haemaphysalis longicornis is an important tick species and pathogen vector characterized by the co-circulation of triploid parthenogenetic and diploid bisexual strains. However, the evolutionary basis of parthenogenesis in this species is unclear. Here we report reference-quality, haplotype-resolved genome assemblies of the parthenogenetic strain and two reference-quality genomes of the bisexual strains. Comparative genomic analysis revealed high collinearity between the parthenogenetic and bisexual genomes, with a stable chromosomal architecture maintained among the three haplotypes of the parthenogenetic strain. The parthenogenetic H. longicornis genome exhibited a major expansion in cell cycle-related gene families, including the inhibitor of apoptosis protein (IAP) family, but was characterized by a contraction in other gene families. Population resequencing of 179 individuals revealed two distinct subpopulations, with chromosome&#x2009;7 harbouring high genetic differentiation and several candidate genes probably associated with parthenogenesis. Functional experiments showed that knockdown of the BIRC5 gene, a member of the IAP family, suppressed oviposition in both strains, with the parthenogenetic strain exhibiting milder adverse effects probably due to a stronger transcriptional response. Overall, our results reveal the genomic and evolutionary features associated with polyploid parthenogenesis in H. longicornis.

Animals↗

Integrated genomic, transcriptomic, and metabolomic analyses of Chrysanthemum aromaticum provide insights into the volatile terpene biosynthesis.

Chrysanthemum aromaticum is renowned for its uniformly emitted strong and attractive scent, primarily attributed to volatile terpenes. Despite its commercial and horticultural significance, the molecular mechanisms underlying volatile terpene production in C. aromaticum remain largely unexplored. Here, we present the haplotype-resolved genome assembly of C. aromaticum, with a total size of 3.10&#x2009;Gb, comprising nine anchored chromosomes with a contig N50 of 30.66&#x2009;Mb and a scaffold N50 of 350.58&#x2009;Mb. Phylogenetic analyses revealed a distant relationship between C. aromaticum and C. indicum, suggesting that C. aromaticum likely represents a distinct species rather than a variety of C. indicum. Through integrated genomic, transcriptomic, metabolomic, and biochemical analyses, we identified seven TPS involved in monoterpene biosynthesis and six TPS for sesquiterpene biosynthesis. Notably, comparative genomic analysis revealed a gene cluster for &#x3b1;-bisabolol biosynthesis in C. aromaticum, which has specifically expanded in Chrysanthemum species through tandem gene duplications, contributing to the elevated accumulation of &#x3b1;-bisabolol in the leaves of C. aromaticum. Our study provides important insights into the biosynthesis of volatile terpenes, highlighting the genetic basis for C. aromaticum's unique aromatic profile.

Chrysanthemum↗

Comparative genomic analysis of Artemisia argyi reveals asymmetric expansion of terpene synthases and conservation of artemisinin biosynthesis.

Artemisia argyi, a perennial herb of the Asteraceae family, possesses significant therapeutic and economic value. We present a 7.88&#x2009;Gb chromosome-level haplotype-resolved genome assembly, revealing its unique evolutionary trajectory. The karyotype (2n&#x2009;=&#x2009;34) of A. argyi is that of an autotetraploid, which underwent gametic chromosome fusion prior to species-specific whole-genome duplication (WGD-3). The genome exhibits pronounced multivalent chromosome pairing and frequent recombination among homologous groups. Asymmetrical evolution following WGD-3 is a hallmark feature, evidenced by imbalanced allelic gene loss and widespread neofunctionalization. The terpene synthase (TPS) gene family exemplifies this pattern, having expanded through four duplication events in A. argyi. Recent tandem duplications and allelic functional differentiation have generated substantial gene functional diversity. Notably, we identified a tandem-duplicated six-copy ADS homolog (AarADS)-a key TPS gene in the artemisinin biosynthetic pathway of Artemisia annua (AanADS)-localized exclusively to a single chromosome in A. argyi. Unlike AanADS, which converts farnesyl pyrophosphate (FPP) to amorpha-4,11-diene, AarADS catalyzes FPP to &#x3b1;-bisabolol. Evolutionary analysis suggested that AanADS acquired its specialized function via a derived mutation in the A. annua lineage. This study elucidates the genomic evolution underpinning A. argyi's distinctive medicinal properties.

Alkyl and Aryl Transferases↗

Haplotype-resolved telomere-to-telomere genome assembly of Populus lasiocarpa unveils retrotransposon-driven centromere evolution.

Centromeres, essential for chromosome segregation, exhibit remarkable evolutionary dynamism in sequence composition and structural organization. Here, we report the first haplotype-resolved, telomere-to-telomere genome assembly of Populus lasiocarpa (PLAS) and precisely map all 38 functional centromeres through CENH3 ChIP-Seq. Unlike classical satellite-rich centromeres in model plants, PLAS centromeres lack abundant satellite arrays but are dominated by retrotransposons, particularly RLG and RIL elements, which form intricate nested TE arrays within the functional centromeric regions, disrupting their structural integrity and driving their evolution. Comparative analysis with P. trichocarpa reveals a conserved retrotransposon-dominated architecture, despite minimal sequence conservation. We propose a cyclic model of centromere evolution in which autonomous retrotransposons destabilize functional centromeres through epigenetic erosion, triggering neocentromere formation at pericentromeric sites enriched in transposable elements (TEs) and tandem repeats (TRs). These neocentromeres either succumb to recurrent retrotransposon invasions or stabilize through KARMA-mediated TR expansion, ultimately giving rise to satellite-rich centromeres. Our work redefines centromeres as dynamic, epigenetically plastic domains shaped by retrotransposon-TR antagonism, challenging the satellite-centric paradigm and offering novel insights into plant genome evolution.

Retroelements↗

Genomic analysis of Dasiphora on the Qinghai-Tibet Plateau provides insights into genetic divergence and flower color variation.

The Qinghai-Tibet Plateau (QTP) harbors diverse alpine flora, including the ecologically significant shrubs Dasiphora fruticosa and D. glabra, for which taxonomic uncertainties remain and adaptive mechanisms are still poorly understood. Based on high-quality genome assembly, population resequencing, and multi-omics integration, we elucidated their evolutionary divergence and flower color genetics. Chromosome-level haplotype-resolved genomes were assembled: autotetraploid D. fruticosa (929.99&#x2009;Mb) and diploid D. glabra (450.89&#x2009;Mb). Phylogenetic analysis showed that the tetraploid D. fruticosa and D. glabra in this study clustered together, while the diploid D. fruticosa sequenced by previous research formed a distinct lineage clustered outside. Consistently, population structure analysis of 55 samples revealed three major clades, with D. fruticosa further subdivided into two divergent branches. Additionally, hybridization events detected by Admixture, coupled with ploidy complexity identified via flow cytometry highlight the intricate genetic relationships within this genus. Adaptive gene families expanded in antioxidant (flavonoid synthesis) and secondary metabolism pathways, adapting to ultraviolet radiation and cold stress. Natural selection analysis identified 193 candidate genes (e.g., TFB5 in the DNA repair pathway), predominantly localized to chromosome 5, which are potential candidates for high-altitude adaptation. Transcriptome and metabolome analyses showed D. fruticosa's yellow petals derive from flavonol (quercetin) accumulation, while D. glabra's white petals result from proanthocyanidin biosynthesis via high LAR/ANR expression. This study provides insights into the taxonomic revision and adaptive genetic divergence of alpine plants, and offers a foundation for horticultural improvement of Dasiphora.

Flowers↗

Haplotype-specific expression of a terpene synthase underlies linalool variation in the grapevine cultivar Riesling.

Grapevine cultivars vary widely in monoterpenoid content, yet the genetic and regulatory mechanisms underlying this variation remain poorly characterized beyond highly aromatic Muscat types. We profiled free volatiles and monoterpenoid glycosides in a Riesling &#xd7; Cabernet Sauvignon F1 mapping population, revealing extensive variation and transgressive segregation consistent with multigenic control. QTL mapping identified 70 significant loci associated with 48 volatile compounds and monoterpene glycosides, including two major QTLs explaining 33.6% and 33.4% of phenotypic variance in (3S)-linalool accumulation. Integration of haplotype-resolved transcriptomics with metabolite data, enabled by a chromosome-scale diploid Riesling genome assembly, resolved a (3S)-linalool/nerolidol synthase cluster on chromosome 10 and identified VviTPS54 as the strongest candidate underlying linalool variation. VviTPS54 exhibited haplotype-specific expression strongly correlated with (3S)-linalool accumulation across genotypes, while no QTL was detected at the 1-deoxy-D-xylulose-5-phosphate synthase 1 (VviDXS1) locus previously identified in Muscat cultivars. In addition, VviDXS1 expression was not correlated with terpene levels, indicating that regulatory variation within terpene synthase clusters, rather than methylerythritol phosphate (MEP) pathway flux, drives monoterpenoid composition in this population. These results establish regulatory variation of terpene synthases as a key mechanism underlying monoterpenoid diversity in grapevine and demonstrate that resolving such variation requires haplotype-phased genome assemblies coupled with haplotype-resolved transcriptomics to detect allele-specific expression differences at complex, heterozygous loci.

Grapevine↗

Comparative genomics reveals lineage-associated structural variation and diversification in a barley fungal pathogen.

Leaf rust, caused by Puccinia hordei, is a major barley disease worldwide. Despite repeated shifts in virulence, contrasting reproductive histories, and emerging fungicide insensitivity, the genomic basis of its diversification and adaptation remains poorly understood. In this study, we generated haplotype-resolved, chromosome-level genome assemblies for two isolates with contrasting virulence and analyzed 41 Australian isolates collected over 54&#x2009;yr (1966-2020), integrating comparative and population genomics, mating-type gene phylogenies, chromosome-specific k-mer profiling, genome-wide copy-number variation (CNV) analysis, and gene-expression analysis. We identified a structurally dynamic chromosome characterized by repeat-associated rearrangements, structural variation, and lineage-associated CNV, representing the first evidence in a rust fungus of chromosome-scale structural diversification of this extent. Population analyses distinguished clonally expanded lineages from recombination-associated lineages, with mating-type gene phylogenies providing further support for lineage differentiation. More recently collected isolates showed increased duplication-associated variation, and CNV boundaries were associated with structural-variant breakpoints. We also identified lineage-associated amplification of Cyp51, with increased copy number associated with higher transcript abundance, supporting a potential role in fungicide adaptation. Overall, our findings highlight structural variation, contrasting reproductive histories, and lineage-associated CNV as important contributors to diversification in P. hordei, providing insights for future rust pathogen surveillance and management strategies.

Cyp51 gene↗

Phased telomere-to-telomere reference genome and pangenome reveal an expansion of resistance genes during apple domestication.

The cultivated apple (Malus domestica Borkh.) is a cross-pollinated perennial fruit tree of great economic importance. Earlier versions of apple reference genomes were unphased, fragmented, and lacked comprehensive insights into the apple's highly heterozygous genome, which impeded advances in genetic studies and breeding programs. In this study, we assembled a haplotype-resolved telomere-to-telomere (T2T) reference genome for the diploid apple cultivar Golden Delicious. Subsequently, we constructed a pangenome based on 12 assemblies from wild and cultivated species to investigate the dynamic changes of functional genes. Our results revealed the gene gain and loss events during apple domestication. Compared with cultivated species, more gene families in wild species were significantly enriched in oxidative phosphorylation, pentose metabolic process, responses to salt, and abscisic acid biosynthesis process. Our analyses also demonstrated a higher prevalence of different types of resistance gene analogs (RGAs) in cultivars than their wild relatives, partially attributed to segmental and tandem duplication events in certain RGAs classes. Structural variations, mainly deletions and insertions, have affected the presence and absence of TIR-NB-ARC-LRR, NB-ARC-LRR, and CC-NB-ARC-LRR genes. Additionally, hybridization/introgression from wild species has also contributed to the expansion of resistance genes in domesticated apples. Our haplotype-resolved T2T genome and pangenome provide important resources for genetic studies of apples, emphasizing the need to study the evolutionary mechanisms of resistance genes in apple breeding.

Malus↗

A chromosome-level genome of the Nicobar pigeon, Caloenas nicobarica.

The Nicobar pigeon (Caloenas nicobarica), the closest living relative of the extinct Dodo (Raphus cucullatus), is endemic to Southeast Asia with a fragmented distribution across numerous small islands. It suffers from habitat loss, hunting, and predation from invasive species, resulting in its classification as Near Threatened by the International Union for the Conservation of Nature. We have generated a haplotype-resolved and chromosome-level genome assembly of the Nicobar pigeon using a combination of PacBio HiFi long-read sequencing and Arima Hi-C chromatin interaction mapping. This assembly includes two haplotypes, each spanning approximately 1.2 Gb. Haplotype 1 has a contig N50 of 25.2&#xa0;Mb and a scaffold N50 of 79.7&#xa0;Mb, whereas haplotype 2 has a contig N50 of 24.7&#xa0;Mb and a scaffold N50 of 107.9&#xa0;Mb. As the first high-quality genome assembly of any bird in the Columbidae Indo-Pacific clade, this resource provides valuable insights for phylogenetic studies. Furthermore, the phylogenetic proximity of the Nicobar pigeon to the Dodo (R. cucullatus) and the Rodrigues Solitaire (Pezophaps solitaria) offers a unique opportunity to study these extinct species, making this assembly a critical resource for evolutionary studies. It also offers a unique model for studying genetic diversity, adaptation, and speciation in island environments. This genomic resource will not only enhance our understanding of the evolutionary history of the Nicobar pigeon but also serve as a valuable tool for future conservation efforts aimed at preserving this unique species and its fragile island ecosystem.

Animals↗

Biosynthesis and heterologous production of the &#x3b1;-agarofuran scaffold of Celangulin V from Celastrus angulatus.

Celangulin V is a widely used biopesticide derived from Celastrus angulatus, and features antifeedant and insecticidal properties as a dihydro-&#x3b2;-agarofuran (DH&#x3b2;AF) sesquiterpenoid. Its biosynthesis remains largely unexplored. Here, we assemble a chromosome-level and haplotype-resolved reference genome of C. angulatus, with each haplotype assembled into 23 pseudochromosomes and achieving scaffold N50 of 14.31 and 14.01&#x2009;Mb, respectively. This high-quality genome reveals that a recent &#x3b2; whole-genome triplication (&#x3b2;-WGT) event occurred ~34.3 million years ago, and that the expansion of sesquiterpene synthases and cytochrome P450s from the CYP71BE family results from whole-genome duplication (WGD) event and tandem duplication, respectively. We identify CaTPS16 as a &#x3b3;-eudesmol synthase, and show that CYP71BE416 further catalyzes the &#x3b3;-eudesmol to tetrahydrofuran ring &#x3b1;-agarofuran for Celangulin V biosynthesis. We further achieve the de novo synthesis of &#x3b1;-agarofuran in Saccharomyces cerevisiae through combined coexpression of these genes. This study has significantly increases the available genomic resources of the Celastraceae family, improves our understanding of the biosynthetic origins and evolution of the tetrahydrofuran ring in DH&#x3b2;AF sesquiterpenoids, and enables its heterologous bioproduction in microbial chassis.

Celastrus↗

Uncovering the mechanism of female restitution in sugarcane hybrids.

Variations of meiosis, which normally halve genetic complements prior to fertilization, can have profound consequences. For example, whole-genome duplications (polyploidy) have shaped the evolution and diversification of most angiosperm lineages. The century-long success of sugarcane interspecific hybrids has been attributed to unusual female restitution-an unreduced maternal gamete fusing with a normal haploid paternal gamete1,2. Here we generated haplotype-resolved genomes of octoploid Saccharum officinarum LA Purple and decaploid Saccharum spontaneum US56-14-4. Eight F1 hybrids between these species exhibited 2:1 maternal to paternal genomic ratios, with 2 assemblies revealing canonical haploid sets of approximately 40 paternal and approximately 80 maternal chromosomes. The maternal chromosomes comprise 40 pairs of duplicated, partially recombined sister chromatids that retain around 62.5% of maternal genetic diversity, characteristic of second division restitution. Using single-molecule long-read sequencing and a novel algorithm that is broadly applicable to polyploid genomes, we identified two classes of recombination breakpoints, including a previously unrecognized configuration supported by both recombinant and non-recombinant reads, across all hybrids and diagnostic of second division restitution. These findings resolve a century-old cytological debate, add new insights into meiotic variations, and offer a genomic approach to accelerate genetic gain in this globally critical sugar and bioenergy crop.

Chimera↗

Pan-genomics and multi-omics for deciphering genetic variation and accelerating genetic improvement in ruminant livestock.

Livestock reference genomes have transformed the discovery of variants associated with production, reproduction, health, and environmental adaptation. Nevertheless, a single linear reference represents only one mosaic haplotype and incompletely captures sequence diversity within a species, particularly structural variants, copy-number changes, repeat-rich regions, and breed-specific sequences. Pangenomes address this limitation by integrating multiple high-quality assemblies or population-scale variants into a unified sequence or graph representation. Concurrently, multi-omics approaches connect genomic variation with transcriptomic, epigenomic, manuscriptproteomic, metabolomic, and microbiome responses, thereby improving biological interpretation of genotype-phenotype relationships. This review synthesizes recent progress in livestock pangenomics and multi-omics, with emphasis on cattle, goats, sheep, water buffalo, and chickens. It describes advances in long-read and haplotype-resolved sequencing, graph construction, structural-variant discovery and genotyping, functional annotation, and integrative analysis. Recent pangenome studies have uncovered substantial non-reference sequence, reduced reference bias, identified breed- and population-specific structural variants, and resolved candidate variants underlying pigmentation, body size, tail morphology, cashmere production, altitude adaptation, and other economically relevant traits. However, translation into routine breeding remains constrained by uneven population representation, inconsistent structural-variant definitions, limited functional annotation, computational demands, and insufficient validation across environments. Future progress will depend on diverse near-complete assemblies, graph-aware imputation and genomic prediction, long-read transcriptomics, single-cell and spatial omics, rigorous causal validation, and open, interoperable resources. Together, these developments can support more accurate, resilient, and biologically informed livestock improvement. Importantly, current dairy-cattle evidence indicates that pangenome-derived structural variants can substantially improve variant discovery and functional interpretation while yielding only marginal average gains in routine genomic prediction, favoring targeted augmentation rather than wholesale replacement of established SNP-based evaluations.

Animals↗