PubMed HealthSearch

SEARCH · PubMed Health

Results for “Hi-C scaffolding”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Assembling genomes of non-model plants: A case study with evolutionary insights from Ranunculus (Ranunculaceae).

Whereas genome sequencing and assembly technologies are improving, cost can still be prohibitive for plant species with large, complex genomes. As a consequence, genomics work on some taxa in evolutionarily pivotal positions in the vascular plant tree of life has been hampered. The species-rich genus Ranunculus (Ranunculaceae) is an important angiosperm group for the study of polyploidy, apomixis, and reticulate evolution. However, neither mitochondrial nor high-quality nuclear genome sequences are available. This limits phylogenomic, functional, and taxonomic analyses thus far. Here, we tested Illumina short-read, Oxford Nanopore Technology (ONT) and PacBio (HiFi) long-read, and hybrid-read assembly strategies. We sequenced the diploid progenitor species R. cassubicifolius (R. auricomus species complex) and selected the best assemblies in terms of completeness, contiguity, and quality scores. We first assembled the plastome (156 kbp, 85 genes) and mitogenome (1.18 Mbp, 40 genes) sequences using Illumina and Illumina-PacBio-hybrid strategies, respectively. We also present an updated plastome and the first mitogenome phylogeny of Ranunculaceae, including studies of gene loss (e.g., infA, ycf15, or rps) with evolutionary implications. For the nuclear genome sequence, we favored a PacBio-based assembly polished three times with filtered short reads and subsequently scaffolded into eight pseudochromosomes by chromatin conformation data (Hi-C). We obtained a haploid genome sequence of 2.69 Gbp, with 94.1% complete BUSCO genes found and 35 482 annotated genes, and inferred ancient gene duplications compared to existing Ranunculales genomes. The genomic information presented here will enable advanced evolutionary-functional analyses for the species complex, but also for the genus and beyond Ranunculaceae.

Ranunculus

Chromosome-level genome assembly of Elaeocarpus petiolatus (Elaeocarpaceae).

Elaeocarpus petiolatus is an ecologically and economically important species in tropical and subtropical forests. Despite its significance, the lack of genomic resources has hindered research on the genetic diversity and adaptive traits of E. petiolatus. To address this gap, we present a comprehensive chromosome-level genome assembly of E. petiolatus generated using advanced PacBio high-fidelity (HiFi) long-read sequencing and Hi-C technology. The assembly spans 322.45 Mb, with a scaffold N50 of 20.58 Mb, indicating that 37.11% of the genome is composed of repetitive elements. We identified 25,295 protein-coding genes, of which 96.74% were functionally annotated. This high-quality genome provides a critical resource for understanding the genetic mechanisms underlying environmental adaptability and biosynthesis of bioactive compounds in E. petiolatus, thereby supporting conservation efforts and sustainable forest management. The assembled genome and associated sequencing data are publicly available, facilitating further evolutionary and functional studies on the Elaeocarpaceae family.

Chromosomes, Plant

The chromosome-level genome assembly and annotation of the silver-lipped pearl oyster, Pinctada maxima.

The silver-lipped pearl oyster (Pinctada maxima) is a valuable tropical aquaculture species, playing a crucial economic role in the global pearl industry. However, the lack of genomic reference limits our in-depth understanding of this species in genome-based breeding, conservation, evolution and adaptation. Here, annotated chromosome-level reference genome for P. maxima was generated by integrating PacBio long-read sequencing, Illumina short-read sequencing, and Hi-C sequencing data. The total genome size is 1,264.93&#x2009;Mb, with contig N50 and scaffold N50 of 649&#x2009;kb and 89.19&#x2009;Mb, respectively. The majority (97.94%) of the assembled genome was anchored to the 14 chromosomes by Hi-C analysis. The relatively high genome completeness was observed, with 97.38% (metazoa_odb10 database) and 95.26% (mollusca_odb10 database) in BUSCO analysis. Genome annotation revealed approximately 65.46% of the repeat sequences and 26,315 protein-coding genes. Comparative genome analysis revealed 28 expanded and 48 contracted families (p&#x2009;<&#x2009;0.05) in P. maxima, with 3.2% of genes (894) being species-specific. This chromosome-level genome serves as an essential resource for research in evolutionary genomics, phylogenetics, and biomineralization.

Animals

Chromosomal level genome assembly of medicinal plant Chrysosplenium macrophyllum.

Chrysosplenium macrophyllum Oliv., a perennial herb native to China, is widely used in traditional medicine for its notable therapeutic properties. However, the absence of a reference genome has constrained its full potential for research and application. This study presents the first chromosome-level de novo genome assembly of C. macrophyllum, constructed by integrating long reads from Oxford Nanopore Technologies (ONT), short reads from BGI, and Hi-C data. The final assembly spans 2.55&#x2009;Gb, with a scaffold N50 of 93.38&#x2009;Mb, and 83.70% of the genome has been assigned to 22 chromosomes. The mapping rate of the BGI short reads to the genome is approximately 97.94%, and BUSCO analysis reveals that 97.94% of the predicted genes are complete. A total of 62,921 protein-coding genes were predicted, with functional annotations for 93.67% of them. This chromosome-level genome assembly represents an important resource for expanding our understanding of Chrysosplenium species and supports future genomic studies and applications.

Genome, Plant

Chromosome-level genome assembly of the Vermilion Snapper (Rhomboplites aurorubens).

Vermilion Snapper (Rhomboplites aurorubens, Lutjanidae) inhabits deep waters (20-300&#x2009;m) from North America to Brazil and supports significant commercial and recreational fisheries. Despite its economic importance, the understanding of its basic biology remains limited. Classified as Vulnerable on the Red List due to overfishing, populations have declined by over 30% in recent generations. We assembled and annotated the first chromosome-scale genome of this species by combining PacBio long reads, Illumina short reads, and Hi-C data. The resulting assembly is 987.5 Mbp, with a scaffold N50 size of 41.3 Mbp, and includes 135 contigs clustered and ordered onto 24 chromosomes with 34,496 predicted genes. The high-quality assembly and annotation contained about 98% complete and single-copy BUSCO genes. It is the most complete, chromosome-level genome assembly of an Atlantic snapper to date. The genome assembly and supporting data are valuable tools for ecological and comparative genomics studies of snappers and other valuable commercial species within the family.

Chromosomes

A chromosomal-level genome assembly of Odontolabis cuvera Hope, 1842 (Coleoptera: Lucanidae).

The stag beetle (Coleoptera: Lucanidae) represents a captivating and evolutionarily significant group, regarded as one of the most basal lineages within the superfamily Scarabaeoidea. Despite their importance for studying beetle evolution and ecology, genomic resources for this family remain scarce. Here, we report a chromosome-level genome assembly of Odontolabis cuvera, generated by integrating PacBio HiFi, Illumina, and Hi-C data. The genome assembly spans 908.07&#x2009;Mb, comprising 66 scaffolds (scaffold N50: 65.36&#x2009;Mb) and 147 contigs (contig N50: 16.39&#x2009;Mb). A total of 99.58% (904.22&#x2009;Mb) of the assembly was anchored to 14 chromosomes. BUSCO analysis (insecta_odb10 dataset, n&#x2009;=&#x2009;1,367) demonstrated high completeness, with 99.1% of conserved insect orthologs identified (98.3% single-copy, 0.8% duplicated). Repetitive elements accounted for 53.00% (281.28&#x2009;Mb) of the genome, and a total of 18,332 protein-coding genes were annotated. This high-contiguity genome provides a critical foundation for uncovering the evolutionary mechanisms and ecological adaptations unique to Lucanidae.

Animals

A telomere-to-telomere reference genome assembly of the red silk cotton tree (Bombax ceiba).

Bombax ceiba, an important ornamental tree and potential fiber resource in the textile industry, is widely distributed in tropical and subtropical regions. In this study, we assembled a nearly gap-free telomere-to-telomere (T2T) genome of B. ceiba using Illumina, PacBio High-fidelity (HiFi), ONT ultra-long, and Hi-C sequencing technologies. The genome spanned approximately 807.89&#x2009;Mb, with a scaffold N50 of 16.58&#x2009;Mb, and 754.68&#x2009;Mb (93.41%) of genomic sequences were anchored onto 48 pseudo-chromosomes. Benchmarking Universal Single-Copy Orthologs (BUSCO) analysis revealed a completeness of 99.40%, identifying 1,378 single-copy and 213 duplicated genes out of 1,614. The genome contained 67.72% (547.11&#x2009;Mb) repeat regions, with 39,708 predicted protein-coding genes. Collectively, our study provides valuable genomic data for investigating the evolutionary history of the Malvaceae family.

Genome, Plant

Chromosome-level genome assembly and annotation of the porcupine fish (Diodon hystrix).

The porcupinefish (Diodon hystrix), a coral reef teleost, is widely distributed in tropical/subtropical waters of the Pacific, Atlantic, Indian Oceans, and Mediterranean Sea. It shares easily recognizable features with pufferfish, such as body inflation and spines. Additionally, its culinary value makes D. hystrix a highly desirable species in many tropical coastal regions, with considerable market potential. However, lack of a high-quality genome hindered further studies on its reproduction, molecular biology, and genomic improvement. Here, we assembled the chromosome-scale genome using PacBio HiFi, ultra-long reads, and Hi-C. Of the 713.62&#x2009;Mb genome, 98.63% anchored to 23 chromosomes (scaffold N50: 31.52&#x2009;Mb) with 39.82% repetitive sequences. The assembled genome achieved a BUSCO completeness score of 97.7%, with 23,171 protein-coding genes predicted, 22,221 of which were functionally annotated. Phylogenetic analysis identified D. hystrix's evolutionary relationships with other species in the Tetraodontiformes. In summary, the high-quality genome of D. hystrix sheds light on valuable insights into genome size evolution, and provides a valuable resource for exploiting genomic study and breeding applications in this species.

Animals

Chromosome level genome assembly and full-length transcriptome of blacktip trevally (Caranx heberi).

Caranx heberi (Bennett, 1830) commonly known as the blacktip trevally belongs to the family Carangidae and is a potential brackishwater aquaculture species. However, the limited genomic resources are hindering the efforts to study its genetic traits and their molecular basis. To bridge this gap, we generated a high-quality reference genome employing multiple sequencing strategies including PacBio Hifi reads (135x), Illumina short reads (150x), and Hi-C chromosome conformation capturing (180x). The high-quality genome assembly consisted of 159 scaffolds summing to 618.71&#x2009;Mb and an N50 value of 26.72&#x2009;Mb. Among these, 24 chromosome level scaffolds covered 97.5% of the total assembly. The genome contained 20.94% of repeat elements and 30,354 protein encoding genes. In addition, full-length transcriptomes were generated using the PacBio IsoSeq approach from seven tissues (gill, kidney, liver, muscle, heart, spleen, and intestine). The comprehensive genomic and transcriptomic resources developed in this study will facilitate the domestication and aquaculture development of C. heberi, as well as support research on its nutritional potential, ecological adaptations, and evolutionary biology.

Animals

Chromosome-level genome assembly of the large carpenter bee Xylocopa dejeanii Lepeletier, 1841 (Hymenoptera: Apidae).

Xylocopinae, a diverse bee subfamily comprising over 1,000 bee species, and also a major model system for studying the pollination and evolution of sociality. The lack of chromosome-level genome assembly resources for the Xylocopinae limits our research of their biology and evolution. Here, we provided the first pseudo-chromosomes genome assembly of the Xylocopa dejeanii combined PacBio CLR long reads, Illumina sequences, and Hi-C data. The final genome is 194.44&#x2009;Mb located in 16 chromosomes. Our assembly includes 141 scaffolds, with a scaffold N50 length of 13.15&#x2009;Mb. BUSCO analysis revealed 99.00% completeness. Genome annotation identified 28.27&#x2009;Mb of repetitive elements, 10,970 protein-coding genes, and 432 ncRNAs. This high-quality X. dejeanii assembly advances our understanding of Xylocopinae genomics and provides new insights into bee evolution.

Animals

A chromosome-level assembly of the alpine snow alga Chloromonas typhlos.

Chloromonas typhlos is a cosmopolitan alpine snow alga distributed across continents, and its blooming accelerates snow melting by decreasing the amount of snow albedo. To elucidate the genetic traits underlying the adaptation of C. typhlos to the alpine habitat, we combined PacBio sequencing and Hi-C to generate a high-quality chromosome-level genome assembly (contig N50: 1.29&#x2009;Mb; scaffold N50: 7.23&#x2009;Mb) with 31 chromosomes and a genome size of 200.86&#x2009;Mb. Repetitive elements constituted 11.05% of the genome, and 16,133 protein-coding genes were predicted, of which 82% were functionally annotated. This study provides a set of omics resources both for snow algae and the genus Chloromonas.

Snow

Chromosome-level genome assembly of starry flounder (Platichthys stellatus).

Starry flounder (Platichthys stellatus) is widely distributed along the coastlines of the North Pacific. As an euryhaline flatfish, it can adapt to a wide range of environmental salinity ranging from freshwater to seawater, and is a promising aquaculture flatfish species in Korea and North China. However, no high-quality starry flounder reference genome has been reported to date, which greatly limits the studies of genetics and functional genomics. Here, we obtained a high-quality chromosome-level starry flounder genome assembly with a&#xa0;length of 643.56&#x2009;Mb (scaffold N50: 26.19&#x2009;Mb, contig N50: 10.00&#x2009;Mb) combining short-reads sequencing, PacBio HiFi sequencing, and Hi-C sequencing. Approximately 94.02% of assembled sequences were anchored into 24 pseudochromosomes, and a total of 18 telomeres were detected. Totally 22,835 protein-coding genes and 227.87&#x2009;Mb repetitive sequences were identified. In summary, the high-quality chromosome-level genome assembly not only provides valuable resources for genetic research in starry flounder, but also advances the development of molecular breeding technology of starry flounder.

Animals

Chromosome-level genome assembly of the hemiparasitic Taxillus sutchuenensis (Loranthaceae).

Taxillus sutchuenensis, an ecologically and medicinally important hemiparasitic plant that parasitizes diverse woody hosts, was sequenced to generate a high-quality chromosome-level genome assembly. PacBio HiFi long reads, RNA-seq transcriptome data, and Hi-C data were used to assemble a 406.32&#x2009;Mb genome anchored onto nine pseudo-chromosomes, with a scaffold N50 of 45.59&#x2009;Mb. The assembly showed high completeness and accuracy, supported by BUSCO (93.6%) and Merqury QV (70.6) assessments. The LTR Assembly Index (LAI) of 13.98 indicated excellent continuity. A total of 21,795 protein-coding genes were predicted, with 94.46% functionally annotated. Repetitive sequences accounted for 50.05% of the genome, primarily LTR retrotransposons. This genome provides a valuable resource for investigating the evolution, functional genomics, and parasitic mechanisms of hemiparasitic plants.

Genome, Plant

A high-quality chromosome-level genome assembly and annotation of the giant freshwater prawn (Macrobrachium rosenbergii).

The giant freshwater prawn, Macrobrachium rosenbergii, is native to Southeast Asia and is used in aquacultural practices worldwide. It is considered advantageous because of its rapid growth, high nutritional value, and economic benefits. As one of the three major freshwater aquaculture shrimp sources in China, a high-quality genome resource is of great significance for promoting the germplasm improvement of varieties. This study presents a high-quality chromosome-level genome assembly of M. rosenbergii that was generated by combining PacBio, MGI, and Hi-C reads. The assembled genome was 2.96&#x2009;Gb in size, with a contig N50 of 0.64&#x2009;Mb and a scaffold N50 of 55.76&#x2009;Mb, which was positioned on 59 pseudo-chromosomes. The Benchmarking Universal Single-Copy Orthologs (BUSCO) analysis for genome assembly reached 94.37%. In total, 27,111 protein-coding genes were identified, of which 25,470 were functionally annotated. These results provide a foundation for future research into adaptive evolution, genomics, and molecular breeding in M. rosenbergii.

Animals

Chromosome-level genome assembly of Ampulex clypecomplana Chen & Li (Hymenoptera: Ampulicidae).

Ampulex clypecomplana Chen & Li, 2010 (Hymenoptera: Ampulicidae) is an important predatory insect in Hymenoptera. However, molecular information about this predatory insect is currently limited. In this study, we employed ONT long-read sequencing, MGI-SEQ short-read sequencing, Hi-C sequencing and transcriptomic data to assemble the high-quality genome of A. clypecomplana. The genome assembly length was 338.43&#x2009;Mb, with a Scaffold N50 length of 19.05&#x2009;Mb. Our BUSCO analysis further confirmed the gene coverage completeness of the genome assembly to be 99.2%. Phylogenetic analysis indicated that A. clypecomplana appeared approximately 132 million years ago. We annotated 110.75&#x2009;Mb of repetitive sequences, accounting for 32.72% of the entire genome. In A. clypecomplana, we identified 180 gene expansions and 1029 genes that underwent contraction or loss. The high-quality genome of A. clypecomplana provides a valuable genetic resource for future research in evolution, molecular biology, and applied studies.

Animals

A chromosome-level genome of the Nicobar pigeon, Caloenas nicobarica.

The Nicobar pigeon (Caloenas nicobarica), the closest living relative of the extinct Dodo (Raphus cucullatus), is endemic to Southeast Asia with a fragmented distribution across numerous small islands. It suffers from habitat loss, hunting, and predation from invasive species, resulting in its classification as Near Threatened by the International Union for the Conservation of Nature. We have generated a haplotype-resolved and chromosome-level genome assembly of the Nicobar pigeon using a combination of PacBio HiFi long-read sequencing and Arima Hi-C chromatin interaction mapping. This assembly includes two haplotypes, each spanning approximately 1.2 Gb. Haplotype 1 has a contig N50 of 25.2&#xa0;Mb and a scaffold N50 of 79.7&#xa0;Mb, whereas haplotype 2 has a contig N50 of 24.7&#xa0;Mb and a scaffold N50 of 107.9&#xa0;Mb. As the first high-quality genome assembly of any bird in the Columbidae Indo-Pacific clade, this resource provides valuable insights for phylogenetic studies. Furthermore, the phylogenetic proximity of the Nicobar pigeon to the Dodo (R. cucullatus) and the Rodrigues Solitaire (Pezophaps solitaria) offers a unique opportunity to study these extinct species, making this assembly a critical resource for evolutionary studies. It also offers a unique model for studying genetic diversity, adaptation, and speciation in island environments. This genomic resource will not only enhance our understanding of the evolutionary history of the Nicobar pigeon but also serve as a valuable tool for future conservation efforts aimed at preserving this unique species and its fragile island ecosystem.

Animals

A chromosome-scale assembly for the genome of southern corn rootworm, Diabrotica undecimpunctata.

Diabrotica undecimpunctata ssp. howardi, the southern corn rootworm or eastern 12-spotted cucumber beetle, is a generalist insect herbivore that causes damage and yield loss to several crops in North America including maize. Unresolved phylogenetic relationships within and among D. undecimpunctata subspecies are impacting current quarantine policies. We report the chromosome-level haploid genome assembly, icDiaUnde3, constructed using HiFi and Hi-C read data from a single male D. undecimpunctata collected and identified as subspecies howardi based on geographic location and morphology. The primary 1.74 Gbp assembly is scaffolded into 11 chromosome-length scaffolds representing 9 autosomes, a single X chromosome and a supernumerary (B) chromosome (scaffold N50&#x2009;=&#x2009;162.8 Mb and L50&#x2009;=&#x2009;5). Ab initio and evidence-based structural reference sequence (RefSeq) annotations predicted 18,959 protein-coding genes, in which 99.2% of the 1,367 Benchmark Universal Single-Copy Orthologs from Insecta were complete. Repeat elements occupy 1.26 Gbp (72.33%) of the icDiaUnde3 assembly, with nearly 36% predicted to be retroelements. Alignment of whole chromosomes from icDiaUnde3 with those previously assembled from Diabrotica spp. predicted 2 and 6 autosomal inversions with D. balteata and D. virgifera virgifera, respectively. The mitochondrial genome had an annotated gene order and orientation conserved among beetles. The icDiaUnde3 reference genome assembly is a vital resource for taxonomic, comparative, and functional studies to enhance sustainable crop production.

agriculture

Chromosome-Level Reference Genome of the Desert Night Lizard Xantusia vigilis.

We present a reference-quality genome assembly for the desert night lizard (Xantusia vigilis). The night lizards (Xantusiidae) are a family of small-bodied lizards found in North America (Xantusia), Central America (Lepidophyma), and Cuba (Cricosaura). The night lizard family has an independent evolutionary history of at least 80 million years from its sister taxa within Scincoidea. The Xantusiids have several unique ecological, behavioral and evolutionary characteristics. For instance, the family contains the only squamate species that form diploid, unisexual, parthenogenic lineages. In addition, most night lizards are viviparous and form stable kin groups that are maintained over multiple years, an unusual life history strategy among lizards. Combining PacBio long-read sequencing, Hi-C, and RNAseq data we developed a reference-quality genome for the desert night lizard, X. vigilis. We assembled a complete mitochondrion and&#x2009;~&#x2009;2.2 Gb nuclear genome, with 20 scaffolds that correlate in size to the X. vigilis karyotype. In addition, we found that X. vigilis chromosome 1 aligns with gene content of both of macrochromosome 1 and microchromosome 9 from a genome assembly of a species in the sister family Cordylidae (Hemicordylus capensis).

Xantusia