PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “PacBio HiFi”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Accurate somatic small variant discovery for multiple sequencing technologies with DeepSomatic.

Somatic variant detection is an integral part of cancer genomics analysis. While most methods have focused on short-read sequencing, long-read technologies offer potential advantages in repeat mapping and variant phasing. We present DeepSomatic, a deep-learning method for detecting somatic small nucleotide variations and insertions and deletions from both short-read and long-read data. The method has modes for whole-genome and whole-exome sequencing and can run on tumor-normal, tumor-only and formalin-fixed paraffin-embedded samples. To train DeepSomatic and help address the dearth of publicly available training and benchmarking data for somatic variant detection, we generated and make openly available the Cancer Standards Long-read Evaluation (CASTLE) dataset of six matched tumor-normal cell line pairs whole-genome sequenced with Illumina, PacBio HiFi and Oxford Nanopore Technologies, along with benchmark variant sets. Across samples, both cell line and patient-derived, and across short-read and long-read sequencing technologies, DeepSomatic consistently outperforms existing callers.

Humans↗

Chromosome-level genome assembly of Triplophysa scleroptera.

Triplophysa scleroptera is an endemic fish species in Qinghai Lake and the upper reaches of the Yellow River. However, studies on conservation and evolutionary genetics were seriously impeded by the absence of a reference genome. Here, by using PacBio HiFi sequencing and Hi-C assembly technology, we assembled a chromosome-level genome of T. scleroptera, with a total length of 660.22 Mb and 99.82% of the sequence anchored to 25 chromosomes. The contig N50 and scaffold N50 were 9.09 Mb and 24.38 Mb, respectively. The evaluation using BUSCO indicated the genome assembly to be 96.40% complete. About 33.41% of the genome consists of repeat elements. We predicted 26,168 protein-coding genes in the genome, and 99.02% of them were functionally annotated. This high-quality reference genome would serve as a valuable genomic resource for advancing evolutionary conservation genetics studies in this species.

Animals↗

Chromosome-level genome assembly with telomeric repeats at scaffold ends for Rhabdosargus sarba.

Rhabdosargus sarba, the goldlined seabream, is a euryhaline marine fish of great aquaculture potential. Genome sequencing and assembly of R. sarba was carried utilizing a multi-platform sequencing strategy that included long-read sequencing (PacBio HiFi), short-read sequencing (Illumina), and chromatin interaction mapping (Hi-C). The final genome assembly size after scaffolding was 764.59 Mb in 31 scaffolds with an N50 length of 33.98 Mb. Repeat profiling of primary assembly showed that 28.71% of the genome comprises of repeat elements. Gene prediction utilising the evidence from ab initio prediction and transcriptome data revealed 26,913 protein encoding genes and functional annotation and pathway analysis showed their participation in 332 pathways. This genome is an excellent resource for future research on genetic improvement and molecular breeding programmes for R. sarba.

Animals↗

The first chromosome-level genome of the lappet moth Trabala vishnou (Lepidoptera: Lasiocampidae).

Trabala vishnou (Lefèbvre, 1827) (Lepidoptera: Lasiocampidae) is a destructive leaf-eating pest that causes severe damage to forest ecosystems, leading to substantial economic losses. Herein, we sequenced and assembled a high-quality chromosome-level genome of T. vishnou using a combination of Illumina reads, PacBio HiFi reads, and High throughput Chromosome Conformation Capture (Hi-C) technologies. The genome size is 561.86 Mb and spans 25 chromosomes, exhibiting a high level of contiguity (scaffold/contig N50 = 21.75 Mb/20.67 Mb). Benchmarking Universal Single-Copy Orthologs (BUSCO) analysis a 99.5% completeness score for this genome assembly. Repeat elements constitute 62.66% of the genome. A total of 1,630 non-coding RNAs and 12,895 protein-coding genes have been identified within the genome. The first chromosome-level genome of T. vishnou serves as a valuable reference for elucidating the evolution of functional traits in Lasiocampidae family and will facilitate the development of strategies for controlling defoliating pests.

Animals↗

Chromosome-level genome assembly and annotation of Spinibarbus caldwelli.

Spinibarbus caldwelli is an economically important freshwater species within the Cyprinidae family, abundant in the middle and lower reaches of the Yangtze River and its adjacent basins. As a promising species suitable for aquaculture in southern China, the lack of genomic resources has hampered the genetic breeding and conservation. Here, we release a chromosome-level genome assembly for S. caldwelli using PacBio HiFi long-reads, Illumina short-reads, and Hi-C sequencing data. The final genome assembly is 1.77 Gb in size, with a contig N50 of 24.27 Mb. Using Hi-C scaffolding, 99.14% of the contigs were successfully anchored to 50 chromosomes, resulting in a scaffold N50 of 35.29 Mb. The final genome assembly shows a BUSCO completeness of 98.27%. The assembled genome contains 49.41% repetitive sequences and 51,505 predicted genes, 90.83% of which have been functionally annotated. This genome provides a genetic basis for S. caldwelli, facilitating the exploration of Cyprinid phylogeny, genetic improvement, and conservation efforts.

Animals↗

A telomere-to-telomere gap-free genome assembly of the endangered humphead wrasse (Cheilinus undulatus).

Humphead wrasse, Cheilinus undulatus, is an endangered fish species with high economic and ecological value as well as natural sex change from female to male, while sexual selection occurs in breeding aggregations. In our present study, we constructed the first gap-free telomere-to-telomere (T2T) genome assembly for humphead wrasse, by integration of PacBio HiFi, ONT Ultra-long and Hi-C sequencing techniques. With 99% of the entire sequences anchored into 24 chromosomes, this haplotypic genome assembly spans approximately 1.25 Gb and presents a complete set of 48 telomeres and 24 centromeres. In terms of correctness (quality value QV: 53.447) and completeness (BUSCO score: 99.3%), this chromosome-scale assembly is indeed of high quality. We predicted 658.03 Mb of repetitive sequences and annotated 26,609 protein-coding genes in the assembled genome. This high-quality T2T genome assembly not only facilitates the genetic conservation of humphead wrasse, but also offers fundamental genomic data for supporting in-depth investigations on functional genomics, genetic diversity, and selective breeding for this economically important teleost.

Animals↗

A chromosome-level reference genome assembly of the Small snakehead (Channa asiatica).

The Small snakehead (Channa asiatica) is an economically important species in both aquaculture and ornamental trade, mainly distributed in South China and Southeast Asia. Despite its significance, limited genomic resources have impeded in-depth genetic studies and breeding programs. In this study, we used PacBio HiFi long-read sequencing, Illumina short-read sequencing, and Hi-C technologies to generate a high-quality chromosome-level genome of the C. asiatica. The final genome spans 659.44 Mb, with an impressive 98.18% anchored to 23 chromosomes. Notably, the contig N50 and scaffold N50 are 23.92 Mb and 29.61 Mb, validated by a BUSCO completeness score of 98.93%. Genome annotation identified 26,603 protein-coding genes, 99.29% of which were confirmed by BUSCO analysis, and 93.68% were functionally annotated. Approximately 27.72% of the genome sequences were classified as repeat elements. This high-fidelity genome assembly provides a robust foundation for advancing molecular breeding, comparative genomics, and evolutionary studies of C. asiatica and related species.

Animals↗

A chromosome-level genome assembly and annotation of Cercis chuniana (Fabaceae).

The genus Cercis L., at the base of the subfamily Cercidoideae of Fabaceae, is known for its ecological adaptability and significant medicinal, ornamental, and economic value. However, the lack of a high-quality genome hinders the understanding of the evolution of Cercis and Fabaceae. In this study, we present a chromosome-level genome of Cercis chuniana by combining Illumina short reads, PacBio HiFi long reads, and Hi-C data. The final genome size is 355.53 Mb, consisting of 12 contigs with a N50 of 42.34 Mb. Notably, 344.24 Mb, corresponding to 96.82% of the genome, was anchored to seven chromosomes. The assembly comprises 24.83% repetitive sequences, including 19.32% long terminal repeats. Additionally, a total of 33,837 protein-coding genes were predicted in the genome, with 32,709 (96.67%) genes successfully annotated. The high-quality genome assembly of C. chuniana not only bridges the existing gap in genomic data and offers important resources for molecular studies of this species, but also provides essential insights for future studies on speciation, functional and comparative genomics within the Fabaceae family.

Genome, Plant↗

A chromosomal level genome assembly of Nguni Sheep, Ovis aries.

Nguni sheep (Ovis aries) are indigenous to the Southern Africa region and common within the smallholder and poor resources farming systems. They are well adapted to different agroecological regions. However, limited genomic resources such as high-quality reference genomes have hindered our understanding of its adaptation and establishment of an effective breeding program. To address this, we assembled a chromosomal-level genome of Nguni sheep using a combination of PacBio HiFi reads and Omni-C reads. The genome size was estimated to be 2.9 Gb with a contig/scaffold N50 74 Mb and 99.6 Mb and a genome completeness of 96.1%, as estimated by the Benchmarking Universal Single-Copy Orthologs (BUSCO) program. The final genome encompassed a total of 25,926 protein-coding genes. The findings of this study provide a valuable genomic resource for understanding the adaptability of the Nguni sheep and the establishment of effective breeding programs.

Animals↗

Chromosome-level genome assembly of starry flounder (Platichthys stellatus).

Starry flounder (Platichthys stellatus) is widely distributed along the coastlines of the North Pacific. As an euryhaline flatfish, it can adapt to a wide range of environmental salinity ranging from freshwater to seawater, and is a promising aquaculture flatfish species in Korea and North China. However, no high-quality starry flounder reference genome has been reported to date, which greatly limits the studies of genetics and functional genomics. Here, we obtained a high-quality chromosome-level starry flounder genome assembly with a length of 643.56 Mb (scaffold N50: 26.19 Mb, contig N50: 10.00 Mb) combining short-reads sequencing, PacBio HiFi sequencing, and Hi-C sequencing. Approximately 94.02% of assembled sequences were anchored into 24 pseudochromosomes, and a total of 18 telomeres were detected. Totally 22,835 protein-coding genes and 227.87 Mb repetitive sequences were identified. In summary, the high-quality chromosome-level genome assembly not only provides valuable resources for genetic research in starry flounder, but also advances the development of molecular breeding technology of starry flounder.

Animals↗

Near-complete reference genome assembly of Hoya carnosa.

Hoya R. Br. is the largest genus in the tribe Marsdenieae (Apocynaceae), comprising 350-450 species. Hoya species are popular in horticulture for their distinctive floral traits and fragrances, primarily sourced from domestication and mutation breeding. However, the lack of molecular analysis for floral morphological traits has limited their cultivation and application. In this study, we assembled a near-complete reference genome for H. carnosa, the model species of the genus, using PacBio HiFi reads and Hi-C method. The genome size was approximately 465.7 Mb with a contig N50 of 39.3 Mb. 99.7% of the sequences were anchored to 11 pseudochromosomes, and the assembly achieved a BUSCO score of 98.5%. We predicted 24,309 protein-coding genes, of which 90.2% (21,927) were functionally annotated. This high-quality genome provides a valuable reference for the research of evolution, conservation and molecular breeding in Hoya.

Genome, Plant↗

A chromosome-level genome assembly of Guimi No. 2 (Actinidia chinensis).

In this study, we report a high-quality chromosome-level genome assembly of Actinidia chinensis var. chinensis 'Guimi No. 2'. This cultivar, discovered in Guizhou karst ecosystems, exhibits resistance to Pseudomonas syringae pv. actinidiae (Psa). Using a combination of MGI short-read sequencing, PacBio HiFi long-read sequencing, and Hi-C technology, we generated a genome assembly of 608.43 Mb with a contig N50 of 20.70 Mb, and 99.70% of the assembly was successfully anchored onto 29 pseudochromosomes. The quality value (QV) and the LTR Assembly Index (LAI) of the assembled genome were 72.23 and 10.10. The BUSCO analysis indicated that the genome assembly and gene model prediction were 98.40% and 96.56% complete, respectively. A total of 251.15 Mb of repetitive sequences and 45,986 protein-coding genes were annotated. This genome assembly provides critical insights into A. chinensis's genomic architecture and serves as a foundational resource for elucidating disease resistance mechanisms against Psa, while enabling comparative phylogenomic studies across the Actinidia genus.

Actinidia↗

A chromosomal-level genome assembly of Odontolabis cuvera Hope, 1842 (Coleoptera: Lucanidae).

The stag beetle (Coleoptera: Lucanidae) represents a captivating and evolutionarily significant group, regarded as one of the most basal lineages within the superfamily Scarabaeoidea. Despite their importance for studying beetle evolution and ecology, genomic resources for this family remain scarce. Here, we report a chromosome-level genome assembly of Odontolabis cuvera, generated by integrating PacBio HiFi, Illumina, and Hi-C data. The genome assembly spans 908.07 Mb, comprising 66 scaffolds (scaffold N50: 65.36 Mb) and 147 contigs (contig N50: 16.39 Mb). A total of 99.58% (904.22 Mb) of the assembly was anchored to 14 chromosomes. BUSCO analysis (insecta_odb10 dataset, n = 1,367) demonstrated high completeness, with 99.1% of conserved insect orthologs identified (98.3% single-copy, 0.8% duplicated). Repetitive elements accounted for 53.00% (281.28 Mb) of the genome, and a total of 18,332 protein-coding genes were annotated. This high-contiguity genome provides a critical foundation for uncovering the evolutionary mechanisms and ecological adaptations unique to Lucanidae.

Animals↗

Chromosome-level genome assembly of Qihe gibel carp.

Qihe gibel carp (Carassius gibelio var. Qihe) is a local population of natural gynogenetic amphitriploid (AAABBB) Carassius gibelio, and has high nutritional and economic value. In this study, we assemble a high-quality chromosome-level genome of Qihe gibel carp through DNBSEQ, PacBio HiFi, and Hi-C sequencing data. The resulting assembly consisted of 350 contigs with the full length of 1.607 Gb and 96.21% (1.515 Gb) of the assembled genome was successfully anchored to 50 chromosomes, with a contig N50 of 28.97 Mb and a scaffold N50 of 29.84 Mb. Repeated sequences accounting for 43.72% (732.494 Mb) of the total were also identified, and gene prediction revealed 46,131 protein-coding genes with an annotation ratio of 96.48%. Furthermore, Benchmarking Universal Single-Copy Orthologue (BUSCO) analysis demonstrated that the genome assembly achieved high completeness, with a score of 97.66%. This high-quality chromosome-level genome lays the foundation for molecular biology research as well as molecular breeding and evolutionary studies of Qihe gibel carp in the future.

Animals↗

A chromosome-level genome assembly of Coffea arabica L. var. 'Kona Typica'.

Coffea arabica L. var. 'Kona Typica' is renowned for its premium cup quality, but its vulnerability to pests and diseases limits production. To accelerate cultivar improvement, we generated a chromosome-level genome assembly of 'Kona Typica' using PacBio HiFi sequencing and Hi-C scaffolding technology. The final assembly spans 1.13 Gb, with a scaffold N50 of 50.50 Mb, organized into 22 chromosomes. BUSCO assessment indicated a high completeness at 99.1%. We annotated 65,458 protein-coding genes and identified 1,073,545 interspersed repeats, accounting for 65.16% of the genome. Analysis of transposon insertion ages revealed that most long terminal repeat retrotransposons proliferated after the polyploidization event. This high-quality genome assembly of 'Kona Typica' provides a valuable resource for exploring coffee genomic evolution and genetic mechanisms of complex traits, facilitating genomics studies and the development of improved coffee cultivars with enhanced disease resistance and quality traits.

Coffea↗

Chromosome-level haplotype-resolved genome assembly of the giant honeycomb oyster, Hyotissa hyotis.

The giant honeycomb oyster, Hyotissa hyotis, a common bivalve inhabitant of tropical and subtropical coastal waters, holds significant ecological and economic importance due to its shell characteristics, rapid growth, and high-quality adductor muscle. However, the lack of high-quality genome has impeded the genetic study and artificial breeding of this species. In this study, we provided the first chromosomal-level haplotype-resolved assembly for the H. hyotis (2n = 20) by combining PacBio HiFi long-read and Hi-C sequencing. We obtained a haplotype-resolved assembly of 3.39 Gb in size, of which 96.69% were anchored to 20 chromosomes. The haplotype A and B genome (HapA and HapB) was 1,639.90 and 1,643.23 Mb in size, respectively. Accordingly, a total of 28,720 and 29,003 protein-coding genes were annotated from HapA and HapB. Through the BUSCO evaluation, the assembly and annotation results exhibited the completeness value of 94.65% and 94.03% for HapA, while 94.13% and 92.98% for HapB. This high-quality genome assembly provides valuable resource for further genetic studies and genetic improvement of the group of oysters.

Animals↗

Chromosomal-level genome assembly of Trypanosoma carassii, the etiologic agent of a recent outbreak of trypanosomiasis in cage-cultured large yellow croaker (Larimichthys crocea) in China.

Trypanosoma carassii, a typical freshwater fish trypanosome, has recently been identified as the etiological agent of a trypanosomiasis outbreak in cage-cultured large yellow croaker (Larimichthys crocea) in China and has been designated as T. c. larimichthys. To date, publicly available genomic data for trypanosomes have been limited to terrestrial species, particularly those of medical importance. Here, we present a chromosome-level genome assembly of T. carassii, the first genome of an aquatic trypanosome, generated using PacBio HiFi long-read sequencing and Hi-C scaffolding technologies. A preliminary genome survey based on Illumina sequencing data estimated the genome size at 56.38 Mb with a heterozygosity of 1.17%. The final assembled genome spans 48.55 Mb, with contig N50 and scaffold N50 values of 139.15 Kb, and achieves 100.00% BUSCO completeness. Hi-C data resolved the assembly into 34 chromosomes and 9 unanchored scaffolds. Repetitive elements account for 53.29% of the genome (approximately 25.87 Mb). A total of 11,584 protein-coding genes were predicted, 95.36% of which were functionally annotated. Synonymous substitution rates analysis of paralogous genes indicates a recent burst of gene duplication, which likely corresponds to a whole-genome duplications. This high-quality genome assembly provides invaluable resources for understanding the evolution and host adaptation of aquatic trypanosomes.

Animals↗

Chromosome-level genome assembly and annotation of the porcupine fish (Diodon hystrix).

The porcupinefish (Diodon hystrix), a coral reef teleost, is widely distributed in tropical/subtropical waters of the Pacific, Atlantic, Indian Oceans, and Mediterranean Sea. It shares easily recognizable features with pufferfish, such as body inflation and spines. Additionally, its culinary value makes D. hystrix a highly desirable species in many tropical coastal regions, with considerable market potential. However, lack of a high-quality genome hindered further studies on its reproduction, molecular biology, and genomic improvement. Here, we assembled the chromosome-scale genome using PacBio HiFi, ultra-long reads, and Hi-C. Of the 713.62 Mb genome, 98.63% anchored to 23 chromosomes (scaffold N50: 31.52 Mb) with 39.82% repetitive sequences. The assembled genome achieved a BUSCO completeness score of 97.7%, with 23,171 protein-coding genes predicted, 22,221 of which were functionally annotated. Phylogenetic analysis identified D. hystrix's evolutionary relationships with other species in the Tetraodontiformes. In summary, the high-quality genome of D. hystrix sheds light on valuable insights into genome size evolution, and provides a valuable resource for exploiting genomic study and breeding applications in this species.

Animals↗