PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “HiFi sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

A telomere-to-telomere reference genome assembly of the red silk cotton tree (Bombax ceiba).

Bombax ceiba, an important ornamental tree and potential fiber resource in the textile industry, is widely distributed in tropical and subtropical regions. In this study, we assembled a nearly gap-free telomere-to-telomere (T2T) genome of B. ceiba using Illumina, PacBio High-fidelity (HiFi), ONT ultra-long, and Hi-C sequencing technologies. The genome spanned approximately 807.89 Mb, with a scaffold N50 of 16.58 Mb, and 754.68 Mb (93.41%) of genomic sequences were anchored onto 48 pseudo-chromosomes. Benchmarking Universal Single-Copy Orthologs (BUSCO) analysis revealed a completeness of 99.40%, identifying 1,378 single-copy and 213 duplicated genes out of 1,614. The genome contained 67.72% (547.11 Mb) repeat regions, with 39,708 predicted protein-coding genes. Collectively, our study provides valuable genomic data for investigating the evolutionary history of the Malvaceae family.

Genome, Plant↗

Chromosome-level genome assembly of the horned turban snail Turbo cornutus.

The horned turban snail (Turbo cornutus) is an ecologically and economically important herbivorous gastropod inhabiting nearshore rocky reef habitats. T. cornutus represents a valuable coastal fishery resource in East Asia. Here, we present a chromosome-level genome assembly for T. cornutus generated using a combination of PacBio HiFi long-read and Illumina short-read sequencing and Hi-C scaffolding. The assembled genome spanned 1.93 Gb and was organized into 18 pseudo-chromosomes, representing 99.50% of the total assembly. The contig and scaffold N50 lengths were 41.02 Mb and 104.01 Mb, respectively, with repeat sequences constituting 59.07% of the genome. A total of 28,920 protein-coding genes were predicted, and genome completeness was assessed at 99.3% using the BUSCO mollusca_odb12 dataset. This chromosome-level genome assembly provides a reference for future studies on the biology of T. cornutus, the organization of the gastropod genome, and comparative genomics.

Animals↗

The UTRs of Leishmania donovani vary in length and are enriched in potential regulatory structures.

Leishmania spp. regulate gene expression largely post-transcriptionally, yet untranslated regions (UTRs) remain poorly delineated. We generated high-quality genome and transcriptome datasets for Leishmania donovani strain 1S2D (Ld1S) by combining PacBio HiFi de novo assembly with Oxford Nanopore direct RNA sequencing of promastigotes and axenic amastigotes. The genome assembly consists of 65 scaffolds totaling ~33.3 Mb. Structural comparisons to LdBPK282A1 revealed numerous rearrangements, including some reshuffling genes among polycistronic transcription units and validated by polycistronic reads from RNA sequencing. Promastigote and amastigote RNA sequencing produced 469,010 and 46,729 monocistronic reads containing a spliced-leader and a polyA tail sequences, defining 8,479 transcripts and supporting 7,415 of the 7,969 annotated protein coding genes, as well as 604 putative long non-coding RNAs. We annotated UTRs for 4,921 genes and observed that putative RNA G-quadruplexes were markedly enriched in UTRs. We also noted that 31.9% and 11.5% were expressed into multiple isoforms in promastigotes and amastigotes, respectively. Collectively, these data provide a comprehensive annotation of L. donovani genes and their UTRs and reveal widespread and stage-specific UTR length polymorphisms, and, overall, points to an important role of 3' UTR in post-transcriptional regulation in L. donovani.

Journal Article↗

Chromosome-Level Genome Assembly and Annotation of the Chinese Lizard Gudgeon (Saurogobio dabryi).

The Chinese lizard gudgeon (Saurogobio dabryi) is an economically important freshwater species within the Cyprinidae family, abundant in the middle and lower reaches of the Yangtze River and its adjacent basins. As a promising species suitable for aquaculture in China, the lack of genomic resources has rendered the genetic breeding and conservation research. Here, we present the first chromosome-level genome assembly of S. dabryi using PacBio HiFi long reads, short reads, and Hi-C sequencing data. The final assembly reaches a total size of 1.09 Gb and Hi-C scaffolding anchors 99.55% of the assembled contigs onto 25 chromosomes, with a scaffold N50 reaching 43.15 Mb. The final genome assembly shows a BUSCO completeness of 98.39%. We annotated 659.55 Mb repetitive sequences and 26,036 protein-coding genes, 99.47% of which are functionally annotated. Comparative phylogenomic analysis clarifies the phylogenetic position of Saurogobio within Gobioninae. This high-quality genome provides a critical genetic basis for exploring cyprinid phylogeny, benthic adaptive evolution, genetic improvement, and conservation efforts of S. dabryi.

Saurogobio dabryi↗

Chromosome-level genome assembly of Manglietia pachyphylla.

Manglietia pachyphylla, an endangered evergreen tree within the Magnoliaceae family, is renowned for its exceptional ornamental value in landscape horticulture. Despite its classification as a Category II nationally protected plant species in China, the genetic basis of its adaptive traits and conservation priorities remains poorly understood. To address this, we present the first chromosome-scale genome assembly of M. pachyphylla utilizing an integrated approach combining PacBio HiFi long-read and Hi-C chromosome conformation capture sequencing technologies. The assembled genome spans 2.15 Gb (contig N50 = 43.57 Mb), exhibiting a heterozygosity rate of 0.78% and repeat content of 78.64%, predominantly comprising long terminal repeat (LTR) retrotransposons (52.86%). Hi-C scaffolding anchored 99.57% of the assembly to 19 pseudochromosomes, achieving a BUSCO completeness score of 96.4%. Annotation revealed 42,505 putative protein-coding genes, with 84.46% of predicted genes were functionally annotated. Phylogenomic analysis positioned M. pachyphylla and Oyama sieboldii clustered together in a well-supported group. This high-contiguity genome assembly enables future investigations into adaptive evolution, functional genomics, and evidence-based conservation strategies for this endangered species.

Chromosomes, Plant↗

Chromosome-level genome assembly of the hemiparasitic Taxillus sutchuenensis (Loranthaceae).

Taxillus sutchuenensis, an ecologically and medicinally important hemiparasitic plant that parasitizes diverse woody hosts, was sequenced to generate a high-quality chromosome-level genome assembly. PacBio HiFi long reads, RNA-seq transcriptome data, and Hi-C data were used to assemble a 406.32 Mb genome anchored onto nine pseudo-chromosomes, with a scaffold N50 of 45.59 Mb. The assembly showed high completeness and accuracy, supported by BUSCO (93.6%) and Merqury QV (70.6) assessments. The LTR Assembly Index (LAI) of 13.98 indicated excellent continuity. A total of 21,795 protein-coding genes were predicted, with 94.46% functionally annotated. Repetitive sequences accounted for 50.05% of the genome, primarily LTR retrotransposons. This genome provides a valuable resource for investigating the evolution, functional genomics, and parasitic mechanisms of hemiparasitic plants.

Genome, Plant↗

Chromosome-level genome assembly and annotation of Spinibarbus caldwelli.

Spinibarbus caldwelli is an economically important freshwater species within the Cyprinidae family, abundant in the middle and lower reaches of the Yangtze River and its adjacent basins. As a promising species suitable for aquaculture in southern China, the lack of genomic resources has hampered the genetic breeding and conservation. Here, we release a chromosome-level genome assembly for S. caldwelli using PacBio HiFi long-reads, Illumina short-reads, and Hi-C sequencing data. The final genome assembly is 1.77 Gb in size, with a contig N50 of 24.27 Mb. Using Hi-C scaffolding, 99.14% of the contigs were successfully anchored to 50 chromosomes, resulting in a scaffold N50 of 35.29 Mb. The final genome assembly shows a BUSCO completeness of 98.27%. The assembled genome contains 49.41% repetitive sequences and 51,505 predicted genes, 90.83% of which have been functionally annotated. This genome provides a genetic basis for S. caldwelli, facilitating the exploration of Cyprinid phylogeny, genetic improvement, and conservation efforts.

Animals↗

Long-read low-pass sequencing enhances variant detection in a peanut MAGIC population.

Accurate genotyping accelerates crop improvement, yet long-read sequencing remains underused in breeding due to cost. We present a scalable long-read low-pass (LRLP) sequencing framework for high-throughput variant discovery and trait mapping. Using PacBio HiFi reads in an allotetraploid peanut (Arachis hypogaea; AABB, 2n = 4x = 40) MAGIC population, we generated both LRLP and short-read low-pass (SRLP) data. At comparable depths, LRLP achieved substantially greater whole-genome and gene-space coverage than SRLP. Data were analyzed using both a single-reference genome and an 18-parent pangenome graph constructed with KhufuPan, a new tool for graph-based genotyping. Across analytical approaches, LRLP consistently identified more SNPs, indels (2-1,000 bp), and structural variants (>1 kb) than SRLP, improving genotype resolution and selection accuracy, particularly for large structural variants. By reducing cost barriers and increasing variant discovery in complex genomes, LRLP provides a practical path for deploying advanced genomics in under-resourced and orphan crops critical to global food security.

Arachis↗

Long-read sequencing reveals putatively mobilizable resistance genes and multi-drug resistance plasmids underestimated by short-read metagenomics.

While shotgun metagenomics is often used to profile antibiotic resistome in gut microbial communities, few studies have investigated if the choice of sequencing platform and assembly strategy affect what mobile genetic elements and antimicrobial resistance genes are recovered. In this study, we compared three platforms (Illumina, Oxford Nanopore, and PacBio HiFi) and seven assembly strategies on gut metagenomes from cattle, pig, and human as case studies. Long-read assemblies recovered 5- to 7-fold more plasmid sequence than Illumina in cattle and pig (mean 17.0 Mb vs. 3.1 Mb), while Illumina performed comparably in the less diverse human gut where high per-species coverage enabled effective short-read plasmid assembly. Long reads also detected more resistance genes on plasmid contigs. Hybrid assembly results depended on the algorithm: scaffolding-based OPERA-MS preserved long-read contiguity and recovered more plasmid-borne resistance genes, while the short-read-centric metaSPAdes hybrid mode produced fragmented assemblies. After collapsing haplotype redundancy, PacBio HiFi identified 2 and 49 unique multi-drug resistance plasmid lineages in cattle and pig, respectively. On the other hand, only 2 and 4 were identified from Illumina. Long reads also placed far more ARGs in a putative mobilization context (50-73%) compared to 14-21% for short reads. Platform and assembly strategy are thus key variables in mobilome and resistome characterization and should be accounted for in antimicrobial resistance surveillance.

Animals↗

In vitro mutational spectrum of cyclopenta[cd]pyrene in the human HPRT gene.

Cyclopenta[cd]pyrene (CPP) is a widely distributed polycyclic aromatic hydrocarbon with potent mutagenic and carcinogenic activity. In order to acquire an understanding of the mutagenic pathways of CPP, we studied mutations induced by this chemical in human cells. Four independent cultures of a human cell line expressing cytochrome P450 CYP1A1 (cell line MCL-5) were treated with CPP, and mutants at the hypoxanthine phosphoribosyltransferase (HPRT) locus were selected en masse by 6-thioguanine (6TG) resistance. The kinds and positions of the mutations were analyzed using the combination of high-fidelity polymerase chain reaction (hifi-PCR) and denaturing gradient gel electrophoresis (DGGE). The third exon of the HPRT gene was amplified from the 6TG-resistant cells using the hifi-PCR and the amplified fragment was subsequently analyzed by DGGE to separate mutant sequences from the wild-type sequence. Mutant bands were excised from the gel, amplified using PCR and sequenced. Sixteen different mutations were identified and consisted mostly of the G to T and A to T transversions. Other mutations identified included G to A and A to G transitions, a G to C transversion, and a single G deletion. Of these mutations, six occurred within a run of six guanines. The predominance of transversions involving a guanine or an adenine observed with CPP is similar to the data previously reported for the racemic mixtures of benzo[a]pyrene (B[a]P), suggesting that the mechanisms of mutation induced by CPP may be similar to those induced by B[a]P.

Base Sequence↗

The complete chloroplast genome sequence and phylogenetic analysis of Amorphophallus gigas.

We sequenced the complete chloroplast genome of Amorphophallus gigas, a perennial monocotyledonous herb in Araceae, using HiFi technology. The genome is 173,034 bp in length with a GC content of 34.96%. It exhibits a typical quadripartite structure: a large single-copy (LSC) region of 95,283 bp, a small single-copy (SSC) region of 15,675 bp, and a pair of inverted repeat (IR) regions of 31,038 bp each. It encodes 130 genes (85 protein-coding, 37 tRNA, 8 rRNA). Phylogenetic analysis revealed that A. gigas is closely related to A. titanum, forming a distinct clade. This study provides valuable genomic resources for understanding the evolution of Amorphophallus and Araceae.

Complete chloroplast genome↗

Applications of constant denaturant capillary electrophoresis/high-fidelity polymerase chain reaction to human genetic analysis.

Constant denaturant capillary electrophoresis (CDCE) permits high-resolution separation of single-base variations occurring in an approximately 100 bp isomelting DNA sequence based on their differential melting temperatures. By coupling CDCE for highly efficient enrichment of mutants with high-fidelity polymerase chain reaction (hifi PCR), we have developed an analytical approach to detecting point mutations at frequencies equal to or greater than 10(-6) in human genomic DNA. In this article, we present several applications of this approach in human genetic studies. We have measured the point mutational spectra of a 100 bp mitochondrial DNA sequence in human tissues and cultured cells. The observations have led to the conclusion that the primary causes of mutation in human mitochondrial DNA are spontaneous in origin. In the course of studying the mitochondrial somatic mutations, we have also identified several nuclear pseudogenes homologous to the analyzed mitochondrial DNA fragment. Recently, through developments of the means to isolate the desired target sequences from bulk genomic DNA and to increase the loading capacity of CDCE, we have extended the CDCE/hifi PCR approach to study a chemically induced mutational spectrum in a single-copy nuclear sequence. Future applications of the CDCE/hifi PCR approach to human genetic analysis include studies of somatic mitochondrial mutations with respect to aging, measurement of mutational spectra of nuclear genes in healthy human tissues and population screening for disease-associated single nucleotide polymorphisms (SNPs) in large pooled samples.

Base Sequence↗

FuFiHLA: a tool for full-field HLA typing from long-read data.

MOTIVATION: Allele typing for Human Leukocyte Antigen (HLA) genes has many important clinical applications. Popular short-read typing can only accurately distinguish alleles at the coding sequence level, which potentially limit our understanding of the effect of variants in non-coding region. Long read data has been proved to be useful in typing HLA alleles in full resolution, but only a few tools are publicly available and with significant limitations in practical application. RESULTS: We developed FuFiHLA, a lightweight open-source software, to type HLA alleles. Currently it supports typing alleles of six HLA genes (HLA-A, HLA-B, HLA-C, HLA-DRB1, HLA-DQA1, and HLA-DQB1) from long reads. Evaluation using 233 PacBio HiFi WGS samples from HPRC shows that FuFiHLA achieves 99.6% accuracy in the full field allele typing and QV as 51.8 for consensus allele sequence construction. Additional testing on four Nanopore R10 reads demonstrates slightly reduced accuracy in the fourth field. AVAILABILITY: FuFiHLA is available at https://github.com/jingqing-hu/FuFiHLA under MIT License.

Humans↗

The first chromosome-level genome of the lappet moth Trabala vishnou (Lepidoptera: Lasiocampidae).

Trabala vishnou (Lefèbvre, 1827) (Lepidoptera: Lasiocampidae) is a destructive leaf-eating pest that causes severe damage to forest ecosystems, leading to substantial economic losses. Herein, we sequenced and assembled a high-quality chromosome-level genome of T. vishnou using a combination of Illumina reads, PacBio HiFi reads, and High throughput Chromosome Conformation Capture (Hi-C) technologies. The genome size is 561.86 Mb and spans 25 chromosomes, exhibiting a high level of contiguity (scaffold/contig N50 = 21.75 Mb/20.67 Mb). Benchmarking Universal Single-Copy Orthologs (BUSCO) analysis a 99.5% completeness score for this genome assembly. Repeat elements constitute 62.66% of the genome. A total of 1,630 non-coding RNAs and 12,895 protein-coding genes have been identified within the genome. The first chromosome-level genome of T. vishnou serves as a valuable reference for elucidating the evolution of functional traits in Lasiocampidae family and will facilitate the development of strategies for controlling defoliating pests.

Animals↗

Near-complete reference genome assembly of Hoya carnosa.

Hoya R. Br. is the largest genus in the tribe Marsdenieae (Apocynaceae), comprising 350-450 species. Hoya species are popular in horticulture for their distinctive floral traits and fragrances, primarily sourced from domestication and mutation breeding. However, the lack of molecular analysis for floral morphological traits has limited their cultivation and application. In this study, we assembled a near-complete reference genome for H. carnosa, the model species of the genus, using PacBio HiFi reads and Hi-C method. The genome size was approximately 465.7 Mb with a contig N50 of 39.3 Mb. 99.7% of the sequences were anchored to 11 pseudochromosomes, and the assembly achieved a BUSCO score of 98.5%. We predicted 24,309 protein-coding genes, of which 90.2% (21,927) were functionally annotated. This high-quality genome provides a valuable reference for the research of evolution, conservation and molecular breeding in Hoya.

Genome, Plant↗

Chromosome-level genome assembly and annotation of the porcupine fish (Diodon hystrix).

The porcupinefish (Diodon hystrix), a coral reef teleost, is widely distributed in tropical/subtropical waters of the Pacific, Atlantic, Indian Oceans, and Mediterranean Sea. It shares easily recognizable features with pufferfish, such as body inflation and spines. Additionally, its culinary value makes D. hystrix a highly desirable species in many tropical coastal regions, with considerable market potential. However, lack of a high-quality genome hindered further studies on its reproduction, molecular biology, and genomic improvement. Here, we assembled the chromosome-scale genome using PacBio HiFi, ultra-long reads, and Hi-C. Of the 713.62 Mb genome, 98.63% anchored to 23 chromosomes (scaffold N50: 31.52 Mb) with 39.82% repetitive sequences. The assembled genome achieved a BUSCO completeness score of 97.7%, with 23,171 protein-coding genes predicted, 22,221 of which were functionally annotated. Phylogenetic analysis identified D. hystrix's evolutionary relationships with other species in the Tetraodontiformes. In summary, the high-quality genome of D. hystrix sheds light on valuable insights into genome size evolution, and provides a valuable resource for exploiting genomic study and breeding applications in this species.

Animals↗

The complete and annotated mitochondrial genome of Hemileia vastatrix Race I, causal agent of coffee leaf rust.

Hemileia vastatrix is the fungal pathogen responsible for coffee leaf rust (CLR), the most economically important disease of Coffea arabica worldwide. Recently, the nuclear genome of this fungus was completely deciphered. However, the mitochondrial genome of H. vastatrix has remained undercharacterized. Here, we present the complete, circularized mitochondrial genome of H. vastatrix Race I (isolate HvRI), assembled using a hybrid approach combining PacBio HiFi long reads and BGIseq short reads. The genome is 173,525 bp in length with a GC content of 33.1% and encodes 41 functional genes, including 15 protein-coding genes, 2 rRNAs, and 24 tRNAs. The assembly reveals significant structural complexity, driven by intron expansion in the cox1 and cob genes. Notably, the atp8 gene contains a group II intron, rare for this locus, whose internal open reading frame displays evidence of pseudogenization via internal stop codons.. We also characterized a putative replication initiation zone (~1.2 kb) defined by a poly-G homopolymer and conserved regulatory motifs. The mitogenome of the HvRI isolate does not contain cob mutations that lead to amino acid substitutions G143A and F129L associated with the quinone outside inhibitor (QoI) fungicide resistance. This high-quality mitogenome is an important resource for comparative mitogenomics, population diversity studies, and the molecular surveillance of QoI fungicide resistance.

Genome, Mitochondrial↗

A chromosome-level genome assembly and annotation of Cercis chuniana (Fabaceae).

The genus Cercis L., at the base of the subfamily Cercidoideae of Fabaceae, is known for its ecological adaptability and significant medicinal, ornamental, and economic value. However, the lack of a high-quality genome hinders the understanding of the evolution of Cercis and Fabaceae. In this study, we present a chromosome-level genome of Cercis chuniana by combining Illumina short reads, PacBio HiFi long reads, and Hi-C data. The final genome size is 355.53 Mb, consisting of 12 contigs with a N50 of 42.34 Mb. Notably, 344.24 Mb, corresponding to 96.82% of the genome, was anchored to seven chromosomes. The assembly comprises 24.83% repetitive sequences, including 19.32% long terminal repeats. Additionally, a total of 33,837 protein-coding genes were predicted in the genome, with 32,709 (96.67%) genes successfully annotated. The high-quality genome assembly of C. chuniana not only bridges the existing gap in genomic data and offers important resources for molecular studies of this species, but also provides essential insights for future studies on speciation, functional and comparative genomics within the Fabaceae family.

Genome, Plant↗