PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “HiFi sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

A chromosome-level genome assembly of Coffea arabica L. var. 'Kona Typica'.

Coffea arabica L. var. 'Kona Typica' is renowned for its premium cup quality, but its vulnerability to pests and diseases limits production. To accelerate cultivar improvement, we generated a chromosome-level genome assembly of 'Kona Typica' using PacBio HiFi sequencing and Hi-C scaffolding technology. The final assembly spans 1.13 Gb, with a scaffold N50 of 50.50 Mb, organized into 22 chromosomes. BUSCO assessment indicated a high completeness at 99.1%. We annotated 65,458 protein-coding genes and identified 1,073,545 interspersed repeats, accounting for 65.16% of the genome. Analysis of transposon insertion ages revealed that most long terminal repeat retrotransposons proliferated after the polyploidization event. This high-quality genome assembly of 'Kona Typica' provides a valuable resource for exploring coffee genomic evolution and genetic mechanisms of complex traits, facilitating genomics studies and the development of improved coffee cultivars with enhanced disease resistance and quality traits.

Coffea↗

Chromosome-level genome assembly of the small-sized Taihang donkey (Equus asinus).

China harbors a rich diversity of donkey breeds, with small-sized donkeys (<110&#x2009;cm) representing a largely underexplored group. Here, we present the first high-quality, chromosome-level genome assembly of a small-sized donkey, generated using PacBio HiFi sequencing (286.7&#x2009;Gb), Hi-C scaffolding (240.47&#x2009;Gb), and annotated with RNA-seq data. The final assembly has a total length of 2.7&#x2009;Gb and comprises 32 chromosomes (including both X and Y chromosomes), in which five chromosomes were fully assembled without gaps. It possesses a scaffold N50 of 106.70&#x2009;Mb and 84 contigs (contig N50&#x2009;=&#x2009;63.60&#x2009;Mb), and captures 99.2% of BUSCO genes. The assembly achieved a consensus quality value (QV) of 77.44, corresponding to an extremely low base-level error rate, indicating exceptional nucleotide accuracy. This high-quality genome provides a valuable resource for investigating genetic variation, adaptive evolution, and domestication processes in small-sized donkeys, and will facilitate the conservation and sustainable utilization of rich donkey genetic resources in China.

Animals↗

Genome analysis of the glycosphingolipid-producing green alga tetraselmis sp. NKG400013.

Microalgae are gaining attention as sustainable resources for the production of valuable compounds, including biofuels, pigments, and bioactive metabolites. To support metabolic engineering and genome editing approaches aimed at enhancing these traits, high-quality genome assemblies are essential; however, genomic information remains limited for many microalgal lineages. Tetraselmis sp. NKG400013 is a green alga known for high glycosphingolipid accumulation with distinctive structural features. Here, we report a draft genome assembly of this strain generated using PacBio HiFi sequencing and transcriptome-supported annotation. The assembled genome spans 423.7&#x2005;Mbp, with 74.5% repetitive sequences and 15,322 predicted protein-coding genes. Comparative analyses across 11 green algal species revealed a positive correlation between genome sizes and repeat contents, indicating that transposable element expansion, particularly long terminal repeat retrotransposons, has substantially contributed to genome enlargement in Tetraselmis. Genome-wide functional annotation and ortholog inference identified core enzymes required for glycosylceramide biosynthesis. Both sphingolipid &#x394;4 and &#x394;8 desaturases were identified in Tetraselmis and their coexistence suggests an expanded capacity for long-chain base modification that may underlie its distinctive glycosphingolipid profile. These results establish a genomic framework for understanding the high glycosphingolipid-producing capacity of NKG400013 and provide insights into the evolutionary diversification of sphingolipid metabolism in green algae.

Chlorophyta↗

Haplotype-aware long-read error correction.

Error correction of long reads is an important initial step in genome assembly workflows. For organisms with ploidy greater than one, it is important to preserve haplotype-specific variation during read correction. This challenge has driven the development of several haplotype-aware correction methods. However, existing methods are based on either ad-hoc heuristics or deep learning approaches. In this paper, we introduce a rigorous formulation for this problem. Our approach builds on the minimum error correction framework used in reference-based haplotype phasing. We prove that the proposed formulation for error correction of reads in de novo context, i.e., without using a reference genome, is NP-hard. To make our exact algorithm scale to large datasets, we introduce practical heuristics. Experiments using PacBio HiFi sequencing datasets from human and plant genomes show that our approach achieves accuracy comparable to state-of-the-art methods. Implementation: https://github.com/at-cg/HALE .

Clustering↗

Chromosome-level genome assembly with telomeric repeats at scaffold ends for Rhabdosargus sarba.

Rhabdosargus sarba, the goldlined seabream, is a euryhaline marine fish of great aquaculture potential. Genome sequencing and assembly of R. sarba was carried utilizing a multi-platform sequencing strategy that included long-read sequencing (PacBio HiFi), short-read sequencing (Illumina), and chromatin interaction mapping (Hi-C). The final genome assembly size after scaffolding was 764.59&#x2009;Mb in 31 scaffolds with an N50 length of 33.98&#x2009;Mb. Repeat profiling of primary assembly showed that 28.71% of the genome comprises of repeat elements. Gene prediction utilising the evidence from ab initio prediction and transcriptome data revealed 26,913 protein encoding genes and functional annotation and pathway analysis showed their participation in 332 pathways. This genome is an excellent resource for future research on genetic improvement and molecular breeding programmes for R. sarba.

Animals↗

A chromosome-level genome assembly of Guimi No. 2 (Actinidia chinensis).

In this study, we report a high-quality chromosome-level genome assembly of Actinidia chinensis var. chinensis 'Guimi No. 2'. This cultivar, discovered in Guizhou karst ecosystems, exhibits resistance to Pseudomonas syringae pv. actinidiae (Psa). Using a combination of MGI short-read sequencing, PacBio HiFi long-read sequencing, and Hi-C technology, we generated a genome assembly of 608.43&#x2009;Mb with a contig N50 of 20.70&#x2009;Mb, and 99.70% of the assembly was successfully anchored onto 29 pseudochromosomes. The quality value (QV) and the LTR Assembly Index (LAI) of the assembled genome were 72.23 and 10.10. The BUSCO analysis indicated that the genome assembly and gene model prediction were 98.40% and 96.56% complete, respectively. A total of 251.15&#x2009;Mb of repetitive sequences and 45,986 protein-coding genes were annotated. This genome assembly provides critical insights into A. chinensis's genomic architecture and serves as a foundational resource for elucidating disease resistance mechanisms against Psa, while enabling comparative phylogenomic studies across the Actinidia genus.

Actinidia↗

Inferring the demographic history of Chinese and Indian rhesus macaque (Macaca mulatta) populations from PacBio HiFi long-read sequencing data.

The rhesus macaque (Macaca mulatta) is one of the most widely used animal models in biomedical research, both as it resembles humans in key biological aspects and as it is characterized by a broad geographic range. Most of the individuals housed in U.S. research colonies have been sampled from either China or India, though notably the source population of these animals has significantly shifted over time. Given the substantial genetic and immunological differences between these populations, a deeper understanding of the underlying population structure is critically important for biomedical interpretation. Despite this, the demographic histories of these two populations remain poorly resolved. Here, we present an analysis of whole-genome, PacBio HiFi long-read sequencing data from ten unrelated individuals of each population, applying four related model- and non-model based demographic inference approaches, in order to reconstruct their ancestral history. We evaluated the fit of the subsequently estimated models against the empirical data, and incorporated underlying uncertainty in the mutation rates used for scaling. We inferred a well-fitting population history characterized by substantial structure between Chinese and Indian populations, with a split time &#x223c;140,000 generations ago from an ancestral population of &#x223c;65,000 individuals. We additionally inferred the subsequent history of size change within, and gene flow between, these populations, reaching the current estimated sizes of &#x223c;220,000 individuals in the Chinese population and &#x223c;14,000 individuals in the Indian population. The robust baseline demographic model established in this study will serve as a valuable resource for future research on this species, including for improved fine-scale recombination mapping, selection inference, and association studies.

Cercopithecidae↗

Chromosome-level genome assembly of Sinocyclocheilus jii based on PacBio HiFi and Hi-C sequencing.

Sinocyclocheilus jii, a cavefish species endemic to China, belongs to the genus Sinocyclocheilus within the family Cyprinidae. Species within this genus exhibit significant morphological differentiation, making it not only the most species-rich genus within Cyprinidae in China but also the most diverse group of cavefishes worldwide. However, the limited availability of genomic resources has limited investigations into the genetic basis of trait variations, phylogenetic relationships, and adaptive evolution in this genus. In this study, we assembled a chromosome-level reference genome for S. jii by integrating PacBio HiFi long reads, Illumina short reads, and Hi-C sequencing data. Flow cytometry was used to estimate the genome size prior to assembly, providing a key step in technical validation. The final genome assembly spans 1.75&#x2009;Gb with a contig N50 of 35.0&#x2009;Mb. Using Hi-C sequencing data, the assembled scaffolds were successfully anchored to 50 chromosomes. The completeness of the chromosome-level assembly was estimated at 98.9% by BUSCO analysis. Genome annotation identified 855.5&#x2009;Mb of repetitive sequences and predicted a total of 52,867 protein-coding genes, of which 51,932 genes were functionally annotated. This study presents a high-quality chromosome-level genome assembly and annotation of S. jii, providing a fundamental genomic resource for future phylogenetic and evolutionary studies.

Animals↗

Chromosome-level de novo assembly of the nuclear and mitochondrial genomes of Arcopilus aureus, a filamentous fungus with multifaceted ecological and economic roles.

The filamentous fungus Arcopilus aureus (Sordariale: Chaetomiaceae) is notable for its multi-domain significance across agriculture, medicine, and industry. In this study, we generated a chromosome-level nuclear genome and a complete circular mitogenome for A. aureus by integrating data from next-generation sequencing, PacBio HiFi, and Hi-C technologies. The final nuclear genome assembly spans 33.77&#x2009;Mb (GC content: 57.67%), and was organized into seven chromosomal-sized scaffolds (only one gap) with an N50 size of 5.09&#x2009;Mb and BUSCO completeness of 95.91%. A total of 10,282 protein-coding genes, 228 non-coding RNAs, and ~1.77&#x2009;Mb of repetitive elements were predicted in the nuclear genome. By contrast, the mitogenome of A. aureus is 33,820&#x2009;bp in length, with a GC content of 25.96%. It harbors 15 typical mitochondrial protein-coding genes, one unidentified ORF, two rRNAs (small subunit rns and large subunit rnl), and 28 tRNAs. This high-quality genome assembly provides a valuable resource for understanding the ecology, genetics, and evolution of A. aureus, which facilitates elucidating its mechanisms of biocontrol, infection, and metabolite synthesis.

Genome, Mitochondrial↗

Long-read, high-coverage reference genome of the nymphalid butterfly Catonephele acontius (Nymphalidae: Biblidinae).

Catonephele acontius (Nymphalidae:Biblidinae:Epicalinii) is a butterfly species with a wide distribution across the Neotropics including the Amazon. Here, we present a long-read high-coverage reference genome for this species to serve as a genomic resource for future studies on Biblidinae butterflies, a group that is the subject of ongoing studies of seasonal adaptation under climate change. We used PacBio HiFi and IsoSeq reads to generate a highly contiguous and well-annotated reference genome. Five libraries were constructed, 4 using RNA from different tissues and 1 using high molecular weight (HMW) DNA from a wild-caught female. The DNA was sequenced using PacBio HiFi technology, and the RNA was sequenced using long read PacBio IsoSeq technology. About 20 Gb of raw HiFi data were generated and assembled to an initial size of 520.7 Mb (39 &#xd7; homozygous coverage) in 90 contigs. The assembly was then polished and decontaminated into 40 contigs with an N50 of 19.927 Mb (BUSCO completeness: 99.0%; duplication: 0.5%; fragmentation: 0.7%; and missing: 0.3%). Final assembly size was 519.2 Mb. Repeats were annotated, showing that the genome consisted of 40.4% transposable elements. IsoSeq transcriptome data from antennae, leg, ovary, and digestive tissue was then used to structurally and functionally annotate gene models for the softmasked genome, uncovering &#x223c;18,500 genes, with 70% of them given functional annotation. This reference assembly joins many published genomes in the Nymphalidae family but represents one of the first high-quality genomes from the Biblidinae subfamily. It provides a valuable resource to study the evolution of plastic and seasonal traits and will help investigate the genetic processes that may influence these species' responses to rapid climate change.

Animals↗

Genetic Adaptation to Brackish Water and Spawning Season in European Cisco.

How species adapt to diverse environmental conditions is essential for understanding evolution and the maintenance of biodiversity. The European cisco (Coregonus albula) is a salmonid that occurs in both fresh and brackish water, and this together with the presence of sympatric spring- and autumn-spawning lacustrine populations provides an opportunity for studying the genetics of adaptation in relation to salinity and timing of reproduction. Here, we present a high-quality reference genome of the European cisco based on PacBio HiFi long read sequencing and HiC-directed scaffolding. We generated low-coverage whole-genome sequencing data from 336 individuals across 12 population samples to explore population structure and genetics of ecological adaptation. We found a major subdivision between two groups of populations most likely reflecting colonisation from different glacial refugia. Within the two major groups, we detected further genetic differentiation between spring- and autumn-spawning populations and between populations from freshwater lakes, rivers and brackish water (Bothnian Bay). A genome-wide screen for genetic differentiation among populations identified a set of outlier SNPs strongly correlated with spawning timing and salinity. Several of the genes associated with spawning time, including BHLHE40, TIMELESS and CPT1A, have previously been shown to have a role in circadian rhythm biology. As many as 17 loci were associated with genetic differentiation between populations reproducing in fresh and brackish water. This study provides insights into the genomic basis of ecological adaptation in European cisco with implications for sustainable fishery management.

Animals↗

Draft genome sequence of the almond red leaf blotch pathogen Polystigma amygdalinum assembled from infected almond leaves collected in California, USA.

We report a draft genome assembly of Polystigma amygdalinum, the causal agent of almond red leaf blotch. DNA extracted from infected leaves was sequenced using PacBio HiFi, and host-derived reads were removed bioinformatically. The 238.7-Mb assembly (90.8% BUSCO completeness) is highly repetitive (82.3%) and unusually large for an ascomycete.

Polystigma amygdalinum↗

nf-core/pacsomatic: a scalable somatic analytic pipeline using PacBio HiFi data.

MOTIVATION: Pacific Biosciences (PacBio) HiFi long-read sequencing enables robust characterization of complex genomic regions, repetitive elements, and structural variants (SVs) that are often inaccessible to short-read technologies. To fully leverage HiFi reads to advance cancer genomics and epigenetics, researchers require an end-to-end, scalable and optimized bioinformatics workflow. The nf-core framework meets this need by providing rigorously tested, community-curated pipelines that ensure reproducibility, transparency, and broad compatibility across computational environments. RESULTS: We present nf-core/pacsomatic, an automated Nextflow DSL2 pipeline designed for comprehensive paired tumor-normal somatic analysis using PacBio HiFi data. The workflow includes steps for read alignments against reference genome, somatic SNV/indel, SV, and CNV calling, CpG methylation profiling and differential methylation region (DMR) detection. Additional downstream modules support functional annotation, mutational signature analysis, tumor purity and ploidy estimation, and homologous recombination deficiency (HRD) assessment. Utilizing nf-core's modular design and containerized execution, nf-core/pacsomatic provides a stable framework for the reproducible discovery of biological insights. AVAILABILITY: nf-core/pacsomatic is available under the MIT License at nf-core (https://nf-co.re/pacsomatic) and github (https://github.com/nf-core/pacsomatic).

Software↗

First clinical diagnosis of FAME3 via commercial Long-Read sequencing reveals mosaic repeat expansion in MARCHF6 gene.

Familial Adult Myoclonic Epilepsy type 3 (FAME3) is a rare autosomal dominant disorder characterized by cortical tremor and epilepsy, caused by a noncoding pentanucleotide repeat expansion (TTTTA/TTTCA)n in the MARCHF6 gene. Conventional genetic testing often fails to detect this expansion due to its repetitive structure and intronic location. We evaluated a 61-year-old woman with refractory myoclonic and generalized tonic-clonic seizures, whose prior genetic testing-including exome and genome sequencing-was non-diagnostic. Using PacBio HiFi long-read whole-genome sequencing and the tandem repeat genotyping tool TRGT, we identified a pathogenic MARCHF6 intronic expansion. The proband harbored one allele with 15 TTTTA repeats and a second allele with a compound expansion of 661 TTTTA and 12 TTTCA repeats. Three affected relatives shared similarly expanded alleles, but with increasing repeat size in the latter generations. Importantly, analysis using TRGT-instability revealed repeat mosaicism in all affected individuals, reflected by variability in motif counts across individual sequencing reads. This somatic heterogeneity may contribute to the phenotypic penetrance, variable expressivity and pleiotropism seen in FAME3 disease expression. To our knowledge, this is the first clinical diagnosis of FAME3 using a commercially available long-read sequencing platform, underscoring its diagnostic utility in resolving complex repeat expansion disorders and uncovering biologically relevant mosaicism.

Humans↗

Accurate somatic small variant discovery for multiple sequencing technologies with DeepSomatic.

Somatic variant detection is an integral part of cancer genomics analysis. While most methods have focused on short-read sequencing, long-read technologies offer potential advantages in repeat mapping and variant phasing. We present DeepSomatic, a deep-learning method for detecting somatic small nucleotide variations and insertions and deletions from both short-read and long-read data. The method has modes for whole-genome and whole-exome sequencing and can run on tumor-normal, tumor-only and formalin-fixed paraffin-embedded samples. To train DeepSomatic and help address the dearth of publicly available training and benchmarking data for somatic variant detection, we generated and make openly available the Cancer Standards Long-read Evaluation (CASTLE) dataset of six matched tumor-normal cell line pairs whole-genome sequenced with Illumina, PacBio HiFi and Oxford Nanopore Technologies, along with benchmark variant sets. Across samples, both cell line and patient-derived, and across short-read and long-read sequencing technologies, DeepSomatic consistently outperforms existing callers.

Humans↗

Integrative Long-Read Multi-Omics of a Patient With GPI Deficiency: A Molecular Case Study of a Candidate Dual-Effect GPI Variant.

The molecular determinants of phenotypic severity in red cell enzymopathies are often obscured by the disconnect between coding sequence variants and their regulatory landscapes. Here we present a single-patient molecular case study that uses an integrative multi-omic approach-combining short-read WGS, PacBio HiFi long-read sequencing, native CpG methylation profiling, and Iso-Seq full-length transcriptomics-to characterize a severe, transfusion-dependent hemolytic anaemia. We identified a compound heterozygous state in the glucose-6-phosphate isomerase (GPI) gene, with no wild-type allele present. One allele (Haplotype 1) carried a missense variant (p.His191Arg); the other (Haplotype 2) carried a distinct missense variant, c.1414C>T (p.Arg472Cys), previously reported as biochemically unstable. Long-read phasing placed the two variants in trans. Allele-resolved transcript counts showed a directionally consistent but statistically non-significant trend toward higher expression of Haplotype 2 across two Iso-Seq replicates. Notably, the c.1414C>T transition abolishes a local CpG dinucleotide; in a small number of haplotype-2 reads spanning this position, the corresponding cytosine on the wild-type/Haplotype-1 background was methylated. We did not measure GPI protein abundance, enzymatic activity, or stability in this patient, and we do not establish that methylation at this site regulates GPI transcription. On the basis of these correlative observations in a single patient, we propose-as a hypothesis for future testing-that a coding variant might simultaneously perturb protein stability and disrupt a local epigenetic mark, and we outline the experiments required to test whether such a dual effect contributes to disease. This case illustrates the value of integrative long-read multi-omics for generating mechanistic hypotheses about variants of uncertain significance, while underscoring that causal claims require dedicated functional validation.

Humans↗

Chromosome-Level Assembly and Annotation of the Grey Reef Shark (Carcharhinus amblyrhynchos) Genome.

To date less than 5% of shark species have nuclear reference genomes, despite next-generation sequencing advances. Particularly for threatened shark species, there is a lack of reliable genomes which are crucial in facilitating research and conservation applications. We assembled the first nuclear reference genome of the endangered grey reef shark (Carcharhinus amblyrhynchos) using long-read PacBio HiFi and Omni-C sequencing to reach chromosome-level contiguity (36 pseudochromosomes; 2.9&#x2005;Gbp) and high completeness (94% complete BUSCOs). BRAKER3 annotated 16,505 protein-coding genes after masking repetitive elements which accounted for 59% of the genome. We identified potential X and Y sex chromosomes on pseudochromosomes 36 and 57, respectively. The quality and completeness of the draft genome of C. amblyrhynchos will enable researchers to investigate genetic variations and adaptations specific to this species as well as across other Carcharhinus spp., opening new venues for comparative genomics and advancing conservation genetic applications.

Animals↗

ALPINE: a scalable pipeline for comprehensive classification of gene-editing outcomes from long-read amplicon sequencing.

SUMMARY: CRISPR genome editing has enabled precise genetic modification for gene and cell therapies, but edits often produce heterogeneous on-target outcomes, including homology-directed repair (HDR) knock-ins, DNA repair template integrations, and structural variants. Existing tools are frequently limited to short reads or lack viral vector-specific integration categories needed for therapeutic development. Here, we present ALPINE (Amplicon Long-read Pipeline for INtegration Evaluation), a scalable and reproducible pipeline for classifying and quantifying gene-editing outcomes from long-read amplicon sequencing supporting both PacBio HiFi and Oxford Nanopore platforms. ALPINE classifies reads into 10+ categories, including DNA repair vector integration subtypes, and performs variant calling near the gene-edited site with batch, multi-sample reporting. Uniquely, ALPINE can distinguish between cells treated with multiple DNA repair vectors and identify distinct molecular features, such as inverted terminal repeats (ITRs), enabling comprehensive characterization of complex gene editing outcomes. Dual-target benchmarking on simulated datasets demonstrated high accuracy for transgene integration events. Independent validation on public crosslinked-HDR dataset confirmed ALPINE's integration detection capabilities, and application to edited T cell samples demonstrated comprehensive gene-editing outcome profiling. AVAILABILITY: ALPINE is available under MIT license at https://github.com/Maggi-Chen/ALPINE and https://doi.org/10.5281/zenodo.20272510. All analysis scripts and visualization code used in this manuscript are available at https://github.com/Maggi-Chen/ALPINE-manuscript-analysis. Simulated datasets are deposited at Zenodo (https://doi.org/10.5281/zenodo.20260865). Public dataset PRJNA913199 is available through NCBI SRA.

Gene Editing↗