PubMed HealthSearch

SEARCH · PubMed Health

Results for “PacBio sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

A chromosome-level assembly of the alpine snow alga Chloromonas typhlos.

Chloromonas typhlos is a cosmopolitan alpine snow alga distributed across continents, and its blooming accelerates snow melting by decreasing the amount of snow albedo. To elucidate the genetic traits underlying the adaptation of C. typhlos to the alpine habitat, we combined PacBio sequencing and Hi-C to generate a high-quality chromosome-level genome assembly (contig N50: 1.29 Mb; scaffold N50: 7.23 Mb) with 31 chromosomes and a genome size of 200.86 Mb. Repetitive elements constituted 11.05% of the genome, and 16,133 protein-coding genes were predicted, of which 82% were functionally annotated. This study provides a set of omics resources both for snow algae and the genus Chloromonas.

Snow

Characterization and analysis of the full-length transcriptome of Frankliniella occidentalis (Thysanoptera: Thripidae).

BACKGROUND: Frankliniella occidentalis, an insect belonging to the order Thysanoptera, causes severe damage to agricultural and horticultural crops, resulting in significant economic losses worldwide. The development of molecular and sequencing technologies has helped elucidate the molecular mechanisms regulating its growth and development as well as its damaging activity. However, much remains to be explored. To further investigate the molecular complexity of this species, we sequenced the full-length transcriptome of mixed samples obtained from specimens at all developmental stages. RESULTS: Of all transcripts, 89.04% matched with the reference genome; additionally, 29,750 alternative splicing events, 2,342 genes with poly(A) sites, and 153 candidate fusion transcript events were identified, and 4,235 long noncoding RNAs were discovered. CONCLUSIONS: This is the first full-length transcriptome of F. occidentalis reported to date. This study greatly contributes to the understanding of the molecular complexity and diversity of this insect, providing a basis to develop specific molecular targets as well as resources for gene function studies in other insects.

Animals

VicMAG, an open-source tool for visualizing circular metagenome-assembled genomes highlighting bacterial virulence and antimicrobial resistance.

Bacterial pathogens spread in clinical and environmental settings, and mobile genetic elements (MGEs), such as plasmids and phages, mediate the transfer of virulence factor genes (VFGs) and antimicrobial resistance genes (ARGs) among bacterial communities. Metagenomic analysis of environmental and wastewater samples using highly accurate long-read sequencing technologies, such as Pacific Biosciences (PacBio) HiFi sequencing, provides valuable insights into monitoring the regional spread of VFGs and ARGs, including dissemination mediated by MGEs. No visualization tool is currently available for the comprehensive display of numerous resulting circular metagenome-assembled genomes (cMAGs) with functional gene annotations. Here, we developed visualization of circular metagenome-assembled genome (VicMAG), a visualization tool for highly complex cMAGs derived from long-read metagenome assemblies annotated using updated databases of VFGs, ARGs, and MGEs. Using 353 cMAGs from PacBio HiFi sequencing of a wastewater sample, we demonstrated the utility of VicMAG for metagenome visualization. VicMAG provides comprehensive, size-aware visualization of cMAGs representing bacterial chromosomes and plasmids, annotated with VFGs, ARGs, and phages. By simultaneously visualizing all cMAGs in a framework, VicMAG facilitates a holistic understanding of the distribution and genomic context of VFGs and ARGs across complex microbial communities. This tool supports integrated surveillance of bacteria associated with virulence and antimicrobial resistance across clinical, environmental, and One Health contexts.

Metagenome

nf-core/pacvar: a pipeline for analyzing long-read PacBio whole genome and repeat expansion sequencing data.

MOTIVATION: Pacific Biosciences (PacBio) single-molecule, long-read sequencing enables whole genome annotation and the characterization of 20 complex repetitive repeat regions, especially relevant to neurodegenerative diseases, through their PureTarget panel. Long-read whole-genome sequencing (WGS) also allows for the detection of structural variants that would be difficult to detect with traditional short-read sequencing. However, the raw unaligned Binary Alignment Map data need to be processed before analysis. There is a need for an intuitive comprehensive bioinformatic pipeline that can analyze these data. RESULTS: We present nf-core/pacvar, a comprehensive pipeline for analyzing both PacBio single-molecule PureTarget and WGS data that demultiplexes and parallelizes pre-processing, variant calling and repeat characterization. nf-core/pacvar is compatible with little configuration and has few dependencies. This pipeline enables rapid end-to-end, parallel processing of PacBio single-molecule whole genome and targeted repeat expansion sequencing. AVAILABILITY AND IMPLEMENTATION: nf-core/pacvar is available on nf-core website (https://nf-co.re/pacvar/) and on github (https://github.com/nf-core/pacvar) under MIT License (DOI: 10.5281/zenodo.14813048).

Software

Patterns of Drug Resistance, Drug Resistance Conferring Mutations and Genomic DNA Methylation Revealed in Mycobacterium tuberculosis From South Africa.

Tuberculosis remains a major public health threat globally, with drug-resistant strains undermining treatment efficacy. We analyzed 126 Mycobacterium tuberculosis (M. tuberculosis) isolates with diverse drug resistance spectra and selected 35 for whole genome sequencing (WGS) using Illumina NextSeq, SMRT PacBio Onso and SMRT PacBio Revio sequencing platforms. The study aimed to characterize drug resistance profiles, compare short- and long-read sequencing performance, identify lineages among South African isolates, detect known drug resistance mutations and their lineage-specific patterns, and utilize long-read SMRT platforms for epigenetic profiling. Multiple drug resistance mutations were identified, some lineage-specific, and notably, East-African-Indian (EAI) Lineage 1 isolates often considered less pathogenic, showed significant potential for multidrug-resistance development, including higher fluoroquinolone resistance as compared to other lineages. Three DNA motifs with methylated adenines, namely CACGCaG, CtCCaG and GaTNNNNRtAC, were detected, with methylation patterns varying by lineage and strain due to mutations in the corresponding methyltransferases (MTases). A particularly notable finding was the stable maintenance of a genetic heterogeneity in the mamB MTase, performing methylation at CACGCaG motifs. These results highlight the combined role of genetic and epigenetic variation in M. tuberculosis adaptive evolution and underscore the value of integrating long-read sequencing into TB surveillance and research.

Mycobacterium tuberculosis

Long-read, high-coverage reference genome of the nymphalid butterfly Catonephele acontius (Nymphalidae: Biblidinae).

Catonephele acontius (Nymphalidae:Biblidinae:Epicalinii) is a butterfly species with a wide distribution across the Neotropics including the Amazon. Here, we present a long-read high-coverage reference genome for this species to serve as a genomic resource for future studies on Biblidinae butterflies, a group that is the subject of ongoing studies of seasonal adaptation under climate change. We used PacBio HiFi and IsoSeq reads to generate a highly contiguous and well-annotated reference genome. Five libraries were constructed, 4 using RNA from different tissues and 1 using high molecular weight (HMW) DNA from a wild-caught female. The DNA was sequenced using PacBio HiFi technology, and the RNA was sequenced using long read PacBio IsoSeq technology. About 20 Gb of raw HiFi data were generated and assembled to an initial size of 520.7 Mb (39 × homozygous coverage) in 90 contigs. The assembly was then polished and decontaminated into 40 contigs with an N50 of 19.927 Mb (BUSCO completeness: 99.0%; duplication: 0.5%; fragmentation: 0.7%; and missing: 0.3%). Final assembly size was 519.2 Mb. Repeats were annotated, showing that the genome consisted of 40.4% transposable elements. IsoSeq transcriptome data from antennae, leg, ovary, and digestive tissue was then used to structurally and functionally annotate gene models for the softmasked genome, uncovering ∼18,500 genes, with 70% of them given functional annotation. This reference assembly joins many published genomes in the Nymphalidae family but represents one of the first high-quality genomes from the Biblidinae subfamily. It provides a valuable resource to study the evolution of plastic and seasonal traits and will help investigate the genetic processes that may influence these species' responses to rapid climate change.

Animals

Inferring the demographic history of Chinese and Indian rhesus macaque (Macaca mulatta) populations from PacBio HiFi long-read sequencing data.

The rhesus macaque (Macaca mulatta) is one of the most widely used animal models in biomedical research, both as it resembles humans in key biological aspects and as it is characterized by a broad geographic range. Most of the individuals housed in U.S. research colonies have been sampled from either China or India, though notably the source population of these animals has significantly shifted over time. Given the substantial genetic and immunological differences between these populations, a deeper understanding of the underlying population structure is critically important for biomedical interpretation. Despite this, the demographic histories of these two populations remain poorly resolved. Here, we present an analysis of whole-genome, PacBio HiFi long-read sequencing data from ten unrelated individuals of each population, applying four related model- and non-model based demographic inference approaches, in order to reconstruct their ancestral history. We evaluated the fit of the subsequently estimated models against the empirical data, and incorporated underlying uncertainty in the mutation rates used for scaling. We inferred a well-fitting population history characterized by substantial structure between Chinese and Indian populations, with a split time ∼140,000 generations ago from an ancestral population of ∼65,000 individuals. We additionally inferred the subsequent history of size change within, and gene flow between, these populations, reaching the current estimated sizes of ∼220,000 individuals in the Chinese population and ∼14,000 individuals in the Indian population. The robust baseline demographic model established in this study will serve as a valuable resource for future research on this species, including for improved fine-scale recombination mapping, selection inference, and association studies.

Cercopithecidae

High-Quality Genome Assembly, Metabolome, Pangenome, and Metabolic Models of Megasphaera hexanoica KCCM 43214T.

Megasphaera hexanoica KCCM 43214T, isolated from cow rumen, is capable of producing medium-chain carboxylic acids such as hexanoate and octanoate. In this study, we present a high-quality genome assembly, along with intracellular metabolomic profiling and pangenomic analysis. Illumina sequencing generated 2.3 Gbp from 15,293,634 reads with a GC content of 49.5%, while PacBio HiFi sequencing produced 331.5 Mbp across 45,266 reads, with an average read length of 7,323 bp and a HiFi read N50 of 8,214 bp. Hybrid assembly of short and long reads resulted in a single 2.88 Mbp contig, containing 2,835 protein-coding genes. Genome-scale metabolic models were constructed to evaluate its metabolic capabilities under specific growth conditions. Intracellular metabolomic analysis of cells grown in medium containing fructose and lactate revealed key metabolic activities associated with chain elongation. Pangenomic analysis across nine annotated genomes identified 6,721 orthologous genes using OrthoMCL, emphasizing the genetic and functional diversity within the Megasphaera genus. This dataset offers valuable insights into the metabolism and biotechnological potential of M. hexanoica KCCM 43214T.

Metabolome

Complete genome sequence and genomic characterization of the probiotic Limosilactobacillus reuteri PSC102.

BACKGROUND: Gut microbiota are potential sources of probiotics and play an essential role in maintaining intestinal health. Limosilactobacillus reuteri PSC102 (L. reuteri PSC102), which was isolated from the feces of healthy pigs, exhibited health-beneficial properties. AIM: We aimed to conduct a whole-genome sequencing analysis of L. reuteri PSC102 to determine its molecular characteristics as a probiotic strain. METHODS: Limosilactobacillus reuteri PSC102 cells were cultured in De Man-Rogosa-Sharpe medium, followed by DNA extraction for genomic analysis using the PacBio-Illumina sequencing platform. The EzBioCloud software was used to perform gene assembly, and the genes were interpreted by the National Center for Biotechnology Information (NCBI) and the Glimmer program. Core and pan-genomic analyses were performed to assess the extent of functional conservation in the genomic sequence. Moreover, the NCBI database and the Basic Local Alignment Search Tool software were used to identify antimicrobial resistance genes and virulence factors. RESULTS: Limosilactobacillus reuteri PSC102 consists of a single circular chromosome with 2,048,626 bp, a guanine- cytosine of 38.9%, 18 rRNA genes, and 69 tRNA genes. Among the 1,846 protein-coding sequences, genes associated with probiotic characteristics were identified, including genes involved in host-microbe interactions, stress tolerance, biogenesis, and defense mechanisms. Furthermore, the genome of L. reuteri PSC102 comprises 2,446 pan-genome and 1,222 core-genome orthologous gene clusters. A total of 74 unique genes were identified in L. reuteri PSC102 genome. These genes mostly encode proteins potentially involved in the transport and metabolism of amino acids and carbohydrates. Moreover, antibacterial resistance genes and virulence factors were absent in L. reuteri PSC102. CONCLUSION: The results of the molecular insight into L. reuteri PSC102 corroborates its use as a probiotic in humans and other animals.

Limosilactobacillus reuteri

De novo clustering of large long-read transcriptome datasets with isONclust3.

MOTIVATION: Long-read sequencing techniques can sequence transcripts from end to end, greatly improving our ability to study the transcription process. Although there are several well-established tools for long-read transcriptome analysis, most are reference-based. This limits the analysis of organisms without high-quality reference genomes and samples or genes with high variability (e.g. cancer samples or some gene families). In such settings, analysis using a reference-free method is favorable. The computational problem of clustering long reads by region of common origin is well-established for reference-free transcriptome analysis pipelines. Such clustering enables large datasets to be split roughly by gene family and, therefore, an independent analysis of each cluster. There exist tools for this. However, none of those tools can efficiently process the large amount of reads that are now generated by long-read sequencing technologies. RESULTS: We present isONclust3, an improved algorithm over isONclust and isONclust2, to cluster massive long-read transcriptome datasets into gene families. Like isONclust, isONclust3 represents each cluster with a set of minimizers. However, unlike other approaches, isONclust3 dynamically updates the cluster representation during clustering by adding high-confidence minimizers from new reads assigned to the cluster and employs an iterative cluster-merging step. We show that isONclust3 yields results with higher or comparable quality to state-of-the-art algorithms but is 10-100 times faster on large datasets. Also, using a 256 Gb computing node, isONclust3 was the only tool that could cluster 37 million PacBio reads, which is a typical throughput of the recent PacBio Revio sequencing machine. AVAILABILITY AND IMPLEMENTATION: https://github.com/aljpetri/isONclust3.

Algorithms

Reference-Guided Chromosome-Scale Genome Assembly With Insights on Population Genomics of the Atlantic Goliath Grouper (Epinephelus itajara), Islas del Rosario, Colombia.

Epinephelus itajara, commonly known as the Atlantic Goliath grouper, is the largest species among the western North Atlantic groupers and is critically endangered. This species plays a crucial ecological, cultural, and economic role and has been the focus of captive breeding efforts at the Oceanario of the Rosario Islands, Colombia. However, despite its ecological and conservation importance, genomic resources and population genomic data for E. itajara remain scarce, particularly in the Colombian Caribbean. This study presents a reference-guided chromosome-scale genome assembly and an analysis of the population genomic structure of E. itajara using PacBio HiFi sequencing and Illumina technologies. The assembled genome has a total size of 1.12 Gb, with a contig N50 of 42.69 Mb and a scaffold N50 of 46.30 Mb. A total of 22,692 protein-coding genes were identified after masking 46% of the genome, which consists of repetitive elements. Comparative genomic analyses revealed a high degree of collinearity with closely related Epinephelus species and identified E. lanceolatus as the closest relative, supporting recent divergence and conserved genome architecture within the genus. Additionally, a population genomics analysis was conducted using 7706 high-quality SNPs to assess the genomic structure of captive populations. The results revealed four distinct genomic lineages, with moderate genetic differentiation among the sampled individuals. In the Colombian Caribbean, two unique lineages were identified, associated with the localities of Bahía Cispatá and Bahía Barbacoas, suggesting possible geographic isolation. These genomic resources provide valuable tools and new opportunities to better understand the genomic diversity, evolutionary history, and reproductive mechanisms of E. itajara. Moreover, they serve as a foundation for conservation strategies, including selective breeding programs aimed at increasing genomic diversity in captive populations and guiding restoration efforts in its natural habitat.

Epinephelus itajara

Chromosome-level genome assembly of Triplophysa scleroptera.

Triplophysa scleroptera is an endemic fish species in Qinghai Lake and the upper reaches of the Yellow River. However, studies on conservation and evolutionary genetics were seriously impeded by the absence of a reference genome. Here, by using PacBio HiFi sequencing and Hi-C assembly technology, we assembled a chromosome-level genome of T. scleroptera, with a total length of 660.22 Mb and 99.82% of the sequence anchored to 25 chromosomes. The contig N50 and scaffold N50 were 9.09 Mb and 24.38 Mb, respectively. The evaluation using BUSCO indicated the genome assembly to be 96.40% complete. About 33.41% of the genome consists of repeat elements. We predicted 26,168 protein-coding genes in the genome, and 99.02% of them were functionally annotated. This high-quality reference genome would serve as a valuable genomic resource for advancing evolutionary conservation genetics studies in this species.

Animals

A chromosome-level genome assembly of Coffea arabica L. var. 'Kona Typica'.

Coffea arabica L. var. 'Kona Typica' is renowned for its premium cup quality, but its vulnerability to pests and diseases limits production. To accelerate cultivar improvement, we generated a chromosome-level genome assembly of 'Kona Typica' using PacBio HiFi sequencing and Hi-C scaffolding technology. The final assembly spans 1.13 Gb, with a scaffold N50 of 50.50 Mb, organized into 22 chromosomes. BUSCO assessment indicated a high completeness at 99.1%. We annotated 65,458 protein-coding genes and identified 1,073,545 interspersed repeats, accounting for 65.16% of the genome. Analysis of transposon insertion ages revealed that most long terminal repeat retrotransposons proliferated after the polyploidization event. This high-quality genome assembly of 'Kona Typica' provides a valuable resource for exploring coffee genomic evolution and genetic mechanisms of complex traits, facilitating genomics studies and the development of improved coffee cultivars with enhanced disease resistance and quality traits.

Coffea

Chromosome-level genome assembly of the small-sized Taihang donkey (Equus asinus).

China harbors a rich diversity of donkey breeds, with small-sized donkeys (<110&#x2009;cm) representing a largely underexplored group. Here, we present the first high-quality, chromosome-level genome assembly of a small-sized donkey, generated using PacBio HiFi sequencing (286.7&#x2009;Gb), Hi-C scaffolding (240.47&#x2009;Gb), and annotated with RNA-seq data. The final assembly has a total length of 2.7&#x2009;Gb and comprises 32 chromosomes (including both X and Y chromosomes), in which five chromosomes were fully assembled without gaps. It possesses a scaffold N50 of 106.70&#x2009;Mb and 84 contigs (contig N50&#x2009;=&#x2009;63.60&#x2009;Mb), and captures 99.2% of BUSCO genes. The assembly achieved a consensus quality value (QV) of 77.44, corresponding to an extremely low base-level error rate, indicating exceptional nucleotide accuracy. This high-quality genome provides a valuable resource for investigating genetic variation, adaptive evolution, and domestication processes in small-sized donkeys, and will facilitate the conservation and sustainable utilization of rich donkey genetic resources in China.

Animals

MHASS: Microbiome HiFi Amplicon Sequencing Simulator.

SUMMARY: Microbiome HiFi Amplicon Sequence Simulator (MHASS) creates realistic synthetic PacBio HiFi amplicon sequencing datasets for microbiome studies, by integrating genome-aware abundance modeling, realistic dual-barcoding strategies, and empirically derived pass-number distributions from actual sequencing runs. MHASS generates datasets tailored for rigorous benchmarking and validation of long-read microbiome analysis workflows, including ASV clustering and taxonomic assignment. AVAILABILITY AND IMPLEMENTATION: Implemented in Python with automated dependency management, the source code for MHASS is freely available at https://github.com/rhowardstone/MHASS along with installation instructions. Our code is also published on Zenodo at https://doi.org/10.5281/zenodo.17486364. The data underlying this article are available on GitHub at https://github.com/rhowardstone/MHASS_evaluation/.

Software

Genome analysis of the glycosphingolipid-producing green alga tetraselmis sp. NKG400013.

Microalgae are gaining attention as sustainable resources for the production of valuable compounds, including biofuels, pigments, and bioactive metabolites. To support metabolic engineering and genome editing approaches aimed at enhancing these traits, high-quality genome assemblies are essential; however, genomic information remains limited for many microalgal lineages. Tetraselmis sp. NKG400013 is a green alga known for high glycosphingolipid accumulation with distinctive structural features. Here, we report a draft genome assembly of this strain generated using PacBio HiFi sequencing and transcriptome-supported annotation. The assembled genome spans 423.7&#x2005;Mbp, with 74.5% repetitive sequences and 15,322 predicted protein-coding genes. Comparative analyses across 11 green algal species revealed a positive correlation between genome sizes and repeat contents, indicating that transposable element expansion, particularly long terminal repeat retrotransposons, has substantially contributed to genome enlargement in Tetraselmis. Genome-wide functional annotation and ortholog inference identified core enzymes required for glycosylceramide biosynthesis. Both sphingolipid &#x394;4 and &#x394;8 desaturases were identified in Tetraselmis and their coexistence suggests an expanded capacity for long-chain base modification that may underlie its distinctive glycosphingolipid profile. These results establish a genomic framework for understanding the high glycosphingolipid-producing capacity of NKG400013 and provide insights into the evolutionary diversification of sphingolipid metabolism in green algae.

Chlorophyta

Genomes of Conopholis americana and Epifagus virginiana: two holoparasitic plants (Orobanchaceae).

Conopholis americana (American cancer-root) and Epifagus virginiana (beechdrops) are sister genera of holoparasitic plants (Orobanchaceae) native to eastern North America, parasitizing oaks and American beech, respectively. Both have served as models for plastid genome reduction, yet no nuclear genomes exist for either genus or any New World holoparasitic Orobanchaceae. Here we present the first nuclear genome assemblies for both species using PacBio HiFi sequencing. The C. americana assembly totals 1.82 Gb and E. virginiana totals 440 Mb, representing an approximately 4-fold difference in genome size between these sister genera. We observed a BUSCO completeness of 79% to 80% in both species, which is typical of holoparasites. While gene prediction identified 33,889 genes in C. americana and 21,031 in E. virginiana, repeat annotation revealed that LTR retrotransposons account for 78% of the genome size difference. These assemblies reveal contrasting mechanisms of genome evolution in sister holoparasitic genera and provide foundational resources for comparative genomics of parasitic plants.

Genome, Plant